What one month of logging Copilot Cowork costs has taught me
- Hamish Sheild
- 6 days ago
- 10 min read
Updated: 2 minutes ago
Common questions I am hearing from clients right now are some version of this: "Are you using Copilot Cowork? Is it worth it? How do I know what it's going to cost?". Copilot Cowork was announced generally available last month which means that the free 'Frontier' period is over and you will now be billed on a usage basis, denominated in Copilot Credits. The most frequent concern we are hearing now is how to budget for Cowork given its variable pricing.

Microsoft provides some high level guidance and ways to predict costs with their simple estimator spreadsheet but it's calculations are "derived from aggregated and anonymized data collected from customers who participated in the Cowork Frontier program.". The spreadsheet also says "The data represents generalized estimates and may vary based on individual customer scenarios."
Predicting the cost is one thing, but measuring the actual cost and value you get in return is where I have been focusing my time recently. Copilot Cowork is not a tool that should be used for everything and I am learning that there are expensive and cheaper ways to use it for the same task. I've built up a way to work with Copilot Cowork to assess the costs of a task, how to optimise usage and identify what tasks Copilot Cowork is cost effective at, and which ones it is not.
The concern is real
Gartner predicts that more than 40% of AI projects will be cancelled by the end of 2027, with escalating costs cited as one of the main drivers alongside unclear value and weak risk controls. That aligns with what I'm hearing. Not because the tools are overpriced, but because you can't see what drives a session's cost, and can't compare with the value of the outcome.
How to get the actual cost
A lot of people are using the /cost command in Copilot Cowork to see what the cost of a task was. This gives a single Copilot Credit number but does not inform where the cost lies and how you might optimise it.
Cowork bills on Copilot Credits, Microsoft's currency for tokens which I assume is to make things simpler to understand. However, to understand costs properly, it is important to understand the token usage breakdown which is split into:
Input: what you type into Copilot Cowork directly in the moment ("draft me a backlog QA review"). This is almost always tiny and almost never the cost driver.
Cache write: the cost of handing it the reference material Copilot Cowork needs for the first time: your project files, the instructions for the task, any documents you've attached. This is the "getting up to speed" cost. It's the most expensive token type per word, because it's new information the Cowork has to absorb.
Cache read: once something has been handed over, the Copilot Cowork keeps a working copy of it for the rest of the conversation. Every time you ask a follow-up question, it re-checks that working copy rather than re-absorbing everything from scratch. This is much cheaper per word than cache write, but it adds up if the conversation runs long, or if you're carrying a large reference document through many back-and-forth turns that don't require it anymore.
Output: what Copilot Cowork actually produces and gives back to you. The written report, the document, the analysis. This is the most expensive token type per word, but it's also the part where you are getting value in return, so high output on a substantial deliverable isn't a red flag by itself.
Update 29 July 2026
Below I describe using the "/usage" command in Copilot Cowork to get the token usage breakdown. However, as of today, this command is no longer working for myself or others.
The instructions and prompts in the rest of this post still work and provide helpful tips about cost management, we just can't get the detailed token breakdown anymore.
To get the token usage breakdown type the /usage command.

If you are like me and are not super technical then this output doesn't make a lot of sense by itself. However, you can ask Copilot Cowork to make sense of it for you! Read on...
How I work with Copilot Cowork
A bit of background first, because it shapes where the cost comes from.
I don't use Copilot Cowork as a chatbot. I use it for a handful of jobs that I have identified it does well, and most of my work runs as a chain. It reads a pile of source material, workshop recordings or documents, and pulls the useful parts into a tidy, structured form. It saves that as a reference it can reuse. Then it works from that reference to do the next job: drafting a document, or reviewing one that already exists.
Two habits sit on top of that. After each job I ask Copilot Cowork to critique its own work and its own cost. I lean heavily on skills, which are written instructions that tells Copilot Cowork it how to do a particular job the way I want it done. As I watch how a skill performs in real work, I refine the skill and feed the better version back to Copilot Cowork for next time.
So the picture is a set of repeatable jobs, each backed by a skill I keep improving. That repetition is the reason the cost is worth understanding. A few dollars saved on something I run once doesn't matter. The same saving on a job I run every day or week does.
What I started doing
Since Copilot Cowork has been generally available and the billing method was announced, I started treating cost as part of my end-of-session routine. It's two steps.
First, I run the /cost and /usage commands myself at the end of a session and note the output. Cowork can't run /usage for you. You need to enter it directly yourself.
Then I ask:
"Based on what you did for this task, give me: (1) which inputs or tasks were the main cost drivers and why, based on your understanding of what was processed; (2) what was inefficient relative to the value it produced; and (3) 2-3 specific, actionable recommendations for how this session could have been cheaper. Format as a structured summary I can paste into a usage log."
I keep a running log: date, task, total cost, model used, what the breakdown said, and the recommendations.
I can then ask Copilot to analyse the log to identify where I can make improvement to my skills files and the way that I work with Copilot Cowork.
Update: I have now turned the prompt into a skill itself so that I don't have to remember the prompt and it ensures that the structure of the summary is the same each time.
What one month of data is showing me
I've logged over 20 sessions so far. Here's what's standing out.
The range is wider than I expected. Sessions have ranged from $0.63 to $11.38 for tasks I'd loosely describe as document-management work.
Populating a new document from three source documents, running on the Sonnet 4.6 model: $0.63.
Updating and populating 13+ project artifacts from a meeting transcript: $11.38.
Both involved reading documents and writing structured output. The difference is explainable once you look at what each session loaded, but you can't see it without running the prompt above.
Model tier has a bigger impact than task complexity. In one session, a trivial new project setup task consisting of read three templates, write three files, nothing to analyse, ran on Opus 4.8 and cost $1.47. A more substantial populate task in the same session (three source files including a Word document, full draft-and-review cycle) ran on Sonnet 4.6 and cost $0.63. The simpler task cost more than twice as much, purely because of model choice. I'm finding that if you don't explicitly set the model then auto seems to use Opus 4.8 most of the time.
It is also interesting, that for some tasks I also prefer what Sonnet produced over Opus even though Opus is the "better" model.

Planning what goes into a session matters. A long run of back-and-forth turns/chats in one session can add up quietly too, especially if there are large reference files attached. Every earlier turn gets carried and resources re-read again on each new chat input in the same session (task). In one case a single reference file ended up re-read in the background multiple times before the conversation was done, simply because the work kept going in one session instead of starting fresh once that reference file's job was finished.
The practical takeaway for someone managing a team using this tool is that the most cost effective sessions are the ones with a clear, bounded task and only the reference material that task actually needs. The expensive ones are usually not expensive because the output was large, they're expensive because something large was kept "open" in the reference files for longer than it needed to be on a multi-turn session.
It is likely that your tasks are completely different to mine. Start using the prompt above, log the results and see what you can learn.
What I think the fixes will save
Based on one month of data from one person's (my) use cases, I estimate that I have identified the following savings through this process:
Switching from Opus 4.8 to Sonnet 4.6 for writing, populate, and QA tasks: probably a 50%+ cost reduction on those tasks
Updating a skill to fix an output re-typing pattern that Copilot Cowork was using (read-and-append rather than writing all content as literal string data): estimated 30–40% reduction in output tokens for story and QA sessions
When reading large reference files only load the sections a skill needs rather than the full document: estimated 20–40% reduction in cache write on input files
I'm working through these fixes now. I don't have validated before-and-after figures yet. What I do have is a clear list of what to change and a baseline to measure against.
Managing at an organisational level
My process gives me user-level visibility. For leaders thinking about this across a team, Microsoft has built controls into the Microsoft 365 Admin Centre.
The Cost Management dashboard lets admins set spending limits at tenant, group, and user level, configure usage alerts, and see usage broken down by user.. Users can see what each task costs as they run it. Admins can set per-user credit caps inside group policies, and users can request additional credits when they need them.
The Microsoft 365 Copilot usage report in the Admin Centre gives you the adoption view. Which users are active across which Copilot experiences, how engagement is trending, and where usage is growing. Together, the two dashboards cover both the adoption picture and the spend picture.
That's the organisational layer. The individual session breakdown, what drove a specific task's cost, still sits inside the session, which is why the end-of-session question matters even with the Admin Centre controls in place. Both the admin and users side of things are required to truly understand what Copilot Cowork costs and also to be able to estimate cost in the future.
Is Copilot Cowork worth it?
I had a client ask me the other day, "would you pay for Copilot Cowork?". The answer is 100% yes, for the right task.
I have identified the right tasks to use Copilot Cowork for and it provides huge value. Some of the tasks it does for me would take me hours manually, or I would do them poorly, or I just would not do them at all because I didn't have the time. For me, it's not just about time saving, Cowork enables me to produce outputs with higher quality or do things that add extra value to my clients but I haven't been able to in the past due to time pressure.
I am spending about $200 per month on Copilot Cowork (just for myself) and I can only see that increasing as I expand to other team members and other use cases.
$200 per month is not a lot, but when you scale that to every person in a large organisation, that adds up very quickly! Rolling out Copilot Cowork in a large organisation needs to be done carefully. Each user will potentially use Copilot Cowork differently depending on their job role. For each user, you will need to find the right tasks that Copilot Cowork provides a good ROI. That is going to take time and effort, it is not just a case of turning Copilot Cowork on and hoping for the best.
It's early days
I've been doing this for just over a month. My use cases are specific: AI design sprint delivery work, processing workshop transcripts, populating structured project artifacts, quality assurance of project artifacts. I don't know whether these patterns hold for different types of work, or how they shift as a project grows.
What I do know is that the cost isn't a black box. The agent knows what it processed, what it loaded, and what it wrote. Ask it what that cost. Write it down. Spot the patterns for refinement.
I'll keep logging. If you're doing something similar, I'd be curious what you're finding.
FAQ
Q: What drives the cost of a Copilot Cowork session?
A: Four inputs determine the cost of each Cowork task: model use, context retrieval, tool calls, and runtime. In practice, the biggest drivers are how much content gets loaded into context at the start of a session (cache write), how often that context gets re-read across multiple turns (cache read), and how much new text the agent generates as output. Sessions doing similar work can cost very different amounts depending on which of these inputs dominates.
Q: How do I find out what drove the cost of a Copilot Cowork session?
A: Run /usage yourself at the end of a session, Cowork can't run it for you. Then paste the output back into the chat and ask the agent to explain what drove the cost, what was inefficient, and what specific changes would make the next similar session cheaper. Keeping a simple log of the task, cost, and explanation is enough to start identifying patterns over time.
Q: Does choosing a different AI model reduce Copilot Cowork costs?
A: Yes, significantly. Copilot Cowork supports multiple models, and the cost difference between them is substantial. A simple templating task that ran on Opus 4.8 cost $1.47 in one session; a more complex populate task in the same session ran on Sonnet 4.6 and cost $0.63. Matching the model to the complexity of the task is one of the highest-leverage cost levers available.
Q: What tools does the Microsoft 365 Admin Centre have for managing Copilot Cowork costs?
A: The Microsoft 365 Admin Centre includes a Cost Management dashboard that lets administrators set spending limits at tenant, group, and user level, configure usage alerts, and view consumption broken down by user, group, and feature. A separate Copilot usage report covers the adoption picture: which users are active, how engagement is trending, and where usage is growing. The two dashboards complement each other, the Admin Centre gives the organisational view, while the per-session /usage command gives the individual task breakdown
Q: What is Copilot Cowork's usage-based billing model?
A: Copilot Cowork bills in Copilot Credits, with charges based on actual task usage rather than a fixed subscription. Each task's cost is calculated from model use, context retrieval, tool calls, and runtime. Admins can choose between pay-as-you-go billing and a prepaid credits model (P3) that offers a discount in exchange for committing to a usage volume in advance.