4 min read
Making sense of AI tokens
itfoundations
Originally posted on September 15, 2026
Last updated on September 15, 2026
What drives AI usage cost? (and why Copilot forgets things)
If you've spent any time using Microsoft Copilot or ChatGPT, you've probably run into one of two frustrations:
The first is trying to understand what using AI is actually going to cost your business as you chew through usage on Claude or Copilot Cowork).
The second is when a conversation starts strongly, but after a while the AI seems to lose track of something you told it earlier.
Although they seem like completely different issues, they're both linked to the same concept: tokens.
- Why does Copilot sometimes forget things?
- Why tokens influence AI costs
- Three simple ways to keep performance and costs under control
- The bottom line
- Next Steps..?
What are AI tokens?

Put simply, tokens are how AI systems process information.
Rather than reading text exactly as we do, AI models break it down into smaller chunks called tokens. These can be whole words, parts of words, punctuation or numbers.
As a rough guide, 1,000 tokens equates to around 750 words, which is around two pages of text. A larger report, contract, or policy document could easily contain tens of thousands of tokens.
The important thing to remember is that both the information you provide and the response you receive consume tokens. Every prompt, document upload, question and answer contributes to the context the AI is working with.
You don't need to understand token counts in detail, but knowing they exist helps explain both the limitations and costs of modern AI tools.
Why does Copilot sometimes forget things?
Every AI model has a limit to how much information it can actively consider at any given moment.
When you're chatting with Copilot, everything contributes to that working context:
- Your prompts
- Previous messages
- Uploaded documents
- Copilot's responses
- Information Copilot retrieves to answer your question
All of this has to fit within the model's available context, and as the conversation grows, that working context fills up. To make room for new information, older content may be condensed, deprioritised or removed.
That's why long conversations can sometimes become inconsistent, why large documents aren't always analysed as expected, and why starting a fresh chat often improves results.
If you've ever found yourself frustrated and typing "As per my previous message..." to Copilot, then there's a good chance you've reached this limit.
Think of it like joining a meeting halfway through - If you're only given the last few minutes of discussion, you're likely to miss some important context from the start.
Why tokens influence AI costs
Tokens don't just affect how AI performs, they're also the main factor that influences cost.
For many organisations using Microsoft 365 Copilot, this isn't something you need to think about day to day. The service is licensed on a straightforward per-user basis, making costs predictable and easy to budget for.
However, things become more interesting when businesses start building custom AI solutions, whether that's creating agents in Copilot Studio or using Cowork, connecting other AI like Claude to business systems, automating workflows or integrating directly with AI services, costs often move from licence-based pricing to usage-based pricing. At that point, the amount of work the AI performs becomes important.
Larger prompts, more retrieved information, more complex reasoning and longer responses all require more resources.
For example, lets take a situation with two Copilot agents connected to SharePoint.
Both agents are asked:
"What is our process for onboarding a new employee?"
One agent searches every document in the HR library before answering, while the other has been configured to focus only on onboarding documentation.
To the user, both provide the same answer.
Behind the scenes, however, one is processing significantly more information than the other. As usage grows, those differences can affect both performance and consumption costs.
This is why good AI design isn't just about getting the right answer, it's about getting the right answer efficiently and from the outset.
This is where organisations and users can get caught out. The solution works exactly as intended, but it's processing far more information than necessary, leading to higher consumption and cost.
In our experience, unexpected AI costs are likely to come from poorly optimised automations and agents.
Three simple ways to keep performance and costs under control
-
Keep conversations focused
Try not to use a single chat for multiple unrelated tasks. When you're moving onto a new topic, start a fresh conversation. You'll often get more accurate responses and reduce the amount of unnecessary context building up over time.
-
Provide relevant information
It's very tempting to upload an entire document when asking a question. However, providing the specific section that's relevant to your query will often generate better results. It also reduces the amount of information the AI needs to process. Less noise generally leads to better answers.
-
Understand what your AI solution is reading
Before deploying an AI-powered process, workflow or agent, ask a simple question: What information is it processing and how often?
That one question can uncover potential performance and cost issues before they become a problem.
The most successful AI projects aren't always the most sophisticated. Instead, they're usually the ones designed with clear objectives, efficient data access, and sensible governance from the outset.
The bottom line
Most businesses don't need to understand the technical details behind large language models. However, having a basic understanding of tokens helps explain two of the questions we hear most often:
"Why is this AI solution costing more than expected?"
and
"Why has Copilot forgotten what I told it?"
In both cases, the answer usually comes back to how much information the AI is working with at any given time.
Understanding that doesn't require a computer science degree, it simply helps you get better results, make smarter decisions and avoid surprises as AI becomes part of your day-to-day operations.
Next Steps..?
AI adoption is moving quickly, but successful deployment still comes down to having the right foundations in place.
Whether you're evaluating Microsoft Copilot, exploring AI agents, or trying to understand the potential costs and benefits for your organisation, a little planning upfront can save significant time and money later.
The good news is that you don't need to become an expert in tokens, context windows, or AI architecture to make informed decisions. You simply need to understand enough to ask the right questions before investing time, money and resources into a new solution.
If you're exploring AI and would like practical advice on licensing, governance, adoption or cost management, we'd be happy to help.
