Token
The small chunks of text an AI model reads and writes, which is how AI usage is measured and priced.
What it is
AI models don't read words; they read tokens, pieces of text roughly three-quarters of a word long. "Unbelievable" might be three tokens. Usage and cost are counted in tokens in and tokens out.
Each model also has a context window: the maximum number of tokens it can consider at once. Paste in a huge document and older parts may be cut off.
Tokens matter when an AI feature is expensive or slow: shorter prompts, smaller inputs, and capped outputs save money.
How to ask for it
“The AI feature is costing too much.”
“Reduce token usage in the summarizer: send only the first 3,000 words of each document, ask for a summary under 120 words, and cache summaries so the same document is never summarized twice.”
The weak prompt states a cost problem; naming tokens points to the levers (input size, output length, caching).
You've seen this in
- PPricing pages for AI APIs ("per million tokens")
- CChat apps warning a conversation is too long
- AUsage dashboards in AI tools