What Tokens Are and How They Drive AI Cost

Perspectives
5 min
April 18, 2026
Bani Chaudhuri

Every AI call is priced and limited in tokens, and a token is not a word. Bani Chaudhuri explains what one actually is, a chunk of text that might be a whole word, part of a word, or a single character, at a rough guide of one token per four characters.

The reason models work this way is storage. There are millions of possible words once you count slang, typos, jargon, and whatever got invented this week, so instead of holding all of them a model keeps a smaller set of reusable pieces and assembles the rest. That design decision is what shows up in your bill and in your agent's memory, because both input and output are billed by the token and the context window limit is counted in them too.

What this video covers

  • What a token is, and why a long word gets split into pieces the model has already seen
  • Why breaking text into reusable chunks keeps a model manageable and lets it handle words it has never met
  • How token counts become cost, with instructions, user messages, and responses all billed
  • Why the context window limit is counted in tokens, and what gets dropped when a conversation hits it
  • Four habits for agent builders, be precise in instructions, know your model's limit, watch your usage, and summarize long histories

Chapters

  • 0:00 What a token actually is
  • 0:50 Why models do not read words
  • 2:11 How tokens affect cost and memory
  • 3:38 Tips for agent builders

More videos

Watching is one thing. Bring us your conversations.

30 minutes, using your real support and acquisition questions instead of a sample agent.

Schedule a Demo