Every AI call is priced and limited in tokens, and a token is not a word. Bani Chaudhuri explains what one actually is, a chunk of text that might be a whole word, part of a word, or a single character, at a rough guide of one token per four characters.
The reason models work this way is storage. There are millions of possible words once you count slang, typos, jargon, and whatever got invented this week, so instead of holding all of them a model keeps a smaller set of reusable pieces and assembles the rest. That design decision is what shows up in your bill and in your agent's memory, because both input and output are billed by the token and the context window limit is counted in them too.