What Happens When You Send a Prompt to an AI
From the text you type to the first character on screen: tokenization, embeddings, attention, and sampling, without the maths.
You type a question, press enter, and a second later an answer streams back. The experience feels like asking a person who happens to type very fast.
Underneath, it is a loop that guesses the next fragment of text and runs again, a few hundred times, with several stages of arithmetic between your sentence and the reply.
Your Words Become Tokens
The model never sees your sentence as words. Before anything else, a tokenizer chops the text into pieces called tokens and swaps each piece for a number. Those numbers come from a fixed list, the vocabulary, which holds around 200,000 entries.
Each model family has its own list and its own splitting rules, so the sizes and the exact splits vary from one model to the next.
Common words get their own entry. Rarer ones are split. Kubernetes becomes two tokens, K and ubernetes. tokenization becomes token and ization.
You can watch this happen:
# pip install tiktoken
import tiktoken
enc = tiktoken.get_encoding("o200k_base") # OpenAI's current encoding
s = "What happens when we send a prompt to an AI?"
ids = enc.encode(s)
print(len(s), "chars,", len(ids), "tokens")
print([enc.decode([i]) for i in ids])
# the same word, with and without a space in front
print([enc.decode([i]) for i in enc.encode("happens")])
print([enc.decode([i]) for i in enc.encode(" happens")])
Which prints:
44 chars, 11 tokens
['What', ' happens', ' when', ' we', ' send', ' a', ' prompt', ' to', ' an', ' AI', '?']
['h', 'app', 'ens']
[' happens']
Eleven tokens for eleven words, and every word happens to survive whole. The question mark is its own token, which is how the model can tell a question from a statement before it reads a single word of the body.
Every token but the first begins with a space, and the space belongs to the token that follows it. With a space in front, the word happens is a single token. With nothing in front of it, the same word breaks into h, app, and ens, because the tokenizer’s training text almost always had a space there. A word’s cost and identity depend on what sits next to it.
Digits are worse. 2 + 2 = is five tokens, one of which is a bare space, and that splitting is part of why models are unreliable at arithmetic they could otherwise perform.
Tokens are the unit you are billed in. Code and numbers break into more pieces than ordinary words do, so code costs more per character than prose does.
The prompt that reaches the model is longer than what you typed. The platform wraps your message in a template with system instructions and role markers, and that whole structure goes through the tokenizer. Your words are one part of a prepared conversation.
Each Token Becomes a Point in Meaning
A number tells the model nothing. So each token ID is looked up in a table and replaced with a vector: a list of a few thousand numbers describing where that token sits relative to the others.
These vectors are learned during training, and the learning arranges them so tokens used in similar ways end up near each other. Direction in that space carries meaning. An old demonstration from the word2vec era showed that king, minus man, plus woman lands closest to queen. Modern models are messier, but the geometry is still there. dog and puppy end up as neighbours; dog and bicycle do not.
Position gets added on top. dog bites man and man bites dog contain the same tokens, and only the positions tell them apart. What arrives at the first transformer layer is a list of vectors, one per token, each one marked with where it appeared.
The Tokens Read Each Other
A modern model is a stack of blocks, often dozens of them, and each block does two things. First, attention. Then a small neural network that processes every token on its own.
Attention is where the tokens exchange information. Each token looks at the tokens before it and works out which ones matter for its own meaning. Take “he paid the bill”. The word bill on its own is ambiguous. If restaurant appeared earlier, the vector for bill shifts toward the food sense; if electric appeared, it shifts the other way. The same token comes out of attention with a different vector depending on its company. Every token does this at once, so the whole prompt is processed in parallel.
Then it repeats up the stack. The early blocks deal with grammar and nearby words. The later ones assemble larger things: what the question is about, what kind of answer is being asked for, and the tone the conversation has established. By the middle of the stack, bank in a sentence about rivers and bank in a sentence about accounts have almost nothing in common.
You can inspect which tokens attended to which, and that tells you where information moved. It does not tell you what the model concluded, or why.
One Token Gets Chosen
After the final block, the model has one vector representing the last position in the sequence, and one more layer converts it into a score for every token in the vocabulary. A high score means the model considers that token a likely continuation.
The scores are turned into probabilities, and then a single token is picked. Always taking the highest-scoring token produces flat, repetitive output, so most systems inject a little randomness, and temperature controls how much. At low temperature the winner is almost always chosen. Turn it up and tokens further down the list get their chance, which makes the output more varied and less dependable. Ask the same question twice and it can come back with two different answers, because a token that lost the draw the first time can win it the second.
At this point the model has not composed a sentence or decided anything about your question. It ranked 200,000 possible fragments and picked one, and nothing in that ranking separates a true token from a likely one.
Then It Happens Again
The chosen token is attached to the end of the input, and the whole process runs again, now with that token as part of the context.
That is why a mistake cannot be taken back. Once a token sits in the input, everything after it is chosen in its context, so a wrong turn early in a reply steers the rest of it. The next pass can only continue from where the last one stopped.
The words appearing on your screen one at a time are this loop, streamed as each token finishes.
The wait before the first word is the prompt being processed, because all of it has to go through the stack before any output can exist. After that, each new token only needs to be appended to work the model has already done, and machines keep that intermediate work rather than recomputing it, so the start of a reply is slow and the rest arrives at a fairly steady pace. Change something near the start of a long prompt and the earlier work is thrown away, because every token after the change now sits in a different context. The model keeps no memory between calls, so every message resends the whole conversation, and the chat gets slower and more expensive as it grows.
The loop ends when the model emits a token that marks the end of the reply, or when it hits the response limit. Reaching the end is just another token it can choose.
Takeaways
- You are billed per token, and the whole conversation is sent again with every message, so a long chat costs more than a short one doing the same work.
- The reply is assembled one token at a time, front to back. A wrong turn early is carried into everything after it, which makes fixing the prompt cheaper than arguing with the answer halfway through.
- The same prompt can come back with a different answer. Treat a single reply as one sample rather than as the answer.
- Nothing along this path checks whether a claim is true, only whether it is likely, so anything that matters needs checking against a source outside the model.