How a language model
actually reads
and writes.
One sentence will stay with us: “The cat sat on the mat.”
What do you think ChatGPT is doing when you hit send?
We walk this path once. Then we stop.
If you remember only this picture, the rest is detail.
NLP is teaching computers to work with human language.
Understand
Classify. Find names. Match meaning.
Generate
Translate. Summarize. Write. Answer.
The hard part
Same word, two meanings. Context lives across sentences.
“I sat by the bank.” River — or money?
Is “bank” a river or money?
A language model predicts the next piece of text.
Running example: “The cat sat on the ___”
Stats era
Count it
Frequency tables. Fast. Almost no long memory.
Neural era
Learn it
Dense vectors. Must run left to right.
Transformers
Attend to it
Every token can look at the others. Parallel.
Which era is ChatGPT in?
An LLM is that same idea, trained at scale.
P
Parameters
billions of weights
C
Compute
the ingredient people forget
Not search. Not a database. Next-token prediction that got large enough to look like reasoning.
Would a smaller model with better data beat a huge one with junk?
The model never sees letters. It sees IDs.
The cat sat on the
→
The
cat
sat
on
the
→
464 · 3797 · 2647 · 322 · 279
Token + ID
A chunk of text, then an integer from a fixed list. That number is what the network can index.
The splitter is frozen
BPE merges frequent pairs until the list is ~32k–200k pieces. Built once. Train and chat must use the same one.
“unhappiness” often becomes un + happi + ness. “ChatGPT” is usually more than one ID.
Can two English models swap this splitter?
An ID becomes a point in a high-dimensional space.
LOOKUP · one ID = one row. “sat” is highlighted.
| ID | token | d0 | d1 | d2 | … | d4095 |
| 464 | The | 0.12 | −0.41 | 0.08 | … | 0.03 |
| 3797 | cat | −0.33 | 0.61 | 0.14 | … | 0.22 |
| 2647 | sat | 0.55 | 0.02 | −0.19 | … | −0.07 |
V×d
the whole table
50k × 4096 ≈ 200M numbers
1×d
one token
a list of numbers, not a word
If d is 4096, what is “sat” actually stored as?
Nobody types the table. Training writes it.
Same training as the rest
The table is just more numbers in the model. No separate “embedding class.”
Meaning is a side effect
Rows that help guess similar next words drift together. Cat sits near dog because that helped predict “sat.”
Then it freezes
In ChatGPT, lookup is only: ID → row. Your prompt does not rewrite the table.
If “mat” never appeared in training, can this table know it well?
Same word. Same vector? That was the bank question.
The lookup row for “bank” is still one row. Context is added in the layers. That is the answer to slide 3.
Which picture fails the river / money test?
Same family. Three ways to look.
Read, then write
Sees the full input, then writes the output.
Understand
Looks both ways. Labels and search. Not a chat model.
Generate
Each token may only look left. This is ChatGPT’s family.
Which one is running when you use ChatGPT?
Three jobs: stack, don’t peek, guess the next ID.
IDs inThe cat sat on the
Lookup + positionthe row, plus where it sits in the sentence
Same block, repeated
Listen: who on the left matters?
Think: update this token’s vector
Scores → percentagesone score per word in the vocab, then they sum to 100%
Pick the next IDappend it. Repeat.
Position is why “dog bites man” is not “man bites dog.” Same words, different order, different vectors.
When the model is writing “mat,” may it look at the word after it?
It writes one token. Then it reads what it just wrote.
Output becomes input
There is no separate “writer.” Generation is just next-token, over and over.
Your prompt is also tokens
ChatGPT is not thinking about your question as a document. It is continuing a token list.
If the first generated token is wrong, what happens to the rest?
First, scores become percentages. Then T reshapes them.
Softmax: every vocab score → a probability, all add to 100%. Temperature only stretches those scores first. The model is unchanged.
If you need a JSON field to always be the same, do you raise T or lower it?
ChatGPT is a next-token model plus manners.
1 · Base
Finish the internet
Raw LLM. It completes text. Ask a question, it may write another question.
2 · Chat training
Answer as an assistant
Show it conversations: user / assistant. Reward replies people actually want.
3 · Your send
Still next-token
The engine did not change. Sampling, temperature, one ID at a time — same as before.
Your chat does not update the weights. It only continues a token list that starts with a hidden system prompt.
If we turned chat training off, what would the model do with “What is a bank?”
From your keystroke to the stop token.
A detailed pass through the machine. Same sentence as the rest of the talk.
Living sequence · prompt stays, reply grows
Waiting for send.