Shubham Pagare
Session · LLMs · Shubham Pagare

How a language model
actually reads
and writes.

One sentence will stay with us: “The cat sat on the mat.”

What do you think ChatGPT is doing when you hit send?

The map

We walk this path once. Then we stop.

Text messy language → Tokens chunks + IDs → Vectors lookup table → Layers look left, think → Next ID pick, append, repeat Green stamp = when you hit send. Amber stamp = how the weights were born. Not happening in the chat.

If you remember only this picture, the rest is detail.

NLP

NLP is teaching computers to work with human language.

Understand

Classify. Find names. Match meaning.

Generate

Translate. Summarize. Write. Answer.

The hard part

Same word, two meanings. Context lives across sentences.

“I sat by the bank.” River — or money?

Is “bank” a river or money?

Language models

A language model predicts the next piece of text.

Running example: “The cat sat on the ___”

HOW FAR CAN IT REMEMBER · schematic Stats n-gram · last few words Neural RNN / LSTM · fading memory Transformer this is our era
Stats era

Count it

Frequency tables. Fast. Almost no long memory.

Neural era

Learn it

Dense vectors. Must run left to right.

Transformers

Attend to it

Every token can look at the others. Parallel.

Which era is ChatGPT in?

What is an LLM
Train time · not the chat

An LLM is that same idea, trained at scale.

P
Parameters
billions of weights
D
Data
web, books, code
C
Compute
the ingredient people forget

Not search. Not a database. Next-token prediction that got large enough to look like reasoning.

Would a smaller model with better data beat a huge one with junk?

Tokens
When you hit send

The model never sees letters. It sees IDs.

The cat sat on the → The cat sat on the → 464 · 3797 · 2647 · 322 · 279

Token + ID

A chunk of text, then an integer from a fixed list. That number is what the network can index.

The splitter is frozen

BPE merges frequent pairs until the list is ~32k–200k pieces. Built once. Train and chat must use the same one.

“unhappiness” often becomes un + happi + ness. “ChatGPT” is usually more than one ID.

HOW THE ALPHABET IS BUILT · once, before training l o w e r lo w e r low er Chars too small. Words too many. Subwords win.

Can two English models swap this splitter?

Embeddings
When you hit send

An ID becomes a point in a high-dimensional space.

LOOKUP · one ID = one row. “sat” is highlighted.

IDtokend0d1d2…d4095
464The0.12−0.410.08…0.03
3797cat−0.330.610.14…0.22
2647sat0.550.02−0.19…−0.07
V×d
the whole table
50k × 4096 ≈ 200M numbers
1×d
one token
a list of numbers, not a word
SKETCH · real space has thousands of axes cat dog sat python the nearby ≈ related

If d is 4096, what is “sat” actually stored as?

How the lookup table is born
Train time · not the chat

Nobody types the table. Training writes it.

ONE TRAINING STEP · “The cat sat on the ___” 1. Random rows · noise at the start 2. Lookup · each ID picks its row 3. Model guesses the next word 4. Wrong? A mistake score is born 5. That mistake nudges those rows only tokens in this sentence move

Same training as the rest

The table is just more numbers in the model. No separate “embedding class.”

Meaning is a side effect

Rows that help guess similar next words drift together. Cat sits near dog because that helped predict “sat.”

Then it freezes

In ChatGPT, lookup is only: ID → row. Your prompt does not rewrite the table.

If “mat” never appeared in training, can this table know it well?

Static vs contextual
When you hit send

Same word. Same vector? That was the bank question.

STATIC · one row, forever bank river? money? AFTER THE LAYERS · the vector has moved bank · river bank · loan

The lookup row for “bank” is still one row. Context is added in the layers. That is the answer to slide 3.

Which picture fails the river / money test?

Three shapes
Architecture

Same family. Three ways to look.

WHO CAN LOOK AT WHOM Encoder–decoder ENC ↔ DEC → T5 · translate Encoder only ENC ↔ ↔ ↔ BERT · classify Decoder only DEC → → → GPT · Llama · chat

Read, then write

Sees the full input, then writes the output.

Understand

Looks both ways. Labels and search. Not a chat model.

Generate

Each token may only look left. This is ChatGPT’s family.

Which one is running when you use ChatGPT?

Decoder-only
When you hit send

Three jobs: stack, don’t peek, guess the next ID.

IDs inThe cat sat on the
Lookup + positionthe row, plus where it sits in the sentence
Same block, repeated
Listen: who on the left matters? Think: update this token’s vector
Scores → percentagesone score per word in the vocab, then they sum to 100%
Pick the next IDappend it. Repeat.
NO PEEKING AT THE FUTURE The cat sat on Green = allowed. Grey = blocked.

Position is why “dog bites man” is not “man bites dog.” Same words, different order, different vectors.

When the model is writing “mat,” may it look at the word after it?

Generation
When you hit send

It writes one token. Then it reads what it just wrote.

THE SENTENCE GROWS BY 1 now The cat sat on the next The cat sat on the mat Then it feeds “mat” back in and guesses again. That is why long answers take time — and why one bad token can derail the rest.

Output becomes input

There is no separate “writer.” Generation is just next-token, over and over.

Your prompt is also tokens

ChatGPT is not thinking about your question as a document. It is continuing a token list.

If the first generated token is wrong, what happens to the rest?

Temperature
When you hit send

First, scores become percentages. Then T reshapes them.

Softmax: every vocab score → a probability, all add to 100%. Temperature only stretches those scores first. The model is unchanged.

“The cat sat on the ___” · schematic T = 0.2 · almost always mat mat ~99% rug floor moon T = 1 · as trained mat ~64% rug ~24% floor ~11% moon ~2% T = 2 · moon is allowed mat ~46% rug ~28% floor ~19% moon ~8% Low T = safe, repetitive. High T = diverse — and more chance the next token derails the rest. T → 0 is always the top word. APIs usually sit near 0.7–1.

If you need a JSON field to always be the same, do you raise T or lower it?

Base model vs chat
Extra training · after the internet

ChatGPT is a next-token model plus manners.

1 · Base

Finish the internet

Raw LLM. It completes text. Ask a question, it may write another question.

2 · Chat training

Answer as an assistant

Show it conversations: user / assistant. Reward replies people actually want.

3 · Your send

Still next-token

The engine did not change. Sampling, temperature, one ID at a time — same as before.

Your chat does not update the weights. It only continues a token list that starts with a hidden system prompt.

If we turned chat training off, what would the model do with “What is a bank?”

Inside one send · schematic IDs
1 / 16

From your keystroke to the stop token.

A detailed pass through the machine. Same sentence as the rest of the talk.

Living sequence · prompt stays, reply grows

Waiting for send.

1 / 15