DATASCI 101: Introduction to AI Applications

Lecture 06: Language, Tokenisation, and Embeddings

Danilo Freire

Department of Data and Decision Sciences
Emory University

Welcome back! 🤓

Recap of last class

  • We covered model evaluation across two domains
  • Traditional ML metrics: accuracy, precision, recall
  • The accuracy paradox: 99.9% accuracy can be useless!
  • Overfitting: models memorise instead of learn
  • Cross-validation: more reliable evaluation
  • LLM evaluation: perplexity, LLM-as-a-Judge, G-Eval
  • Hallucinations and RAG for grounding answers
  • Today: How do LLMs actually understand language? 🤔

Lecture overview

Today’s agenda

Part 1: How LLMs “See” Text

  • The translation problem: Text to numbers
  • The processing pipeline

Part 2: Tokens and Tokenisation

  • What tokens are (hint: not always words!)
  • Tokenisation methods and quirks
  • Token limits and API costs

Part 3: Embeddings

  • Words as vectors in space
  • Semantic similarity and the king-queen example

Part 4: Parameters

  • What are the billions of parameters?
  • Embeddings, weights, and biases
  • Temperature and creativity dials

Meme of the day

Source: Reddit r/Singularity

New study by Anthropic on how AI impacts coding skills

How LLMs “See” Text 👁️

The translation problem

LLMs don’t read English!

  • Computers only understand numbers
  • Type “Hello, how are you?” and the LLM does not see letters
  • It sees something like: [15496, 11, 703, 527, 499, 30]
  • This creates a translation problem:
    • How do we convert text into numbers?
    • How do we preserve meaning in those numbers?
    • How do we capture relationships between words?
  • Two concepts solve it:
    • Tokenisation: break text into pieces
    • Embeddings: turn pieces into meaningful vectors

From text to output

The LLM processing pipeline

  1. Input text: “The cat sat on the mat”
  2. Tokenisation: break into tokens → ["The", "cat", "sat", "on", "the", "mat"]
  3. Token IDs: map to numbers → [464, 3797, 3332, 319, 262, 2603]
  4. Embeddings: each token becomes a list of ~4,096 numbers!
  5. Processing: pass through neural network layers
  6. Output: generate next token probabilities
  7. Decoding: convert back to text

Your prompt goes through all of these steps before the model produces a single word

Source: NanoBanana

Why this matters for you

Practical implications

The pipeline helps you:

  • Write better prompts: know what the model “sees”
  • Understand costs: API pricing is per token, not per word
  • Avoid surprises: some words use more tokens than others
  • Debug issues: why did the model misunderstand?
  • Use context wisely: token limits are real constraints

Fun fact: The word “everything” is 1 token, but “ChatGPT” is 3 tokens in OpenAI’s (old) tokeniser! 🤯

Source: OpenAI

Tokens and Tokenisation 🧩

What is a token?

The basic unit of LLM processing

  • A token is the basic unit an LLM reads
  • Tokens are not always words
  • A token can be:
    • A full word: “hello” → 1 token
    • Part of a word: “unhappiness” → “un” + “happiness” (2 tokens)
    • A single character: “!” → 1 token
    • A space: ” ” → often included with the next word
  • Rule of thumb: 1 token ≈ 4 characters in English, or 100 tokens ≈ 75 words
  • Other languages often use more tokens per word
  • Let’s try “Hello, it’s Danilo here!” in OpenAI’s tokenizer

Why use tokens instead of words?

A clever engineering choice

Three approaches to breaking up text:

Method Example: “Evergreen” Pros Cons
Word-based 1 token Intuitive Huge vocabulary
Character-based 9 tokens Tiny vocabulary Loses meaning
Subword (BPE) 2 tokens Best of both! Less intuitive
  • Modern LLMs use Byte Pair Encoding (BPE)
    • Start with single characters (bytes) as tokens
    • Find the pair that appears together most often in the training text, e.g. “u” + “n”
    • Merge it into a new token (“un”) and repeat until the vocabulary is big enough
  • Common subwords become single tokens; rare words split into known pieces
  • It also saves space: “unbelievable”, “unhappy”, “undo” and “unknown” share the same “un”. The model stores “un” only once!

Why subwords win:

  • Handles unknown words gracefully
  • Keeps vocabulary manageable (~50,000 to 200,000 tokens)
  • Common words stay whole, rare words split sensibly
  • Works across languages

Source: Hugging Face

Tokenisation in action

Try it yourself!

OpenAI’s Tokeniser: platform.openai.com/tokenizer

Token counts change from one version to another!

Some surprising examples:

Text Tokens Count
“Hello” [“Hello”] 1
“hello” [“hello”] 1
” hello” (with space) [” hello”] 1
“Hello!” [“Hello”, “!”] 2
“everything” [“everything”] 1
“ChatGPT” [“Chat”, “G”, “PT”] 3
“São Paulo” [“S”, “ão”, ” Paulo”] 3
“🎉” Multiple bytes 2+

Capitalisation, spacing and punctuation all affect tokenisation!

Source: OpenAI Platform

Token limits and context windows

Why your prompt has a maximum length

  • Every LLM has a context window: the maximum tokens it can process
    • Longer inputs need much more computing power and memory, so providers set a limit
  • The limit covers both your prompt and the response
  • Output tokens usually cost 5-6x more than input tokens!
Model Context Window
GPT-3.5 4,096 or 16,384 tokens
GPT-4 8,192 or 128,000 tokens
Gemini Pro / Claude / GPT-5.5 1,000,000+ tokens
  • A 3,500-token prompt against a 4,096 limit leaves 596 tokens for the response!
  • Exceeding limits gives an error or truncated output
  • Non-English languages often use more tokens
    • Higher costs and shorter effective context, though newer tokenisers narrow the gap

Source: Anthropic

Same meaning, different tokens:

Language Text GPT-3 GPT-4o
English “Hello” 1 1
Chinese “你好” 4 1
Arabic “مرحبا” 6 2
Japanese “こんにちは” 6 1

Embeddings 🌌

What are embeddings?

Words as points in space

  • Each token gets converted to an embedding
  • An embedding is a vector: a list of numbers, typically 768 to 4,096 per word
  • Example: “cat” → [0.23, -0.45, 0.12, -0.89, ... (4,096 numbers)]
  • These numbers capture the meaning of the word
  • Similar words have similar vectors
    • “cat” and “dog” sit close together
    • “cat” and “aeroplane” sit far apart
  • The model learns these representations during training, and this is where the “understanding” happens
    • Each word gets one vector, fixed once training ends
  • Remember \(\vec{\text{king}} - \vec{\text{man}} + \vec{\text{woman}} \approx \vec{\text{queen}}\)

Source: TensorFlow Projector

Similar words cluster together in the embedding space

Word2Vec: A brief history

Where modern embeddings began

  • Word2Vec (2013, Google’s Mikolov et al.) revolutionised NLP
  • CBOW (Continuous Bag of Words): predict the word from its context
  • Skip-gram: predict the context from the word
  • Key insight: words in similar contexts have similar meanings
    • “I love my _____” → cat, dog, hamster are all likely
    • so cat, dog and hamster should sit near each other!
  • Trained on billions of words from the web
  • Created the embedding approach modern LLMs use
  • Watch: Word2Vec Explained (10 min)

Source: Chris McCormick

“You shall know a word by the company it keeps” — J.R. Firth (1957)

Word2vec illustrated

How embeddings capture meaning

  • Here are the embeddings for a few words
  • Each row is one word’s 50 numbers, from blue (-2) to red (+2)
  • See how “woman” and “girl” are similar, but “queen” is further away?
  • We don’t know what each dimension codes for, yet patterns still show
  • In some places “king” and “queen” match each other and differ from all the rest
  • The model has learned something about gender and royalty!
  • Note that “water” has few connections to other words

Data: GloVe. Layout after Jay Alammar

How big are real embeddings?

From 50 numbers to thousands

  • GloVe (2014) stores each word as 50 numbers, like the rows on the last slide
  • LLMs store each token as thousands of numbers: 4,096 in Llama 3 8B, 12,288 in GPT-3
  • More numbers leave room for finer shades of meaning
  • Nobody reads them one by one: we study the patterns instead

Embeddings in practice

Real-world applications

How embeddings power modern AI:

  • Semantic search: search by meaning, not keywords
    • “cheap flights” finds “budget airfare”
  • Recommendations: find similar products, articles or content
    • “Users who liked X also liked Y”
  • Clustering: group similar documents automatically
    • topic modelling without labels
  • RAG (Retrieval-Augmented Generation): find relevant documents for LLM context
    • the power behind ChatGPT + browsing/search

Similarity search example:

Query: “How to fix a broken phone screen?”

Most similar (by embedding): 1. “Repairing cracked smartphone displays” 2. “DIY phone screen replacement guide” 3. “Mobile repair services near me”

Meaning match, not keyword match!

Source: Pinecone

Parameters: The LLM’s Brain 🧠

What is a parameter?

What are these numbers, exactly?

  • Think back to algebra: \(y = 2x + 3\)
  • The numbers 2 and 3 are parameters: they set the behaviour of the function
  • In an LLM, parameters are numbers that:
    • get set during training
    • control how the model processes text
    • determine what the model “knows”
  • GPT-3: 175 billion parameters
  • GPT-4: estimated 1+ trillion parameters
  • Each GPT-3 parameter was adjusted about 100,000 times during training
  • Training means finding the numbers that make predictions accurate

Source: Medium

Each arrow represents parameters that are learned during training

How many parameters?

The scale of modern LLMs

Model Total parameters Used per token Training data
GPT-2 (2019) 1.5 billion All 40 GB text
GPT-3 (2020) 175 billion All ~300 billion tokens
Llama 3.1 (2024) 405 billion All 15 trillion tokens
Kimi K2 (2025) 1 trillion 32 billion 15.5 trillion tokens
DeepSeek-V4-Pro (2026) 1.6 trillion 49 billion 32+ trillion tokens
GPT-5, Claude, Gemini Undisclosed Undisclosed Undisclosed

To put this in perspective:

  • 1.6 trillion parameters = 1,600,000,000,000 individual numbers
  • Counting 1 per second would take ~50,000 years!
  • Most new models use Mixture of Experts (MoE): many smaller “expert” networks
    • Each token goes to only a few experts: DeepSeek-V4-Pro uses ~3% of its parameters per token
  • Training data grew ~100x since GPT-3

The three types of parameters

Embeddings, weights, and biases

1. Embedding parameters

  • Store the learned vector for each token
  • Vocabulary (~50,000) × dimensions (~4,096) = ~200 million parameters just for embeddings!

2. Weight parameters

  • Control connections between neurons, and how strongly parts influence each other
  • Like volume knobs: amplify or damp signals
  • Like regression coefficients: almost all of an LLM’s parameters are weights

3. Bias parameters

  • Adjust thresholds for neuron activation
  • Like the intercept in a regression, they shift the output up or down
  • Many modern LLMs (Llama, DeepSeek) drop most biases

Source: Medium

Simple analogy:

  • Embeddings = vocabulary
  • Weights = grammar rules
  • Biases = intuition adjustments

Training: Finding the right numbers

How parameters get their values

The training process:

  1. Start with random parameter values
  2. Show the model text: “The cat sat on the ___”
  3. Model predicts: “table” (wrong!)
  4. Calculate the error
  5. Backpropagation: adjust parameters to reduce the error
  6. Repeat billions of times

In short:

  • The model learns by predicting the next word, and every error teaches it something
  • Patterns emerge from massive repetition
  • Training GPT-3 cost ~$4.6 million in compute!

Source: Daniel McKee and Maghav Kumar

Parameters are adjusted step by step to minimise errors

Model size vs. capability

Bigger isn’t always better

The scaling hypothesis:

  • More parameters generally means more capability
  • But diminishing returns set in: a 10x bigger model isn’t 10x smarter

Small models fighting back:

  • Efficient training: better data, longer training
  • Distillation: teach small models from big ones
  • Quantisation: reduce precision (32-bit → 4-bit)
  • Mixture of Experts: use only the relevant parts

Example:

  • LLaMA 2 (7B) can beat GPT-3 (175B) on some tasks!
  • Why? Better training data and techniques
  • Quality beats quantity in training data, although more parameters help with generalisation

Source: OpenAI Scaling Laws Paper

Hyperparameters: The creativity dials

Temperature, top-p, and top-k

Not all parameters are learned. Hyperparameters are settings you control:

  • Temperature: controls randomness
    • low (0.1-0.3): focused, deterministic
    • high (0.7-1.0): creative, varied
    • 0.0: always picks the most likely word
    • (OpenAI and Google AI Studio go up to 2.0)
  • Top-p (nucleus sampling): keep the fewest words whose probabilities add up to p
    • 0.9: the most likely words until they cover 90% (often just a handful)
  • Top-k: only consider the top k most likely words
    • k=50: choose from the 50 most likely words

Temperature in action:

Source: Medium

Putting It All Together 🔗

The full picture

From your prompt to the response

The attention mechanism (recap)

How LLMs understand context

  • Attention is the key innovation of transformers
  • Every word “looks at” every other word to find which are relevant to each other

Example:

“The animal didn’t cross the street because it was too tired”

  • What does “it” refer to? Attention connects “it” to “animal”
  • Not all words matter equally for understanding

How it works:

  • Each word asks: “who should I pay attention to?”
  • Weighted connections between all words capture long-range dependencies

Source: Jay Alammar

Darker = more attention between those words

Two ways of using context

Word2vec and attention are not the same thing

Word2vec

  • Happens once, during training
  • Looks at neighbours across billions of sentences
  • Averages everything it saw into one point per word, then freezes it
  • “cheap” keeps that same vector for ever, in every sentence you ever write

Attention

  • Happens every time you run the model
  • Looks at the neighbours in this sentence only
  • Rebuilds each word’s vector from scratch, every time
  • “cheap flights” and “cheap wine” give “cheap” two different vectors

When you add context to a prompt, attention is the part that reads it

Practical tips for prompting

Using your new knowledge

Now that you understand the internals:

  • Be concise: fewer tokens mean lower cost and more room for the response
  • Use clear language: common words tokenise efficiently
  • Provide context: help the attention mechanism
  • Be specific: the model uses your exact words
  • Test different temperatures: match creativity to the task
  • Mind the context limit: plan for both prompt and response

When things go wrong:

  • Model confused? → rephrase with different words
  • Model doesn’t understand? → give more context
  • Running out of context? → summarise earlier content

Quick reference:

Task Temperature
Factual Q&A 0.0-0.3
General chat 0.5-0.7
Creative writing 0.7-1.0
Brainstorming 0.8-1.2

Source: Microsoft Education

Summary 📚

Main takeaways

  • LLMs don’t read text; they process tokens, numerical representations of text chunks

  • Tokenisation breaks text into pieces; ~100 tokens ≈ 75 words; affects cost and limits

  • Embeddings convert tokens to vectors, capturing meaning in ~4,096 dimensions

  • The famous king − man + woman ≈ queen shows embeddings encode semantic relationships

  • Parameters (billions of them!) are the numbers learned during training

  • Hyperparameters like temperature let you control creativity vs. determinism

  • Attention mechanisms help models understand context and word relationships

… and that’s all for today! 🎉