Lecture 15 - Retrieval-Augmented Generation and Fine-Tuning
ModelfileSource: Wikipedia
LLMs have read most of the public internet, but none of your documents: your notes, your data dictionary, this course’s files
How can we change that?
1. Embeddings, again
2. Retrieval-augmented generation
3. Fine-tuning
4. Choosing
embeddinggemma has 300 million parameters, a quarter of llama3.2:1bWord2Vec embeddings, with “dog” and its neighbours highlighted. Explore it yourself at projector.tensorflow.org
In Lecture 12 this picture explained how models work. Today we use it directly
Our script writes the formula out by hand, so you can see each step
To find the best match, score the question against every passage and take the highest score
Real scores for “What does Quiz 02 cover?”:
| Best passage from | Score |
|---|---|
| Lecture 12 README | 0.62 |
| Quiz 02 README | 0.43 |
| Anything, for “capital of Nepal?” | 0.20 |
The best match is a paragraph about Quiz 02 in the Lecture 12 README, not the quiz’s own file
pip install ollama:localhost:11434/api/embed, as in Lecture 14If the pull fails, update Ollama, or use the older nomic-embed-text (274 MB)
“Grounded in the information you trust” is RAG as marketing. Source: notebook.google
RAG can cite sources and use information newer than the model
Sometimes you should! If your notes fit in the context window, pasting them is simpler.
RAG is worthwhile when:
Where RAG fails:
A 300-page handbook is about 120,000 tokens. You ask 1,000 questions over a term:
| Strategy | Input per question | 1,000 questions |
|---|---|---|
| Paste everything | 120,000 tokens | 120M tokens ≈ $18.00 |
| RAG, top 3 chunks | ~600 tokens | 0.6M tokens ≈ $0.09 |
Long prompts also hurt quality: models remember the start and end of a long text better than the middle (Liu et al., 2024)
Why this corpus:
You can use your own notes instead
A chunk is a piece of text that we embed and retrieve. Common ways to cut:
| Strategy | How it cuts | Gains | Costs |
|---|---|---|---|
| By paragraph | On blank lines | The author’s own units of meaning | Sizes vary wildly |
| Fixed window | Every ~500 tokens, with overlap | Uniform, nothing lost at edges | Cuts mid-thought |
| By structure | On headings and sections | Sections stay whole | Chunks can be huge |
Our script cuts at blank lines and drops pieces under 80 characters, which removes headings and short bullets
demo/, so Python finds corpus/chunks starts as an empty list.md file in corpus/, in alphabetical orderread_text reads the file, and split("\n\n") cuts it at blank linesstrip removes extra spaces, and the if skips headings and short lines(file name, paragraph). The file name makes citations possible later('course-readme.md', 'Welcome to [DATASCI 350]...')For eight files, we do not need LangChain or a vector database. A Python list is enough
embed turns text into vectors, and cosine compares two vectors:import math
import ollama
EMBED_MODEL = "embeddinggemma"
def embed(texts):
"""Turn a list of texts into one vector per text."""
response = ollama.embed(model=EMBED_MODEL, input=texts)
return response["embeddings"]
def cosine(a, b):
"""Cosine similarity between two vectors."""
dot = 0.0
length_a = 0.0
length_b = 0.0
for i in range(len(a)):
dot = dot + a[i] * b[i]
length_a = length_a + a[i] * a[i]
length_b = length_b + b[i] * b[i]
return dot / (math.sqrt(length_a) * math.sqrt(length_b))embed returns 768 numbers per textcosine multiplies the pairs, adds them up, and divides by the two lengthsThe full script is in Appendix 02
TOP_K = 3
QUESTION = "What does Quiz 02 cover?"
chunk_texts = []
for name, text in chunks:
chunk_texts.append(text)
chunk_vectors = embed(chunk_texts)
question_vector = embed([QUESTION])[0]
scores = []
for vector in chunk_vectors:
scores.append(cosine(vector, question_vector))
ranked = []
for i in range(len(chunks)):
ranked.append((scores[i], chunks[i]))
ranked.sort(reverse=True)
top = ranked[:TOP_K]QUESTION at the top of the file to ask something elseembed needs plain stringsranked holds (score, chunk) pairs. Python sorts pairs by the first item, so reverse=True puts the highest score first[:TOP_K] keeps the top threeThat is all of retrieval. The rest is prompting
CHAT_MODEL = "llama3.2:1b"
passages = []
for score, chunk in top:
passages.append(chunk[1])
context = "\n\n".join(passages)
prompt = (
"Answer the question using ONLY the context below. "
"If the answer is not in the context, "
"say you do not know.\n\n"
f"Context:\n{context}\n\nQuestion: {QUESTION}"
)
reply = ollama.chat(
model=CHAT_MODEL,
messages=[{"role": "user", "content": prompt}])
print(reply.message.content)join puts them togetherf inserts context and QUESTION into the textollama.chat sends the prompt, as in Lecture 14CHAT_MODEL can be any model you have pulledPTCF applies: the context is C, the task is T, and we skipped the persona
$ python rag.py # QUESTION = "What does Quiz 02 cover?"
Corpus: 132 chunks from 8 files
Retrieved chunks:
[0.621] lecture-12-readme.md: Next class is Quiz 02: Literate Programming, worth 6%. It covers lectu...
[0.429] quiz-02-readme.md: There are two bonus tasks at the end of the quiz README. Attempt them ...
[0.429] quiz-01-readme.md: There are two bonus tasks at the end of the quiz README. Attempt them ...
Answer:
Based on the context, Quiz 02: Literate Programming covers topics such as:
1. Lectures 10 and 11: Quarto, Markdown, citations, `freeze`, and building a site.
2. Open notes, open slides, open web, and AI allowed.Retrieval gives the same result every time. Generation does not
In four runs, the scores never changed
But the answers did. Once, the model said “I do not know” with the right chunk in front of it 😅
Write questions where you know the source file, then check whether it is in the top 3:
| Question | Expected file | Rank |
|---|---|---|
| What does Quiz 01 cover? | quiz-01-readme.md | 2 |
| Which tutorials does the course have? | tutorials-readme.md | 1 |
| What is Lecture 12 about? | lecture-12-readme.md | 1 |
| When is Quiz 02? | quiz-02-readme.md | 3 |
| What does Lecture 10 cover? | lecture-10-readme.md | 1 |
| What does Lecture 11 cover? | lecture-11-readme.md | 1 |
| How much is the final project worth? | course-readme.md | miss |
As with Lecture 14’s human_label: write the answer key first
The hit rate here is 6 of 7. You do not need a chat model to test it
demo/ folder from the course repositoryollama pull embeddinggemmapip install ollamarag.py, set QUESTION to “What does Quiz 02 cover?”, and run python rag.py from inside demo/QUESTION twice more, to questions the corpus can answerQUESTION to “What is the capital of Nepal?” and run againrag.py, change TOP_K from 3 to 1What to look for:
TOP_K = 1, questions that need two files fail firstExpected behaviour and troubleshooting:
WHERE queriesLangChain and LlamaIndex wrap these same five stages
You wrote them yourself, so their documentation will make sense 😉
For a course project, a list is enough. Switch when the search gets slow
Source: daxa.ai
Try it: add that line to a corpus/ file, ask a question that retrieves it, and see what your model does
The pipeline also runs on OpenRouter with your Lecture 14 key. Only two calls change:
:free models, so your $0.00 key limit still worksAs with classify.py, you can switch between local and hosted models
So far we changed what the model reads. Fine-tuning changes what the model is.
| RAG / prompting | Fine-tuning | |
|---|---|---|
| What changes | The context | The weights |
| New facts | Immediately | Only by retraining |
| Sources | Citable | Gone, absorbed |
| Cost shape | Per question | Up front |
| Undo | Delete a file | Keep the old weights |
To fix a RAG mistake, you edit a file. To fix a fine-tuning mistake, you train again
Training data is JSONL: one example per line (split here to fit the slide):
MESSAGE, but with thousands of examplesYou have already used a model made this way:
Modelfile changes a model’s behaviour with words. LoRA does it with trained weightsFine-tune for behaviour:
Do not fine-tune for facts:
Most requests to fine-tune are really RAG problems. Check before paying for GPUs
You will not fine-tune anything in this course. It is good to know these notebooks exist
llama3.2:1b partly by distilling larger Llama modelsHere, a model writes the training examples, not people
@AnthropicAI, 23 February 2026
prompt → RAG → fine-tune
Each step costs more and is harder to undo:
The AI module is complete. Next class we start cloud computing
We will rent a computer from AWS and control it from the terminal
Quiz 03 covers the AI module and the cloud module
Before then:
.envTry rag.py on your own notes this week
Step 5, questions the corpus answers well:
Steps 7 and 8, the Nepal question:
llama3.2:1b refused seven times and twice answered “Japan”, copied from a Quarto chunk (-P country:Japan)Steps 9 and 10, TOP_K = 1:
Going further:
CORPUS at a folder of your own notesErrors and fixes are in Appendix 04
"""A minimal RAG pipeline over the course's own notes."""
import math
from pathlib import Path
import ollama
EMBED_MODEL = "embeddinggemma"
CHAT_MODEL = "llama3.2:1b"
TOP_K = 3
QUESTION = "What does Quiz 02 cover?"
CORPUS = "corpus" # the folder of notes, next to this script
def embed(texts):
"""Turn a list of texts into one vector per text."""
response = ollama.embed(model=EMBED_MODEL, input=texts)
return response["embeddings"]
def cosine(a, b):
"""Cosine similarity between two vectors of the same length."""
dot = 0.0
length_a = 0.0
length_b = 0.0
for i in range(len(a)):
dot = dot + a[i] * b[i]
length_a = length_a + a[i] * a[i]
length_b = length_b + b[i] * b[i]
return dot / (math.sqrt(length_a) * math.sqrt(length_b))
# Split every markdown file into paragraph chunks
chunks = []
for path in sorted(Path(CORPUS).glob("*.md")):
text = path.read_text(encoding="utf-8")
for paragraph in text.split("\n\n"):
clean = paragraph.strip()
if len(clean) > 80:
chunks.append((path.name, clean))
files = set()
for name, text in chunks:
files.add(name)
print(f"Corpus: {len(chunks)} chunks from {len(files)} files")
chunk_texts = []
for name, text in chunks:
chunk_texts.append(text)
chunk_vectors = embed(chunk_texts)
question_vector = embed([QUESTION])[0]
scores = []
for vector in chunk_vectors:
scores.append(cosine(vector, question_vector))
ranked = []
for i in range(len(chunks)):
ranked.append((scores[i], chunks[i]))
ranked.sort(reverse=True)
top = ranked[:TOP_K]
print("\nRetrieved chunks:")
for score, chunk in top:
name, text = chunk
print(f" [{score:.3f}] {name}: {text[:70]}...")
passages = []
for score, chunk in top:
passages.append(chunk[1])
context = "\n\n".join(passages)
prompt = (
"Answer the question using ONLY the context below. "
"If the answer is not in the context, say you do not know.\n\n"
f"Context:\n{context}\n\nQuestion: {QUESTION}"
)
reply = ollama.chat(model=CHAT_MODEL, messages=[{"role": "user", "content": prompt}])
print(f"\nAnswer:\n{reply.message.content}")Replace the two Ollama calls in rag.py with these, using the client and .env setup from Lecture 14:
import os
from dotenv import load_dotenv
from openai import OpenAI
load_dotenv()
client = OpenAI(base_url="https://openrouter.ai/api/v1",
api_key=os.environ["OPENROUTER_API_KEY"])
def embed(texts):
response = client.embeddings.create(
model="nvidia/nemotron-3-embed-1b:free", input=texts)
vectors = []
for item in response.data:
vectors.append(item.embedding)
return vectors
# ...and replace the ollama.chat call with:
reply = client.chat.completions.create(
model="google/gemma-4-31b-it:free", # any live :free id
messages=[{"role": "user", "content": prompt}])
print(f"\nAnswer:\n{reply.choices[0].message.content}")Both calls use :free models, so a $0.00 key limit works. Free models come and go: check openrouter.ai/models for live ones
ollama pull embeddinggemma fails
Your Ollama is too old for this model. Update the application, or switch EMBED_MODEL to nomic-embed-text and pull that instead.
Connection refused on localhost:11434
Ollama is not running. Open the app, or run ollama serve.
ModuleNotFoundError: No module named 'ollama'
Run pip install ollama in the Python you use (which python3).
Corpus: 0 chunks from 0 files
You ran the script from the wrong folder, so it finds no notes, retrieves nothing, and the model says “I do not know”. cd demo first.
The first run takes ages
The embedding model is loading into memory. Every run re-embeds all 132 chunks, because the script keeps nothing between runs
The answer is nonsense
Read the retrieved chunks. Wrong chunks: rephrase, or check the corpus has the answer. Right chunks, wrong answer: try a larger chat model.
Every score is low
Scores near 0.2 mean the corpus has nothing close to your question.
The model answers from outside the corpus
Small models leak training memory past the grounding instruction. Try a larger model and note the difference