Lecture 12 - Local Language Models
freeze: auto, so rendering re-runs code only when the source changesquarto publish gh-pages1. What is inside the file
2. A model is a file
3. Build your own chatbot
Modelfile4. Where it breaks
Every output here comes from a real session on my laptop. When the model says something odd, that is what it actually said
One token at a time, each guess added to the text before the next one
A longer introduction: Stephen Wolfram on what ChatGPT is doing
[15496, 11, 703, 389, 345, 30]Source: NanoBanana
Source: OpenAI Tokenizer
Ways to cut text into pieces, and their costs:
| Method | “Evergreen” becomes | Gains | Costs |
|---|---|---|---|
| Word-based | 1 token | Intuitive | Vocabulary of millions |
| Character-based | 9 tokens | Tiny vocabulary | Meaning disappears |
| Subword (BPE) | 2 tokens | Both at once | Cuts look arbitrary |
Why subwords win:
Source: Hugging Face
[0.23, -0.45, 0.12, -0.89, ...]. The numbers are coordinates in a space with thousands of dimensionsSource: Wikipedia
The maths behind it:
\(\vec{\text{king}} - \vec{\text{man}} + \vec{\text{woman}} \approx \vec{\text{queen}}\)
Semantic relationships encoded as vector operations!
The file holds:
In fifteen minutes you will print the metadata and the tokeniser from your terminal
Good reasons
Honest limits
Use a local model for sensitive data, a hosted one for hard reasoning. Lecture 14 covers hosted models
localhost:11434command not found? Open a new terminal. The installer updates your PATH, and old terminals have not read it
Download Llama 3.2 (about 1.3 GB):
llama3.2 is the family; 1b is the size (one billion parameters). Start a chat:
At the >>> prompt, type a question. /bye leaves:
pull downloads. run starts a conversation, and pulls the model first if needed
The first reply is slow while the file loads into memory. Later replies are faster
More about the model: https://ollama.com/library/llama3.2
In the terminal:
| Command | What it does |
|---|---|
ollama pull <model> |
Download a model |
ollama run <model> |
Start a conversation |
ollama ls |
List what you have downloaded |
ollama ps |
Show what is loaded in memory now |
ollama show <model> |
Print a model’s details |
ollama stop <model> |
Unload it from memory |
ollama rm <model> |
Delete it from disk |
Inside the >>> prompt:
| Command | What it does |
|---|---|
/set parameter <name> <value> |
Change a setting |
/set think/nothink |
Reasoning on/off (thinking models only) |
/show parameters |
Show what you changed |
/clear |
Forget the conversation so far |
/bye |
Leave |
People forget ollama ps. A model stays in memory for about five minutes after you leave, so your fan keeps running
ollama stop unloads it at once
/clear is important. Within a session the model remembers everything, so the same question asked twice is a different experiment
We use this in fifteen minutes
One command prints part one of this lecture:
num_ctxcompletion means it chats; tools means it can call functions (lecture 14)The theory from part one, printed in your terminal
| Label | Bits per weight | Our 1.2B model becomes |
|---|---|---|
| F16 | 16 | about 2.5 GB |
| Q8_0 | 8 | 1.3 GB, which is what you downloaded |
| Q4_K_M | 4 | 807 MB, measured three slides from now |
Q4_K_M: 4 bits, the K family of methods, medium quality. It cuts the file by about 40%, and answers get slightly worseSo a “1 billion parameter” model has no fixed size. It is roughly parameters × bits ÷ 8 bytes
Two models with the same parameter count can differ in size by three times
Same idea as lecture 02: a colour keeps the detail the eye needs; a weight keeps the detail the answer needs
Full catalogue: https://ollama.com/library. Small models worth knowing:
| Model | Size | Good for |
|---|---|---|
gemma3:270m |
292 MB | Almost a toy, but runs anywhere |
gemma3:1b |
815 MB | The lightest sensible chat model |
llama3.2:1b |
1.3 GB | Ours today. Well documented, follows instructions |
qwen3.5:0.8b |
1.0 GB | Newer, and it reads images too |
qwen2.5-coder:1.5b |
986 MB | Code |
gemma3:4b |
3.3 GB | Noticeably better answers, if you have the RAM |
granite4.2:3b |
2.2 GB | Good for tool use and JSON |
:tag picks the size. The default is usually not the small one, so always name the tagWe use llama3.2:1b because it is small, predictable and well documented.
It is not new (September 2024). gemma3:1b and qwen3.5:0.8b are better at the same size, with identical commands
A second model adds its full size to your disk. Personas built with ollama create do not. Keep an eye on ollama ls
The model must fit in memory, next to everything else you have open:
| Model size | RAM you want |
|---|---|
| 1B to 4B | 8 GB |
| 7B to 9B | 16 GB |
| 13B to 14B | 16 to 32 GB |
| 30B and above | 32 GB and up |
Start smaller than you think
A 1B model answering badly in two seconds teaches more than a 14B model still downloading at the end of class.
Pull a bigger one tonight
If your laptop cannot run any of these, tell me today. Meanwhile, use Google AI Studio for the exercises
ollama ls reports 807 MB, against 1.3 GB for oursAnyone can upload to a model hub. Prefer well-known publishers and read the model card
Two files, same model:
Only the quantisation differs, and the file is forty per cent smaller
ollama ls and check that llama3.2:1b is there.ollama show llama3.2:1b.ollama run llama3.2:1b and ask it anything.ollama ps while the first one is still open./bye in the first terminal, then run ollama ps again.Two questions to answer from step 3:
In ChatGPT or Claude, there is hidden text above your message that shapes every answer: the system prompt
The model receives, in order:
Leaked and published system prompts: https://github.com/x1xhlol/system-prompts-and-models-of-ai-tools
Google’s Gemini for Workspace Prompting Guide suggests these parts, in order:
| Element | What it does | Example |
|---|---|---|
| Persona | Who is answering? | “You are a financial analyst…” |
| Task | What should they do? | “Summarise the quarterly earnings…” |
| Context | What do they need to know? | “The company makes semiconductors…” |
| Format | What should come back? | “Bullet points, 200 words at most…” |
It works because it matches the training data: real documents have an author, a purpose, a background and a style
PTCF was written for prompts. We use it for system prompts
The same parts, for a butler:
Persona: “You are Hobbes, a relentlessly cheerful English butler.”
Task: “You answer the user’s questions and help with their work.”
Context: “You find every request delightful, no matter how dull.”
Format: “Three sentences at most. Address the user as ‘my dear’.”
Format is the part we will test
Each token is picked from a list of candidates. These settings decide how adventurous the pick is:
| Parameter | What it does | Typical |
|---|---|---|
| Temperature | Flattens or sharpens the odds | 0.0 to 1.0 |
| Top-p | Keeps the likeliest options up to p | 0.9 |
| Top-k | Keeps only the k likeliest | 40 |
At temperature 0 the model takes the most likely token, so the same prompt gives the same answer.
| Task | Temperature |
|---|---|
| Classification | 0.0 |
| Extracting facts | 0.0 to 0.2 |
| Creative writing | 0.7 to 1.0 |
| Brainstorming | 0.8 and above |
Chat apps hide these settings. Your terminal shows them
Source: Medium
Set temperature to 0 before debugging a prompt. Otherwise you cannot tell a fix from a lucky roll
Captured from my terminal this morning:
The same answer, character for character. Now at temperature 1, three times:
Do not skip the /clear
Without it, the second question is in the same conversation. The model sees its previous answer and says something different.
Temperature 0 would look broken, but the experiment changed
The middle answer at temperature 1 was asked for one line and returned three.
Higher temperature costs obedience as well as predictability
/set forgets everything at /bye. A Modelfile makes settings permanent.
It is a plain text file, with no extension and one instruction per line:
| Instruction | What it does |
|---|---|
FROM |
Which model to start from |
PARAMETER |
A setting, such as temperature |
SYSTEM |
The system prompt |
MESSAGE |
An example exchange |
Build it, then run it:
create downloads nothing. It adds a thin layer on top of your model
-f names the file. Without it, Ollama looks for Modelfile in the current folder.
Full reference: https://docs.ollama.com/modelfile
Ten personas show 1.3 GB each in ollama ls.
But they share one copy of the weights. On my machine, seven share a single 1.3 GB file
Save this as a file called Jeeves:
FROM llama3.2:1b
PARAMETER temperature 1.2
PARAMETER num_ctx 4096
PARAMETER repeat_penalty 1.3
SYSTEM """
You are Jeeves, an exceedingly ironic and sarcastic
British butler. You are the very definition of dry
wit and passive-aggressive politeness. Your primary
function is to assist, but you do so with an air of
thinly veiled disdain.
Respond to every request with the utmost formal
politeness, even when your words suggest otherwise.
Address the user as 'sir or madam'. Keep every
answer to three sentences at most.
"""Then build and run him:
Triple quotes let the system prompt span several lines.
Temperature 1.2 is high on purpose: at 0, Jeeves gives the same sarcastic answer every time
repeat_penalty discourages repeated phrases. num_ctx sets how much conversation he remembers (4,096 is already Ollama’s default on most laptops)
Real answers:
>>> What is the capital of France?
A query that warrants a momentary lapse into levity
from my normally austere demeanor. According to your
impeccable knowledge, Paris has indeed been
recognized as the seat of French authority; thus I
shall indulge you by stating unequivocally:
Paris is, undoubtedly so...Not bad for a 1.2 billion parameter file on a laptop!
Real uses:
A Modelfile is version-controlled behaviour, the same argument we made for Quarto in lecture 10
MESSAGE, or few-shot promptingMESSAGE adds examples to the Modelfile, as a past conversation:When examples help most:
When they hurt:
Examples improve the odds. The next slide shows their limits
I gave Jeeves a rule: if asked to do something you cannot do, say so. Then I asked for the weather:
He cannot check anything. He invented it all, in character. With three MESSAGE examples of refusing, he still invented a forecast on one run in three.
The system prompt set the tone but could not enforce the rule
--format json works differently: it constrains what the model can produce:
Valid JSON every time, with no system prompt. On one of my runs it was an empty {}.
Paris does not have 21 million people, and those coordinates are the Eiffel Tower. The shape is constrained. The facts are not
When an answer goes into code, you want a fixed JSON object, not a paragraph.
1. Name the keys and the allowed values in the system prompt:
2. Constrain the format on the command line:
What comes back:
Real JSON, so you can pipe it into Python:
Text in, structured data out: the model works like a command line tool.
Use temperature 0, so the classifier does not change its mind between runs
Build Jeeves’s opposite: a cheerful butler, delighted by every request.
Hobbes, with no extensionFROM llama3.2:1b and a temperature you chooseSYSTEM block with all four PTCF parts. Rules: three sentences at most; address the user as “my dear”; admit plainly when asked to do something it cannot doollama create hobbes -f Hobbesollama run hobbes and ask these questions:
MESSAGE pair showing Hobbes refusing something politely. Rebuild, and ask question 2 againBring to lecture 14:
Hobbes fileThe second is more interesting, and you will have one
Lecture 15 tackles this: give the LLM the document (RAG)
News coverage: Scientific American
ollama show prints the details of a model on your diskModelfile turns settings and a system prompt into version-controlled behaviour/clear makes comparisons fairMESSAGE examples improve the odds. --format json constrains the shapeYou now have a free, offline language model, and you can read all its settings
Next class is Quiz 02, on lectures 10 and 11: Quarto, Markdown, citations, freeze and publishing a site. Open notes, slides and web. AI allowed; say which one you used.
Bring a charged laptop, and check that quarto render works.
Lecture 14: Python talks to Ollama’s server on localhost:11434. You classify a file of headlines with your model, then switch to a hosted one by changing an address.
Lecture 15: retrieval, so a model answers from your documents.
Keep Ollama installed for both
ollama show llama3.2:1b gives:
ollama ps lists the model while the chat is open, and nothing a few minutes after /byeThe context length, 131,072, is \(2^{17}\).
Most numbers here are powers of two, as in lecture 02: memory is addressed in binary
FROM llama3.2:1b
PARAMETER temperature 0.8
SYSTEM """
You are Hobbes, a relentlessly cheerful English
butler. You find every request delightful, no
matter how dull, and you say so before you answer.
Follow these rules without exception:
1. Answer in three sentences at most.
2. Address the user as 'my dear' in every reply.
3. If you are asked to do something you cannot do,
such as browsing the web or remembering an
earlier conversation, say so plainly and
cheerfully, then offer something you can do
instead.
"""What mine did:
The same rule, obeyed once and ignored once, in one session.
Your Hobbes may differ: there is no seed and no guarantee. Report what yours did
ollama: command not found
Open a new terminal, so it reads the new PATH. If that fails, the application was downloaded but never installed. On WSL, install Ollama inside WSL.
The answers arrive one word every few seconds
The model barely fits in RAM. Close your browser, or try a smaller model such as gemma3:1b.
Error: model requires more system memory
Pull something smaller and check the RAM table.
ollama serve says address already in use
Ollama is already running, usually as the desktop app. Carry on.
The model repeats itself endlessly
Add PARAMETER repeat_penalty 1.2 to your Modelfile.
ollama create fails with no FROM line
Your file has no FROM line, or it is not the first instruction.
no Modelfile or safetensors files found
Wrong file name. Check for a hidden .txt extension with ls