Lecture 25 - Docker for Data Science
==venv and pip gave us an isolated Python
/opt/venvpython:3.14-slim in about two minutesdanilofreire/datasci350-example1. The project starter
2. The Dockerfile, instruction by instruction
FROM to CMD.dockerignore, docker history, and the cold build timed on this laptop3. Run, change, rebuild
-v flag that decides whether you ever see the report4. What breaks, and the wider world
| Criterion | Weight |
|---|---|
Reproducibility: the container builds and the report renders from a clean docker run |
30% |
| Analysis quality: the research question, the method, the interpretation | 30% |
| Communication: clarity of the writing and the visualisations | 20% |
| Code quality: organised, readable, no dead code | 10% |
| Git workflow: meaningful commits from all members, sensible repository structure | 10% |
A report that calls the API at render time fails on the day the API is down. Pull once, commit data/raw/, and leave it alone
scripts/pull_data.py talks to the API, and report.qmd reads only data/raw/| Path | What it does |
|---|---|
scripts/pull_data.py |
Calls the World Bank API once and saves the response |
data/raw/ |
The saved snapshot, committed on purpose |
report.qmd |
The Quarto report, headings matching the rubric |
requirements.txt |
Some Python packages, pinned |
Dockerfile |
The recipe that builds the container |
.dockerignore |
What stays out of the image |
data/raw/ is committedscripts/pull_data.py asks the World Bank for one indicator and writes two files:.gitignore contains a note:Add data/raw/ to .gitignore and your build still succeeds. The failure arrives later, in my clone, when the report tries to read a file that never left your laptop: see it happen
requirements.txt for a report# Talking to web APIs
requests==2.32.5
# Data processing (use either, or both)
# DuckDB runs SQL over your CSV files and
# hands the answer back as a pandas
# DataFrame when you call .df() on it.
pandas==3.0.5
duckdb==1.5.5
# Plotting
matplotlib==3.10.8
# Quarto needs these to run Python code
# inside report.qmd
ipykernel==7.2.0
jupyter-client==8.8.0
nbclient==0.10.4
nbformat==5.10.4
pyyaml==6.0.3| Package | Why it is there |
|---|---|
requests |
The pull script calls the API with it |
pandas |
The pull script tidies the API response with it, and DuckDB hands results back as pandas DataFrames |
duckdb |
Lecture 22’s SQL over files, used in the demo chunk |
matplotlib |
Draws the figure |
ipykernel |
Quarto starts a Python kernel through it |
jupyter-client |
Talks to that kernel |
nbclient, nbformat |
Execute and store the notebook Quarto builds |
pyyaml |
Quarto’s Python helper imports it |
pip freeze | wc -l printed 53, because every dependency brings its ownreport.qmd and what it needs from the imageimport duckdb
import matplotlib.pyplot as plt
data = duckdb.sql(
"SELECT country, year, value "
"FROM 'data/raw/life_expectancy.csv' "
"ORDER BY country, year"
).df()
print(f"Loaded {len(data)} rows covering "
f"{data['country'].nunique()} countries.")
fig, ax = plt.subplots(figsize=(7, 4))
for country, subset in data.groupby("country"):
ax.plot(subset["year"], subset["value"])
[...]
plt.show()output/report.html
duckdb, pandas and matplotlib, and a kernel Quarto can startquarto binary itselfdata/raw/life_expectancy.csv at that exact relative pathembed-resources: true puts the figures inside the HTML
output/report.html is one file you can email, 1.3 MB for mineLoaded 96 rows covering 4 countries.echo: true, which shows your code in the report
# 1. Base image
FROM ubuntu:24.04
# 2. Build-time settings
ENV DEBIAN_FRONTEND=noninteractive
SHELL ["/bin/bash", "-c"]
# 3. System packages
RUN apt-get update && apt-get install -y [...]
# 4. Quarto
RUN ARCH=$(dpkg --print-architecture) && [...]
# 5. Python environment
RUN python3 -m venv /opt/venv
ENV PATH="/opt/venv/bin:$PATH"
COPY requirements.txt /project/requirements.txt
RUN pip install --no-cache-dir -r /project/requirements.txt
# 6. Project files
WORKDIR /project
COPY . /project
# 7. Default command
ENV QUARTO_PYTHON=/opt/venv/bin/python
CMD ["bash", "-c", "quarto render report.qmd && [...]"][1/8] to [8/8]ENV, SHELL, and CMD set metadata, so BuildKit does not number themFROM, ENV, SHELL# We pin the tag to "24.04" rather than "latest".
# "latest" changes over time, and a base image
# that changes breaks reproducibility.
FROM ubuntu:24.04
# DEBIAN_FRONTEND=noninteractive stops apt-get
# from asking questions during the build.
# Nobody is at the keyboard to answer them.
ENV DEBIAN_FRONTEND=noninteractive
# Use bash for the RUN steps below. Docker uses
# the simpler "sh" shell by default.
SHELL ["/bin/bash", "-c"]latest on the Ubuntu repository resolved to 24.04 last year, and today to 26.04FROM ubuntu:latest in September builds a different operating system in Decembertzdata is the package that would otherwise stop and ask you for a time zoneubuntu:24.04 is about 29MB compressed, and 108MB once unpackedENV persists into the running container, so docker run sees it tooSHELL makes every RUN below it use bash instead of the simpler sh-c lets you pass a string of commands to the shellubuntu image has no Python, so block 3 installs it.deb file, which needs a Debian-based system like UbuntuRUN apt-get installRUN is one layer
RUN lines here would give six layers and a bigger imageRUN: a later layer cannot shrink an earlier one--no-install-recommends skips optional extras| Package | Why |
|---|---|
ca-certificates |
HTTPS downloads fail without it |
wget |
Fetches the Quarto .deb |
git |
Some Quarto extensions want it |
python3.12 |
The interpreter, version 3.12.3 here |
python3.12-venv |
Ubuntu ships venv separately |
python3-pip |
Installs into the virtual environment |
dpkg --print-architecture and a .debRUN ARCH=$(dpkg --print-architecture) && \
wget -q "https://github.com/quarto-dev/quarto-cli/releases/download/v1.10.18/quarto-1.10.18-linux-${ARCH}.deb" && \
apt-get update && \
apt-get install -y --no-install-recommends ./quarto-1.10.18-linux-${ARCH}.deb && \
rm quarto-1.10.18-linux-${ARCH}.deb && \
apt-get clean && rm -rf /var/lib/apt/lists/*.deb from GitHub./ in front of the filename tells apt to install a local file rather than search the archiveapt-get update runs a second time because block 3 deleted the package listsdpkg --print-architecture printed arm64 on my Mac, amd64 on most Windows and Linux laptops
pip# Ubuntu protects its own Python installation,
# so we create a virtual environment at
# /opt/venv and install our packages there.
RUN python3 -m venv /opt/venv
# PATH tells the shell where to look for
# programs. Putting /opt/venv/bin first means
# "python" and "pip" refer to the virtual
# environment, without activating it by hand.
ENV PATH="/opt/venv/bin:$PATH"
# We copy requirements.txt on its own, before
# the rest of the project.
COPY requirements.txt /project/requirements.txt
RUN pip install --no-cache-dir -r /project/requirements.txtvenv, inside an image: ENV PATH replaces source .venv/bin/activatepip install and tells you to make a virtual environment
externally-managed-environment: Appendix 02requirements.txt is the hand-written file in your repository, the one we read a few slides ago
COPY takes it from the build context and puts a copy inside the image, and pip reads that copy--no-cache-dir stops pip keeping a copy of every wheel inside the imageCOPY requirements.txt comes before COPY . /project, so editing report.qmd keeps this slow step cached
WORKDIR, COPY ., CMD# WORKDIR sets the folder that later
# instructions run inside.
WORKDIR /project
# Copy everything else from your repository into
# the image, including report.qmd and the saved
# data snapshot in data/raw/.
# The .dockerignore file next to this one lists
# what stays out.
COPY . /project
# QUARTO_PYTHON points Quarto at the virtual
# environment's Python.
ENV QUARTO_PYTHON=/opt/venv/bin/python
CMD ["bash", "-c", "quarto render report.qmd && mkdir -p output && cp -a report.html output/ && echo 'Rendered report is in output/report.html'"]COPY . /project copies the build context: the folder named by the . in docker build
ENV PATH already points Quarto at /opt/venv. QUARTO_PYTHON makes it explicitCMD is three commands joined with &&: render, make output/, copy the HTML in/project and copying afterwards avoids thatCMD ["python", "hello.py"] used the same exec form: a list of strings, with no shell
bash, which gives us the shell we need for &&.dockerignore# Version control
.git
# Python environments and caches
.venv/
venv/
__pycache__/
# API keys never enter the image
.env
# Quarto working files and rendered output
.quarto/
*.quarto_ipynb
report.html
report_files/
output/
# Operating system clutter
.DS_Store
# NOTE: data/raw/ is deliberately absent. The
# snapshot must reach the image, because
# report.qmd reads it.Dockerfile, and is committed with it.git, the data, and a rendered report.gitignore keeps files out of your repository; .dockerignore keeps them out of your image
data/raw/ belongs in neitherreport.html out of the list and every render changes the context
COPY . /project even when your source did not movetime docker build --no-cache --progress=plain -t datasci350-project .
#5 [1/8] FROM docker.io/library/ubuntu:24.04
#5 CACHED
#6 [2/8] RUN apt-get update && apt-get install [...]
#6 DONE 36.0s
#7 [3/8] RUN ARCH=$(dpkg --print-architecture) [...]
#7 91.74 Setting up quarto (1.10.18) ...
#7 DONE 92.0s
#8 [4/8] RUN python3 -m venv /opt/venv
#8 DONE 1.3s
#9 [5/8] COPY requirements.txt /project/requirements.txt
#9 DONE 0.0s
#10 [6/8] RUN pip install --no-cache-dir -r [...]
#10 50.70 Successfully installed asttokens-3.0.2 [...]
#10 DONE 51.4s
#11 [7/8] WORKDIR /project
#11 DONE 0.1s
#12 [8/8] COPY . /project
#12 DONE 0.1s
#13 exporting to image
#13 DONE 10.5s
docker build [...] 3:11.77 total--progress=plain prints the whole log instead of the animated summary
--no-cache forces every step to run again, which is how I timed this[1/8] says CACHED because ubuntu:24.04 was already on my laptop
docker history shows what each instruction costdocker history --format "{{.Size}}\t{{.CreatedBy}}" \
datasci350-project
0B CMD ["bash" "-c" "quarto render [...]
0B ENV QUARTO_PYTHON=/opt/venv/bin/python
81.9kB COPY . /project
0B WORKDIR /project
421MB RUN pip install --no-cache-dir -r [...]
4.1kB COPY requirements.txt [...]
0B ENV PATH=/opt/venv/bin:/usr/local/sbin:[...]
15.6MB RUN python3 -m venv /opt/venv
459MB RUN ARCH=$(dpkg --print-architecture) [...]
169MB RUN apt-get update && apt-get install [...]
0B SHELL [/bin/bash -c]
0B ENV DEBIAN_FRONTEND=noninteractive
108MB ADD file:0387b3d029de8fa08[...] in /
0B [...]docker history lists one row per layer, newest at the topubuntu:24.04 base is only 108MBENV, SHELL and CMD only set metadata. WORKDIR gets a step number but adds 0BCACHED step skips the download as well as the workdocker images reports 1.52GB on disk and 347MB content, because content size is measured compressedtime docker run --rm \
-v "$(pwd)/output:/project/output" \
datasci350-project
Starting python3 kernel...Done
Executing 'report.quarto_ipynb'
Cell 1/1: 'fig-demo'...Done
[...]
Output created: report.html
Rendered report is in output/report.html
docker run [...] 3.237 total
ls -la output
-rw-r--r-- 1 dafreir staff 1314625 report.html| Part | What it does |
|---|---|
--rm |
Deletes the container when it exits |
-v host:container |
Connects a folder on your machine to one inside the container |
datasci350-project |
The image to start |
-v is short for volume: left of the colon is your laptop, right of it is the container$(pwd) is your current folder
${PWD}, the old Windows terminal %cd%$(pwd) gives an absolute path, which works everywhere. Recent Docker also accepts -v ./output:...docker run creates a container from the image and starts it
--rm deletes it again before you finish reading the output--rm then deleted the container, and the file went with itA render inside a container that has already been deleted leaves nothing on your laptop. Check ls output after every run, and check the timestamp on the file
report.qmd and built again:time docker build --progress=plain -t datasci350-project .
#4 transferring context: 4.95kB done
#5 [1/8] FROM docker.io/library/ubuntu:24.04
#5 DONE 0.0s
#6 [3/8] RUN ARCH=$(dpkg --print-architecture) [...]
#6 CACHED
#7 [2/8] RUN apt-get update && apt-get install [...]
#7 CACHED
#8 [5/8] COPY requirements.txt /project/requirements.txt
#8 CACHED
#9 [4/8] RUN python3 -m venv /opt/venv
#9 CACHED
#10 [6/8] RUN pip install --no-cache-dir -r [...]
#10 CACHED
#11 [7/8] WORKDIR /project
#11 CACHED
#12 [8/8] COPY . /project
#12 DONE 0.0s
docker build [...] 0.666 totalCOPY . /project saw new input, because only report.qmd changedrequirements.txt was untouched, so the 51-second pip install above it stayed cached# numbers are log IDs. The [n/8] labels give the real orderrequirements.txt instead, and the cache misses at [5/8]: pip and everything below run againdata/raw/ to .gitignore, committed, and cloned the result into a fresh folder:Test from a fresh clone in an empty folder. The folder where you wrote the code already has everything, including the files you forgot to commit
docker run --rm -v "$(pwd)/output:/project/output" \
datasci350-project-nodata
Starting python3 kernel...Done
Executing 'report.quarto_ipynb'
Cell 1/1: 'fig-demo'...
An error occurred while executing the following cell:
[...]
data = duckdb.sql("SELECT country, year, value "
"FROM 'data/raw/life_expectancy.csv' [...]").df()
[...]
IOException: IO Error: No files found that match
the pattern "data/raw/life_expectancy.csv"
WARN: Error encountered when rendering files
exit code: 1git add data/raw/ and git commitgit ls-files data/raw, which should print two file namesduckdb==1.5.5 to duckdb==1.5.9 and built:#10 [6/8] RUN pip install --no-cache-dir -r [...]
#10 0.818 Collecting pandas==3.0.5 [...]
#10 1.302 ERROR: Could not find a version that
satisfies the requirement duckdb==1.5.9
(from versions: [...] 1.5.4, 1.5.5, 1.5.6.dev36,
1.6.0.dev379, 2.0.0.dev2609121639)
#10 1.302 ERROR: No matching distribution found
for duckdb==1.5.9
Dockerfile:85
--------------------
84 | COPY requirements.txt /project/requirements.txt
85 | >>> RUN pip install --no-cache-dir -r /project/requirements.txt
--------------------
ERROR: failed to build: failed to solve: process
"/bin/bash -c pip install [...]" did not complete
successfully: exit code: 1[6/8] names the instructionDockerfile:85 with the >>> marker names the linedocker images shows no new entry[1/8] to [5/8] stay cached, so fixing the pin costs you the pip step and nothing moreNo matching distribution found. Here it is in a real builddocker run -it and what it tells youCMD, so this starts a shell instead of rendering:-i keeps input open and -t gives you a terminal: an interactive sessiondocker exec -it <container> bash does the same on a container already runningrequirements.txt and rebuild insteaddocker logs <container> prints the output of a container you started without watching$ docker ps -a
CONTAINER ID IMAGE STATUS
5594bb583fbd hello-world Exited (0) 2 hours ago
$ docker system df
TYPE TOTAL ACTIVE SIZE RECLAIMABLE
Images 7 1 10.88GB 4.029GB (37%)
Containers 1 0 0B 0B
Build Cache 37 0 3.046GB 1.403GB
$ docker rm 5594bb583fbd
$ docker rmi datasci350-project-nodata
[...]
$ docker system prune
[...]
Total reclaimed space: 1.931GB
$ docker system df
Images 4 0 3.217GB 1.97GB (61%)
Containers 0 0 0B 0B
Build Cache 8 0 1.643GB 0B| Command | What it removes |
|---|---|
docker ps -a |
Nothing, it lists containers including exited ones |
docker rm <id> |
One stopped container |
docker rmi <image> |
One image |
docker system df |
Nothing, it reports what is using space |
docker system prune |
Stopped containers, unused networks, dangling images, build cache |
docker system prune -a |
All of the above plus every image not in use |
--rm on docker run is what stops the container list growing in the first placedocker system prune -a deletes every image no running container is using, including the one you spent three minutes building. Read the prompt before you type y
quay.io/jupyter/scipy-notebook is 5.28GB on disk (1.24GB of content), and took 23 minutes to pull herer-ver, rstudio, tidyverse, verse, geospatialThe stack builds on itself: ubuntu → docker-stacks-foundation → base-notebook → minimal-notebook → scipy-notebook. Docs
For the project, write your own Dockerfile: a 5GB image for one CSV reader is a poor trade
buildx command, which builds for both kinds of chip:docker pull and docker run it, on Windows, Linux or any Mac, with no buildDo this at home with your clone of the starter. Step 2 takes about a minute.
seaborn==0.13.2 under duckdb==1.5.5 in requirements.txtdocker build -t datasci350-project .CACHEDreport.qmd, add import seaborn as sns under import duckdbprint(f"seaborn {sns.__version__}") under the Loaded printdocker run --rm -v "$(pwd)/output:/project/output" datasci350-projectoutput/report.html and find the line seaborn 0.13.2Solution and captured output: Appendix 01
What to look for:
[1/8] to [4/8] stay CACHED: Ubuntu, apt, Quarto, and the empty venv[5/8] COPY requirements.txt, because that file changedpip install runs again for 50.0 secondsCOPY . /project, so it finishes in secondspip freeze inside the image goes from 53 packages to 54CMD.dockerignore: keeps .git, output/, and .env out of the image.env stays out of the image, through .dockerignoreCOPY requirements.txt above COPY . /project, so a rebuild takes a second-v: without it the report dies with the containerdocker run -it ... bash: look inside when a render failsdocker system prune: keeps the disk under controlThe whole lecture fits in one sentence: git clone, docker build, docker run -v, open the report. If that works from a fresh folder, the container part of your project is finished
Thursday 26 November is Thanksgiving. No class
Lecture 26, on 1 December, is revision. Bring your project repository and your questions, and we will build containers together in class
Quiz 05, on 3 December, covers Lectures 21, 22, 23, and 25. Assignment 10 is due the same day
Before then:
git ls-files data/raw. It should print two file namesgit shortlog -sn. Every group member should appearStep 1, in requirements.txt:
Steps 4 and 5, in the demo chunk of report.qmd:
import duckdb
import seaborn as sns
import matplotlib.pyplot as plt
data = duckdb.sql(
"SELECT country, year, value "
"FROM 'data/raw/life_expectancy.csv' "
"ORDER BY country, year"
).df()
print(f"Loaded {len(data)} rows covering "
f"{data['country'].nunique()} countries.")
print(f"seaborn {sns.__version__}")Step 7, in the rendered report:
CACHED, back to 347MB and 53 packagesSteps 2 and 3, the rebuild on my laptop:
#5 [1/8] FROM docker.io/library/ubuntu:24.04
#5 DONE 0.0s
#6 [3/8] RUN ARCH=$(dpkg --print-architecture) [...]
#6 CACHED
#7 [2/8] RUN apt-get update && apt-get install [...]
#7 CACHED
#8 [4/8] RUN python3 -m venv /opt/venv
#8 CACHED
#9 [5/8] COPY requirements.txt /project/requirements.txt
#9 DONE 0.0s
#10 [6/8] RUN pip install --no-cache-dir -r [...]
#10 18.77 Downloading seaborn-0.13.2-[...].whl (294 kB)
#10 49.45 Successfully installed [...] seaborn-0.13.2 [...]
#10 DONE 50.0s
#11 [7/8] WORKDIR /project
#11 DONE 0.1s
#12 [8/8] COPY . /project
#12 DONE 0.1s
docker build [...] 57.726 total#9 forced everything below it to run againIOException: No files found that match [...] life_expectancy.csv
The snapshot never reached the repository. Run git add data/raw/, commit, and check with git ls-files data/raw.
No matching distribution found for duckdb==1.5.9
That version does not exist on PyPI. Run pip index versions duckdb and pin one it lists.
ls: output: No such file or directory
You ran the container without -v. Add -v "$(pwd)/output:/project/output", or ${PWD} in PowerShell. On Linux the message reads ls: cannot access 'output'.
error: externally-managed-environment
You ran pip install against Ubuntu’s own Python. Create a virtual environment and put /opt/venv/bin first in PATH.
failed to read dockerfile: open Dockerfile
You are in the wrong folder. The . in docker build -t name . is where Docker looks, so cd into the folder with the Dockerfile.
NoSuchKernel: No such kernel named python3
ipykernel is missing from requirements.txt, so Quarto cannot start a kernel. Put ipykernel==7.2.0 back and rebuild.
toomanyrequests: pull rate limit
Docker Hub allows 100 anonymous pulls every six hours from one address, and 200 once you log in. Run docker login and build again.
# 1. Base image
# FROM chooses the starting point: Ubuntu 24.04,
# the system you saw in the containers lecture.
# We pin the tag to "24.04" rather than "latest".
FROM ubuntu:24.04
# 2. Build-time settings
ENV DEBIAN_FRONTEND=noninteractive
SHELL ["/bin/bash", "-c"]
# 3. System packages
# We do not pin apt versions here. Ubuntu replaces
# a version whenever it ships a security fix, and a
# pin that no longer exists stops the build. The
# base image tag and requirements.txt carry the pins.
RUN apt-get update && apt-get install -y --no-install-recommends \
ca-certificates \
wget \
git \
python3.12 \
python3.12-venv \
python3-pip && \
apt-get clean && rm -rf /var/lib/apt/lists/*
# 4. Quarto
# Quarto is not in the Ubuntu package list, so we
# download the .deb from GitHub. The version is
# written out in full, so the image you build is
# the image everyone else builds.
RUN ARCH=$(dpkg --print-architecture) && \
wget -q "https://github.com/quarto-dev/quarto-cli/releases/download/v1.10.18/quarto-1.10.18-linux-${ARCH}.deb" && \
apt-get update && \
apt-get install -y --no-install-recommends ./quarto-1.10.18-linux-${ARCH}.deb && \
rm quarto-1.10.18-linux-${ARCH}.deb && \
apt-get clean && rm -rf /var/lib/apt/lists/*# 5. Python environment
# Ubuntu protects its own Python installation, so
# we create a virtual environment at /opt/venv and
# install our packages there.
RUN python3 -m venv /opt/venv
# PATH tells the shell where to look for programs.
ENV PATH="/opt/venv/bin:$PATH"
# We copy requirements.txt on its own, before the
# rest of the project. Docker caches each layer.
COPY requirements.txt /project/requirements.txt
RUN pip install --no-cache-dir -r /project/requirements.txt
# 6. Project files
# WORKDIR sets the folder that later instructions
# run inside.
WORKDIR /project
# Copy everything else from your repository into
# the image, including report.qmd and the saved
# data snapshot in data/raw/.
# The .dockerignore file next to this one lists
# what stays out.
COPY . /project
# 7. Default command
# We render first and copy afterwards. Quarto
# clears its output folder before writing, and it
# cannot clear a folder Docker mounted from outside.
# QUARTO_PYTHON points Quarto at the venv's Python.
ENV QUARTO_PYTHON=/opt/venv/bin/python
CMD ["bash", "-c", "quarto render report.qmd && mkdir -p output && cp -a report.html output/ && echo 'Rendered report is in output/report.html'"]eb4b6f8docker run --rm -v "$(pwd)/output:/project/output" datasci350-project twenty times is how typos happen:compose.yaml next to your Dockerfile, though docker-compose.yml still works:./output for you, so $(pwd) never appearsdown removes the container and the network it createddocker compose with a space is the current command; docker-compose with a hyphen is the legacy one