[ 0.567 -0.359 0.094]
[ 2.073 -0.402 1.267]
Lecture 10 - Reproducible Research and Literate Programming
gh CLIA new module starts today: can anyone check your work?
git pullgit pull is enoughAmy Cuddy’s “power posing”: 78 million TED views, but its hormone effects failed to replicate
Medicine
Psychology
R files from 2,109 replication packages on Harvard DataverseMarkowetz (2015) argues reproducibility helps you, too:
Reason 2 will save you hours in this course. Reason 1 returns in a few slides, with two famous economists
Reproducibility is fully in your control, so this course grades you on it
harder
▲
│
│
easier
replicable
reusable
reproducible
re-runnable
In one large study, most deposited R files did not clear the first step
A result is reproducible only if all four of these travel together:
Leave one out and it fails:
data/raw/ holds your inputs. Everything else can be rebuilt from itdata/raw/ is read-only: download once, never edit01-clean.py, 02-analyse.py) show the run orderdata/clean/ and output/ are disposable: the scripts rebuild them.gitignore and commit only sourcesREADME.md explains the project to the next person, usually you in six monthsDelete everything except data/raw/ and scripts/. Can you rebuild the project?
# Rainfall and crop yields in Brazil
Analysis of rainfall and maize yields, 2000-2024.
## Data
- data/raw/rainfall.csv: INMET, downloaded 12 Sep 2026
- data/raw/yields.csv: IBGE, downloaded 12 Sep 2026
## How to rebuild
Run the scripts in order:
python scripts/01-clean.py
python scripts/02-analyse.py
quarto render report.qmd
## Environment
Python 3.13, pandas 3.0, matplotlib 3.10A top reason code fails on another machine:
Anyone else who runs it gets:
/ and names one spot on one machine/Users/ana, so the script fails everywhere elsepwd in the shellWrite paths relative to the project folder:
Or build them with pathlib:
pathlib picks the right separator on macOS, Linux and Windows.qmd fileos.chdir(): it changes state mid-script and makes the code hard to followIf your code contains your username, it is not reproducible
Without a seed, every run gives new numbers:
[ 0.567 -0.359 0.094]
[ 2.073 -0.402 1.267]
With a seed, everyone gets the same sequence:
np.random is not truly random. It is a pseudorandom generator: a formula that walks a fixed sequencerandom has its own generator. scikit-learn needs random_state=At minimum, note them in your README:
pip lists every package. python -m uses the Python you actually run:
A version list helps diagnose failures. Module 08 adds tools that prevent them: virtual environments, uv and containers
**bold**, - list, [link](url).qmd renders to HTML, PDF, Word, slides, websites, books and dashboardsformat: line changes%%{
init: {
"theme": "dark",
"themeCSS": ".label foreignObject, .cluster-label foreignObject { font-size: 90%; overflow: visible; }"
}
}%%
flowchart LR
A1[qmd] --> C{"knitr<br>(R)"}
A1[qmd] --> B{"Jupyter<br>(Python)"}
A2[ipynb] --> B{"Jupyter<br>(Python)"}
B --> D[md]
C --> D[md]
D --> E{Pandoc}
E --> F[pdf]
E --> G[docx]
E --> H[html]
E --> I[...]
subgraph engine [Engine]
B
C
end
.qmd files and Jupyter notebooks (.ipynb). A YAML cell on top is optional.qmd → engine (Jupyter or knitr) → .md → Pandoc → outputQuarto 1.9.38
[✓] Checking environment information...
[✓] Checking versions of quarto binary dependencies...
Pandoc version 3.8.3: OK
Dart Sass version 1.87.0: OK
[✓] Checking Quarto installation......OK
Version: 1.9.38
[✓] Checking tools....................OK
TinyTeX: v2026.04
[✓] Checking LaTeX....................OK
[✓] Checking basic markdown render....OK
[✓] Checking Python 3 installation....OK
Version: 3.13.13
Jupyter: 5.9.1Run this first when something breaks. Your versions will differ from mine
.qmd file is plain text: VS Code, RStudio, Jupyter Lab, Neovim, even Notepad workquarto --help lists every command. These do most of the work:
Terminal
render builds the document once, when you are donepreview keeps a live version open while you writereport.qmd → report.html), unless a project sets output-dir--- lineskey: value pairs, where indentation creates nesting. Use spaces: tabs are illegal for indentationformat: html is the short form. Indent under it to add optionsexecute: sets defaults for every chunktitle: "Quarto: a first look"_quarto.yml, GitHub Actions and Docker Compose use the same syntax{python} open a chunk; three backticks close it#| lines are chunk options, in YAML syntax| Option | Effect |
|---|---|
echo: false |
hide the code, keep the output |
eval: false |
show the code, do not run it |
include: false |
run the code, show nothing |
warning: false |
hide warnings |
label: fig-rain |
name the chunk for cross-references |
fig-cap: "..." |
put a caption under the figure |
layout-ncol: 2 |
arrange the outputs in columns |
#| options apply to one chunkfig-, tbl-, eq-, sec-, lst-@fig-rain then prints “Figure 1”, numbered for youlayout-ncol: 2 puts two plots side by side; echo: false hides the code:
What the reader sees:


Both figures come from the code at render time. Change the data and they update
jupyter: python3---
title: "Palmer Penguins Demo"
format:
html:
code-fold: true
jupyter: python3
---
## Meet Quarto
Quarto enables you to weave together content and
executable code into a finished document. To learn
more about Quarto see <https://quarto.org>.
```{python}
#| echo: false
import seaborn as sns
from palmerpenguins import load_penguins
sns.set_style('whitegrid')
penguins = load_penguins()
g = sns.lmplot(x="flipper_length_mm",
y="body_mass_g",
hue="species",
height=7,
data=penguins,
palette=['#FF8C00','#159090','#A034F0'])
g.set_xlabels('Flipper Length')
g.set_ylabels('Body Mass')code-fold: true hides code behind a toggle, but echo: false removes this chunk’s code## headingpip install palmerpenguins)Create a Quarto document from scratch and render it to HTML.
weather.qmd## heading and a short paragraphUse this data so everyone gets the same answer:
Stuck? See the chunk options slide. quarto check diagnoses a broken installation
.qmd, one render command, many formatsfreeze: rendering without re-running everythingThe complete weather.qmd:
---
title: "Rainfall report"
author: "Your name"
format: html
jupyter: python3
---
## Rainfall in the first four months
Rainfall fell steadily from January to April, with April
receiving less than half of January's total.
```{python}
import pandas as pd
rain = pd.DataFrame({
"month": ["Jan", "Feb", "Mar", "Apr"],
"mm": [241, 215, 180, 96]
})
rain
```
```{python}
#| echo: false
import matplotlib.pyplot as plt
plt.bar(rain["month"], rain["mm"], color="#1B3A6B")
plt.ylabel("Rainfall (mm)")
plt.show()
```Then render it:
#| echo: false hides the code and keeps the output, as task 5 asked.
The chart your second chunk produces:
