DATASCI 350 - Data Science Computing

Lecture 10 - Reproducible Research and Literate Programming

Danilo Freire

Department of Data and Decision Sciences
Emory University

Nice to see you all again! 😊

Recap of our last lectures

Before the quiz, we covered

  • The command line: navigation, file management, text tools, pipes, and scripts
  • Git: repositories, staging, commits, branches, and merges
  • GitHub: remotes, forks, pull requests, issues, and the gh CLI
  • Last class you put all of it to work in quiz 01
  • Congratulations, that was the hardest setup work of the course 🎉

I hope the image doesn’t bring back bad memories! :)

Today’s lecture

A new module starts today: can anyone check your work?

  • The reproducibility crisis: why published research often fails when re-run
  • Habits that prevent it: project structure, raw data, paths, seeds, environments
  • Literate programming: text, code and results in one file
  • Quarto: installing it and writing your first document
  • Next class: PDFs, citations, websites and parameterised reports

Keep your copy of the course up to date

  • Did you fork the course repository? Your fork does not update by itself
  • Please sync it often! 🤓
  • On GitHub: open your fork, click Sync fork, then Update branch
  • Then, in your terminal, inside your local copy: git pull
  • Did you clone the repository directly? Then git pull is enough
  • Keep your own work in a separate folder, so a sync never clashes with it

The same from the terminal, with the gh CLI:

# Update your fork on GitHub
gh repo sync your-username/datasci350

# Update the copy on your laptop
cd datasci350
git pull

Tweets of the day

Why is reproducible research so important? 🤔

The reproducibility crisis

  • Reproducibility: someone using your data and code gets your results
  • A low bar, yet many published findings fail it
  • The reproducibility crisis hit psychology, medicine and economics hardest. No empirical field escaped
  • Causes: small samples, p-hacking, selective reporting, unshared code
  • This course focuses on code and data. Most computational results fail for a mundane reason

Amy Cuddy’s “power posing”: 78 million TED views, but its hormone effects failed to replicate

How bad is it? Two famous papers

Medicine

  • Ioannidis (2005), one of the most cited papers in medicine
  • For most study designs, a positive finding is more likely false than true
  • Causes: small samples, flexible analyses, financial interests, many teams chasing one question

Psychology

  • Open Science Collaboration (2015): 270 researchers re-ran 100 psychology experiments
  • Originals: 97% significant. Replications: 36%, with effects halved
  • Journals responded with pre-registration, open data and registered reports

We all write code, so we are fine, right?

  • Trisovic et al. (2022) re-ran 9,078 R files from 2,109 replication packages on Harvard Dataverse
  • 74% failed on the first run. 56% still failed after automated fixes
  • Common causes: missing packages, hardcoded paths, changed functions, missing data
  • Authors chose to deposit these files with their papers
  • Imagine the code they did not publish 😅

Five selfish reasons to work reproducibly

Markowetz (2015) argues reproducibility helps you, too:

  1. Avoid disaster: catch your errors before a reviewer, a retraction or a job interview does
  2. Easier writing: when the data change, every number and figure updates
  3. Reviewers see it your way: they run your analysis instead of guessing
  4. Continuity: a new collaborator, or future-you, can pick it up tomorrow
  5. Reputation: working, shared code shows you know what you are doing

Reason 2 will save you hours in this course. Reason 1 returns in a few slides, with two famous economists

What does it take to be reproducible?

The reproducibility ladder

  • These words mean different things:
  • Re-runnable: the code runs again on your machine, today
  • Reproducible: same code and data give the same results, on any machine
  • Reusable: others can apply your work to new questions
  • Replicable: new, independent data give the same finding
  • Each step is harder than the one below it
  • Replicability needs data you may never get, so it is mostly outside your control

Reproducibility is fully in your control, so this course grades you on it

harder

▲

│

│

easier

replicable

reusable

reproducible

re-runnable

In one large study, most deposited R files did not clear the first step

What a reproducible project needs

A result is reproducible only if all four of these travel together:

  1. Code: every step from raw data to final number, runnable
  2. Data: the raw inputs, or a script that fetches them
  1. Environment: the language and package versions
  2. Documentation: what the project does and the run order

Leave one out and it fails:

  • No data: readers cannot check your code
  • No environment: it runs today and breaks next year
  • No documentation: nobody knows which script runs first

Project structure

  • One folder per project, with the same layout every time
  • We use it for the rest of the course, and the final project requires it
my-project/
├── data/
│   ├── raw/          # never edited by hand
│   └── clean/        # built by scripts
├── scripts/          # the code, in run order
├── output/           # figures and tables
└── README.md         # what and how
  • data/raw/ holds your inputs. Everything else can be rebuilt from it
  • data/raw/ is read-only: download once, never edit
  • Numbered scripts (01-clean.py, 02-analyse.py) show the run order
  • data/clean/ and output/ are disposable: the scripts rebuild them
  • Put them in .gitignore and commit only sources
  • README.md explains the project to the next person, usually you in six months

Delete everything except data/raw/ and scripts/. Can you rebuild the project?

Every project needs a README

  • GitHub shows a README on every repository page
  • It answers what is this, how do I rebuild it and in what order
  • List each data source and its download date, so stale data is easy to spot
  • List the rebuild commands in run order, ready to paste
  • Note the language and package versions (more in a minute)
  • Ten lines are enough. Start early and keep it current
# Rainfall and crop yields in Brazil

Analysis of rainfall and maize yields, 2000-2024.

## Data
- data/raw/rainfall.csv: INMET, downloaded 12 Sep 2026
- data/raw/yields.csv: IBGE, downloaded 12 Sep 2026

## How to rebuild
Run the scripts in order:

    python scripts/01-clean.py
    python scripts/02-analyse.py
    quarto render report.qmd

## Environment
Python 3.13, pandas 3.0, matplotlib 3.10

The most famous spreadsheet error in economics

  • Reinhart and Rogoff (2010), “Growth in a Time of Debt”: above 90% debt-to-GDP, growth turns negative
  • After 2008, it became the standard citation for austerity
  • In 2013, graduate student Thomas Herndon could not reproduce it and asked for the spreadsheet
  • One formula averaged rows 30 to 44, not 30 to 49, dropping five countries
  • With two other fixes (excluded data, odd weighting), growth above 90% was +2.2%, not -0.1%
  • The analysis sat in a private spreadsheet, so three years of policy debate passed before anyone noticed
  • A public script and dataset would have caught it in an afternoon

The absolute path problem

A top reason code fails on another machine:

analysis.py
import pandas as pd

df = pd.read_csv("/Users/ana/Desktop/project/data/raw/rainfall.csv")

Anyone else who runs it gets:

FileNotFoundError: [Errno 2] No such file
or directory:
'/Users/ana/Desktop/project/data/raw/rainfall.csv'
  • An absolute path starts at / and names one spot on one machine
  • Only Ana has /Users/ana, so the script fails everywhere else
  • A relative path starts from the working directory, like pwd in the shell
  • A program inherits it from where you launch it, so “it works for me” proves little
  • Hardcoded paths were a top cause of failure in the Trisovic study

Relative paths fix it

Write paths relative to the project folder:

analysis.py
import pandas as pd

df = pd.read_csv("data/raw/rainfall.csv")

Or build them with pathlib:

analysis.py
from pathlib import Path

raw = Path("data") / "raw" / "rainfall.csv"
df = pd.read_csv(raw)
  • pathlib picks the right separator on macOS, Linux and Windows
  • Quarto runs each document from the folder holding the .qmd file
  • Avoid os.chdir(): it changes state mid-script and makes the code hard to follow

If your code contains your username, it is not reproducible

Randomness needs a seed

Without a seed, every run gives new numbers:

import numpy as np

print(np.random.normal(size=3).round(3))
print(np.random.normal(size=3).round(3))
[ 0.567 -0.359  0.094]
[ 2.073 -0.402  1.267]

With a seed, everyone gets the same sequence:

np.random.seed(350)
print(np.random.normal(size=3).round(3))

np.random.seed(350)
print(np.random.normal(size=3).round(3))
[ 1.697 -0.533 -0.311]
[ 1.697 -0.533 -0.311]
  • np.random is not truly random. It is a pseudorandom generator: a formula that walks a fixed sequence
  • The seed sets the starting point, which fixes the whole sequence
  • Newer code uses a generator object:
rng = np.random.default_rng(350)
rng.normal(size=3)
  • Python’s random has its own generator. scikit-learn needs random_state=
  • Randomness hides in sampling, train-test splits and model starting weights
  • Seed every random draw

Record your environment

  • Packages change: new arguments, flipped defaults, removed methods
  • Code from 2024 may give different numbers in 2026, or not run. This is version drift
  • Record your versions so others can match them

At minimum, note them in your README:

import sys, numpy, pandas

print("python:", sys.version.split()[0])
print("numpy :", numpy.__version__)
print("pandas:", pandas.__version__)
python: 3.13.13
numpy : 2.4.3
pandas: 3.0.2

pip lists every package. python -m uses the Python you actually run:

Terminal
python -m pip list
Package         Version
--------------- -------
jupyter_client  8.6.3
matplotlib      3.10.8
numpy           2.4.3
pandas          3.0.2

A version list helps diagnose failures. Module 08 adds tools that prevent them: virtual environments, uv and containers

Literate programming

Literate programming

  • Now a tool for the biggest habit: text, code and results in one file
  • Donald Knuth coined literate programming in 1984: write for humans first, machines second
  • Two operations split the file:
    • Weaving: the readable document (HTML, PDF)
    • Tangling: the executable code
  • Explaining your steps catches bad logic early
  • Numbers come from the code, so they cannot go stale

From LaTeX to Quarto

  • LaTeX came first: beautiful output, verbose syntax, steep learning curve (I’ve been there 😅)
  • Markdown kept the ideas with a simpler syntax: **bold**, - list, [link](url)
  • R Markdown added R chunks, and Jupyter did the same for Python
  • Both struggle with mixed languages or output formats
  • Quarto (2022, Posit) rebuilt R Markdown for any language: Python, R, Julia, Observable JS
  • One .qmd renders to HTML, PDF, Word, slides, websites, books and dashboards
  • Pandoc does the final conversion
  • These slides, the course website and every course PDF use Quarto

Part of a LaTeX template. This is why Markdown happened

What Quarto can produce

  • Clockwise from top left: a company report, an interactive dashboard, a website, and a book
  • All four are plain text, so they work well with Git
  • Only the format: line changes
  • Click any image to zoom in

How does Quarto work?

%%{
  init: {
    "theme": "dark",
    "themeCSS": ".label foreignObject, .cluster-label foreignObject { font-size: 90%; overflow: visible; }"
  }
}%%
flowchart LR
  A1[qmd] --> C{"knitr<br>(R)"}
  A1[qmd] --> B{"Jupyter<br>(Python)"}
  A2[ipynb] --> B{"Jupyter<br>(Python)"}
  B --> D[md]
  C --> D[md]
  D --> E{Pandoc}
  E --> F[pdf]
  E --> G[docx]
  E --> H[html]
  E --> I[...]

  subgraph engine [Engine]
  B
  C
  end

  • Quarto reads .qmd files and Jupyter notebooks (.ipynb). A YAML cell on top is optional
  • Pipeline: .qmd → engine (Jupyter or knitr) → .md → Pandoc → output
  • Pandoc turns Markdown into HTML, PDF, Word and dozens of other formats
  • Quarto calls Pandoc for you, with the options from your YAML header

Getting Quarto running 🛠️

Installing Quarto

  • Download the installer on the left, or use a package manager
  • On macOS, with Homebrew:
Terminal
brew install --cask quarto
  • On Windows, install Quarto inside WSL, next to your shell and Git
  • PDFs also need LaTeX. Install the lightweight TinyTeX:
Terminal
quarto install tinytex

Did it work? Ask Quarto itself

Terminal
quarto check
Quarto 1.9.38
[✓] Checking environment information...
[✓] Checking versions of quarto binary dependencies...
      Pandoc version 3.8.3: OK
      Dart Sass version 1.87.0: OK
[✓] Checking Quarto installation......OK
      Version: 1.9.38
[✓] Checking tools....................OK
      TinyTeX: v2026.04
[✓] Checking LaTeX....................OK
[✓] Checking basic markdown render....OK
[✓] Checking Python 3 installation....OK
      Version: 3.13.13
      Jupyter: 5.9.1

Run this first when something breaks. Your versions will differ from mine

Write in whatever editor you like

  • A .qmd file is plain text: VS Code, RStudio, Jupyter Lab, Neovim, even Notepad work
  • We use VS Code with the Quarto extension: highlighting, completion, a render button and live preview
  • Install it from the Extensions marketplace
  • Its visual editor renders Markdown as you type. Try both modes

Rendering from the terminal

quarto --help lists every command. These do most of the work:

Terminal
## Render a file to its default format
quarto render report.qmd

## Render to a specific format
quarto render report.qmd --to pdf

## Render, open in the browser, and re-render on every save
quarto preview report.qmd
  • render builds the document once, when you are done
  • preview keeps a live version open while you write
  • Output lands next to the source (report.qmd → report.html), unless a project sets output-dir

Your first Quarto document

Anatomy of a Quarto document

  • A Quarto document has a YAML header, code chunks and Markdown text
  • We look at each in turn. Markdown in depth comes next lecture

The YAML header

---
title: "Rainfall report"
author: "Your name"
date: 2026-10-06
format:
  html:
    toc: true
    code-fold: true
execute:
  warning: false
jupyter: python3
---
  • YAML stands for “YAML Ain’t Markup Language” (a recursive joke). It describes data
  • It sits at the top, between two --- lines
  • key: value pairs, where indentation creates nesting. Use spaces: tabs are illegal for indentation
  • format: html is the short form. Indent under it to add options
  • execute: sets defaults for every chunk
  • Quote values that contain a colon: title: "Quarto: a first look"
  • _quarto.yml, GitHub Actions and Docker Compose use the same syntax

Code chunks

```{python}
#| label: fig-rainfall
#| echo: false
#| warning: false

import matplotlib.pyplot as plt

plt.bar(rain["month"], rain["mm"])
plt.ylabel("Rainfall (mm)")
plt.show()
```
  • Three backticks and {python} open a chunk; three backticks close it
  • On render, the chunk runs and its output goes into the document
  • #| lines are chunk options, in YAML syntax
  • Here they label the figure and hide the code and warnings

The chunk options you will use most

Option Effect
echo: false hide the code, keep the output
eval: false show the code, do not run it
include: false run the code, show nothing
warning: false hide warnings
label: fig-rain name the chunk for cross-references
fig-cap: "..." put a caption under the figure
layout-ncol: 2 arrange the outputs in columns
  • #| options apply to one chunk
  • Document-wide defaults go in the YAML header:
execute:
  echo: false
  warning: false
  • Chunk options override these, so you can hide all code but show one chunk
  • Cross-references need a prefixed label: fig-, tbl-, eq-, sec-, lst-
  • @fig-rain then prints “Figure 1”, numbered for you
  • Full list in the Quarto documentation

Chunk options in action

layout-ncol: 2 puts two plots side by side; echo: false hides the code:

```{python}
#| echo: false
#| layout-ncol: 2

plt.bar(rain["month"], rain["mm"])
plt.show()

plt.plot(rain["month"], rain["mm"])
plt.show()
```

What the reader sees:

Both figures come from the code at render time. Change the data and they update

Engines: what actually runs your code

  • Quarto runs no code itself. It hands chunks to an engine
  • Jupyter runs Python and any language with a Jupyter kernel. knitr runs R
  • This course uses Jupyter. These slides render with jupyter: python3
  • Each render starts one fresh session and runs the chunks top to bottom
  • Chunks share state like notebook cells: a variable from chunk one exists in chunk ten
  • So every render is a restart-and-run-all test
  • A document that renders has just run its analysis from scratch
---
title: "My report"
format: html
jupyter: python3
---

Quarto can usually detect the engine. Declaring it tells readers what to install.

To use R, write {r} chunks and Quarto picks knitr

A minimal Python document

---
title: "Palmer Penguins Demo"
format:
  html:
    code-fold: true
jupyter: python3
---

## Meet Quarto

Quarto enables you to weave together content and
executable code into a finished document. To learn
more about Quarto see <https://quarto.org>.

```{python}
#| echo: false

import seaborn as sns
from palmerpenguins import load_penguins
sns.set_style('whitegrid')

penguins = load_penguins()

g = sns.lmplot(x="flipper_length_mm",
               y="body_mass_g",
               hue="species",
               height=7,
               data=penguins,
               palette=['#FF8C00','#159090','#A034F0'])
g.set_xlabels('Flipper Length')
g.set_ylabels('Body Mass')
  • Thirty lines, with everything from today’s lecture
  • The YAML header sets the title, format and engine
  • code-fold: true hides code behind a toggle, but echo: false removes this chunk’s code
  • The prose is Markdown, with a ## heading
  • One chunk loads penguin data and draws a figure (needs pip install palmerpenguins)
  • The next slide shows the result

The rendered page

  • Title from the YAML, heading and prose from the Markdown, figure from the chunk
  • echo: false wins over code-fold: true: only the figure shows
  • Change the data and render again: the figure and numbers follow
  • The page can only show figures its own code produced

Try it yourself!

Create a Quarto document from scratch and render it to HTML.

  1. Make a new file called weather.qmd
  2. Start it with this YAML header:
---
title: "Rainfall report"
author: "Your name"
format: html
jupyter: python3
---
  1. Add a ## heading and a short paragraph
  2. Add a chunk that builds the table below and prints it
  3. Add a second chunk that draws a bar chart, with its code hidden
  4. Render the file and open the result

Use this data so everyone gets the same answer:

import pandas as pd

rain = pd.DataFrame({
    "month": ["Jan", "Feb", "Mar", "Apr"],
    "mm": [241, 215, 180, 96]
})

Stuck? See the chunk options slide. quarto check diagnoses a broken installation

Answer: Appendix 01

Summary

  • The reproducibility crisis is real. Most computational results fail because the code does not run
  • The ladder: re-runnable → reproducible → reusable → replicable. This course gets you to step two
  • Code, data, environment and documentation travel together
  • Habits: untouched raw data, cleaning in code, relative paths, seeds, recorded versions, a README
  • Literate programming keeps text, code and results in one file
  • Quarto is our tool: one plain-text .qmd, one render command, many formats

Next class

  • Quarto in practice: Markdown, PDFs and automatic citations
  • freeze: rendering without re-running everything
  • Slides (like these) and websites, published from GitHub
  • Parameterised reports: one file, one report per country
  • Bring a working Quarto installation. Today’s exercise is the test 🤓

Next class: websites like this one, built with Quarto

Additional materials

And that’s all for today! 🥳

Thank you very much and see you soon! 😊 🙏🏼

Appendix 01: Solution to the exercise

The complete weather.qmd:

---
title: "Rainfall report"
author: "Your name"
format: html
jupyter: python3
---

## Rainfall in the first four months

Rainfall fell steadily from January to April, with April
receiving less than half of January's total.

```{python}
import pandas as pd

rain = pd.DataFrame({
    "month": ["Jan", "Feb", "Mar", "Apr"],
    "mm": [241, 215, 180, 96]
})

rain
```

```{python}
#| echo: false

import matplotlib.pyplot as plt

plt.bar(rain["month"], rain["mm"], color="#1B3A6B")
plt.ylabel("Rainfall (mm)")
plt.show()
```

Then render it:

Terminal
quarto render weather.qmd

#| echo: false hides the code and keeps the output, as task 5 asked.

The chart your second chunk produces:

Back to the main text