Lecture 19 - Working with APIs in Practice
2xx worked, 4xx is your mistake, 5xx is theirs200, so look at the data too{ } and arrays [ ]current → temperature_2mcurl -o saved a reply to a filecurlToday we do it all in Python
1. Reading JSON in Python
2. requests, step by step
requests.getrequests build the URL3. From JSON to a DataFrame
4. Your project’s pull script
pull_data.py, part by partdict) is Python’s JSON object: {"key": value}course["code"][a, b, c]years[0]course["instructor"]["name"] is the path instructor → name from last classwith open(...) as f: meansThe file is last class’s curl -o download: wb_gdp_bra.json. Save it in a data/ folder next to your notebook
open("data/wb_gdp_bra.json") finds the file and prepares it for readingas f gives the open file a short name, fjson.load(f) reads the JSON text and turns it into Python objectsjson.load and json.loadsjson.load(f) reads JSON from an open filejson.loads(text) reads JSON from a string. The s stands for “string”false became Python Falser.json() does the same for an API replywb[0] is the metadata, and wb[1] is the list of recordsmeta and records, makes the code easy to readrecords[0] is the newest year, 2025first["country"]["value"] is the path country → valuewb[1][0]["country"]["value"] in one go. It is the same pathKeyError means the dictionary has no such keydate, not yearIndexError means the list is shorter than you thinkwb has two items, at positions 0 and 1curl last class, or download wb_pop_ury.jsonwith open(...) as f: and json.load(f)What to look for
Stuck, or want to compare your code with mine?
requests, step by step 🌐requests.getrequests.get returns a response, which we call rr.status_code is the status coder.text is the reply as plain text. [:60] shows the first 60 charactersr.json() turns that text into a dictionary, like json.loadstimeout=30 gives up after 30 seconds. Without it, a silent server can make your script wait foreverrequests once with pip install requestsrequests build the URLparamsurl holds only the path. No ? and no ¶ms becomes one option in the query stringrequests adds the ? and the & for you: became %3Ar.url shows the full URL that was sent. Paste it into your browser to check itraise_for_status() stops the script on a bad statusA typo in the path: countryy instead of country
r.raise_for_status() checks the status code4xx or 5xx, it raises an error: Python stops and prints a message2xx, it does nothing and the script carries onrequests.get200 that carries an error messageA country code that does not exist, XYZ:
200, so raise_for_status() lets it passif len(payload) < 2: catches itimport json
import requests
url = "https://api.worldbank.org/v2/country/BRA/indicator/NY.GDP.PCAP.KD"
params = {"format": "json", "date": "2014:2025"}
r = requests.get(url, params=params, timeout=30)
r.raise_for_status()
payload = r.json()
if len(payload) < 2:
print("No data:", payload)
else:
with open("data/wb_gdp_bra.json", "w") as f:
json.dump(payload, f, indent=2)
print("Saved", payload[0]["total"], "records")curl -oopen(..., "w") opens the file for writing. It creates the file, or replaces itjson.dump is the reverse of json.load: it turns Python objects into JSON text in a fileindent=2 makes the file readable, like json.toolelse: block runs only when the check passes, so a bad reply is never savedURY, indicator SP.POP.TOTLdata/wb_pop_ury.jsonr.url, and paste it into your browserXYZ. What does your script print?What to look for
?format=json&date=2014%3A2025XYZ, the “No data” message, and no file savedStuck, or want to compare your code with mine?
rows = [] starts an empty listfor record in records: runs the indented lines once per recordrows.append(row) adds it to the end of the listint(record["date"]) turns the year from text, "2025", into a number, 2025pd.DataFrame(rows) makes one row per dictionarycountry, year and value, become the column namessort_values("year") puts the oldest year firstreset_index(drop=True) renumbers the rows from 0 after sortingdf.shape is (rows, columns): 12 years, 3 columnspd.json_normalizejson_normalize keeps all the fields, ten columns herecountry → value becomes country.valuedate is still text here. The loop turned it into a number;Life expectancy for five countries, 2000 to 2024
Then the same loop and pd.DataFrame(rows):
BRA;IND;JPN;NGA;USA asks for five countries at onceSP.DYN.LE00.IN is life expectancy at birthper_page=1000 gets them all in one reply, so there is only one page to readwb_life_expectancy_5.jsonThe same loop, with BRA;NGA and the years 2022:2025
"value": nullnull becomes Python None, and pandas shows it as NaN (“not a number”)dropna() removes the rows with missing valuesThe CSV starts like this:
df[df["year"] == 2024] keeps only the rows for 2024ascending=False puts the highest value firstto_csv writes the table to a file your report can readindex=False leaves out those row labels, which are not dataNY.GDP.PCAP.KD) for Argentina, Brazil, and Uruguay (ARG;BRA;URY), 2014 to 2025raise_for_status() and the length checkcountry, year, valuedf["value"].isna().sum()data/gdp_south_america.csvWhat to look for
Stuck, or want to compare your code with mine?
pull_data.py, part 1: the settings# The World Bank indicator code. This one is
# life expectancy at birth.
INDICATOR = "SP.DYN.LE00.IN"
# The countries you want, as ISO three-letter
# codes, joined by ";".
COUNTRIES = "BRA;IND;NGA;USA"
# The years to cover, written as "first:last".
YEARS = "2000:2023"
# A short, readable name for the output files.
OUTPUT_NAME = "life_expectancy"scripts/COUNTRIES uses the ; you saw todaypull_data.py, part 2: the requestrequests section, in one functiondef fetch_indicator(indicator, countries, years):
url = f"{BASE_URL}/country/{countries}/indicator/{indicator}"
params = {
"format": "json",
"date": years,
"per_page": 20000,
}
response = requests.get(url, params=params, timeout=60)
response.raise_for_status()
payload = response.json()
if not isinstance(payload, list) or len(payload) < 2:
raise RuntimeError(f"Unexpected response from the API: {payload}")
return payloadComments and type hints removed to fit the slide. The file itself explains every line
def creates a function: a named block of code you can run laterreturn payload hands the reply back to whoever called the functionf"..." string fills in the values inside { }. BASE_URL is set near the top of the file: https://api.worldbank.org/v2per_page=20000 fits everything on one pageraise RuntimeError(...) stops the script with a message, like raise_for_status()pull_data.py, part 3: the tableThe starter’s version
Today’s version, which does the same
for goes inside the bracketsdropna(subset=["value"]) removes rows with no valuepull_data.py, part 4: running itRun it once, from the top folder of your repository:
The output, with the long folder paths shortened:
Requesting https://api.worldbank.org/v2/country/BRA;IND;NGA;USA/indicator/SP.DYN.LE00.IN
Saved the raw response to data/raw/life_expectancy_raw.json
Saved 96 tidy rows to data/raw/life_expectancy.csv
First few rows:
country_code country indicator year value
0 BRA Brazil SP.DYN.LE00.IN 2000 69.584
1 BRA Brazil SP.DYN.LE00.IN 2001 69.980
2 BRA Brazil SP.DYN.LE00.IN 2002 70.396
Now commit both files in data/raw/ to your repository.if __name__ == "__main__":, runs the script only when you call it directly.env file, and put .env in .gitignoreper_page=20000. One request is simpler than a loop| File | What it does |
|---|---|
scripts/pull_data.py |
Runs once: python scripts/pull_data.py |
data/raw/<name>_raw.json |
The untouched API reply |
data/raw/<name>.csv |
The tidy table your report reads |
report.qmd |
Reads only the saved copy, 1,500 to 2,500 words |
Dockerfile |
Builds the image your report renders in |
requirements.txt |
Pinned versions, so the build is the same next month |
data/raw/. Do not put it in .gitignoreThe starter repository: github.com/danilofreire/datasci350-project-starter
wdi_panel.parquetThe eight indicators include GDP per capita, population, life expectancy, and CO2 emissions per person
The script that built it, and the full list, are in the appendix
with open(...) as f: opens a file and closes it for youjson.load reads JSON, and json.dump writes itwb[1][0]["value"]requests.get(url, params=..., timeout=30) sends the requestraise_for_status() and a length check catch the errorsfor loop picks the fields you want from each recordpd.DataFrame(rows) turns the list into a table; asks for many countries, and per_page gets them in one replynull becomes NaN, and dropna() removes itto_csv(..., index=False) saves the table for your reportpull_data.py does all of this, and you can now read every lineBefore then
python scripts/pull_data.py oncewb[0] is the metadata, and total is one of its keys5 (2025 is at 0)1 → 5 → value becomes wb[1][5]["value"]wb[1][0] works as well as any otherimport json
import requests
url = "https://api.worldbank.org/v2/country/URY/indicator/SP.POP.TOTL"
params = {"format": "json", "date": "2014:2025"}
r = requests.get(url, params=params, timeout=30)
r.raise_for_status()
payload = r.json()
if len(payload) < 2:
print("No data:", payload)
else:
with open("data/wb_pop_ury.json", "w") as f:
json.dump(payload, f, indent=2)
print(r.url)
print("Saved", payload[0]["total"], "records")URY and SP.POP.TOTL changed. The params are the same for every indicatorXYZ, the length check prints the “No data” message, and no file is writtendata/ folder does not exist, open fails. Create the folder firstimport requests
import pandas as pd
url = "https://api.worldbank.org/v2/country/ARG;BRA;URY/indicator/NY.GDP.PCAP.KD"
params = {"format": "json", "date": "2014:2025", "per_page": 1000}
r = requests.get(url, params=params, timeout=30)
r.raise_for_status()
payload = r.json()
rows = []
for record in payload[1]:
rows.append({
"country": record["country"]["value"],
"year": int(record["date"]),
"value": record["value"],
})
df = pd.DataFrame(rows)
print(df.shape)
print(df["value"].isna().sum())
latest = df[df["year"] == 2025]
latest = latest.sort_values("value", ascending=False)
print(latest)
df.to_csv("data/gdp_south_america.csv", index=False).env, one NAME=value per line, with no quotesNASA_API_KEY=abc123def456
.env to your .gitignore before your first commitpython-dotenv, as in Lecture 14GDP per capita for every country in 2023, 100 records per page
import time
import requests
url = "https://api.worldbank.org/v2/country/all/indicator/NY.GDP.PCAP.KD"
records = []
page = 1
while True:
params = {"format": "json", "date": 2023,
"per_page": 100, "page": page}
r = requests.get(url, params=params, timeout=30)
r.raise_for_status()
payload = r.json()
records.extend(payload[1])
print("page", page, "of", payload[0]["pages"])
if page >= payload[0]["pages"]:
break
page = page + 1
time.sleep(0.5)
print(len(records), "records")while True: repeats until a breakpages there are, so the loop knows when to stoprecords.extend(...) adds a whole page of records to the listtime.sleep(0.5) waits half a second between requests, to be polite to the server429 status means you are asking too fast. Wait and try againdata/ (wb_gdppc_2023_page1.json to page3)data/wdi_panel.parquet holds eight indicators for every country (regional and income groups such as “World” are dropped) from 1990 to 2023, in long format
| Column | Type | Meaning |
|---|---|---|
country |
category | Country name |
iso3 |
category | Three-letter country code |
indicator |
category | Short name, one of the eight |
year |
int16 | 1990 to 2023 |
value |
float64 | The observation, NaN where missing |
| Short name | WDI code |
|---|---|
gdp_per_capita |
NY.GDP.PCAP.KD |
population |
SP.POP.TOTL |
life_expectancy |
SP.DYN.LE00.IN |
co2_per_capita |
EN.GHG.CO2.PC.CE.AR5 |
internet_users_pct |
IT.NET.USER.ZS |
urban_pop_pct |
SP.URB.TOTL.IN.ZS |
fertility_rate |
SP.DYN.TFRT.IN |
primary_enrolment_net |
SE.PRM.NENR |
The script is data/build_wdi_panel.py. It saves the raw replies in data/raw/, so a second run makes no network requests