0
1
4
9
16
25
Lecture 21 - Parallel Computing Fundamentals
curl and requests fetched the datafor loop turned the reply into a DataFrameLecture 19’s course panel, wdi_panel.parquet, uses the parquet format we meet later today
map functionjoblib and concurrent.futures: parallel work on your laptopdask.delayed for your own functionsmap functionbar on each number, one at a time:map applies one function to every element of a listmap:joblib for parallel computingbar is independent, so the calls can run at the same timejoblib sends tasks to worker processes: separate copies of Python, one per coreParallel(n_jobs=6) runs up to six tasks at once. n_jobs=-1 uses all coresdelayed(bar) wraps bar, so joblib runs it later, on a workerdelayed(bar)(x) for x in foo makes one task per elementfor loop inside bracketsbar function and foo array from before:foocalculation runs ten heavy steps on ten million random numbers%timeit runs a line several times and reports the average time. It works in Jupyter only%%timeit (two % signs) times a whole cell_ is the usual name for a loop variable we do not useconcurrent.futures: the standard-library optionconcurrent.futures comes with Python: nothing to installProcessPoolExecutorThreadPoolExecutorwith starts the workers and stops them at the endbar from a .py file. Workers cannot see functions defined in a cell. joblib avoids thisfrom concurrent.futures import ThreadPoolExecutor
import requests
base = "https://api.worldbank.org/v2/country/"
end = "/indicator/SP.POP.TOTL?format=json&per_page=100"
urls = [base + "BRA" + end,
base + "IND" + end,
base + "USA" + end]
def fetch(url):
return requests.get(url, timeout=30).json()
with ThreadPoolExecutor(max_workers=3) as executor:
results = list(executor.map(fetch, urls))n_jobs=-1load_digits() loads 1,797 small images of handwritten digitsfrom sklearn.datasets import load_digits
from sklearn.model_selection import GridSearchCV
from sklearn.svm import SVC
import time
digits = load_digits()
param_grid = {"C": [0.1, 1, 10], "gamma": [0.001, 0.01]}
start = time.time()
search = GridSearchCV(SVC(), param_grid, cv=3, n_jobs=-1)
search.fit(digits.data, digits.target)
elapsed = time.time() - start
print(f"Best score: {search.best_score_:.3f}")
print(f"Time: {elapsed:.2f}s (all cores)")Best score: 0.976
Time: 1.93s (all cores)
concurrent.futures suits downloads, and machines where you cannot install packages| Scenario | Tool |
|---|---|
| sklearn / numeric loops | joblib (n_jobs=-1) |
| Download many files | ThreadPoolExecutor |
| CPU-heavy, no dependencies | ProcessPoolExecutor |
| Bigger than RAM, clusters | Dask (next section) |
Every tool follows the same steps: split the work, run it on separate workers, then collect the results
calculation on n inputs is O(n): a hundred calls take a hundred times as long as oneimport matplotlib.pyplot as plt
import numpy as np
# Simulate processing times
num_images = np.array([1, 10, 50, 100, 200])
sequential_time = num_images * 2 # 2 seconds per image
parallel_time = (num_images * 2) / 4 # 4 cores, ideal speedup
plt.figure(figsize=(8, 5))
plt.plot(num_images, sequential_time, 'o-', label='Serial O(n)', linewidth=2, markersize=8)
plt.plot(num_images, parallel_time, 's-', label='Parallel on 4 cores (still O(n))', linewidth=2, markersize=8)
plt.xlabel('Number of Images', fontsize=12)
plt.ylabel('Time (seconds)', fontsize=12)
plt.title('O(n) Scaling: Serial vs Parallel', fontsize=14, fontweight='bold')
plt.legend(fontsize=11)
plt.grid(True, alpha=0.3)
plt.show()
p cores, an O(n) task takes about n/p time plus overhead. It is still O(n)n separate inputsParallelism never changes the Big O class. It only divides the time by a constant, so O(n²) on eight cores is still O(n²)
n = 100_000 gives 10 billion. Eight cores still leave 1.25 billion eachpip install joblib numpy if you have not done so alreadynp.arange(1000000) makes an array of the numbers 0 to 999,999import matplotlib.pyplot as plt
import numpy as np
cores = np.arange(1, 65)
serial_fracs = [0.05, 0.2, 0.5]
labels = ["5% serial", "20% serial", "50% serial"]
colours = ["#1B3A6B", "#E07A24", "#2CA02C"]
plt.figure(figsize=(8, 5))
for s, lab, c in zip(serial_fracs, labels, colours):
speedup = 1 / (s + (1 - s) / cores)
plt.plot(cores, speedup, label=lab, linewidth=2, color=c)
plt.xlabel("Number of cores", fontsize=12)
plt.ylabel("Speed-up (×)", fontsize=12)
plt.title("Amdahl's law: speed-up vs core count",
fontsize=14, fontweight="bold")
plt.legend(fontsize=11)
plt.grid(True, alpha=0.3)
plt.tight_layout()
plt.show()
| Task | Use | Why |
|---|---|---|
| Number crunching | Processes (joblib, ProcessPoolExecutor) |
Each has its own GIL |
| Downloading many URLs | Threads (ThreadPoolExecutor) |
GIL released while waiting |
| NumPy | Threads can help | GIL released in C code |
| Polars / DuckDB | Neither | Already use all cores |
sklearn with n_jobs=-1 |
Processes (via joblib) |
Each fold trains independently |
joblib splits the work into function calls. Dask also splits the data into chunks, so the data can be bigger than your RAM
.compute().compute() runs the plan:+, *, exp, log) and summaries (sum(), mean(), std()) work as in NumPyNot a rule: here NumPy uses one core and Dask uses several. On small data, Dask is often slower
dask.datasets.timeseries() makes random example data: one row per second from 1 to 30 January 2000..., because no rows have been computed yethead() is not lazy: it computes the first rows straight away| name | id | x | y | |
|---|---|---|---|---|
| timestamp | ||||
| 2000-01-01 00:00:00 | Yvonne | 999 | -0.64 | 0.35 |
| 2000-01-01 00:00:01 | Charlie | 1015 | -0.67 | -0.09 |
| 2000-01-01 00:00:02 | Charlie | 912 | 0.22 | 0.03 |
| 2000-01-01 00:00:03 | Quinn | 988 | -0.85 | -0.53 |
| 2000-01-01 00:00:04 | Xavier | 968 | 0.20 | 0.48 |
y > 0, then take the standard deviation of x for each nameDask Series Structure:
npartitions=1
float64
...
Dask Name: getitem, 8 expressions
Expr=(((Filter(frame=ArrowStringConversion(frame=Timeseries(815c770)), predicate=ArrowStringConversion(frame=Timeseries(815c770))['y'] > 0))[['name', 'x']]).std(ddof=1, numeric_only=False, split_out=None, observed=True))['x']
df3 still shows no numbers.compute() runs the plan and returns a pandas Series.persist() keeps a computed result in memoryx and the maximum of y, for each name:.compute() returns an ordinary pandas DataFrame for the last stepsLoad and filter in Dask, aggregate down to something small, call .compute(), then finish in pandas
import dask
import dask.dataframe as dd
df = dask.datasets.timeseries()
# Step 1: filter with Dask
positive = df[df.x > 0]
# Step 2: aggregate with Dask
summary = positive.groupby("name").agg({"x": "mean", "y": "std"})
# Step 3: bring the small result to pandas
pdf = summary.compute()
# Step 4: use pandas normally
pdf.sort_values("x", ascending=False).head(5)| x | y | |
|---|---|---|
| name | ||
| Bob | 0.5 | 0.58 |
| Zelda | 0.5 | 0.58 |
| Norbert | 0.5 | 0.58 |
| Oliver | 0.5 | 0.58 |
| Jerry | 0.5 | 0.58 |
pip install dask if you have not alreadynumpy and dask.arraysize = 10_000_000. Use a smaller number if your laptop is short on memorychunks=100_000da.sqrt(x**2).mean().compute() with %timeitchunks=2_000_000chunks=5_000_000import numpy as np
import dask.array as da
size = 10_000_000
# Dask with SMALL chunks
da_data_small = da.random.random(size, chunks=100_000) # 100 chunks
%timeit da.sqrt(da_data_small**2).mean().compute()
# Dask with MEDIUM chunks
da_data_medium = da.random.random(size, chunks=2_000_000) # 5 chunks
%timeit da.sqrt(da_data_medium**2).mean().compute()
# Dask with LARGE chunks
da_data_large = da.random.random(size, chunks=5_000_000) # 2 chunks
%timeit da.sqrt(da_data_large**2).mean().compute()* in 'data/*.csv' with each namename(i) turns partition i into a date; hides the file list that to_csv returnsdata folder now has 30 CSV files, 182 MB in totaldd.read_csv accepts a glob pattern: 2000-*-*.csv matches them all| Feature | CSV | Parquet |
|---|---|---|
| Storage | Row-based | Column-based |
| Compression | None | Snappy/gzip |
| Column selection | Reads all | Reads only needed |
| Data types | Text only | Typed (int, float, date) |
| Our January 2000 data | 182 MB | 86 MB |
CSV is not obsolete. Use parquet for files over 100 MB, or with many columns when you need only a few. Keep CSV for small files and anything a colleague will open in Excel
dask.delayed parallelises functions you already have@ line above a def. It changes what the function does@dask.delayed, calling the function creates a task instead of running it@dask.delayed
def delayed_calculation(size=10000000):
arr = np.random.rand(size)
for _ in range(10):
arr = np.sqrt(arr) + np.sin(arr)
return np.mean(arr)
results = []
for _ in range(4):
results.append(delayed_calculation())
# Compute all results at once
%timeit final_results = dask.compute(*results)515 ms ± 12.1 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)
dask.compute(*results) runs all four tasks. The * passes the list items one by one@dask.delayed
def generate_data(size):
return np.random.rand(size)
@dask.delayed
def transform_data(data):
return np.sqrt(data) + np.sin(data)
@dask.delayed
def aggregate_data(data):
return {
'mean': np.mean(data),
'std': np.std(data),
'max': np.max(data)
}
sizes = [1000000, 2000000, 3000000]
# Build the tasks (nothing runs yet)
dask_results = []
for size in sizes:
data = generate_data(size)
transformed = transform_data(data)
stats = aggregate_data(transformed)
dask_results.append(stats)
%timeit dask.compute(*dask_results)26.8 ms ± 495 μs per loop (mean ± std. dev. of 7 runs, 10 loops each)
dask.visualize() draws the plan without running itDask sees all nine tasks first. The branches are independent, so they can run on separate workers
localhost:8787from dask.distributed import Client, then client = Client(), opens a dashboard at localhost:8787pip install "dask[distributed]" bokehgenerate_data, teal is transform_data, and green is aggregate_data%timeit and cProfile find the slow part. Amdahl’s law says what fixing it is worth.compute() and .persist() rarely: each call runs the whole planchunks='auto'Most common mistake: parallelising before measuring. More cores will not fix a slow disk or an \(O(n^2)\) loop
map, embarrassingly parallel problems, joblibconcurrent.futures: processes for CPU work, threads for waitingdask.delayed parallelises your own functions, and dask.visualize shows the plan| Engine | Language | Parallel? | Out-of-core? |
|---|---|---|---|
| pandas | Python/C | Single-threaded | No |
| Polars | Rust | All cores | Yes (streaming) |
| DuckDB | C++ | All cores | Yes |
| Dask | Python | All cores + clusters | Yes |
tutorials/04-duckdb-sql-tutorial.qmd46.8 ms ± 2.44 ms per loop (mean ± std. dev. of 7 runs, 10 loops each)
1.9 s ± 16.8 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)
square(x) finishes so fast that the cost of sending a million tiny tasks to the workers swamps the workmean(sqrt(x^2)) at four chunk sizes44.1 ms ± 498 μs per loop (mean ± std. dev. of 7 runs, 10 loops each)
13.8 ms ± 89.2 μs per loop (mean ± std. dev. of 7 runs, 100 loops each)
20.2 ms ± 235 μs per loop (mean ± std. dev. of 7 runs, 10 loops each)
33.2 ms ± 222 μs per loop (mean ± std. dev. of 7 runs, 10 loops each)