DATASCI 101: Introduction to AI Applications

Lecture 04: Supervised, Unsupervised, and Reinforcement Learning

Danilo Freire

Department of Data and Decision Sciences
Emory University

Welcome back! 🤓

Recap of last class

  • Good data is the foundation of AI systems: garbage in, garbage out!
  • Different AI/ML tasks need different labels (classification, regression, etc.)
  • Selection bias is dangerous and hard to fix after the fact
  • Clear guidelines and multiple annotators improve labelling quality
  • Inter-annotator agreement (Cohen’s Kappa) measures labelling reliability
  • Today: how do machines learn from data? 🤔

You can’t out-train bad data

Source: Programmer Humor

Lecture overview

Today’s agenda

  • The three paradigms of machine learning
  • Supervised learning: learn from labelled examples
    • Linear regression, decision trees, neural networks
  • The bias-variance trade-off
  • Unsupervised learning: find hidden structure
    • Clustering and dimensionality reduction
  • Reinforcement learning: learn from interaction
    • Agents, rewards, and policies

The three paradigms of ML

Source: Data Science Dojo

Tweet of the day 😄

The three learning paradigms 🎓

How do machines learn?

Three different approaches

  • AI algorithms learn patterns from data
  • But what kind of data and what kind of feedback?
    1. Supervised learning: learn from labelled examples (last lecture)
    2. Unsupervised learning: find structure in unlabelled data
    3. Reinforcement learning: learn from rewards and punishments
  • Your choice depends on the data you have and the goal you want

Learning paradigms overview and examples

Source: Medium

Supervised learning 📚

What is supervised learning?

Learning from examples with answers

  • Supervised learning: learn from labelled examples
  • Training data: input-output pairs \((x_i, y_i)\)
  • Goal: learn a function \(f\) such that \(f(x) \approx y\)
  • The “supervision” is a teacher giving the correct answers: the known labels
  • Most common paradigm in real-world applications
  • Examples:
    • Email → Spam/Not spam
    • Image → Cat/Dog/Bird
    • Patient data → Disease risk

A hard classification problem 😂

Source: Memedroid (!)

Classification vs regression

The two main supervised tasks

Classification 🏷️

  • Predict a discrete category
  • Binary: Yes/No, Spam/Ham
  • Multi-class: Cat/Dog/Bird/Fish
  • Output: class label (or probabilities)
  • Examples:
    • Fraud detection
    • Disease diagnosis
    • Sentiment analysis

Regression 📈

  • Predict a continuous value (a number)
  • Examples:
    • House price prediction
    • Stock price forecasting
    • Temperature prediction
    • Age estimation from photo

The same algorithm family often handles both tasks with minor modifications!

Linear models

The simplest supervised learners

  • Linear regression: predict y as a weighted sum of features

\[\hat{y} = w_0 + w_1 x_1 + w_2 x_2 + \ldots + w_n x_n\]

  • Learn weights \(w\) that minimise prediction error
  • Logistic regression: classification via the sigmoid function

\[P(y=1|x) = \frac{1}{1 + e^{-(w_0 + w_1 x_1 + \ldots)}}\]

  • Simple, interpretable, fast to train
  • Works well when relationships are roughly linear
  • A strong baseline before trying complex models

Linear regression fits a line

Source: Medium

A logistic function converts a line into a curve

Source: Wikipedia

Decision trees

Learning rules from data

  • Decision trees: learn a hierarchy of yes/no questions
  • Each node splits the data on a feature; leaves hold the predictions
  • Easy to interpret: “If income > $50k AND age > 30, then approve loan”
  • Capture non-linear relationships
  • Prone to overfitting (memorising training data)
  • Fix: random forests combine many trees
    • Each tree sees different data/features
    • Averaging their predictions raises accuracy

Decision tree structure

Source: Medium

Neural networks for supervised learning

Learning complex patterns

  • Neural networks: layers of interconnected nodes, each transforming the input
  • Learn arbitrarily complex functions
  • Deep networks (many layers) = deep learning
  • Need more data than simpler models
  • Less interpretable (“black box”)
  • State-of-the-art for:
    • Image classification (CNNs)
    • Speech recognition
    • Natural language processing
  • Backpropagation makes training possible (lecture 02)

Neural network architecture

Source: 3Blue1Brown

The bias-variance trade-off

Too simple or too complex?

  • Every model makes errors! Two sources:
  • Bias: error from oversimplified assumptions
    • Underfitting: model too simple, misses real patterns
  • Variance: error from sensitivity to training data
    • Overfitting: model too complex, memorises noise and fails on new data

\[\text{Total Error} = \text{Bias}^2 + \text{Variance} + \text{Noise}\]

  • Trade-off: reducing one often increases the other
  • Goal: find the sweet spot where total error is smallest

Bias-variance trade-off

Source: Vizuara

Cross-validation

Testing your model properly

  • A single train/test split can be misleading: it ignores random variation
  • Cross-validation: use multiple train/test splits
  • K-fold CV: split the data into K parts
    • Train on K-1 folds, test on 1 fold
    • Repeat K times, rotating the test fold
    • Average the results
  • Benefits:
    • Every data point is tested once, so the estimate is more reliable
    • Helps detect overfitting
  • Common choice: K = 5 or K = 10

5-fold cross-validation

Source: Vizuara

Unsupervised learning 🔍

What is unsupervised learning?

Finding structure without labels

  • Unsupervised learning: learn from data without labels
  • No correct answers given. The goal is to find hidden structure
  • Like learning without a teacher
  • Why use it?
    • Labels are expensive: experts have to make them
    • Labels sometimes don’t exist (“How many customer types?”)
    • Exploration: understand the data before modelling
    • Pre-training: learn representations, then fine-tune
  • Two questions to ask first: are there natural groups? Can the data be compressed?

Unsupervised learning finds structure

Source: University of Cambridge

Clustering

Grouping similar data points

  • Clustering: partition data into groups (clusters)
  • Points in a cluster are similar, points across clusters are dissimilar
  • K-means, the most popular algorithm:
    1. Choose K cluster centres at random
    2. Assign each point to the nearest centre
    3. Move each centre to the mean of its points
    4. Repeat until convergence
  • Hierarchical clustering: builds dendrograms without pre-specifying K

Real-world uses:

Domain Application
Marketing Customer segmentation
Biology Cell types, disease subtypes
Finance Fraud detection
Healthcare Patient risk groups

K-means clustering

Source: Machine Learning CoBan

Dimensionality reduction

Compressing information

  • Curse of dimensionality: high-dimensional data is hard to work with
    • Distances lose meaning, data needs grow exponentially, plots become impossible
  • Dimensionality reduction: find a lower-dimensional representation
  • Keep the structure, drop the noise
  • Linear: PCA (Principal Component Analysis)
  • Non-linear: t-SNE, UMAP
  • Uses: visualisation, preprocessing, compression
  • Cool example, more in lecture 06: https://projector.tensorflow.org/

Reducing dimensions while preserving structure

Source: Medium

Dimensionality reduction techniques

PCA, t-SNE, and UMAP

Linear: PCA (Principal Component Analysis)

  • Find the directions of maximum variance
  • Keep the top K components to reduce dimensions
  • Fast, well understood, keeps global structure
  • Limitation: only captures linear relationships

Non-linear: t-SNE and UMAP

  • Keep local structure: nearby points stay nearby
  • Excellent for visualisation
  • Reveal clusters that PCA misses
  • Caution: cluster sizes and distances can mislead

Principal Component Analysis

Source: Vizuara

Reinforcement learning 🎮

What is reinforcement learning?

Learning from interaction

  • Reinforcement learning (RL): learn by trial and error
  • An agent interacts with an environment
  • It takes actions and receives rewards (or punishments)
  • Goal: learn a policy that maximises cumulative reward
  • Like training a dog (or kids!): reward good behaviour!
  • Unlike supervised: no “correct” action is given
  • Unlike unsupervised: there IS a goal (maximise reward)

The RL loop

Source: Wikipedia

Key concepts in RL

Concept Definition Example (Chess)
Agent The learner/decision-maker The chess-playing AI
Environment What the agent interacts with The chess board and opponent
State Current situation Board position
Action What the agent can do Move a piece
Reward Feedback signal +1 win, -1 lose, 0 otherwise
Policy Strategy: state → action “In this position, move queen”
Value Expected future reward from a state How good is this position?

Fun fact: Magnus Carlsen once said he “can’t beat his phone in chess” 😅 With RL, AlphaZero taught itself chess by playing against itself

The exploration-exploitation trade-off

Should you try something new or stick with what works?

  • Exploration: try new actions to find better strategies
  • Exploitation: use known good actions to maximise reward
  • Too much of either costs you: time wasted on bad actions, or better strategies missed
  • Restaurant choice: exploit your favourite, or explore a new one (might be better!)
  • To balance both: ε-greedy explores with probability ε, exploits otherwise
  • The classic version of this dilemma is the multi-armed bandit problem

Exploration vs exploitation

Source: Lilian Weng

Q-learning

Learning action values

  • Q-learning: learn which action is best in each situation
  • Keep a score card: \(Q(s, a)\) = “How good is action \(a\) in state \(s\)?”
  • The agent learns by doing:
    1. Try an action, see the reward → “Go right… hit a wall, ouch!”
    2. Update the score → “Going right here is bad”
    3. Repeat thousands of times → “Eventually: always go left here!”

The update rule: \[\text{New Score} = \text{Old Score} + \text{Small Correction}\]

Source: Medium

RLHF: RL for language models

From AlphaGo to ChatGPT

RLHF pipeline

Source: Simform

RLHF

RLHF from Anthropic

Source: Anthropic’s hh-rlhf dataset (check it out, it’s really cool!)

Jailbreaking RLHF

If you can make it, you can jailbreak it!

AI safety

Making sure AI does what we actually mean

  • Alignment problem: make models do what we mean, not what we literally say
  • Specification gaming: models find perverse ways to maximise rewards (e.g., reward hacking)
  • Scalable oversight: how to supervise models more capable than their human evaluators
  • Red teaming: hunt for vulnerabilities and jailbreaks before others do
  • Build safety into training, don’t bolt it on later

AI alignment

Source: NanoBanana

Comparing the three paradigms 📊

Side-by-side comparison

Supervised Unsupervised Reinforcement
Data Labelled (X, y) Unlabelled (X) States, actions, rewards
Goal Predict y from X Find structure Maximise reward
Feedback Correct answers None Reward signal
Analogy With a teacher Exploring alone Trial and error
Evaluation Against known labels Internal metrics Cumulative reward
Examples Spam detection, diagnosis Customer segments Game playing, robotics
Difficulty Medium (if labels exist) Hard to evaluate Hard to train

Summary 📚

Main takeaways

  • Three paradigms: Supervised (labels), unsupervised (no labels), reinforcement (rewards)

  • Supervised learning is most common: Learn from labelled examples to predict

  • The bias-variance trade-off: Simple models underfit, complex models overfit

  • Unsupervised learning discovers hidden structure: Clustering, dimensionality reduction

  • Reinforcement learning learns from interaction: Exploration vs exploitation

  • RLHF powers modern LLMs like ChatGPT: RL to align with human preferences

… and that’s all for today! 🎉