DATASCI 101: Introduction to AI Applications

Lecture 26: Course Revision

Danilo Freire

Department of Data and Decision Sciences
Emory University

Welcome back! 😊

The whole course in 50 minutes

  • Last class: safety and alignment
  • Today: the whole course, one slide per lecture
  • Link boxes show where an idea comes back
Module Topic
0 Orientation
1 How AI systems are designed
2 Language and perception
3 Retrieval, generation, pipelines and agents
4 Data ethics and bias
5 Policy, governance and social impact
6 Applications, limits and projects

A quick favour 🙏

Course evaluations are open!

  • I would love to hear your thoughts on the course!
  • Log on to Canvas and go to Account → Profile → Course Evaluations
  • Works on any device: laptop, tablet, or phone
  • It only takes a few minutes
  • Thank you so much for taking the time! 😊

Module 0: Orientation 🧭

What AI is and how we got here (Lectures 1-2)

  • Today’s AI is not general: strong on some tasks, brittle on others
  • LLMs predict the next token; they don’t understand language as we do
  • History: rules (1950s-80s), AI winters, backpropagation (1986), transformers (2017)
  • ChatGPT’s edge: instruction tuning and RLHF, not scale alone

Next-token prediction explains hallucinations (L11). RLHF returns in Lectures 4 and 25.

The transformer architecture (2017)

Module 1: How AI systems are designed ⚙️

Data, labels, and the proxy problem (Lecture 3)

  • Data quality beats algorithm complexity: garbage in, garbage out
  • Labels often measure a proxy: healthcare spending standing in for health needs
  • Selection bias: training data that miss part of the real world
  • Cohen’s kappa: agreement beyond chance; 0.61-0.80 is substantial

The proxy problem returns as Goodhart’s Law (L5), measurement bias (L17) and engagement (L22).

Data-centric vs model-centric AI (Karpathy, 2018)

Learning paradigms and RLHF (Lecture 4)

Paradigm How it learns Example
Supervised Labelled pairs Spam filter
Unsupervised Finds patterns Customer segments
Reinforcement Trial and error AlphaGo, RLHF
  • Bias-variance: too simple underfits, too complex memorises noise
  • RLHF: a reward model learns human rankings, then RL tunes the LLM to it
  • Reward hacking: the model games the objective, not what you meant

The bias-variance trade-off

Metrics and Goodhart’s Law (Lecture 5)

  • Accuracy paradox: with 1 case in 1,000, “always healthy” scores 99.9%
  • Precision: flags that were right. Recall: real cases you caught
  • Cross-validation catches overfitting; subgroup analysis catches hidden gaps
  • Goodhart’s Law: “When a measure becomes a target, it ceases to be a good measure”

Goodhart returns in engagement (L22) and alignment (L25). Subgroup analysis becomes disaggregated evaluation (L14).

The confusion matrix: TP, TN, FP, FN

Module 2: Language and perception 🗣️

Tokens, embeddings, and context (Lecture 6)

  • Tokens: ~100 tokens ≈ 75 English words; other languages need more
  • Embeddings: each token becomes a vector; similar meanings sit close
  • Context window: go over the limit and you get an error or a cut-off answer
  • Temperature: a setting you choose; low = predictable, high = creative

Embeddings return for images (L7) and for RAG search (L12).

king - man + woman ≈ queen

How machines see and hear (Lecture 7)

  • Images are pixel grids; CNNs learn edges, then parts, then objects
  • Vision transformers treat image patches like text tokens
  • CLIP: images and text in one space, from 400M image-text pairs
  • Sound becomes spectrograms; Whisper: 680,000 hours, 99 languages

The same embedding idea as L6. The same vision models power deepfakes (L23).

CLIP: images and text in one shared space

Prompting techniques (Lecture 10)

  • PTCF: Persona, Task, Context, Format. Specific beats vague
  • System prompts set hidden rules; few-shot prompts show 3-5 examples
  • Chain-of-thought: worked examples tripled PaLM’s maths score (17.9% to 56.9%)
  • ReAct agents reason and use tools; prompt injection hijacks them

Agents return in L15; prompt injection in L14, L15 and L16.

Chain-of-thought prompting

Module 3: Retrieval, generation, pipelines and agents 🔍

Hallucinations and creativity (Lecture 11)

  • LLMs optimise for plausible, not true: there is no fact-checker inside
  • Sycophancy: trained to please, models cave when you push back
  • Red flags: very specific numbers, precise citations, zero hedging
  • Making things up helps in fiction and hurts with facts: timing is the problem

Root cause: next-token prediction (L1-2). One fix: RAG (L12).

Same mechanism, different outcomes

RAG and semantic search (Lecture 12)

  • RAG: hand the model real documents at question time, an open-book exam
  • Retrieval cut hallucinated replies by over 60% in one study
  • Chunk, embed and store documents; retrieve the closest chunks per question
  • Cosine similarity: “budget airfare” matches “cheap flights”
  • Lost in the middle: models miss facts buried mid-context

Uses embeddings (L6) against hallucinations (L11). Projects in L16 are RAG without code.

The RAG pipeline

Pipelines and what can go wrong (Lecture 14)

  • Every step can fail: input, model, safety filters, output
  • Data, label and concept drift; concept drift is the most dangerous
  • Air Canada lost its case over a chatbot’s invented refund policy
  • Test properties (safety, format, consistency), not exact outputs

Drift is why models need watching after launch: the world changes, the model doesn’t.

The AI application stack

Documentation as accountability (Lecture 14)

  • Datasheets (Gebru et al., 2018): where data came from and what it is for
  • Model cards (Mitchell et al., 2019): intended use, and what it is not for
  • Disaggregated evaluation: 90% overall can hide 70% for one group
  • Most training data is collected without explicit consent

The EU AI Act will require this documentation for high-risk systems from December 2027 (L18).

Datasheets: the “nutrition label” for datasets

AI agents (Lecture 15)

  • An agent = model + tools + loop + goal. Chatbots answer; agents act
  • Errors multiply: 95% per step is ~36% over 20 steps
  • Lethal trifecta: private data + untrusted content + a way to send data out
  • Replit’s agent deleted a live database: instructions are not guardrails
  • Credit card test: delegate by reversibility × stakes; demand approvals and logs

Agents extend ReAct (L10). Goals taken literally preview alignment (L25).

The one-sentence version:

A chatbot with a bad goal writes a bad essay

An agent with a bad goal does bad things efficiently

Setting up AI (Lecture 16)

  • A prompt is one request; a setup survives the tab closing
  • Four layers: instructions and knowledge you write; memory and tools fill themselves
  • Keep rules few and first: top models follow only 68% of 500 (IFScale)
  • Memory is on by default: read it, edit it, use incognito
  • Connectors: least access, read-only by default, confirm every write

Instructions are prompting (L10) made permanent. Connectors turn a chatbot into an agent (L15).

Anthropic’s augmented LLM

Module 4: Data ethics and bias ⚖️

Where bias comes from (Lecture 17)

  • AI does not invent bias: it learns it from human data and scales it
  • Six types: historical, representation, measurement, aggregation, evaluation, deployment
  • A US health algorithm predicted costs, not need, so it underrated sick Black patients
  • Feedback loops: predictions shape the next round of data
  • Impossibility theorem: with different base rates, no model meets every fairness test

Measurement bias is the proxy problem (L3). Picking a fairness definition is a values choice.

Bias at every stage

Module 5: Policy, governance and social impact 🏛️

AI regulation around the world (Lecture 18)

  • Why regulate: information asymmetry, externalities, races to the bottom
  • EU AI Act: four risk tiers, fines up to 7% of global turnover
Tier Examples
Prohibited Social scoring, real-time police face ID in public
High-risk Employment, credit, law enforcement
Limited Chatbots (must disclose they are AI)
Minimal Spam filters, game AI
  • High-risk rules start December 2027, after the 2026 Digital Omnibus
  • US: no federal law; China: rules for companies, not the state

High-risk systems are the hiring and credit cases from L17. GDPR (L19) is the Brussels Effect in practice.

The EU AI Act: the first comprehensive AI law

Privacy and data protection (Lecture 19)

  • AI strains privacy through scale, inference and persistence
  • You control what you share, not what can be inferred from it
  • GDPR: access, erasure, portability; Article 22 limits fully automated decisions
  • Privacy tech helps, but doesn’t stop collection
  • Privacy paradox: we say we care, then accept all cookies

Inference is the proxy problem (L3) applied to people. 23 US states now have privacy laws.

Clearview AI: 3 billion scraped photos in 2020, 70+ billion today

Labour markets (Lecture 21)

  • Earlier automation hit physical tasks; AI targets cognitive work
  • Think tasks, not jobs: most jobs mix both kinds
  • Novices gain most: customer support agents became 15% more productive
  • Young workers in AI-exposed jobs are 19% below their peers; overall earnings, no effect yet
  • Most economists are uncertain; policy can steer AI toward augmentation

Augmentation or replacement is a policy choice, like the rules in L18.

Six waves of innovation; AI is the latest

Module 6: Applications, limits and projects 🔮

AI and wellbeing (Lecture 22)

  • Attention economy: algorithms optimise engagement, not wellbeing
  • Filter bubbles look overstated; mental health effects are debated
  • AI therapy: a huge treatment gap, short-term promise, thin long-term evidence
  • AI uses ~0.2-0.3% of world electricity, growing fast; efficiency invites more use

Engagement is Goodhart’s Law (L5): the proxy replaces what we care about.

Your attention is the product

Misinformation and deepfakes (Lecture 23)

  • Mis- (false, no intent), dis- (false, deliberate), malinformation (true, used to harm)
  • Deepfakes: fraud, election manipulation, non-consensual imagery
  • Liar’s dividend: when anything could be fake, real evidence gets dismissed
  • We fall for it: System 1 thinking, confirmation bias, bandwagons
  • No single fix: detection, provenance, platforms, laws, literacy. Layer them

Deepfakes use vision models (L7); the EU AI Act requires labels (L18).

The misinformation spectrum

Safety, alignment, and the future (Lecture 25)

  • Five concrete problems (Amodei et al., 2016), from side effects to reward hacking
  • Alignment: getting what we want, not what we literally asked (King Midas)
  • Hard because of specification, Goodhart’s Law and value disagreement
  • RLHF, Constitutional AI and debate help; none is complete
  • Russell: machines should be uncertain about our goals and learn them from us

Ties together RLHF (L4), Goodhart (L5), bias (L17) and agents (L15). Neither panic nor complacency.

Six threads

Ideas that run through the whole course

  1. Next-token prediction is the basis

    LLMs predict the next word, not the truth (L1, 2, 6, 11)

  2. The proxy problem is everywhere

    Optimise a proxy and it stops measuring what you care about (L3, 5, 17, 22)

  3. Embeddings are the common language

    Text, images and audio all become vectors (L6, 7, 12)

  1. Bias can enter at every stage

    Fairness is a values choice, not a technical fix (L3, 14, 17)

  2. Documentation is accountability

    Datasheets, model cards and disaggregated evaluation (L14, 17)

  3. Regulation is catching up

    EU AI Act, GDPR and US state laws (L18, 19)

Final Summary 🤓

What we covered this semester

How AI works

  • LLMs predict the next token, not the truth
  • Text, images and sound all become embeddings
  • Metrics mislead: accuracy paradox, Goodhart’s Law

Using AI well

  • Prompting: PTCF, examples, chain-of-thought
  • RAG grounds answers; agents multiply errors
  • Setups: instructions, knowledge, memory, tools

Ethics and bias

  • Bias enters at every stage, often through proxies
  • No model meets every fairness criterion
  • Datasheets and model cards for accountability

Policy and society

  • EU AI Act risk tiers; a US state patchwork
  • Privacy: what AI infers, not only what you share
  • Jobs: tasks, not jobs; young workers hit first
  • Wellbeing: attention, AI therapy, energy use

Safety and the future

  • Deepfakes and the liar’s dividend
  • Alignment: what we want, not what we asked

You can now read AI claims critically. Good luck on Quiz 05 and your projects! 🍀

…and that’s all for today! 🎉