DATASCI 101: Introduction to AI Applications

Lecture 08: Quiz 01 Review Session

Danilo Freire

Department of Data and Decision Sciences
Emory University

Welcome back! 🧠

Today’s plan

  • Quiz 01 is on Thursday! (50 minutes, in class)
  • The quiz tests understanding and critical evaluation, not memorisation
  • Today is your session: I ask, you answer
    1. A rapid recap of the seven lectures
    2. One realistic scenario, broken down together
    3. Practice questions, alone or in pairs
    4. Open Q&A
  • Stop me whenever something is unclear! 😉

Quiz logistics

  • Format: 4 questions plus one bonus, covering lectures 01 to 07
  • Length: 50 minutes, in person
  • Submission: one Word or PDF file, uploaded to Canvas
  • Allowed: laptops, notes, lecture slides, web search
  • AI assistants: use them sparingly, only when a question needs one
  • Honour code: write your own answers; if you used an AI, say which one
  • What is graded: your reasoning, not the polish of your AI’s prose
  • Your job is to spot problems, judge trade-offs, and write something useful, better than the AI can

Comic of the day

Hey, wait a minute 😅

Source: https://xkcd.com/2228/

Part 1: quick recap 🚀

How this part works

  • I name a few concepts from one of the seven lectures
  • One of you explains it in your own words (about 30 seconds)
  • The rest of us push back, add nuance, or correct
  • Wrong answers are useful; silence is not 😉

Concept 1: hallucination

  • In one sentence: what is an AI hallucination, and why does it happen?
  • Any volunteer or victim? 🙋
  • Is a hallucination the same as a lie? The same as a factual error?
  • Why might an open-laptop quiz punish someone who copy-pastes an LLM answer? 😂

Concept 2: dataset and labels

  • Let’s pick someone new!
  • What does it mean to “label” a dataset, and why is it often one of the most expensive steps?
  • Room poll: is the label easy, hard, or contested?
    • “Is this email spam?”
    • “Is this tweet toxic?”
    • “Is this CT scan showing pneumonia?”
    • “Is this résumé a good fit for the role?”
  • Connect to: garbage in, garbage out

Concept 3: the three learning paradigms

  • Three students, one paradigm each, one sentence per person:
    • Supervised learning?
    • Unsupervised learning?
    • Reinforcement learning?
  • Room: name one real product that uses each
  • Trick question: where does training a chatbot with thumbs up / thumbs down fit?

Concept 4: precision vs recall

  • Volunteer to draw the confusion matrix on the board (or describe it)
  • Precision (of what it flagged, how much was right) or recall (of what it should have flagged, how much it caught)?
    • A detector that flags student essays as AI-written
    • The smoke alarm in your flat
    • YouTube blocking an upload for copyright. Answer it once as the creator, then again as the rights holder

Concept 5: overfitting and validation

  • A model predicts daily passenger numbers at Atlanta airport
    • Trained on 2019 to 2020: 99% of predictions within 5% of the truth
    • Validated on 2021, which it never saw: 97% within 5%
  • Room: would you use it today?
  • Hint: what was air travel like in 2021?

Concept 6: tokens and embeddings

  • Pair up with the person next to you. One minute. 🗣️
    • One of you explains tokenisation
    • The other explains embeddings
  • Then I pick two pairs to share
  • Rapid-fire:
    • Why does OpenAI’s API charge by the token and not by the word?
    • Why does the order of words matter, even after embedding?
    • What is a context window, and why does it cost so much to make it bigger?

Concept 7: multimodal

  • Last concept of the recap
  • What does “multimodal” mean for an AI system?
  • Name three modalities we discussed
  • Why is audio harder than text in some ways and easier in others?
  • When you upload a photo to ChatGPT, what is happening under the hood?

Part 2: a real scenario 🏥

The setup

A fictional startup, ClinicNote, sells an AI tool to hospitals. The tool listens to doctor-patient conversations and writes the clinical note.

  • The model is a fine-tuned LLM
  • Training data: 80,000 doctor-patient transcripts
    • All from one hospital network in California
    • Collected between 2019 and 2024
  • The labels are the final clinical notes doctors wrote and signed off
  • ClinicNote reports 92% agreement with doctor-written notes on a held-out test set
  • They now sell the product to hospitals across the country

Question A

  • What problems can you spot with the training data?
  • Hints:
    • Where did the data come from? Who is in it? Who is not in it?
    • The labels are doctor-written notes: the truth, or one person’s interpretation?
    • The data is from 2019-2024: does medicine change over five years?
  • Try to name at least three distinct issues

Question B

  • ClinicNote reports 92% agreement with doctor-written notes
  • The doctor writes something wrong and the model copies it. What does it score?
  • Is the 8% spread evenly across patients?
  • The same doctor writes that note again a week later. Do the two versions agree 100%?

Question C

  • Your hospital is considering buying ClinicNote
  • The room votes by show of hands:
    • Yes, buy it
    • No, do not buy it
    • Buy it, but with conditions
  • One defender of each position argues for 60 seconds
  • Then: what evidence would change your mind?

Part 3: practice questions 📝

How this part works

  • I show you three practice questions, one at a time
  • Work alone or with the person next to you, as you prefer
  • You have a few minutes per question
  • Then I pick someone to share, and the room critiques
  • These are not the real quiz questions, but the format is similar

Practice question 1

Lecture 06 said a model never sees words, only tokens.

Someone asks an older model how many times the letter r appears in “strawberry”. It answers two. They then ask it to spell “strawberry” one letter at a time, and it does so perfectly.

Your task:

  1. One answer is wrong, one is right. Explain how both come out of the same machinery
  2. The model has seen the word millions of times. Explain why that alone would not teach it the spelling, and say what kind of text would

Practice question 2

Lecture 06 showed you semantic search: a query for “cheap flights” finds a page titled “budget airfare”, because those words sit close in embedding space.

Someone checks the same model and finds this:

cheap and expensive also sit close together

  • The lecture said embeddings capture meaning. Two opposites, side by side
  • First: how did that happen?
  • Harder: why does a search for “cheap flights” still work?
  • Hint: Lecture 06 compared word2vec and attention. When does each one look at the neighbours?

Practice question 3

A delivery app introduces a bonus for its best couriers. Six months later, couriers avoid the hilly part of the city, refuse orders from restaurants that are slow to cook, and several have been caught running red lights. 🛵

Nobody designed any of this.

Your task:

  1. Work backwards. What is the bonus most likely measuring?
  2. Pick one of the three behaviours and trace it back to that measure
  3. Is any of the three understandable? Say which, and who you would blame

Part 4: open Q&A 💬

Anything unclear

  • Your last chance before Thursday
  • If you are unsure, others are too
  • Topics most often asked about in past semesters:
    • Embeddings: what does the vector “mean”?
    • Validation vs test set: what is the difference?
    • Multimodal: how does the model “see” an image?
    • The proxy problem: when is a metric a bad proxy?

Many thanks and see you soon! 👋