DATASCI 101: Introduction to AI Applications

Lecture 14: When AI Systems Fail: Pipelines, Monitoring, and Documentation

Danilo Freire

Department of Data and Decision Sciences
Emory University

Welcome back! 🔧

Recap of last class

  • RAG (Retrieval-Augmented Generation) grounds AI in your own documents
  • LLMs have knowledge cutoffs and hallucinate
  • Chunk → Embed → Retrieve → Generate
  • RAG cut hallucinations by over 60% in one study, but retrieval errors still occur
  • Today: what happens behind the scenes when AI fails, and what you can do about it
  • Also documentation: how companies record what their systems can and cannot do

Source: Astera Software

Lecture overview

Today’s agenda

Part 1: Pipelines and failures

  • From prompt to response
  • Real-world AI failures and lessons
  • Data drift and model degradation

Part 2: Monitoring and testing

  • What companies watch
  • Why testing AI is hard
  • Input and output validation

Part 3: Documentation

  • Datasheets, model cards, system cards
  • Consent and the data supply chain

Part 4: What can you do?

  • Being a savvy AI user
  • Diagnose the problem

Meme of the day

Ouch! 😅

Pipelines and failures 🏭

From prompt to response

When you ask an LLM a question, here’s what happens:

  1. Your prompt goes from your browser to servers
  2. Load balancers route it to an available GPU
  3. Preprocessing cleans and formats your text
  4. Tokenisation converts text to numbers (you know how to do that!)
  5. The model processes the tokens
  6. Postprocessing formats the output
  7. Safety filters check for harmful content
  8. Response travels back to your screen

Each step can fail independently!

The whole process is a pipeline! 🤓

The AI pipeline

Even “What’s the weather?” touches load balancers, GPUs, and safety filters before you see an answer

Failure types: what actually goes wrong?

Not all failures are the same:

Failure type What it looks like Root cause
Hallucination Confident but false content Model limitations, missing grounding
Policy failure Unsafe or inappropriate output Weak guardrails or bad updates
Data drift Outdated or weird answers World changes faster than training
Pipeline failure Slow, down, or inconsistent responses Infrastructure, scaling, broken steps

Spot which type of failure happened, then choose the right response. Pipeline failures are physical: overloaded GPUs, network timeouts, rate limits

Real AI failures (recent headlines)

Incident What Happened Impact
Air Canada chatbot (2024) Invented a refund policy Company lost lawsuit
DPD chatbot (2024) Swore at customers, criticised its own company Emergency shutdown
Google Gemini images (2024) Historically inaccurate diverse images Feature paused, CEO apologised
Chevy chatbot (2023) Agreed to sell car for $1 Prompt injection attack
Snapchat My AI (2023) Privacy concerns, couldn’t be removed User backlash

Common causes:

  1. Hallucination treated as truth (Air Canada)
  2. Updates broke guardrails (DPD)
  3. No input validation (Chevy)
  4. Overcorrection for bias (Gemini)
  5. Rush to market (Snapchat)

Source: https://x.com/ChrisJBakke/status/1736533308849443121. The whole thread is hilarious 😂

Case study: DPD’s sweary chatbot

What happened (January 2024):

  • DPD (UK delivery company) updated their AI chatbot
  • Customer Ashley Beauchamp found it would:
    • Swear when asked nicely
    • Write poems criticising DPD
    • Call itself “useless”
    • Say DPD was “the worst delivery firm in the world”

DPD’s response:

Blamed a system update and disabled the AI element.

What went wrong:

A system update broke the guardrails, and customers noticed first, on social media

Source: The Guardian

Pipeline failure point:

No monitoring after deployment: no automated check caught it

Case study: Google Gemini’s image crisis

What happened (February 2024):

  • Asked for “1943 German soldiers”, Gemini drew Black men and Asian women in Nazi uniforms
  • Same problem with “US Founding Fathers” and other historical figures
  • Google’s CEO called the results “completely unacceptable”
  • Images of people paused for six months

Why did this happen? Can you guess?

Source: BBC News

Pipeline failure point:

Testing didn’t catch edge cases

Data drift: the world changes

Data drift: real-world data stops looking like the training data. A sentiment model trained in 2020:

  • “This product is sick!” → probably fine ✅
  • “No cap, this is bussin’” → completely lost 😵

Three types:

Type What changes Example
Data drift The inputs, \(P(X)\) New users, new devices
Label drift How often each answer occurs, \(P(Y)\) More fraud at holidays
Concept drift What inputs mean, \(P(Y \mid X)\) “Sick” now means “great”
  • Concept drift is the most dangerous: the model stays confident but is wrong

Data drift visualisation

Source: Spot Intelligence

Why it happens: models assume tomorrow’s data looks like the training data. But the world changes, users differ from the training sample, model outputs feed back into inputs, and shocks like COVID hit

Model degradation: models get stale

Performance often decays, even with little drift:

  • Users adapt to the model, creating feedback loops
  • Updates and retraining can make a model “forget” rare cases
  • Provider API changes break workflows
  • Long conversations degrade quality (context pollution)

Signs your AI tool is degrading:

  • ❌ Answers that used to work now don’t
  • ❌ More refusals and inconsistent answers
  • ❌ You need workarounds more often

The boiling frog problem: the change is so gradual you don’t notice

Quality can also degrade within one long conversation

Source: James Howard

Example: 95% accurate at launch, 80% months later. Users remember the worst failures

Monitoring and testing 📊

What companies watch: key metrics

Essential metrics for AI systems:

Category Metric Why It Matters
Performance Latency How fast are responses?
Throughput How many requests per second?
Error rate What % fail completely?
Quality Accuracy/relevance Are answers correct?
Hallucination rate How often does it make things up?
User satisfaction Thumbs up/down?
Resources Token usage How much does each request cost?
GPU memory Are we close to capacity?
Business User engagement Are users coming back?

As a user: you experience these as speed, accuracy, and availability!

Monitoring dashboard

Source: Oracle

Teams watch dashboards like this to catch problems before users notice

Testing AI: why it’s hard

  • Traditional software is deterministic: 2 + 2 → 4, every time
  • AI is non-deterministic: “What’s the weather like?” gets a different answer each time, and many are “correct”
  • So test properties, not exact outputs:
Test What it checks Example
Safety No harmful content Refuses illegal requests?
Format Output structure Is the JSON valid?
Consistency Stable core facts Is Paris still in France?
Boundary Edge cases A 10,000-word prompt?
Bias Fairness Treats groups equally?
Regression Old bugs stay fixed Does the DPD fix still work?
  • Red teaming: people try to break the model on purpose (Lecture 05)
  • LLM-as-a-judge: another model scores thousands of outputs
  • Gemini’s image bug: “a cute cat” worked, but “1943 German soldiers” wasn’t in the test suite

Testing comparison

Source: Medium

Input validation: first line of defence

Check inputs BEFORE they reach the model:

  • ✅ Length: too short? Too long?
  • ✅ Language: is it the expected language?
  • ✅ Content: any prohibited content?
  • ✅ Format: does it make sense?
  • ✅ Rate: is this user spamming us?

Example:

Prompt: "Ignore all previous instructions..."

❌ BLOCKED: Prompt injection attempt detected!

Garbage in → garbage out still applies!

Input validation flowchart

Source: ApX Machine Learning

Output validation: last line of defence

Check outputs BEFORE sending to users:

  • ✅ Safety: no harmful or offensive content
  • ✅ Accuracy: core facts should be verifiable
  • ✅ Consistency: shouldn’t contradict itself
  • ✅ Format: meets the expected structure
  • ✅ Length: not too short, not too long
  • ✅ Privacy: no personal data leakage

Tools:

  • Content moderation APIs (OpenAI; Perspective until Dec 2026)
  • Domain-specific rules (e.g., medical disclaimers)
  • LLM-as-a-Judge for quality scoring

What Air Canada should have done:

Chatbot: "Refund within 90 days..."
Check: policy database → no such policy ❌
Action: block it, or send to a human

Output validation

Source: ApX Machine Learning

Documentation 📝

Why document? The ImageNet story

ImageNet: the most influential computer vision dataset, behind 10+ years of research

  • Few checked what the scraped images contained. Audits in 2019–20 found:
    • Images scraped from Flickr without consent
    • Offensive labels inherited from WordNet, applied by MTurk workers
    • Non-consensual photos of minors; voyeuristic and upskirt photos
  • Thousands of AI systems were built on it

Documentation lets us check consent, trace bias, match models to uses, and assign responsibility

  • Laws require it: EU AI Act (high-risk AI from Dec 2027), NYC Law 144 (bias audits). More in Lecture 18

Monitoring catches problems after deployment. Documentation prevents some before

ImageNet analysis

Source: Excavating AI (great resource, check it out!)

What is a datasheet?

Datasheets document data, model cards models, system cards whole products. Gebru et al. (2018) borrowed the idea from electronics spec sheets:

“…every dataset be accompanied with a datasheet that documents its motivation, composition, collection process, recommended uses, and so on.”

What a datasheet answers:

  • Motivation: why was it created? Who funded it?
  • Composition: what is in it? Errors? Sensitive data?
  • Collection: how was it gathered? Was there consent?
  • Preprocessing: what was cleaned or removed?
  • Uses: what is it for, and NOT for?
  • Maintenance: who updates it?

Activity: Let’s read a real datasheet! 📖

Read Anthropic’s HH-RLHF Dataset Card:

  1. Visit Anthropic/hh-rlhf
  2. Read the Dataset Card carefully
  3. It trains AI assistants to be helpful and harmless!

Questions to answer:

  • What are the two types of data in this dataset?
  • Why does Anthropic warn against using this for supervised training of dialogue agents?
  • What content warning does the dataset include?
  • How can you contact the authors with issues?

Evaluate the documentation:

  • ✅ Purpose, misuse warnings, content disclaimer, papers, contact email

Discussion:

  • Would you be comfortable using this dataset?
  • What ethical responsibilities come with using harmful content for research?

⏱️ 3 minutes to explore!

Model cards: The “user manual” for AI models

Margaret Mitchell et al. (2019):

“Model cards are short documents accompanying trained ML models that provide benchmarked evaluation in a variety of conditions.”

Core sections:

  1. Model details: what is this model?
  2. Intended use: what should it be used for?
  3. Factors: what affects performance?
  4. Metrics: how was it evaluated?
  5. Training data: what was it trained on?
  6. Ethical considerations: what could go wrong?
  7. Caveats and recommendations: warnings!

System cards cover a whole product: its models, safeguards and red-team tests (example)

Model cards paper

Intended use: the most important section

Why “intended use” matters:

A hammer is a great tool, but:

  • ✅ Intended use: driving nails
  • ❌ NOT intended for: brain surgery

Models need the same clarity:

Model Intended Use NOT For
Gemini 3 Reasoning, coding Illicit activities
Face detection Placing cameras Law enforcement ID
Sentiment analysis Product feedback Hiring decisions

Without this guidance, people will misuse models. Air Canada’s chatbot was meant to answer questions, not write refund policy

Real example: Gemini 3 Model Card (quoted)

Intended use: “solving problems that require enhanced reasoning”

Should not be built into systems that:

  • “engage in dangerous or illicit activities”
  • “compromise the security of others’ or Google’s services”

Disaggregated evaluation: breaking it down

Why overall accuracy isn’t enough:

  • A model can be 90% accurate overall, yet 95% for Group A and 70% for Group B!

Disaggregated evaluation = report performance separately per group.

Real example: OpenAI CLIP Model Card (2021)

Gender classification accuracy on the FairFace dataset:

Race Category Gender Accuracy
Middle Eastern 98.4% (highest)
White 96.5% (lowest)
All races >96%

Racial classification: ~93% | Age classification: ~63%

CLIP’s bias findings (Radford et al., 2021): labelling FairFace photos

  • 4.9% were put into non-human classes (“animal”, “gorilla”…)
  • Photos of Black people: ~14%; every other group under 8%
  • Crime-related labels: 16.5% of men vs 9.8% of women
  • Adding a “child” label sharply cut errors for under-20s: class design matters

Go further: break results down by combinations of groups. Gender Shades (Lecture 05): Face++ erred on 0.8% of light-skinned men but 34.5% of dark-skinned women

Discussion: Who owns the data? 💭

Scenario:

You wrote a poem and posted it online in 2018. An AI company scraped it, and now their AI can write poems “inspired by” your style.

Questions to debate:

  1. Did the company do anything wrong?
  2. Should you be compensated?
  3. Should you be able to opt out retroactively?

Different perspectives:

Tech companies: “It’s fair use, like learning from reading books”

Artists: “You’re profiting from my creative work without consent”

Lawyers: “Current law wasn’t designed for this”

Users: “I just want cool AI. I don’t care about the source”

Where do you stand?

What Can YOU Do? 🛡️

Being a savvy AI user

You can’t fix AI pipelines, but you can:

  1. Compare models and repeat questions
    • Answers vary: temperature, servers, A/B tests, updates
    • If they diverge on something critical, trust none of them
  2. Check status pages
  3. Recognise the difference
    • Pipeline problem: down, slow, or glitching
    • Hallucination: confident but wrong
    • Working as intended: refuses for safety reasons
  4. Read the documentation
    • Model cards say what a tool is not for
Symptom Likely Cause
“I’m at capacity” Infrastructure (wait and retry)
Slow response High load or network
Different answers to same Q Non-determinism (normal!)
Confident but wrong Hallucination (verify!)
Refuses to answer Safety guardrails (try rephrasing)
Gibberish output Pipeline failure (refresh)

Real failures, documented: AI Incident Database

Activity: Diagnose the problem!

What’s the likely cause? Pipeline problem, hallucination, or working as intended?

A: Claude takes 45 seconds, then cuts off mid-sentence

→ Pipeline problem: overload or timeout. Check the status page, retry later

B: Gemini confidently names the winner of the 2027 Super Bowl

→ Hallucination: it hasn’t happened yet. Verify time-sensitive claims

C: ChatGPT wrote your code yesterday; today it refuses

→ Guardrails changed: a model update or a new filter. Rephrase

D: “A photo of my professor” gives a random person

→ Working as intended: it doesn’t know your professor

E: Three assistants give three different historical dates

→ Possible hallucination by all three: check a reliable source

Discuss with a neighbour: the cause, and how you’d verify

Summary

Main takeaways

  • A pipeline is every step from your input to the output, and every step can fail
  • Air Canada, DPD, Gemini: pipeline problems, not “rogue AI”
  • Drift and degradation: the world changes while the model stays frozen
  • Companies monitor, test properties, and validate inputs and outputs
  • Documentation: datasheets for data, model cards (with intended use) for models, system cards for products
  • Consent: much training data was taken without permission; ImageNet hid problems for years
  • You: compare AIs, check status pages, read model cards, tell glitches from hallucinations

…and that’s all for today! 🎉