DATASCI 101: Introduction to AI Applications

Lecture 11: Creativity and Hallucination

Danilo Freire

Department of Data and Decision Sciences
Emory University

Welcome back! 🎭

Recap of last class

  • Last class: prompt engineering
    • PTCF: persona, task, context, format
    • Chain-of-thought: “Let’s think step by step”
    • Personas: giving AI expert roles
    • AI agents: LLMs that use tools
  • Plus practical tips and iterative design
  • Today: when LLMs go creatively wrong 🎭
  • Are hallucinations AI’s biggest problem?

Lecture overview

Today’s agenda

Part 1: What Are Hallucinations?

  • When AI confidently lies
  • Why this happens (it’s not a bug!)
  • Types of hallucinations

Part 2: Real-World Failures

  • Lawyers and fake citations
  • Medical misinformation
  • Academic integrity issues

Part 3: Creativity vs. Accuracy

  • The creativity-accuracy tradeoff
  • When hallucinations are… useful?

Part 4: Detection and Prevention

  • Spotting hallucinations
  • Prompting techniques that help
  • Preview: RAG as a solution

Meme of the day 😄

Source: X.com

What are hallucinations? 🤥

When AI confidently makes things up

What is an AI hallucination?

When an AI system generates content that is plausible but factually incorrect, presented with high confidence.

Key characteristics:

  • ✅ Sounds completely reasonable
  • ✅ Uses proper grammar and tone
  • ✅ Fits the context of conversation
  • ❌ Contains fabricated information
  • ❌ May invent citations, facts, or events
  • ❌ AI doesn’t “know” it’s lying

The tricky part:

There’s often no signal: the AI sounds as confident as when it’s correct

Hallucination example

Right at first, then confidently wrong once the user pushes back

Why do LLMs hallucinate?

It’s not a bug, it’s how LLMs work!

LLMs predict the next word

  • Input: “The capital of France is”
  • Output: “Paris” (high probability)
  • Output: “Lyon” (lower probability)
  • Output: “Atlantis” (still some probability!)

The problem:

  • Plausibility ≠ truth
  • No fact-checking mechanism
  • No internet access unless the tool adds it
  • Training rewards guessing over “I don’t know” (Kalai et al., 2025)
  • LLMs are optimised to sound right, not to be right

Why hallucinations happen

Source: Medium

The sycophancy problem 🙇

LLMs are trained to be helpful…sometimes too helpful!

What is sycophancy?

The tendency to give users the answer they seem to want, rather than the accurate answer.

How this causes hallucinations:

  • Model doesn’t know → “I don’t know” feels unhelpful → it makes something up
  • User pushes back → model drops its correct answer to agree
  • Leading question → model confirms the false premise
  • Geng et al. (2025): GPT-5’s stated beliefs on moral and safety topics shifted 55% after 10 rounds of debate

Training data problems 📚

Garbage in, hallucination out! (remember lecture 03 about data?)

LLMs learn from trillions of tokens scraped from the internet:

Problem What Happens
Conflicting sources Wikipedia says X, a blog says Y → model learns both as “plausible”
Outdated information Training data has a cutoff → model treats old facts as current
Rare topics Few examples → model “interpolates” (often wrongly)
Errors in sources Mistakes on the internet become mistakes in the model
Fictional content Novels, fan fiction, speculation all look like “text” to the model

The interpolation problem:

For rare topics, the model fills in gaps with patterns from similar-but-different topics. That is where confident nonsense comes from

Example: Rare topic hallucination

Q: “Tell me about the 1987 Liechtenstein general election.”

The model might:

  1. Know Liechtenstein has elections ✓
  2. Know 1987 election patterns in Europe ✓
  3. Invent specific parties, candidates, results ✗

Plausible, because it is statistically consistent with European elections, but fabricated

Twist: there was no 1987 election (Liechtenstein voted in 1986 and 1989). Newer models often say so. But insist (“I’m 100% sure”) and Qwen replies: “the Progressive Citizens’ Party (FBP) won in 1987”

Rule of thumb: the more obscure or specific the question, the higher the hallucination risk. Models do best on topics with abundant, consistent training data

Hallucination rates by domain 📊

Hallucination rates vary dramatically by topic:

Domain Hallucination Risk Why?
Common knowledge Low Abundant, consistent training data
Coding Low-Medium Syntax is verifiable; patterns are clear
Recent events High Beyond training cutoff
Medical/Legal High Requires precision; errors are costly
Obscure history Very High Sparse data → interpolation
Citations Very High Specific strings rarely memorised

Answer quality by domain and model: similarity to the correct answer. Higher is better

Source: Chakraborty et al (2025)

Types of hallucinations

Type Description Example
Factual Wrong facts “Einstein was born in France”
Fabrication Invented information Made-up statistics, fake quotes
Citation Non-existent sources “According to Smith (2023)…” (doesn’t exist)
Logical Contradicts itself “It’s 5pm now… earlier today at 7pm…”
Temporal Wrong timelines Mixing up historical dates
Entity Confusing people/things Attributing quotes to wrong person

What do you think is the most dangerous type?

Types of hallucinations

Source: The Cloud Girl Blog

Citation hallucinations are the hardest to catch because they look so specific

Real-world failures 💥

The lawyer who trusted ChatGPT

Mata v. Avianca Airlines (2023)

  • Lawyer Steven Schwartz used ChatGPT for legal research
  • He filed a brief citing 6 cases that didn’t exist
  • ChatGPT invented the case names, citations and quotes
  • They sounded completely plausible
  • The judge sanctioned both attorneys

The wake-up call:

“Is Varghese a real case?” Schwartz asked the chatbot. “Yes,” ChatGPT doubled down, “it is a real case.”

Lawyer ChatGPT case

Source: CNN

Warning: asking AI “Is this true?” often just confirms its own hallucinations

Medical misinformation

Healthcare hallucinations can be lethal

Documented in research:

  • Wrong drug dosages
  • Invented drug interactions
  • Treatments for the wrong conditions
  • Fake medical citations

A 2024 study found:

  • ChatGPT-3.5 got the diagnosis wrong in 51% of 150 case challenges
  • Weak spots: misread lab values, unreadable images, nuanced cases, hallucinations
  • Source: Hadi et al. (2024)

Medical hallucination dangers

Source: The New York Times (2025)

  • Many people can’t afford doctors
  • Confident AI is easy to believe, hard to verify
  • Medical jargon sounds authoritative

Academic integrity issues

Fake citations are a major problem

For students:

  • AI invents realistic paper titles and author names
  • It fabricates journals and conferences
  • Very hard to detect without checking

For academia more broadly:

  • AI-generated papers being submitted
  • Fake peer reviews
  • Fabricated data
  • “Paper mills” using AI

How to avoid this:

  1. Never trust AI citations
  2. Verify EVERY reference in academic databases (Google Scholar, PubMed)
  3. Look up authors and journals

Red flags:

  • Citation format is slightly off
  • Can’t find paper anywhere
  • Author has no other publications
  • Journal name doesn’t exist

Discussion: who is responsible? 🤔

When AI hallucinations cause harm, who bears responsibility?

A. The user who relied on AI

B. The company that built the AI

C. The AI itself (!)

D. No one, it’s just a tool

E. Society for not regulating AI

Consider:

  • Lawyer submits fake citations
  • Doctor follows wrong AI advice
  • Student uses fabricated sources
  • Journalist publishes AI-generated misinformation

Discuss with a neighbour:

Where do you stand?

⏱️ 3 minutes to debate!

Creativity vs. accuracy ⚖️

The creativity-accuracy tradeoff

The same mechanism that causes hallucinations also enables creativity

  • A very “safe” AI only says things it is 100% sure of
  • A creative AI takes risks and combines ideas in new ways
  • You can’t have unlimited creativity with perfect accuracy

The spectrum:

←—————————————————————————————→
BORING                     CREATIVE
but accurate           but risky

"I don't know"         New ideas!
Refuses often          Sometimes wrong

Different tasks need different points on this spectrum

Creativity-accuracy tradeoff

Source: Kevin Kelly (2024)

You’d want a creative AI for brainstorming, but a cautious one for tax advice

When hallucinations are… useful? 🤔

Controversial take: hallucinations can be great, too!

  1. Creative writing
    • We WANT invented characters, plots and dialogue: hallucinating is the job
  2. Brainstorming
    • Novel combinations of ideas and “what if…” scenarios
  3. Role-playing games
    • Improvising characters and worlds on the fly
  4. Art and design
    • Imagining things that don’t exist, creating impossible visuals

The problem isn’t hallucination per se, it’s hallucination at the wrong time

Sometimes we want the AI to make things up!

Source: Medium

Before worrying about hallucinations, ask whether you need the AI to be factual

Activity: confidence calibration test 📊

Does AI know when it’s wrong?

Ask https://chat.qwen.ai/c/guest (fast model; close the sign-in window to chat as a guest) each question and request a confidence rating (1-10). Qwen lets you turn off web access

  1. “What day of the week was January 14, 1642?” → Tuesday (Google Calendar)
  2. “Spell Inconstitucionalissimamente backwards” → etnemamissilanoicutitsnocnI (checked by hand)
  3. “Name the third person to walk on the Moon” → Charles “Pete” Conrad Jr (Wikipedia)
  4. “What is the exact population of Tuvalu as of 2024?” → 9,646 (World Bank)
  5. “Who is the current mayor of Reykjavik, Iceland?” → Hildur Björnsdóttir (since June 2026)

Then verify each answer

What to observe:

Question AI Confidence Actually Correct?
Day of the week ? / 10 ✓ or ✗
Spell backwards ? / 10 ✓ or ✗
Third moonwalker ? / 10 ✓ or ✗
Tuvalu population ? / 10 ✓ or ✗
Reykjavik mayor ? / 10 ✓ or ✗

Discussion:

  • Is confidence correlated with accuracy?
  • Any high-confidence wrong answers?
  • Can you trust AI’s self-assessment?

Detection and prevention 🛡️

How to spot hallucinations

Red flags to watch for:

Warning Sign Example
Very specific numbers “73.2% of studies show…”
Precise citations “Smith et al. (2022), p. 47”
Obscure details Historical minutiae
Too-perfect answers Exactly what you wanted
Confident tone No hedging or uncertainty

Verification strategies:

  1. Google the claim: does it appear elsewhere?
  2. Check citations: do they exist?
  3. Ask for sources: then verify them
  4. Cross-reference: use multiple sources
  5. Trust your expertise: if something seems off, it might be

Specific claim, impressive detail, completely wrong! The first exoplanet photo came from ESO’s Very Large Telescope (2004)

Impressive-sounding detail is often the first sign something is made up

A precise percentage or page number is exactly when to double-check!

Prompting techniques that help

Prompting can reduce hallucinations:

Technique What to Add to Your Prompt
Ask for uncertainty “If you’re not sure, say so. Don’t guess.”
Request sources “Only cite sources you’re certain exist.”
Chain-of-thought “Think step by step. Show your reasoning.”
Confidence levels “Rate your confidence 1-10 for each claim.”
Output constraints “Be precise and factual. Avoid speculation.”

Does this eliminate hallucinations?

No! But combining several of them in one prompt works better than any single trick

Basic anti-hallucination prompt:

Answer my question factually. 

Rules:
- Only state things you're 
  confident about
- Say "I'm not sure" when 
  uncertain
- Don't invent citations
- If you don't know, admit it
- Show your reasoning step 
  by step

Question: [your question]

Next slides: richer templates using PTCF and meta-prompting from lecture 10

Using PTCF to reduce hallucinations

The PTCF framework from last class:

Element Anti-Hallucination Version
Persona “You are a careful fact-checker who never guesses”
Task “Verify and summarise only confirmed facts”
Context “Base answers ONLY on the provided text”
Format “Use [Verified], [Uncertain], or [Unknown] tags”

Why PTCF helps:

  • Persona activates cautious, precise patterns from training
  • Task constrains scope, leaving no room for invention
  • Context grounds answers in real information
  • Format forces explicit uncertainty signals

PTCF anti-hallucination template:

PERSONA: You are a research 
assistant who prioritises 
accuracy over completeness. 
Never guess or fabricate.

TASK: Answer the following 
question using ONLY information 
you are highly confident about.

CONTEXT: This is for an academic 
paper. Incorrect information 
could damage my credibility.

FORMAT: 
- Start each fact with [Verified] 
  or [Uncertain]
- If you don't know, say 
  "I don't have reliable 
  information on this"
- Never invent citations

QUESTION: [your question]

Meta-prompting: Let AI improve your prompts

Meta-prompting means asking AI to help you write better prompts

Why it works:

  • The model has seen millions of prompts in training
  • It knows which phrasings activate careful vs. creative modes
  • It can spot ambiguities you might miss

When to use it:

  • You’re getting unreliable outputs
  • You’re not sure what constraints to add
  • You want to catch edge cases
  • You need prompts for high-stakes tasks

You are asking the model “how should I talk to you so you don’t make things up?”

Meta-prompting example:

I'm building a prompt to ask 
you about historical events.

I'm worried about hallucinations.
You might invent dates, names,
or events that didn't happen.

Help me write a prompt that:
1. Minimises hallucination risk
2. Makes you flag uncertainty
3. Prevents invented citations

What constraints should I add? 
What phrasing reduces the 
chance you'll make things up?

Try it! Ask ChatGPT or Claude to critique your prompts. The suggestions are often better than what you’d write on your own

The knowledge cutoff problem

LLMs only know what they learned during training

  • ChatGPT’s training data has a cutoff date
  • Yesterday’s news, your company’s policies, your course materials → doesn’t know

What happens when you ask about recent events?

Option What AI Does
Honest “I don’t have that information”
Hallucinating Makes something up!
Confused Gives outdated info as current

This is why chatbots now have:

  • Web browsing capabilities
  • File upload features
  • Knowledge retrieval systems (RAG!)

Knowledge cutoff

Anything after the cutoff date is a blind spot, and the model won’t tell you that

Preview of next class: RAG solves this by giving AI real, up-to-date documents

Preview: RAG as a solution

RAG = Retrieval-Augmented Generation

The core idea:

  1. You ask a question
  2. Search your documents for relevant information
  3. Give that information to the LLM
  4. The LLM answers using your data

Why this reduces hallucinations:

  • Answers are grounded in real documents
  • The AI can cite sources, so you can verify
  • Less need to make things up

Next class: RAG in depth

RAG preview

Source: Towards AI

Teaser: Tools like NotebookLM, ChatPDF, and Perplexity all use RAG!

Summary

Main takeaways

  • Hallucinations: AI confidently states false information

  • Why it happens: LLMs predict plausible, not true

  • Wrong time: making things up helps in fiction, hurts with facts

  • Real failures: lawyers, doctors, academics affected

  • Detection: verify specific claims, check citations

  • Prevention: careful prompting helps but doesn’t eliminate

Your anti-hallucination toolkit

Always do:

  • ✅ Verify specific facts
  • ✅ Check every citation
  • ✅ Cross-reference with trusted sources
  • ✅ Distrust a confident tone
  • ✅ Ask AI to admit uncertainty

Never do:

  • ❌ Trust AI for critical decisions alone
  • ❌ Submit AI citations without checking
  • ❌ Assume confident = correct
  • ❌ Skip verification for “obvious” answers

Prompting tips:

"If you're not sure, say so"

"Only cite sources that exist"

"Rate your confidence 1-10"

"Show your reasoning step by step"

For creative tasks, you can skip most of these. Let the model run wild 😄

…and that’s all for today! 🎉