DATASCI 101: Introduction to AI Applications

Lecture 19: Privacy and Data Protection

Danilo Freire

Department of Data and Decision Sciences
Emory University

Welcome back!

Recap of last class

  • EU AI Act: first comprehensive AI law; banned, high, limited and minimal risk
  • Fines up to 7% of global turnover; high-risk rules delayed to December 2027
  • US approach: sectoral regulation with state-level initiatives
  • A 2025 executive order lets federal lawyers challenge state AI laws
  • China: content control, innovation goals, algorithm registration
  • No global consensus yet, but convergence on key principles is likely
  • Today: how does AI intersect with privacy?

Oh, well 😅

Lecture overview

What we will cover today

Part 1: Privacy in the AI age

  • Why AI makes privacy harder
  • Data collection at unprecedented scale
  • Inference and prediction of sensitive attributes

Part 2: Legal frameworks

  • GDPR and data protection principles
  • Rights of individuals
  • US privacy landscape

Part 3: Technical approaches

  • Differential privacy
  • Federated learning
  • Privacy-preserving machine learning (synthetic data)

Part 4: Challenges and tensions

  • Case study: Clearview AI
  • Discussion: where’s your line on privacy?
  • What can you do?

Meme of the day

Source: CanIPhish

Good news of the day

Source: Financial Times

Privacy in the AI age

Why AI makes privacy harder

AI breaks privacy in ways older laws did not anticipate:

Scale of collection:

  • Old surveillance was limited by staffing; AI collection runs around the clock at almost no cost
  • Every interaction becomes “data exhaust”, a byproduct of normal use

Inference capabilities:

  • AI predicts sensitive information from innocuous data: shopping → health, typing → mood, friends → politics
  • Things you never disclosed can be inferred

Persistence:

  • Data doesn’t decay: models trained today affect you forever, and decisions follow you across contexts
  • “Right to be forgotten” is technically hard

Source: AI Multiple

“If you’re not paying for the product, you are the product.” This understates it! Even when you pay, your data is often the real product!

What can be inferred from your data?

Research has shown AI can predict:

Data source What can be inferred
Facebook likes Political views, sexuality, personality
Smartphone sensors Depression, anxiety, Parkinson’s
Typing patterns Age, gender, emotional state
Purchase history Pregnancy, health conditions
Location data Home address, workplace, religion
Voice recordings Emotional state, health, age

Famous example: Target and pregnancy

  • Target’s algorithm spotted a pregnant teenager and posted baby coupons to her home
  • Her father complained to the store: she was “still in high school”
  • Days later he apologised to the manager: she was pregnant
  • The algorithm knew before her family did (NYT)

Source: Time

You can control what you share, but you cannot control what can be inferred from it

Training data and privacy

LLMs have a training data problem:

  • Trained on internet-scale data, which includes personal information
  • Models can memorise and regurgitate training data
  • Your name, address, phone number might be in there!

Demonstrated attacks:

  • Carlini et al. (2021) extracted verbatim training data from GPT-2: names, phone numbers, emails
  • Extraction attacks keep improving

The consent problem:

  • Most people don’t know what’s in training sets
  • “Publicly available” ≠ “consented to AI training”

Source: The Hacker News

When you ask an LLM about yourself, it might actually know things from training data you never shared with it directly

Discussion: would you share?

Quick poll (raise your hand):

Would you share your data if…

  1. A health app predicts disease risk but sells data to insurers?
  2. A smart home device improves comfort but records all conversations?
  3. A job search site personalises results but shares with employers?
  4. A social app connects you with friends but builds a profile for advertisers?
  5. An AI tutor helps you learn but reports to your school?

The usual pattern:

  • People say they care about privacy but don’t act like it
  • This is the privacy paradox

Why?

  • Benefits are immediate; harms are distant and abstract
  • Default settings favour sharing
  • Terms of service are unreadable
  • “Everyone does it” normalisation
  • We’re not good at probabilistic thinking

GDPR fundamentals

The General Data Protection Regulation (2016/679, effective 2018) is the EU’s comprehensive privacy law.

Main principles:

  1. Lawfulness, fairness, transparency
  2. Purpose limitation: use data only for the stated purpose
  3. Data minimisation: collect only what’s necessary
  4. Accuracy and storage limitation: keep it correct, and no longer than needed
  5. Integrity and confidentiality: protect data properly
  6. Accountability: demonstrate compliance

Six lawful bases: consent, contract, legal obligation, vital interests, public task and legitimate interests

  • Legitimate interests is the most contested: firms use it to skip consent, even for AI training

Like the EU AI Act, GDPR covers any organisation processing EU residents’ data, wherever it is located. A US startup with European users must comply with no office in the EU.

Individual rights under GDPR

You have the right to:

Right What it means
Access Get a copy of your data
Rectification Correct inaccurate data
Erasure “Right to be forgotten”
Portability Move data to another service
Restriction Limit processing
Object Stop certain processing
Automated decisions Ask for a human in significant ones

Article 22: automated decisions

  • You have the right not to be subject to purely automated decisions
  • Consequential decisions need human involvement
  • Exceptions exist for contracts and explicit consent

Source: GDPR.eu (Proton)

In practice: exercising these rights is often difficult. Companies hide the forms, respond slowly, or claim exemptions.

GDPR and AI tensions

  1. Purpose limitation vs model training
    • You gave data for one purpose. Can it train a model for another?
    • AI companies claim legitimate interest
  2. Data minimisation vs big data
    • AI works better with more data; GDPR says collect only what’s necessary
    • What’s “necessary” for a foundation model? Everything?
  3. Right to explanation vs black boxes
    • You can ask why an AI decided about you
    • Many AI systems can’t explain themselves
  4. Right to erasure vs model training
    • Can you demand removal from a trained model?
    • “Unlearning” is technically very difficult

Source: CertPro

GDPR was written before the LLM era. Applying 2018 law to 2026 technology creates interpretation challenges

US privacy landscape

No comprehensive federal privacy law. Instead, a patchwork of sector-specific rules:

Law Scope Protects
HIPAA Healthcare data Patient medical records
FERPA Education records Student academic data
COPPA Children’s data Under-13 online activity
GLBA Financial data Bank and loan records
FCRA Credit reporting Credit scores and history
ECPA Electronic comms Emails, calls, stored data

State laws filling the gap:

  • California (CCPA/CPRA): most comprehensive, GDPR-like rights
  • Virginia, Colorado, Connecticut: similar frameworks
  • Illinois BIPA: biometric data, with private right of action
  • 23 states now have comprehensive privacy laws (Vermont in June 2026)

Source: IAPP

GDPR enforcement

Major fines (selected):

Company Fine What they did wrong
Meta (2023) €1.2B Sent EU user data to the US without adequate safeguards
Amazon (2021) €746M Targeted ads without valid consent (fine annulled in 2026, breach upheld)
Meta (2022) €405M Exposed children’s contact details on Instagram
Google (2022) €150M Made rejecting cookies harder than accepting them
TikTok (2023) €345M Failed to protect children’s privacy settings and data

Patterns:

  • Big tech companies are the primary targets
  • Regulators focus on data transfers, consent and children
  • Fines are getting larger
  • Enforcement varies by country; Ireland’s regulator is often called slow

Source: Secureframe

€1.2 billion sounds big, but Meta made ~$135 billion in 2023. Is that a fine or a cost of doing business?

Technical approaches to privacy

Differential privacy

Differential privacy is a mathematical framework for privacy-preserving data analysis (Dwork & Roth, 2014).

  • Add calibrated random noise to query results before releasing them
  • “How many people here have diabetes?” The true answer is 137; the system returns 134 or 141
  • Individual records stay hidden; aggregate patterns stay visible

Formal guarantee:

  • The output is nearly the same whether or not any single individual is in the dataset
  • Epsilon (ε) measures privacy loss: small ε means strong privacy and noisier results, large ε the reverse
  • This is provable, not just a promise

Source: Flower AI

Real-world use: Apple for emoji suggestions, Google for Chrome usage stats, the US Census for 2020 data

Federated learning

Federated learning trains models without pooling the data (McMahan et al., 2017)

How it works:

  1. A server sends the same model to many devices
  2. Each device trains it on its local data
  3. Each sends back only what it learned (weights), not the data
  4. The server averages the updates and sends the better model back

Why it matters for privacy:

  • Your raw data never leaves your device: the server sees only model updates
  • Not perfect: updates can still leak information
  • Coordinating thousands of devices is complex

Source: Wikipedia

Example: Google’s Gboard improves predictions this way. Your typing stays on your phone; only model improvements are shared

Synthetic data for privacy

Synthetic data is data that looks real but isn’t (Jordon et al., 2022).

How it works:

  • Train a generative model to learn the real data’s patterns
  • Generate new, fake data points that look real
  • Train your AI on those instead

Privacy benefits and limitations:

  • Fewer direct links to real people, but generators can still leak training records
  • Easier to share for research and collaboration
  • It may not capture rare cases well

Source: GOV.UK

Real-world use: healthcare bodies use synthetic patient data for research; banks test fraud detection on synthetic transactions

Do these techniques actually help?

What they can do:

  • Allow research on sensitive medical and financial data while protecting privacy
  • Give engineers concrete tools
  • Differential privacy gives mathematical guarantees
  • Federated learning keeps data on the device

What they cannot do:

  • Stop companies collecting data in the first place
  • Fix the power imbalance between users and platforms
  • Make users understand how their data are used

However…

The best protection is not collecting data at all. But AI companies have the opposite incentive: more data = better models = more profit

Technical fixes work within the system, not on the system itself

Watch out for “privacy washing”: companies announce differential privacy with large epsilon values (weak privacy), or federated learning that still collects metadata

Challenges and tensions

Case study: Clearview AI

  • January 2020: the NYT revealed Clearview had scraped 3 billion photos from social media (70+ billion today)
  • It built a face search engine without consent: people never knew they were in it
  • It sold access to thousands of US police agencies, which ran nearly a million searches by 2023

The legal fallout:

  • The ACLU sued under Illinois BIPA (2020): a 2022 settlement bans sales to most private companies
  • France (CNIL) and Italy: €20M fine each
  • Netherlands: €30.5M fine (2024)
  • Australia and Canada: ordered to delete data; the UK case is still in court

Source: Library of Congress

Clearview said the service was only for law enforcement, but gave accounts to retailers, banks and investors’ friends. Critics call it a “perpetual police line-up”.

Quick quiz: what’s wrong and how would you fix it?

  1. A social media company uses your messages to train an AI without telling you

  2. A shopping app collects your exact location every 30 seconds, even when not in use

  3. A company keeps customer data “just in case” with no deletion policy

  4. An AI hiring tool rejects candidates without any human review

Where’s your line?

Scenario:

A new app offers free health monitoring. It tracks:

  • Your heart rate and sleep
  • Your location and activity
  • What you eat and drink
  • Your social interactions

In exchange, it provides:

  • Personalised health advice
  • Early warning of health issues
  • Discounts from health insurers
  • Connection with others like you

Let’s discuss together:

  1. Would you use this app?
  2. What would make you change your mind?
  3. What data would be “too much”?
  4. Does it matter who runs it (tech company, hospital, government)?

There’s no right answer. The point is to identify your own values and understand what trade-offs you’re willing to make

What can you do?

Individual actions:

  • Review privacy settings and use privacy-focused tools (Signal, DuckDuckGo, Firefox)
  • Limit location sharing and use different email addresses for different services
  • Exercise your GDPR/CCPA rights (access, deletion requests) and be sceptical of “personalisation”

Limitations of individual action:

  • Opting out often means losing service. Privacy is collective
  • Your data can be inferred from others’ data, and power imbalance is structural

Collective action:

  • Support privacy legislation and demand transparency from companies
  • Support organisations fighting for privacy (EFF, EPIC, noyb)
  • Choose privacy-respecting services; vote for candidates who protect privacy

Using Signal instead of WhatsApp helps, but it won’t change how insurers use your health data. That takes legislation

Summary

Main takeaways

AI and privacy

  • Collection at unprecedented scale
  • Inference makes “non-sensitive” data sensitive
  • Models can memorise and leak personal data

Legal frameworks

  • GDPR: comprehensive and rights-based; US: sectoral patchwork
  • AI development and data protection pull against each other

Technical approaches

  • Noise, decentralised training, artificial data
  • All have trade-offs

Key insights

  • Collection is the core problem
  • The same data enables benefits and surveillance
  • Privacy is collective, not just individual
  • Technical fixes don’t address power imbalances
  • Rules and oversight decide whether data helps or harms

… and that’s all for today! 🎉