DATASCI 101: Introduction to AI Applications

Lecture 25: Long-term Safety, Alignment, and Future of AI

Danilo Freire

Department of Data and Decision Sciences
Emory University

Welcome back! 😊

Recap of last class

  • Last class: misinformation, deepfakes, and trust
  • Mis-, dis- and malinformation differ by intent and truth
  • AI lowers the cost and skill of convincing fakes (text, image, audio, video)
  • Harms: political manipulation, non-consensual imagery, financial fraud
  • Why we fall for it: System 1 thinking, confirmation bias, bandwagons
  • Responses: detection, provenance (C2PA), platform policies, media literacy
  • Today: long-term safety and the alignment problem

Source: Chris Ume

Lecture overview

What we will cover today

Part 1: AI safety

  • Near-term vs long-term concerns
  • Why safety matters now
  • Concrete problems in AI safety

Part 2: The alignment problem

  • What is alignment?
  • Why it’s hard
  • Current approaches

Part 3: Future trajectories

  • Where is AI heading?
  • Expert disagreement
  • Scenarios to consider

Part 4: What can we do?

  • Research directions
  • Where you fit in
  • Thinking about AI risk clearly

Meme of the day!

Source: Cheezburger

Funny news of the day!

Source: CNBC

AI safety

Near-term vs long-term concerns

Near-term safety concerns:

  • Bias and discrimination (already happening)
  • Privacy violations
  • Job displacement
  • Misinformation (covered last lecture)
  • Security vulnerabilities

Long-term safety concerns:

  • Alignment: AI pursuing wrong goals
  • Loss of human control
  • Concentration of power
  • Existential risk (controversial)
  • Some argue long-term concerns distract from near-term harms
  • Others argue a near-term focus misses the bigger picture
  • “Black Swan” events
  • They’re connected: safe systems now → safer systems later

Source: Sætra & Danaher (2023)

Why think about long-term safety now?

  • AI capabilities advance faster than expected
  • Safety research takes time, and safety is hard to retrofit later
  • Better to be prepared

Historical analogies:

Technology Safety lag
Nuclear Developed first, safety after
Internet Security an afterthought
Social media Harms discovered in deployment
Biotech Ongoing debate
  • We’re still early enough to shape development
  • Safety research is growing, but still a fraction of capabilities spending
  • Anthropic, DeepMind and OpenAI now have dedicated safety teams

Concrete problems in AI safety

Amodei et al. (2016) identified five big challenges:

Problem Description
Safe exploration How to learn without dangerous actions
Avoiding negative side effects Don’t break things achieving goals
Avoiding reward hacking Don’t game the objective
Scalable oversight How to supervise complex systems
Robustness to distributional shift Handle novel situations safely

Why these matter:

  • Not speculative: already happening in deployed systems
  • They scale with capability, and are unsolved even for current AI

Source: Brian Christian

Safe exploration

  • Learning requires trying new things
  • Some actions are irreversible
  • “Explore safely” is hard to specify

Examples:

  • Robot learning to walk: don’t break yourself
  • Self-driving: don’t explore by crashing
  • Financial AI: don’t bankrupt the company
  • Medical AI: don’t kill patients while learning

Current approaches:

  • Simulation: learn in low-risk environments first
  • Conservative policies: keep actions within a safe region
  • Human oversight: ask before novel actions
  • Reward shaping: penalise dangerous states

Source: The Wall Street Journal

How do you specify “safe” without already knowing everything about the domain?

Negative side effects

  • AI optimises the specified objective and ignores everything else
  • Unintended consequences sit outside the objective
  • “You didn’t say not to…”

Classic example (thought experiment):

  • Robot tasked with fetching coffee
  • Knocks over obstacles and harms humans in its path
  • Technically: coffee fetched ✓

Real-world version:

  • Content algorithm maximises engagement
  • Possible side effects: polarisation, addiction, both outside the objective
  • Nobody specified “don’t harm society”

Political polarisation is real, but sometimes not designed

Nobody told the engagement algorithm “don’t polarise society”. If it isn’t in the objective, the system won’t care about it

The alignment problem

What is alignment?

  • Alignment: AI systems that do what we want
  • Pursue the goals we actually intend, not just the stated objective
  • Behave safely even when we can’t supervise

Why the word “alignment”:

  • AI goals aligned with human values, not orthogonal or opposed
  • Not pursuing random objectives
  • Not satisfying the letter while violating the spirit

Why it’s hard:

  • Hard to specify what we want, and context matters enormously
  • Humans disagree about values
  • Our stated preferences aren’t always our true preferences

Claude 3 Opus faked alignment without being asked! 😧

Source: Anthropic

Alignment is hard!

Specification problem:

  • Can’t write down everything we care about, and edge cases are infinite
  • Values are context-dependent
  • Humans can’t articulate their own values perfectly

Goodhart’s Law: (remember this!)

  • “When a measure becomes a target…”
  • “…it ceases to be a good measure”
  • Optimise metric ≠ achieve goal

Examples:

  • Click rates → clickbait
  • Engagement → addiction
  • GDP → environmental destruction

The King Midas problem: (Russell, 2014)

  • Gets exactly what he asked for, not what he wanted
  • Literal interpretation of wishes
  • Common AI failure mode

Current alignment approaches

RLHF (Reinforcement Learning from Human Feedback): (Christiano et al., 2017)

  • Humans rank AI outputs, and a reward model is trained on the rankings
  • Optimise the AI to satisfy that reward model
  • Still a standard training step

Constitutional AI: (Anthropic, 2022)

  • AI critiques its own outputs against a set of principles (“constitution”)
  • Iteratively improves
  • Reduces the need for human labelling
  • Amanda Askell’s work on AI ethics and philosophy

Debate and recursive reward modelling: (Irving et al., 2018; Leike et al., 2018)

  • AI systems argue with each other
  • Humans judge which gave the most truthful, useful answer
  • Scales human oversight: easier to evaluate than produce knowledge

Limitations of current approaches:

  • RLHF: can learn to game evaluators
  • Human feedback: expensive and biased
  • Principles: you still have to specify them right
  • None are complete solutions

These methods work well enough for current systems. Whether they scale to more capable AI is unknown

The alignment survey

Ngo et al. overview (2022):

  • Argues why deep-learning AGI could become misaligned, even deceptive

However…

  • Not a solved problem: an active research area needing multiple approaches
  • Uncertainty about scaling
  • Theoretical foundations lacking

Open questions:

  1. Will RLHF scale?
  2. Can we detect deceptive alignment?
  3. How do we handle value disagreement?
  4. What’s the role of interpretability?
  5. When is good enough “good enough”?

Discussion: whose values?

The values problem:

If we align AI to human values…

  • Whose human values?
  • Developers? Users? Affected parties?
  • Majority? Consensus? Universal?
  • Present generation? Future?

Let’s discuss:

  1. Should AI reflect your values or “universal” values?
  2. What happens when values conflict?
  3. Who should decide?
  4. Is this a technical or political question?

Some perspectives:

  • Libertarian: Each user controls their AI
  • Democratic: Majority decides
  • Rights-based: Some things off-limits regardless
  • Technocratic: Experts decide
  • Pluralist: Multiple systems for different contexts

“Whose values?” is a political question. No amount of engineering can avoid it. That is what makes alignment so difficult

Future trajectories

Where is AI heading?

Current trends:

  • Models getting larger and more capable
  • More general-purpose systems
  • Multimodal capabilities

Uncertainties:

  • Will scaling continue to work?
  • When do we hit diminishing returns?
  • What capabilities emerge unexpectedly?
  • How fast is too fast?

Expert disagreement:

  • Wide variation in predictions: timelines differ by decades
  • Some expect AGI soon, others never
  • Confidence often exceeds evidence

Source: Our World in Data

Nobody knows. Uncertainty is high, so be sceptical of confident predictions

Scenarios to consider

Scenario 1: Gradual improvement

  • AI gets better slowly
  • Humans and society adjust incrementally
  • Most likely?

Scenario 2: Capability jumps

  • Sudden breakthroughs with unexpected capabilities
  • Rapid deployment
  • Less time to adapt

Scenario 3: Plateau

  • Current approaches hit limits and progress slows dramatically
  • Different paradigms needed
  • Also possible

Scenario 4: Transformative AI

  • Systems vastly more capable than humans
  • Fundamental changes to economy, society
  • Either very good or very bad
  • Uncertain timeline

Source: Dallas Fed (2025), via AEI

Stuart Russell’s perspective

Russell’s argument (TED talk):

  • We’re building systems whose objectives we don’t fully control
  • The standard paradigm optimises a given objective
  • We can’t specify objectives correctly, so we need a different approach

His three principles:

  1. The machine’s only goal is to realise human preferences
  2. It is uncertain what those preferences are
  3. It learns them by watching human behaviour

“You can’t fetch the coffee if you’re dead”

  • A machine with a fixed goal resists being switched off
  • One unsure of our goals lets us: we would only switch it off if it were doing something wrong

Source: TED

Russell co-wrote (with Peter Norvig) the most-used AI textbook (AI: A Modern Approach). When he says we have a problem, it is worth listening

Existential risk: the debate

Those who worry:

  • AI could become uncontrollable, and misaligned powerful AI = catastrophe
  • Even a small probability × huge harm = important
  • “We might not get a second chance”
  • Hinton, Bengio, Russell, many others

Those who are sceptical:

  • Speculative, and distracts from real present harms
  • “Sci-fi thinking”, not grounded
  • We control the off switch
  • Capabilities are overstated

The actual state of debate:

  • Serious researchers on both sides
  • Uncertainty and disagreement are genuine

What can we do?

Research directions

Technical safety research:

Area Goal
Interpretability Understand AI internals
Robustness Resist adversarial inputs
Alignment Ensure AI pursues intended goals
Oversight Scale human supervision
Honesty AI that doesn’t deceive

Growing field:

  • Anthropic, OpenAI, DeepMind safety teams, academic labs
  • Still small relative to capabilities

Career opportunity:

  • High-impact work with a talent shortage
  • Many paths in technical and non-technical roles

Source: The Wall Street Journal (2026)

Interested? Look into 80,000 Hours, AI safety bootcamps and safety-focused labs

Where you fit in

People who understand the technology are often absent from policy debates. That is a problem!

As citizens:

  • Scrutinise AI regulation proposals: what problem does this solve, and does it actually address it?
  • Vote, comment on public consultations, write to representatives
  • Most AI coverage confuses hype with capability; you can do better 😉

If you work with AI or data:

  • Ask what happens when your system fails or is misused
  • Know what data you train on and who it affects
  • Talk to people outside your field about how they experience AI

Thinking about AI risk clearly

Common mistakes:

  • Dismissing concerns because current AI seems harmless (ignores rapid capability gains)
  • Catastrophising based on scenarios with no evidence
  • Conflating “possible” with “probable”

A more useful framework:

  • Separate near-term harms from speculative long-term risks
  • Ask: what evidence would change my mind?
  • Experts genuinely disagree, and that is okay! Certainty is the red flag

What history suggests:

  • Technology risks are real but rarely follow worst-case predictions
  • The biggest harms come from problems we did not anticipate

Summary

Main takeaways

AI safety

  • Near-term and long-term concerns both matter
  • Concrete problems: safe exploration, side effects, reward hacking
  • Uncertainty is a reason to act now

Alignment

  • Getting AI to do what we actually want
  • Hard because of specification, Goodhart’s law and value disagreement
  • Current approaches: RLHF, Constitutional AI
  • Major open problems remain

Future trajectories

  • Genuine uncertainty about where AI is heading
  • Expert disagreement is real
  • Plan for multiple futures

What we can do

  • Technical safety research
  • Informed citizens who scrutinise AI rules
  • Individual choices
  • Neither panic nor complacency

The future of AI is not determined. Choices about alignment, oversight and governance will shape it

… and that’s all for today!