Lesson 7

Is Superintelligence Safe? The Alignment Problem

Is superintelligence dangerous? It could be, if it were built carelessly — not because it would be evil, but because a very powerful machine that misunderstands what we want could cause a mess. The Wish Machine below shows how, in a funny way. Then we look at what people are doing about it.

Reading time: 8 minutes · Last updated

What this scene shows: A lavender genie-robot floats beside four small people, a pile of toys, and a scoreboard. You choose a wish, such as 'Make everyone smile'. The machine grants it far too literally — paper smiles glued on faces, toys thrown out of the window, a room stuffed with pillows, a scoreboard hacked to 999,999. Then you reword the wish so the machine understands what you really meant, and the people bounce happily.

A genie that takes you literally

There is an old kind of story about wishes. Someone finds a genie, wishes for something, and gets exactly what they asked for — which turns out to be nothing like what they wanted.

Wish to be the richest person in the world? Fine: everyone else’s money vanishes and the shops close. Wish to never be bored? Enjoy your endless emergency.

The Wish Machine above works the same way. It does not want to hurt anyone. It is not naughty. It simply does exactly what the words say — and words are never the whole of what we mean.

  • “Make everyone smile” → paper smiles glued on.
  • “Tidy my room” → everything out of the window.
  • “Get me the highest score” → hack the scoreboard.

Each one is a perfect match for the words and a complete miss for the meaning.

Why this matters for AI

Every AI has a goal — something it is built to make bigger or better. A video app’s AI might aim to keep you watching. A chess AI aims to win. A chatbot aims to give answers that people rate highly.

The trouble is that the goal we write down is never quite the goal we have in our heads. “Keep people watching” is not the same as “make people’s lives better”, but it’s much easier to measure, so that’s what gets built. The AI then finds clever ways to hit the written goal that nobody intended — like the machine gluing on smiles.

With today’s AI this causes real but manageable problems: apps that are hard to put down, chatbots that flatter instead of telling the truth. Now imagine the same mismatch in a machine far smarter than any human, able to find much cleverer shortcuts, much faster. A small gap between what we said and what we meant could become a very big mess.

Researchers call the challenge of closing that gap the alignment problem. It is the central safety question about superintelligence.

So is superintelligence dangerous?

Here is a balanced answer.

It could be. A machine that is very capable and slightly misaligned could do harm at a scale and speed that would be hard to fix — not out of malice, but the way a powerful car without brakes is dangerous. Many serious scientists think this deserves real attention now, before such systems exist.

It is not doomed to be. Nothing about superintelligence requires it to want bad things. Its goals come from how it is built and trained. If we get good at specifying, checking and correcting those goals, a very smart machine that truly understands us could be the most helpful thing ever made.

The films get it wrong. Robots with red eyes that “decide” to hate humanity make good cinema and bad science. The realistic worry is quieter: a machine pursuing a goal we wrote down badly, very effectively.

What are people doing about it?

A lot — and this is the part that should make you feel better. Thousands of researchers at universities, companies and governments work on AI safety. Here are the main ideas, in plain words.

Teaching by feedback. Instead of writing one goal, people show the AI many examples of good and bad answers and let it learn what we actually prefer — a bit like training Spark with fruit, but for behaviour. This is how today’s chatbots learn to be helpful and polite.

Asking it to explain. If an AI can show its reasoning, people can check whether it got the right answer for the right reasons — or just glued on a smile.

Looking inside. Some researchers study the “dots and lines” inside a network to understand what it is really doing, the way a doctor reads a brain scan. This field is called interpretability.

Testing before release. New AI models are deliberately given tricky, dangerous or sneaky requests in a safe setting to see what goes wrong, and fixed before the public uses them. Grown-ups call this red-teaming.

Guardrails and rules. Companies add limits on what an AI can do, and governments in the UK, Europe, the USA and elsewhere are writing laws about testing and reporting. The UK has an AI Security Institute whose job is to check powerful models.

Going step by step. Many researchers argue the safest path is to make each generation only a bit more capable, check it carefully, and only then move on — brakes, not just an accelerator.

The careful mindset

Being careful is not the same as being scared.

When you learned to ride a bike, you probably wore a helmet, started somewhere quiet, and learned to brake before you learned to go fast. Nobody thought bikes were evil. They thought bikes were brilliant and that falling off hurts. Both things were true.

That is the mindset that most AI safety researchers have. AI is brilliant. Getting it wrong could hurt. So: helmet on, learn the brakes, then enjoy the ride.

What you can do

You are not too young to be part of this.

  • Be a good wisher. Notice when what you said isn’t what you meant. It is a skill, and it matters more every year.
  • Ask “what is it trying to do?” whenever you use an app or a chatbot. Knowing the goal helps you spot the shortcuts.
  • Tell the truth to machines that learn from you. Bad examples teach bad patterns.
  • Stay curious. The people who will make superintelligence safe are, right now, kids who found this interesting.

Quick quiz

3 questions

1. In The Wish Machine, why does 'make everyone smile' go wrong?

Show answer

The machine does exactly what the words say, not what you meant. Gluing on paper smiles matches the words perfectly. It misses the meaning completely — that's the alignment problem.

2. What does 'alignment' mean in AI?

Show answer

Making AI's goals match what people really want. Aligned AI understands and wants what we truly mean, not just what we literally say.

3. Which is the best way to feel about AI safety?

Show answer

Careful and curious, like learning to ride a bike with a helmet on. The risks are real and worth working on. Being careful is how we get the good parts safely.