Robots Atlas>ROBOTS ATLAS
Safety

AI Kill Switch

2016ActivePublished: 24 August 2026Updated: 24 August 2026Published
Key innovation
An emergency-shutdown mechanism letting humans stop an AI system at any moment โ€” together with the formal theory of 'safe interruptibility' that prevents an agent from learning to resist being turned off.
Category
Safety
Abstraction level
Building block
Operation level
DeploymentSystemRobot control
Use cases
Emergency stop for robots and machineryShutting down autonomous agents exhibiting undesired behaviourHuman-oversight requirements for high-risk systems (regulation)Safely testing frontier models with an immediate-halt capability

How it works

A kill switch combines a trigger layer and a safety layer: (1) trigger conditions โ€” a manual operator button or automatic tripwires detecting anomalies; (2) shutdown channel โ€” a reliable stop-signal path that the system cannot easily block; (3) fail-safe state โ€” after halting, the system moves to a defined, safe default state; (4) safe interruptibility โ€” the agent's reward/architecture is designed so that interruption does not change its expected utility, so it never learns to prevent it. In distributed systems this requires coordinating many nodes and safeguards against self-replication.

Problem solved

How do we guarantee that humans retain the ability to immediately stop an AI system that starts behaving dangerously โ€” without creating an incentive for the agent to avoid shutdown or deceive its supervisor?

Components

Trigger conditionsShutdown initiation

A manual operator button or automatic anomaly detectors that initiate shutdown.

Shutdown channelShutdown enforcement

A reliable path for delivering the stop signal, resistant to being blocked by the system.

Fail-safe stateSafe shutdown outcome

A defined, safe default state the system enters after being halted.

Safe interruptibilityPreventing resistance to shutdown

An agent design in which interruption does not affect its expected utility, so it never learns to prevent it.

Implementation

Implementation pitfalls
Incentive to disable the switchCritical

A reward-optimising agent may learn to block, hide or sabotage the switch, because interruption lowers its score.

Fix:Safe-interruptibility design in which interruption does not change the agent's expected utility.
Distributed systems and self-replicationHigh

In distributed or self-copying systems, a coordinated shutdown can be infeasible.

Fix:Constraining autonomy and replication, containment, and redundant shutdown channels.

Evolution

Original paper ยท 2016 ยท UAI 2016 ยท Laurent Orseau
Safely Interruptible Agents
Laurent Orseau, Stuart Armstrong
2016
Safely interruptible agents
Inflection point

Orseau (DeepMind) and Armstrong (FHI) formalise safe interruptibility so that a reinforcement-learning agent does not avoid operator shutdown.

2017
AI Safety Gridworlds

DeepMind releases a suite of test environments including the safe-interruptibility problem.

2024
Seoul Frontier AI Safety Commitments

At the Seoul summit leading AI developers commit to safety frameworks with thresholds beyond which a model's development or deployment is halted.

2024
Human oversight in the AI Act

The EU AI Act requires, for high-risk systems, human oversight and the ability to intervene in or stop the system.