Robots Atlas>ROBOTS ATLAS
Other

LSI

2026ActivePublished: 17 May 2026Updated: 17 May 2026Published
Category
Other
Abstraction level
Paradigm
Operation level
SystemTrainingPost-training
Use cases
AI safety policy — calibrating expectations about RSI and singularity timelinesAI lab strategy — planning investments in ML4ML tools (autoresearch) with realistic ROI forecastsFraming the "fast takeoff" vs "slow takeoff" debate in the AI safety communityAnalyzing the effectiveness of specific self-improvement systems (AI Scientist, DGM, AlphaEvolve) against LSI frictionCounter-argument to deterministic intelligence-explosion narratives (Yudkowsky 2007/2008)

How it works

LSI identifies three classes of friction in the AI self-improvement loop: (1) narrowness of automatable research — agents are good at optimizing single metrics (test loss) but weak at balancing multiple objectives simultaneously, which is the actual job of top researchers; (2) diminishing returns from parallel agents — analogously to Amdahl's law, scaling the number of agents on a task saturates quickly because they all sample from a similar solution distribution; (3) resource and political bottlenecks — allocation of billions of dollars in compute remains under human organizational control. Each friction translates into "loss" of gain per cycle, bringing the hypothetically exponential RSI down to a more linear trajectory.

Problem solved

The narrative of an impending intelligence explosion via RSI ignores empirical frictions in the AI model development process. LSI provides the language to precisely discuss those frictions and forecast a trajectory that is more linear than exponential.

Components

Narrow-research friction

AI agents are good at optimizing local metrics (test loss, single reward) but not at balancing many competing objectives that characterize real research.

Scale friction (Amdahl's law in AI)

Adding more agents to a problem yields rapidly saturating speedup — all sample from a similar solution distribution and are bottlenecked by human supervision.

Resource and political friction

Allocation of billion-dollar compute budgets is an organizational decision, not an algorithmic one — AI cannot autonomously draw on those resources.

Complexity brake (Paul Allen)

The more we understand intelligence, the harder further progress becomes — a law of diminishing returns for the entire research system.

Implementation

Implementation pitfalls
Confusing linear with exponential growthHigh

Early stages of a sigmoid look exponential. LSI warns that the spectacular 2023-2026 improvements need not be sustained.

Ignoring political resource constraintsHigh

RSI scenarios assuming autonomous access to billion-dollar compute budgets are unrealistic — allocation remains under human organizational control.

Overlooking the human supervision bottleneckMedium

A single researcher cannot meaningfully supervise hundreds of agents per day — this is a fundamental productivity-growth limit.

Over-generalizing from narrow successesMedium

AI successes at optimizing local metrics (Karpathy autoresearch, kernel writing) do not automatically translate to the ability to design entire models.

Evolution

Original paper · 2026 · Nathan Lambert
Lossy self-improvement
Nathan Lambert
2007
Eliezer Yudkowsky — "Levels of Organization in General Intelligence" formalizes the Seed AI and RSI concept (the thesis LSI will later contest).
2011
Paul Allen — essay "The Singularity Isn't Near" introduces the "complexity brake" — a conceptual ancestor of LSI.
Inflection point
2026
Nathan Lambert (Ai2) — essay "Lossy self-improvement" (March 22, 2026) on the Interconnects blog introduces the term LSI and three friction classes.
Inflection point
2026
IEEE Spectrum — article "AI Is Starting to Build Better AI" (Matthew Hutson, May 7, 2026) popularizes LSI in mainstream technical press.
2026
Jason Weston and Jakob Foerster (Meta) — proposal that instead of full self-improvement we should maximize co-improvement; a complementary proposal in the same discourse.

Execution paradigm

Primary mode
Mixture
Activation pattern
Stage dependent

Parallelism

Parallelism level
Partially parallel
Scope
TrainingAcross devices
Constraints
!Amdahl's law — structural cap on maximum speedup from parallel agent count.
!Human supervision as a bottleneck: a single researcher coordinates at most a handful of agents per day.

Hardware requirements

Iterative training with internally generated data requires GPU for both inference (data generation) and training — often on the same cluster.