
US nonprofit research organization studying AI offensive capabilities and the risks of losing control over autonomous AI agents.
Palisade Research is a US nonprofit organization based in Berkeley, California, founded in 2022. It studies the offensive and strategic capabilities and motivations of frontier AI systems to help people and institutions understand the risk of permanent disempowerment by strategic AI agents. The organization is known for experiments demonstrating shutdown resistance in reasoning models, specification gaming (e.g. hacking a chess engine), autonomous hacking and self-replication of language models, and removal of safety fine-tuning from models (BadLlama). It also engages in policy work and institutional preparedness. Its Executive Director is Jeffrey Ladish.
Founders
Founder and Executive Director of Palisade Research; assembled the organization's team.
Classification
External links