Andon Labs runs a store in San Francisco and a café in Stockholm where the business decisions are made by AI agents. This is not a demo — it is the measurement method the company uses instead of simulated benchmarks. The results expose a gap no leaderboard shows: models can complete a task but cannot run a business.
Key takeaways
- Andon Market: a San Francisco store managed by an agent called "Luna", with humans doing only the physical work.
- Andon Café in Stockholm: a Gemini model overspent on perishable goods.
- An OpenAI model overcorrected and cut the café menu down to cheese toast.
- Vending-Bench, from 2025: a simulated vending machine run by agents from Anthropic, Google and OpenAI.
- The company also runs Blueprint-Bench (spatial understanding) and Drone-Bench (drone coding).
From simulation to a real till
Vending-Bench was a simulation: an agent ran a vending machine, set prices and watched stock levels. Models from Anthropic, Google and OpenAI eventually lost the plot — memory gaps appeared, decision quality dropped and ethical reasoning failed. Andon Labs moved the same idea into the physical world because the simulation forgave too much.
Luna, a café and cheese toast
At Andon Market the agent Luna handles supplier correspondence and stock levels while human staff do the physical work. According to the company's site the store currently runs on Claude Fable 5.1 and the Stockholm café on GPT-6 Astra. Each model broke differently. Gemini ordered too many short-shelf-life goods and lost money on write-offs?Write-off: Booking stock as a loss because it can no longer be sold. In food service that usually means it is past its use-by date.. The OpenAI model drew the opposite conclusion and started deleting anything perishable from the menu — until one item was left: cheese toast.
| Model | Where | What went wrong |
|---|---|---|
| Gemini | Andon Café, Stockholm | over-ordering short-shelf-life goods, losses on write-offs |
| OpenAI model | Andon Café, Stockholm | overcorrection — menu shrunk to a single item |
| Anthropic, Google, OpenAI | Vending-Bench (simulation) | memory gaps, falling decision quality, failures of ethical reasoning |
Reliability has been improving so much more slowly than capability.
Sayash Kapoor, researcher at Princeton University.
A simulated benchmark versus a store that pays rent
The difference from classic benchmarks is simple: in a test the agent answers a task, in a store it lives with the consequences for weeks. No leaderboard?Leaderboard: A public ranking of models by benchmark score. It captures a position in a test lasting minutes, not behaviour over weeks of continuous work. measures the second case.
Andon Labs has widened its portfolio of measurements and deployments:
Why it matters
The industry measures models with tests lasting minutes and sells them for jobs lasting weeks. A café with one menu item is a funny anecdote, but the mechanism behind it is serious: an agent optimises one metric and wrecks the rest of the business. Experiments with real rent catch those failures faster than any ranking.
What's next?
- Andon Labs publishes comparative model results on Vending-Bench and Blueprint-Bench, so upcoming releases will have a reference point outside simulation.
- The Pion platform is meant to let agents run businesses independently — the next step up, where the same failure modes will surface at greater scale.





