Amazon has released Strands Decider 2B, an open model that scores predefined answer options instead of generating text. The weights and training code landed on Hugging Face on 1 October 2026 under Apache 2.0, with no hosted API. It is the first free counterpart to TypeSafe AI's closed Jev.
Key takeaways
- The base is Qwen3.5-2B-Base, with a rank-16 LoRA adapter adding just over a million parameters
- JevBench, public split: 167 correct answers out of 231 tasks, or 72.3%
- Brier score 0.348 and expected calibration error 0.050 — the model returns probabilities, not bare labels
- Median latency of 106 ms on an RTX 3090, 296 ms at the 95th percentile
- Training took roughly 70 minutes on eight H100s, or 11 hours on a single RTX 3090
A pointer head instead of next-token prediction
Decider strips the next-token layer off Qwen3.5 and replaces it with a small head that scores every allowed answer option. The agent gets a probability distribution over the candidate list rather than a sentence to parse.
The whole model fits into four lines from its card:
With it goes the whole intermediate layer that wraps every language-model call today: forcing JSON, validating the schema, re-asking for a valid format.
TypeSafe AI demonstrated the same idea in September with Jev, but that model stayed closed and API-only. Amazon shipped the weights, the data and the training scripts, so the result can be reproduced on local hardware.
The numbers, without the marketing
An accuracy of 72.3% is useful but far from certain — one decision in four is wrong. Calibration is the more interesting figure.
What a Brier score is, and why it matters here
The Brier score?Brier score: The mean squared gap between a stated probability and what actually happened. Zero is perfect, 0.25 is coin-flip territory. measures not accuracy so much as the honesty of a model's confidence.
Symbol meaning
- …
- probability the model stated for decision i
- …
- actual outcome: 1 if the option was correct, 0 if not
- …
- number of scored decisions
A calibration error of 0.050 means the model's stated probability tracks reality, so a cutoff threshold can be set deliberately. On narrow tasks the results are clearly better — 0.884 on MuSiQue and 0.872 on ContractNLI.
{
"jevbench_public": { "accuracy": 0.723, "correct": "167/231" },
"calibration": { "brier": 0.348, "ece": 0.050 },
"musique": { "accuracy": 0.884, "n": 1199 },
"contractnli": { "accuracy": 0.872, "n": 1026 },
"held_out_short": { "accuracy": 0.641, "n": 6000 }
}Cost matters as much as quality here. A 106 ms median on a single RTX 3090 means decisions happen locally, inside the agent loop, with no external API call.
Where Decider stands against its rivals
| Model | Size | Median latency | Availability |
|---|---|---|---|
| Strands Decider 2B | 2B | 106 ms | Apache 2.0, weights + training |
| Clef-flash | 9B | 38.8 ms | Apache 2.0 + Workers AI |
| Clef | 27B | 209.3 ms | Apache 2.0 + Workers AI |
| Jev | undisclosed | 524.1 ms | closed, API only |
Why it matters
Agents spend a lot of time on small rulings: whether to call a tool, whether a document is relevant, whether to close the loop. Handing those to a large text-generating model is expensive and unreliable. An open, calibrated 2B model changes the economics of that layer and undercuts the case for a paid API built around it.
What's next
- Amazon launched no hosted API, so deployment requires in-house infrastructure
- The 72.3% figure covers the public JevBench split, the private split has not been released
- The
hobson-v19variant name on the model card points to further adapter iterations
Sources
- VentureBeat — Amazon unveils a free, fast, open source Jev killer: Strands Decider 2B makes decisions in fractions of a second
- Hugging Face — StrandsAgents/strands-decider-2B-hobson-v19
- Strands Agents — Strands Agents SDK





