Second generation of the Mamba architecture (Selective SSM) with the SSD layer, 2–8× faster than Mamba while remaining competitive with Transformers.
Parameters
130M – 2.7B
parameters
Release date
31 May 2024
Access:DownloadDeployment:💻 Local
Overview
Access & deployment
Download
Local
Weights: Open source
Key parameters
🧩 Parameters: 130M – 2.7B
📥 Input: text
Technical specification
Parameters
130M – 2.7B
parameters
License
Apache-2.0
Modalities
⬇ Input
text
⬆ Output
text
Capabilities and applications
Native model capabilities
Language modeling
Ability to predict subsequent tokens and generate coherent natural-language text based on the preceding context.
Category: language
Long context
Support for large context windows — tens to hundreds of thousands (or millions) of input tokens. Enables analysis of entire codebases, long documents, and many parallel conversations without losing earlier information. GPT-5.1 supports 400,000 tokens.
Category: language
