DF
A lossless block-diffusion drafter for speculative decoding (Inco AI). A 2B draft model predicts a whole token block at once, giving up to 3.43x speedup without changing the target model's output.
โ Activeโ Public accessโ Open sourceSpecialized AI
Parameters
2B (drafter)
parameters
Release date
1 January 2026
Access:DownloadDeployment:๐ป Local
Overview
Classification
Specialized AI
Access & deployment
Download
Local
Weights: Open source
Key parameters
๐งฉ Parameters: 2B (drafter)
๐ฅ Input: text
Technical specification
Parameters
2B (drafter)
parameters
License
Apache 2.0
Hardware requirements
Lightweight 2B draft model (BF16), paired with a target model in speculative decoding; 7-8 draft tokens per verification step.
Modalities
โฌ Input
text
โฌ Output
text
Capabilities and applications
Native model capabilities
Real-time inference
The model's ability to generate responses with very low latency (>1000 tokens/sec) on specialized inference hardware (e.g. Cerebras WSE), enabling interactive, turn-by-turn collaboration with a human.
Category: coding
Language modeling
Ability to predict subsequent tokens and generate coherent natural-language text based on the preceding context.
Category: language
Benchmark results
1 benchmark
GSM8K (przyspieszenie speculative decoding)
concurrency 1
3.43ร
๐ technical_report
Lossless speedup versus the target model's decoding.
Technical architecture
Core Architecture