
GPT-BERT
Hybrid language model merging a causal (GPT) objective with masked language modeling (BERT) in a single transformer; winning architecture of the 2024 BabyLM Challenge.
🔬 Research✓ Public access⚖ Open sourceLLM
Parameters
119M (base) / 30M (small)
parameters
Release date
31 October 2024
Access:DownloadDeployment:💻 Local
Overview
Classification
LLM
Access & deployment
Download
Local
Weights: Open source
Key parameters
🧩 Parameters: 119M (base) / 30M (small)
✓ Fine-tuning
📥 Input: text
Technical specification
Parameters
119M (base) / 30M (small)
parameters
License
MIT
Hardware requirements
Small model (~30M–119M parameters); training and inference feasible on a single GPU.
Features:✓ Fine-tuning
Modalities
⬇ Input
text
⬆ Output
text
Capabilities and applications
Native model capabilities
Language modeling
Ability to predict subsequent tokens and generate coherent natural-language text based on the preceding context.
Category: language
Classification
Assigning an observation to one of predefined classes (binary or multi-class). Output: class label and optionally probabilities.
Category: other
Technical architecture
Core Architecture
Model Form
Training Techniques