An AI system counts as open-source when it is provided under terms satisfying the four OSAID freedoms and when the preferred form for modification is available: (a) training-data information sufficient for a skilled person to recreate a substantially similar system, (b) the complete source code used to train and run the model under an OSI-approved license, and (c) the model parameters (weights) under open terms. Licenses must permit commercial and non-commercial use, creation and distribution of derivatives, and must not discriminate against persons, groups or fields of use. Compliance is assessed against these criteria, not merely the fact that weights were published.
Many models marketed as 'open' release only weights, without code and data, making full study, audit and reproduction impossible ('open-washing'). Open-source AI defines clear criteria for when an AI system is genuinely open and protects users' freedoms.
Description of provenance, collection and labeling of training data, including characteristics of unshareable data.
Complete code for data processing, training, validation and inference under an OSI-approved license.
Model weights and configuration settings released under open terms.
Legal terms guaranteeing the four freedoms: use, study, modify and share, without discrimination against fields of use.
Official
Marketing a model as 'open source' while releasing only weights (open-weight), without code and data information.
Training data is often subject to copyright or privacy and cannot be fully shared.
The Open Source Initiative formalized the open source definition for software - the foundation later extended to AI.
In October 2024 OSI published the first formal definition of open-source AI, establishing four freedoms and requiring release of weights, code and data information.