Robots Atlas>ROBOTS ATLAS
Data

Active Learning

1994ActivePublished
Key innovation
Shifts control over which data gets labeled from a fixed, pre-collected set to the model itself: the algorithm selects the most informative instances for annotation, reaching high accuracy with far fewer labels.
Category
Data
Abstraction level
Paradigm
Operation level
DataTraining
Use cases
Medical imaging and diagnostics (expensive expert annotation)Low-resource NLP and text classificationSpeech recognitionObject detection and segmentation in computer visionAnomaly detection and rare classesDrug discovery and computational chemistryRobotics and selecting demonstrations for policy learningContent moderation and spam detection

How it works

Active learning operates as a loop. (1) A base model is trained on a small initial labeled seed set. (2) A query strategy scores the informativeness of unlabeled instances via a utility function — e.g. prediction uncertainty (least confidence, margin, entropy), committee disagreement (query-by-committee), expected gradient/model change, expected error or variance reduction, optionally weighted by density/diversity to avoid selecting near-duplicate points within a batch. (3) The most informative instances (one at a time or in a batch of size B) are sent to the oracle for labeling. (4) The new labels are added to the training set and the model is retrained. The loop repeats until the labeling budget is exhausted or a stopping criterion is met. Scenarios differ in the candidate source: pool-based scores an entire pool and picks the best, stream-based decides on each incoming instance individually, and membership query synthesis generates synthetic instances to query.

Problem solved

Labeling data is often the most expensive and slowest part of building ML models, while raw unlabeled data is cheap and plentiful. Passive learning (labeling a random sample) wastes the annotation budget on uninformative examples. Active learning minimizes the number of labels required by directing annotator effort toward the examples that will most improve the model.

Components

Oracle (annotator)Provides labels on demand

An authoritative source of labels — usually a human expert, sometimes another system. It answers the model's queries by providing ground-truth labels for selected instances.

Official

Learner (base model)Learns and produces the informativeness signal

The ML model trained on the growing labeled set; its state (uncertainty, gradients, predictions) drives the selection of the next queries.

Official

Unlabeled pool / streamSource of query candidates

A set or stream of unlabeled instances from which candidates for labeling are selected.

Query strategyDecides what to ask the oracle

The rule selecting instances: uncertainty sampling, query-by-committee, expected model change, expected error/variance reduction, density/diversity-based methods.

Uncertainty samplingSelects instances the model is least certain about (least confidence, margin, entropy).
Query-by-committeeA committee of models votes; instances with maximal disagreement are selected.
Expected model changeSelects instances that would change the current model the most (e.g. largest expected gradient).
Diversity / core-setRepresentativeness- and diversity-aware methods, important in batch-mode AL.

Official

Informativeness measureQuantifies the value of a query

A utility (acquisition) function assigning each instance an expected information-gain score used to rank candidates.

Official

Stopping criterion / budgetControls cost and termination

The condition that ends the loop: exhausted labeling budget, validation-accuracy plateau, or a confidence threshold.

Official

Implementation

Implementation pitfalls
Sampling bias and distribution shiftHigh

An actively collected set is not iid with respect to the target distribution, which can hurt other models trained on that data and complicate fair evaluation.

Fix:Use density/representativeness-weighted strategies and validate on an independent random set.
Cold-start problemMedium

With a very small seed set the model is weak, so its uncertainty signal is unreliable and early queries are poorly chosen.

Fix:Enlarge the seed set, use a pretrained model, or start with diversity-based sampling.
Batch redundancyHigh

Pure uncertainty sampling in batch mode selects highly similar, mutually redundant instances, wasting the labeling budget.

Fix:Combine uncertainty with diversity (core-set, BADGE, clustering) when forming a batch.
Noisy or expensive oracleHigh

The assumption of an infallible oracle is often unrealistic; label noise and variable annotation cost erode active-learning gains.

Fix:Model the noise, use redundant annotations, and apply cost-aware (cost-sensitive) AL strategies.
Miscalibrated uncertainty in deep networksHigh

Softmax in deep networks tends to be overconfident, so raw uncertainty is a poor measure of informativeness.

Fix:Use Bayesian uncertainty (MC-dropout, BALD, ensembles) instead of raw softmax.

Evolution

Original paper · 1994 · Machine Learning, 15(2), 201-221 · David A. Cohn
Improving Generalization with Active Learning
David A. Cohn, Les E. Atlas, Richard E. Ladner
1988
Formalization of query-based learning (membership queries)

Dana Angluin formalizes membership-query learning — the theoretical basis for the membership query synthesis scenario.

Queries and Concept Learning (D. Angluin) (paper)
1992
Query by Committee
Inflection point

Seung, Opper and Sompolinsky introduce a committee of models voting on instances; points of maximal disagreement are selected.

1994
Coining of the term active learning in ML
Inflection point

Cohn, Atlas and Ladner introduce the term and show improved generalization from selective instance selection.

1994
Uncertainty sampling (pool-based) for text classification
Inflection point

Lewis and Gale propose a sequential algorithm for training text classifiers based on uncertainty sampling.

2009
Unifying literature survey
Inflection point

Burr Settles publishes the Active Learning Literature Survey, organizing scenarios and query strategies into a single, widely cited taxonomy.

2017
Deep Bayesian active learning
Inflection point

Gal, Islam and Ghahramani combine active learning with Bayesian deep-network uncertainty (MC-dropout, BALD) for image data.

2018
Core-set approach (diversity/representativeness) for batch AL
Inflection point

Sener and Savarese cast batch selection as a core-set problem, emphasizing representativeness over uncertainty alone in active learning with CNNs.

Hyperparameters (configurable axes)

Query strategyCritical

Choice of acquisition function (uncertainty, QBC, expected model change, diversity/core-set) — governs sampling effectiveness.

Labeling budgetCritical

Maximum number of oracle queries; determines cost and usually the accuracy ceiling.

Query batch size (B)High

Number of instances labeled per round; large B requires a diversity strategy to avoid redundancy.

Base learnerHigh

Type of model (e.g. logistic regression, SVM, deep net); affects the quality and calibration of the uncertainty signal.

Initial seed set sizeMedium

Number of labels at the start; too small causes the cold-start problem and poor early queries.

Stopping criterionMedium

Condition for ending the loop (budget, accuracy plateau, confidence threshold).

Computational complexity

Time complexity: O(R · (C_train + |U|·c_score)).

Parallelism

Parallelism level
Sequential

The canonical loop is sequential (each round depends on the previous one). Batch-mode active learning parallelizes the labeling of many instances within a single round.

Scope
Training

Hardware requirements

Primary

Active learning is a hardware-independent data-selection paradigm; it works with any base model and infrastructure.