Robots Atlas>ROBOTS ATLAS
Architecture

Embeddings (vector representations)

mature
Category
Architecture
Abstraction level
Primitive
Operation level
ModelInferenceTrainingData
Use cases
First layer of language models (token embeddings + positional embeddings)Semantic search and document similarity (cosine similarity)Retrieval-Augmented Generation (RAG) โ€” backbone of the retrieval stageVector databases (Pinecone, Weaviate, Qdrant, pgvector, Milvus)Recommender systems (user/item embeddings)Document clustering and visualization (t-SNE, UMAP)Transfer learning โ€” pre-trained embeddings as a starting pointMultimodal embeddings (CLIP) joining text and image in one space

How it works

Static embeddings (Word2Vec Skip-gram): 1. For each central word in a context window, predict surrounding words (Skip-gram) or vice versa (CBOW). 2. Training via negative sampling: for each true (word, context) pair, increase P(context|word) and decrease P(negative|word) for random negatives. 3. The resulting embedding vectors are the rows of the input-layer weight matrix.

Contextual embeddings (Transformer): 1. Token Embedding Layer: each token ID is mapped to a d_model-dimensional vector from a [|V| ร— d_model] matrix. 2. A positional embedding (sinusoidal or RoPE) is added to each token embedding. 3. Successive Transformer layers contextually transform these representations โ€” each token "sees" all others via self-attention. 4. Output pooling: a CLS token, mean pooling or last-token pooling yields a sentence/document embedding.

Semantic search: 5. A query and documents are encoded into the embedding space using the same model. 6. Similarity is computed as cosine similarity (or dot product for normalised vectors) โ€” HNSW or IVF for ANN search.

Problem solved

ML models cannot operate directly on discrete objects (words, tokens, categories) โ€” they require a numerical representation of input. One-hot encoding creates very sparse, high-dimensional vectors with no semantic similarity information (cosine distance between any two words is zero). Embeddings solve both problems: they represent objects as dense vectors in a continuous space where geometric proximity reflects semantic similarity.

Key mechanisms

Embedding matrix: a table of vectors of shape [|V|, d], indexed by token ID
Cosine similarity and Euclidean distance as semantic measures
Contrastive training โ€” pull similar objects together, push dissimilar apart
Word2Vec Skip-gram: predict context from a central word
Word2Vec CBOW: predict the central word from context
Negative sampling โ€” efficient approximation of softmax over a large vocab
GloVe: factorization of a global co-occurrence matrix (log-bilinear)
Contextual embeddings โ€” the vector comes from hidden-layer activations in a given context
Pooling (mean, CLS, last-token) to obtain a sentence/document embedding
Matryoshka โ€” embeddings with nested dimensionality levels

Strengths & limitations

Strengths
โœ“Continuous, dense representations โ€” efficient memory and vector operations
โœ“Preserve semantic relations (synonymy, analogies)
โœ“Universal โ€” apply to text, image, audio, graph, multimodal data
โœ“Pre-trained embeddings drastically reduce downstream-data requirements
โœ“Operations on embeddings (similarity search) are very fast โ€” dot product
โœ“Scalability โ€” vector databases handle billions of embeddings at <100 ms latency
โœ“Compositional โ€” arithmetic operations have semantic interpretation
Limitations
โœ—Static embeddings cannot handle polysemy (one meaning per word)
โœ—Sensitive to training data โ€” they reflect social and cultural biases
โœ—Out-of-vocabulary โ€” Word2Vec/GloVe cannot represent unseen words
โœ—High dimensionality increases memory cost (e.g. 1536-D ร— millions of documents)
โœ—Cosine similarity is not a perfect semantic measure โ€” sensitive to anisotropy
โœ—Contextual embeddings require costly pre-training of a large model
โœ—Individual dimensions are not interpretable
โœ—Embeddings from different models are not interchangeable (different spaces)

Implementation

Implementation pitfalls
Curse of dimensionality in cosine similarityMedium

At very high dimensions (e.g. 4096D) cosine differences between vectors shrink โ€” all vectors appear similar. Requires normalization and optional dimensionality reduction (PCA, UMAP).

Domain shift โ€” embeddings from one domain do not transfer to anotherMedium

An embedding model trained on general text (e.g. Wikipedia) generates poor representations for specialized text (medicine, law, code). Requires fine-tuning or a dedicated model.

No embedding refresh after data changesMedium

Embeddings are static after generation โ€” changing a source document does not automatically update its embedding in the vector store. Requires an invalidation and re-embedding system.

Evolution

Original paper ยท 2013 ยท Tomas Mikolov
Efficient Estimation of Word Representations in Vector Space (Word2Vec)
Tomas Mikolov, Kai Chen, Greg Corrado, Jeffrey Dean
1986
Rumelhart, Hinton and Williams introduce distributed representations as an alternative to one-hot encoding.
2003
Bengio et al. publish "A Neural Probabilistic Language Model" โ€” the first neural language model with learned word embeddings.
2013
Mikolov et al. release Word2Vec (CBOW + Skip-gram) โ€” embeddings become a standard NLP tool.
2013
Mikolov et al. publish the Word2Vec extension with negative sampling and hierarchical softmax โ€” drastic training speedup.
2014
Pennington, Socher, Manning publish GloVe (Global Vectors) โ€” embeddings based on a global co-occurrence matrix.
2016
Bojanowski et al. (Facebook AI) release fastText โ€” n-gram embeddings that handle OOV words.
2018
ELMo (Peters et al.) and BERT (Devlin et al.) introduce contextual embeddings โ€” a vector that depends on the surrounding context.
2019
Reimers and Gurevych publish Sentence-BERT (SBERT) โ€” efficient sentence embeddings for retrieval and clustering.
2021
OpenAI releases its first public embedding API (text-embedding-ada-001/002) โ€” the start of the commercial embedding-model era.
2024
Matryoshka Representation Learning (Kusupati et al.) and OpenAI text-embedding-3 expose adjustable-dimension embeddings.

Computational complexity

Computational characteristics
โ†’Memory: |V| ร— d ร— 4 B (e.g. 50,000 ร— 1024 ร— 4 B โ‰ˆ 200 MB)
โ†’Embedding-lookup inference: O(1) per token (table indexing)
โ†’Similarity search: O(N) naively, O(log N) with ANN (HNSW, IVF, PQ)
โ†’Word2Vec training on a ~1B-token corpus: a few hours on CPU
โ†’Contextual-model training (BERT-base): a few days on 16 TPU/GPU
โ†’Quantization (int8, binary) reduces memory 4โ€“32ร— with minor quality loss
โ†’Typical dimensionality: 50โ€“300 (static), 384โ€“1024 (sentence), 1024โ€“4096 (LLM)
Benchmark notes

Standard benchmarks: word-analogy (Google Analogy, 19,558 pairs; Mikolov 2013), word-similarity (WordSim-353, SimLex-999), MTEB (Massive Text Embedding Benchmark, ~58 tasks). Word2Vec 300-D reaches ~72% top-1 on Google Analogy. SBERT improves the Spearman correlation on the STS Benchmark to ~0.85. In MTEB (2024) leading models (Cohere Embed v3, OpenAI text-embedding-3-large, BGE-M3) score >65 on average, while classical TF-IDF and Word2Vec-mean trail significantly (~40).

Hardware requirements

Generating embeddings for large document collections (batch encoding) is significantly faster on GPU โ€” models like text-embedding-3 or BGE-M3 leverage CUDA.