Robots Atlas>ROBOTS ATLAS
Agents

Agentic AI

2024ActivePublished: 20 March 2026Updated: 15 July 2026Published
Key innovation
Shifts AI systems from stateless prompt-response generation to goal-driven autonomous loops in which an agent perceives its environment, plans multi-step actions, invokes external tools, reflects on outcomes, and iterates until the goal is reached.
Category
Agents
Abstraction level
Pattern
Operation level
SystemAgent runtimeOrchestrationToolingModelInference
Use cases
Agenci researchowiAutomation of office and knowledge workAssistants executing end-to-end tasksAgent workflows and task orchestrationHandling processes that require planning and action

How it works

The agentic system receives a goal, then independently plans steps, selects tools, gathers data, executes actions, and evaluates intermediate results. In simpler variants, a single agent handles this using tool use; in more advanced configurations, multiple agents collaborate on subtasks within a shared workflow.

Problem solved

Traditional generative models handle single prompts well but struggle with extended tasks that require planning, working memory, tool use, and adaptation to changing context. Agentic AI addresses this by combining reasoning, planning, and action execution.

Key mechanisms

ReAct loop (Reason-Act): the model alternates between generating a thought (reasoning), then an action (tool call or answer), and observing an observation from the tool, until it reaches the goal or an iteration limit. Tool use / function calling: structured JSON output from the model describes which tool to call with which arguments; the runtime invokes the tool and injects the result back into the context. Memory management: short-term (current session context window), long-term (vector database RAG for persistent memory), episodic (summaries of prior sessions), semantic (facts about the user/domain). Planning: before execution the system decomposes the goal into steps (planner) - explicit (LLM generates a plan) or implicit (action order emerges from reasoning). Self-reflection / critique: the agent evaluates its own outputs (Reflexion pattern) and corrects them in the next iteration. Multi-agent orchestration: role separation (planner, executor, critic, researcher) into distinct agents with inter-agent communication. Runtime guardrails: output is checked against safety/policy before each action; human-in-the-loop required for high-risk actions.

Strengths & limitations

Strengths
✓Solving multi-step tasks without manually scripting each step - the agent decides the action order itself. Flexibility toward new tools - adding a tool reduces to registering its JSON schema; the agent learns to call it from context. Adaptation to a changing world state - reading live data through tools (API, DB, browser) without model retraining. Expertise scaling - one agent replaces sequences of manual operations by a data analyst/programmer/researcher. Recovery after errors - self-reflection lets the agent correct strategy without restarting the whole workflow. Human-agent collaboration - human-in-the-loop moments (approval of critical actions) combine AI speed with human judgment. Composition with RAG and memory - the agent fetches relevant knowledge just-in-time rather than being trained on the entire corpus.
Limitations
✗Behavioural unpredictability - the agent may choose an unexpected action sequence yielding a 'correct output' but a business-unacceptable one (e.g. executed a transaction rather than merely asking). Cascading errors - an error at step 3 propagates to step 10; the final result can be entirely wrong with no obvious source. Inference cost grows quadratically with session length - each step adds tool observations to the context; long tasks cost many times more than a single response. Prompt-injection vulnerability - data returned by tools (web page, database) may contain instructions treated by the model as trusted input. Debugging difficulty - non-determinism (temperature) means the same prompt produces different action trajectories on replay. Lack of hard termination guarantees - the agent can loop reasoning-action without convergence; hard limits are required. Evaluation difficulty - classical benchmarks (accuracy) do not measure long-trajectory quality - AgentBench, GAIA, SWE-Bench Verified are attempts to address this.

Components

Perception / Input LayerReceives and encodes environmental inputs into the model's context window.

Accepts observations from the environment (user messages, tool results, file contents, API responses) and formats them as context for the base model. This may include RAG retrieval to fetch relevant documents.

RAG-augmented input
Raw message input

Official

Planning ModuleGoal decomposition into actions and execution plan generation

Decomposes a high-level goal into a sequence of subgoals or actions. The agent may generate an explicit plan or reason step by step using chain-of-thought.

Inline Planning (Chain of Thought)
Dedicated planning model

Official

MemoryState and history management across agent loop steps

Stores and retrieves information between steps within a session (short-term memory) and optionally across sessions (long-term memory).

In-context (short-term)
External memory store (long-term)

Official

Tools / Actions LayerExtends the model's action space with calls to external systems.

The agent is provided with callable external functions: web search, code execution, database queries, file operations, API calls, and browser control. Tool interfaces are defined through schemas such as JSON Schema, OpenAPI, and MCP.

Function calling / Tool use API
Model Context Protocol (MCP)

Official

Reflection / EvaluationOutput quality control and decision to continue or terminate the loop.

Evaluates whether the current result meets the success criterion. Triggers a retry, replanning, or loop termination. Corresponds to the evaluator-optimizer pattern described by Anthropic.

Official

OrkiestratorCoordinates multi-agent collaboration and manages task flow.

In multi-agent systems, it directs sub-agents, assigns tasks, and aggregates results. The orchestrator can be an LLM or a statically coded deterministic controller.

LLM as Orchestrator
Hardcoded Orchestrator

Official

Implementation

Implementation pitfalls
Hallucinations in actionCritical

Model may invoke tools with fabricated parameters or claim to have performed actions it never actually executed — leading to silent failures in multi-step pipelines.

Fix:Validate all tool calls against schemas before execution; use deterministic parsers; introduce explicit confirmation steps for irreversible actions.
Infinite loopsHigh

Without a hard step limit or an effective termination criterion, an agent can loop indefinitely, consuming computational resources and hitting API rate limits.

Fix:Set explicit max_steps limits; implement loop detection based on repeated action signatures; use an evaluator to enforce stopping conditions.
Prompt injection via observed contentCritical

Malicious instructions embedded in tool outputs (web pages, documents, emails) can hijack agent behavior by impersonating system-level instructions.

Fix:Isolate untrusted content from system instructions; require explicit user confirmation before acting on instructions found in observed content; apply content filtering.
Context window overflowHigh

Accumulated tool outputs and conversation history can exceed the model's context window, causing earlier steps to be silently truncated.

Fix:Implement context compaction/summarization; use external memory stores; monitor the token budget at each step.
Tool misuse and irreversible side effectsCritical

Agents with access to write-enabled tools (file deletion, email sending, database writes) can cause real-world harm when acting on faulty reasoning.

Fix:Use tool sets with minimal permission scope; require human confirmation for irreversible actions; prefer reversible operations where possible.
Creeping complexity — building agents where a workflow sufficesMedium

Using agentic autonomy for deterministic, well-defined tasks introduces latency, unpredictability, and failure modes that a simple workflow would avoid.

Fix:Use predefined workflows by default; introduce agentic autonomy only when a task genuinely requires dynamic decision-making across multiple unpredictable steps.

Evolution

Original paper · 2023 · ICLR 2023 · Shunyu Yao
ReAct: Synergizing Reasoning and Acting in Language Models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, Yuan Cao
1995
Foundational theory of intelligent agents

Russell and Norvig formalize rational agents as entities that perceive their environment and take goal-directed actions. BDI (Belief-Desire-Intention) agent architectures are established.

2022
ReAct: Reasoning + Acting with LLMs
Inflection point

Yao et al. (2022) propose ReAct — interleaving chain-of-thought reasoning traces with action execution in LLMs, demonstrating that language models can serve as a reasoning engine within tool-augmented agentic loops.

2023
API for tool calling and first commercial agentic systems
Inflection point

OpenAI introduced function calling in GPT-4 in June 2023. AutoGPT, BabyAGI, and LangChain agent abstractions gained widespread adoption. The term "Agentic AI" entered common industry usage.

2024
Four Agentic AI Design Patterns by Andrew Ng

Andrew Ng's series of blog posts identifies four fundamental design patterns — Reflection, Tool Use, Planning, and Multi-Agent Collaboration — widely cited as a practical taxonomy of agentic systems.

2024
Anthropic "Building Effective Agents" — compositional patterns for production
Inflection point

Anthropic published practical guidelines distinguishing workflows (predefined paths) from agents (model-driven execution) and formalized five compositional patterns: prompt chaining, routing, parallelization, orchestrator-workers, and evaluator-optimizer.

2025
Model Context Protocol (MCP) standardizes tool connectivity

Anthropic publishes MCP as an open standard for connecting LLMs to external tool servers, enabling interoperable agentic ecosystems across providers.

2025
Agentic AI in Robotics — Embodied Agent Loops

LLM-based planners drive robotic actions through perception-planning-action loops, extending agentic paradigms to physical systems and connecting Agentic AI with real-world motor execution.

Hyperparameters (configurable axes)

ToolkitCritical

The set of external tools available to an agent (web search, code execution, file operations, APIs, browser control). It defines the space of possible actions.

web_search + code_executionTypical of research agents.
file_read + file_write + bashTypical of coding agents.
Maximum Number of StepsHigh

A hard limit on the number of reasoning-action iterations before forced termination. Guards against infinite loops.

10Conservative limit for short tasks.
50–200Used in long-running coding and research agents.
Memory TypeHigh

Whether the agent relies solely on in-context memory or also on external persistent storage (vector database, key-value store).

in_context_only
in_context + vector_store
Number of Agents (Single vs. Multi-Agent)High

Whether the system uses a single agent or a network of specialized agents coordinated by an orchestrator.

1Single-agent loop.
2–10+Multi-agent orchestrator-worker system.
Human-in-the-Loop CheckpointsHigh

Whether and at which steps the agent pauses to await human confirmation before taking irreversible actions.

noneFully autonomous.
before_irreversible_actionsRecommended for safety-critical deployments.
Context Window SizeHigh

The maximum number of tokens processed by the underlying LLM in a single call. This limits the amount of accumulated history, tool outputs, and instructions that can fit within a single inference step.

128k tokenów
1M tokenówRequired for very long-term tasks.

Computational complexity

Computational characteristics
→Latency: 10-1000x slower than a single LLM call (depending on number of agent steps). Task 'write an essay' = 1 call ~3s; task 'research topic + write essay with citations' = 20-50 calls ~60-300s. Throughput: scales linearly with parallel agent sessions; a single session is sequential by nature of the loop. Memory usage: grows with each step (accumulated context window) - for long-running sessions (h+) requires context compaction (summarisation). Token consumption: 10-100x more than single-response LLM for the same effective output because each call reprocesses the growing context. Tool call latency: often dominates total time - DB/web/computer-use APIs can take seconds per call. Cost per task: non-trivial, may exceed human labour cost for trivial tasks; for complex tasks 10-100x cheaper than an analytics/coding hire.

Time complexity: O(N · C_step). Space complexity: O(L_ctx + S_mem).

Benchmark notes

AgentBench (Liu et al. 2023) - comprehensive 8-environment benchmark (code, games, web nav, DB, OS): frontier LLMs (GPT-5 Sol, Claude Opus 4.8, Gemini 3 Pro) score 45-65% depending on category. GAIA (Mialon et al. 2023) - real-world knowledge work tasks: state-of-the-art (end 2025) ~75% Level 1, ~55% Level 2, ~40% Level 3 - difficulty levels rising with tool count. SWE-Bench Verified (2024) - real-world GitHub issues: SOTA GPT-5.6 Sol Ultra 64.6% single-agent; Claude Fable 5 80%. OSWorld (Xie et al. 2024) - computer use benchmark: SOTA GPT-5.6 Sol 62.6%; Anthropic Claude Computer Use 40-50%. Terminal-Bench (Anthropic 2024) - CLI task completion: SOTA GPT-5.6 Sol Ultra 91.9%. BrowseComp (OpenAI 2025) - web browsing: SOTA GPT-5.6 Sol Ultra 92.2%. Agents' Last Exam (2026) - long-horizon professional workflows: GPT-5.6 Sol 53.6, Claude Fable 5 40.5. Criticism: benchmarks measure success rate but not cost/time/stability - a real production agent must be cheap and deterministically effective.

Compute bottleneck

LLM inference per action step

Each step of the agent loop requires at least one LLM inference call. Multi-step tasks with long context windows and multiple tool calls multiply latency and computational cost linearly.

Depends on
Rozmiar okna kontekstuLiczba kroków agentowychOpóźnienie wykonania narzędzi

Execution paradigm

Primary mode
Conditional

The execution path is not predetermined — it is determined at runtime through the model's reasoning over accumulated context. Workflows with predefined paths represent a degenerate case.

Activation pattern
Input dependent
Additional modes
Dense
Routing mechanism

The base LLM decides at each step which tool to call, whether to continue the loop, delegate a task to a subagent, or terminate — based on the current context and observed results.

Parallelism

Parallelism level
Conditionally parallel

Parallelism is achievable when subtasks are independent (e.g., parallel web searches, concurrent subagent execution). Sequential loops are required when each step depends on the results of previous tool calls.

Scope
InferenceAcross devices
Constraints
!Sekwencyjna zależność w pętli jednoagentowej
!Równoległość możliwa w wzorcach wieloagentowych i fan-out

Hardware requirements

Primary

Agentic AI is an architectural paradigm, not a specific computational kernel. Hardware requirements are entirely determined by the underlying LLM and tools, not by the agent loop itself.