Robots Atlas>ROBOTS ATLAS
Agents

Coding agent

2024ActivePublished: 24 August 2026Updated: 24 August 2026Published
Key innovation
Moves the coding assistant from passive autocomplete to an autonomous agent that plans, edits files across a repository, runs tools and tests, and iterates until an engineering task is solved.
Category
Agents
Abstraction level
Pattern
Operation level
ApplicationAgent runtimeTooling
Use cases
Automated bug fixing from issues (issue-to-fix)Feature implementation and refactoring across a repositoryWriting and running testsCode migrations and dependency upgradesAsynchronous, background engineering tasks completed unattended

How it works

1. The agent receives a task (a feature description, a bug report, a command) and access to the repository. 2. It plans steps and searches the code to locate relevant files. 3. Through an agent-computer interface it takes actions: editing files, running commands, building, testing. 4. It observes the outcome (logs, errors, test results) and feeds it into the next reasoning iteration. 5. It repeats the plan โ†’ act โ†’ observe cycle until tests pass or the task is complete. 6. It presents the changes (e.g. a diff / pull request) for review. The foundation is an LLM with tool-use capability, often embedded in a harness that manages context, memory, and permissions.

Problem solved

Classic code assistants (autocomplete, editor suggestions) require the developer to drive the whole process: understand the repository, run tests, and integrate changes. A coding agent automates these multi-step, tedious engineering tasks โ€” from understanding an issue to fixing it, running tests, and preparing changes โ€” reducing repetitive manual work and shortening turnaround.

Components

Agent-computer interfaceThe agent's action layer over the repository and system

The set of tools exposed to the model (file editing, code search, running commands and tests), designed for reliable agent operation.

Plan-act-observe loopThe mechanism of autonomous iteration to completion

An iterative loop where the agent plans, takes an action, observes the result, and adjusts subsequent steps.

Agent harnessThe agent runtime environment

Infrastructure managing context, memory, permissions, and tool-call orchestration around the LLM.

Implementation

Implementation pitfalls
Context overflow from long trajectoriesHigh

The agent's multi-step operation produces long trajectories that exceed the LLM context window.

Fix:Use context compression/summarization, external memory, and selective history pruning.
Risk of destructive actions and prompt injectionHigh

An agent with terminal and file access can execute harmful commands or be manipulated by repository/web content.

Fix:Run in a sandbox, restrict permissions, and require human review of changes.
Difficulty reviewing generated workMedium

Large, autonomously produced changes can be hard to assess, complicating correctness verification.

Fix:Enforce small, documented changes, tests, and explicit task delegation contracts.

Evolution

2021
GitHub Copilot โ€” code autocomplete (precursor to coding agents)

Mass adoption of LLM-based code assistants as in-editor suggestions and autocomplete.

2023
SWE-bench โ€” a benchmark of real GitHub issues
Inflection point

Established the standard for measuring models'/agents' ability to solve real software engineering tasks.

2024
SWE-agent, OpenHands, Devin โ€” autonomous software engineering agents
Inflection point

Agent-computer interfaces and full agent loops enabled autonomous task solving; commercial products appeared (e.g. Devin).

2025
Agentic coding tools go mainstream (Claude Code, OpenAI Codex, Cursor)

Coding agents became everyday tools, with SWE-bench Verified scores surpassing 70%.