What is a coding agent? A plain-English guide

How AI coding agents read, plan, edit, run and check code, how they differ from autocomplete, and how to keep one safe on your machine.

Updated October 2026

A coding agent is an AI program that works on a codebase by itself: it reads files, edits them, runs commands such as tests, looks at the results and keeps going until the task is done. Autocomplete suggests the next few lines; a coding agent carries out a whole task and checks its own work.

Last checked: October 2026. This page is published by the makers of Darce, an open source terminal coding agent.

What is an AI coding agent, in plain English?

Think of the difference between a spellchecker and an assistant you hand a job to. Autocomplete tools, like the early versions of AI code completion, watch you type and suggest what comes next. You stay in the driver's seat for every line.

A coding agent takes a goal instead: "add input validation to the signup form and a test for it". It then decides which files to open, what to change, which commands to run, and whether the result works. You review what it did, or approve risky steps along the way.

The word "agent" just means the model is allowed to take actions (read a file, write a file, run a command) and see the results, in a loop, rather than only producing text.

How does a coding agent work? The agent loop

Almost every coding agent runs the same loop:

  1. Read. It searches and opens the files that look relevant. Many agents also read instruction files such as AGENTS.md or CLAUDE.md for project rules.
  2. Plan. The model decides what needs to change. Some agents show a plan first or have a read-only plan mode.
  3. Edit. It writes changes to files, usually as small targeted edits rather than whole rewrites.
  4. Run. It runs commands: tests, the build, a linter, a code generator, a package install.
  5. Check. It reads the output. If a test fails or the build breaks, it goes back to step 1 with that new information.

The loop ends when the model decides the task is done, when it needs your input, or when you stop it. Good agents finish with a summary of what changed. Darce, for example, ends each task with a receipt listing files changed, commands run, the highest risk taken, models used and cost.

That run-and-check step is what separates an agent from a chat window. A chatbot can write code that looks right; an agent can find out whether it actually works.

Harness vs model: what's the difference?

A coding agent has two parts:

  • The model is the large language model doing the reasoning: Claude, GPT, Gemini, Qwen, DeepSeek, Kimi and so on. It decides what to do next.
  • The harness is the program around the model. It gives the model its tools (read, edit, run), decides what context to send, asks you for approval, keeps history, and handles undo.

The same model can behave very differently in two harnesses, and the same harness can run many models. Some tools tie the two together: Claude Code runs Claude models, and Codex CLI is built around OpenAI's. Others are model-agnostic: OpenCode, Pi and Aider work with your own keys for many providers, and Darce gives you 250+ models through one plan.

When you compare agents, compare the harness features (safety, undo, parallel work, Windows support) separately from the model quality. You can often swap the model later; the harness is what you live in.

What kinds of coding agents are there?

IDE agents

These live inside your editor as an extension or a full editor. You see diffs inline and accept or reject them. Examples include Cursor, the Cline and Kilo Code extensions, and Google's Antigravity desktop app. They suit people who want to watch every change in a familiar editor.

CLI (terminal) agents

These run in a terminal, in your project folder, and work with whatever editor you use. They are easy to script, run over SSH, and use in CI. Examples include Claude Code, OpenAI Codex CLI, OpenCode, Antigravity CLI (which replaced Gemini CLI for individual users in 2026), Aider, Pi, Crush, Qwen Code, Goose and Darce. We compare twelve of them in the best CLI coding agents.

Cloud and background agents

These run on someone else's machine. You hand off a task, close your laptop, and come back to a branch or pull request. Examples include OpenAI's Codex Web at chatgpt.com/codex and Amp's cloud agents ("orbs"). They are good for long, well-defined jobs and for running many tasks at once, but you see less of the work as it happens.

The lines blur: many tools now ship more than one surface. Claude Code has a terminal app, a desktop app and IDE extensions; Codex has a CLI, a desktop app and Codex Web; Cline has an editor extension and a CLI.

Coding agent vs Copilot-style autocomplete

AutocompleteCoding agent
InputThe code around your cursorA task in plain language
OutputA suggested line or blockEdits across many files, plus commands run
Runs commandsNoYes: tests, builds, installs
Checks its workNoYes, by reading command output
Who drivesYou, line by lineThe agent, with your approval on risky steps
Main riskA bad suggestion you acceptA bad action taken on your machine

Many products now offer both, so the useful question is less "which product" and more "am I asking for a suggestion or handing off a task?"

What are coding agents good at?

  • Well-scoped changes: add a field, write tests for a module, fix a failing test with a clear error.
  • Tedious work across many files: renames, migrations from one API to another, adding types, updating call sites.
  • Getting oriented: explaining an unfamiliar codebase, finding where something happens, tracing a bug.
  • Grinding through feedback loops: run the tests, fix, run again, until green.
  • Boilerplate and glue: config files, scripts, small UI components.

What are coding agents bad at?

  • Vague goals. "Make it better" leads to sprawling changes. Specific tasks get specific results.
  • Knowing what they don't know. Agents can confidently use an API that doesn't exist or misread a requirement. Tests catch some of this; review catches the rest.
  • Big design decisions. They will happily build the wrong architecture well.
  • Staying in scope. Some agents "helpfully" touch files you didn't ask about.
  • Anything outside your machine. An agent that can run commands can also deploy, migrate a database or push to a remote, and those can't be taken back by any undo button.

How do you keep a coding agent safe?

A coding agent runs real commands as your user. Three kinds of protection matter.

Approvals

The agent asks before certain actions. Tools differ in how they decide: some use per-tool rules (allow, ask, deny), some use approval modes, and Darce scores every command by risk (read-only, changes the project, reaches outside, destructive) and asks for the risky ones with the reason. "YOLO" or full-auto modes skip approvals entirely; useful in a throwaway container, risky on your laptop.

Sandboxes

The agent's commands run with restricted access to the file system or network. Codex CLI sandboxes commands at the OS level, including on Windows; Claude Code supports sandboxing on macOS, Linux and WSL 2. A sandbox limits the damage a command can do, but it doesn't reverse changes inside the folder you gave it.

Checkpoints and undo

The harness snapshots your files so you can roll back. The important detail is what gets snapshotted. Claude Code's own docs say its checkpoints don't track files changed by Bash commands, so if the agent ran rm or a code generator, rewinding leaves that in place. Darce snapshots the working tree before every step, so /undo also brings back files that shell commands created, changed or deleted (in a git repository, for the current session; gitignored files like node_modules aren't snapshotted).

No undo reaches outside your machine. A deploy, a database migration or a git push stays done. That's why approvals and undo work best together: undo for local changes, a pause before anything remote. Our guide to coding agent checkpoints compares how different tools handle this, and Darce's safety docs explain its approval modes in detail.

How do you pick a coding agent?

Ask these questions in order:

  1. Which models do you want? If you already pay for Claude or ChatGPT, their first-party agents (Claude Code, Codex CLI) are the obvious start. If you want to mix vendors, pick a model-agnostic harness.
  2. Editor, terminal or cloud? Pick the surface that matches how you work today.
  3. How do you want to pay? A subscription, your own API keys, or a plan that bundles many models.
  4. What happens when it goes wrong? Check what undo actually restores, and whether risky commands ask first.
  5. Do you need open source? If you want to audit or fork the harness, check the license.
  6. Does it run where you work? Windows support varies; some tools recommend WSL.

Then try two on the same real task and compare the diffs. If you want a side-by-side of the terminal options, read the best CLI coding agents. To try Darce, run npx darce-cli in a project folder, no account needed; see getting started.

Sources

Frequently asked questions

What is an AI coding agent?

An AI coding agent is a program that uses a language model to work on code by itself: it reads files, makes edits, runs commands like tests and checks the results in a loop until the task is done.

What is the difference between a coding agent and Copilot-style autocomplete?

Autocomplete suggests the next line while you type. A coding agent takes a whole task, edits many files, runs commands and checks whether its changes work, asking you before risky steps.

What is a CLI coding agent?

A CLI coding agent runs in your terminal inside a project folder rather than in an editor. Examples include Claude Code, Codex CLI, OpenCode, Aider, Pi and Darce. They work with any editor and are easy to script or use over SSH.

Are autonomous coding agents safe to run on my computer?

They run real commands as your user, so use approvals for risky commands, a sandbox where available, and an agent whose undo covers what commands changed. No agent can undo remote actions like deploys, migrations or a git push.

Is a coding agent the same as the model it uses?

No. The model (Claude, GPT, Gemini, Qwen and others) does the reasoning; the harness gives it tools, asks for approvals and handles undo. The same model can behave quite differently in different harnesses.

Related guides