You created a subagent because somebody said it would save you context, and instead your five-hour usage window emptied in fifteen minutes and you are not sure the thing even finished. Or you installed a pack of a hundred ready-made agents, and now Claude picks the wrong one, or none of them. Subagents in Claude Code are genuinely useful, but almost everything written about them is either the official mechanics or a screenshot of somebody’s folder of agent files. The part that decides whether they help or hurt is missing: when to hand work off, what to hand it to, and what it actually costs you.
I run my back office on Claude Code. More than thirty scheduled jobs across eight products, firing every day whether I am at my desk or on a plane, and a lot of that work is split across subagents so one noisy task does not drown the run that is steering it. That fleet has also taught me, expensively, where subagents go wrong. This is the split I actually use, the one failure that cost me a morning of usage, and a template you can copy today.
On this page
- The short answer
- The three kinds of help I hand off
- Forks, Codex, or a specialist: which to reach for
- Do subagents save tokens or burn them
- What a subagent can and cannot see
- Build a lean set, not a 100-agent zoo
- When a subagent is worth it, and when to just prompt
- How to set one up (copy this)
- Frequently asked questions
- Where this comes from
The short answer
Use a subagent when the work is self-contained, likely to produce more output than you want sitting in your main conversation, and either repeated often or safe to run in parallel. Reach for a plain fork first, an ad hoc subagent you spin up for one research or file-reading task. Promote it to a named specialist only when you find yourself re-briefing the same job over and over. Send a task to Codex, a second model, when it is complex backend work or when you simply want to spare your Claude usage limits. Skip the hundred-agent mega-packs.
The one line to keep: a subagent saves your main thread’s context, not your token bill. The work still runs and still costs tokens. It just runs somewhere that does not clog the conversation you are trying to steer. Once you internalise that, most of the confusion around subagents clears up, and the token-burn horror stories start to make sense.
The three kinds of help I hand off
People say “subagent” and mean three different things. I treat them as three separate tools, because they fail in different ways and I reach for them at different moments.
A fork is an ad hoc subagent I spin up inside a session for one job: read these forty files and tell me where the auth check lives, search the whole repo for every place we format a date, draft three versions of this section. It runs on the same model, it reports back a short answer, and then it is gone. Its entire value is that the forty files it read never land in my main context. If several forks need to edit files at the same time, I run each in its own git worktree so they do not collide.
A Codex handoff sends work to a different model altogether. I run both Claude and Codex every day, and the split I have settled on is simple: Claude for interface and UI work and for carrying a feature end to end, Codex for complex backend and for a steadier second-pass review. I also route bulk, context-heavy research to Codex as a parallel subagent specifically to conserve my Claude budget. The full breakdown is in my Codex vs Claude Code comparison.
A specialist is a saved agent definition that lives in your agents folder: a name, a description that tells Claude when to call it, a restricted tool list, and its own system prompt. This is the one you build when a fork keeps doing the same job. I have a handful: a code reviewer, an SEO drafter that knows each site’s rules, a research scout, and an orchestrator that fans the others out. The reusable ones ship free at Locul so you can install them and read the raw files.
Forks, Codex, or a specialist: which to reach for
Lane rule on this blog: when I compare things, I name a pick. Here is how the three line up on the decisions that actually come up, and what I reach for in each case.
| What you are deciding | Fork (ad hoc) | Codex handoff | Named specialist | My pick |
|---|---|---|---|---|
| One-off research or file reading | Spin up, read, summarise, done | Overkill to set up | Not worth defining | Fork |
| A job you repeat every day | Re-briefing it gets old | Fine, but you repeat the setup | Define once, trigger by name | Specialist |
| Complex backend or a second opinion | Same model, same blind spots | Different model, catches more | Only if you wrap Codex in one | Codex |
| Protecting your Claude usage limits | Still spends your window | Spends the other provider instead | Depends what runs inside | Codex |
| Many file edits at once | One worktree each, runs clean | Harder to coordinate | Harder to coordinate | Fork in worktrees |
If you only take one rule from the table: start with a fork, graduate to a specialist when you are tired of repeating yourself, and keep Codex for the backend and for the days your Claude limits are the bottleneck. Everything else is premature.
Do subagents save tokens or burn them
This is the complaint I see most, and both halves of it are true at once. People report the same task costing a couple of thousand tokens in the main thread and a hundred and fifty thousand through a subagent, and people report spawning five or six in parallel and draining a whole plan in minutes. Neither is a bug. A subagent carries its own context, so a fork that reads forty files genuinely spends a lot of tokens in its own window. What you buy is that those forty files never touch your main window, which is the context you care about keeping clean.
A subagent moves the mess out of the room you are working in. It does not make the mess smaller.
Parallelism is where this turns into a bill. Running subagents side by side buys you wall-clock speed, and it scales token burn linearly with no natural ceiling. I learned this the expensive way. One morning an SEO audit job I had written fanned out into thirty Opus subagents and drained a five-hour usage window in about fifteen minutes, which starved every other scheduled job behind it. The fix was not a cleverer prompt. I rewrote that audit as a plain six-second script with zero subagents and moved it to a cheaper model. The rule that came out of it: fan-out is a power tool, so bound it, put the workers on Sonnet or Haiku rather than Opus, and never let an unattended job decide on its own to spawn thirty of anything.
If you want hard limits rather than discipline, Claude Code gives you the knobs. Set CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS to cap how many run at once, CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH to stop a subagent from spawning its own subagents without end, and maxBudgetUsd to put a dollar ceiling on a run so it stops itself before it empties your window. The SDK docs on capping depth, concurrency, and spend list the current defaults. More on how I keep scheduled runs from running up a bill is in my Claude Code best practices from running it unattended.
What a subagent can and cannot see
Here is the misconception that quietly ruins results: people assume a subagent inherits the conversation. It does not. Every subagent starts blank and sees only the prompt string you hand it. If the real objective lives in three messages of chat history the subagent never receives, it will confidently answer the wrong question. So the brief is the whole job. Spell out the goal, the constraints, and exactly what you want back, as if you were handing the task to a contractor who has never seen your project.
The return path is the other half. A subagent does its work and reports back once, at the end, and subagents cannot talk to each other. There is no live conversation to interrupt and no shared memory between them. That shapes how you should design them: tell each one to return a tight summary and the handful of facts you need next, not a transcript of everything it did. If one agent’s output has to feed the next, the coordinator, your main session, is what carries that result forward and passes it in. Think of them as workers filing a one-page report, not colleagues in a meeting.
Build a lean set, not a 100-agent zoo
The GitHub packs with a hundred agents look like a shortcut and are usually a trap. When a dozen agents all have overlapping, job-title descriptions like “backend expert” and “API specialist,” the router has to guess which one you meant, and it guesses wrong. Most of those agents never fire, and the few that do would have worked just as well as a plain prompt.
A lean set beats a big one because routing is a matching problem, and matching gets harder as the options blur together. The specialists I keep each have a narrow, non-overlapping trigger: one reviews code, one drafts articles under a specific site’s rules, one scouts the web and comes back with sources, one orchestrates the others for a big multi-step job. That is enough to cover most of what I delegate. When I want to share a capability rather than keep it private, I publish it as a skill. You can install my set in one command with Claude Code’s plugins and marketplaces: /plugin marketplace add mkhalid1/locul-skills, then /plugin to pick the bundles you want. Every one of them is also readable as a plain file, which is the fastest way to learn the format.
When a subagent is worth it, and when to just prompt
A subagent is not free. Defining one, or spinning up a fork, adds real startup overhead, and for a quick one-off that overhead is slower and more expensive than just asking Claude directly. The honest test is three questions. Is the job self-contained, meaning it can be finished from a clear brief without back-and-forth? Will it produce more output than you want in your main window? Will you do it again, or can it safely run next to other work? A subagent earns its place only when the job is repeated, self-contained, and noisier than your main window should be. If you cannot say yes to at least two of those, prompt the main agent and move on.
There is one more rule I will not run a fleet without, and it also came from a subagent going wrong. A subagent once made unrequested production writes and deploys to Murkuz and Hydori, two of my live products, during a task that was only supposed to pull SEO data. Nothing was lost, but it was changing live systems nobody had cleared. The standing rule since: no subagent touches a production database or ships a deploy without explicit confirmation first. If you let agents run unattended, scope their tools so the destructive ones are simply not in the toolbox, and require a human yes for anything that changes production. Autonomy is fine. Silent autonomy on production is not.
How to set one up (copy this)
Two ways in. Interactively, run /agents in Claude Code and let it walk you through creating one, which is the easiest way to get the description and tool list right. Or create the file yourself. A named specialist is just a markdown file in your project’s .claude/agents/ folder, or ~/.claude/agents/ for ones you want in every project, and Claude reads the description and routes to it automatically whenever a task matches. Every specialist I keep has this shape, and you can paste this and edit it:
---
name: research-scout
description: Use for open-ended web research. Finds and reads sources, returns a short brief with links. Trigger when the task needs current facts the repo does not contain.
tools: WebSearch, WebFetch, Read, Grep
---
You are a research scout. Given a question, find the best 3 to 5 sources,
read them, and return: a 5-bullet summary, each bullet with its source URL,
and one line on what is still uncertain. Do not write prose beyond that.
Never invent a statistic; if a number is not in a source, say so.
The two fields that decide whether it ever gets used are the name and the description. Keep the description about when to call it, not what it is, and keep it from overlapping with your other agents. Restrict tools to only what the job needs, both for safety and because a smaller toolbox makes the agent more predictable.
When a job has several steps that each deserve their own agent, I define a thin orchestrator whose only job is to call the others in order and stitch the results together. Same file shape, different intent:
---
name: article-runner
description: Orchestrates a full blog run. Calls research-scout, then a drafter, then a reviewer, and returns the final draft. Trigger when asked to produce a complete article end to end.
tools: Task, Read, Write
---
You coordinate specialists; you do not do their work yourself.
1. Call research-scout with the topic and collect its brief.
2. Pass that brief to the drafter and get a draft back.
3. Send the draft to the reviewer and apply what it returns.
Return only the final draft and one line on what the reviewer changed.
The orchestrator pattern is what keeps a big job out of one bloated context: each step runs clean in its own window and hands back a summary, and the coordinator never sees the forty files the research step read. One more field worth knowing: a specialist can be given persistent memory so it carries notes across conversations, which the official subagents guide covers. Most of mine do not need it, but a long-running reviewer that should remember past decisions is a good candidate.
If you are building programmatically rather than in the CLI, the Agent SDK subagents docs cover the same idea in code. For wiring external tools into a subagent, the MCP documentation is the place to start, and Skills are the lighter-weight alternative when you want reusable instructions without a full agent.
FAQ
Do Claude Code subagents share my conversation history?
No. A subagent starts with a blank context and receives only the prompt you pass it. It cannot see your chat history or what other subagents are doing, so the brief has to contain everything it needs to finish the job.
What is the difference between a subagent and the Agent SDK?
A subagent is a delegated helper you invoke from inside Claude Code, usually defined as a markdown file or spun up ad hoc. The Agent SDK is the programmatic way to build agents in code, including subagents, when you are building your own application rather than working in the CLI. Same concept, different surface.
Where can I find good Claude Code subagents examples?
The official course and docs are the cleanest starting point, and reading real, working agent files teaches the format faster than any guide. I published the ones I run as 14 Claude Code subagents examples and skills, each with the full file, when to use it and a one-paste install, and a separate set of writing and marketing skills at locul.ai/skills. I would learn from a small, focused set before installing any hundred-agent pack.
Do subagents make Claude Code slower?
For a single quick task, yes, because of the startup overhead, which is why one-offs are better handled by prompting directly. For large or parallel work they make the wall-clock time shorter, at the cost of spending more tokens across the workers. Speed up, bill up.
Where this comes from
I build and run eight products solo, and the back office runs on more than thirty scheduled Claude Code jobs, so the split above is what I actually operate on, not a demo. The failures are real and each cost me something specific: a morning when an audit spawned thirty Opus subagents and emptied a five-hour window in about fifteen minutes, and the time a subagent wrote to production during a task that should only have read. The reusable specialists I mention are free to install from Locul. If you reduce all of this to one line: delegate the work that is repeated, self-contained, and noisy, keep the set small enough that Claude can tell the agents apart, and never let one touch production without a yes.
