News

Claude can now run up to 1,000 AI agents on one job

Anthropic added dynamic workflows to Claude Managed Agents: a lead agent plans the work and runs up to 1,000 agents per run, 64 at a time. In Anthropic's test it found 66 of 70 planted bugs.

Pencil sketch of a robot conductor planning work while many small robots carry code to a funnel

For most of the past two years, an "AI agent" has meant one model working through a task step by step. It reads a file, runs a command, looks at the result, and decides what to do next. That works well for small jobs, but it gets slow and patchy when the job is wide: reviewing hundreds of documents, auditing a large codebase, or checking many sources against each other. A single agent runs out of room, forgets earlier findings, or simply misses things.

On October 9, 2026, Anthropic shipped its answer to that problem. Claude Managed Agents, the company's hosted service for running Claude as an agent in the cloud, now supports what Anthropic calls dynamic workflows. The feature is a public beta. In short, a lead agent no longer has to do all the work itself. It can write a small program, a workflow, that splits a big job into phases, hands the pieces to many other agents, and then combines what they return. One run can start up to 1,000 agents.

That number made the headlines, but the details matter more than the headline. Below is what the feature actually does, what the real limits are, what Anthropic's own test showed, and why this is a meaningful step for anyone who builds with AI agents.

Pencil sketch of a robot conductor with a clipboard at a drafting table while a large crowd of small robots carries sheets of code toward a funnel where the results are merged

The three things you need to know

  • What is new. Claude Managed Agents can now run dynamic workflows. A lead agent writes a workflow, a program that runs many agents in phases and merges their results. The server runs it in the background while the lead agent keeps talking to the user or ends its turn.
  • The scale. A single workflow run can start up to 1,000 agents over its whole life, but only 64 threads work at the same time, and Anthropic says that number can change. A run lasts up to 24 hours by default.
  • The test. Anthropic hid 70 bugs in a codebase of about 116,000 lines. A single agent found 14, 15 and 27 bugs in three runs. The dynamic workflow found 66 in each of its three runs.

What dynamic workflows actually are

Anthropic's documentation describes three ways an agent can hand work to other agents in Managed Agents. The first is subagents: the main agent delegates a task to another agent, waits for its report, and can send follow up messages to that same subagent later. The second is an advisor: the main agent consults a stronger model for guidance, for example to plan an approach or review its work, but keeps doing the work itself. The third, and the new one, is dynamic workflows.

With a workflow, Claude does not coordinate every step by hand. Instead, it writes a program that orchestrates the other agents without Claude's direct involvement. Context and results are passed from one agent to the next by that program. The main session stays free to talk with the user and to check in on one or several workflow runs to report progress. The documentation is clear about the trade off: the lead agent cannot send follow up messages to the threads inside a run, and the server archives each of those threads by the end of the run.

The agents inside a run can be of two kinds. Predefined agents are agents a developer has already created, each with its own model, system prompt, tools, servers and skills. Inline agents are not saved anywhere: the workflow defines them on the fly and writes a system prompt for each one. Inline agents use the model and tools of the agent that runs the session. If a developer wants some agents in a run to use a different, perhaps cheaper model, they create those agents up front and list them, up to 20 per list.

Turning the feature on is a configuration change. Developers set the agent's multiagent type to a new version introduced this month, and workflows are enabled by default with that type. The agent itself decides when a run makes sense. Anthropic recommends telling the agent in its system prompt when to use a run, and its example is a contract reviewer: start a workflow when asked to review more than a few contracts, but review one or two contracts directly.

The numbers behind "1,000 agents"

Much of the early coverage said Claude can now run 1,000 agents in parallel. That is not quite what the limits say, and the difference matters if you plan a real job.

According to Anthropic's limits for workflow runs, as summarized by several independent outlets, 1,000 is the total number of agents a workflow can start over the full life of one run. Once a run reaches that cap, the server starts no more agents and the run ends with an error. The number of agents that work at the same moment is much lower: 64 threads per run, a figure Anthropic explicitly says is not guaranteed and may change. On top of that, a run lives for up to 24 hours by default, or a shorter time the agent sets, and a session can keep up to 10 unfinished runs open at once by default.

So the right mental model is not a thousand workers in one room. It is a team of 64 working shifts, with a budget of 1,000 shifts in total and a day to finish. That is still a big jump from one agent working alone, and it is the kind of capacity that changes which jobs are worth handing to an AI system at all.

What Anthropic's own test showed

To show why splitting the work helps, Anthropic ran an internal experiment. The team planted 70 bugs in a codebase of roughly 116,000 lines and then asked Claude to find them, three times with a single agent and three times with a dynamic workflow.

The single agent found 14 bugs in the first run, 15 in the second, and 27 in the third. That is between 20 and about 39 percent of the planted bugs, and the spread between runs is large. The workflow found 66 bugs in every one of its three runs, about 94 percent, and with no variation at all between runs.

Two caveats belong next to that result. First, it is Anthropic's own test on a task Anthropic designed, not an independent benchmark. At least one outlet that covered the launch noted that it could not find the experiment in Anthropic's published documentation and treated the figures as a vendor result. Second, bug hunting in a large codebase is close to the ideal case for this approach: the job splits naturally into many independent pieces, and a second pass by other agents can catch what the first pass missed. Whether the gains hold for other kinds of work remains to be seen.

Pencil sketch of a long code printout on a table with many small robots searching it with magnifying glasses and many bugs circled in red, while a single robot on the side has circled only a few

The cost question

The obvious catch is tokens. Every agent in a run reads, thinks and writes, and every one of those steps costs money. Anthropic itself says that workflows can use a lot of tokens and recommends starting with a scoped task. Its documentation advises developers to set a session budget when they create a session, which caps the spend of the whole session including all its runs. When a session hits that budget, its runs pause, and they continue if the budget is raised or removed.

This is also where the debate about agent swarms comes in. Some engineers argue that throwing many agents at a problem mostly burns tokens without improving results. Anthropic's bug hunting test is its counterargument: for wide tasks, more agents found far more problems and found them consistently. The honest answer is probably that both sides are right in different situations. For a narrow task, one careful agent is cheaper and good enough. For a wide audit where missing a problem is expensive, paying for many agents can be the cheaper option overall.

How this fits Anthropic's agent push

Dynamic workflows are not a new idea inside Anthropic's products. Earlier versions ran locally inside Claude Code, the company's coding tool, where a lead agent could fan out work on a developer's machine. What changed on October 9 is that the feature now runs on Anthropic's servers as part of Managed Agents. Each run reports its progress as events on the session's event stream, so a developer's application can follow the phases and read each agent's thread while the work happens.

That matters because it moves large multi agent jobs from a developer's laptop into infrastructure that can run for hours in the background. A company can now, in principle, ask a hosted agent to audit a codebase overnight, cross check a stack of contracts, or run a deep research task across many sources, and come back to merged results in the morning.

Why it matters

The first wave of AI agents was limited less by intelligence than by attention. One model, one context window, one line of work. Dynamic workflows attack that limit directly by turning one agent into a coordinator of many, with the plan written as code rather than improvised step by step. If Anthropic's test result even partly carries over to real work, tasks that used to be too big or too unreliable for a single agent, such as full code audits or large document reviews, become practical.

It also raises the stakes for oversight. A run with hundreds of agents doing things in parallel is harder to watch than one agent in a chat window. Budgets, permission policies for the tools those agents call, and clear rules about when a run should start become important safety and cost controls, not just nice to have settings.

Dany's take

I like that Anthropic published real limits instead of just the big number. "Up to 1,000 agents" sounds like marketing, "64 at a time, 1,000 in total, 24 hours max" sounds like something you can actually plan with. The bug test is impressive, 66 out of 70 every single time versus a wildly swinging single agent, but it is still Anthropic grading its own homework. I would love to see someone run the same idea on a real open source project and publish the bill. My guess: for big audits this will be worth every token, for everyday tasks one good agent stays the smarter choice.

Sources

Source: Anthropic Claude Docs

Newsletter

The AI news that matters, in your inbox.