Anthropic’s Labs team runs on two-week product reviews, a 20-person roster and a high failure rate

Anthropic’s Labs team runs on two-week product reviews, a 20-person roster and a high failure rate

N
News Editor
2026-09-08 02:46:08
Anthropic’s internal Labs group, led by co-founder Ben Mann, has emerged as a small product incubator inside the AI company. According to a Business Insider report cited in the source material, the roughly 20-person team reviews prototypes about every two weeks and decides whether to keep building, change direction, merge the work into other efforts, or shut it down. Mann said only 20% to 30% of ideas succeed on their own, while the rest may still contribute features, technical lessons or product insight elsewhere. The team was created in 2024 to give new product ideas room to develop and has already produced Claude Code, Model Context Protocol, or MCP, and Claude Design. Mann described Labs as a true startup incubator, with projects “graduating” into independent teams once they usually grow beyond four people. That structure is designed to keep Labs small and mobile while mature products move into longer-term engineering and maintenance. The report also ties Labs’ role to Anthropic’s commercial push. Business Insider said the company is preparing for an IPO, raising the stakes for turning model advances into products people keep using. The source also notes a growing tension between Anthropic as a model platform and Anthropic as a product maker, especially as tools such as Claude Design move into categories occupied by software companies like Adobe and Figma.

Anthropic’s Labs unit has around 20 people. That team has already produced Claude Code, Model Context Protocol, or MCP, and Claude Design.

Anthropic’s Labs team runs on two-week product reviews, a 20-person roster and a high failure rate 2

But Ben Mann, an Anthropic co-founder who leads Labs, said most attempts in the group are supposed to fail. According to Business Insider, Labs reviews its product prototypes on roughly a two-week cycle and decides whether to keep moving, change direction, stop the work, or fold parts of it into other projects. Products that do work eventually “graduate” from Labs and move to independent teams inside the company.

Mann put the success rate for ideas at 20% to 30%. Some efforts that do not become standalone products are still absorbed into other prototypes or existing products. That low hit rate, set against the success of Claude Code, captures the operating logic of Labs: run many experiments, accept that most will not last, and keep resources available for the few directions worth backing.

That system matters more as Anthropic pushes deeper into commercialization. Business Insider reported that the company is preparing for an IPO. In that context, Labs faces a practical question: how to turn advances in frontier models into products users want to keep using.

Labs was set up to work like a startup incubator

Mann said in the interview that a few years ago he still had to persuade the company to start shipping products. At the time, one question under discussion was whether Anthropic could rely on philanthropic funding and keep advancing research on powerful AI systems.

He created Labs in 2024 to give new product ideas room for open-ended exploration. Since then, the team has incubated Claude Code, MCP, which connects AI agents to outside data and tools, and Claude Design, which launched this year. Instagram co-founder Mike Krieger was also involved.

Mann described Labs as a real startup incubator. He pointed to clear precedents for that structure, including Bell Labs, Google’s internal incubator Area 120, which he joined in 2018, and Google X. Each offered a version of the same basic model: use dedicated teams to explore higher-uncertainty ideas, then spin mature projects out into more independent operations.

What makes Labs different is its distance from model research. Frontier model capabilities keep changing. A product may fall short of usability today and then cross that threshold after the next model upgrade. If a product team knows in advance which capabilities are improving, it can make earlier calls on what is worth trying and what should move from theory into a prototype.

How Claude Code emerged

Claude Code came out of that setup. Mann said that before development formally began in late 2024, researchers had already told his team that newer models were showing promise in agentic programming. That gave Labs a signal that a coding product could ask the model to do more real work, not just assist at the margins.

Mann handed the effort to Boris Cherny, who had just joined the company. Cherny’s first idea was a code analysis tool, but Mann felt the scope was too small. He said in the interview that new employees in particular need to raise the size of their ambitions, because there was still so much uncertainty around what agents could actually do.

Anthropic’s Labs team runs on two-week product reviews, a 20-person roster and a high failure rate 3

That speaks to a real problem in AI product development. Teams that only design around capabilities they already understand can underestimate the jobs models may soon be able to handle.

The prototype Cherny built later gained traction inside Anthropic. In February 2025, the company released Claude Code in preview. The product uses the underlying model to write, edit and run code, and later model upgrades kept improving what it could do. Claude Code has since become an important driver of growth for Anthropic, and Cherny now leads the product.

A system built to let most ideas stop early

Seen in retrospect, Labs’ advantage was not a single flash of creativity. The process linked several steps closely together: researchers spotted capability changes, product teams turned those signals into prototypes, internal use tested demand, and later model improvements sharpened the experience. The tighter those steps connect, the better the chance a product has of catching technical progress at the right time.

Still, seeing new capabilities early does not mean every idea deserves long-term investment. Labs runs what Mann described as a “persevere or pivot” review about every two weeks. The cycle is for checking direction, not a rule that every product must be finished in two weeks.

If an idea performs poorly, the team may kill it or take the useful parts and merge them into other efforts. People on the project then move to new tasks. Mann said some attempts ultimately fed into products such as Claude’s Chrome extension.

For that reason, a 20% to 30% success rate should not be read as if the other 70% to 80% produced nothing. A prototype that never becomes its own product may still leave behind reusable features, technical know-how, or better judgment about user demand.

Another rule is just as important. When a project team usually grows beyond four people, it graduates from Labs. Claude Code and Claude Design both moved into independent teams that way. The arrangement keeps Labs small enough to keep exploring new directions. It also explains why the group sees frequent movement: once a project matures, some members leave with it, and Labs brings in new people for the next round of experiments.

Feedback runs both ways between research and product

Projects that have been validated need more engineering work and long-term maintenance. Small teams need room to change direction. Moving mature products out at the right time helps keep Labs from being swallowed by day-to-day operations.

That comes with management costs. Staffing keeps changing, new members need to understand research progress quickly, and the team has to rebuild working relationships over and over. Whether the two-week review works depends on how well people judge prototype performance.

Anthropic’s Labs team runs on two-week product reviews, a 20-person roster and a high failure rate 4

From Mann’s description, the relationship between Labs and the research organization works in both directions. Product developers look for openings created by new models, while gaps that show up in real tasks push researchers toward new capability targets.

Mann summed up one goal for Labs as expanding AI’s “action space,” meaning giving stronger models the ability to do more in the real world. In practice, that breaks down into concrete tasks: programming requires models that can operate on code, design work requires better visual outputs, and multilingual use demands systems that can understand more kinds of expression. Whether a product stands up in use exposes what the models still lack.

Commercialization raises new ecosystem tensions

As these products move into more markets, Anthropic is likely to face more complicated business relationships. Forrester analyst Mike Gualtieri told Business Insider that tools launched by Anthropic could compete directly with software already used by enterprise customers. Claude Design is one example, entering design software territory occupied by companies such as Adobe and Figma.

That creates an ongoing tension. Anthropic supplies model capabilities to developers and enterprises, but it is also building its own products on top of those capabilities. The more functionality the platform adds, the more likely it is to overlap with other software products.

For users, more built-in capability may mean less switching between tools. For ecosystem partners, it means reassessing what differentiated value they can still offer. How Anthropic balances those two sides will shape the appeal of its platform.

If the company does go public in the future, Labs would also need to preserve room for higher-uncertainty exploration under more explicit operating goals. The source material does not provide a timeline for any listing, and it does not say how Anthropic might adjust research and development spending after that.

Mann’s long-term ambitions go beyond coding and office tools

At least in Mann’s view, Labs is not meant to stop with workplace or developer products. He said he wants the team to contribute in the future to breakthroughs in biological research, clean energy storage and other real-world applications. He also said that accelerating basic scientific research would help AI deliver practical benefits to more people.

Those remain ambitions rather than verified outcomes. Still, they follow the same operating pattern Labs already uses: watch the edge of model capability, look for worthwhile problems, and test ideas through prototypes.

Claude Code shows that one path can work. The next test for Labs is whether it can keep finding the next task worth funding as model capability shifts and user demand changes with it. For a team of about 20 people, the key may never have been making every project succeed. It may be whether the group can stop weaker bets quickly enough and give stronger products the resources to keep going.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.