a16z Says OpenAI Is Taking the AI App Layer, but Startups Still Have Room Beyond the Yellow Brick Road

a16z Says OpenAI Is Taking the AI App Layer, but Startups Still Have Room Beyond the Yellow Brick Road

N
News Editor 01
2026-07-23 18:45:16
a16z partner Joe Schmidt argues that startups should avoid head-on competition with model labs in horizontal AI tools and focus on vertical workflows, governance, compliance, and system integration.
a16zOpenAIAI applicationsenterprise softwarevertical workflows

a16z partner Joe Schmidt argues that the AI application market is not one battlefield. In his framework, there is the “Yellow Brick Road,” where model labs such as OpenAI and Anthropic are moving directly into product categories like code generation, writing, image creation, general-purpose agents, and horizontal workplace assistants. Then there is everything outside that path: the messy, vertical, compliance-heavy parts of real business operations. He says the stronger startup opportunity is in the second group.

His core point is simple. Enterprises do not keep paying for a smarter chat window on its own. They pay for systems that can own outcomes, handle messy customer data, multi-party approvals, edge cases, audits, and governance, and keep running as models change underneath them. Foundation models will keep improving and will become more replaceable over time; what is harder to replace is the operating layer built around a specific workflow, industry, and customer environment.

Why the “Yellow Brick Road” is a dangerous route

Schmidt describes the most obvious startup path as the riskiest one. Many founders build a product by connecting a strong model to common enterprise tools such as Google Drive, Slack, Salesforce, Notion, and GitHub, then adding an orchestration layer for agents. It looks powerful. The trouble is that this is also the exact direction model labs are already pushing with products such as Codex and Claude Code.

The challenge is not only product overlap. Labs own the model, the margin structure, the distribution engine, and the architectural choices that shape what problems the product is built to solve. Horizontal, low-step tasks fit naturally into a “model plus tool calling” setup. A startup using the same pattern, but without deeper configuration, domain-specific agent structures, or strong distribution, can end up as a thin wrapper in a category the labs are better positioned to control.

Schmidt also points to what OpenAI and Anthropic are doing in practice. He says their large frontline deployment and joint-project efforts amount to an admission that a single general AI coworker will not solve every enterprise problem. If the next model release were enough to handle those operational realities, there would be little reason to put massive resources into custom enterprise configuration.

Vertical workflows create a different kind of moat

The opportunity, in his view, lies in building systems for complex workflows inside specific industries or functions. The value in those products does not come only from the model. It comes from the scaffolding around the model: context gathering across systems, routing tasks to the right people, permission controls, audit logs, exception handling, legacy integrations, and the engineering work needed to make outputs reliable inside production environments.

That matters because real business work rarely resolves in a single prompt. A task may span multiple tools, require deterministic outputs, involve little tolerance for error, and connect directly to a financial result. Labs understand these markets are valuable, but Schmidt argues they are structurally less suited to sit inside each industry long enough to absorb undocumented norms, company-specific practices, and the tacit knowledge held by operators.

He says two compounding loops can form here. One is across customers: as a company sees more variations of the same problem, its pattern recognition improves. The other sits within each customer: reasons behind decisions, unstated exceptions, and internal operating logic only show up once users interact with the system in production. Much of that knowledge is not on the public internet and is not captured in general training data.

Guardrails, model routing, and migration work matter

The article highlights three pieces that are often underestimated. The first is guardrails. Different industries and job functions require very different rules around what an agent may access, say, or do. Legal work has professional rules and litigation standards. Healthcare has HIPAA. Financial services involve SEC, FINRA, and state-level insurance regulation. A horizontal product can gesture at these constraints, but going deep on them across many industries is far harder.

The second is model routing. Schmidt argues that strong application companies will not send every request to the same frontier model. They will route the hardest tasks to frontier systems, many everyday tasks to mid-tier models, and proven narrow tasks to smaller custom or fine-tuned models. That is not just a technical preference; it is a margin decision. He puts it bluntly in the piece: sending every query to Opus 4.7 is the fastest path to negative gross margins.

The third is migration. Model labs will keep shipping new systems, but they usually do not absorb the enterprise work of rerunning evaluations, recalibrating prompts around customer edge cases, and rolling upgrades into production without breakage. A vertical AI company that takes on that burden is not selling API access alone. It is selling continuity.

11x and FurtherAI as examples from the field

To ground the thesis, the piece includes operating views from 11x CEO Prabhav Jain and FurtherAI CEO Aman Gour. Jain says 11x starts from an outcome customers care about: generating more pipeline. From there, the company breaks the problem into concrete tasks such as custom-signal prospecting, lead enrichment, deep account research, pulling CRM context, writing messages for different channels, lead qualification agents, and email deliverability systems. Some of those are agentic tasks. Some are not. The point is that serious workflow products require deep engineering, not just one-shot prompting.

He argues that in any real workflow, roughly half the work is non-agentic, and that half does not hand model labs an automatic advantage. The other difficulty is messy data. A company may already be a customer and should not be contacted again, but parent-child company structures, stale CRM records, and bad matching fields can easily create mistakes. Jain says 11x addresses that by building systems around the actual shape of the problem rather than pointing a generic copilot at the CRM.

Jain also says that although market expectations around what “AI-written” and “human-like” outreach look like change every few months, 11x has seen its positive reply rate increase 4x over the past few months and has created hundreds of millions of dollars in pipeline for customers. In another case, working with a Fortune 1000 organization on consent-based voice outreach to its SMB customer base, he says 11x now generates more daily sales opportunities than that company’s full sales team creates in a month within that segment.

FurtherAI offers a parallel argument from insurance operations. Aman Gour says much of the intelligence in insurance lives inside the workflow itself rather than inside the model alone. Two insurers may appear to follow the same path from submission to review to quote to underwriting, but the real differences sit inside the details: which risks must be escalated, which loss signals matter, how conflicting underwriting preferences are prioritized, when human signoff is mandatory, what outside data must be pulled, and how final decisions are recorded.

Those rules are often spread across operating procedures, manager reviews, underwriting philosophy, carrier-specific risk appetite, and years of institutional experience. Many are not written in a clean form a model can simply read. Gour says that is why FurtherAI builds “agentic workflows,” where workflows provide repeatability, auditability, and cost control; agents handle variability and recover when the ideal path breaks; and humans stay in the loop where judgment and accountability are required. His conclusion is that the first workflow delivered is not the moat. The moat comes from repeated production usage and the operating memory it creates over time.

Customers buy outcomes, not benchmark scores

Schmidt’s closing distinction is sharp. Model labs are measured by benchmarks. Companies building outside the Yellow Brick Road are measured by the customer’s profit-and-loss statement. Buyers do not pay because a model scores better on SWE-Bench or MMLU. They pay because an agent closed business, redlined a contract correctly, or underwrote the right policy.

If a product is selling general capability, it is easier for a seat-based offer from Claude or Codex to replace it. If the product has become the system through which work actually gets done, with data capture, workflow orchestration, governance, and compliance built in, replacement gets much harder. Schmidt’s argument is not that the labs will stop winning. It is that the next generation of enterprise software is more likely to be built outside the main road they dominate.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.