Goldman Sachs said its latest Silicon Valley research found that AI is entering a stage where systems are expected not just to answer questions, but to carry out tasks. The bank said commercialization is shifting from seat-based subscriptions toward pricing based on consumption, transaction volume, and outcomes. At the same time, agents are moving from support tools to workflow executors, and industry value is moving from the model itself toward proprietary data, business context, and domain expertise.
That changes where competition sits. Model capability still matters, but Goldman said the more decisive edge may be whether a system can enter a production environment, understand business context, and complete tasks consistently.
Goldman completed its third straight year of Silicon Valley field checks
The assessment came after Goldman Sachs conducted on-the-ground research across the Silicon Valley AI chain. On Aug. 18 and 19, the bank visited AI startups, leading venture capital firms, and researchers at Stanford University, the University of California, Berkeley, and the University of California, San Francisco for the third consecutive year.
Goldman said value allocation across frontier models, open-source models, world models, enterprise software, and proprietary data is likely to change as agents move deeper into real-world use.
Enterprise adoption hinges less on capability and more on control
If earlier AI tools were designed to help people finish tasks, agents are now being built to finish the tasks themselves. Goldman said that in large-scale enterprise deployment, the central constraint may no longer be model capability alone. Responsibility, traceability, and control over execution are becoming equally important.
Citing Stanford researchers, the report said most companies are still operating under human supervision models. In fields such as legal services, risk control, insurance, and auditing, questions over who is accountable when a model makes a mistake, how the process can be traced, and whether errors can be corrected quickly carry weight comparable to the model's raw performance.
Goldman said workflows most likely to be automated first usually share three traits:
- clear decision boundaries
- verifiable outcomes
- errors that can be rolled back
Invoice processing was cited as a typical case. AI extracts fields and performs checks, low-confidence cases are sent to human review, and booking is then completed through a reversible ERP process.
On that basis, Goldman said information service providers with trusted content, validated domain models, and established regulatory relationships may be better positioned to enter enterprise production environments first.
Frontier and open-source models may split by workload type
On the debate over open versus closed models, Goldman said the signals from this round of meetings did not point to a simple either-or outcome. Instead, different models may serve different layers of enterprise workflows.
Supporters of frontier models argued that enterprise benchmark testing often understates what those models can do. In real production settings, losses caused by lower accuracy can outweigh savings in inference cost. Goldman said that while many AI-native companies describe themselves as using multi-model strategies, they still rely heavily on frontier models in core production environments.
Another view is that most enterprise workflows do not require frontier-level intelligence. As open-source models continue to improve, customers are becoming more willing to accept limited performance trade-offs in return for lower inference cost. One venture capital firm estimated that about 90% of inference tokens will flow to open-source models over the next 12 to 18 months.
Goldman said this points to a clearer market split ahead: frontier models handling high-value, high-reliability complex tasks, and open-source models taking on larger volumes of standardized work and most token consumption.
World models are emerging as a second compute growth curve
Over the past 18 months, researchers have increasingly shifted attention from large language models, or LLMs, to world models.
Unlike LLMs, which are trained mainly on internet data, world models need to understand environments, causal relationships, physical rules, and dynamic interactions in the real world. Their data is more likely to come from physical systems, specific industries, and actual operating scenarios. Goldman said that raises the importance of proprietary data even further.
The bank added that the problem space across physics, industry, science, and robotics is much larger than pure text generation, and those workflows often demand greater compute intensity. As AI moves from the digital world deeper into the physical world, training, simulation, and inference demand could open a new growth curve for computing power.
Goldman estimated that compute demand could rise about 24 times over the next five years, with supply tightness lasting longer. It said cloud and compute infrastructure companies including Microsoft, Oracle, and CoreWeave stand to benefit directly.

