Laminar, an observability platform for AI agents, has introduced flow-1, a model built to inspect agent execution traces and identify where an agent failed and why. The model runs inside Laminar’s Signals agent and reads model calls, tool calls, and returned outputs across a full trace before drilling into individual steps when needed.
According to Laminar, flow-1 was trained first with supervised fine-tuning on synthetic investigation data, then with reinforcement learning focused on tool use and analysis of complex traces. In the company’s in-house benchmark of 523 difficult traces, flow-1 posted an error-detection F1 score of 0.835, compared with 0.816 for GPT-6-sol. Laminar said GPT-6-sol delivered higher recall, meaning it found more real errors, while flow-1 achieved higher precision and generated fewer false positives.
Laminar also compared inference costs on traces under 100,000 LLM tokens. It said flow-1 averaged about $0.0011 per trace, versus roughly $0.026 for GPT-6-sol. On that basis, the company estimated that $1 would cover about 888 trace analyses with flow-1 and about 38 with GPT-6-sol, a gap of roughly 23x. Laminar added that flow-1 is about 25% cheaper than GPT-6-luna. The results come from Laminar’s self-built benchmark and have not yet been independently reproduced. The company said about 48% of the training data came from software engineering tasks, and some community members have questioned whether the model can maintain the same performance on novel failures in real production settings.
Laminar, an observability platform for AI agents, has released flow-1, a model designed to inspect agent execution traces. The model runs inside Laminar’s Signals agent and reads model calls, tool calls, and returned results to identify where an agent failed and why.
Full-trace search and step-level review
Laminar said flow-1 can search across an entire trace, or the agent’s full execution record, and then inspect specific steps as needed.
How the model was trained
For training, Laminar said it first used supervised fine-tuning on synthetic investigation data. It then applied reinforcement learning to train tool use and analysis of complex traces.
Benchmark results against GPT-6-sol
In Laminar’s in-house test set of 523 difficult traces, flow-1 recorded an error-detection F1 score of 0.835, compared with 0.816 for GPT-6-sol. Laminar said GPT-6-sol had higher recall and was able to catch more real errors, while flow-1 showed higher precision and produced fewer false positives.
Cost comparison
For traces below 100,000 LLM tokens, Laminar said flow-1 costs about $0.0011 per analysis on average, versus about $0.026 for GPT-6-sol. By Laminar’s calculation, $1 can analyze about 888 traces with flow-1 and about 38 with GPT-6-sol, a difference of roughly 23x.
The company also said flow-1 costs about 25% less than GPT-6-luna.
Limits of the current results
All of the figures were produced on Laminar’s self-built benchmark, and no third-party replication is available at this stage. Laminar said the training data for flow-1 mainly came from synthetic workflows, with about 48% tied to software engineering tasks.
Some community members have already questioned whether the model can deliver the same level of performance when it faces new types of failures in real production environments that were not predesigned in advance.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.