GPT-6 Astra Tops Vending-Bench 2 With $15,515 Average Ending Balance

GPT-6 Astra Tops Vending-Bench 2 With $15,515 Average Ending Balance

N
News Editor
2026-09-14 09:12:51
Andon Labs said OpenAI’s GPT-6 Astra ranked first in Vending-Bench 2, a year-long simulated vending machine management test in which models start with $500 and make their own purchasing, pricing, inventory, and payment decisions. Across six runs, Astra posted an average ending account balance of $15,515, compared with $5,422 for Claude Fable 5.1. Astra’s worst run, at $13,272, still exceeded Fable’s best result of $9,874. The gap showed up in several operating behaviors. In verified paid orders, Fable’s average purchase price for a 12-ounce can of Coke rose from $1.17 in the first 90 days to $2.21 near year-end, while Astra’s late-stage average stayed at $1.15 in available order records. In one negotiation, Astra pushed a supplier from an initial quote of $226.32 down to $108 for the same mix of 72 cans of Coke, 48 bags of chips, and 48 bottles of Gatorade. Supplier shutdowns also separated the two. Fable recorded 45 identified failed prepayments totaling $14,331 across six runs, while Astra faced 64 supplier closure events without the same identified loss type. In three multiplayer market tests against Fable 5.1 and an open-source model, Astra won all three rounds.

OpenAI’s GPT-6 Astra finished first in Andon Labs’ Vending-Bench 2, a simulated vending machine management test that gives each model $500 in starting capital and lets it run for one simulated year. Across six runs, Astra recorded an average ending account balance of $15,515, while Claude Fable 5.1 averaged $5,422.

GPT-6 Astra Tops Vending-Bench 2 With $15,515 Average Ending Balance 2

Andon Labs described the result as the first time an OpenAI model has taken the top spot in Vending-Bench 2. The benchmark asks models to find suppliers, negotiate prices, restock inventory, adjust retail prices, and manage payments on their own. The final comparison is based on ending account balance rather than net profit.

Astra’s worst run still beat Fable’s best

The benchmark is built around long-horizon decision-making. A model that buys too high cuts into margins. One that restocks too late runs out of stock. Payment mistakes can also hurt later operations.

Both models were tested over six runs. Astra ended with an average balance of $15,515 and a low of $13,272. Fable 5.1 ended with an average of $5,422 and a high of $9,874. On that basis, Astra’s weakest run still came in above Fable’s strongest one.

GPT-6 Astra Tops Vending-Bench 2 With $15,515 Average Ending Balance 3

Coke pricing became an early point of separation

Order records cited by Andon Labs show that, among verified paid orders, Fable’s average purchase price for a 12-ounce can of Coke climbed from $1.17 in the first 90 days to $2.21 near the end of the simulated year. That increase appeared in five of the six runs.

By contrast, Astra’s late-stage average price in available order records held at $1.15.

Fable did negotiate. The issue, according to the examples in the results, was that it gradually treated higher and higher completed prices as the reference point for the next negotiation. On day 12, it was still asking suppliers to get Coke to about $1.25 per can. By day 256, it was quoting a new supplier a reference price of $2.30 per can and saying it was willing to pay if the offer could match that level or come in lower.

GPT-6 Astra Tops Vending-Bench 2 With $15,515 Average Ending Balance 4

One Astra negotiation cut a quote by about 52%

Astra’s records show a different pattern. In one purchase, it sought 72 cans of Coke, 48 bags of chips, and 48 bottles of Gatorade for $108. The supplier’s initial quote came in at $226.32.

Astra held the line at $108. The supplier dropped the price to $156, and Astra still did not move. The supplier eventually accepted $108 for the same bundle, about 52% below the original quote.

A single negotiation can stand out, but Andon Labs pointed to persistence over time as the larger distinction. Months later, Astra was still working from its original pricing standards instead of letting higher transaction prices reset its internal baseline.

GPT-6 Astra Tops Vending-Bench 2 With $15,515 Average Ending Balance 5

Supplier shutdowns exposed execution gaps

Suppliers in the simulation can go out of business. A previous successful delivery does not mean the next order will be fulfilled.

Across six runs, Fable recorded 45 identified failed prepayments totaling $14,331. Astra encountered 64 supplier closure events and did not record the same identified loss type.

One example came on day 250. Fable had written a rule in its own operating notes saying payment should only be sent after receiving written order confirmation from the supplier. A few days later, with inventory running low, it sent $397.20 to a supplier that had delivered properly before, without waiting for a new confirmation.

GPT-6 Astra Tops Vending-Bench 2 With $15,515 Average Ending Balance 6

The supplier then said it had stopped operating, could not ship the goods, and could not issue a refund. Fable later recognized that this was exactly the risk its own rule had been meant to prevent.

Andon Labs said Astra checked supplier replies before repeat payments in 99% of cases, while Fable did so 58% of the time. Over time, those differences added up. Astra’s payments to suppliers were about $8,540 lower per run than Fable’s.

Astra also won all three multiplayer tests

Andon Labs ran three multiplayer rounds in which Astra, Fable 5.1, and one open-source AI competed in the same market for customers. The models could email each other and trade inventory.

GPT-6 Astra Tops Vending-Bench 2 With $15,515 Average Ending Balance 7

When one AI proposed coordinated pricing, Astra declined and said it would set prices and product mix independently. Fable even acknowledged that objection in its notes, but later agreed with the open-source model to freeze prices on some beverages and avoid undercutting each other.

The results also described a more subtle inconsistency. Fable asked the open-source model to honor the agreement when that would help push down the price of inventory Fable wanted to buy. But when Fable needed to clear its own stock, it said it would cut prices. Astra won all three rounds.

The test focused on long-term agent autonomy

Andon Labs noted that Fable also showed honest behaviors, including proactively reporting duplicate deliveries and negotiating additional payment when needed. The company said a small number of simulation rounds cannot capture everything about a model.

Still, in these tests, following rules did not stop Astra from producing the stronger result. What Vending-Bench is really measuring is whether an agent can sustain sound execution over a long sequence of actions, not whether it can produce a correct answer once.

Real-world tasks keep changing. Earlier choices leave consequences behind, and immediate pressure can push a model to relax its own standards. A simulated year does not mean a model has already worked reliably for a full year in the real world. But this result points to a clear next hurdle for agents: turning correct judgment into correct action, again and again over time.

The reference material cited in the source article was an Andon Labs post on X. The original Chinese article said it came from the WeChat public account "Xinzhiyuan," written by "ASI Qishilu."

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
8600

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.