Roboflow, a third-party vision benchmarking firm, said GPT-5.6 Sol delivered the strongest visual performance yet from OpenAI after the company ran the full GPT-5.6 lineup through its VLM benchmark. Roboflow project lead SkalskiP said, 「GPT-5.6 Sol is the strongest vision model in OpenAI’s history so far」.
Object detection posted the largest jump
Roboflow’s benchmark covered practical image tasks, including object detection, counting, text recognition, and information extraction. Object detection has long been a weak point for GPT vision models because the task requires the model to place boxes around each item in an image.
On that metric, GPT-5.5 scored 13.8. GPT-5.6 Sol reached 46.2, more than tripling the previous result. Terra and Luna followed with scores of 44.7 and 43.3.

Roboflow also noted that Luna, the lowest-priced tier in the new generation, still outperformed the prior flagship model on detection.
Counting and document layout parsing also improved
Counting scores moved higher as well. Sol’s accuracy rose to 73% from 64.9% for GPT-5.5. Terra scored 67.6%, and Luna came in at 66.2%.
One of the clearest gains came in document layout recognition. Roboflow said Sol could cleanly mark out titles, body text, tables, illustrations, and signatures. That matters for workflows involving contracts, invoices, and reports, where systems often need to locate the relevant region before reading the text.

Roboflow said GPT-5.6 also held up better in dense scenes where many similar objects appear together, such as pills or eggs. It described a more constrained test as well: counting bullet holes only within a specified scoring ring on a target sheet. According to the benchmark, Sol handled the task correctly, identifying not just what to count but where the counting rule applied.
OCR results were mixed
The model did not improve across every vision task. Roboflow’s results showed a split between full-text OCR and targeted extraction.

In full OCR transcription, Sol scored 90.7%, slightly below GPT-5.5 at 91.2%. The gap was small. In targeted extraction, where the model only needs to pull a specific field rather than transcribe everything, Sol scored 82.5%. GPT-5.5 scored 87.6%, leaving Sol about 5 points lower.
Roboflow said Sol could transcribe handwritten notes and correctly extract dates. It also read size markings printed on curved tire surfaces and formatted live hockey broadcast scores as requested. One failure case came from the expiration date printed on a blister pack, where the text was small, vertical, low-contrast, and reflective.
OpenAI acknowledged instability on large images
Roboflow said that in some examples, Sol returned detection boxes that drifted to unrelated parts of an image, with almost no overlap with the actual object. Those boxes were often arranged in unnaturally regular patterns, such as a straight line or an evenly spaced group.

After Roboflow shared those samples with OpenAI, the company responded that Sol can become unstable on images around 2000×2000 pixels or larger. It also said the issue becomes more pronounced at lower reasoning settings.
Roboflow outlined two ways to address that limitation. One is to increase the reasoning tier, though that raises token usage, latency, and cost. The other is simpler: resize or crop the image before sending it to the API.

Cost and speed remained part of the trade-off
Roboflow’s testing put Sol at about 2.5 cents per image with a turnaround time of about 10 seconds. Terra came in at about 1 cent per image and about 6 seconds. Luna was priced at under 0.5 cents per image and took about 5 seconds.
- Sol: about $0.025 per image, about 10 seconds
- Terra: about $0.01 per image, about 6 seconds
- Luna: under $0.005 per image, about 5 seconds
For comparison, Gemini 3.5 Flash cost about 0.8 cents per image, cheaper than Sol by roughly two-thirds, and still led this benchmark on detection and counting. Roboflow’s takeaway was that Gemini Flash remained the better value for high-frequency, large-volume detection and counting workloads.
The competition is shifting from image understanding to execution
The main surprise in Roboflow’s results was not simply that GPT-5.6 could interpret images better, but that its visual stack is moving into tasks usually associated with specialized vision systems: locating targets, drawing boxes, counting, handling spatial constraints, and parsing document layouts.

That said, the benchmark also drew clear limits around the model’s current performance. Sol improved sharply over GPT-5.5 in several categories, but targeted OCR extraction, stability on very large images, and cost efficiency remain weaker spots.
The original article was published by the WeChat account Xinzhiyuan and credited to ASI Qishilu.

