OCR

OpenAI
2026-08-18 10:50:11

Roboflow says GPT-5.6 Sol is OpenAI’s strongest vision model so far

Roboflow, a third-party computer vision evaluation firm, said GPT-5.6 Sol posted the strongest visual performance yet from OpenAI in its VLM benchmark. The biggest jump came in object detection, where Sol scored 46.2 versus 13.8 for GPT-5.5, while Terra and Luna followed at 44.7 and 43.3. Counting accuracy also improved, with Sol rising from 64.9% to 73%. Roboflow highlighted document layout parsing as another area of progress, saying Sol could cleanly identify titles, body text, tables, illustrations, and signatures in documents. The gains were not universal. In full OCR transcription, Sol scored 90.7%, slightly below GPT-5.5’s 91.2%. In targeted extraction tasks such as pulling a date from an invoice, Sol fell to 82.5% from 87.6% for GPT-5.5. Roboflow also reported that detection boxes could become unstable on images around 2000×2000 pixels or larger, a limitation OpenAI acknowledged. According to Roboflow’s testing, Sol cost about 2.5 cents per image and took around 10 seconds, compared with about 1 cent and 6 seconds for Terra, and under 0.5 cents and 5 seconds for Luna. The firm said Gemini 3.5 Flash remained ahead on detection and counting in this benchmark while also costing less per image.

30
Roboflow says GPT-5.6 Sol is OpenAI’s strongest vision model so far
Whale Movemen
2026-07-29 02:34:00

Crypto and AI Roundup on July 29: Regulation, listings, hacks and market stress

A broad set of crypto and AI developments landed over the past day, spanning regulation, market structure, fundraising, protocol upgrades and security incidents. Kenya cut the paid-in capital requirement for stablecoin issuers by 40% to about $2.32 million while keeping strict reserve and redemption rules in place. Russia’s central bank published its first draft framework for organized trading in digital assets, and Myanmar passed a cybercrime law that allows life sentences for crypto-related fraud. In the U.S., Senate Republicans are still trying to move the Clarity Act before the August recess, though ethics provisions and bank lobbying remain major obstacles. On the corporate side, PayPal posted better-than-expected second-quarter results and did not address a previously reported buyout approach. Luno and Visa both outlined layoffs tied to restructuring and capital allocation, while Morgan Stanley Investment Management rolled out exchange-traded products tied to Ethereum and Solana. Zcash activated its Ironwood NU6.3 upgrade, Layer 2 TVL on Ethereum fell to its lowest level since 2023, and Bitcoin briefly dropped below $63,000 as AI and semiconductor weakness spilled into crypto. Security reports also stayed in focus, with Blockaid saying crypto losses from hacks topped $1 billion in the first half of 2026 and several fresh token incidents reported across the market.

990
Crypto and AI Roundup on July 29: Regulation, listings, hacks and market stress
Kimi
2026-07-28 16:02:00

Kimi open-sources PerceptionBench as no model tops 60% accuracy in visual perception test

Kimi has released PerceptionBench, an open-source benchmark designed to measure visual perception in multimodal large language models by breaking the task into 10 atomic capabilities. The benchmark covers areas including visual relations, counting, attributes, depth and 3D, localization, comparison, fine-grained recognition, context integration, OCR, and hallucination detection. According to PANews, the dataset was built from model failure cases collected across 42 existing evaluation sets and contains 3,000 manually verified questions. Each question is designed to test only one visual skill and does not require reasoning or external knowledge. Results across 16 leading multimodal models showed that none achieved an overall accuracy above 60%. GPT-5.6-Sol ranked first with 59.7%, followed by Kimi K3 at 58.5%, Claude-Fable-5 at 57.2%, Gemini-3.1-Pro at 56.2%, and GPT-5.5 at 55.8%. The report said hallucination remained the weakest area across models, indicating that core visual perception performance still has significant room for improvement.

710
Kimi open-sources PerceptionBench as no model tops 60% accuracy in visual perception test
SparkKitty
2026-07-27 11:37:13

SparkKitty malware slipped into Apple and Google app stores to steal crypto wallet seed phrases

Security firm Check Point said on July 27 that a cross-platform malware strain called SparkKitty has been distributed through Apple’s App Store, Google Play, and third-party Android app stores. The malware uses optical character recognition, or OCR, to scan images stored in a user’s photo library for crypto wallet seed phrases, passwords, and QR codes. Once installed, the app requests access to photos, keeps scanning both existing and newly added images, and uploads extracted text along with device information to an attacker-controlled server. Check Point described SparkKitty as an upgraded version of the earlier infostealer SparkCat. On iOS, malicious code was found inside a crypto app called “币 coin,” which reportedly used code obfuscation to pass Apple’s review process. On Android, researchers said a chat and crypto trading app called “SOEX” carried the malware and was removed from Google Play only after surpassing 10,000 downloads. The report also said the malware spread through third-party app stores, sideloaded APKs, modified TikTok versions, and entertainment apps. Check Point warned users not to store seed phrases in phone photo galleries and recommended offline backups instead.

660
SparkKitty malware slipped into Apple and Google app stores to steal crypto wallet seed phrases
Chainlink
2026-07-22 16:50:13

What Is Chainlink (LINK)? How the Oracle Network Powers Web3 Data

Chainlink is a decentralized oracle network linking smart contracts to off-chain data. It solves the oracle problem through multiple independent nodes, supporting DeFi, NFTs, and enterprise applications with services like Data Feeds, VRF, CCIP, and Proof-of-Reserve.

670
What Is Chainlink (LINK)? How the Oracle Network Powers Web3 Data
Ghost Font
2026-07-13 07:42:10

Ghost Font’s ‘human-only’ message was cracked in a day after a single prompt

Ghost Font, a browser-based experiment by developer Eric Lu, briefly looked like a fresh way to hide text from AI systems while keeping it readable to people. The tool turns typed text into a noisy video: pixels forming letters move upward while the background noise moves downward. Humans can spot the message through motion, but frame-by-frame analysis leaves only static snow. Initial tests appeared to support the idea. According to the article, Claude Fable and GPT-5.6 Sol Ultra both failed to recover the real hidden message and instead reported decoy text embedded in the video. ChatGPT 5.5 Pro reportedly spent 19 minutes and still hallucinated a message that was not there. Gemini 3.1 Pro also returned a planted decoy. That edge did not last. Prompt engineer Riley Goodside gave GPT-5.6 Sol a single instruction explaining the motion directions of the letter pixels and the background. After 1 minute and 56 seconds, the model produced the correct message: “RILEY WAS HERE.” The episode turned Ghost Font from a showcase of human perceptual advantage into a test of how close current multimodal AI is to handling motion once the right cue is provided.

690
Ghost Font’s ‘human-only’ message was cracked in a day after a single prompt