GLM-5

Private AI
2026-07-14 01:32:50

IOSG says private AI is gaining ground as open models close the gap in cost and accuracy

IOSG argues that private AI is moving from a niche concern to a practical choice for both enterprises and consumers, as companies grow more wary of sending sensitive data and proprietary knowledge into closed-model systems. In a long-form analysis by Jeff @IOSG, the firm lays out the tradeoff now facing the market: frontier labs still lead in general capability, but open models are improving quickly and, in some specialized domains, already outperform frontier systems on both accuracy and cost. The report traces several privacy approaches, from contractual zero-data-retention and Oblivious HTTP to trusted execution environments, end-to-end encryption, fully homomorphic encryption, and local inference. It argues that only some of these offer verifiable privacy, and those routes largely depend on open models rather than proprietary ones. IOSG also points to a recent case from Bridgewater-backed AIA Labs and Thinking Machines, where a fine-tuned Qwen3-235B model beat frontier models on expert financial tasks. Even so, the report says major gaps remain. Tool use in agent workflows, private post-training, and encrypted search are still hard to deliver at scale. IOSG’s conclusion is that privacy inference is becoming cheaper and more deployable, but the most defensible opportunities lie in the unsolved layers around training loops, tool execution, and search infrastructure.

1240
IOSG says private AI is gaining ground as open models close the gap in cost and accuracy
IOSG
2026-07-14 01:32:50

IOSG says private AI is moving from theory to deployment as open models gain ground in cost and accuracy

IOSG argues that private AI is no longer a niche technical preference but an emerging requirement for enterprises and power users that do not want proprietary data, internal workflows, or high-value judgment calls exposed to model providers. In its report, the firm lays out how the privacy problem starts the moment a prompt leaves a user’s device and reaches a server in plaintext, and why contractual protections such as zero-data-retention terms can only go so far. The piece links that risk to corporate restrictions on ChatGPT, shadow AI leaks, and a series of legal cases in which user chats became discoverable evidence. The report also maps the trade-offs across today’s privacy stack, from contract-based retention promises and OHTTP relays to trusted execution environments, end-to-end encryption, fully homomorphic encryption, and local inference. Its central case study comes from Bridgewater’s AIA Labs and Thinking Machines, which showed that a fine-tuned open model, Qwen3-235B, beat frontier models on both accuracy and cost in financial judgment tasks. IOSG’s conclusion is narrow but clear: for execution-heavy agent workflows, trust-based setups still dominate because tool calls expose plaintext to downstream services; for high-value strategic reasoning and domain-specific alpha, verified private infrastructure around open models is becoming a practical path.

1150
IOSG says private AI is moving from theory to deployment as open models gain ground in cost and accuracy
Private AI
2026-07-14 01:32:50

Why firms are reconsidering private AI as open models narrow the gap

A new report from IOSG argues that the core debate in AI is shifting from model capability alone to a harder question: who gets to see the data, and whether privacy claims can actually be verified. The piece points to a string of examples showing why that matters. Palantir CEO Alex Karp said companies are paying a token premium to frontier labs while letting proprietary knowledge leak out through plaintext requests. Wall Street banks restricted ChatGPT use within months of its launch, Samsung banned generative AI across its network after engineers exposed chip source code, and court orders later forced OpenAI to retain and disclose consumer chat records in litigation. The report maps the current privacy stack, from contractual zero-data-retention and anonymous relays to trusted execution environments, end-to-end encryption, fully homomorphic encryption and local inference. It argues that verifiable privacy is still mostly limited to open models, because frontier labs have little incentive to expose model weights or serving code. At the same time, the economics are changing: enclave-based inference is getting cheaper, and in some cases can match or undercut plaintext API pricing. IOSG also highlights a June 30 case from Bridgewater-backed AIA Labs and Thinking Machines, where a fine-tuned open model beat frontier systems on both accuracy and cost in financial tasks. The report’s broader point is that private AI remains incomplete, especially for agentic workflows and tool use, but it is no longer hypothetical.

1580
Why firms are reconsidering private AI as open models narrow the gap
Zhipu AI
2026-07-13 07:49:00

Zhipu AI says post-IPO focus will stay on AGI research, not short-term monetization

Zhipu AI founder Tang Jie used an internal letter dated July 11 to tell employees that the company will keep directing resources toward artificial general intelligence research over the next two years instead of prioritizing near-term monetization. The message arrived at a sensitive moment. After listing in Hong Kong on Jan. 8 at HK$116.2, Zhipu’s shares at one point climbed to HK$2,980, more than 24 times the IPO price, before falling more than 19% around the lock-up expiry, which fueled debate over valuation and whether the stock had entered bubble territory. Rather than address the sell-off directly, Tang laid out what he called a long-term strategy built around four engines: long-horizon tasks, autonomous agent systems, fully self-training AI, and strict safety governance. He said the company plans to devote resources at the scale of tens of billions of yuan to mechanistic interpretability research and argued that safety work must advance in parallel with more powerful AI systems. The letter also linked Zhipu’s open-source push to that strategy, highlighting GLM-5.2, which the company described as its strongest open-source model so far, with million-token context support and release under the MIT license. The original report said Zhipu’s MaaS platform ARR grew 60-fold over the past year after the company shifted more resources toward coding and reasoning following DeepSeek R1’s release in early 2025.

1110
Zhipu AI says post-IPO focus will stay on AGI research, not short-term monetization
Prime Intelle
2026-07-13 02:31:09

Prime Intellect Raises $130 Million at $1 Billion Valuation, Says ARR Has Surpassed $100 Million

Prime Intellect, a decentralized AI infrastructure network founded in 2024, said it raised a $130 million Series A at a $1 billion valuation on July 8, 2026. The round was led by Radical Ventures and included NVIDIA Ventures, Intel Capital, and Dell Technologies Capital, bringing total funding to more than $150 million. At the same time, the company said its annual recurring revenue has climbed past $100 million in less than a year and that it now serves more than 6,000 enterprise and startup customers. The company’s recent progress spans distributed training, reinforcement learning, inference orchestration, and hosted infrastructure. Prime Intellect highlighted milestones including INTELLECT-1, INTELLECT-2, and INTELLECT-3, along with products such as Prime Intellect Lab, prime-rl, and Sandboxes. It also disclosed deeper ties with NVIDIA across both hardware and software, including the use of Blackwell systems and deployment of NVIDIA Dynamo in production workflows. Foresight’s report also pointed to a shift in the company’s public positioning. Language in official documentation that previously referenced Base Sepolia, a future proprietary chain, and token rewards contracts has been removed. The report argues that while Prime Intellect still uses a distributed network design, its messaging has moved away from a crypto-first framing and toward an AI infrastructure business aimed at enterprise use cases.

450
Prime Intellect Raises $130 Million at $1 Billion Valuation, Says ARR Has Surpassed $100 Million
Cerebras
2026-07-11 06:32:13

Cerebras says AI compute is sold out with $25 billion backlog as Black Forest Labs links generative video to robotics

Cerebras CEO and co-founder Andrew Feldman said demand for AI compute has already been booked out, with the company holding a $25 billion backlog and seeing orders tied to a global buildout of data centers. Speaking on the All-In Podcast aired on July 10, 2026, Feldman argued that this cycle differs from earlier tech booms because customers are not waiting for capacity to appear before committing. He said buyers including OpenAI, Anthropic, SpaceX, Google, Microsoft and AWS are already lining up for more infrastructure. Feldman also framed reasoning as the next major compute sink. He said reasoning workloads consume massive amounts of tokens, making inference speed a central competitive factor, and claimed Cerebras can run 15 times faster in some cases. On open-source AI, he said enterprises increasingly want control, especially in regulated sectors, and described sovereign deployment as a growing priority. Feldman went even further on AGI, saying that by definitions commonly used 20 years ago, the industry has already passed the threshold. In the same conversation, Black Forest Labs CEO and co-founder Robin Rombach said the company is building multimodal models that span image, video, audio and action prediction. He described work with Martin Scorsese on AI-assisted visual ideation and said the longer-term opportunity is not just filmmaking. In his view, the same multimodal model used to create films could also serve as the brain of a robot.

1110
Cerebras says AI compute is sold out with $25 billion backlog as Black Forest Labs links generative video to robotics
Zhipu AI
2026-07-10 21:26:13

Zhipu AI Fixes Critical Bugs in GLM-5 Coding Agent System, Boosts Throughput by 132%

Zhipu AI resolved two severe bugs in GLM-5 models causing garbled text and repetition under high concurrency and long contexts. Fixes include race condition in PD-separation architecture and sync issue in HiCache. Additionally, a novel anomaly detection method and up to 132% throughput gain via LayerSplit KV Cache optimization.

240
Zhipu AI Fixes Critical Bugs in GLM-5 Coding Agent System, Boosts Throughput by 132%