GLM-5

AI Security
2026-07-21 09:01:16

AISI says open-weight AI cyber capability gap has narrowed to 4-7 months

The UK AI Security Institute said in a July 17 report that leading open-weight AI models are now 4 to 7 months behind frontier closed models in cyberattack capability, narrowing from a 6 to 10 month gap seen in internal testing last year. The institute used two evaluation systems: a 70-task benchmark covering vulnerability research, reverse engineering, web exploitation and cryptography, and a Cyber Range environment designed to test multi-step autonomous attack chains in a simulated enterprise network. Across those tests, GLM-5.2 matched Opus 4.6 on narrow tasks with a four-month lag, and reached Opus 4.5-level performance on Cyber Range with a seven-month lag, while DeepSeek V4-Pro tracked Opus 4.5 on narrow tasks with a five-month gap. The report also found a much wider pricing gap than the capability gap. In a Cyber Range run with a 100 million token budget, Opus 4.5 or 4.6 cost about $85, GLM-5.2 about $46, and DeepSeek V4-Pro $1.19. AISI said the shrinking lead time leaves defenders with less preparation time and sharpens a policy question now moving to the center: above what capability level should model weights no longer be openly released?

280
AISI says open-weight AI cyber capability gap has narrowed to 4-7 months
Policy Regula
2026-07-21 02:36:40

CAISI chief Chris Fall exits after three months as U.S. AI standards unit sees third leadership change in half a year

The U.S. Center for AI Standards and Innovation, or CAISI, has lost another leader, with director Chris Fall stepping down after only three months in the role. The Commerce Department did not give a reason for the resignation and said only that the appointment had always been temporary. Fall is the third person to cycle through the job in a little over six months, following Collin Burns, who lasted less than a week, and David Sacks, the White House AI and crypto czar who left in March. CAISI, housed under the National Institute of Standards and Technology, is meant to set technical standards for AI models, test capabilities, and identify cybersecurity risks. Yet several high-profile U.S. AI oversight actions in recent months unfolded without the agency playing a visible role, including the temporary restrictions on Anthropic’s Mythos and Fable models and the White House’s Gold Eagle vulnerability coordination mechanism. The latest leadership turnover also comes as debate intensifies in Washington over Chinese open-source AI models after Moonshot released a new version of Kimi that reportedly matched leading frontier systems in benchmark performance.

1020
CAISI chief Chris Fall exits after three months as U.S. AI standards unit sees third leadership change in half a year
Moonshot AI
2026-07-18 12:44:50

Moonshot AI launches Kimi K3, a 2.8-trillion-parameter model set for open-weight release on July 27

Moonshot AI has introduced Kimi K3, its latest flagship large language model, with 2.8 trillion parameters, a Mixture of Experts design, a native 1 million-token context window and built-in multimodal support across text, images and video. The company’s official X account said full model weights will be released on July 27, alongside four product endpoints: Kimi.com, Kimi Work, Kimi Code and the Kimi API platform. The article says Kimi K3 is among the world’s largest frontier models with open weights. It also outlines two named architectural changes, Kimi Delta Attention and Attention Residuals, which Moonshot says improve long-context decoding speed and training efficiency. The model uses 896 experts, activates 16 per inference step, and stores weights in MXFP4 format, putting total storage needs at roughly 1.4 TB. Beyond the model release, the report revisits Moonshot AI’s corporate backdrop. Bloomberg previously reported that the company’s annual recurring revenue had exceeded $200 million as of April 2026. Moonshot is now seeking to raise $2 billion at a $30 billion valuation and is restructuring for a Hong Kong IPO, according to the article. Benchmark data cited from ChainCatcher and Artificial Analysis places Kimi K3 ahead in coding and agent-style tasks, while still trailing Claude Fable 5 and GPT-5.6 Sol in general-purpose user experience.

1710
Moonshot AI launches Kimi K3, a 2.8-trillion-parameter model set for open-weight release on July 27
Zhipu
2026-07-17 03:20:50

Zhipu acquires Zhongke Jiahe to strengthen model infrastructure and compiler stack

Chinese large-model company Zhipu has spent several hundred million yuan to acquire Zhongke Jiahe, an AI heterogeneous computing software infrastructure company, according to a report cited by ChainCatcher from AI Technology Review. The deal is aimed at closing gaps in Zhipu’s lower-layer model engineering and compiler capabilities as it faces structural computing shortages and high-concurrency inference demands driven by rapid user growth. Zhongke Jiahe’s technology originated from the compiler laboratory at the Institute of Computing Technology under the Chinese Academy of Sciences. Its founder is Dr. Cui Huimin, and its core team previously worked on compiler development for several domestic chip projects, including Loongson, Sunway, Cambricon and Huawei Ascend. The company’s key strength is a virtual instruction set technology that uses a software middle layer to unify different chip brands and models into a large-scale cluster. The report said Zhongke Jiahe’s SigInfer inference engine claims it can cut large-model inference latency by as much as 74x. It also said Zhipu’s recently launched GLM-5.2 model saw average daily token calls jump 27x in its first week on an aggregation platform, exposing engineering bottlenecks in high-concurrency and long-context inference scenarios.

1090
Zhipu acquires Zhongke Jiahe to strengthen model infrastructure and compiler stack
Zhipu
2026-07-17 02:26:51

Zhipu shares fell more than 17% intraday as Kimi K3 launch and Anthropic criticism weighed

Zhipu’s Hong Kong-listed shares dropped sharply on July 17, falling more than 17% at one point during the session before trimming losses to nearly 15%, according to Bitget market data. The move came as fresh pressure emerged on the news front. Moonshot AI, described as Zhipu’s direct competitor, released Kimi K3 and said the model outperformed Claude Opus 4.8 and GPT-5.5 in some coding and agent tests. Separately, Anthropic named Zhipu for the first time a day earlier and accused GLM-5.2 of distilling Claude and OpenAI models. The combination of a rival product launch and direct allegations from Anthropic coincided with the stock’s intraday decline.

1110
Zhipu shares fell more than 17% intraday as Kimi K3 launch and Anthropic criticism weighed
Policy Regula
2026-07-14 12:30:00

All-In podcast weighs Anthropic and OpenAI IPOs, AI ROI, China model curbs and Trump Accounts

Episode 280 of the All-In Podcast pulled together four big debates now shaping the AI market: whether Anthropic and OpenAI should rush to IPO, how much real earnings lift AI is producing, whether open-source models can actually pull enterprise spending away from frontier labs, and what China’s reported plan to curb overseas access to top domestic AI models could mean. Jason Calacanis hosted the discussion with Chamath Palihapitiya, Brad Gerstner and David Sacks, with Brad filling in while Friedberg was away. The panel split sharply on valuation and timing. Brad argued Anthropic could become a blockbuster listing if annualized revenue tops $100 billion, while Chamath said companies should go public as soon as possible if the market has not fully absorbed weak downstream ROI. That skepticism carried into a broader argument over AI economics: Chamath said token costs at one of his companies are doubling every 45 days while productivity gains are capped at roughly 5%, leaving what he sees as only 0% to 2% real AI ROI for the S&P 493. The second half shifted to policy and market structure. Sacks said Chinese labs appear to be following a familiar pattern of going open while catching up, then closing once near the frontier. The episode closed with Brad’s extended defense of Trump Accounts, a plan that gives each U.S. newborn $1,000 in an S&P 500 investment account and reportedly opened 1.5 million accounts within 24 hours of launch.

1070
All-In podcast weighs Anthropic and OpenAI IPOs, AI ROI, China model curbs and Trump Accounts
IOSG
2026-07-14 01:32:50

IOSG says private AI is gaining ground as open models close the gap in cost and accuracy

IOSG argues that demand for private AI is rising across both enterprises and consumers as concerns over intellectual property leakage, data retention, and legal discovery become harder to ignore. The report maps the current privacy stack, from contract-based zero data retention and anonymous proxies to trusted execution environments, end-to-end encryption, fully homomorphic encryption, and local inference. Its main point is that the tradeoff is no longer as simple as privacy versus performance. A central example comes from Bridgewater’s AIA Labs and Thinking Machines. In a June 30 case study, an expert-tuned open model, Qwen3-235B, outperformed frontier models on financial judgment tasks while also delivering much lower inference cost. The model scored 84.7% on an independent test set, above an 80% threshold set by investment professionals. Frontier models averaged about 50% with simple prompts and reached 78.2% with expert prompting. By the report’s framing, the fine-tuned Qwen made 29.8% fewer mistakes than the best frontier baseline and ran at 13.8x lower inference cost. IOSG also says infrastructure for private inference and post-training is starting to mature. Enclave-based services from companies such as Phala, Tinfoil, and NEAR AI are pushing privacy costs down, in some cases to parity with or below plain-text routes. Still, major gaps remain in tool calling, agent workflows, and encrypted search, where privacy guarantees often break once requests leave the model layer.

1150
IOSG says private AI is gaining ground as open models close the gap in cost and accuracy
Private AI
2026-07-14 01:32:50

IOSG says private AI is gaining ground as open models close the gap in high-value enterprise work

IOSG argues that private AI is moving from a niche concern to a practical deployment choice for both enterprises and consumers. The report says the core issue is no longer abstract model safety, but where plaintext prompts, internal data, and company-specific judgment end up once they leave a user’s device. It reviews the current privacy stack, from contractual zero-data-retention and anonymous relays to trusted execution environments, end-to-end encrypted inference, fully homomorphic encryption, and local inference, and finds that costs and performance penalties are falling for several of these approaches. A central example comes from a June 30 case study by Bridgewater’s AIA Labs and Thinking Machines. In that work, an expert-tuned open model based on Qwen3-235B outperformed frontier models in both accuracy and inference cost on investment-related tasks, scoring 84.7% versus 78.2% for the best frontier setup using expert prompts, while cutting inference cost by 13.8x. IOSG’s argument is not that privacy AI is solved. Tool calls in agent workflows, encrypted search, and private post-training remain major gaps. But the report says the infrastructure needed to train and run open models inside controlled, attestable environments is arriving piece by piece, giving companies a clearer path to keep their own alpha inside their own boundary.

1670
IOSG says private AI is gaining ground as open models close the gap in high-value enterprise work