GLM-5

Policy and Re
2026-08-24 05:03:10

Iran sanctions details due Monday as Nvidia server prices may rise more than 15% next year

PANews’ latest Wall Street morning briefing tracks several market threads moving at once: U.S. August composite PMI came in above expectations and lifted the major indexes on Friday, yet the weekly trend remained soft. Long-dated Treasury yields stayed near the top of their one-year range, while Bridgewater founder Ray Dalio warned that a U.S. debt crisis could emerge within one to five years if deficits and debt keep worsening. Another focal point is U.S. Treasury Secretary Bessent’s plan to release detailed Iran sanctions during Monday trading hours. According to the report, the package targets purchases of Iranian oil, money transfers and ship-to-ship transfers through secondary sanctions. In parallel, the AI trade is facing a cost test: some of Nvidia’s largest customers were told that servers using Nvidia AI chips could cost more than 15% extra next year because of surging memory chip prices. The article also reviews sharp stock moves. Tesla rose 5.14% and led the “Magnificent Seven” after Nevada approved Robotaxi operations in Las Vegas involving Tesla, Waymo and Uber. Alibaba dropped after announcing an HK$80 billion placement to fund AI expansion. Crypto-related stocks moved higher as Bitcoin approached $80,000 and expectations for liquidity and crypto-friendly policy improved.

620
Iran sanctions details due Monday as Nvidia server prices may rise more than 15% next year
UC Berkeley
2026-08-23 10:55:51

UC Berkeley and UT Austin researchers release edge-native MoE inference engine FreeToken

Researchers from the University of California, Berkeley and the University of Texas at Austin have introduced FreeToken, an edge-native mixture-of-experts inference engine designed to turn personal computers into a unified elastic inference platform. According to the release cited by Techub, the system can run the 753B-parameter GLM-5.2 model on a single workstation GPU, a 284B model on a gaming desktop, and a 35B model at interactive speed on a laptop GPU with 8GB of VRAM. FreeToken has been open-sourced under the Apache-2.0 license on GitHub and published on PyPI. The team also provides one-click desktop applications for Windows and Linux. Its command-line interface supports Linux x86_64 systems and NVIDIA GPUs, and the ft serve command can expose an API endpoint on port 1919 that is compatible with OpenAI and Anthropic. The project is aimed at individual developers, startups, and engineering teams at small and medium-sized businesses, with a focus on privacy-sensitive use cases such as healthcare, legal work, defense, finance, and intellectual-property-heavy R&D. Example applications include local coding agents, private code review, offline contract analysis, and synthetic data generation.

610
UC Berkeley and UT Austin researchers release edge-native MoE inference engine FreeToken
Ox Alpha
2026-08-23 02:15:10

Anonymous model Ox Alpha draws attention after coding tests place it near top-tier systems

An anonymous model called Ox Alpha has quickly become a focal point in the AI community after appearing on OpenRouter with a 1 million-token context window, multimodal input support for text, images, and video, tool use, and free access for now. What pushed it into the spotlight was not its listing, but its coding performance. Developer Ben Davis tested the model on 10 DeepSWE tasks and reported that it solved eight, for an 80% pass rate. In the comparison he shared, Fable 5 Max scored 65%, GLM-5.3 Max and Grok 4.6 xhigh each scored 62%, and GPT-5.6 Sol Max came in at 52%. A later run by other developers on a different DeepSWE subset produced a result of about 63%, which left Ox Alpha’s exact standing unresolved because the task sets and runtime configurations were not identical. At the same time, speculation about the model’s identity has centered on Zhipu. Analysts pointed to matching video-token behavior with GLM-5V-Turbo, a consistent 75-token gap versus GLM-5.3 across 25 prompts, and other product traits that resemble GLM routing and agent behavior. Ben Davis said he was 99% sure the model was GLM-5.x, but neither OpenRouter nor Zhipu had publicly responded as of publication. Separate debate has also formed around another anonymous model, korrine, now being tested on Code Arena.

850
Anonymous model Ox Alpha draws attention after coding tests place it near top-tier systems
Z.ai
2026-08-19 07:28:20

Z.ai founder Jie Tang says bigger parameter counts no longer tell the full story of model strength

Jie Tang, founder of Z.ai and a professor at Tsinghua University, argues that asking only how many parameters a model has no longer says much about how strong it is. In his review of the evolution of scaling laws—from GPT-3 to Chinchilla and then Mixture of Experts (MoE)—he says model capability depends on more than parameter count. Training data volume, where compute is spent, and how a model is actually used all matter. Tang’s point is that the old training-first view of scaling is less useful once commercial AI systems are deployed and called billions of times a day. Under that setup, inference cost changes the optimization target. A smaller model trained for longer may make more sense than a larger one trained less efficiently. He cited Llama-2-7B and Gemma-2-9B as examples of models trained far beyond the classic Chinchilla ratio. He also said MoE makes headline parameter numbers even less informative, because total parameters and activated parameters describe different things. For reasoning-heavy workloads, Tang argued that effective depth in a single inference pass and post-training may now be more important scaling dimensions. He described GLM-5.3 as a controlled test of that idea, keeping the base model and parameter counts unchanged from GLM-5.2 while expanding long-horizon environments and reinforcement learning over a month.

740
Z.ai founder Jie Tang says bigger parameter counts no longer tell the full story of model strength
U.S. debt
2026-08-19 02:07:00

SEC Unveils New Crypto Asset Rules as U.S. Debt, OpenAI Safety Measures, and Bitcoin News Dominate the Tape

PANews’ August 18-19 digest was packed with policy, markets, and AI updates. The U.S. debt burden may cross $40 trillion sooner than expected after tariff-related revenue losses accelerated Treasury borrowing. The SEC also proposed “Regulation Crypto Assets,” a new framework that includes startup and fundraising exemptions plus a safe harbor for digital assets that stop all managerial activity. OpenAI said it is tightening safeguards on its unreleased models, adding sandboxing and faster alerts after recent security concerns, while also pausing parts of its latest reinforcement learning training for two weeks. Elsewhere, Bhutan’s government-linked address moved 300 BTC, USDC Treasury minted 250 million USDC on Solana, and Metaplanet said it will use 2,100 BTC and $2.5 million in cash to acquire Super League and build a U.S. Bitcoin treasury platform called Superplanet. Cash App expanded beyond Bitcoin and USDC by integrating MoonPay, Ripple Prime sold $275 million of senior unsecured notes, and FASB proposed treating qualifying stablecoins as cash equivalents. The digest also covered NoOnes’ shutdown plan, a Maya Protocol exploit, Solana’s slot-time reduction, and major AI and IPO updates from OpenAI, Anthropic, Temporal, and Zhiyu/Unitree-related market listings.

1010
SEC Unveils New Crypto Asset Rules as U.S. Debt, OpenAI Safety Measures, and Bitcoin News Dominate the Tape
J-Space
2026-08-18 15:31:29

J-Space Faces Fabrication Questions After Community Retest Contradicts Its Published AI Benchmark Claims

J-Space Cognition Suite, an AI project that gained rapid traction on X, is facing community scrutiny after a retest challenged the benchmark results it had promoted for DeepSeek V4. According to MaxForAI, the project had claimed that pairing V4 Flash with J-Space could match GLM-5.3, while V4 Pro could outperform Fable 5 across multiple agent benchmarks. It also advertised a 2.53x speed improvement and a 2.21x gain in token efficiency. A GitHub user, GoForceX, said they reran the test with an 87-question subset from Terminal Bench 2.1 under high concurrency and confirmed that J-Space-related modules were loaded. The retest reportedly showed the opposite direction from J-Space’s claims: benchmark scores fell slightly after adding J-Space, while token usage and costs increased. The community has since called on the project to release its full evaluation setup, per-question results, execution logs, raw timing data, and token consumption records. So far, the project has mainly published aggregated results, with no complete raw experimental records available to verify the precise figures. The project author had previously said the data was 「indeed exaggerated」 and estimated actual gains at roughly 1.6x to 3x. As of now, there has been no formal response to the latest criticism, and some related issue threads have been deleted.

560
J-Space Faces Fabrication Questions After Community Retest Contradicts Its Published AI Benchmark Claims
Z.ai
2026-08-14 20:02:10

Z.ai launches GLM-5.3 and calls it the strongest open-weight coding model

Chinese AI lab Z.ai on Thursday introduced GLM-5.3, a 743-billion-parameter coding model the company describes as the strongest open-weight coder available. The model is already live through the GLM Coding Plan subscription and ZCode, while API access and downloadable weights are scheduled to roll out in stages after safety review. According to Z.ai, the main work behind GLM-5.3 was scaling post-training on the stack built for GLM-5.2, with more environments, more varied tasks, and more compute over the past month. The company said the new model was designed with token efficiency in mind rather than raw score chasing. On Z.ai Code Bench at Max effort, GLM-5.3 posted 34.5% while using about 75,000 output tokens per task, compared with GLM-5.2’s 23.4% at 96,000. It also showed stronger cybersecurity results, including an 84.5% score on CyberGym and 2,436 flagged vulnerabilities across 269 open-source projects. Even so, some leading U.S. closed models still rank higher on major coding benchmarks, while Z.ai says the model’s lower pricing and upcoming public weights remain central draws.

550
Z.ai launches GLM-5.3 and calls it the strongest open-weight coding model
Databricks
2026-08-14 09:45:00

Databricks closes $5 billion financing as DeepSeek opens Harness v0.1 developer preview

PANews’ daily roundup on Aug. 14 collected a wide spread of crypto, AI, regulatory and market developments, led by Databricks closing a $5 billion strategic financing and DeepSeek opening global testing for the developer preview of DeepSeek Harness v0.1 under the MIT license. The report also said Tether completed its first full independent financial statement audit, receiving an unqualified opinion from KPMG U.S. for Tether International, S.A. de C.V.’s 2025 accounts. In U.S. regulation, JPMorgan was reported to have ended its banking relationship with Polymarket last year over regulatory concerns, while the CFTC scheduled its Innovation Advisory Committee’s first meeting for Aug. 20 to discuss crypto assets, AI and prediction market oversight. The project and corporate section included Binance Alpha’s planned Aug. 14 listing of KiiChain (KII), SharpLink staking $200 million in ETH through Lido, and DeepSeek’s API price update that will take effect on Aug. 17. The funding and market data portion covered Kalshi’s talks for a new $750 million round at a $40 billion valuation, AMD’s potential bond sale of up to $5 billion, Bitcoin spot ETF net outflows of $131 million on Aug. 13, Reddit’s upcoming addition to the S&P 500, several crypto company earnings releases, and whale address activity tracked on-chain.

870
Databricks closes $5 billion financing as DeepSeek opens Harness v0.1 developer preview