UC Berkeley and UT Austin researchers release edge-native MoE inference engine FreeToken

UC Berkeley and UT Austin researchers release edge-native MoE inference engine FreeToken

N
News Editor
2026-08-23 10:55:51
Researchers from the University of California, Berkeley and the University of Texas at Austin have introduced FreeToken, an edge-native mixture-of-experts inference engine designed to turn personal computers into a unified elastic inference platform. According to the release cited by Techub, the system can run the 753B-parameter GLM-5.2 model on a single workstation GPU, a 284B model on a gaming desktop, and a 35B model at interactive speed on a laptop GPU with 8GB of VRAM. FreeToken has been open-sourced under the Apache-2.0 license on GitHub and published on PyPI. The team also provides one-click desktop applications for Windows and Linux. Its command-line interface supports Linux x86_64 systems and NVIDIA GPUs, and the ft serve command can expose an API endpoint on port 1919 that is compatible with OpenAI and Anthropic. The project is aimed at individual developers, startups, and engineering teams at small and medium-sized businesses, with a focus on privacy-sensitive use cases such as healthcare, legal work, defense, finance, and intellectual-property-heavy R&D. Example applications include local coding agents, private code review, offline contract analysis, and synthetic data generation.

Researchers from the University of California, Berkeley and the University of Texas at Austin have released FreeToken, an edge-native mixture-of-experts inference engine built to turn personal computers into a unified elastic inference platform.

The system is described as capable of running the 753B-parameter GLM-5.2 model on a single workstation GPU, a 284B model on a gaming desktop, and a 35B model at interactive speed on a laptop GPU with 8GB of VRAM.

FreeToken has been open-sourced on GitHub under the Apache-2.0 license and has also been published on PyPI. The project includes one-click desktop applications for Windows and Linux.

Its CLI supports Linux x86_64 systems and NVIDIA GPUs. With the ft serve command, users can expose an API endpoint on port 1919 that is compatible with OpenAI and Anthropic.

The engine is aimed at individual developers, startups, and engineering teams at small and medium-sized businesses, especially in settings with strict data privacy requirements. The release highlights healthcare, legal, defense, finance, and intellectual-property-intensive research and development as target scenarios. Example use cases include local coding agents, private code review, offline contract analysis, and synthetic data generation.

Techub cited MarkTechPost in its summary of the release.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
270

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.