Prime Intellect, an AI infrastructure company, has introduced Prime Inference, a new inference platform built for frontier open-source models. The service offers two deployment options: serverless endpoints for variable demand and reserved capacity for steady workloads. According to the company, the platform runs on its own multi-data-center GPU clusters and was already processing nearly 1 trillion tokens per day internally before its public release.
Prime Inference serves as the inference layer of Prime Intellect’s open training stack and supports an OpenAI-compatible API. The stack combines NVIDIA Dynamo, vLLM, Mooncake, and FlashInfer, with the stated goal of optimizing agent workloads. Prime Intellect also said its GLM-5.3 endpoint leads on speed on OpenRouter, posts a near-zero tool-calling error rate, and has maintained 100% uptime since launch. The platform currently uses NVIDIA Blackwell hardware, with support for Vera Rubin planned later. Team-wide consolidated billing and usage tracking are already live, while model-specific pricing has not yet been fully disclosed.
Prime Intellect has launched Prime Inference, an inference platform for frontier open-source models that offers serverless endpoints and reserved capacity.
The company said the platform runs on its own multi-data-center GPU clusters and was handling nearly 1 trillion tokens per day internally before the public release.
Prime Inference is the inference layer in Prime Intellect’s open training stack and supports an OpenAI-compatible API. Its stack combines NVIDIA Dynamo, vLLM, Mooncake, and FlashInfer, which the company said is designed to optimize agent workloads.
Prime Intellect also reported that its GLM-5.3 endpoint leads on speed on OpenRouter, has a near-zero tool-calling error rate, and has maintained 100% uptime since launch.
The platform comes with two service modes: serverless endpoints for variable demand and reserved capacity for sustained workloads. The hardware currently uses NVIDIA Blackwell, with support for Vera Rubin planned in the future.
Team-level consolidated billing and usage tracking are already available, though model-specific pricing has not been fully disclosed.
The item was cited by Techub and attributed to MarkTechPost.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.