Multiverse Computing Launches Open-Source Pulsar 16B Inference Model: 16.1B Parameters, Only 3.1B Activated, 4808 Tokens/s on Blackwell GPU

Multiverse Computing Launches Open-Source Pulsar 16B Inference Model: 16.1B Parameters, Only 3.1B Activated, 4808 Tokens/s on Blackwell GPU

N
News Editor
2026-06-27 16:40:57
Multiverse Computing has introduced the Pulsar 16B inference model, based on Nvidia's Nemotron architecture with 16.15 billion total parameters but only ~3.1 billion activated during inference thanks to CompactifAI compression technology. The model achieves 4,808 tokens per second on Nvidia Blackwell GPUs. Released under Apache 2.0 license, it is available on Hugging Face for commercial use. Founded in 2019, Multiverse Computing raised $215 million in Series B funding last year and previously collaborated with the Bank of Canada on quantum computing for financial simulations. This release signals the company's expansion from quantum to classical AI, targeting cost-efficient real-time inference for crypto and fintech applications.
Multiverse ComputingPulsar 16Bopen-source inference modelCompactifAIBlackwell GPUquantum computingfinancial AIsparse activation

On June 27, Multiverse Computing officially launched Pulsar 16B, an open-source inference model designed for efficient deployment in production environments. Built on Nvidia's Nemotron architecture, the model packs 16.15 billion total parameters but leverages the company's CompactifAI compression technique to activate only about 3.1 billion parameters during inference, dramatically reducing compute requirements. On Nvidia Blackwell GPUs, Pulsar 16B achieves a throughput of 4,808 tokens per second, making it suitable for real-time inference tasks such as high-frequency trading data processing and smart contract analysis.

Technical Highlights and Open-Source Licensing

The core innovation of Pulsar 16B lies in its sparse activation architecture: while possessing a 16.15 billion-parameter “brain,” each inference call only triggers roughly one-fifth of the parameters. CompactifAI achieves this balance through knowledge distillation and structured pruning, preserving accuracy while significantly boosting speed. The model is released under the Apache 2.0 open-source license, allowing developers to download weights from Hugging Face and integrate them directly into commercial products without royalty concerns. This lowers the barrier for AI adoption in financial technology, where latency and cost are critical.

Benchmark data shows Pulsar 16B outperforms comparable sparse models in token throughput (4,808 tokens/s) on Blackwell hardware. For crypto trading strategies, smart contract auditing, and on-chain data analytics that demand low-latency responses, this model offers a practical alternative to dense large language models. Although Multiverse Computing's primary focus remains quantum computing, the Pulsar 16B demonstrates how their optimization techniques—originally developed for quantum circuits—can enhance classical neural networks.

Company Background and Quantum-AI Strategy

Founded in 2019 and headquartered in Spain, Multiverse Computing specializes in applying quantum computing to finance and industrial use cases. In June 2025, the company closed a $215 million Series B round backed by European venture funds and strategic investors. Earlier, Multiverse collaborated with the Bank of Canada to explore quantum-enhanced financial risk simulations and portfolio optimization.

The release of Pulsar 16B signals the company’s pivot from pure quantum computing to hybrid solutions that combine quantum-inspired compression (CompactifAI) with classical AI models. Given the financial industry’s sensitivity to inference speed and operational cost, the sparse activation architecture could become a competitive advantage for high-frequency trading, risk management, and algorithmic compliance. The open-source strategy further aims to build a developer ecosystem around Multiverse’s technology, accelerating adoption in both traditional finance and decentralized finance (DeFi).

In the crypto space, lightweight inference models like Pulsar 16B are well-suited for on-chain anomaly detection, MEV strategy optimization, and DeFi protocol parameter predictions. Multiverse Computing’s interdisciplinary background—spanning quantum physics, AI, and finance—positions it uniquely to deliver tailored inference solutions for blockchain applications in the future.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
400

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.