Inco open-sources Splash for Mac local inference, showing up to 144 token/s

Inco open-sources Splash for Mac local inference, showing up to 144 token/s

N
News Editor
2026-09-19 08:22:59
Inco AI has released Splash, an open-source local inference engine for Mac, and said Qwen3.8-27B reached a peak of 144 token/s on an M5 Max MacBook Pro. LM Studio has already added support, with direct access available in version 0.4.25. The project builds on Inco’s earlier DFlash approach, which lets a smaller model guess multiple tokens in parallel before a larger model verifies them in batches, cutting the cost of token-by-token generation. Splash applies that optimization across the full local inference engine rather than as a narrower acceleration layer. For now, support is limited to Qwen3.8-27B and Qwen3.6-35B-A3B, with each model getting its own GPU kernel, memory setup, and DFlash 2 small model. Inco also shared benchmark figures from a 48GB M5 Pro used for side-by-side testing: Qwen3.8-27B hit 74 token/s in a single short-context stream, while aggregate throughput reached 170 token/s under four-way concurrency, or 3.9 times the runner-up. The company said that profile makes Splash better suited to Agent workloads, where total concurrent throughput can matter more than the speed of a single chat response.

Inco AI has open-sourced Splash, a local inference engine for Mac, and said Qwen3.8-27B reached a peak of 144 token/s on an M5 Max MacBook Pro. LM Studio moved quickly to add support, and the engine is already available directly in version 0.4.25.

From DFlash acceleration to a full inference engine

Inco’s earlier DFlash work has already been integrated into SGLang, vLLM, TensorRT-LLM, and llama.cpp. Meta, NVIDIA, Xiaomi, and Poolside have also paired DFlash with their own models to speed up smaller models.

The method works by having a smaller model predict multiple tokens in parallel, then handing those predictions to a larger model for batch verification. That reduces the compute burden of generating output one token at a time. Splash takes that same idea and extends it across the full local inference engine.

Model support is narrow, but heavily tuned

At this stage, Splash supports only Qwen3.8-27B and Qwen3.6-35B-A3B. Inco said each model has its own GPU kernel, memory scheme, and DFlash 2 small model. The trade-off is limited model coverage in exchange for more aggressive performance tuning.

Benchmark figures on M5 Max and M5 Pro

The 144 token/s figure is the highest demo result Inco showed on M5 Max hardware. On the 48GB M5 Pro system the company used for side-by-side testing, Qwen3.8-27B reached 74 token/s in a single short-context stream. Under four-way concurrency, total throughput rose to 170 token/s, which Inco said was 3.9 times the second-place result.

That is also why Splash is presented as a better fit for Agent workloads. When one Agent spins up several subtasks at the same time, aggregate concurrent throughput can matter more than the response speed of a single conversation.

Open-source release and system requirements

Splash is open-source and does not require LM Studio. Current requirements are a Mac with an M3 chip or newer, macOS 26.4 or later, and at least 36GB of unified memory.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
1500

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.