Tether has introduced a new AI framework through its QVAC Fabric platform, bringing BitNet LoRA training to smartphones, GPUs, and other consumer devices. The company said the update allows billion-parameter models to run on phones and GPUs, aiming to cut training costs and widen access to AI tools on everyday hardware.
Cross-platform support spans desktop and mobile chips
The update adds cross-platform BitNet LoRA fine-tuning, letting models run across different hardware and operating systems. Tether said the framework supports GPUs from AMD, Intel, and Apple, along with mobile chipsets, using Vulkan and Metal backends for compatibility. According to the company, this is the first time BitNet LoRA has worked across such a broad device range, allowing users to train models on standard consumer hardware.
Lower memory demands and faster mobile inference
The system combines BitNet and LoRA to reduce memory and compute requirements. BitNet compresses model weights into simplified values, while LoRA limits the number of trainable parameters. Tether said this cuts hardware demands sharply. In its benchmarks, VRAM usage dropped by as much as 77.8% versus comparable systems, while GPU inference on mobile devices ran 2x to 11x faster than CPU inference.
Tests included Samsung S25 and iPhone 16
Tether also showed smartphone-based fine-tuning. Tests indicated that 125 million-parameter models could be trained in minutes on devices such as the Samsung S25. The company also reported successful fine-tuning of models up to 13 billion parameters on the iPhone 16, extending edge AI use cases to larger models. Support for mobile GPUs including Adreno, Mali, and Apple Bionic broadens development beyond specialized hardware.
CEO Paolo Ardoino said AI development often depends on expensive infrastructure, and described the framework as a shift toward local devices. Tether added that the system reduces reliance on centralized platforms and allows users to train models and process data directly on their own devices.

