Vitalik Buterin said the low-memory release of DeepSeek V4 no longer depends on centralized servers, allowing users to run advanced AI models directly on their own machines. He framed that shift as an important step for open AI systems. Hardware still matters a lot: with 90 GB of VRAM, Apple systems can generate 35 tokens per second on the DeepSeek V4 2-bit model, while AMD-based systems reach about 7 tokens per second.
Hardware support shapes how open local AI can become
Buterin argued that broad support across hardware manufacturers is what separates a genuinely open AI ecosystem from one that is only decentralized in name. The point is practical. If advanced models run well only on a narrow set of machines, users still face limits even without a centralized server layer. The 2-bit quantized version is relevant here because compressing model weights cuts memory needs and makes local deployment more attainable.
LuceBox Hub posted faster results in Buterin’s tests
He also highlighted LuceBox Hub, a tool built to run heavy AI models faster and more efficiently on local hardware. In tests on an RTX 5090, Buterin said LuceBox Hub performed nearly twice as fast as llama.cpp, one of the most widely used local solutions. He added that the project is still under active development.
ZK links private model access with Ethereum RPC privacy
Buterin said local AI infrastructure could directly support Ethereum’s privacy goals. In his view, zero-knowledge proofs can be used both for paid remote large language model calls and for private RPC reads on Ethereum, because the two areas share technical foundations. Progress on one side can benefit the other. That creates a path where privacy-preserving AI access and privacy-preserving blockchain reads develop on parallel rails.
Ethereum-tuned AI models may speed up verification work
He also said AI models fine-tuned for the Ethereum ecosystem could improve smart contract and protocol security. As an example, Buterin pointed to the Leanstral model, which can generate Lean code at 38 tokens per second and remains competitive with much larger models. By extension, Ethereum-specific models could accelerate code verification and security audits across decentralized applications.
Buterin called for more systematic and automated approaches to smart contract auditing, saying artificial intelligence and blockchain should be developed in tandem to improve security and reliability.

