Vitalik Buterin Sees Local AI as Key to Balancing Privacy and Speed
Ethereum co-founder Vitalik Buterin recently shared his thoughts on local AI and its potential to protect user privacy without sacrificing speed. According to Buterin, laptop AI is reaching a practical turning point. He noted that recent improvements in Qwen 3.8 Flash and llama.cpp have brought local models close to handling a 'large share' of tasks on his Strix Halo laptop.
Buterin's comments were made on September 17th, and he highlighted the ability of local AI to improve privacy while keeping the power to move funds behind separate, enforceable controls. He also mentioned that the benchmark image attached to the post showed impressive results, with input-processing rates ranging from 109.82 to 373.22 tokens per second, and output generation ranging from 18.42 to 33.37 tokens per second.
Buterin's views on local AI have evolved over time. In an April account of his local AI setup, he described a narrower role for laptop models, limiting them to tasks such as transcription and summarization. However, in the more recent post, he suggested that local models can now become the main interface for a larger share of activity and decide when remote models are necessary.
Qwen3.8-Flash-Next, released by Alibaba's Qwen team, is an open-weight multimodal mixture-of-experts model with 125 billion parameters. The official repository documents local text and vision inference through llama.cpp using quantized GGUF builds. Buterin emphasized that the capability benchmarks leave wallet authority unresolved.