posts

Making local MoE viable with asymmetric 2-bit quantization

Compressing routed experts down to 2-bit weights while preserving precision on shared attention paths is the kind of disciplined engineering local inference desperately needs. Salvatore Sanfilippo, the creator of Redis, has surfaced ds4, an inference project built around asymmetric quantization to squeeze routed Mixture of Experts models onto standard machines.

The real issue with MoE architectures at the edge has never been compute; it is raw memory footprint. Even if only a handful of experts activate per token, the entire model still has to live in RAM. Bludgeoning the entire network with uniform low-bit quantization tends to degrade coherence fast. Isolating the routed experts for aggressive 2-bit compression while protecting the shared routing and attention layers directly attacks the bottleneck without destroying output quality.

It is refreshing to see pragmatic systems thinking applied to local AI infrastructure, which has grown bloated under layers of unwieldy runtimes. Local, agentic software will never be practical if it requires a workstation-class GPU cluster just to idle, and this kind of mechanical sympathy is exactly how we shrink the footprint.

Source: dwarfstar.sh

← all posts