🚨 Chutes releases Parallax, a blueprint for decentralized MoE training that cuts compute needs and keeps sensitive data local 🚨
Chutes (SN64) has published a draft technical report on Parallax, a decentralized training method for mixture-of-experts models. The idea is to split expert routing across participants so training becomes lighter, faster to onboard, and privacy-aware by design.
🔑 Key highlights:
🔹️ Parallax splits MoE experts across participants instead of forcing every GPU to hold and synchronize the full model.
🔹️ The design cuts VRAM and FLOPS per worker while reducing the need for high-bandwidth GPU networking.
🔹️ Privacy is built into the architecture, with sensitive data staying on the owner’s hardware and workers only seeing narrow training slices or compressed sketches.
🔹️ Chutes tested the approach on 20B parameter models and says the results stayed close to baseline training, including a 4-GPU H100 run and an 8-node cheaper-hardware run.
🔹️ The report also includes a 176B feasibility test on four B300 nodes, showing the system can start, route across 640 experts per layer, sync across nodes, and reduce loss.
🔹️ The team says the bigger goal is not just decentralized training, but building model architectures that also make sense for inference and serving.
🎯 Bottom Line: Parallax is Chutes’ attempt to make decentralized MoE training practical by lowering compute costs, preserving privacy, and proving the system can still train large models.
https://taodaily.io/chutes-sn64-releases-parallax-a-new-blueprint-for-decentralized-moe-training/