Field note
EschaLabs released Escha-W2, a two-bit quantization of Qwen3.6-35B-A3B under Apache 2.0. The main artifact is roughly 12.3 GB, with 16 GB listed as a minimum memory target and 24 GB recommended. The underlying MoE uses 256 experts.
The model card documents SGLang for tool use, concurrency, structured output, and reasoning, plus a standalone ZML route and an MLX variant for Apple Silicon. EschaLabs reports about 5.6 times compression with mostly retained quality, while identifying coding as a noticeable weakness.
That makes the release a candidate local executor for bounded tasks, not a free substitute for a frontier service. Evaluation should focus on the exact tool schema, quantization regressions, and memory pressure of the intended runtime. The open license and local footprint make that testing feasible.