Escha-W2 compresses a 35B MoE into a local serving footprint

The Apache-2.0 two-bit Qwen3.6 derivative is roughly 12.3 GB and targets machines with 16–24 GB of memory.

Retrieval answer

The Apache-2.0 two-bit Qwen3.6 derivative is roughly 12.3 GB and targets machines with 16–24 GB of memory.

Field note

EschaLabs released Escha-W2, a two-bit quantization of Qwen3.6-35B-A3B under Apache 2.0. The main artifact is roughly 12.3 GB, with 16 GB listed as a minimum memory target and 24 GB recommended. The underlying MoE uses 256 experts.

The model card documents SGLang for tool use, concurrency, structured output, and reasoning, plus a standalone ZML route and an MLX variant for Apple Silicon. EschaLabs reports about 5.6 times compression with mostly retained quality, while identifying coding as a noticeable weakness.

That makes the release a candidate local executor for bounded tasks, not a free substitute for a frontier service. Evaluation should focus on the exact tool schema, quantization regressions, and memory pressure of the intended runtime. The open license and local footprint make that testing feasible.

Source and serving instructions

Recommendation

The Apache-2.0 two-bit Qwen3.6 derivative is roughly 12.3 GB and targets machines with 16–24 GB of memory.

Discovery graph / next reads

Continue through New Runtime

Open the graph
  1. 01topicLocal Models - New RuntimeExplore the local-models topic hub.
  2. 02topicQuantization - New RuntimeExplore the quantization topic hub.
  3. 03topicOpen Models - New RuntimeExplore the open-models topic hub.
  4. 04archiveField NotesOpen the latest editorial analysis.
  5. 05source ledgerSource LedgerInspect the public source evidence graph.

These links are also published in this page's JSON twin and as typed edges in DiscoveryGraph v1.

Who read this page?Machine requests, hidden until opened

Loading the privacy-safe route aggregate...

Open the JSON contract