Field note
Wafer's Kimi K3 post should not be treated as just another Kimi model item.
The interesting part is the infrastructure economics. Wafer frames the run around AMD MI355X, prefill optimizations, throughput, and performance per dollar. That shifts the discussion from model leaderboard novelty to the hardware-and-memory path that determines whether open models can be served competitively.
For New Runtime, the useful angle is memory as an inference moat. Long-context and agentic workloads make prefill, KV-cache handling, bandwidth, and node-level throughput more important than simple tokens-per-second headlines. The model matters, but the serving architecture is where the practical advantage can appear.
