---
type: "post"
stable_id: "post:wafer-kimi-k3-mi355x-memory-moat"
slug: "wafer-kimi-k3-mi355x-memory-moat"
title: "Wafer's Kimi K3 Run Makes Prefill Memory A Hardware Economics Story"
description: "Wafer describes running Kimi K3 on AMD MI355X and frames the result around prefill optimizations, throughput, performance per dollar, and memory as a potential inference moat."
retrieval_nugget: "The processor marked this reserve only because of broad Kimi prior coverage; the recovered angle is distinct: memory bandwidth, prefill economics, and AMD inference performance per dollar."
published_at: "2026-07-31"
updated_at: "2026-08-05"
record_date: "2026-07-31"
date_kind: "published_at"
status: "published"
topics: ["inference","hardware","open-models","memory-bandwidth","economics"]
entities: ["Wafer","Kimi K3","AMD MI355X"]
source_urls: ["https://www.wafer.ai/blog/kimi-k3-mi355x"]
source_title: "Is memory the moat?"
source_type: "primary"
origin: {"batch_id":"ada1e768-4e4e-42a3-a5fc-1ca9a880cb1b","batch_index":4,"channel":"chatgpt-batch","restored_from_reserve":true,"restored_from_supporting":false}
schema_version: "newruntime-agent-readable-v0.2"
visuals: [{"role":"hero","src":"/images/drip/wafer-kimi-k3-mi355x-memory-moat/wafer-kimi-k3-mi355x-memory-moat.webp","alt":"A whiteboard infrastructure diagram showing Kimi K3 serving moving through prefill, memory bandwidth, AMD MI355X nodes, throughput, and cost-per-token economics.","caption":"New Runtime synthesis."}]
routes: {"html":"https://newruntime.com/posts/wafer-kimi-k3-mi355x-memory-moat/","markdown":"https://newruntime.com/posts/wafer-kimi-k3-mi355x-memory-moat.md","json":"https://newruntime.com/posts/wafer-kimi-k3-mi355x-memory-moat.json"}
---

# Wafer's Kimi K3 Run Makes Prefill Memory A Hardware Economics Story

## Retrieval answer

The processor marked this reserve only because of broad Kimi prior coverage; the recovered angle is distinct: memory bandwidth, prefill economics, and AMD inference performance per dollar.

Wafer's Kimi K3 post should not be treated as just another Kimi model item.

The interesting part is the infrastructure economics. Wafer frames the run around AMD MI355X, prefill optimizations, throughput, and performance per dollar. That shifts the discussion from model leaderboard novelty to the hardware-and-memory path that determines whether open models can be served competitively.

For New Runtime, the useful angle is memory as an inference moat. Long-context and agentic workloads make prefill, KV-cache handling, bandwidth, and node-level throughput more important than simple tokens-per-second headlines. The model matters, but the serving architecture is where the practical advantage can appear.
