Why a GPU is still essential, even when a CPU runs 2.8 trillion parameters
TL;DR
kimi-k3-in-c runs a 2.8-trillion-parameter model on a CPU with 8 GB of RAM. The price: 10 to 32 seconds to produce a single word. That is not a flaw a future release will fix, it is the direct consequence of reading a model from a disk rather than from memory built for the job. A GPU remains, and will remain, the only way to run this computation at a usable speed. The democratization this project demonstrates has a hard limit, and that limit is hardware.
What the feat proves
kimi-k3-in-c loads a 1.56 TB model on an 8 GB machine by keeping in memory only what serves the current instant: a compact dense trunk, and the 16 experts out of 896 the router just selected. The rest waits on disk. It is a solid proof of feasibility, tested and documented. It is not proof that the GPU has become optional.
The number that breaks the dream
32.7 seconds per word with 8 GB of RAM. 10.7 seconds per word with 128 GB. Writing a 100-word answer takes between 18 minutes and an hour of computation. A GPU produces the same answer in seconds.
No setting closes that gap. Multiplying the allocated RAM by sixteen, from 8 to 128 GB, buys a factor of three on speed. Returns fall off fast, because RAM does not attack the real cause of the slowdown.
Why this is structural, not an optimization gap
Generation time is dominated by disk reads, not by computation. An NVMe SSD reads at around 7 gigabytes per second. A high-end graphics card reads its own memory, VRAM, at more than 3,000 gigabytes per second, a factor of 400.
A GPU does not merely speed up multiplications, it runs them in parallel across thousands of cores, with the data already within reach in memory designed for that. A CPU streaming its weights from a disk, however well written the code, stays bound by the speed of that disk. Smarter code does not solve this, it is a hardware difference that code can only work around as best it can.
What a GPU does that this hack cannot
A GPU keeps the whole model, or the bulk of its active weights, in memory as fast as its compute units. It handles several requests in parallel without losing per-request latency. It powers a chatbot, a production API, an agent chaining dozens of calls per task: uses where the answer must arrive in seconds, not tens of minutes.
kimi-k3-in-c targets none of these uses, and the repository makes no such claim. The name of its slowest preset, laptop, describes a demonstration, not a workstation.
Where this approach belongs
Streaming a model from disk earns its keep when latency is not the problem to solve: checking that an open-source model answers correctly without renting a cluster, auditing a checkpoint in an isolated environment with no network, or simply understanding an inference engine's architecture by reading it rather than importing it as a black box. Those uses exist, and kimi-k3-in-c serves them well.
A hardware limit, not a software one
The fastest disk will always be slower than a GPU's memory, because it is not built for the same task. No software layer, however clever, closes a 400-fold bandwidth gap. kimi-k3-in-c democratizes access to a model's weights, to reading them, to understanding them. It does not democratize its speed, and nothing in its architecture claims otherwise. Producing an answer in seconds will still take a GPU.