← All posts

Local LLM on an NPU mini-PC: what Ryzen AI 9 365 actually delivered

A Habr field report: Qwen3.6–35B on FastFlowLM, RAM traps, one-request-at-a-time NPU limits, and payback versus cloud APIs.

Local LLM on an NPU mini-PC: what Ryzen AI 9 365 actually delivered
Contents

In brief

A Habr author bought a mini-PC with a Ryzen AI 9 365 and 64 GB of RAM to run a large local model on the NPU. The promised NPU+iGPU hybrid on Windows never worked for the model he wanted. On Proxmox with FastFlowLM he finally ran Qwen3.6–35B‑A3B around the clock: about 200 tok/s prefill and 13–17 tok/s decode. Along the way he learned the NPU sees roughly half of system RAM, serves one request at a time, and can crash when a client cancels a long call.

What happened

His older Ryzen 7 5800U box without a discrete GPU spent almost all of a long prompt on reading, not answering. AMD marketing pointed at the NPU for that phase, so he bought a FIREBAT mini-PC with Ryzen AI 9 365. Box plus memory landed near $825. The usual llama.cpp / Ollama / LM Studio stack ignores the NPU. AMD’s Lemonade promised a hybrid path, but only for small prepared models on Windows. For Qwen3.6–35B he needed FastFlowLM (flm) and its NPU model catalog.

On Windows the NPU driver was a separate quest. Once the model ran, Task Manager showed the NPU pool at about half of the RAM Windows could see, with more carved out for the iGPU in BIOS. A ~20 GB model plus KV cache and a second request already spilled over — fixed by setting UMA Frame Buffer to the minimum. With the hybrid dream gone, he moved to Proxmox VE 9. The NPU does not pass through to a VM cleanly, so flm runs on the host next to the hypervisor. Guest RAM was pinned hard, without ballooning, so inference would not starve the host of pages.

When home services pushed about 6 million tokens a day, the hardware limits showed up. One NPU behaves like a single checkout lane: a queue exists, parallelism does not. Clients with a 120 s timeout cancelled during prefill and flm died with a double free, taking the queue with it. An embedding model on the same NPU sometimes unloaded the 35B model to answer with a tiny one. The working setup: long client timeouts, Restart=on-failure, an explicit --ctx-len, a proxy with a global concurrency cap, and embeddings on CPU via llama.cpp.

Why it matters

A local large model promises quiet hardware, a predictable bill, and independence from someone else’s API. For this author’s traffic, ~$825 of hardware pays back in months against expensive cloud tiers — and in a year or two against the cheapest. The NPU is still not a GPU substitute: against an RTX 3090 it loses by a large margin on both prefill and decode. The win is power, noise, and a cookie-tin form factor.

For production the peak single-request speed matters less than queue behavior. One outstanding request, crashes on cancel, and model eviction by embeddings are service-design problems, not “install and forget.” Without a single front door in front of the NPU, several polite clients still create a storm one chip cannot absorb.

In practice

This is a measured field report, not a blanket reason to buy any mini-PC with an “AI” sticker. It fits a large MoE model running 24/7 in a quiet box, background jobs, long prompts, and short answers. It fits poorly for interactive chat at 13–17 tok/s or true multi-request parallelism.

  1. Budget NPU memory as roughly half of RAM minus the BIOS iGPU reservation; 32 GB is usually too little for a ~20 GB model.
  2. Do not count on NPU+iGPU hybrid for large models; for 35B the author ended on pure NPU via flm.
  3. On Proxmox keep flm on the host — NPU passthrough is still immature.
  4. Put one proxy with a global concurrency limit in front of the NPU; make client timeouts longer than the worst queue wait.
  5. Move embeddings and ultra-long prompts off the NPU so they do not evict the main model or block the queue for minutes.

Takeaway

An NPU mini-PC can run a serious local LLM around the clock if you treat the chip as one queue, not a farm. The author reached a stable Proxmox + FastFlowLM + Qwen3.6–35B setup with clear speeds on long prompts. The price of admission is RAM planning, drivers, client discipline, and a separate path for jobs that break the NPU.

Comments

Loading comments…