Contents
In brief
A Habr author bought a mini-PC with a Ryzen AI 9 365 and 64 GB of RAM to run a large local model on the NPU. The promised NPU+iGPU hybrid on Windows never worked for the model he wanted. On Proxmox with FastFlowLM he finally ran Qwen3.6–35B‑A3B around the clock: about 200 tok/s prefill and 13–17 tok/s decode. Along the way he learned the NPU sees roughly half of system RAM, serves one request at a time, and can crash when a client cancels a long call.
What happened
His older Ryzen 7 5800U box without a discrete GPU spent almost all of a long prompt on reading, not answering. AMD marketing pointed at the NPU for that phase, so he bought a FIREBAT mini-PC with Ryzen AI 9 365. Box plus memory landed near $825. The usual llama.cpp / Ollama / LM Studio stack ignores the NPU. AMD’s Lemonade promised a hybrid path, but only for small prepared models on Windows. For Qwen3.6–35B he needed FastFlowLM (flm) and its NPU model catalog.
On Windows the NPU driver was a separate quest. Once the model ran, Task Manager showed the NPU pool at about half of the RAM Windows could see, with more carved out for the iGPU in BIOS. A ~20 GB model plus KV cache and a second request already spilled over — fixed by setting UMA Frame Buffer to the minimum. With the hybrid dream gone, he moved to Proxmox VE 9. The NPU does not pass through to a VM cleanly, so flm runs on the host next to the hypervisor. Guest RAM was pinned hard, without ballooning, so inference would not starve the host of pages.
When home services pushed about 6 million tokens a day, the hardware limits showed up. One NPU behaves like a single checkout lane: a queue exists, parallelism does not. Clients with a 120 s timeout cancelled during prefill and flm died with a double free, taking the queue with it. An embedding model on the same NPU sometimes unloaded the 35B model to answer with a tiny one. The working setup: long client timeouts, Restart=on-failure, an explicit --ctx-len, a proxy with a global concurrency cap, and embeddings on CPU via llama.cpp.
Why it matters
A local large model promises quiet hardware, a predictable bill, and independence from someone else’s API. For this author’s traffic, ~$825 of hardware pays back in months against expensive cloud tiers — and in a year or two against the cheapest. The NPU is still not a GPU substitute: against an RTX 3090 it loses by a large margin on both prefill and decode. The win is power, noise, and a cookie-tin form factor.
For production the peak single-request speed matters less than queue behavior. One outstanding request, crashes on cancel, and model eviction by embeddings are service-design problems, not “install and forget.” Without a single front door in front of the NPU, several polite clients still create a storm one chip cannot absorb.
In practice
This is a measured field report, not a blanket reason to buy any mini-PC with an “AI” sticker. It fits a large MoE model running 24/7 in a quiet box, background jobs, long prompts, and short answers. It fits poorly for interactive chat at 13–17 tok/s or true multi-request parallelism.
- Budget NPU memory as roughly half of RAM minus the BIOS iGPU reservation; 32 GB is usually too little for a ~20 GB model.
- Do not count on NPU+iGPU hybrid for large models; for 35B the author ended on pure NPU via
flm. - On Proxmox keep
flmon the host — NPU passthrough is still immature. - Put one proxy with a global concurrency limit in front of the NPU; make client timeouts longer than the worst queue wait.
- Move embeddings and ultra-long prompts off the NPU so they do not evict the main model or block the queue for minutes.
Takeaway
An NPU mini-PC can run a serious local LLM around the clock if you treat the chip as one queue, not a farm. The author reached a stable Proxmox + FastFlowLM + Qwen3.6–35B setup with clear speeds on long prompts. The price of admission is RAM planning, drivers, client discipline, and a separate path for jobs that break the NPU.



Comments