
Qwen 3.8 27B on Cerebras at ~1500 tokens/s
Qwen 3.8 27B on Cerebras public endpoints: ~1500 tokens/s, 64k/128k context, original unpruned weights in production.
Developer news without the noise — My Dev News in Telegram.

Tag
All blog posts with this tag.

Qwen 3.8 27B on Cerebras public endpoints: ~1500 tokens/s, 64k/128k context, original unpruned weights in production.

vLLM on Blackwell: 96.9% prefix cache hit rate, 452:1 input/output, and why teams keep inference in-perimeter for control—not token savings.