← All posts

AI hardware: CPU, GPU, TPU, NPU, VRAM, and what neural networks need

Why GPUs matter for neural networks, how CPU, TPU, and NPU differ, why VRAM and bandwidth matter, how multi-GPU setups work, quantization, and the stack from PyTorch down to silicon.

AI hardware: CPU, GPU, TPU, NPU, VRAM, and what neural networks need
Contents

Why can’t you train a neural network on a plain CPU as efficiently as a “normal” program? Because the network is a sea of similar matrix operations that parallelize well. Ordinary code often follows sequential logic; AI hardware is built for massive parallelism, memory layout for tensors, and fast data movement. Below is a full walkthrough: CPU and GPU, CUDA and Tensor Cores, VRAM and bandwidth, TPU/NPU, Apple Silicon, multi-GPU, quantization, machine tiers, and a choice map for OCR, YOLO, and LLMs. Software context: PyTorch vs TensorFlow and ML frameworks overview.

Key takeaways

A neural network is massive parallelism. Matrices and convolutions benefit more from thousands of simple cores than from a handful of “smart” CPU cores.

A GPU is useless without a CPU. The processor prepares data, orchestrates the loop, and handles post-processing; the GPU runs tensor math.

VRAM often matters more than “teraflops.” Weights, activations, gradients, and optimizer state must fit in accelerator memory—especially during training.

Training ≠ inference. Training is memory-hungry and often multi-GPU; inference can be shrunk with quantization, NPU, CPU, or in-browser WebGPU.

The software ecosystem matters as much as the chip. The CUDA stack around NVIDIA is a huge practical advantage; alternatives (ROCm, Metal, TPU) are alive, but check your framework’s runtime.

What a neural network computes

A typical layer:

Y = XW + b

Inside: matrix multiplies, adds, activations, convolutions, attention, normalization. All of this lives in tensors—multidimensional arrays of numbers:

Tensor → operations → huge stream of numbers → compute hardware

A normal program: a few complex sequential steps. A network: billions of similar operations you can hand to thousands of execution units.

CPU vs GPU: different architectures

CPU — a few powerful cores, complex control logic, great for general-purpose and sequential work.

GPU — a huge number of simpler blocks, massive parallelism, strong at matrices and tensors.

CPU:  [Core] [Core] [Core] [Core]

GPU:  [ ][ ][ ][ ][ ][ ][ ][ ]
      [ ][ ][ ][ ][ ][ ][ ][ ]
      [ ][ ][ ][ ][ ][ ][ ][ ]

The myth “neural networks need only a GPU” breaks on reality:

CPU → data loading, preprocessing, orchestration, control
         ↓
GPU → tensor ops, matmul, layers, training / heavy inference

Without fast batch preparation, even a top-tier accelerator sits idle.

Inside the GPU, CUDA, and Tensor Cores

A modern AI graphics card is not “RGB decoration”—it is a compute node: execution units, matrix acceleration, VRAM, memory controllers, cache, interconnect.

Typical software path:

PyTorch / TensorFlow → CUDA → NVIDIA GPU

Around NVIDIA grew CUDA kernels, libraries, cuDNN, TensorRT, NCCL. The edge is not just silicon but a mature software ecosystem.

Tensor Cores (and similar matrix units) target mixed precision: FP32, FP16, BF16, TF32, INT8, and more. Large network → matrices → specialized blocks → faster math when you pick numeric formats carefully.

VRAM, bandwidth, and the data path

During training, memory holds more than weights:

VRAM → weights · activations · gradients · optimizer states · buffers

So a model that infers on 8 GB may not train on the same card: gradients and Adam inflate the budget.

Rough estimate:

Memory ≈ weights + gradients + activations + optimizer + overhead

Numeric format (FP32 / FP16 / BF16 / INT8) and tricks like LoRA/QLoRA change the number a lot.

VRAM capacity ≠ the whole story. Also weigh:

FLOPS  vs  Memory Bandwidth  vs  VRAM Capacity

A job may be compute-bound (limited by math) or memory-bound (limited by how fast tensors move through memory).

Data path:

Storage (SSD) → RAM → CPU → PCIe → GPU VRAM

Slow disk, narrow PCIe, or a starving DataLoader leaves the GPU waiting. With multiple GPUs, fast interconnect (NVLink and peers), NCCL, and collective ops matter—otherwise distributed training hits communication, not teraflops.

Multi-GPU, TPU, NPU, Apple, and others

When one card is not enough:

  • data parallelism — different batches on model replicas;
  • model / tensor / pipeline parallelism — model pieces on different GPUs.

More GPUs ≠ linear speedup: synchronization cost grows.

TPU — Google’s specialized tensor hardware (often TensorFlow/JAX). Not “another GPU” but another point on the spectrum: CPU is general, GPU is massive parallelism, TPU is narrowly focused tensor compute.

NPU — AI accelerator in laptops and phones, usually for inference: speech, camera, blur, local embeddings, light LLMs. Energy efficiency beats peak FLOPS.

Apple Silicon combines CPU, GPU, and Neural Engine via unified memory—no classic “separate VRAM island.” Handy for local AI; stack is Metal/MPS, MLX, and related tools.

Nearby: AMD (compute + ROCm), Intel (CPU/GPU/NPU), Google (TPU). The market is wider than NVIDIA, but practical choice is often dictated by framework support.

Storage, machine tiers, and power

Full training loop:

Dataset → SSD → RAM → CPU preprocess → PCIe → GPU VRAM
        → forward → loss → backprop → optimizer → new weights → repeat

Priority map:

Large model / training?        → VRAM / GPU
GPU idle?                      → data pipeline / CPU / SSD / RAM
Small local model?             → CPU / NPU / unified memory
Huge dataset?                  → SSD + RAM + prefetch
Task What to watch
Simple classifier CPU or modest GPU
OCR / CRNN GPU speeds training; inference often on CPU (CRNN)
YOLO / detection GPU, VRAM, throughput (YOLO)
LLM inference memory + bandwidth; quantization
LLM training lots of VRAM, multi-GPU, fast interconnect
Smartphone NPU, power, latency
Browser WebGPU + compact weights (WebGPU)

Configuration classes (no SKU shopping list):

  • Learning / start: CPU, 16–32 GB RAM, modest GPU, SSD.
  • Computer Vision: stronger GPU and VRAM, fast SSD, adequate CPU.
  • Local LLM: available memory first (VRAM / unified); quantization.
  • Serious training: multi-GPU, fast interconnect, large RAM, power and cooling.

Finding the bottleneck:

GPU ~100%?  → compute-bound
GPU ~20%?   → data pipeline / CPU / I/O
VRAM 100%?  → batch / model / precision
Disk busy?  → storage

Watch GPU/CPU utilization, VRAM/RAM, disk and PCIe throughput, samples/sec.

Common buying mistakes: FLOPS only; ignore VRAM and bandwidth; weak CPU with a strong GPU; undersized PSU and cooling; expect linear speedup from N cards; confuse training and inference hardware; treat an NPU as a datacenter GPU replacement.

Selection algorithm: model → training or inference → memory → throughput → precision → framework → supported backends → budget / power / cooling.

Ownership tiers (no price list—numbers move):

Laptop → workstation → one strong GPU → multi-GPU → AI server → cluster

Tasks, memory, power draw, and cost of ownership grow with tier. A GPU needs power and cooling: TDP, PSU, airflow, noise. A data center is racks, network, power, and cooling at scale—not an “RGB gaming PC.”

Training vs inference, quantization, and why it “doesn’t fit”

Training Inference
GPU Often critical Depends on model
VRAM High Lower
Gradients / optimizer Required No
Latency Less critical Often critical
Power Important Often very important
NPU Less common Especially interesting

Typical reasons a “powerful PC” still struggles: too little VRAM/RAM, model or context too large, heavy activations, large batch, fragmentation. Fixes: quantization, smaller model, LoRA, sharding, offloading, smaller batch, inference tuning (ONNX).

Quantization:

FP32 → FP16/BF16 → INT8 → INT4

Smaller size and memory footprint, sometimes faster—at some quality risk. For local LLMs it is often the only way to fit the model.

Stack: from application to silicon

Application
   ↓
PyTorch / TensorFlow / JAX
   ↓
ML libraries
   ↓
CUDA / ROCm / Metal / TPU runtime
   ↓
GPU / TPU / NPU / CPU
   ↓
Memory · PCIe / fabric · Storage / Network

A framework does not live in a vacuum: its capabilities hit the hardware stack. YOLO, CRNN, and Transformers are different workloads on the same “tensors → accelerator” principle.

FAQ

What should I build for OCR and YOLO?

SSD + enough RAM + NVIDIA GPU with VRAM headroom for training. Inference for compact OCR models often stays on CPU.

Is a Mac with unified memory enough?

For local inference and experiments—often yes. For heavy training in the CUDA ecosystem, NVIDIA + drivers is still the usual reference.

Why is GPU utilization stuck at 30%?

Usually data starvation: disk, DataLoader, CPU preprocessing, narrow PCIe. Inspect the pipeline, not only the model.

Do I need a TPU?

If you are already on Google Cloud / JAX/TF stack and the job maps to TPU—yes. Otherwise it is not a mandatory upgrade over a CUDA workstation.

Conclusion

AI hardware is a system: the CPU prepares and orchestrates, GPU/TPU/NPU run tensors, memory and bandwidth decide whether things “fit,” storage feeds the pipeline, and power and cooling let the machine run for hours. Training and inference need different budgets; quantization and export lift some limits. Choose the chain “task → framework → runtime → accelerator,” not an isolated “most powerful card.”

Network math → Framework → Runtime (CUDA/…) → Accelerator → Memory/IO

Comments

Loading comments…