Contents
Why can’t you train a neural network on a plain CPU as efficiently as a “normal” program? Because the network is a sea of similar matrix operations that parallelize well. Ordinary code often follows sequential logic; AI hardware is built for massive parallelism, memory layout for tensors, and fast data movement. Below is a full walkthrough: CPU and GPU, CUDA and Tensor Cores, VRAM and bandwidth, TPU/NPU, Apple Silicon, multi-GPU, quantization, machine tiers, and a choice map for OCR, YOLO, and LLMs. Software context: PyTorch vs TensorFlow and ML frameworks overview.
Key takeaways
A neural network is massive parallelism. Matrices and convolutions benefit more from thousands of simple cores than from a handful of “smart” CPU cores.
A GPU is useless without a CPU. The processor prepares data, orchestrates the loop, and handles post-processing; the GPU runs tensor math.
VRAM often matters more than “teraflops.” Weights, activations, gradients, and optimizer state must fit in accelerator memory—especially during training.
Training ≠ inference. Training is memory-hungry and often multi-GPU; inference can be shrunk with quantization, NPU, CPU, or in-browser WebGPU.
The software ecosystem matters as much as the chip. The CUDA stack around NVIDIA is a huge practical advantage; alternatives (ROCm, Metal, TPU) are alive, but check your framework’s runtime.
What a neural network computes
A typical layer:
Y = XW + b
Inside: matrix multiplies, adds, activations, convolutions, attention, normalization. All of this lives in tensors—multidimensional arrays of numbers:
Tensor → operations → huge stream of numbers → compute hardware
A normal program: a few complex sequential steps. A network: billions of similar operations you can hand to thousands of execution units.
CPU vs GPU: different architectures
CPU — a few powerful cores, complex control logic, great for general-purpose and sequential work.
GPU — a huge number of simpler blocks, massive parallelism, strong at matrices and tensors.
CPU: [Core] [Core] [Core] [Core]
GPU: [ ][ ][ ][ ][ ][ ][ ][ ]
[ ][ ][ ][ ][ ][ ][ ][ ]
[ ][ ][ ][ ][ ][ ][ ][ ]
The myth “neural networks need only a GPU” breaks on reality:
CPU → data loading, preprocessing, orchestration, control
↓
GPU → tensor ops, matmul, layers, training / heavy inference
Without fast batch preparation, even a top-tier accelerator sits idle.
Inside the GPU, CUDA, and Tensor Cores
A modern AI graphics card is not “RGB decoration”—it is a compute node: execution units, matrix acceleration, VRAM, memory controllers, cache, interconnect.
Typical software path:
PyTorch / TensorFlow → CUDA → NVIDIA GPU
Around NVIDIA grew CUDA kernels, libraries, cuDNN, TensorRT, NCCL. The edge is not just silicon but a mature software ecosystem.
Tensor Cores (and similar matrix units) target mixed precision: FP32, FP16, BF16, TF32, INT8, and more. Large network → matrices → specialized blocks → faster math when you pick numeric formats carefully.
VRAM, bandwidth, and the data path
During training, memory holds more than weights:
VRAM → weights · activations · gradients · optimizer states · buffers
So a model that infers on 8 GB may not train on the same card: gradients and Adam inflate the budget.
Rough estimate:
Memory ≈ weights + gradients + activations + optimizer + overhead
Numeric format (FP32 / FP16 / BF16 / INT8) and tricks like LoRA/QLoRA change the number a lot.
VRAM capacity ≠ the whole story. Also weigh:
FLOPS vs Memory Bandwidth vs VRAM Capacity
A job may be compute-bound (limited by math) or memory-bound (limited by how fast tensors move through memory).
Data path:
Storage (SSD) → RAM → CPU → PCIe → GPU VRAM
Slow disk, narrow PCIe, or a starving DataLoader leaves the GPU waiting. With multiple GPUs, fast interconnect (NVLink and peers), NCCL, and collective ops matter—otherwise distributed training hits communication, not teraflops.
Multi-GPU, TPU, NPU, Apple, and others
When one card is not enough:
- data parallelism — different batches on model replicas;
- model / tensor / pipeline parallelism — model pieces on different GPUs.
More GPUs ≠ linear speedup: synchronization cost grows.
TPU — Google’s specialized tensor hardware (often TensorFlow/JAX). Not “another GPU” but another point on the spectrum: CPU is general, GPU is massive parallelism, TPU is narrowly focused tensor compute.
NPU — AI accelerator in laptops and phones, usually for inference: speech, camera, blur, local embeddings, light LLMs. Energy efficiency beats peak FLOPS.
Apple Silicon combines CPU, GPU, and Neural Engine via unified memory—no classic “separate VRAM island.” Handy for local AI; stack is Metal/MPS, MLX, and related tools.
Nearby: AMD (compute + ROCm), Intel (CPU/GPU/NPU), Google (TPU). The market is wider than NVIDIA, but practical choice is often dictated by framework support.
Storage, machine tiers, and power
Full training loop:
Dataset → SSD → RAM → CPU preprocess → PCIe → GPU VRAM
→ forward → loss → backprop → optimizer → new weights → repeat
Priority map:
Large model / training? → VRAM / GPU
GPU idle? → data pipeline / CPU / SSD / RAM
Small local model? → CPU / NPU / unified memory
Huge dataset? → SSD + RAM + prefetch
| Task | What to watch |
|---|---|
| Simple classifier | CPU or modest GPU |
| OCR / CRNN | GPU speeds training; inference often on CPU (CRNN) |
| YOLO / detection | GPU, VRAM, throughput (YOLO) |
| LLM inference | memory + bandwidth; quantization |
| LLM training | lots of VRAM, multi-GPU, fast interconnect |
| Smartphone | NPU, power, latency |
| Browser | WebGPU + compact weights (WebGPU) |
Configuration classes (no SKU shopping list):
- Learning / start: CPU, 16–32 GB RAM, modest GPU, SSD.
- Computer Vision: stronger GPU and VRAM, fast SSD, adequate CPU.
- Local LLM: available memory first (VRAM / unified); quantization.
- Serious training: multi-GPU, fast interconnect, large RAM, power and cooling.
Finding the bottleneck:
GPU ~100%? → compute-bound
GPU ~20%? → data pipeline / CPU / I/O
VRAM 100%? → batch / model / precision
Disk busy? → storage
Watch GPU/CPU utilization, VRAM/RAM, disk and PCIe throughput, samples/sec.
Common buying mistakes: FLOPS only; ignore VRAM and bandwidth; weak CPU with a strong GPU; undersized PSU and cooling; expect linear speedup from N cards; confuse training and inference hardware; treat an NPU as a datacenter GPU replacement.
Selection algorithm: model → training or inference → memory → throughput → precision → framework → supported backends → budget / power / cooling.
Ownership tiers (no price list—numbers move):
Laptop → workstation → one strong GPU → multi-GPU → AI server → cluster
Tasks, memory, power draw, and cost of ownership grow with tier. A GPU needs power and cooling: TDP, PSU, airflow, noise. A data center is racks, network, power, and cooling at scale—not an “RGB gaming PC.”
Training vs inference, quantization, and why it “doesn’t fit”
| Training | Inference | |
|---|---|---|
| GPU | Often critical | Depends on model |
| VRAM | High | Lower |
| Gradients / optimizer | Required | No |
| Latency | Less critical | Often critical |
| Power | Important | Often very important |
| NPU | Less common | Especially interesting |
Typical reasons a “powerful PC” still struggles: too little VRAM/RAM, model or context too large, heavy activations, large batch, fragmentation. Fixes: quantization, smaller model, LoRA, sharding, offloading, smaller batch, inference tuning (ONNX).
Quantization:
FP32 → FP16/BF16 → INT8 → INT4
Smaller size and memory footprint, sometimes faster—at some quality risk. For local LLMs it is often the only way to fit the model.
Stack: from application to silicon
Application
↓
PyTorch / TensorFlow / JAX
↓
ML libraries
↓
CUDA / ROCm / Metal / TPU runtime
↓
GPU / TPU / NPU / CPU
↓
Memory · PCIe / fabric · Storage / Network
A framework does not live in a vacuum: its capabilities hit the hardware stack. YOLO, CRNN, and Transformers are different workloads on the same “tensors → accelerator” principle.
FAQ
What should I build for OCR and YOLO?
SSD + enough RAM + NVIDIA GPU with VRAM headroom for training. Inference for compact OCR models often stays on CPU.
Is a Mac with unified memory enough?
For local inference and experiments—often yes. For heavy training in the CUDA ecosystem, NVIDIA + drivers is still the usual reference.
Why is GPU utilization stuck at 30%?
Usually data starvation: disk, DataLoader, CPU preprocessing, narrow PCIe. Inspect the pipeline, not only the model.
Do I need a TPU?
If you are already on Google Cloud / JAX/TF stack and the job maps to TPU—yes. Otherwise it is not a mandatory upgrade over a CUDA workstation.
Further reading
Conclusion
AI hardware is a system: the CPU prepares and orchestrates, GPU/TPU/NPU run tensors, memory and bandwidth decide whether things “fit,” storage feeds the pipeline, and power and cooling let the machine run for hours. Training and inference need different budgets; quantization and export lift some limits. Choose the chain “task → framework → runtime → accelerator,” not an isolated “most powerful card.”
Network math → Framework → Runtime (CUDA/…) → Accelerator → Memory/IO



Comments