Contents
Three listings sit on the table. A 30-square-meter apartment costs 3.0 million, 50 square meters costs 4.2, 70 costs 5.4. You need a price for a fourth flat without calling an appraiser.
Those listings run through the whole piece. Area and price are the data. The rule “price per square meter times area, plus a fixed add-on” is the function that turns area into an answer. The price per meter and the add-on are the parameters, two numbers you can turn. The miss, in millions, is the loss. “Nudge the price per meter up” or “nudge it down” is the gradient. A listing from another neighborhood, absent while you tuned the rule, is a new example. The mathematics of machine learning is not a shelf of formulas. It is a way to write a flat as numbers, set the numbers in the rule, measure the miss, and check whether the price still holds on listings that were not there during tuning.
Key takeaways
Training means setting the price per meter and the add-on from the listings. You choose the shape of the rule first. The numbers inside it move until the miss on known listings gets smaller.
A straight line does not draw a bend. A linear layer can weigh features. Bends, forks, and yes-or-no answers show up when layers stack and a nonlinear cut sits between them.
Error is a direction, not a verdict. The derivative says whether to raise or lower the price per meter. Backpropagation carries that advice from the final miss back to the earliest numbers.
Hitting your own listings does not prove the price in another neighborhood. The sample, the loss, and the metric decide which miss you treat as expensive. Accuracy on lopsided data is a cheap victory.
The same mathematics sits inside a classifier, a convolution, a recurrence, and attention. The shape of the rule changes: a pixel window, a memory of the previous step, or a weighted average by meaning.
Training a model means setting the price per meter
Three different objects get called “a neural network,” though for these listings they are different objects.
An algorithm is a fixed procedure. “Multiply area by five hundredths and add one million” is a finished rule. The numbers were set ahead of time. Your three listings do not move them.
A mathematical model is a rule shape with empty numbers. For price from area the shape is: price = weight × area + bias. The weight is the price per meter, the bias is the add-on. The slots exist; the numbers do not.
A trained model is that same shape with the numbers filled in. Weight 0.06 and bias 1.2 hit the three listings exactly: 0.06 × 30 + 1.2 = 3.0; 50 square meters gives 4.2; 70 gives 5.4. Training, here, means finding those two numbers.
Set the weight to 0.05 and the bias to 1.0 and the predictions become 2.5, 3.5, and 4.5. The misses are 0.5, 0.7, and 0.9 million. The mean of the squared misses is 0.52. The square is not decoration. A large miss weighs more than a small one, and the signs “too cheap” and “too expensive” do not cancel when you add them. Zero error at weight 0.06 says something only about these three cards. A fourth flat with a renovation, a floor, and a neighborhood was never in the rule.
The chain is the same for any number of listings. The data are the cards. The parameters are the numbers in the rule. The function computes a price. The loss says how wrong that price is. Training moves the numbers toward a smaller miss. Held-out listings answer whether you memorized only these three cards.
areas = [30, 50, 70]
prices = [3.0, 4.2, 5.4]
def predict(area, weight, bias):
return weight * area + bias
def mean_squared_error(weight, bias):
errors = [
predict(area, weight, bias) - price
for area, price in zip(areas, prices)
]
return sum(error * error for error in errors) / len(errors)
print(round(mean_squared_error(0.05, 1.0), 2)) # 0.52
print(mean_squared_error(0.06, 1.2)) # 0.0
Nudge the weight by one hundredth and watch the error swell. A one-number model is enough to see that training is a search over numbers, not a magic program.
The numbers you need along the way
You do not need a year of analysis before the first model. You need a short set of words, each tied to these listings.
A variable is a name for a number that can change: area, weight, error. A function takes arguments and returns a number. Price from area is a function. The argument is what you put in, not what the function “thinks.”
Percents, powers, and logarithms show up earlier than they look. A probability of 0.01 is one chance in a hundred. A power grows or shrinks a number by repeated multiplication. A logarithm will later punish a confident mistake harder than a shy one: the negative log of 0.1 is larger than the negative log of 0.9. For now it is enough that the log gets large when its argument nears zero.
A graph is the same function laid on a plane. The price per meter runs horizontally, the miss vertically. A minimum is a dent where the miss is smallest. A plateau is a shelf where the miss barely changes, even though the price is still bad.
A sum and a mean compress a pile of numbers. For 2, 4, 4, 4, 5, 5, 7, 9 the mean is 5. Spread has two common versions. Describe this exact list, divide the sum of squared deviations by 8, and the standard deviation is 2. Treat the list as a short pack drawn from a longer list and divide by 7, and you get about 2.14. Distance between two objects is the length of the segment between their numbers. Areas 30 and 70 are 40 meters apart; with several features, distance uses every axis at once.
A derivative is how fast the miss changes when you nudge the price per meter. Positive: the miss grows as that price grows, so lower it. Near zero: a small move in the price per meter barely changes the miss.
A probability is not “the model feels unsure.” It is a number from 0 to 1 that describes a share of outcomes. A random variable is a quantity that jumps according to a distribution when you repeat the trial.
Ahead of time, be comfortable reading a function, a mean, a graph, and the sentence “nudge the argument.” Pick up logarithms, partial derivatives, and Bayes in the section where the task first needs them. Do not save them for the end of a textbook. They are short when an example sits beside them.
values = [2, 4, 4, 4, 5, 5, 7, 9]
mean = sum(values) / len(values)
spread = (sum((value - mean) ** 2 for value in values) / len(values)) ** 0.5
distance = abs(70 - 30)
print(mean, spread, distance) # 5.0 2.0 40
An object becomes a row of numbers, a layer a product
A network does not see an apartment, a letter, or a frame. It sees a row of numbers. Area, floor, and “has a balcony” are already a feature vector: (50, 4, 1). Several rows make a matrix. A batch of rows is a tensor. The word adds no magic. A tensor is an array whose axes have an agreed order.
A dot product weighs features and adds them. Features (1, 2) and weights (0.5, −0.1) give 0.5 × 1 + (−0.1) × 2 = 0.3. That is one output adding a weighted sum of the input. A weight matrix does it for several outputs at once. A linear map rotates, stretches, and shifts a cloud of points. By itself it does not bend a line into an arc.
A grayscale image is a matrix of brightness. A color image is a tensor with a channel axis. Text, after it is cut into pieces, is a sequence of integers and then rows of real vectors. A batch of eight images adds another axis. Mismatched sizes are the most common break: you multiply a row by a row of a different length. The lesson on an image as a tensor shows it on a text crop: pixels are already numbers, and convolution is still ahead.
A linear layer is y = Wx + b. Here x is the input vector, W the weight matrix, b the bias, y the output. Bias exists so a zero input is not forced to a zero output. A zero-area flat would still “cost” 1.2 in our formula. That is not a market fact. It is a property of the shape.
Take input (1, 2), a matrix with rows (0.5, −0.1) and (0.2, 0.3), and bias (0.1, −0.2). The first output is 0.5 × 1 + (−0.1) × 2 + 0.1 = 0.4. The second is 0.2 × 1 + 0.3 × 2 − 0.2 = 0.6. The same layer in PyTorch is torch.nn.Linear. Weights are stored as “out by in,” and the result matches if you copy the same numbers.
import torch
x = torch.tensor([1.0, 2.0])
layer = torch.nn.Linear(2, 2)
with torch.no_grad():
layer.weight.copy_(torch.tensor([[0.5, -0.1], [0.2, 0.3]]))
layer.bias.copy_(torch.tensor([0.1, -0.2]))
print(layer(x)) # tensor([0.4000, 0.6000])
The framework around those layers is a separate conversation. PyTorch, TensorFlow, and JAX differ in how the layer is written in code, not in the multiplication table.
A straight line does not draw a bend
Stack two linear layers with nothing between them and you still have one linear layer. Two “multiply and add” corrections in a row are still one such correction. So a function sits between layers and cuts, squashes, or softly damps.
ReLU keeps a positive number and zeros a negative one. On −2, −0.5, 0, 0.5, 2 it returns 0, 0, 0, 0.5, 2. The dead zone on the left is the price of simplicity: while the sum before the cut is negative, that branch passes no gradient.
sigmoid squashes any number into the range from 0 to 1. The same inputs give about 0.119, 0.378, 0.5, 0.622, 0.881. Useful when you want to read the output as a share. On the far ends it is almost flat, so changing the weight barely helps.
tanh is similar, on the range from −1 to 1: about −0.964, −0.462, 0, 0.462, 0.964. Zero stays zero. Large inputs stick to the ends.
GELU is smoother than ReLU: negative numbers are not chopped to zero, they are shrunk. The usual approximation on the same inputs gives about −0.045, −0.154, 0, 0.346, 1.955. Recent language blocks often use it not because it is “smarter,” but because a smooth cut stays usable in a deep stack.
softmax turns raw scores into shares that add to one. Scores 2.0, 1.0, and 0.1 become about 0.659, 0.242, and 0.099. That is not a fact about the world. It is a way to spread a lead across competitors so the shares sum to 1.
One linear operation is not enough when the boundary between classes bends, or when the answer depends on a combination rather than a single feature. “Large area and ground floor” is not the sum of two independent bonuses if the ground floor spoils only the large flat. A composed function — layer, cut, layer — builds those combinations. Depth is not ornament. Each cut adds a kink the previous straight line could not draw.
Plot the four activations on one axis from −2 to 2. The eye sees where the function goes silent, where it cuts, and where a small change in the weight is still heard.
The derivative says which way to turn
Return to the dent on the graph of the miss. For a round picture, take one number w and the error L(w) = (w − 3)². The minimum is at 3, where the error is zero. This is not the price per meter 0.06 from the listings: it is the same idea on a dent whose bottom you can see at once. The derivative is 2(w − 3). Left of 3 it is negative: the error falls as the number grows. Right of 3 it is positive: the error falls as the number shrinks. The gradient is that derivative collected over every number. It points to the steepest rise of the error. To shrink the miss, you step against it.
The step is θ(t+1) = θ(t) − η · ∇L(θ(t)). θ is every number in the rule at once, η is the learning rate, the size of the step. Start at zero with rate 0.1. The error is 9, the gradient is −6, the new number is 0.6. Then it goes 1.08, 1.46, 1.77, 2.02, 2.21, 2.37. Over those steps the error falls from 9 to about 0.40. Three is still ahead: the step size is fixed, the dent narrows, and the tail of the approach is long.
Too large a step jumps the dent and can climb out. Too small a step crawls until patience or data runs out. A local minimum is a dent that is not the deepest on the whole graph, but a small step cannot leave it. A plateau is a place where the gradient is almost zero while the miss is still large: a small move of the number barely says which way to go. In practice you watch whether the error falls, not only the final digit, and whether it saws over the rim.
The picture for this section is the points (w, L) by step. They should slide into the dent, not draw a saw.
The miss has to travel back through the layers
In a network of many layers, a number in the first layer is not visible in the final price. It changes an intermediate calculation, that changes the answer, and the answer changes the miss. To know how to move an early number, you carry the error backward along the same chain. That is the chain rule: the derivative of a composed function is the product of the derivatives along the way.
A small network, every number visible. Input x = 1, target 0. First layer: z = 0.5 × x − 0.2 = 0.3. ReLU leaves 0.3. Second layer: y = 0.8 × 0.3 + 0.1 = 0.34. The error is the squared miss: (0.34 − 0)² = 0.1156.
The derivative of the error with respect to the output is 2 × 0.34 = 0.68. With respect to the second weight: 0.68 × 0.3 = 0.204. With respect to its bias: 0.68. With respect to the hidden number: 0.68 × 0.8 = 0.544. The cut, being to the right of zero, passes the gradient unchanged; to the left it would zero it. The input is 1, so the first weight and the first bias both receive 0.544.
PyTorch autograd computes the same thing if you mark the numbers as values that need a gradient and run the backward pass. A disagreement in the third digit is a reason to check the formula, not a reason to blame the framework. A hand calculation on two layers is worth doing once. Deeper networks keep the same meaning. The miss in the answer reaches the first number as a product of factors. Many factors smaller than one fade the signal. Many factors larger than one blow it up. That is why depth, initialization, and cuts that go silent at the ends need care.
import torch
w1 = torch.tensor(0.5, requires_grad=True)
b1 = torch.tensor(-0.2, requires_grad=True)
w2 = torch.tensor(0.8, requires_grad=True)
b2 = torch.tensor(0.1, requires_grad=True)
hidden = torch.relu(w1 * 1.0 + b1)
output = w2 * hidden + b2
loss = (output - 0.0) ** 2
loss.backward()
print(round(w2.grad.item(), 3)) # 0.204
print(round(w1.grad.item(), 3)) # 0.544
A sample is not the world, and not every miss costs the same
Three apartments are not a market. They are a short pack of cards. A random variable describes how a number jumps from trial to trial. A distribution says which jumps are common. Mean and variance compress that into a center and a width. A conditional probability is a share given something you already know: price given area, not price “in general.”
Bayes’ theorem swaps the condition. A disease shows up in 1 person out of 100. A test catches it 99 times out of 100 and falsely fires for 5 healthy people out of 100. A positive test is still more often a healthy person: the probability of disease given a positive test is about 0.17, not 0.99. A model that emits a share is easy to confuse with a frequency in the world. Bayes is the reminder that a rare class and a sensitive test do not grant automatic confidence.
Maximum likelihood walks into the same task from another door. You choose numbers that make the observed prices the most expected. For many tasks that is the same move as minimizing a particular loss. Cross-entropy belongs to that family: it is large when the model assigned a small share to the true class.
Sample bias is a skew in the cards. If every listing comes from one neighborhood, the weight “a meter costs six hundredths of a million” does not have to survive another neighborhood. Representativeness is how much the pack resembles the neighborhood the new listings will come from. Overfitting is a rule that memorized today’s cards, oddities included. Underfitting is a shape too crude for the bend you actually have. Quality on the training cards does not promise quality on new ones. So you split examples: one part moves the numbers, another watches whether the error rose. The cut itself changes the estimate. A random split inside one neighborhood looks cheerful and useless if the new listings come from somewhere else.
The loss decides which miss is expensive. Mean squared error, MSE, pulls hard toward an outlier: one apartment at 20 million drags the line. Mean absolute error, MAE, treats misses more evenly and tolerates a rare absurd price. For a yes-or-no answer, binary cross-entropy looks at the share given to the true answer. For several classes, the same idea runs across all of them. The negative log of the true-class share, for scores 2.0, 1.0, and 0.1 with the first class correct, is about 0.42. Equal shares would cost about 1.10. A confident mistake, almost all the mass on the wrong class, costs still more.
The same architecture with different losses diverges on data that contains an outlier. The square chases the outlier. The absolute value stays with the main mass. The smallest error on the training cards still does not promise new listings: the loss was chosen for the training cards.
Draw two samples from different centers, plot the histograms, and fit a line on one while scoring it on the other. A gap in the means matters more than a pretty training curve.
One pass: price, miss, correction
The full cycle puts the earlier pieces into one pass.
First you prepare the input: numbers on a comparable scale, a shared length, an honest split into training, validation, and held-out test. The forward pass computes the prediction: layers, cuts, and softmax when you need shares. The loss compares the named price with the real one. The backward pass writes a gradient for each number. The optimizer takes a step. Then the pass repeats.
An epoch is a full pass over the training listings. A batch is the pack that produces one gradient: not the whole neighborhood, a few cards. An iteration is one such step. Stochastic gradient descent uses a small pack, so the direction is noisy and the step is cheap. momentum adds inertia: the correction does not twitch at every odd card. Adam keeps a separate pace per number and adapts the step. AdamW decays the weights themselves so that penalty does not get mixed into the adaptation. A first straight line only needs an ordinary step against the gradient. Optimizer names matter when there are many numbers on different scales.
Initialization sets the numbers before the first pass. Zeros in every weight make symmetric neurons twins: the gradient cannot tell them apart. Starting numbers that are too large drive a sigmoid onto its flat end. Normalization brings a pack to a stable scale so the next layer does not receive a whisper and then a shout. Regularization is a penalty for numbers in the rule that grow too large: weight decay, randomly zeroing some connections (dropout), or stopping early when validation error rises while training error still falls.
Training diverges when the step is too large or the loss is computed on an awkward scale. It stalls when the gradient dies in a flat cut, the data carry no signal, or the numbers have been pushed flat by the regularizer. A wrapper such as Trainer hides this pass. The useful minimum is to see it without the wrapper.
On synthetic prices, with area divided by ten, the true slope is about 0.6 and the bias about 1.2, plus a little noise. Start from zero with rate 0.02. The first loss is on the order of 18: the prediction is zero, the prices are not. After a few hundred steps the loss falls toward hundredths, the slope nears 0.6, and the bias catches up more slowly. The feature was not centered, so the price per meter and the add-on live on different scales. That is not a broken optimizer. It is the geometry of the rule. Fix the generator seed, and write down the rate and the step count, or “it converged for me” cannot be repeated.
Frameworks differ in how this pass is written down. A comparison of the two writings is in PyTorch and TensorFlow.
Metrics, a pixel window, and a memory of the last step
When the answer is a label rather than a price, the line becomes a boundary. Logistic regression is a linear score plus a sigmoid: one side of a threshold is “yes.” Several classes mean several scores and a softmax. A linear boundary is flat. A bent boundary needs the cuts from earlier.
A confusion matrix spreads the misses out. Ten photos: six cats and four dogs. The model correctly names five cats and three dogs, calls one cat a dog, and one dog a cat. Accuracy is 0.8. Precision for cat is 5 of 6 alarms. Recall is 5 of 6 real cats. F1 folds precision and recall into one number; here it is about 0.83. Call everything a cat and accuracy drops to 0.6, cat recall becomes 1, and precision becomes 0.6. A single accuracy does not show that split. A rare disease is the same trap as the medical test above.
For a box around an object, look at intersection over union, IoU. Two 2×2 squares shifted by one unit overlap on area 2 with union 6: IoU is one third. Mean average precision, mAP, gathers those hits across thresholds and classes into one detection summary. The model’s confidence is the share after softmax, and it should be checked against how often that share is right.
Convolution is the same multiplication, with a window moving across the picture. A 2×2 kernel with entries (1, 0; 0, 1) on the picture
1 2 3
0 1 2
1 0 1
produces four responses: 2, 4, 0, 2. The kernel looks for a local pattern. It does not weigh every pixel with one vector. A feature map is the new “photo” of responses. Output size depends on window size, stride, and edge padding. A frame classifier, a detector, and digit recognition differ in the head of the network and in the metric, not in the multiplication table. A practical entry to a frame as numbers is the CNN and CRNN series and the overview of vision in PyTorch and TensorFlow.
A recurrent step remembers the previous step. The hidden state is a number or a vector carried to the next token. Reduce it to one number: new state = tanh(0.5 × previous + 1 × current input), start at zero, inputs 1, 0, 1. The states are about 0.762, then 0.363, then 0.828. A zero on the second step does not wipe memory: 0.363 is the trace of the one. LSTM and GRU add gates that decide what to keep, what to forget, and what to release. Without gates, a long chain of factors fades or blows up — the same effect as a deep stack of layers.
Build the 3×3 convolution by hand, then three recurrent steps on paper. Both are shorter than a library call, and both show where the parameters live.
Attention is a weighted average by meaning
A language model feeds on a sequence of vectors, embeddings: a piece of text became a row of numbers. A positional pattern adds the place in the line. Otherwise “the cat chased the dog” and the reverse order would look the same to a bag of vectors with no order.
Attention asks which other pieces to lean on while building the meaning of the current one. Query, key, and value are three matrices, usually produced from the same inputs by different linear layers. The query is “what I am looking for,” the key is “what I can answer to,” the value is “what I hand over if I answer.”
The formula is Attention(Q, K, V) = softmax(QKᵀ / √dₖ) V. Dot products of queries and keys are a table of similarity. Dividing by the square root of the key size keeps those products from swelling. Without it, softmax sticks to a single winner and the gradient goes quiet. Row-wise softmax turns similarities into shares. Multiplying by the values builds a weighted average of the vectors that answered.
Numbers for two pieces and size 2. Queries are the rows (1, 0) and (0, 1). Keys are (1, 0) and (1, 1). Values are (1, 2) and (3, 4). After dividing by √2 and applying softmax, the first weight row is 0.50 and 0.50, the second is about 0.33 and 0.67. The outputs are (2, 3) and about (2.34, 3.34). The first piece finds both values equally interesting. The second prefers the second key. Change the input vectors and the shares move. That is the whole exercise.
Several heads are several such triples in parallel, concatenated afterward. One head can watch the neighboring word, another the subject at the start of the sentence. A residual link adds the block’s input to its output, so a deep stack is not forced to relearn the identity. Normalization holds the scale before the next block.
Training a language model usually predicts the next piece. The loss is cross-entropy on the true index in the vocabulary. Teacher forcing, feeding the true previous piece during training, keeps an early mistake from dragging the whole sentence off: during training the model is shown the previous piece. At generation there is no prompt. The model continues from its own text. The context window is how many pieces fit in one pass. Comparing every piece with every other piece costs about the square of the window length: twice the text, about four times the similarity table.
Temperature, applied before the next piece is chosen, stretches or squeezes the shares. Below one, the winner is more often picked. Above one, the tail is picked more often. It is not a newly trained number. It is a way to read scores you already computed.
A step-by-step code pass of the same scheme is in the attention-from-scratch lesson. The series itself starts in the first lab lesson.
New listings arrive after tuning
Generalization is whether the price holds on listings that were not there while you tuned it. The bias–variance trade-off names two different ways a rule goes bad. Bias is a shape so stiff that you miss even on average: one straight line on a bent market. Variance is a shape so flexible that a new pack of cards resets the numbers past recognition. A simple model is wrong in a stable way. A model that is too rich memorizes the oddities.
Calibration asks whether a stated share matches how often you are right. Nine correct out of ten where the model said “0.9” is an honest sign. Six out of ten under the same sign is a costume of confidence. Class imbalance makes accuracy a cheap win: always saying “healthy” when five people out of a hundred are sick scores 0.95 and helps nobody. Compute the metric on a held-out test pack that was touched neither while setting the numbers nor while choosing the step size.
Quantization stores numbers more coarsely. On the three apartments, weight 0.06 and bias 1.2 hit the prices. Round the weight to tenths — you get 0.1 — and the predictions become 4.2, 6.2, and 8.2. Saving digits breaks the rule because the important price per meter lived in the hundredths. In a large network many weights tolerate rounding and some do not. That is why a smaller file does not shrink quality in proportion: particular axes suffer, not “the model in general.” Numerical stability is the same story during training. Adding a huge number and a tiny one in limited precision drops the tiny one. Hence the scaling inside attention, careful losses, and formats such as half precision on an accelerator.
A local run is limited by memory, bandwidth, and whether the context window fits. The mathematics does not cancel the case and the disk. It explains why, after rounding, one task survives and another drifts. The hardware that does the arithmetic is covered separately: CPU, GPU, and neighboring chips. Moving an already trained network to another runner is in the ONNX note.
The project that closes the article is short and repeatable.
- A synthetic price sample with a known slope, bias, noise, and a fixed seed.
- A linear model in NumPy, with no training library.
- Mean squared error written by hand.
- The gradient for weight and bias written by hand.
- A loop of steps and a plot of the loss.
- The same task as a two-layer network with a
ReLUcut. - A check of that gradient against
autograd. - Error on a held-out third the loop never saw.
- Three step sizes: tiny, workable, and too large.
- A dependency file, a run command, and a table: rate, final loss, held-out error.
After that, go deeper on the part your actual task breaks, not on “all of mathematics.”
| Where the work goes | What to pick up first |
|---|---|
| Regression and classification | Linear algebra, samples, optimization |
| Convolutions and detection | Windows, tensor sizes, IoU |
| Sequence memory | Derivatives through time, gates |
| Transformers and language models | Matrices, softmax, entropy, gradients |
| Long training runs | Step size, batch statistics, numerical stability |
| Compression and local runs | Rounding, axis scale, memory |
Frequently asked questions
Do I need a full university course before the first model?
No. A line on one feature needs a function, a mean, and the idea of stepping against the rise of the error. Bring in the logarithm, Bayes, and the chain rule on the evening when the loss or the backward pass is unreadable without them.
How is an algorithm different from a trained network?
An algorithm already contains its numbers and does not move them from your examples. A trained network is a chosen shape plus numbers set from error on data. A shape without those numbers filled in is not yet a network that can do anything with your listings.
Why not just search every weight?
On the two numbers of an apartment price a grid search is still thinkable. On millions of weights the grid fits neither in time nor in memory. The gradient replaces the grid with local advice: which way each number reduces the error on this pack.
Why square the error instead of taking the absolute value?
The square pulls harder toward rare huge misses and has a smooth derivative at zero. The absolute value is calmer around an outlier and less smooth at zero. The choice is about which miss you treat as expensive, not about which formula is “more mathematical.”
Why can a deep network be worse than a short one?
Depth adds kinks, and it also adds factors on the backward pass. Extra layers on a small sample memorize the oddities in the cards. A short line is more honest when the market relationship has no bend.
What breaks in softmax if you drop the division by the square root of the size?
Dot products grow with the number of coordinates. Without the division, shares stick to one piece, the others get almost zero, and the early numbers stop hearing the error.
Is accuracy 0.95 a good model?
Only if the classes are not lopsided and the costs of the mistakes are comparable. Five sick people out of a hundred, and a constant answer “healthy,” scores 0.95 and helps nobody. Look at recall on the rare class and at the confusion matrix.
Is a hand-computed gradient still worth it?
Once, on two layers, yes. That is how you earn the right to distrust a mismatch with autograd, and how fading gradients become a mechanism instead of a slogan. Nobody writes the full gradient of a language model by hand. The mechanism is the same; the sizes are not.
Further reading
These sit next to the map: lab write-ups and comparisons of frameworks.
- Frameworks: PyTorch, TensorFlow, JAX — where the framework stops and the mathematics of a layer begins.
- PyTorch and TensorFlow — the same training pass written in two frameworks.
- An image as a tensor — pixels, axes, and sizes before convolution.
- The CNN and CRNN lab — an entry to vision and line recognition.
- Attention from scratch — query, key, and value on small matrices.
- The language-model lab — where the series starts once the attention formula looks familiar.
- ONNX and the runtime — what happens to trained numbers outside the training loop.
Conclusion
The mathematics here is an apartment estimate you can check. Area becomes a number, the rule becomes a function with a price per meter and an add-on, the miss in millions becomes the error, “nudge it up or down” becomes the gradient, and a listing from another neighborhood is the only honest test. A classifier, a convolution, a recurrence, and attention change the shape of the rule, not that order.
This week, put the three apartments into twenty lines, bring the weight to 0.06, and run one backward pass on a two-layer network, checking four gradients against PyTorch. Until those digits match, the next architecture is a list of names.



Comments