The Book · written in public

Counting to Intelligence

How Simple Mathematics Becomes Machine Learning

The math is smaller than the magic. It is just counting, repeated until it looks like thought.

About the book

Long before machine learning became a web of dense matrices and hundred-billion-parameter neural networks, it began with an operation every human understands: counting.

Counting to Intelligence does not strip the mathematics out of modern AI—it puts it where it belongs. By leading with physical intuition, equations arrive only as the precise tools needed to solve failures you’ve already seen unfold.

Written for curious readers, technical leaders, and hands-on builders who prefer clear mental models over dense notation, this book charts the path from simple ratios of counts to systems that seem to think. It shows how the most basic mathematical principles scale into the technology reshaping our world.

The math is smaller than the magic. It is just counting, repeated until it looks like thought.

From the book blog

Latest posts

All book posts →
01

Why before what

Every idea starts with the problem it solves, before any method gets a name.

02

Maths where it belongs

Equations arrive as the precise tools for failures you’ve already watched unfold — never as the opening move.

03

Foundations to MLOps

From the maths you need, through deep learning, to running models in production.

Contents

7 parts, 33 chapters

13 of 33 chapters drafted. Open a chapter to see its sections. Chapters link to the blog once their section posts go live.

  • Published on the blog
  • Drafted
  • Planned
Part 1 Before the Code: Building Intuition 3 ch.
  1. 01 What is Machine Learning, Really? Drafted
    1. 1.1 AI, ML, Deep Learning — drawing the map The three circles · The map: continent, country, capital · The shift: from rules to patterns
    2. 1.2 Learning from examples vs programming the rules The rule writer: a spam filter · The pattern finder · The shift: from deduction to induction
    3. 1.3 The three types of machine learning Supervised learning: the teacher · Unsupervised learning: the explorer · Reinforcement learning: the puppy
    4. 1.4 How a child learns The "that is a dog" moment · Generalisation vs memorisation · Where the analogy breaks
    5. 1.5 Why now? The perfect storm The brilliant architects · The connectionists' bet · The two AI winters · Data: the fuel · Compute: the engine · Algorithms: the blueprint · The ImageNet moment
    6. 1.6 What is inside the box? The panel of knobs · The four-part scaffold
  2. 02 How Machines "See" the World: Data and Features Drafted
    1. 2.1 All of it is numbers The data that already looks like data · Everything else becomes numbers too · Why unstructured data is the harder problem
    2. 2.2 Features, labels, and observations Three words for one table · A row becomes a point · When features were the whole job
    3. 2.3 The honest scorecard: training, validation, test The exam you wrote yourself · Why two held-out piles, not one · Ratios, and when to break them
    4. 2.4 Data quality over quantity The outlier problem · Label noise · Class imbalance, and why accuracy lies · Data leakage: the silent killer
  3. 03 The Learning Loop: How a Model Improves Drafted
    1. 3.1 Predictions, errors, and correction The prediction: what the model currently believes · The error: how wrong was that? · The update: nudge, don't lurch
    2. 3.2 Loss: a single number for wrongness Why one number · Building one: the regression case · Building one: when the answer is a probability · The landscape: where the loop lives
    3. 3.3 Overfitting and underfitting A perfect score and a useless model · Too simple, too flexible, just right · Reading the curves
    4. 3.4 The bias-variance tradeoff Two kinds of wrong · Why it is a tradeoff
  4. I Interlude · Anatomy of an ML Project Planned

    An unnumbered interlude closing Part I: the minimal end-to-end project lifecycle — business problem, data, baseline, iterate, deploy, monitor — grounding every later technique in a business context.

    1. I.1 From business problem to ML problem
    2. I.2 Data collection and exploratory analysis
    3. I.3 Baseline first: feature engineering and a simple model
    4. I.4 Iterate, deploy, monitor

    Planned sections — from the working outline.

Part 2 The Math You Actually Need 4 ch.
  1. 04 Probability — The Language of Uncertainty Drafted
    1. 4.1 Trials, sample spaces, outcomes, and events
    2. 4.2 Joint, marginal, and conditional probability
    3. 4.3 Bayes' theorem — the most important formula in ML
    4. 4.4 Likelihood — reading the same grid downwards
  2. 05 Linear Algebra: The Geometry of Data Planned
    1. 5.1 Scalars, vectors, matrices — what they mean physically
    2. 5.2 Dot products as similarity
    3. 5.3 Matrix multiplication as transformation
    4. 5.4 Eigenvalues and eigenvectors
    5. 5.5 SVD — every matrix is rotate · stretch · rotate
    6. 5.6 Norms and distances
    7. 5.7 Extension: LU decomposition, triangular systems & pivoting
    8. 5.8 Extension: Matrix multiplication methods & the cost of matrix operations

    Planned sections — from the working outline.

  3. 06 Calculus: The Math of Change Drafted
    1. 6.1 Functions, slopes, and the idea of change Functions, the math kind · Slope, or the rate of change · Average vs instantaneous rate of change · The limit, and the trick that makes calculus work · When the derivative doesn't exist: jumps and corners · The derivative, in symbols
    2. 6.2 Commonly used functions in machine learning Two big families: regression and classification · Squared and absolute error: scoring numerical predictions · Cross-entropy: scoring probabilistic predictions · A brief tour of activation functions · The forms, in symbols
    3. 6.3 Second derivatives, curvature, and the Hessian The slope of the slope · Curvature, and why the optimiser cares · Many directions, and the saddle problem · The Hessian: the full curvature ledger · Why nobody writes the ledger down · The second derivative and the Hessian, in symbols
    4. 6.4 Partial derivatives, gradients, and the Jacobian One knob at a time: the partial derivative · Assembling the compass: the gradient · When outputs multiply: the Jacobian · The gradient and the Jacobian, in symbols
  4. 07 Statistics: Making Sense of Data Drafted
    1. 7.1 Mean, median, mode — and when each lies Three different answers to "what is typical" · Where the average lies · The spread around the centre
    2. 7.2 Hypothesis testing basics What a test actually asks · The trap of many tests · Allowed to be wrong, by a known amount
    3. 7.3 The central limit theorem Why so many things wear the same shape · One question deeper: why bells are the default view of high-dimensional data
    4. 7.4 Extension: The bias of the MLE variance and Bessel's correction The sample mean cheats in its own favour · The fix, and why a known fix still matters
Part 3 Classical Machine Learning 7 ch.
  1. 08 Linear Regression: Fitting a Line to Life Drafted
    1. 8.1 The equation of a line The target the machine is chasing · The trick that makes the line bend
    2. 8.2 Least squares: minimising squared errors The shadow on the wall · How good is the answer?
    3. 8.3 Multiple regression and interaction terms
    4. 8.4 Regularisation: Ridge and Lasso The penalty that selects
    5. 8.5 Extension: Generalised Additive Models
    6. 8.6 Extension: Bayesian linear regression and the predictive ellipsoid
  2. 09 Logistic Regression: When the Output Is a Probability Drafted
    1. 9.1 Why a straight line cannot answer a yes-or-no question
    2. 9.2 The sigmoid — squashing a score into a probability
    3. 9.3 Log loss in practice
    4. 9.4 Decision boundaries
    5. 9.5 When there are more than two classes — softmax and one-vs-rest
    6. 9.6 Extension: LDA, QDA, RDA and the generative-versus-discriminative divide
    7. 9.7 Extension: classifiers for the high-dimensional regime — diagonal LDA and nearest shrunken centroids
    8. 9.8 Extension: Newton-Raphson and iteratively reweighted least squares
  3. 10 Decision Trees: Asking the Right Questions Drafted
    1. 10.1 How a Tree Splits Data
    2. 10.2 Information Gain and Gini Impurity
    3. 10.3 How Deep Is Too Deep
    4. 10.4 Pruning Strategies
    5. 10.5 Extension: PRIM — Bump Hunting
  4. 11 Ensemble Methods: The Wisdom of the Crowd Drafted
    1. 11.1 Bagging and Random Forests The correlation ceiling · What the forest knows about your data
    2. 11.2 Boosting: AdaBoost and Gradient Boosting Gradient descent in the space of functions
    3. 11.3 XGBoost and LightGBM: The Second-Order Engine
    4. 11.4 Stacking
    5. 11.5 Why Ensembles Almost Always Win
    6. 11.6 Extension: Bumping
    7. 11.7 Extension: Mixtures of Experts and Hierarchical Mixtures
  5. 12 Support Vector Machines: Finding the Widest Street Drafted
    1. 12.1 The Maximum-Margin Classifier
    2. 12.2 The Kernel Trick: Lifting to Higher Dimensions
    3. 12.3 Soft Margins for Noisy Data
    4. 12.4 The Dual Problem and the KKT Conditions
    5. 12.5 Extension: The Gamma Knob, Sparsity, SVR, and Multiclass SVMs
    6. 12.6 Extension: Why Kernel SVMs Don't Scale — and the Linear-SVC Escape Hatch
    7. 12.7 Extension: Gaussian Processes — the Kernel Ridge Cousin
    8. 12.8 Extension: Relevance Vector Machines
  6. 13 Nearest Neighbours and Naive Bayes: The Power of Doing Almost Nothing Drafted
    1. 13.1 Distance-Based Learning: k-Nearest Neighbours
    2. 13.2 Naive Bayes and the Independence Assumption
    3. 13.3 Extension: Kernel Smoothing and the Smoother Matrix
    4. 13.4 Extension: Local Regression and Local Likelihood
    5. 13.5 Extension: Kernel Density Estimation and Bandwidth Selection
    6. 13.6 Extension: Approximate Nearest Neighbours
  7. 14 Time Series Analysis & Forecasting Planned
    1. 14.1 Thinking in time: trend, seasonality, and stationarity
    2. 14.2 Autocorrelation and lag features
    3. 14.3 Classical models: ARIMA, SARIMAX, and Exponential Smoothing
    4. 14.4 Modern approaches: Prophet, XGBoost with lag features, Temporal Fusion Transformers
    5. 14.5 Validation in time: why standard K-Fold CV fails
    6. 14.6 Forecast evaluation metrics: MASE and sMAPE
    7. 14.7 Extension: Modern deep forecasting — DeepAR and N-BEATS

    Planned sections — from the working outline.

Part 4 Unsupervised Learning & Recommender Systems 4 ch.
  1. 15 Clustering — Finding Tribes in Data Planned
    1. 15.1 k-Means: the algorithm and its assumptions
    2. 15.2 Choosing k: the elbow method and silhouette scores
    3. 15.3 DBSCAN: clusters of arbitrary shape
    4. 15.4 Hierarchical clustering & dendrograms
    5. 15.5 Gaussian Mixture Models
    6. 15.6 Extension: Identifying and validating latent variables
    7. 15.7 Extension: K-Medoids, PAM, CLARA
    8. 15.8 Extension: Self-Organizing Maps (SOM)
    9. 15.9 Extension: Vector Quantization and mixed-variable clustering
    10. 15.10 Extension: Association rules — Apriori, confidence, lift
    11. 15.11 Extension: Spectral clustering — arbitrary shapes via eigenvectors
    12. 15.12 Extension: Semi-supervised learning — a few labels, lots of data

    Planned sections — from the working outline.

  2. 16 Dimensionality Reduction — Seeing the Big Picture Planned
    1. 16.1 The curse of dimensionality
    2. 16.2 PCA: projecting to the directions that matter
    3. 16.3 What eigenvalues tell us about variance
    4. 16.4 t-SNE for visualisation
    5. 16.5 UMAP: fast and scalable
    6. 16.6 Extension: Factor Analysis — continuous latent variables
    7. 16.7 Extension: Probabilistic PCA, Kernel PCA, and Bayesian PCA
    8. 16.8 Extension: Independent Component Analysis (ICA)
    9. 16.9 Extension: Multidimensional Scaling (MDS) and Principal Curves
    10. 16.10 Extension: Supervised PCA
    11. 16.11 Extension: Non-negative Matrix Factorization (NMF)

    Planned sections — from the working outline.

  3. 17 Anomaly Detection — Finding the Oddball Planned
    1. 17.1 Statistical methods: z-score and IQR
    2. 17.2 Isolation forests
    3. 17.3 Autoencoders for anomaly detection
    4. 17.4 Applications: fraud, fault detection
    5. 17.5 Evaluation without labels
    6. 17.6 Extension: Density-based anomaly detection with KDE
    7. 17.7 Extension: Local Outlier Factor (LOF)

    Planned sections — from the working outline.

  4. 18 Recommendation Systems & Search Planned
    1. 18.1 Collaborative filtering: user–user vs. item–item
    2. 18.2 Content-based filtering and the cold-start problem
    3. 18.3 Matrix Factorization: SVD and ALS
    4. 18.4 Two-Tower Neural Networks
    5. 18.5 Candidate generation vs. re-ranking pipelines
    6. 18.6 Extension: The matrix factorization loss — explicit vs. implicit feedback

    Planned sections — from the working outline.

Part 5 Neural Networks & Deep Learning 7 ch.
  1. 19 The Neuron — Nature's Simplest Computer Planned
    1. 19.1 Biological neuron → perceptron analogy
    2. 19.2 Weights, bias, and activation
    3. 19.3 Why non-linearity is everything
    4. 19.4 Common activations: ReLU, Sigmoid, Tanh, GELU
    5. 19.5 Universal approximation theorem

    Planned sections — from the working outline.

  2. 20 Feedforward Networks — Layers of Abstraction Planned
    1. 20.1 Input, hidden, output layers
    2. 20.2 Forward pass: data flowing through
    3. 20.3 Backpropagation: credit assignment
    4. 20.4 Vanishing & exploding gradients — the full treatment
    5. 20.5 Why sigmoid saturates and kills gradients
    6. 20.6 Why ReLU largely solved the problem
    7. 20.7 Weight initialisation strategies
    8. 20.8 Extension: A worked code example for forward and backward pass

    Planned sections — from the working outline.

  3. 21 Training Deep Networks — The Practitioner's Guide Planned
    1. 21.1 Mini-batch gradient descent in practice
    2. 21.2 Optimisers: SGD, Momentum, Adam revisited
    3. 21.3 Learning rate schedules and warm-up
    4. 21.4 Dropout: accidental regulariser that works
    5. 21.5 Batch normalisation and layer normalisation
    6. 21.6 Early stopping
    7. 21.7 Common failure modes and how to diagnose them
    8. 21.8 Data augmentation: more data for free
    9. 21.9 Extension: Bayesian Neural Networks and uncertainty
    10. 21.10 Extension: Theoretical frontier — lottery tickets, NTK, double descent, ensembling
    11. 21.11 Extension: Why deep networks are hard to train — eigenvalues, singular values, gradients
    12. 21.12 Extension: Why deep learning works — trainability, generalization & depth

    Planned sections — from the working outline.

  4. 22 Convolutional Networks — How Machines See Planned
    1. 22.1 The convolution operation visualised
    2. 22.2 Filters as feature detectors
    3. 22.3 Pooling and spatial hierarchy
    4. 22.4 Classic architectures: LeNet, VGG, ResNet
    5. 22.5 Transfer learning and fine-tuning
    6. 22.6 Extension: Object detection
    7. 22.7 Extension: Semantic segmentation and U-Net
    8. 22.8 Extension: Component mix-and-match — the CNN building-block catalogue
    9. 22.9 Extension: Vision Transformers — image patches as tokens

    Planned sections — from the working outline.

  5. 23 Recurrent Networks — Memory in Machines Planned
    1. 23.1 Sequences and time steps
    2. 23.2 Vanilla RNN and the vanishing gradient problem
    3. 23.3 LSTMs: gates explained intuitively
    4. 23.4 GRUs: simpler but powerful
    5. 23.5 Applications: time series, language modelling
    6. 23.6 Extension: Hidden Markov Models and Linear Dynamical Systems (Kalman)

    Planned sections — from the working outline.

  6. 24 Attention — Teaching Machines to Connect the Dots Drafted
    1. 24.1 Order without recurrence
    2. 24.2 What a token is looking for
    3. 24.3 Scoring, routing, retrieving
    4. 24.4 More than one relationship at a time
    5. 24.5 Why this is not a recurrent network
    6. 24.6 Communication, then computation
    7. 24.7 Writing the output, one token at a time
    8. 24.8 Content and position
  7. 25 Graph Neural Networks — Learning on Nodes & Edges Planned
    1. 25.1 The three challenges of learning on graphs
    2. 25.2 Types of graphs
    3. 25.3 Representing a graph: adjacency matrix, nodes, and edges
    4. 25.4 Walks, paths, and powers of the adjacency matrix
    5. 25.5 Node ordering, permutation invariance, and equivariance
    6. 25.6 Node embeddings and the GNN processing pipeline
    7. 25.7 Graph-level, node-level, and edge-prediction tasks
    8. 25.8 Spatial graph convolution and message passing
    9. 25.9 Higher-order (n-hop) aggregation
    10. 25.10 Aggregation functions and Graph Isomorphism Networks
    11. 25.11 Variants of aggregation and combination
    12. 25.12 Training on graphs: batching, inductive vs transductive settings
    13. 25.13 Edge embeddings

    Planned sections — from the working outline.

Part 6 Modern Machine Learning 4 ch.
  1. 26 Large Language Models — Predicting the Next Word at Scale Planned
    1. 26.1 Pre-training and self-supervised learning
    2. 26.2 From words to numbers: bag-of-words, TF-IDF, and n-grams
    3. 26.3 Meaning as geometry: static word embeddings — Word2Vec and GloVe
    4. 26.4 Tokenisation and vocabulary
    5. 26.5 Scaling laws: bigger is better, until it isn't
    6. 26.6 Fine-tuning vs prompting vs RAG
    7. 26.7 Emergent abilities: what happens at scale
    8. 26.8 Hallucinations and why they're structural
    9. 26.9 Extension: Agents, tool use, and chain-of-thought reasoning
    10. 26.10 Extension: Mixture of Experts — how frontier models scale beyond dense

    Planned sections — from the working outline.

  2. 27 Generative Models — Making New Things Planned
    1. 27.1 VAEs: encoding meaning into a latent space
    2. 27.2 GANs: the generator vs discriminator game
    3. 27.3 Diffusion models: denoising step by step
    4. 27.4 Sampling and temperature
    5. 27.5 Image, audio, code, video generation
    6. 27.6 Extension: Mixture Density Networks
    7. 27.7 Extension: Probabilistic graphical models — undirected (MRFs)
    8. 27.8 Extension: Probabilistic graphical models — directed (Bayesian Networks)
    9. 27.9 Extension: Inference in graphical models — factor graphs, sum-product, junction tree
    10. 27.10 Extension: Variational Inference and the reparameterisation trick
    11. 27.11 Extension: Generative-model taxonomy & evaluation — FID, IS, β-VAE
    12. 27.12 Extension: Normalizing Flows
    13. 27.13 Extension: Laplace approximation and Bayesian inference machinery
    14. 27.14 Extension: MCMC for sampling from the posterior

    Planned sections — from the working outline.

  3. 28 Reinforcement Learning — Learning by Doing Planned
    1. 28.1 Agents, environments, rewards, episodes
    2. 28.2 The exploration-exploitation tradeoff
    3. 28.3 Q-learning intuitively
    4. 28.4 Policy gradients
    5. 28.5 RLHF: aligning language models with humans
    6. 28.6 Extension: Model-based RL — AlphaZero, MuZero, world models
    7. 28.7 Extension: Reinforcement learning — MDPs, value methods & policy gradients

    Planned sections — from the working outline.

  4. 29 Multi-Modal Models — Seeing, Hearing and Reading Together Planned
    1. 29.1 Embedding different modalities into shared space
    2. 29.2 CLIP: connecting images and text
    3. 29.3 Vision-language models
    4. 29.4 Speech + text + image pipelines
    5. 29.5 The trend toward unified models

    Planned sections — from the working outline.

Part 7 ML in Practice 4 ch.
  1. 30 Feature Engineering — The Craft Behind the Science Planned
    1. 30.1 Scaling, normalisation, standardisation
    2. 30.2 Encoding categoricals: ordinal, one-hot, target encoding
    3. 30.3 Handling missing data
    4. 30.4 Feature crosses and polynomial features
    5. 30.5 Feature selection: filter, wrapper, embedded methods
    6. 30.6 Imbalanced Data & Cost-Sensitive Learning
    7. 30.7 Common leakage pitfalls
    8. 30.8 Extension: When p ≫ N — the high-dimensional regime

    Planned sections — from the working outline.

  2. 31 Experiment Design & Model Evaluation Planned
    1. 31.1 Cross-validation strategies
    2. 31.2 Confusion matrix deep dive
    3. 31.3 Calibration: do your probabilities mean what you think?
    4. 31.4 A/B testing basics
    5. 31.5 Statistical significance and practical significance
    6. 31.6 The Bootstrap and confidence intervals
    7. 31.7 Translating Model Metrics to Financial Value
    8. 31.8 Extension: Model selection — the why and the how
    9. 31.9 Extension: Optimism of the training error and the in-sample criteria (Cp / AIC / BIC / MDL)
    10. 31.10 Extension: Effective degrees of freedom
    11. 31.11 Extension: VC dimension and Structural Risk Minimization
    12. 31.12 Extension: Bayesian model comparison and the evidence approximation
    13. 31.13 Extension: Active learning
    14. 31.14 Extension: Causal inference — beyond correlation

    Planned sections — from the working outline.

  3. 32 MLOps — Keeping Models Alive Planned
    1. 32.1 The model lifecycle
    2. 32.2 Data drift and concept drift
    3. 32.3 Model monitoring and alerts
    4. 32.4 CI/CD for ML pipelines
    5. 32.5 Experiment tracking: MLflow, Weights & Biases
    6. 32.6 Feature stores and model registries
    7. 32.7 Parameter-Efficient Fine-Tuning (PEFT) & Quantization
    8. 32.8 Online learning and incremental updates
    9. 32.9 Extension: Distributed training
    10. 32.10 Extension: Model serving — quantisation, distillation, caching
    11. 32.11 Extension: The continuous-learning loop
    12. 32.12 Extension: AutoML, NAS, and HPO

    Planned sections — from the working outline.

  4. 33 Ethics, Fairness & Responsible AI Planned
    1. 33.1 Bias in data and in models
    2. 33.2 Fairness metrics and their tradeoffs
    3. 33.3 Model explainability: SHAP, LIME, PDP and ALE
    4. 33.4 Privacy: federated learning, differential privacy
    5. 33.5 Regulation and responsible deployment
    6. 33.6 The practitioner's moral checklist
    7. 33.7 Adversarial ML & LLM Security
    8. 33.8 Extension: Mesa-optimisation and goal misgeneralisation
    9. 33.9 Extension: From attribution to circuits — a tour of mechanistic interpretability

    Planned sections — from the working outline.

  1. E Epilogue · The Road Ahead Planned

    Where the field is going: ML as an instrument of science, and the open problems between today's models and general intelligence.

    1. E.1 AI for scientific discovery
    2. E.2 Open problems on the path to AGI and societal implications

    Planned sections — from the working outline.