Machine Learning · Statistics · Linear Algebra
Working notes on machine learning, one chapter at a time.
I’m Amit—Head of Engineering, currently leading the build of an agentic AI platform that reads, understands and validates documents at enterprise scale. My background is in ML and data science, and these pages contain my chapter-by-chapter notes and complete derivations for the foundational textbooks behind modern AI: Bishop, Strang, ISLR, and Think Stats.
Currently, I’m turning these working notes into my own book, Counting to Intelligence.
The Book · writing in public
Counting to Intelligence
How Simple Mathematics Becomes Machine Learning
The math is smaller than the magic. It is just counting, repeated until it looks like thought. I’m publishing it here section by section, as short posts.
01 — Series
Read a book with me
Each series follows one textbook in order. Start at chapter one, or jump to the topic you need.
Counting to Intelligence
How Simple Mathematics Becomes Machine Learning
Pattern Recognition and Machine Learning
Probability, linear models, neural networks, kernels, graphical models and EM.
Linear Algebra
Subspaces, projections, least squares, determinants, eigenvalues and the SVD.
An Introduction to Statistical Learning
Regression, classification, resampling, trees, SVMs and unsupervised learning — with exercises.
Think Stats
Distributions, estimation and hypothesis testing for programmers.
02 — Recent notes
From the notebook
- The Mistake That Makes Gradient Descent Work Counting to Intelligence
- Data Cleaning for Machine Learning Systems: A Survey Data Cleaning for Machine Learning Sysytems: A Survey Note
- Mixture Models and Expectation Maximization - The EM Algorithm in General Bishop · Chapter 9 — Mixture Models and Expectation Maximization Bishop
- Mixture Models and Expectation Maximization - An Alternative View of EM Bishop · Chapter 9 — Mixture Models and Expectation Maximization Bishop
- Mixture Models and Expectation Maximization - Mixtures of Gaussians Bishop · Chapter 9 — Mixture Models and Expectation Maximization Bishop
- Mixture Models and Expectation Maximization - K-means Clustering Bishop · Chapter 9 — Mixture Models and Expectation Maximization Bishop
- Graphical Models - The Sum-product Algorithm, The Max-Sum Algorithm Bishop · Chapter 8 — Graphical Models Bishop
- Graphical Models - Inference in Graphical Models Bishop · Chapter 8 — Graphical Models Bishop
- Graphical Models - Markov Random Fields Bishop · Chapter 8 — Graphical Models Bishop
- Graphical Models - Conditional Independence Bishop · Chapter 8 — Graphical Models Bishop
- Graphical Models - Bayesian Networks Bishop · Chapter 8 — Graphical Models Bishop
- Sparse Kernel Methods - Maximum Margin Classifiers: Relation to Logistic Regression, Multiclass SVMs, SVMs for Regression Bishop · Chapter 7 — Sparse Kernel Methods Bishop
- Sparse Kernel Methods - Maximum Margin Classifiers: Overlapping Class Distributions Bishop · Chapter 7 — Sparse Kernel Methods Bishop
- Sparse Kernel Methods - Maximum Margin Classifiers Bishop · Chapter 7 — Sparse Kernel Methods Bishop
- Sparse Kernel Methods - Lagrange Multipliers Bishop · Chapter 7 — Sparse Kernel Methods Bishop
- Kernel Methods - Gaussian Process Bishop · Chapter 6 — Kernel Methods Bishop
- Kernel Methods - Constructing Kernels & Radial Basis Function Networks Bishop · Chapter 6 — Kernel Methods Bishop
- Kernel Methods - Dual Representations Bishop · Chapter 6 — Kernel Methods Bishop
- Neural Networks - Mixture Density Networks & Bayesian Neural Networks Bishop · Chapter 5 — Neural Networks Bishop
- Neural Networks - Regularization in Neural Networks Bishop · Chapter 5 — Neural Networks Bishop
- Neural Networks - The Hessian Matrix Bishop · Chapter 5 — Neural Networks Bishop
- Neural Networks - Error Backpropagation Bishop · Chapter 5 — Neural Networks Bishop
- Neural Networks - Network Training Bishop · Chapter 5 — Neural Networks Bishop
- Neural Networks - Feed-forward Network Functions Bishop · Chapter 5 — Neural Networks Bishop
- Linear Models for Classification - The Laplace Approximation & Bayesian Logistic Regression Bishop · Chapter 4 — Linear Models for Classification Bishop
- Linear Models for Classification - Probabilistic Discriminative Models Bishop · Chapter 4 — Linear Models for Classification Bishop
- Linear Models for Classification - Probabilistic Generative Models (Maximum Likelihood Solution) Bishop · Chapter 4 — Linear Models for Classification Bishop
- Linear Models for Classification - Probabilistic Generative Models Bishop · Chapter 4 — Linear Models for Classification Bishop
- Linear Models for Classification - The Perceptron Algorithm Bishop · Chapter 4 — Linear Models for Classification Bishop
- Linear Models for Classification - Fisher’s Linear Discriminant Bishop · Chapter 4 — Linear Models for Classification Bishop
- Linear Models for Classification - Least Squares for Classification Bishop · Chapter 4 — Linear Models for Classification Bishop
- Linear Models for Classification - Discriminant Functions (Part 2) Bishop · Chapter 4 — Linear Models for Classification Bishop
- Linear Models for Classification - Discriminant Functions Bishop · Chapter 4 — Linear Models for Classification Bishop
- Linear Models for Regression - Evidence Approximation & Limitations of Fixed Basis Function Bishop · Chapter 3 — Linear Models for Regression Bishop
- Linear Models for Regression - Bayesian Model Comparison Bishop · Chapter 3 — Linear Models for Regression Bishop
- Linear Models for Regression - Bayesian Linear Regression Bishop · Chapter 3 — Linear Models for Regression Bishop
- Linear Models for Regression - Bias-Variance Decomposition Bishop · Chapter 3 — Linear Models for Regression Bishop
- Linear Models for Regression - Linear Basis Function Models : Part 2 Bishop · Chapter 3 — Linear Models for Regression Bishop
- Linear Models for Regression - Linear Basis Function Models : Part 1 Bishop · Chapter 3 — Linear Models for Regression Bishop
- Probability Distributions - Nonparametric Methods Bishop · Chapter 2 — Probability Distributions Bishop
- Probability Distributions - The Exponential Family Bishop · Chapter 2 — Probability Distributions Bishop
- Probability Distributions - The Gaussian Distribution: Part 5 Bishop · Chapter 2 — Probability Distributions Bishop
- Probability Distributions - The Gaussian Distribution: Part 4 Bishop · Chapter 2 — Probability Distributions Bishop
- Probability Distributions - The Gaussian Distribution: Part 3 Bishop · Chapter 2 — Probability Distributions Bishop
- Probability Distributions - The Gaussian Distribution: Part 2 Bishop · Chapter 2 — Probability Distributions Bishop
- Probability Distributions - The Gaussian Distribution: Part 1 Bishop · Chapter 2 — Probability Distributions Bishop
- Probability Distributions - Multinomial Variables Bishop · Chapter 2 — Probability Distributions Bishop
- Probability Distributions - Binary Variables Bishop · Chapter 2 — Probability Distributions Bishop
- Introduction - Information Theory Bishop · Chapter 1 — Introduction Bishop
- Introduction - Decision Theory Bishop · Chapter 1 — Introduction Bishop
- Introduction - Model Selection & Curse of Dimensionality Bishop · Chapter 1 — Introduction Bishop
- Introduction - Probability Theory Bishop · Chapter 1 — Introduction Bishop
- Introduction - Polynomial Curve Fitting Bishop · Chapter 1 — Introduction Bishop
- Left, Right and Pseudo Inverses Strang · Chapter 28 Strang
- Linear Transformations, Change of Basis and Image Compression Strang · Chapter 27 Strang
- Singular Value Decomposition Strang · Chapter 26 Strang
- Similar Matrices Strang · Chapter 25 Strang
- Positive Definite Matrices Strang · Chapter 24 Strang
- Complex Matrices and Fourier Transform Strang · Chapter 23 Strang
- Symmetric Matrices and Positive Definiteness Strang · Chapter 22 Strang
- Markov Matrices and Fourier Series Strang · Chapter 21 Strang
- Differential Equations and Matrix Exponentials Strang · Chapter 20 Strang
- Diagonalization and Powers of a Matrix Strang · Chapter 19 Strang
- Eigenvalues and Eigenvectors Strang · Chapter 18 Strang
- Formula for $A^{-1}$ and Cramer's Rule Strang · Chapter 17 Strang
- Determinant and Cofactors Strang · Chapter 16 Strang
- Determinant Strang · Chapter 15 Strang
- Orthonormal Vectors, Orthogonal Matrices and Gram-Schmidt Method Strang · Chapter 14 Strang
- Projection Matrices and Least Squares Strang · Chapter 13 Strang
- Projection of a Matrix Strang · Chapter 12 Strang
- Orthogonal Vectors and Orthogonal Subspaces Strang · Chapter 11 Strang
- Graphs, Networks and Incidence Matrices Strang · Chapter 10 Strang
- Matrix Spaces Strang · Chapter 9 Strang
- Four Fundamental Subspaces Strang · Chapter 8 Strang
- Matrix Independence, Span, Basis & Dimension Strang · Chapter 7 Strang
- Algorithm for solving $Ax=b$ Strang · Chapter 6 Strang
- Algorithm for solving $Ax=0$ Strang · Chapter 5 Strang
- Vector Space and Subspace Strang · Chapter 4 Strang
- Inverse of a Matrix & Factorization into $A=LU$ Strang · Chapter 3 Strang
- Elimination & Permutation with Matrices Strang · Chapter 2 Strang
- Geometry of Linear Equations & Matrix Multiplications Strang · Chapter 1 Strang
- The Wilcoxon Signed-Rank Test The Wilcoxon Signed-Rank Test: Derivation of Mean and Variance Note
- Logistic Regression Logistic Regression: Derivation Note
- Hypothesis Testing (Part 6) Tests for Variances and Power of a Test Note
- Hypothesis Testing (Part 5) Tests with Categorical Data & Tests for Homogeneity and Independence Note
- Hypothesis Testing (Part 4) Distribution-Free Tests Note
- Hypothesis Testing (Part 3) Tests for the Difference Between Two Means (Large and Small Samples) and Tests with Paired Data Note
- Hypothesis Testing (Part 2) Tests for a Population Proportion Note
- Hypothesis Testing (Part 1) Tests for a Population Mean (Large and Small Samples) Note
- Confidence Intervals (Part 3) Confidence Intervals with Paired Data and Population Variance/ Prediction Intervals Note
- Confidence Intervals (Part 2) Confidence Intervals for Proportions and the Difference Note
- Confidence Intervals (Part 1) Confidence Intervals for a Population Mean Note
- Commonly used Distributions (Part 2) Commonly used Distributions Note
- Commonly used Distributions (Part 1) Commonly used Distributions Note
- Measurement and Propagation of Error (Part 2) Measurement and Propagation of Error Note
- Measurement and Propagation of Error (Part 1) Measurement and Propagation of Error Note
- Random Variables (Part 3: Jointly Distributed Random Variables) Jointly Distributed Random Variables Note
- Random Variables (Part 2: Continuous Random Variables) Continuous Random Variables Note
- Random Variables (Part 1: Discrete Random Variables) Discrete Random Variables Note
- Performance Metrics for Classification Algorithms Performance Metrics for Classification Algorithms Note
- Hypothesis testing Hypothesis testing Note
- Maximum Likelihood Estimation Estimation Note
- Naive Bayes Classifier Classification Note
- Correlation Think Stats · Chapter 9 Think Stats
- Estimation Think Stats · Chapter 8 Think Stats
- Hypothesis Testing Think Stats · Chapter 7 Think Stats
- Operations on Distributions Think Stats · Chapter 6 Think Stats
- Probability Think Stats · Chapter 5 Think Stats
- Continuous Distributions Think Stats · Chapter 4 Think Stats
- Cumulative Distribution Functions Think Stats · Chapter 3 Think Stats
- Descriptive Statistics Think Stats · Chapter 2 Think Stats
- Statistical Thinking for Programmers Think Stats · Chapter 1 Think Stats
- Content Based Movie Recommendation Engine Content based recommendation engine Note
- Unsupervised Learning: Applied Exercises ISLR · Chapter 10 — Unsupervised Learning ISLR
- Unsupervised Learning: Conceptual Exercises ISLR · Chapter 10 — Unsupervised Learning ISLR
- Hierarchical Clustering ISLR · Chapter 10 — Unsupervised Learning ISLR
- K-Means Clustering ISLR · Chapter 10 — Unsupervised Learning ISLR
- Principal Components Analysis: More on PCA ISLR · Chapter 10 — Unsupervised Learning ISLR
- Principal Components Analysis ISLR · Chapter 10 — Unsupervised Learning ISLR
- Support Vector Machines: Applied Exercises ISLR · Chapter 9 — Support Vector Machines ISLR
- Support Vector Machines: Conceptual Exercises ISLR · Chapter 9 — Support Vector Machines ISLR
- Support Vector Machines and Kernels ISLR · Chapter 9 — Support Vector Machines ISLR
- Support Vector Classifiers ISLR · Chapter 9 — Support Vector Machines ISLR
- Maximal Margin Classifier ISLR · Chapter 9 — Support Vector Machines ISLR
- Tree-Based Methods: Applied Exercises ISLR · Chapter 8 — Tree-Based Methods ISLR
- Tree-Based Methods: Conceptual Exercises ISLR · Chapter 8 — Tree-Based Methods ISLR
- Bagging, Random Forests, Boosting ISLR · Chapter 8 — Tree-Based Methods ISLR
- Decision Trees ISLR · Chapter 8 — Tree-Based Methods ISLR
- Moving Beyond Linearity: Applied Exercises ISLR · Chapter 7 — Moving Beyond Linearity ISLR
- Moving Beyond Linearity: Conceptual Exercises ISLR · Chapter 7 — Moving Beyond Linearity ISLR
- Local Regression, Generalized Additive Models ISLR · Chapter 7 — Moving Beyond Linearity ISLR
- Smoothing Splines ISLR · Chapter 7 — Moving Beyond Linearity ISLR
- Regression Splines ISLR · Chapter 7 — Moving Beyond Linearity ISLR
- Polynomial Regression, Step Functions, Basis Functions ISLR · Chapter 7 — Moving Beyond Linearity ISLR
- Linear Model Selection and Regularization: Applied Exercises ISLR · Chapter 6 — Linear Model Selection and Regularization ISLR
- Linear Model Selection and Regularization: Conceptual Exercises ISLR · Chapter 6 — Linear Model Selection and Regularization ISLR
- Dimension Reduction Methods ISLR · Chapter 6 — Linear Model Selection and Regularization ISLR
- Shrinkage Methods ISLR · Chapter 6 — Linear Model Selection and Regularization ISLR
- Subset Selection ISLR · Chapter 6 — Linear Model Selection and Regularization ISLR
- Resampling Methods: Applied Exercises ISLR · Chapter 5 — Resampling Methods ISLR
- Resampling Methods: Conceptual Exercises ISLR · Chapter 5 — Resampling Methods ISLR
- The Bootstrap ISLR · Chapter 5 — Resampling Methods ISLR
- Cross-Validation ISLR · Chapter 5 — Resampling Methods ISLR
- Classification: Applied Exercises ISLR · Chapter 4 — Classification ISLR
- Classification: Conceptual Exercises ISLR · Chapter 4 — Classification ISLR
- Linear Discriminant Analysis ISLR · Chapter 4 — Classification ISLR
- Logistic Regression ISLR · Chapter 4 — Classification ISLR
- Linear Regression: Applied Exercises ISLR · Chapter 3 — Linear Regression ISLR
- Linear Regression: Conceptual Exercises ISLR · Chapter 3 — Linear Regression ISLR
- Other Considerations in the Regression Model ISLR · Chapter 3 — Linear Regression ISLR
- Multiple Linear Regression ISLR · Chapter 3 — Linear Regression ISLR
- Simple Linear Regression ISLR · Chapter 3 — Linear Regression ISLR
- Statistical Learning: Applied Exercises ISLR · Chapter 2 — Statistical Learning ISLR
- Statistical Learning: Conceptual Exercises ISLR · Chapter 2 — Statistical Learning ISLR
- Assessing Model Accuracy ISLR · Chapter 2 — Statistical Learning ISLR
- What Is Statistical Learning? ISLR · Chapter 2 — Statistical Learning ISLR
- Introduction to Statistical Learning ISLR · Chapter 1 — Introduction ISLR