Machine Learning · Statistics · Linear Algebra

Working notes on machine learning, one chapter at a time.

I’m Amit—Head of Engineering, currently leading the build of an agentic AI platform that reads, understands and validates documents at enterprise scale. My background is in ML and data science, and these pages contain my chapter-by-chapter notes and complete derivations for the foundational textbooks behind modern AI: Bishop, Strang, ISLR, and Think Stats.

Currently, I’m turning these working notes into my own book, Counting to Intelligence.

Latest · Counting to Intelligence

The Mistake That Makes Gradient Descent Work

Concept explained

Counting to IntelligenceGradient DescentOptimization

The Book · writing in public

Counting to Intelligence

How Simple Mathematics Becomes Machine Learning

The math is smaller than the magic. It is just counting, repeated until it looks like thought. I’m publishing it here section by section, as short posts.

01 — Series

Read a book with me

Each series follows one textbook in order. Start at chapter one, or jump to the topic you need.

02 — Recent notes

From the notebook

  1. The Mistake That Makes Gradient Descent Work Counting to Intelligence
  2. Data Cleaning for Machine Learning Systems: A Survey Data Cleaning for Machine Learning Sysytems: A Survey Note
  3. Mixture Models and Expectation Maximization - The EM Algorithm in General Bishop · Chapter 9 — Mixture Models and Expectation Maximization Bishop
  4. Mixture Models and Expectation Maximization - An Alternative View of EM Bishop · Chapter 9 — Mixture Models and Expectation Maximization Bishop
  5. Mixture Models and Expectation Maximization - Mixtures of Gaussians Bishop · Chapter 9 — Mixture Models and Expectation Maximization Bishop
  6. Mixture Models and Expectation Maximization - K-means Clustering Bishop · Chapter 9 — Mixture Models and Expectation Maximization Bishop
  7. Graphical Models - The Sum-product Algorithm, The Max-Sum Algorithm Bishop · Chapter 8 — Graphical Models Bishop
  8. Graphical Models - Inference in Graphical Models Bishop · Chapter 8 — Graphical Models Bishop
  9. Graphical Models - Markov Random Fields Bishop · Chapter 8 — Graphical Models Bishop
  10. Graphical Models - Conditional Independence Bishop · Chapter 8 — Graphical Models Bishop
  11. Graphical Models - Bayesian Networks Bishop · Chapter 8 — Graphical Models Bishop
  12. Sparse Kernel Methods - Maximum Margin Classifiers: Relation to Logistic Regression, Multiclass SVMs, SVMs for Regression Bishop · Chapter 7 — Sparse Kernel Methods Bishop
  13. Sparse Kernel Methods - Maximum Margin Classifiers: Overlapping Class Distributions Bishop · Chapter 7 — Sparse Kernel Methods Bishop
  14. Sparse Kernel Methods - Maximum Margin Classifiers Bishop · Chapter 7 — Sparse Kernel Methods Bishop
  15. Sparse Kernel Methods - Lagrange Multipliers Bishop · Chapter 7 — Sparse Kernel Methods Bishop
  16. Kernel Methods - Gaussian Process Bishop · Chapter 6 — Kernel Methods Bishop
  17. Kernel Methods - Constructing Kernels & Radial Basis Function Networks Bishop · Chapter 6 — Kernel Methods Bishop
  18. Kernel Methods - Dual Representations Bishop · Chapter 6 — Kernel Methods Bishop
  19. Neural Networks - Mixture Density Networks & Bayesian Neural Networks Bishop · Chapter 5 — Neural Networks Bishop
  20. Neural Networks - Regularization in Neural Networks Bishop · Chapter 5 — Neural Networks Bishop
  21. Neural Networks - The Hessian Matrix Bishop · Chapter 5 — Neural Networks Bishop
  22. Neural Networks - Error Backpropagation Bishop · Chapter 5 — Neural Networks Bishop
  23. Neural Networks - Network Training Bishop · Chapter 5 — Neural Networks Bishop
  24. Neural Networks - Feed-forward Network Functions Bishop · Chapter 5 — Neural Networks Bishop
  25. Linear Models for Classification - The Laplace Approximation & Bayesian Logistic Regression Bishop · Chapter 4 — Linear Models for Classification Bishop
  26. Linear Models for Classification - Probabilistic Discriminative Models Bishop · Chapter 4 — Linear Models for Classification Bishop
  27. Linear Models for Classification - Probabilistic Generative Models (Maximum Likelihood Solution) Bishop · Chapter 4 — Linear Models for Classification Bishop
  28. Linear Models for Classification - Probabilistic Generative Models Bishop · Chapter 4 — Linear Models for Classification Bishop
  29. Linear Models for Classification - The Perceptron Algorithm Bishop · Chapter 4 — Linear Models for Classification Bishop
  30. Linear Models for Classification - Fisher’s Linear Discriminant Bishop · Chapter 4 — Linear Models for Classification Bishop
  31. Linear Models for Classification - Least Squares for Classification Bishop · Chapter 4 — Linear Models for Classification Bishop
  32. Linear Models for Classification - Discriminant Functions (Part 2) Bishop · Chapter 4 — Linear Models for Classification Bishop
  33. Linear Models for Classification - Discriminant Functions Bishop · Chapter 4 — Linear Models for Classification Bishop
  34. Linear Models for Regression - Evidence Approximation & Limitations of Fixed Basis Function Bishop · Chapter 3 — Linear Models for Regression Bishop
  35. Linear Models for Regression - Bayesian Model Comparison Bishop · Chapter 3 — Linear Models for Regression Bishop
  36. Linear Models for Regression - Bayesian Linear Regression Bishop · Chapter 3 — Linear Models for Regression Bishop
  37. Linear Models for Regression - Bias-Variance Decomposition Bishop · Chapter 3 — Linear Models for Regression Bishop
  38. Linear Models for Regression - Linear Basis Function Models : Part 2 Bishop · Chapter 3 — Linear Models for Regression Bishop
  39. Linear Models for Regression - Linear Basis Function Models : Part 1 Bishop · Chapter 3 — Linear Models for Regression Bishop
  40. Probability Distributions - Nonparametric Methods Bishop · Chapter 2 — Probability Distributions Bishop
  41. Probability Distributions - The Exponential Family Bishop · Chapter 2 — Probability Distributions Bishop
  42. Probability Distributions - The Gaussian Distribution: Part 5 Bishop · Chapter 2 — Probability Distributions Bishop
  43. Probability Distributions - The Gaussian Distribution: Part 4 Bishop · Chapter 2 — Probability Distributions Bishop
  44. Probability Distributions - The Gaussian Distribution: Part 3 Bishop · Chapter 2 — Probability Distributions Bishop
  45. Probability Distributions - The Gaussian Distribution: Part 2 Bishop · Chapter 2 — Probability Distributions Bishop
  46. Probability Distributions - The Gaussian Distribution: Part 1 Bishop · Chapter 2 — Probability Distributions Bishop
  47. Probability Distributions - Multinomial Variables Bishop · Chapter 2 — Probability Distributions Bishop
  48. Probability Distributions - Binary Variables Bishop · Chapter 2 — Probability Distributions Bishop
  49. Introduction - Information Theory Bishop · Chapter 1 — Introduction Bishop
  50. Introduction - Decision Theory Bishop · Chapter 1 — Introduction Bishop
  51. Introduction - Model Selection & Curse of Dimensionality Bishop · Chapter 1 — Introduction Bishop
  52. Introduction - Probability Theory Bishop · Chapter 1 — Introduction Bishop
  53. Introduction - Polynomial Curve Fitting Bishop · Chapter 1 — Introduction Bishop
  54. Left, Right and Pseudo Inverses Strang · Chapter 28 Strang
  55. Linear Transformations, Change of Basis and Image Compression Strang · Chapter 27 Strang
  56. Singular Value Decomposition Strang · Chapter 26 Strang
  57. Similar Matrices Strang · Chapter 25 Strang
  58. Positive Definite Matrices Strang · Chapter 24 Strang
  59. Complex Matrices and Fourier Transform Strang · Chapter 23 Strang
  60. Symmetric Matrices and Positive Definiteness Strang · Chapter 22 Strang
  61. Markov Matrices and Fourier Series Strang · Chapter 21 Strang
  62. Differential Equations and Matrix Exponentials Strang · Chapter 20 Strang
  63. Diagonalization and Powers of a Matrix Strang · Chapter 19 Strang
  64. Eigenvalues and Eigenvectors Strang · Chapter 18 Strang
  65. Formula for $A^{-1}$ and Cramer's Rule Strang · Chapter 17 Strang
  66. Determinant and Cofactors Strang · Chapter 16 Strang
  67. Determinant Strang · Chapter 15 Strang
  68. Orthonormal Vectors, Orthogonal Matrices and Gram-Schmidt Method Strang · Chapter 14 Strang
  69. Projection Matrices and Least Squares Strang · Chapter 13 Strang
  70. Projection of a Matrix Strang · Chapter 12 Strang
  71. Orthogonal Vectors and Orthogonal Subspaces Strang · Chapter 11 Strang
  72. Graphs, Networks and Incidence Matrices Strang · Chapter 10 Strang
  73. Matrix Spaces Strang · Chapter 9 Strang
  74. Four Fundamental Subspaces Strang · Chapter 8 Strang
  75. Matrix Independence, Span, Basis & Dimension Strang · Chapter 7 Strang
  76. Algorithm for solving $Ax=b$ Strang · Chapter 6 Strang
  77. Algorithm for solving $Ax=0$ Strang · Chapter 5 Strang
  78. Vector Space and Subspace Strang · Chapter 4 Strang
  79. Inverse of a Matrix & Factorization into $A=LU$ Strang · Chapter 3 Strang
  80. Elimination & Permutation with Matrices Strang · Chapter 2 Strang
  81. Geometry of Linear Equations & Matrix Multiplications Strang · Chapter 1 Strang
  82. The Wilcoxon Signed-Rank Test The Wilcoxon Signed-Rank Test: Derivation of Mean and Variance Note
  83. Logistic Regression Logistic Regression: Derivation Note
  84. Hypothesis Testing (Part 6) Tests for Variances and Power of a Test Note
  85. Hypothesis Testing (Part 5) Tests with Categorical Data & Tests for Homogeneity and Independence Note
  86. Hypothesis Testing (Part 4) Distribution-Free Tests Note
  87. Hypothesis Testing (Part 3) Tests for the Difference Between Two Means (Large and Small Samples) and Tests with Paired Data Note
  88. Hypothesis Testing (Part 2) Tests for a Population Proportion Note
  89. Hypothesis Testing (Part 1) Tests for a Population Mean (Large and Small Samples) Note
  90. Confidence Intervals (Part 3) Confidence Intervals with Paired Data and Population Variance/ Prediction Intervals Note
  91. Confidence Intervals (Part 2) Confidence Intervals for Proportions and the Difference Note
  92. Confidence Intervals (Part 1) Confidence Intervals for a Population Mean Note
  93. Commonly used Distributions (Part 2) Commonly used Distributions Note
  94. Commonly used Distributions (Part 1) Commonly used Distributions Note
  95. Measurement and Propagation of Error (Part 2) Measurement and Propagation of Error Note
  96. Measurement and Propagation of Error (Part 1) Measurement and Propagation of Error Note
  97. Random Variables (Part 3: Jointly Distributed Random Variables) Jointly Distributed Random Variables Note
  98. Random Variables (Part 2: Continuous Random Variables) Continuous Random Variables Note
  99. Random Variables (Part 1: Discrete Random Variables) Discrete Random Variables Note
  100. Performance Metrics for Classification Algorithms Performance Metrics for Classification Algorithms Note
  101. Hypothesis testing Hypothesis testing Note
  102. Maximum Likelihood Estimation Estimation Note
  103. Naive Bayes Classifier Classification Note
  104. Correlation Think Stats · Chapter 9 Think Stats
  105. Estimation Think Stats · Chapter 8 Think Stats
  106. Hypothesis Testing Think Stats · Chapter 7 Think Stats
  107. Operations on Distributions Think Stats · Chapter 6 Think Stats
  108. Probability Think Stats · Chapter 5 Think Stats
  109. Continuous Distributions Think Stats · Chapter 4 Think Stats
  110. Cumulative Distribution Functions Think Stats · Chapter 3 Think Stats
  111. Descriptive Statistics Think Stats · Chapter 2 Think Stats
  112. Statistical Thinking for Programmers Think Stats · Chapter 1 Think Stats
  113. Content Based Movie Recommendation Engine Content based recommendation engine Note
  114. Unsupervised Learning: Applied Exercises ISLR · Chapter 10 — Unsupervised Learning ISLR
  115. Unsupervised Learning: Conceptual Exercises ISLR · Chapter 10 — Unsupervised Learning ISLR
  116. Hierarchical Clustering ISLR · Chapter 10 — Unsupervised Learning ISLR
  117. K-Means Clustering ISLR · Chapter 10 — Unsupervised Learning ISLR
  118. Principal Components Analysis: More on PCA ISLR · Chapter 10 — Unsupervised Learning ISLR
  119. Principal Components Analysis ISLR · Chapter 10 — Unsupervised Learning ISLR
  120. Support Vector Machines: Applied Exercises ISLR · Chapter 9 — Support Vector Machines ISLR
  121. Support Vector Machines: Conceptual Exercises ISLR · Chapter 9 — Support Vector Machines ISLR
  122. Support Vector Machines and Kernels ISLR · Chapter 9 — Support Vector Machines ISLR
  123. Support Vector Classifiers ISLR · Chapter 9 — Support Vector Machines ISLR
  124. Maximal Margin Classifier ISLR · Chapter 9 — Support Vector Machines ISLR
  125. Tree-Based Methods: Applied Exercises ISLR · Chapter 8 — Tree-Based Methods ISLR
  126. Tree-Based Methods: Conceptual Exercises ISLR · Chapter 8 — Tree-Based Methods ISLR
  127. Bagging, Random Forests, Boosting ISLR · Chapter 8 — Tree-Based Methods ISLR
  128. Decision Trees ISLR · Chapter 8 — Tree-Based Methods ISLR
  129. Moving Beyond Linearity: Applied Exercises ISLR · Chapter 7 — Moving Beyond Linearity ISLR
  130. Moving Beyond Linearity: Conceptual Exercises ISLR · Chapter 7 — Moving Beyond Linearity ISLR
  131. Local Regression, Generalized Additive Models ISLR · Chapter 7 — Moving Beyond Linearity ISLR
  132. Smoothing Splines ISLR · Chapter 7 — Moving Beyond Linearity ISLR
  133. Regression Splines ISLR · Chapter 7 — Moving Beyond Linearity ISLR
  134. Polynomial Regression, Step Functions, Basis Functions ISLR · Chapter 7 — Moving Beyond Linearity ISLR
  135. Linear Model Selection and Regularization: Applied Exercises ISLR · Chapter 6 — Linear Model Selection and Regularization ISLR
  136. Linear Model Selection and Regularization: Conceptual Exercises ISLR · Chapter 6 — Linear Model Selection and Regularization ISLR
  137. Dimension Reduction Methods ISLR · Chapter 6 — Linear Model Selection and Regularization ISLR
  138. Shrinkage Methods ISLR · Chapter 6 — Linear Model Selection and Regularization ISLR
  139. Subset Selection ISLR · Chapter 6 — Linear Model Selection and Regularization ISLR
  140. Resampling Methods: Applied Exercises ISLR · Chapter 5 — Resampling Methods ISLR
  141. Resampling Methods: Conceptual Exercises ISLR · Chapter 5 — Resampling Methods ISLR
  142. The Bootstrap ISLR · Chapter 5 — Resampling Methods ISLR
  143. Cross-Validation ISLR · Chapter 5 — Resampling Methods ISLR
  144. Classification: Applied Exercises ISLR · Chapter 4 — Classification ISLR
  145. Classification: Conceptual Exercises ISLR · Chapter 4 — Classification ISLR
  146. Linear Discriminant Analysis ISLR · Chapter 4 — Classification ISLR
  147. Logistic Regression ISLR · Chapter 4 — Classification ISLR
  148. Linear Regression: Applied Exercises ISLR · Chapter 3 — Linear Regression ISLR
  149. Linear Regression: Conceptual Exercises ISLR · Chapter 3 — Linear Regression ISLR
  150. Other Considerations in the Regression Model ISLR · Chapter 3 — Linear Regression ISLR
  151. Multiple Linear Regression ISLR · Chapter 3 — Linear Regression ISLR
  152. Simple Linear Regression ISLR · Chapter 3 — Linear Regression ISLR
  153. Statistical Learning: Applied Exercises ISLR · Chapter 2 — Statistical Learning ISLR
  154. Statistical Learning: Conceptual Exercises ISLR · Chapter 2 — Statistical Learning ISLR
  155. Assessing Model Accuracy ISLR · Chapter 2 — Statistical Learning ISLR
  156. What Is Statistical Learning? ISLR · Chapter 2 — Statistical Learning ISLR
  157. Introduction to Statistical Learning ISLR · Chapter 1 — Introduction ISLR
View the full archive (157 notes)