<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Optimization on Amit Rajan</title><link>https://amitrajan012.github.io/tags/optimization/</link><description>Recent content in Optimization on Amit Rajan</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Sun, 04 Oct 2026 06:28:22 +0000</lastBuildDate><atom:link href="https://amitrajan012.github.io/tags/optimization/index.xml" rel="self" type="application/rss+xml"/><item><title>The Mistake That Makes Gradient Descent Work</title><link>https://amitrajan012.github.io/book/blog/the-mistake-that-makes-gradient-descent-work/</link><pubDate>Sun, 04 Oct 2026 06:28:22 +0000</pubDate><guid>https://amitrajan012.github.io/book/blog/the-mistake-that-makes-gradient-descent-work/</guid><description>&lt;p&gt;Machine learning starts with a simple idea: find the model that makes the fewest mistakes. We measure those mistakes with a loss function, and training is the process of driving that loss downward. The natural conclusion is that we should keep descending until we reach the lowest possible point—the deepest valley in the landscape.&lt;/p&gt;&#10;&lt;p&gt;There is a problem with this idea. The deepest valley is not always the best place to be.&lt;/p&gt;</description></item></channel></rss>