Gradient descent by hand
Introduction Deep learning is black magic. We throw data into PyTorch and we get a model that seemingly understands language, speech, and vision. However, one fundamental technique for deep learning has remained the same. That is gradient descent. The modern optimisers such as Adam and AdamW are refinements of gradient descent. So, I believe that gradient descent is still a core technique worth learning. I've written a post on the relationship between differentiation and optimisation ( https://yasufumimoriya.blogspot.com/2026/04/differentiation-and-optimisation.html ). In this post, there was a single parameter to optimise or to find its slope. Gradient descent is essentially finding slopes of multiple parameters at once. This post focuses on how computation of gradient descent proceeds to update parameters, taking a function with two parameters as an example. Refresher: Differentiation and Optimisation This section is a refresher of the relations...