Skip to main content

Command Palette

Search for a command to run...

ML from Scratch — Episode 2: Linear Regression, The Implementation

Updated
1 min readView as Markdown
A
A learner building ML algorithms from scratch to actually understand what's happening under the hood. Writing about the math, the mistakes, and the process and not just the results.

The Math

Linear regression learns to predict a continuous value from input data, by finding the best-fit line through the training examples.

Cost Function

To measure how wrong our model is, we use MSE (Mean Squared Error):

$$J(w, b) = \frac{1}{2m} \sum_{i=1}^{m} (\hat{y}^{(i)} - y^{(i)})^2$$

Where m is the number of training examples, ŷ is the predicted value, and y is the actual label.We square the differences to penalize large errors more heavily-an error of 10 becomes 100, while an error of 2 becomes 4.

Gradients

To minimize the cost, we compute how it changes with respect to each parameter:

$$\frac{\partial J}{\partial w} = \frac{1}{m} \sum_{i=1}^{m} (\hat{y}^{(i)} - y^{(i)}) \cdot x^{(i)}$$

$$\frac{\partial J}{\partial b} = \frac{1}{m} \sum_{i=1}^{m} (\hat{y}^{(i)} - y^{(i)})$$

Gradient Descent

We update the parameters on every iteration:

$$w = w - \alpha \cdot \frac{\partial J}{\partial w}$$

$$b = b - \alpha \cdot \frac{\partial J}{\partial b}$$

Here α is the learning rate —> it controls how big each step is toward the minimum. Too large and you overshoot. Too small and training takes forever. Too large and you overshoot.

The goal: reduce J as much as possible with each iteration.

15 views