Fit the line by squaring every vertical miss
Ordinary least squares picks the coefficients of a linear regression by minimising the sum of squared gaps between observed responses and the model’s predictions. Geometrically those gaps run parallel to the dependent-variable axis; smaller squared distances mean a tighter fit, and a simple formula often delivers the estimator.
Write each observation as a response equal to a linear combination of regressors plus an unobserved error, or in matrix form y = Xβ + ε. A column of ones usually supplies an intercept; without it the fitted surface is forced through the origin. Regressors need not be mutually independent—polynomials of the same variable still yield a model linear in parameters—yet perfect multicollinearity blocks unique coefficient estimates and consistency for those entangled terms. Short of that singularity, rising multicollinearity mainly inflates standard errors.
The estimator solves a quadratic minimisation: choose β-hat to minimise the squared Euclidean norm of y − Xβ. Under exogenous regressors and a full-rank (no perfect collinearity) design, OLS is consistent for the level-one fixed effects. Finite fourth moments help consistency for residual variance. The Gauss–Markov theorem then crowns OLS as the best linear unbiased estimator when errors are homoscedastic and serially uncorrelated—minimum variance among linear unbiased competitors when error variances are finite.
Add normally distributed zero-mean errors and OLS becomes the maximum-likelihood estimator, outperforming nonlinear unbiased rivals under that stronger model. Some texts simply equate OLS with linear regression itself. In practice analysts still watch leverage, heteroscedasticity, and near-collinear predictors, because the tidy optimality theorems travel only as far as their assumptions. When a unique minimiser exists, the algebra yields a closed-form estimator especially tidy in simple regression with one right-hand regressor. The method’s endurance owes less to mystique than to a transparent loss function: punish big vertical mistakes harder than small ones, then solve.
Source: Ordinary least squares