Lagrange multipliers find the best point when you must stay on a path
Imagine hiking along a fixed trail and wanting its highest point. At the summit of your trail, the contour lines of the landscape just graze the path, running parallel to it. That simple picture is the heart of Lagrange multipliers, a method for optimising a quantity while obeying exact constraints.
Named after Joseph-Louis Lagrange, the method tackles problems where you want the largest or smallest value of a function, but the variables must satisfy one or more equations exactly. Rather than solving the constraint for one variable and substituting, it folds the constraints into a new combined expression called the Lagrangian, which brings in an extra unknown, the multiplier, for each constraint. The derivative tests used for unconstrained problems can then be applied as usual.
The intuition runs like this. Walk along the curve where the constraint holds. If the function were still increasing in some direction along that curve, you could keep walking and climb higher, so you were not at a maximum. At a true peak, the function barely changes as you step along the constraint, which happens exactly where a contour line of the function touches the constraint curve and their tangents line up. In gradient terms, the gradient of the function becomes a combination of the gradients of the constraints, with the multipliers as the weights.
Its big advantage is that you never need to describe the constraint explicitly with parameters, which is why the method is so widely used on difficult optimisation problems. The Karush–Kuhn–Tucker conditions extend it to inequality constraints as well.
There are caveats. The method gives only a necessary condition, so not every candidate it finds is a genuine solution, and a technical requirement called constraint qualification must hold. The solution to the original problem shows up as a saddle point of the Lagrangian, which can be spotted using a bordered Hessian matrix. Even points that pass the sufficient tests are only guaranteed best locally, so finding the overall optimum means comparing the objective function across all the surviving candidates.
Source: Lagrange multiplier