The gradient is an arrow pointing straight up the steepest hill
Stand anywhere on a hillside and one direction climbs faster than every other. The gradient is the arrow that points that way, and its length tells you how steep the climb is. The same idea lets machine learning systems improve by stepping repeatedly in the opposite direction, downhill toward lower error.
Mathematically, the gradient takes a function of several variables, such as temperature throughout a room, and assigns an arrow to every point. In a heated room, each arrow shows the direction in which it gets warmer most quickly, and its magnitude says how fast. On a landscape described by height above sea level, the arrows point up the steepest slope. Where the arrow shrinks to nothing, at a peak, a valley floor or a flat pass, the point is called stationary.
Gradients also answer questions about other directions. Suppose the steepest route up a hill has a 40 percent slope. A road angled 60 degrees away from straight uphill climbs more gently, and the dot product gives its slope as 40 percent times the cosine of 60 degrees, which comes to 20 percent. This combination of arrow and direction is known as the directional derivative.
In ordinary x, y, z coordinates, the gradient is just the list of partial derivatives, one per variable. For the function 2x plus 3y squared minus the sine of z, it works out to 2, then 6y, then minus the cosine of z. The symbol is an upside-down triangle called nabla and read as del, though grad f is also used. That simple recipe holds only for orthonormal axes; other coordinate systems need a correction from the metric tensor.
There is a subtle trap. A function can have partial derivatives in every direction yet still fail to be differentiable, like x squared times y divided by the sum of x squared and y squared at the origin. There the recipe breaks: rotating the axes changes the answer, and the result may not point toward the steepest ascent.
Source: Gradient