Matrix calculus has two rival spelling systems, and mixing them breaks answers
Matrix calculus bundles hundreds of partial derivatives into tidy vectors and matrices, which is why statisticians and machine learning engineers lean on it. The catch: practitioners split into camps that lay those derivatives out differently, and formulas copied from one author into another's framework can quietly produce serious errors.
The core idea is bookkeeping. A function may depend on many inputs, or produce many outputs, and you want every rate of change of every output with respect to every input. Rather than writing each partial derivative separately, matrix calculus stacks them into a single object you can manipulate as one thing. Inputs and outputs can each be a scalar, a vector or a matrix, giving nine combinations in all; six of them fit neatly into a matrix, while the rest, such as differentiating a vector by a matrix, need higher-rank tensors.
Familiar ideas fall out as special cases. Differentiate a single number by a vector and transpose, and you get the gradient, the arrow pointing uphill on a surface; the electric field, for instance, is the negative gradient of electric potential. Differentiating position by time yields velocity, a tangent vector, and differentiating again gives acceleration. That compactness makes it easier to find the peaks and troughs of functions of many variables and to solve systems of differential equations.
The trouble is layout. One convention writes the derivative of a scalar by a vector as a row, the other as a column, commonly called numerator and denominator layout. There are really more than two choices, because authors can pick independently for each kind of derivative and many mix them. A field such as econometrics may settle on one, yet authors within it still disagree, and each tends to write as if theirs were standard. The practical advice is to identify the layout a source uses and stick with it rather than forcing everything into one style.
Physicists often prefer tensor index notation with the Einstein summation convention, which handles high-rank tensors easily. In estimation theory the index count becomes unmanageable, so matrix calculus wins, powering derivations such as the expectation-maximization algorithm for Gaussian mixtures.
Source: Matrix calculus