The F-score folds precision and recall into a single grade
A search engine or spam filter can fail in two ways: flagging things it shouldn't, or missing things it should catch. The F-score blends both kinds of error into one number between 0 and 1, and it does so with a harmonic mean, which punishes lopsided performance far more than an ordinary average would.
Two ingredients go in. Precision asks, of everything the system labelled positive, what share really was; recall asks, of everything truly positive, what share the system found. Doctors know recall as sensitivity and precision as positive predictive value. The balanced version, F1, is their harmonic mean, which in raw counts equals twice the true positives divided by twice the true positives plus the false positives and false negatives. A perfect score is 1.0, and if either precision or recall is zero the score collapses to 0.
Not every job values both equally. A weighted version adds a factor, beta, describing a user who cares beta times as much about recall as precision. Common choices are 2, favouring recall, and one half, favouring precision. The idea descends from an effectiveness measure in a book by the information scientist Van Rijsbergen, and the name F-measure is thought to come from a different F function in that book when the metric was introduced at the Fourth Message Understanding Conference in 1992.
Its home turf is information retrieval: judging search results, document sorting and query classification, especially where the positive cases are rare among many negatives. Natural language processing leans on it heavily too, for tasks such as spotting named entities and segmenting words. As giant search engines spread, attention shifted from the balanced F1 toward versions tilted to precision or to recall. Set-minded mathematicians may recognise F1 as the Dice coefficient between the retrieved items and the relevant ones.
It has blind spots. The measure ignores true negatives entirely, so for judging a binary classifier some prefer alternatives such as the Matthews correlation coefficient or Cohen's kappa. It also shifts with class balance, making comparisons across problems with different proportions of positives tricky. A lazy classifier that always answers yes gets perfect recall, and its F1 creeps toward 1 as positives become more common.
Source: F-score