Precision, Recall & F1 Score Calculator
Enter the four numbers from your confusion matrix — true positives, false positives, false negatives and true negatives — and get every common classification metric at once, with the formula for each.
Confusion matrix
1,000 samples · 110 actual positives (11.00% of the data)
F1 score
0.8095
2 × P × R ÷ (P + R)
Precision
85.00%
TP ÷ (TP + FP)
Recall (sensitivity)
77.27%
TP ÷ (TP + FN)
Accuracy
96.00%
(TP + TN) ÷ total
Specificity
98.31%
TN ÷ (TN + FP)
F2 score
0.7870
(1 + β²)PR ÷ (β²P + R)
Balanced accuracy
87.79%
(Recall + Specificity) ÷ 2
MCC
0.7884
Matthews correlation, −1 to 1
Negative predictive value
97.22%
TN ÷ (TN + FN)
False positive rate
1.69%
FP ÷ (FP + TN)
β = 2 weights recall higher; β = 0.5 weights precision higher.
Every classification metric from one confusion matrix
Whenever a model sorts things into two groups — spam or not spam, fraud or legitimate, relevant or irrelevant — its results fit into a confusion matrix of four numbers. From those four numbers come all the standard evaluation metrics: precision, recall, F1, accuracy and more. This calculator computes them all at once, shows each formula, and warns you when a high accuracy is hiding a weak model.
How to use the calculator
- Count your model’s predictions against the true labels on a test set.
- Enter TP (positives correctly found), FP (negatives wrongly flagged), FN (positives missed) and TN (negatives correctly rejected).
- Read F1, precision, recall and the other metrics, updated as you type.
- Adjust β to weight recall or precision more heavily in the F-beta score.
What each metric tells you
- Precision = TP ÷ (TP + FP). When the model says “positive”, how often is it right?
- Recall (sensitivity) = TP ÷ (TP + FN). Of all real positives, how many did it catch?
- F1 = 2 × precision × recall ÷ (precision + recall). The harmonic mean punishes a weakness in either, so both must be good for F1 to be high.
- Accuracy = (TP + TN) ÷ total. Simple, but misleading when one class is rare.
- Specificity = TN ÷ (TN + FP). How well negatives are recognized.
- Balanced accuracy = the average of recall and specificity, robust to class imbalance.
- MCC (Matthews correlation coefficient) uses all four cells and ranges from −1 to 1, where 0 means no better than random guessing.
Worked example
A fraud model is tested on 1,000 transactions, 110 of them fraudulent. It flags 100: 85 are real fraud (TP) and 15 are not (FP). It misses 25 frauds (FN) and correctly clears 875 legitimate transactions (TN). Precision is 85 ÷ 100 = 85%, recall is 85 ÷ 110 = 77.3%, and F1 is 0.81. Accuracy looks excellent at 96%, but mostly because legitimate transactions are easy and plentiful — which is exactly why F1 and MCC are preferred for imbalanced problems like this.
Choosing the right trade-off
Precision and recall pull against each other: lowering a model’s decision threshold catches more positives (higher recall) but raises false alarms (lower precision). Decide which error is more expensive. Missing a disease or a fraud is usually worse than a false alarm, so recall — and F2 — matter most. For spam filters and content moderation, wrongly blocking good content is costly, so precision and F0.5 carry more weight.
Preparing training data? Convert spreadsheets with the CSV to JSONL Converter, size hardware with the GPU VRAM Calculator, or work out percentage changes between model versions with the Percentage Calculator.
Frequently asked questions
What is the F1 score?
F1 is the harmonic mean of precision and recall: 2 × precision × recall ÷ (precision + recall). It's high only when both are high, which makes it useful when you care about finding positives and about avoiding false alarms.
What's the difference between precision and recall?
Precision asks: of everything the model flagged as positive, how much was right? Recall asks: of all the real positives, how many did the model find? A spam filter with high precision rarely flags good email; one with high recall catches most spam.
Why is accuracy misleading for imbalanced data?
If only 1% of cases are positive, a model that always predicts negative is 99% accurate but useless. F1, MCC and balanced accuracy account for the imbalance and give a truer picture.
When should I use F-beta instead of F1?
Use β above 1 when missing positives is worse than false alarms, such as disease screening, and β below 1 when false alarms are costlier, such as spam filtering. F2 weights recall twice as much as precision.
What is MCC?
The Matthews correlation coefficient uses all four cells of the confusion matrix and ranges from −1 (always wrong) through 0 (no better than chance) to 1 (perfect). It's considered one of the most balanced single metrics for binary classifiers.
Related tools
CSV to JSONL Converter (Fine-Tuning)
Turn a spreadsheet into JSONL training data in chat-message format for fine-tuning.
Percentage Calculator
Work out percentages, percentage change, increases and decreases.
GPU VRAM Calculator for LLMs
Estimate the GPU memory needed to run or fine-tune an LLM, and which GPUs fit.