Part 1 Chapter 2 Last verified 2026-06-03

Offline metrics and the threshold trade-off

Precision, recall, F1, and the operating point — why a single accuracy number lies on imbalanced data, and how to choose a threshold from the business cost of each error. You build the metrics in mini_eval and feel the trade-off in an interactive demo.

On this page
  1. Would you ship it?
  2. The confusion matrix is the ground truth
  3. Feel the trade-off
  4. How this is graded
  5. Industry variation
  6. Stretch: when it isn’t classification

Offline metrics and the threshold trade-off

Would you ship it?

Your team trained a fraud classifier on last quarter’s transactions. The notebook reports 99.2% accuracy. A PM asks, in the interview you’re half-imagining, “Great — ship it?”

Sit with that for a second before reading on. What would you ask before answering? Write down one number you’d want to see.

That move — refuse the single number, ask which error matters — is the whole chapter.

The confusion matrix is the ground truth

Every classification metric is a summary of four counts at a chosen threshold tt (predict positive when scoret\text{score} \ge t):

| | predicted + | predicted − | | ------------ | ----------- | ----------- | | actual + | TP | FN | | actual − | FP | TN |

In mini_eval you build this directly — no library, so you can explain every line cold:

from mini_eval import confusion_counts, precision, recall, f1

c = confusion_counts(y_true, y_score, threshold=0.5)
precision(c)  # TP / (TP + FP) — of what we flagged, how much was right
recall(c)     # TP / (TP + FN) — of the real positives, how much we caught
f1(c)         # harmonic mean — punishes a lopsided trade-off

The three headline metrics, in symbols:

precision=TPTP+FP,recall=TPTP+FN,F1=2PRP+R.\text{precision}=\frac{TP}{TP+FP},\quad \text{recall}=\frac{TP}{TP+FN},\quad F_1=\frac{2\,PR}{P+R}.

Feel the trade-off

Here is the same imbalanced, fraud-like classifier from the opener — 2,000 transactions, ~8% fraud. Before you touch the slider: predict what happens to precision and recall as you drag the threshold from 0.5 toward 0.9.

Threshold explorerFraud-like classifier (synthetic, imbalanced) · 167/2000 positive (8.3%)
Precision
22.0%
Recall (TPR)
88.6%
F1
0.352
Accuracy
72.8%
FPR
28.6%
Confusion matrix
pred +pred −
actual +TP 148FN 19
actual −FP 525TN 1308
Score distribution (▢ neg · ▰ pos) + threshold
PR curve (recall → precision)

Now explain what you saw: the two score distributions overlap, so there is no threshold that cleanly separates them — every choice trades a false-positive rate against a miss rate. At t=0.5t=0.5 recall is high but precision is ~22%: you’d flag about three or four legitimate transactions for every real fraud. That may be exactly right (blocking fraud is worth annoying some customers) or exactly wrong (a declined legitimate transaction churns a customer) — the data cannot tell you which. The business cost of each error does.

How this is graded

Against the four-dimension rubric, a strong answer here shows:

  • Technical Correctness — precision ≠ recall ≠ accuracy; you state each precisely and know recall = TPR.
  • Trade-off Awareness — you name both errors and tie the threshold to their relative cost, rather than defaulting to 0.5 or “maximize F1”.
  • Evaluation Rigor — you report a metric with an operating point and a guardrail (e.g. “recall ≥ 0.9 subject to precision ≥ 0.3”).
  • Communication — “I’d run at recall 0.9 because a missed fraud costs ~30× a false alarm” beats “the F1 is 0.35”.

Industry variation

  • Fraud / risk — recall-heavy, but a false-positive budget (declined-txn rate) is a hard guardrail.
  • Content moderation — precision and recall both matter; the usual answer is a tiered threshold (auto-act on the confident tails, human-review the middle).
  • Healthcare — recall-dominant on the screen, with calibrated probabilities (Chapter 4) so a clinician can reason about the risk.

Stretch: when it isn’t classification

You’re handed a ranking problem instead — a retriever returning the top-k documents for a query. Precision and recall still apply (precision@k, recall@k), but the order within the top-k now matters, which a threshold can’t capture. What single number would you reach for, and what does it reward that F1 ignores? (We pick this up as NDCG and retrieval metrics in Chapter 9 — RAG evaluation.)