Methodology
How detection accuracy is measured
Written and maintained by Hendrik Schneider · Last reviewed · How we check this
A scanner that says it reduces false positives without publishing a number is making the same unfalsifiable claim as every other scanner. Korthex ships an Accuracy Engine that measures detection against a ground-truth corpus, and it is the same engine that gates its own releases. This page is the method. The figures it produces are at the bottom, including the ones that are not flattering.
What is being measured
Three ratios, computed from three counts. A true positive is a finding the scanner reported that the ground truth also holds. A false positive is one it reported that the ground truth does not. A false negative is one the ground truth holds that the scanner did not report.
| Metric | Formula | What a high value means | What it hides |
|---|---|---|---|
| Precision | tp / (tp + fp) | Findings you see are real | A scanner that reports almost nothing scores well |
| Recall | tp / (tp + fn) | Real findings are not missed | A scanner that reports everything scores well |
| F1 | harmonic mean of the two | Neither is being traded away | A single number hides which of the two moved |
What ground truth is here
A ground-truth file is a hand-verified inventory of the cryptography in a corpus: for each occurrence, the algorithm, the file, and the line. It is JSON, it is checked in, and it is the thing the scanner is graded against - which makes it the weakest link in the whole method, because a ground truth that is wrong grades a correct scanner as broken.
That is why the corpus is named in every published figure. A precision number without a named corpus is not a measurement of a scanner, it is a measurement of a corpus nobody can inspect.
The four modes
- Scanner: the default. Scan output against the ground truth, one corpus, one run.
- Migration: the same comparison across a migration - the before state in the expected slot, the after state in the actual one. It measures whether the plan worked, not whether the scan was right.
- False positives: run against a corpus that is known to be clean. Every finding is by definition a false positive, so this mode isolates the rate without a ground truth having to enumerate the true ones.
- Corpus: an aggregate across a whole corpus rather than a single scan, with a wider line window.
The tolerances, and the one that is deliberately absent
Source code moves. A refactor that shifts a call two lines down is not a detection regression, so matching allows a configurable line tolerance. Algorithm names are normalised as well: MD5Sum, md5 and MD-5 collapse to one canonical match, because a scanner is not wrong for spelling it differently than the ground-truth author did.
One collapse is refused on purpose: two findings both labelled Unknown never match each other. It would be trivially easy to let them, and it would inflate every precision number on this page. An unidentified algorithm on both sides is two open questions, not one answer.
What these numbers do not show
- They do not generalise to your codebase. A ratio measured on a named corpus describes that corpus. A codebase with different languages, different libraries or unusual crypto wrappers will score differently, and possibly worse.
- They do not measure severity. A missed hardcoded private key and a missed MD5 checksum count the same in recall, and they are not the same finding.
- They do not measure the compliance mapping. Whether a finding is graded against the right BSI or NIST rule is a separate question, held by a different build gate.
- A single run is a point, not a trend. The engine keeps a run history precisely because one measurement cannot tell you whether the scanner is improving.
Why the engine gates the release?
Cryptographic detection is full of edge cases, and a change to the AST analysis can fix five findings while breaking twelve. That failure is silent unless something measures it, so the accuracy run is a pass/fail step in the release pipeline by default: a run whose recall falls under 0.8 exits non-zero without any threshold being passed on the command line. Measuring without gating is possible, but it has to be asked for.
The run history that supports this is deliberately narrow. Each record holds the timestamp, the mode, the line tolerance and the metrics - and never a finding, a file path or a source excerpt. That is what makes it safe to keep on a shared CI runner, and it is the same on-premise reasoning the rest of the product follows.
Current measurements
No measurement has been published yet. The engine and the method are described above; the figures appear here once a run against a stable corpus has been reviewed and released. An unpublished number is not a withheld number - it is one that has not been checked yet.
Frequently asked questions
What is the false-positive rate of Korthex?
It is measured per corpus rather than claimed as one number, using a corpus known to be clean, so that every finding in that run is by definition a false positive. The published figures and the corpus each one was measured on are at the bottom of this page.
How is scanner accuracy calculated?
Precision is tp/(tp+fp), recall is tp/(tp+fn), and F1 is their harmonic mean, where a true positive is a finding that the hand-verified ground truth also holds at that file and line, within a configurable line tolerance.
Do these accuracy numbers apply to my codebase?
No. A ratio measured against a named corpus describes that corpus. Different languages, libraries or crypto wrappers will produce different numbers, which is why the corpus is named next to every figure rather than left implicit.
Does Korthex publish measurements that got worse?
Yes - that is the point of publishing them. A methodology page that can only show improvements is a brochure. Publication is a deliberate step for each measurement, but the criterion is that the run was valid, not that the result was good.