Docs / ANALYSIS ENGINES
ANALYSIS ENGINES
Accuracy Engine
Written and maintained by Hendrik Schneider · Last reviewed · How we check this
The Accuracy Engine measures how well the scanner is finding what it should be finding. It compares actual scan output against ground-truth corpora and emits standard precision / recall / F1 metrics so users (and Korthex itself) can answer "is the scanner getting better or worse over time?"
Measurement Modes
One command, four modes, selected with --mode . There is no measure subcommand and the command takes no positional arguments at all: every input arrives through a named option. The mode names below are the values --mode accepts, not the internal engine class names. Inputs are JSON, not .kxr . Both files are read as text and parsed, so hand an encrypted report to --actual and the run fails on the read. Produce the input with korthex scan . --format json --output scan.json . And note the direction of the names: --actual is what gets read, --output is where the result gets written. They are not two spellings of the same thing.
| Mode | Inputs | Output |
|---|---|---|
| scanner | The default. --expected is the ground-truth JSON, --actual the scan report JSON. | True positives / false positives / false negatives + precision/recall/F1 + the actual misses[] and false-alarms[] for review. |
| migration | The before-snapshot goes in --expected , the after-snapshot in --actual . | Findings resolved, remaining, newly introduced + success rate. |
| fp | A known-clean corpus in --clean-manifest , plus --actual . Refuses to run without the manifest. | False-positive rate plus the individual false alarms. |
| corpus | Same two inputs as scanner , aggregated across many files. | Micro + macro precision/recall/F1 + per-file breakdown. |
Tolerance Knobs
Source code shifts: whitespace changes, refactors, statement reformatting. Accuracy measurement absorbs these via configurable line tolerance so a perfect detection isn't flagged as a regression just because the target moved two lines. Override it with --line-tolerance . Zero is not "no tolerance": it means "use the mode default", which is why the defaults above are reachable without passing the option at all. Algorithm-name normalization is built in: MD5Sum , md5 , MD-5 all collapse to the same canonical match. Two "Unknown" labels never collapse into a match - accuracy stays honest.
| Mode | Default tolerance |
|---|---|
| scanner | 2 lines (absorbs whitespace / statement-formatting drift). |
| migration | 5 lines (absorbs refactor drift). |
Using the Engine
# Scanner accuracy against a ground-truth file (the default mode) korthex accuracy \ --expected .korthex_ground_truth.json \ --actual scan.json # Migration success rate: before goes in --expected, after in --actual korthex accuracy --mode migration \ --expected before.json --actual after.json # False-positive rate against a known-clean corpus korthex accuracy --mode fp \ --clean-manifest tests/clean-manifest.json --actual scan.json # Aggregate across an entire corpus, with a wider line window korthex accuracy --mode corpus \ --expected .korthex_ground_truth.json --actual scan.json --line-tolerance 5 # Human-readable instead of JSON, written to a file korthex accuracy --expected gt.json --actual scan.json --format table --output accuracy.txt
CI Gates and Exit Codes
This command gates by default, whether or not you ask it to. If you drop it into a pipeline expecting a measurement tool, it is a pass/fail step: a scanner or corpus run whose recall falls below 0.8 exits 10 even though no threshold was passed on the command line. Set --min-recall 0 to measure without gating. Two of those reuses are worth reading twice. --min-recall is the migration success floor, and --min-precision is the false-positive ceiling, where a lower number is stricter. The names describe the scanner mode they were built for, not the mode you may be running.
| Mode | What is compared | Threshold option and default |
|---|---|---|
| scanner | Recall, precision and F1, all three at once. One below its floor fails the run. | recall --min-recall (0.8), precision --min-precision (0.0), F1 --min-f1 (0.0) |
| corpus | The same three, taken from the micro-averaged figures. | same as scanner |
| migration | Success rate against a single floor. | --min-recall (0.8), reused as the migration floor |
| fp | False-positive rate against a ceiling. Smaller is better here, so the comparison is inverted. | --min-precision reused as the maximum allowed rate; 0.1 when it is left unset |
| Exit code | Meaning |
|---|---|
| 0 | Measured, and every threshold in play was met. |
| 2 | The run failed: an input was missing or unreadable, or the result could not be parsed. |
| 3 | The current license tier does not include the Accuracy Engine. |
| 10 | Measured successfully, and a threshold was missed. This is the gate firing, not an error. |
Run History
Every run appends a record to disk unless you pass --no-history . One JSON file per run, named for its UTC timestamp and mode, under %APPDATA%/korthex/accuracy-history/ on Windows and ~/.config/korthex/accuracy-history/ elsewhere. Point it somewhere else with --config-root . The record is a fixed whitelist: timestamp, mode, line tolerance and the metrics for that mode. Findings, file paths and source excerpts are never written, so the history is safe to keep on a shared runner. Nothing prunes it automatically, and a write that fails warns on stderr rather than failing the measurement.
Why This Engine Exists
Cryptographic scanning is full of edge cases. Without measurement, releases can silently regress: a small change to the AST analyzer might fix five findings but break twelve others. The Accuracy Engine is how Korthex itself ensures every release improves rather than degrades detection - and is the engine you reach for when a customer says "you missed something" to prove the regression or confirm the original detection.