KORTHEXDocumentation

Docs / ANALYSIS ENGINES

ANALYSIS ENGINES

Accuracy Engine

Written and maintained by Hendrik Schneider · Last reviewed · How we check this

The Accuracy Engine measures how well the scanner is finding what it should be finding. It compares actual scan output against ground-truth corpora and emits standard precision / recall / F1 metrics so users (and Korthex itself) can answer "is the scanner getting better or worse over time?"

Measurement Modes

One command, four modes, selected with --mode . There is no measure subcommand and the command takes no positional arguments at all: every input arrives through a named option. The mode names below are the values --mode accepts, not the internal engine class names. Inputs are JSON, not .kxr . Both files are read as text and parsed, so hand an encrypted report to --actual and the run fails on the read. Produce the input with korthex scan . --format json --output scan.json . And note the direction of the names: --actual is what gets read, --output is where the result gets written. They are not two spellings of the same thing.

ModeInputsOutput
scannerThe default. --expected is the ground-truth JSON, --actual the scan report JSON.True positives / false positives / false negatives + precision/recall/F1 + the actual misses[] and false-alarms[] for review.
migrationThe before-snapshot goes in --expected , the after-snapshot in --actual .Findings resolved, remaining, newly introduced + success rate.
fpA known-clean corpus in --clean-manifest , plus --actual . Refuses to run without the manifest.False-positive rate plus the individual false alarms.
corpusSame two inputs as scanner , aggregated across many files.Micro + macro precision/recall/F1 + per-file breakdown.

Tolerance Knobs

Source code shifts: whitespace changes, refactors, statement reformatting. Accuracy measurement absorbs these via configurable line tolerance so a perfect detection isn't flagged as a regression just because the target moved two lines. Override it with --line-tolerance . Zero is not "no tolerance": it means "use the mode default", which is why the defaults above are reachable without passing the option at all. Algorithm-name normalization is built in: MD5Sum , md5 , MD-5 all collapse to the same canonical match. Two "Unknown" labels never collapse into a match - accuracy stays honest.

ModeDefault tolerance
scanner2 lines (absorbs whitespace / statement-formatting drift).
migration5 lines (absorbs refactor drift).

Using the Engine

# Scanner accuracy against a ground-truth file (the default mode) korthex accuracy \ --expected .korthex_ground_truth.json \ --actual scan.json # Migration success rate: before goes in --expected, after in --actual korthex accuracy --mode migration \ --expected before.json --actual after.json # False-positive rate against a known-clean corpus korthex accuracy --mode fp \ --clean-manifest tests/clean-manifest.json --actual scan.json # Aggregate across an entire corpus, with a wider line window korthex accuracy --mode corpus \ --expected .korthex_ground_truth.json --actual scan.json --line-tolerance 5 # Human-readable instead of JSON, written to a file korthex accuracy --expected gt.json --actual scan.json --format table --output accuracy.txt

CI Gates and Exit Codes

This command gates by default, whether or not you ask it to. If you drop it into a pipeline expecting a measurement tool, it is a pass/fail step: a scanner or corpus run whose recall falls below 0.8 exits 10 even though no threshold was passed on the command line. Set --min-recall 0 to measure without gating. Two of those reuses are worth reading twice. --min-recall is the migration success floor, and --min-precision is the false-positive ceiling, where a lower number is stricter. The names describe the scanner mode they were built for, not the mode you may be running.

ModeWhat is comparedThreshold option and default
scannerRecall, precision and F1, all three at once. One below its floor fails the run.recall --min-recall (0.8), precision --min-precision (0.0), F1 --min-f1 (0.0)
corpusThe same three, taken from the micro-averaged figures.same as scanner
migrationSuccess rate against a single floor.--min-recall (0.8), reused as the migration floor
fpFalse-positive rate against a ceiling. Smaller is better here, so the comparison is inverted.--min-precision reused as the maximum allowed rate; 0.1 when it is left unset
Exit codeMeaning
0Measured, and every threshold in play was met.
2The run failed: an input was missing or unreadable, or the result could not be parsed.
3The current license tier does not include the Accuracy Engine.
10Measured successfully, and a threshold was missed. This is the gate firing, not an error.

Run History

Every run appends a record to disk unless you pass --no-history . One JSON file per run, named for its UTC timestamp and mode, under %APPDATA%/korthex/accuracy-history/ on Windows and ~/.config/korthex/accuracy-history/ elsewhere. Point it somewhere else with --config-root . The record is a fixed whitelist: timestamp, mode, line tolerance and the metrics for that mode. Findings, file paths and source excerpts are never written, so the history is safe to keep on a shared runner. Nothing prunes it automatically, and a write that fails warns on stderr rather than failing the measurement.

Why This Engine Exists

Cryptographic scanning is full of edge cases. Without measurement, releases can silently regress: a small change to the AST analyzer might fix five findings but break twelve others. The Accuracy Engine is how Korthex itself ensures every release improves rather than degrades detection - and is the engine you reach for when a customer says "you missed something" to prove the regression or confirm the original detection.