KORTHEXkorthex.io

Published model

How long does a scan actually take?

Written and maintained by Hendrik Schneider · Last reviewed · How we check this

Enter your project and your machine. You get a span for the first scan and a span for every scan after it, together with the measurements the span is built from. Where the model runs out of evidence it says so, instead of extrapolating.

What the number is built from

The calculator is a cost model over measured runs, not a marketing figure with a slider attached. Every factor is either a measurement with its input shape named, or a declared assumption, and the page states which is which for each one. The interactive version needs JavaScript; the model itself is below either way.

  • Project size: the code, context, AST, TLS, config and database path grows as files raised to about 1.37. Measured at two sizes: 900 files across 12,101,499 bytes and all 18 languages in 11.052 s, and 25,684 files with about 7 million lines in 1,080.5 s. Two points fix the exponent exactly and say nothing about the shape between them.
  • History depth: about 18.9 ms per commit, linear. Measured once, at 674,780 commits in a 5.7 GB pack, taking 12,759 s. For most repositories this is the single largest term in the estimate.
  • CPU cores: Amdahl scaling with a 13.4 % serial share, fitted through one measured ratio. Two workers to sixteen gave a 3.02 times wall-clock speedup while CPU efficiency fell to 38 %. The consequence is a ceiling: an unlimited core count buys at most 1.41 times over the sixteen-core reference machine.
  • RAM: a threshold, not a factor. Demand is about 56 kB per file plus 27 kB per commit, derived from a 110 MB peak on the small corpus and a 19.5 GB peak on the large one. Above the demand, more RAM changes nothing. Below it the scan degrades, and below half of it the calculator declines to give a figure, because at that ratio the reference run aborted with an allocation failure rather than merely slowing down.
  • Storage: it acts almost exclusively on the history phase. The code path cannot be storage-bound, since it consumes source at about 1.1 MB/s, which even a hard disk outruns by two orders of magnitude. Reading a multi-gigabyte pack with one random access per commit is a different regime. Only the reference NVMe was measured; the other device classes are declared figures.
  • Follow-up scan: not a fixed percentage. A re-run with nothing changed still costs a floor of roughly 17 % of the first scan, and the changed share of the files decides the rest. Measured at both ends: 40 parse-heavy modules with one changed went from 1.626 s to 0.553 s, and the reference corpus re-run unchanged went from about 1,080 s to about 180 s.
  • GPU: no factor and deliberately no control.

Why there is no GPU slider?

There is no GPU control here because there is no GPU effect to control. The two phases that dominate a scan are a git tree diff and a set of language parsers. Both are branch-heavy work with unpredictable memory access, which is precisely what a GPU is worst at. Korthex does dispatch one workload to the GPU, a delta hash that is bit-identical to the CPU path and falls back to it on any failure, and it sits on neither of those two critical paths. A slider here would advertise a lever that does not exist.

The reference machine and the reference corpus

Every timing in the model comes from one host and two corpora. Naming them is the point: a duration without its input shape is an illustration that looks like a specification.

  • Host: 8 physical and 16 logical cores, 30.9 GB RAM, a Samsung 980 PRO NVMe, Windows 11.
  • Large corpus: LibreOffice core, 25,684 source files, about 7 million lines, 674,780 commits, a 5.7 GB pack. Cold and complete it took 3 h 50 min, split into git history 12,759 s, context 1,009 s, AST 67 s, TLS 1.2 s, config 1.0 s and database 2.3 s.
  • Small corpus: the committed performance baseline of 900 files, 12,101,499 bytes across all 18 languages. Code engine wall-clock 11.052 s, peak resident set 110 MB.

Where this model is thin

Stated plainly, because a calculator that hides its weak spots is a claim with a slider attached.

  • Two measured project sizes fix the exponent exactly and say nothing about the shape between them. Anything outside 900 to 25,684 files is extrapolation.
  • The per-commit cost comes from a single repository with a single pack. A history with wide trees, or with many tiny commits, will not match it.
  • Only the reference NVMe was measured. SATA, hard disk and network figures are declared device-class values, and the hard-disk history estimate is a floor, since it assumes only one pack access per commit.
  • The core curve is fitted through one measured ratio. It reproduces that ratio exactly and is unverified everywhere else.
  • For RAM, the demand and the threshold are measured. The slope of the slowdown below the threshold is not.
  • The span itself is a declared assumption. The project performance gate treats a drop to half throughput on the same corpus and the same machine as still inside tolerance, so a narrower span would claim more stability than the gate assumes.

Frequently asked questions

How long does a Korthex scan take?

It depends on three things, and the calculator on this page shows all three: how much source there is, how deep the git history is, and what the machine is. For a 500,000 line codebase with a short history on a current sixteen-core workstation the model puts a first scan at roughly one and a half to three minutes. The same codebase with 50,000 commits of history is a quarter of an hour, because history is usually the dominant term.

Is a scan of 50,000 to 500,000 lines under two minutes?

For the source analysis alone, comfortably: the model at korthex.io/scan-duration puts 500,000 lines at about half a minute on the reference machine, and 50,000 lines at a second or two. Whether the whole run stays under two minutes is decided by the git history, which is not part of that figure and is frequently larger than it. That is why the flat promise was replaced by the korthex.io/scan-duration calculator.

Does more RAM make the scan faster?

Only up to the point where there is enough. RAM is a threshold, not a throttle: above what the working set needs, adding more changes nothing measurable. Below it the scan degrades, and far below it the run can fail outright rather than just take longer, which is what happened on the reference corpus at a 19.5 GB peak.

Does a GPU speed up a Korthex scan?

No, and the calculator deliberately offers no GPU control. Git tree diffing and language parsing are branch-heavy work with unpredictable memory access, which is the workload class a GPU handles worst. Korthex does run one hash workload on the GPU, bit-identical to the CPU path and falling back to it, and it is not on the critical path of either dominant phase.

How much faster is a follow-up scan?

It is not a fixed percentage, which is why the calculator asks how much changed. A re-run with nothing changed still pays a floor of roughly 17 % of the first scan. From there the cost rises with the share of files that changed, and with the number of new commits, up to the cost of a cold scan when everything changed.