How qa turns a running app into a verdict you can defend, and the math under each step. Every figure below runs the same formula the code does, cited by file.
The five layers
L1
Detectors
Three detectors read your app and write the canonical inputs. inventory maps routes (from Next.js, React Router, TanStack Router, Remix, SvelteKit, Astro, Nuxt and Expo Router, or from a list you declare), components, queries and mutations. classify runs the 23 static rules over every JavaScript and TypeScript source file. sweep drives Chromium through each route as each persona, at each width and in each mode.
writes .verify/{inventory,classify,sweep}.json
Node owns everything browser-facing: navigation, auth state, personas, widths, forced states, journeys, network and DOM observation.
L2
Rust kernel
The kernel folds every observation into one graph, then scores routes and plans what to measure next. Identical inputs produce identical bytes: no hash-map iteration ever reaches a written file.
writes nodes.jsonl · edges.jsonl · run-manifest.json (sha256 of each)
An unmeasured dimension is a first-class value, never a zero. The static plan sets reached = compared = agreed = 0 rather than inventing runtime counts.
L3
CLI
Eight commands. qa setup gets a fresh install ready. The rest archive every run, rank findings into lanes, and write one brief at a time for the coding agent. Exit codes are the contract: 0 clean within scope, 1 gating findings, 2 could not run.
qa setup · init · verify · measure · show · run · fix · dev
The CLI never launches an AI agent or edits your source. It measures; the host agent fixes.
L4
Marathon
qa run --hours N walks eight phases: serve, preflight, study, baseline, deep, handoff, closing and report. State is written before every transition, so a run can resume after a crash.
writes .verify/marathon.json before every transition
At handoff the host agent repairs and commits. The gate then reruns: exit 0 escalates depth, and exit 2 stops and names the blocker.
L5
Proving the instrument
The instrument fails more often than the app, so it is tested harder. Every static rule must have a planted instance, or the suite fails. On every run, a generated held-out corpus scores the detectors at precision ≥ 0.95 and recall ≥ 0.9.
qa dev selftest --fast · --all
A check that cannot fail is not a check. Every guard is shown failing before it is trusted.
Rendered live with WebGL. Amber marks a finding: the brief in L3, the handoff in L4, a planted defect caught in L5.
WebGL is unavailable here. The five layers, bottom to top: detectors → Rust kernel → CLI → marathon → proving the instrument.
A
The math, interactive.
Six figures, each one claim from the code. Move the controls; nothing here is an illustration of an idea the tool does not implement.
Fig. 1
Route confidence never punishes what was not measured
From agent/src/scoring.rs. Confidence blends four measured dimensions at fixed weights. When a route was never swept, the runtime weight leaves the denominator instead of counting as a zero.
static coverage = 0.45 + 0.15·min(queries, 3)
confidence ·
qa
if unmeasured counted as 0
static · 0.25
runtime · 0.35
state · 0.20
role · 0.20
Fig. 2
The obligation space, and a funnel that cannot be inflated
From coherence-math.md §12 and §14. What could be checked is a product of axes. Each stage below can only shrink, and a static plan reports runtime stages as unmeasured instead of inventing them.
bound
Fig. 3
Unknown is not false
From coherence-math.md §6. Whether an invalidation covers a cache key is three-valued: true, false, or unknown when a helper cannot be resolved. Collapse unknown into false and the tool starts proving things it never saw.
Fig. 4
The planner is greedy, and says so
From coherence-math.md §17 and its counterexample. qa spends a budget on whichever candidate covers the most obligation weight per unit cost. That is fast, and sometimes wrong. The exact optimum is computed here by trying every subset.
cand.
cost
weight
w / cost
picked by
Fig. 5
One missing invalidation, many symptoms
From coherence-math.md, the join witness J[m,s] = 1 iff Σₑ W[m,e]·R[s,e] > 0. Toggle what each mutation writes and what each surface reads. Pick a mutation to see every surface that must agree after it.
Fig. 6
Jev screens with a band, not a coin flip
From skills/qa/jev.md. Jev, TypeSafe's System One model, returns a calibrated probability per check. In 15 identical runs, one question wandered between 0.43 and 0.53. A 0.5 cutoff flips on noise; the band does not.
Replay values are synthetic, drawn inside the 0.43–0.53 spread recorded in jev.md. The band edges are starting points until tuned on labelled verdicts.