Five layers. Every number traceable.

How qa turns a running app into a verdict you can defend, and the math under each step. Every figure below runs the same formula the code does, cited by file.

The five layers

L1

Detectors

Three detectors read your app and write the canonical inputs. inventory maps routes (from Next.js, React Router, TanStack Router, Remix, SvelteKit, Astro, Nuxt and Expo Router, or from a list you declare), components, queries and mutations. classify runs the 23 static rules over every JavaScript and TypeScript source file. sweep drives Chromium through each route as each persona, at each width and in each mode.

writes .verify/{inventory,classify,sweep}.json

Node owns everything browser-facing: navigation, auth state, personas, widths, forced states, journeys, network and DOM observation.

L2

Rust kernel

The kernel folds every observation into one graph, then scores routes and plans what to measure next. Identical inputs produce identical bytes: no hash-map iteration ever reaches a written file.

writes nodes.jsonl · edges.jsonl · run-manifest.json (sha256 of each)

An unmeasured dimension is a first-class value, never a zero. The static plan sets reached = compared = agreed = 0 rather than inventing runtime counts.

L3

CLI

Eight commands. qa setup gets a fresh install ready. The rest archive every run, rank findings into lanes, and write one brief at a time for the coding agent. Exit codes are the contract: 0 clean within scope, 1 gating findings, 2 could not run.

qa setup · init · verify · measure · show · run · fix · dev

The CLI never launches an AI agent or edits your source. It measures; the host agent fixes.

L4

Marathon

qa run --hours N walks eight phases: serve, preflight, study, baseline, deep, handoff, closing and report. State is written before every transition, so a run can resume after a crash.

writes .verify/marathon.json before every transition

At handoff the host agent repairs and commits. The gate then reruns: exit 0 escalates depth, and exit 2 stops and names the blocker.

L5

Proving the instrument

The instrument fails more often than the app, so it is tested harder. Every static rule must have a planted instance, or the suite fails. On every run, a generated held-out corpus scores the detectors at precision ≥ 0.95 and recall ≥ 0.9.

qa dev selftest --fast · --all

A check that cannot fail is not a check. Every guard is shown failing before it is trusted.

Rendered live with WebGL. Amber marks a finding: the brief in L3, the handoff in L4, a planted defect caught in L5.

WebGL is unavailable here. The five layers, bottom to top: detectors → Rust kernel → CLI → marathon → proving the instrument.

The math, interactive.

Six figures, each one claim from the code. Move the controls; nothing here is an illustration of an idea the tool does not implement.

Fig. 1

Route confidence never punishes what was not measured

From agent/src/scoring.rs. Confidence blends four measured dimensions at fixed weights. When a route was never swept, the runtime weight leaves the denominator instead of counting as a zero.

static coverage = 0.45 + 0.15·min(queries, 3)

confidence ·
qa
if unmeasured counted as 0
  • static · 0.25
  • runtime · 0.35
  • state · 0.20
  • role · 0.20

Fig. 2

The obligation space, and a funnel that cannot be inflated

From coherence-math.md §12 and §14. What could be checked is a product of axes. Each stage below can only shrink, and a static plan reports runtime stages as unmeasured instead of inventing them.

bound

Fig. 3

Unknown is not false

From coherence-math.md §6. Whether an invalidation covers a cache key is three-valued: true, false, or unknown when a helper cannot be resolved. Collapse unknown into false and the tool starts proving things it never saw.

Fig. 4

The planner is greedy, and says so

From coherence-math.md §17 and its counterexample. qa spends a budget on whichever candidate covers the most obligation weight per unit cost. That is fast, and sometimes wrong. The exact optimum is computed here by trying every subset.

cand.costweightw / costpicked by

Fig. 5

One missing invalidation, many symptoms

From coherence-math.md, the join witness J[m,s] = 1 iff Σₑ W[m,e]·R[s,e] > 0. Toggle what each mutation writes and what each surface reads. Pick a mutation to see every surface that must agree after it.

Fig. 6

Jev screens with a band, not a coin flip

From skills/qa/jev.md. Jev, TypeSafe's System One model, returns a calibrated probability per check. In 15 identical runs, one question wandered between 0.43 and 0.53. A 0.5 cutoff flips on noise; the band does not.

droppedread queuelead → verify 0.5 00.300.701

Replay values are synthetic, drawn inside the 0.43–0.53 spread recorded in jev.md. The band edges are starting points until tuned on labelled verdicts.