Docs

Run once, read the evidence, then wire qa into CI and into your agent.

Install

qa is one executable. The script downloads it, checks its sha256 against the manifest and installs it to ~/.local/bin/qa. No sudo. Running it again updates in place.

$ curl -fsSL https://qakraken.com/install.sh | sh
$ qa setup
$ npx --yes --allow-remote=root --package=https://qakraken.com/qa.tgz qa setup
$ bunx --bun --package=https://qakraken.com/qa.tgz qa setup

Download the archive for your platform, then check it and unpack it.

PlatformArchiveSizesha256
darwin-arm64qa-1.0.0-darwin-arm64.tar.gz43.3 MB5d445192d15fa6b7b5e04b0d0a559c24470c64c568928df2ed13a7d488181ba0
darwin-x64qa-1.0.0-darwin-x64.tar.gz46.7 MB0a91eb801a5ca21c1cf529c121f0a8ded32b1a1ddd8af592128e7bf23c6dca37
linux-arm64qa-1.0.0-linux-arm64.tar.gz47.9 MBd92264e1fe239deda0047c0efe1cf33357b34de7030ebccd41e022e7f0c0c72a
linux-x64qa-1.0.0-linux-x64.tar.gz47.9 MBa89cf1cf461a39491bdb91d094719c3f4eac90c3b378b41880eb5e25d58e5bf7
$ echo "<sha256>  <archive>" | shasum -a 256 -c
$ tar -xzf <archive> && mkdir -p ~/.local/bin && mv qa-1.0.0-*/qa ~/.local/bin/qa
$ qa setup

Manifest: manifest.json. On macOS a file saved by a browser is quarantined; run xattr -d com.apple.quarantine ~/.local/bin/qa. A curl download is not.

Install qa on this machine, then verify one app repo.
1. Run: curl -fsSL https://qakraken.com/install.sh | sh
   It downloads one executable, checks its sha256 against the manifest, and installs it to ~/.local/bin/qa. No sudo. If ~/.local/bin is not on PATH, add it.
2. Run: qa setup
   It installs the browser qa drives. This is the only step that needs Node.
3. From the app repo: qa init <repo>
   then: qa verify <repo> --preflight
Report each exit code. 0 is clean within the measured scope, 1 is findings, 2 is could not run and is never a pass. Do not edit the app's source to make a check pass.
macOS arm64 and x64Linux x64 and arm64, glibc with libatomic1No Windows. No Alpine.Node for npx and browser setup; Bun runs the shim directlyUpdate: run the install command againUninstall: remove ~/.local/bin/qa
NeedsVersionFor
macOSarm64 or x64the executable
Linuxx64 or arm64, glibc, libatomic1the executable. No Alpine (musl). No Windows.
Node.js18 or newerqa setup only, which installs the browser
Chromiuminstalled by qa setupthe runtime sweep only
Your appa web app served at a URLthe runtime sweep, in any framework
Your sourceJavaScript or TypeScript (.js .jsx .ts .tsx)the static rules and the data-layer model

What it reads

PartCovers
RoutesNext.js (app and pages routers), React Router, TanStack Router, Remix, SvelteKit, Astro, Nuxt and Vue Router, Expo Router, hash routing. Any other app: declare the routes in verify.routes.json.
Static rules23 rules over .js, .jsx, .ts and .tsx. Some rules are React-specific; two are Next.js app router only.
Data layerTanStack Query and SWR: unread error and loading state, shared cache keys, and writes that leave another page stale.
BrowserChromium through Playwright, per route, width, colour scheme and persona, including apps behind a service worker. A subset of accessibility rules and in-page performance observers (LCP, layout shift, long tasks).
Sign-inFive evidence lanes. You provide the launcher; qa never logs in on its own.
OutputJSON under .verify/, SARIF, GitHub annotations, JUnit, and a local MCP server.

Next

Run your first session, read the exit codes, and wire qa into your agent and CI on Using qa.

CLI reference

Generated from qa help --json at build time, so it matches the binary you install.

Prepare a run

qa setup
qa setup [--dry-run] [--yes] [--json] [--no-browser] [--no-skills] [--mcp]

Gets a fresh install ready. Run it once after installing qa. Each step prints what it will do first; every step is safe to run again. 1 Playwright looks where the sweep looks (repo, .playwright, cwd, PLAYWRIGHT_HOME, the per-user home). If missing, installs playwright-core into the per-user home with npm, then runs its CLI to fetch Chromium. Launches Chromium headless once and prints its version and path. Needs node and npm on PATH; if they are absent it prints the exact commands and exits 2. 2 Jev key checks whether TYPESAFE_API_KEY or TYPESAFE_JEV_KEY is set. Presence only; the value is never printed. The CLI never needs it; only the HUNT screens in the qa skill do. Absent is not a failure. 3 skills copies the packaged skills into ~/.claude/skills and ~/.codex/skills for the agents that exist. A directory you made is never touched; a differing copy is reported, not replaced. 4 MCP with --mcp, registers qa with Claude Code and Codex when their CLIs exist (printed first, run only with --yes); otherwise prints the config snippet. Without --mcp, a one-line hint. 5 next prints the first commands to run. --dry-run print every step's plan, change nothing --yes run without asking. Required when there is no terminal. --no-browser skip step 1 --no-skills skip step 3 --json one document on stdout: { ok, dryRun, steps[], next[] } Exit codes: 0 everything requested is ready; 2 invalid usage, no terminal without --yes or --dry-run, or a requested step could not complete. A skipped or failed step is never reported as ready. Never 1.

qa init
qa init <repo> [flags]
  --study|--analysis|--roles|--crud|--ignore|--workflow|--hook|--all

Scaffolds declarations without overwriting them. --study writes source-derived research; --analysis additionally writes .verify/analysis.json and analysis.md with an evidence ledger, feature/context/risk/story/mutation/proof/fix view. It never invents business requirements or executes mutations.

qa verify
qa verify <repo> [any verify.sh flag] [--lane LANE] [--autonomous] [--json]
       qa verify <repo> [--no-archive] [--keep N] [--link-adapters
       crud,experiments]
       qa verify <repo> --base URL [--console-allow REGEX] [--action-ms MS]

Runs verify.sh with the flags given, unchanged. verify.sh --help is the authority on those; two that matter to a noisy app: --console-allow REGEX declares a console error expected (matches are counted and the patterns republished under consoleAllowed, never dropped in silence; an unparseable pattern exits 2), and --action-ms MS is one locator action's budget -- a visibility wait, a click, a menu dismissal -- default 2000.

Then opens a private pending run directory and atomically writes immutable report snapshots plus normalized roles and the exact validated CRUD snapshots, with their sha256 digests. The agent receives only those explicit paths and never reads the live role or CRUD files. It always attempts a compatible qa-agent (a verified prebuilt binary, or a cargo build fallback) with those explicit paths, the run ID, lane and mode via --emit-result -- optional evidence by default (unmeasured when no compatible binary exists), never gating unless --autonomous makes it required. Only after the returned digests and schema are verified does the pending directory publish to .verify/runs/<id>/: a reader sees no run or a complete run, never a partial one. run.json carries the embedded result under agent, plus lane and runId; the archive also gets agent-result.json (its canonical bytes). Adapter evidence (qa measure crud, qa measure experiments) is NEVER folded in automatically: name the adapters explicitly with --link-adapters, and their findings travel with their own denominator rather than entering this run's counts. Appends to .verify/trajectory.jsonl and prints the triage. Exit is derived: 2 dominates (the gate itself invalid, OR --autonomous with an absent/invalid agent), else 1 if the gate found gating evidence, else 0. --json output separates gate (verify.sh's own exit) from agent (the embedded result) so a consumer never conflates the two verdicts. --no-archive runs the same snapshot and agent sequence into a private pending directory and still emits gate/agent on stdout (--no-archive --autonomous still enforces require-agent) but publishes nothing: the pending directory is discarded, only after the result is validated, so a failure still leaves a diagnosis on stdout.

--autonomous also produces the canonical three-file graph contract (nodes.jsonl, edges.jsonl, run-manifest.json) via the cargo-built qa-agent into .verify/autonomous/ and validates its manifest. This is reported under the JSON output's autonomous field, distinct from agent. Needs cargo; absent cargo is reported unmeasured. A required kernel or manifest failure is a hard exit 2. The earlier 42-file surface is retired.

Measure the app

qa measurekind: coherence, crud, experiments
qa measure <kind> <repo> ...   kind: coherence, crud, experiments

Dispatches to one of the three deep-measurement adapters. Each kind keeps its own flags and subwords. crud and coherence --mutate perform real, declared mutation writes under the repo's own opt-in (verify.crud.json); every other kind is read-only. See qa help measure coherence / measure crud / measure experiments for what each one measures and its exit codes.

Read results

qa show
qa show [kind] <repo> ...
  kind: triage (default), next, report, ledger, analysis, graph

Reads results. With no kind, prints ranked findings and coverage (triage). report's own sub-verbs stay reachable both nested (qa show report diff|runs|ci <repo>) and directly (qa show diff|runs|ci <repo>). See qa help show triage / show next / show report / show ledger for each one's flags, output, and exit codes. qa show analysis <repo> renders the generated analysis; --format json emits its structured artifact.

Fix and automate

qa run
qa run <repo> [--base URL | --serve] [--hours N]
              [--lane LANE] [--lanes N] [--widths 390,1440] [--no-states]
              [--no-dark]
              [--stability-runs 3] [--checkpoint-minutes 45] [--resume]
              [--fresh] [--keep-serving] [--no-autonomous]
              [--once] [--dry-run] [--json] [passthrough verify flags]
       qa run [kind] <repo> ...   kind: serve (default: below)

The one long-running "do everything" command. With no kind, every phase below is a child qa <cmd> process; this command reads what each one wrote to .verify/ and decides whether the next one runs: 0 --base is used as given. Otherwise this command builds and starts a production server itself (qa run serve --prod; its base becomes --base, stopped at the end unless --keep-serving). If that fails without an explicit --serve, the run is static-only and the report says so; with --serve the failure is INVALID. 1 qa verify --preflight -- exit 2 stops everything and prints its fix lines. When a --base is known, its probe of it decides --prod automatically for every phase below: prod -> add --prod, dev -> continue without it and say timing findings will be P3, unknown -> continue without it. 2 qa init --study; qa init --roles --ignore -- init never overwrites. With more than one principal and no roles file, init scaffolds one role per principal (each owning "/") and the run continues, saying so. 2b Writes turn on when verify.crud.json declares every flow "production": false and the base is loopback: the crud and coherence phases pass --mutate themselves. The sweep gets --mutate only when an app-owned journey declares mutates: true -- generic guessed replay is disabled, so without such a journey the flag could only invalidate the run, and the decision log says it was withheld. An explicit --lane wins; otherwise real-integration is selected when verify.auth-launcher.mjs exists and every persona in verify.roles.json declares "expect". Draft journeys whose every field and submit control resolved from source are promoted to .verify/journeys.mjs under that same mutation opt-in. Each decision and its reason is printed and lands in the exit report. 3 qa verify --ratchet (the baseline) -- exit 2 stops. No --base means a static-only run, said plainly in the report. 3b deep measurement, once, in two groups: Group A (concurrent -- read-only against the app, or writing only to an isolated target) is static coherence. Group B (strictly serial -- these can all write to the SAME dev database) is qa measure crud (under the mutation opt-in), qa measure experiments (qa run generates the Node-owned default queue from the baseline's Rust route priority when none exists, so it needs cargo; pass --queue <path> for a supplied plan; wall time is a quarter of what is left, at most 30 minutes), then runtime coherence (under its own repo-declared opt-in). Whatever could not run is a named line under "what was not measured" with how to enable it. Whichever adapter left measured evidence is folded into the very next qa verify via --link-adapters automatically. 4 qa verify --runs N on the baseline's runtime P0/P1 routes (skipped when there are none) -- ghost findings are written into .verify/loop.json's blocked list so the fix loop never chases them. 5 qa fix plan --lanes N (one worktree + brief per class, disjoint batches), then waits for the host agent to execute the briefs. --once hands off immediately with exit 1. 6 every chunk: qa show ledger, qa show report, one line appended to .verify/trajectory.jsonl. If something was fixed and time remains, qa verify runs again and qa show diff compares it -- new findings send it back to step 4 with whatever time is left. After the depth ladder is exhausted, runtime measurements repeat at --checkpoint-minutes cadence without --resume; source changes wake the coordinator sooner. Static-only runs wait for source changes. --once finishes after the first full pass. 7 exit report: .verify/marathon-report.md (routes total/clean/fixed with commit hashes/open fids/needs-reproduction ghost fids, the ratchet trajectory, every phase's command and exit code verbatim, what stayed unmeasured, which adapters got linked, the depth ladder's rungs with their --resume reused/rejected counts) plus a qa show report HTML. Its first word is PASS, FAIL, or INVALID, matching this command's own exit code. --status reads .verify/marathon.json and prints phase, running/stopped/finished, minutes left, base, lane, writes on/off, depth rung, last command and decision, the last checkpoint counts, wait state, and the report path -- for an agent polling a background run. Exit 0 running or finished, 1 stopped mid-run (resume it), 2 nothing has run here. A bare qa run <repo> on a repo whose .verify/marathon.json is not "done" resumes it -- this covers both a killed process and a --once handoff to open lanes; --fresh discards that state and starts over. A separate marathon.lock refuses a second live coordinator, including --fresh. --resume is still accepted and is a no-op. A recorded server that died is restarted only if --serve was given, otherwise that is INVALID. --dry-run prints the plan above with this run's actual flags and runs nothing. --mutate is added only by the declaration rule in step 2b, never by --hours or any other flag. Default --hours is 10, default --lanes is 4. Exit: 0 the final gate is clean, 1 it has findings (or lanes are ready/open, or time ran out with findings open), 2 a phase could not run.

--fixer, --fixer-protocol, --fixer-timeout, --measure-only, and --commit are retired for qa run: the host agent owns source changes and commits. Use --once for a copy-paste-ready brief. serve is reachable by name (qa run serve ...) -- see qa help run serve for its own flags and exit codes.

qa fix
qa fix <sub> ...
  sub: plan, gate, close, rebase, board (lanes) · fp add|list|rm

Parallel class-level remediation (qa fix plan <repo>, qa fix gate <class> <repo>, ...) and the cross-repo false-positive memory (qa fix fp add|list|rm ...). See qa help fix plan / fix fp for flags, output, and exit codes.

Maintain the harness

qa devkind: atlas, capabilities, selftest, validate, version
qa dev <kind> ...   kind: atlas, capabilities, selftest, validate, version

Maintaining the harness itself, not the target repo under test: qa dev atlas --strict, qa dev selftest --fast, qa dev validate <repo>, qa dev capabilities --json, qa dev version --json. See qa help dev atlas / dev capabilities / dev selftest / dev validate / dev version for flags and exit codes.

Safety

  • Never mutates by default. Writes need an explicit --mutate, a committed flow declaring "production": false, and a non-production base. Production-looking hosts are always refused.
  • No secrets in artifacts. Credentials, cookies, tokens, payload values and user data never enter a report, SARIF file or fixture.
  • Signed-in runs use your own launcher. Login state lives in a private lease that is deleted after every run.
  • Untrusted inputs fail closed with exit 2 and write nothing.

Report security issues privately to Arkash Jain. Use the Email link on his contact page. Include the affected version and steps to reproduce. Do not post secrets or security reports in public.