test-strategy
test-strategy
When an agent or user needs to decide what a change must prove, which test lane to run, or whether an existing check is real. Also use when the user says "what should I test," "which tests do I run," "is this covered," "why is the suite slow," "did the tests actually run," "add a test for this," or "this passed but I don't trust it." Use this whenever the question is about proof rather than code. For measuring runtime, see perf-audit. For reading a diff, see diff-review. For whether the proof is enough to ship, see release-gate.
Test Strategy
You decide what a change must prove and whether the proof can fail. A green
suite that cannot go red is theatre.
The loop, in order
Write the red-first test. For a fix, name the test that goes RED on the
old code and GREEN on the new. State that order explicitly. A test added
after a fix that passes on both is documentation, not a gate.
Pick the lane by the MOMENT, and say which. The suite is ~1317 files.
Its cost is 113.55s uncontended and 421.43s contended — measured
2026-09-11/13 by vitest's own Duration on a suite that changed 1.7%
between those runs. Quote the floor; the spread is queueing, not tests.
(The older "~1135 files / 87.3s" is stale — 2026-09-01.)
| Moment |
Run |
Why |
| edit |
verify:fast |
54s when it gets a slot, 837s when it does not (measured 2026-09-13, same lane, deploy-runs.json). Nothing cheaper can still go red. |
| land |
the fast gate, in the branch's worktree, with dev merged in |
a red here has three owners — branch, dev, or the seam |
| review |
verify (FULL) |
the only lane that counts before a ship |
| deploy |
nothing new — the FULL lane's memo, reported REUSED |
release.sh ship must never re-run what promote proved |
Take FULL, never fast, when: it is the last cycle of a plan · the change
touched schema/, packages/sdk/, or auth/authority code · a file was
renamed or deleted · a fast pass just went red and you are confirming the
fix. A fast pass is never reported as a full pass.
What a memo already answers, ask it instead of running it. A cache hit
needs no slot, so probing is free and queueing for one is not: measured
2026-09-13, a deploy gate queued 300s to deliver two cache HITs needing
0s of compute. TEST_CACHE_PROBE=1 / TEST_CACHE_KEY_ONLY=1 and
tsc-cached.sh --probe answer without running anything — see
gate-throughput § Admission control.
What is NEVER worth running on this box:
- The full suite to answer a question about one file. It costs 113–421s
and a 2GB slot; vitest related <file> plus the pins costs 152M.
- A re-run of a lane whose memo is already green for this exact tree —
that is the 300s above.
- The real-TypeDB lane on a saturated box. It is non-blocking by design
(typedb lane FAILED (rc=1) — non-blocking) and its failures track the
box, not the code: every run ≤138s was 22/22 GREEN and every run ≥206s
failed, measured 2026-09-13 across seven runs. Running it under load
manufactures red.
- Any gate you cannot wait for. A gate abandoned mid-run still holds its
slot until run_bounded reaps it.
Keep the fast lane's two invariants. Any change to it must preserve
both, or the lane becomes a liar:
- An empty diff falls back to the FULL suite — never to a pass. A lane
that goes green by selecting zero tests is worse than the slow one.
- The pinned suites always run.
vitest related walks the import graph;
config, parity, and boundary gates import nothing from what they guard —
one reads a .toml.
Apply the docblock rule. A test touching the DOM carries
// @vitest-environment jsdom on line 1. environment defaults to node
because only ~230 of ~1316 files need a DOM. Forget it and you get
ReferenceError: document is not defined immediately — loud and
self-correcting, which is why there is no central list.
Prove the checker can fail. Break the thing the gate guards and confirm
RED before trusting it. A checker that stays green against gutted code has
proven nothing about the code.
Hard rules
Never reach for environmentMatchGlobs — vitest 4 removed it and ignores
it silently, so the config looks right and does nothing.
Never mock TypeDB in integration tests — use real TypeDB or skip.
An exit code is not the gate result. Read the last line. (0 test) is a
load failure. 144 is unrun. 141 is SIGPIPE from producer | grep -q under
pipefail — which returns 141 when it matches, on a large producer.
An unrun gate is not a pass — AND it is not a failure. These two lines
together are a starved gate, not a broken one:
[gate-run] verify-fast: no slot after 1800s — machine saturated.
✗ typecheck FAILED
The check never ran: gate_slot timed out at GOVERN_QUEUE_WAIT (default
1800s), the wrapper exited 1, and the caller printed its failure line for a
command that was never executed. Measured 2026-09-13: 8 of 10 land runs
ended this way, 8,413s of wall clock with ZERO phases recorded — the
phases: [] in deploy-runs.json is the tell, because a gate that ran
always leaves a phase. Report it as unrun, name the wait that bounded it,
and re-run with GOVERN_QUEUE_WAIT=5400 rather than believing the word
FAILED. Never file a finding against code on this evidence.
Cannot run
You apply that last rule to everyone else's gates. It applies to you.
Say cannot-run — never "the lane is fine" — when the diff is empty or was
never handed over, when the suite could not load ((0 test), 144, 141), or when
the box was too saturated to trust a timing. Return
{ ok: false, reason: "<which of those, and the line that says so>" } and close
with warn. A lane you could not observe is not a lane you approved.
Out of scope
- Tuning worker counts or measuring wall clock — see perf-audit.
- Writing the implementation the test covers.
- Declaring a release shippable — see release-gate.