← Skills

test-strategy

test-strategy

When an agent or user needs to decide what a change must prove, which test lane to run, or whether an existing check is real. Also use when the user says "what should I test," "which tests do I run," "is this covered," "why is the suite slow," "did the tests actually run," "add a test for this," or "this passed but I don't trust it." Use this whenever the question is about proof rather than code. For measuring runtime, see perf-audit. For reading a diff, see diff-review. For whether the proof is enough to ship, see release-gate.

Test Strategy

You decide what a change must prove and whether the proof can fail. A green
suite that cannot go red is theatre.

The loop, in order

  1. Write the red-first test. For a fix, name the test that goes RED on the
    old code and GREEN on the new. State that order explicitly. A test added
    after a fix that passes on both is documentation, not a gate.

  2. Pick the lane by the MOMENT, and say which. The suite is ~1317 files.
    Its cost is 113.55s uncontended and 421.43s contended — measured
    2026-09-11/13 by vitest's own Duration on a suite that changed 1.7%
    between those runs. Quote the floor; the spread is queueing, not tests.
    (The older "~1135 files / 87.3s" is stale — 2026-09-01.)

    Moment Run Why
    edit verify:fast 54s when it gets a slot, 837s when it does not (measured 2026-09-13, same lane, deploy-runs.json). Nothing cheaper can still go red.
    land the fast gate, in the branch's worktree, with dev merged in a red here has three owners — branch, dev, or the seam
    review verify (FULL) the only lane that counts before a ship
    deploy nothing new — the FULL lane's memo, reported REUSED release.sh ship must never re-run what promote proved

    Take FULL, never fast, when: it is the last cycle of a plan · the change
    touched schema/, packages/sdk/, or auth/authority code · a file was
    renamed or deleted · a fast pass just went red and you are confirming the
    fix. A fast pass is never reported as a full pass.

    What a memo already answers, ask it instead of running it. A cache hit
    needs no slot, so probing is free and queueing for one is not: measured
    2026-09-13, a deploy gate queued 300s to deliver two cache HITs needing
    0s of compute. TEST_CACHE_PROBE=1 / TEST_CACHE_KEY_ONLY=1 and
    tsc-cached.sh --probe answer without running anything — see
    gate-throughput § Admission control.

    What is NEVER worth running on this box:

    • The full suite to answer a question about one file. It costs 113–421s
      and a 2GB slot; vitest related <file> plus the pins costs 152M.
    • A re-run of a lane whose memo is already green for this exact tree —
      that is the 300s above.
    • The real-TypeDB lane on a saturated box. It is non-blocking by design
      (typedb lane FAILED (rc=1) — non-blocking) and its failures track the
      box, not the code: every run ≤138s was 22/22 GREEN and every run ≥206s
      failed, measured 2026-09-13 across seven runs. Running it under load
      manufactures red.
    • Any gate you cannot wait for. A gate abandoned mid-run still holds its
      slot until run_bounded reaps it.
  3. Keep the fast lane's two invariants. Any change to it must preserve
    both, or the lane becomes a liar:

    • An empty diff falls back to the FULL suite — never to a pass. A lane
      that goes green by selecting zero tests is worse than the slow one.
    • The pinned suites always run. vitest related walks the import graph;
      config, parity, and boundary gates import nothing from what they guard —
      one reads a .toml.
  4. Apply the docblock rule. A test touching the DOM carries
    // @vitest-environment jsdom on line 1. environment defaults to node
    because only ~230 of ~1316 files need a DOM. Forget it and you get
    ReferenceError: document is not defined immediately — loud and
    self-correcting, which is why there is no central list.

  5. Prove the checker can fail. Break the thing the gate guards and confirm
    RED before trusting it. A checker that stays green against gutted code has
    proven nothing about the code.

Hard rules

  • Never reach for environmentMatchGlobs — vitest 4 removed it and ignores
    it silently, so the config looks right and does nothing.

  • Never mock TypeDB in integration tests — use real TypeDB or skip.

  • An exit code is not the gate result. Read the last line. (0 test) is a
    load failure. 144 is unrun. 141 is SIGPIPE from producer | grep -q under
    pipefail — which returns 141 when it matches, on a large producer.

  • An unrun gate is not a pass — AND it is not a failure. These two lines
    together are a starved gate, not a broken one:

    [gate-run] verify-fast: no slot after 1800s — machine saturated.
    ✗ typecheck FAILED
    

    The check never ran: gate_slot timed out at GOVERN_QUEUE_WAIT (default
    1800s), the wrapper exited 1, and the caller printed its failure line for a
    command that was never executed. Measured 2026-09-13: 8 of 10 land runs
    ended this way, 8,413s of wall clock with ZERO phases recorded
    — the
    phases: [] in deploy-runs.json is the tell, because a gate that ran
    always leaves a phase. Report it as unrun, name the wait that bounded it,
    and re-run with GOVERN_QUEUE_WAIT=5400 rather than believing the word
    FAILED. Never file a finding against code on this evidence.

Cannot run

You apply that last rule to everyone else's gates. It applies to you.

Say cannot-run — never "the lane is fine" — when the diff is empty or was
never handed over, when the suite could not load ((0 test), 144, 141), or when
the box was too saturated to trust a timing. Return
{ ok: false, reason: "<which of those, and the line that says so>" } and close
with warn. A lane you could not observe is not a lane you approved.

Out of scope

  • Tuning worker counts or measuring wall clock — see perf-audit.
  • Writing the implementation the test covers.
  • Declaring a release shippable — see release-gate.