Skip to content

QA and UAT — how testing is run ​

Read before changing the QA suite

  1. This page · 2. the case library tools/qa/cases-*.mjs (the SSOT; the workbook is a BUILD OUTPUT of npm run qa:workbook) · 3. Test suite. Gates: check:qa-suite (shape, uniqueness, workbook freshness — never whether a case is GOOD). Never hand-edit a definition column in the spreadsheet; the next regeneration discards it.

THE ONE-LINE ANSWER

Case definitions live in git. Execution lives in a spreadsheet. The spreadsheet is generated from git, never hand-authored. Everything below follows from that split.

Why not "one Excel workbook as the source of truth" ​

The owner asked for a single Google Drive workbook as the SSOT for QA, and asked for honest pushback. Here it is: a spreadsheet is the right tool for execution and the wrong tool for definitions, and using it for both is what produces the problems it was meant to solve.

The stated fearA spreadsheet-only SSOTDefinitions in git
Lost test casesOne accidental sort or delete is unrecoverable; no history per caseEvery case has a commit, an author and a diff
Different testers keeping different versionsCopies proliferate the moment someone downloads itOne branch, one review, one merge
Missing regression coverageNothing can compare the suite against the productA gate can fail a PR that changes a screen and not its cases
Outdated casesA case describing a screen that changed looks identical to a current oneThe case sits beside the code; the diff makes staleness visible
DuplicatesTwo people add the same scenario with different idscheck:qa-suite Q1 refuses a duplicate id

⚠ And the failure already happened here. The previous suite was delivered in a chat message. Fifty-two of its sixty-six cases were never written down anywhere, so a context reset destroyed them, and the owner's own device findings had to be recovered from a screenshot to be analysed. That is the "lost test cases" risk, realised, before any tooling decision was made.

What a spreadsheet is genuinely better at, and why we keep it: many people editing at once, per-row ownership, dropdown discipline, filtering, and a shape non-engineers already know. Execution is exactly that shape. So:

  • tools/qa/cases-*.mjs — the case library. Engineering-owned. Versioned. Reviewed.
  • qr-setu-qa-workbook.xlsx — generated by npm run qa:workbook. QA-owned. One copy per cycle.
  • npm run check:qa-suite — refuses a malformed case, a duplicate id, and a stale workbook.

When this stops being enough: the day three or more people execute concurrently on more than one platform, or the day we need per-case history across releases. Then the answer is a real test management tool (TestRail, Qase, Xray if we ever adopt Jira) — and because the library is already structured data, exporting into one is a script, not a re-authoring project. We are not there yet, and adopting one now would add a tool nobody has time to administer.

1 · How the master repository is structured ​

tools/qa/
  cases-auth.mjs        startup · sign up · sign in · OTP · address · app lock · session
  cases-consumer.mjs    Home · navigation · account · empty surfaces · QR makers
  cases-biodata.mjs     create · editor · visibility · publish · link/QR · public page · deep links
  cases-platform.mjs    business · authorization · backend · UI/UX conformance · regression
  emit.mjs              → the workbook + the portal page
  hash.mjs              the freshness stamp, ONE implementation

Every case carries: id · module · persona · priority · type · state · pre · steps[{do,expect}] · data · notes · ref.

⚠ steps is a list of PAIRS, and the gate refuses a step with an action and no expectation. That single constraint is what makes "verify login works" unrepresentable in this suite.

⚠⚠ state is the field that stops noise, and it is measured against the build rather than assumed. ready · blocked-not-built · blocked-not-deployed · known-defect · watch. Without it a tester files a dozen defects for features that were never built and misses the real ones, and the suite loses its credibility in one cycle.

2 · Ownership ​

RoleOwnsNever does
DeveloperAdding and updating cases for what they build, in the SAME change as the code. Running the developer-validation checklist before handing over.Marks their own work QA-passed
QAExecution, defect filing, retest, the cycle workbook, the regression selectionEdits a definition column in the spreadsheet
Owner / UATJourney-level acceptance, product judgement, release sign-offDiscovers basic visual defects — that is a process failure upstream

The rule that makes this real: a developer hands over a build with the cases already updated. QA's first question is "what changed and which cases cover it", and the answer must exist before the build does.

3 · How a design change becomes test cases ​

Approved design → design:spec extraction → parity contract rows → QA cases → implementation

⚠ The contract comes BEFORE the code, and the cases come from the contract. npm run design:spec prints every styled element and every binding in an artboard section, so the enumeration is mechanical rather than remembered. Each enumerated state becomes a parity-contract row; each row a tester can observe becomes a case.

Measured justification: on one screen, seven of eleven visual defects were plainly "the design says X and the code says Y" — radius 24 vs 40, e1 vs e2, a 72 px QR vs 64, missing letter-spacing, the wrong font face, a missing icon, and a line the design renders nowhere. All of them passed type-check, lint and the parity gate. The one time the artboard was extracted mechanically, ten defects fell out in a single pass.

4 · Regression coverage ​

  • Every case tagged Regression runs every cycle. No exceptions and no sampling.
  • A fixed defect becomes a case. Not a note in a ticket — a numbered case with steps, in the library, forever. BIO-005 is the worked example: it reproduces the exact ordering that poisoned an idempotency key, and it will now be run on every build for the life of the product.
  • A fix is re-verified on every surface, not just the one that reported it. Three consecutive defects here were introduced by fixes verified only where the symptom appeared.
  • Anything touching packages/tokens/**, apps/*/src/ui/**, theme plumbing or press handling re-runs REG-002 on both personas, because those are shared and the blast radius is all surfaces.

5 · Defect traceability ​

Every Fail carries a Defect ID, and every defect carries a case id. A failure with no defect id is treated as not executed — that is deliberate, because an untraceable failure cannot be retested and cannot be proven fixed.

The defect record is the tracker (dev-tracker/tracker.md, permanent QRS-###). The workbook's Ref column already carries the tracker id for cases that exist because of a defect, so the link runs both ways.

6 · Blocked cases and environment problems ​

Blocked is not Fail, and conflating them is the most common way a QA report misleads. Fail means the product is wrong. Blocked means we could not find out.

Blocked requires a reason in Comments, and the reason names which of these it is: the feature is not built · a dependency is not deployed · the environment is wrong · the device or account is missing · another defect blocks the path. Blocked cases are pre-seeded by the generator for the blocked-* states, so a tester never has to decide those.

Two cases run first, every cycle, and they exist to stop a whole cycle being void:

  • AND-004 — adb shell pm dump … | grep lastUpdateTime proves the device runs the intended binary. A whole session was once spent verifying a fix against the previous day's bundle.
  • API-005 — npm run check:ef-drift proves the backend matches the repo. Testing an app against a drifted backend produces failures nobody can act on.

7 · UI/UX validation, incorporated early instead of discovered late ​

This is the owner's sharpest observation, and the honest diagnosis is that the current process reviews UI at the wrong moment: implementation happens from a remembered design, and the owner is the first person to compare it against the artboard. That makes the product owner the visual QA function, which is both expensive and demoralising.

The fix is three controls, in this order:

  1. Extract, never transcribe. npm run design:spec on the artboard before writing the screen. Radius, elevation, type, spacing and every binding come out as numbers.
  2. A parity contract per screen, enumerated from the design's own states — every row gets an explicit pass · gap · blocked, never blank. npm run check:design-parity gates that.
  3. reachableFrom on every pass row. ⚠ This is the standing recommendation and it is not yet mandatory. Evidence in a component proves the behaviour was written; a token in a route proves something renders it. A green contract once covered a PIN gate that no user could reach for a month, and six Home CTAs shipped completely dead with every gate green.

⚠ What no script can do, stated so nobody expects it: nothing can diff an artboard against React Native and decide they agree. UX-001 is the human comparison, and it is written as a method — same width, side by side, in a fixed order — because an unstructured "does it look right" pass is what lets a radius of 24 read as 40.

8 · Preventing duplicate, missing and outdated cases ​

RiskControl
Duplicatecheck:qa-suite Q1 — a duplicate id fails the gate
MalformedQ2 — no preconditions, no steps, or a step without an expected result fails
OrphanedQ3 — a case whose prefix maps to no sheet would vanish from the workbook; refused
OutdatedQ4 — the workbook carries a hash of the library; an edit without a regeneration fails
Missing⚠ Not automatable, and saying otherwise would be false. Coverage is judged at contract time: a design state with no case is visible in the contract, not in the suite

9 · Release readiness ​

A build is signed off when all of these are true, each produced by a command or a filled column, never asserted:

  1. Every P0 case executed. Not "sampled" — executed.
  2. Zero open P0/P1 failures.
  3. Nothing left Not Run. Blocked is acceptable only with a stated reason.
  4. Every Fail has a Defect ID, and every fixed defect has a Retest Status of Passed.
  5. Every Regression-typed case run on this build.
  6. check:qa-suite, check:ef-drift, check:design-parity, check:screens, lint, type-check and the unit suites green, with output pasted.
  7. Parity: the same cycle executed on iOS. ⚠ There is no exception path — a feature ships on Android, iOS and web together. This cycle is Android only, so the current suite cannot sign off a release on its own.
  8. Owner UAT on the journey, not the screens: sign up → address → create → fill → publish → the recipient reads it.

⚠ The Summary tab computes 1-4 from the dropdowns. It cannot compute 5-8, and it says so.

10 · How this scales ​

  • A new feature adds a case file and a sheet group. Two lines in emit.mjs; the workbook grows a tab. That is the whole cost.
  • A new persona (Enterprise) adds a persona value and its own files. The three-user-category rule already says these are different products; the suite should mirror that rather than blend it.
  • When a second platform runs concurrently, the cycle workbook gets a platform column and one copy per platform — not one merged sheet, because a merged sheet hides which surface failed.
  • The migration path out is already open. The library is structured data, so exporting to a test management tool is a script. Adopt one when concurrent execution, not case count, becomes the bottleneck.

The lifecycle, end to end ​

Approved design
  → design:spec extraction            (mechanical, not remembered)
  → parity contract rows              (every state, explicit verdict)
  → QA cases added in the SAME change as the code
  → implementation
  → developer validation              (AND-004 + API-005 + the touched cases)
  → QA execution in the cycle workbook
  → defect logged with a case id
  → fix
  → retest against the same case id
  → regression sweep
  → owner UAT on the journey
  → release sign-off against the nine criteria above

THE ONE THING THAT WILL DECIDE WHETHER THIS WORKS

Not the tooling. Whether cases are updated in the same change as the code. Every control here is downstream of that, and the only enforcement that exists today is review discipline plus Q4. Making check:docs-impact demand a case change when a screen changes is the obvious next gate, and it is not built.