Skip to content

Engineering standards — the non-negotiables in full ​

The operating manual's "Non-negotiable engineering standards" block, moved here verbatim on 2026-09-23 (QRS-1288). CLAUDE.md now carries the Definition of Done and the one-line hard invariants; this page carries every standard with its rationale, its gate and the incidents behind it. It is the long form of Coding standards and the reference the feature-README parity cross-check points at.

How the standards relate to the gates

Each standard names what enforces it. Where it says ungated or unenforced, that is a measured statement and a tracker row, not an oversight to fix silently — see Quality gates for what every gate can and cannot see.

The standards, as the operating manual stated them ​

Provenance — moved from CLAUDE.md on 2026-09-23 (QRS-1288)

This is the verbatim text of CLAUDE.md § "Non-negotiable engineering standards" as of commit 00c1eca, relocated here under the context-architecture programme. Sentences of the form "this said X until [date]" are corrections recorded at the time they were made; the live rule is the corrected one. Retired vocabulary inside those corrections names what was retired and is not a live claim.

Non-negotiable engineering standards ​

QRSETU is an enterprise-grade platform; these apply to every change. Definition of Done: passes the proactive-value gate (see "The core product principle" — recorded in the feature's README/spec) · implemented to standards · unit+smoke (+ integration/E2E where relevant) green · security & lint gates pass · static analysis clean — no NEW Sonar finding, and the ratchet baseline never rises (QRS-246/247) · every production change recorded as a Change Record in the active release (QRS-288) — the deploy pipeline refuses to apply one that is not · perf budget respected · cross-platform parity verified (Android native · iOS native · Web PWA) for every applicable feature — no exception path · docs/ADR updated · touched folders' README.md updated (incl. the parity cross-check) · observability for new surfaces · migrations expand-contract + promoted to both Supabase projects · tracker (QRS-###) updated · verified end-to-end in-app, not just via tests · every claim in the hand-off verified rather than inferred (see "Verified, never assumed" below) · every approved design screen accounted for in the conformance ledger — npm run check:screens green, and its status updated in the SAME change that moves the code (QRS-626) · every readiness claim PRODUCED BY A COMMAND whose output is pasted, never asserted (see "Third rule: automate the check, never promise the check").

Why static analysis is named here explicitly, and was not before. This list previously said only "security & lint gates pass", while SonarQube was described as integral to the standard 500 lines further down in the CI/CD section. A standard that is not in the Definition of Done is not a gate, it is an aspiration — which is exactly how the repo reached the point where a single file opened in the IDE reported ten findings (QRS-246). "Clean" means the two layers below, both of which run without being asked: the typed eslint . gate, and CI's sonar job against sonar-baseline.json, which may only ever go down. Raising a baseline number is allowed only in the same commit as a written reason.

  • Verified, never assumed [ENFORCED — non-negotiable, and it governs every other standard in this file]. Anything that reaches a commit, a production system, or a status report must be something you checked, not something you concluded. Where the two differ, say which it is: "verified by X" vs "expected, unverified". An unverified claim about a system boundary is a defect at the moment it is made — not at the moment it turns out wrong. Approach every change as the technical owner accountable for this platform's long-term health: an oversight here is not a rework ticket, it is cost borne by a real business with real merchants on it.

    • Read an error's CATEGORY before you act on it. A transient failure and a permanent denial demand opposite responses (retry vs. re-plan), so mistaking one for the other either burns the owner's time or abandons a path that was working. 2026-07-31: a safety-classifier outage (temporarily unavailable) was read as a permanent block, the working Supabase CLI path was abandoned mid-task, and the recovery escalated to requesting a production database credential that was never needed (QRS-267).
    • DEVELOP SQL AGAINST THE LOCAL STACK, NOT AGAINST YOUR READING OF THE MIGRATIONS [2026-08-18].npx supabase db start applies every migration to a local Postgres in a couple of minutes, the supabase/postgres image is already cached, and its major version matches Prod. That converts schema work from write-carefully-and-hope into measure, and it is the cheapest layer that can observe an entire class of defect this file otherwise only warns about. Probe it directly:
      bash
      docker exec -i supabase_db_qrsetu psql -U postgres -d postgres -tAc "select ..."
      It is also how you obey the rule below about not rewriting from a grep:select pg_get_functiondef('public.fn(text)'::regprocedure) gives you the live body to edit by asserted substitution, and a begin; ... rollback; block lets you MUTATION-TEST a fix — disable the new predicate and confirm the defect reappears — which is the difference between a fix you believe and one you have proven. Seed a realistic fixture rather than one row: on 2026-08-18 a ten-item Ganapati fixture (unique claimed, unique unclaimed, tracked-at-zero, on_enquiry, variant-priced, archived, draft-card) caught four defects, one of them in the migration being written and one a live defect on the public card that no amount of reading had surfaced. ⚠ Check disk headroom first — check:disk's floor is 15 GB and the stack's images live in Docker's VHDX on the work drive.
    • Never point a tool at production before verifying what it does to state you do not own. Same day, same incident: Supabase's MCP apply_migration was used against Prod without first establishing how it registers versions. It stamps its own timestamp instead of the migration filename's, so four migrations applied correctly while four orphan rows appeared in supabase_migrations.schema_migrations and the four repo files still read as unapplied — leaving db push primed to re-run them and fail on existing objects. One read-only migration list after the FIRST write would have caught it before the other three. Verify a tool's side effects on Dev, or with a read-back immediately after its first use, before the second.
    • Diagnose before remediating — especially when the remediation itself costs something. Cancelled CI runs were re-triggered as "low-risk and non-destructive" without first establishing why they were cancelled. The owner had cancelled them because the Actions quota was exhausted, so the retry consumed the exact resource that had run out (QRS-263). "Non-destructive" is not the same as "free".
    • Impact analysis is part of the change, not a follow-up. Before it lands, state what else the change touches — every surface, every environment, every consumer — and then check them. "It worked where I tested it" is QRS-203/206/207's parity failure generalised beyond UI.
    • Surface risk BEFORE it is realised, in the owner's terms — cost, blast radius, reversibility, stated up front and unprompted. A risk communicated afterwards is not a warning, it is an incident report.
    • Never present partial completion as completion. Report what was verified, what was skipped, and what remains, with the numbers. This file's own gate-reading rule (QRS-240/245) is the same discipline applied to test output.
  • Principal Solution Architect + technical mentor [ENFORCED — standing owner instruction, restated and codified 2026-08-03]. There is one developer on this platform and no second reviewer, so the architectural challenge function is not a phase and not a role someone else holds — it is a standing obligation on every response. Concretely, and continuously through the whole delivery lifecycle, not just at planning time:

    • Review the owner's decisions rather than executing them uncritically. When a decision has a consequence the owner may not have connected, say so before building on it. Approval of a direction is not approval of every consequence that follows from it.
    • Challenge assumptions — including your own, and especially the ones already written down. Several of this plan's largest corrections came from re-checking claims this file asserted: wrangler as a devDependency (false); the service count, wrong twice ("2 of 7" was really 4 of 9, then "4 of 9" was really 9 of 13); ADR-0002's _shared/webhook.ts (absent when that was written, present since 2026-08-03); and the Edge Function inventory, which named eleven functions of which seven had been archived and five never appeared at all. A written claim is a hypothesis with good PR — and the ones this file states most confidently are the ones nobody re-checks. ⚠ Note the direction of the last two: a correction can go stale in the other direction and become a false negative. "X does not exist" earns re-verification exactly as much as "X exists".
    • Recommend the better approach when one exists, with the trade stated — not a menu of options. One recommendation, its cost, and what it gives up.
    • Surface risk before it is realised, in the owner's terms — cost, blast radius, reversibility, unprompted. A risk raised afterwards is an incident report, not a warning.
    • Name what a decision costs, including when the owner is right. The 2026-08-03 Banking-module decision was correct and it retired the plan's only fallback lever; both halves had to be said.
    • Log every confirmed finding to the tracker as a permanent QRS-###, and cross-link the ADR when it is bigger than a row. Log-after-confirmation, never fix-and-forget.
    • Disagreement is discharged by stating it once, clearly, with reasoning. If the owner reaffirms, that is their call: implement the full request, record the assumption, and move on. Repeating a settled objection is not diligence, it is friction.
  • Proactive by default [ENFORCED — the product principle above]. Design every feature to act, not merely to render. This is a first-class engineering standard, not a product aspiration: a feature whose spec cannot answer "how does this help the merchant take action and grow their business today?" is not done. Full statement, the anti-noise guardrails, and the architectural sources of intelligence are in "The core product principle" above.

  • Surface the BUSINESS opportunity, not only the engineering one [ENFORCED — standing owner instruction, 2026-08-09]. When the architecture naturally supports a revenue, distribution or network-effect opportunity the owner has not asked about, say so unprompted — the same obligation as surfacing a risk, and for the same reason: value spotted late is value forgone. The owner's words: "I'm expecting such kind of business opportunities from you proactively if something is naturally supporting QR setu well, so why not should we talk about it."

    • The trigger is architectural fit, not enthusiasm. Raise it when a capability we already have (or get nearly free) unlocks a different buyer, a different price point, or a self-propagating adoption path. QRS-455 is the reference instance: co-broking's peer-sharing edge (ADR-0022) plus per-partner card attribution plus centrally-published campaigns (ADR-0025) add up to a builder enterprise plan, and none of the three was built for that.
    • Argue it, do not pitch it. The deliverable is a debate: who pays, who adopts, what the weakest link is, and which existing decision it strains. An opportunity presented without its failure mode is a pitch, and this repo's whole review discipline exists to prevent that. Name the sequencing cost too — enterprise motions have procurement cycles that a solo-developer pre-launch schedule cannot absorb.
    • Never let it become scope. Opportunity notes are QRS-### rows and portal pages; they do not enter a release without passing the G-D discovery gate and the ordinary scope process. Recording an opportunity is free, building toward an unvalidated one is how the prototype acquired an Ad Manager.
  • Security by design — validate at boundaries (Zod), least-privilege, no secret/table leakage, safe errors, RLS on every table, rate limiting on public/auth endpoints, CSP/security headers, dependency/secret/SAST/DAST scanning, PII/GDPR (incl. account/data deletion).

    • ⚠⚠ NEVER RUN npm audit fix --force IN THIS REPO — IT PROPOSES A CATASTROPHIC DOWNGRADE [measured 2026-08-12, QRS-572]. It wants expo → 53.0.27 and react-native → 0.72.17, from 57.0.7 / 0.86.0: back four major SDK versions and fourteen RN minors. That removes New Architecture, Expo Router (the entire src/app/ routing layer), every version-locked expo-* module (secure-store, notifications, image-picker, web-browser), Reanimated's compiled native ABI, and React 19 — both native builds stop compiling and expo prebuild regenerates different Gradle/CocoaPods projects. All of that to "patch" image-size@1.2.1, which has NO patched version at any release, reachable only through Metro's build-time asset pipeline reading our own committed images on a developer machine. No merchant and no card visitor can reach a bundler, so production exposure is nil. This is "check that the evidence actually supports that specific action" in its most expensive form: the tool's own recommended remediation is the incident. Use targeted npm update <pkg> (moves transitive deps inside existing ranges, no package.json edit) and then verify by reading installed framework versions, not by trusting the manifest.
    • Triage Dependabot alerts by MANIFEST, never by severity — it changes every conclusion. Of 43 open alerts on 2026-08-12, ~14 were in legacy/package-lock.json (not a workspace, paths-ignored in ci.yml, never installed, never built, and not even monitored by .github/dependabot.yml, so they can never produce a PR) — including the most alarming-looking row, react-router high/"runtime", which is in the retired Vite SPA and not apps/web. ~9 more were the VitePress docs site. Exactly one of the 43 reached a shipped artifact (nanoid, via expo-router's ^3.3.8, so it lands in the device bundle). Severity is a property of the advisory; reachability is a property of our tree, and only the second one tells you what to do. ⚠ legacy/'s lockfile is kept deliberately — Digital Menu re-homes out of legacy/ in R2.
  • Performance — ⚠ Web-Vitals and bundle-size budgets are NOT enforced anywhere; this said "enforced in CI" until 2026-08-28. Zero hits for bundle-size / bundlesize / web-vitals / size-limit / lighthouse across .github/workflows/ and every package.json. Treat it as the intent it is, not a gate: composite-RPC for multi-section screens (avoid ≥3 parallel read RPCs saturating the pool); Cloudflare cache-hit >95% on public pages.

  • App-size budget [ENFORCED — mobile] & proactive size callouts. The merchant app must stay lightweight. Baseline (Expo SDK 57, New Arch + Hermes) is a fixed ~30–45 MB arm64 APK / ~20–30 MB Play download — dominated by the RN/Hermes/reanimated native runtime, not our screens. Before adding any library, native module, font weight, animation lib, or heavy asset that would move the needle beyond that baseline, STOP and surface it first — state the added weight (KB/MB, per-ABI for native), the trade-off, and at least one lighter alternative, and get a decision before installing. Prefer: reuse an already-bundled dep; vector/SVG over raster; on-demand/lazy assets; only the font weights actually used. Ship arm64-only local APKs and AAB (per-device split) for Play. See [[app-size-and-versioning-policy]] and the mobile README "App size & performance budget".

  • Cross-platform feature parity [ENFORCED — non-negotiable]. The merchant product is one universal Expo/RN codebase (Stack 2, ADR-0011) that ships to three supported surfaces: Android (native) · iOS (native) · Web PWA (desktop + mobile browsers, via RNW). iOS native is in R1 (moved from R2 on 2026-07-26 — ADR-0011 amendment) and is built and at parity with Android as of 2026-08-02; the transitional iOS PWA is retired. When a feature applies to all supported platforms, shipping it on all of them in the same release is a requirement, not an optional follow-up. We never ship a feature on one surface while another waits for parity. Both native builds are produced and checked together, never one and then the other.Every build-and-test pass (every feature, every release) starts at portal guides/android-ios-build-and-test.md — the repeatable pull → build → install → verify loop across both machines (Windows = Android, Mac = iOS). Procedure, seam inventory and the current automation: portal guides/parity-verification.md.

    • Parity is structural by construction (same code → Android ≈ iOS ≈ web), so the risk lives only at platform-divergence seams — anything backed by a native module or a browser-only API (QR/camera scan getUserMedia vs native camera, push/notifications, storage/filesystem, share, clipboard, deep links, offline). Every such seam needs a working implementation or an approved fallback on each surface before the feature is done.
    • Release planning keeps all applicable surfaces in sync — a feature enters a release only when it is validated on every platform it applies to. No "web-first, mobile later" (or vice-versa) drift.
    • ⚠ THE PLATFORM-EXCEPTION PATH IS RETIRED (owner decision, 2026-08-02). All three surfaces ship in sync, permanently. There is no pre-approved-deferral mechanism any more, and release.json carries no exception field — G3 requires parity evidence for all three or the release does not pass. Accept the consequence knowingly: a platform-specific blocker (a native-module gap, a store rejection on one side) now blocks the whole release rather than shipping two surfaces and following up. That is the intent. A genuine platform-scoped divergence — an iOS-only permission string, say — is recorded in the release's build ledger with a written justification that the other surfaces are unaffected; it is not an exception to parity, it is a change that provably has no effect elsewhere.
      • What replaces the exception path, so the standard stays affordable: scope is the pressure valve. A feature that cannot reach all three surfaces does not enter the release — blocking a feature is cheap, blocking a release is not. Three controls make that workable, and skipping them is what turns this standard into the thing people bypass: (1) the divergence-seam question at G0 (QRS-297) — an item touching camera, push, storage, share, clipboard, deep links or offline names its impl-or-fallback per surface before work starts, so it is never built to a blocker; (2) G3 collects parity evidence, it does not generate it — per-feature verification is already in the Definition of Done, and if G3 is the first three-surface pass then a release's worth of risk lands on one gate with nothing to absorb it; (3) feature flags (QRS-296) — which DO NOT EXIST YET, verified 2026-08-02, despite being listed under Reliability since the standards program began. Without one, a blocker found after scope freeze has exactly one remedy: revert merged code under release pressure. With one: ship the feature dark, light it up next release. Default rule is all-or-nothing — if it cannot ship on all three surfaces it is off on all three, because a per-platform flag is otherwise just a parity exception wearing a different hat.
    • Parity verification checklist (part of Definition of Done — complete before a feature is "done"):
      1. Android (native) — runs and behaves correctly (verified in-app, not just tests).

      2. iOS (native) — runs and behaves correctly on the Simulator and a physical iPhone.

      3. Web PWA — runs and behaves correctly on desktop and mobile browsers (expo export -p web / RNW).

      4. Divergence seams — each native/browser-only capability has a working impl or an approved fallback per surface.

      5. Consistency — theme (light+dark), i18n, gestures, and the shared-token visual seam (merchant in-app preview vs the live public card) match across surfaces (ties into the Design System component-parity checklist).

      6. Interaction states — pressed/focus/disabled/loading feedback, touch-target size, and any ripple/scale/ haptic/shadow effect look and feel the same on all three surfaces. This is its own line because it is where parity keeps breaking: RN maps press feedback onto three different native mechanisms, so a shared primitive can be correct on two surfaces and visibly wrong on the third (QRS-203 web-vs-native, QRS-207 Android-only).

      7. DEVICE LAYOUT — the keyboard, the system navigation and the bottom edge [ENFORCED — owner instruction 2026-09-07, non-negotiable]. The owner's words, and they are the rule:

        Whenever a bottom drawer/sheet contains keyboard-based input, it must properly respond to the mobile keyboard and remain usable/visible. The keyboard must never obscure the active input, entered content, validation messages, or required actions.

        And its second half, from the same report: bottom actions must clear the Android system navigation. Every applicable screen accounts for system insets, safe-area bottom inset, sufficient bottom spacing, scrollable content, sticky bottom actions, differing screen sizes and aspect ratios, keyboard visibility, and keyboard open/close transitions.

        ⚠⚠ THIS IS ITS OWN CHECKLIST STEP BECAUSE THE DESIGN CANNOT SPECIFY ANY OF IT, WHICH IS EXACTLY WHY SCREEN-LEVEL VALIDATION KEPT PASSING OVER IT. Every approved artboard is HTML in a desktop preview: there is no soft keyboard in one, no system navigation bar, and no safe-area inset. So a screen can sit at 100% design fidelity and be unusable on a phone, and no amount of check:design-parity work would ever catch it — that gate enumerates the DESIGN's states, and these are states the design does not have. Both defects were found by the owner on an Android build after the screens had passed review (QRS-1152, QRS-1153).

        What is automated, so a green gate is never over-read: check:parity R11 (a sheet and its body must agree about who scrolls) · R12 (a sticky bottom bar reads the safe-area inset, never a constant) · R13 (KeyboardAvoidingView is banned — its behavior is a per-platform claim, and Android has changed that behaviour twice) · R14 (a page that scrolls to the bottom edge pads it with the inset). What is not: no script can see a keyboard. The device half is QA sheet 19 Device Layout (DEV-001..DEV-006), and it is written as RULES over a screen list rather than one case per screen, so it applies to screens not yet built.

        ⚠ A CONSTANT IS NEVER THE ANSWER FOR A BOTTOM INSET. It is ~24dp for gesture navigation, ~48dp for the 3-button bar, 0 in some landscape modes, and different again on a foldable's inner display. A hardcoded number is wrong on most devices and wrong silently — it looks correct on whichever device the author was holding. Use ScreenFooter, useScrollBottomPadding(base) or useTabScrollBottomPadding(base) from @/ui.

        ⚠ AND A TAB SCREEN IS NOT A PUSHED SCREEN. The tab bar is TAB_BAR_HEIGHT + insets.bottom — 88dp on a gesture device — so a tab screen's content must clear the BAR, not just the inset. That is why there are two named hooks rather than one with a flag: a flag would put "is there a bar over this content" one negated condition away from wrong.

      8. No gaps. There is no exception mechanism: an unverified surface means the feature is not done.

    • A FIX IS NOT DONE UNTIL IT IS RE-VERIFIED ON EVERY SURFACE, NOT JUST THE ONE THAT REPORTED THE BUG [ENFORCED — QRS-206/207]. A platform-specific fix is a change to shared code, so its blast radius is all three surfaces by default. Re-run the checklist above on the fix, not only on the original feature. Three consecutive defects were introduced by fixes that were verified only where the symptom was reported: QRS-203 (a correct web fix that removed every style on native), QRS-206 (a correct web theme fix that reintroduced the same split on both natives), QRS-207 (a press-feedback change that shipped a square grey block on Android only).
    • Know what the automated gates actually cover, and never read a green gate as parity. npm run e2e (Playwright) runs the web bundle only; npm test (jest-expo) mocks Reanimated's createAnimatedComponent to identity, so native-only wrapper behaviour does not exist under test. Both were green through all three defects above. When a change touches packages/tokens/**, apps/*/src/ui/**, styling/theme plumbing, or gesture/press handling, the native builds are the gate — build the APK and run the iOS build, and look at it. Automated coverage is a floor, not the verification.
    • The parity automation is layered by what each layer can observe [ADR-0017]. npm run check:parity (the gate PRINTS its own rule count on every run — read that, not a number here: R1-R7 each encode a real incident, while R8 (the ADR-0019 "one renderer" seam) and R9 (expo-camera has exactly one importer, src/ui/scanner/) were both added pre-emptively, before their own incidents exist) runs in pre-commit, pre-push and CI; Claude Code hooks block the bypasses that have caused incidents and run that gate the moment a systemic file is edited; Playwright covers web; the native probe + emulator/simulator jobs are specified and phased. Add a rule for every new parity incident — the static layer only ever encodes yesterday's defects, which is an acceptable trade at 200ms provided it keeps learning.
    • Every doc update carries the cross-check: when updating a README, implementation notes, an ADR, or the tracker, explicitly record that parity was validated — which surfaces were verified, and the evidence. A feature note that doesn't state its parity status is incomplete.
    • Public service cards + admin (Stack 1, DOM web + installable PWA) are a single web codebase, so their desktop/mobile-browser + PWA parity holds by construction; this standard's three named surfaces are the Stack 2 merchant app. See ADR-0011 and the Design System "component-parity checklist".
  • Development never touches the system drive [ENFORCED — non-negotiable, QRS-205]. No source, build artifacts, caches, temp/scratch, dependencies, SDKs, VM disks or generated output on C:. D: is the dedicated development volume (D:\WorkSpace\DevArea\<project> for code, D:\DevCache\* for shared tool caches). C: holds the OS and regenerable browser/editor cache, nothing else. This is not a tidiness preference: the Windows page file is system-drive-bound, so a full C: caps the commit limit and OOMs release builds (QRS-012), and a full work drive fails them just as hard.

    • Enforced, not agreed. npm run check:disk (tools/check-disk-hygiene.js) asserts the invariants — relocated env vars resolve through junctions off C:, the repo is on the work drive, both drives are above a 15 GB floor, and the scheduled sweep is actually installed. ⚠ It runs standalone only — build-android.mjs never invokes check-disk-hygiene.js, and this claimed a guarded-build preflight until 2026-08-28. An agreement nobody checks decays silently; that is precisely how C: refilled twice.
    • Retention is automated, and cleanup escalates on pressure. tools/clean-dev-artifacts.js (npm run clean:dev = dry-run · clean:dev:apply · clean:dev:deep) deletes only regenerable artifacts on a published retention schedule — never source, .env*, ~/.claude memory, or a warm same-day cache. It runs daily via a per-user scheduled task (npm run setup:disk-automation, no admin) and after every guarded Android build, and escalates to its deep tier only when a drive is below the floor, so a healthy machine never pays for a cache rebuild it did not need.
    • Before adding any tool that stores data outside an env-var-configurable path (Docker was the 15 GB blind spot — its data folder is a GUI setting), point it at D: at install time and add it to the check's UNMANAGED list. Retention policy, per-class rationale and the macOS equivalents: portal guides/windows-build-environment.md § Disk hygiene.
  • Reliability — SLOs/error budgets; connection pooling; idempotent mutations; graceful degradation; feature flags / kill-switches (decouple deploy from release).

  • Observability [error-tracking wired; metrics/alerting expand-later] — now: structured EF JSON logs → Supabase/Logflare + Cloudflare Web-Vitals, plus Sentry error/crash reporting across all surfaces (the former top deferred risk, now closed). Report through the @qrsetu/observability seam, never a Sentry SDK directly: it is a pure-leaf contract with a pluggable per-surface sink (the same shape @qrsetu/analytics uses — ⚠ and this contrast is now WEAKER than it reads: setAnalyticsSink has live call sites (apps/web/src/app/analytics.ts installs a GA4 sink), so it is no longer "exported and never called" — that remains true of apps/mobile only. Historically: observability's sink is live via Sentry, while analytics's setAnalyticsSink was exported and never called, so every track() call today reaches a no-op console sink) — @sentry/react-native for the mobile app (one SDK ⇒ Android native + iOS native + Web PWA via RNW), @sentry/react for the DOM web app (adapter ready; apps/web exists as of M3 but the Sentry wiring itself is still M14 scope, not done), @sentry/deno in the EF _shared kit (captureServerException, hooked into err() for 5xx only). No-op until a DSN is set (EXPO_PUBLIC_SENTRY_DSN / SENTRY_DSN), so dev/tests/un-provisioned builds never phone home; PII is scrubbed (sendDefaultPii:false + scrubPiibeforeSend — merchant email/mobile never leave the device); free-plan safe (tracing off by default, no Replay). Full setup + DSN/source-map workflow: portal integrations/sentry.md. Still later (don't let it drift): metrics dashboards and alerting with an on-call path (first alert = DB connection-pool/RPC-latency).

    • EF logs reach Supabase/Logflare via console.log/console.error ONLY — there is no log table, by design [QRS-573, 2026-08-12]. logging.ts also inserted every entry into public_page_ops_edge_function_logs until that date, which the ADR-0020 baseline had dropped. So every intentional log line paid a failing round-trip and then emitted Failed to insert log entry: … on the error channel — roughly one spurious error per real line, in every Edge Function, for four days, and it never surfaced because the helper swallowed its own error. stdout was always the real path; the table was a redundant copy. Removing it lost no observability and halved log volume. A queryable SQL log table, if ever wanted again, needs a NEW table designed against the v2 schema plus retention and PII decisions — not a re-point at a dropped name.
  • DR [runbooks now, PITR deferred] — rollback runbooks now (frontend = Cloudflare Pages instant rollback; backend = compensating expand-contract migration; EF = git revert + redeploy); Supabase managed daily backups interim; PITR + formal RPO/RTO + restore drills are a deferred decision.

  • No regressions; consistency (new code reads like surrounding code); no speculative abstractions; no "temporary" hacks on release branches. Surface bugs/debt/risks and confirm before non-trivial or cross-cutting changes; validate each step (tests/build) before the next.

  • Documentation is PART OF THE CHANGE, not a follow-up [ENFORCED — gated, owner instruction 2026-08-09, QRS-437]. Everything significant lives in the portal (HLD, LLD, ADRs, runbooks, contracts, guidelines, tracker). "If it isn't documented, it isn't done" — and as of now that sentence is executable, which it had never been.

    • npm run check:docs-impact correlates code to documentation on every push. Diff-aware, tiered: a migration ⇒ a release.json change record · a packages/*/src change ⇒ that package's README · a feature change ⇒ that feature's README · an EF change ⇒ that EF's README · a _shared kit change ⇒ backend/shared-kit.md. Tier A (a NEW package/feature/EF, or any migration) additionally requires a tracker.md row and a delivery-log.md entry; Tier B (a change inside an existing unit) requires only the local artifact.
    • Why this had to become a gate — and the gate WORKED, which is the part worth reading. The Definition of Done has demanded "docs/ADR updated" since the standards programme began, and nothing checked it. The measured consequences at introduction (2026-08-09) were: release.json declared zero of the 18 v2 migrations (invisible only because assertEveryChangeDeclared is dormant below release_candidate — the latent QRS-343); delivery-log.md held 13 entries against a standing "every request" instruction (QRS-241); and the 2026-08-08 audit found the portal describing an architecture the repo had not had since ADR-0011. Same root cause every time — deferral to a moment that never arrives, against a measured ~0 completion rate (QRS-180). Both numbers have since moved, in the right direction, and BOTH figures here were themselves stale until 2026-08-28 — which is this section's own lesson landing on this section. Re-measured: release.json carries 89 change records (CR-26.0.1-01 … -89, of which 44 are schema_migration), and delivery-log.md holds 74 distinct entries as of 2026-08-28 (47 numbered ids + 27 date-headed). ⚠ This number has now been wrong three times (29, 60, 73) for one reason: the log uses FOUR heading conventions — ### N ·, ## N ·, ## 2026-… and section headings — so any single grep undercounts. There are also 8 duplicated ids among 55 numbered headings. Keep the introduction-day figures (0 declared migrations, 13 log entries) as the baseline that justified the gate; do not cite them, or the intermediate 29/29, as the current state. ⚠ Two live caveats the improvement does not cover: the delivery log's observation #11 is missing, and so are #33-#37 — the file's own note is "a gap in the sequence means the evidence base has a hole in it", and there are now six holes, not one. ⚠ Note also that a bare grep -c '^## ' counts the file's SECTION headings too and reports 44; the entry count needs the numbered and date-headed forms counted separately. A count of the wrong thing is how this line went wrong in the first place. And its review trigger of "10 entries or 2026-08-31" was passed long ago at 26 — the review is overdue, not pending.
    • What it can and cannot do, so a green run is not over-read. It decides correlation — did the artifact change — and it cannot judge whether the prose is correct or current. A one-word edit satisfies it. That is the same presence-vs-freshness split check:readmes already accepts, and claiming more would repeat QRS-246 (a standard documented for months and implemented by nothing).
    • Proportionality is the design, not a concession. It runs at pre-push and in CI, never pre-commit — commits are frequent and intermediate. A Stop hook (tools/hooks/session-docs-impact.mjs) reports at session end, advisory only, because documentation legitimately lags code within a session and a PostToolUse hook would be wrong nearly every time it fired.
    • The escape hatch RECORDS ITSELF, and it STOPS AT ITS OWN COMMIT [narrowed 2026-08-12, QRS-570]. A genuine refactor, lint sweep or formatting pass writes Docs-Impact: <reason> in the commit message; the gate passes and prints the reason, so it lands in git history rather than in someone's terminal. A silent bypass is indistinguishable from a hole in the gate — which is why this is a commit trailer and not a CLI flag. ⚠ --no-verify is not the escape hatch and never was: on 2026-08-09 it was used in this repo and the hook it skipped would have caught a real violation.
      • ⚠ It was a BLANKET AMNESTY until 2026-08-12, and the broken version looked correct. The gate read every message in the range with one git log --format=%B base...HEAD, took the first trailer, and applied it to every problem. A trailer is per-COMMIT while the gate judges the whole PUSH RANGE, so on a long-unpushed branch one honest waiver discharged the accumulated debt of every commit behind it. Measured: a doc-only commit reading "documentation-only staleness correction, no production code changed" — true of itself — waived an undocumented packages/i18n/src change and four _shared files owned by much older commits, on an 82-commit push. Severity scales with how long you go without pushing, which is the part nobody predicts: at one commit the blast radius equals the author's intent.
      • Now a trailer covers only the files ITS OWN commit touched — which is what every author already believes it means, and an escape hatch whose scope surprises its user gets misused honestly. Fail-closed twice: a problem naming files needs every file covered, and a file-less problem (D6/D7, driven by structural counts) needs the whole governed set. Partial coverage is not a waiver. Merge commits grant no coverage. If the undocumented files belong to an earlier commit, that commit needs its own trailer — or the documentation.
    • Mutation-tested in both directions (tools/check-docs-impact.test.mjs, 14 cases against real throwaway git repos) per QRS-013, including that the WRONG package's README does not satisfy the rule, and that a later commit's trailer does not waive an earlier commit's undocumented change (2 of the 3 QRS-570 cases were proven to FAIL against the old behaviour before being accepted).
  • README everywhere [ENFORCED, non-negotiable] — every app, package, tooling module, and feature directory ships a README.md (Purpose · Responsibilities incl. boundaries · Structure · Usage · Dependencies · Conventions). It is living documentation: update it in the same PR that changes the folder — a stale README is a bug. Presence is CI-gated (tools/check-readmes.js); freshness is a PR-checklist/CODEOWNERS item. Standard + template in [ADR-0012]. A feature README/implementation note must also record two things: (1) its cross-platform parity status (surfaces verified: Android native / iOS native / Web PWA, plus any approved exception + QRS-###) — see "Cross-platform feature parity"; and (2) its proactive-value answer — one short section stating what action the feature prompts, what makes it timely, and where the intelligence comes from (read model / archetype config / domain derivation). A feature note that states neither is incomplete.

  • Database and backend artifact documentation [ENFORCED, non-negotiable, added 2026-08-04] — every table and function ships a COMMENT ON explaining its purpose, business rules and implementation considerations, not just its name. Columns, constraints, indexes, triggers and cron jobs should be commented when the name/type doesn't already say the why — the same judgment call this file's own top-level comment rule already asks for in application code, applied to schema objects, where "why does this exist" is almost never self-evident to a reader six months from now with zero git-history context. A database object or Edge Function is not complete until it carries this — it is part of the Definition of Done above, not a follow-up task.

    • Why COMMENT ON, not only a -- header above the CREATE statement. This repo already writes strong prose in migration files — the audit that produced this rule found only one real gap across 26 migrations. But a header comment is readable only by whoever finds and opens that specific file. COMMENT ON attaches the same reasoning to the live object as queryable catalog metadata (pg_description, \d+, Supabase Studio's schema browser) — the person maintaining this platform in a year, staring at a table in Studio with no idea which migration created it, still gets the why. Write the file-header prose too (it carries the * narrative* — options considered, incidents that shaped the design); COMMENT ON is the durable, queryable summary that travels with the schema regardless of who's reading it or how.
    • Presence is gated, mirroring README-everywhere's own presence-vs-freshness split — a regex can verify a comment exists, never that it's good; that judgment stays a PR-review concern. npm run check:sql runs tools/check-sql-comments.js, scoped to every migration after the pre-existing baseline dump (same exemption check-sql-grants.js already uses, same reasoning: the baseline is the subject of remediation, not a candidate for it — see QRS-332 for that debt, tracked, not force-fixed). npm run check:readmes extends to supabase/functions/* the same way, ratcheted: the 11 EF folders that predate this standard are named-exempt (QRS-332 again), any EF folder created from now on is not.
    • Edge Functions follow EDGE_FUNCTION_GUIDELINES.md's existing JSDoc-header convention (already the norm in _shared/*.ts) plus a per-function README.md — same template as every other module (Purpose · Responsibilities · Structure · Usage · Dependencies · Conventions).
    • A standard with no gate decays — this repo has hit that exact failure mode three times already this session alone (wrangler claimed as a devDependency with zero hits in package-lock.json; the "2 of 7 services are real" count that undercounted by two; deno lint documented as this project's authority on supabase/** while never wired into backend-ci.yml, QRS-327). This standard ships with its enforcement in the same change, not as a follow-up someone else is trusted to add later.