Skip to content

WhatsApp OTP — implementation trail ​

The live record of executing the approved plan. Status is updated as work lands, and every finding is logged here and in the tracker rather than decided silently. Architecture and reasoning: the approved plan. Meta-side setup: the playbook.

THE EXECUTION PRINCIPLE

Architect and build the backend foundation correctly → validate it thoroughly → build the client experience on top of it → perform complete end-to-end validation → close all findings.

Not complete when the happy path works. Complete when backend, client, docs, tests and end-to-end validation are reviewed and every finding is resolved or consciously deferred with a reason.

Status ​

PhaseObjectiveState
Gate 0Answer the factual unknowns before committing🟢 COMPLETE — Route A confirmed by measurement
Phase 1Backend + core architecture, production-ready🟢 all 9 steps addressed — one deferral recorded (F20)
Phase 2Client experience on the finalised backend⚪ not started
Phase 3End-to-end revalidation and closure⚪ not started

Gate 0 · Answer before committing ​

Objective. Establish, by measurement, who owns the OTP lifecycle, and clear the ground.

#TaskDeliverableAcceptanceState
0aRoute probe: external.phone on, auth.hook.send_sms at a stub, no providerA recorded answerRoute A or Route B chosen on evidence🟢 DONE — ROUTE A
0bDoes signInWithOtp({ phone, options: { data } }) write raw_user_meta_data?A recorded answerIf no, the plan's step 9 changes🟢 DONE — YES
0cClear the ground so nothing legacy masks a defectA measured deletionTesting cannot be confounded by old rows🟢 DONE, NARROWED

0a · Route A is confirmed — QRS-934 ​

POST /auth/v1/otp returned 200 in 2382 ms with no provider configured, and the function logs carry POST | 200 | .../functions/v1/send-auth-otp. GoTrue never consults sms_provider when a Send SMS Hook is enabled — the "twilio" value on this project is GoTrue's default enum value, not a configuration, and every credential field under it is null.

So Supabase owns the code and we own delivery only. That is the preferable division: GoTrue already solves expiry, replay and attempt limits correctly and sets phone_confirmed_at. ⚠ We must not add our own attempt counter — that would be a second source of truth for "is this code still valid", which is the QRS-249 class. Our limits are about cost and portfolio capacity and sit before Supabase, never beside its verification.

The payload is now measured rather than guessed, which is the probe's second dividend:

headerswebhook-id · webhook-signature · webhook-timestamp
schemeStandard Webhooks — base64 over {id}.{timestamp}.{body}
bodymetadata{ip_address,name,time,uuid} · user{…,phone,user_metadata,app_metadata} · sms{otp,phone}
otp6 characters

⚠⚠ That is NOT Meta's scheme. X-Hub-Signature-256 is hex over the raw body; _shared/webhook.ts implements that one and cannot verify this one. Hence _shared/standardWebhooks.ts as a separate module, with tests asserting each scheme against the other's fixtures. C2's idempotency key auth-otp:<webhook-id> is directly supported.

0b · The metadata channel works, and the key is primary_context ​

primary_context: 'individual' → provisioned individual, display_name set, 0 workspace memberships. Exactly what the three-category rule requires, so no SQL change is needed and step 9 is unaffected.

⚠ Found on the way — QRS-935. Sending the wrong key (account_type) provisioned the user as business with no error anywhere, because handle_new_user falls through to else 'business'. The client is correct today (account_type has zero occurrences in apps/**), so this is a latent trap rather than a live defect — but the next signup path added has a one-word way to provision every consumer as a merchant, silently.

⚠ Also found — QRS-936. GoTrue stores the phone without the leading + (919999999998). Our ledger specifies E.164 with it. Two formats, one number, and the mismatch is on the table that decides identity. A decision is owed before the send path is written, because retro-normalising a high-volume ledger is a migration on a table that will not be small.

0c · Narrowed deliberately, and the reasoning is the record ​

The plan said wipe the Dev auth users; the owner authorised it as "if required". Measured first: a full wipe destroys 19 orders · 23 order_items · 18 payments · 5 workspaces · 5 setu_cards · 7 catalog_items · 4 reminders — the money-path and Ganapati fixtures.

It is not required. C4's invariant binds new accounts regardless, and "nothing legacy masks a defect" is bought for free by testing on fresh numbers. Only the three probe users created by this session were deleted (verified after: 7 auth users, 0 phone users, 0 orphans, workspaces/orders/payments unchanged at 5/19/18). ⚠ C4 is therefore enforced going forward, not retroactively — a recorded choice, not an oversight.

✅ Gate 0 no longer blocks anything. Steps 6, 7 and 8 are unblocked and their shape is settled by Route A.


Phase 1 · Backend and core architecture ​

Objective. A production-ready WhatsApp send-and-receive foundation with OTP as its first consumer, wired end to end, before any client work begins.

#TaskDepends onDeliverableAcceptanceState
1Identity primitive — userId as the signed-in predicate; email/phone nullable—AuthSession widened; identityOf + auth/phoneA phone-only session survives a cold start🟢 DONE — 938 tests green
2Migration: five tables + RLS + grants + indexes—One migration, COMMENT ON throughoutcheck:sql green; every table service-role only🟢 DONE, applied + verified on Dev
3pgTAP2communications_whatsapp_test.sqlEvery constraint proven to reject; rank monotonicity; idempotent replay🟡 written + 26/26 verified on Dev; harness unrun — F7
4_shared/whatsapp.ts · _shared/communication.ts · _shared/standardWebhooks.ts2Three modules + Deno teststest:ef zero failures🟢 DONE — 336 passed, 0 failed (62 new)
5whatsapp-webhook EF + registry seed4The EF, config.toml entry, seed migrationA signed event lands; a replay changes nothing🟢 DONE — proven live, F9 closed
6The send EF: three limits, idempotency, in-request retry4, 0aThe EFA replayed hook sends exactly one message🟢 DONE — 200 in 3517ms, replay sent nothing
7Wire the hook, deploy, align OTP expiry to the template's 10 minutes5, 6Live configA real OTP arrives through the hook⚪
8AuthService.sendPhoneOtp / verifyPhoneOtp + stub + E.164 primitive1, 7, 0bSeam + testsProvider-agnostic; the stub can express a consumer🟢 DONE — 950 tests green
9Identity invariant — no account without a verified phone8Linking flowBoth routes land in one account⚪

Phase 2 · Client ​

Objective. Build the approved Claude Design screens on the finalised backend. The client consumes; it does not drive architecture. Planned in detail once Phase 1 is validated.

Screens are designed and validated through design round 14 — journey validation.


Phase 3 · End-to-end revalidation and closure ​

Objective. Recheck the whole chain, then close every finding.

WhatsApp → OTP delivery → verification → Supabase identity → account creation/lookup → session → consumer onboarding → client experience → error and edge cases → production readiness.

Closure requires every row in the findings log below to be resolved or consciously deferred with a written reason.


Findings log ​

Everything discovered during implementation. Nothing here is decided silently.

F1 · 🟢 RESOLVED 2026-09-01 · The environment token and the stored CLI login pointed at a THIRD Supabase account ​

Measured 2026-09-01, before any write. SUPABASE_ACCESS_TOKEN in the environment and the stored npx supabase login both resolve to an account holding digious_site_prod and digious_site_dev — the Digious website projects. Neither can see qr-setu-dev (dyhjofjjuazhyqcvlrkx).

⚠ CLAUDE.md documents a TWO-account trap (dev vs prod). This is a third, and it is the one that makes npm run sb -- dev and every npx supabase command target the wrong product entirely. The preflight in tools/supabase-as.mjs exists for exactly this and would have caught it; the value of running it first is that nothing was written before the mismatch surfaced.

✅ Not a hard block: the Supabase MCP connector authenticates separately and CAN see qr-setu-dev (and qr-setu-legacy-bkp, but not qr-setu-prod — which is the safe posture, since it makes production unreachable from this session by construction).

What the MCP can do: apply migrations, execute SQL, deploy Edge Functions, read advisors. What it cannot do: change GoTrue auth configuration, or set Edge Function secrets.

So the blocked work is precisely: gate 0a (enabling external.phone and the send-SMS hook), therefore 0b, and later the WHATSAPP_* secrets. Needs one of: the owner performing those in the dashboard, or a SUPABASE_ACCESS_TOKEN scoped to the QR Setu account.

✅ Resolved. A QR Setu-scoped token was granted and placed in the Windows User environment. ⚠ The distinction that mattered: this process inherited SUPABASE_ACCESS_TOKEN at launch, so process.env still held the old value and reported the wrong account long after the new one was set. The live value lives in the registry (HKCU), which is exactly why tools/supabase-as.mjs reads it from there — "so an already-open shell does not report it missing". Every script in this work reads the registry, never process.env.

Verified: the granted token sees qr-setu-dev (dyhjofjjuazhyqcvlrkx) and cannot see qr-setu-prod — the safe posture, since it makes production unreachable from this session by construction rather than by discipline.

Consequence, and it improved the work: with a working CLI the migration went through supabase db push, which registers the version from the filename. The MCP connector stamps its own timestamp and would have produced a QRS-267 orphan row. Reaching for the blocked-path workaround would have been worse than the path itself.

Tracker: QRS-933.

F2 · 🟠 A wrong metadata key silently provisions a consumer as a merchant ​

handle_new_user reads raw_user_meta_data ->> 'primary_context' and falls through to else 'business' for anything else. Sending account_type instead produced a business account with no error, no log line and no failing test. The client is correct today, so this is a latent trap, not a live defect — which is exactly why it is written down before the next signup path (this one) is added. Tracker: QRS-935.

F3 · 🟠 auth.users.phone has no +; our ledger does ​

GoTrue stores 919999999998 for +919999999998. communication_messages.recipient_phone is specified as E.164, which carries the +. Two formats for one number, across the table that decides identity and the table that decides who we messaged. Decision owed before the send path is written. Tracker: QRS-936.

F4 · 🟠 CLAUDE.md overstates the anonymous Edge Function surface ​

It claims config.toml carries seven verify_jwt = false entries; there is one. Corrected in config.toml itself in this change by replacing the hand-maintained count with a per-function list of which and why plus the re-measure command — a count kept by hand beside the thing it counts has now drifted twice. Tracker: QRS-937.

F5 · 🟢 Four updated_at columns had no touch trigger — fixed before applying ​

Caught by checking the repo convention rather than by a gate: 17 tables attach public.touch_updated_at() and the four new mutable tables did not. An updated_at populated at insert and then frozen is worse than no column — it reads as maintained. Every registry here is replace-on-sync, so "when did this last change" is precisely the question being asked of them. Fixed in the same migration; check:sql went from 32 to 36 commented triggers.

F6 · 🟢 The identity predicate was wrong in FIVE places, not one ​

The plan named resolveEntryRoute and getCurrentAuthContext. The compiler, once AuthSession was widened, enumerated three more: app/index.tsx, OnboardingSetup/index.tsx, and ReminderAlertsProvider — where Boolean(email) meant a phone-only merchant would receive no reminder alerts at all, with the provider sitting in its idle branch and nothing indicating a whole proactive surface had switched itself off.

🔎 Widening the type was the enumeration. Rather than grepping for email and judging each hit, making the field nullable turned every site that misused it into a compile error. That is the cheapest layer that could observe this defect class, and it found more than the plan's own list.

⚠ Knowingly left email-keyed: markOnboardingCompleted(sessionEmail) skips for a phone-only account. Harmless today (the real impl is a documented no-op, and v2 derives completion from workspace membership) but recorded rather than swept in, because widening it is a separate decision about whether that method should still exist.

F7 · 🟠 test:db cannot run locally — both drives are under the floor with nothing reclaimable ​

check:disk exits 1: C: 10.4 GB, D: 4.3 GB against a 15 GB floor. clean:dev then finds 0 MB reclaimable and says so plainly: "the remaining space is in things this tool must not delete."

⚠ The automation worked exactly as designed and still could not help. Retention sweeps regenerable artifacts; the drives are full of things that are not. So supabase db start is not attemptable — its images live in Docker's VHDX on the work drive, and a full work drive fails builds as hard as a full system drive.

Worked around for this change, not solved. The 26 schema guarantees were verified against live Dev inside a do $$ … $$ block that ends by raising, so the rollback is structural rather than remembered. All 26 passed; a follow-up count confirmed all five tables back at zero rows. pgtap is available but not installed on Dev, and installing an extension on a live project purely to run a test was declined as a change nobody asked for.

⚠ Read the state honestly: the schema's guarantees are verified; the pgTAP HARNESS is not exercised. A defect in the test file itself would not yet be visible. Those are different claims, and conflating them is exactly the shape of QRS-246.

Tracker: QRS-938.

F8 · 🟢 Three layers, and the boundary is what makes the foundation extensible ​

META           _shared/whatsapp.ts        knows Meta · knows NOTHING about QRSETU domains
MESSAGING      _shared/communication.ts   knows messages · knows NOTHING about channels
ADAPTERS       send-auth-otp              ← today.  communication-dispatch ← later

Authentication depends on the messaging layer and nothing in the messaging layer knows what an OTP is. That is the property that makes every future sender another thin adapter rather than a second send path, and it is asserted by tests rather than asserted in prose.

Three decisions inside it that are easy to get backwards:

  1. The sender is an ARGUMENT, never read from env. Configuring it in env would cap the platform at one number forever and make whatsapp_phone_numbers decorative — the very table the tenant programme depends on. Env holds the credential; the registry holds the routing. A test asserts two senders reach two endpoints, so this cannot regress silently.
  2. Failure classification is by HTTP status first, with a numeric-code override. Status-based classification cannot rot; Meta's codes drift. ⚠ An unrecognised 4xx defaults to permanent, and that asymmetry is deliberate: retrying a certain failure burns the user's wait and can bill per attempt, while the inverse costs one avoidable failure that is visible in the ledger.
  3. The core holds no budgets. countRecent is generic and the adapter chooses the numbers. An OTP rule in the core would leak authentication into the messaging layer — an order confirmation legitimately follows an OTP to the same number within seconds.

⚠ retry.ts is deliberately unused here. Its 1s + 2s waits alone eat a third of the Send SMS Hook's ~10-second wall clock before the third attempt starts, and it retries every error — including the permanent ones, where a second attempt is a second charge for a certain rejection.

⚠ The limit we CANNOT have, named rather than implied: there is no per-IP or per-source limit, because Supabase calls the hook, not the client. Implying one would be worse than having none.

F9 · 🔴 WHATSAPP_APP_SECRET is not set, and this session cannot set it ​

It is a Meta credential, from the app's Settings → Basic page. Three of the four secrets are set (WHATSAPP_ACCESS_TOKEN, SEND_SMS_HOOK_SECRET, WHATSAPP_WEBHOOK_VERIFY_TOKEN); this one needs the owner.

Consequence, stated rather than glossed: whatsapp-webhook fails closed with 401 on every POST. That is the correct posture, and it is indistinguishable from a bad signature — so a missing value looks like an attack in the logs rather than a configuration gap.

⚠ This is why the live probe's POST results must not be over-read. Unsigned and garbage-signed requests both returned 401, but the handler returns 401 before verifying when the secret is absent. The probe proves fail-closed; it does not prove signature verification. The logic is the same verifyHmacSignature covered by 13 tests, so the risk is low — but "low risk" and "verified" are different claims and only the first is true today.

Owed, and cheap: set the secret, then re-run the probe with a correctly signed body. Meta's webhook subscription also needs pointing at https://dyhjofjjuazhyqcvlrkx.supabase.co/functions/v1/whatsapp-webhook with the verify token.

F10 · 🟢 The webhook is live and every branch was probed ​

proberesult
GET correct verify token200, CHAL123 echoed as plain text
GET wrong token403
POST no signature401
POST garbage signature401
PUT405

🔎 The responses carry OUR strings, not the gateway's, which is what confirms verify_jwt = false actually took effect — a claim that would otherwise rest on the config file saying so. check:fn-config cannot verify this (QRS-643), so the probe is the only evidence available.

⚠ The registry seed is what makes X1 real rather than theoretical. resolveSender() now returns the live default sender from whatsapp_phone_numbers; before the seed it threw. Verified by reading the rows back: 1 WABA active, 1 connected default sender, 3 templates APPROVED in en/mr/hi.

F9 · 🟢 CLOSED — WHATSAPP_APP_SECRET was supplied and the signature path is now proven ​

The owner supplied it; it was pushed from the gitignored .env.whatsapp without being printed.

What changed is the CLASS of evidence, not just the state. Before, the webhook returned 401 on every POST — correct fail-closed behaviour, and indistinguishable from a bad signature, so the probe could only demonstrate fail-closed. Now:

proberesult
correctly signed Meta delivery200, recorded
redelivery of the identical body200, no second row — the constraint absorbed it
body tampered by one word, original signature401

⚠ Still owed and outside this session: Meta's webhook subscription must be pointed at https://dyhjofjjuazhyqcvlrkx.supabase.co/functions/v1/whatsapp-webhook with the verify token. The function is ready; nothing is sending to it yet.

F11 · 🟢 The send path, proven against the live database rather than a response code ​

A correctly-signed hook request returned 200 in 3517 ms and Meta returned a wamid. The ledger row proves five separate invariants at once:

columnvalueproves
statussentthe send completed and was recorded
recipient_phone+919999999901E.164 with the plus — QRS-936 normalisation, live
phone_number_id1203827669491594resolved from the registry, not env — X1, live
variables{}no OTP persisted — the security invariant, verified in the database
attempt_count / next_attempt_at0 / nullthe OTP path never writes the retry columns — C1

A replay of the same webhook-id returned 200 and sent nothing; exactly one row survives.

⚠ The probe used a number that is not on WhatsApp, deliberately, so the whole chain ran without messaging a real person. ⚠ And note Meta accepts then reports delivery failure later by webhook — "accepted is not delivered", so a 200 from Meta is not proof of delivery.

F12 · 🔴→🟢 Three defects, each found because a probe asked a different question ​

None was on the plan's list, and each was invisible to the check before it.

QRS-939 — String(e) renders a PostgrestError as [object Object]. The first failing probe logged a phase and nothing else. A PostgrestError is a plain object with no toString, so the idiomatic catch-block stringify discards code, details and hint — exactly the fields that identify the fault. 🔎 This is auth/context.ts's own recorded lesson, repeated by me in the same session I read it: an error saying only that a read "failed twice" is why a missing function looked like a network blip for five days. Fixed with describeError.

QRS-940 — a ledger FK violation could block AUTHENTICATION. The moment the error became readable it said 23503 · recipient_user_id … is not present in table "users". The probe's synthetic id caused it, but the consequence is real: if a public.users row is ever missing, the insert throws, the adapter returns 503, and nobody with that account can sign in. The ledger is observability and must never gate authentication. Now degrades to a null identity, loudly logged.

QRS-941 — message_id was never populated, even on success. Every earlier check asked "did the request succeed" and every answer was yes. This one asked "can I reconstruct what happened to this message" — and a join returned zero rows, because applyStatusEvent returned only an outcome string. An audit trail that could not be joined to what it audits, and a partial index that could never be used.

🔎 The pattern across all three: they were found by verifying the STATE, not the RESPONSE. Every one of them sat behind a 200 or a plausible-looking failure.

F13 · 🟢 The complete loop, end to end ​

send → wamid → webhook → correlated event → status, driven with the wamid the send actually returned and delivered deliberately out of order:

  • read (ts 400) arrives first → applied
  • delivered (ts 300) arrives late → ignored, message stays read
  • both events linked to the message, both stamped status_update:out_of_order

⚠ The out-of-order event is stamped processed, not dead-lettered — out-of-order arrival is normal and the ladder handled it correctly. Dead-lettering it would fill the exception queue with the system working as designed.

F14 · 🔴 The hook secret CANNOT be read back — and a green probe hid it ​

sms_otp_exp is now 600 (was 60), matching the template's declared 10 minutes. A 60-second server-side expiry meant the code died while the user was still reading the message telling them it lasted ten. sms_otp_length is 6, matching the template.

Then the real path failed, and the failure is the most instructive thing in this phase.

POST /auth/v1/otp returned 500 · "Invalid payload sent to hook", and our log said no_matching_signature — while a probe that signed its own requests had been passing all along.

⚠⚠ The cause: the Management API returns a MASKED SHA-256 DIGEST of hook_send_sms_secrets. PATCH it with v1,whsec_<44 base64 chars>, GET it back, and you receive 64 characters matching ^[0-9a-f]{64}$. The documented workflow is to put that secret in the function's environment, and the obvious way to obtain it is to read the config — so every key was being derived from a hash.

🔎 Why the test suite could not catch it. The probe signed with the same read-back value the verifier used, so the two agreed with each other while GoTrue disagreed with both. That is verbatim the trap standardWebhooks.test.ts warns about in its own header — "a fixture copied from the implementation can only prove the implementation matches itself" — written by the same hand that then fell into it. For a signature scheme, the only meaningful test is one where the counterparty produced the signature.

⚠ And the diagnostic was fooled too. An inspection reported "base64-ish: true, decodes to 48 bytes" — because hex is a subset of the base64 alphabet, so a character-class check matches a digest happily and it "decodes" to plausible bytes. A shape test cannot tell an encoding from a hash.

⚠ The lesson that generalises past this bug: when a signature will not verify, suspect the KEY MATERIAL before the ALGORITHM. Three encodings were tried against a hash of the secret, which is the most thorough possible way to be wrong. Once the real value was set on both sides in one operation, the ORIGINAL derivation was confirmed correct: key_form: "base64-decoded".

Tracker: QRS-942. The correct provisioning sequence is now a danger box in the playbook.

F15 · 🔴→🟢 The QRS-940 guard turned out to be load-bearing on the PRIMARY flow ​

The real path logged Key (recipient_user_id)=(78319ff0-…) is not present in table "users" — for a user present in both auth.users and public.users moments later.

Supabase documents that auth hooks "are run in a transaction" — GoTrue's own, uncommitted — while the Edge Function connects on a separate session and cannot see the row. ⚠⚠ So this is permanent and universal: it happens on EVERY new user's FIRST OTP, the single most common path in the product.

🔎 QRS-940 was written as a hypothetical guard — "if the row is ever absent…" — against a defect a probe surfaced with a synthetic id. It turns out the row is always absent on first sign-in. Without that guard, every new consumer's first OTP would have returned 503, and nothing short of running the real thing would have revealed it.

⚠ Both the migration comment and the module comment asserted the opposite. Plausible, confidently written, wrong — corrected in place. The log also dropped from warn to info with the cause named, because a warning on every signup trains people to ignore the channel.

Tracker: QRS-943.

F16 · 🟢 The hook used 2822ms of a hard 5000ms ceiling ​

Client-observed totals reached 5094 ms. Instrumenting showed ~1.2 s was Meta and the rest was sequential Postgres round trips from the edge — template lookup, sender lookup and two limit counts, each waiting on the last despite being mutually independent.

Collapsed into one Promise.all: 2822 ms → 2215 ms, warm client total 1846 ms.

⚠ The checks were deliberately not parallelised. Parallelising the reads does not license parallelising the decisions: the category assertion and the limit comparison still run, in order, before anything is enqueued or sent.

🔎 General point about this platform: an Edge Function is REMOTE from its database, so a round trip costs 100-300 ms and four of them is most of a second. Against a hard external deadline that is not a micro-optimisation — it is the difference between fitting and not.

Tracker: QRS-944.

F17 · 🟢 SIX REAL OTPs DELIVERED TO TWO REAL HANDSETS — owner-confirmed ​

The first validation by a human rather than by a query. Three sends to each of two numbers, at 60-75 second intervals, all through the real signInWithOtp path:

batchsendsclient latencyowner confirmed
number 134201 / 2057 / 2081 ms✅ "I got three otps"
number 233708 / 2439 / 2858 ms✅ "got 3 otp's on another number as well"

All six ledger rows: status=sent, wamid present, variables={} (no OTP persisted), template_language=en, no error_code.

🔎 AND THE LEDGER CONFIRMED QRS-943 WITHOUT BEING ASKED TO. The recipient_user_id column across the first batch:

sendtimeidentity linked
112:58:03false ← the account was being created; the row was uncommitted
212:59:21true
313:00:39true

Exactly the predicted pattern, in production data. Without the QRS-940 guard, send 1 would have been a 503 and the owner would never have received a first code. A guard written as a hypothetical, validated as load-bearing, by an experiment run for a different reason.

⚠ A digit was worth checking first. The requested number differed from the production SENDER by one digit (9273373367 vs 9270373367), and the playbook carries an explicit warning that three numbers in this Meta account differ by a digit or two. Flagged before sending rather than after.

F18 · 🔴→🟢 Every hand-listed authService mock had drifted ​

Three screen tests mocked the service with an object literal, and one named signInWithGoogle — a method that left the interface in P3, and whose absence service.ts documents in a fourteen-line comment. So the fixture listed a method that does not exist while omitting two that do.

⚠⚠ jest.mock's factory returns any, so nothing in the repo could see it. Not the type-checker, not the tests, not lint. A screen calling an omitted method gets undefined and fails at runtime with authService.sendPhoneOtp is not a function — a message that reads like a broken import rather than an incomplete fixture.

🔎 It is CR-26.0.1-88's lesson in a new place — "a test that enumerates what it tests cannot see anything added after it was written" — applied to a mock instead of an assertion. Same fix both times: derive the list from the thing itself.

⚠ Latent, not live: no current screen calls the missing methods, so nothing was broken today. It would have broken the moment a Phase 2 phone screen was wired, and it would have looked like a bad import. Fixed with a two-sided compile-time guard, so adding an interface method now fails to compile until it is listed.

Tracker: QRS-945.

F19 · 🟢 The linking mechanism, measured before it was designed ​

The plan named updateUser({ phone }) and explicitly marked it unverified (§1.1a): "Confirm the flow before designing the screen state, or the screen will be drawn for the wrong number of steps." Probed on the live path with an email+password session — the same shape a Google sign-in produces:

observationconsequence
same user id returnedit LINKS, it does not fork ✅ C4's core claim
phone empty, new_phone setthe number is pending, not verified
an OTP went through our WhatsApp hook — real ledger row, status=sent, wamidlinking carries a FULL verification
verification needs type: 'phone_change'not type: 'sms'. GoTrue keys the code by type, so the wrong one reports an invalid code for a valid one — a failure that looks like the user mistyping and is entirely ours

🔎 The warning was right, and checking cost one probe. Designing the screen first would have produced a one-step affordance for a two-step flow.

F20 · 🟠 The invariant is EXPRESSIBLE and TESTED but NOT ENFORCED — deliberately ​

⚠ Measured: all 7 Google accounts have no phone. Every account created through the currently-working merchant sign-in path violates the invariant.

Built: needsPhoneVerification (16 tests) and the two-step linkPhone / verifyPhoneLink seam, with 8 tests pinning that the account id and email survive.

⚠⚠ Deliberately NOT built: the resolveEntryRoute branch. It could return /onboarding/setup — that route exists — but the wizard has no phone step, so a set-up merchant signing in with Google would re-enter a wizard they cannot complete and be sent back on every launch. That is QRS-730 exactly, recreated by the fix for a different defect.

Routing to a screen that cannot discharge the condition is a worse defect than the one it addresses. So the enforcement point is the screen, and the gate lands with it in one change.

🔎 Recorded as a tracker row rather than a comment because "expressible" and "enforced" are different claims, and this repo has a measured history of conflating them: SonarQube documented for months and implemented by nothing (QRS-246), a lint gate passing as a green no-op (QRS-013), deno lint documented as authoritative and never wired (QRS-327).

Tracker: QRS-946.

F21 · 🔴→🟢 The stub nearly proved the OPPOSITE of what it claimed ​

Its first version returned userId: stub-user-<phone> from verifyPhoneLink — an id derived from the number rather than the account being linked to. That is a fork, the exact defect C4 forbids, and every test written against it would have passed while demonstrating the reverse.

Caught while writing the tests, not by a gate. 🔎 A stub that answers plausibly is more dangerous than one that throws, because the plausible answer is the one nobody re-checks — and this one would have been the foundation for the screens.

It now tracks the signed-in account and preserves both id and email; signOut clears it so a stale link cannot attach to nobody.

F22 · 🔴→🟢 GET /secrets masks EVERY value, not just the hook secret ​

Measured across all 21 secrets on Dev: every one returns 64 characters matching ^[0-9a-f]{64}$, whatever its real length or format. This generalises QRS-942 from one field to the whole endpoint — a much larger trap, because any script reading a secret back to configure something else configures it with a hash.

⚠ The earlier read-back was already evidence, misread. Setting the four WhatsApp secrets printed "value returned 64 chars" and that was recorded as confirmation the secret was stored. It was confirmation a digest was returned. A read-back that cannot distinguish the value from a hash of the value is not a read-back.

The tell is the length: a service-role JWT is 219 characters, and 64 is not a truncation of it. ✅ API keys have a real read path (/api-keys); every other secret has none.

Tracker: QRS-947.