Appearance
26.0.1 — Production change log
Change records
| CR | Class | Target | Env | Reversible | Depends on | Blast radius | Evidence |
|---|---|---|---|---|---|---|---|
| CR-26.0.1-01 | app_build | branch reconciliation | prod | no | — | repository only | #cr-01 |
| CR-26.0.1-02 | data_migration | 20260803094212_capture_prod_reserved_slugs.sql | dev, prod | yes | — | reference data only; no-op on Prod | #cr-02 |
| CR-26.0.1-03 | schema_migration | the 5 migrations Prod is behind | prod | yes | CR-01 | +4 tables, +5 functions, +1 trigger, +4 policies on Prod | #cr-03 |
| CR-26.0.1-04 | ef_config | verify_jwt=true on 5 Type A functions | dev, prod | yes | — | restores the platform JWT gate | #cr-04 |
| CR-26.0.1-05 | schema_migration | money data model | dev, prod | yes | — | new tables only | #cr-05 |
| CR-26.0.1-06 | schema_migration | profile_items extension + item_media | dev, prod | yes | — | additive columns on a live table | #cr-06 |
| CR-26.0.1-07 | rpc_function | get_public_profile_by_slug | dev, prod | yes | — | changes what the public route resolves | #cr-07 |
| CR-26.0.1-08 | ef_code | 6 new Edge Functions + _shared/webhook.ts | dev, prod | yes | CR-05, CR-06 | new functions only | #cr-08 |
| CR-26.0.1-09 | cron_job | reconciler + money_health + project-URL fix | dev, prod | yes | CR-08 | scheduled jobs; pg_cron absent on Dev | #cr-09 |
| CR-26.0.1-10 | cloudflare | Pages project + DNS + SSL + Turnstile | prod | yes | — | public DNS for qrsetu.com | #cr-10 |
| CR-26.0.1-11 | third_party | Razorpay Route live payments | prod | no | CR-08 | real money to real bank accounts | #cr-11 |
| CR-26.0.1-12 | schema_migration | 20260801143000_app_release_policy_and_kill_switch.sql | dev, prod | yes | — | min-supported-version kill switch; seeded permissive | #cr-12 |
| CR-26.0.1-13 | schema_migration | 20260803111957_payment_kill_switch.sql | dev, prod | yes | — | payments kill switch; seeded closed | #cr-13 |
| CR-26.0.1-14 | schema_migration | 20260803124128_backfill_reminders_model_comments.sql | dev, prod | yes | — | COMMENT ON only; no structural or data change | #cr-14 |
| CR-26.0.1-15 | ef_code | onboarding-persistence + sign-out fix (e88fa35) | dev, prod | yes | — | the two functions on the signup critical path | #cr-15 |
| CR-26.0.1-16 | schema_migration | catalog_items.pricing_mode replaces a magic price_minor = 0 | dev, prod | yes | — | the anon-callable public catalogue projection, widened not narrowed | #cr-16 |
| CR-26.0.1-17 | schema_migration | orders + order_items, and the orders feature on ledger | dev, prod | yes | — | first consumer-readable merchant row in the schema | #cr-17 |
| CR-26.0.1-18 | schema_migration | payments + append-only payment_events, payments as universal | dev, prod | yes | — | money invariants enforced in the database; inert until payout onboarding exists | #cr-18 |
| CR-26.0.1-19 | schema_migration | payments.provider accepts cash and upi_manual | dev, prod | yes | — | superset CHECK, nothing to backfill; unblocks recording the advances this vertical actually takes | #cr-19 |
| CR-26.0.1-20 | schema_migration | Setu Card template registry re-authored, with the retirement guard | dev, prod | yes | — | governance was lost in the 2026-08-08 drop while switching kept working; Template Gallery is now in R1 | #cr-20 |
Why the record count grew from 11 to 15 (QRS-343). This defect would have refused the production deploy, and it was measured rather than reasoned about.
The prose target strings (10 of the original 11) were unusable by machine. deploy-prod.yml derives the declared migration set with target.split('/').pop().match(/^\d{14}/), so a target reading "supabase/migrations/ — the five migrations Prod is behind: …" yields no version at all. Simulated against Prod's live schema_migrations (21 applied, 9 genuinely pending), the committed manifest produced one declared version against eight pending migrations:
::error:: 20260801143000 PENDING on Prod but NOT declared … and 7 more
=> deploy REFUSES (exit 1)Three corrections to how this was first characterised, recorded because the first two were wrong:
- The blocking gate is
deploy-prod.yml, notcheck:release.check-release.jsnever readsbase_refat all —resolveBase()uses--base, then$GITHUB_BASE_REF, thenorigin/develop. Socheck:releasepolices a pull request's diff, not the release's range. base_refis documentary only. It was8e8514a(the repository's initial-import commit, 137 commits back) and is now291215c— the commit whose migration set matches what Prod actually held when the release began. That is a truthfulness fix, not the bug fix.- An earlier draft of this note claimed ~70
UNDECLAREDerrors at G2 from a 79-file diff. That was wrong, and a deliberately-failing negative test is what exposed it: removing a declared migration from the manifest did not fail the gate, because locally the diff was empty. The real fault was eight pending migrations invisible to the deploy comparison.
Two additions to the framework, both minimal:
targetnow accepts a string OR an array of strings. One deployable change routinely spans several files (an Edge Function isindex.tsplushelpers.ts). This does not weaken the control — every changed production file must still be claimed by some record — and four new validator tests, one a mutation guard, hold that line.deploy-prod.ymlflattens arrays too; without that it would comma-collapse each group onto one line and silently drop every path but the last.applied_atper change. A release is explicitly not atomic, and the framework said so fortargets[]but had no way to say it per change. CR-03's five migrations reached Prod on 2026-08-04 during S3 while CR-05's are still pending, so the comparison read a completed record as a stale one and refused the deploy.applied_at.proddistinguishes the two; only a genuinely stale record blocks. Post-fix: 9 declared, 9 pending, 0 errors, deploy proceeds.
Five of these fifteen are PROBE_REQUIRED classes — ef_config, cron_job, cloudflare, third_party and app_build are invisible to every automated gate, so their evidence must be a live functional probe, never a claim that a value was set. QRS-273 was closed on exactly such a claim and the auth log later proved sign-in had never once worked.
CR-26.0.1-02 — reserved_slugs Prod↔Dev reconciliation
Already applied to Dev on 2026-08-03; still pending on Prod (where it is a no-op). Prod held 752 rows, Dev held 8, the repo declared 8 — and all 752 were unversioned, created directly in the live database. A db reset would have discarded 744 of them silently.
Verified functionally, not by row count: after applying, is_slug_reserved returns true for privacy, mumbai and zomato, and false for the real merchant slug hotel-krushna. Dev ended at 753 — the extra is www, which the repo seed has and Prod does not, because that seed is one of the five migrations in CR-03.
⚠ QRS-267 re-confirmed with a measurement. The file was authored as 20260803094500; MCP apply_migration recorded it as 20260803094212 — its own timestamp, 2m48s off. The file was renamed to match. Never let MCP name a migration version: read it back with list_migrations and reconcile. Twice-observed now, so treat it as behaviour rather than anecdote.
CR-26.0.1-03 — the five migrations Prod is behind
Measured 2026-08-03. The drift runs Dev → Prod, which is the opposite of the intuitive direction:
| Prod | Dev | |
|---|---|---|
| Tables | 59 | 63 |
| DB functions | 27 | 32 |
| RLS policies | 118 | 122 |
| Triggers | 16 | 17 |
| Migrations applied | 16 | 21 |
Migration histories are a clean prefix — Prod is a strict subset of Dev — so there is no conflict resolution to do. Prod is missing the slug-integrity trigger, the reserved-slug seed, and the entire ADR-0009 taxonomy (vertical_archetypes, capabilities, archetype_default_capabilities, profile_capabilities, get_business_domains, get_my_capabilities).
This is why onboarding's industry step is broken in production right now. It is also why a "clone Prod into Dev" would have destroyed five migrations' worth of work.
CR-26.0.1-04 — verify_jwt correction
Prod currently has verify_jwt = false on manage-profile, manage-account, manage-settings, manage-reminders and get-dashboard-data, while the three public_page_ops_* cron functions are true. config.toml declares the exact opposite.
Stated precisely, because the alarming reading is the wrong one: every affected function reads the Authorization header and calls auth.getUser() before doing any work — verified by reading the recovered source, not inferred. The in-function gate is what has been enforcing auth, so this is a removed defense-in-depth layer plus config drift, not an open endpoint. It still matters: the gateway is the cheap layer and it is off on the function that owns account deletion.
Order of operations: confirm each function enforces auth in-function before flipping, flip, then probe live. See QRS-303.
CR-26.0.1-12 — app_release_policy + the minimum-supported-version kill switch
Creates app_release_policy and the anon-readable get_app_release_policy() RPC that the app reads on launch (QRS-291). Not yet on Prod — verified against Prod's live schema_migrations, which holds 21 migrations ending at 20260801091104.
Seeded permissive (min_supported_version_code = 0), so its arrival blocks nobody. It must be applied before the first store build is live, because the app reads it at startup and the honest failure mode of a missing RPC on that path has not been characterised.
⚠ Never raise a platform's min_supported_version_code above its own live build — that bricks every user on that platform. The release gate cross-checks this against production state, which currently records no live build on either store, so the floor is null and unbounded today.
CR-26.0.1-13 — Payments kill switch
Creates payment_kill_switch and the anon-readable get_payment_kill_switch() RPC (QRS-315). Seeded CLOSED (enabled = false), and that direction is load-bearing rather than a default — it is the entire mechanism behind shipping payment code dark on 15 Aug and switching it on around 20 Aug once a live rupee has settled into a real bank account. Flipping it needs no deploy, no gate run and no store review.
⚠ A CDN-cached card can defeat this switch. The public card is edge-cached by design, so a cached page could keep showing a live pay button after the switch is flipped. That is a safety control a CDN can outlive, and it is unresolved — it must be settled before the 20 Aug switch-on, either by keeping the pay CTA out of the cached HTML or by purging on flip.
CR-26.0.1-14 — COMMENT ON backfill for the reminders model
COMMENT ON statements only: no structural change, no data change, no grant change. Backfills queryable documentation onto reminder_categories and three trigger functions, satisfying the database-documentation standard established in QRS-331. Verified against pg_description on Dev rather than by reading the migration file.
CR-26.0.1-15 — Onboarding persistence + sign-out fix
Commit e88fa35: saveProfileStep now actually persists, and sign-out honours result.ok. Touches manage-profile (index + helpers) and validate-user-input.
CR-26.0.1-16 — pricing_mode, and the price leak it closes
Adds catalog_items.pricing_mode ('fixed' | 'on_enquiry') with a default 'fixed', so no backfill is needed and every existing row keeps exactly its current meaning.
Why it exists. The Store design was forbidden from inventing a price-type field, so to express "price on enquiry" it used the only mechanism available: price_minor = 0, plus a form hint reading "Leave price at 0 to show 'Price on enquiry' on your card." That sentinel is ambiguous with a genuinely free item, and it ranks every undisclosed item as the cheapest thing the merchant sells, so "sort by price" and "under ₹X" return wrong answers in the vertical where price filtering is the primary query. QRS-456 has the full reasoning.
The consumer-visible part is not the column, it is the projection. get_public_catalogue now emits price_minor: null and compare_at_price_minor: null when the mode is on_enquiry. A column alone would have been a half-fix: that function is anon-callable and returns JSON, so a renderer merely declining to display the number would leave it readable in the response body — "price on enquiry" as a UI convention over a published figure, which for a property is the negotiating position. Withholding belongs at the boundary that publishes, the same reasoning that keeps stock_quantity out of that function entirely.
So publicCatalogItemSchema.price_minor became nullable, and a client typed non-null would throw on the first on-enquiry item. requires_min_app_build is null because no shipped app build consumes either catalogue function yet — this is a widening of the response type, not a narrowing, and it is free to make today for the same reason QRS-427's shape change was.
get_my_catalogue gains the key but does not null the price: the merchant set it and still edits it while it is undisclosed to visitors. That asymmetry between the two functions is the feature.
Declared separately from CR-08 despite sharing the ef_code class, because it is unrelated to the money loop and carries a different and larger blast radius: both functions sit on the signup critical path, so a regression here blocks new-vendor onboarding outright rather than degrading a feature nobody can reach yet.
CR-26.0.1-01 — Branch reconciliation
main and uat both sit at 8e8514a (2026-07-10), 48 commits behind develop, and the working branch is 66 ahead of develop. There are zero git tags. deploy-prod.yml refuses a real deploy unless the ref is main, so the pipeline cannot run until this is done — which is why it is CR-01 rather than a housekeeping task.
Sequence: feat/mobile-nativewind-b2 → develop → uat → main, then tag v26.0.1 at the cut.
Irreversible, with a forward fix rather than a rollback: shared branches only move forward. A bad merge is corrected by a follow-up merge, never by rewriting history that others may have pulled.
Build ledger
| Build | Platform | Number | Commit | Submitted | Outcome | Functional delta |
|---|---|---|---|---|---|---|
| none yet |
apps/mobile/app.json is stamped at 26.0.1 / 26000100 for both platforms. If either store rejects, the resubmission takes a new build number on that platform only, and the numbers diverge permanently — see production state for why that is expected rather than a defect.
CR-26.0.1-17 — orders, and why it lands on the ledger primitive
Creates orders and order_items and registers the orders feature. It exists because the first G-D discovery brief found that the v2 schema is a publishing platform while the release plan assumed a transacting one (QRS-480): festival_stall, the launch vertical, resolved to ten features of which not one was transactional, so it could publish a catalogue and count stock and could not take a single order.
Why ledger. That primitive is defined as "money and quantity movement, manual or online", which is what an order is. The placement was chosen for its consequences and then verified against a real reset database rather than reasoned about:
| Archetype | Industries | orders applicable? |
|---|---|---|
goods | boutique, dairy, festival_stall, kirana, sweet_shop, tiffin | yes, all six |
time | salon, tutor, yoga_fitness | no — they book |
expertise | car_sales, direct_seller, electrician_plumber, photographer, real_estate | no — they enquire |
So festival_stall gains orders with no composition change at all, and an estate agent cannot be sent an order, which is correct. The rejected alternative was catalogue, which would have handed "order this" to every industry with a catalogue including real estate.
⚠ The riskiest part is orders_select. It is the first policy in this schema that lets a non-member read a merchant-owned row, on the relationship buyer_user_id = auth.uid(). That is what makes a registered consumer's order history possible, and it is exactly the clause that leaks if subtly wrong — so the pgTAP suite asserts the negatives rather than the positives: a consumer sees the order they placed and not the anonymous order beside it in the same workspace.
Deliberately absent: advance/deposit handling, and any collection date. Both are unanswered discovery questions (9, 13, 15), and modelling them now would be inventing the answers.
CR-26.0.1-18 — payments, and why it is inert on purpose
Creates payments and the append-only payment_events, and registers payments as universal — every business can take money, so applicability is the wrong axis to gate it on. Gating lives on the other two: entitlement (a paid-plan capability, per the owner's decision of 2026-08-10) and availability (a workspace-scope grant written when Razorpay reports the linked account activated).
⚠ That second gate needs no new mechanism. The owner's requirement — a merchant with incomplete banking details must not be able to enable Pay Now, validated merchant-side so the button can never render on a public card — is ADR-0026 D3 reused verbatim, with a webhook as the writer instead of an ops human. The public card cannot show Pay Now for an un-onboarded merchant because the feature does not resolve, not because a screen remembered to check a flag.
Money invariants are enforced by the database, not trusted from the write path: gross = commission + vendor; refunds bounded by the amount collected; a captured row must carry both a capture time and a provider payment id. payment_events carries a UNIQUE (provider, provider_event_id) replay guard, because a provider will redeliver and a non-idempotent handler double-credits an order.
⚠ This ships INERT. payout_accounts does not exist, so nothing can write the availability grant and payments resolves false for every workspace until the Banking module lands. That ordering is deliberate: the schema that records money must exist before the onboarding that authorises it, because the reverse onboards merchants into a system that cannot record what it collects.
CR-26.0.1-19 — cash, because the notebook is full of it
Widens payments.provider to ('razorpay','cash','upi_manual') and adjusts two CHECK constraints (QRS-484).
Why, one day after the table was created. payments was constrained to razorpay on the reasonable-sounding assumption that a payment row records an online collection. The festival-stall discovery session then established the facts: advances are minimum 10%, taken in cash or UPI, hand to hand at the stall, tracked in a notebook, with a paper slip as the buyer's proof. Replacing that notebook is the product's concrete win for this vertical, so a payments table that cannot record cash records none of its contents — and payment_status = 'partly_paid', a value the creating migration deliberately reserved for exactly this case, would have been unreachable.
Two new values rather than one, because only upi_manual has an external reference (a UTR) to reconcile against; collapsing both into offline would erase the one distinction reconciliation needs.
⚠ The half that was easy to miss. payments_captured_is_complete demanded a provider payment id unconditionally, so a cash advance could never have reached captured and the offline path would have been unusable in practice. It is now narrowed to razorpay rather than removed — an online captured payment still requires a provider id, and that is asserted separately.
⚠ And a business fact stated plainly: an offline payment earns QRSETU nothing. The new payments_offline_has_no_commission enforces zero commission on offline rows, because the invariant being protected is not "the numbers add up" — a mis-set rate whose arithmetic balances would pass payments_split_adds_up — but "we only take a cut of money we actually moved". A merchant recording only cash advances therefore uses the product and pays no transaction fee. That is the honest position, and it means the revenue case for online collection has to stand on its own convenience — which per QRS-482 it may not, for this vertical, at all.
Contracting changes
| CR | Removes | requires_min_app_build | Oldest live build | Safe? |
|---|---|---|---|---|
| none |
Nothing has ever shipped, so there is no compatibility floor yet and no contracting change is possible in this release by definition.
CR-26.0.1-20 — switching survived the drop; governance did not
setu_card_templates was dropped on 2026-08-08 with the rest of the pre-v2 schema (QRS-411), and template switching kept working — because ADR-0019 D1 keeps manifest content in the repo and only metadata in the database. That resilience is real and it masked what was actually lost: no status, so a template could not be withdrawn from new selection while live cards kept rendering; no retirement guard, which was the entire "a live vendor card never breaks" guarantee; no seasonal window; and no feature_code hook. The owner has since moved the Template Gallery into R1, and the gallery's screens are a projection of this table's columns, so designing them first would have meant designing against a guess.
Three things this version can do that the pre-v2 one could not. archetype_keys is validated against business_archetypes by a trigger mirroring industries_validate_primitives — and there are three archetypes now, not the five the old registry assumed, because ADR-0021 established that ecommerce_cart and catalog_informational were compositions (QRS-392). feature_code is a real foreign key to features, so ADR-0007's ban on gating a template by a hardcoded tier string is enforced by the database rather than by a comment; the pre-v2 table's required_tier CHECK was that exact violation. And is_selectable is computed at read time, never stored, because status plus the seasonal window plus the clock are three inputs and storing their conjunction is a duplicate source of truth that disagrees the moment a window closes.
⚠ There is deliberately NO foreign key from setu_cards to this table, and the reason is evidence rather than preference. 2026-08-08 is the experiment that already ran: the registry vanished and every public card kept rendering. An FK would have converted a metadata outage into an outage of the platform's only public surface. The registry governs SELECTION; the repo manifest governs RENDERING. Integrity for new selections belongs at the write path, which is where this platform already enforces everything else.
⚠ And one defect found by the gates rather than by review, worth recording because the class recurs: the first version of this migration enabled RLS and omitted the explicit REVOKE. Supabase ships ALTER DEFAULT PRIVILEGES ... GRANT ALL ON TABLES TO anon, authenticated, so every newly created table in public arrives with 7 privileges already granted to both roles, and an earlier blanket REVOKE ALL ON ALL TABLES does nothing for a table created later. v2_isolation_test.sql §A caught it as have: 7, want: 0 — that assertion earning its place, and see QRS-510 for making it a commit-time gate instead of a test-time one.
CR-26.0.1-21 — the state four surfaces needed and the schema could not store
orders.status was pending | confirmed | completed | cancelled. The Round 6 chat design filtered on ready, and so did a conversation-row chip, a notification kind and the consumer's order axis. ready did not exist. confirmed means we will make it and completed means you have it; for a festival stall the gap between them is days, and it is the moment that earns the one push notification this product will ever send a buyer.
⚠ The provenance is the finding, not the fix. The gap came from a design prompt of mine that asserted the triage pills were "derived from the customer's ORDER STATE, which the product already knows" — written without opening the orders migration first. Claude Design implemented it faithfully. A design brief is an unreviewed specification, and a false claim inside one propagates into an implementation that looks correct; no gate in this repo watches that channel. See QRS-537.
Widening a CHECK is a pure expand, so nothing can fail validation and no reader breaks, and ready is neither the default nor emitted by any write path yet. Now is the cheapest this will ever be: no environment has a live order and no store build is live, so assertCompatibilityFloor returns early and contraction is still unbounded. Once a build ships, narrowing this column becomes bounded by the oldest live app build; adding a value never does.
The value is ready, not ready_for_collection, because one orders table serves all three archetypes and "collection" is a Goods word — a yoga class is not collected. Human wording resolves through FULFILMENT_LABEL_KEYS[archetype][status] in @qrsetu/domain, which is a lookup table and not a branch on archetype (ADR-0021 D4).
CR-26.0.1-22 — chat, and the authorization layer that fails silently
Seven tables across this and CR-23, but the schema is the easy part. The part worth reading twice:
RLS on
public.messagesdoes not protect the Realtime channel.
They are two surfaces. Policies on the table govern the REST read; policies on realtime.messages govern who may attach to a websocket topic. Ship only the first and any authenticated user can subscribe to chat:<any-uuid> and watch strangers' conversations arrive live — and nothing fails to reveal it, because the subscription simply succeeds. There is no error, no denied request, and no log line. It is the single most omittable control in this release.
It is sharper here than it would be in most products: category 3 makes authenticated the logged-in general public, and by orders of magnitude the largest population on the platform. So no policy in either file grants by role; every one scopes by participation in the conversation. chat_topic_conversation_id() fails closed by construction — a malformed or hostile topic resolves to NULL, which no conversation matches, rather than raising and turning a crafted topic name into an error in the subscription path.
⚠ Deploy-order constraint, deliberate: this migration fails to apply if realtime.messages is absent. An environment without Realtime cannot run chat, and failing at apply time is vastly preferable to shipping an unauthenticated firehose. Do not wrap the policy in an existence guard — a migration that silently skips a security policy is indistinguishable from one that never had it.
Two model decisions that are not stylistic. A conversation is between a USER and a WORKSPACE, not two users, because several staff may legitimately answer for one vendor and a user-to-user thread breaks the first time a merchant hires someone. And consumer_user_id is NOT NULL, the opposite of orders.buyer_user_id — an order may be anonymous, a message may not, because sending the first message is the registration gate and an unidentified sender is unmoderatable. Oversight is deliberately not granted: an ancestor workspace may read a descendant's orders, never its message content.
CR-26.0.1-23 — the controls the client could not enforce
⚠ The only non-additive touch in this set is an ALTER on public.workspaces — away_reply_enabled boolean not null default false and away_message text. Both are nullable-or-defaulted so it remains an expand and every existing reader is unaffected, but workspaces is the tenancy root and the busiest table in the schema, so it is called out here rather than buried in a table list.
Three things move from the client into the database, because a client-only control is not a control:
| Was | Now |
|---|---|
doBlock changed a header label and left the composer able to send — so the action's own subtitle, "They cannot message you again", was false | messages_reject_when_blocked() refuses the insert, in both directions, with system messages exempt because the away reply is the product speaking |
PIN_CAP = 3 was a JS constant, so a second device exceeded it and the cap silently stopped being one | a trigger raising P0001 with a message naming what to unpin |
| pin order came from an array index, which does not survive becoming rows | pinned_at |
away_reply_enabled defaults to false on purpose: a merchant who configured nothing has opted into nothing, and the product speaking on behalf of a silent business is the bot-pretending-to-be-a-person failure the away reply exists to avoid. It is also inert without hours — openStatus() returns unknown when setu_card_hours is empty, and isOutsideBusinessHours() returns false for unknown.
And the one thing the design got right that the schema had to keep right: manually_unread_at is a separate column from messages.read_at. Rewinding read_at to implement "mark as unread" would flip the other party's read receipt back to delivered and assert something false about what they actually saw.
CR-26.0.1-24 — the read API, and the grant that should never have existed
Two RPCs: get_my_conversations() and get_conversation_messages(uuid, timestamptz, integer). Additive, no table or policy altered. Why it exists is more useful than what it contains.
The first draft of CR-22/CR-23 granted 22 table privileges to authenticated, and v2_isolation_test.sql §A rejected it: have: 22, want: 0. That assertion is not a hardening preference — it encodes the data-access architecture:
The app never queries a table directly.
Reads go through a SECURITY DEFINER RPC, writes through an Edge Function on the service role, and supabase.from() is banned in app code so table names never reach the network wire. Every other v2 table already followed it — orders, payments, setu_cards, catalog_items all revoke from authenticated and grant it nothing, giving grant execute only to functions. The grants were the defect; the test was right, and it was written before this feature existed.
⚠ But revoking them leaves the tables unreachable, and an unreachable table is exactly the pressure that makes the next person re-add a grant. So the invariant needs a sanctioned path beside it, which is this file. Fixing the symptom without providing the path is how the same grant comes back in three months with a plausible justification.
One consequence would have been very hard to debug in the field. The realtime.messages policy originally read public.conversations with a direct subquery. An RLS policy expression is evaluated with the caller's privileges, so once authenticated holds no grant that subquery matches nothing and every channel join is denied — silently. Chat would never deliver, with no error anywhere to point at the cause, and the obvious debugging instinct (checking the table's own policies, which are correct) leads away from it. It now routes through my_conversation_ids(), which is precisely what the my_* SECURITY DEFINER family exists for, and a pgTAP assertion greps the policy expression to keep it that way.
⚠ The risk this CR carries is the SECURITY DEFINER pattern itself: these functions bypass RLS, so the policies sitting next to them look like they apply and do not. Each therefore enforces participation in its own body. get_conversation_messages raises rather than returning empty, because an empty page is indistinguishable from a new thread and a caller must not be able to use that difference as an existence oracle for other people's conversations.
get_my_conversations() is also the answer to QRS-546: one round trip returning the counterparty, the derived category label, the unread count, the per-participant organisation state and the order context behind the merchant triage pills — instead of the seven separate aggregates the design prototype computed on every list open. It returns viewer_side, so one function serves both halves of the feature and neither app branches on who is asking.
CR-26.0.1-25 — the chat feature, four fields, and two corrections on the record
Registers chat and adds orders.collect_on, catalog_items.is_unique, workspaces.advance_pct/advance_terms and industries.item_attribute_schema. All additive or defaulted; no backfill, no reader changes behaviour.
⚠ Two claims I made to the owner were wrong, and the file header records both rather than quietly fixing them.
First: orders was reported missing from the feature registry. It had existed since 20260810100000_v2_orders.sql:57 — same category, same ledger primitive, same depends_on, same sort order as the draft then spent a turn "deciding" on. The mechanical error was grepping one seed file when this repo's convention is that a feature is registered by the migration that creates its table. The duplicate insert then failed on the primary key and took the chat row with it, so nothing applied at all, and the runner reported success because cmd | tail; echo $? reports tail's status (QRS-553).
Second: item_attribute_schema was reported absent. It exists on business_archetypes. The real defect is the grain — three archetypes cannot carry the different attribute sets six goods industries need, so this adds an industry-level override resolved industry-first then archetype (QRS-556).
⚠ The blocker this file deliberately does not close. orders and payments have no plan entitlement grants, so resolve_features reports default:denied and no merchant on any plan can take an order today. Six scoped features are in that state. The mechanism is complete — feature_grants carries limit_value/limit_period and store/free already demonstrates the shape at 25/total. What is missing is the plan mapping: platform_plans holds free/pro/ enterprise while the commercial model is three paid tiers at 100/500/unlimited orders. Seeding a wrong cap silently blocks real merchants, so it waits on an owner decision (QRS-554).
And what did not change: festival_stall.enabled_primitives is untouched, as is the fulfilment feature. Collection is an order state transition (ready → completed). Giving a stall the fulfilment primitive would have dragged in schedule and party through fulfilment → bookings → customers, handing it an empty Bookings screen and an empty CRM.
CR-26.0.1-26 — the features that were born denied, and the Pay Now gate that was never closed
Six feature_grants rows and no schema change, but two of them matter more than the count suggests.
The defect (QRS-559). The grant seeds in 20260808150000_v2_resolve_features.sql are set queries over the registry as it stood on 8 August — select f.key … where f.applicability = 'universal', and again for pro/enterprise over 'scoped'. They ran once. The entitlement axis is deny by default, so every feature registered afterwards resolved entitlement_source = 'default:denied' on every plan including Enterprise: orders, payments, and chat — registered by CR-26.0.1-25 the previous day and therefore dark for everybody.
⚠ On chat the error was mine and it has a name. I argued that applicability = 'universal' meant no composition edit could silence an inbox. True, and irrelevant: I reasoned about the applicability axis and never inspected the entitlement one. Same root cause as QRS-553.
⚠ The count I reported was also wrong in the other direction. I named six features; four of them (customers, bookings, campaigns, assets) predate the seed and hold Pro and Enterprise grants. What they lack is a free grant, which is deliberate. Three features dark on all plans is a defect; four absent from the free tier is a price list.
No free grant for orders — owner decision of 11 August, to be tuned later through the admin Feature Control panel rather than by migration. Encoded as absence rather than an explicit deny row, because entitlement_source cannot express a deny honestly (QRS-561). The consequence, stated so nobody rediscovers it from a blank screen: a free merchant can publish a card, list 25 items, chat and record cash sales, but cannot create an order record — so Collections and OrderDetail are paid surfaces.
The sixth row is the real find (QRS-560). 20260810110000_v2_payments.sql asserts that "the public card cannot render a Pay Now button for an un-onboarded merchant because the feature does not resolve." It resolves. The availability axis allows by default for an active feature, and payments was registered without a status. So the five entitlement grants above would, on their own, have rendered Pay Now for every paid merchant with no linked account and no bank details — the precise case that file claims is impossible. Closed with a platform-scope availability deny: workspace precedence (60) beats platform (10), so the Razorpay webhook's per-workspace grant overrides it on activation.
⚠ The hazard that ships with it, carried in the row's own reason so it is visible in the admin panel: a plan-scope availability grant (rank 40) would also override this deny, for an entire plan at once. Availability for payments may only ever be granted at workspace scope.
Why there is no auto-grant trigger, having proposed one. A trigger writing the platform entitlement grant for every universal feature is the obvious guard, and payments disproves it — universal in applicability, deliberately paid by owner decision, so the trigger would have reversed a commercial decision silently. Replaced by a pgTAP completeness assertion: every feature must hold at least one entitlement grant somewhere, reported by name. A check that fails loudly and forces a decision beats a trigger that makes one.
Verified: 30 migrations from scratch, RESET_EXIT=0; pgTAP Files=6 Tests=188 PASS. Mutation-run in both directions, which caught a weak assertion in the new test file itself — see 04-test-evidence.
CR-26.0.1-27 — reminders schema, feature registration and grants
supabase/migrations/20260811150000_v2_reminders.sql · schema_migration · reversible · dev + prod
Three tables, one predicate, one trigger function, two triggers, four policies, two indexes, four seeded categories, the reminders feature row and its platform entitlement grant. Purely additive.
This closes an inverted feature (QRS-564), and the inversion is the point. Reminders arrived screen-first and got stranded. What already existed: a 704-line screen, a 308-line hook, a complete notification port layer with 7 test files, 198 lines of Zod naming these exact column names, and ~1,000 lines of DST-correct domain logic with ~117 node --test cases. What did not exist: the tables, both RPCs, the Edge Function, and a row in public.features. That last one is what made it a defect rather than a gap — resolve_features could not return reminders on any axis, so the one capability the design calls universal was ungateable, unentitleable, and impossible to ship dark.
The model is a RULE plus SPARSE occurrence EXCEPTIONS, never a materialised series. The legacy table carried one is_completed boolean per rule, so completing one Tuesday completed every following Tuesday — unfixable by constraint, which is why the schema was replaced rather than repaired. Absence of an exception row means pending.
Owned by a USER, not a workspace — the only v2 feature that is. workspace_id is a nullable tag. Three reasons: a category-3 consumer has zero workspaces and must still hold reminders; an employee's personal notes must stay outside their employer's ADR-0023 oversight path, so these policies carry no oversight clause and pgTAP asserts their absence; and packages/schemas already specified it.
Two corrections to my own work, recorded rather than quietly fixed
⚠ QRS-564's own note was wrong. It said to register reminders with primitive_key = 'recurrence'and applicability = 'universal'. That combination is refused by features_universal_has_no_primitive, correctly: a universal feature bypasses the applicability axis, so its primitive would never be read and would imply a derivation that never runs. primitive_key is NULL, as all seven pre-existing universal rows have it.
⚠ The first application left a real hole. TRUNCATE, TRIGGER and REFERENCES remained on all three tables for anon and authenticated, because Supabase's default privileges grant every new table and 20260808210000's blanket revoke cannot reach a table created later (the QRS-214 class). Measured, not theorised: these were the only three tables in the schema where anon held anything. TRUNCATE to anon would let a caller that reaches SQL destroy every reminder on the platform. Explicit revoke added, and asserted in pgTAP §A so a future table in this feature fails loudly instead.
check:sql then caught two more genuine defects in this file, both fixed rather than waived: four policies with no TO clause (Postgres defaults to PUBLIC, i.e. anon — the exact invisible default behind ADR-0014), and four missing COMMENT ON objects.
Verified: dropped and cleanly re-applied from scratch, EXIT=0; pgTAP reminders_test.sql66 assertions, 0 failures.
CR-26.0.1-28 — reminders read API
supabase/migrations/20260811160000_v2_reminders_read_api.sql · schema_migration · reversible · dev + prod
get_reminders and get_reminder_categories, both SECURITY DEFINER, both granted to authenticated only. Split from CR-26.0.1-27 because one migration is one Change Record, and a function rolls back by CREATE OR REPLACE while a table does not.
get_reminders takes no user parameter, deliberately. SECURITY DEFINER bypasses RLS, and scoping to auth.uid() is unsteerable by the caller; accepting a user id would be an IDOR for no benefit, since there is exactly one correct answer per caller. Contrast get_my_catalogue(p_workspace_id), which must accept a workspace and therefore checks the relationship in its body.
Four behaviours worth naming, each chosen against a specific failure:
- It raises when
auth.uid()is NULL (i.e. under the service role) rather than returning an empty page — an empty result that means "wrong caller" gets debugged in the client for an hour first. - A partial cursor is rejected: keyset on
(starts_at, id)with one half missing silently skips or repeats rows sharing a timestamp. p_limitis clamped to 1..200, not rejected, so a client asking for 10,000 still gets a screen.- The window bounds occurrences, not rules — a two-year daily reminder holds ~730 exception rows and one screen needs the few in view. pgTAP proves the rule is still returned when its occurrences are filtered out.
⚠ It does not expand recurrence, and must never learn to. Expansion is pure @qrsetu/domain code with ~117 node --test cases and no database, which is what makes the iOS scheduled-notification budget testable at all (ADR-0016). An expand-window parameter here would move DST arithmetic into SQL where none of those tests reach.
get_reminder_categories deliberately does not project key, even though the column exists: reminderCategorySchema declares {id, name, color_token, icon}, and adding an undeclared column is how a strict parser starts failing on a migration that looked additive.
Verified end to end against real Postgres — shape, keyset paging including the probe-row trim and a null cursor on the last page, the window filter, owner isolation, all three rejections, both clamp directions, and the service-role raise. pgTAP §G.
CR-26.0.1-29 — v2_chat could never apply on Supabase
supabase/migrations/20260811100000_v2_chat.sql · schema_migration · reversible · dev + prod
One statement removed and replaced by an assertion, inside an already-declared migration (CR-26.0.1-21).
⚠⚠ The original could never succeed. alter table realtime.messages enable row level security fails must be owner of table messages (SQLSTATE 42501) on the local stack and on qr-setu-dev, aborting db push at this migration. It was found on 2026-08-12, on the first push attempt — the file had passed review, check:sql and deno check, none of which can see it.
The blast radius was far larger than chat. Dev was behind by thirteen migrations: six applied, seven stranded behind this one. So orders, payments, chat and reminders were all absent from Dev, which is why the reminders backend appeared undeployed — it was queued behind a migration that could not run.
Diagnosed on Dev, not reasoned about: realtime.messages is owned by supabase_realtime_admin, postgres is not a member of that role, and RLS is already enabled on the table because Supabase enables it as part of Realtime Authorization. The statement was therefore simultaneously impossible and redundant. CREATE POLICY on the same table was probed separately and succeeds as postgres — so the policy is untouched and only the ALTER is gone.
The security intent is strengthened, not weakened. It is replaced by a DO block that raises if realtime.messages lacks RLS. Attempting to enable RLS proves nothing about the end state; checking it proves the end state directly and still fails closed. The migration's own comment forbidding "an existence guard that silently skips a security policy" is honoured — nothing is skipped.
Editing rather than superseding is correct here, against the usual rule: the file had never applied successfully in any environment, so no deployed state depends on the old text, and a later migration cannot repair an earlier one that aborts the run.
The lesson, and it is the expensive one: a migration is not verified until it has been APPLIED somewhere. Three gates and a review passed on a file that could not execute. A completed supabase db reset is the cheapest control that would have caught it.
Verified: db push applied all 13 pending migrations to qr-setu-dev, EXIT=0, both reminders versions read back present, and a live round trip on Dev created a weekly rule, deduped a double-completed occurrence to one row, and rejected a malformed recurrence.
CR-26.0.1-30 — slug availability becomes a pure SQL read
supabase/migrations/20260813130000_v2_setu_card_slug_status.sql adds public.resolve_setu_card_slug_status(text), returning exactly available | reserved | taken. SECURITY DEFINER, granted to authenticated only.
Why it exists: the onboarding slug step's CTA could never enable. The availability check returned HTTP 500 on every call, for two independent reasons, either one sufficient, both measured against qr-setu-dev with curl rather than reasoned about:
validate-user-inputcalledrpc('is_slug_reserved', { check_slug: value })while the function isis_slug_reserved(p_slug text). PostgREST binds RPC arguments by NAME, so the call 404'd withPGRST202— its own hint read "Perhaps you meant to call the function public.is_slug_reserved(p_slug)".- Its uniqueness half read
setu_cardsthrough an anon client, andanonholds noSELECTgrant on that table. That returns42501 permission denied— an exception, not an empty result. Fixing (1) alone would still have 500'd, which is why the probe matrix mattered more than the first hypothesis.
Not fixed by a grant, deliberately. GRANT SELECT ON public.setu_cards TO anon — which is what PostgREST's own error hint suggests — would publish every merchant's draft card. SECURITY DEFINER keeps both underlying tables unreadable, which is the pattern is_slug_reserved already documents for reserved_slugs (whose 753-entry list is withheld because it reveals which brands we treat as impersonation risks and which categories as prohibited).
This is an affordance, not the enforcement layer, and the distinction is load-bearing: setu_cards.slug is UNIQUE and the slug_not_reserved + slug_write_once triggers reject a bad handle at write time, so a caller who skips or ignores this check still fails closed with a 409 at provisioning. The RPC exists to tell the merchant before they finish the wizard.
One function rather than two halves the latency and — the part that matters — makes reserved reachable. SlugStatus has carried it since it was written and @qrsetu/i18n has separate copy for it, but the EF's binary isValid could not express it, so the client documented collapsing it to taken: telling a merchant that someone else holds a handle that is in fact permanently unclaimable.
Verified on Dev in both directions (QRS-013). As authenticated: available/reserved/taken all correct, including mixed-case and whitespace-padded input (citext + btrim), with taken proven using a real draft card inside a rolled-back transaction (0 rows persisted afterwards). As anon: permission denied for function. Registered under its filename version 20260813130000, so no QRS-267 orphan-version drift.
⚠ One self-inflicted delay worth recording: the first push failed with "type citext does not exist". citext lives in the extensions schema, so an explicit ::citext cast cannot resolve under set search_path = public. The live get_public_setu_card already does this identical comparison with no cast, relying on implicit text-to-citext coercion. Copying the proven pattern was the fix; the un-cast form also keeps the comparison index-usable.
CR-26.0.1-31 — the validate-user-input Edge Function is deleted
Removed from the repo and undeployed from qr-setu-dev. Owner decision, 2026-08-13: "I suggest to delete that EF validate user input to avoid further confusions."
Enumerated rather than assumed before deleting — the owner's own condition was that removal must not break anything that legitimately depends on it. Its only runtime caller anywhere in the repo was packages/data/src/onboarding/service.supabase.ts's checkSlug, replaced by CR-26.0.1-30 in the same change. Its other four validation types (email, password, url, text_length) had zero callers and never had any.
Why delete rather than repair. CLAUDE.md's own decision rule already assigned this correctly and was not followed: pure SQL read + no secrets + no external HTTP → RPC; everything else → Edge Function. An availability check is a pure SQL read. Routing it through an EF also added a Deno cold start (~200-500ms) to a check that fires on debounced typing.
It is additionally a security improvement. The function took tableName and columnNamefrom the request body and passed them straight to .from(); its own comment acknowledged the hazard and could only mitigate it by relying on RLS. Deleting it removes the last endpoint in the repo shaped that way, and the last verify_jwt = false function — config.toml now has none.
Not archived to _archive_pre_v2/, on purpose: that directory means "written against the pre-ADR-0020 schema", and this was a v2-era retirement. Filing it there would misdescribe it. Git history is the record.
Verified: supabase functions delete returned "Deleted Function validate-user-input"; a follow-up anonymous POST returns HTTP 404 NOT_FOUND where it previously returned 500; and the shipped web bundle was grepped to confirm the string validate-user-input no longer appears in it.
⚠ A gate blind spot this exposed, logged as QRS-641-adjacent: check:docs-impact demanded a README.md for supabase/functions/validate-user-input/ — a directory this change deletes. It reads the diff's changed-file list without distinguishing a deletion (D) from a modification, so it asks for documentation of something that no longer exists. Discharged with a Docs-Impact: trailer, which records the reason in git history rather than in a terminal.
CR-26.0.1-32 — chat media: voice notes, photos and documents
Design round 22 added audio messages and attachments to both halves of Chat. This is the schema for them, and the whole record here is one avoided defect plus one decision that must not be re-litigated later as if it were an assumption.
The defect: the first draft locked every consumer out of chat media by construction. public.media already existed for business assets, so its ownership column was workspace_id NOT NULL. Chat media has a different owner — the conversation — and a category-3 consumer has zero workspace memberships. A buyer's voice note was therefore unrepresentable. CLAUDE.md's three-user-category principle states the test verbatim: "any resolver, RPC or policy that requires a workspace to answer a question cannot serve category 3 — that is a design defect, not an edge case."
The fix is not a second attachment table. workspace_id becomes nullable, conversation_id is added, and ownership is exactly one scope, enforced as an XOR:
sql
constraint media_scope_exactly_one
check ((workspace_id is not null) <> (conversation_id is not null))Two nullable columns plus a convention would have allowed a row owned by both, readable through two unrelated policies at once. media_purpose_matches_scope then pins each purpose to its scope, so a chat_voice cannot be filed as a business asset.
The audio format is stored as DATA, and that is what makes it revisable. codec (aac_lc|he_aac|opus) and container (m4a|ogg|caf|webm) are CHECK-constrained columns, alongside duration_ms, peaks smallint[], sample_rate and bitrate_bps. So AAC-LC in .m4a is a default, not a schema commitment — see QRS-654 for why Opus loses today and why the reason is a container gap rather than a codec one. peaks is pinned at exactly 32 non-null samples in 0..100, because the waveform renders from stored peaks on both the merchant and consumer side and a variable-length array is how the two renderers start disagreeing about the same recording.
Two defects caught by controls rather than by reading, both before apply:
check:sqlrejected the newmediaSELECT policy for having noTOclause. Postgres defaults that toPUBLIC, i.e. includinganon— every chat attachment row exposed. Same invisible default behind ADR-0014, and the same defect CR-26.0.1-27 hit on four policies at once.- The peaks range check was first written as a subquery inside a
CHECKconstraint, which Postgres refuses outright. It would have failed on apply. Rewritten as0 <= all (peaks) and 100 >= all (peaks).
Verified: applied to qr-setu-dev; local and remote migration versions agree with no orphan rows (QRS-267); pgTAP chat_test.sql §F, 14 new media assertions, PASS.
CR-26.0.1-33 — per-message actions: forwarding and Favourites
The same design round added a per-message action sheet. Two things it needs from the backend, and one thing it deliberately does not get.
messages.forwarded is a boolean, not a provenance chain. Storing the source message id would let a recipient learn that a conversation they cannot see exists — an existence oracle for another buyer's thread, in a two-sided marketplace. So the schema records that a message was forwarded and never from where. This is the same reasoning get_conversation_messages already applies by raising rather than returning an empty page (CR-26.0.1-24).
Favourites is per-account state, so it gets its own table shaped exactly like conversation_states — message_states (message_id, user_id, starred_at, updated_at), PK on the pair. get_conversation_messages is DROP/CREATEd once for both of today's rounds (media columns and forwarded and starred), because two signature changes to one function in one day is two requires_min_app_build floors and two probes for no benefit. starred projects through left join public.message_states ms on ms.message_id = m.id and ms.user_id = auth.uid().
The defect: the first draft granted four table privileges to authenticated. v2_isolation_test.sql §A refused it — have: 4, want: 0. Same class as CR-26.0.1-24's 22 grants, and the test is right for the same reason: the app never queries a table directly, reads go through a SECURITY DEFINER RPC, writes go through an Edge Function on the service role, and supabase.from() is banned in app code so table names never reach the network wire. A grant to authenticated is not a convenience, it is a hole in that architecture. Grants removed; the message_states_own policy is kept as documented-unreachable defence in depth, and the pgTAP assertion was reworded to claim only what it proves.
⚠ Consequence to know, stated rather than discovered later: message_states has no reachable write path until the chat Edge Function exists, so Favourites is schema-complete and functionally dark. That is deliberate. Shipping the table with a client-writable grant is exactly the pressure the isolation test exists to resist.
check:sql additionally caught a missing COMMENT ON TRIGGER.
Verified: applied to qr-setu-dev, versions read back with no orphan drift; pgTAP chat_test.sql §G — including that a merchant sees starred = false on a message the buyer starred, which proves the per-account guarantee from the other side rather than from the owner's. Suite 268 → 277 assertions, PASS.
CR-26.0.1-34 — get_my_catalogue projects is_unique
supabase/migrations/20260814090000_catalogue_project_is_unique.sql · reversible · no compatibility floor
One column added to one projection, and the reason it needed a migration at all is the interesting part.catalog_items.is_unique has existed since CR-26.0.1-25, with a CHECK tying it to stock_quantity <= 1. The column was never missing. What was missing was the READ: get_my_catalogue projects an explicit key list rather than row_to_json(i.*), which is a deliberate choice — the merchant and public catalogue shapes differ on purpose, and an implicit projection would leak whichever column arrived next onto the public surface. The cost of that choice is that a new column is invisible to the client until someone names it here.
What could not be drawn without it. The dashboard's stock_split widget carries two denominators, and the design states why in its own words: "the bar counts listings; these count pieces, split into one of a kind and countable stock so nothing adds an idol to a prasad box." One-of-a-kind versus countable is is_unique. With every line treated as countable, both figures collapse into one number and a stall holding 40 individually carved murtis alongside 200 prasad boxes reads as 240 interchangeable pieces.
⚠ What this does NOT fix, stated so it is not later read as closed. The same widget's middle segment is reserved, and that is a relationship rather than a column: orders carry a text item_summary and no per-item line rows, so no query can attribute a reserved piece to a listing. reserved therefore stays 0 everywhere — which is the correct answer for a product that records no reservations, not a placeholder. Order line items are QRS-661.
Expand-contract in both directions, which is why there is no requires_min_app_build: an older app ignores a key it does not know, and catalogItemSchema.is_unique carries .default(false) so a newer app running against a database where this has not been applied still parses the catalogue rather than failing the whole read.
CR-26.0.1-35 · The Business plan, and the resolver that finally answers something other than free
supabase/migrations/20260815172134_v2_business_plan_and_workspace_subscriptions.sql · schema_migration · reversible
Owner decision, 2026-08-15, taken after visiting twelve live Ganapati stall vendors: one paid plan for this MVP — Business, ₹9,999 per festival session plus a 5% platform commission on online payments.
It lands at rank 200, which the 2026-08-08 plans migration deliberately left vacant for exactly this tier — so this is the seat being taken rather than a tier being squeezed in.
workspace_subscriptions is the smallest table that lets resolve_workspace_plan() do its job. That function has returned a hardcoded 'free' for every workspace since 2026-08-08, and its own comment prescribed the remedy: "Wave 2 replaces this body with: the workspace's own subscription… the resolver is untouched." That is precisely what happened — one function body, and resolve_features did not change at all.
Two deliberate non-actions, recorded so neither reads as an oversight:
- "Per session" does not fit
billing_period IN (month, year, none). It is recorded asnoneplus a description that carries the human meaning. Extending that vocabulary is its own decision, not something to smuggle in beside a plan seed. - No plan-scope
availabilitygrant forpayments. CR-26.0.1-30 is explicit that availability for payments may only ever be granted at workspace scope, when that specific merchant's Razorpay linked account activates. A plan-scope row would light Pay Now up for every Business merchant with no bank account — the exact incident that constraint exists to prevent.
This supersedes the open "orders on free" question. free keeps no orders row (the 2026-08-11 owner decision stands); vendors reach order-taking by buying Business.
CR-26.0.1-36 · get_my_order_book — the read without which orders were unreachable
supabase/migrations/20260815180000_v2_orders_read_api.sql · schema_migration · reversible
orders, order_items and payments shipped on 2026-08-10/11 with a reviewed schema and no way for any client to reach them. authenticated holds zero table privileges by design, so Collections, Order detail and Payments were all built against a stub fixture — the order model was real and reviewed, and every number on those screens was a fixture.
One composite read, not three. The day list, the awaiting-a-date queue and the money position all render on one screen. Three parallel RPCs against QRS-290's measured 130-160 ms floor is half a second of blank on the coldest screen in the app, and one payload is also the only thing that makes a header total unable to disagree with the sections beneath it.
It reproduces orders_select as its first statement, minus the consumer clause. SECURITY DEFINER bypasses RLS, so the policy is enforced here or nowhere. The consumer direction is deliberately absent: this is the merchant's book, and a buyer reading their own order is a different function with a different shape and no other buyer's PII in it. It raises no_data_found rather than "forbidden", so the error cannot be used to enumerate workspace ids.
⚠ today and every date are computed in the WORKSPACE's timezone. collectionsByDay classifies every row as overdue / today / tomorrow against an injected today. If the client supplied its own date, a merchant with a wrong phone clock would see a different working day from the one their stall is actually having — and "overdue" would be wrong in the direction that loses an uncollected idol.
CR-26.0.1-37 · complete_order_with_settlement — the two facts that must not be separable
supabase/migrations/20260815190000_v2_complete_order_with_settlement.sql · schema_migration · reversible
This function exists for one reason: the failure that lands between two writes. The design's CTA reads "Mark collected, balance received" — a status change and a payment. An Edge Function issuing two PostgREST calls cannot make them atomic, so the failure mode is the payment recorded and the status not, or the reverse, with real money on one side.
It locks the order FOR UPDATE (two counter staff settling the same order at a busy stall is not hypothetical), derives payment_status from the ledger rather than accepting one, refuses a terminal order, and refuses a razorpay settlement by name — a gateway payment is written by the webhook from a signed event, so a client-triggered one is money that never moved.
Verified against Dev in both directions. ⚠ Worth recording that the first razorpay test was a false pass: it ran against an already-completed order, hit the terminal guard, and never reached the provider check at all. Re-run against a live order it genuinely refuses. A guard tested only in its passing direction is not a guard, and a guard that was never reached is not even tested.
CR-26.0.1-38 · manage-order — every merchant-side order write
supabase/functions/manage-order/index.ts · ef_code · reversible
Six actions: create_manual · set_collection_date · set_fulfilment_status · complete · cancel · record_payment.
No total is ever taken from the client. Line totals and the order subtotal are computed from unit price × quantity. orders_total_is_subtotal_plus_tax is a database CHECK, and a client supplying its own totals can satisfy it with a consistent set of wrong numbers — which is exactly why orders_write_member's own comment keeps order writes out of client hands.
Guard order is a security property, mirroring manage-item: workspace and membership are established before the idempotency claim, because key is the sole primary key of the v2 ledger, so a replay lookup is global — checking it first would let anyone holding a key read another workspace's stored response.
Verified live against Dev with a real JWT over HTTP — 11 of 11: totals derived server-side; an idempotent replay returning the first result and creating no second order; a client-supplied payment_status ignored; razorpay refused (400); cross-tenant read and write both refused; atomic settlement reaching completed/paid; a terminal order refusing further writes (409); and the feature gate refusing on the free plan with a message naming entitlement=default:denied — the exact axis the locked UI reads, so the lock affordance and the server refusal now provably agree.
Ships 22 co-located Deno tests; the full Edge Function suite is 160 passing.
CR-26.0.1-39 · config.toml declares manage-order and manage-reminder
supabase/config.toml · ef_config · reversible
manage-reminder had no [functions.*] entry (QRS-643), and the consequence turned out to be measurable rather than theoretical: the live function on Dev is deployed verify_jwt = false. It is a Type-A user-facing function whose every action calls requireAuth, so the body still rejected unauthenticated callers — defence in depth lost, not an open hole — but the posture was an accident rather than a decision.
An absent entry is a silent default. That is what check:fn-config exists to catch, and exactly what it cannot catch while it exits with usage unless handed --project <ref>. Both functions are now declared verify_jwt = true; the next deploy of manage-reminder tightens it. Fixing the gate itself remains QRS-643.
CR-26.0.1-40 · order_items snapshots hsn_sac, tax_rate_bp and price_includes_tax
supabase/migrations/20260817090000_v2_order_items_tax_snapshot.sql · schema_migration · reversible
The one genuinely lost-forever item of QRS-707, and the reason it landed before the festival while the provider-fee columns did not.
Every other missing financial fact is recoverable: payment_events stores Razorpay's verbatim payload before processing, so a fee, a transfer id or a settlement reference can be backfilled from rows we already hold. No provider payload anywhere contains our tax fields. They live only in catalog_items, which is merchant-mutable, so the classification and rate at the instant of sale become irrecoverable the first time a vendor edits the item, silently and with no trace. 20260808180000_v2_catalogue.sql had already said this in the repo's own words: "added after orders exist, historical invoices cannot be reconstructed."
⚠ And order_items.tax_rate_bp had existed since 20260808180000 and was written by nothing. That is the worst version of the defect: a reader inspecting the schema sees the field and reasonably concludes the fact is being captured. It was not. The two columns added here are what make it usable, and the write path lands in CR-41.
Two nullable columns, no default and no inline CHECK, so this is catalog-only in PG11+ at any row count and takes no validation pass under ACCESS EXCLUSIVE on a live money table. There is no closing window and therefore no reason to rush a constraint.
orders.tax_minor stays hardcoded 0 deliberately, and the blocker is not the CHECK: place-public-order mints the Razorpay link on subtotalMinor while derivePaymentStatus compares captured against total_minor. They coincide only because tax is zero. The moment it is not, every fully-paid order derives as partly_paid and the Route transfer is computed on the wrong base.
CR-26.0.1-41 · place-public-order writes the tax snapshot and reads commission through a shared seam
supabase/functions/place-public-order · ef_code · reversible · depends on CR-26.0.1-40
Two changes.
The write half of CR-40. The three tax fields are selected from catalog_items and copied onto every order line. Snapshotted, never joined: a join would restate a past sale under today's classification, and a GST return filed against a restated HSN is a statutory error rather than a display bug. A NULL honestly records "the merchant had not classified this when it sold", which is distinguishable from a zero-rated classification.
A commission seam, not a commission table. The inline commissionBp() helper becomes _shared/commission.ts resolveCommissionBp(client, workspaceId), whose body still reads PLATFORM_COMMISSION_BP today. Nothing about the rate changes. What changes is that the source is replaceable in one body instead of at every call site, so the wave-2 platform_commission_rates work is a single-file change. It is async and takes the two parameters a per-merchant lookup needs, deliberately, so that migration has no call-site churn.
⚠ It is called outside the try that saves the order. Today it only reads an environment variable and cannot fail transiently. When it becomes an RPC that distinction is load-bearing: a transient lookup failure must not be swallowed as payment_link_failed and quietly cost the platform its commission, while a genuinely missing rate term must be a hard 500. Two error categories, and collapsing them is how a commission silently becomes zero.
The snapshot half is unchanged: payments.commission_rate_bp is still written per row, so a later rate change cannot rewrite the history of money that already moved.
⚠ Deploy after CR-40 or the insert references columns that do not exist.
CR-26.0.1-42 · razorpay-webhook correlation rewrite, monotonic refunds, provider-fact capture
supabase/functions/razorpay-webhook · ef_code · reversible
Closes QRS-710 and the second half of QRS-706 in one edit.
Correlation becomes a fallback chain over provider_payment_id, provider_link_id, provider_order_id and, only when it is genuinely a uuid, order_id. It replaces an if / else if carrying three independent defects: a provider string bound into a uuid column (22P02, thrown, 5xx), .maybeSingle() on a query the money model deliberately allows to return several rows, and an exclusive chain that abandoned an event one key could not match even when another would have. Where several rows match, the choice is deterministic and the first rule is load-bearing: payments has a unique (provider, provider_payment_id), so attaching a capture to the wrong attempt row burns that pay_id and turns the correct row's own event into a permanent 23505 loop.
Refunds derive their status from the amount, never the event name. refund.processed mapped to refunded unconditionally, which is terminal rank 5, so a second partial refund failed the monotonic guard and its amount was dropped: the money left the account and the ledger did not move. ⚠ Removing that status is itself a money-losing regression on its own, because with status null the guard is skipped entirely and Razorpay gives no ordering guarantee. Monotonicity is restored on the amount axis with Math.max, which is the axis that carries the money. Both halves were required.
processed_at is now set only on success, so a dead letter stays visible to payment_events_unprocessed_idx instead of looking handled while carrying an error nobody would query for. payment_events.payment_id is written for the first time — it was written by nothing, 3 rows and 0 linked, so the append-only provider ledger could not be joined to the money it describes.
⚠ The subtlest change is the stale-transition branch. payment.captured arrives about 1.4s afterpayment_link.paid has already advanced the row, and it is the only event carrying fee and tax. So the event that carries the fee is precisely the one the monotonic guard rejects. That branch now applies provider facts before returning. Adding fee extraction without it would have passed every unit test and captured nothing.
Provider facts fill a NULL and never overwrite a differing present value: changing a recorded fee because a later event reports another number destroys the evidence that the two ever disagreed, and that evidence is the only thing separating a provider correction from a bug weeks later.
CR-26.0.1-43 · /consumer/order-code renders an availability state instead of a fictional QR
apps/mobile/src/app/consumer/order-code.tsx · app_build · reversible
QRS-708. OrderCodeScreen is complete and good, but packages/data/src/collection is stub-only and the stub returns invented data: order GS-1041 from "Shree Ganapati Arts" with codeSignature: 'sig-1041-server-issued'. A buyer reaching the route was shown a QR encoding another, fictional person's order reference and a fabricated signature, with the full confidence of a working feature.
⚠ Nothing could have rejected it. The merchant scanner (QRS-645) is not built, so there is no scanner to refuse a bad code and no moment at which the fiction surfaces. The failure is silent and lands at a counter, in a queue, with money involved.
Gated at the route, not by deleting the screen, which is finished wave-2 work. Not a redirect: a deep link, a notification tap or a back-navigation lands here, and bouncing the buyer to Home tells them nothing and looks broken. Not an upgrade prompt: this is the availability axis, not entitlement, and nothing is purchasable here, so a CTA would be a second falsehood on top of the first. CLAUDE.md's fifth rule says an unavailable capability is presented as not ready yet, with the real next step where one exists — and one does: quote the order reference at the stall, which is the paper-slip workflow this feature replaces.
New i18n leaves orderCode.notReady / notReadyBody in all three locales, em-dash-free per the gated copy rule.
CR-26.0.1-44 · Reconcile seven drifted migration version stamps on Dev
supabase_migrations.schema_migrations (qr-setu-dev) · data_migration · reversible
QRS-705, extending QRS-267. The MCP apply_migration tool stamps its own timestamp rather than the migration filename's, so seven migrations were registered under versions the repo does not contain — including catalogue_project_is_unique, in the repo as 20260814090000 and registered as 20260815172258.
The DDL ran correctly every time. Only the registered version was wrong, which is exactly what makes it invisible: the schema is right, the app works, and nothing looks broken until the next supabase db push, which would try to re-run all seven files and fail on "already exists". QRS-267 documented this hazard for a single migration and prescribed calling list_migrations after the first apply_migration. That check was not carried forward, and the drift accumulated silently over three days while every gate stayed green.
⚠ This is the sixth rule's promotion-checklist line 2 in its exact form — "every migration in the repo is applied on the target, and every migration applied on the target exists in the repo, including the reverse direction" — and it had never once been executed as a command.
Corrected in place; relative order is preserved by every rename, so the sequence is unchanged. Then verified in both directions: 43 repo files against 43 registered versions, with both comm sets empty.
The durable fix is not this correction, it is a gate. Nothing in the repo compares a live environment's migration set against the repo's, so until one exists, every claim that "Dev matches the repo" is an inference. Folded into QRS-696.
CR-26.0.1-45 · 20260817100000_v2_payment_reconciliation.sql
supabase/migrations/20260817100000_v2_payment_reconciliation.sql · schema_migration · reversible
Two tables plus two functions, all new and empty; no existing object altered. runs is the heartbeat an OUT-OF-BAND watchdog reads, exceptions is the human work queue - a design with only exceptions cannot distinguish nothing is wrong from nothing has run since Thursday. get_payment_reconciliation_health is the first ever READER of payment_events' two dead-letter indexes, which had been correct and unqueried for days; on its first call it reported 3 dead letters nobody knew about. RLS on with no policy, matching payment_events: platform operations data, service role only. EXECUTE revoked from anon AND authenticated on both functions, verified false.
CR-26.0.1-46 · 20260817110000_v2_merchant_payment_alerts.sql
supabase/migrations/20260817110000_v2_merchant_payment_alerts.sql · schema_migration · reversible
get_merchant_payment_alerts, workspace-membership scoped, for the in-app payment alert. NOT YET APPLIED to Dev - flagged by the migration parity check as local-with-no-remote, which is that gate working on its first real opportunity.
CR-26.0.1-47 · 20260817120000_v2_payments_provider_method.sql
supabase/migrations/20260817120000_v2_payments_provider_method.sql · schema_migration · reversible
payments.provider_method, nullable, backfilled from payment_events payloads. Recorded for financial traceability only: the gateway fee varies by instrument, so without it margin is a blended average and a mix shift is indistinguishable from a pricing change. No CHECK on the values, deliberately - the provider owns that vocabulary and a CHECK would turn a new payment method into a failed webhook write on the money path.
CR-26.0.1-48 · 20260817130000_v2_provider_charge_policy_comments.sql
supabase/migrations/20260817130000_v2_provider_charge_policy_comments.sql · schema_migration · reversible
Comments only; no object altered. Restates the provider-charge columns under the owner rule of 2026-08-17: Razorpay charges are provider-controlled commercial terms on a renegotiable plan, so no fee may be hardcoded and no payment logic may depend on one - capture actuals for reconciliation instead. A new migration rather than an edit to 20260816190000, which is applied: never edit a merged migration. Also removes specific paise figures from the schema comments, because a number in a COMMENT ON acquires authority it has not earned.
CR-26.0.1-49 · reconcile-payments
supabase/functions/reconcile-payments · ef_code · reversible
The reader the dead-letter indexes never had. DETECT-ONLY: it never writes payments, orders, order_items or payment_events - only its own run row and exceptions via the RPC. Scope was cut from replay-and-repair after three adversarial reviews found 25 defects in that build, two of which could book money that never arrived (a critical conflict that did not gate the write, marking a FAILED payment captured; and an order-cache check comparing status strings only, so a doubled capture passed clean). Both required a write, so removing the write capability makes them unreachable rather than fixed. Proven live: 401 without the secret, 12 candidates examined, 6 exceptions opened, and a second sweep bumped all 6 to seen_count 2 rather than duplicating.
CR-26.0.1-50 · config.toml
supabase/config.toml · ef_config · reversible
Declares reconcile-payments verify_jwt = true. Also corrects a heading that read THERE ARE NO verify_jwt = false FUNCTIONS ANY MORE while razorpay-webhook was declared false forty lines above it - true when written for QRS-640, false from the moment the webhook landed. A stale claim about the JWT posture, in the file that IS the JWT posture, which check:fn-config could not catch because it verifies an entry exists and never that the prose is still true.
CR-26.0.1-51 · the two anon doors the buyer journey needs
supabase/migrations/20260817160000_v2_public_buyer_read_path.sql · schema_migration · reversible
The backend money loop was deployed and proven on live money while the buyer had no way to reach it: place-public-order had zero client callers, and both questions a public card must answer — one before a purchase, one after — were unanswerable by an anonymous caller.
resolve_setu_card_payment_readiness(text) — a slug-keyed wrapper that delegates to resolve_workspace_payment_readiness. That function was already security definer, already granted to anon, and its own header already named the public card as an intended caller — but it is keyed on workspace_id, and no anon-callable function in this schema returns one, because get_public_setu_card deliberately exposes no uuid at all ("NEVER resolve features or entitlements here"). A correct, granted, intended-for-this-caller answer was unreachable by the only caller that wanted it, for want of a slug-shaped door. Delegating rather than reimplementing keeps the three conditions and the order the unmet one is reported in inside one body: a second copy would drift into a card rendering a Pay Now button the write path then refuses, after the buyer has already decided to pay. It also keeps the workspace-scope-only availability rule in one place — payments availability may only ever be granted at workspace scope (precedence 60 over the platform deny at 10), because a plan-scope allow (40) would override the deny for an entire tier of merchants with no bank account. Resolves only status = 'published', so a draft card returns the same unknown_workspace a nonexistent slug does and unlaunched slugs cannot be enumerated.
get_public_order_status(uuid, text) — one order for the anonymous buyer who placed it, for the page they land on after a Razorpay link. Both the order uuid and the printed reference are required, because neither is solely ours: the uuid is handed to Razorpay as the payment link's reference_id (so it lives in a third party's systems, dashboard and emails), and the reference is 8 Crockford characters that a buyer reads aloud at a stall counter. Together they mean "obtained both, for the same order". ⚠ It seeks by primary key and filters on reference, never the reverse — no index has reference as its leading column, so a reference-first probe would sequentially scan every order on the platform on a hot post-payment page.
⚠ The projection is the security boundary, and the shortcut it avoids is a mass-PII leak.get_my_order_book must never be granted to anon: it projects buyerName, buyerPhone and buyerNote for every order in the workspace with where o.workspace_id = p_workspace_id as its only filter, so one workspace uuid would dump every customer's name and phone number. Its own comment says so. The sanctioned mechanism — named in 20260810110000's policy comment for the analogous consumer case — is a narrowed RPC projection, never a widened grant and never a second table. Deliberately absent: commission_rate_bp / commission_minor / vendor_minor (the merchant's terms with Digious), provider_fee_minor / provider_fee_tax_minor (what Razorpay charges Digious — not disclosed even to the merchant, because the gap between the commission and the gateway fee is the platform margin), buyer PII (the caller supplied it; echoing it back turns a leaked secret pair into a disclosure), and every internal id that would enable a second query.
⚠ payment_status is read, never re-derived. It is already a ledger-derived cache maintained through payment_counts_as_collected(), whose sole gateway writer is razorpay-webhook. A second definition of "the money arrived" is QRS-706 exactly — one rule with four spellings, where the raw status = 'captured' ones discarded the positive leg of a refunded payment and wrote unpaid onto orders that had genuinely been paid.
⚠ Deliberately does not require the card to still be published. The publication gate belongs on the write path. Retroactively hiding a real customer's paid order because the merchant unpublished their card is worse than the disclosure it would prevent — the buyer already holds both secrets for an order they placed.
No existing object is altered and no grant is widened. Both functions are new; resolve_workspace_payment_readiness is untouched.
CR-26.0.1-52 · state and city become reference data
supabase/migrations/20260817170000_v2_location_reference_data.sql · schema_migration · reversible
Standing owner requirement, 2026-08-17, non-negotiable: state and city are never free-text fields. Adds public.states (36 rows — 28 states + 8 Union Territories, each carrying its 2-digit GST code) and public.cities (178 rows across all 36), plus get_states() and get_cities(p_state_key) granted to anon and authenticated — because onboarding collects a state before any workspace exists, and a consumer picking the areas they browse has no workspace at all. The state → city filter the owner requires is the state_key FK, enforced server-side rather than by a client filtering a full list; an unknown state returns an empty array rather than every city, so a client that forgets to pass one shows nothing instead of offering Chennai to a Punjab merchant.
Why this is structural and not cosmetic. Free text yields "Pune" / "pune" / "PUNE" / "Pune City" as four cities, and the damage lands in three places: the city becomes a marketplace URL path segment, so four spellings are four indexed URLs of the same content — index dilution on the primary SEO surface — and one typo mints a city page holding a single vendor; facet counts are computed over it, so "12 stalls in Pune" is silently wrong the moment two spellings exist; and the state joins platform_tax_identity.state_code, the GST code already load-bearing on the money path.
⚠ And this was the cheapest moment it will ever be, which is why it happened now. Measured first: state and city are declared z.string().optional() in packages/schemas/src/onboarding.ts and no onboarding step collects either, so the three free-text column pairs are essentially empty. Once vendors have typed cities and the marketplace has indexed the resulting URLs, the same change is a mapping table plus 301s plus a re-crawl.
Identifier decisions, each awkward to reverse. key is the URL segment — one field, not a key plus a slug, because a second column carrying the same fact drifts on the first rename (QRS-249 class). City keys are globally unique because the URL forces it: the taxonomy is /marketplace/<category>/<city>, a bare segment with no state above it, and Indian city names repeat across states — so the second occurrence carries a state suffix (bilaspur / bilaspur-hp, hamirpur-up / hamirpur-hp), a convention demonstrated in the seed rather than only described. gst_code is seeded and iso_code deliberately is not: the ISO 3166-2:IN assignments for at least Odisha (OD/OR) and Uttarakhand (UT/UK) have historical variants, and writing an uncertain value into a table that becomes authoritative is worse than leaving it null. Retired GST codes 25 (Daman and Diu, merged into 26 in 2020) and 28 (pre-bifurcation Andhra Pradesh, now 37) are absent, so no merchant can select a state that no longer exists for GST purposes.
Expand only. Nullable state_key/city_key FK columns are added beside the existing free-text pairs on workspaces, setu_cards and locations. Nothing is dropped, no reader changes, and ADD COLUMN of a nullable column with no default is catalog-only on PG11+ — so this cannot break a running client. The contract step (backfill, switch every reader, drop the text columns) is a later migration that must declare requires_min_app_build, because an old build still sending prose would silently record no location at all.
RLS enabled with no policy and every client grant revoked on both tables: reference data is read through the RPCs and is never writable by a merchant or a consumer, so adding a city is a migration — which is what keeps the URL vocabulary reviewable.
The seed was validated by parsing the migration, not by reading it: 36 states (28 + 8), 36 unique GST codes with 25 and 28 correctly absent, 178 cities with no duplicate keys, no orphan state_key, every key matching its table's own CHECK, and Maharashtra = 27 cross-checked against the already-seeded platform_tax_identity.
Contract side: profileSetupSchema gains state_key/city_key (locationKeySchema), with the free-text state/city marked @deprecated rather than removed — the expand half on the client too.
CR-26.0.1-53 · an event nobody handled must stay replayable
supabase/migrations/20260817180000_v2_webhook_handler_provenance.sql · schema_migration · reversible
A live data-loss path, measured on Dev, not a tidiness issue. razorpay-webhook records every event before deciding whether it can act on it — correct — but the not-handled branch then calls finish(null), which stamps processed_at exactly like a full apply. So an event no handler touched fell outside both replay selectors:
payment_events_unprocessed_idx WHERE processed_at IS NULL→ does not see itpayment_events_failed_idx WHERE process_error IS NOT NULL→ does not see it
Measured consequence: three real transfer.processed rows reported as 3/3 processed with no transfer branch in the function at all. "Recorded and ignored" was indistinguishable from "handled" for a query, the reconciler, the watchdog and a support reader.
⚠ Urgent rather than merely wrong. The owner has just subscribed the four product.route.* events plus the dispute set, transfer.failed, payment_link.expired and payment.authorized — none of which has a handler yet. Every one arriving before its handler ships would have been stamped processed and become permanently unrecoverable, including the event that says a merchant's Route account went live, which is the gate for showing their Pay control. A provider retry is no remedy here: redelivery hits unique (provider, provider_event_id) and returns duplicate before any processing, so a replay is the only remedy and replayability is the thing to protect.
payment_events.handler records which branch acted; NULL means none existed. A three-way split replaces a two-way one:
| Condition | Meaning |
|---|---|
processed_at IS NULL | genuinely stuck mid-flight, a real alarm |
process_error IS NOT NULL | failed, needs a replay, a real alarm |
handler IS NULL AND processed_at IS NOT NULL | recorded, no handler existed: the replay pool |
⚠ Deliberately does NOT skip the processed_at stamp. That would fill payment_events_unprocessed_idx with every delivered-but-unimplemented event forever, so the watchdog's stuck_unprocessed > 0 assertion would fire permanently. An alarm that is always on is an alarm that is off. The new partial index payment_events_unhandled_idx is the replay pool, disjoint from the other two sets.
⚠ The backfill is deliberately narrow — only event types that genuinely had a branch get legacy. A boolean not null default true would have asserted that those three transfer.processed rows were handled, inventing the exact fact this change exists to record.
Second change: activation_status gains rejected and activated_kyc_pending, both unrepresentable while their events were already subscribed, so a rejection could not be stored at all. ⚠ activated_kyc_pending is recorded but not permitted to take money: whether a kyc-pending linked account can receive Route transfers is unconfirmed, and both silent choices are wrong. Collapsing it into activated risks holds and reversals on money that has already moved; collapsing it into under_review misreports a state the provider explicitly distinguishes. Recording it truthfully with the gate shut is the only option that is not a guess. resolve_workspace_payment_readiness tests activation_status = 'activated' exactly, so every value added here is fail-closed by construction and no gate needed editing. Permitting it later is a one-predicate change, never a data migration. This is CLAUDE.md's fourth rule applied to a contract rather than a screen: never silently narrow a provider's enum, because a state that cannot be stored can never be noticed as missing.
Read back on Dev: 3 transfer.processed rows now in the replay pool · 18 genuinely-handled rows marked legacy · 10 errored rows correctly excluded because processed_at IS NULL · all 7 activation values admitted · index present. razorpay-webhook redeployed so handler is written going forward.
CR-26.0.1-54 · product.route.* handling, the gate behind "no online orders until validated"
supabase/functions/razorpay-webhook · ef_code · reversible
Handles the four product.route.* events, which drive payout_accounts.activation_status — the field resolve_workspace_payment_readiness tests, and therefore the gate behind the owner's rule that no online order is offered while a merchant's Route enablement is incomplete. Cash and self-managed UPI are untouched, because neither goes through Route.
⚠ These events do not correlate to a payment. They resolve to a payout_accounts row keyed on the linked-account id, so the branch sits before the payment correlation. Falling through to correlatePayment would dead-letter every one as "no matching payment row", which reads as a payment defect and is not one. Account (account.*) and product (product.route.*) are different objects with different lifecycles, and payout_accounts mirrors that with two columns — driving the gate from account.* would gate money on the wrong signal.
⚠ The state comes from the event name, not the payload. The event name is the one part of the delivery whose shape cannot drift. A payload activation_status that disagrees is logged as a provider-contract finding, never used as an override.
⚠ The guard is RECENCY, not RANK, and conflating it with the payment path would be a real defect. A payment only moves forward, so canTransition refuses a backwards step. A Route product legitimately moves backwards: activated → needs_clarification when Razorpay later wants another document, or → suspended. A monotonic rank would pin an account at its high-water mark and permanently hide a merchant's outstanding requirement. So it applies only when the provider's own event timestamp is at least as new as the stored last_verified_at; a missing timestamp applies rather than discarding a real transition over an optional field.
⚠ The payload shape is not yet observed in this project — payment_events on Dev holds zeroproduct.route.* rows — so the extractor is written from documentation, tries four plausible paths for the account id, and dead-letters when none resolves rather than guessing. A dead letter keeps the full payload, lands in payment_events_failed_idx, is visible to the reconciler, and is replayable once the true shape is known from that very delivery. An unknown linked account is also a dead letter, permanently: a retry cannot conjure the row.
7 new Deno tests (29 in the suite), deno check exit 0.
🟡 Not yet deployed. The CLI token flipped to the prod/Nefoxx account mid-session and the deploy returned 403. The consequence is contained by CR-53: Route events arriving before this deploy land in the payment_events replay pool (handler IS NULL) instead of being silently swallowed — which is exactly what that change was built for, now validated in the real gap rather than in theory.
CR-26.0.1-55 · stock commitment, so a one-of-a-kind idol is sold exactly once
supabase/migrations/20260817190000_v2_stock_commitment.sql · schema_migration · reversible
The launch blocker of the set, because Ganapati idols are the named use case. Stock was never committed anywhere: place-public-order read stock_quantity to validate and never wrote it, no trigger decremented it, and is_unique's quantity clamp was per request rather than per item lifetime. So two concurrent buyers both passed the check, and the same single murti could be ordered by an unlimited number of buyers, each told it was theirs. At a stall that is one idol and several people arriving to collect it.
Two mechanisms, because is_unique and track_inventory are independent columns and both default false. A one-of-a-kind idol whose merchant never configured inventory has no stock to decrement, so arithmetic alone cannot protect the case that matters most.
| Mechanism | |
|---|---|
| A. Exclusivity | order_items.holds_unique_claim, denormalised from the item by trigger (an index predicate may only reference its own table), plus a partial unique index on (item_id) where holds_unique_claim. ⚠ That is a constraint, not a check-then-act — a trigger that SELECTed "already claimed?" then inserted is a race two concurrent buyers both win. The second insert now gets 23505. |
| B. Arithmetic | An atomic UPDATE … WHERE stock_quantity >= quantity, which serialises two buyers on the row lock: the second re-evaluates against the decremented value, matches nothing, and raises. |
⚠ Commits at placement, not at payment — a deliberate divergence from the design. The design says "stock is committed on a successful payment", which is right for an online-pay order and wrong for the launch cohort: the owner's rule lets a merchant without Route enablement still take orders paid in cash, and a cash order has no payment event, so committing on payment would never commit it and two buyers would again be promised one idol. It lives in the database rather than in place-public-order so manage-order's counter sales get it for free and a future third writer cannot forget it.
Release on cancel restores tracked stock and clears the claim without deleting the line, so the record of what was ordered and then cancelled survives. ⚠ The trigger's WHEN (old.status <> 'cancelled') is correctness, not optimisation: a second firing would credit stock the order never took and inflate inventory out of thin air.
⚠ Safe to raise, verified before writing: both write paths already delete the order when the line insert fails (place-public-order:272, manage-order:255), so a raise is a clean failure rather than an order with no lines.
Tested: new pgTAP suite stock_commitment_test.sql — 28 assertions across arithmetic, exclusivity, release and the no-ops (untracked item, null item_id). Read back on Dev: 53 migrations, column present, unique index present, both triggers live, and 16 pre-existing order lines correctly defaulted false so no historical order holds a claim.
⚠ Deliberately not done: expiring an abandoned online order. A buyer who opens the pay sheet and walks away holds the idol until someone cancels. Releasing it automatically needs a scheduler, and pg_cron is not installable here (postgres is not superuser), so it is a GitHub Actions sweep like payments-watchdog.yml. Tracked, not silently omitted — holding stock for a real buyer is the safe direction to fail.
CR-26.0.1-56 · the return leg, and a design that survives being wrong about it
supabase/functions/_shared/razorpay.ts + place-public-order + apps/web · ef_code · reversible
Razorpay was sent no callback_url, so it redirected the buyer nowhere after they paid. The confirmation page was not partly built, it was unreachable. createPaymentLink now sends callback_url + callback_method: 'get', place-public-order builds the URL from PUBLIC_WEB_BASE_URL, and apps/web gains the route it lands on.
The whole shape of this change is surviving being wrong about two field names
Razorpay refuses an unknown key by rejecting the entire request ("extra fields sent"). So one wrong name would not degrade the return leg — it would break every payment link creation on the money path. That is exactly what partial_payment vs accept_partial did on 2026-08-16, invisibly, because the tests assert the body we send rather than the body Razorpay accepts.
So a field rejection retries without the two fields: the buyer still gets a payable link, and the failure is logged at ERROR so a wrong name cannot hide. ⚠ Only an unambiguous 4xx retries. A 5xx or a timeout may have created the link and lost the response — retrying that would mint a second payment link against one reference_id, and a buyer able to pay twice.
The URL is a bearer token, so the headers are the security
Our identifiers go in the path, because Razorpay appends its own query parameters. That means the URL carries both secrets, so the route serves Cache-Control: no-store and X-Robots-Tag: noindex. ⚠ /:slug is cached for seven days — inheriting that would hand one buyer's confirmation to the next visitor. Both headers were verified as actually sent against the real built server, because the QRS-569 trap (a loader's headers never reaching a document response without a headers export) would here be a disclosure rather than a caching inefficiency.
The query string is never trusted
Displayed state comes entirely from get_public_order_status, so no signature verification is needed and a forged parameter cannot make an unpaid order read as paid. Proven: a forged razorpay_payment_link_status=paid against an unpaid ledger yields zero occurrences of "Payment received".
The one use of that parameter is choosing reassuring copy for the paid-but-not-yet-reflected window. The webhook is a separate network hop and a browser is often faster; telling someone who just paid that they have not is the worst thing this page could say. An untrusted value picking a gentler sentence is safe in a way that using it as the source of truth would not be.
New seam
packages/data/src/publicOrder — factory only, no bound singleton, because SSR loaders have no long-lived process for one to be safe in. ⚠ A thrown error and a miss stay distinct: null renders not-found, an error propagates. Swallowing a failed read as null would tell a buyer who genuinely paid that their order does not exist.
Verified: 5 new Deno tests (236 in the EF suite) · route driven through the real SSR server — renders, noindex meta, no-store + x-robots-tag + vary all sent · a wrong reference and a wrong uuid both return byte-identical 404s (4817 bytes), so the page is not an enumeration oracle · type-check 0 errors · 919 mobile tests · 390 pgTAP.
⚠ Dependency, stated plainly: PUBLIC_WEB_BASE_URL has no real value yet, because apps/web is not deployed (Cloudflare not provisioned, QRS-306). So the graceful degradation is the live path today — no base URL means no callback_url, the order still lands, the link is still payable, and an ERROR log records the missing return leg. The return page cannot function until the web app has a public origin.
CR-26.0.1-57 · three ledgers for money that moves after capture
supabase/migrations/20260817200000_v2_money_movement_ledgers.sql + razorpay-webhook · schema_migration · reversible
transfer.processed × 3 were sitting in payment_events with handler IS NULL — recorded, acknowledged, and acted on by nothing, because there was nowhere to put them. The owner has since subscribed transfer.failed, settlement.processed and the four payment.dispute.* events too. All of it describes money moving after the buyer's capture, which is the half of the money story this schema could not express.
Three questions the platform could not answer, and now can:
| Did the merchant actually get paid? | payments.captured says the buyer paid us. It says nothing about whether the Route split reached the vendor's linked account. A transfer.failed with nowhere to land meant we would believe a merchant was paid when they were not, and discover it only when they asked. |
| What hit Digious's bank, and which UTR? | The finance reconciliation anchor. Without it, matching Razorpay payouts to our own bank statement is manual forever. |
| Is the money being clawed back? | A dispute on a Route payment is money that may already have been transferred to the merchant, so a lost one is a real commercial exposure. |
Three tables rather than columns on payments, because each has its own lifecycle that outlives the payment's: a transfer moves created → pending → processed → failed → reversed independently of the capture; a settlement covers many payments and belongs to Digious rather than any workspace; a dispute resolves weeks later. Folding any of them in would mean a row whose status means three different things depending on which column you read.
⚠ Named platform_settlements, not settlements. A merchant plausibly wants that name too — Razorpay settles to their linked account directly, and they will want to see it. That is CLAUDE.md's "will a second thing plausibly want this name?" test, the same one that produced setu_cards, platform_plans and process_primitives.
⚠ payments.provider_transfer_id stays as a convenience POINTER; payout_transfers is authoritative. The handler fills the pointer only when it is NULL, so a disagreement between the two is preserved as evidence rather than overwritten — two places holding a transfer id is the duplicate-source-of-truth shape of QRS-249/284/287, and the resolution is that one is a pointer.
Service-role only on all three — RLS on, zero policies, every client grant revoked, verified on Dev. A merchant read ("was I paid?") is a legitimate future RPC with a narrowed projection; a widened grant is not that RPC, and shipping them readable "for now" is how a projection nobody designed becomes the contract. The disclosure boundary extends unchanged: a transfer's provider_fee_minor is our cost and is never shown to the merchant.
Every write is an UPSERT on unique (provider, provider_*_id), so a redelivery updates rather than duplicating and the constraint is the guarantee. workspace_id and payment_id are nullable deliberately: a transfer we cannot correlate is still money moving in our own Razorpay account, and dropping it would hide exactly the discrepancy worth finding.
⚠ None of these payload shapes is observed in this project — payment_events holds three transfer.processed rows and nothing at all from the other two families — so extraction is written from documentation, and a missing provider id dead-letters rather than writing a half row. The event keeps its payload, lands in payment_events_failed_idx, and is replayable the moment the true shape is known.
⚠ A latent bug found by deno lint flagging an unused import. The dispute branch was reached by elimination, so any family later added to isMoneyMovementEvent without its own branch would have been silently written as a dispute — a settlement-adjacent event landing in payment_disputes with a fabricated status. It is now explicit, and an unbranched family falls through to the replay pool instead. The cheapest possible way to have found that.
8 new Deno tests (37 in the webhook suite). Read back on Dev: all three tables present · RLS enabled on all three · zero policies · zero anon/authenticated grants · 12 indexes.
🟡 Function deploy pending — the CLI PAT flipped to the prod/Nefoxx account mid-session (403). The consequence is contained by CR-53: events arriving before the deploy land in the payment_events replay pool with handler IS NULL rather than being swallowed.
CR-26.0.1-58 · four transfer defects, found by reading the first real payload
supabase/migrations/20260817210000_v2_transfer_shape_corrections.sql + razorpay-webhook · schema_migration · reversible
CR-57 shipped payout_transfers one hour earlier with a banner on it saying, in as many words, that no transfer payload had ever been observed in this project and the extraction was written from documentation. Three real transfer.processed payloads had been sitting in payment_events the whole time, unread, because they had handler IS NULL and nothing had ever needed them. Reading one broke four assumptions.
| the assumption | what the payload says | |
|---|---|---|
| 1 | source is a payment id | "source": "order_TQhwuqYioBjVdZ" — an ORDER id, on all three deliveries, never pay_ |
| 2 | error_description is a flat string | "error": { "code": null, "reason": null, "description": null, … } — an object |
| 3 | status says whether the vendor was paid | on_hold exists, and a transfer can be processed AND held |
| 4 | amount_minor is the amount that moved | amount_reversed exists, and a partial reversal stays processed |
⚠ Defect 1 was unfindable by any means other than this. The handler looked source up in payments.provider_payment_id, which holds pay_…, so it could never match — and a null payment_id is a legitimate recorded state in this table by design. Every transfer would have been written, looked correct, and been unjoinable forever. The fix stores the provider's field verbatim in provider_source_id and correlates against both provider_payment_id and provider_order_id: both columns rather than an order_/pay_ prefix branch, because trusting a prefix convention is the same category of assumption that caused the defect in the first place.
⚠ Defect 2 was visible only by TYPE, not by value. All three observed payloads are processed, so every sub-field of error is null. failure_reason would have been NULL on every single transfer.failed — losing the one field that row exists to carry.
Also in this change: processed_at now comes from the provider's own payload timestamp instead of our clock, so a redelivery days later cannot restamp a transfer that moved long ago; the transfer log line escalates to ERROR for on_hold and reversal as well as failure, because all three mean the vendor does not have the money; and payments gains a partial index on (provider, provider_order_id), without which every transfer webhook seq-scans the payments table.
Two cross-checks now proven on live data, not asserted:
- For all three transfers,
amountequals the matchingpayments.vendor_minorexactly — 76000 / 199500 / 380000 against amounts of 80000 / 210000 / 400000. The split we computed is the split that moved, so a future divergence is a real reconciliation exception rather than a rounding artifact. - The handler's two-key correlation was simulated in SQL against the real source ids and resolves all three to a payment, a denormalised
pay_…id and a workspace. The old code returnsUNRESOLVEDon all three.
⚠ Provider fees are TAX-INCLUSIVE, and that is now recorded where someone will find it. A real delivery carried fees: 224 with tax: 34, and 34/(224−34) is 17.9% — so provider_fee_minor already contains provider_fee_tax_minor and adding the two double-counts the tax. Nothing computes on either, per the owner's standing decision; the note exists for whoever reads the columns.
5 new Deno tests built from the real payload as a byte-faithful fixture (43 in the suite), and mutation-proven: 3 of the 5 fail against the old extractor, the 4th fails at type-check because providerSourceId did not exist, and the 5th passes both ways — labelled a regression guard rather than a defect pin, because a test that passes before and after is not evidence (QRS-013).
Why four defects cost nothing. The events were kept whole with their payloads, and the ledger did not exist yet to be wrong. Four things that would each have been a silently-incorrect production row became four corrections to an unused table, for the price of one SELECT. The fail-closed dead-letter posture and CR-53's replay pool are what made that true — and the lesson generalises past this table: a banner admitting "this shape is unobserved" is a task, not a disclaimer. It sat there for an hour with the answer already in the database.
CR-26.0.1-59 · the money alarm was guaranteed red, so it meant nothing
supabase/migrations/20260817220000_v2_health_present_tense.sql + razorpay-webhook · schema_migration · reversible
⚠ payments-watchdog.yml is the only money-safety alarm in this system, and it had been unconditionally red since the first real payment. Measured on Dev:
dead_letter_events = 7 -> A4 fires
stuck_unprocessed = 4 -> A4 fires again, on the SAME seven rowsA4 asserts both are zero. Both were permanently non-zero, so the workflow could never report anything new. This is QRS-013's green no-op in mirror image: an alarm that fires unconditionally carries exactly as little information as a gate that passes unconditionally, and it is worse in one respect, because it teaches the reader to ignore it. CLAUDE.md states that rule for merchant nudges — "a nudge the merchant learns to ignore does not degrade to neutral" — and it applies with more force to an alarm about money.
⚠ The defect is a TENSE error, not a threshold error, and that distinction is the whole fix.payment_events.process_error is a point-in-time record: at the moment that delivery arrived we genuinely could not correlate it, and writing it was correct. Reading it later as a present-tense statement of "this is broken now" is what was wrong, because nothing ever re-evaluated it.
The measured discrimination, which is why this earned a migration:
| 6 of 7 | payment.captured whose payload pay_… is now in payments | the documented arrival-order race — payment.captured carries neither payment_link_id nor reference_id, and its siblings land the same money ~230 ms later. It happens on every online transaction, so the alarm was destined to be red from transaction one, by construction. |
| 1 of 7 | payment.failed for pay_TQhxgHqpnLqeWR, still absent from payments | the one real finding, invisible at 1-in-7 |
The true signal was not missing. It was diluted six to one.
⚠ And one problem was raising two alarms. finish() writes processed_at = null whenever process_error is set, so every dead letter tripped dead_letter_events and stuck_unprocessed simultaneously — which hid the fact that stuck is genuinely zero. Those must name different failures to be worth two counters: a dead letter finished, badly, and needs a decision; stuck means we began and never came back — a runtime failure, and the more urgent of the two.
After: dead_letter_events 1 (real) · stuck_unprocessed 0 (true) · plus two new keys — stale_dead_letters 6 (reported, never alerted, and expected to grow with volume) and unhandled_events 5 (the CR-53 replay pool, a third state meaning an event class is unimplemented rather than unmatched).
⚠ Every existing key name and source expression is preserved verbatim. The ten keys the watchdog reads were grepped out of the workflow rather than recalled — after a first draft of this migration renamed five of them (newest_run_id, newest_status, newest_examined, …) and would have silently broken the alarm while looking like a tidy-up. Two keys added; none renamed, retyped or removed.
The second half, in the webhook: the dead-letter test is now DID MONEY MOVE, not did it correlate.
- An unplaceable
payment.captured,order.paid,payment_link.paidorrefund.processedis a genuine dead letter — money exists and a human must find where it belongs. An unplaceable refund is the worst of them: we have paid someone back against no record of having charged them. - An unplaceable
payment.failedis an ordinary abandoned checkout with nothing to reconcile, ever. It is recorded withhandler = 'payment_lifecycle:unmatched_no_money', counted inunhandled_events, and not alerted.
Without that split the alarm would re-break within a day of launch, because during a festival week with twelve vendors abandoned checkouts are the most common event class there is. ⚠ This suppresses an alert, never evidence: the row keeps its full payload and its handler names the decision.
1 new Deno test pinning the classification of every handled event, including that a newly-added handled event cannot silently default to "no money moved". Both halves deployed on Dev.
CR-26.0.1-60 · the stuck index now matches its only reader
supabase/migrations/20260817230000_v2_stuck_index_matches_reader.sql · schema_migration · reversible
CR-59 corrected stuck_unprocessed to mean we began and never came back, which left the index and its only consumer disagreeing:
index predicate processed_at is null
the only query processed_at is null AND process_error is nullA superset index still serves that query, so this is not a performance defect. It was caught by a pgTAP assertion that deliberately pins the count against the index predicate itself — "the function is the reader those indexes never had" — and that invariant is worth keeping true rather than loosening, because it is the thing that stops the two drifting apart again.
⚠ The two failure indexes now partition the failure space, disjointly and completely:
| index | predicate | what it means | what it needs |
|---|---|---|---|
payment_events_failed_idx | process_error is not null | finished, badly | a decision about money we could not place |
payment_events_unprocessed_idx | processed_at is null and process_error is null | never came back | a look at the runtime — the more urgent of the two |
Before this a dead letter appeared in both, which is exactly how one problem raised two alarms and stuck was never observably zero. The partition is now enforced by the schema rather than remembered by whoever writes the next query. Secondary benefit at volume: dead letters are evidence and are never deleted, so the old index grew without bound while only ever being read for the rows it now excludes.
Read back on Dev: predicate confirmed; health reports dead_letter_events 1 · stuck_unprocessed 0 · stale_dead_letters 6 · unhandled_events 5. pgTAP grew 390 → 396 across 10 suites, all green — including the two obsolete assertions rewritten to state the new rule, and four new ones proving a dead letter stops alerting once its payment exists, with no row edited.
CR-26.0.1-61 · the buyer can finally place an order
apps/web — OrderPanel + /:slug/order + the publicOrder seam's write side · app_build · reversible
place-public-order existed, was tested, and was live on Dev with zero client callers anywhere in the repo. The money path was complete end to end and unreachable — every test order in this project had been placed by a raw fetch() against the Edge Function, which proved the backend and nothing about the journey. This closes the blocker named in the payment status report.
⚠ One constraint shaped the entire feature
/:slug is served s-maxage=604800, so its HTML is shared by every visitor for seven days. An idempotency key baked into it would be identical for all of them, and buyer #2's submit would replay buyer #1's order through idempotency_keys and return someone else's reference and total. That is a cross-buyer disclosure, not merely a bug. Consequences, all of them forced:
- the write lives on a separate, uncached route rather than an action on the card;
- the key is written to the input by a
useRef+ direct DOM write in an effect — notuseState's initialiser (which also runs during SSR), and notsetStatein an effect (which re-renders for a value nothing renders, andreact-hooks/set-state-in-effectis right to flag it). The key is a form value, not UI state, and saying so in code is what makes it correct; - the action mints its own only when the field arrives empty (no-JS), choosing the recoverable failure — a double-tap makes two orders — over the catastrophic one;
- the first test asserts the SSR output is key-free, and it is mutation-proven: moving the mint into
useState's initialiser, the exact "simplification" a future reader would make, fails it.
Native <details> + native <form>
No JavaScript is required to open, fill or submit. That is a decision about this audience, not a preference: festival buyers on cheap Android phones and congested evening networks routinely interact before hydration finishes, and a visit is often only seconds long. The Edge Function call stays server-side, so no order-shaped request and no publishable key ever appears in the page.
What deliberately does not cross the boundary
No price and no commission. The form sends item_id and quantity; every price is derived server-side from catalog_items and the split is computed inside the function. A client-supplied price on a public unauthenticated endpoint is a buyer choosing what to pay — a test asserts the markup carries no price/amount/total field at all.
The control is never removed when payments are unavailable
CLAUDE.md's fifth rule. A merchant whose banking/KYC/Route enablement is pending still takes cash orders — the owner's standing rule — so canTakePayments changes the copy ("Order and pay" vs "Place order") and sends pay=later. It is an affordance only: place-public-order re-resolves readiness itself, so a stale cached page offering pay-now cannot mint a link for an unbanked merchant. The cash copy states what does happen ("Pay at the stall when you collect"), never what is missing. payment: null is success, and both paths land on the same confirmation URL.
⚠ Both e2e gates were silently skipping the whole form
The touch-target walk drops zero-sized elements and axe-core excludes hidden ones — both correct in isolation. But a collapsed <details> makes its entire subtree zero-sized, so the buyer's five inputs were measured by nothing while both specs reported green and axe-core reported zero violations. QRS-013's shape again: not a gate that is wrong, a gate that never looked — and it surfaced only by checking the selector instead of believing a green run.
Both now expand every disclosure first, synchronising on details[open] input:not([type="hidden"]). ⚠ That exclusion is load-bearing: the form's first two inputs are hidden fields and a hidden input is never "visible" to Playwright, so the unqualified selector waited out a full 30s timeout on every project. The right observable condition needs the right observable. Mutation-proven: fields at h-6 (24px) now fail the 44px assertion where they previously passed — and the suite got faster (27s vs 45s), because a real condition beats a fixed wait.
Also in this change: the watchdog had never once authenticated
payments-watchdog.yml read secrets.RECONCILER_WORKER_SHARED_SECRET_DEV/_PROD while the only secret that existed was PAYMENT_RECONCILER_SECRET. A missing GitHub secret resolves to the empty string rather than failing, so WORKER_SECRET was blank and every sweep call got 401. Renamed to PAYMENT_RECONCILER_SECRET_DEV/_PROD, mirroring the Edge Function's own env var name. An ef_secret change is one of the nine classes invisible to every gate in this repo, and the failure read as a reconciler problem rather than a naming one.
Verified
Driven through the real built SSR server (.claude/skills/run-web/driver.mjs, taught the readiness RPC and the order EF). Payment path → 302 to the Razorpay URL; cash path → 302 to the confirmation page with pay: false. The request the form actually built carries a client-minted idempotency_key and zero price fields. 7 new vitest cases (38 total), 21 Playwright, axe-core zero violations with the form expanded, lint and type-check clean.
Parity: apps/web SSR DOM only. ADR-0019 gives the public card exactly one renderer and never an RN one, so the three-surface checklist does not apply. The buyer-side RN twin is the consumer tier's ItemView — separate work, and it must consume this same packages/data seam rather than re-implementing the order call.
CR-26.0.1-62 · the order form took the whole card down on any non-HTTPS origin
apps/web — idempotencyKey.ts · app_build · reversible
⚠ crypto.randomUUID() is restricted to secure contexts. It is defined on https://… and on localhost, and undefined on a plain-http LAN address. CR-61 called it directly inside a mount effect, so on the owner's first real device test the effect threw a TypeError, React unmounted the tree, and the buyer saw "Something went wrong" — over a server response that was a clean 200 throughout.
Measured with a real browser at the LAN IP:
isSecureContext false
crypto.randomUUID undefined
crypto.getRandomValues function⚠ Why no existing layer could catch it — the durable lesson
localhost is a secure context and a LAN IP is not, so the difference is invisible to every test this repo runs: vitest in node, apps/web/e2e on 127.0.0.1, apps/mobile/e2e on 127.0.0.1, and the run-web driver on 127.0.0.1. Only a real device on a real network reaches it.
That is exactly the gap CLAUDE.md already names for native builds ("the native builds are the gate") and had never named for origins. The missing-API case is now simulated in unit tests rather than waited for, which is the only way a secure-context difference becomes testable at all.
The fix does not weaken the randomness
Only randomUUID and crypto.subtle are secure-context-restricted — getRandomValues is not. So the fallback assembles an RFC 4122 v4 from CSPRNG bytes, setting the version and variant nibbles exactly as randomUUID would. Never Math.random(): a weak key is worse than no key here, because the key's whole job is uniqueness across buyers, and a collision replays a stranger's order back to them.
And it never throws. A throwing effect destroys the page, not the feature — so with no safe randomness at all it returns null and the action mints its own key server-side. A double-tap can then create two orders, which is recoverable, where blanking the card is not. A webview that throws on property access rather than returning undefined is covered too.
Audited the same class elsewhere: SetuCardActions guards navigator.share and navigator.clipboard properly (both also secure-context-restricted) and degrades rather than throwing; place-order.tsx runs crypto.randomUUID in Node, where it is always available.
5 new vitest cases (43 total). Verified with a real browser at http://192.168.31.178:8788: 3 forms, 3 distinct v4 keys hydrated, zero page errors, all three priced items rendering.
CR-26.0.1-63 · consumer onboarding was a trap with no exit
supabase/migrations/20260817230500_v2_set_my_primary_context.sql + the client wiring · rpc_function · reversible
⚠ Measured before writing anything, and it changed the plan. The server was already built:
public.users.primary_context text not null default 'business'
check (primary_context in ('business','individual'))
handle_new_user() ALREADY sets it, from raw_user_meta_data
get_my_context() ALREADY projects it
packages/schemas context.ts ALREADY carries itNothing in the client read it except one test fixture. So the gap was never a missing column — it was entirely client-side, and my earlier "four call sites plus a migration" was wrong in both directions.
Two defects compounding
| 1 | resolveEntryRoute had no consumer branch. An individual who finished signing in landed on /dashboard — the merchant console, keyed on a workspace they do not have and can never have. apps/mobile/src/app/consumer/ has held nine routes with zero inbound navigation, reachable only by typing the URL. |
| 2 | hasCompletedOnboarding derives completion from workspace membership, which is zero for a consumer by definition (CLAUDE.md's three categories make membership count the discriminator). So case 1 never fired for them at all: they fell through to the wizard on every launch, where provisionWorkspace threw for want of a brand name they are never asked for. |
The exit condition was unreachable. No amount of work inside the wizard could have freed them.
The one genuinely missing server piece
set_my_primary_context exists because OAuth cannot carry metadata. handle_new_user reads raw_user_meta_data — which now covers email OTP via signInWithOtp's options.data — but signInWithOAuth takes none, because the provider owns it. So a Google sign-up is always created with the default 'business', and Google is currently the only working auth path (email OTP fails at SMTP, QRS-076). Without this RPC the consumer fork would be unrecordable for every real user.
Idempotent (changed: false on a no-op, so the client may call it unconditionally), validates explicitly to 22023 rather than leaving an opaque 23514, audited, anon revoked, authenticated granted, PUBLIC closed — all five verified on Dev, plus a live rolled-back transaction proving the transition, the idempotent second call, the audit row and the rejection.
⚠ It grants nothing. Every merchant-data policy scopes by relationship (workspace_members), never by this column — a routing preference, not a permission. And it is deliberately re-settable by its own subject: this is not archetype immutability. CLAUDE.md calls category "the DEFAULT EXPERIENCE, never a permanent exclusion".
⚠ A defect in my own first draft, hidden by its own safety net
The audit insert named (subject_type, subject_id, metadata) from memory; audit_log actually carries (target_table, target_id, before, after) plus a NOT NULL actor_kind. Every insert would have raised, been swallowed by the surrounding exception handler as a warning, and produced no audit trail at all while the function reported success — a green no-op inside the mechanism meant to make it safe. Columns re-read off information_schema; the live probe now shows {business} → {individual}.
The pure functions moved to @qrsetu/domain
Forced by a real constraint, not tidiness: packages/data deliberately has no Node types — the guardrail that makes a stray fs/Buffer a compile error — so a co-located node:test file cannot type-check there, and adding types: ["node"] would have deleted that guardrail across a package both apps consume. That is the exact trade CLAUDE.md records as the wrong one. domain already carries the sanctioned exception and is already where CLAUDE.md places derivation. packages/data re-exports, so every importer is unaffected and there is still exactly one implementation.
⚠ Three tests encoded the bug as expected behaviour
They passed type: 'individual' purely because it was the shortest walk to the celebrate step, then asserted /dashboard and provisionWorkspace — a merchant outcome on a consumer path. That is precisely how this stayed invisible. Now one test per category. A fourth needed fixing too: the failure-path test mocked provisionWorkspace rejecting while passing individual, which after this change never calls it, so the assertion would have passed for the wrong reason. Per-test mock clearing was also absent, letting call counts leak between tests.
Verified: 5 new node --test cases in domain (466 total), mutation-proven — removing the consumer branch fails 1 of 5. 4 new entryRoute cases (10 total), also mutation-proven. 924 mobile across 128 suites. lint, type-check, check:naming green.
🔴 Still open, named rather than implied: check:rpc reports four get_consumer_* RPCs called from packages/data/src/consumer/service.ts that no live migration defines. The consumer screens are stub-fed and cannot show real data yet. That is the next chunk, not this one.
CR-26.0.1-64 · an item that cannot be bought must not appear on a buying surface
supabase/migrations/20260818090000_v2_unorderable_items_are_absent.sql · schema_migration · reversible
⚠ Reported as a quantity bug. It is not one. Reproducing all three catalogue items through the real server, before changing anything:
Modak box tracked, stock 10 qty 2 -> 302, order total 160000 ok
Shadu idol untracked qty 2 -> 302, order total 420000 ok
Signature murti unique, tracked, STOCK 0 qty 2 -> 502
Signature murti unique, tracked, STOCK 0 qty 1 -> 502 <-- fails at ONE tooQuantity was never the variable. The arithmetic was correct throughout — line total, order total and the payment row all carried exactly 2× unit. The third item cannot be ordered at any quantity.
The root cause: two columns disagreed, and the vocabulary already existed
catalog_items.availability is CHECKed against ('available','out_of_stock','discontinued','coming_soon','draft'), and the murti held availability = 'available' with stock_quantity = 0. Nothing has ever written 'out_of_stock': the CR-55 trigger decremented the count and left the flag alone, so the flag drifts from the arithmetic on the first sale that empties a bin. place-public-order reads the arithmetic and refuses correctly; the public card reads the flag and renders an Order button. Both internally consistent, contradicting each other — QRS-249 with one copy stored and one derived.
Two fixes, for different readers. The trigger now maintains availability (only ever forward to out_of_stock, so a deliberate discontinued/coming_soon survives an unrelated sale), and the public RPC filters on derived truth rather than the flag — which fixes every existing wrong row with no data migration and survives a future writer forgetting the flag.
⚠ The RPC body is the live definition with exactly two predicates changed, taken from pg_get_functiondef. A first draft was written from a grep and would have silently deleted the function's organisation-sharing logic, which the grep never showed. The predicate appears twice — categories and items — and a hand-rewrite would likely have caught one.
Proven live in a rolled-back transaction: last unit sold → stock=0, availability=out_of_stock; gone from the card immediately; a deliberate coming_soon not overwritten.
⚠ coming_soon stays visible — an open question (QRS-733), not a decision.
CR-26.0.1-65 · the error the buyer saw was the opposite of the truth
apps/web + packages/data + place-public-order · app_build · reversible
Defect 1, and the worst. A clean 422 was rendered as a 502 "we could not confirm your order, check with the stall before trying again". On a money path that is the most expensive mistranslation available: it implies we may have taken an order we definitively refused, and it buries the function's own accurate message. CLAUDE.md names the rule — read an error's CATEGORY before acting on it.
A typed PlacePublicOrderError now carries status, serverMessage and maybeCreated. A 4xx refusal (excluding 408/429, which may have been applied after a write) returns 409 with the server's own words and invites an immediate retry; 5xx and network failures keep the careful unconfirmable copy.
⚠ Defect 2 made defect 1 invisible, and was the real silencer. The route's loader threw a 405, on the reasoning that a GET there has no meaning. Individually defensible, a defect in combination: when the action returns data(…, { status }) rather than a redirect, React Router renders the route to show it — and rendering a route runs its loader. The loader threw, the error boundary won, and the buyer got React Router's default "An unexpected error occurred."
So the out-of-stock text and every field-validation message were written and unreachable. The route worked only on its happy path, which redirects and never renders. Both now verified rendering on the real server — 409 with the stock message, 400 with "Please enter your name so the stall knows who is collecting." A one-assertion unit test pins that the loader must not throw.
Defect 3, the one that was reported: no post-payment redirect. buildReturnUrl read only PUBLIC_WEB_BASE_URL — a Supabase secret not set on Dev — so it returned null, no callback_url was sent, and the buyer stayed on Razorpay's page. It now falls back to the request's own Origin, which is the address the card was actually served from and is correct on a LAN IP, on localhost and in production with nothing to configure. The explicit value still wins when present, and production should set it: an Origin is caller-supplied, and although the exposure is minimal (the callback carries only the order id and printed reference, both already in the response to that same caller), minimal is not none.
The categorisation was extracted to a helper after sonarjs/cognitive-complexity flagged the method at 18 against 15 — extracted rather than suppressed, because the rule was right that three nested try/catch levels plus the status arithmetic made the money branch the hardest thing in the file to read. It is duck-typed rather than instanceof Response, because packages/data declares no DOM types by design and naming that global would have meant widening shared config to silence a symptom.
🟡 Defects 1 and 2 are client-side and live. Defect 3 needs place-public-order deployed and the CLI in this shell cannot see qr-setu-dev (403 on every call, with no SUPABASE_ACCESS_TOKEN set — a credential-store difference between Git Bash and the owner's cmd.exe).
CR-26.0.1-66 · the Origin fallback could never have fired, and coming_soon now says pre-order
apps/web + packages/data + place-public-order · app_build · reversible
⚠ This corrects my own fix from one commit earlier. CR-65 made buildReturnUrl fall back to the request's Origin when PUBLIC_WEB_BASE_URL is unset. That fallback cannot fire on the path that matters: Origin and Referer are browser-set headers, and the SSR action calls the Edge Function server-to-server, where Node's fetch sets neither. The function would have read nothing and the buyer would still have been stranded on Razorpay's success page — the exact symptom reported, with a fix in place that looked like it addressed it.
Caught by reasoning about who sets Origin before telling the owner to re-test, rather than after.
The correct shape: the SSR action is the only party that knows its own public origin, because it is serving the buyer's own request. It now passes new URL(request.url).origin explicitly as return_base_url. The function prefers PUBLIC_WEB_BASE_URL, then the body field, then a header — the header path retained as a courtesy for a future direct browser caller rather than deleted. Production should still set the env var: a body field is caller-supplied, and while the exposure is minimal (the callback carries only the order id and printed reference, both already in the response to that same caller), minimal is not none.
coming_soon says Pre-order (QRS-733, owner decision: keep it)
Keeping them orderable is right — "not ready yet, order now, collect later" is a real festival flow and collect_on already models it. Measured live across all three shapes:
| shape | on the card |
|---|---|
coming_soon + untracked | visible — a real pre-order |
coming_soon + tracked, stock 5 | visible |
coming_soon + tracked, stock 0 | hidden — the stock trigger would refuse it |
So the decision needed no filter change: it is already orderable wherever it can be fulfilled, and hidden only where the arithmetic makes an order impossible — the same rule as everything else.
⚠ But the control said "Order and pay", identical to an in-stock item, so a buyer would believe they were buying something ready to collect. That misrepresents stored state — the same class as the defect CR-64 fixed one layer down. Now labelled from the availability field the RPC already projects, with copy stating what does happen ("Not ready yet. The stall will confirm your collection date.") rather than what is missing. Labels extracted to named constants after sonarjs/no-nested-conditional correctly refused three interacting booleans inline.
Verified: 46 web tests (2 new — the pre-order label, and that an in-stock item is unaffected, so the exception cannot become the default), lint, type-check, deno check. The built server bundle carries return_base_url. The live card still shows exactly two "Order and pay" controls and no pre-order wording, so nothing leaked.
🟡 The client half is live. The Edge Function needs one more deploy — the owner deployed the CR-65 version and this supersedes it.
CR-26.0.1-67 — the consumer discovery read path, and a card defect found while building it
Class: rpc_function · Depends on: CR-26.0.1-64 · Reversible · QRS-731
packages/data/src/consumer/service.ts had declared four RPC names since the consumer tier was built, and its own header said so plainly: "⚠ NO BACKEND. None of these RPCs exists." That was honest while nothing could reach those screens. resolveEntryRoute gained its /consumer branch on 2026-08-17 — so for one day a real buyer signing in browsed invented vendors and invented idols.
⚠ The defect that mattered more than the feature (QRS-740)
Found while writing the item feed, on the public card, not the consumer surface:
order_items_unique_claim_idx = UNIQUE (item_id) WHERE holds_unique_claim
get_public_catalogue mentions holds_unique_claim → NOSo a one-of-a-kind murti already claimed by one buyer still rendered with an Order control, and the second buyer chose it, submitted, and hit a unique-index violation — on the most valuable item in a festival catalogue, after deciding to pay. CR-64 fixed the counted case; this is the claimed one, which CR-55's own trigger comment already distinguished ("a one-of-a-kind piece is CLAIMED, not counted"): a unique item need not track inventory, so nothing about the row changes when it sells.
Mutation-proven: with the claim predicate stubbed to false, the claimed murti reappears on the card.
One predicate, not four
catalog_item_orderable() extracts the buyability rule CR-64 had inlined twice; the two new feed bodies would have made four copies of the rule deciding whether a buyer sees something they can pay for. It carries no SET clause deliberately, so it still inlines and get_public_catalogue's plan on the money path is unchanged. That function's body came from pg_get_functiondef and was edited by asserted mechanical substitution — the generator proves share_catalogue, the on_enquiry withholding, storage_key and the category dedup all survive.
Five things measurement changed before any SQL was written
| Finding | Consequence |
|---|---|
nearest has no geometry — setu_cards has city_key, no coordinates; the buyer sends a city key | exact-city-first; the label should read "Nearby" (QRS-735) |
No facet schema is seeded anywhere — all three archetype schemas are {}, zero industry overrides | facets derived from values items actually carry (QRS-738) |
| the orderable predicate was about to exist four times | extracted (above) |
openStatus already owns open/closed, timezone-correct | not re-derived in SQL — a now()-between would be UTC-evaluated against an IST shop and wrong for most of the working day |
| no public media base URL exists anywhere in this repo | RPCs project storage keys; mediaUrl() fails closed, so every image is null until a bucket exists |
And one correction after pulling the design
The first version derived masonry tile geometry from each image's aspect ratio. The approved design (consumer-data.js → tileGeometry) decides it from feed position and column count, for a better reason than mine — "a vendor cannot make their tile taller by cropping taller, so nobody has to" — and a server cannot know a viewport breakpoint anyway. Wrong in rule and wrong in layer. Removed from the RPC; now packages/domain/src/consumer/tiles.ts, shared by the RN feed and the DOM marketplace so the two cannot lay out differently (QRS-739).
Verified against a local Postgres 17.6.1.155 (same major as prod)
A Ganapati fixture: two published stalls in two cities, one draft stall, ten items covering tracked-in- stock · unique-unclaimed · unique-claimed · on_enquiry · tracked-at-zero · coming_soon · variant-priced · archived · cross-city · draft-card.
- Card shows exactly the right five; feed shows six; claimed, sold-out, archived and draft-card items all absent.
on_enquirycarries no price key at all; the variant item readsfromat the minimum.- Facet counts exclude their own key while honouring the others; two filters intersect; a no-match filter returns
[], not an error. filters_worth_showingis computed on the unfiltered universe — proven by pushing the universe to 18 and narrowing to 1: it staystrue, where the filtered total would have hidden the controls the buyer had just used.- 477/477 domain tests, including 11 new tile-rule cases.
🟡 Not yet applied to Dev — the CLI in this shell is authenticated to the prod account and cannot see qr-setu-dev. Needs supabase db push against dyhjofjjuazhyqcvlrkx.
CR-26.0.1-68 — slug governance, before the first public signup
Class: schema_migration · Depends on: CR-26.0.1-52 · Reversible · QRS-607 · QRS-752
⚠ Now-or-never, and the reason is a trigger pair. cards_slug_not_reserved_trg rejects a reserved slug at claim time, but setu_cards_slug_write_once means a slug is never rewritten — so the day a vendor registers /pune, that URL is unrecoverable short of manual intervention against a live merchant's card. QRS-591 recorded the principle: indexed URLs are the least reversible artifact this platform will ever ship. Twelve vendors are being onboarded now.
Measured before writing anything
| Before | |
|---|---|
| cities reserved | 3 of 178 (mumbai, delhi, bangalore) |
| states reserved | 1 of 36 |
| industry category slugs | 0 of 14 (salon, kirana, dairy, tiffin, boutique, …) |
| phishing / trust words | none (verify, kyc, secure, claim, reward) |
| two-character slugs | 28 of 1,332 |
reserved_until | 0 rows — built, never used |
Three metros reserved and the rest of urban India not is what an ad-hoc list looks like, which is why QRS-607 required the policy be decided in the same change or it recurs.
The policy — three tests, in the migration header so it travels with the data
- Could the platform ever want it as a URL?
- Would one merchant owning it harm other merchants or the platform?
- Would a bad actor owning it harm a user?
Test 3 is not in QRS-607's framing and is the one where a missed word has a victim.
Deliberately not reserved, because over-reservation has a real cost: place names below our own cities list (~8,000 towns — and unnecessary, since the taxonomy is /marketplace/<category>/<city> where the city is a third segment that cannot collide with a root slug; cities are reserved for optionality and anti-squatting, not collision) · compounds (pune-sarees, sharma-sweets stay claimable — the test is CATEGORY vs ENTITY) · bare surnames and bare common given names · single characters, already forbidden by the card regex.
Section 19 is what makes this governance rather than a sweep
After-insert triggers on cities, states and industries reserve the slug in the same transaction. A later migration seeding 400 cities inherits the rule for free. Enforced at the INSERT rather than in a build gate because a gate only runs where the repo runs, while the insert is the authoritative moment — and additive only, because deleting a city must never release a slug that may already be indexed or printed on a QR code.
Founder protection publishes nothing — checked, not assumed
reserved_slugs has no grant to anon or authenticated, RLS is on, and there is no SELECT policy, so the table is not enumerable. The only public surface is is_slug_reserved(text), a boolean about a slug the caller already guessed. Bare given names and the bare surname are excluded on purpose: balaji is both a very common given name and a deity name, and lahade would block every unrelated family of that name.
Verified
753 → 2,572 reserved · cities/states/industry-slug/two-char gaps all 0 · idempotent across repeated applies · the trigger fires for a new city and a new industry (bus_travel → bus-travel, the bus-vertical case QRS-607 raised) · claiming /pune now fails with slug pune is reserved and cannot be claimed. Existing cards are unaffected: the guard is BEFORE INSERT OR UPDATE **OF slug**, so publishing or editing never fires it.
🔴 NOT APPLIED TO DEV. Every figure above is from a local Postgres stack. The CLI in this session is authenticated to the prod account and cannot see qr-setu-dev, so the push is the owner's action.
CR-26.0.1-69 — a marketplace facet can accept more than one value
supabase/migrations/20260820180000_v2_consumer_multi_select_filters.sql · applied to Dev · schema migration · reversible · not contracting
consumer_attributes_match() compared a single scalar per attribute, so the facet rail could only ever apply one value: a buyer who wanted shadu or fibre had to pick one and lose the other.
The failure mode is the part worth recording. A JSON array arrived and matched nothing, so the page came back EMPTY rather than erroring. An empty result reads to a buyer as "this vendor has nothing" — it blames the vendor for a broken filter, and nobody reports it as a bug.
The function now branches on jsonb_typeof:
| value | meaning |
|---|---|
| scalar | exactly its previous meaning, so every existing call site is untouched |
| array | any-of |
| empty array | no constraint |
That last row is the trap. "Match nothing" is the intuitive reading of an empty list and it would blank the page the moment a buyer unchecked the last box in a group. It uses jsonb_each rather than jsonb_each_text, so a non-string value is not coerced before comparison.
Verified by reading pg_proc.prosrc back off Dev after applying, not by assuming the push landed — QRS-693 is why that distinction is written down rather than trusted.
CR-26.0.1-70 — the item feed can be scoped to one category
supabase/migrations/20260820190000_v2_consumer_feed_category_scope.sql · applied to Dev · schema migration · reversible · not contracting
The desktop marketplace browse route is /marketplace/<category>/<city> and the feed had no way to honour the category, so it returned every published item in the city. /marketplace/ganapati-idols/pune and /marketplace/sarees/pune would have rendered the same list under two different headings — worse than a 404, because it looks like it worked.
Two decisions in it:
- The category resolves ONCE into a
uuid[]and that array is reused at all five predicate sites. Five copies of one lookup is the duplicate-source-of-truth class (QRS-249), and here it would also have cost five extra scans per request. DROP+CREATE, notALTER, because adding a parameter changes the signature. The parameter is added last, with a default, so the consumer tier and the RN twin are unaffected and no client has to ship first.
CR-26.0.1-71 — the marketplace URL word stops being an internal key
supabase/migrations/20260821120000_v2_industry_marketplace_slug.sql · applied to Dev · reversible
/marketplace/<category>/<city> derived its category from industries.key, so the launch vertical's indexed URL was /marketplace/festival-stall/pune. An industry key is an internal taxonomy and a URL is a marketing asset, and they were the same string. Nobody searches "festival stall"; a buyer searches "ganapati idols pune", and the path segment is the strongest on-page signal a category page has.
No existing URL changed. Every row backfills to the hyphenated key — byte-for-byte what the URL already was — and only festival_stall carries a hand-written slug. The legacy form still resolves and 301s, because a shared link must not break (QRS-591).
⚠ The status filter here is deliberately looser than get_industries(), and the owner's request for an industry Active/Inactive control is what exposed it:
| question | filter |
|---|---|
may a merchant CHOOSE this? (get_industries) | status = 'public' |
| does this category PAGE exist? (this resolver) | anything not retired |
An industry set to private stops taking new merchants while the ones already on it are live. Had this filtered to public, flipping a status would have 404'd their category page instantly — breaking an indexed URL and making live merchants undiscoverable with no code change to blame. Caught before it was applied.
CR-26.0.1-72 — provisioning writes the location keys
supabase/migrations/20260821130000_v2_provision_location_keys.sql · applied to Dev · reversible
CR-52 added state_key/city_key and nothing wrote them, so every workspace since has null keys. The client now sends keys and the server derives the display names: accepting both would allow a city_key of pune beside a city of "Mumbai", two individually-valid values nothing downstream would flag.
The city is validated against its state — something no foreign key can do, and precisely the mistake a client makes by not clearing the city when the state changes.
Two mechanical traps worth recording:
- The old 8-arg overload is dropped explicitly. Adding defaulted parameters creates a second function; PostgREST resolves by argument name and two candidates fail with "could not choose the best candidate function".
citextis qualified asextensions.citext. A signature type resolves at CREATE time against the session search_path, not the function's ownset search_path = public. The unqualified form failed this exact push — the trap20260813130000documents in its own header, which I read and then walked into anyway.
CR-26.0.1-73 — the Edge Function reads them
supabase/functions/provision-workspace/ · deployed to Dev
The reader the client seam was explicitly waiting for. The keys join the idempotency request hash: omit them and two requests differing only by location look identical to the ledger, so a merchant who corrected their city and retried would be served the first result and never see the fix.
Shape is validated at the boundary, existence in the RPC — a malformed key is a client bug worth a 400 naming the field, while whether a well-formed key exists is a question only the reference tables can answer.
CR-26.0.1-74 — manage-media, the upload path
supabase/functions/manage-media/ · not deployed yet · reversible
Before this, _shared/r2.ts — the presigner — had zero live importers. No client could put a byte into the bucket, so get_public_catalogue's ordered images[] was always empty and every surface rendered its no-photo state permanently.
Four properties are security controls rather than validation, and are worth reading before changing anything here:
| control | why it is not cosmetic |
|---|---|
| content-type allow-list | R2 serves back the type that was signed, on a public origin — image/svg+xml would be executable content on our own media host |
| random uuid in the key | the bucket is public and unauthenticated; a guessable key walks a competitor's photos |
purpose derived from target | otherwise: upload as avatar, attach to an item — QRS-249 inside one request |
| the signed URL is never logged | a presigned URL is a bearer token, and a log sink outlives its expiry |
Collection targets link at issue; single-valued ones at confirm. Position allocation must be atomic against the unique (item_id, position) index — but pointing logo_media_id at a pending row would blank the merchant's current logo the moment they begin an upload that might fail.
Registered in config.toml in this same change, because an absent entry is a silent default and check:fn-config is still broken (QRS-643).
CR-26.0.1-75 — R2 buckets validated, MEDIA_BASE_URL wired
This is one of the nine change classes no gate can observe, so the evidence is a live probe, not a claim that a value was set.
MEDIA_BASE_URL was na on web-dev — the single reason no image could have resolved even with a correct bucket. Now set on all three environments. UAT deliberately points at the dev bucket, because UAT and Dev share one Supabase project and a different bucket would show broken images for items Dev created.
Public read proven on both public buckets by put → fetch → delete → confirm-404. Both private buckets have r2.dev disabled, which is the permanently correct state.
⚠ Prod's Supabase R2 secrets are UNVERIFIED. The stored PAT is a Dev-account token, so supabase-as.mjs refused the prod preflight by design. Reported as unverified rather than assumed.
CR-26.0.1-76 — the render path
⚠ The bug this fixes would have hit only Dev and UAT — the two environments the owner tests on — while production worked. Measured, same object, both origins:
| origin | direct | /cdn-cgi/image/… |
|---|---|---|
media.qrsetu.com (zone) | 200 | 200, and the PNG's IHDR really is 192px |
pub-…r2.dev (dev/uat) | 200 | 404 |
CatalogBlock now renders the first projected image and the designed no-photo state — a camera glyph, which CatalogueSection.dc.html specifies explicitly: "A missing photo is a DESIGNED state, never a broken image look." The aspect box sits outside the branch so a mixed catalogue does not jump.
apps/mobile gains startMediaBootstrap(), because EXPO_PUBLIC_MEDIA_BASE_URL was read by nothing — the owner had set it and it did nothing.
CR-26.0.1-77 — the merchant photo picker
apps/web desktop catalogue (DOM) + apps/mobile/src/lib/itemPhotoPicker.ts · Dev
The first client caller of manage-media, so this is the change that makes an image reachable at all. Two slots on the desktop catalogue: quick add at 56px, each table row at 40px, both per the design.
Deferred vs immediate is the one design-shaped decision. The design keys its quick-add slot on a pre-minted item id so the photo is dropped before the name is typed. Our catalog_item_media.item_id is a foreign key, so nothing can attach to an item that does not exist. The quick-add slot therefore normalises and previews locally and uploads the instant createItem returns — the merchant's order of operations is identical, only the moment of upload differs. The alternative (a client-minted primary key on create) is a decision about who owns identity and needs its own ADR, not a side effect of an upload feature.
⚠ A failed photo does not report a failed save. The listing was saved, and saying otherwise makes the merchant re-enter a row that already exists — on a 1,400-item run a duplicate costs more than a missing photo. The message names the recovery instead.
⚠ The re-encode is a privacy control, not an optimisation. A camera photo carries GPS, which for a home-run business is the merchant's home address on a public card — and with presigned direct-to-R2 uploads the client is the only place it can happen, because no Edge Function sees the bytes (QRS-813). createImageBitmap with imageOrientation: 'from-image' rather than an <img>, because drawImage from an element silently loses EXIF orientation and would store every portrait phone photo rotated 90°: a bug that appears only for camera picks, never for the screenshot used while testing.
Not at parity, therefore not release-ready. The native helper type-checks but is wired into no screen (QRS-811); adding a photo control to the phone catalogue needs a pull of mobile-console/Catalogue.dc.html rather than a guess. Per CLAUDE.md the scope valve is the feature, not the release.
CR-26.0.1-78..80 — unblocking the merchant journey on Dev
Driven by a merchant demo with the CI quota exhausted until 31 Aug. Four blockers, three of them invisible from the code alone.
| # | blocker | how it presented |
|---|---|---|
| 1 | get_consumer_item_feed had never existed | Browse threw on every request (404 PGRST202) |
| 2 | provisioning creates the card draft, nothing could publish | scanned QR 404s, listing invisible |
| 3 | deploy:manual had never worked (cwd) | died on ENOENT after running every gate |
| 4 | my own false claim about get_my_catalogue | photo uploaded, merchant's table showed nothing |
⚠ check:rpc would have caught #1 and is wired into no hook and no workflow (QRS-742). Second defect through that hole, and the first that would have killed a live demo.
⚠ #3 is the one worth sitting with: the escape hatch built for an emergency was broken until the emergency arrived. Every path in the script is repo-root-relative while the invocation CLAUDE.md documents sets cwd to apps/web. It burned type-check, tests and build before failing.
⚠ The publish control is a FLAGGED design deviation. Publishing belongs on desktop-console/CardEditor.dc.html, which is not built. It sits on the Catalogue screen — where the design already puts the card-facing action — and moves when CardEditor ships.
Also corrected: the card route's 404 under the run-web driver was the mock, not the app. Against real Dev data the card returns 200 with its display name, items and prices. The driver now accepts RUN_WEB_REAL_SUPABASE_URL, which is what settled it.
CR-26.0.1-79 — the publish control
packages/data/src/setuCardEditor/publish.ts · SetuCardLivePanel.tsx · Dev
provision_merchant_workspace creates every card status = 'draft'. Both reads the journey depends on require 'published' — get_public_setu_card (so a scanned QR 404s) and get_consumer_item_feed (so the listing is invisible). manage-setu-card has had a publish action all along and nothing in the app called it. A merchant could finish onboarding, build a catalogue, and have nothing a buyer could reach.
⚠ Flagged design deviation. Publishing belongs on desktop-console/CardEditor.dc.html, which is not built. The control sits on the Catalogue screen because the design already puts the card-facing action ("See it as a buyer") in that header region. It moves to CardEditor when that ships — it is not a second publishing surface.
⚠ Not an implementation of SetuCardEditorService. My first attempt conformed to that interface with the unbuilt methods throwing, and it needed an as unknown as cast because SetuCardEditorResult carries a whole SetuCardEditorState. Casting past that would have been a lie in the type system to make a narrow function look broad.
⚠ It shows no status before the first publish — a real limitation, not an oversight: there is no live merchant-side read for card status (get_my_auth_context is archived, getMySetuCard unimplemented). Publishing twice is harmless; the EF is idempotent on the action.
CR-26.0.1-80 — deploy:manual cwd fix, and the deploy itself
apps/web/scripts/deploy-manual.mjs → devv.qrsetu.com
⚠ The escape hatch built for an emergency was broken until the emergency arrived. Every path in the script is repo-root-relative, while the invocation CLAUDE.md documents (npm run -w @qrsetu/web deploy:manual dev) sets cwd to apps/web — so it resolved apps/web/apps/web/src/app/entry.server.tsx and died on ENOENT after running type-check, tests and build. Nobody had hit it because the manual path is a quota-window measure that had not been used.
Fixed by anchoring cwd to the script's own location via import.meta.url, so it is correct from the repo root, from apps/web, or anywhere else. Deployed with all three vars bound and the script's own four smoke tests passing.
Live: merchant sign-in, onboarding, console home, catalogue, /<slug>/setu-card and marketplace Browse all 200; the legacy festival-stall segment 301s to ganapati-idols.
CR-26.0.1-81 — the merchant Catalogue's remaining design gaps
apps/web merchant catalogue + ConsoleShell + @qrsetu/domain · Dev · 21/42 → 37/42 pass
Sixteen contract rows, all built from Catalogue.dc.html and vendor-core.js at round 30, transcribed rather than paraphrased.
The centrepiece is the facet-completeness loop, which this screen had none of: the nudge names the lost sales → the chip (or the nudge's CTA) filters to the incomplete items → a row chip opens its editor → the bulk bar fills the rest. It is also why the buyer-side facet filter has anything to filter on.
⚠ The bulk vocabulary replaced a stopgap that failed exactly where it mattered. Options were derived from values the catalogue already contained — so a brand-new catalogue offered nothing to set, and bulk-filling 200 idols was impossible precisely when most needed.
⚠ The schema disagreed with the design in four ways, and the worst was casing: 'idols' vs the design's 'Idols' would have matched nothing, making every row read "Not asked for Idols" — a facet mechanism that appears to work and asks for nothing.
⚠ Clearing a facet now deletes the key rather than storing ''. attributeGaps counts a present key as answered, so an empty string would clear the warning colour while leaving the item just as invisible.
Not closed, each named on its row: row_clone (needs a second media write), row_delete (recorded divergence — archive, not delete), photo_reaches_card_and_phone (needs the native picker), toast_undo (blocked, no restore action), row_stock_state (blocked — get_my_catalogue projects no order facts, so RESERVED is unreachable; QRS-828).
CR-26.0.1-82 — two catalogue-write defects, found from a live log
manage-item + the desktop catalogue - Dev
⚠ Desktop item creation had never once succeeded. The client sent stock_quantity with no track_inventory, and manage-item refused every create. The EF was right, and the DB agrees since catalog_items_stock_requires_tracking would otherwise refuse it with a bare constraint name.
A unique piece now sends neither field, which follows catalogItemState's own model: a unique item's sold-ness comes from the order path, a tracked line's from stock. Tracking a unique item at 1 would put its singularity in two places, free to disagree.
⚠ is_unique was silently dropped on every create - absent from OPTIONAL_FIELDS, so the "One of a kind" choice reached the database as the false default with no error anywhere. Three readers were consuming a value no writer set: get_my_catalogue (via a migration written specifically to expose it), stockLabel, and the mobile dashboard's piece count. The lines-vs-pieces distinction was computed from a column that was always false.
The failure mode is why it survived: an absence, not an error. A rejected field gets noticed in a week; a dropped one looks like a merchant who never chose.
⚠ This was the THIRD distinct cause behind "I cannot add a catalogue item", after the disabled fields and the workspace_id key. One symptom, three independent causes.
CR-26.0.1-83 — R2 CORS: why media upload could never have worked
supabase/storage/r2-cors-media.json - Dev
⚠ The bucket had NO CORS configuration. The upload PUTs from the browser to <account>.r2.cloudflarestorage.com, which is cross-origin, so the preflight failed and no byte was ever sent. Invisible server-side, because manage-media succeeded — it issued a URL that was then unusable.
⚠ manage-media's own header made this inevitable and I missed it. It explains at length why bytes go direct to Cloudflare rather than through the function — which is correct, and is exactly what makes bucket CORS a hard requirement. I built the presign path, verified the credentials, probed public read end to end, and never tested a cross-origin write. Verifying the half you built is not verifying the journey.
Explicit origin allow-list rather than *. Verified both ways: a browser-shaped OPTIONS from devv returns 204 with the right allow headers; an unlisted origin returns 403 with no allow-origin.
Committed as a file, so uat and prod are reproducible instead of a console click — this is one of the nine change classes no gate can observe.
CR-26.0.1-84 — the purge log stops crying wolf
⚠ Every card write logged WARN Card cache purge failed, because CLOUDFLARE_* is unset on Dev (QRS-306) so the purge path is reachable only through its failure branch. A warning that fires constantly and cannot be fixed by its reader devalues every other warning — and this one was actively misleading, read as the cause of a media failure it had nothing to do with.
Not provisioned is a deployment state, not a fault — the distinction r2Config already draws for R2. Now CLOUDFLARE_NOT_CONFIGURED, logged at INFO under its own event, naming which vars are absent and never a value.
⚠ A test pinned the old WARN behaviour and failed — the test doing its job. Replaced with two that assert the INFO branch and prove the failure event is not also emitted. The genuine-failure branch is a named placeholder, because reaching it needs a configured zone that then refuses: visible missing coverage beats silent absence.
CR-26.0.1-85 — media upload keys must be bare UUIDs
packages/data/src/media + platform-globals.d.ts - Dev
⚠ Every upload 400'd on its first call. uploadImage appended -issue / -confirm / -cleanup to the caller's key to make three distinct ones — which stops each being a UUID, and validateIdempotencyKey accepts a UUID and nothing else.
The reasoning behind it was sound and still wrong. The ledger keys on key alone, so two actions sharing one key would make confirm replay the issue response — a real constraint, correctly identified, argued at length in a comment. The scheme satisfied it and violated a second constraint already enforced two files away. Two valid requirements; one checked — and the comment defending the choice made it read as considered rather than unverified.
Now: the caller's key reaches issue_upload unchanged (the step where a duplicate costs an orphan row and a gallery position), and confirm/cleanup get fresh UUIDs. confirm is guarded by state, not the ledger, so nothing is lost.
⚠ crypto.randomUUID had to be declared, with a caveat. Unlike fetch, it does not pass platform-globals.d.ts's own admission test unaided — Hermes defines no crypto at all (QRS-279). It works because this repo polyfills it and imports that polyfill first in _layout.tsx. If that import moves, the declaration becomes a lie and a native build throws at the first key.
⚠ The deploy gate caught a type error in my own test that vitest had passed — npm test transpiles without type-checking. The two-speed split doing exactly what it is for.
CR-26.0.1-86 — every published vendor was invisible in the marketplace
set_workspace_location · manage-setu-card · get_my_context · LocationPanel - Dev
Measured, not inferred. get_consumer_item_feed with no area returned 4 items; with p_area set it returned 0 for all eight cities tried. Both live vendors had city_key null. Browse rendered 200 and listed nothing, the cards said published, the QR path worked - nothing looked broken anywhere.
Three independent facts had to line up: location is optional by design (correct, unchanged); provision_merchant_workspace was the only writer of city_key and runs once; and manage-setu-card appeared to let a merchant fix it, as free text on setu_cards while the feed filters workspaces.city_key.
⚠ That third one was worse than useless, and it was proven rather than argued. Against a local Postgres, setting setu_cards.city='Pune' leaves the feed's own predicate unmatched; the same field allowed card city Mumbai beside discovery city_key='pune'. It also contradicted the rule 20260817170000 states in its own words: "state is never collected as free text."
The fix: one RPC writing both tables transactionally, service_role only (it is SECURITY DEFINER and takes the user id as a parameter, so authenticated would be privilege escalation), free-text city/state removed, four keys added to get_my_context so the console can see the gap, and a conditional nudge. 16/16 mutation tests on real Postgres, 17 Deno tests.
⚠ A near-miss worth keeping. The get_my_context migration was first written by re-typing the body from a partial read, which invented the entire user block. Applying it would have silently broken the one live merchant context read. Regenerated from pg_get_functiondef by asserted substitution - CLAUDE.md's rule, now with an incident behind it.
⚠ Two earlier "findings" this session were both wrong: an image 403 that was Cloudflare error 1010 refusing a Python user-agent, and a /marketplace 404 that its own route module documents as deliberate. Isolating the variable is the only reason neither was reported as an outage.
Not backfilled, deliberately. The correct city for a real stall is not derivable from the database, and a wrong city is indistinguishable from a right one to everyone except the vendor.
CR-26.0.1-87 — get_my_context projects the location
Four additive keys (state_key, city_key, state, city) inside the existing workspaces[] objects, so the console can SEE the discoverability gap, pre-fill the picker, and confirm a save. The write path is useless without the matching read. Read from workspaces - the tenant, and the column discovery filters on - never the setu_cards mirror. Generated from pg_get_functiondef by asserted substitution; see the near-miss noted above.
CR-26.0.1-88 — set_location action, and free-text city/state removed
Contracting, floor 26000100. The narrowed fields had zero senders in any shipped build, so the floor is declared because a contraction must name one, not because a caller was found. A test that pinned five actions over a hardcoded array was rewritten to derive from MANAGE_SETU_CARD_ACTIONS - set_location became the sixth and the old test still passed, because "these five are accepted" stays true when a sixth is added. New tests also pin the contraction, so re-adding free-text city fails loudly.
CR-26.0.1-89 — LocationPanel and the location write on the location seam
The write lives on the location seam because that seam already owns the vocabulary: getStates / getCities produce the very keys it consumes, so the picker and the save cannot disagree about what a valid key is. The panel is a flagged deviation from CardEditor.dc.html (not built) and moves there when it ships. The nudge is conditional - a merchant with a city sees one quiet line, only an undiscoverable one sees the warning - because a prompt that cannot be acted on trains the merchant to ignore every prompt.
CR-26.0.1-90 — the communications envelope and the WhatsApp registries
Five tables, one immutable function, nine indexes, four triggers. Additive only — nothing existing is altered or dropped, which is why requires_min_app_build is null rather than a floor picked defensively. RLS on all five, revoke all … from public, anon, authenticated on every one, verified against the live database after apply rather than asserted: zero rows in information_schema.role_table_grants for anon/authenticated/PUBLIC.
That posture is load-bearing, not cautious. CLAUDE.md's three-category rule makes authenticated the logged-in general public, and these tables carry phone numbers and message content — a TO authenticated policy here would expose every consumer's messages to every other consumer.
Applied via supabase db push, not the MCP connector: db push registers the version from the filename, while apply_migration stamps its own and would have left a QRS-267 orphan. A both-directions diff of repo files against applied versions showed zero drift either way.
Four review amendments are encoded in the schema rather than left to the code: retry columns exist but the OTP path never writes them and the comment says outright that no drain worker, cron or scheduler exists anywhere (QRS-885), with a partial index on queued as the compensating control; idempotency_key is NOT NULL UNIQUE so a retried enqueue is a no-op by constraint; campaign_id is present now because adding a column to a high-volume ledger later is a migration on a table that will not be small; phone_number_id and template_language are recorded per message because the registry is replace-on-sync and mutable.
⚠ A gap caught before apply that no gate would have seen: the four mutable tables had updated_at and no touch trigger, while 17 existing tables attach one. An updated_at written at insert and then frozen reads as maintained while being stale — and every registry here is replace-on-sync, so "when did this last change" is exactly the question asked of it.
CR-26.0.1-91 — send-auth-otp, deployed to Dev as a probe
It sends nothing. Recorded as a change record precisely because a deployed-but-inert function reads as a working feature — config.toml's own comment calls that "worse than absent", and declaring it is what stops it being a silent state.
It answers one question nobody had verified, which decided the architecture: can Supabase phone auth run with a Send SMS Hook and no provider? Yes (CR-26.0.1-93). No client can reach it — there is no EDGE_FN entry and none may be added; GoTrue is the only caller. The OTP is never logged, only its length. Zero imports, so nothing else could explain a failure of the one thing it measures.
Replaced by the real adapter in step 6, which stays a thin adapter: no Meta knowledge, no ledger logic, both in the shared core so the next sender is another adapter and not a second send path (ADR-0029 D2). Its README states that boundary explicitly, because "the OTP sender" as a name is what invites the second path.
CR-26.0.1-92 — verify_jwt = false for send-auth-otp, and the heading it forced
Necessary, not convenient: GoTrue authenticates with a Standard Webhooks signature, never a JWT, so with the gateway check on its call is rejected at the edge — indistinguishable from the hook never being reached. Declared in the same change that first deploys the function, because manage-reminder shipped with no entry and turned out to be verify_jwt = false by accident rather than by decision. check:fn-config cannot help: it exits with usage unless handed --project (QRS-643).
⚠ The heading read "EXACTLY ONE", and before that "THERE ARE NO … ANY MORE" — wrong from the moment razorpay-webhook landed. A hand-maintained count beside the thing it counts has now drifted twice, because whoever adds the next entry is never whoever wrote "exactly one". The fix is not a better adjective: the block now names which and why per function, and carries the awk one-liner to re-measure. It also records that the three exceptions use three different schemes — so "webhook" is not one exemption but three arguments, and _shared/webhook.ts cannot verify the Supabase one.
CR-26.0.1-93 — Dev GoTrue: phone auth on, Send SMS Hook wired, no provider
An auth_setting is invisible to every automated gate, so the evidence is a live functional probe, not a claim that a value was set. Every auth incident in this project came from this class, and QRS-273 was closed on exactly such a claim while the auth log later proved sign-in had never once worked.
POST /auth/v1/otp → 200 in 2382 ms, with POST | 200 | .../send-auth-otp in the function logs. The point is what was not set: sms_provider reads "twilio" — GoTrue's default enum value, not a configuration, with every credential field null and no such account existing. It is never consulted when a hook is enabled.
No third-party messaging or authentication provider is involved: Official WhatsApp/Meta → QR Setu-owned integration → Supabase auth/session, per the owner's standing requirement.
Route A is therefore confirmed — Supabase owns the code lifecycle and sets phone_confirmed_at; we own delivery. ⚠ We must not add our own attempt counter: that is a second source of truth for "is this code still valid" (the QRS-249 class). Our limits are about cost and portfolio capacity and sit before Supabase, never beside its verification.
⚠ Known residual while the hook is a stub: a phone sign-in on Dev returns success and delivers nothing. Acceptable on Dev and closed by step 7 — recorded rather than left to be rediscovered as a bug. ⚠ Still owed: sms_otp_exp is 60 seconds against the template's declared 10 minutes, which would expire a code before a user could plausibly type it. That lands with the real sender in step 7, not while the hook is inert.
CR-26.0.1-94 — the WhatsApp registry seed
Three rows so the send path can resolve a sender instead of reading one from env. Without them resolveSender() throws and nothing sends — the correct failure, and why this is a migration rather than a manual insert nobody can reproduce.
The ids are identifiers, not credentials. The only secret is WHATSAPP_ACCESS_TOKEN, which appears nowhere in this repo. Env holds the credential; the registry holds the routing.
⚠ The same WABA serves Dev and Prod, deliberately: QRSETU has one WhatsApp Business Account and Meta offers no sandbox for a verified business number. A test send from Dev is a real message, at real cost, against the real quality rating. That constrains test volume; it is not a reason to fabricate a second account.
⚠ A seed, not an authority. status, quality_rating and messaging_limit_tier change without us acting. The latter two are left NULL rather than guessed — a fabricated starting value would be indistinguishable from a real one that had been synced.
Verified by reading the live rows back: 1 WABA active, 1 connected default sender, 3 AUTHENTICATION templates APPROVED in en/mr/hi.
CR-26.0.1-95 — whatsapp-webhook, deployed to Dev
It shipped before the sender on purpose: it is the only thing that makes a failed send visible. Without it a message sits at sent forever whether it arrived or bounced, and building the sender first would mean debugging delivery with no delivery data.
Meta's scheme is hex over the raw body with a sha256= prefix — the same family as Razorpay's, so _shared/webhook.ts is right here and standardWebhooks.ts is not. Three anonymous functions now exist and each authenticates by a different mechanism.
The response policy is inverted from an ordinary API because Meta treats non-2xx as retry: 401 bad signature · 200 duplicate (the constraint is the dedupe) · 200 + dead letter uncorrelated · 5xx only when the database failed, the one case where the event happened and we lost it.
⚠ The event id is status:{wamid}:{status}:{timestamp}, not the wamid. Keying on the wamid alone makes delivered and read the same event, so the second is dropped as a duplicate and the read receipt vanishes with no trace — while a genuine redelivery still produces an identical id, so the constraint still dedupes it. Both directions pinned by test.
Live probe against the deployed function: GET with the correct token echoes hub.challenge as plain text (200) · wrong token 403 · POST unsigned 401 · POST garbage signature 401 · PUT 405. The responses carry our strings, which confirms verify_jwt = false took effect.
⚠ The POST 401s prove fail-closed, not signature verification. WHATSAPP_APP_SECRET is unset, so the handler returns 401 before verifying, and the two are indistinguishable from outside. The verification logic is the same verifyHmacSignature covered by 13 tests, but the live end-to-end signature proof is owed and is not claimed here.
CR-26.0.1-96 — verify_jwt = false for the webhook, and a second stale count
Necessary, not convenient: Meta cannot send a Supabase JWT, and its HMAC over the raw body is a stronger control than the gateway check.
⚠ A second stale count was found and fixed in the same file. The heading above razorpay-webhook read "THE ONE verify_jwt = false FUNCTION IN THE PROJECT" — true when written, stale the moment send-auth-otp landed. That is the third drift of a hand-maintained count here ("THERE ARE NO…" → "EXACTLY ONE" → "THE ONE…"). It now names a position ("the first of three"), which cannot go stale when a fourth is added.
CR-26.0.1-97 — Dev Edge Function secrets
WHATSAPP_ACCESS_TOKEN, SEND_SMS_HOOK_SECRET and WHATSAPP_WEBHOOK_VERIFY_TOKEN are set. No value was printed, logged or placed on a command line.
⚠ SEND_SMS_HOOK_SECRET was read back from the project's auth config, not generated — it must be the same value GoTrue signs with. A fresh one produces a hook that rejects every real call, which reads as a broken algorithm rather than a mismatched secret and sends someone rewriting correct code.
⚠⚠ WHATSAPP_APP_SECRET is NOT set and cannot be by this session — it is a Meta credential from the app's Settings → Basic page. Consequence, stated rather than glossed: whatsapp-webhook fails closed with 401 on every POST. Correct, but indistinguishable from a bad signature, so a missing value looks like an attack in the logs. Until it is set, no delivery receipt can be accepted.
check:env-drift could not confirm parity: it requires both project refs and this session's token deliberately cannot see production. Noted rather than worked around.
CR-26.0.1-98 — send-auth-otp, the real adapter, proven end to end
Route A: GoTrue owns the code lifecycle and sets phone_confirmed_at; this function only delivers. ⚠ We do not add our own attempt counter — that would be a second source of truth for "is this code still valid". Our limits are about cost and portfolio capacity and sit before Supabase.
⚠⚠ The 5-second budget is half what the plan assumed. Supabase documents 5s for the whole invocation including its own retries. Hence a 2.5s per-attempt ceiling, two attempts, and a deadline that refuses to start an attempt it cannot finish.
The response contract is a control, not a report. 200 proceeds · 429/503 make Supabase retry inside that same budget · 400/403 surface as a 500. ⚠ A rate limit must not return 429 — it reads as "try again shortly" and would burn the remaining budget re-hitting a limit that cannot clear in two seconds. Refusing once honestly beats refusing four times.
Proven live, not asserted. A correctly-signed request returned 200 in 3517 ms and Meta returned a wamid. The ledger row reads status=sent · +919999999901 (E.164 with the plus) · phone_number_id=1203827669491594 resolved from the registry, not env · variables={} (no OTP persisted — verified in the database) · attempt_count=0, next_attempt_at=null (the OTP path never writes the retry columns). A replay of the same webhook-id returned 200 with no second send, and exactly one row survives.
⚠ The probe used a number that is not on WhatsApp, so the whole chain ran without messaging a real person — and note that Meta accepts and reports delivery failure later by webhook, so a 200 from Meta is not delivery.
Three defects were found by probing and fixed here: QRS-939 (a PostgrestError rendered as an unhelpful object string), QRS-940 (a ledger FK violation could block authentication), QRS-941 (message_id was never populated even when correlation succeeded).
CR-26.0.1-99 — WHATSAPP_APP_SECRET, closing F9
Supplied by the owner and pushed from the gitignored .env.whatsapp; never printed or logged.
It converts a claim into a proof. Before it, the webhook returned 401 on every POST — the correct fail-closed posture, and indistinguishable from a bad signature, so the earlier probe could only demonstrate fail-closed. With it set: a correctly-signed delivery is accepted and recorded, a redelivery of the identical body is absorbed by the unique constraint with no second row, and a body tampered by one word against the original signature is refused with 401.
All four WhatsApp-path secrets are now set.
⚠ Still owed, outside this session: Meta's webhook subscription must be pointed at the deployed whatsapp-webhook URL with the verify token, or no real delivery receipt will ever arrive. The function is ready; nothing is sending to it yet.
CR-26.0.1-100 — OTP expiry aligned, and the hook secret re-provisioned
sms_otp_exp 60 → 600, matching the template's declared 10 minutes. A 60-second server-side expiry meant the code died while the user was still reading the message telling them it lasted ten — the system and the message disagreed, and the message is the one the user believes.
⚠⚠ The important half: hook_send_sms_secrets is returned MASKED by the Management API. PATCH it with v1,whsec_<44 base64 chars>, GET it back, and you receive 64 characters matching ^[0-9a-f]{64}$ — a SHA-256 digest. The documented workflow is to place that secret in the function's environment and the obvious way to obtain it is to read the config, so every key was being derived from a hash and verification failed on every request with no_matching_signature — an error pointing at the algorithm, the encoding, or an attacker, and never at the value.
Why no test caught it: the probe signed with the same read-back value the verifier used, so the two agreed with each other while GoTrue disagreed with both. That is verbatim the trap standardWebhooks.test.ts warns about in its own header. For a signature scheme, the only meaningful test is one where the counterparty produced the signature.
⚠ The diagnostic was fooled too: "base64-ish: true, decodes to 48 bytes" — because hex is a subset of the base64 alphabet. A character-class check cannot distinguish an encoding from a hash.
Fixed by generating once and writing the same in-memory value to both places in one operation, never reading it back. The original derivation was then confirmed correct on the live path: key_form: "base64-decoded". Evidence is a real signInWithOtp, not a self-signed request: 200, Meta returned a wamid, ledger row status=sent.
CR-26.0.1-101 — the FK guard proven load-bearing, and 600ms recovered
QRS-943 — auth hooks run inside GoTrue's uncommitted transaction. The hook fires with a user.idnot yet visible in public.users from any other connection — logged for a user present in both tables moments later. ⚠⚠ Permanent and universal: it happens on every new user's first OTP, the most common path in the product.
🔎 The QRS-940 guard was written as a hypothetical against a probe artifact. It turns out the row is always absent on first sign-in — without it, every new consumer's first sign-in would have returned 503. Both the migration and module comments asserted the opposite ordering: plausible, confidently written, wrong. Corrected. The log dropped warn → info with the cause named, because a warning on every signup trains people to ignore the channel.
QRS-944 — 2822 ms of a hard 5000 ms ceiling. ~1.2 s was Meta; the rest was sequential Postgres round trips from the edge for four mutually independent reads. One Promise.all: 2822 → 2215 ms, warm client total 1846 ms. ⚠ The checks were not parallelised — the category assertion and limit comparison still run in order before anything is enqueued or sent.
🔎 An Edge Function is remote from its database: a round trip is 100–300 ms and four is most of a second. Against a hard external deadline that is the difference between fitting and not.
The multi-derivation key fallbacks were removed once base64-decoded was measured. Three code paths on the authentication hot path when one is proven correct is speculative generality; the reasoning survives as a comment and a test.
CR-26.0.1-102 — the phone-OTP client seam
⚠ The channel is deliberately absent from the method names. sendPhoneOtp, not sendWhatsAppOtp: the caller asks for "a code to this number" and the channel is a server-side decision. Naming it here would put a delivery detail into a contract with no business knowing it, and every screen would then encode the choice.
⚠ A separate method rather than a widened verifyOtp — the underlying call differs (type: 'sms' vs type: 'email'), and one method taking "an identifier" would have to guess which from the string's shape. That is how an email with a leading + becomes a phone number.
Two failure distinctions, both load-bearing:
| pair | why |
|---|---|
rate_limited vs send_failed | "Try again" in front of a rate limit invites hammering the button — exactly what the limit refuses, and every attempt that gets through costs a real message |
expired_code vs invalid_code | Opposite actions. Collapsing them tells somebody holding a code that was right ten minutes ago that it is wrong |
Both classified by status code, never by GoTrue's message text.
The stub can now express a CONSUMER, defaulting to individual — the email stub's own comment explains why the opposite default "would make a consumer routing bug invisible to exactly the tests meant to catch it", and this is the consumer route. It also models an existing account keeping its category, and every spelling of one number being one account.
QRS-945 fixed here too: three screen tests mocked authService by hand and one named signInWithGoogle, gone from the interface since P3, while omitting two real methods. jest.mock's factory returns any, so nothing could see it. Replaced with a derived factory carrying a two-sided compile-time guard — adding an interface method now fails to compile until it is listed. useEmailAuth's error union is likewise derived rather than re-listed; the hand-copied version broke when the service widened, which is the good outcome.
CR-26.0.1-103 — six real OTPs, two real handsets, owner-confirmed
The first validation by a human rather than by a query. Three sends to each of two numbers, 60–75 seconds apart, through the real signInWithOtp path. Latencies 4201/2057/2081 ms and 3708/2439/2858 ms; all six accepted, receipt confirmed on both handsets.
All six ledger rows: status=sent, wamid present, variables={} (no OTP persisted — verified in the database), template_language=en, no error_code.
🔎 The ledger confirmed QRS-943 without being asked to.recipient_user_id was null on the first send of a batch and set on the second and third — exactly the predicted pattern, because the account was being created and the row was still uncommitted inside GoTrue's transaction. Without the QRS-940 guard the first send would have been a 503 and no first code would ever have arrived. A guard written as a hypothetical, validated as load-bearing by an experiment run for a different reason.
⚠ A digit was flagged before sending: the requested number differed from the production sender by one digit, and the playbook warns that three numbers in this account differ by a digit or two. Raised before the first batch, not after.
Irreversible by nature — a delivered message cannot be recalled — hence no rollback section. The forward fix for a wrong recipient is to stop sending, which is why the digit check came first.
CR-26.0.1-104 — the C4 identity invariant
No account may exist without a verified phone number.
Supabase treats phone and OAuth as separate identities unless explicitly linked, and signInWithOAuth returns no phone at all — so two accounts for one person can never be matched by data afterwards. A consumer who signs up by WhatsApp and later falls back to Google lands in a second account with an empty product and concludes their data was deleted.
The mechanism was measured before anything was designed, because the plan named updateUser({ phone }) and marked it unverified:
| observation | consequence |
|---|---|
| same user id returned | it links, it does not fork ✅ |
phone empty, new_phone set | the number is pending |
a real OTP went through our hook (status=sent, wamid) | linking carries a full verification |
verification needs type: 'phone_change' | not 'sms' — GoTrue keys the code by type |
⚠ The predicate is a SESSION predicate, not a row predicate, and the distinction is load-bearing: the database legitimately holds phone set with phone_confirmed_at null — two such rows exist on Dev, from OTPs delivered and never entered. Applied to a row it would call those verified. It is sound for a session because a session only exists after verification.
⚠ The stub nearly proved the opposite of what it claimed. Its first version returned a userId derived from the phone rather than the account being linked to — a fork, the exact defect the invariant forbids, and every test against it would have passed while demonstrating the reverse. Caught while writing the tests, not by a gate. 🔎 A stub that answers plausibly is more dangerous than one that throws.
The derived mock's exhaustiveness guard was also mutation-tested — removing one method produces a compile error, restoring it passes — because an untested control is a belief.
CR-26.0.1-105 — PhoneAuthScreen, the round-12/14 sign-in spine
Additive: one leaf route, one screen, one hook, 26 copy keys per locale. Nothing existing changes behaviour — /sign-in, the wizard and the email path are untouched.
No sign-up / log-in tabs, and their absence is the design. Round 14: sign up and sign in are the same gesture. The email screen's ModeTabs were always cosmetic — AuthService already decided whether setup runs — so removing them deletes a control that could only ever mislead. Asserted as an absence, so a future edit re-adding them fails a test.
⚠ The channel is a prop, so an SMS fallback is a prop change and not a second catalog. ⚠ SHOWN_ATTEMPTS is a display count, not an authority — GoTrue owns the real limit, and a second counter would be a second source of truth for "is this code still valid" (QRS-249).
⚠ It is on /phone-sign-in and does not replace /sign-in yet, deliberately: cutting over in the same change that first builds the screen would mean both are reviewed as one thing, so a problem with either reads as a problem with both.
⚠ The splash and boot path are untouched, verified rather than asserted: app/index.tsx, _layout.tsx, BrandSplash.tsx, Wordmark.tsx and app.json were hashed before the work and are byte-identical after it.
Three things the gates caught, each a compile error rather than a runtime surprise: AppText has no h2 and FieldState is idle|valid|error; the analytics map is closed, so method gained 'phone' deliberately; and the locale store key is lang, not locale — a test set locale, a silent no-op, so three locales were "tested" at the device language. t() also interpolates {vars} itself, so a hand-rolled .replace() chain was removed.
Verified: 970 tests, type-check clean, lint zero, and the built route driven in headless Chromium with a screenshot — not merely a 200 from an SPA shell, which proves nothing about a screen.
CR-26.0.1-106 — set_my_display_name, the write path that did not exist
The sibling of set_my_primary_context, deliberately the same shape. handle_new_user sets display_name exactly once, at account creation from raw_user_meta_data — and the design's name step comes after the code is verified. Supabase does not rewrite metadata on a later OTP send, and Google arrives already named. The one path with nowhere to write is the consumer journey.
⚠ It accepts null deliberately — the step is skippable, and a person may later want the name removed; a setter that can only ever set makes deletion an escalation. Blank normalises to null so "absent" has one spelling, which matters because resolveAccountState reads exactly that field. 80 characters is an explicit refusal, not a truncation: silently shortening somebody's name is worse than refusing it.
⚠ Not audited, unlike primary_context. That column decides which product a person sees, so a record is answerable support. A display name decides nothing, and auditing renames would put PII in a second table that the DPDP erasure question makes strictly worse.
Verified in a rolled-back block: 3/3 guard rails hold — authenticated has the grant, anon cannot execute, and an unauthenticated call raises 42501 rather than no-opping.
CR-26.0.1-107 — the round-14 branch, and a flag that could never have carried it
⚠⚠ QRS-948 is the headline, and it is a silent skip avoided rather than fixed. isNewAccount is !hasCompletedOnboarding(ctx), and that returns true for every individual the moment they exist. So it is structurally false for the entire consumer population — and a host gating the name step on it would have skipped that step for 100% of consumers, with no error and no failing test. Consumers are the primary R1 audience.
🔎 Caught by asking what the session lets me ask, not which flag looks right — having just learned from QRS-918 that a missing field makes a bug untestable.
Fixed by adding the field, not reinterpreting one. AuthSession.displayName: projected from the context already in hand, carried on the cold-start read so a relaunch does not re-ask, cached for routing, and modelled in the stub so "signed in, no name yet" is a state a fixture can hold. resolveAccountState deliberately does not accept isNewAccount, with the omission documented so it cannot be reintroduced.
PhoneNameScreen is a single field whose skip is live even when empty — anonymous-first is a hard requirement, and a mandatory name here is a signup wall one step further in. Both actions disable during a save, because a live skip lets someone leave mid-write and the store would then update for a screen nobody is on.
Verified: 976 tests, and the built route driven in English and Marathi with screenshots — Devanagari renders throughout, so the three-locale copy is real rather than a silent fallback.
CR-26.0.1-108 — the choice first, the consumer story whole, the phone spine as the door
Three owner decisions of 2026-09-01 and one omission the owner found by opening the app.
The fork moved ahead of the story. Splash → For my business / For myself → story → sign-up. setAccountType records the fork without marking the story seen, so a kill on the choice screen cannot skip the first run forever — the one call that did both (markWelcomeSeen(type)) was right only while the fork was the story's last scene.
Phone became the entry for everyone. Measured before the change: entryRoute returned /sign-in and nothing navigated to /phone-sign-in, so the WhatsApp OTP spine was unreachable by any real user. Stated plainly: an account with no verified phone now has no way in, which is safe only because the owner declared every existing account test data. Google is absent from the phone screen for R1 by owner decision.
⚠⚠ QRS-950 — the consumer welcome story had never been built.WelcomeStory.dc.html exposes variant: Business | Consumer and branches every scene's art and copy on it; we had shipped the Business branch only, so somebody who tapped For myself was told about menus, payments and bookings. Now: five consumer scenes (the design's count, verified against the fetched file), verbatim copy in three locales, and the design's own illustrations ported to RN — openingArt(PEOPLE), problemArtConsumer, hubArtConsumer (kinds from consumer-data.js), buildCardConsumer, phoneRunArtConsumer. The Aadhaar scene is withheld from consumers on the design's own reasoning (an identity-document analogy applied to a person), and the business tagline hero is gated to the business variant. Open: the business story carries a sixth scene the current design lacks — QRS-951.
Verified: story.test.ts pins five scenes, _consumer art ids only and no aadhaar; scenes.test.tsx mounts all eleven illustrations; the real web export was driven through the consumer story with screenshots of every scene.
CR-26.0.1-109 — sign-up and sign-in made legible, and every designed error state
The design's answer to "am I signing up or signing in" is a state, not a tab.Onboarding.dc.html's own comment: "the difference is an answer that arrives AFTER the code, so it is a branch here and not a tab." The RECOGNISED state now names which account was found, whether it is complete or unfinished, and — when the tapped tile disagrees with the server — that the server's kind wins. One component for business and consumer; only PHONE_COPY vs PHONE_COPY_BIZ differs, by prop.
Every state the design's otpStage machine draws is implemented: invalid number · rate limited (wait) · delivery failed (this number may not be on WhatsApp — the SMS half is not built, no channel exists, QRS-952) · offline and no-route (CTA disabled with the design's note; on verify the typed code is kept and Try again resends it) · wrong with attempts left · locked (attempts exhausted or server throttle — cells removed, not greyed) · expired (own panel, cooldown lifted).
⚠ Two type gaps, found by asking what the seam could express. A throttled verify (429 over_request_rate_limit) collapsed into invalid_code, telling a rate-limited person they had mistyped — QRS-953; VerifyResult gained rate_limited and network, OtpResult gained delivery_failed and network, classified by status and GoTrue code, never message text. And a new business was indistinguishable from an abandoned one, so a first-time merchant landed on /dashboard with no workspace — QRS-954; resolveAccountState takes the account type and the host routes business → wizard.
Supabase's own docs establish that a failed send hook reaches the client as a 500, so why delivery failed is not knowable client-side — stated, not papered over.
Verified: 1056 mobile tests (PhoneAuthScreen 25 · route 8 · seam 23 · choice 6 · scenes 16), 519 domain tests, lint 0, type-check 0; consumer and business phone steps screenshotted in the real export. Open: send-auth-otp has zero tests (QRS-955); the attempt lock is client-side only (QRS-958); the design's purpose copy conflicts with CLAUDE.md's reassurance rule (QRS-956, owner).
CR-26.0.1-110 — five owner corrections to sign-in, fixed as one pass
Framing. Sign up and Sign in are explicit again: the email screen's ModeTabs are shared with the phone screen as FRAMING (owner decision over the design's round-14 removal — QRS-962). The heading and one short line follow the tab; the account's own state still decides the outcome, and the RECOGNISED state still names the account. No step indicator, no biodata-specific purpose text (QRS-956, QRS-957).
Validation, three layers (QRS-960). "Send code" had lit up on six digits. The field keeps digits only and caps at ten; Send is disabled until a valid Indian mobile; toE164 refuses incomplete and non-mobile numbers; and the send-auth-otp Edge Function refuses any recipient outside +91[6-9]\d{9} before touching Meta — because GoTrue's own check would have created an auth user with a bogus number and sent a real message.
Resend policy (QRS-961). 60 seconds between resends, three resends per window, then a five-minute wait with the minutes counted down and "{n} resends left" shown. The EF's per-recipient ledger limit is now four sends per five minutes.
Incorrect vs expired (QRS-959). GoTrue returns otp_expired for a wrong code too — Supabase's own troubleshooting guide says so — which is why a mistyped code read "This code has expired". classifyOtpRejection decides by time against the code's 600-second lifetime: inside it, "Incorrect code. Please try again." with attempts; after it, the expired panel. Three wrong codes lock the step for five minutes with a live countdown — client-side only, since GoTrue has no per-code attempt cap (QRS-963).
Verified: otpPolicy.test.ts 10 · phone.test.ts +5 · PhoneAuthScreen.test.tsx 31 under fake timers · full suite 1062, domain 534, lint 0, type-check 0; sign-up and sign-in framings screenshotted in the real export. ⚠ The hardened EF's Dev deploy and the GoTrue sms_max_frequency = 60 setting are recorded in the delivery log with what was actually applied (QRS-964).
CR-26.0.1-111 — the kill switch's server half, re-authored for v2 (QRS-965)
The product owner found a console error on /phone-sign-in: PGRST202 · "Could not find the function public.get_app_release_policy(p_platform)".
The table and RPC existed only under _archive_pre_v2/ — the ADR-0020 baseline dropped them and v2 shipped no replacement — while UpdateGate is mounted in app/_layout.tsx, so every cold start of the app issued a failing request, on the first screen a new consumer ever sees.
⚠ The read fails OPEN, and that is precisely why this mattered rather than why it did not. The service returns null, evaluateReleasePolicy reads null as ok, and the app carries on — so nobody was ever blocked and nothing user-facing was wrong. What was broken is the control: the kill switch could not fire. QRS-288's whole argument for shipping it before the first release is that a retired build is reachable only through code it already contains, so a switch that is inert on the day it is needed is worth nothing. A gate that fails open and logs a warning is indistinguishable from one that works.
Re-authored, not restored. updated_by references public.users (the v2 actor convention) rather than auth.users; the iOS seed URL is null rather than the archived id0000000000 placeholder, which would have sent a blocked merchant to a 404. Seeded permissive on all three platforms, so the mechanism is provably inert before it is armed. RLS on with no policy (the RPC is the access path), security definer with a pinned search_path, granted to anon by design — a retired build may be unable to sign in, which can be exactly why it was retired — and updated_by is withheld from the projection because it is a user id.
Applied to Dev and verified by command: the RPC returns the seeded android row, and the browser now reports get_app_release_policy 200 with the warning gone. The known-live-defect allowance in check-rpc-contract.js is deleted, because the gate's own stale-allowance rule would now fail on it.
⚠ Two caveats, stated rather than smoothed over. QRS-267 recurred: MCP apply_migration stamped its own version and the reconciling UPDATE was refused by the sandbox, so the repo file still reads as unapplied — self-healing, because the migration is idempotent by construction. And the 18-assertion pgTAP suite is written and unrun: both drives sit below the 15 GB floor, so the local stack was not started.
CR-26.0.1-112 — consumer onboarding, visual-consistency pass (QRS-966, QRS-967, QRS-968, QRS-969)
Four defects from the owner's review. Three of the four were shared by both personas — the owner had found shared components from the consumer side, which is worth recording because it changes what "fix the consumer screens" means.
The overlapping chip (QRS-966). RadialHub rendered height: size, but its ring nodes are positioned from the box's centre, so the lowest node's label escaped the box and was painted over by the tricolour chip beneath it — measured 5px, on the business story too. Height is now derived (hubHeight, plus a per-art exported constant the scene registry consumes), so the hand-written literals that went stale are gone.
The broken pill (QRS-967). NativeWind drops className on a LinearGradient — silently, with no error — so the offer/preferences chip rendered flexDirection: column with zero padding on both stories, stacking the icon on its label. The giveaway had been sitting in the code: borderRadius re-specified inline beside a rounded-xl class that was doing nothing. Layout moved to style; new parity rule R10 fails any className on that component, and it belongs in the parity gate because cssInterop registration here is already platform-conditional.
The missing gradient (QRS-968). All five consumer scenes, in all three locales, carried no brand-gradient emphasis while the business story beside them did. The transcription was faithful to the design and wrong for the product: the design renders a flat headline for both personas, and the business copy's markers were added in-repo under the brand rule. Marked per language now, and gated in the catalogue — GradientHeadline erases markers under jest, so no renderer-level test could ever have seen this. The consumer finale also takes the design's own larger final size, which it had never got.
The 37x19 target (QRS-969). The fork screen's "Sign in" link — the only escape hatch a returning person has — failed the 44x44 minimum. hitSlop looked like the fix and is not: it never reaches the DOM box on RNW.
Verified in the real export before and after each item: zero colliding labels on both stories, chip flexDirection: row with 8/12 padding, gradient word SVGs rendering, touch targets < 44x44: none, no console errors. Gates: lint 0 · type-check 0 · parity 10 rules · 1069 mobile tests.
⚠ QRS-639 was reproduced during this pass and deliberately NOT fixed here. Dark mode's imperative colour channel returns light values, and the new evidence is that it mis-colours backgrounds and brand ink, not only text — the welcome story paints near-white words on a cream ground. It is systemic (src/ui + tokens), ADR-0015 makes that design-first with the native builds as the gate, and its own row says not to fix it until the cause is established.
CR-26.0.1-111 addendum — the ledger reconciled, and a defect in the test that proved nothing yet
The owner restored the Dev CLI token (setx + supabase login), which closed both caveats above.
The orphan version is gone, through the sanctioned tool rather than the hand UPDATE the sandbox refused: migration repair --status reverted 20260902071602, then --status applied 20260902100000. Read back and diffed programmatically: 0 mismatched rows across all 74 migrations, local and remote agreeing in both directions — the bidirectional check the sixth rule asks for and that no gate here automates.
⚠ A trap worth knowing, because it makes a good credential look bad. setx writes the USER environment and does not update already-running shells, while supabase-as.mjs reads process.env.SUPABASE_ACCESS_TOKEN before the registry. A session opened before the setx therefore keeps serving the old token and fails with the identical "CANNOT SEE qr-setu-dev" message as a genuinely wrong PAT. Measured here as two different 44-character sbp_ values (registry tail 2dc0, process env tail f416). Compare the two sources before blaming the token.
The pgTAP suite is still unrun locally — the disk floor has not moved — but every security property it asserts is now verified on Dev by direct catalog reads: 3 rows seeded, 0 arming the switch, RLS on with 0 policies, anon/authenticated may execute, PUBLIC may not, neither role holds table SELECT, security definer with search_path pinned, argument named p_platform.
⚠⚠ And measuring it found a defect in the unrun test itself. §F asserted the projection with bag_eq over information_schema.columns — which does not list a set-returning function's OUT parameters (zero rows for this function), so it would have failed on the suite's first run, inside the file written to prove that very property. Replaced with pg_get_function_result(), exact about names, types and order, plus a legible assertion that updated_by is absent; both re-verified against the live schema before being trusted. An assertion nobody has executed is a hypothesis — and a suite blocked from running hides its own defects as well as the code's.
CR-26.0.1-113 — the grant pgTAP caught, and the Postgres version nobody had compared
Unblocking the local stack (QRS-970) let the pgTAP suite run for the first time, and it immediately earned its keep.
CR-111's table granted anon everything. It enabled RLS with no policy and omitted the revoke, so Supabase's ALTER DEFAULT PRIVILEGES left anon and authenticated holding DELETE/INSERT/REFERENCES/SELECT/TRIGGER/TRUNCATE/UPDATE. RLS hides the rows, not the privilege. It was the only table in public that anon held anything on, and it broke v2_isolation_test.sql, whose entire job is asserting that count is zero.
⚠ Dev looked clean and a fresh environment would not have been. On Dev the same file left only postgres and service_role, because it was applied through a different role whose default privileges differ. The local CLI-applied stack showed the grants. The same migration produces different grants depending on who applies it — so a CLI-applied Prod promotion would have inherited the wrong ones, and checking Dev alone would have certified it safe.
Fixed with the convention all 61 other post-baseline tables already carry, and gated: new check-sql-grants rule 4 fails any new table lacking a series-wide anon revoke. Rule 1 could never have caught this — it bans an explicit GRANT ALL … TO anon, and the dangerous grant is the one nobody writes down.
And the local stack was Postgres 15 while Dev and Prod are 17.6.1.155 (QRS-972). CLAUDE.md asserts the local major version "matches Prod" in the same section that tells you to develop SQL against it; the claim was false, so local results proved behaviour on a version the code would never meet. Now 17, verified by a full db reset: all 74 migrations apply and the whole suite passes — 12 files, 459 tests, Result: PASS.
CR-26.0.1-114 — the auth function finally has tests, and the split that made them possible
send-auth-otp was the only Edge Function on the authentication path and the only one with no tests/ folder (QRS-955). The reason was structural rather than neglect: every rule lived inside Deno.serve, so exercising the recipient rule meant standing up a signed HTTP request, a Supabase client and a Meta endpoint.
The split line is that a function belongs in helpers.ts if its answer is a function of its arguments alone. The signature verification, the ledger counts, the template resolution and the Meta call stay in index.ts, where _shared/tests/ and a live probe already cover them. index.ts imports every constant from helpers.ts rather than keeping a copy, so a limit changed in the tests and not in the handler would not compile.
No behaviour change — constants, comments and comparison semantics moved intact, and the limit comparison now routes through exceedsLimit rather than a second inline >= that could drift by one.
The 20 assertions encode incidents rather than shapes: the six malformed numbers GoTrue itself would accept (including the measured six-digit typo that would create an auth user with a bogus number and send a real message), anchoring escapes, the idempotency key's stability across a retry (the double-charge guard), the limit ceiling proven at 3/4/5, sms.phone winning over user.phone (they differ mid phone-change), and the language fallback being appended rather than substituted — a template can be approved in en and rejected in mr, and a sign-in must not fail over a copy problem.
⚠ The type checker caught a real signature error during the extraction: user.phone is string | null, not string | undefined. Verified: deno check clean, full EF suite 373 passed, mutation-tested (widening the recipient regex fails a test; reverting restores green). Redeployed to qr-setu-dev — version 12, ACTIVE, new bundle hash — so the repo and the environment agree even though the behaviour is identical, because the sixth rule's point is that inferring agreement is what goes wrong.
CR-26.0.1-115 — the namespace gets its own table (consumer plan 1.1)
The consumer identity page needs an address, and there was no consumer slug anywhere in the schema: setu_cards.slug is workspace_id-scoped, and a consumer has zero workspaces by definition. D1 decided a registry rather than a wider card table; this is that migration.
One primary key on the name is the whole argument. A consumer and a merchant cannot collide. Two tables with two unique indexes would need a cross-table uniqueness trigger, which is race-prone and needs advisory locks — one PK removes the class outright.
⚠ ON DELETE SET NULL, never CASCADE, and D1 records this as a real defect in its own first draft. Cascade is the house idiom here (reminders, conversations, workspace_members), so writing it is the natural mistake. The chain is what makes it silent: public.users cascades from auth.users, and manage-account calls auth.admin.deleteUser() — a hard delete. With cascade, a closed account would hand /priya-sharma back to the pool. Presence is permanent; ownership is not. The num_nonnulls(...) <= 1 check is <= 1 rather than = 1 for the same reason: a released row legitimately has no owner.
RLS is on with no policy and every privilege revoked from public/anon/authenticated — this table maps names to owners, so a client that could read it directly could enumerate who exists. Reads will reach clients through the 1.3 RPCs.
⚠ A real finding while pushing (QRS-979): a bare citext applied cleanly locally and was REFUSED by Dev with type "citext" does not exist (42704). The extensions migration asserts a bare reference is safe and says it was verified by probe — the probe was real, but it proved the LOCAL session's search_path carries extensions, and the session db push opens does not. Fixed with extensions.citext, which resolves in both. The failed push left Dev clean because the migration is wrapped in begin/commit.
Scope is the table only. 1.2 backfills merchant slugs and FKs setu_cards.slug into the registry; 1.3 adds the claim and availability RPCs, where the rule lives that a consumer-held slug and a reserved slug must return the same availability status — otherwise the check is an enumeration oracle for who exists.
CR-26.0.1-116 — the slug registry becomes the source of truth (consumer plan 1.2)
Every existing merchant slug is now registered, and setu_cards.slug carries a foreign key into public.slugs, so no slug exists in one place only — the plan's own acceptance wording.
⚠ A bare foreign key would have broken merchant provisioning on the next signup.provision_merchant_workspace inserts into setu_cards directly, so a FK with no registry row fails on a constraint whose name means nothing to whoever hits it. The bridging trigger therefore ships in the same transaction as the constraint: adding an invariant and the thing that satisfies it together is the difference between a tightening and an outage. It is a trigger rather than a patch to the one inserting function I found by grep, because a trigger does not depend on that grep having been complete.
⚠ No ON DELETE clause, and both directions are the point (verified confdeltype='a'): deleting a card leaves the slugs row standing, because presence is permanent and a name must not return to the pool because a merchant unpublished; deleting a slugs row while a card points at it is refused, which is the invariant a registry exists for.
Verified on Dev by reading the catalog: 5 cards, 5 registry rows, 0 unregistered, FK and trigger present, ledger head 20260902190000. Locally: db reset from scratch, check:sql 75 migrations, pgTAP 459 tests PASS.
🔎 Found while measuring, and not a defect: balaji-lahade is a published card whose slug is also in reserved_slugs under founder_protection — the founder legitimately holding their own reserved name. It is exactly the case 1.3 must not get wrong: reserved and taken have to look identical from outside while being different inside.
⚠ The trigger is a bridge and 1.3 retires it; its own comment says so. And its behaviour has no pgTAP test yet (QRS-983) — existence and backfill completeness were measured, behaviour was not.
CR-26.0.1-117 — the namespace gets a front door, and the oracle closes (consumer plan 1.3)
Four objects: resolve_slug_status (the namespace-wide availability answer), claim_slug (atomic, service-role only), get_my_slug (ConsumerHome's identity-card read), and resolve_setu_card_slug_status rewritten as a thin wrapper over the first.
⚠⚠ The load-bearing behaviour, which is decision D1: a consumer-held address and a platform-reserved one return the identical status. Answering taken for a consumer would move the account-existence oracle out of the public route, which D5 closes, and into this RPC over a dictionary of names. Consumer addresses are name-seeded and claimed at onboarding, so that would enumerate the largest population on the platform. taken now means a merchant address, whose existence a public Setu Card already publishes.
⚠ resolve_setu_card_slug_status had been reading the wrong table since CR-116, which is what made this urgent rather than tidy: it answered taken from setu_cards, so once public.slugs became the namespace a consumer-held address returned available and a merchant would have finished the whole wizard before failing on a primary key.
⚠ A defect in CR-116 is fixed here (QRS-985): its registration trigger used on conflict do nothing, so a card inserted on an address a consumer held attached itself to that person's registry row. Mutation-proven in both directions — the 1.2 body reproduces it inside a rolled-back transaction, the 1.3 body refuses it. The trigger is hardened rather than retired, reversing what CR-116's own comment promised: retiring it would put correctness back on having found every writer by grep, which is the dependency CR-116 explicitly rejected.
🔎 Two measurements worth keeping. The ::text cast in slugs_format is load-bearing, not stylistic: citext's ~ operator is case-insensitive, so the obvious spelling would accept MixedCase while reading as though it forbade it (measured: uncast t, cast f). And one of my own Dev probes was a false positive — prosrc like '%on conflict%' matched the comment in which I describe the line I had just deleted. A probe that matches your own documentation is not evidence about the code.
⚠ No client can claim an address yet. claim_slug is service_role-only by design; manage-account is its home, in plan 1.4.
CR-26.0.1-118 — an ownerless address becomes honest
118 adds slugs_release_on_owner_loss, and it lands before any client can claim an address rather than after. slugs.user_id is ON DELETE SET NULL, and manage-account still hard-deletes, so the first account deletion after a claim would have left a row reading user_id = NULL, state = 'ACTIVE' — a lie in the state vocabulary, and a silent one, because resolve_slug_status would still keep the name out of the pool. A trigger rather than a fix in the deleter, for the reason CR-116 already settled: it holds against every writer, including plan 1.4's soft-delete rewrite.
⚠ Not a substitute for that soft delete. A hard-deleted account releases its address with no reclaim_hmac, so that person can never prove a claim to their own name if they return.
CR-26.0.1-119 — the address becomes claimable
119 gives manage-account the claim_slug action — the only path to an RPC granted to service_role alone. user.id comes from the verified JWT, never the body. 23505 maps to one 409 for three different causes on purpose: held, reserved, or you-already-have-one. Telling them apart would rebuild the oracle decision D1 closes.
⚠ Deployed to Dev and NOT exercised end to end. The validator is unit-tested and the RPC beneath it is covered by pgTAP, but the two have not been driven together with a real session. ef_code is one of the nine change classes invisible to every gate, so this is the record that it was read rather than probed.
CR-26.0.1-120 — the sign-in flow learns whether it has met you before
get_my_context projects is_returning, derived from auth.users.last_sign_in_at against created_at. Additive projection only.
⚠ What it fixes: resolveAccountState inferred "new" from the absence of a display name and a workspace, both legitimately empty for a real returning account. Three live Dev accounts — including both the owner created while testing — classified brand new on every sign-in, permanently. The person was pushed back through sign-up, the recognition panel could never render, and each loop spent an OTP until the rate limit fired and blamed the OTP.
🔎 This reverses the assessment's own recommendation, and the reasoning is worth keeping. A users.signup_completed_at column only pays off if hasCompletedOnboarding is restructured to read it — and that returns true for every individual deliberately, because resolveEntryRoute routes on it. Changing that is exactly how QRS-730 shipped. Complete-vs-not decides a destination and stays exact; new-vs-returning decides only which words appear. A derivation is right for the second and would have been wrong for the first.
CR-26.0.1-121 — sign-out actually signs you out, and the next person is asked who they are
Three owner-reported defects and one request.
Sign-out now tears down locally whatever the server says, and navigates on both paths, with a (user) route guard behind it so a deep link and back navigation are covered too. scope: 'local' is a client operation: refusing the part we control because the part we do not control failed leaves somebody signed in who believes they are out.
Persona is no longer inherited. signOut clears it, a new personaDeclared flag distinguishes a tap from a reset default, and primary_context — which handle_new_user writes once — is sent only when somebody actually chose. Every signed-out launch lands on the fork.
Theme defaults to light rather than system, on a device with no stored preference. The dark pass has never been done (QRS-639), so a fresh install on a dark phone opened in the theme the app is least ready to show.
⚠ Seven tests that encoded the old behaviour were rewritten rather than deleted. Two of them asserted the sign-out bug as desirable — "does NOT clear the session or navigate" — which is why the suite was green throughout the incident. Each now says what the contract is and why the previous one was wrong.
CR-26.0.1-122 · Reserved route segments — bind the SQL seed to the QR registry
Class: schema_migration · Scope: QRS-1020 · Migration:20260904085905_v2_reserved_route_segments.sql · Reversible: yes · Contracting: no
The first migration of the re-based consumer release, and it is first for one reason: it is the only one where every day it is absent costs something that cannot be bought back. qrsetu.com/<slug> is one flat namespace shared by merchants and consumers, and a slug is write-once by trigger — so a first segment claimed by a vendor is gone permanently.
The design registry declares nine reserved first segments and its matchers consult that list to decide whether segment zero is a business slug at all. It was declared in two places with nothing binding them. Measured against the live local database at Dev head: six were reserved, three were not — o, item, c. intro and photos are added ahead of their features (decisions D-l and the deferred Photos maker) because reserving costs a row and reclaiming is impossible.
⚠ The first measurement was wrong, and that is the transferable part. A grep of the original seed migration reported eight of nine missing; the table holds 2,572 rows because later migrations added to it, and that one file seeds 753. A grep over one file reads exactly like a measurement until the answer matters.
The durable half is not the rows, it is the binding: the seed block is delimited by markers and packages/domain/src/qr/reservedSegments.test.ts parses this migration and asserts set equality with RESERVED_ROUTE_SEGMENTS in both directions, naming which side to edit. Mutation-tested three ways and proven clean on restore.
Verified on the local stack after applying: all eleven present, the five new ones carrying platform_route, is_slug_reserved('item') and is_slug_reserved('Item') both true — the column is citext, so a case-sensitive client check would disagree with the database — and is_slug_reserved('ganesh-idols-pune') still false.
CR-26.0.1-123 · media owner scope — a consumer can own a file
Class: schema_migration · Scope: QRS-1016 · QRS-1019 · Migration: 20260904090822_v2_media_owner_scope.sql · Reversible: yes · Contracting: no
media was scoped by a workspace OR a conversation, and a consumer has neither — so there was no media row a consumer could legally insert, and every consumer photo path in this release sat behind one column.
⚠ The schema was already contradicting itself, which is how it went unnoticed.users.avatar_media_id has carried a live FK to media since the v2 baseline while media_purpose_matches_scope required purpose avatar to be workspace-scoped. A user avatar was referenced by a column and forbidden by a constraint at the same time; nothing failed because nothing had tried.
Both constraints moved together. Widening the scope XOR alone applies cleanly and then fails on the first insert — worse than failing to apply, because it fails at run time in front of a user rather than at deploy time in front of us.
⚠⚠ The constraint is num_nonnulls(...) = 1, not a chained <>. Boolean inequality is exactly-one for two operands and parity for three, so a <> b <> c is true when all three are set. The obvious extension would have permitted exactly the state the constraint exists to forbid, and no ordinary test would have gone looking. There is now a pgTAP assertion for that single case.
Verified on the local stack: a 19-case mutation probe inside a rolled-back transaction, 19/19, database confirmed unchanged afterwards. Made permanent as media_owner_scope_test.sql — pgTAP is now 14 files / 531 tests, up from 13 / 510.
Not included: manage-media still does not accept an owner scope, so a consumer cannot yet upload. That is the other half of QRS-1016, and keeping it separate is what let the constraint be proven before any code depended on it.
CR-26.0.1-124 · feature_grants scope documentation — and the change that was not made
Class: schema_migration (COMMENT ON only) · Scope: QRS-1018 · Migration: 20260904101758_v2_feature_grants_scope_documentation.sql · Reversible: yes · Contracting: no · Assessment:feature_grants scope XOR
QRS-1018 was scoped into this release asking for check (num_nonnulls(seven scope columns) = 1), on the grounds that the exactly-one rule was "stated in the table comment and enforced by nothing". Both halves are false.
The rule is enforced, by feature_grants_scope_target_matches_kind — a CASE over scope_kind that is stronger than an XOR, since an XOR would accept an industry_key on a plan-scoped row. There is a third layer besides, feature_grants_live_unique_idx.
And the prescribed CHECK fails to apply. Attempted in a rolled-back transaction against the live table it returns ERROR 23514: check constraint is violated by some row: nine of fifty-five rows are scope_kind = 'platform', which correctly sets zero scope columns. workspace_member sets two. On a freshly reset database it would have applied cleanly and quietly removed both shapes.
Root cause is a sentence. The table comment said "Polymorphic scope is exactly-one-non-null FK COLUMNS", which is not the rule. A tracker row was written from that phrase, inherited its imprecision, and then presented it as a measurement.
Fixed: the table comment, a comment on scope_kind and on all seven scope columns (they had none — user_id and member_user_id are both uuid references users(id) and mean different things), and a comment on the constraint itself.
The durable half is the tests. A comment can drift again; the whole cost of QRS-1018 was that nobody had a failing test to contradict it. pgTAP now pins the constraint and both branches that make the naive reading wrong. 537 tests, up from 531.
CR-26.0.1-125 · account deletion becomes soft, and the status column starts to bite
Class: schema_migration · Scope: QRS-909 · QRS-1012 · Migration: 20260904102614_v2_account_soft_delete.sql · Reversible: yes · Contracting: no
manage-account called auth.admin.deleteUser(), which destroyed the row and with it any possibility of the "stop serving now, erase within 72 hours" window decision C3 asks for.
⚠⚠ A soft delete that nothing enforces is cosmetic, and that was the real work. users.status has held active | suspended | deleted since the v2 baseline and is read by exactly one thing — get_my_context projects it. No RLS policy, no RPC and no function tests it. Writing status = 'deleted' and stopping there gives an account that reads as deleted and carries on signing in, holding its workspace and writing rows.
Enforcement is therefore at the identity layer: manage-account sets auth.users.banned_until in the same request, so a banned principal cannot obtain or refresh a session and every downstream surface is covered by construction. A check in requireAuth was considered and rejected — one indexed lookup on every authenticated request forever, to defend a state a handful of rows will ever be in, and it would leave RLS and the RPC surface untouched anyway.
⚠ One honest gap: an already-issued access token survives until it expires (~1h). The global sign-out revokes refresh tokens and closes renewal; a live token cannot be recalled. Acceptable for a deletion the person initiated seconds ago, not acceptable for a suspension imposed on a hostile account — that case needs the requireAuth check this deliberately does not add.
⚠⚠ The slug is not released, which reverses a note in the project-state record. That note said releasing was "already true by construction". Measured: slugs_release_on_owner_loss is BEFORE INSERT OR UPDATE ON public.slugs — it releases a row that is already ownerless and does not make one ownerless. A hard delete nulls slugs.user_id via the FK and that update fires it; a soft delete leaves the row in place, so nothing fires. Releasing would be wrong anyway: state = 'released' keeps the name out of the pool because slug is the primary key, so it would strand the name while the owning account still exists and can return.
Also removed: purgeAvatarStorage, which swept a Supabase Storage bucket (profile-pictures) that no migration declares — media lives in R2 behind public.media. It swept nothing and reported success. Deleted rather than disabled, because dead code that looks like an erasure step is worse than its absence.
Verified: applied locally, deno check clean, test:ef 379 passed, pgTAP 15 files / 554 tests (up from 14 / 537) — including idempotence, and one assertion that deliberately pins that the SQL half does not ban, so a green suite is never read as proof that deletion is enforced.
CR-26.0.1-126 · the marriage-biodata field registry, in the database and in the domain
Class: schema_migration · Scope: QRS-1033 · Migration:20260904104653_v2_biodata_field_registry.sql · Reversible: yes · Contracting: no · Applied to Dev: yes, read back live
The design says one module decides visibility, and its own note adds that the screen never decides "which is why the two surfaces cannot drift". That is right, and it is an argument about two clients. The public marriage profile is read by strangers, so the tier projection has to happen before the data leaves the database: an RPC that returned every field and trusted the renderer to hide four would put dob, phoneNo, income and address on the wire for anybody holding a forwardable link. Structurally unreleasable is only true if the filter runs before transmission.
Measured from the design, not from the plan: 60 fields · 8 groups · basic 24 · released 32 · private 4 · required 9. ⚠ The plan of record said twelve required fields. Carrying that forward would have become a validation rule refusing valid profiles.
Generated, not transcribed. The SQL seed and the TS port come from one parse of the design file, so they cannot disagree at birth; fields.test.ts keeps them from diverging later.
⚠⚠ The failure mode is asymmetric, so the test does more than compare. A field wrongly marked basic publishes something private; one wrongly marked private merely hides something. Only the second is recoverable. The test therefore pins the four private fields by name — an equality check alone would pass if both copies were wrong in the same direction, and they share a parse.
Verified locally, then on Dev by reading the live database: 60 rows and every count matching. Mutation-tested two ways, including a private field promoted to basic in SQL alone, which is caught and named.
CR-26.0.1-127 · ADR-0031 — the first domain schema
Class: schema_migration · Scope: QRS-1033 · Decision:ADR-0031 · Migration:20260904112217_v2_biodata_schema.sql · Applied to Dev: yes, read back live
The owner asked for domain-based naming so that ownership is obvious from the name, with the stated goal of an enterprise-grade backend rather than a generic one. The goal was adopted; the proposed mechanism was not.
Flat audience prefixes (consumer_biodata_*) were rejected on two measured grounds. They are factually wrong on day one — a biodata is owned by a person, primary_context is "PREFERENCE ONLY — never an authorization input", so a merchant can own one for their sister with no schema change — and they land on exactly 63 bytes, where Postgres truncates silently.
What shipped instead: a schema per bounded context. biodata.fields, with anon and authenticated holding no USAGE — which is the part a prefix could never give, since REVOKE ALL ON SCHEMA is a boundary and a prefix is a label.
It costs the client nothing because RPCs stay in public and zero client code names a table (the from() ban, measured). A table in a domain schema is therefore not REST-reachable at all.
Verified on Dev: 60 rows moved, private tier still 4, public.biodata_fields gone rather than shadowed, zero leaked table grants, zero functions in the domain schema. pgTAP 16 files / 566 tests, up from 15 / 554.
CR-26.0.1-128 · the marriage-biodata record
Class: schema_migration · Scope: QRS-1033 · Migration:20260904113123_v2_biodata_subjects_and_profiles.sql · Applied to Dev: yes, read back live
biodata.subjects (the person, usually not the account holder) and biodata.profiles (the record), plus biodata.reference_seq. Typed columns for what the server acts on; jsonb bags keyed by biodata.fields.id for what it only renders — so a design round adds a registry row, not a migration.
⚠⚠ Every vocabulary was fetched from the design rather than taken from the plan, and one was wrong. life is open | discussion | concluded; the plan's prose implied in_discussion, which this migration nearly shipped as a CHECK. The client would never have sent that value, so every attempt to move a profile into discussion would have failed at the database — and nothing in this repo would have caught it. There is now an assertion pinning the rejection of in_discussion by name, so a later "correction" explains itself.
Modelling decisions worth the read: subjects.relation is free text, not the design's RELATIONS enum — that vocabulary describes the family members listed inside a biodata, and the owner-to-subject space is wider (a daughter, a nephew), so the enum would only fail on somebody's real family. There is no owner column on profiles: ownership lives on subjects and the policy resolves through the join, because a second copy of the fact that decides access is the QRS-249 class. Plan caps are entitlements in feature_grants; MAX_PEOPLE, MAX_PHOTOS and MAX_CUSTOM are constraints, because they are limits of the layout rather than of the plan.
⚠ The Feistel key is deliberately absent from the database. The sequence is here; the D6 permutation is minted in manage-biodata, because adjacent references would tell anybody holding two of them roughly how many families are on the platform. A pgTAP assertion proves no reference-minting function exists in the schema — one would mean the key came with it.
Verified on Dev: 3 tables, 24 CHECK constraints, 3 policies, the one-live-per-subject index, zero leaked grants, authenticated still without USAGE. pgTAP 17 files / 593 tests, up from 16 / 566 — against the design's own fixture rather than a minimal row, including that Devanagari content round-trips byte-for-byte.
CR-26.0.1-129 · marriage-biodata access: shares, requests, and a status nobody stores
Class: schema_migration · Scope: QRS-1045 · Migration:20260904120702_v2_biodata_shares_and_requests.sql · Applied to Dev: yes, read back live
biodata.shares (one grant of read access: either THE forwardable link or one named recipient), biodata.access_requests (somebody asking the family for more, and the family's answer), and public.resolve_biodata_share_status().
⚠⚠ The modelling decision is the whole change, and the obvious schema was wrong.biodata-core.js exports a five-value shareState() (active · basic · released · withdrawn · expired) and its demo rows each carry one, so transcribing it as a state column is the natural move. Read those rows against each other and the five decompose exactly into three orthogonal facts: kind (basic|person), tier (basic|released — the authorisation fact), and lifecycle (withdrawn_at). The one-column version fails twice:
activeandbasicare the same underlying fact, split only bykind. A stored column would need a trigger keeping a chip label in step with another column.- "withdrawn AND ALSO lapsed" becomes unrepresentable. A family who withdrew a release would silently start reading as merely lapsed on the day its ninety days ran out. The two mean opposite things to the recipient — a lapse offers a re-ask, a withdrawal is a decision — and the design flags that asymmetry itself (spec flag 7). There is a test asserting exactly that row.
⚠⚠ expired is derived at read time, and that is a safety property rather than a convenience. This release has no scheduler, so a stored flag needs a job to write it, and any day the job did not run a lapsed release would keep serving released content: a disclosure failure caused by an absent cron. The clock is the enforcement. The boundary is pinned to the second (<=, so exactly at expiry is expired) and mutation-proved — remove expires_at from a lapsed row and the answer flips expired → released, which is precisely what a missing scheduler would have produced.
⚠ The forwardable link can never be released. SHARE_TIER = 'basic' is a design constant, not a default, and the share sheet's own copy says so. Without shares_basic_kind_is_basic_tier, one mis-scoped Edge Function action would turn the one link built to survive landing with a stranger into a released one — publishing a family's photographs and phone number to everybody it had ever reached.
⚠ The token never reaches Postgres. token_hash is a 32-byte SHA-256 digest, hashed in the Edge Function, so the plaintext is never a bind parameter and cannot appear in pg_stat_statements (installed here), a log_statement line or an error detail. The length CHECK earns its place on a specific typo: encode(digest(...),'hex')::bytea yields 64 bytes, is still unique, and still looks up fine — against a column nothing else compares, so it would surface only as "the link does not work" for whichever half of the code used the other encoding.
⚠ The function is in public, and the first draft had it in biodata.domain_schema_boundary_test.sql refused it, and the function moved rather than the assertion — relaxing a gate to fit new code is how gates erode. ADR-0031's own line settles it: public holds what more than one surface depends on, and three callers share this one comparison (the recipient's page, the owner's overview, the biodata-read EF). So it is interface, not module internals. Revoked from PUBLIC and anon per ADR-0014; verified live on Dev that anon cannot execute it and that zero functions exist in biodata.
Two columns ship with no producer, stated rather than hidden. open_cities (the design's link card shows cities) needs a geo-IP source that has not been discussed or approved — a Supabase Edge Function sees x-forwarded-for and no city, while Cloudflare's cf-ipcity reaches the Worker rather than the function — so nothing writes it and the client renders the open count alone. phone_verified defaults false and has no writer at all: a request arrives from an anonymous visitor, and send-auth-otp provisions an account for every number it verifies (QRS-998). ⚠ The client copy must therefore not claim verification while it is false (QRS-1046) — an owner releasing a family's photographs on the strength of a checkmark we never earned is worse off than one shown an unverified number. Both columns exist now because adding either later is a migration against a table holding live shares.
Other decisions worth the read. Requests carry three states, not four (the design draws Release · Decline; withdrawing a release afterwards changes the share's lifecycle, never the ask's, so the request keeps its dated answer). asked is a jsonb array with no vocabulary CHECK, because the design has no REQUEST_ASKS constant — the three asks are drawn inline in the artboards — and a CHECK transcribed from an artboard is a constraint with no authority behind it. requester_user_id is nullable by construction (anonymous is the normal case, since a signup wall would sit inside the one flow the feature exists to serve), which is what makes on delete set null legal rather than aspirational — the lesson from the proposal that recommended SET NULL for a NOT NULL column. opens is a count with last_open_at, never a per-open log: a row per open would be an append-only write on a public read path and a reader-level trail of who looked at a stranger's family. Caps stay entitlements (PLAN.personLinkCap = 6, forwardable link deliberately not counted).
A recorded divergence. biodata-core.js exports a LIFT_CONTRACT proposing profile_field_visibility, profile_field_tier and profile_photo as normalised tables. biodata.profiles uses jsonb bags instead, for the reason CR-128 gives. Noted so a later reader sees the contract was read and answered rather than missed; profile.theme_id — its one scalar — was adopted as written.
Evidence. pgTAP 18 files / 640 tests PASS (up from 17 / 593; this file adds 47, every constraint asserted in both directions per QRS-013) · check:sql green over 86 migrations · check:naming green · migration list clean in both directions before the push · every object read back off Dev.
CR-26.0.1-130 · an absent account context is a decision, an invented one is a bug
Class: schema_migration · Scope: QRS-935 · Migration:20260904124512_v2_provisioning_context_is_loud.sql · Applied to Dev: yes, read back live
handle_new_user ended its CASE with else 'business', so a wrong key and an absent key produced the same row — and the function's own comment described that as a safeguard: "an unrecognised value falls back to the default instead of failing the signup." That sentence was the defect, not a description of one.
| input | before | after |
|---|---|---|
| absent / NULL | business | business — unchanged |
business · individual | as sent | as sent — unchanged |
| anything else | business, silently | raises 22023, naming the accepted values |
⚠ Why it was worth closing while the client is correct. Measured 2026-09-04: primary_context is the key at every call site and account_type has zero occurrences in apps/** or packages/**. So this was a latent trap — which is exactly what makes it worth closing before more signup paths exist. The value is written once at account creation and never recomputed, so a one-word mistake in the next path would have provisioned every consumer as a merchant with no error, no log line and no failing test: the account simply lands on /dashboard, keyed to a workspace it can never have. That is QRS-730's trap re-armed through a different door, and the population it would hit is the one this release exists for.
⚠ Raising versus logging was the real choice, and the asymmetry decided it. A raise warning would not fail the signup, which sounds kinder and is wrong: nobody reads a warning for an event that has never happened, and the account is already wrong by the time anyone could. A failed signup is loud, immediate and retryable. A silently mis-provisioned account is invisible at creation and permanent afterwards.
⚠ The absent case stays a default deliberately. An OAuth signup has no fork to declare, so "the client said nothing" is a legitimate state and raising there would break Google sign-in entirely. That is also precisely why this does not close QRS-917: the email path omits the key, which arrives as NULL, and no database change can distinguish that from OAuth. Only the client can send what it knows. provisioning_context_test.sql §D asserts that limit, so no later reader concludes mis-provisioning is solved.
⚠ The expand-contract obligation is enforced, not created. users.primary_context already carries a CHECK, so an unrecognised value could never be stored — the old code merely rewrote it on the way in. Widen the accepted set before shipping a client that sends a new value; skipping that step now fails in development instead of mis-provisioning every account the new build creates.
⚠ A defect in the diagnostic itself, caught by the probe. The first draft wrote %L, borrowing format()'s literal-quoting spec, which raise does not have — so it substituted on the % and left the L, reading primary_context consumerL is not a recognised account context. Found only because the probe printed the message rather than checking the SQLSTATE alone. The test now asserts the text, not just the code: a diagnostic nobody reads back is a diagnostic that lies.
Evidence. pgTAP 19 files / 652 tests PASS (up from 18/640). The new provisioning_context_test.sql covers a function that every other test file relied on and none asserted — they all treat it as a fixture side effect. Mutation-proved locally: restore the old body and "consumer" is silently provisioned as business again. Read back off Dev — the swallow is gone, the NULL fallback kept, anon and authenticated still cannot execute.
CR-26.0.1-131 · the recipient's side, and a policy that must not exist
Class: schema_migration · Scope: QRS-1047 · Migration:20260904131044_v2_biodata_kept_and_reports.sql · Applied to Dev: yes, read back live
biodata.kept (the recipient's own shelf) and biodata.reports (append-only abuse reports). These are the only two tables in the schema whose rows do not belong to the profile's owner, and each is modelled wrong in a different direction if written by analogy with shares.
⚠ A report has a target, and it is not always the profile. The obvious shape is (profile_id, reason). The design disproves it in one line: it exports openingReportUrl(slug) → /report/<slug>/opening, a report against one emblem, beside the note that "QR setu can remove a reported picture without removing the profile." A table that could only name the profile would turn every picture complaint into a complaint about a family. So target is profile | opening_emblem | photo, and target_ref is a jsonb key, not an FK — an emblem and a photograph live in profiles.opening / profiles.photos bags rather than tables (CR-128), so there is nothing to reference. The constraint is an equivalence, so both failure directions are refused.
⚠ There is no reason vocabulary, measured. biodata-core.js ships no report-reason constant at all; the five reasons the plan mentions belong to ScanVerify, a different surface with a different model. So reason carries no CHECK, exactly as access_requests.asked does — a vocabulary transcribed from an artboard is a constraint with no authority behind it, and here there is not even an artboard list to transcribe.
⚠ Append-only, and the moderation policy is deliberately absent. No state, no outcome, no reviewed_at. The design's own round-8 flag records what is undecided: "whether review is pre-publication or post-report, who reviews, and what happens to the profile when an emblem is removed… that is a policy call." A report is a fact that never changes; an action taken is a different fact with its own actor and time. Half-specifying the second here would invent the policy that flag is waiting on (QRS-803). A test asserts that neither column exists.
⚠⚠ reports has zero policies, unlike every other table in this schema, and the absence is the control. The design is explicit that the family is "not told who" reported them — and an owner-readable row leaks the reporter by correlation: a family holding one released share and seeing one report knows exactly who it was. A reporter-scoped policy would be worse, letting one person read other people's reports about the same family. RLS on, revoked, deny-all outside a definer function. Verified on Dev: all seven biodata tables give anon and authenticated no SELECT.
⚠ kept.user_id is the recipient, not the owner — the one table here where the family has no claim on the row, so a policy joining through subjects (which every other table correctly does) would have been exactly backwards. user_id is NOT NULL, so keeping requires an account while reading does not: that is the right place for this surface's only registration prompt, whereas a wall in front of the profile itself would destroy the forwarding mechanic the feature runs on. share_id is provenance only and on delete set null, so a shelf entry outlives its share — a kept profile whose access ended must still render, because that is what tells the reader the link was pulled.
biodata.templates was deliberately not built
The plan lists it as the D7 sibling of setu_card_templates. Strip from that shape the things a biodata does not have — archetype_keys (a biodata is a personal record, so the archetype axis does not apply), a seasonal availability window, a feature_code entitlement gate — and what remains is a registry holding exactly one row, default/1, with no second candidate anywhere in the design, whose variation axis is its 12 themes: a column that already exists. Its one genuinely valuable part, the trigger refusing to retire a version a live profile still pins, cannot fire while there is one template. And the FK it would add needs an i18n display_name_key and a manifest-validating gate that the plan itself sequences into P7.
Tracked as QRS-1048 against P7's existing acceptance line, which already reads "check:setu-card-templates covers the biodata manifest". ⚠ Accepted cost, stated rather than left to be discovered: profiles.template_key / template_version are unvalidated until then, so a typo is storable and would fail at render.
Evidence. pgTAP 20 files / 672 tests PASS (up from 19/652) · check:sql green over 88 migrations · every object read back off Dev.
CR-26.0.1-132 · notification read state, keyed by the person and not by the persona
Class: schema_migration · Scope: QRS-1050 · Migration:20260904134022_v2_notification_reads.sql · Applied to Dev: yes, read back live
The notification list stays derived and is never stored. This table holds the one fact derivation cannot produce, because it is about a person rather than about the activity.
⚠ It is not consumer_notification_reads, which is what the plan calls it. The plan predates ADR-0031, and re-deriving the name under that ADR gives a different answer: the axis is the bounded context, never the audience. Two reasons, the second decisive:
- There are already two notification surfaces in the client —
tiers/user/features/notificationsandtiers/consumer/features/notifications— and they are two readers of one idea. Read state is "who, and which", and neither half of that is a persona. - ⚠⚠ A person can be both. This platform's own three-category rule says a consumer who later starts a business gains a workspace membership, never a second account and never a data migration, and that it must be modelled so that is free. Split read state by audience and that person immediately has two read-state tables for one pair of eyes: notifications they dismissed as a consumer come back when they open the merchant surface. An audience-keyed table makes the platform's central promise cost a migration.
notification_key is a synthetic id, not an FK (there is no notifications table, deliberately), and carries no prefix CHECK because the kind vocabulary grows per feature — a CHECK would make adding a notification kind a migration. No kind column: the design's controls are per-item and global, never per-kind.
⚠ Growth is unbounded and there is no reaper, stated rather than discovered. Accepted: narrow rows, of the order of hundreds per user per year. A watermark was considered and rejected — the design clears notifications individually, so "read the third but not the second" (the ordinary case, since each links to its own object) would be unrepresentable, and a hybrid buys space at the cost of two sources of truth for one question.
CR-26.0.1-133 · the rate-limit bucket, and why a hashed phone number is not anonymous
Class: schema_migration · Scope: QRS-921 · Migration:20260904140833_v2_rate_limits.sql · Applied to Dev: yes, read back live
⚠ This does not close QRS-921. That row says there is NO rate limiting on any endpoint, and that is still true: no endpoint consumes this yet. The mechanism exists; the enforcement is _shared/rateLimit.ts and its call sites. Conflating the two would be the presence-versus-enforcement error this repo keeps paying for.
⚠⚠ A plain SHA-256 of a phone number is not anonymisation, and that is the security content of this migration. Reusing the share-token discipline — hash the subject, store the digest — is the right instinct and an insufficient one, because the entropy differs by orders of magnitude. A share token is 128 random bits and its digest is irreversible in practice. An E.164 Indian mobile number is one of about 10⁹ possibilities and an IPv4 address one of 2³², both exhaustively enumerable in seconds on a laptop. A plain digest would make a copy of this table yield the phone number of every person who has ever requested a code. So subject_hash is an HMAC under a secret pepper (RATE_LIMIT_SUBJECT_KEY, an Edge Function secret), and the key must never reach the database — the same rule as the D6 reference permutation, because a pepper stored beside the data it peppers is decoration.
Two independent limits are the caller's composition, not the schema's. QRS-921 requires per-phone and per-source limits on OTP send, because per-phone alone lets one attacker walk a list of numbers and per-source alone lets a botnet flood one number.
The fail posture is deliberately not encoded. The plan's R10 needs three postures on one mechanism — an anonymous read fails closed, an authenticated owner action fails open, a timeout allows once and logs — so the function only counts and reports. One that decided as well as counted would force one posture on all three.
The increment is one SQL statement, because a read-then-write limiter is a race a flood wins: two concurrent requests both read a count one below the limit and both proceed. A denied attempt still counts, so pausing does not re-admit a flooder. Fixed window, not sliding — up to 2× the limit across a boundary, accepted, since sliding needs per-attempt rows (write amplification on the exact path being flooded) or a compaction job that does not exist.
⚠⚠ check:sql caught the table revokes missing from the first draft, and the near-miss is worth recording. Supabase's default privileges grant anon ALL privileges on every new table in public, so without them anon could have read every bucket and, worse, UPDATEd one — zeroing their own count makes the limiter a no-op for exactly the caller it exists to stop. It was also missed on the first read of the gate's own output, by piping it through tail: the visible line was benign and the failure was above it. That is the trap CLAUDE.md documents by name, and it caught me on a gate I had just run.
Evidence. 15 behavioural probes green — the ladder, denials counting, per-subject and per-scope independence, window alignment, an expired bucket never resurrected, 200 consecutive consumes counted exactly — plus rate_limits_test.sql, and anon/authenticated verified to hold nothing on the target.
CR-26.0.1-134 · a shop or not, without reopening an oracle somebody deliberately closed
Class: schema_migration · Scope: QRS-1052 · Migration:20260904143507_v2_resolve_slug_owner_kind.sql · Applied to Dev: yes, proven on live rows
⚠⚠ The plan's specification for this function would have reopened a deliberately closed oracle, and that is the whole design content of this change. The plan asks for workspace | user | null. resolve_slug_status collapses a consumer-held address into reserved on purpose, and its own comments say why: taken is for MERCHANT addresses only, and that asymmetry is D5's, not an oversight: a business address is meant to be found… Indistinguishability is a consumer-only requirement, and THE ORACLE-CLOSING BRANCH… A caller therefore cannot distinguish "the platform holds this word" from "a person exists at this name".
A sibling answering user hands back exactly that, so the two functions would disagree about whether consumer existence is disclosable and an attacker would simply ask whichever one answers. The plan was right to collapse reserved / released / never-claimed, and wrong to make the person a distinguishable fourth value.
Shipped: workspace for an active workspace-owned slug, other for everything else — a person, a reserved word, a released row, or nothing at all, in one bucket. Routing on other costs nothing, because the greeting page is byte-identical for a claimed, an unclaimed and a private address (decision D-t/D5): opening a greeting for a name nobody holds is the mechanism, not a defect.
⚠ It is also not person, which is what the design's client tests — and the plan's user was wrong about the value name too, the third plan constant corrected against the design this session. qr-registry.js matches ownerKind(slug) === 'person'. The client needs a one-line change, and that change is the point: person is the value that would disclose person-existence, while other is informationally equivalent for every routing decision the client actually makes.
⚠⚠ The polarity must not be inverted when that change lands. The design states the guard on both sides and explains why: the identity entry above returns false while the lookup is unavailable, so the guard has to be stated on both sides. Keep the positive test on the identity branch and the negative one on the card, so a null — a failed lookup, an offline device — still resolves to the shop. Inverting it is strictly worse than today's defect: a person resolving as a shop is an annoyance, while a shop resolving as a personal greeting breaks the most common code in the wild: a sticker on a counter, a poster, a printed bill.
Evidence, on live Dev rows rather than a fixture. chai (a real merchant) → workspace, and case-insensitively; meera (a real person's address), admin (reserved) and an unclaimed name are all other. The pair is consistent: both functions disclose the merchant, neither discloses the person. slug_registry_test.sql §H.
CR-26.0.1-135 · meetings, written twice
Class: schema_migration · Scope: QRS-1053 · Migration:20260904144500_v2_meetings_schema.sql · Applied to Dev: yes, read back live
Five tables (meetings · meeting_occurrences · meeting_participants · meeting_invite_states · meeting_joins), the resolution rule as a function, and one existing function replaced by a delegate. The participant half only: the design is explicit that receiving a meeting "needs no provider account, ever. This is most people's entire relationship with the feature."
The one part with reach beyond meetings
⚠⚠ is_valid_reminder_recurrence now delegates to a neutrally-named public.is_valid_recurrence, under a live CHECK on reminders.recurrence. Copying the validator for meetings would have been the duplicate-source-of-truth bug class applied to a scheduling rule — and a drift between two copies does not raise an error, it fires something on the wrong weekday. Signature unchanged, so the CHECK is untouched; expand-contract, so the old name can be retired later.
Agreement is asserted across a 21-case battery, and the battery is asserted to contain both verdicts — "no disagreements" over an all-true battery would be vacuous. Verified on Dev that the delegate still refuses {"freq":"daily","intervall":3} and accepts the weekdays mapping.
The tables are in public, and ADR-0031 said meetings
That row's premise — "one surface owns them" — is false. Measured: the design has a merchantMeetings screen (round 35) whose mobile-console/meetings-core.js the consumer module imports, and it says why — "so a merchant's session and a personal meeting can never disagree about what 'today' or 'live now' means." Two surfaces sharing one module is verbatim the reason the same ADR table keeps conversations in public.
ADR-0031 is amended (section A1) rather than deviated from, since it states its context vocabulary changes only by amendment. biodata stands, D1 is unchanged, and D8's Edge Function context prefix is unaffected.
Written twice, and three of the four corrections would have been silent
| # | The first draft | The design |
|---|---|---|
| 1 | mode in ('online','in_person') | MODES gives the id offline with the label "In person". Every in-person meeting the client could send would have been refused — the same shape as life = 'in_discussion'. |
| 2 | Carry-forward: the latest answer at or before an occurrence | A series default plus an exact-match override. Under carry-forward, declining Tuesday declines Wednesday and every morning after, so "decline a single morning" becomes unrepresentable. |
| 3 | A comment asserting the design had no reschedule | Contract 1: an entry is "cancelled, or moved to another date or time", and occurrencesFor implements it. |
| 4 | Three recur kinds | Four. weekdays exists because otherwise "a five day batch has to be stored as daily, which claims a Sunday class that does not happen." All four map onto the existing shared shape with no additions. |
⚠ Correction 2 is the one worth dwelling on, because the probe passed while encoding the wrong behaviour. The assertion read "the days AFTER inherit the decline" and went green — it was the expectation that was wrong, and it only became visible because the assertion printed its consequence in words. A test that states what it believes in prose is what turns a passing run into a readable claim.
Other decisions
Server-side reminder preferences (plan R6). The prototype keeps reminderOff and reminderMins in localStorage, so somebody who silences a 6.30am batch on their phone is woken by their tablet the next morning. A preference about being woken up belongs to the person, not the device. Local notifications are still scheduled on the device — there is no push sender — so the column is what every device reads to agree.
Deliberately absent. No producer (decision C1): nothing creates a meeting this release. No group_id — a meeting may target a batch, but groups are a merchant concept (capacity, a hall, enrolment) arriving with hosting. No marked_attended: contract 3 requires it be a separate, differently-labelled fact from a join tap — "a row in joins means the person tapped Join inside QR setu, which is not the same as having been in the room" — so folding it into meeting_joins is precisely what that contract forbids.
⚠ A plan-versus-design conflict is recorded, not resolved. Contract 2: "there is no meeting SDK in any phase. Joining is always a deep link out to the meeting app or browser." The plan's P9 proposes a Zoom Meeting SDK inside a WebView. This migration takes neither side — it stores a link and nothing about how it opens — and QRS-1014 owns the decision.
One process note
⚠ The file was renamed before pushing. Its first timestamp sorted before an already-applied migration and db push refused it. Renamed rather than forced through with --include-all, so the repo's migration order keeps matching the application order — which is what makes migration list meaningful in both directions, and the absence of which is the QRS-267 orphan-version hazard.
Evidence. pgTAP 22 files / 757 tests PASS (up from 21/709; meetings_test.sql adds 48 and asserts each of the four corrections by name) · 27 behavioural probes green · check:sql green over 92 migrations · every object read back off Dev.
CR-26.0.1-136 · the cascade the soft delete was missing
Class: schema_migration · Scope: QRS-1056 · Migration:20260904150210_v2_soft_delete_cascade.sql · Applied to Dev: yes, read back live
⚠⚠ soft_delete_account set users.status = 'deleted', cleared the display name and avatar, and stopped. manage-account then banned the identity and signed the person out. None of it touched biodata.
Because a soft delete leaves the public.users row in place, nothing cascaded. Every profile stayed published, every share token stayed live, and the public biodata page kept serving after the account was deleted. The person could not sign in — while their sister's photographs, community details and the family's phone number carried on being served to anyone holding a link.
How it happened, and the ordering is the lesson. soft_delete_account shipped at 10:26 today (CR-127), biodata.profiles at 11:31 (CR-128), biodata.shares at 12:07 (CR-129). The cascade could not have covered tables that did not exist yet, and nothing re-examined it when they arrived. P1 is sequenced irreversibles-first for good reasons; this is its cost.
⚠ Not a live exposure — no biodata row exists on any environment and the consumer feature is unshipped. A latent defect, which is why it was worth closing now rather than discovering it from a deletion request.
⚠ Found by writing the erasure runbook, not by any gate. Describing what deletion does forces you to read what it does — which is the argument for the runbook existing at all.
Fixed: withdraw every share, then retire every profile. That order is deliberate — a partial failure must leave the content unreachable rather than leave the keys live, and the page closes either way, so nobody would notice the reverse order. retired rather than unpublished because profiles_one_live_per_subject is partial where status <> 'retired', so only retired frees the subject — and a soft delete is reversible inside the 72-hour window. The slug stays untouched for the original reason: the account still exists, so a returning person keeps their address.
⚠ Media is still not erased and cannot be from SQL — the bytes are in R2 and there is no drain (QRS-1012). That half is the runbook.
Evidence. Proven locally (published → retired, 2 live tokens → 0), asserted in account_soft_delete_test.sql §Z in both directions including that retiring frees the subject, and read back off Dev.
CR-26.0.1-137 · the switch that had nowhere to be stored
Class: schema_migration · Scope: QRS-1059 · Migration:20260904152500_v2_biodata_field_hidden.sql · Applied to Dev: yes, read back live
The design has two per-field controls, and says so in its own words above canHide():
"A SECOND AND SIMPLER QUESTION THAN THE TIER. The tier says WHO may read a field; this says whether the field appears AT ALL."
Its LIFT_CONTRACT names them as two stores — profile_field_visibility (…, hidden) and profile_field_tier (…, tier). CR-128 shipped only field_tiers. So the field sheet's visibility switch — a designed control with its own section summary, hiddenIn() — had nowhere to be stored. No client work could have produced it, and nobody reading a screen could have seen it was absent. That is the fourth rule's exact shape: a missing state in the contract, not a missing branch in a component.
Found in P2, not P1. Porting biodata-core.js meant reading its lift contract line by line against the schema. The P1 review did not catch it and no gate can — a column that does not exist has nothing to assert on.
⚠⚠ Why it could not be folded into field_tiers, which is the tempting fix. Expressing "hidden" as field_tiers[id] = 'private' looks equivalent — a private field is public at no tier. But the design gates the switch on the registry tier: canHide(f) = !f.required && f.tier !== 'private'. Hide a basic field that way and canHide answers false afterwards, so the switch disappears and the field can never be unhidden. A one-way door. It also conflates two different statements a family makes — private is "held so you can use it inside QR setu", hidden is "not on the page" — and the design is explicit that "hiding is not the same as deleting, so a hidden field keeps its value and keeps counting as answered."
Shape: a jsonb object keyed by field id, mirroring field_tiers. Deliberately not the contract's two side tables, by the design's own photo argument ("one rule instead of two"): small, always read whole, written by the owner in one save. The combined bag CHECK is replaced to cover it rather than joined by a second constraint — one rule with two spellings is the QRS-706 class — which is safe only because the table has zero rows and this is the same change that adds the column.
Evidence. Mutation-tested locally (scalar and array refused, object accepted, people still guarded after the replace); biodata_record_test.sql §G2 asserts column, type, not-null, default and both directions; read back on Dev; check:sql green at 94 migrations.
CR-26.0.1-138 · the owner's three biodata reads
Class: schema_migration · Scope: QRS-1075 · Migration:20260905083000_v2_biodata_owner_reads.sql · Applied to Dev: yes, 2026-09-05, read back
get_my_biodata_overview · get_my_biodata · get_my_biodata_shares, plus one shared predicate, is_biodata_value_answered. No table is touched and nothing is written.
Three functions, because there are three surfaces with three lifetimes. The overview is read on every app open and has to be one round trip, since plan finding R34 measured that get_my_context carries no slug, no features and no biodata; the record is read when the editor opens; the share list only on the people-and-access view. One combined read would make Home fetch a share list it never renders, and would carry that list on every save.
It derives no completion percentage, and that is the point. biodataJourney() in @qrsetu/domain weights nine sections, resolves the required gaps and picks the next one to fill. Re-expressing that in SQL would put the product's completion model in two languages, and they would disagree the first time a section weight changed. The overview returns the content bags and the service derives from the same module the editor renders with: one round trip, one definition, and it obeys the standing rule that services return stored rows while derivation lives in domain.
⚠ p_expiry_warn_days is required and has no default, for the same reason. BIODATA_EXPIRY_WARN_DAYS = 10 already drives the design's "an expiry approaching" chip. A literal here would be a second copy of a product constant whose drift is silent — move the design to 14 and the chip moves while the count does not, so the overview reports two releases lapsing beside a list showing three. A default would hide exactly that bug behind a caller who forgot to pass it.
Scoping. Every function joins through biodata.subjects.owner_user_id, the only place ownership lives. A profile that is not the caller's answers null or [] rather than raising, so not-found and not-yours are indistinguishable — otherwise any authenticated caller could learn whether a given uuid is a real biodata.
Evidence. A full local supabase db reset replayed all 95 migrations clean, this one included; check:sql green; check:rpc rule R4 fired the moment the functions landed, and three forward-declared allowances were retired in the same change.
CR-26.0.1-139 · the tier projection, in the database
Class: schema_migration · Scope: QRS-1075 · Migration:20260905091500_v2_get_public_biodata.sql · Applied to Dev: yes, 2026-09-05, read back
get_public_biodata, plus effective_biodata_tier and biodata_relation_field. Granted to nobody: the only caller is the biodata-read Edge Function as service_role, which is what makes the rate limit unbypassable — a client that could reach this directly could read a family's page as fast as it liked.
Why the projection is here and not in the renderer. biodata.fields was built for this and its own table comment says so. The design is right that one module should decide visibility, but that argument is about two clients agreeing with each other. The public page is read by strangers, so an RPC returning every field and trusting the page to hide four would put dob and phoneNo on the wire, where a renderer mistake is a disclosure rather than a layout bug.
⚠ The private tier is guarded twice, and that is measured rather than asserted. A 41-assertion fixture plus a ten-case mutation suite showed that removing f.tier <> 'private' alone leaks nothing (the viewer-set test catches it) and removing the viewer-set test alone leaks nothing (the explicit test catches it). dob and phoneNo reach a stranger only when both are removed, which the double mutation confirmed — so the guarantee is falsifiable rather than merely untested, and the apparently redundant test must not be tidied away. A third guard is load-bearing on its own: effective_biodata_tier treats the registry tier as a floor, so an owner-supplied {"income":"basic"} cannot free a private field.
The other properties the mutation suite proved are real: a token is bound to its own profile (without that predicate a valid token for one family reads another family's page at whatever slug it is presented against — the bearer-token confused deputy, one line to close); a forwardable link gets the cover photograph and nothing else; a hidden field is suppressed, but hidden cannot suppress a required one; removal outranks every other state and serves no first name; a withdrawn reader is told withdrawn rather than concluded, because somebody is owed the reason they lost access; and an unclaimed name, a business address and an unpublished profile all answer null, so none of them can be told apart.
It takes the token's digest, never the token — a plaintext statement parameter reaches pg_stat_statements, the slow-query log and every EXPLAIN somebody later pastes into a ticket. It records nothing and mints no URL: the open count and the derivative presign belong to biodata-read (plan finding R1 — a public read that wrote could not be stable, and would hand an anonymous visitor a write).
Evidence. Applied on the local stack; 41 of 41 assertions green with psql exit 0; the mutation suite's output read in full; the fixture rolled back and the local database verified unchanged afterwards. check:sql, check:naming, check:docs and check:rpc all green.
CR-26.0.1-140 · the only door into the biodata schema
Class: schema_migration · Scope: QRS-1075 · Migration:20260905141000_v2_biodata_write_api.sql · Applied to Dev: yes, 2026-09-05, read back
Eight functions in public, granted to service_role alone.
Why it exists at all. manage-biodata cannot write these tables. Edge Functions reach Postgres through PostgREST, config.toml exposes schemas = ["public", "graphql_public"], and the ADR-0031 migration already stated the consequence in its own header: "a table in this schema is NOT REST-REACHABLE AT ALL. Defence in depth, for free." That was written as a property acquired rather than designed — and this is the bill for it, paid gladly.
The result is better than table access would have been. The write surface of an entire bounded context is now a short, named, enumerable list with each transition's precondition in one place, instead of whatever .update() the next Edge Function happens to call.
⚠⚠ These functions take a user id and trust it, because auth.uid() is NULL under the service role. That is safe only because service_role is their sole grantee. Widening any of these grants would be the most expensive one-line change in this schema.
The decisions inside, and each one has a failure behind it:
update_biodata_profiletests ownership and version in the UPDATE's ownWHEREclause rather than reading first and checking in plpgsql. A family editing on two phones is this feature's normal case, not a rare race.- It raises two distinguishable errors:
40001(stale version → reload or overwrite) andP0002(missing or not yours → a reload could never succeed). Collapsing them would make the client offer an action that cannot work. publish_biodata_profilesetspublished_atonce and never resets it — it is "live since", which the public page prints, so an unpublish-then-republish must not make a profile look newly created to a family who has been reading it for a month. The reference is minted only at first publication for a sharper version of the same reason: it is what a family quotes to a bureau, so a second one strands every conversation that used the first. It returnsfirst_publicationso the subject notice is not sent twice, and it refuses while a removal request is outstanding.conclude_biodata_profilewithdraws nothing. Concluding is an announcement, and somebody holding a valid link is owed the concluded page rather than a withdrawal they were never given. Reopening is what withdraws every release, because the reopen copy promises a search that starts at zero.reopen_biodata_profileretires before it inserts.profiles_one_live_per_subjectis a partial unique index, so inserting first fails on an index name rather than on the rule — the kind of error that sends the next person rewriting something that was correct.
Evidence. Applied on the local stack; a 21-assertion state-machine probe green at psql exit 0, covering create, a stale version, a foreign owner, a republish keeping both the reference and published_at, a removal blocking publication, conclude not withdrawing, reopen retiring and withdrawing and carrying content into a new draft, one-live-per-subject still holding afterwards, a retired profile refusing edits, and a draft refusing to conclude. Fixture rolled back.
CR-26.0.1-141 · the biodata feature and its two caps
Class: schema_migration · Scope: QRS-1075 · Migration:20260905150000_v2_biodata_feature_registry.sql · Applied to Dev: yes, 2026-09-05, read back
A gap found in P4, not P1. The plan's P1 list included "rows for consumer.biodata (subject cap 1, person-link cap 6)" and they were never written — measured as zero rows in public.features and zero in public.feature_grants matching biodata.
⚠ It would have failed closed and looked like an Edge Function bug. assertFeature refuses when the resolver does not return a feature at all — deny-by-default, correctly — so the first family to tap make a marriage profile would have got a 403 naming a feature key that exists nowhere. The failure would have pointed at manage-biodata, which would have been innocent.
⚠ The key is biodata, not consumer.biodata, and that is a deliberate correction to the plan. Two independent reasons:
- ADR-0031 already settled that the axis is the bounded context, never the audience.
consumer_biodata_*was proposed for the schema, evaluated, and rejected as factually wrong: a biodata is owned by a person,primary_contextis "PREFERENCE ONLY — never an authorization input", so a merchant can own one for their sister with no schema change, and the design already contains a marriage bureau managing them for families. A feature key naming the audience would reintroduce exactly the error the schema avoided. - A dot in this table means parent and child, not a namespace.
store.mediaandstore.stockare children ofstoreviaparent_key.consumer.biodatawould imply aconsumerfeature that does not exist and should not — being a consumer is not a capability anybody grants.
Two keys, because there are two caps and a row carries one. feature_grants holds one limit_value per (feature, axis), so biodata (one subject) and biodata.shares (six person links) are a parent/child pair. The split is meaningful rather than mechanical: one governs how many people you may make a profile for, the other how many families may read one. The forwardable basic link is not counted against it, and manage-biodata applies that rule — the grant carries only the number.
applicability = 'universal' because applicability is derived from a workspace's primitive composition and a consumer has no workspace, so a scoped feature could never apply to the audience it exists for. risk_tier = 'high', alongside orders and payments: this holds a third party's personal data under a removal right the owner cannot refuse.
Evidence. Applied locally, then resolve_features(<consumer uuid>, null) returned biodata enabled with limit 1 and biodata.shares enabled with limit 6 — entitlement_source=grant:platform, applicability=universal, availability=default:active.
CR-26.0.1-142 · manage-biodata, the record write path
Class: ef_code · Scope: QRS-1075 · Deployed: yes, 2026-09-05, verify_jwt read back · config.toml: verify_jwt = true, declared in the same change
Six actions — create, update, publish, unpublish, conclude, reopen. Sharing and access are the second slice and are refused by name, with a sentence, so a client written against the full seam gets an explanation rather than a mystery.
⚠ The entry in config.toml is written in the same change on purpose. check:fn-config still requires --project <ref> and exits with usage without one, so it runs nowhere (QRS-643), and manage-reminder is the standing lesson: an absent entry is a silent default, and that function turned out to be deployed verify_jwt = false by accident rather than by decision.
The validation is the point of the function. Content lives in jsonb bags keyed by biodata.fields.id, which is what lets a design round add a field as a row rather than a migration. The price is that the database cannot check the contents, so this is where an unregistered key is refused: a field nobody registered has no tier, and a value with no tier cannot be filtered — it would escape the disclosure model entirely.
The write-side disclosure rules: a per-field tier override may only ever tighten; a required field cannot be hidden; a private field cannot be hidden because it is already on no page. The first of those is not made redundant by get_public_biodata flooring a loosening override — that stops it being served, while storing it would leave the row asserting something false about the family's decision, and the next reader written might believe it.
Unknown patch keys are refused rather than dropped. Silently ignoring a key the client believed it had saved is how a family retypes the same paragraph three times and blames the app.
The reference (decision D6) is a four-round Feistel permutation of biodata.reference_seq, keyed by BIODATA_REFERENCE_KEY with no default — a default key is a published key, and anybody reading the repository could then invert it and recover the counter the permutation exists to hide. A permutation rather than a random string means collisions are impossible by construction, so there is no retry loop under a unique index to misbehave at the worst possible moment; a gap in the sequence is normal and tells nobody anything.
⚠ The D3 subject notice reports failed today, honestly, for two reasons. The utility template does not exist — whatsapp_message_templates holds qrsetu_otp only, and the submission is owner-side (QRS-1013, launch critical path). And the notice must carry a removal link to be the D3 notice at all: it is what makes the subject's removal right real rather than declared, and the signed removal token is slice two's work. A notice telling somebody a profile about them exists, with no way to act on it, is worse than none — because it looks like the platform informed them. not_applicable is returned when nobody is owed one (the owner is the subject, no number was given, or it is a republish).
Evidence. deno check clean; 21 Deno tests green, with every refusal paired with an acceptance of the same input minus the offending part (QRS-013) — including a bijection test over 4,000 sequence values proving the permutation cannot collide. check:readmes and check:naming green.
CR-26.0.1-143 · the biodata sharing and access API
Class: schema_migration · Scope: QRS-1075 · Migration:20260905170000_v2_biodata_share_api.sql · Applied to Dev: yes, 2026-09-05, read back
Eleven functions, service_role only, same posture as CR-140.
The decisions, each with the failure it prevents. Minting always creates at the basic tier, because share and release are two decisions the design draws separately. Asking for the forwardable link twice returns the existing row rather than a unique violation naming an index — opening the share sheet twice is not a mistake. Release refuses the forwardable link outright, and refuses a withdrawn share, because quietly restoring access would make a withdrawal reversible without the family saying so again. Extend counts from now, not from the old expiry, or extending a lapsed release produces a share that is both extended and still expired. Withdraw is idempotent and keeps the first timestamp, because the moment access ended is the fact a recipient is owed.
⚠⚠ The sharpest one: answering a request with a release mints a NEW person share and never promotes the share the asker arrived through. They arrived on the forwardable link, so promoting it would release the entire basic audience at once — the single worst outcome this feature can produce, and one line away from happening.
⚠ request_biodata_removal takes no owner and checks no ownership, by design. The subject has no account and must not need one, so the Edge Function's signed single-purpose token is the entire authorisation — which is exactly why service_role is the sole grantee. It stops service immediately (decision C3) and withdraws every share; the 72 hours count to erasure, not to a window during which the profile keeps serving.
Evidence. Applied locally; a 22-assertion probe green at psql exit 0, covering idempotent minting, cross-owner refusal, the forwardable link refusing release, extend-from-now, withdraw idempotency and slot release, the new-share-not-promotion rule, double-answer refusal, the anonymous request path, a guessed digest learning only a refusal, keep/unkeep idempotency, anonymous report acceptance, and removal unpublishing plus blocking a re-publish. It also caught a real constraint doing its job (shares_expiry_follows_release) — a fixture error, not a code one.
CR-26.0.1-144 · what biodata-read needs beside the projection
Class: schema_migration · Scope: QRS-1075 · Migration:20260905183000_v2_biodata_read_support.sql · Applied to Dev: yes, 2026-09-05, read back
record_biodata_open and get_biodata_photo_keys — the other half of plan finding R1's split. get_public_biodata is stable, records nothing and mints no URL, because a public read that wrote could not be cached and would hand an anonymous visitor a write path.
An open counts against the share it came through, or against the forwardable link when the page was reached without a token, because the tokenless address is that link. ⚠ An unknown digest is a no-op, never an error, or a guessed digest would learn from the difference between the two outcomes. A count and a last time, never a per-open log: a row per open is an append-only write whose volume an anonymous visitor controls.
⚠ get_biodata_photo_keys resolves ids the projection has already authorised, and must never be asked to authorise anything itself.
CR-26.0.1-145 · biodata-read
Class: ef_code · Scope: QRS-1075 · Deployed: yes, 2026-09-05, verify_jwt read back · config.toml: verify_jwt = true, declared in the same change
Four steps, and the order is the design: rate limit (failing closed) → read → record the open → presign.
⚠ Two rate budgets, deliberately. Per-token stops one recipient hammering one family's page; per-source stops one visitor walking many families. Either alone leaves the other attack wide open. The source is the first x-forwarded-for hop — taking the last would key every visitor behind one proxy to a single bucket, turning a per-source limit into a global outage the first time a mobile carrier NATs its users — and it is hashed before it is stored, because a rate-limit table is not a place personal data should accumulate. The limiter failing is a refusal, never an allow: a limit you can remove by inducing an error is not a limit.
⚠ A basic reader is served the derivative or nothing. Falling back to the original when no derivative exists would quietly hand a forwardable link the full-resolution photograph — exactly what the derivative was introduced to prevent (C4). The fallback that looks helpful is the bug.
Failure behaviour is graded rather than uniform: a missed open count never fails the read (the visitor is owed the page), and a presign failure hides one photograph, not the page — returning url: null rather than a guessed origin, because a fabricated base gives a torn-page icon that a reader blames on the family's photograph rather than on the platform.
Evidence. deno check clean; 5 Deno tests over the two rules that have real failure modes behind them; test:ef 414/414.
CR-26.0.1-146 · a conversation is between principals
Class: schema_migration · Scope: QRS-1077 · Migration:20260905193000_v2_principal_conversations.sql · Applied to Dev: yes, 2026-09-05, read back
The schema half of ADR-0032. New conversation_participants; conversations drops workspace_id, consumer_user_id and the two blocked_by_* columns; messages.sender_kind and conversation_reports.reported_by widen from the audience vocabulary to the principal one; five policies rewritten.
⚠ The two-role assumption was a vocabulary, not a column pair. It ran through the participant columns, the blocking columns, messages.sender_kind, conversation_reports.reported_by and five policies — so scoping this from the participant columns alone would have under-counted it by four objects.
The original reasoning is preserved, not reversed. The table comment said a user-to-user model breaks the moment a merchant hires a second person. That is right, and a principal model keeps it: the workspace is the participant, while messages.sender_user_id still records who typed.
⚠ The one invariant that had to survive is unique (workspace_id, consumer_user_id) — what made tapping Message twice continue one thread. It becomes conversations.dyad_key: a canonical sorted key over the participant references, maintained by trigger, exact for two-party threads and absent for a group. Losing it to the rewrite would have been a silent product regression.
Three tables were already right and are untouched. conversation_states, conversation_labels and message_states key on user_id, which stays the correct grain — each staff member keeps their own pins on a workspace thread.
⚠ The insert policy on conversations is dropped, not rewritten. Creating a thread goes through manage-chat as the service role, where the request gate and the block check live.
Also adds the updated_at triggers that plan finding R18 measured missing on conversations and conversation_states — both had the column and nothing maintained it.
Safe only because every chat table holds zero rows and no Edge Function wrote them. Evidence: a 17-assertion probe green at psql exit 0, proving consumer↔consumer and business↔business are now representable, that the same pair cannot fork a second thread, the XOR in both directions, the block generalisation (including the system exemption and the blocker writing in their own thread), the retired vocabulary being refused, updated_at being server-controlled, and the cascade.
CR-26.0.1-147 · the three readers that knew the old shape
Class: schema_migration · Scope: QRS-1077 · Migration:20260905195000_v2_principal_conversations_readers.sql · Applied to Dev: yes, 2026-09-05, read back
⚠⚠ Postgres does not track column dependencies inside a $$-quoted function body. Dropping the two participant columns therefore succeeded while leaving three functions that read them — they compile, they deploy, and they fail at call time.
Found by asking the catalog rather than by reading the migration. Twelve names matched the string; nine were unrelated hits on c.workspace_id in catalogue and order queries, and exactly three genuinely read the dropped columns. Enumerate, then verdict — a grep whose hits you do not classify is a confirmation device, not a measurement.
messages_reject_when_blocked, the block trigger, generalised: any other participant's block stops a message, while the sender's own does not — blocking means I do not want to hear from you, and writing in a thread you blocked is unblocking by action.systemstays exempt.get_my_conversations, rewritten so the counterparty is derived as the participant that is not me, which is what lets one query serve all three pairings. Three columns are appended (counterparty_kind,counterparty_user_id,is_request) and every existing column keeps its name and meaning, so the client compiles unchanged. ⚠ Its unread count now compares the author rather than the sender kind — the old version could use the kind only while there were exactly two sides.get_my_order_book, whosechatIdmatched the two dropped columns inside ajsonb_build_object— invisible to every dependency check Postgres performs, so it would have failed on the merchant's order screen. Re-emitted from the live body with one asserted substitution, so nothing else in a hundred-line projection was disturbed.
Verified afterwards by the same catalog query: zero functions still reference the dropped columns.
CR-26.0.1-148 · open_conversation
Class: schema_migration · Scope: QRS-1077 · Migration:20260905200000_v2_open_conversation.sql · Applied to Dev: yes, 2026-09-05, read back
One function, service_role only. It is SQL rather than Edge Function code for one reason: find-or-create plus two participant inserts is four statements, and doing them from outside leaves a window in which two taps create two threads — exactly the invariant the old unique index protected. A forked thread splits the history and the unread count, and neither side can tell which one the other is reading.
⚠ It carries the request gate's asymmetry (ADR-0032 D5). A workspace peer is accepted on creation, because a card exists to be messaged. A person peer is a request, because capability does not imply trust — the biodata basic link is forwardable by design.
It also refuses a principal opening a thread with itself, which would otherwise produce a conversation whose two participants are one principal that every counterparty query would have to defend against.
CR-26.0.1-149 · manage-chat
Class: ef_code · Scope: QRS-1015 · Deployed: yes, 2026-09-05, verify_jwt read back · config.toml: verify_jwt = true, declared in the same change
Closes QRS-1015: the RLS comment named "the chat Edge Function" as the enforcement layer and no Edge Function touched the tables, so every consumer chat write was a stub. Ten actions.
⚠ Sequenced deliberately after QRS-1077. Writing it against the old (workspace_id, consumer_user_id) pair would have encoded the two-role shape into the write path and then had to be rewritten — the entire reason the identity model was decided while chat was unbuilt.
⚠ Unlike manage-biodata it writes its tables directly, because conversations and messages live in public — ADR-0031 applied rather than overridden, since a principal-based conversation spans consumers, merchants and business-to-business.
The rules enforced here rather than trusted: acting as a workspace is checked against workspace_members, never taken from the body · the pin cap is server-side, because two devices pinning at once each pass their own client check and leave four pins · mark_read touches only what somebody else sent, or the other side's receipts go wrong · mark_unread is a flag rather than a rewind · muted stores an until rather than a since, or muting for an hour mutes forever · an item message is snapshotted rather than joined, or a price change rewrites a family's own conversation · delete_label is scoped by user_id as well as id · and a conversation that is not yours answers 404, never 403, so a caller cannot learn that an id exists.
Idempotency is two different things here. The request key makes a retried request a no-op; client_key makes a retried send produce no second bubble, refused by a unique constraint rather than by a check that can race with itself.
Evidence. deno check clean; 6 Deno tests pinning the action list, the exactly-one-principal rule, the retired vocabulary being refused, the muted-until mapping and the body bounds; test:ef 420/420.
CR-26.0.1-150 · 20260905220000_v2_consumer_prefs_and_activity
Class: schema_migration · Scope: QRS-1011 · Applied to Dev: not yet
Two things Account and Notifications both need, and neither had.
Preferences are server-side (decision D-w). users.prefs jsonb not null default '{}'. A consumer reaches QR setu from a phone and from a browser, and a quiet-hours window that applies on one and not the other is worse than no quiet hours at all: the person believes they are unreachable and they are not. The device stores stay a cache of this, never the record.
There is deliberately no notifications table. The list is DERIVED from what already happened — a waiting biodata request, a subject's removal request, unread conversations — rather than written as rows when those things occur. A stored notification is a second copy of a fact that can drift from the fact, and it needs a writer at every producer. What is stored is the read state, because "have I seen this" is not derivable from anything else.
⚠ Read state is per ACCOUNT, not per device (decision D-c, QRS-1011). Clearing the bell on a phone must clear it in the browser, or the count is a property of the device and the person is told about the same request twice. It is keyed by a synthetic text id the deriver reproduces, not a polymorphic foreign key: the thing being marked read may live in three different tables, one of them in another schema, and a polymorphic FK would need a type column and could still dangle.
⚠ Unread is summed from get_my_conversations() rather than re-derived. A second count(*) over messages would be a second definition of "unread", free to drift from the one the Chats tab renders. It is also already right about the cases a hand-rolled count gets wrong: a manually-unread thread, and the sender's own messages never counting against them.
⚠ The biodata half crosses a schema boundary, which is why the function is SECURITY DEFINER with search_path left at public and every biodata.* table qualified. Widening the path inside a definer function to reach a domain schema is a privilege-escalation vector (ADR-0031).
Marketplace kinds (orders, payments, new items) are absent while the marketplace is off: off means absent, and a notification for a feature the person cannot reach is noise.
Verified: 11 assertions green on the local stack. ⚠ The cross-account isolation assertion is mutation-proven — dropping the ownership predicate turns it red — so it is falsifiable rather than merely untested. check:sql green across 104 migrations.
CR-26.0.1-151 · 20260905233000_v2_every_account_has_an_address
Class: schema_migration · Scope: QRS-1078 · Applied to Dev: not yet
Closes a hole in the identity spine, not in a screen. Under ADR-0032 D2 the slug is the address, so a principal without one cannot be reached by any capability: no QR, no link, no invite, no conversation. handle_new_user referenced slug zero times and INDIVIDUAL_STEPS was ['auth','name','celebrate'], so every consumer reached the authenticated product unaddressable. Merchants were already consistent — their slug is a mandatory onboarding step. Raised by the owner, who spotted the inconsistency against the locked chat model.
⚠ The seed is never the phone or the email. Both are in scope on the trigger row, both are credentials, and a slug is printed on a QR code and handed to strangers. The locked model's first move is that the phone is a credential and never an address. Mutation-proven: seeding from new.phone makes the address literally 919812345678 and turns the assertion red.
⚠ A business signup is seeded opaquely, which avoids a measured regression: a merchant whose display name matched their brand had already reserved that brand from themselves by the time they reached their own slug step. The name is the primary key, so the row cannot be released and re-inserted for the workspace — transferring it belongs in provision_merchant_workspace (QRS-1087). A merchant's personal address is their chat address, not their shopfront, so an opaque one costs nothing today and is renameable.
⚠ Releasing means nulling the owner, not flipping state. slugs_one_active_per_user is (user_id) WHERE user_id IS NOT NULL — one row per user ever — so a released row that keeps user_id still holds the person's slot and the next claim fails on the index. The schema was built for the other way: slugs_one_owner permits an unowned row and a trigger forces such a row to released. The name stays in the registry, which is what stops it being sniped the instant it is vacated and keeps a printed QR from re-pointing at a stranger.
Verified: 11 assertions green on the local stack. The privacy rule and the every-account invariant are both mutation-proven. check:sql green across 105 migrations.
CR-26.0.1-152 · 20260907120000_v2_consumer_read_state_consolidated
Class: schema_migration · Scope: QRS-1134, QRS-1135 · Applied to Dev: yes, 2026-09-07, verified by live functional probe · Contracting: yes, requires_min_app_build: 26000100
Two tables existed for one fact, and the activity RPC was deriving a list the domain also derives.
The duplicate. notification_reads (09-04, 0 rows, no reader) and consumer_notification_reads (09-05, 0 rows, read by the activity RPC) modelled the same thing. CR-26.0.1-150 created the second one day after the first, under if not exists, so it never failed and nothing complained. The 09-04 migration had already argued against that name under ADR-0031: the axis is the bounded context, never the audience. A consumer who starts a business gains a workspace membership and never a second account, so read state split by persona gives one pair of eyes two read states, and a notification dismissed as a consumer reappears on the merchant surface.
⚠ The design settles the name, and it was fetched rather than reasoned about.prototype/mobile-console/notification-reads.js — a module nobody had read — says "On lift this becomes a notification_reads table", and it is the merchant module that also serves "the bell dot on Chats". prototype/consumer/Notifications.dc.html agrees from the other side: "READ STATE IS DEVICE LOCAL HERE, AND THAT IS A PROTOTYPE SHORTCUT. On lift it must be PER ACCOUNT on the server." One read-state module, both surfaces, unprefixed.
⚠ notification_reads was also the strictly better table on its own merits: its policy is FOR ALL, while the prefixed one was SELECT-only and could never have backed the write path this migration adds. Both were empty, which is the only cheap moment the duplicate-identity class (QRS-249) ever has.
⚠ Why this is contracting and why the floor is nevertheless the existing one. It is a DROP, so it declares a build floor by rule. No shipped client ever read the dropped table: the consumerActivity seam was stub-bound from the day the table was created until today, so the barrel never issued a call against it. 26000100 — the existing floor — is therefore safe, and this is recorded rather than asserted because a contraction bounded by nothing is exactly the claim that deserves its reasoning written down. drop ... restrict, never cascade, so an unmeasured dependant fails loudly instead of being dropped alongside.
The derivation moved. get_my_consumer_activity() returned a finished notification list with its own kind vocabulary (biodataRequest, message) while @qrsetu/domain's consumerNotifications() derived the same list from raw activity with eleven snake_case kinds. One fact, two implementations, two vocabularies that disagreed. CLAUDE.md decides which one moves: services return stored rows as-is; derivation lives in @qrsetu/domain. It now projects activity. ⚠ The old shape could not have driven the screen anyway — it returned a single total unread while the design renders a per-row dot.
orders and savedVendorsWithNewItems are projected present and empty by decision: marketplace is off and off means absent (D-a), and no saved table exists in any schema. Meetings are now projected; the superseded comment claimed they had no producer, which had gone stale.
⚠⚠ A bug in this migration was caught by the functional probe and by nothing else.timestamptz - timestamptz yields an interval, which cannot cast to int (42846). check:sql passed and db push succeeded, because a plpgsql body is not parsed until it runs. Corrected with migration repair --status reverted followed by a re-push, so the repo and Dev hold one correct version rather than a broken one plus a patch.
CR-26.0.1-153 · 20260907121000_v2_consumer_prefs_write
Class: schema_migration · Scope: QRS-1135 · Applied to Dev: yes, 2026-09-07, verified by live functional probe
Preferences could be read and never written. get_my_prefs() has read users.prefs since CR-26.0.1-150; nothing has ever written it. manage-account accepts exactly four actions and none touches prefs, so the delivery goal's D1 condition — "a preference survives a reinstall" — was unreachable from the client. Adds set_my_prefs(jsonb) and jsonb_booleans_only(jsonb).
⚠ An RPC rather than an Edge Function action, which is a deliberate deviation from CLAUDE.md's writes-go-to-an-EF rule. The deciding reason is a duplicate-source-of-truth hazard on a rule with a security consequence: the locked-topic rule — order and payment alerts can never be switched off — lives in mergeConsumerPrefs in @qrsetu/domain, and an Edge Function runs Deno and cannot import that package. An EF write path would need a second implementation of a rule whose failure mode is silently silencing somebody's payment alerts. In SQL there is one enforcement point and it is unbypassable, because authenticated holds no table privileges on users.
Precedent: set_my_display_name and set_my_primary_context are already definer functions granted to authenticated. What an EF would have added, itemised: authorization no (auth.uid() is the whole of it), idempotency no (setting a preference twice is setting it once), external HTTP none, secrets none, multi-step none.
⚠ The one thing it would have added is rate limiting, and that is named rather than glossed. The two precedents carry the identical exposure and the platform answer is public.rate_limits (QRS-921, unbuilt), so this is a consistent pre-existing gap rather than a new one. Mitigated locally by a 4 KB ceiling on the stored payload, because this is granted to authenticated and prefs sits on a core table read at every session bootstrap.
channels and topics deep-merge per key — jsonb || is shallow, so a top-level merge would drop every key the patch did not mention and the loss would be silent. quietHours and areaSlugreplace, and an explicit null is meaningful. It enforces a type and never a vocabulary: the closed channel and topic sets stay in @qrsetu/domain, uncopied.
The probe is the evidence. Sending topics: {orders: false, payments: false, messages: false} with channels: {push: false, email: "not-a-bool"} stored orders: true and payments: true (coerced), messages: false — a legitimate opt-out still applies, so the coercion is surgical rather than a blanket reset — kept push: false, and dropped the non-boolean email. Also proven: authenticated is refused direct table access to notification_reads (42501), so the definer functions are genuinely the only door.
CR-26.0.1-154 · 20260907122000_v2_chat_labels_read
Class: schema_migration · Scope: QRS-1135 · Applied to Dev: yes, 2026-09-07, verified by live functional probe
Read-only, additive, no table change. Adds get_my_chat_labels().
There were three writes and no read. conversation_labels and conversation_label_members shipped 2026-08-11 and manage-chat exposes create_label, delete_label and set_label_member — but no RPC projected either table, and authenticated holds zero table privileges, so a user could create a list and never see it again. It blocked two states the design declares: Chats.dc.html's listState carries "No custom lists yet" and "Six custom lists", fetched live 2026-09-07.
🔎 This is the second occurrence of one defect shape, which is why it is recorded and not merely fixed. The ChatService interface's own header records the same missing-read omission one layer up, in TypeScript: "listLabels was ADDED AFTER THE FACT … three write methods look complete enough that the missing fourth does not stand out." It was fixed there and then recurred here. A CRUD surface reviewed write-by-write reads as complete.
Labels and membership return in one call because they render one control: fetched separately they can arrive skewed, and a label whose membership has not landed renders as an empty list rather than as loading. Ordered by position, not created_at, so a rename cannot reshuffle the pill row. A label with no members is absent from the map rather than present-and-empty, so the client reads members[id] ?? []. The membership join re-checks label ownership, so a guessed label_id cannot leak another account's conversation ids.
⚠ quick_replies is deliberately not included: it is workspace_id-scoped, so it is a business's saved replies. A consumer honestly has none, and the merchant half is still owed a read.
CR-26.0.1-155 · 20260907130000_v2_biodata_photo_keys_owner_scoped
Class: schema_migration · Scope: QRS-1146 · Applied to Dev: yes, 2026-09-07, verified by live functional probe · Contracting: yes, floor 26000100
Replaces get_biodata_photo_keys(uuid[]) with get_biodata_photo_keys(uuid[], text).
The old version resolved a signable storage key for ANY owner's photograph. It filtered m.purpose = 'biodata_photo' and m.status = 'ready' and nothing else, so it returned the bucket, key and derivative for any media id passed to it. Its own comment said, correctly, "IT DOES NOT AUTHORISE ANYTHING, and must never be asked to" — the tier projection belongs in get_public_biodata.
That reasoning was sound and it was the only guard, because the write path had none either (QRS-1145): validatePhotos checked that a photograph's mediaId was a 1-64 character string. So family A could put family B's photograph id in A's own photos array, publish, and A's released readers would be served a stranger's private portrait — with every gate green, because the id is well-formed and the row it names is a real ready photograph. Both layers are closed in this one change.
It takes the slug rather than the owner's uuid, and that is the decision worth recording. The obvious signature was (p_media_ids, p_owner_user_id), with get_public_biodata returning the owner so biodata-read could pass it back down. But get_public_biodata is a public read whose result reaches an anonymous visitor, so that would have put an internal principal id on the wire with nothing but a delete in an Edge Function between it and exposure. The address is already public and names the same owner, so nothing new crosses the boundary.
state = 'active' and user_id is not null are copied from get_public_biodata deliberately: somebody who takes over a released address must not inherit the previous holder's page, and therefore must not inherit their photographs.
The old signature is dropped, not overloaded. Leaving the unfiltered version callable is the opposite of the point, and config.toml's own warning that a stale artifact "reads as a working feature" applies to an overload too. Grants unchanged: service_role only, anon and authenticated none — verified on Dev with has_function_privilege.
Deploy order is migration first, then biodata-read, and the intermediate window fails closed. The old function is gone, the call errors, presignPhotos returns url: null for every photograph and the page renders as not-loaded. No disclosure in the gap.
Verified on Dev by live functional probe, inside a do block ending in raise so the whole thing rolls back. Two real accounts, one photograph each:
| probe | result | wanted |
|---|---|---|
| family A's address, both ids | 1 key, b_in_a = false | 1 |
| family B's address, both ids | 1 key | 1 |
| the unfiltered predicate, same ids | 2 keys | 2 — the hole, measured rather than argued |
a released address | 0 | 0 |
a pending row | 0 | 0 |
a biodata_emblem id | 0 | 0 |
public.media row count back to 0 afterwards, slugs unchanged.
CR-26.0.1-156 · 20260907150000_v2_my_biodata_photo_keys
Class: schema_migration · Scope: QRS-1149 · Applied to Dev: yes, 2026-09-07, probed
Read-only, additive, no table change. Adds get_my_biodata_photo_keys(uuid[], uuid) and a my_photo_urls action on biodata-read.
A family could complete a photograph upload and see an empty grey slot. get_my_biodata projects 'photos', p.photos — the raw jsonb, holding media ids and no urls and no storage keys. Presigning needs the R2 secret, which only an Edge Function holds, so there was no path by which the editor could ever have rendered a photograph. The upload worked and showed nothing.
This is the fourth instance of one shape in a single day, and the only one that is a missing DIRECTION rather than a wrong line. The write path (QRS-1145) and the stranger's read (QRS-1146, QRS-1147) were both built the same day, and nobody asked what the owner sees — because the public reader's view had been designed so carefully that "the read" felt finished. Completion of parts never bounds the whole.
A second function rather than a branch on the first, and that is the day's other lesson applied on purpose. get_biodata_photo_keys confines by slug, because that is what a public caller can prove; this confines by owner, because that is what a session can prove. One function taking both and choosing would put private photographs one branch from the wrong confinement — the exact shape of bucket: ownerScoped ? 'private' : 'media' that got manage-media reverted.
It goes on biodata-read, not manage-biodata, for two reasons and the second is decisive: presigning is the shared rule (the R2 secret, bucketEnvFor, the rate limit), so duplicating it across two functions is the "share the RULES, never the DISPATCH" violation the context split exists to prevent; and every manage-biodata action claims an idempotency key, whose replay returns the stored result — a presign read would serve expired urls for ever after its first call, and a family's photographs would stop rendering fifteen minutes into a session and never recover.
Two deliberate differences from the public function: both purposes (an owner editing the opening needs their own emblem, which the public projection never serves) and no status = 'ready' filter (the owner is the uploader, so a visible pending slot is how a stalled upload stops being silent; status is projected so the editor can say which it is).
The owner branch gets its own per-user rate-limit scope (biodata_my_photos) rather than reusing the IP-keyed biodata_read_source: the first version shared it, which made an owner's editor spend the budget strangers behind the same carrier NAT depend on.
CR-26.0.1-157 · biodata-read: the rate limiter's return value was misread
Class: ef_code · Scope: QRS-1150 · Deployed to Dev: yes, 2026-09-07, proven by probe
⚠⚠ P0. The public biodata page returned 429 to every visitor from the moment this function was deployed.
consume_rate_limit is declared returns TABLE(allowed boolean, used integer, resets_at timestamptz), so PostgREST hands back an array of one row. consume() did return data === true. An array is never === true, so every call returned false and every read answered 429 — including the first request in a fresh window.
Severity: a total outage of the page a family shares with a prospective match, which is the entire point of the consumer launch.
Why it survived: the bug wears its own guard's clothes. The failure mode is a 429 from a rate limiter, which is indistinguishable from a working rate limiter to anyone outside the function — and the fail-closed comment directly above the broken line is correct, and is what made the symptom look designed.
It was found by accident, while verifying an unrelated rate-limit scope change:
| observed | measured | verdict |
|---|---|---|
| HTTP 429 | sum(used) = 2 | contradiction |
| — | limit = 120 | the limiter had not been reached |
A contradiction between a measurement and an observation is the only instrument that could have found it. Nothing else could: no test in the repo asserts a real RPC's return shape; deno check cannot, because the EF client is any by necessity; and a mocked client returns whatever the mock's author believed, which is the same belief that wrote the bug.
Fix: read the row. const row = Array.isArray(data) ? data[0] : data; return row?.allowed === true.
⚠ Boolean(data) would have been the wrong fix, and it is worth naming: it is true for an array and for [], so a limiter answering with no row would read as allowed — the fail-open direction, which is worse than the bug being fixed.
Swept: consume_rate_limit has exactly one caller, and every other === true in the Edge Function tree is a body or row boolean. This is the only instance.
Proven on Dev after the fix: an unknown slug returns 200 {"biodata":null} and a real person slug returns 200, where both returned 429 before.
The standing lesson: a TABLE-returning function is an ARRAY over PostgREST. When an RPC's result decides a branch, read its pg_get_function_result once rather than assuming the shape.
CR-26.0.1-158 · the three biodata Edge Function secrets, set on Dev
Class: ef_secret · Scope: QRS-1080 · Dev: set 2026-09-08, names read back
Found by the owner, on a device, as a 500 on publish — and it could not have been found any other way, because an ef_secret is one of the nine change classes no automated gate in this repository can see. Every gate was green while publish was unusable.
BIODATA_REFERENCE_KEY is not set. Publishing mints a reference and there is deliberately no
default key: a default is a published key, and anybody could then invert the permutation and
recover the counter it exists to hide.
at mintBiodataReference (manage-biodata/reference.ts:88)
at publish (manage-biodata/index.ts:150)The throw is correct and stays. What was missing was provisioning.
⚠⚠ The second key is the one the first failure does not name, and it fires on the primary case
Setting only BIODATA_REFERENCE_KEY would have moved the same 500 one line down:
| step | needs | what happens unset |
|---|---|---|
publish mints the reference | BIODATA_REFERENCE_KEY | 500 — the error the owner saw |
publish awaits sendSubjectNotice (index.ts:206) | — | not wrapped in try/catch |
| the notice mints a removal token, before resolving the template | BIODATA_REMOVAL_SIGNING_KEY | 500 on publish, not notice: failed |
And it fires on the primary case rather than an edge one: a family making a profile for a sister, where the owner is not the subject and a subject phone exists. The owner-is-subject path returns not_applicable before reaching it, so a fix verified on the wrong test profile would have looked complete.
This was established by reading the call graph, not by retrying — which matters, because a retry-driven fix would have shipped one key, reported success, and handed the owner the same 500.
What was set, and how the values were kept out of everything
| name | value | why it cannot default |
|---|---|---|
BIODATA_REFERENCE_KEY | 32 random bytes, hex | a default key is a published key |
BIODATA_REMOVAL_SIGNING_KEY | 32 random bytes, hex | the subject acts with no account; the token is the only proof |
PUBLIC_WEB_BASE_URL | https://devv.qrsetu.com (measured 200 first) | a wrong origin looks right and points at the wrong environment (QRS-992) |
Generated into a scratchpad env file, passed with --env-file, then the file was overwritten and unlinked. No value reached a command line, a log, this repository, or any output. The read-back is by name: secrets list returns a SHA digest in its value field, so even the verification exposes nothing.
Measured: 21 secrets before, 24 after.
⚠ What is machine-checked here, and what is not
| claim | how |
|---|---|
| the three names exist on Dev | secrets list read-back — machine |
| the mint is a bijection and throws when unset | 51 Deno tests green — machine |
| publish works for a family | the owner retrying on the device — NOT machine-checked |
Those are three different claims and only the first two were produced by a command. Saying "publish is fixed" would be the third rule's category error: what is proven is that the reason it failed is gone.
⚠ Prod needs its own distinct values at promotion. A key shared across environments makes any Dev reader a Prod oracle, which is the whole property the permutation exists to protect.
CR-26.0.1-159 · manage-biodata: one concept had two environment variable names
Class: ef_code · Scope: QRS-1165 · Deployed to Dev: yes, 2026-09-08, deployed source downloaded and grepped
manage-biodata read PUBLIC_WEB_ORIGIN at three sites. .env.example documents PUBLIC_WEB_BASE_URL, with a nine-line comment, and place-public-order reads that name.
So an owner provisioning Dev from the documentation would have set the documented variable and still had a broken mint_share and a broken subject notice — with an error naming a variable that appears nowhere in the documentation. That reads as a bug in the code rather than a gap in the configuration, which is the expensive part: it is worse than a plain omission, because the fix looks applied.
This is the QRS-249 duplicate-identity class applied to configuration. It was found only because CR-26.0.1-158 forced a read of what the function actually reads rather than of what the documentation says it reads.
Renamed toward the documented name — four occurrences: two Deno.env.get sites, one error string, one tracker row. The alternative, setting a second secret with the same value, would have made the divergence permanent and left two names to keep in step for ever.
.env.example now also documents both BIODATA_* keys, which it named neither of, and that absence is part of why they were never set.
Evidence: deno check clean · 51 Deno tests green · deployed at 183 kB, then downloaded back with functions download --use-api --workdir into a directory outside the repository (so a read-only check cannot overwrite the working tree) and grepped: PUBLIC_WEB_BASE_URL at three sites, PUBLIC_WEB_ORIGIN at none. Deployed state measured, not inferred from the deploy message.
CR-26.0.1-160 · biodata.profiles.custom_values dropped
Class: schema_migration · Scope: QRS-1171 · Dev: applied 2026-09-08, read back
A column written by nothing and read by nothing, whose NAME asserted it held a custom field's answer. The answer lives INLINE in custom_schema, which is what get_public_biodata reads. A client that believed the name would have published every custom field as nothing, at every tier, silently (QRS-1166).
Contraction, and safe because it was MEASURED rather than assumed:
| measured on Dev | value |
|---|---|
rows with custom_values data | 0 |
| profiles total | 2 |
requires_min_app_build owed | none — no shipped client writes the key |
⚠⚠ The generator caught two bugs in its own transform, and both would have produced valid SQL
The first attempt removed the column by filtering lines and mending the comma that filtering left behind. Its own guards rejected it twice:
- A line filter deleted seven real columns.
reopen_biodata_profilenames the column inline:values, field_tiers, hidden, people, photos, opening, custom_schema, custom_values. Dropping any line that mentionscustom_valuestakes the whole list with it, and the result is valid SQL that inserts a profile with no content. - The comma mend used
\s*, which spans newlines, so it could reach across statements.
Both failures are silent and syntactically valid — the worst kind available to a migration. The edits became four asserted substitutions that throw if an anchor moves, plus an invariant that fails if any other bag disappeared. That is what "edit by asserted substitution" is for.
Read back on Dev, not inferred from the push
| check | result |
|---|---|
custom_values column | absent |
custom_schema | present, and newly carrying a COMMENT ON that says the answer is inline |
| functions naming it | 0 (was 3) |
| constraints naming it | 0 (was 1) |
| bag CHECK | still guards the remaining seven bags |
get_my_biodata_overview | executed, returning the honest first-run shape |
| profiles | 2, still readable |
⚠ The last row is the one a definition read-back cannot give you: replacing a function body proves nothing about whether it still runs.
CR-26.0.1-161 · manage-biodata stops accepting the key
Class: ef_code · Scope: QRS-1171 · Deployed to Dev: yes, 2026-09-08
custom_values leaves the PATCHABLE allow-list and its validator is removed. The ordering is the point: an unknown patch key is refused by name rather than ignored, so a client that still sends it gets a 400 that says which key, instead of a write error against a column that no longer exists.
Evidence: deno check clean · 51 Deno tests green · deployed at 183 kB.
CR-26.0.1-162 · qr_reports and verify_qr_lookups
Class: schema_migration · Scope: QRS-1237 · Applied to Dev: yes, 2026-09-09
Purely additive: one table, one function, no existing object touched.
⚠ It is what makes the verified verdict reachable at all. verifyCanonical in @qrsetu/domain/scan emits qrsetu-verified only when a business lookup answers, and nothing could answer, so every scan of our own domain resolved to slug-unknown (caution). The safety screen could show three of its four designed verdicts and never the one that justifies the product's claim.
It discloses nothing new. Every business field returned is already anon-readable through get_public_setu_card, and the same status = 'published' gate applies, so it adds no oracle over draft cards. The one genuinely new datum is the report count, which the design renders to the scanner deliberately. qr_reports holds no table grant to anon or authenticated: the only read is the definer function's count, and the only write is an Edge Function's service role.
Two fields are deliberate constants, and saying so is the point: code_status is always unknown (a cancelled-sticker answer needs the minted-code registry, QRS-576) and vpa_owner is always null (no VPA is stored anywhere in the platform, measured). Both make the engine fall closed — no lookup, no tick.
Evidence, read back on Dev:
| check | result |
|---|---|
| registered migration version | 20260909170000 — matches the filename, no orphan row (QRS-267) |
qr_reports | present, RLS on |
table grants to anon / authenticated | none (postgres + service_role only) |
verify_qr_lookups signature | p_slug text, p_code text, p_vpa text, p_url text |
| function grants | anon, authenticated (+ postgres, service_role) |
| verified path | returns the seeded business, and normalises case and surrounding whitespace |
| unknown slug · stranger VPA · all-null args | business: null, fail-closed |
| two distinct reporters | count 2 |
| a third insert by an existing reporter | refused by the unique index, count still 2 |
check:sql | green — 113 migrations, every table, function, policy and trigger commented |
⚠ test:db (pgTAP) could not run. Both drives sit under check:disk's 15 GB floor (C: 5.1 GB, D: 12.6 GB), so the local Supabase stack will not start. Recorded as unrun, not green.
⚠ A fourth parameter where the plan of record specified three. p_url is required by the engine's own behaviour: verifyUrl raises the reported signal for any host, so a reported phishing page — the highest-value case for a stranger's code — needs its count reachable. With three parameters that path could never have been warned about.
CR-26.0.1-163 · manage-account gains report_code
Class: ef_code · Scope: QRS-1237 · Deployed to Dev: yes, 2026-09-09
One new action; the four existing ones are untouched. It is the only way into qr_reports.
⚠ A repeat report of the same target by the same person returns SUCCESS, not a conflict. The unique index refuses the duplicate row and the handler swallows 23505, because telling somebody who is trying to warn other people that their warning failed is wrong — and the alternative leaks whether they had reported it before. The reporter id comes from the verified JWT and never from the body.
Why here rather than a new function: the same reason claim_slug lives here. It needs a verified principal plus a service-role write, and this dispatcher already has both.
Evidence: deno check clean on index.ts and helpers.ts · 19 Deno tests green (up from 11) · deployed at 124 kB · probed live: unauthenticated report_code returns 401 with a clean error body.
⚠ The authenticated round trip is NOT proven. It needs a real session, so the device pass is what closes it. The validator and the SQL layer are both covered independently.
CR-26.0.1-164 · manage-biodata redeployed, so the release length finally applies
Class: ef_code · Scope: QRS-1193 · Deployed to Dev: yes, 2026-09-09
No code change in this release. The fix was committed earlier and had never been deployed, so the live function still hardcoded 90 days and the design's 30 / 90 / 180 choice reached nothing.
⚠ This is the sixth rule's other half in miniature: the repo was right, the declaration was right, and the environment disagreed with both — which no gate that reads the repo can see. Deployed opportunistically while manage-account was going out for CR-163.
Evidence: deployed at 184 kB. ⚠ The behaviour itself is not re-verified: no share was released at 30 or 180 days against the new bundle, so this records a deploy and not a functional confirmation.
CR-26.0.1-165 · The published biodata carried no age, and the preview did
Class: schema_migration · Scope: QRS-1269 · Deployed to Dev: yes, 2026-09-20
20260920120000_v2_public_biodata_derives_age.sql.
⚠⚠ This is the owner's own reported symptom in its sharpest form. age is row 146 of the field registry — ('age', 'about', 'basic', 'derived', true, false, 5) — so it is required and visible at the basic tier, which is the forwardable link a family prints on a QR code. It is also derived, and QRS-1158 established that a derived value is never written. get_public_biodata builds its values object by reading v_profile.values -> f.id, so a field nothing ever writes can never appear; its source dob is private and reaches nobody. Net effect: one of the first three things anybody reads on a marriage biodata was visible in the owner's in-app preview and absent from the published page, the forwarded link and the QR-scanned view.
🔎 Both halves were individually correct, which is the whole lesson. Not storing a derived value is right. Not projecting a private date is right. The defect lived in the seam, so no test on either side could see it — and that is the argument for a shared resolver plus a paired test rather than for more review.
What changed: two functions — biodata_age_from(text, timestamptz) and biodata_derived_values(jsonb, jsonb, text[], timestamptz) — and one merge line in the projection.
⚠ The date still never leaves the database. Only the completed-years integer does, and only when age's own effective tier is one the viewer holds. Projecting dob so a renderer could do the arithmetic would publish the date; a renderer that can compute an age is one that was handed a birth date.
⚠ A derived field carries its own tier, never its source's. Gating the derivation on dob would hide every age from everybody — the same bug one layer down. An owner who narrows age to released still gets that, because the derivation runs through effective_biodata_tier like every stored field.
⚠ Nothing is backfilled, because nothing was ever stored. Read-path only: every already-published profile gains its age on the next read, with no write and no cache stamp to bump.
Evidence: supabase/tests/database/biodata_projection_test.sql — new, 39 assertions, supabase test db PASS against the local stack with every migration applied. Mutation-proven both directions: replacing biodata_derived_values with select '{}' reddens six assertions, including both halves of the TypeScript/SQL pair. The two functions were additionally probed on a throwaway supabase/postgres:17.6.1.155 container across the birthday boundary (28 the day before, 29 the day of), every malformed-date case (ISO, garbage, impossible month, impossible day, future — all NULL, never a raise), and the tier gate. The derived age equals the TypeScript resolver's answer for the same record, to the year.
✅ APPLIED TO DEV AND PROVEN ON THE LIVE PAGE, 2026-09-20. migration list showed exactly one pending version and no orphan in either direction before the push. Read back afterwards, on Dev: both functions exist, pg_get_functiondef on the LIVE get_public_biodata contains the biodata_derived_values merge, and the newest applied version is 20260920120000.
The functional probe, which is the half that matters. Both real published profiles on Dev now return an age through the RPC — sunil 23, vedaalahade 21 — with stores_age: false (nothing ever wrote one) and dob_leaked: false (the private date appears nowhere in the returned document). Then the live pages themselves: https://devv.qrsetu.com/sunil/biodata and /vedaalahade/biodata both 200, both rendering the age in the facts tile, and zero month names anywhere in either document — so the integer left the database and the date did not.
⚠ The CLI's stored login was the PROD account and could not see qr-setu-dev at all. Used npm run sb -- dev, which supplies the Dev identity per invocation and refuses unless the token can actually see the intended ref. This is the documented trap and it is still live — check projects list before any write, every time.
Forward fix: revert the migration; the age disappears from every public reading again. No data is lost, because the age was never stored.
CR-26.0.1-166 · manage-biodata accepts both the retired and the new conflict codes
The Edge Function half of QRS-1277, and it must deploy before CR-26.0.1-167. That ordering is the only user-visible risk in the pair: if the migration lands first, this function meets a code it does not know and falls through to throw e, which _shared/response.ts turns into a 500 — masking the message above status 499, so the family loses the actionable sentence, and firing captureServerException, so every ordinary version conflict also becomes a Sentry alert.
No client change. Every conflict branch still returns 409, and packages/data/src/biodata/service.supabase.ts:66 derives conflict from the HTTP status alone, never from the Postgres code. That is what makes the codes movable at all.
⚠ 40001 is retained permanently, not left behind. Production runs its own migration timeline, so an environment will exist that still raises it. This is the expand half of expand-contract; the contract half is deleting the branch once no live environment raises it, which is a decision with a measurement behind it rather than a cleanup.
⚠ Both rethrow helpers moved into a new codes.ts, for a mechanical reason. index.ts calls Deno.serve at module scope, so anything defined there cannot be imported by a test without starting an HTTP server. That is much of why this mapping had zero coverage while two copies of it disagreed about what 40001 meant — one rendering it "Somebody else saved this… reload or overwrite", the other "That is not possible right now".
Evidence. deno check clean; test:ef 473/473, including 13 new cases in manage-biodata/tests/codes.test.ts — both mappers, both new codes, the legacy code, not-found, and the case that makes the rest mean anything: an unknown code is rethrown untouched, without which a mapper returning 409 for everything would pass. Deployed to Dev and verified inside the deployed bundle rather than off the success line.
Forward fix: redeploy the previous bundle. No state is written, so nothing needs restoring.
CR-26.0.1-167 · Seven write-API functions stop raising SQLSTATE 40001
⚠⚠ This closes an incident that cost 1,378,564,796 aborted transactions. 40001 is the SQL standard's serialization_failure — "this transaction failed for concurrency reasons, retry it" — and these seven functions used it for deterministic refusals: a withdrawn share, a concluded profile, a stale version. None can succeed on retry, so the retry never stopped. Four pooled connections spun from 15 September, CPU pinned at 100% for seven days, against zero HTTP requests for these RPCs in the same window.
⚠ The reported hypothesis — that unaccessed biodata profiles were being processed — was disproven. release_biodata_share is a single-row UPDATE … WHERE sh.id = p_share_id, there is no pg_cron on this project at all, and every biodata trigger is touch_updated_at. No profile was ever scanned. The trigger was identified to 17 milliseconds: share 05227baa… was withdrawn at 11:44:52.672 and the stuck backend started at 11:44:52.660.
Proven by controlled experiment, not inferred. Three throwaway functions with identical bodies differing only in SQLSTATE, each called once through the real PostgREST 14.5:
| errcode | HTTP | executions from ONE call |
|---|---|---|
40001 | never returned | 9,215 over 96 s |
QRS09 | 400 | 1 |
QRS10 | 400 | 1 |
P0002 | 500 | 1 |
⚠ The retry outlives the request — the client gave up at 08:03:36 and the last execution landed at 08:03:52. That is how one request poisons a pooled connection for a week with nothing calling it. All probes dropped afterwards, verified zero remain.
Two codes rather than one. QRS09 is a stale version (reload genuinely helps); QRS10 is a precondition (it cannot). ⚠ publish, unpublish and conclude test both a version and a precondition in one WHERE clause and raise one message, so they are genuinely ambiguous and take QRS10 knowingly: the old copy asserted "Somebody else saved this" even for a removal guard, which is a fabrication, and a neutral sentence is the honest answer to a cause the function cannot determine. Giving those three the re-read update_biodata_profile already has is a tracker row, not scope in a migration whose purpose is to stop a retry loop.
⚠ The migration was generated, not transcribed. Each body was extracted from its defining migration and only the errcode literal substituted; a verification script then proved all seven bodies byte-identical apart from that literal. A wrong predicate in release_biodata_share would be a disclosure, not a bug, so "generated carefully" was not good enough.
⚠ update_biodata_profile and reopen_biodata_profile were taken from 20260908150000, not the original 20260905141000 — they are defined twice, and the earlier body would resurrect a superseded one on db reset.
No schema, data, grant or signature change. create or replace preserves grants and comments, both verified live afterwards: service_role only, anon and authenticated false, comments intact. The read path is untouched, so public rendering, QR access, the ?embed=1 work, R2 and Cloudflare are unaffected.
Evidence. migration list showed exactly one pending and no orphan version; all seven read back live with no 40001 and the right code each; check:sql green over 115 migrations. Containment measured before and after: 631 rollbacks/sec → 0.
Forward fix: re-apply the previous bodies with create or replace — preserved verbatim in 20260905141000, 20260905170000 and 20260908150000. ⚠ Reverting this migration without also reverting CR-26.0.1-166 is safe, because the function still handles 40001. Reverting the function without the migration is not, and that is the ordering depends_on records.
CR-26.0.1-168 · A family can choose the design its biodata reads in
template_key on biodata.profiles has existed since 20260904113123, and the public read has honoured it since the three designs shipped: get_public_biodata projects it and the web page's registry renders that design. Nothing could write it. update_biodata_profile patched theme_id and family_layout and never template_key, so every profile read default (measured on Dev: 8 of 8). The owner's "How it reads" picker needs the write (QRS-1366).
Expand only. A constraint, profiles_template_key_known (default, parichay, chitra, patrika), and one line in the RPC's patch, written against the latest body (20260922090000) so QRS09/P0002 are unchanged. An absent key leaves the design alone.
⚠ Order: this migration, then CR-26.0.1-169. The deployed function refuses template_key as an unknown key, which is a safe 400. The reverse order would return success for a choice the old RPC ignored.
Evidence. Local reset of all 116 migrations clean; check:sql green; pgTAP biodata_template_key_test.sql 10/10. On Dev, read back: the constraint exists, the live body writes the column and still raises QRS09, no row violates; a rolled-back call on a real profile wrote the design inside the transaction and nothing persisted.
Forward fix: drop the constraint and re-apply the 20260922090000 body. No row is rewritten.
CR-26.0.1-169 · manage-biodata accepts and validates template_key
One patch key, validated against the same four designs, so an unknown design is a 400 with a sentence rather than a constraint violation. Deployed to Dev after CR-26.0.1-168; check:ef-drift --project dev shows no drift. deno test manage-biodata/ 69/69 with a new case for every registered design and the refusals.
Forward fix: redeploy the previous bundle, which refuses the key as unknown. Nothing is destructive.
CR-26.0.1-170 · every reader of a published biodata sees what the family enabled (D10 (a), D11.1)
The owner's MVP decision: no access requests, so a published biodata shows every field the family has not made private or hidden, on every link. get_public_biodata's default reader tier now comes from one policy function, biodata_public_reader_tier(), which answers released; the rest of the body is unchanged. Accepted knowingly: a printed QR code now shows the released fields too (full name, parents, shown address and map place, references, employer, LinkedIn, place of birth, the family contact's phone). private fields reach nobody, hidden still suppresses, a withdrawn or expired person link still closes the page. biodata-read then serves originals to every reader, which zoom needs.
Evidence. Clean local reset; pgTAP biodata_projection_test.sql passes, asserting the MVP policy with the real function and the tier machinery with it pinned to basic. Applied to Dev: the function answers released, the live read calls it, and the three published Dev profiles read anonymously at released with no date of birth and no phone number.
Forward fix: make the function answer basic again. Nothing is lost.
CR-26.0.1-171 · D6's nine-digit Biodata ID, stored (D10.2, D11.2)
biodata.profiles.reference_no bigint, unique, nine digits, set once. Expand only: the seven-character reference stays until no installed build reads it. Two service-role functions store the number and report an older profile still without one; the owner overview returns it. The minting key never enters the database.
Evidence. check:sql green; new pgTAP biodata_reference_no_test.sql passes (type, privileges, owner scope, set once, drafts excluded, range and uniqueness refused). Applied to Dev and read back.
Forward fix: drop the column and the two functions and restore the previous overview body.
CR-26.0.1-172 · manage-biodata mints the nine-digit ID
D6's 30-bit keyed Feistel with cycle-walking, and the inverse of the seven-character permutation, so one formula serves every path: the number is D6's permutation of the profile's sequence value. Assigned at publish, and for a profile published earlier on its next save. A failure is logged (biodata.reference_no.failed) and never fails the family's save.
Evidence. deno test 83/83 (range, bijection over 4,000, determinism, key dependence, round trip). Deployed to Dev after CR-26.0.1-171; check:ef-drift --project dev no drift. The live mint waits on a signed-in owner's save or publish.
Forward fix: redeploy the previous bundle; numbers already stored stay.
CR-26.0.1-173 · RATE_LIMIT_SUBJECT_KEY on Dev (QRS-1430)
A new Edge Function secret, generated locally and never printed. It must exist before CR-26.0.1-174 runs, or the limiter refuses every public read. Evidence: the secret list names it; the live page renders after CR-174. Forward fix: rotate by setting a new value; never unset while CR-174 is live.
CR-26.0.1-174 · biodata-read keys the rate-limit subject (QRS-1430)
The stored subject was a bare SHA-256 of the visitor's IP, which the IPv4 space reverses; the table's migration required an HMAC. Now it is one, under CR-173's key, fail closed without it. The 95 bare-digest rows on Dev were deleted with the owner's approval. Evidence: deno test biodata-read/ 12/12, no drift after deploy, the live page renders, one keyed row after the purge. Forward fix: a new bundle.