Appearance
Architecture assessment: 10M users, N industries
Critical review requested 2026-08-07. Scope: product-vision alignment, documentation accuracy, horizontal/vertical scalability, performance, and whether the design survives 10 million Indian users across many industries. Written to find foundational problems while they are still cheap.
Verdict up front, so the rest can be read against it: the identity/tenancy spine and the public-card boundary are sound and I would not change them. Two things do not survive contact with the vision: the claim that a new vertical is configuration rather than code (§A1 — the dairy test disproves it), and the performance model for the public card at 10M vendors (§B1 — the >95% cache-hit target assumes traffic concentration that 10M cards do not have). Both are fixable now and expensive later. Six smaller findings follow, then what is genuinely fine.
Part A · Can the platform absorb a new industry without isolated implementations?
A1 — ⚠ FOUNDATIONAL: "new vertical = config, not code" is true only for CATALOGUE-SHAPED variation
ADR-0009's promise, restated in CLAUDE.md and in the data architecture, is that a new vertical is a row plus configuration. Run the owner's dairy example against it and it fails — not at the edges, at the centre.
| Dairy requirement | Supported by the current design? |
|---|---|
| Dairy product catalogue | ✅ catalog_items + attributes |
| Inventory management | 🟡 track_inventory + stock_quantity is a decrementing counter. Dairy stock is "today I have 40 litres" — a per-period quantity, replenished daily and expiring. A counter models the wrong thing. |
| Daily sales tracking | ❌ Nothing. orders models an online, buyer-initiated purchase. A dairy records walk-up cash sales with no buyer identity and no online flow. That is a sales ledger entry, not an order. |
| Subscription-based customers | ❌ Nothing. "1 litre daily, skip Sundays, paused 3–7 Aug" is a vendor→customer recurring commitment. subscriptions in the design means QRSETU plans. I correctly said vendor subscriptions are vendor commerce and then never modelled them. |
| Recurring billing | ❌ Nothing, and it is a different money flow from plan billing — vendor→customer, quantity-variable, month-end reconciled against actual deliveries, settled via e-mandate/UPI Autopay. Not a Razorpay plan subscription. |
| CRM | ❌ leads (26.4.0) models a prospect. A dairy has customers — with a delivery address, a standing order, and a running balance. And critically, a dairy's customer is usually not a QRSETU user at all, so users cannot represent them. |
| Business-specific workflows | ❌ The deepest gap. archetypes.item_attribute_schema declares data shape. A daily delivery run, a stylist-and-chair appointment, a tuition batch with a syllabus are processes. Nothing in the design expresses a process. |
The honest architectural statement, which should replace the current promise: you cannot express arbitrary domain workflows as configuration — that is a known limit, not a gap to engineer away. What is achievable, and what the design must actually deliver, is: decompose workflows into a small set of composable process primitives, and let an archetype compose them. Then a genuinely novel workflow is a new primitive (code, once, reusable) rather than a per-vertical implementation.
The primitive set the dairy example reveals — and every one of these serves ≥3 other verticals:
| Primitive | What it owns | Also serves |
|---|---|---|
| Catalogue | What the business offers | all |
| Party | A workspace-scoped counterparty: customer, lead, student, patient, tenant. Not a users row. | every vertical; subsumes leads/contacts as stages, not tables |
| Schedule | "A thing happens at a time" | dairy delivery run · salon appointment · tuition batch · doctor slot · property viewing |
| Recurrence | The rule behind a repeating schedule | dairy standing order · gym membership · AMC · tuition term · reminder |
| Fulfilment | An obligation discharged (delivered / served / completed) | dairy delivery · service call · class attended |
| Ledger | Money and quantity movement, online or manual | daily cash sales · online orders · adjustments |
| Balance | Running account per Party | dairy month-end bill · cafe tab · retainer drawdown |
Six of the current design's seven concepts survive; Party, Schedule, Fulfilment, Ledger and Balance are genuinely absent. The existing capabilities list (catalog, media, leads, appointments, orders, subscriptions) is not a primitive set — it is a feature list with leads and appointments as siblings of catalog, which is a category error: appointments is Schedule × Party, and leads is Party at an early stage. Left as-is, the 20th vertical costs as much as the 3rd, which is precisely the outcome ADR-0009 exists to prevent.
A2 — The single largest piece of unused leverage in the repo: @qrsetu/domain's recurrence engine already solves the dairy subscription
packages/domain/src/reminders/recurrence.ts exists, is pure, has 117 node --test cases, handles DST, and was hardened during the QRS-247 refactor (which also closed a latent bug by making an unhandled frequency a compile error). ADR-0016's model — a recurrence rule plus sparse exception rows, with occurrences computed on read — is exactly, precisely the right model for "1L daily, skip Sundays, paused 3–7 Aug." The exception row is the pause; the rule is the standing order.
This is not an analogy. It is the same computation, and reusing it means a dairy standing order, a gym membership, an AMC and a tuition term are all one tested engine rather than four hand-rolled date loops. ADR-0016 is therefore validated far beyond reminders and should be promoted from "the reminders domain model" to "the platform recurrence model" — the strongest single finding in this review's favour, and it is currently invisible because the code lives under a feature-shaped folder name.
A3 — Archetypes need a declared PROCESS composition, not only an attribute schema
Concretely: archetypes gains enabled_primitives text[] and per-primitive configuration, so:
dairy→ Catalogue + Party + Schedule(Recurrence) + Fulfilment + Ledger + Balancesalon→ Catalogue + Party + Schedule + Fulfilment + Ledgerfestival_stall→ Catalogue + Ledger (the R1 vertical — deliberately the thinnest composition)tutor→ Catalogue + Party + Schedule(Recurrence) + Balance
An archetype then composes rather than unlocks, and adding dairy is genuinely configuration — because the primitives it needs already exist. This is the change that makes the ADR-0009 promise true rather than aspirational.
A4 — Party must be workspace-scoped and identity-optional, or the model excludes most Indian SMB customers
A dairy's customers, a tutor's students and a salon's regulars are overwhelmingly not QRSETU account holders. If a customer must be a users row, the CRM is empty for every vertical that matters.
So: parties(workspace_id, name, phone, address, …, user_id uuid NULL) — the optional link fires when a consumer (category 3) recognises themselves. That nullable FK is the join between vendor CRM and the consumer product, and it is also what makes the marketplace possible later: a dairy's 200 offline customers become addressable the day they register. Getting this column in from the start costs nothing; adding it to a populated CRM later is a reconciliation project.
Part B · Performance and UX at 10M users
B1 — ⚠ FOUNDATIONAL: the >95% cache-hit target assumes traffic concentration that 10M cards do not have
The card strategy is SSR on Cloudflare Pages behind an edge cache, targeting >95% hit rate. That target is achievable when traffic concentrates on relatively few URLs. With 10M vendor cards, the distribution is a long tail: most cards get a handful of views per month, so most requests arrive at a COLD edge and go to origin. A 95% hit rate on aggregate views can coexist with a majority of cards being served cold, and the cold path is the one a first-time visitor to a small vendor experiences — the exact moment that decides whether QRSETU feels premium.
This is not an argument against the cache. It is an argument that the ORIGIN path must be designed as the common case for the tail, not as the fallback. Consequences to build in now:
- The cold card render must be one indexed query. The
cardstable (P2) already gives us this — oneSELECT … WHERE slug = $1on a unique index, plus its content children. Do not let it become a multi-table composite read, and never resolve features/entitlements on the public path. - Declare a p95 origin-render budget and gate it, the same way CWV is gated. Untargeted, this regresses invisibly.
- Keep the public read off the primary. A read replica or Cloudflare Hyperdrive for the card route means vendor writes and visitor reads never contend. Cheap to adopt now, awkward later.
- Long-lived cache + tag purge, not short TTL. The design already purges by tag via the outbox, so TTL can be long — which is what makes the tail warm for repeat visitors.
B2 — The RLS pattern in the design is the #1 known Supabase performance trap, unmitigated
Every merchant-data policy is specified as workspace_id in (select workspace_id from workspace_members where user_id = auth.uid()) (P2a). Written naively, auth.uid() is re-evaluated per row, and the subquery can be too. On a table with millions of rows this turns a seek into a scan.
Required, and it must be a written convention rather than folklore: wrap volatile calls as (select auth.uid()) so Postgres hoists them into an InitPlan; index workspace_members(user_id, workspace_id); index every tenant table's workspace_id leading column. A pgTAP test cannot catch this — it needs an EXPLAIN-based check or a documented review step, and I would add one micro-benchmark per hot table to CI rather than trusting the convention.
B3 — resolve_features is a hot path with no caching or latency budget
Three axes × eight scopes × a precedence join, potentially on every screen. At 10M users this must be provably O(small). Missing from the design: an index plan, a latency budget, and an invalidation story.
Recommendation: a resolved_features cache keyed by (workspace_id) and (user_id) for consumers, written by the outbox on any grant mutation (the mechanism already exists), with the live resolver as the authority and the cache as the read path. Do not put resolved features in JWT claims — a plan downgrade or an incident kill-switch would take up to the token lifetime to apply, and the emergency lever must be immediate (the QRS-373 lesson).
B4 — No retention or partitioning policy beyond analytics_events
analytics_events is declared monthly-partitioned; that is correct and insufficient. At 10M vendors: orders, payment_events, audit_log, outbox and media all grow without bound. Each needs a declared strategy in the baseline, because retrofitting partitioning onto a populated append-only table is a rewrite:
analytics_events— monthly partitions + a stated retention (drop partitions after N months; the rollup is permanent, the events are not). No retention was stated.payment_events,audit_log— partition by month, retain long (compliance), never delete silently.outbox— delete on success. It is a queue, not a log; an unbounded outbox is an outage.media— ~10M vendors × ~20 assets ≈ 200M rows. Partition byworkspace_idhash, or accept it as the largest table and index accordingly. Decide deliberately.
B5 — Connection pooling is named as a standard and specified nowhere
CLAUDE.md lists "connection pooling" under Reliability. Nothing states the mode. At scale, SSR isolates plus Edge Functions plus the mobile app all connect, and Supabase's direct connection ceiling is low. Supavisor in transaction mode is mandatory for the SSR path, which in turn forbids session-level features (prepared statements, SET LOCAL outside a transaction) — and the design's governor-escape-flag pattern for the archetype trigger uses exactly a session-local variable. That interaction needs resolving in the baseline, not discovered in production.
B6 — Zero-downtime DDL has no documented pattern
Expand-contract is mandated. But at scale, CREATE INDEX must be CONCURRENTLY, which cannot run in a transaction — so it cannot live in a normal Supabase migration file without care. Adding a NOT NULL column with a default rewrites the table on older Postgres. Neither is documented, and both will first be hit under live traffic. One page in the promotion runbook now.
Part C · Product-vision alignment
C1 — The four user categories are correctly represented, with one asymmetry worth naming
Solo owner, enterprise (admin + seat user), consumer and QRSETU staff all have a home in the model after the 2026-08-07 corrections, and the membership-count formulation (0 / 1 member-owned / 1 org-owned) is the right invariant. The asymmetry: enterprise is modelled thoroughly and consumers are modelled thinly. Consumers currently have an identity, a nullable buyer FK and a user grant scope — which is correct for R1 — but the workflows that force them to register (chat, booking, consultation, order tracking, subscriptions held) have no tables. Party (§A4) is the missing bridge, and specifying it now is what stops the consumer product from being a second platform later.
C2 — Subscriptions are modelled for the wrong half of the business
plans/subscriptions/seats/billing_accounts covers QRSETU→vendor billing well, including the paying-≠-owning insight. Vendor→customer recurring commerce is absent entirely (§A1), and it is:
- the dairy's core business model,
- the gym's, the tutor's, the AMC electrician's,
- the thing the
subscriptionscapability key has implied since 2026-08-01 while pointing at nothing, - and a different provider integration (e-mandate / UPI Autopay, quantity-variable month-end billing) from plan billing.
Two distinct money flows share one word today. That is the QRS-249/284/287 duplicate-meaning class, and it will be much cheaper to name them apart before either is built.
C3 — Feature control is the right shape and the tier vocabulary is still blocking
Nine→eight scopes × three axes with per-axis provenance is the correct model, and putting plan features in feature_grants rather than plans.features jsonb is the decision that makes the Access Control matrix renderable. It cannot be built while plans.key is undecided — every entitlement grant references it, and three vocabularies exist in-repo. This is now the longest-standing open blocker.
Part D · Documentation accuracy — verified, and the finding is systemic
The three-category principle appears in exactly five files: the four written on 2026-08-07 plus CLAUDE.md. It has reached no ADR and no overview page. Specifically:
| Document | State |
|---|---|
overview/target-end-users.md | Stale. Calls category 3 "Individuals — generic-feature users (no business profile)"; does not present enterprise as a category at all. |
overview/product-vision.md, overview/platform.md | Predate the three-category framing. |
| ADR-0001 (tenancy) | Superseded in part by ADR-0020 D1/D4 with no note on the ADR itself. |
| ADR-0007 (entitlements) | Storage superseded; still 🟡 Proposed as though open. |
| ADR-0009 (archetypes) | Says capability layer is BUILT — true, and now scheduled for replacement. Also carries the "config not code" promise §A1 disputes. |
| ADR-0014 (public exposure) | Amended by ADR-0020 D2; no note. |
| ADR-0011 (frontend) | Missing the fourth tier (enterprise org-admin portal). |
| ADR-0016 (reminders) | Understated — should be promoted to the platform recurrence model (§A2). |
The structural problem, not the individual gaps: a decision recorded in one place and cited from nowhere is indistinguishable from a decision never taken. This repo already knows this — it is why check:release, check:readmes and check:design fail closed — but ADRs have no freshness gate, so supersession relies on memory. A minimal, cheap control: require every ADR to carry Superseded-by:/Amends: in its frontmatter, and add a check:adr that fails when an ADR named as superseding does not name its target back. Bidirectional, mechanical, and it would have caught all six rows above.
What I would change in the design BEFORE writing any SQL
Ranked by cost-of-delay:
- Add the five missing primitives — Party, Schedule, Fulfilment, Ledger, Balance — and re-express archetypes as compositions of primitives (§A1/A3). This is the difference between a platform and a collection of verticals.
- Name the two subscription concepts apart (§C2): QRSETU
plans/subscriptionsvs vendor→customer recurring commerce. Different tables, different provider integration. - Reuse
@qrsetu/domain's recurrence engine for every recurring commitment, and promote ADR-0016 accordingly (§A2). - Make Party identity-optional from day one (§A4).
- Design the cold-card path as the common case, with a p95 origin budget, a read replica, and a one-indexed-query rule (§B1).
- Write the RLS performance convention into the baseline —
(select auth.uid()), theworkspace_memberscomposite index,workspace_idleading everywhere (§B2). - Declare partitioning + retention for all six unbounded tables (§B4).
- Resolve Supavisor transaction mode vs the session-local governor flag (§B5).
- Propagate the three-category principle into the overview pages and every affected ADR, and add the bidirectional
check:adrgate (§D).
Cost: ~+3.0d of design and ~+2.5d of implementation on top of the ~14.0d baseline (the primitives are mostly tables and RLS; the recurrence engine is reuse, not new code). Doing items 1–4 after the baseline ships costs multiples of that, because Party and Ledger are referenced by orders, analytics and CRM.
What is genuinely sound and should not be relitigated
Stated explicitly so the criticism above is calibrated rather than uniform:
- The identity/tenancy spine. Principal = user, workspace = business tenant, membership count distinguishing the categories. It absorbs solo, enterprise and consumer without special cases, and the pure-
INSERTtransition test is a real invariant. - The public-card table boundary (P2). Separating
cardsfromworkspaceseliminates a whole defect class rather than testing it harder, and it is what makes anonymous-first consumer access structural. - Feature resolution as three axes with provenance. Correct decomposition; the gap is caching, not shape.
- Money in integer minor units, snapshotted order lines, append-only provider ledgers. Standard and right.
- The outbox. It closes G1/G2 properly and is the natural home for cache invalidation of the resolver cache too.
- Text keys for reference data. Removes the QRS-249 hazard class permanently.
- Re-baselining Dev rather than migrating. The greenfield directive is the cheapest correct moment, and it is now.
One scheduling note, stated once and not re-argued
The 26.0.1 plan targets a 15 Aug public launch. The redesign directive plus these findings put the critical path at roughly 20 work-days before Sections C/D/E/F of that plan begin. I am proceeding on the assumption that the launch date is superseded by the redesign decision — asking for a first-principles schema, a user-ecosystem definition and this assessment is not compatible with shipping in eight days, and I would rather state the assumption than quietly miss the date. Correct me if the date still stands, in which case the scope conversation is a different one.