Skip to content

AI strategy: what to build, what is not AI at all, and what it costs ​

Part of car_sales — the dealership operating layer.

⚠ Three findings that reframe the question, before any feature list

1 · AI inference is not the cost problem. 🧮 A busy franchise outlet running AI across conversations, classification, drafting and management questions costs ≈₹2,100/month — ₹25,500 a year of Claude API spend. That is ~17% of one Growth subscription. The cost anxiety is misplaced.

2 · ⚠ The expensive part is the CHANNEL, not the intelligence. 🧮 A five-turn AI conversation on Haiku 4.5 costs ₹1.24. 🌐 Meta's own Business-Agent interactions bill at roughly ₹3.5-4.5 per interaction — so buying Meta's AI is 14-18x the cost of running our own inference for an entire conversation. Run our own model, on our own channel — ⚠ and that channel is QR Setu native chat, not WhatsApp, for reasons that go well beyond cost (§4a).

3 · ⚠⚠ Most of the proposed "AI agents" are DETERMINISTIC work wearing an AI label. Overdue follow-up detection, test-drive reminders, going-cold flags, "which rep has the most overdue follow-ups" — every one is a SQL query or the recurrence engine that already exists and is live. Using an LLM there costs money and is less reliable. See §3. QRS-859.

1 · Verified model pricing ​

🌐 Read from Anthropic's own pricing page, 2026-08-23. Per million tokens, USD.

ModelInputCache readOutputBatch in/out
Claude Haiku 4.5$1$0.10$5$0.50 / $2.50
Claude Sonnet 5$2$0.20$10$1 / $5
Claude Opus 5$5$0.50$25$2.50 / $12.50

⚠ One correction to what a search result told me: Sonnet 5's $2/$10 was announced as introductory through 31 Aug 2026. 🌐 Anthropic's page now states this is the standard price and "the previously scheduled increase to $3/$15 will not occur." So there is no September repricing risk to plan around.

🔎 Cache read is 10% of base input, a 5-minute cache write is 1.25x, a 1-hour write is 2x. A dealer's catalogue, policies and FAQ are the same on every call, so they are cached and cost a tenth — which is what makes the numbers below small.

2 · What each QR Setu task actually costs ​

🌐 Anthropic's own published benchmark: ~3,700 tokens per support conversation, ~$37 per 10,000 on Haiku 4.5 = ₹0.31 per conversation. Our own model, with the dealer context cached:

TaskModelCost
Lead classification, one callHaiku 4.5₹0.13
Review reply draftHaiku 4.5₹0.15
Reply draft for a rep, in native chatHaiku 4.5₹0.18
Customer-concierge turn, grounded in the catalogueHaiku 4.5₹0.25
Lead summary before a callHaiku 4.5₹0.31
Indic mixed-script intent extractionSonnet 5₹0.44
GM question over pre-computed metricsSonnet 5₹0.81
Root-cause analysis, rareOpus 5₹9.45

A busy franchise outlet, per month ​

VolumeCost
AI conversations, 5 turns each1,000₹1,239
Lead classifications2,000₹269
Reply drafts1,000₹176
Rep pre-call summaries600₹184
GM / Team Leader questions200₹161
Opus root-cause10₹94
Total4,810 actions₹2,124/month · ₹25,487/year

🧮 Blended ₹0.44 per AI action. 🔎 That is the number the whole commercial design rests on, and it means an AI package priced at ₹30,000-2,00,000/yr carries 60-70% gross margin — comparable to the software itself.

3 · ⚠⚠ What is NOT AI, and must never be built as AI ​

The single most valuable output of this analysis

📘 The brief calls the AI Follow-Up Agent "potentially one of the most valuable capabilities." 🔎 It is roughly 80% deterministic, and the deterministic 80% is cheaper, faster, auditable and correct every time.

Proposed as AIActuallyWhy deterministic wins
Follow-ups due · overdueSQLAn exact answer. An LLM can be wrong about a date
Test-drive reminder at T-24h🟢 The reminders engine, already liveRules + sparse occurrence exceptions, expansion in @qrsetu/domain, already unit-tested
Service due by date or odometer🟢 The same live engine
Leads going cold — no contact in N daysSQL
"Which rep has the most overdue follow-ups?"One query⚠ Asking an LLM to count is paying for a worse answer
"Which leads were not contacted in 24 hours?"One query
Customers repeatedly engaging with a cardA scan-count threshold
Review-request eligibilitySQL — interaction done, not yet asked
Hot / warm / cold when the fields are explicitRulesOnly free text needs a model
Round-robin assignment, availability routingRules

The genuinely AI part of "follow-up" is exactly two things: reading intent out of a free-text conversation, and drafting the message. Everything else is a query. Build the query layer first; the LLM sits on top of it and is cheap because it is doing so little.

The rule to apply to every future AI request

If a deterministic rule can produce the answer, an LLM is a more expensive way to be less certain. Use the model only where the input is unstructured language or the output is prose.

4 · Where AI genuinely earns its cost ​

🔎 Ranked by value per rupee of inference, and the first one is India-specific.

#CapabilityWhy a model is required
1Reading a mixed-script conversation into structured lead fields⚠ The highest-value AI task in the vertical, and no US vendor has it. Real Indian dealership chat is transliterated Hindi/Marathi/English — "gaadi ka on-road price kya hai, exchange me meri Swift". Rules cannot parse it; a model does it easily. It fills budget, timeline, model interest, exchange and finance intent from a conversation nobody had to structure. ⚠ CORRECTED 2026-08-23: the source is QR SETU NATIVE CHAT, not WhatsApp — see §4a
2Reply drafting in the customer's own language and registerProse output. A rep edits and sends
3Pre-call summary of a long historyCompression of unstructured text
4Grounded catalogue Q&A on the Setu CardNatural-language question against retrieved rows
5Review sentiment and recurring-complaint themesClustering prose into themes
6Explaining WHY a metric moved — over numbers SQL computed⚠ Narration, never calculation. See §6
7Tele-call QA in Indic languages🌐 One listed group runs 262 tele-callers against 11 IT staff. Scoring calls is the highest-leverage automation in the building. ⚠ Needs audio, so it is later

4a · ⚠ What the AI actually reads, and the correction that forced this section ​

⚠ I wrote "reading a WhatsApp thread" as the #1 AI capability WITHOUT ESTABLISHING THAT WE CAN READ ONE. We cannot

📘 The owner asked the right question: "how exactly would QR Setu access each Sales Representative's WhatsApp conversations?" The answer is that it could not, and the claim was a design error, not a detail.QRS-860.

The six questions, answered directly ​

QuestionAnswer
1What data source is the AI actually reading?⚠ Two, and I had collapsed them: the anonymous enquiry note (no account, high volume) and the registered chat thread (conversations, account required). Both ours. See §4b
2Does QR Setu have access to a salesperson's WhatsApp?❌ No, and it never can. 🌐 The Cloud API delivers only messages sent to a business phone number you control, by webhook. There is no access to personal accounts and no historical-conversation read at all. A rep's personal WhatsApp is E2E-encrypted and outside any API
3Privacy, consent and API limits?Even for a business number: inbound only from the moment of connection, no backfill, a 24-hour service window governing replies, template approval for anything outside it, an opt-in ledger, and DPDP purpose-limitation on every message we process
4Why build around WhatsApp when native chat exists?⚠ We should not. See the comparison below — native chat wins on five grounds, three of which I had not considered
5Can the same AI capability be delivered from our own data?✅ Better, and in two tiers. The language problem is identical in any text box. A registered thread gives the complete history with card, rep and party attached; an anonymous enquiry gives one message — lower fidelity, far higher volume
6Where does WhatsApp stay useful?Reach and notification, never record. See §4c

⚠ And there was a second, worse problem with the WhatsApp version ​

🔎 To read a conversation on a business number, the customer has to message the business number instead of the rep's personal one. That is a behaviour change on the customer side — and 📘 the entire roadmap of this vertical was re-ordered around behaviour change required being the thing that kills adoption (value proposition §6). I proposed the highest-value AI capability on top of the exact failure mode the roadmap exists to avoid.

Why native chat is the better system of record ​

QR Setu native chatWhatsApp business number
Channel cost per message✅ Zero₹0.115 utility, ₹0.8631 marketing
Thread completeness✅ Complete from the first message⚠ Inbound only from connection. No backfill, ever
Attribution✅ Which card was scanned → which rep → which party, by construction⚠ A phone number arriving at a shared inbox. Rep attribution must be reconstructed by routing rules
Per-rep threads✅ Free — the card identifies the rep⚠ Needs a number per rep or shared-inbox routing
Consent provenance✅ The customer initiated by scanning. Strongest possible under DPDPAn opt-in ledger we maintain and must defend
Policy surface✅ None. Our surface, our rules24-hour window, template approval, quality rating, pooled limits, ❓ TRAI's draft OTT rules
AI grounding✅ Retrieval and history in the same database as the answerTwo stores to reconcile

🔎 Three of those I had not weighed and they are the decisive ones: the thread is complete rather than partial, attribution is free rather than reconstructed, and there is no third-party policy surface on a capability the whole product depends on.

The decision

QR Setu native chat is the system of record for customer conversations, lead intelligence and AI grounding. WhatsApp is a reach and notification channel, and an add-on.

📘 Recorded as D7.

4b · ⚠ The one thing that can kill native chat, and it is not adoption ​

Notification delivery. If the customer closes the tab, they never learn the rep replied

🔎 This is the strongest argument for WhatsApp and it has nothing to do with preference: a customer already has WhatsApp installed, with push working. A web card cannot rely on that — mobile web push needs a permission grant and is unreliable across iOS Safari and in-app browsers. Without a dependable "you have a reply" signal, native chat degrades into a contact form with extra steps, and the conversation the AI depends on never happens.

⚠⚠ AND THE FLOW I DREW HERE WAS IMPOSSIBLE. conversations.consumer_user_id IS NOT NULL

I wrote "sends the first message with NO account, then asked only for a phone number" and called anonymous-first preserved. 🧮 The live schema forbids it:

sql
consumer_user_id  uuid not null references public.users (id) on delete cascade
unique (workspace_id, consumer_user_id)

An anonymous conversation is not policy-gated, it is UNREPRESENTABLE. ⚠ And the migration's own header comment states the contrast I had just contradicted, verbatim: "consumer_user_id IS NOT NULL, WHICH IS THE OPPOSITE OF orders.buyer_user_id." The answer was written in the schema and I did not read it.

📘 CLAUDE.md is equally explicit: the public card must be usable with no account, and registration is demanded only where identity is genuinely required — and it names chat as one of those cases. So chat requires registration, full stop. QRS-861.

The distinction I had collapsed: an anonymous ENQUIRY is not a registered CONVERSATION ​

Track A · Anonymous enquiryTrack B · Registered chat
Account required❌ No✅ Yes — phone + OTP creates a users row
Schema precedent🧮 orders.buyer_user_id is nullable — anonymous transactions are already a supported pattern🧮 conversations.consumer_user_id is NOT NULL
ShapeOne-way: name, phone, free-text note → a leadTwo-way, threaded, persistent
Volume⚠ The majority, and it must stay frictionlessA minority early, growing
Reply reaches them by⚠ WhatsApp or a phone call — there is no in-app thread to notify intoIn-app, push via the consumer app
Channel cost₹0.115 per utility messageZero
What the AI readsOne free-text message — abundantThe complete thread — richer
Exists today❌ No enquiry table at all. Wave 1🟢 Schema shipped 2026-08-11; no dealership surface

The corrected architecture: two tracks, and the goal is to MIGRATE between them

text
TRACK A — the entry point, and it must never have a signup wall
  scan → browse the card anonymously → "Enquire" → name + phone + free text
       → lead created, rep attributed, NO account
       → rep replies by WhatsApp or phone          [₹0.115, or a call]
       → AI parses the SINGLE enquiry message → intent, model, budget, timeline

TRACK B — the destination, entered when the customer has a reason to register
  order tracking · test-drive management · service history · a real conversation
       → phone + OTP → a users row → conversations thread
       → in-app, zero channel cost, complete history, push via the consumer app
       → AI parses the FULL thread

🔎 This is what makes native chat strategically important in the way the owner intends — as the DESTINATION, not the entry point. A signup wall at first contact destroys the growth mechanic; a signup offered in exchange for order tracking on a ₹15 lakh purchase is a fair trade a buyer will take.

⚠ And it narrows the AI claim honestly: "the AI reads the conversation" is true of Track B and a minority of volume. The high-volume AI task is parsing a single anonymous enquiry — which is still mixed-script, still unparseable by rules, and still ours.

4c · Where WhatsApp genuinely stays ​

UseKeep on WhatsApp?Why
Reply notification✅ Yes, the load-bearing useInstalled base and reliable push. ₹0.115
Outbound utility — test-drive confirm and remind, service due, insurance renewal✅ YesTime-critical, and the customer need not be in our surface
Campaigns✅ Yes, as an add-onReach. Nothing native competes with it
A customer who refuses anything else✅ Yes, as a fallbackSome will. Losing the enquiry is worse than a partial thread
The conversation of record❌ NoIncomplete, unattributed, policy-exposed, and costs per message
AI grounding source❌ NoWe would be reasoning over a partial copy of a conversation we do not own

⚠ The failure mode to avoid is drift: every individual decision to "just reply on WhatsApp because the customer is there" moves the record out of the system, and the AI, the attribution and the management rollups all read from the record. A conversation that happened on a rep's personal phone is invisible to every management screen in this product, which is the same defect as a shared Setu Card one layer up.

5 · The verdict matrix ​

BUILD NOW · BUILD LATER · ADD-ON · EXPERIMENT · AVOID

#Proposed capabilityVerdictModelNote
1Reply drafting for a rep, in native chat🟢 BUILD NOWHaiku⚠ The only AI feature whose data already exists — the chat schema shipped 2026-08-11. Human sends, so no autonomy risk
2AI-assisted free reputation report🟢 BUILD NOWSonnet🔎 Summarises a dealer's public Google reviews. Zero schema dependency, sellable before the product exists (Podium teardown)
3Lead field extraction from an ANONYMOUS ENQUIRY🟢 BUILD LATER — Wave 1-2Sonnet⚠ The high-volume version. One free-text note, no account. Needs leads and an enquiry table
3bLead field extraction from a REGISTERED THREAD🟡 BUILD LATER — Wave 2SonnetRicher, scarcer. The differentiator, on a smaller base (§4b)
4Lead qualification / hot-warm-cold🟡 BUILD LATERHaiku⚠ Only the free-text part. Explicit fields stay rules
5Pre-call summary for a rep🟡 BUILD LATERHaikuNeeds interaction history
6AI Customer Concierge on the Organisation Setu Card🔵 ADD-ONHaiku⚠ Cheapest place to put AI, because the channel is OUR web surface with no per-message fee. Needs the catalogue and media first
7AI Lead Response Agent in native chat🔵 ADD-ONHaiku + Sonnet escalationThe flagship demo, and ⚠ cheaper than the WhatsApp version because our channel is free (§4a). Autonomy risk — see §8
8AI Review agent — eligibility, drafts, sentiment, escalation🔵 ADD-ONHaiku⚠ Eligibility is SQL; drafting and sentiment are AI
9AI Test-Drive Agent🟡 ADAPT, mostly deterministicHaiku for the conversation only⚠ Availability, slots, booking, confirmation and reminders are rules + the live reminder engine. AI only reads the request and writes the messages
10AI Receptionist for FAQs🔵 ADD-ONHaikuGrounded in dealership policy rows. Escalates on anything uncertain
11AI Follow-Up Agent🟡 SPLITHaiku for drafts🔎 80% SQL. Build the queries; the AI writes the message
12AI Team Leader assistant🔵 ADD-ONSonnet⚠ Question → pre-built metric, never generated SQL
13AI GM / management agent🔵 ADD-ONSonnet, Opus for root-cause⚠ See §6. Highest-risk design on the page
14Tele-call QA in Indic languages🟠 EXPERIMENTSonnet + transcriptionThe biggest prize and the biggest build. Needs audio capture nobody has agreed to
15AI-generated reports🔴 AVOID—📘 The brief rules it out and is right. A report nobody asked for, in prose nobody checks
16AI negotiating price or discount🔴 AVOID—⚠ Commercial and legal exposure on a ₹10-45 lakh purchase
17Buying Meta's Business-Agent AI🔴 AVOID—🧮 14-18x our own inference cost per interaction
18AI voice bot answering the dealership phone🔴 AVOID—📘 We refuse to own telephony. A voice agent is a different company
19Text-to-SQL over live dealer data🔴 AVOID—⚠ A correctness and RLS-security hazard. §6

🧮 2 BUILD NOW · 3 BUILD LATER · 5 ADD-ON · 2 SPLIT/ADAPT · 1 EXPERIMENT · 5 AVOID.

6 · ⚠ The management agent: the model must never compute the number ​

Text-to-SQL over a multi-tenant dealership database is the worst idea in the brief, and it is the most tempting one

📘 "How did the Pune sales team perform this week?" is exactly the question a GM wants to type. The naive implementation — hand the schema to a model and let it write SQL — fails three ways at once:

FailureConsequence
Wrong number, confidently phrasedA GM acts on it. ⚠ Worse than no answer, because it is unfalsifiable to the reader
RLS bypassGenerated SQL running with more privilege than the asker, across a tenant boundary
Cost and latency driftAn unbounded generated query on the largest tables in the schema

The safe design, and it is barely more work:

text
GM question in natural language
      ↓
model classifies it against a CLOSED SET of pre-built metric queries      [Sonnet]
      ↓
the metric runs as an ordinary RPC, under the asker's own RLS             [SQL — the number]
      ↓
model phrases the answer and CITES the metric it used                     [Sonnet]
      ↓
no matching metric?  →  "I cannot answer that yet" + the nearest dashboard

🔎 The model chooses and narrates; SQL computes. Every answer is reproducible, every number is auditable, and the failure mode is "I do not have that metric" rather than a plausible fabrication. ⚠ Opus is worth its ₹9.45 only for the rare "why did test-drive conversion fall" question, and only over numbers already computed.

7 · Grounding architecture ​

📘 The brief is right that this belongs in the design from the beginning.

RuleWhy
Retrieval only. No model-remembered facts about the dealerPrevents invented specs, prices and policies
Price, availability and discount must be a RETRIEVED FIELD or the answer is a refusal⚠ A wrong price on a ₹20 lakh car is a commercial exposure, not a bug
Retrieval runs under the ASKER's RLS, not a service roleAn AI path must not become a privilege-escalation path
Uncertain → escalate, never hedge📘 "Never fabricate insight" is already a platform principle
Every AI message is audited — prompt, retrieved rows, model, output, recipient⚠ The dealer's brand sent it. They need to know what was said
PII minimisation into the prompt📘 The observability seam already scrubs PII; the AI path needs its own rule, and DPDP purpose-limitation applies
Dealer-configurable tone, never dealer-authored promptsA free-text system prompt is an injection surface

8 · Where AI must never act alone ​

BoundaryRule
Price, discount, negotiation⚠ Never. Retrieved list price only, no arithmetic on it
Commitments — delivery dates, allotment, finance approvalNever. The DMS owns these and we do not touch it
Anything after two failed grounding attemptsHand to a human with the transcript
A customer who asks for a personImmediate handover, no retry
Complaints and negative sentiment⚠ Route to a manager. Never let AI handle an angry customer under the dealer's brand
Legal, warranty, insurance advice📘 Licence-gated territory. Refuse
Sending outside consentOpt-in ledger governs. AI does not decide who may be messaged

9 · Metering, in units a dealer understands ​

📘 The brief is right: meter business units, never tokens.

One unit, so the dealer has one number to think about

An "AI action" is one completed piece of AI work — a drafted reply, a classified lead, a summarised history, one turn of a customer conversation, one answered management question.

🔎 Internal model routing is ours, not the dealer's problem. Haiku, Sonnet and Opus all consume one action, and 🧮 the blended cost of ₹0.44 is what makes that simplification safe.

PackageAI actions / month🧮 Our cost / yrPrice / yr, ex-GSTMargin
Included with Growth300₹1,600₹0 — discovery—
Included with Complete1,000₹5,300₹0 — discovery—
QR Setu AI Assist2,000₹10,800₹29,99964%
QR Setu AI Agent6,000₹32,400₹99,99968%
QR Setu AI Agentic15,000₹81,000₹1,99,99960%
Enterprise AInegotiated—quoted—
Top-up recharge5,000-action pack₹2,250₹7,50070%

⚠ Why a small allowance is INCLUDED rather than everything being an add-on

📘 CLAUDE.md's fifth rule: gating controls access, never discovery. An AI capability nobody has ever seen is an AI capability nobody upgrades to. 🔎 300 actions a month is roughly ten drafted replies a day — enough for a rep to form a habit and a GM to notice, nowhere near enough to run the dealership on.

⚠ And the packages must show usage against allowance in the product, because an add-on whose consumption is invisible generates a billing dispute rather than an upgrade.

10 · Model routing ​

Use caseHaiku 4.5Sonnet 5Opus 5Recommended
FAQ / policy answer, grounded✅✅✅Haiku
Reply drafting✅✅✅Haiku
Lead classification from explicit fields⚠ use rules, not a modelNo model
Lead classification from free text✅✅✅Haiku, Sonnet if mixed-script
Indic mixed-script intent extraction, over native chat🟡 marginal✅✅Sonnet — ⚠ the one place the cheaper model is a false economy
Conversation summarisation✅✅✅Haiku
Customer conversation, multi-turn grounded✅✅✅Haiku, escalate to Sonnet on a second grounding failure
Follow-up recommendation⚠ mostly rules✅ for the proseRules + Haiku
Management question → metric selection🟡✅✅Sonnet
"Why did this metric move" root-cause❌🟡✅Opus, rare, over pre-computed numbers
Review sentiment and themes✅✅✅Haiku, batched at 50% off

Two cost levers that matter more than model choice

  1. Prompt caching. The dealer's catalogue, offers, policies and FAQ are identical on every call — cached, they cost 10%. 🧮 This is most of why a concierge turn is ₹0.25 rather than ₹2.
  2. Batch for anything not interactive — review sentiment, nightly lead scoring, weekly digests. 🌐 50% off both input and output.

11 · Sequencing: AI cannot precede the data ​

⚠ An AI layer over this schema today would have almost nothing to be grounded in

🧮 Five of the seven primitives this vertical declares have zero tables — party, schedule, resource, asset, campaign (architecture validation). No leads, no customers, no interaction history, no catalogue media. "Summarise this customer's history" has no history to read.

So two AI capabilities can ship early and the rest cannot:

ShipsWhy it can
AI reply draftingWave 1The chat schema shipped 2026-08-11. It needs a conversation and nothing else
AI reputation reportNow, as a sales assetReads a dealer's public Google reviews. No QR Setu schema at all
Everything elseWave 2+Waits on parties, leads, interaction_events, catalogue media

12 · The pitch, and what it must not become ​

What a dealer should hear

"Your consultant opens a customer's chat and the next message is already written, in the customer's own language, using your catalogue and your offers. Your enquiries get a first reply in seconds instead of hours. And your GM asks a question in plain language instead of reading five dashboards."

🔎 Every clause is a workload reduction or a response-time claim — never "you will sell more cars." 📘 Same discipline the rest of this vertical follows, and 🌐 the same discipline Podium follows at $3 billion.

⚠ What it must not become

  • "QR Setu has AI" as a slide. 📘 The brief rules this out and is right.
  • An AI chatbot on the card. The concierge is grounded retrieval with a refusal path, not a chat toy.
  • A reason to raise the base price. 🔎 AI is an ARPU layer above the subscription, priced on metered actions, precisely so the core book stays comparable.
  • An in-app upsell on native. 📘 ADR-0002 / Apple 3.1.3(d). Plans and AI packages convert on the web.
  • An autonomous voice on the dealer's brand with no audit trail. §7 is not optional.