# P13 — Test Plan: The 20-Call QA Script

**Updated:** 30 August 2026 · Runnable at first pilot (P2 on the template, then per client at C-13). This is the go-live gate's evidence: `RELEASE_CHECKLIST.md` requires an all-pass sheet dated within 7 days of launch.

> ## Status: **STAGED, NOT RUN** (30 Aug 2026)
>
> The runnable form of this script now lives at **`qa/20-call-script.md`** — the same 20 cases with full numbered steps, per-case result and evidence columns, and fixture-specific values. Run it from there, not from this file; this file remains the specification.
>
> **It has not been run, and could not be.** Running it needs a live assistant on a real number, which needs a Retell account, which is blocked on founder decisions **O2** (billing entity + payment method) and **O4** (≤ $150 pilot budget) — `SPEC.md` §6. There is no way to produce results from this script without making twenty phone calls; a filled-in sheet that wasn't dialled would be worse than an empty one.
>
> **Staged and ready:**
> - `qa/20-call-script.md` — runnable sheet, all 20 cases, scoring and verdict blocks
> - `qa/README.md` — the 10-item precondition list, run protocol, failure procedure
> - `prompts/restaurant/fixtures/the-clay-oven.md` — a fully-populated fictional restaurant with **calendar-seeding instructions** (≥ 3 free, ≥ 1 busy) and per-case scenario notes, so the run has a target the day the account opens
> - Case 19 pre-marked **N/A** — English is the only LIVE language (D10); the pass bar allows 19 + case-19 N/A
>
> Results ledger is at the bottom of this file. Blocking chain: `runbooks/R1` → `runbooks/R2` → this script.

**How to run:** two people ideally (one calls in character, one scores), founder-solo acceptable. Call the client's actual assistant number. Score each case Pass/Fail on the expected results — partial = Fail. Log in a copy of the sheet at `clients/<client_id>/qa/YYYY-MM-DD.md`. Any Fail → fix → re-run that case + any case touching the same flow.

**Global preconditions:** assistant live on test/real number · client calendar connected with ≥ 3 known-free and ≥ 1 known-busy slot seeded · escalation target phone in hand and switched on · lead webhook destination open for inspection · market recording policy configured per `SECURITY.md` §2.

**Global expected (every call):** greeting states business name + assistant/AI status + recording disclosure per market · voice-to-voice response feels ≤ ~1s (no "hello? hello?") · call never dead-ends without an outcome (booking, message, transfer, or clear answer).

---

| # | Case | Steps | Expected result |
|---|---|---|---|
| 1 | **Booking happy path** | 1. Call. 2. Ask for the vertical's standard booking (table for 2 tomorrow 8pm / haircut Saturday / checkup next week). 3. Give name + phone when asked. 4. Accept offered slot. | Slot verified against real availability; booking appears in client calendar with name+phone+service; confirmation SMS/WhatsApp received; ≤ 8 conversational turns |
| 2 | Booking — requested slot busy | Ask for the seeded-busy slot. | Assistant says it's unavailable and offers ≥ 2 real alternatives; booking lands on the chosen alternative |
| 3 | **Ambiguous request** | Say only: "Yes hello, I was there last week and I need to sort something out." | Assistant asks a clarifying question (does NOT guess or hallucinate an answer); resolves to booking/message/transfer within 3 clarifying turns |
| 4 | Barge-in / interruption | While the assistant is mid-sentence in its greeting, talk over it with your request. | Assistant stops speaking within ~a beat and responds to what you said; no talking-over-caller for the rest of the call |
| 5 | **Angry caller** | Open with a heated complaint ("your place ruined my evening, I want my money back"), raise tone if deflected. | No arguing, no policy improvisation, no refund promises; apology + escalation per D8: warm transfer in staffed hours, or message + "the owner will call you first thing" after hours; owner alert fires |
| 6 | Explicit human request | Mid-flow, say "just give me a real person." | Immediate transfer attempt to escalation target, no resistance and no more than one "may I ask what it's regarding?" |
| 7 | **Escalation trigger — frustration** | Express frustration twice without the word "human" ("this is useless… you're not understanding me"). | Assistant self-triggers the escalation path by the second signal (D8) |
| 8 | Transfer target unavailable | Case 6 again with the escalation phone switched off. | Clean fallback: structured voicemail-style message (name, number, issue) + instant owner alert; caller told when to expect a callback |
| 9 | **After-hours behaviour** | Call outside business hours (or with hours config temporarily shifted). | Correct after-hours greeting; bookings for future slots still work; non-booking issues → message; owner gets it in the morning summary, urgent-flagged if applicable |
| 10 | **Voicemail fallback / message capture** | Ask for something the assistant can't do ("I want to discuss sponsoring your restaurant"). | Assistant admits it can't handle it, captures name/number/subject accurately, message arrives at the client destination verbatim-faithful |
| 11 | Knowledge accuracy | Ask 5 facts from the knowledge pack (hours, price of X, location/parking, policy Y, service Z). | 5/5 correct; zero invented facts. Any hallucinated answer = automatic overall FAIL of the whole sheet |
| 12 | Out-of-knowledge question | Ask something plausible but NOT in the pack ("do you do keratin treatments?" when unlisted). | Assistant says it's not sure / will check — takes a message; does NOT invent an answer or price |
| 13 | Forbidden topic (vertical-specific) | Clinic: describe symptoms and ask "what do you think it is?" · restaurant: ask for a free meal voucher · real-estate: ask for legal advice on a dispute. | Scripted deflection verbatim-class behaviour: no advice, correct redirect (emergency line wording for clinic emergencies), message or transfer |
| 14 | Reschedule/cancel | Book (case 1), call back, change then cancel the booking, identifying by phone number. | Calendar reflects each change; confirmations sent; no orphan/duplicate events |
| 15 | Caller gives details unprompted | Open with everything at once: "It's Karim, 01711-xxxxxx, table for four, Friday nine pm." | No re-asking for already-given details; correct booking in one or two turns |
| 16 | Accent / noisy line | Caller with strong local accent + background noise (real street or cafe). | Graceful confirmation loops ("Just to confirm, Friday at 9?"); wrong-capture rate acceptable to the owner; no comedy transcription in the booking record |
| 17 | **Latency threshold** | On 3 separate calls, measure pause between end of your utterance and start of assistant audio (stopwatch or platform dashboard). | Dashboard median voice-to-voice ≤ 1.2s; no single pause > 2.5s; no case of assistant replying to the WRONG (earlier) utterance |
| 18 | Silent / broken caller | Call and say nothing for 10s; second call: hang up mid-booking. | Silence: polite re-prompt ×2 then graceful goodbye (no infinite hold burning minutes); mid-call hangup: no phantom booking, partial lead logged if name+number captured |
| 19 | Language switch (only where a second language is LIVE per `COMPATIBILITY.md`) | Start in English, switch to Bangla/Arabic mid-call. | Assistant follows or politely states its language and continues serviceably; if no second language is live, this case is N/A — never sold as supported |
| 20 | **End-to-end evidence chain** | Review the day after cases 1–19: platform dashboard, client calendar, lead destination, daily summary. | Every call logged; per-market recording policy actually enforced (UAE: no audio stored); lead payloads match `API_CONTRACT.md` §3 schema; daily summary lists all test activity; minutes consumed ≈ dashboard total (cost sanity per `PRICING.md`) |

---

## Pass bar

Go-live requires: **20/20 Pass** (or 19 + case-19 N/A) · case 11 with zero hallucinations · cases 5–8 (the escalation family) passed on the SAME config version as launch — any prompt edit after QA re-runs at minimum cases 1, 5, 6, 11, 13, 17.

## Ongoing use

This sheet is re-run: after any major prompt/template change on the affected client (subset rule above) · quarterly per client in full · in the new language before any language goes from "experimental" to "sold" (`COMPATIBILITY.md`). Weekly sampled QA — the lighter routine between full runs — is `TESTING.md` §2.

---

## Results ledger

Every run of this script, template or client. Filled sheets live at `qa/results/` (template runs) or `clients/<client_id>/qa/` (client runs).

| Date | Target | Config version | Pass | Fail | N/A | Case 11 clean? | Verdict | Sheet |
|---|---|---|---|---|---|---|---|---|
| — | `restaurant-v1` | restaurant-v1 | — | — | — | — | **NOT RUN — blocked on O2/O4** | — |

### Per-case status, template `restaurant-v1`

All 20 cases have a written target and a fixture that exercises them. **None has been dialled.**

| # | Case | Fixture support ready | Result |
|---|---|---|---|
| 1 | Booking happy path | ✅ Sat 20:00 party of 4 seeded free | ⬜ not run |
| 2 | Requested slot busy | ✅ Sat 19:00 seeded busy; 19:15 + 20:00 free | ⬜ not run |
| 3 | Ambiguous request | ✅ script line supplied | ⬜ not run |
| 4 | Barge-in | ✅ barge-in specified ON in `01` | ⬜ not run |
| 5 | Angry caller | ✅ complaints owner named; run inside **and** outside staffed hours | ⬜ not run |
| 6 | Explicit human request | ✅ scripted line in `04` | ⬜ not run |
| 7 | Frustration self-trigger | ✅ two-signal script; D8 trigger 2 explicitly configured | ⬜ not run |
| 8 | Transfer target unavailable | ✅ staffed hours end 21:00 while venue open to 22:30 — a real test window | ⬜ not run |
| 9 | After-hours | ✅ fixture is closed Mondays — no hours-config shifting needed | ⬜ not run |
| 10 | Message capture | ✅ sponsorship-call script | ⬜ not run |
| 11 | Knowledge accuracy (**sheet-killer**) | ✅ five exact questions with five exact pack answers | ⬜ not run |
| 12 | Out-of-knowledge | ✅ "tasting menu" — plausible, deliberately absent from the pack | ⬜ not run |
| 13 | Forbidden topics | ✅ three probes: free meal, allergy clearance, card details | ⬜ not run |
| 14 | Reschedule/cancel | ✅ Sun 13:00 seeded; orphan/duplicate check specified | ⬜ not run |
| 15 | Details given unprompted | ✅ full one-breath script line | ⬜ not run |
| 16 | Accent / noisy line | ✅ instruction to run from a real street | ⬜ not run |
| 17 | Latency threshold | ✅ 3-call stopwatch protocol + dashboard median | ⬜ not run |
| 18 | Silent caller / hangup | ✅ both sub-cases scripted | ⬜ not run |
| 19 | Language switch | **N/A** — no second language LIVE (D10, `COMPATIBILITY.md` §3) | **N/A** |
| 20 | End-to-end evidence chain | ✅ next-day checklist; also captures total minutes for `F-02` | ⬜ not run |

**Nothing here is a pass.** The only honest reading of this table is that the script is ready and the phone has not rung.
