Step-by-step guide for operations and customer teams who want to automate customer replies with AI without reputation risk. Assumes shared inbox or tickets — hospitality, B2B services, local commerce. For process framing see AI process automation guide; for agent architecture see What is an AI Operating Layer.
Step 1 — 20-minute audit: one flow, one metric
Pick one channel (e.g. info@) and one metric (hours/week on replies + current response time). Current work starts with a Web & eCommerce quote.
Bring: screenshot or description of the painful queue, rough volume (messages/day), systems (mail, CRM), peak constraints.
Leave with: fit yes/no, hour order-of-magnitude, next step (Discovery, connector, wait).
Step 2 — Taxonomy with your team
Agree 6–10 real categories: enquiry, complaint, cancellation, quote, spam, supplier. Mark categories that never auto-send.
| Category | Agent action | Human required? |
|---|---|---|
| General enquiry | Draft from FAQ | Yes — approve send |
| Cancellation today | Urgent tag + summary | Yes — always |
| Spam | Archive per rules | No — weekly audit |
| Supplier invoice | Route to back-office | Validate amounts |
Without taxonomy, the queue fills with noise.
Step 3 — Templates and policies
AI drafts from your tone and rules — not invented refund policy. Fix contradictory FAQs first.
Deliverable: approved template set + forbidden replies (legal, unauthorised discounts).
| Template type | Fixed content | Agent fills |
|---|---|---|
| Acknowledgement | SLA wording | Name, ticket ref |
| Standard FAQ | Verified facts | Language |
| Escalation | Apology + timeframe | Reason code |
Review quarterly — cancellation rules change more often than teams notice.
Step 4 — Tools with minimum privilege
Read: mail/tickets, CRM, doc store. Write: drafts, labels, internal notes — not unsupervised external send in phase 1.
Prefer MCP/API per agent automation practice and Operating Layer.
Step 5 — Open default, frontier when justified
Classification and standard drafts → open models. Delicate complaints or long chains → frontier if delta pays.
Document which categories use which tier — monthly review. Illustrative 6–12× frontier premium: open vs frontier cost.
Step 6 — Trial with success criteria
One flow on your accounts, 2–4 weeks real traffic:
- ↓ time to first useful action
- Target illustrative: majority of drafts approved with minor edits (agreed in Discovery — not a guarantee)
- Auto-escalate when context missing
Discovery (€1,950) first if no written plan. Weekly 30-minute reviews: why were drafts rejected?
Step 7 — Human queue and operation
| Draft state | Operator action |
|---|---|
| Good as-is | Send |
| Minor edit | Send after tweak |
| Missing context | Escalate |
| Policy doubt | Reject, write manually |
| Out of scope | Archive |
Roles: approver, escalator, pause authority. Training: 1–2 sessions. Run retainer from €1,850/month when multiple flows need care.
Step 8 — Measure weeks 3–4
Track rejection reasons — usually taxonomy, template, or tool gap, not “longer prompt.”
Compare before/after thinking on triage: minutes vs hours is the common win — see the process automation guide.
Costs (published + illustrative)
| Item | Starting point |
|---|---|
| Discovery | €1,950 |
| Pilot | from €4,800 |
| Inference (open, illustrative) | tens €/month |
Full breakdown: AI agent implementation cost.
Mistakes to avoid
- Skipping taxonomy and templates
- Vanity metric “conversations handled”
- Sending because text “sounds right”
- One expensive model for all volume
- No internal owner — queue dies
vs connectors
Pure “form → spreadsheet” without free text? See agents vs RPA vs Zapier — connector may win.
Ask for a Web & eCommerce quote — we will tell you if reply automation is the right tool, or something duller and cheaper.
Frequently asked questions
Can we send replies without human review?
We do not recommend unsupervised external sends in phase one. Very narrow subsets (e.g. document receipt acks) may automate after evaluation — not general consultative replies.
What systems do we need?
Minimum: shared inbox or tickets + curated templates. Better: CRM history. APIs or MCP agreed in Discovery — avoid brittle screen scraping.
How long until a pilot is live?
Weeks after Discovery on your accounts — not months of abstract consulting. Legacy integrations extend the calendar.
How do we know it works?
Time to first useful response, approval-without-major-edit rate, escalations for missing context, abandoned queue rate.
