This is the question we hear most in audits: what does AI agent implementation actually cost in production, not in a demo. We use published starting points from Yerbabuena Digital and ranges clearly labelled illustrative where your number depends on systems and volume. For the full agent framework, see our Spanish-market pillar guide (ES) or the EN Operating Layer definition; for tool choice, see agents vs RPA vs Zapier.
What each budget line includes (and excludes)
Before numbers, what market quotes often omit:
- Usually in a serious governed pilot: flow design, bounded tool connections, human gates, deployment on your accounts, agreed metrics, basic runbook.
- Usually not included without explicit scope: third-party licences (CRM, PMS, ERP), internal IT hours, master-data cleanup, company-wide training programmes.
Ask for deliverables per phase in writing — not only a total.
Phase breakdown (published starting points when this note was written)
Current public work is a Web & eCommerce quote or US → EMEA. The table is background, not a live offer.
| Phase | Includes | Starting point |
|---|---|---|
| Entry | 20-minute Efficiency Audit + written follow-up | Free |
| Diagnosis | Discovery Workshop — scoped plan, data, security, success criteria | €1,950 fixed |
| Governed pilot | Operating Layer Sprint — one production workflow with human oversight | from €4,800 |
| Full implementation | Multiple flows, integrations, operator training | from €12,000 (custom after Discovery) |
| Ongoing operation | Run retainer — monitoring, tuning, reliability | from €1,850/month |
No row is a closed quote without Discovery. US buyers: illustrative USD equivalents ~$2,100 Discovery, ~$5,200 pilot entry at typical FX — not a quote.
Inference cost per workflow (illustrative)
Variable model spend sits on top of project fees:
| Scenario | Assumptions | Order of magnitude / month |
|---|---|---|
| Inbox triage — medium volume | ~2,000 messages/month, mostly open models | Tens of € |
| Same flow — all frontier | Same volume, no routing | 6–12× open inference (see open stack) |
| Agent with 5–8 MCP tools | Includes logging + gateway | Add gateway hosting and trace storage |
Not measured savings guarantees — sizing guidance before budget sign-off.
Worked example (illustrative)
2,000 messages/month, 80% open / 20% frontier routing:
| Line | Illustrative / month |
|---|---|
| Open inference (80%) | ~€30–50 |
| Frontier inference (20%) | ~€40–80 |
| Gateway + logs | ~€10–30 |
| Total inference | ~€80–160 |
Route 100% frontier without criteria and multiply the open row by 6–12× — pilots die in finance review even when the demo dazzled.
Hidden costs Discovery should surface
- Data: contradictory FAQs, stale templates, inherited folder permissions
- Integration: no API, external IT or ERP vendor dependency
- Human review: displaced, not deleted — budget approval time
- Training: queue usage, escalation, emergency pause
- Model drift: re-evaluate when request types shift
Skipping these in the first proposal is why “it worked in the demo and failed in September.”
How to reduce total cost without cutting governance
- Open-by-default routing — frontier only where quality delta pays.
- One workflow in the Sprint — not three half-live flows.
- MCP/APIs over brittle UI automation.
- Phased scope — triage before auto-draft; draft before assisted send.
- Agreed metrics — do not pay for unused “autonomy.”
The AI Operating Layer centralises routing, traces, and gates so you do not re-pay the same integration for every experiment.
Year-one scenarios (illustrative, not quotes)
| Profile | Assumptions | Illustrative year-one order of magnitude |
|---|---|---|
| Minimum viable | Discovery + Sprint, one flow, open inference, no retainer | ~€7,000–9,000 + inference |
| Stable operation | Above + 6 months Run, second flow in Q3 | ~€18,000–25,000 (scope-dependent) |
| Discovery only | Written plan, pilot deferred | €1,950 |
Executive summary range: Discovery €1,950 + Sprint €4,800 + inference ~€600–1,200/year (open, medium volume) + optional Run €1,850/month for multi-flow continuity. Not a closed budget — a conversation anchor.
This article is refreshed quarterly alongside the Spanish version; the Updated date in the header reflects the last pricing review.
For numbers applied to your systems, ask for a Web & eCommerce quote or start with US → EMEA. The figures above are background from when this note was written.
Frequently asked questions
What does a production agent cost per month?
Beyond project fees: inference (often tens of €/month on open routing for medium volume), gateway hosting, and optionally a Run retainer from €1,850/month when multiple live flows need continuity. Illustrative until volume and routing are confirmed in Discovery.
Why do vendor quotes vary so much?
Because they mix demo, production, legacy integration, governance, and operations. A single number without written scope usually omits data prep, human review time, or API spend.
Does open source reduce total cost?
It lowers per-token inference when routing is disciplined; it does not remove integration, supervision, or operations. Frontier APIs can run 6–12× open for equivalent task classes — illustrative hybrid-stack reference.
When should we postpone spend?
When volume is below ~20 cases/month, there is no operational owner for an approval queue, or data access is undefined. An honest first conversation often returns 'connector first' or 'wait' — that is a valid outcome.
