Practice · Intelligent Content Management

From folder chaos to a corpus AI can trust

Open-source ECM, metadata and workflows — so documents are findable, auditable and safe to use with agents. Built for mid-market teams that outgrew email attachments.

Misty mature forest canopy with soft filtered light

Architecture diagram from capture and classification into a metadata repository, workflows, agents over a trusted corpus, and ERP CRM search integrationsIngestCaptureClassify / extractRepositoryMetadata · RetentionWorkflowsAgent-augmented contentTrusted corpus onlyHuman validation queuesERPCRMSearch & audit
Capture → classify → governed repository → workflows → agents over a trusted corpus.
Irrigation system watering a green crop field

Yerbabuena Digital

We design content architectures and deliver on best-fit open-source ECM stacks (including Nuxeo and Alfresco): document types, metadata, capture, classification, retention and workflows. Representative engagements across regulated and operations-heavy organisations taught us the same lesson — AI only helps when the corpus is governed. This practice is heritage and referral-led; websites and stores start on Web & eCommerce.

Who this practice is for

Organisations where shared drives and inboxes are still the system of record — and audits or AI projects are catching up.

Regulated mid-market

Pain: Auditors ask for retention proof; staff cannot find contracts or claims files.

Fit: Metadata, retention and access control aligned to your sector — without a multi-year SI theatre.

Public sector & healthcare operations

Pain: Citizen or patient files spread across scans, email and legacy apps.

Fit: Capture workflows, ENI-aware integrations where required, and audit trails you can defend.

Operations-heavy industry

Pain: Approvals stuck in inboxes; no single view of document status.

Fit: Process automation with e-signature, SLA tracking and clean handoffs to ERP or CRM.

Govern the corpus, then let agents help

Enterprise content management is a core capability of our technical team. Representative prior-role work delivered document platforms for regulated and operations-heavy organisations — on open-source stacks including Nuxeo and Alfresco, and on commercial ECM when that was the right fit. Details stay under NDA; the patterns travel.

We help you design content architectures that scale: document types, metadata, retention, search and process automation — whether you extend an existing ECM, migrate off legacy shares, or stand up a modern repository.

The goal is not another shelf for PDFs. It is findable, auditable content that moves work forward — with capture, classification, approval and publication wired to how teams actually operate.

Once the corpus is trusted, agents can capture, classify, route and summarise with human validation — the same Operating Layer standards we apply to other practices. Websites and ecommerce packages are a different path; choose this practice when documents are the bottleneck.

Two current ways to start

This page describes a capability. Current work starts with Web & eCommerce packages for Southern European companies, or US → EMEA market entry for teams crossing the Atlantic.

See Web & eCommerceSee US → EMEA

Return on investment

Where the numbers come from

Illustrative scenarios based on typical ECM engagements — results vary by volume, legacy state and adoption.

Architecture diagram from capture and classification into a metadata repository, workflows, agents over a trusted corpus, and ERP CRM search integrationsIngestCaptureClassify / extractRepositoryMetadata · RetentionWorkflowsAgent-augmented contentTrusted corpus onlyHuman validation queuesERPCRMSearch & audit
Capture → classify → governed repository → workflows → agents over a trusted corpus.

Financial services — contract retrieval

Legal and operations spent hours locating signed contracts across email and file shares. Metadata model + migration to a governed repository.

Time to find a contract
−70% — Illustrative after search + metadata
Audit prep
Days → hours — Retention reports automated
Validation queue
< 48h SLA — Human-in-the-loop classification

Structured metadata pays back the first time audit or legal stops asking where the file is.

Public administration — citizen files

Expedientes split between scan batches, email and legacy apps. Unified capture and workflow with ENI-aware integrations.

Cycle time
−35% — Routing + notifications
Duplicate scans
−50% — Single ingestion pipeline
Go-live waves
Phased — No big-bang cutover

Phased migration avoids the lift-and-dump trap while teams see value early.

Talk with us

Capabilities

What we design and deliver

Right-sized for mid-market — architecture first, then platform, then agents.

Content architecture

Models that match how your business thinks about information.

  • Document types and folder taxonomy
  • Metadata schemas and controlled vocabularies
  • Search and discovery design
  • Migration planning from legacy shares

Capture & classification

Get content in reliably — then know what it is.

  • Scan, email and API ingestion patterns
  • AI-assisted classification and extraction
  • Human validation queues
  • Barcode and batch capture workflows

Process automation

Route content through approval, review and publication.

  • Workflow and task assignment
  • E-signature and approval chains
  • Notifications and SLA tracking
  • Integration with ERP, CRM and line-of-business apps

Agent-augmented content

Agents that capture, classify and route — with human validation.

  • Agent-driven capture and ingestion pipelines
  • Classification and extraction with validation queues
  • Routing and triage for inbound documents
  • Search and summarisation over the trusted corpus

Platform & integration

Build on what you have — or help you choose what fits.

  • Open-source ECM deployment and configuration
  • Nuxeo, Alfresco and current best-fit stacks
  • API and webhook integrations
  • SSO, directory sync, training and admin playbooks

What you walk away with

Deliverables by engagement

Open-source ECM, metadata and workflows — findable, auditable content for operations and compliance.

Discovery Workshop

  • Content source and volume inventory
  • Metadata model sketch
  • Compliance and retention drivers map
  • Migration or greenfield outline

Governed PoC / Pilot

  • Repository slice with capture pipeline
  • AI-assisted classification with validation queue
  • Search and retrieval proof
  • Retention rules draft

Full Implementation

  • Production ECM deployment (open-source stack)
  • Workflow automation and integrations
  • Phased migration with metadata mapping
  • Admin playbooks and training materials

Governed Retainer

  • Metadata governance and taxonomy updates
  • Classification tuning and queue health
  • Audit report preparation
  • Phase-two backlog planning

Use cases

Where this practice lands

Financial services

Contract and claims corpus

Challenge: Signed agreements scattered across email and shares; audit prep took days.

Outcome: Governed repository with metadata and retention reports (illustrative; details under NDA).

Public sector

Citizen file capture

Challenge: Scans and email created duplicate expedientes with no single status view.

Outcome: Unified capture, workflow and ENI-aware integrations in phased waves.

Industry & operations

Quality and supplier docs

Challenge: Approvals stalled; ERP had no reliable link to the latest PDF.

Outcome: Workflow + e-signature + ERP handoff; agents triage inbound packs with human validation.

Engagement

How a content management engagement runs

Structured like an enterprise programme, delivered with studio discipline.

Assess

We inventory content sources, volumes, compliance drivers and pain points. You get a gap analysis and a target architecture sketch.

Design

Metadata model, retention rules, workflow maps and migration waves. You approve before we touch production content.

Build & migrate

Repository configuration, capture pipelines and phased migration. Early waves prove search and process before full cutover.

Operate

Admin playbooks, monitoring, classification tuning and optional agent layers over the trusted corpus.

Before & after

What changes with governed content

Before

  • Email attachments and shared drives as the system of record
  • Hours spent searching for the latest signed version
  • Audits become fire drills
  • AI projects blocked because nobody trusts the corpus

After

  • Repository with metadata, retention and access control
  • Search that returns the right document first
  • Exportable audit evidence
  • Agents that capture and summarise only over trusted content
Diagram showing content flow from email and folders into a classified repository with search and retentionEmail & sharesUnclassifiedRepositoryMetadataSearchSearch+ retention
Governed repositories replace ad-hoc folders with findable, retained content.

Outcomes

What organisations typically see

Findability

Contracts, expedientes and claims files retrieved in seconds — not hunted across drives.

Audit readiness

Retention and access evidence you can export when auditors ask — without a fire drill.

Faster cycles

Approvals leave inboxes; SLAs and notifications keep work moving.

AI-ready corpus

Agents only see governed content — with validation queues and logs.

Works with your platforms

Nuxeo, Alfresco and best-fit open-source ECM today — plus ENI-aware integrations for public sector. We integrate with existing repositories; we do not resell vendor licenses.

Frequently asked questions

Frequently asked questions

Is this the same as your website packages?

No. Web & eCommerce is for marketing sites, shops and booking systems with fixed prices. This practice is for document platforms, metadata and workflows. If you need both, we sequence them — usually the public site first when cash flow matters.

Which ECM platforms do you work with?

We are strongest on open-source stacks including Nuxeo and Alfresco, and we work with commercial ECM when your estate already depends on it. We pick what fits — we do not force a rewrite.

Can mid-market teams afford this?

Yes when scoped as waves: assess, design a thin slice, migrate a pilot corpus, then expand. We avoid multi-year SI programmes that never reach production.

How do agents fit?

After the corpus is governed. Agents help with capture, classification, routing and summarisation — always with human validation on consequential actions.

Do you migrate from file shares?

Yes. We map metadata, run phased waves and keep a rollback path. Big-bang lift-and-dump is rarely the right move.

How does this connect to the AI Operating Layer?

A trusted corpus is often the missing input for production agents. We design content so Operating Layer workflows can use it safely — same guardrails, same evidence habits.

Yerbabuena Digital

Ready to plant something good together?

Tell us what season you are in. We will walk the field with you — honestly — and suggest the smallest next step that can take root.