Most teams that say "we're building AI" are actually building a pipeline, and the distinction matters enormously for cost and timeline. Choosing a model is an afternoon. Building the pipeline around it is the project.
The model is the smallest part
When a demo works in a notebook and then dies on contact with production, it is almost never the model's fault. The model did what it always did. What changed is that real inputs arrived — malformed PDFs, empty fields, a customer name with an apostrophe, a 400-page contract, a request at 3am when the upstream API was down.
A pipeline is the machinery that handles all of that so the model only ever sees inputs it can actually reason about, and so the output is checked before anyone acts on it.
Rule of thumb. In the engagements we run, model selection and prompt work account for roughly 10–15% of effort. Data access, transformation, evaluation, and integration account for the rest.
The stages of an AI pipeline
Almost every production pipeline we build has the same seven stages, whether it is classifying support tickets or grading collectible cards with industrial lasers.
1. Ingestion
Getting the data out of wherever it actually lives — a CRM, a shared drive, an SFTP dump, a database replica, a scanner, a camera. This stage is where most timelines slip, because access is a political problem as often as a technical one.
2. Normalization
Turning heterogeneous input into one predictable shape. PDFs become text with layout preserved. Dates become ISO strings. Currencies get a unit. Images get resized and colour-corrected. Nothing downstream should ever have to ask what format it is looking at.
3. Enrichment and retrieval
Attaching the context the model needs to answer correctly. For document work this usually means retrieval-augmented generation — pulling the handful of relevant passages out of a large corpus and putting them in the prompt. For operational work it might mean joining a customer record, an order history, and an entitlement.
4. Inference
The actual model call, or several. Production pipelines frequently use more than one model: a small cheap one to classify and route, a larger one only for the cases that need it. This is the single biggest lever on inference cost.
5. Validation
Checking the output before anybody trusts it. Schema validation, business-rule checks, confidence thresholds, and cross-checks against the source. A pipeline without a validation stage is a demo.
6. Action
Writing the result somewhere it does work — creating the ticket, posting the journal entry, flagging the claim, updating the record. This is where value is actually realized, and where governance matters most.
7. Observation
Logging every input, output, decision, and cost so you can debug failures, measure quality over time, and prove what happened when someone asks.
Batch, streaming, or on-demand
The stages stay the same; the trigger changes.
| Pattern | Triggered by | Typical latency | Good for |
|---|---|---|---|
| Batch | A schedule | Minutes to hours | Back-office document processing, nightly reconciliation, bulk enrichment |
| Streaming | An event arriving | Seconds | Transaction monitoring, alerting, live moderation |
| On-demand | A user or system request | Under a few seconds | Copilots, search, chat, anything a human is waiting on |
Teams routinely over-engineer for real-time when the business process is daily. If the humans downstream work a morning queue, a batch pipeline that runs at 6am is cheaper, simpler, and easier to audit.
Why pipelines fail
Across the projects we have been brought in to rescue, the same five causes come up.
- No evaluation set. The team cannot tell whether a change made things better or worse, so every iteration is a guess. Build the eval set before you tune anything.
- Data access was assumed. The pipeline design presumed clean API access to a system that only exports a nightly CSV.
- No human-in-the-loop path. The pipeline handles the 85% it can and has nowhere to put the other 15%, so it blocks entirely.
- Cost discovered late. Running one large model over every record was fine for the 500-row sample and ruinous at 4 million rows.
- No owner. The pipeline shipped, the consultants left, the prompt drifted, and nobody was accountable for quality.
Designing one that survives
Start from the action, not the data. Ask what decision or task you want automated, then work backwards to the minimum context required to make that decision correctly. Teams that start from "we have a lot of data, what can AI do with it" tend to build impressive systems nobody uses.
Then apply three constraints from day one:
- Define done numerically. "Correctly extracts the invoice total on 98% of a held-out set of 500 real invoices" is a specification. "Improves invoice processing" is a wish.
- Route by confidence, not by hope. Decide in advance what score sends a case to a human, and staff for that volume.
- Instrument before you scale. You cannot optimize cost or quality on a pipeline you cannot observe.
Frequently asked questions
What is the difference between an AI pipeline and an ML pipeline?
They describe the same architecture, but the terms carry different assumptions. "ML pipeline" usually implies you are training a model on your own data, so the pipeline includes feature engineering, training, and model versioning. "AI pipeline" more often refers to orchestrating one or more pretrained foundation models, where the work shifts toward retrieval, prompting, validation, and routing rather than training.
How long does it take to build an AI pipeline?
A narrow, well-scoped pipeline against accessible data typically takes six to twelve weeks to reach production. The variance is almost entirely in data access and evaluation, not model work. If the data requires new integrations, legal review, or cleansing, expect three to six months.
Do I need a vector database for an AI pipeline?
Only if the pipeline needs to search unstructured content semantically. If your retrieval step is a straightforward lookup by ID, customer, or date, a normal database is faster, cheaper, and easier to reason about. Vector search earns its complexity when users ask questions in natural language over a large document corpus.
Can one pipeline serve multiple use cases?
The ingestion and normalization stages usually can and should be shared — that is where the expensive, reusable engineering lives. The retrieval, prompting, validation, and action stages are specific to each use case. Teams that try to build one universal pipeline for everything generally end up with something that serves nothing well.
What does an AI pipeline cost to run?
Ongoing cost splits into inference, infrastructure, and human review. Inference dominates for high-volume text work and is highly controllable through model routing and caching — see our breakdown of LLM API pricing and reducing inference cost. Human review is the cost teams forget to budget.
