From rented APIs to owned AI infrastructure, one workflow at a time.
The audit compares closed API, open-weight, private cloud, hybrid, and self-run paths against your actual workflow — then names which one wins on cost, risk, quality, and operational reality. Sometimes the right answer is no AI at all.
Savings model · Infrastructure path · EU-only
When each option fits
A quick filter. The audit walks the same dimensions in depth on your workflow and makes the ownership trade-off explicit.
| Option | When it fits | Risk | Cost | Complexity |
|---|---|---|---|---|
| Closed API | Best model quality, low setup | Data & vendor exposure | Variable | Low |
| Open-weight | More control, cheaper at scale | Quality varies | Medium | Medium |
| Private cloud | Sensitive data, ownership path | Ops burden | Medium / high | High |
| Self-run | High volume, in-house capability | Reliability & staffing | Lower at scale | Highest |
| No AI | Workflow does not justify AI | Lowest risk | Lowest | Low |
What each option actually is
Plain definitions. We use these consistently in the report.
Closed API
You call a frontier provider (OpenAI, Anthropic, Google, hosted Mistral, Azure OpenAI) over the network. Data leaves your perimeter under the provider's terms. Lowest day-one operational load; per-token billing; least control over model lifecycle.
Open-weight
You run an open-weight model family (Llama, Mistral, Qwen and similar) on infrastructure you choose. Higher ops load; fixed-ish infrastructure cost; full control over version and change cadence. Often the first serious step toward owning more of your AI stack.
Private cloud
A managed deployment of an open or closed model inside a dedicated tenant (e.g. Azure AI dedicated, EU sovereign clouds, on-prem appliances). Data stays inside a contractual boundary you can audit. Ops load depends on the offering, but the business gains a clearer path away from pure per-token dependency.
Self-run infrastructure
Your team runs the model serving layer, retrieval stack, monitoring, updates, and capacity planning on cloud or on-prem infrastructure. This can lower unit costs and increase control at scale, but only if volume, skills, and reliability needs justify the operational burden.
Hybrid
Different steps of the same workflow use different architectures. Often: sensitive summarisation private; long-tail edge cases via closed API; deterministic pre- and post-processing in your own code. Usually the right answer for non-trivial workflows and a practical bridge toward private infrastructure.
No AI
The workflow does not justify a model. A rule, a template, a removed step, or a better-designed form does the job for less cost and less risk. The audit will say so in writing when that is the honest answer.
What we do not claim
- Private or self-run deployment is always cheaper. Below a volume threshold, closed APIs often win fully loaded — we show the math.
- Owning infrastructure is automatically strategic. If it adds operational burden without lowering cost, risk, or dependency, we will say so.
- Open-weight is always higher quality. Frontier models still lead on some reasoning and language coverage; we benchmark what you actually run.
- Sensitive data always means private. Low-sensitivity workflows are often fine on closed APIs — we will say so.
- A substitute for your policies and legal review. Architecture is one input; your DPIA, contracts, and processes still govern outcomes.
- That AI is the right answer at all. Sometimes the recommendation is simplify or stop — replace AI with a rule, template, or removed step.
Turn the infrastructure question into a first-pass report.
Describe the workflow and get immediate fit, risk, and architecture direction before any call.