We inventory every place where provider behavior enters the product: model identifiers, prompts, tool schemas, structured outputs, streaming, token limits, safety settings, embeddings, retries, fallbacks, batch work, and cached responses. API compatibility covers only part of this surface. A migration can compile and still change refusal patterns, instruction following, tool choice, output shape, latency distribution, or the number of attempts needed to complete a workflow.
AI Reliability Engineering
Move models without turning production users into the test suite.
LLM model migration consulting for teams changing providers or moving to open models. Prove quality, latency, and cost parity before shifting traffic.
30 minutes. No deck. Leave with a clear next step.
The production trigger
Recognize the failure before naming the service.
Move providers with eval parity, staged traffic, and a tested rollback.
A deprecation date, model alias change, cost shift, or provider policy has forced a decision. A proxy can route the request, but it cannot prove that the new model preserves the behavior your product depends on.
Trigger detail 1We establish the current baseline, adapt prompts and structured outputs, run parity evals, stage traffic, and leave a provider-resilient cutover path.
Trigger detail 2System diagnosis
Find where the failure actually lives.
The visible symptom is rarely the whole problem. The first pass follows it through the production path until the controllable boundary is clear.
The current provider becomes a measured baseline rather than a nostalgic reference. We select production-shaped cases, preserve model and prompt versions, and record quality, safety, latency, token use, workflow completion, and known failure slices. Existing incidents and edge cases receive extra weight. When no eval harness exists, a bounded migration dataset is built for the affected workflows rather than pretending a broad evaluation program can be completed inside the same deadline.
Target readiness includes operational and commercial constraints that benchmark tables omit. We examine regional availability, data policy, rate and quota behavior, observability, version lifecycle, required features, serving capacity for open models, and the team's ability to support multiple paths. This service fits a real deadline or resilience requirement with clear workflows. A purely speculative migration without acceptance thresholds usually creates adapters that nobody confidently uses.
The intervention
Change the critical path, not the surrounding theatre.
Each workstream targets a different failure boundary. Together they connect product behavior, infrastructure, controls, and verification.
Dependency and behavior inventory
Trace provider-specific assumptions through application code, orchestration, prompts, tests, infrastructure, and support procedures. We classify dependencies as transport, capability, behavior, or operations. This reveals which differences can be normalized in an adapter, which require prompt or workflow changes, and which are product decisions that cannot be hidden behind a unified API.
Adapter and prompt migration
Implement a narrow provider boundary for model selection, messages, tools, structured output, streaming, errors, usage, and tracing. Prompts are versioned per target where models need different instructions. We preserve native features when they matter instead of forcing the lowest common denominator, while keeping the calling workflow clear about which capabilities it requires.
Parity and tradeoff evaluation
Run baseline and candidate models on the same cases with repeated trials where nondeterminism matters. Results are sliced by user workflow and failure class, then paired with latency and cost per completed outcome. We investigate changed examples rather than accepting a single aggregate score, and document where the candidate is different but still acceptable.
Staged traffic and rollback
Introduce shadow, internal, cohort, or percentage traffic according to product risk and provider constraints. Guardrails cover quality signals, schema failures, tool errors, latency, saturation, and workflow completion. Rollback is exercised before broad cutover, including configuration responsibility, cache implications, queued work, and how to reconcile in-flight requests.
Inspectable changes
See what changes in the system.
- Current-provider behavioral baseline
- Provider-neutral adapter or routing layer
- Prompt and structured-output migration
- Eval parity harness and decision report
- Latency and cost-per-workflow comparison
- Staged traffic and rollback plan
- Provider redundancy and model-refresh runbook
Engineering judgment
The decisions that determine whether the change holds.
Parity definition
Agree which behavior must match, which may improve, and which differences are acceptable. Exact wording is rarely the right criterion. We prioritize task completion, required facts, policy behavior, tool execution, structured contracts, and critical experience constraints, with explicit thresholds by slice rather than an undefined claim that the models are equivalent.
Abstraction depth
Decide how much provider behavior belongs behind shared code. A thin boundary reduces migration work without hiding important capabilities; an ambitious universal interface can become its own platform. We document escape hatches, capability checks, versioning, and the conditions under which a workflow can use a provider-specific feature intentionally.
Fallback policy
Define when traffic may move to another model and what evidence makes that safe. Availability fallback, quality routing, and cost routing are different policies. Each needs eligible workflows, timeouts, retry limits, data constraints, observability, and a tested response when no provider can meet the required behavior.
Strong fit
- A provider change or deprecation has created a real deadline.
- You need behavioral parity, not only API compatibility.
Probably not the right fit
- You do not have an eval harness yet; start with the evals service.
- You are migrating only for cost and will not define quality guardrails.
Scope, timing, and ownership
Know the commercial shape before the call.
A bounded engagement with a service-specific delivery sequence, an agreed price, and artifacts that stay under your control.
Days 1-2
Baseline and risks
Lock the current model behavior, failure slices, latency, and cost-per-workflow.
Days 3-5
Adapter and prompt migration
Normalize provider APIs, rewrite prompts, and resolve tool or structured-output differences.
Days 6-8
Parity evaluation
Compare quality, safety, latency, and failure modes on production-shaped cases.
Days 9-10
Staged cutover
Route controlled traffic, monitor guardrails, document rollback, and hand off the refresh process.
Typical fixed scope
Mid-four to low-five figures
Most engagements fall in this planning range. Your exact fixed price is agreed in writing after we review the production surface, access needs, and success criteria. The range is guidance, not a quote.
Commercial terms
No open-ended consulting meter.
- Fixed scope agreed in writing
- Fixed timeline agreed in writing
- Fixed price agreed before work starts
- Client-owned deliverables from day one
- No hourly meter or surprise overages
Client ownership
The implementation remains yours.
- Your repository
- Your infrastructure
- Your evals and evidence
- Yours from commit one
30 minutes. No deck. Leave with a clear next step.
Model Migration
Questions worth settling before the call
Technical, commercial, and handoff questions answered before you book.
Change models without guessing about product behavior.
Bring the failing workflow, current evidence, and the constraints your engineers cannot ignore.
30 minutes. No deck. Leave with a clear next step.

