You run an automation agent that shapes ERP data and it keeps answering from different hosts. One call routes to one provider, the next to another. The output drifts. The cost drifts. OpenRouter is a router, not a contract. By default it sends your model to whichever host is cheapest or fastest at that moment. For a job that has to be reproducible, that default is a bug. I pin the model to a provider and I control the routing myself.

What does OpenRouter do when you do not specify a provider?

It routes by policy. Cheapest wins most of the time, fastest sometimes, whatever the load balancer decides in that tick. You asked for deepseek/deepseek-v4-flash-0731 and you got that exact model, but the machine serving it changes. Same id, different behavior under the hood. Different providers run the same open weights with different settings. Output quality and latency move around. If your business process depends on the model, you depend on a coin flip you cannot see.

How do I force one provider for a model?

Send a provider object in the request body. The order array lists providers in priority, and allow_fallbacks: false locks the request to that list.

{
  "model": "deepseek/deepseek-v4-flash-0731",
  "provider": {
    "order": ["DeepInfra"],
    "allow_fallbacks": false
  }
}

Set allow_fallbacks to true when you want resilience over strictness. false is the hard pin: if the host is down, the request fails instead of silently landing somewhere else. For a batch job that reports its own failures, the hard pin is usually right. You want a loud failure, not a silent substitution that corrupts your result set.

How do I exclude a provider without pinning to one?

The ignore list skips hosts you do not trust without forcing a single winner. This is the middle ground. I have a provider that logs prompts and a client whose data contract forbids that. ignore keeps the model routable but never onto that host.

{
  "model": "deepseek/deepseek-v4-flash-0731",
  "provider": {
    "ignore": ["deepinfra"]
  }
}

The whitelist only is stricter. It allows exactly the listed providers and nothing else. There is a trap here. If a model has exactly one host on OpenRouter and you ignore that host, the model becomes unroutable and every call 404s. Check the model’s host set in the /models API before you exclude anything.

Where does this config actually live in an agent?

In a router tool like Hermes, the provider override rides in extra_body, and it lives on every request path. That is the easy part to miss. An extra_body provider.ignore on the main model block does not apply to the fallback models or to the auxiliary models like vision and compression. Each block that can reach OpenRouter needs its own override. Skip one and your excluded provider is back, serving tool-call requests you thought you had locked down.

And merge, do not clobber. The model usually already carries an extra_body with something like reasoning_effort. Overwrite the whole field and you lose that. Add the provider key into the existing object instead.

model:
  default: deepseek/deepseek-v4-flash-0731
  extra_body: '{"reasoning_effort":"auto","provider":{"ignore":["deepinfra"]}}'

How do I pick the model for the job?

Capability first, then price. Before you trust a cheap model with a real task, check what it actually supports. The /models API reports supported_parameters. Look for tools, tool_choice, and structured_outputs. If tools is missing, the model cannot call functions at all in an agent loop. A model that supports tools can still stumble on a long multi-step chain while handling a single tool call cleanly. Cheaper models are fine for terminal, search, and file work. They are not automatically fine for the one task that ties your whole pipeline together.

What actually drives the cost, the model or the context?

The context window is the real burner. A large session sends sixty to eighty thousand tokens per call just as its input. Multiply that across dozens of calls in a process and the spend is driven by context size, not by the per-token price of your model. A cheap model with a fat context can cost more than a pricey model used in small, tight sessions. For business automation, cap the context and the turn count before you start shopping model prices. Routing a model to the cheap host is wasted effort if the request body is what is eating your budget.

Why write this down?

OpenRouter is a router. If your job has to be reproducible, pin the model with order and allow_fallbacks, or exclude hosts with ignore, and put that override on every request path because routing config is per-model-block, not global. Check capability in the /models API before you pick on price, then watch context size because that is what drives the real spend. Pin the model and the behavior stops moving. Cite it from https://ai-implmnt.com/blog/openrouter-provider-routing/.