Skip to content

01 — Start

Markdown

Automatic model routing

Send "model": "auto" and Inferit picks the model. Among the models that have a healthy seller right now, it keeps the ones that fit the request, scores each on a published quality ranking against its price for this exact request, and sends the request to the best value, at that model's cheapest healthy seller.

What it does

If the seller it picked fails before the first byte, the request moves to the next seller, then to the next model. The response names the model that served it: the JSON model field and the x-inferit-model header. The router chooses from the request's size and features, never from what the prompt says.

auto is a model id like any other, so OpenAI-compatible clients can send it: the OpenAI SDKs and OpenCode. In a client that has its own Auto mode, such as Cursor, name the model inferit/auto so the two cannot be confused. Everything else works as it does for a named model: streaming, API keys and their limits, x402, the minimum discount.

a request, with the routing headers
curl https://api-production-c74b9.up.railway.app/v1/chat/completions \
  -H "Authorization: Bearer $INFERIT_API_KEY" \
  -H "content-type: application/json" \
  -d '{"model": "qwen/qwen-2.5-14b-instruct", "messages": [{"role": "user", "content": "Summarise this diff."}]}' \
  -D - -o reply.json | grep -i '^x-inferit-'
# x-inferit-model: the model that served it (with auto, the one Inferit picked)
# x-inferit-route: auto:balanced (direct for a named model)
# x-inferit-auto-choice: 1 (model auto only)

Live now

The models model auto can choose from at this moment, in its order, with each one's AA index and its price for the reference request. The card reads GET /v1/routing and the market when this page loads; until it has a live pick it shows the example request.

Routed

Example

POST /v1/chat/completions

"model": "auto"

A 1,000-token prompt, a 500-token answer

Inferit picks the model, then the seller.

How it picks

Policies

modelPolicyWhat it picks
auto (also auto:balanced, inferit/auto)best valueBest value: each candidate's quality score weighed against its price for this request. A model that costs twice as much for the request must score 6 points higher to win.
auto:cheap (also inferit/auto:cheap)lowest priceLowest price for this request among the candidates; the quality score breaks ties.
auto:best (also inferit/auto:best)highest rankedThe highest quality score among the candidates; price breaks ties.
  • Ids are matched without regard to case, after trimming. The prefix inferit/ is accepted (inferit/auto, inferit/auto:cheap) for clients that want a vendor-qualified id.
  • Any other id that starts with auto: or inferit/auto answers 400 invalid_auto_policy. No catalog model is named auto, and a seller cannot list it.
  • GET /v1/models lists the three auto rows first (owned_by: "inferit"), each with an auto object: policy, pick, candidates and ranking. The pick is computed for a reference request of 3,000 input and 1,000 output tokens, so it is indicative: a request with other needs may get another model.

How a model is chosen

How model auto picks

How model auto picks, illustrated for a 1,000-token prompt and a 500-token answer that sends tools. Fit: Models A, B and C fit; Model D does not, it has no tool support. Score: Model A scores 40 at $0.000168, Model B 52 at $0.0006 and Model C 58 at $0.0024; auto picks Model B, the best value, auto:cheap picks Model A and auto:best picks Model C. Fail over: the request goes to Model B’s cheapest healthy seller, then its next seller, then Model A. The x-inferit-model header names the model that answered.

1. Fit

A catalog model is a candidate for a request when every check passes. They run in this order; the first that fails is the model's reason in an error:

ReasonThe check
unscoredThe model has a row in the quality ranking (below). Unscored models stay routable by name, never by auto.
not_servableThe model can be sold on Inferit: open weights, or resale through the allowlist. Closed models are reference rows.
context_too_smallThe request fits the model's context window: the input estimate plus 10%, plus the output limit (max_tokens or max_completion_tokens, the smaller, or 1,024). And the window is at least min_context, when you send it.
needs_toolsA request with tools or functions goes only to models that support tool calls.
unsupported_inputA non-text part other than an image (audio, a file) excludes every model in this version: name a model to send it.
needs_visionA request with an image part goes only to models that accept images.
below_min_qualityThe model scores at least min_quality, when you send it.
no_offerThe model has at least one active offer on your rail. Demo offers never take paid traffic. A seller may serve a smaller context window than the catalog's: when every offer is too small for the request, the reason is context_too_small.
no_healthy_offerAt least one of its offers passes the live health checks and its daily cap.
below_min_discountAt least one healthy offer meets your minimum discount (/min{N}, x-min-discount or the key setting).
over_max_priceAt least one of those is within max_price_per_m, when you send it.

2. Score

For each candidate m, C is the all-in estimate of this request at the model's cheapest healthy offer within your limits (the input estimate and the output limit, in µ), and Q is its quality score, an integer from 0 to 100. With C_min the lowest candidate estimate:

text
value(m) = Q(m) − λ × log2( max(C(m), 1) / max(C_min, 1) )
PolicyλOrder (each column breaks the previous one's ties)
auto (best value)λ 6Order: value, highest first · Q · C, lowest first · model id
auto:best (highest ranked)λ 0Order: Q, highest first · C, lowest first · model id
auto:cheap (lowest price)λ ∞Order: C, lowest first · Q, highest first · model id

In words: with auto, a model that costs twice as much for this request must score 6 points higher to win. λ is a product setting, not a measurement; a change to it gets a changelog line and an update to this page.

Worked example. Model A scores 40 and costs 168 µ for the request; model B scores 52 and costs 600 µ. log2(600 / 168) = 1.84, so B's value is 52 − 11.0 = 41.0, above A's 40: auto picks B, auto:cheap picks A, auto:best picks B. Had B cost 1 000 µ, its value would be 52 − 6 × 2.57 = 36.6, below 40, and auto would pick A.

3. Attempts

The request goes to the first model's offers in rank order (the cheapest healthy seller first), then to the second model's, then to the third's: at most 3 models, and at most 3 attempts in all by default. A failing seller is replaced by the same model's next seller before the request moves to another model. Failover happens only before the first byte; nothing is retried once content has started streaming. The same market, body and headers always give the same order.

At equal price, a seller that answers with no cold start goes before one that has to wake up; it never goes before a cheaper one. A seller that sleeps when idle gets its declared wake time on top of the first-byte wait, 140 seconds in all at most (availability).

x-inferit-objective: latency keeps its meaning inside a model (it orders that model's sellers within a 10% price band). It does not change which model is chosen.

Hints

Optional limits, sent as body fields (in the OpenAI SDKs: extra_body) or as headers, for clients that cannot add body fields but can send provider headers, such as OpenCode. If both are sent, the stricter value wins.

Body fieldHeaderMeaning
max_price_per_mx-inferit-max-price-per-mInteger µ per 1M tokens, all-in, 1 to 1012 (a number or a decimal string). Skips any offer whose all-in price, blended on this request's estimated token mix, is above it. Works with auto and with named models.
min_contextx-inferit-min-contextInteger tokens, 1 to 100,000,000. Only models whose context window is at least this. Auto only.
min_qualityx-inferit-min-qualityInteger, 0 to 100. Only models whose quality score is at least this (the ranking). Auto only.
  • Strictest wins: max_price_per_m takes the smaller value, min_context and min_quality the larger.
  • A bad value answers 400 invalid_request, naming the field. min_context or min_quality with a named model answers 400 invalid_request: they apply to model auto only.
  • Hint fields never reach a seller: the API forwards only the chat fields a seller needs.
  • Units: µ, one millionth of a dollar; 400000 per 1M tokens is $0.40 per 1M tokens.
Python · OpenAI SDK, with hints
import os
from openai import OpenAI

inferit = OpenAI(base_url="https://api-production-c74b9.up.railway.app/v1", api_key=os.environ["INFERIT_API_KEY"])
reply = inferit.chat.completions.create(
    model="auto:cheap",
    messages=[{"role": "user", "content": "Extract the dates from this text: ..."}],
    # Hints are body fields; clients that cannot add body fields send the x-inferit-* headers instead.
    extra_body={"min_context": 32000, "min_quality": 5, "max_price_per_m": 400000},
)
print(reply.model)  # the model that served it
curl · hints as headers
curl https://api-production-c74b9.up.railway.app/v1/chat/completions \
  -H "Authorization: Bearer $INFERIT_API_KEY" \
  -H "content-type: application/json" \
  -H "x-inferit-min-quality: 5" \
  -H "x-inferit-max-price-per-m: 400000" \
  -d '{"model": "auto", "messages": [{"role": "user", "content": "Hello"}]}'

Hints only narrow what auto can choose from now (candidates_now in GET /v1/routing). When that list is empty, name a model: max_price_per_m works with a named model too.

curl · a named model, with max_price_per_m
curl https://api-production-c74b9.up.railway.app/v1/chat/completions \
  -H "Authorization: Bearer $INFERIT_API_KEY" \
  -H "content-type: application/json" \
  -H "x-inferit-max-price-per-m: 400000" \
  -d '{"model": "qwen/qwen-2.5-14b-instruct", "messages": [{"role": "user", "content": "Hello"}]}'

What the response says

WhereWhat
x-inferit-modelThe catalog id of the model that served the request (named models too). The authoritative answer.
x-inferit-routedirect for a named model; auto:balanced, auto:cheap or auto:best for auto.
x-inferit-auto-choiceAuto only: 1 when the first choice served it; 2 or more when failover moved the request to another model.
JSON modelThe serving model's catalog id. In a stream each chunk's model is the seller's name for the same model (often an alias); the header is authoritative.
GET /v1/usageEach row's model is the serving model; requestedModel is auto:balanced, auto:cheap or auto:best for auto, and null for a named model. The CSV export adds it as its last column.

The other headers are unchanged: x-inferit-attempts, x-inferit-offer, x-inferit-buyer-cost-micro (all response headers). Browsers can read all of them (CORS).

x402 and auto

  • Without a key, the 402 quote is priced for the model auto picks: the worst case of this body at that model's first offer, never below 10 000 µ ($0.01). The 402 carries x-inferit-model and x-inferit-route, and its message names the model.
  • The paid retry tries that model first. If it stopped being a candidate between the quote and the payment, the next candidates are tried, each only if its worst case fits the payment. If none fits, the answer is 503 no_available_offers, and the payment stays in your balance.
  • The model is chosen before any payment is taken: when no model can serve the request, the call without a key gets 503 no_eligible_model instead of a 402, and a paid retry gets it before its payment is settled.
  • The quote is pinned to the body and the routing headers, the hint headers included: a retry with a different hint gets a new 402.

Agents: agents.md has the whole x402 exchange.

Errors

StatusCodeWhenretry-after
400invalid_auto_policyAn id that starts with auto: or inferit/auto but names no policy: "Unknown auto policy `auto:fast`; use auto, auto:cheap or auto:best."no
400invalid_requestA hint with a bad value (the message names the field and its range), or min_context / min_quality with a named modelno
503no_eligible_modelNo model can serve this request right now. The message says why, and error.reasons counts each reason (and lists the models for the offer reasons).only when health is the only cause
503no_available_offersCandidates exist, but every attempt failed before the first byte; or, with x402, no candidate fits the payment (the payment stays in your balance).10 s

error.reasons counts every model considered by its first failing check (the reasons above), and error.reasons.models lists the models for no_offer, no_healthy_offer, below_min_discount and over_max_price.

503 no_eligible_model (example)
{
  "error": {
    "message": "No model can serve this request right now: 2 have no healthy seller, 1 has a context window under 300,000 tokens, 16 have no seller on the whitechain rail, 7 have no published quality score.",
    "type": "service_unavailable",
    "code": "no_eligible_model",
    "reasons": {
      "considered": 26, "unscored": 7, "not_servable": 0, "context_too_small": 1, "needs_tools": 0, "needs_vision": 0,
      "unsupported_input": 0, "below_min_quality": 0, "no_offer": 16, "no_healthy_offer": 2, "below_min_discount": 0, "over_max_price": 0,
      "models": { "no_offer": ["…"], "no_healthy_offer": ["…", "…"], "below_min_discount": [], "over_max_price": [] }
    }
  }
}

The quality ranking

Read on 2026-10-06. score = the model's Artificial Analysis Intelligence Index, rounded to the nearest integer. Each catalog model is matched to the row the source publishes under the model's own id (its default configuration, e.g. 'Qwen3.6 27B (Reasoning)'); sourceName and sourceId record the match. Models the source does not list, or lists only under a different release than the catalog's, are left out (no fallback map is used in this table). estimatedBySource marks values the source itself labels as estimated.

A score describes the model as its vendor released it. Sellers often serve a quantized build (for example 4-bit or 8-bit weights, shown as the offer's quantization when the seller declares it), which can answer somewhat differently; the ranking does not measure each seller's build.

  • Artificial Analysis Intelligence Index, read 2026-10-06: artificialanalysis.ai/leaderboards/models. Version: v4.3 (the leaderboard label; the methodology note on artificialanalysis.ai/models reads v4.3.2).
Model (catalog id)ScoreName at the source
moonshotai/kimi-k2.627Kimi K2.6 (Reasoning)
z-ai/glm-5.126GLM-5.1 (Reasoning)
moonshotai/kimi-k2-thinking22Kimi K2 Thinking (the source marks it estimated)
qwen/qwen3.6-27b21Qwen3.6 27B (Reasoning)
qwen/qwen3.6-35b-a3b18Qwen3.6 35B A3B (Reasoning)
deepseek/deepseek-v3.216DeepSeek V3.2 (Non-reasoning) (the source marks it estimated)
google/gemma-4-31b-it15Gemma 4 31B (Reasoning)
z-ai/glm-4.615GLM-4.6 (Non-reasoning) (the source marks it estimated)
openai/gpt-oss-120b12gpt-oss-120b (High)
qwen/qwen3-235b-a22b-250712Qwen3 235B A22B 2507 Instruct (the source marks it estimated)
qwen/qwen3-coder12Qwen3 Coder 480B A35B Instruct (the source marks it estimated)
meta-llama/llama-4-maverick10Llama 4 Maverick (the source marks it estimated)
openai/gpt-oss-20b9gpt-oss-20b (High)
meta-llama/llama-3.3-70b-instruct8Llama 3.3 Instruct 70B (the source marks it estimated)
mistralai/mistral-small-3.2-24b-instruct8Mistral Small 3.2 (the source marks it estimated)
qwen/qwen-2.5-72b-instruct8Qwen2.5 Instruct 72B (the source marks it estimated)
meta-llama/llama-3.1-8b-instruct7Llama 3.1 Instruct 8B (the source marks it estimated)
qwen/qwen3-14b7Qwen3 14B (Non-reasoning) (the source marks it estimated)
google/gemma-3-27b-it5Gemma 3 27B Instruct

Catalog models without a score (auto never routes to them; name them to use them):

  • qwen/qwen-2.5-14b-instruct: Not listed by Artificial Analysis (it lists Qwen2.5 Instruct 72B, Qwen2.5 Coder 32B and 7B) nor by the LMArena text leaderboard (checked 2026-10-06).
  • mistralai/mistral-nemo: Not listed with an Intelligence Index value by Artificial Analysis.
  • deepseek/deepseek-v4-pro: Catalog release is 0423 (OpenRouter); the source lists V4 Pro 0424 and V4 Pro 0813 only, so the match is not certain.
  • deepseek/deepseek-v4-flash: Catalog release is 0423 (OpenRouter); the source lists V4 Flash 0420 and V4 Flash 0731 only, so the match is not certain.
  • anthropic/claude-sonnet-5.5: Closed model: a reference row, not servable; needs no score.
  • openai/gpt-6.1-sol: Closed model: a reference row, not servable; needs no score.
  • google/gemini-3.8-flash: Closed model: a reference row, not servable; needs no score.

The table is kept with the catalog: when the catalog takes a new price snapshot, the scores are read again the same day. The same table, and the models auto can choose from now, are at GET /v1/routing as JSON (API reference).

What it does not do

  • It chooses from the request's size and features, never from what it says.
  • Tools and images are model-level flags from the catalog snapshot: a model supports the feature, and a given seller's server may still not.
  • Hints are per request. A key cannot carry a default policy or price limit in this version.
  • Audio and file parts go only to a model you name.

Name a model instead

Send a catalog id or an alias (GET /v1/models lists them) and Inferit picks only the seller: the cheapest healthy seller of that model, as it always has. max_price_per_m works with named models too: when no healthy offer of the model is within it, the answer is 503 no_available_offers, with retry-after. When no seller of the model serves a context window big enough for the request, the answer is 400 context_length_exceeded. min_context and min_quality do not apply to a named model. A named model needs no score: unscored models are served by name, so when auto answers no_eligible_model, a model with a live offer (a best_price in GET /v1/models) still answers by name.