01 — Start
MarkdownAutomatic model routing
Send "model": "auto" and Inferit picks the model. Among the models that have a healthy seller right now, it keeps the ones that fit the request, scores each on a published quality ranking against its price for this exact request, and sends the request to the best value, at that model's cheapest healthy seller.
What it does
If the seller it picked fails before the first byte, the request moves to the next seller, then to the next model. The response names the model that served it: the JSON model field and the x-inferit-model header. The router chooses from the request's size and features, never from what the prompt says.
auto is a model id like any other, so OpenAI-compatible clients can send it: the OpenAI SDKs and OpenCode. In a client that has its own Auto mode, such as Cursor, name the model inferit/auto so the two cannot be confused. Everything else works as it does for a named model: streaming, API keys and their limits, x402, the minimum discount.
a request, with the routing headers
curl https://api-production-c74b9.up.railway.app/v1/chat/completions \
-H "Authorization: Bearer $INFERIT_API_KEY" \
-H "content-type: application/json" \
-d '{"model": "qwen/qwen-2.5-14b-instruct", "messages": [{"role": "user", "content": "Summarise this diff."}]}' \
-D - -o reply.json | grep -i '^x-inferit-'
# x-inferit-model: the model that served it (with auto, the one Inferit picked)
# x-inferit-route: auto:balanced (direct for a named model)
# x-inferit-auto-choice: 1 (model auto only)Live now
The models model auto can choose from at this moment, in its order, with each one's AA index and its price for the reference request. The card reads GET /v1/routing and the market when this page loads; until it has a live pick it shows the example request.
RoutedOne request, routed
Example
POST /v1/chat/completions
"model": "auto"
A 1,000-token prompt, a 500-token answer
Inferit picks the model, then the seller.
How it picksLive example: the page asks the API which models auto can choose from now and which it picks for its reference request (GET /v1/routing), and shows each one's quality score, its price for that request and the pick's saving against the list price. Current prices: models.
Policies
| model | Policy | What it picks |
|---|---|---|
auto (also auto:balanced, inferit/auto) | best value | Best value: each candidate's quality score weighed against its price for this request. A model that costs twice as much for the request must score 6 points higher to win. |
auto:cheap (also inferit/auto:cheap) | lowest price | Lowest price for this request among the candidates; the quality score breaks ties. |
auto:best (also inferit/auto:best) | highest ranked | The highest quality score among the candidates; price breaks ties. |
- Ids are matched without regard to case, after trimming. The prefix
inferit/is accepted (inferit/auto,inferit/auto:cheap) for clients that want a vendor-qualified id. - Any other id that starts with
auto:orinferit/autoanswers400 invalid_auto_policy. No catalog model is namedauto, and a seller cannot list it. GET /v1/modelslists the three auto rows first (owned_by: "inferit"), each with anautoobject:policy,pick,candidatesandranking. The pick is computed for a reference request of 3,000 input and 1,000 output tokens, so it is indicative: a request with other needs may get another model.
How a model is chosen
How model auto picks
A 1,000-token prompt, a 500-token answer
How model auto picks, illustrated for a 1,000-token prompt and a 500-token answer that sends tools. Fit: Models A, B and C fit; Model D does not, it has no tool support. Score: Model A scores 40 at $0.000168, Model B 52 at $0.0006 and Model C 58 at $0.0024; auto picks Model B, the best value, auto:cheap picks Model A and auto:best picks Model C. Fail over: the request goes to Model B’s cheapest healthy seller, then its next seller, then Model A. The x-inferit-model header names the model that answered.
How model auto picks: it keeps the ranked models with a live seller that fit the request; weighs each one's quality score against its price for this request (auto picks the best value, auto:cheap the lowest price, auto:best the highest ranked); and if the chosen seller fails before answering, tries the next seller, then the next model. The response header x-inferit-model names the model.
1. Fit
A catalog model is a candidate for a request when every check passes. They run in this order; the first that fails is the model's reason in an error:
| Reason | The check |
|---|---|
unscored | The model has a row in the quality ranking (below). Unscored models stay routable by name, never by auto. |
not_servable | The model can be sold on Inferit: open weights, or resale through the allowlist. Closed models are reference rows. |
context_too_small | The request fits the model's context window: the input estimate plus 10%, plus the output limit (max_tokens or max_completion_tokens, the smaller, or 1,024). And the window is at least min_context, when you send it. |
needs_tools | A request with tools or functions goes only to models that support tool calls. |
unsupported_input | A non-text part other than an image (audio, a file) excludes every model in this version: name a model to send it. |
needs_vision | A request with an image part goes only to models that accept images. |
below_min_quality | The model scores at least min_quality, when you send it. |
no_offer | The model has at least one active offer on your rail. Demo offers never take paid traffic. A seller may serve a smaller context window than the catalog's: when every offer is too small for the request, the reason is context_too_small. |
no_healthy_offer | At least one of its offers passes the live health checks and its daily cap. |
below_min_discount | At least one healthy offer meets your minimum discount (/min{N}, x-min-discount or the key setting). |
over_max_price | At least one of those is within max_price_per_m, when you send it. |
2. Score
For each candidate m, C is the all-in estimate of this request at the model's cheapest healthy offer within your limits (the input estimate and the output limit, in µ), and Q is its quality score, an integer from 0 to 100. With C_min the lowest candidate estimate:
value(m) = Q(m) − λ × log2( max(C(m), 1) / max(C_min, 1) )| Policy | λ | Order (each column breaks the previous one's ties) |
|---|---|---|
auto (best value) | λ 6 | Order: value, highest first · Q · C, lowest first · model id |
auto:best (highest ranked) | λ 0 | Order: Q, highest first · C, lowest first · model id |
auto:cheap (lowest price) | λ ∞ | Order: C, lowest first · Q, highest first · model id |
In words: with auto, a model that costs twice as much for this request must score 6 points higher to win. λ is a product setting, not a measurement; a change to it gets a changelog line and an update to this page.
Worked example. Model A scores 40 and costs 168 µ for the request; model B scores 52 and costs 600 µ. log2(600 / 168) = 1.84, so B's value is 52 − 11.0 = 41.0, above A's 40: auto picks B, auto:cheap picks A, auto:best picks B. Had B cost 1 000 µ, its value would be 52 − 6 × 2.57 = 36.6, below 40, and auto would pick A.
3. Attempts
The request goes to the first model's offers in rank order (the cheapest healthy seller first), then to the second model's, then to the third's: at most 3 models, and at most 3 attempts in all by default. A failing seller is replaced by the same model's next seller before the request moves to another model. Failover happens only before the first byte; nothing is retried once content has started streaming. The same market, body and headers always give the same order.
At equal price, a seller that answers with no cold start goes before one that has to wake up; it never goes before a cheaper one. A seller that sleeps when idle gets its declared wake time on top of the first-byte wait, 140 seconds in all at most (availability).
x-inferit-objective: latency keeps its meaning inside a model (it orders that model's sellers within a 10% price band). It does not change which model is chosen.
Hints
Optional limits, sent as body fields (in the OpenAI SDKs: extra_body) or as headers, for clients that cannot add body fields but can send provider headers, such as OpenCode. If both are sent, the stricter value wins.
| Body field | Header | Meaning |
|---|---|---|
max_price_per_m | x-inferit-max-price-per-m | Integer µ per 1M tokens, all-in, 1 to 1012 (a number or a decimal string). Skips any offer whose all-in price, blended on this request's estimated token mix, is above it. Works with auto and with named models. |
min_context | x-inferit-min-context | Integer tokens, 1 to 100,000,000. Only models whose context window is at least this. Auto only. |
min_quality | x-inferit-min-quality | Integer, 0 to 100. Only models whose quality score is at least this (the ranking). Auto only. |
- Strictest wins:
max_price_per_mtakes the smaller value,min_contextandmin_qualitythe larger. - A bad value answers
400 invalid_request, naming the field.min_contextormin_qualitywith a named model answers400 invalid_request: they apply to model auto only. - Hint fields never reach a seller: the API forwards only the chat fields a seller needs.
- Units: µ, one millionth of a dollar;
400000per 1M tokens is $0.40 per 1M tokens.
Python · OpenAI SDK, with hints
import os
from openai import OpenAI
inferit = OpenAI(base_url="https://api-production-c74b9.up.railway.app/v1", api_key=os.environ["INFERIT_API_KEY"])
reply = inferit.chat.completions.create(
model="auto:cheap",
messages=[{"role": "user", "content": "Extract the dates from this text: ..."}],
# Hints are body fields; clients that cannot add body fields send the x-inferit-* headers instead.
extra_body={"min_context": 32000, "min_quality": 5, "max_price_per_m": 400000},
)
print(reply.model) # the model that served itcurl · hints as headers
curl https://api-production-c74b9.up.railway.app/v1/chat/completions \
-H "Authorization: Bearer $INFERIT_API_KEY" \
-H "content-type: application/json" \
-H "x-inferit-min-quality: 5" \
-H "x-inferit-max-price-per-m: 400000" \
-d '{"model": "auto", "messages": [{"role": "user", "content": "Hello"}]}'Hints only narrow what auto can choose from now (candidates_now in GET /v1/routing). When that list is empty, name a model: max_price_per_m works with a named model too.
curl · a named model, with max_price_per_m
curl https://api-production-c74b9.up.railway.app/v1/chat/completions \
-H "Authorization: Bearer $INFERIT_API_KEY" \
-H "content-type: application/json" \
-H "x-inferit-max-price-per-m: 400000" \
-d '{"model": "qwen/qwen-2.5-14b-instruct", "messages": [{"role": "user", "content": "Hello"}]}'What the response says
| Where | What |
|---|---|
x-inferit-model | The catalog id of the model that served the request (named models too). The authoritative answer. |
x-inferit-route | direct for a named model; auto:balanced, auto:cheap or auto:best for auto. |
x-inferit-auto-choice | Auto only: 1 when the first choice served it; 2 or more when failover moved the request to another model. |
JSON model | The serving model's catalog id. In a stream each chunk's model is the seller's name for the same model (often an alias); the header is authoritative. |
GET /v1/usage | Each row's model is the serving model; requestedModel is auto:balanced, auto:cheap or auto:best for auto, and null for a named model. The CSV export adds it as its last column. |
The other headers are unchanged: x-inferit-attempts, x-inferit-offer, x-inferit-buyer-cost-micro (all response headers). Browsers can read all of them (CORS).
x402 and auto
- Without a key, the 402 quote is priced for the model auto picks: the worst case of this body at that model's first offer, never below 10 000 µ ($0.01). The 402 carries
x-inferit-modelandx-inferit-route, and its message names the model. - The paid retry tries that model first. If it stopped being a candidate between the quote and the payment, the next candidates are tried, each only if its worst case fits the payment. If none fits, the answer is
503 no_available_offers, and the payment stays in your balance. - The model is chosen before any payment is taken: when no model can serve the request, the call without a key gets
503 no_eligible_modelinstead of a 402, and a paid retry gets it before its payment is settled. - The quote is pinned to the body and the routing headers, the hint headers included: a retry with a different hint gets a new 402.
Agents: agents.md has the whole x402 exchange.
Errors
| Status | Code | When | retry-after |
|---|---|---|---|
| 400 | invalid_auto_policy | An id that starts with auto: or inferit/auto but names no policy: "Unknown auto policy `auto:fast`; use auto, auto:cheap or auto:best." | no |
| 400 | invalid_request | A hint with a bad value (the message names the field and its range), or min_context / min_quality with a named model | no |
| 503 | no_eligible_model | No model can serve this request right now. The message says why, and error.reasons counts each reason (and lists the models for the offer reasons). | only when health is the only cause |
| 503 | no_available_offers | Candidates exist, but every attempt failed before the first byte; or, with x402, no candidate fits the payment (the payment stays in your balance). | 10 s |
error.reasons counts every model considered by its first failing check (the reasons above), and error.reasons.models lists the models for no_offer, no_healthy_offer, below_min_discount and over_max_price.
503 no_eligible_model (example)
{
"error": {
"message": "No model can serve this request right now: 2 have no healthy seller, 1 has a context window under 300,000 tokens, 16 have no seller on the whitechain rail, 7 have no published quality score.",
"type": "service_unavailable",
"code": "no_eligible_model",
"reasons": {
"considered": 26, "unscored": 7, "not_servable": 0, "context_too_small": 1, "needs_tools": 0, "needs_vision": 0,
"unsupported_input": 0, "below_min_quality": 0, "no_offer": 16, "no_healthy_offer": 2, "below_min_discount": 0, "over_max_price": 0,
"models": { "no_offer": ["…"], "no_healthy_offer": ["…", "…"], "below_min_discount": [], "over_max_price": [] }
}
}
}The quality ranking
Read on 2026-10-06. score = the model's Artificial Analysis Intelligence Index, rounded to the nearest integer. Each catalog model is matched to the row the source publishes under the model's own id (its default configuration, e.g. 'Qwen3.6 27B (Reasoning)'); sourceName and sourceId record the match. Models the source does not list, or lists only under a different release than the catalog's, are left out (no fallback map is used in this table). estimatedBySource marks values the source itself labels as estimated.
A score describes the model as its vendor released it. Sellers often serve a quantized build (for example 4-bit or 8-bit weights, shown as the offer's quantization when the seller declares it), which can answer somewhat differently; the ranking does not measure each seller's build.
- Artificial Analysis Intelligence Index, read 2026-10-06: artificialanalysis.ai/leaderboards/models. Version: v4.3 (the leaderboard label; the methodology note on artificialanalysis.ai/models reads v4.3.2).
| Model (catalog id) | Score | Name at the source |
|---|---|---|
moonshotai/kimi-k2.6 | 27 | Kimi K2.6 (Reasoning) |
z-ai/glm-5.1 | 26 | GLM-5.1 (Reasoning) |
moonshotai/kimi-k2-thinking | 22 | Kimi K2 Thinking (the source marks it estimated) |
qwen/qwen3.6-27b | 21 | Qwen3.6 27B (Reasoning) |
qwen/qwen3.6-35b-a3b | 18 | Qwen3.6 35B A3B (Reasoning) |
deepseek/deepseek-v3.2 | 16 | DeepSeek V3.2 (Non-reasoning) (the source marks it estimated) |
google/gemma-4-31b-it | 15 | Gemma 4 31B (Reasoning) |
z-ai/glm-4.6 | 15 | GLM-4.6 (Non-reasoning) (the source marks it estimated) |
openai/gpt-oss-120b | 12 | gpt-oss-120b (High) |
qwen/qwen3-235b-a22b-2507 | 12 | Qwen3 235B A22B 2507 Instruct (the source marks it estimated) |
qwen/qwen3-coder | 12 | Qwen3 Coder 480B A35B Instruct (the source marks it estimated) |
meta-llama/llama-4-maverick | 10 | Llama 4 Maverick (the source marks it estimated) |
openai/gpt-oss-20b | 9 | gpt-oss-20b (High) |
meta-llama/llama-3.3-70b-instruct | 8 | Llama 3.3 Instruct 70B (the source marks it estimated) |
mistralai/mistral-small-3.2-24b-instruct | 8 | Mistral Small 3.2 (the source marks it estimated) |
qwen/qwen-2.5-72b-instruct | 8 | Qwen2.5 Instruct 72B (the source marks it estimated) |
meta-llama/llama-3.1-8b-instruct | 7 | Llama 3.1 Instruct 8B (the source marks it estimated) |
qwen/qwen3-14b | 7 | Qwen3 14B (Non-reasoning) (the source marks it estimated) |
google/gemma-3-27b-it | 5 | Gemma 3 27B Instruct |
Catalog models without a score (auto never routes to them; name them to use them):
qwen/qwen-2.5-14b-instruct: Not listed by Artificial Analysis (it lists Qwen2.5 Instruct 72B, Qwen2.5 Coder 32B and 7B) nor by the LMArena text leaderboard (checked 2026-10-06).mistralai/mistral-nemo: Not listed with an Intelligence Index value by Artificial Analysis.deepseek/deepseek-v4-pro: Catalog release is 0423 (OpenRouter); the source lists V4 Pro 0424 and V4 Pro 0813 only, so the match is not certain.deepseek/deepseek-v4-flash: Catalog release is 0423 (OpenRouter); the source lists V4 Flash 0420 and V4 Flash 0731 only, so the match is not certain.anthropic/claude-sonnet-5.5: Closed model: a reference row, not servable; needs no score.openai/gpt-6.1-sol: Closed model: a reference row, not servable; needs no score.google/gemini-3.8-flash: Closed model: a reference row, not servable; needs no score.
The table is kept with the catalog: when the catalog takes a new price snapshot, the scores are read again the same day. The same table, and the models auto can choose from now, are at GET /v1/routing as JSON (API reference).
What it does not do
- It chooses from the request's size and features, never from what it says.
- Tools and images are model-level flags from the catalog snapshot: a model supports the feature, and a given seller's server may still not.
- Hints are per request. A key cannot carry a default policy or price limit in this version.
- Audio and file parts go only to a model you name.
Name a model instead
Send a catalog id or an alias (GET /v1/models lists them) and Inferit picks only the seller: the cheapest healthy seller of that model, as it always has. max_price_per_m works with named models too: when no healthy offer of the model is within it, the answer is 503 no_available_offers, with retry-after. When no seller of the model serves a context window big enough for the request, the answer is 400 context_length_exceeded. min_context and min_quality do not apply to a named model. A named model needs no score: unscored models are served by name, so when auto answers no_eligible_model, a model with a live offer (a best_price in GET /v1/models) still answers by name.