Errors
Every Inference API error code, where each comes from, exact response bodies, and retry guidance.
A request to the Inference API crosses three surfaces, and each fails differently. The network between you and our infrastructure stops malformed or oversized traffic before it arrives. The middleware checks every request before it reaches a model: authentication, routing, limits. The model itself handles everything that needs a model's judgment, like parameter validation. This page covers all three: every status code, the exact response bodies, and what to do about each.
All error responses carry an X-Request-Id header. Quote it in a support ticket and we can find the exact request. See Request IDs.
The two error shapes
Errors raised by the middleware, meaning authentication, routing, limits, and budgets, look like this:
{
"detail": {
"message": "model 'vendor/model-name': not found",
"error_type": "not_found"
}
}message is human-readable and safe to log. error_type is a stable snake_case string listed in the table below. Some responses include an extra details object; when there's nothing to detail, it's omitted.
Errors raised by the model itself follow the OpenAI error format:
{
"error": {
"message": "This model's maximum context length is 1048576 tokens...",
"type": "invalid_request_error",
"param": null,
"code": null
}
}Errors from the network surface are plain text, usually a single line.
Reading errors in the OpenAI SDK
The OpenAI SDK raises APIStatusError for all shapes. e.status_code and e.message work the same either way, so matching on the status code is the reliable path in code.
Status codes
| Code | Layer | Body | What it means |
|---|---|---|---|
| 400 | network, middleware, model | varies | Malformed request, invalid parameters, or the prompt failed a content policy check |
| 401 | middleware | JSON detail | Missing, invalid, or revoked API key |
| 403 | middleware | JSON detail | The key isn't allowed to use the model |
| 404 | middleware | JSON detail | Model ID doesn't exist, or isn't enabled for your project; or unknown path |
| 405 | middleware | empty | Wrong HTTP method for the endpoint |
| 408 | network | plain text | The client stalled while sending the request |
| 413 | network | plain text | Request body over 500 MB |
| 415 | model | OpenAI error | The endpoint requires a different Content-Type, like multipart/form-data on audio and OCR endpoints |
| 422 | middleware | JSON detail | A parameter failed validation |
| 429 | middleware | plain text or JSON | Per-minute rate limit, or the project's monthly budget is exhausted |
| 499 | network | none; the client is already gone | The client closed the connection before the response completed |
| 500 | middleware, model | JSON | Something failed on our side |
| 502 | network, middleware | varies | A service on our side failed or refused the connection |
| 503 | middleware, model | JSON | A safety check couldn't run, or the model is overloaded |
| 504 | network | plain HTML | The model sent nothing for 10 minutes |
The 499 row is the odd one: no client ever sees that status, because it's recorded after the client has already disconnected. It matters anyway, and it's explained below, because it's what a client-side timeout looks like on our side.
The codes in detail
400 Bad Request
Four common causes:
Malformed JSON. A body that can't be parsed is rejected before anything else runs, with a plain-text message:
failed to parse route metadata from request body: unexpected end of JSON inputModel validation. A request that parses but has invalid parameters reaches the model and gets the OpenAI-shaped error, usually with type: "invalid_request_error". The message names the offending parameter, for example a badly formed tools array or a messages array that's empty.
Content policy. A prompt that fails the content policy check is denied before it reaches the model, with a message saying the prompt was blocked. Adjust the prompt; if the prompt is legitimate and still blocked, contact support.
Oversized headers. Individual header lines beyond a few kilobytes are rejected with plain text before the request is processed. A normal client never produces this; it shows up with hand-rolled HTTP clients that stuff data into headers.
401 Unauthorized
Two messages cover almost every case, and both mean the same thing: the key never got authenticated.
Missing key:
{
"detail": {
"message": "Either Authorization or X-API-Key header needs to be set",
"details": {},
"error_type": "unauthorized_error"
}
}Invalid or revoked key:
{
"detail": {
"message": "Unauthorized",
"details": {},
"error_type": "unauthorized_error"
}
}Check that the header is Authorization: Bearer sk-..., with no quotes around the key and no whitespace after Bearer. Keys are scoped to one project; if the key was revoked in AI Studio, this is the response you get.
403 Forbidden
The key authenticated, but isn't allowed to use the requested model, typically because the model's access list excludes the project. A project owner can change the model's access in AI Studio; see Model Catalog for which models your project can use.
404 Not Found
The model ID is resolved against what your API key can route to, so two situations produce the same response: the model doesn't exist anywhere, or it exists but isn't enabled for your project.
{
"detail": {
"message": "model 'vendor/model-name': not found",
"error_type": "not_found"
}
}Fix the ID first: model IDs include the vendor, like zai-org/GLM-5.3, and a typo in the vendor prefix reads as a missing model. If the ID is right, the model is probably not enabled for the project that issued the key. Compare against the Model Catalog, or list what your key can actually use:
curl https://api.inference.nebul.io/v1/models \
-H "Authorization: Bearer $NEBUL_API_KEY" | jq -r '.data[].id'An unknown URL path returns plain 404 page not found, which means the endpoint name is wrong, not the model.
405 Method Not Allowed
The path is right, the verb isn't, like a GET to /v1/chat/completions. The response has an empty body. Check the method list for the endpoint in the Quick Start or the API Reference.
408 Request Timeout
The connection opened, but the client stopped sending, either headers or body. The connection is closed after about a minute of a stalled upload. This is a client-side problem in practice: a proxy between you and us dropping the connection, or an upload that died halfway.
413 Payload Too Large
The request body exceeded 500 MB. The response is a plain-text error page. If you need to send more than 500 MB of input to a model, split the request or preprocess the input down. Note that vision inputs should be URLs or reasonably sized base64 strings, not raw multi-megabyte images inline.
415 Unsupported Media Type
The endpoint requires a different Content-Type than the one sent. The audio and OCR endpoints expect multipart/form-data, documented per endpoint in Speech to Text and OCR. Sending application/json to those endpoints produces this error, in the OpenAI error shape.
429 Too Many Requests
Two different limits produce a 429:
Per-minute rate limit. The number of requests per minute for your key or model tier was exceeded. The response is short and plain: rate limit exceeded. Wait for the current minute and retry. Limits and the recommended retry pattern are in Rate Limits.
Monthly project budget. The project that issued the key has spent its monthly limit. The response body is the limit status:
{
"project_id": "0196b4a2-ec3a-7c1e-9d2a-3f5e7a9c1b21",
"limit_eur_cents": 50000,
"consumed_eur_cents": 50000,
"remaining_eur_cents": 0,
"exceeded": true,
"period": "2026-10"
}Requests stay denied until the next billing period or until a project owner raises the budget in AI Studio.
499 Client Closed Request
You'll never receive a 499, but you may cause one, and it explains a confusing situation. 499 is what our side records when the client closes the connection before the response completes. If your client's timeout is shorter than the generation time, your side shows a timeout while our logs show 499, and the request may still have completed on our side. When you report a timeout to support, we'll likely see 499 next to your request ID; that's the fingerprint of a client-side timeout, and the fix is streaming or a longer client timeout, same as for 504.
500 Internal Server Error
Unexpected failures on our side. Retry with backoff. If the same request fails repeatedly, include the X-Request-Id from a failed response when you contact support.
502 Bad Gateway
A service on our side refused or dropped the connection. This is transient infrastructure failure, not something in your request. Retry with backoff.
503 Service Unavailable
The request passed authentication and routing, but a required safety check couldn't run: the check service was unreachable or timed out, and the check is configured to deny on failure rather than let the request pass unchecked. This is a deliberate fail-closed design: the request is not checked, so it is not forwarded. Retry with backoff; if it persists, the check service needs attention on our side, so include the X-Request-Id.
504 Gateway Timeout
Our infrastructure waited 10 minutes for the model to send anything and gave up. This happens on very long non-streaming requests: the model is still generating, but nothing flows to the client until the whole response is ready. Two fixes:
- Stream the response. With
stream: true, tokens flow as they're generated and the connection stays active. See Examples. - Raise your client timeout above your worst-case generation time. If your HTTP client gives up before 10 minutes, you'll see a client-side timeout instead, and our side records a 499.
Which errors to retry
| Retryable | Codes |
|---|---|
| Yes, with exponential backoff | 408, 429, 500, 502, 503, 504 |
| No, the request must change | 400, 401, 403, 404, 405, 413, 415, 422 |
Retrying a 400 or a 404 with the same body will fail forever; fix the request instead. For 429, honor the retry rhythm from Rate Limits, which includes a working backoff example in Python. For 408 and 504, check your client's timeout before retrying the same way, or the retry will hit the same wall.
Getting help
When an error doesn't match anything on this page, grab three things before contacting support: the X-Request-Id response header, the status code, and the exact timestamp. With those, we can trace the request end to end.