Nebul Docs

LiteLLM

Use Nebul models from LiteLLM, directly from the Python SDK or through a self-hosted LiteLLM proxy with budgets, virtual keys, and spend tracking.

LiteLLM is an open-source SDK and gateway that presents 100+ providers behind the OpenAI API format. It works with Nebul in two ways:

  • Python SDK: call Nebul models directly from Python code, with unified error handling, retries, and cost tracking.
  • Proxy server: run a local or team gateway in front of Nebul, hand out virtual keys with budgets, and point any OpenAI-compatible tool at it.

LiteLLM does not ship a nebul provider, so use its OpenAI-compatible routing: prefix model IDs with openai/ and set api_base to the Nebul endpoint.

Prerequisites

Option 1: Python SDK

Install the SDK:

pip install litellm

Call a Nebul model by prefixing the model ID with openai/ and pointing api_base at Nebul:

import os
from litellm import completion

response = completion(
    model="openai/zai-org/GLM-5.3",               # openai/ prefix routes via the OpenAI client
    api_key=os.environ["NEBUL_API_KEY"],
    api_base="https://api.inference.nebul.io/v1",
    messages=[{"role": "user", "content": "Say OK"}],
)

print(response.choices[0].message.content)

base URL must end in /v1

LiteLLM's openai/ routing uses the official OpenAI client, which appends the endpoint path itself. Pass https://api.inference.nebul.io/v1 exactly. Without /v1 the request hits a 404.

The same pattern works for embeddings. Nebul's embedding models, for example BAAI/bge-m3, are available through litellm.embedding() with model="openai/BAAI/bge-m3".

Option 2: LiteLLM proxy

The proxy is useful when you want one gateway for a team, per-user virtual keys with spend limits, or a stable local URL for tools you can't reconfigure often.

1. Install the proxy

pip install "litellm[proxy]"

2. Define your models

Create a config.yaml. Each entry maps a short model_name, which is what clients pass, to a Nebul model:

model_list:
  - model_name: glm-5.3
    litellm_params:
      model: openai/zai-org/GLM-5.3
      api_base: https://api.inference.nebul.io/v1
      api_key: os.environ/NEBUL_API_KEY
  - model_name: gpt-oss
    litellm_params:
      model: openai/openai/gpt-oss-120b
      api_base: https://api.inference.nebul.io/v1
      api_key: os.environ/NEBUL_API_KEY

os.environ/NEBUL_API_KEY reads the key from the environment when the proxy starts, so the key stays out of the file.

3. Start the proxy

export NEBUL_API_KEY=sk-your-api-key-here
litellm --config config.yaml --port 4000 --host 127.0.0.1

--host 127.0.0.1 keeps the proxy reachable only from your machine. The proxy takes over the terminal, so open a new one and verify it serves your aliases:

curl http://localhost:4000/v1/models

You should see glm-5.3 and gpt-oss in the response.

4. Point clients at the proxy

Any OpenAI SDK or OpenAI-compatible tool now uses http://localhost:4000/v1 as its base URL:

curl http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-5.3",
    "messages": [{"role": "user", "content": "Say OK"}]
  }'
import openai

client = openai.OpenAI(
    api_key="anything",                  # no master key set; see below
    base_url="http://localhost:4000/v1",
)

print(client.chat.completions.create(
    model="glm-5.3",
    messages=[{"role": "user", "content": "Say OK"}],
).choices[0].message.content)

Secure the proxy

Without a master key, anyone who can reach the proxy can spend on your Nebul account. Before exposing it beyond localhost, set one:

general_settings:
  master_key: os.environ/LITELLM_MASTER_KEY
export LITELLM_MASTER_KEY=sk-$(openssl rand -hex 16)
litellm --config config.yaml --port 4000 --host 127.0.0.1

The master key is a value you invent, not a Nebul key. Clients pass it as their API key. For per-user keys, budgets, and spend tracking, add a database and create virtual keys.

Track spend with model metadata

LiteLLM estimates cost from its built-in model price map, which does not know Nebul pricing. Override it per model with the values from /v1/model/info:

model_list:
  - model_name: glm-5.3
    litellm_params:
      model: openai/zai-org/GLM-5.3
      api_base: https://api.inference.nebul.io/v1
      api_key: os.environ/NEBUL_API_KEY
    model_info:
      input_cost_per_token: 0.00000147     # EUR per token, from /v1/model/info
      output_cost_per_token: 0.00000462
      max_tokens: 1048576                   # context window, from /v1/model/info

Claude Code through the proxy

If you prefer a single gateway for every tool, the proxy can also serve the Anthropic Messages API to Claude Code. That said, Nebul natively speaks the Anthropic protocol. See Claude Code for the direct setup with ANTHROPIC_BASE_URL=https://api.inference.nebul.io, which needs no proxy at all.

Troubleshooting

  • 404 Not Found on requests: api_base is missing the /v1 suffix, or the model entry spells the Nebul model ID differently from the catalog.
  • This model isn't available: Nebul scopes /v1/models to the key's project. Check that the model is enabled for the project that issued the key.
  • Costs show as 0: LiteLLM does not know Nebul prices. Add the model_info cost fields as shown above.
  • Unauthorized against the proxy: you set a master key, so clients must pass it instead of the Nebul key.

Data and telemetry

LiteLLM collects no anonymous telemetry on its own. The SDK and proxy only send data to external services when you configure integrations, for example success_callback: ["langfuse"]. Review the callbacks in your config.yaml if you want nothing leaving your infrastructure.

Requests the proxy forwards to Nebul are covered by Nebul's data handling: prompts are not used for training, and processing stays in the EU.

On this page