LiteLLM
Use Nebul models from LiteLLM, directly from the Python SDK or through a self-hosted LiteLLM proxy with budgets, virtual keys, and spend tracking.
LiteLLM is an open-source SDK and gateway that presents 100+ providers behind the OpenAI API format. It works with Nebul in two ways:
- Python SDK: call Nebul models directly from Python code, with unified error handling, retries, and cost tracking.
- Proxy server: run a local or team gateway in front of Nebul, hand out virtual keys with budgets, and point any OpenAI-compatible tool at it.
LiteLLM does not ship a nebul provider, so use its OpenAI-compatible routing: prefix model IDs with openai/ and set api_base to the Nebul endpoint.
Prerequisites
- Python 3.9 or later
- A Nebul AI Studio API key
- A model ID from the Model Catalog
Option 1: Python SDK
Install the SDK:
pip install litellmCall a Nebul model by prefixing the model ID with openai/ and pointing api_base at Nebul:
import os
from litellm import completion
response = completion(
model="openai/zai-org/GLM-5.3", # openai/ prefix routes via the OpenAI client
api_key=os.environ["NEBUL_API_KEY"],
api_base="https://api.inference.nebul.io/v1",
messages=[{"role": "user", "content": "Say OK"}],
)
print(response.choices[0].message.content)base URL must end in /v1
LiteLLM's openai/ routing uses the official OpenAI client, which appends the endpoint path itself. Pass https://api.inference.nebul.io/v1 exactly. Without /v1 the request hits a 404.
The same pattern works for embeddings. Nebul's embedding models, for example BAAI/bge-m3, are available through litellm.embedding() with model="openai/BAAI/bge-m3".
Option 2: LiteLLM proxy
The proxy is useful when you want one gateway for a team, per-user virtual keys with spend limits, or a stable local URL for tools you can't reconfigure often.
1. Install the proxy
pip install "litellm[proxy]"2. Define your models
Create a config.yaml. Each entry maps a short model_name, which is what clients pass, to a Nebul model:
model_list:
- model_name: glm-5.3
litellm_params:
model: openai/zai-org/GLM-5.3
api_base: https://api.inference.nebul.io/v1
api_key: os.environ/NEBUL_API_KEY
- model_name: gpt-oss
litellm_params:
model: openai/openai/gpt-oss-120b
api_base: https://api.inference.nebul.io/v1
api_key: os.environ/NEBUL_API_KEYos.environ/NEBUL_API_KEY reads the key from the environment when the proxy starts, so the key stays out of the file.
3. Start the proxy
export NEBUL_API_KEY=sk-your-api-key-here
litellm --config config.yaml --port 4000 --host 127.0.0.1--host 127.0.0.1 keeps the proxy reachable only from your machine. The proxy takes over the terminal, so open a new one and verify it serves your aliases:
curl http://localhost:4000/v1/modelsYou should see glm-5.3 and gpt-oss in the response.
4. Point clients at the proxy
Any OpenAI SDK or OpenAI-compatible tool now uses http://localhost:4000/v1 as its base URL:
curl http://localhost:4000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5.3",
"messages": [{"role": "user", "content": "Say OK"}]
}'import openai
client = openai.OpenAI(
api_key="anything", # no master key set; see below
base_url="http://localhost:4000/v1",
)
print(client.chat.completions.create(
model="glm-5.3",
messages=[{"role": "user", "content": "Say OK"}],
).choices[0].message.content)Secure the proxy
Without a master key, anyone who can reach the proxy can spend on your Nebul account. Before exposing it beyond localhost, set one:
general_settings:
master_key: os.environ/LITELLM_MASTER_KEYexport LITELLM_MASTER_KEY=sk-$(openssl rand -hex 16)
litellm --config config.yaml --port 4000 --host 127.0.0.1The master key is a value you invent, not a Nebul key. Clients pass it as their API key. For per-user keys, budgets, and spend tracking, add a database and create virtual keys.
Track spend with model metadata
LiteLLM estimates cost from its built-in model price map, which does not know Nebul pricing. Override it per model with the values from /v1/model/info:
model_list:
- model_name: glm-5.3
litellm_params:
model: openai/zai-org/GLM-5.3
api_base: https://api.inference.nebul.io/v1
api_key: os.environ/NEBUL_API_KEY
model_info:
input_cost_per_token: 0.00000147 # EUR per token, from /v1/model/info
output_cost_per_token: 0.00000462
max_tokens: 1048576 # context window, from /v1/model/infoClaude Code through the proxy
If you prefer a single gateway for every tool, the proxy can also serve the Anthropic Messages API to Claude Code. That said, Nebul natively speaks the Anthropic protocol. See Claude Code for the direct setup with ANTHROPIC_BASE_URL=https://api.inference.nebul.io, which needs no proxy at all.
Troubleshooting
404 Not Foundon requests:api_baseis missing the/v1suffix, or the model entry spells the Nebul model ID differently from the catalog.This model isn't available: Nebul scopes/v1/modelsto the key's project. Check that the model is enabled for the project that issued the key.- Costs show as 0: LiteLLM does not know Nebul prices. Add the
model_infocost fields as shown above. Unauthorizedagainst the proxy: you set a master key, so clients must pass it instead of the Nebul key.
Data and telemetry
LiteLLM collects no anonymous telemetry on its own. The SDK and proxy only send data to external services when you configure integrations, for example success_callback: ["langfuse"]. Review the callbacks in your config.yaml if you want nothing leaving your infrastructure.
Requests the proxy forwards to Nebul are covered by Nebul's data handling: prompts are not used for training, and processing stays in the EU.