Integrations Overview
Connect any OpenAI- or Anthropic-compatible tool to the Nebul Inference API. Base URLs, model discovery, and telemetry opt-outs.
Most AI tools already know how to talk to an OpenAI API. Coding agents, editor assistants, and gateways all understand it. The Nebul Inference API accepts those same requests, plus the Anthropic format for Claude-style tools. You won't find a Nebul plug-in for your favorite tool, and you don't need one: connecting a tool usually comes down to typing three values into its settings.
This section walks you through that, tool by tool. If your tool is not on the list but has an option called OpenAI Compatible, Custom OpenAI endpoint, or similar, the same three values apply. Any guide here will show you the pattern.
The three values every tool asks for
| Value | What to enter |
|---|---|
| Base URL | https://api.inference.nebul.io/v1, the address the tool sends requests to |
| API key | A key from your active Nebul AI Studio project (sk-...) |
| Model ID | The model to run, for example zai-org/GLM-5.3 |
Two details matter before you start.
Which base URL you need depends on the tool. Tools that speak the OpenAI format, which is most of them, use the address above. Tools that speak the Anthropic format, such as Claude Code, use https://api.inference.nebul.io without the /v1. Each guide in this section says which one its tool needs.
Model IDs include the vendor, like vendor/model-name. Only models enabled for your project will work. The next section shows how to look up exactly what your key can use, so you never have to guess.
Not sure which model to pick?
Every guide uses zai-org/GLM-5.3 as its example. It's our recommended model for coding agents: strong reasoning and tool use with a 1M-token context. Swap in any model your project can access. The Model Catalog lists what is available and what each model costs.
Find out which models you can use
Most tools ask you to type a model ID by hand, and a typo is the most common reason a setup fails. Two endpoints tell you exactly what your key can use.
GET /v1/models returns a plain list of model IDs:
curl https://api.inference.nebul.io/v1/models \
-H "Authorization: Bearer $NEBUL_API_KEY" | jq -r '.data[].id'GET /v1/model/info returns the same models with the details a tool needs to use them well: context window, pricing, tool calling, vision, and the reasoning levels each model accepts.
curl https://api.inference.nebul.io/v1/model/info \
-H "Authorization: Bearer $NEBUL_API_KEY" | jq '.data[0]'Both endpoints narrow to your API key's project when you authenticate. When a tool asks for a context window or an output limit, take the number from here instead of estimating. A correct number means the tool manages its context properly instead of truncating or over-sending. A few ready-made queries:
# Context window and pricing for one model
curl -sS https://api.inference.nebul.io/v1/model/info \
-H "Authorization: Bearer $NEBUL_API_KEY" | jq '.data[]
| select(.model_name == "zai-org/GLM-5.3")
| {context: .model_info.max_input_tokens,
input_eur_per_1m: .model_info.input_cost_per_1m_tokens,
output_eur_per_1m: .model_info.output_cost_per_1m_tokens}'
# Every model that supports tool calling (what coding agents need)
curl -sS https://api.inference.nebul.io/v1/model/info \
-H "Authorization: Bearer $NEBUL_API_KEY" | jq -r '.data[]
| select(.model_info.supports_function_calling == true) | .model_name'
# Reasoning models and the effort levels they accept
curl -sS https://api.inference.nebul.io/v1/model/info \
-H "Authorization: Bearer $NEBUL_API_KEY" | jq -r '.data[]
| select(.model_info.supports_reasoning == true)
| "\(.model_name): \(.model_info.reasoning_efforts)"'The Model Catalog page explains every field this endpoint returns.
What works out of the box
Tools written for OpenAI endpoints sometimes send requests in a slightly unusual shape, and strict gateways then force you to enable compatibility switches. Nebul accepts the common variations, so every guide in this section works with the tool's default settings:
- the
developerandsystemmessage roles - both
max_tokensandmax_completion_tokensoutput limits reasoning_efforton reasoning models; see Reasoning Models for what this controls- streaming, tool calling, structured output, and image input; see Examples
The full endpoint list, including embeddings, reranking, audio, OCR, and images, is in the Quick Start. To call the API from your own code, see OpenAI SDKs and plain HTTP.
Coding agents and tools
Each guide is self-contained:
| Tool | Type | Guide |
|---|---|---|
| Claude Code | Terminal agent | Claude Code |
| OpenCode | Terminal agent | OpenCode |
| Kilo Code | VS Code extension | Kilo Code |
| Kilo CLI | Terminal agent | Kilo CLI |
| Codex CLI | Terminal agent | Codex CLI |
| DeepSeek Harness (dsh) | Desktop / web agent | DeepSeek Harness |
| Pi | Terminal agent | Pi |
| Hermes Agent | Terminal / gateway agent | Hermes Agent |
| OpenClaw | Gateway agent | OpenClaw |
| NemoClaw | Sandboxed agent runner | NemoClaw |
| Cline | VS Code extension | Cline |
| Continue | VS Code / JetBrains extension | Continue |
| Zed | Editor | Zed |
| LiteLLM | SDK / gateway | LiteLLM |
Data and telemetry
Connecting a tool to Nebul keeps your prompts and code off third-party model providers. The tools themselves are a different story: many collect their own telemetry, meaning anonymous usage statistics, crash reports, and error logs sent to the tool's maker rather than to Nebul. None of it is required for the tool to work. Every guide in this section ends with the exact switch that turns it off.
| Tool | Opt-out |
|---|---|
| Claude Code | DISABLE_TELEMETRY=1, DISABLE_ERROR_REPORTING=1 |
| OpenCode | "share": "disabled", "autoupdate": false in opencode.json |
| Kilo Code | "Allow error and usage reporting" toggle off |
| Kilo CLI | "experimental": {"openTelemetry": false} in kilo.jsonc |
| Codex CLI | not signed in to ChatGPT; custom provider only. See the guide |
| DeepSeek Harness | Web UI collects no product events. See the guide |
| Pi | PI_OFFLINE=1, or PI_SKIP_VERSION_CHECK=1 for the version check only |
| Hermes Agent | setup metrics are opt-in at first run; declining sends nothing |
| OpenClaw | openclaw telemetry off, or update.checkOnStart: false to also stop update checks |
| NemoClaw | model traffic goes to the endpoint you configure; network policy controls egress |
| Cline | "Cline Telemetry" toggle in Cline settings |
| Continue | "Allow Anonymous Telemetry" toggle off |
| Zed | "telemetry": {"diagnostics": false, "metrics": false} in settings.json |
| LiteLLM | no built-in anonymous telemetry |
For how Nebul treats the requests it receives, see Privacy & Security: no training on your data, EU-only processing, project-scoped keys.
Calling the API directly
Writing your own code instead of using a pre-built tool? Any language with an OpenAI SDK will do, and a plain HTTP client is enough. OpenAI SDKs and plain HTTP has working examples for Python, TypeScript, JavaScript, Go, Rust, Java, C++, and the shell.