Nebul Docs

Integrations Overview

Connect any OpenAI- or Anthropic-compatible tool to the Nebul Inference API. Base URLs, model discovery, and telemetry opt-outs.

Most AI tools already know how to talk to an OpenAI API. Coding agents, editor assistants, and gateways all understand it. The Nebul Inference API accepts those same requests, plus the Anthropic format for Claude-style tools. You won't find a Nebul plug-in for your favorite tool, and you don't need one: connecting a tool usually comes down to typing three values into its settings.

This section walks you through that, tool by tool. If your tool is not on the list but has an option called OpenAI Compatible, Custom OpenAI endpoint, or similar, the same three values apply. Any guide here will show you the pattern.

The three values every tool asks for

ValueWhat to enter
Base URLhttps://api.inference.nebul.io/v1, the address the tool sends requests to
API keyA key from your active Nebul AI Studio project (sk-...)
Model IDThe model to run, for example zai-org/GLM-5.3

Two details matter before you start.

Which base URL you need depends on the tool. Tools that speak the OpenAI format, which is most of them, use the address above. Tools that speak the Anthropic format, such as Claude Code, use https://api.inference.nebul.io without the /v1. Each guide in this section says which one its tool needs.

Model IDs include the vendor, like vendor/model-name. Only models enabled for your project will work. The next section shows how to look up exactly what your key can use, so you never have to guess.

Not sure which model to pick?

Every guide uses zai-org/GLM-5.3 as its example. It's our recommended model for coding agents: strong reasoning and tool use with a 1M-token context. Swap in any model your project can access. The Model Catalog lists what is available and what each model costs.

Find out which models you can use

Most tools ask you to type a model ID by hand, and a typo is the most common reason a setup fails. Two endpoints tell you exactly what your key can use.

GET /v1/models returns a plain list of model IDs:

curl https://api.inference.nebul.io/v1/models \
  -H "Authorization: Bearer $NEBUL_API_KEY" | jq -r '.data[].id'

GET /v1/model/info returns the same models with the details a tool needs to use them well: context window, pricing, tool calling, vision, and the reasoning levels each model accepts.

curl https://api.inference.nebul.io/v1/model/info \
  -H "Authorization: Bearer $NEBUL_API_KEY" | jq '.data[0]'

Both endpoints narrow to your API key's project when you authenticate. When a tool asks for a context window or an output limit, take the number from here instead of estimating. A correct number means the tool manages its context properly instead of truncating or over-sending. A few ready-made queries:

# Context window and pricing for one model
curl -sS https://api.inference.nebul.io/v1/model/info \
  -H "Authorization: Bearer $NEBUL_API_KEY" | jq '.data[]
  | select(.model_name == "zai-org/GLM-5.3")
  | {context: .model_info.max_input_tokens,
     input_eur_per_1m: .model_info.input_cost_per_1m_tokens,
     output_eur_per_1m: .model_info.output_cost_per_1m_tokens}'

# Every model that supports tool calling (what coding agents need)
curl -sS https://api.inference.nebul.io/v1/model/info \
  -H "Authorization: Bearer $NEBUL_API_KEY" | jq -r '.data[]
  | select(.model_info.supports_function_calling == true) | .model_name'

# Reasoning models and the effort levels they accept
curl -sS https://api.inference.nebul.io/v1/model/info \
  -H "Authorization: Bearer $NEBUL_API_KEY" | jq -r '.data[]
  | select(.model_info.supports_reasoning == true)
  | "\(.model_name): \(.model_info.reasoning_efforts)"'

The Model Catalog page explains every field this endpoint returns.

What works out of the box

Tools written for OpenAI endpoints sometimes send requests in a slightly unusual shape, and strict gateways then force you to enable compatibility switches. Nebul accepts the common variations, so every guide in this section works with the tool's default settings:

  • the developer and system message roles
  • both max_tokens and max_completion_tokens output limits
  • reasoning_effort on reasoning models; see Reasoning Models for what this controls
  • streaming, tool calling, structured output, and image input; see Examples

The full endpoint list, including embeddings, reranking, audio, OCR, and images, is in the Quick Start. To call the API from your own code, see OpenAI SDKs and plain HTTP.

Coding agents and tools

Each guide is self-contained:

ToolTypeGuide
Claude CodeTerminal agentClaude Code
OpenCodeTerminal agentOpenCode
Kilo CodeVS Code extensionKilo Code
Kilo CLITerminal agentKilo CLI
Codex CLITerminal agentCodex CLI
DeepSeek Harness (dsh)Desktop / web agentDeepSeek Harness
PiTerminal agentPi
Hermes AgentTerminal / gateway agentHermes Agent
OpenClawGateway agentOpenClaw
NemoClawSandboxed agent runnerNemoClaw
ClineVS Code extensionCline
ContinueVS Code / JetBrains extensionContinue
ZedEditorZed
LiteLLMSDK / gatewayLiteLLM

Data and telemetry

Connecting a tool to Nebul keeps your prompts and code off third-party model providers. The tools themselves are a different story: many collect their own telemetry, meaning anonymous usage statistics, crash reports, and error logs sent to the tool's maker rather than to Nebul. None of it is required for the tool to work. Every guide in this section ends with the exact switch that turns it off.

ToolOpt-out
Claude CodeDISABLE_TELEMETRY=1, DISABLE_ERROR_REPORTING=1
OpenCode"share": "disabled", "autoupdate": false in opencode.json
Kilo Code"Allow error and usage reporting" toggle off
Kilo CLI"experimental": {"openTelemetry": false} in kilo.jsonc
Codex CLInot signed in to ChatGPT; custom provider only. See the guide
DeepSeek HarnessWeb UI collects no product events. See the guide
PiPI_OFFLINE=1, or PI_SKIP_VERSION_CHECK=1 for the version check only
Hermes Agentsetup metrics are opt-in at first run; declining sends nothing
OpenClawopenclaw telemetry off, or update.checkOnStart: false to also stop update checks
NemoClawmodel traffic goes to the endpoint you configure; network policy controls egress
Cline"Cline Telemetry" toggle in Cline settings
Continue"Allow Anonymous Telemetry" toggle off
Zed"telemetry": {"diagnostics": false, "metrics": false} in settings.json
LiteLLMno built-in anonymous telemetry

For how Nebul treats the requests it receives, see Privacy & Security: no training on your data, EU-only processing, project-scoped keys.

Calling the API directly

Writing your own code instead of using a pre-built tool? Any language with an OpenAI SDK will do, and a plain HTTP client is enough. OpenAI SDKs and plain HTTP has working examples for Python, TypeScript, JavaScript, Go, Rust, Java, C++, and the shell.

On this page