← All posts

A Drop-In OpenAI Compatible API — Migration in Two Lines

Aug 20, 2026 · 6 min read · api, migration

unbleep is an OpenAI compatible API, which in practice means the migration is two lines and no new dependency. You keep the official SDK, your retry logic, your streaming loop, your token accounting and your error handlers. What changes is where the requests go and which key authenticates them. This post covers the swap, then the four things that differ enough to break something if you do not know about them.

Migrating to the OpenAI compatible API in two lines

python
import os

from openai import OpenAI

client = OpenAI(
    base_url="https://unbleep.ai/v1",
    api_key=os.environ["UNBLEEP_API_KEY"],
)

resp = client.chat.completions.create(
    model="unbleep",
    messages=[{"role": "user", "content": "Summarise this incident report."}],
)
print(resp.choices[0].message.content)

If you would rather not touch the code at all, the SDK reads both values from the environment, so a config change is enough:

bash
export OPENAI_BASE_URL="https://unbleep.ai/v1"
export OPENAI_API_KEY="ub_live_9f2c..."

That same pair covers most of the ecosystem built on top of the SDK — LangChain, LlamaIndex, Instructor, the Vercel AI SDK, anything that exposes a base-URL setting. Keys are prefixed so leaks are obvious to secret scanners: ub_live_ and ub_test_ both bill against the same prepaid credit at the same per-token rate. A test key is a separate, revocable credential with a lower ceiling — 15 requests/minute instead of the account's 60 — not a free tier. Keep both server-side.

GET /v1/models works, so tooling that enumerates models to populate a dropdown does not need special-casing.

Choosing a model

Three tiers. The dated ids — unbleep-250811, unbleep-high-250811, unbleep-mini-250811 — are accepted as aliases, but today they resolve to the same build as the bare id, and the response reports the bare id back. Treat them as forward-compatible spelling, not as a reproducibility guarantee; if an evaluation has to be repeatable, record the outputs, not the model id.

| Model | Context | Price in / out per 1M | Notes | | --- | --- | --- | --- | | unbleep | 256K | $3.00 / $3.00 | Default. Reasoning tier. | | unbleep-high | 1M* | $5.00 / $5.00 | Largest jobs. Reasoning tier. | | unbleep-mini | 32K | $1.00 / $1.00 | Cheap and fast. Answers directly, no reasoning trace. |

*The request body is capped at 2,000,000 bytes — roughly 500k tokens — so a single call cannot actually fill the 1M window; anything larger comes back 413 payload_too_large.

The 32K ceiling on unbleep-mini is the one that catches people migrating from a 128K-context model: a prompt that fit before will now be rejected. If you route by cost, route by length too.

python
def pick_model(prompt_chars: int) -> str:
    """~4 chars/token is a deliberate under-estimate; leave room for the completion."""
    est_tokens = prompt_chars // 4
    if est_tokens < 24_000:
        return "unbleep-mini"
    return "unbleep" if est_tokens < 200_000 else "unbleep-high"

Streaming

Set stream=True and you get standard Server-Sent Events: each event is a chat.completion.chunk carrying a delta, and the stream terminates with a literal data: [DONE]. Your existing loop works unchanged.

bash
curl -N https://unbleep.ai/v1/chat/completions \
  -H "Authorization: Bearer $UNBLEEP_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "unbleep",
    "messages": [{"role": "user", "content": "Explain the residual stream."}],
    "stream": true,
    "stream_options": {"include_usage": true}
  }'

You get the token counts whether or not you ask for them: unbleep always requests usage from the backend and relays it, and only an explicit stream_options: {"include_usage": false} strips it on the way out. Either way the stream ends with one final chunk that has an empty choices array — carrying the populated usage object unless you opted out — which is why the loop below checks choices before touching it.

The reasoning_content field

This is the one genuine addition to the schema. unbleep and unbleep-high think before they answer, and that chain of thought is returned in reasoning_content, a sibling of content on the message (non-streaming) or the delta (streaming). Upstream backends disagree about whether to call it reasoning or reasoning_content; the API normalises to reasoning_content so you only ever handle one shape.

Because it is not part of the OpenAI schema, it is not in the SDK's type stubs. The response models allow extra fields, so attribute access works at runtime — but read it with getattr so a mini response, which has no trace, does not raise:

python
import sys

stream = client.chat.completions.create(
    model="unbleep",
    messages=[{"role": "user", "content": "Why did this detection rule misfire?"}],
    stream=True,
    stream_options={"include_usage": True},
)

usage = None
for chunk in stream:
    if not chunk.choices:          # final usage-only chunk
        usage = chunk.usage
        continue
    delta = chunk.choices[0].delta
    thought = getattr(delta, "reasoning_content", None)
    if thought:                    # trace to stderr, answer to stdout
        sys.stderr.write(thought)
    if delta.content:
        sys.stdout.write(delta.content)

if usage:
    details = usage.completion_tokens_details
    print(f"\nreasoning tokens: {getattr(details, 'reasoning_tokens', 0)}")

Three practical consequences:

python
resp = client.chat.completions.create(
    model="unbleep",
    messages=[{"role": "user", "content": "phishing or benign?"}],
    max_tokens=4,
    extra_body={"thinking": False},   # spend the budget on the answer, not the trace
)

Never feed reasoning_content back into a later turn as assistant content. It is diagnostic output, not conversation history, and replaying it degrades the next response.

Errors and the two real gotchas

Errors use the OpenAI envelope — {"error": {"type", "code", "message"}} — so your existing except blocks keep working. Status codes map as you would expect: 401 bad key, 422 blocked by policy: strict, 429 rate limit, 5xx retryable upstream.

The status you have probably never handled is 402, out of credit. Accounts are prepaid, so there is no overage and no invoice — requests simply stop until you top up. OpenAI signals quota exhaustion as a 429, which means the naive migration path is for your backoff logic to retry a 402 forever. Treat it as terminal and alert on it.

The second gotcha: system_fingerprint is not returned. It identifies the serving backend, so it is stripped along with the other vendor fields. If you key a cache or a reproducibility check on it, you will need your own version marker: the dated model ids are aliases for the current build, not frozen snapshots, so they will not tell you when the backend changes.

Rate limits come back as headers on every response — x-ratelimit-limit-requests, x-ratelimit-remaining-requests, x-ratelimit-reset-requests — so a batch job can pace its request rate instead of discovering that ceiling by hitting it. There is a second ceiling the headers do not describe: at most 8 requests in flight per account, and a ninth comes back 429 with code too_many_concurrent_requests. A streaming call holds its slot until the stream ends, so cap your own worker pool at 8.

What you are pointing at

Worth being explicit: the models behind this endpoint are abliterated, meaning their refusal behaviour has been removed at the weight level. That is the point — it is a developer API aimed at security research, red-teaming and evaluation, where a refusal is a measurement error. It also means the usual guardrails are not there to catch a bad prompt, so keep a human accountable for outputs and read the acceptable use policy. Lawful use is on you.

Get an API key — the migration is genuinely two lines.