API documentation

The API is compatible with the OpenAI Chat Completions API, so existing SDKs and tools work by changing the base URL. It is available on Plus, Pro, Max, Ultra plans; create keys on the account page.

Base URL and authentication

https://api.cyberdeck.rosy.tools/v1

Send your key as Authorization: Bearer sk-.... Keys are shown once when created; we store only a hash. Revoke a leaked key from the account page. Keys only work while a plan that includes API access is active; buying a credit pack does not enable them.

Endpoints

  • GET /v1/models — models available to your plan.
  • POST /v1/chat/completions — chat completion, streaming ("stream": true) or not.

Supported request fields: model, messages, stream, stream_options, max_tokens / max_completion_tokens, temperature, top_p, top_k, min_p, repetition_penalty, stop, presence_penalty, frequency_penalty, seed, response_format, tools, tool_choice, and the optional web_search extension below. Other fields are ignored; n is always 1. Reasoning models stream their thinking in delta.reasoning_content.

Optional web search

Search is off unless you explicitly set web_search: true. It is available on every chat plan, including Free; API-key access still requires an API-enabled plan. Free includes one automatically planned query for 25,000 credits. Paid plans include up to 3 related queries for one fixed 50,000-credit fee, even when fewer queries are needed. Planning is included; final-answer and evidence tokens are charged at normal model rates. The fee is retained after the first successful lookup, including empty results, even if the answer later fails or is stopped.

{
  "web_search": true
}

The selected model derives queries from your question and bounded recent conversation context; there is no separate query field. Generated queries contain at most 400 characters each. Only these queries, not the raw conversation, go to the search service, but they may contain details from your messages. Avoid search for sensitive content. The Python OpenAI SDK accepts the extension through extra_body. A very large current question is rejected rather than silently truncated for planning.

Before planning, the server reserves the fixed fee plus bounded evidence and at least 512 answer tokens. An explicit smaller output cap, insufficient context or insufficient credits rejects the request before paid work. Invalid planning or an entirely failed search batch has no search fee and does not generate an answer. Partial retrieval uses the available sources. Planning and lookups are never retried automatically; no paid pages are scraped.

Responses include search: {queries, completedQueries, results, searchedAt, credits}. The query list is the plan; completedQueries counts successful lookups, including those with no results. A lower completed count means partial or interrupted retrieval. Results are URL-deduplicated across queries and bounded as one packet. Each result contains title, url, content (a snippet) and date (possibly empty); searchedAt is Unix seconds. Streaming sends choices: [] metadata after each successful lookup, before answer text. Replace the previous snapshot rather than appending it; credits is the one fee, not an additional charge per frame. Preserve the latest snapshot if the answer fails or is stopped. Non-streaming errors also return any paid snapshot alongside error.

Retrieved snippets are supplied as untrusted user-level context with your question, not promoted to system instructions. API-key callers keep their own system prompts. A numbered citation such as [1] refers to the first entry in that response's search.results; [1, 2] refers to both entries. Chat renders valid numbers as source links. API clients can resolve them against the returned snapshot themselves. A citation attributes a claim; it does not verify that the claim is correct.

Search retrieves snippets, not full pages. Retrieval time is not a publication date, and a model can still misinterpret a source. Verify important claims at the linked source. Saved evidence in a later prompt does not perform or charge for another lookup; setting web_search: true again does.

Coding agents

Specter is built for code and works with OpenAI-compatible agents (Cline, Aider, OpenCode, Continue…): set the base URL above, your key and model specter. Things worth knowing:

  • max_tokens limits one reply, not the task. Agents work in many short turns, so the per-reply cap (16K for Specter) is rarely reached; if it is, the reply ends with finish_reason: "length" and agents continue automatically.
  • Specter reasons before answering and that thinking counts toward max_tokens. Use at least 1000–2000; with very small values the whole budget can go to thinking and the answer comes back empty.
  • Agents resend the conversation and code every turn, so the context size matters most (Free and Lite have no API): Plus (32K) suits small scripts, Pro (64K) typical projects, Max (128K) and Ultra (200K) large codebases. Prompt tokens are cheap (a fraction of a credit each), which is where agents spend most. Prompt tokens the provider serves from its cache (the unchanged start of a long conversation) are cheaper still on Plus and above: up to half price on Ultra. Streaming responses report them in usage.prompt_tokens_details.cached_tokens.

Models

Model idContextCredits / output tokenCredits / prompt tokenLowest plan
specter200K10.175Free
quill200K1.250.4Free

Examples

curl

curl https://api.cyberdeck.rosy.tools/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "specter",
    "messages": [{"role": "user", "content": "Hello!"}],
    "stream": true
  }'

Python

from openai import OpenAI

client = OpenAI(base_url="https://api.cyberdeck.rosy.tools/v1", api_key="sk-...")

stream = client.chat.completions.create(
    model="specter",
    messages=[{"role": "user", "content": "Hello!"}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

JavaScript / TypeScript

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.cyberdeck.rosy.tools/v1", apiKey: process.env.API_KEY });

const stream = await client.chat.completions.create({
  model: "specter",
  messages: [{ role: "user", content: "Hello!" }],
  stream: true,
});
for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}

Credits and limits

Each request costs ceil(prompt_tokens × prompt weight + completion_tokens × output weight) credits, using the token counts reported by the model server (or an estimate of the prompt and generated output when usage is absent). Before dispatch, a conservative prompt/reply budget is reserved from your available credits; unused credits are returned at settlement. Concurrent requests cannot spend the same credits. A request that cannot fund its prompt and a usable reply is rejected. The effective context is the smaller of your plan's and the model's; max_tokens is capped by context, remaining credits and provider capacity.

PlanCreditsContextParallel requestsAPI
Free50K credits / UTC day, 200K credits / rolling 7 UTC days16K1No
Lite5M credits / 30-day cycle32K1No
Plus40M credits / 30-day cycle32K2Yes
Pro150M credits / 30-day cycle64K3Yes
Max350M credits / 30-day cycle128K4Yes
Ultra650M credits / 30-day cycle200K6Yes

Paid requests have priority in each model's bounded queue. A full queue fails immediately with capacity. Otherwise a request waits up to 5 minutes on paid plans (1 minute on Free) for a slot and the provider's rolling RPM/TPM allowance. These are shared across users. A cold model may need additional startup time; streaming responses send SSE comment keep-alives while waiting, so use generous client timeouts.

Errors

Errors use {"error": {"message": "...", "code": "..."}}. Streaming starts with HTTP 200 before admission completes; later errors arrive as a data: event with the same shape. HTTP 200 alone is not proof of a successful generation. Inspect the stream for errors and completion.

CodeHTTPMeaning
bad_request400Malformed JSON or an invalid field.
context_too_long400Prompt exceeds the context of your plan or the model.
unauthorized401Missing or invalid API key.
plan_required403The model or API access needs a higher (active) plan.
model_not_found404Unknown model id — see GET /v1/models.
quota_exceeded429A credit window is used up; resetsAt says when it frees up.
concurrency_limit429Too many simultaneous generations for your plan.
rate_limited429Too many attempts from your network; retry later.
capacity503The model queue is full, or a slot/provider rate allowance did not become available in time.
upstream_error502The model backend or explicitly requested web search failed or timed out.