API documentation
The API is compatible with the OpenAI Chat Completions API, so existing SDKs and tools work by changing the base URL. It is available on Plus, Pro, Max, Ultra plans; create keys on the account page.
Base URL and authentication
https://api.cyberdeck.rosy.tools/v1Send your key as Authorization: Bearer sk-.... Keys are shown once when created; we store only a hash. Revoke a leaked key from the account page. Keys only work while a plan that includes API access is active; buying a credit pack does not enable them.
Endpoints
GET /v1/models— models available to your plan.POST /v1/chat/completions— chat completion, streaming ("stream": true) or not.
Supported request fields: model, messages, stream, stream_options, max_tokens / max_completion_tokens, temperature, top_p, top_k, min_p, repetition_penalty, stop, presence_penalty, frequency_penalty, seed, response_format, tools, tool_choice, and the optional web_search extension below. Other fields are ignored; n is always 1. Reasoning models stream their thinking in delta.reasoning_content.
Optional web search
Search is off unless you explicitly set web_search: true. It is available on every chat plan, including Free; API-key access still requires an API-enabled plan. Free includes one automatically planned query for 25,000 credits. Paid plans include up to 3 related queries for one fixed 50,000-credit fee, even when fewer queries are needed. Planning is included; final-answer and evidence tokens are charged at normal model rates. The fee is retained after the first successful lookup, including empty results, even if the answer later fails or is stopped.
{
"web_search": true
}The selected model derives queries from your question and bounded recent conversation context; there is no separate query field. Generated queries contain at most 400 characters each. Only these queries, not the raw conversation, go to the search service, but they may contain details from your messages. Avoid search for sensitive content. The Python OpenAI SDK accepts the extension through extra_body. A very large current question is rejected rather than silently truncated for planning.
Before planning, the server reserves the fixed fee plus bounded evidence and at least 512 answer tokens. An explicit smaller output cap, insufficient context or insufficient credits rejects the request before paid work. Invalid planning or an entirely failed search batch has no search fee and does not generate an answer. Partial retrieval uses the available sources. Planning and lookups are never retried automatically; no paid pages are scraped.
Responses include search: {queries, completedQueries, results, searchedAt, credits}. The query list is the plan; completedQueries counts successful lookups, including those with no results. A lower completed count means partial or interrupted retrieval. Results are URL-deduplicated across queries and bounded as one packet. Each result contains title, url, content (a snippet) and date (possibly empty); searchedAt is Unix seconds. Streaming sends choices: [] metadata after each successful lookup, before answer text. Replace the previous snapshot rather than appending it; credits is the one fee, not an additional charge per frame. Preserve the latest snapshot if the answer fails or is stopped. Non-streaming errors also return any paid snapshot alongside error.
Retrieved snippets are supplied as untrusted user-level context with your question, not promoted to system instructions. API-key callers keep their own system prompts. A numbered citation such as [1] refers to the first entry in that response's search.results; [1, 2] refers to both entries. Chat renders valid numbers as source links. API clients can resolve them against the returned snapshot themselves. A citation attributes a claim; it does not verify that the claim is correct.
Search retrieves snippets, not full pages. Retrieval time is not a publication date, and a model can still misinterpret a source. Verify important claims at the linked source. Saved evidence in a later prompt does not perform or charge for another lookup; setting web_search: true again does.
Coding agents
Specter is built for code and works with OpenAI-compatible agents (Cline, Aider, OpenCode, Continue…): set the base URL above, your key and model specter. Things worth knowing:
max_tokenslimits one reply, not the task. Agents work in many short turns, so the per-reply cap (16K for Specter) is rarely reached; if it is, the reply ends withfinish_reason: "length"and agents continue automatically.- Specter reasons before answering and that thinking counts toward
max_tokens. Use at least 1000–2000; with very small values the whole budget can go to thinking and the answer comes back empty. - Agents resend the conversation and code every turn, so the context size matters most (Free and Lite have no API): Plus (32K) suits small scripts, Pro (64K) typical projects, Max (128K) and Ultra (200K) large codebases. Prompt tokens are cheap (a fraction of a credit each), which is where agents spend most. Prompt tokens the provider serves from its cache (the unchanged start of a long conversation) are cheaper still on Plus and above: up to half price on Ultra. Streaming responses report them in
usage.prompt_tokens_details.cached_tokens.
Models
| Model id | Context | Credits / output token | Credits / prompt token | Lowest plan |
|---|---|---|---|---|
specter | 200K | 1 | 0.175 | Free |
quill | 200K | 1.25 | 0.4 | Free |
Examples
curl
curl https://api.cyberdeck.rosy.tools/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "specter",
"messages": [{"role": "user", "content": "Hello!"}],
"stream": true
}'Python
from openai import OpenAI
client = OpenAI(base_url="https://api.cyberdeck.rosy.tools/v1", api_key="sk-...")
stream = client.chat.completions.create(
model="specter",
messages=[{"role": "user", "content": "Hello!"}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)JavaScript / TypeScript
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.cyberdeck.rosy.tools/v1", apiKey: process.env.API_KEY });
const stream = await client.chat.completions.create({
model: "specter",
messages: [{ role: "user", content: "Hello!" }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}Credits and limits
Each request costs ceil(prompt_tokens × prompt weight + completion_tokens × output weight) credits, using the token counts reported by the model server (or an estimate of the prompt and generated output when usage is absent). Before dispatch, a conservative prompt/reply budget is reserved from your available credits; unused credits are returned at settlement. Concurrent requests cannot spend the same credits. A request that cannot fund its prompt and a usable reply is rejected. The effective context is the smaller of your plan's and the model's; max_tokens is capped by context, remaining credits and provider capacity.
| Plan | Credits | Context | Parallel requests | API |
|---|---|---|---|---|
| Free | 50K credits / UTC day, 200K credits / rolling 7 UTC days | 16K | 1 | No |
| Lite | 5M credits / 30-day cycle | 32K | 1 | No |
| Plus | 40M credits / 30-day cycle | 32K | 2 | Yes |
| Pro | 150M credits / 30-day cycle | 64K | 3 | Yes |
| Max | 350M credits / 30-day cycle | 128K | 4 | Yes |
| Ultra | 650M credits / 30-day cycle | 200K | 6 | Yes |
Paid requests have priority in each model's bounded queue. A full queue fails immediately with capacity. Otherwise a request waits up to 5 minutes on paid plans (1 minute on Free) for a slot and the provider's rolling RPM/TPM allowance. These are shared across users. A cold model may need additional startup time; streaming responses send SSE comment keep-alives while waiting, so use generous client timeouts.
Errors
Errors use {"error": {"message": "...", "code": "..."}}. Streaming starts with HTTP 200 before admission completes; later errors arrive as a data: event with the same shape. HTTP 200 alone is not proof of a successful generation. Inspect the stream for errors and completion.
| Code | HTTP | Meaning |
|---|---|---|
bad_request | 400 | Malformed JSON or an invalid field. |
context_too_long | 400 | Prompt exceeds the context of your plan or the model. |
unauthorized | 401 | Missing or invalid API key. |
plan_required | 403 | The model or API access needs a higher (active) plan. |
model_not_found | 404 | Unknown model id — see GET /v1/models. |
quota_exceeded | 429 | A credit window is used up; resetsAt says when it frees up. |
concurrency_limit | 429 | Too many simultaneous generations for your plan. |
rate_limited | 429 | Too many attempts from your network; retry later. |
capacity | 503 | The model queue is full, or a slot/provider rate allowance did not become available in time. |
upstream_error | 502 | The model backend or explicitly requested web search failed or timed out. |