257 Commits

Author SHA1 Message Date
6770f6ba0b fix(dashboard): ReferenceError getProviderLabel in RecentRequests
getProviderLabel was a useCallback inside the UsageStats component, but
RecentRequests (a sibling module-level component) called it too, causing
"getProviderLabel is not defined" at runtime.

Extract a module-level resolveProviderLabel() (registry lookup by
id/alias/display prefix) used by RecentRequests; keep the component-scoped
getProviderLabel (which adds connected-provider nodeName/name lookup) for
the main table render paths.
2026-08-17 16:30:26 +07:00
c61dc6de46 fix(dashboard): show provider names instead of node ids in Usage by Provider
The Usage by Provider table (and related provider cells) rendered raw
provider keys, so requests routed through a custom OpenAI-compatible node
appeared as "openai-compatible-chat-096baf9a-6433-4f22-b079-20371092555a"
instead of the node's display name.

- UsageStats: add getProviderLabel() resolving a provider key (built-in id,
  alias, or custom node id) to a friendly name via the providers state
  (nodeName > connection name) and AI_PROVIDERS/getProviderByAlias fallback
- Apply the label in group headers, detail rows, badges, and recent requests
  (raw key kept as hover title)
- UsageTable: accept providerLabel prop, use it for provider group headers
2026-08-17 14:24:13 +07:00
b5c0f10610 feat(settings): restore per-provider connect timeout overrides
Re-apply the settings/UI layer of the per-provider timeout feature that
was dropped during the origin/master merge (core providerTimeout.js +
executor wiring survived; the settings keys and dashboard UI did not):

- settingsRepo: providerTimeouts:{} + defaultTimeoutMs:null defaults
- Profile page: "Default Connect Timeout" card (global fallback, ms)
- Provider detail page: per-provider "Connect Timeout" input, saved to
  providerTimeouts[providerId].timeoutMs, applied via
  resolveProviderTimeoutMs() priority: per-provider > global > registry > env
2026-08-17 09:49:44 +07:00
1256f29d92 Merge remote-tracking branch 'origin/master' into gitea/new_feature
# Conflicts:
#	open-sse/handlers/chatCore.js
#	open-sse/services/combo.js
#	src/app/(dashboard)/dashboard/profile/page.js
#	src/app/api/v1/models/route.js
#	src/lib/db/repos/settingsRepo.js
2026-08-17 00:21:51 +07:00
de9e00c66d feat(settings): runtime log level + free provider enable/disable
- Add LOG_LEVEL env + runtime setLogLevel (dashboard Settings → Logging),
  applied immediately, persisted across restarts; WARN/ERROR quiet production
  INFO lines (▶ POST / 📊 DONE / [COMBO] / [CHAT])
- Allow toggling free/noAuth providers (gemini-cli, kilo, etc.) off via
  providerStrategies.enabled from Providers page and provider detail page
- auth.js: honor disabled override before noAuth/connection branches
- CompatibleModelsSection: parallel model testing
2026-08-17 00:17:55 +07:00
decolua
699edac327 # v0.5.55 (2026-08-14)
## Features
- **Auth**: native SAML 2.0 SSO alongside OIDC — AuthnRequest generation, ACS
  assertion handling, SP metadata export, admin config test, replay-protected
  via a `saml_state` cookie matched against `InResponseTo`
- **Providers**: add Alibaba Token Plan (`token-plan.ap-southeast-1`) — the
  fourth Alibaba key type, Singapore-only and OpenAI-compatible transport only
- **Providers**: add `glm-5.3` to GLM Coding and GLM (China)
- **Providers**: Kimchi accepts API keys as well as OAuth (dual auth), with a
  working Test Connection for both modes
- **Antigravity**: add Gemini 3.7 Flash and its tiered high/medium/low variants
  (also in the Gemini registry) with pricing and quota tracking
- **TTS**: add Fish Audio — model id travels in an HTTP `model` header, voice
  is a `reference_id` (preset or cloned voice model)
- **OpenCode-Go**: route by request format via declared transports instead of
  forcing every client into `/messages` — Codex/OpenAI clients no longer pay a
  lossy Responses→OpenAI→Claude double translation. Per-model `supportedFormats`
  guard; the bespoke executor is gone (its shared `_lastModel` cache could cross
  auth headers between concurrent requests)
- **Usage**: dedup + cache Claude quota calls (120s TTL keyed by access token,
  in-flight promise dedup, last-good read on soft failure) to stop multiple
  tabs tripping 429; manual refresh (↻) sends `force=1` to bypass the cache

## Fixes
- **Docker**: ship `sql.js` in the image so the pure-JS DB fallback can start —
  file tracing carried the package's JS without `dist/sql-wasm.wasm`, so a
  container with no native driver aborted with ENOENT and never got a database
  (#3248)
- **Usage**: read Gemini `usageMetadata` out of the antigravity `{ response }`
  envelope — every non-streaming antigravity request logged `IN 0 | OUT 0`
  (#3260)
- **Claude**: re-anchor passthrough cache breakpoints — the client's own
  `cache_control` markers point at pre-normalization offsets, so the tail was
  re-cached every request. Last system block and last tool pinned at 1h TTL,
  last assistant turn at 5m, mid-conversation system messages folded into the
  neighbouring user turn instead of hoisted into `body.system`
- **Combos**: detect images from Hermes and attachment payloads (`images[]`,
  `experimental_attachments`, message-level `image_url`/`audio_url`, inline
  `data:` URIs) so the Vision Adapter auto-switch fires for Hermes/Ollama/
  Vercel AI SDK shapes
- **Kiro**: intercept chat via `x-amz-target` — Kiro IDE 1.0.228+ moved
  `GenerateAssistantResponse` to `POST /` + header, bypassing MITM. Also emit
  the now-mandatory initial-response frame and map the `auto` model slot
- **Kiro**: report real output tokens and stop discarding usable turns
- **Qoder**: detect billing blocks at stream start and return a synthetic 403
  so combo/account fallback triggers instead of leaking the error into chat
- **Antigravity**: strip competitive system prompts (Zed IDE's Claude-agent
  prompt) that Antigravity flags with a 429 Quota Exhausted
- **OpenCode**: send the official client fingerprint on free-tier requests so
  the Console stops classifying traffic as unidentified and rate-limiting it;
  session id resolves conversation-stable to preserve prompt caching
- **Responses**: don't close the message on an empty `tool_calls` array — some
  providers attach one to every chunk, and the truthy check ended the message
  on the first content token (#3234)
- **Translator**: preserve `prompt_cache_key` when converting chat to responses
- **Models**: expose snake_case token limits on `/v1/models`
- **Combos**: strip `stream_options` from the Fusion panel fan-out to avoid a
  DeepSeek 400 (#3024); raise the dashboard model-test probe budget to 1024 and
  soft-pass reasoning-only responses (#3010)
- **Headroom**: the toggle reflects the `headroomEnabled` setting even when the
  proxy is down — it previously showed OFF while the engine kept calling
  `/v1/compress`; proxy status stays visible via the status chip
- **Hermes**: add the `api_key` parameter to the model block in YAML config
- **Providers**: add llm7 to provider test support

## Docs
- **i18n**: add Spanish, French, and Brazilian Portuguese README translations

## Security
- **Real IP**: `x-9r-real-ip` and the Host fallback were trusted from
  client-controlled headers whenever `custom-server.js` was not in the request
  path (`npm run start`, `start:bun`), letting a remote caller pose as local to
  skip API key auth and reach `LOCAL_ONLY_PATHS` (`/api/mcp/*`,
  `/api/tunnel/enable`, `/api/auth/reset-password`). The server now stamps a
  per-process `x-9r-peer-token` on every request it sanitizes and only trusts
  `x-9r-real-ip` behind it — falling back to Host in development and failing
  closed in production (GHSA-pjm4-8fpg-f9p6). Also fixes IPv6 loopback
  detection (`::1`, `::ffff:127.0.0.1`) and routes `npm run start` /
  `start:bun` through `custom-server.js`
- **Search**: `resolveBaseUrl()` rejects client-supplied non-public baseUrls
  (SSRF guard on `/v1/search`)
- **Login**: fresh-install remote login with the default password returns 403
  without issuing a JWT
- **Usage**: `/api/usage/request-details` redacts request/response payloads
2026-08-14 17:08:02 +07:00
decolua
540ebbe682 test(baseline): regenerate provider snapshot for opencode-go transports 2026-08-14 16:53:04 +07:00
KiMelody
e1115e2839 feat(opencode-go): route by request format via transports + per-model guard
opencode-go hard-coded targetFormat: claude per model, so every client
format was force-routed to /messages (Codex/OpenAI clients paid a lossy
Responses->OpenAI->Claude double translation). Declare the existing
upstream multi-endpoint transports [openai, claude, openai-responses]
and guard per model via registry supportedFormats: kimi/glm/mimo only
support /chat/completions, minimax/qwen add /messages, deepseek adds
/responses. Undeclared models keep the upstream default.

Drop the bespoke OpenCodeGoExecutor (its shared _lastModel cache could
cross auth headers between concurrent requests); DefaultExecutor already
consumes runtimeTransport and injects reasoning content.
2026-08-14 16:52:37 +07:00
Nguyen Thanh Dat
27f3710c8b fix(docker): ship sql.js so the pure-JS DB fallback can start
Next file tracing follows JS imports, and sql.js loads dist/sql-wasm.wasm by
path at runtime, so the standalone output carries the package's JS without its
wasm binary. When both native drivers fail the last-resort adapter then aborts
with ENOENT on the missing binary and the container never gets a database.

The CLI bundle already guards this explicitly (build-cli.js step 3b,
ensureModuleInBundle("sql.js")); the image just never got the same treatment.
Copy the package the same way node-forge and next already are.

Fixes #3248
2026-08-14 16:40:53 +07:00
Nguyen Thanh Dat
59d858b639 fix(usage): read Gemini usageMetadata out of the antigravity response envelope
Antigravity and gemini-cli wrap their payload in { response: {...} }.
extractUsageFromResponse only tested top-level usageMetadata, so every
non-streaming antigravity request logged zero usage (IN 0 | OUT 0) and
zeroed rows in the usage dashboard. Read the envelope the same way
usageTracking.js and nonStreamingHandler.js already do; top-level
metadata keeps priority and the OpenAI/Claude branches are untouched.

Fixes #3260
2026-08-14 16:34:54 +07:00
Nguyen Thanh Dat
92259214db fix(security): require proof that x-9r-real-ip came from the socket (GHSA-pjm4-8fpg-f9p6)
x-9r-real-ip and the Host fallback were trusted from client-controlled
headers whenever custom-server.js was not in the request path (npm run
start, start:bun), letting a remote caller pose as local to skip API key
auth and reach LOCAL_ONLY_PATHS (/api/mcp/*, /api/tunnel/enable,
/api/auth/reset-password).

custom-server.js now generates a per-process secret at boot and stamps it
as x-9r-peer-token on every request it sanitizes. hasTrustedPeerHeaders()
(src/lib/auth/trustedPeer.js) gates trust in x-9r-real-ip on that secret;
otherwise the guard falls back to Host only in development, and fails
closed in production. Same gate on loginLimiter.getClientIp() so a spoofed
header cannot rotate the login lockout bucket.

Also: fix isLoopbackHostname for IPv6 (::1, ::ffff:127.0.0.1) which the
old split(":")[0] reduced to empty string; route npm run start /
start:bun through custom-server.js (postbuild copies it into
.next/standalone, build-cli.js fails without it) so documented deployments
keep passwordless local access.
2026-08-14 16:33:58 +07:00
Nguyen Thanh Dat
b04c03c6b5 feat(providers): add Alibaba Token Plan (token-plan.ap-southeast-1)
Fourth Alibaba key type — Coding Plan (alicode/alicode-intl) and Model Studio
(alims-intl) both reject Token Plan keys. Registry entry only; PROVIDER_MODELS
builds from providers/registry so no executor or translator work is needed.

Singapore-only (eu-central-1 answers IllegalEndpoint) and OpenAI-compatible
transport only (the Anthropic surface is not authorized for this plan).

Closes #2754
Closes #2806
2026-08-14 16:32:54 +07:00
AlexNoVibe
8b2b2fefb5 docs(i18n): add Spanish and French README translations
Add i18n/README.es.md and i18n/README.fr.md mirroring the English
README structure, and link both from the language switcher.
2026-08-14 16:29:43 +07:00
Azriel Akbar Ferry Ardiansyah Kusumawardhana
86694ed8d0 feat(antigravity): add Gemini 3.7 Flash models (#3286, #3281)
Add gemini-3.7-flash and its tiered high/medium/low variants to the
Antigravity and Gemini registries, with matching capabilities, pricing
and Antigravity quota tracking.

extractModel now recognises gemini-3.7-flash-tiered alongside 3.6 and
derives the version from the request, so thinkingLevel still maps to the
right tiered alias.

Closes #3286
Closes #3281
2026-08-14 16:27:10 +07:00
Nguyen Thanh Dat
8af5e752da feat(tts): add Fish Audio as a text-to-speech provider
Registry entry plus one config-driven FORMAT_HANDLERS handler. The model id
travels in an HTTP `model` header rather than the JSON body, and the voice is
a reference_id (preset or cloned voice model).

Closes #2411
2026-08-14 16:21:11 +07:00
zmf
8ed9da7165 feat(providers): add glm-5.3 to GLM Coding and GLM (China) registries
Zhipu released GLM-5.3 on both api.z.ai and open.bigmodel.cn coding
endpoints. Verified live against both, returning model:"glm-5.3" with
native reasoning_content.

No other changes needed: the '*glm-5*' family pattern in capabilities.js
and 'glm-5*' in pricing.js already cover it.
2026-08-14 16:16:46 +07:00
decolua
7e5f5a8813 fix(claude): re-anchor passthrough cache breakpoints with 1h TTL
Passthrough kept the client's own cache_control markers, which point at
pre-normalization offsets. Once normalize/dedupe reshaped system and tools,
the breakpoints landed mid-array and the tail was re-cached every request.

- Pin the last system block and last tool at ttl 1h (was the client's 5m)
- Anchor the last assistant turn at 5m, falling back to the final message
  so a first turn still gets a breakpoint
- Fold mid-conversation system messages into the neighbouring user turn
  instead of hoisting them into body.system, where the volatile token
  counters invalidated the prefix on every request
- Run the anchoring after every token saver, at the final body

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-14 16:08:30 +07:00
Bertho Joris
345cdcf6a5 fix(combo): detect images from Hermes and attachment payloads for Vision Adapter
Inspect images[], experimental_attachments/attachments, message-level
image/image_url/audio_url, and inline data:image|audio|pdf URIs on trailing
user turns so Vision Adapter auto-switch fires for Hermes/Ollama/Vercel AI
SDK shapes. stripOpenAI now also drops msg.images and image attachments when
the active model lacks vision support.
2026-08-13 18:30:44 +07:00
Duc Nguyen
65197ad11c feat(auth): add native SAML 2.0 SSO integration
Add SAML 2.0 as a second SSO protocol alongside OIDC under a unified
authMode/ssoType model. SP flows via @node-saml/node-saml: AuthnRequest
generation, ACS POST assertion handling, SP metadata export, and admin
config test endpoint. Replay-protected via saml_state cookie (httpOnly,
SameSite=Lax) matched against InResponseTo; wantAssertionsSigned enforced.

- src/lib/auth/saml.js: SAML instance builder, X.509 cert formatter, claim pickers
- 4 routes under src/app/api/auth/saml/: start, acs, metadata, test
- settingsRepo: ssoType + saml* defaults; login/status routes dispatch by type
- profile page: SSO protocol switcher, IdP metadata XML + cert uploaders
- login page: dynamic SAML sign-in button; Header: SAML user badge
2026-08-13 17:56:34 +07:00
Fadjrir Herlambang
e02bde4a70 feat(providers): add Kimchi API key support (dual OAuth + API key)
Kimchi's transport is OpenAI-compatible (Authorization: Bearer) but the
registry declared it OAuth-only, so the dashboard, /api/providers, and
the connection test all rejected API keys. Enable dual auth
(authModes: ["oauth", "apikey"]) and add a kimchi case to
testApiKeyConnection so the Test Connection button works for both modes.
Regenerate the golden snapshot with the Kimchi entries (+ other
previously-missing providers).
2026-08-13 17:53:44 +07:00
haumanto
30fec4318e fix(models): expose snake_case token limits on /v1/models 2026-08-13 12:19:06 +07:00
CyrixJD115
67271d859e fix(opencode): send official client headers on free-tier requests
Mirror the official opencode CLI fingerprint (User-Agent, x-opencode-session, x-opencode-request, x-opencode-project) on free-tier requests so the Console no longer classifies traffic as an unidentified client and rate-limits it with FreeUsageLimitError / HTTP 429.

Session id resolves conversation-stable via resolveSessionId (client session to assistant-text hash to connection) to preserve prompt caching, normalized into opencode ses_ format with a generated fallback. When the downstream client is already opencode, its headers are forwarded as-is.
2026-08-13 12:13:51 +07:00
stoXmod
b566b20ade fix(antigravity): strip competitive system prompts to prevent 429 quota errors
Zed IDE injects a Claude-agent system prompt that Antigravity flags as
competitive, blocking the request with a 429 Quota Exhausted response.
Scan systemInstruction.parts and remove the prompt before dispatch.
2026-08-13 12:11:30 +07:00
Clayton Tavares
6d30ce6de5 fix: Fusion strip stream_options + reasoning model test probe
- combos: strip stream_options from Fusion panel fan-out to avoid DeepSeek 400 (#3024)
- dashboard: raise model-test probe budget to 1024 + soft-pass reasoning-only responses (#3010)
2026-08-13 11:56:43 +07:00
rm1dev
5b417f9bf2 fix(kiro): intercept chat via x-amz-target and prepend initial-response frame
Kiro IDE 1.0.228+ moved GenerateAssistantResponse from path
/generateAssistantResponse to POST / + x-amz-target header, so chat turns
bypassed MITM. The SmithyMessageDecoderStream also now requires an
initial-response frame at stream start, and agent/vibe mode sends
modelId "auto" which had no mappable slot.

- Add isChatRequest() header-based match for kiro in mitm/config.js
- Add buildInitialResponseFrame/withInitialFrame to emit the mandatory
  initial-response once per stream (kiro.js)
- Add "auto" model slot and update mitmDomain to runtime.us-east-1.kiro.dev
2026-08-13 11:56:31 +07:00
Cokky Turnip
b57c041345 fix(providers): add llm7 to provider test support 2026-08-13 11:56:06 +07:00
zmf
8a527fec91 fix(security): SSRF guard on search baseUrl, default-password remote login, and request-details redaction
- resolveBaseUrl() rejects client-supplied non-public baseUrls via assertPublicUrl (SSRF guard on /v1/search)
- fresh-install remote login with default password returns 403 without issuing a JWT
- /api/usage/request-details redacts request/providerRequest/providerResponse/response payloads
- declare chalk and prop-types in package.json (used but previously undeclared)
2026-08-13 11:50:25 +07:00
Nguyen Thanh Dat
70ba0024b0 fix(translator): preserve prompt_cache_key when converting chat to responses 2026-08-13 11:46:49 +07:00
brimob-sowax
80afb59907 fix(qoder): detect billing blocks at stream start, return 403 for failover
Peek the first SSE frame in wrapQoderSSE; if statusCodeValue != 200 and the
body carries a billing signature (code 112/10605 or pricingUrl), return a
synthetic 403 so chatCore marks the connection unavailable and triggers
combo/account fallback instead of leaking the error text into chat.

wrapQoderSSE becomes async; consumed peek bytes are re-processed in the
stream start() seed loop so nothing is dropped.
2026-08-13 11:43:11 +07:00
chisewaguri
10a923da11 fix(responses): don't close message on empty tool_calls array
Some providers (e.g. codebuddy/cbcn) attach an empty tool_calls array to every streaming chunk. An empty array is truthy in JS, so the guard 'if (delta.tool_calls)' closed the message on the first content token and emitted response.output_text.done early, dropping the remaining deltas. Guard on a non-empty array; finish_reason still closes the message and real tool calls still close it before emitting function_call items.

fixes #3234
2026-08-13 11:40:45 +07:00
yusei21
01858feca0 docs(i18n): add Brazilian Portuguese documentation 2026-08-13 11:40:26 +07:00
Moein Arabi
e2a4fe048f fix(hermes): add api_key parameter to model block in YAML configuration 2026-08-13 11:35:02 +07:00
nguyenha935
b44bb09f72 fix(kiro): report real output tokens and stop discarding usable turns 2026-08-13 11:33:41 +07:00
decolua
456f2a2635 feat(usage): wire force flag through client + usage route
Manual refresh (↻) sends ?force=1 so it bypasses the Claude quota cache (dedup + TTL) added in cd4003bc. Auto-refresh and multi-tab stays cached, so Anthropic's usage endpoint is no longer hammered.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-13 11:31:07 +07:00
decolua
cd4003bc8b feat(usage): dedup + cache Claude quota calls to avoid 429
Multiple tabs/accounts/auto-refresh funneled straight to Anthropic and tripped 429. Add a 120s TTL cache keyed by access token with in-flight promise dedup, serve the last good read on soft failure, and thread a force flag through getUsageForProvider for manual refresh. Also lower the dashboard poll cadence (180s to 600s) and stable group-by-provider so connection order stops jumping.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-13 11:27:57 +07:00
decolua
71dcdc1053 fix(headroom): toggle reflects enabled setting even when proxy is down
Toggle was checked={headroomEnabled && headroomRunning} and disabled when the proxy was down, so a downed proxy showed OFF while headroomEnabled stayed true in the DB. The engine only checks headroomEnabled, so it kept calling /v1/compress. Toggle now reflects the user setting; proxy up/down stays visible via the status chip.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-13 11:27:49 +07:00
a3182a7265 merge: integrate origin/master (v0.5.50) into gitea/new_feature
- Resolve conflicts in chatCore handlers: keep apiKey/streamErrorPatterns
  from the details-filters feature, adopt origin's stripContinuityFields,
  customToolNames, cache-inclusive usage accounting, and Responses-API
  SSE→JSON conversion
- Adopt origin's provider usage handlers (codebuddy-intl, qoder creds)
  and modality detection (audio/video inputs)
- Keep requestDetails apiKey column (schema v2) + masked key persistence

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-08-06 09:54:31 +07:00
386b25ff7f feat(dashboard): add model/status/account/api-key filters to usage details tab
- Add apiKey column to requestDetails (schema v2 + migration 002)
- Persist masked API key via buildRequestDetail across chatCore handlers
- Add getRequestDetails apiKey filter + distinct models/apiKeys/statuses helpers
- New /api/usage/filters endpoint returning grouped connections per provider
- Group account dropdown by provider using <optgroup>; add Status column
  with success/error badge to the details table

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-08-06 09:41:03 +07:00
decolua
15223724c3 # v0.5.50 (2026-08-05)
## Features
- **Providers**: add TokenRouter (300+ models via OpenAI-compatible gateway) with
  exact per-model pricing for 110 models and `reasoning_effort` thinking config
- **Providers**: add Self-hosted STT / TTS / Embedding — point 9Router at your own
  OpenAI-compatible speech and embedding servers (whisper.cpp, faster-whisper,
  Kokoro-FastAPI, llama-server, vLLM, Infinity). Unlike the named cloud providers
  these read `baseUrl` per connection, so one provider can front several machines
- **Combos**: default-enable vision/audio capacity adapter (auto-routes to a
  vision/audio-capable model when the target lacks that capability, falling back
  to `oc/mimo-v2.5-free`), wired into chat handler routing
- **Endpoint**: auto-provision a "Default Key" for first-time users so `/v1`
  works without a manual dashboard step
- **Codex**: support GPT-5.6 Max/Ultra reasoning-level overrides (cx/ routes only)
- **Qoder**: support PAT (Personal Access Token) connections end-to-end, alongside
  OAuth device flow
- **CLI tools**: add OpenDesign (manalkaff/opendesign) support
- **Headroom**: report effective payload savings (tool schema/history bytes broken
  out, byte-savings % reflects actual outbound reduction)
- **Ollama**: Cloud quota tracker (session + weekly) + proactive background OAuth
  token refresh scheduler for all providers

## Fixes
- **Providers**: remove Qwen (OAuth flow stopped working reliably)
- **Passthrough**: detect codex-tui/Codex Desktop as native Codex client — they
  were falling through to the translator and losing fields like `reasoning.summary`
- **OAuth**: scope antigravity header fixes to loadCodeAssist/onboardUser only
- **OAuth**: keep `open` external in the build so xAI/Grok token refresh works on
  Windows
- **OAuth**: declare missing `searchParams` in register-session handler (was a
  500 instead of JSON on error)
- **DB**: `ENABLE_REQUEST_LOGS` env var now overrides the UI setting correctly;
  observability defaults to off (opt-in)
- **Translator**: preserve Codex Responses Lite tool use across chat-native
  OpenAI-compatible providers
- **Translator**: don't drop image-only user messages in `prepareClaudeRequest`
- **Translator**: drop JSON Schema keywords Gemini rejects (`uniqueItems`,
  `contains`, `multipleOf`, `unevaluatedProperties`, `unevaluatedItems`,
  `contentSchema`)
- **Claude**: remove global header cache that leaked one client's identity
  headers onto another client/account sharing the server; gate `anthropic-beta`
  by model instead
- **Antigravity**: drop retired Gemini 3.0 quota tiers, show Gemini 3.6 Flash
  usage bars
- **Cloudflare AI**: declare API key authentication (dashboard showed "No
  connections" despite an active key)
- **GitHub Copilot**: hold monthly-exhausted accounts until UTC month reset
  instead of only cooling down 120s
- **CodeBuddy**: dodge Tencent CN content filter, add usage tracking, normalize
  codebuddy-intl messages
- **Usage**: stop losing cached prompt tokens in the forced-SSE→JSON path
- **Grok CLI**: display the public subscription tier from the OAuth token claim
- **Providers**: count apikey connections for Ollama free-tier card; free-tier/
  apikey providers without `authModes` now default to apikey (were treated
  oauth-only)
- **Build**: include static/public assets in standalone output (login page hung
  on 404s when run via PM2)
- **Server**: support IntelliJ IDEA OpenAI-compatible clients over HTTP (h2c
  upgrade handling)
- **Auth**: redirect already-logged-in sessions away from `/login`
- **CLI tools**: enable Apply button for dynamic OpenAI/Anthropic-compatible
  provider connections
- **CLI**: include complete API artifacts in the CLI package
- **TTS**: a bare self-hosted model name is the MODEL, not the voice — `kokoro`
  was parsed as a voice against a default model, 404ing or synthesising with the
  wrong one
  endpoint that drops packets never returns headers, so the request previously
  hung indefinitely
2026-08-05 16:49:14 +07:00
decolua
35f86e5828 fix(oauth): scope antigravity header fixes to loadCodeAssist/onboardUser only
Google fingerprints User-Agent/Client-Metadata on loadCodeAssist and
onboardUser, silently refusing to provision a cloudaicompanionProject
when they don't match the real IDE. Split antigravity's headers out of
the shared gemini-cli constants instead of overwriting them, so the fix
doesn't touch gemini-cli or any other provider.

Inspired by #3000 (thanks @stoXmod for flagging the resource-exhausted
issue), rewritten to keep gemini-cli untouched.
2026-08-05 16:40:09 +07:00
Dasep Moch Luay
41588bea01 feat(providers): add TokenRouter accurate pricing + thinking config
Adds exact per-model rates for 110 TokenRouter models (pulled from
TokenRouter's own pricing API) plus a dedicated thinkingFormat case
(reasoning_effort enum low/medium/high/xhigh/max) and the provider
logo. Provider registration itself already landed in a prior commit;
this fills in what PR #3043 added on top.
2026-08-05 16:31:03 +07:00
decolua
03f8487cc7 test(baseline): regenerate provider/alias snapshots
Sync alias-baseline.json and providers-baseline.json with the current
registry (poolside, tokenrouter, selfhosted-* providers already added;
stale claudeOverlay hook and Kiro X-Amz-Target header already removed).
2026-08-05 16:27:17 +07:00
decolua
99639c0540 test(capacity-adapter): remove unit test file 2026-08-05 16:25:47 +07:00
decolua
e41d85037d test(capacity-adapter): add unit coverage for the capacity adapter service
Covers pool flattening, model augmentation for required capabilities,
context-window history stripping, and the withCapacityAdapterStripping
wrapper.
2026-08-05 16:25:12 +07:00
decolua
02c66fe2bd feat(endpoint): auto-provision default API key for first-time users
- Endpoint page auto-creates a "Default Key" when no keys exist yet,
  so /v1 works out of the box without a manual dashboard step
- Show/copy key buttons stay visible instead of opacity-0 by default
2026-08-05 16:22:06 +07:00
decolua
dcdd4628b3 fix(providers): remove Qwen provider support
Qwen OAuth flow (portal.qwen.ai) stopped working reliably; drop the
executor, registry entry, OAuth provider/service, token refresh
profile, usage handler, and related test coverage and baselines.
2026-08-05 16:17:26 +07:00
decolua
6498b3122f feat(combos): wire capacity adapter into chat handler routing
- detectRequiredCapabilities: infer audioInput/videoInput from block
  type and embedded mime, not just vision/pdf
- handleChat / handleSingleModelChat: augment combo and single-model
  routing with capacity-adapter models when the target lacks a
  required capability, wrapped with history stripping for the
  adapter model's context window
2026-08-05 16:12:55 +07:00
decolua
8e59093db7 feat(combos): default-enable vision/audio adapter with mimo fallback
- Enable vision + audioInput capacity-adapter pools by default for new
  and existing users (mergeWithDefaults backward-compat)
- Fall back to oc/mimo-v2.5-free when an enabled pool has no models
  configured, both in the backend resolver and the combos UI (auto
  refill on removing the last model from a pool)
- Hide PDF/Video from the Vision Adapter UI (PDF never implemented,
  Video lacks translator support) while keeping the settings shape
- Exclude combos from the model picker when opened from the Vision
  Adapter section
- mimo-v2.5 registry entry now declares audioInput/videoInput
- Simplify combo strategy and Vision Adapter descriptions
2026-08-05 16:09:42 +07:00
decolua
cd13d904d7 fix(passthrough): detect codex-tui/Codex Desktop as native Codex client
detectClientTool only matched the legacy "codex-cli" User-Agent, so the
current codex-tui CLI and Codex Desktop (UA "Codex Desktop", originator
"codex_work_desktop") fell through to null and lost native passthrough —
their requests got re-translated, stripping/overwriting fields like
reasoning.summary instead of forwarding the client body as-is.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 15:53:20 +07:00
omar-nahhas
fe547f4dc0 feat(providers): self-hosted OpenAI-compatible STT, TTS and embedding providers
Add Self-hosted STT/TTS/Embedding providers that read baseUrl per connection
instead of a fixed registry endpoint, so 9Router can point at whisper.cpp,
faster-whisper, Kokoro-FastAPI, llama-server, vLLM, Infinity, and similar
OpenAI-compatible local servers.

Self-hosted Embedding refuses to run without a baseUrl rather than falling
back to api.openai.com like openaiCompatNode does, since that fallback would
silently send input text and the API key to OpenAI under a provider named
"Self-hosted". Also fixes embeddingsCore to catch adapter build errors as a
400 instead of letting them escape uncaught, and bounds the upstream fetch
with FETCH_CONNECT_TIMEOUT_MS to avoid hanging forever on a dead endpoint.

Self-hosted TTS treats a bare model value as the model rather than the voice,
since the generic OpenAI TTS convention (bare = voice) is backwards for a
provider where the model is the variable part.
2026-08-05 13:38:13 +07:00
Ubuntu
b480892952 feat(providers): add TokenRouter provider
OpenAI-compatible gateway exposing 300+ models (OpenAI, Claude, Gemini,
Qwen, DeepSeek, Kimi, GLM, and more). Registered as p116, append-only.
2026-08-05 13:32:46 +07:00
RobertsXML
3fab15ae3e fix(db): implement ENABLE_REQUEST_LOGS env var override
- Add config priority chain: ENABLE_REQUEST_LOGS > UI setting > OBSERVABILITY_ENABLED fallback
- Fix transaction callback syntax from arrow to function
- Update saveRequestDetail guard to early return instead of semicolon
- Default enableObservability to false (opt-in)
2026-08-05 13:30:28 +07:00
nguyenha935
d06e0d26c6 fix(translator): preserve Responses Lite tools across Chat providers
Codex Responses Lite clients routed to a chat-native OpenAI-compatible
provider lost tool use in three places: non-streaming Chat responses
leaked the raw chat.completion envelope instead of Responses output
items, internal reasoning continuity fields leaked into the outbound
Chat body causing some upstreams to reject the request, and the
Responses to Chat request translator ignored additional_tools,
custom_tool_call, and custom_tool_call_output items entirely.

Also fixes apiType (chat vs responses) for openai-compatible nodes
being resolved from the immutable provider ID instead of the stored
node config, so editing a node's API Type had no runtime effect.
2026-08-05 13:27:25 +07:00
decolua
b11be8be0a fix(antigravity): drop retired Gemini 3.0 tiers from quota tracker
gemini-3-flash-agent, gemini-3-flash, and gemini-3-pro-image are retired;
they no longer need a quota bar in the tracker.
2026-08-05 13:24:57 +07:00
whale9820
42c691b3ea feat(antigravity): show Gemini 3.6 Flash usage bars in quota tracker
getAntigravityUsage filtered the fetchAvailableModels response through a
hardcoded importantModels list that only contained 3.5 Flash, silently
dropping the 3.6 Flash quota buckets so no usage bar rendered.
2026-08-05 13:21:38 +07:00
cardinusantara
a7941ddab4 fix(translator): don't drop image-only user messages in prepareClaudeRequest
hasValidContent() only treated text/tool_use/tool_result blocks as valid
content, so a user message containing only an image block was filtered
out as empty. When it was the only non-system message, this left an
empty messages array and Anthropic rejected the request.
2026-08-05 13:20:27 +07:00
Sutarto Jordan Chrisfivo
646b3b9ba3 fix(cloudflare-ai): declare API key authentication
Cloudflare AI registry entry was missing authType/authModes, causing
the dashboard to report "No connections" despite an active API-key
connection. Closes #2969
2026-08-05 13:13:54 +07:00
huuanh20
baebc9a06e docs(i18n): fix port typo and add RTK Token Saver features
Fix port 201281 -> 20128/v1 typo in README.zh-CN.md diagram, add
RTK Token Saver mention to README.vi.md/README.zh-CN.md, update
Tier 3 free providers to 2026 lineup, and drop hardcoded /tmp
node_modules path from tests/package.json and tests/README.md.
2026-08-05 11:55:02 +07:00
Matt Van Horn
25e4bf1c6c fix(cli): include complete API artifacts in CLI package
Merge the complete generated .next-cli-build/server tree into the
packaged CLI after the standalone copy, since Next's standalone output
is trace-pruned and can omit route modules (e.g. /api/v1/messages) or
chunks loaded dynamically. Add a post-copy integrity check for the
required API route artifacts so an incomplete package fails during
pack:cli instead of at runtime.

Fixes #2945
2026-08-05 11:54:16 +07:00
MiQieR
c570fe33ae feat(tts): add Xiaomi MiMo text-to-speech support
Adds mimo-v2.5-tts as a Media Provider TTS through the existing
OpenAI-compatible chat-completions endpoint. Voice is selected via the
top-level audio.voice field, and an optional style/language hint is
threaded through tts.js -> ttsCore.js -> the new adapter.
2026-08-05 11:46:23 +07:00
ryanngit
d0751bcff7 fix(grok-cli): display public subscription tier
Map the Grok OAuth access-token tier claim to the public plan label
and prefer it over internal /user entitlement names. Fails open to
existing plan detection for opaque or malformed tokens; upstream
remains authoritative for access and quota enforcement.
2026-08-05 11:43:40 +07:00
alfep
948dd8f89b fix(oauth): declare searchParams in register-session POST handler
Missing declaration caused a ReferenceError -> 500 HTML response instead
of JSON when clients called POST .../register-session.
2026-08-05 11:41:10 +07:00
seakleang.nhak
86131b9ca4 feat(codex): support GPT-5.6 Max and Ultra overrides
Add "ultra" reasoning level for Codex GPT-5.6 Sol and Terra, and expose
Max for Luna (Luna falls back Ultra to Max since it is not supported
upstream). Scoped to cx/ routes only; Kiro and generic OpenAI routing
unchanged.
2026-08-05 11:39:59 +07:00
Diwak4r
651df2f0e2 feat(cli-tools): add OpenDesign (manalkaff/opendesign) support
Adds a guide-type CLI Tools entry for OpenDesign, the open-sourced
claude.ai/design skills pack. It has no standalone config - it inherits
the host agent's model/provider config - so once the host (Claude Code,
Cursor, Codex, Gemini CLI, OpenCode) points at 9Router, /opendesign
sessions route through automatically.
2026-08-05 11:36:37 +07:00
minhnhat166
da8691f866 feat(headroom): report effective payload savings
Break out tool schema and tool-history bytes in the size snapshot and
add an effective byte-savings percentage so token-saved logs reflect
the actual outbound payload reduction, not just processed content.
2026-08-05 11:32:38 +07:00
decolua
13ed14568d fix(claude): remove global header cache, gate anthropic-beta by model
The global claudeHeaderCache singleton overlaid the last-seen Claude Code
client's identity headers onto every subsequent request, leaking one
client's headers (anthropic-beta, user-agent, x-stainless-*, etc.) onto
another client/account sharing the same server. Removed the singleton and
the claudeOverlay hook entirely, falling back to static per-provider
headers. anthropic-beta is now computed per-request from the requested
model, gating heavy-agent flags (advanced-tool-use, effort) to
opus/sonnet only.
2026-08-05 11:32:14 +07:00
decolua
1eb37db32d refactor(qoder): dedupe PAT exchange logic, validate PAT keys properly
PAT-to-job-token exchange was duplicated between the executor and the model service, each with its own cache. Consolidate into qoderModels.js and have the executor import it.

Also add a qoder case to the API-key validate route - the generic OpenAI-compat probe cannot validate a PAT (needs job-token exchange + COSY signing first), so bulk-add always reported unknown for qoder keys.
2026-08-05 11:26:22 +07:00
mannnrachman
d433c0b295 feat(qoder): support PAT (Personal Access Token) connections end-to-end
Adds pt-... token auth as an alternative to OAuth device flow. A PAT can't
sign COSY requests directly, so it's exchanged for a short-lived job token
(jt-...) plus userId via openapi.qoder.sh, then used for signing.

Also fixes job-token traffic (jt-...) being rejected by api3.qoder.sh with
403 "Login expired" — the official qodercli serves jt- traffic from
api2.qoder.sh instead, so buildUrl/model-list routing now branches on it.

Quota usage and the dashboard add-key modal are updated to resolve PAT
credentials and label the field correctly, and bulk-add now validates
each key so it gets a real testStatus instead of a hardcoded "unknown".
2026-08-05 11:16:50 +07:00
ryanngit
3292dfc102 fix(github): hold monthly-exhausted accounts until reset
Lock GitHub Copilot connections account-wide until 00:00 UTC on the
first of next month when the upstream 402 response indicates the
monthly additional-usage-limit was hit, instead of only cooling down
the requested model for 120s. Other GitHub 402 responses keep the
existing model-scoped cooldown.
2026-08-05 11:00:40 +07:00
Rafi Mahardika
9138c99391 fix(codebuddy): dodge Tencent filter for CN, add usage tracking & normalize messages for INT
Neutralize CLI-agent system prompts that trigger CodeBuddy CN's content filter, add usage/quota tracking for codebuddy-intl sharing CN's logic, and normalize codebuddy-intl request messages to the shape it expects.
2026-08-05 10:52:03 +07:00
omar-nahhas
41606a37a3 fix(usage): don't lose cached tokens in the forced-SSE->JSON path
handleForcedSSEToJson dropped cached prompt tokens in two ways: the
Responses branch summed only input_tokens, which excludes cache_read
and cache_creation on cache-capable upstreams (measured 2012 reported
vs ~5344 actual, 5332 from cache); and the Chat Completions branch
computed usage correctly but it didn't always reach the client (an
Anthropic response with cache_read_input_tokens: 11022 arrived with no
usage field at all). Now folds cache counters into prompt_tokens,
surfaces them via prompt_tokens_details, and re-attaches usage before
serialisation.
2026-08-05 10:45:55 +07:00
Muhammad Usama
2abe8b855c fix(translator): drop JSON Schema keywords Gemini has no field for
Tool schemas carrying uniqueItems, contains, multipleOf,
unevaluatedProperties, unevaluatedItems, or contentSchema get rejected
by the Gemini API with "Unknown name ...: Cannot find field", failing
the whole request. Add them to UNSUPPORTED_SCHEMA_CONSTRAINTS alongside
the existing stripped keywords (minItems, maxItems, format, ...).
2026-08-05 10:42:54 +07:00
Tomauskasz
c06cc08453 fix(oauth): keep open external so xAI/Grok token refresh works on Windows
`open` derives its own directory from import.meta.url at module scope.
Webpack replaces that with the build machine's absolute path as a
string literal, so a release built on macOS ships a file:///Users/...
URL that fileURLToPath rejects on Windows (no drive letter), throwing
on import. refreshXaiToken dynamic-imports the xAI OAuth service, which
imports open eagerly, so every Grok token refresh silently failed and
was swallowed by a catch that only logs a warning.

Add open to serverExternalPackages so it keeps its real import.meta.url
at runtime, and bundle it into the CLI package via ensureModuleInBundle
(same guard already used for sql.js) since externalizing it means
webpack no longer traces/copies it automatically.
2026-08-05 10:38:25 +07:00
Cokky Turnip
d6df6576c5 fix(providers): count apikey connections for ollama freeTier provider
ollama's registry entry lacked authModes, so dualAuthTypes on the providers page defaulted to oauth and its apikey connections showed as No connections on the freeTier card.
2026-08-05 10:34:44 +07:00
DaDecky
786b3013ba fix(build): include assets in standalone output
With output: "standalone", next build writes server.js under
.next/standalone but leaves generated static/public assets in the
project root, so starting the standalone server directly (e.g. via PM2)
404s on JS/CSS/font/favicon requests and /login stays stuck loading.
Add a postbuild step that copies .next/static and public into the
standalone directory, skipping the workspace-traced CLI build which
already copies its own assets.
2026-08-05 10:32:55 +07:00
dajinglingpake
0648e9e420 fix(server): support IntelliJ IDEA OpenAI clients over HTTP
JetBrains Runtime (JBR 25+) sends an h2c upgrade on OpenAI-compatible requests, which the HTTP/1.1 server would otherwise close. Intercept the upgrade, replay the buffered request through the existing handler, and respond over HTTP/1.1.
2026-08-05 10:32:32 +07:00
DaDecky
ae4f76c433 fix(auth): redirect active sessions from /login
/api/auth/status did not expose whether the auth cookie corresponds to a
valid dashboard session, so /login could only detect "auth disabled"
(requireLogin === false) and not "already logged in". Add authenticated
to the status response and redirect from /login when it's true.
2026-08-05 10:27:27 +07:00
lazysaltyfish
918b3c87a1 fix(cli-tools): enable Apply button for dynamic OpenAI/Anthropic-compatible providers
getAllAvailableModels() only consulted the static PROVIDER_MODELS catalog,
which has no entry for dynamically-registered compatible providers
(id like openai-compatible-chat-uuid). Fall back to the connection's
own defaultModel/customModels/placeholder, mirroring ModelSelectModal.js.
2026-08-05 10:23:59 +07:00
techysy
0e5da70cb1 fix: freeTier/apikey providers without authModes default to apikey in dualAuthTypes
Free-tier and apikey providers (e.g. cloudflare-ai, byteplus, ollama, vertex) whose registry entry omits authModes were treated as oauth-only, hiding their apikey connections on the providers grid card.
2026-08-05 10:22:57 +07:00
2a37a4085e chore: normalize formatting (2-space → tabs) in stream-error-patterns files
Re-tab only — no logic changes. Follows the repo's tab-based formatting
for these files, matching the CommandCode executor/translator style.

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-08-05 09:03:41 +07:00
e2f8323ab1 docs(open-sse): in-stream error handling recipe + CHANGELOG 2026-08-04 23:35:57 +07:00
9bd7adc556 feat(dashboard): Stream Error Patterns editor on provider page
Per-provider textarea (one pattern per line: plain text or /regex/flags)
saved via /api/settings as streamErrorPatterns. Mirrors the existing
providerTimeouts load/save pattern.
2026-08-04 23:35:25 +07:00
34a78f579e feat(open-sse): mark streaming requestDetail as error on stream-error pattern match 2026-08-04 23:34:21 +07:00
c8b96a61e7 feat(open-sse): non-streaming stream-error pattern match → 502 fallback 2026-08-04 23:34:00 +07:00
e438a03f96 feat(open-sse): early-peek stream error detection + fix UTF-8 loss in CommandCode peek
- new open-sse/utils/streamErrorPeek.js: bounded peek of the first bytes of a
  200 stream; configured pattern match → 502 so account/combo fallback can run
  before any byte reaches the client (streaming included). Re-emits RAW bytes
  so split multi-byte UTF-8 sequences survive the peek (never re-encode
  decoded text — TextDecoder flush corrupts a lone leading byte to U+FFFD).
- chatCore: run the peek after executor.execute when the provider has
  streamErrorPatterns configured; chat.js passes the settings through.
- commandcode executor: same raw-bytes fix in peekForUpstreamError + regression
  test that fails against the old flush-based re-encode.
2026-08-04 23:33:12 +07:00
008e0ef311 feat(settings): default streamErrorPatterns key 2026-08-04 23:28:17 +07:00
5058a402f1 feat(open-sse): streamErrorPatterns matching util (text + regex) 2026-08-04 23:27:49 +07:00
9b27ee2611 fix(open-sse): treat CommandCode in-stream error events as request failures
Upstream emits AI SDK v5 {"type":"error"} events inside an HTTP 200 stream.
The translator turned them into fake success content ([CommandCode error: ...]
+ finish_reason stop), so account/model fallback never fired and logs showed
Status: success.

- translator: error events now emit an OpenAI-shaped error chunk (chunk.error)
  instead of content; parseSSEToOpenAIResponse already detects chunk?.error
- executor: peek the first events before committing the response; an early
  error event returns 502 so fallback runs before any byte reaches the client
2026-08-04 23:26:37 +07:00
c5ce1ef140 feat(commandcode): quota usage dashboard, CLI-parity request, and connect timeout fixes
- Add CommandCode usage handler mirroring the official CLI /usage
  (whoami → credits/subscriptions → summary) with 5h/weekly/monthly
  quota rows, registered in services/usage.js and registry config
- Parse commandcode quota rows in ProviderLimits (remainingPercentage + $ unit)
- Match official CLI request shape: x-command-code-version 1.10.0,
  User-Agent cli, and config.environment '${platform}-${arch}, Node.js ${version}'
- Fix connect timeout unit confusion: both profile and provider pages
  now use ms with a 1s minimum guard (prevents 60ms footgun)
- Fix fetchT0 ReferenceError in base.js error path and log fetch
  diagnostics only on upstream failure
- Quota Tracker defaults to the Active account filter
- Ignore .commandcode/ CLI local state
2026-08-04 22:34:13 +07:00
fcd3dcb409 feat(translator): add vision support for commandcode provider
Map OpenAI image_url / Claude-style image blocks to the {type:"image",
image:"<data URI|url>"} shape the command-code CLI sends to /alpha/generate
instead of dropping them to "[image omitted]". Handles data URIs, raw base64
(with media_type / image/png fallback), and remote URLs.

- Promote bugs-gemini-cursor-commandcode "image content is preserved" from
  it.fails to a real assertion (bug fixed)
- Add vision unit tests to openai-to-commandcode.test.js
- Add test-commandcode-vision.sh curl helper for live verification

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-08-03 17:06:32 +07:00
067f18aaa1 fix(usage): ReferenceError in byProvider lastUsed overlay broke daily-summary periods
The provider lastUsed overlay loop referenced histRows before its const
declaration (temporal dead zone), throwing ReferenceError for 7d/30d/60d/all
which use the usageDaily summary path — leaving the overview cards and tables
empty regardless of the selected period. Merged the provider overlay into the
existing histRows loop.
2026-08-02 22:48:15 +07:00
B1nh M1nh
f260a1817b feat: Ollama Cloud quota tracker + proactive background OAuth refresh
Ollama: replace informational stub with real quota tracker hitting ollama.com/api/usage (session 5h + weekly 7d, 0..1 ratio) and /api/me plan label; bind handler to apiKey + add features.usageApikey so apikey connections work.

Token refresh: add backgroundTokenRefresh scheduler that refreshes OAuth connections within max(provider lead, 30min) of expiry, independent of inbound traffic (10s after boot, then every 5min, unref'd timers, DISABLE_BACKGROUND_TOKEN_REFRESH kill-switch, fail-open per tick/connection). Registered from custom-server.js (listening) and initializeApp.js. checkAndRefreshToken gains opt-in {force} for the scheduler; request path unchanged.
2026-08-02 09:34:27 +07:00
88faba150a feat(dashboard): model test-all with connection selector, usage by provider, combo enable toggle and showOnlyComboModels setting
- provider detail: Test All Models button runs every model (built-in + custom)
  with an optional connection selector; failed models can be disabled in bulk
- bulk selected-model test now runs all connections in parallel (Promise.all)
- usage overview: new 'Usage by Provider' table view (default) and a Provider
  column in Recent Requests; byProvider now tracks lastUsed
- combos: per-combo enable/disable toggle; disabled combos are skipped by the
  routing engine (getComboModels/getComboModelsFromData) and model listing
- settings: 'Only show combo models' toggle filters ModelSelectModal and the
  /v1/models response to models present in enabled combos
2026-07-31 10:37:45 +07:00
0dbae80930 Merge branch 'master' into gitea/new_feature 2026-07-30 23:14:05 +07:00
decolua
6fcd27337a # v0.5.45 (2026-07-30)
## Features
- **Providers**: add Poolside (OpenAI-compatible)
- **Providers**: add api-airforce, baidu, bazaarlink, bluesminds, kilo-gateway, llm7, morph, sambanova, tencent
- **OAuth**: zed / trae / windsurf providers + harden callback proxies
- **CLI tools**: set Claude Code max context tokens
- **Qoder**: PAT auth + refresh model list
- **Gemini**: Gemini 3.6 Flash tier routing + Gemini 3.5 Flash Lite
- **Claude**: bump default Opus to `claude-opus-5`
- **Kiro**: add Claude Opus 5 models
- **Usage**: Kimi and DeepSeek usage handlers
- **Usage**: SuperGrok weekly pool via gRPC-web

## Fixes
- **Refresh**: rotate `refresh_token` between retry attempts
- **Kiro**: canonicalize tool history and route API keys correctly
- **Kiro**: normalize dashboard thinking intensity models
- **Cursor**: stop leaking agent tool errors as text
- **Gemini**: fill empty tool schemas after `$ref` strip
- **Antigravity**: strip `stream_options` from non-stream requests
- **Jina-reader**: recover after transient errors, use JSON POST API
- **Usage**: record exact embedding tokens
- **Tunnel**: preserve successor cloudflared PID
- **Console-log**: initialize capture at server boot + prevent SSE proxy buffering
- **Dashboard**: count dual-auth, free-tier OAuth and API-key connections correctly
- **Dashboard**: flex quota rows, thin global scrollbars, no hidden-row overflow

## Docs
- **i18n**: expand pt-BR translation to 986 terms
- README: Indonesian translation
2026-07-30 09:43:55 +07:00
decolua
9be6588cc8 chore: drop source-attribution comments from provider code
Remove "Ported from OmniRoute" and cockpit-tools attribution comments.
User-Agent strings and README/landing credits are left intact.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 21:05:23 +07:00
decolua
1319dea620 fix(providers): count apikey connections for freeTier providers
openrouter, nvidia, gemini lack authModes, so dualAuthTypes on the
providers page defaulted to "oauth" and their apikey connections showed
as "No connections" on the freeTier card.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 20:21:27 +07:00
whale9820
31df0635aa feat(providers): add Poolside provider (OpenAI-compatible)
Adds Poolside (inference.poolside.ai) as an API-key provider using the default OpenAI transport. Registers three Laguna models with reasoning capabilities (262K context, 32K max output).
2026-07-29 20:17:26 +07:00
decolua
baf3356583 fix(ui): count free-tier oauth connections on providers list
Free-tier cards (e.g. kimchi, oauth-only) hardcoded "apikey" for stats and
toggle, so oauth connections were invisible on /dashboard/providers despite
showing on the detail page. Use dualAuthTypes per provider instead.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 20:06:20 +07:00
ridwan kulu
24fd165b0d docs(readme): add Indonesian translation
Add an Indonesian README and link it from the root README language switcher.
2026-07-29 19:38:30 +07:00
Sutarto Jordan Chrisfivo
a8313cd322 feat(kiro): add Claude Opus 5 models
Register Opus 5 and its thinking/agentic variants with 1M context
and adaptive-thinking capabilities.
2026-07-29 19:34:14 +07:00
Fábio A.
f8e8039446 i18n(pt-BR): expand partial translation to 986 terms
Add ~793 new pt-BR UI strings for the dashboard and settings.
2026-07-29 19:31:16 +07:00
Cokky Turnip
e3e3e235f6 fix(gemini): fill empty tool schemas after $ref strip
Vertex rejects orphan {} left when $ref/$defs are removed from function declarations. Promote empty nodes to object+reason placeholder in addPlaceholders.
2026-07-29 19:31:15 +07:00
Nurwanda Romadhon
0afe949387 fix(antigravity): strip stream_options from non-stream requests
OpenAI clients may send stream_options with stream=false; Google
generateContent rejects that combination. Drop it when not streaming.
2026-07-29 19:30:44 +07:00
Kyle Welsworth
5e59790824 fix(cursor): stop leaking agent tool errors as text
Emit SSE error frame for unsupported Cursor AgentService IDE tools
instead of assistant content, and drop frames after the turn finishes
to avoid double-closing the stream controller.
2026-07-29 19:29:27 +07:00
nguyenha935
16cb40fda1 fix(kiro): canonicalize tool history and route API keys correctly
Route API-key inference through Amazon Q first, enforce adjacent
one-to-one tool use/result pairs after session replay, and treat
payload-invalid HTTP 400 as terminal.
2026-07-29 19:27:41 +07:00
decolua
44c7b34837 fix(ui): count dual-auth provider cards correctly
Share oauth+apikey/api_key stats for dual-auth providers (incl. kiro)
so card totals match the detail page.
2026-07-29 19:27:31 +07:00
decolua
15dfd86416 fix(ui): flex quota rows and thin global scrollbars
Replace the fixed-table quota layout with flex rows that shrink cleanly,
keep the hidden-quota chip row from overflowing, and use thin mac-like
scrollbars app-wide.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 18:10:33 +07:00
decolua
8b0fcf4b16 feat(cli-tools): allow setting Claude Code max context tokens
Add a context-window selector on ClaudeToolCard that writes
CLAUDE_CODE_MAX_CONTEXT_TOKENS into settings.json (nudged 2K under the
labeled cap), and clear it on reset/default.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 18:10:31 +07:00
decolua
6d96e24bd9 chore(providers): refresh catalogs, free tiers, and hide stale ones
Update model lists and context lengths across free/apikey providers,
move bazaarlink, kilo-gateway, and kimchi into freeTier, demote llm7 to
apikey, and hide bluesminds, sambanova, zed, and mimo-free.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 18:10:09 +07:00
decolua
3b14bf4a49 feat(devin-cli): bridge client tools via MCP and use full agent
Default to the full agent with built-in tools, expose client function
tools as an MCP server, surface tool calls as OpenAI tool_use, resolve
workspace cwd from the request, and bump context windows.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 18:10:03 +07:00
decolua
9c9dd7b191 feat(qoder): support PAT auth and refresh model list
Exchange Personal Access Tokens for short-lived job tokens, close the
SSE stream on terminal frames so non-streaming clients do not hang,
re-enable OAuth plus API-key auth modes, and replace the model catalog
with the current Qoder aliases.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 18:09:57 +07:00
decolua
f17a68aaee feat(usage): fetch SuperGrok weekly pool via gRPC-web
Decode GetGrokCreditsConfig frames when REST billing returns empty
caps, so SuperGrok weekly quota shows in the usage dashboard.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 18:09:43 +07:00
decolua
6eaa9f8369 feat(usage): add Kimi and DeepSeek usage handlers
Wire /v1/usages for Kimi (OAuth + API key) and balance API for DeepSeek,
flag both providers with usage/usageApikey, and normalize their quotas
in the dashboard ProviderLimits parser.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 18:09:35 +07:00
decolua
65ac9b3cec fix(ui): prevent hidden quota row from overflowing
Use w-full instead of min-w-0/flex-1 + overflow-x-auto so the hidden
quota chips wrap cleanly instead of stretching the row.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-26 10:03:31 +07:00
decolua
72ec06a81d feat(cli-tools): add Devin CLI provider with ACP stdio executor
Wire Devin CLI as a routed provider that spawns the local `devin acp`
binary. Add the DevinCliExecutor, register it in the executor map, expose
its status through the cli-tools batch endpoint and devin-settings route,
and document setup in cliTools constants.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-26 10:03:22 +07:00
decolua
de2da19a9e feat(providers): add api-airforce, baidu, bazaarlink, bluesminds, kilo-gateway, llm7, morph, sambanova, tencent
Register 9 new upstream providers with logos and update the auto-generated
registry index. Refresh providers/alias baselines and extend the alias
token allowlist so verify-alias stays green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-26 10:03:10 +07:00
decolua
aa0448f7e2 fix(refresh): rotate refresh_token between retry attempts
Rotating-RT providers (xAI/grok-cli) issue a new refresh_token on every
refresh; mutate credentials in-place so refreshWithRetry reuses the fresh
RT instead of the already-consumed one.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-25 17:30:11 +07:00
decolua
41c9e6be87 feat(claude): bump default Opus to claude-opus-5
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-25 17:30:02 +07:00
decolua
8e04fe1734 feat(oauth): zed/trae/windsurf providers + harden callback proxies
- zed live model discovery; codebuddy-intl handler; remove duplicate workbuddy
- split oauth providers.js into per-provider files (facade re-export)
- fold 5 standard refresh providers into config-driven generic
- hide trae/windsurf from registry (no tool calling support)
- fix login-CSRF + SSRF on trae/windsurf/zed local callback proxies
  via loopback-origin guard + strict state validation + apiOrigins allowlist
- move zed RSA private key transit to POST body; redact proxy logs

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-25 17:25:19 +07:00
Phuong Lambert
783e271c16 feat(gemini): add Gemini 3.6 Flash tier routing and 3.5 Flash Lite
Add gemini-3.6-flash tiered (high/medium/low) for Antigravity routing
via upstreamModelId "gemini-3.6-flash-tiered(level)" + thinkingLevel,
plus gemini-3.6-flash and gemini-3.5-flash-lite direct API models.

- getModelUpstreamId: split (level) suffix before lookup, re-append after
- Antigravity executor: preserve transformed body.model
- MITM extractModel: parse thinkingLevel for tiered model (default medium)
- Isolate Cloud Code endpoints: discovery (loadCodeAssist/onboardUser/
  quota) on PROD cloudcode-pa, chat transport on daily-cloudcode-pa
  to bypass prod 429
2026-07-23 16:34:24 +07:00
jacardl
3c17d3406b fix(jina-reader): recover after transient errors and use JSON POST API
Clear stale provider error code and account lock after a successful web
fetch (the core fetch handler never consumed the onRequestSuccess
callback), switch Jina Reader to its documented JSON POST request, and
parse the Title: metadata line before falling back to a Markdown heading.
2026-07-23 16:28:37 +07:00
ankit1324
007d372724 fix(kiro): normalize dashboard thinking intensity models
Strip the generic dashboard model(level) suffix before resolving Kiro
synthetic -thinking/-agentic variants so the upstream request no longer
carries an invalid parenthesized model id. Map explicit levels to native
Kiro effort fields only for supported Claude/GPT model families, and stop
advertising native levels for unsupported legacy Kiro models. Applies to
both OpenAI→Kiro and direct Claude→Kiro routes.
2026-07-23 16:27:06 +07:00
ryanngit
e45bd73d6e fix(tunnel): preserve successor cloudflared PID
Make PID cleanup conditional on the exiting child still owning the PID file so a stale exit cannot erase a replacement tunnel's PID. Only null the in-memory process when the exiting child is current. Explicit disable keeps unconditional cleanup.
2026-07-23 16:22:47 +07:00
zie
c85a5c57ba fix(usage): record exact embedding tokens 2026-07-23 16:07:57 +07:00
Duc Nguyen
57b3b2c175 fix(console-log): initialize capture at server boot + prevent SSE proxy buffering
Initialize initConsoleLogCapture() via Next.js instrumentation register()
hook so logs are captured from startup in headless/Docker deployments, and
add X-Accel-Buffering/Cache-Control headers to the SSE stream route to
prevent reverse proxies from buffering the initial payload.
2026-07-23 16:06:21 +07:00
Biuzai OpenClaw Agent
53a8b5ed55 feat(providers): add Gemini 3.6 Flash and Gemini 3.5 Flash Lite models 2026-07-23 15:48:23 +07:00
decolua
039c4dbc72 feat(providers): add trae/windsurf/zed/workbuddy/codebuddy-intl + icons
- New providers: trae, windsurf, zed, workbuddy, codebuddy-intl
  (registry + executor, wired into executors/index.js + registry/index.js)
- zed: port hosted cloud proxy from OmniRoute — RSA access-token → short-lived
  LLM token exchange (shared/zedAuth.js) + NDJSON {event}/{status}/[DONE]
  stream translated back to OpenAI via Claude/Gemini/OpenAI-Responses translators
- Provider icons (128x128 png) for the 5 new providers
- qoder + tokenRefresh provider tweaks

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-23 15:37:24 +07:00
decolua
79918c7830 # v0.5.40 (2026-07-20)
## Features
- **i18n**: add Khmer (km) translations
- **CLI tools**: configure Grok Build subagent models
- **Kimi**: merge OAuth into dual-auth provider, add K3 / K2.7 models
- **Dashboard**: ProviderTopology flow animation

## Fixes
- **DB**: resolve better-sqlite3 parameter binding crash
- **Translator**: pass `service_tier` through OpenAI → Responses conversion
- **Kiro**: map GPT-5.6 reasoning effort fields
- **Kiro**: validate terminal streams before emitting output
- **Kiro**: map GPT reasoning effort fields
- **Codex**: current `client_version` + refresh-aware model sync
- **Alicode-intl**: split into Coding Plan + Model Studio providers
- **Cursor**: HTTP/2 AgentService support + version bump 3.12.17
- **Dashboard**: cut duplicate API/icon spam, lazy-load provider assets
2026-07-20 17:21:41 +07:00
long2ice
6994cd1f70 fix(cursor): HTTP/2 AgentService support + version bump to 3.12.17
Real Cursor IDE now uses AgentService at agent.api5.cursor.sh (HTTP/2-only)
while 9router still spoke the retired ChatService at api2.cursor.sh with
outdated headers, producing HTTP 429 "Update Required". Add an executeAgent
path that builds an agent.v1.RunRequest Connect RPC over a raw http2 stream
and fetches the account-specific usable model catalog via GetUsableModels.

Also implement MCP tool calling over AgentService: encode OpenAI tools as
AgentRunRequest.mcp_tools (McpToolDefinition with google.protobuf.Value
input_schema), decode McpArgs tool calls, and forward them to the client as
OpenAI tool_calls so the client runs the tool and resumes in the next turn.
Reply to request_context_args with a non-empty RequestContext, to server
heartbeats with client_heartbeat, and to KV blob get/set with empty results,
so action queries no longer stall the stream. Fold the client system prompt
into the user message (custom_system_prompt makes the server return an empty
turn). Bump clientVersion to 3.12.17 and add the x-cursor-client-commit
header so the gateway identifies as a current Cursor IDE release.
2026-07-20 15:39:55 +07:00
Lê Tấn Thắng
4f48ab8c7f fix: resolve better-sqlite3 parameter array binding crash
Spread params into better-sqlite3 Statement.run/get/all so positional ? placeholders bind correctly. better-sqlite3 accepts positional args, not an array, so binding crashed whenever a query had parameters. Matches the bun:sqlite and node:sqlite adapters.
2026-07-20 12:08:58 +07:00
Rafli Ahmad Zulfikar
c97963c4fb fix(translator): pass service_tier through OpenAI→Responses conversion
Forward the service_tier field from OpenAI requests into the Responses API payload so clients can select priority/default/flex tiers instead of the field being silently dropped.
2026-07-20 11:24:37 +07:00
Edison42
cef5dd4d61 fix(kiro): map GPT-5.6 reasoning effort fields
Route GPT-5.6 reasoning effort through Kiro's native reasoning.effort field instead of the legacy Claude output_config.effort path. GPT-5.6 models now emit reasoning.effort for low/medium/high/xhigh, with max mapped to the xhigh wire value.

Preserve the Responses API reasoning.effort through the OpenAI intermediate by copying it to reasoning_effort before the field is dropped. Skip legacy thinking_mode prompt tags when a supported native GPT effort is emitted, while keeping the legacy fallback for unsupported values (auto/minimal/ultra) and explicit disable semantics (none/off/disabled). Claude adaptive effort continues to use thinking plus output_config.effort.
2026-07-20 11:11:37 +07:00
Nur Ad-Duja
d587b2a487 fix(codex): current client_version + refresh-aware model sync
Bump client_version to 0.144.6 (above the 0.144.0 gate in codex CLI's
manifest) so /codex/models no longer returns 200 with newest entries
silently filtered out. Add the originator: codex_cli_rs header used by
every other codex call site, and move the entry onto buildOAuthResolver
so token refresh on 401/403 and a warning field on empty results are
wired in, matching gemini-cli and grok-cli.
2026-07-20 10:55:42 +07:00
Edison42
7c7fae3955 fix(kiro): validate terminal streams before emitting output
Validate AWS EventStream framing, header bounds, CRCs, error frames,
and terminal stop metadata before exposing Kiro output. Classify stop
reasons into dispositions (complete / retryable / terminal_incomplete /
refusal) and retry once when the stream ends with a malformed tool call,
ellipsis-only output, or a short future-action sentence.

Fail closed: propagate streaming failures as error SSE (502) instead of
collapsing them into a successful stop, so incomplete responses no longer
leak as final answers.

Detect the observed evidence-prefixed trailing progress final without
broadening the Chinese heuristic to completed findings.
2026-07-20 10:55:33 +07:00
seakleang.nhak
9ba8f37486 feat(i18n): add Khmer language support
Register Khmer (km) locale, add complete 1,394-entry dashboard translation,
and show the Cambodia flag in the language selector. Place Khmer immediately
before Thai in the selector and keep the header locale in sync after switching.
2026-07-19 16:30:37 +07:00
decolua
55628eea02 fix(alicode-intl): split into Coding Plan + Model Studio providers
8b9cac1 swapped alicode-intl to the DashScope compatible-mode endpoint to
fix #2591 for standard DashScope keys, but that broke Coding Plan keys
(sk-sp-...) which only work on coding-intl.dashscope.aliyuncs.com. The two
key types use two different hosts and are not interchangeable.

- alicode-intl: revert to coding-intl endpoint (Coding Plan keys)
- alims-intl: new provider for dashscope-intl/compatible-mode (standard keys)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-19 16:29:55 +07:00
Don Tanggang
c4a120af8f docs(readme): update free-tier provider status for 2026
Outdated free-tier info was misleading new users. Sync corrections into
README.md and README.zh-CN.md, and force IPv4-first DNS resolution in the
CLI launcher to avoid undici IPv6 connect timeouts (502).

- Kiro AI: free tier now ~50 credits/month (was "Unlimited FREE")
- Qwen Code / Gemini CLI: free tiers discontinued in 2026
- OpenCode Free: note free model list fluctuates
- Vertex AI: Gemini API no longer uses $300 free credits since Mar 2026
- cli/cli.js: spawn server/tray with --dns-result-order=ipv4first
2026-07-19 14:12:58 +07:00
Edison42
eb00222c4f fix(kiro): map GPT reasoning effort fields
GPT-5.6 via Kiro needs reasoning.effort while Claude uses output_config.effort.
Resolve the effort path per-model schema (like Kiro CLI/KAS) so GPT-5.6
receives the correct structured thinking level. Claude path unchanged.

- Add resolveKiroEffortPath returning "reasoning" | "output_config" | null
- buildKiroAdditionalModelRequestFields emits schema-specific shape
- Keep prompt tags for backward compatibility
- Add OpenAI/Claude translator coverage for GPT-5.6 effort mapping
2026-07-19 13:53:30 +07:00
tuanminhhole
43d4abbcf2 docs(README): add Vietnamese OpenClaw Zalo video guide 2026-07-19 13:45:20 +07:00
rixzkiye
e0ba667450 feat(cli-tools): configure Grok Build subagent models
Add separate model selectors for Grok Build main, general-purpose,
explore, and plan agents. Each override gets an independent 9Router
custom-model slot and context_window derived from 9Router model
capabilities. Preserve and restore pre-existing config on reset.
2026-07-19 13:35:38 +07:00
decolua
0513bf393f Flow animatopn 2026-07-19 13:19:13 +07:00
d826e39008 merge origin/master into gitea/new_feature
Bring local branch up to v0.5.35 while keeping xAI image/edit, SuperGrok
quota tracking, per-provider timeouts, and pinned model-test actions.
2026-07-17 15:31:47 +07:00
2897cc3972 feat(providers): pin model tests to a specific account
Add Test action on each connection and Test Selected for multi-select.
Model pings go through /api/models/test with x-connection-id so the call
uses only the chosen account, with no round-robin fallback.
2026-07-17 15:23:35 +07:00
decolua
ccb0842d0a fix(dashboard): cut duplicate API/icon spam, lazy-load provider assets
Share one /api/models fetch via useModelCaps cache, mount ModelSelectModal
only when open, stop double fetchModelAliases on CLI tool cards, and resolve
provider icons through a session 404 cache with missing PNGs + loading=lazy.
Also include Claude Exa MCP toggle (claude-settings + ClaudeToolCard) that
was already in the working tree.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-17 12:12:20 +07:00
decolua
68566f53dc feat(kimi): merge OAuth into dual-auth provider, add K3/K2.7 models
Gộp kimi-coding vào kimi (oauth+apikey), parity CLIProxyAPI device flow/headers/refresh.
Thêm K3 + K2.7 Code (+ Kimi Code ids), pricing/caps vision, cập nhật baseline.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-17 12:09:14 +07:00
decolua
bc252ea802 # v0.5.35 (2026-07-16)
## Features
- **xAI**: Grok Imagine video generation (`/v1/videos`) + CLI
- **CLI tools**: Grok Build setup — writes `[model.9router]` to `~/.grok/config.toml`
- **GitHub Copilot**: route Claude models through Copilot's native `/v1/messages`
- **Kiro**: add GPT-5.6 model family (#2596)
- **RTK**: `X-9Router-Token-Saver` header to bypass token savers per request
- **Providers**: quota visibility settings
- **Translator**: drop temperature for all Claude models
- **i18n**: Thai (th) + Persian (fa) translations / README

## Fixes
- **Providers**: bulk-add API keys no longer overwrite existing keys (gap-fill `Key N`)
- **Anthropic**: lowercase `anthropic-version` header to prevent duplication on `/v1/messages`
- **Alicode-intl**: use DashScope compatible-mode endpoint so standard keys work
- **Grok CLI**: align Grok Build with current subscription protocol (#2590)
- **Grok CLI**: surface `expiresAt` so proactive token refresh fires (#2546)
- **Kiro**: improve direct session cache reuse
- **Models**: populate capabilities for live-catalog LLM models
- **Models**: list compatible provider models in `/v1/models`
- **Thinking**: send explicit `thinking:{type:adaptive}` alongside `output_config.effort`
- **Translator**: strip `client_metadata` when converting openai-responses → openai

## Improvements
- **Perf**: skip inactive background services on startup
2026-07-16 18:13:51 +07:00
asynx6
de680e789f fix(providers): bulk-add API keys no longer overwrite existing keys
Bulk-add named auto-generated keys by paste-line index, blind to existing
connection names. The backend upserts apikey connections by exact name
(connectionsRepo), so a colliding generated name silently replaced an
existing key instead of inserting a new one.

Add a collision-aware planner (src/shared/utils/bulkAdd.js) that gap-fills
the smallest free "<base> <n>" against both existing connection names and
names assigned earlier in the same batch, so a generated name is never
reused and the backend always inserts. Applies to auto-named lines, custom
name|apiKey lines, and Cloudflare name|apiKey|accountId lines.

Wire the planner into AddApiKeyModal and pass existing connection names
from the provider detail page. Add unit tests covering gap-fill, custom
names, Cloudflare 3-part format, and robustness.
2026-07-16 17:20:10 +07:00
Tuan Do
6acc3bb965 fix(anthropic): lowercase anthropic-version header key to prevent duplication on /v1/messages 2026-07-16 16:15:24 +07:00
Ella CEO
8b9cac180e fix(alicode-intl): use DashScope compatible-mode endpoint so standard keys work
Switch baseUrl from coding-intl.dashscope.aliyuncs.com (Coding Plan keys
only) to dashscope-intl.aliyuncs.com/compatible-mode so ordinary DashScope
API keys authenticate. Path /v1/chat/completions and preserveCacheControl
quirk unchanged.

Fixes #2591
2026-07-16 15:56:19 +07:00
YasharSL
30d0f6d3d8 docs(README): Add Persian youtube video tutorial 2026-07-16 15:47:40 +07:00
ryanngit
59b7828237 fix(grok-cli): align Grok Build with current subscription protocol (#2590) 2026-07-16 15:33:19 +07:00
ann
d6761c6fb0 feat(xai): add Grok Imagine video generation (/v1/videos) + CLI
Async video job proxy mirroring the existing image-generation layer split:
Next routes → src/sse/handlers/videoGeneration.js (auth gate, account
fallback loop, refresh persistence) → open-sse/handlers/videoCore.js
(transparent upstream proxy, 401 refresh-once/retry-once, secret sanitization).

- POST /v1/videos/{generations,edits,extensions}: byte-exact body forward
  (JSON + multipart), request_id passthrough, Idempotency-Key forwarded
- GET /v1/videos/{request_id}: status/progress/video.url passthrough
- Register grok-imagine-video (kind: "video"); add "video" to MODEL_TYPE_TO_KIND
  so video models stay out of chat lists (also fixes runwayml leak)
- 9router xai video CLI: submit → poll → atomic MP4 download
- No auto-retry of creation POSTs (billable jobs); rotate accounts only on
  401/403/429; sanitize Bearer tokens + credential values from errors/logs

Closes #1285
2026-07-16 15:29:52 +07:00
M0nt
02ccdc2d22 i18n: add Persian (fa) translations for README and UI
Add full Persian (Farsi) translation of README and sync UI literals
to match zh-CN key set (1389 keys). No runtime code changes.
2026-07-16 15:19:38 +07:00
Edison42
9c58ba645e fix(kiro): improve direct session cache reuse
Reshape Kiro direct requests so resumed client sessions reuse Kiro's
cache-affinity fields instead of starting unrelated CodeWhisperer
conversations.

- keep conversationState.conversationId stable when the client sends an
  explicit session id (x-session-id, session_id, conversation_id, Claude
  Code session metadata)
- add a stable conversationState.agentContinuationId per Kiro session
- send conversationState.agentTaskType: "vibe" and agentMode: "vibe",
  matching the normal Kiro CLI/KAS chat path
- move Kiro thinking instructions into Kiro-compatible systemPrompt /
  additionalModelRequestFields instead of generic top-level thinking
- keep volatile timestamp context out of the top-level systemPrompt; it
  remains only in user content fallback
- suppress additionalModelRequestFields for legacy 4.5-era Claude/Kiro
  models that reject it, while defaulting future Claude/Kiro model ids
  to supported
- preserve Kiro meteringEvent credit usage internally for accounting
  without leaking provider-specific fields into OpenAI-compatible usage
- prevent unrelated headerless Kiro requests from sharing one
  connection-wide continuation
- cap/evict continuation sessions so long-running processes do not grow
  the continuation map unbounded
- treat generated headerless Kiro sessions as one-shot so they do not
  evict real explicit-session continuations
- keep credit-only Kiro metering valid for internal persistence when
  token metrics are unavailable
2026-07-16 15:15:05 +07:00
rixzkiye
70e8dc4974 feat(cli-tools): add Grok Build setup
Add Grok Build to Dashboard → CLI Tools. Apply writes a [model.9router]
custom model to ~/.grok/config.toml and sets [models].default, routing
the xAI Grok TUI through 9Router. Reset removes the slot and restores
the previous default.
2026-07-16 14:47:20 +07:00
Edison42
b94685b80d feat(kiro): add GPT-5.6 model family (#2596)
Add GPT-5.6 Sol/Terra/Luna and their synthetic thinking/agentic/
thinking-agentic variants to the Kiro static catalog with the observed
272k context window and credit multipliers (2.4/1.2/0.6), register MITM
mapping slots for the new base ids, and override runtime capabilities so
the GPT-5.6 family reports the 272k window instead of the generic GPT-5
profile.
2026-07-16 14:38:08 +07:00
decolua
0248dd5348 feat(i18n): add Thai language translation (#2581) 2026-07-16 14:31:23 +07:00
hungtrinh
27b37705b3 perf(startup): skip inactive background services 2026-07-16 12:09:48 +07:00
Ella CEO
7dfb346667 fix(grok-cli): surface expiresAt so proactive token refresh fires (#2546) 2026-07-16 12:00:14 +07:00
decolua
a6a41dfb3c Merge remote-tracking branch 'upstream/master'
# Conflicts:
#	.gitignore
#	open-sse/handlers/chatCore.js
2026-07-16 11:59:46 +07:00
luoyide
2629218b04 fix(models): populate capabilities for live-catalog LLM models 2026-07-16 11:50:15 +07:00
joachimBrindeau
c9926897ba feat(rtk): add X-9Router-Token-Saver header to bypass token savers per request 2026-07-16 11:27:42 +07:00
liamgnc
88a8c72d2d fix(models): list compatible provider models in /v1/models
Replace the overly-broad UPSTREAM_CONNECTION_RE regex (which matched all
provider IDs with UUID suffixes) with an x-9r-internal-models-fetch
header to detect cross-instance recursive /models fetches.

fetchCompatibleModelIds now sends the header when fetching upstream
/models; the GET handler detects it and skips dynamic fetching, breaking
the recursion loop while letting compatible providers (MLX, Ollama, vLLM)
list their models. Fixes #2626.
2026-07-16 11:18:09 +07:00
luoyide
ba508f2506 fix(thinking): send explicit thinking:{type:adaptive} alongside output_config.effort 2026-07-16 11:17:08 +07:00
decolua
a077ee85bd gitignore 2026-07-16 11:16:57 +07:00
qianze
e567ba800f fix(translator): strip client_metadata when converting openai-responses to openai
client_metadata is an OpenAI Responses API-specific field. When translating
openai-responses requests to openai (Chat Completions), it leaked through
to providers like NVIDIA, which rejected it with a 400 "Unsupported
parameter". Strip it in the Responses-specific cleanup block alongside
input, instructions, store, and reasoning.
2026-07-15 17:38:53 +07:00
luoyide
542a088c04 feat(github): route Claude models through Copilot's native /v1/messages
GitHub Copilot's /chat/completions and /responses endpoints never surface
prompt-cache token counts for Claude models. Route Claude models (detected
by name pattern) to Copilot's Anthropic-native /v1/messages shim via a new
executeWithMessagesEndpoint(), translating OpenAI-shape requests to Claude
natively so cache_control gets injected and cached_tokens surface.

Also fixes translateRequest()'s internal _toolNameMap being sent upstream,
which made Anthropic's strict schema reject tool-call requests with a 400 —
now stripped and threaded through response state. Removes the now-dead
response_format Claude JSON-mode workaround.
2026-07-15 17:34:08 +07:00
Moradii.Mohammadreza
9173c29b66 feat(translator): drop temperature for all Claude models
Broaden strip rule from /claude-opus-4/i to /claude/i so temperature is
removed for every Claude model, not just opus-4. Fixes Anthropic 400 on
OpenAI-compatible routes. #1748
2026-07-15 17:08:41 +07:00
decolua
eceac9d7ae gitignore 2026-07-15 16:40:41 +07:00
ab9a3c1d43 feat(xai): track SuperGrok weekly limit + API usage quota
Fetch OAuth quota from cli-chat-proxy billing, GetGrokCreditsConfig weekly
window, and settings plan label so the dashboard matches grok.com usage.
2026-07-13 23:44:36 +07:00
minnyww
837cfec5a9 feat(i18n): complete Thai translation (1389 keys) + README.th.md 2026-07-13 17:08:49 +07:00
b1d368d960 feat: xAI image generate/edit, API key import, and per-provider timeouts
- Add dedicated xAI image adapter with generate + edit (multi-image) via
  /v1/images/generations and /v1/images/edits, plus aspect_ratio/resolution UI
- Support importing existing API keys and exposing connection api-key routes
- Add global/per-provider connect timeout overrides from settings
- Keep unrelated provider UX improvements on this branch; no Grok quota tracking
2026-07-13 16:52:22 +07:00
minnyww
f89ba32d79 feat(i18n): add Thai language translation 2026-07-13 16:50:09 +07:00
decolua
9845a1702f # v0.5.30 (2026-07-10)
## Features
- **Perplexity**: add Agent API provider (#2492)
- **Grok CLI**: add Grok CLI / Grok Build provider with OAuth device-code flow (#2502)
- **Featherless**: add OpenAI-compatible provider presets
- **SearXNG**: configure endpoint via SEARXNG_URL env (#2499)
- **Providers**: add max thinking level for gpt-5.6-sol (#2500)
- **Headroom**: add extras detection and install UI (#2403)
- **Headroom**: activate/uninstall extras + fix interpreter detection
- **PXPipe**: PXPIPE token saver — multimodal prompt compression (#2465)
- **Proxy-Pools**: auto-rotate strategy for no-auth providers (#2409)

## Fixes
- **Cloudflare-AI**: support accountId in bulk key import (#2449)
- **DB**: backup on schema change, MCP child cleanup, codex models, usage providers OOM
- **Codex**: avoid bare-email OAuth dedup (#2477)
- **CLI**: allow staged app bundle builds (#2479)
- **Headroom**: compress Kiro conversation state (#2488)
- **Gemini-CLI**: raise output floor for thinking and add validated toolConfig (#2486)
- **GitHub**: label Copilot profiles by account identity (#2498)
- **OpenAI-to-Claude**: unwrap bare {function:{…}} tools without parent type (#2473)
- **Translator**: clamp thinking effort max->xhigh for OpenAI format (#2466)
- **RTK/find**: detect and group Windows backslash-style find output (#2448)
- **Codex**: handle fast tier and capacity SSE (#2452)
- **Volcengine-ark**: clamp Kimi max_tokens to 32768 endpoint cap
- **Antigravity**: align provider fingerprint with IDE Desktop 2.1.1 (#2389)
- **Pricing**: update Claude/Codex model rates and add new models

## Improvements
- **i18n(zh-CN)**: complete Chinese translations for all UI strings (#2436)
- **API**: caching for tunnel and version status endpoints
- **Perf**: faster dev startup and lighter bundle
2026-07-10 18:12:07 +07:00
decolua
a625ea9fd8 refactor(log): unify request lifecycle logging with session-colored tags
Collapse scattered per-request console lines (request/routing/auth/pending/
usage/stream-usage/stream) into 3 correlated lines: request, transform,
done. Add stable per-session color tag so concurrent request lines are
easy to follow, surface thinking intent, always-on full error logging
for debug, re-enable warn level, and uppercase keyword labels. Also fix
usage overview cards wrapping (5 cards -> grid-cols-5).

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-10 18:01:20 +07:00
decolua
b61c50cbb7 # v0.5.29 (2026-07-10)
## Features
- **Perplexity**: add Agent API provider (#2492)
- **Grok CLI**: add Grok CLI / Grok Build provider with OAuth device-code flow (#2502)
- **Featherless**: add OpenAI-compatible provider presets
- **SearXNG**: configure endpoint via SEARXNG_URL env (#2499)
- **Providers**: add max thinking level for gpt-5.6-sol (#2500)
- **Headroom**: add extras detection and install UI (#2403)
- **Headroom**: activate/uninstall extras + fix interpreter detection
- **PXPipe**: PXPIPE token saver — multimodal prompt compression (#2465)
- **Proxy-Pools**: auto-rotate strategy for no-auth providers (#2409)

## Fixes
- **Cloudflare-AI**: support accountId in bulk key import (#2449)
- **DB**: backup on schema change, MCP child cleanup, codex models, usage providers OOM
- **Codex**: avoid bare-email OAuth dedup (#2477)
- **CLI**: allow staged app bundle builds (#2479)
- **Headroom**: compress Kiro conversation state (#2488)
- **Gemini-CLI**: raise output floor for thinking and add validated toolConfig (#2486)
- **GitHub**: label Copilot profiles by account identity (#2498)
- **OpenAI-to-Claude**: unwrap bare {function:{…}} tools without parent type (#2473)
- **Translator**: clamp thinking effort max->xhigh for OpenAI format (#2466)
- **RTK/find**: detect and group Windows backslash-style find output (#2448)
- **Codex**: handle fast tier and capacity SSE (#2452)
- **Volcengine-ark**: clamp Kimi max_tokens to 32768 endpoint cap
- **Antigravity**: align provider fingerprint with IDE Desktop 2.1.1 (#2389)
- **Pricing**: update Claude/Codex model rates and add new models

## Improvements
- **i18n(zh-CN)**: complete Chinese translations for all UI strings (#2436)
- **API**: caching for tunnel and version status endpoints
- **Perf**: faster dev startup and lighter bundle
2026-07-10 17:51:48 +07:00
decolua
baafc74c1f Bump version 2026-07-10 17:48:53 +07:00
decolua
bb314118f2 Bump version 2026-07-10 17:38:40 +07:00
decolua
2d515c8abc # v0.5.28 (2026-07-10)
## Features
- **Perplexity**: add Agent API provider (#2492)
- **Grok CLI**: add Grok CLI / Grok Build provider with OAuth device-code flow (#2502)
- **Featherless**: add OpenAI-compatible provider presets
- **SearXNG**: configure endpoint via SEARXNG_URL env (#2499)
- **Providers**: add max thinking level for gpt-5.6-sol (#2500)
- **Headroom**: add extras detection and install UI (#2403)
- **Headroom**: activate/uninstall extras + fix interpreter detection
- **PXPipe**: PXPIPE token saver — multimodal prompt compression (#2465)
- **Proxy-Pools**: auto-rotate strategy for no-auth providers (#2409)

## Fixes
- **Cloudflare-AI**: support accountId in bulk key import (#2449)
- **DB**: backup on schema change, MCP child cleanup, codex models, usage providers OOM
- **Codex**: avoid bare-email OAuth dedup (#2477)
- **CLI**: allow staged app bundle builds (#2479)
- **Headroom**: compress Kiro conversation state (#2488)
- **Gemini-CLI**: raise output floor for thinking and add validated toolConfig (#2486)
- **GitHub**: label Copilot profiles by account identity (#2498)
- **OpenAI-to-Claude**: unwrap bare {function:{…}} tools without parent type (#2473)
- **Translator**: clamp thinking effort max->xhigh for OpenAI format (#2466)
- **RTK/find**: detect and group Windows backslash-style find output (#2448)
- **Codex**: handle fast tier and capacity SSE (#2452)
- **Volcengine-ark**: clamp Kimi max_tokens to 32768 endpoint cap
- **Antigravity**: align provider fingerprint with IDE Desktop 2.1.1 (#2389)
- **Pricing**: update Claude/Codex model rates and add new models

## Improvements
- **i18n(zh-CN)**: complete Chinese translations for all UI strings (#2436)
- **API**: caching for tunnel and version status endpoints
- **Perf**: faster dev startup and lighter bundle
2026-07-10 17:36:21 +07:00
decolua
74d5fedf79 feat(headroom): activate/uninstall extras + fix interpreter detection
- find interpreter next to headroom binary so extras/version read correctly
- add on/off toggle to activate [code]/[ml] via proxy restart
- add uninstall action + live install log progress + ~1GB confirm modal
2026-07-10 17:32:45 +07:00
Elio Bonfim Júnior
dcf1927f22 feat(pxpipe): PXPIPE token saver — multimodal prompt compression (#2465)
Add pxpipe as an experimental fifth Token Saver: Claude-format request
bodies above a configurable size threshold are rendered as dense PNGs
via the pxpipe-proxy library API (transformAnthropicMessages) before
dispatch, cutting estimated input tokens by ~35-60% on token-dense
contexts. Integration follows the Headroom pattern: applied to the final
body in chatCore just before dispatch, fail-open on any error/timeout.

Managed npm install into DATA_DIR/pxpipe, dynamic loader with per-version
cache-bust, JSONL event log with rotation, /api/pxpipe/* endpoints, Token
Saver card (marked experimental) + /dashboard/pxpipe page, and per-request
Activated/Skipped annotation in Request Details. Disabled by default.
2026-07-10 16:10:42 +07:00
Fadjrir Herlambang
e1f3399b73 feat(proxy-pools): auto-rotate strategy for no-auth providers (#2409)
Add round-robin/random proxy pool rotation for no-auth free providers
(e.g. OpenCode Free) to distribute load across all active pools and
avoid per-IP rate limits. Rotation strategy is selectable per provider
in NoAuthProxyCard and persisted to settings.providerStrategies.
2026-07-10 16:05:07 +07:00
KunN-21
f1f9d27061 feat(headroom): add extras detection and install UI (#2403)
- add Headroom extras status + install endpoints
- show Headroom version + code/ml extras in Token Saver UI
- fix Windows interpreter selection to read from env with headroom-ai
2026-07-10 16:05:04 +07:00
decolua
90df008f0c chore(release): v0.5.25
Update CHANGELOG, bump version, trim usage overview cards.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-10 16:04:37 +07:00
decolua
d2599ebf17 fix(pricing): update Claude/Codex model rates and add new models
Add claude-fable-5, gpt-5.6 family; correct gpt-5/5.1/5.2/5.3-codex rates to official pricing.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-10 16:02:18 +07:00
Rizkal
5cdcf67484 fix(cloudflare-ai): support accountId in bulk key import (#2449)
Bulk import now parses name|apiKey|accountId lines for Cloudflare AI and
forwards accountId via providerSpecificData, with a provider-specific
placeholder and format hint.
2026-07-10 13:10:29 +07:00
decolua
b25e10160d fix: DB backup on schema change, MCP child cleanup, codex models, usage providers OOM
- Backup DB only on real SCHEMA_VERSION change, not every app version bump
- Kill idle MCP stdio bridge children to prevent orphan process leaks
- Add getDistinctProviders to avoid loading every row JSON blob (OOM fix)
- Update codex model list (gpt-5.6 sol/terra/luna, drop 5.3 codex variants)
- Reorder Claude default models (fable first)

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-10 13:08:58 +07:00
decolua
0270f6ea70 perf: faster dev startup and lighter bundle
- Switch dev default to Turbopack (5-14x faster compile); keep webpack as dev:webpack
- Tailwind v4 source() base so JIT scans identically under both bundlers
- Lazy-load @xyflow/react via next/dynamic to keep it out of the shared bundle
- optimizePackageImports for heavy barrel imports (xyflow, dnd-kit, material-symbols, marked)
- Replace blind setTimeout waits with TCP health-check (waitServerReady)
- Run checkForUpdate in parallel instead of blocking server spawn
- Background MITM/tunnel/cloudflared kills off the critical path

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-10 13:08:02 +07:00
newnol
ce6bdf7fc2 feat(perplexity): add Agent API provider (#2492)
Add perplexity-agent provider using OpenAI-compatible Responses API,
routing third-party models (GPT, Claude, Gemini, Grok, GLM, Kimi, Sonar)
through one endpoint. Expose /v1/models discovery, add chat-search wrapper
via web_search tool. Existing Sonar provider unchanged.
2026-07-10 11:57:15 +07:00
Fadjrir Herlambang
a11937cdd6 feat(grok-cli): add Grok CLI / Grok Build provider with OAuth device-code flow (#2502)
New OAuth provider routing through cli-chat-proxy.grok.com (OpenAI Responses
API), distinct from xai (api.x.ai) and grok-web (cookie SSO):

- Registry + GrokCliExecutor: Chat Completions -> Responses transform, CLI
  fingerprint headers, virtual effort models grok-4.5-{low,medium,high}
- OAuth device-code flow (auth.x.ai) with no-PKCE, shared xAI token refresh
- store=false multi-turn continuity via reasoning encrypted_content
- Quota tracker: on-demand window + prepaid balance on dashboard
- Connection test: 402 spending-limit = soft success (auth OK, out of credits)
- Alias/oauth/provider baselines + unit tests
2026-07-10 11:47:08 +07:00
Hermes Hunter
c73c419d09 fix(codex): avoid bare-email OAuth dedup (#2477)
Only update an existing Codex OAuth row when both rows share the same
chatgptAccountId, so a second Codex login no longer overwrites the first
account's rotated token pair. Also fall back to
workspaceId || chatgptAccountId || accountId for the chatgpt-account-id header.
2026-07-10 11:41:43 +07:00
ryanngit
a3b267a5cb fix(cli): allow staged app bundle builds (#2479)
Write the standalone app bundle and MITM bundle to NINEROUTER_CLI_APP_DIR
when set, so staged deploys can build to a separate destination before
swapping. Default output stays at cli/app when the env var is unset.
2026-07-10 11:40:03 +07:00
Edison42
65c65a0f56 fix(headroom): compress Kiro conversation state (#2488)
Project conversationState history/currentMessage into OpenAI-style
messages for /v1/compress, then write compressed text back into the
original Kiro fields while preserving provider payload shape. Fail open
when the proxy returns malformed or reordered messages.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-10 11:33:21 +07:00
DOMANHDUC
7610f28f42 fix(gemini-cli): raise output floor for thinking and add validated toolConfig (#2486)
Gemini CLI requests with small max_tokens spend the whole output budget on
thoughts after reasoning_effort maps to thinkingConfig, returning blank
content or finish=length. Raise maxOutputTokens floors per thinking level/
budget (clamped to caps.maxOutput). Also emit toolConfig
functionCallingConfig.mode=VALIDATED for Gemini CLI tool requests to avoid
MALFORMED_FUNCTION_CALL.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-10 11:32:11 +07:00
newnol
0d4d4bc261 feat(featherless): add OpenAI-compatible provider presets
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-10 11:29:48 +07:00
ryanngit
3a7a878f91 fix(github): label Copilot profiles by account identity (#2498)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-10 11:29:12 +07:00
zie
e79f9eddb4 feat(searxng): configure endpoint via SEARXNG_URL env (#2499)
Add SEARXNG_URL runtime override for the built-in SearXNG web-search
provider, defaulting to http://localhost:8888/search. Enables Docker
and remote SearXNG deployments without changing existing behavior.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-10 11:29:06 +07:00
Rafli Ahmad Zulfikar
b9e2611045 feat(providers): add max thinking level for gpt-5.6-sol (#2500)
Expose max in the Codex thinking dropdown for gpt-5.6-sol only (maps to
xhigh on wire; live probe rejected ultra). Include custom/kilo models
when computing provider thinking options so manually added gpt-5.6-sol
contributes its max level.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-10 11:17:32 +07:00
Samir Abis
ddd5509e97 fix(openai-to-claude): unwrap bare {function:{…}} tools without parent type (#2473)
Translator only unwrapped tool.function when both tool.type==="function"
and tool.function were truthy. Loose/legacy OpenAI clients emit the bare
{ function: { name, parameters } } shape (no parent type), which fell
through and forwarded name: undefined upstream, rejected by strict
Anthropic-compatible gateways (MiniMax M3) as (2013) invalid tool type.

Unwrap tool.function whenever present. Built-in tools stay pass-through.
Adds regression coverage for the 4 tool shapes. See #2435.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-10 11:11:37 +07:00
thienpv
288940960a fix(translator): clamp thinking effort max->xhigh for OpenAI format (#2466)
Claude Code sends reasoning_effort "max" (its top level); OpenAI enum caps
at "xhigh" and rejects "max" with HTTP 400 "max effort not support".
applyFormat case "openai" now clamps "max"->"xhigh" before assigning
body.reasoning_effort; other levels pass through unchanged.

Add regression test covering client output_config.effort, direct
reasoning_effort, passthrough of xhigh/high, and budget_tokens capping.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-09 15:13:24 +07:00
Diwak4r
d75471bbbc fix(rtk/find): detect and group Windows backslash-style find output (#2448)
isPathLike rejected any line with a colon, so Windows absolute paths
(C:\Users\me\a.js) were never recognized and find dumps went uncompacted.
find.js also split only on "/", mis-grouping backslash paths.

- autodetect: treat drive-letter prefix (X:\ or X:/) as path-like before
  the general colon rejection.
- find.js: split on the last "/" or "\" separator and normalize emitted
  directory labels to forward slashes.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-09 15:12:52 +07:00
ryanngit
0c55d49ab6 fix(codex): handle fast tier and capacity SSE (#2452)
- map service_tier=fast to upstream priority; drop unsupported tiers
- normalize reasoning effort max to xhigh (codex-only)
- convert 200-SSE model-capacity errors into 503 so account fallback rotates
- keep normal SSE output intact after peeking

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-09 15:12:46 +07:00
whale9820
cfbdf06047 fix(volcengine-ark): clamp Kimi max_tokens to 32768 endpoint cap
VolcEngine Ark caps the Kimi family at max_tokens <= 32768, but the
model's advertised ceiling is far higher (Kimi-K2.7-Code resolves to
maxOutput 262144), so clampToModelMaxOutput alone leaves it uncapped and
the request 400s. Add a Kimi-scoped rule with an explicit maxOutputCap of
32768, combined with the model ceiling via min(). Covers max_tokens,
max_completion_tokens, max_output_tokens.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-09 15:10:05 +07:00
decolua
a4c5fa4e14 refactor(api): implement caching for tunnel and version status endpoints 2026-07-09 15:08:30 +07:00
qianze
20b442b708 i18n(zh-CN): complete Chinese translations for all UI strings (#2436)
Add 551 new translations covering previously untranslated areas
(landing, CLI tools, MITM, skills, combos, token-saver, OIDC,
relay deploy, provider details, proxy pools, quota tracker,
OAuth modals). Total: 838 -> 1389 entries.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-09 15:07:12 +07:00
nguyenha935
71cd5b2f23 fix(antigravity): align provider fingerprint with IDE Desktop 2.1.1 (#2389)
Match captured official Antigravity IDE traffic: cloudcode-pa host,
antigravity/ide/2.1.1 User-Agent, IDE-shaped agent requestId, and drop
router-only stream/usage headers plus the legacy double system prompt.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-08 10:17:47 +07:00
decolua
b10b807063 # v0.5.20 (2026-07-07)
## Features
- **Thinking**: per-model thinking level picker on provider page — appends `(level)` suffix to copied model names for forced reasoning effort across all formats (openai, claude, gemini, deepseek, kimi, qwen, zai, minimax, hunyuan, step)
- **RTK**: add JS-native git-log filter (#2423)
- **Caveman**: add targeted upstream-aligned style rules (#2424)
- **i18n**: add Farsi (fa) language support (#2385)

## Fixes
- **Thinking**: strip `(level)` suffix from upstream `body.model` so providers no longer reject requests
- **Translator**: preserve developer instructions in openai-responses conversion (#2434)
- **count_tokens**: count structured Anthropic blocks (#2419)
- **Volcengine-ark**: clamp GLM-5 max_tokens to model output ceiling (#2428)
- **Kimi**: normalize reasoning_effort to backend enum (#2427)
- **Claude**: reconcile max_tokens vs thinking budget and lift per-model ceiling (#2381)
- **Kiro**: deliver system prompt natively, add Opus 4.5/4.7/4.8, tolerate dash version ids (#2366)
- **Headroom**: proxy dashboard through app (#2372)
- **MITM**: recover from stale lock file on server start
2026-07-07 16:29:11 +07:00
baibiao
081c6f2aff fix(count_tokens): count structured Anthropic blocks (#2419)
Estimate tokens for tool_use, tool_result, thinking, system, and tools
blocks instead of text only, so count_tokens no longer returns 0 for
structured content and breaks Claude Code auto-compaction (#2337).

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-07 12:06:09 +07:00
KunN-21
19281b5524 feat(rtk): add JS-native git-log filter (#2423)
Compress git log output via dedicated RTK filter: keep commit headers,
Author/Date, subject; drop body padding, decoration, embedded diff lines.
Wire into autodetect (git-log prioritized before git-diff) and registry.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-07 12:02:56 +07:00
whale
bbae990b92 fix(volcengine-ark): clamp GLM-5 max_tokens to model output ceiling (#2428)
Ark rejects max_tokens above 128000 for GLM-5.2. Add a config-driven STRIP_RULES entry that clamps max_tokens, max_completion_tokens and max_output_tokens down to the model maxOutput before the upstream call.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-07 11:57:04 +07:00
whale
8c068a1f5c fix(kimi): normalize reasoning_effort to backend enum (#2427)
Map auto→high, minimal→low, xhigh→max and whitelist low/medium/high/max
so Kimi/kimchi SGLang backends no longer receive invalid effort values.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-07 11:56:39 +07:00
KunN-21
97a6708651 feat(caveman): add targeted upstream-aligned style rules (#2424)
Add four shared Caveman prompt fragments (no invented abbreviations,
preserve user language, no self-reference, no decoration) across all six
levels, and remove ULTRA contradictions around abbreviations/arrow
shorthand. Adds regression tests for the prompt rules.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-07 11:55:45 +07:00
deranalabs
a3cd7c82bc fix(translator): preserve developer instructions in openai-responses conversion (#2434)
Map role="developer" messages to top-level instructions alongside
role="system" in openaiToOpenAIResponsesRequest. Previously developer
messages matched no branch and were silently dropped from the Responses
request, losing GPT-5/Codex system-level prompts.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-07 11:54:07 +07:00
decolua
bf7da67859 docs(readme): swap in Vietnamese tutorial video; chore(pricing): minor update
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-07 11:44:30 +07:00
decolua
da0149de97 fix(mitm): recover from stale lock file on server start
Detect dead PID in lock file and reclaim it instead of failing, and drop unused fs dependency.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-07 11:44:20 +07:00
MuhammadHamidRaza
1885ad7f64 docs(readme): add English and Urdu/Hindi video tutorials (#2305)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-05 17:47:17 +07:00
Sutarto Jordan Chrisfivo
481e7e467b fix(headroom): proxy dashboard through app (#2372)
Add a 9Router-side proxy so the Headroom dashboard and its data
endpoints (/stats, /health, /stats-history, /transformations/feed)
stay same-origin when opened remotely through the 9Router app, and
add an "Open Headroom Dashboard" link in the Token Saver modal.

Gate /api/headroom/proxy as LOCAL_ONLY (loopback + CLI token) to
match start/stop, and strip cookie/authorization when the Headroom
target is non-loopback to avoid leaking viewer credentials.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-05 17:45:04 +07:00
Mohammed Faheem
008de32c06 docs: add CLAUDE.md guidance for Claude Code (#2354)
Top-level guide for Claude Code / AI coding agents working in this repo,
complementing docs/ARCHITECTURE.md and open-sse/AGENTS.md.

Docs-only: PR's incidental code changes were dropped as they reverted #2366.

Co-Authored-By: Mohammed Faheem <mohammed.faheem@adbsafegate.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-05 17:44:04 +07:00
Arash Kadkhodaei
b6454d84da feat(i18n): add Farsi (fa) language support (#2385)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-05 17:38:17 +07:00
thienpv
46e6c01a01 fix(claude): reconcile max_tokens vs thinking budget and lift per-model ceiling (#2381)
On the translated OpenAI->Claude path, adjustMaxTokens capped max_tokens
before applyThinking set thinking.budget_tokens, so max-effort budget
(128000) could exceed a 64k-clamped max_tokens -> Anthropic 400.
prepareClaudeRequest now reconciles after the budget is known: prefer
raising max_tokens, only shrink budget when it meets/exceeds the ceiling.

Also lift the global 64000 cap: the ceiling is now the model's real
maxOutput, so high-output models (fable/mythos, opus-4.8/sonnet-4.6) get
their full budget. adjustMaxTokens gains an optional ceiling arg (default
unchanged, callers untouched); openai-to-claude passes the model maxOutput.

Native Claude Code passthrough is unaffected.

Co-Authored-By: Claude <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-05 17:38:16 +07:00
VitzS7
5041494e1c fix(kiro): deliver system prompt natively, add Opus 4.5/4.7/4.8, tolerate dash version ids (#2366)
- Pass system prompt via native systemInstruction field (+ <instructions> fallback)
  so Claude models stop treating it as info-only <system-reminder>
- Add Opus 4.5/4.7/4.8 (base/thinking/agentic/thinking+agentic) to Kiro registry
- normalizeModelId(): dash->dot version separator, scoped to Kiro provider only
- Replace <system-reminder> with <instructions> in claude-to-openai/openai-to-kiro

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-05 17:34:04 +07:00
nguyenha935
4dadab9d5f feat: add provider quota visibility settings 2026-07-04 23:26:49 +07:00
decolua
7f436e2792 # v0.5.18 (2026-07-03)
## Features
- **Usage**: track cached tokens + correct input/output/cache cost (#2209) — hodtien
- **Codex**: show reset credit expiry details (#2290) — Rafli Ahmad Zulfikar
- **NVIDIA**: add new models and capabilities — decolua
- **ClinePass**: add provider support — sternelee

## Fixes
- **Usage**: dedupe streaming request-details log entries — Qin Li
- **Claude**: drop foreign thinking signatures in passthrough — decolua
- Prevent non-SSE stream pipe crash and cross-IdP account overwrites (#2244) — KunN-21
- **Kiro**: route IdC auth to regional CodeWhisperer surface (#2297) — Volodymyr Saakian
- **Kiro**: add Claude Sonnet 5 model support (#2264) — Edison42
- **Xiaomi-tokenplan**: region selector, key validation, multi-connection (#2251) — MiQieR
- **Translator**: strict Anthropic content block compliance (#2225) — Sahrul Ramadhan Hardiansyah
- **Kimchi**: strip reasoning_content echo to bound multi-turn input tokens — KunN-21
- **Kimchi**: bump User-Agent to kimchi/0.1.40 (#2256) — Ansh7473
- **Codebuddy-cn**: strip empty tool_calls arrays to preserve reasoning — zmf
- **Antigravity**: preserve Claude tool delta index (#2223) — Sutarto Jordan Chrisfivo
- **MITM**: generate root CA on server startup (#2228) — Sutarto Jordan Chrisfivo
2026-07-03 15:37:17 +07:00
hodtien
54e3245ace feat(usage): track cached tokens + correct input/output/cache cost (#2209)
Normalize every provider to one cache-inclusive convention via
canonicalizeUsage() before persist, and price cached + cache_creation as
subsets of prompt_tokens in calculateCostFromTokens() to stop
double-counting. usageRepo now delegates cost math to a single source.
Surface Cached tokens/cost across dashboard (overview, tokens, cost,
details). Merge Claude message_start cache with message_delta output so
cache counts survive. Compatible LLM nodes now allow multiple API-key
connections (key pool).

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-03 15:18:27 +07:00
Qin Li
960f8a0379 fix(usage): dedupe streaming request-details log entries
handleStreamingResponse and buildOnStreamComplete each generated their
own streamDetailId for what should be one logical record — the
placeholder row (0 tokens) and the final row (real usage) never shared
an id, so the DB's ON CONFLICT(id) upsert never merged them, leaving a
permanent 0-token stub for every streaming request.

Share the id from buildOnStreamComplete with handleStreamingResponse
so both writes hit the same row.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-03 15:14:49 +07:00
Rafli Ahmad Zulfikar
5cc4f222f8 feat(codex): show reset credit expiry details (#2290)
Add read-only GET to inspect per-credit reset inventory (status, granted,
expiry, remaining) with a Quota Tracker modal. DRY the route via shared
connection/refresh helpers; keep existing consume POST unchanged.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-03 15:07:45 +07:00
decolua
cd557a2552 fix(claude): drop foreign thinking signatures in passthrough
Combo mixes models, so non-Claude thinking signatures leak into
conversation history. Native passthrough forwarded them verbatim and
Anthropic rejected the request. Validate signatures and drop invalid
thinking blocks, re-inserting a placeholder when tool_use requires one.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-03 15:06:19 +07:00
decolua
ced51ed62f feat(nvidia): add new models and capabilities for NVIDIA provider
- Updated capabilities for NVIDIA models to enforce OpenAI-compatible reasoning formats.
- Added new models: MiniMax M3, GLM 5.2, DeepSeek V4 Pro, DeepSeek V4 Flash, Kimi K2.6, and Nemotron 3 Ultra to the NVIDIA registry.

This enhances the provider's functionality and aligns with OpenAI standards.
2026-07-03 12:15:58 +07:00
KunN-21
cb0135b695 fix: prevent non-SSE stream pipe crash and cross-IdP account overwrites (#2244)
- streamingHandler: when upstream returns non-SSE/JSON (e.g. Cloudflare
  5xx HTML), read body, sanitize <title>, notify streamController and
  return a clean JSON error instead of crashing the pipe.
- connectionsRepo: dedup OAuth connections on (email + username) so
  cross-IdP accounts sharing an email no longer overwrite each other;
  workspace providers keep workspace-id matching.
- kimchi: bump User-Agent to 0.1.50, add svg asset + browser-login
  service, and 21 unit tests.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-03 11:11:07 +07:00
Volodymyr Saakian
abc0add031 fix(kiro): route IdC auth to regional CodeWhisperer surface (#2297)
IAM Identity Center (authMethod=idc) tokens failed every request with 403
"bearer token invalid". Treat idc like api_key/external_idp:

- executors/kiro.js: route idc to *.amazonaws.com CodeWhisperer surface,
  region-aware from credentials.region instead of hardcoded us-east-1.
- openai-to-kiro.js / claude-to-kiro.js: send resolved profileArn or empty
  for idc/external_idp, never the shared builder-id placeholder ARN.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-03 11:06:40 +07:00
MiQieR
9102c4c6d8 fix(xiaomi-tokenplan): region selector, key validation, multi-connection (#2251)
- Add top-level regions array so Add/Edit modals render region <Select>
- EditConnectionModal: load/persist region generically for region-aware providers
- validate: accept 403 for xiaomi-tokenplan valid keys, add 8s fetch timeout
- Remove single-connection guard for compatible/embedding nodes

Co-authored-by: MiQieR <122154116+MiQieR@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-03 11:03:18 +07:00
Sahrul Ramadhan Hardiansyah
ce6120ce7b fix(translator): strict Anthropic content block compliance (#2225)
Filter empty text blocks from thoughtSignature-only parts, preserve
tool_calls when functionResponse and functionCall coexist in the same
content, and skip empty regular text parts before they reach Claude.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-03 10:58:39 +07:00
KunN-21
7afaecd617 fix(kimchi): strip reasoning_content echo to bound multi-turn input tokens
Clients echo full message history each turn including reasoning_content,
which the Kimchi OpenAI gateway counts as input tokens. Multi-turn convos
balloon to 100k+ tokens and the model returns empty content.

KimchiExecutor.transformRequest now strips reasoning_content from assistant
messages when it exceeds an 8-char threshold, preserving the 1-char
placeholder injectReasoningContent sets and keeping content intact.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-03 10:58:31 +07:00
Edison42
a5363b83b5 fix(kiro): add Claude Sonnet 5 model support (#2264)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-03 10:54:15 +07:00
sternelee
b08751c4ea feat(clinepass): add ClinePass provider support
Register clinepass provider (OAuth + API-key) using Cline's
OpenAI-compatible API with 10 curated models, live /v1/models
resolver, refreshCline-based token refresh with workos: prefix,
and dashboard OAuth login handler.

Reference: https://github.com/jellydn/pi-clinepass-provider
Closes #2261

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-03 10:53:41 +07:00
Ansh7473
76752a4396 fix(kimchi): bump User-Agent to kimchi/0.1.40 (#2256)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-03 10:49:53 +07:00
zmf
602ee4054b fix(codebuddy-cn): strip empty tool_calls arrays to preserve reasoning
CodeBuddy CN includes "tool_calls": [] in every SSE streaming delta.
@ai-sdk/openai-compatible checks delta.tool_calls != null — an empty
array passes ([] != null is true in JS), triggering premature
reasoning-end on every reasoning chunk (0/1ms durations in OpenCode).

Strip empty tool_calls arrays in passthrough before hasValuableContent.
Zero side-effect: real tool_calls always have at least one element.

Closes #2176

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-01 09:39:31 +07:00
Sutarto Jordan Chrisfivo
8f81f17b99 fix(antigravity): preserve Claude tool delta index (#2223)
Gemini response translation wrote OpenAI-shaped bookkeeping into the
shared state.toolCalls map, which the downstream openai-to-claude
translator uses for Claude block metadata. That pre-population skipped
blockIndex creation, so Anthropic input_json_delta events lost index.

Track Gemini function calls via state.geminiToolCallCount instead,
leaving state.toolCalls clean for the Claude translator.

Closes #2218

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-01 09:36:43 +07:00
Sutarto Jordan Chrisfivo
182c849979 fix(mitm): generate root ca on server startup (#2228)
Direct MITM server startup read rootCA.key/.crt immediately and exited
when either was missing, bypassing the manager.js CA setup path.

- generate Root CA from server.js when key/cert is missing
- make generateRootCA()/generateCert() synchronous to avoid a startup
  race before readFileSync
- add unit test covering synchronous Root CA creation

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-01 09:30:52 +07:00
decolua
0b3c794075 # v0.5.15 (2026-06-29)
## Features
- Add Kimchi OAuth provider — Nant361
- Refine Qwen vision/video + thinking model patterns — decolua
- Opt-in Codex auto-ping quota keep-alive — Emirhan

## Fixes
- **Responses**: handle response.done terminal events (#2142) — rifuki
- **Headroom**: skip unsafe responses tool history (#2132) — Sutarto Jordan Chrisfivo
- **Translator**: map mid-conversation system message to user (claude→openai) — decolua
- **Gemini**: normalize contents to prevent 400 invalid_argument (#2192) — warelik
- **Gemini**: backfill thoughtSignature + suppress stream done sentinel — WARELIK
- **Alicode**: preserve cache_control for DashScope providers (#2069) — Rex
- **Antigravity**: strip deprecated/readOnly/writeOnly from tool schemas — iletai, Yudhistira-Official
- **CodeBuddy CN**: show bonus packs as one-time, not monthly-replenishing — whale9820
- **Kiro**: strip leaked <thinking> tags from content stream (#2158) — hamsa0x7
- **Tray**: make Windows context menu DPI-aware — Emirhan
- **Kilocode**: expose full gateway catalog in combo model picker — jellylarper
- **OpenCode**: fix Go GLM — decolua
2026-06-29 16:23:33 +07:00
rifuki
a9785a5f70 fix(responses): handle response.done terminal events (#2142)
Treat response.done as a terminal OpenAI Responses stream event so
passthrough streams ending with response.done are not flagged incomplete
and no synthetic response.failed is emitted. Restore the data: [DONE]
sentinel for same-format Responses passthrough streams.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-29 16:05:29 +07:00
Sutarto Jordan Chrisfivo
373850ee36 fix(headroom): skip unsafe responses tool history (#2132)
Guard openai-responses compression: skip Headroom when body.input
contains non-message items (function_call, function_call_output,
reasoning) to preserve the Responses contract instead of collapsing
them into chat messages.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-29 15:59:21 +07:00
decolua
749c2e3f9c fix(translator): map mid-conversation system message to user in claude-to-openai
Claude Code chèn role:system cuối messages[], trước đây bị map thành assistant
khiến hội thoại không kết thúc bằng user → provider OpenAI-compat (LiteLLM)
dịch ngược Anthropic trả 400 "assistant message prefill". Map system -> user
và wrap <system-reminder> để giữ ngữ nghĩa instruction.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-29 15:53:26 +07:00
decolua
7fa2e7f029 feat(capabilities): refine Qwen vision/video and thinking model patterns
Add qwen omni (audio/video input), qwen3.5/3.6/3.7 (native vision/video),
and mark qwen coder & max as text-only reasoning models.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-29 15:51:56 +07:00
warelik
8d1db46beb fix(gemini): normalize contents to prevent 400 invalid_argument (#2192)
Merge adjacent same-role blocks and strip empty parts before sending to
Gemini, avoiding 400 INVALID_ARGUMENT on consecutive same-role messages.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-29 15:38:03 +07:00
Rex
9e3866658a fix(alicode): preserve cache_control for DashScope providers (#2069)
Opt-in quirk preserveCacheControl keeps cache_control on content blocks
for alicode/alicode-intl, enabling DashScope prompt caching. signature
is always stripped; all other providers unchanged.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-29 15:29:49 +07:00
Nant361
8a664d619d feat(kimchi): add Kimchi OAuth provider support
Add Kimchi as a browser-token OAuth provider routed through its
OpenAI-compatible gateway. Discover live models for /v1/models and
provider models, normalize Claude-compatible requests, and wire up
provider connection tests.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-29 15:29:17 +07:00
WARELIK
2d94fffe3b fix(gemini): backfill thoughtSignature and suppress stream done sentinel
Backfill DEFAULT_THINKING_AG_SIGNATURE onto functionCall parts missing it
(client history replay) and on Claude tool_use blocks, fixing 400
INVALID_ARGUMENT from Gemini-family APIs. Suppress the OpenAI-style
data: [DONE] sentinel for antigravity/gemini/vertex to avoid parser crashes.

Fixes #2193.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-29 15:28:03 +07:00
Yudhistira-Official
319caa2d7b fix(antigravity): strip 'deprecated' from tool schemas before Gemini
Gemini rejects the non-standard 'deprecated' keyword in nested tool
schemas with INVALID_ARGUMENT (400). Add it to UNSUPPORTED_SCHEMA_CONSTRAINTS
alongside 'optional' so it gets stripped during translation.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-29 15:28:02 +07:00
whale9820
95bfc64f06 fix(codebuddy-cn): show bonus packs as one-time, not monthly-replenishing
CodeBuddy CN bonus packs ("Bonus Pack N") are one-shot credits whose
CycleEndTime equals DeductionEndTime — they expire for good and never
replenish. The dashboard rendered their resetAt as "Reset in Xd",
implying a monthly refill.

Tag bonus packs recurring:false (refill packs recurring:true) in the
usage handler, forward the flag through parseQuotaData, and word the
quota table / progress bar as "Expires in" / "Expires at" for
one-shot packs.
2026-06-29 15:22:47 +07:00
hamsa0x7
eff81b1242 fix(kiro): strip leaked <thinking> tags from content stream (#2158)
CodeWhisperer leaks literal <thinking> blocks into assistantResponseEvent,
duplicating reasoning already routed via reasoningContentEvent. Track
inThinking state to strip these tags during SSE transform, handling split
chunks across tag boundaries.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-29 15:14:15 +07:00
Emirhan
b66b5c68ce feat(quota): add opt-in Codex auto-ping
Generalize Claude 5h auto-ping into a provider-generic scheduler and add
opt-in Codex auto-ping that warms the next 5h window via a tiny gpt-5.5
request when session.resetAt slides. Default off, per-connection toggle,
failure cooldown, blocking-quota skip, drains stream before success.

Closes #2107

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-29 15:13:17 +07:00
Emirhan
fc8722e897 fix(tray): make Windows context menu DPI-aware
Set process DPI awareness (Per-Monitor V2 with fallbacks) and enable
WinForms visual styles before creating the tray NotifyIcon, so the
Windows context menu renders sharply on DPI-scaled displays.

Closes #2161

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-29 15:12:24 +07:00
jellylarper
713c563765 fix(kilocode): expose full gateway catalog in combo model picker
Add modelsFetcher + passthroughModels so the dynamic Kilo Gateway
catalog surfaces in the combo model picker, matching openrouter.js.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-29 15:12:18 +07:00
iletai
3d20a4ccd2 fix(antigravity): strip deprecated/readOnly/writeOnly from tool schemas
Gemini/Antigravity generateContent rejects the JSON Schema annotation
keywords deprecated, readOnly, writeOnly with a 400 INVALID_ARGUMENT.
MCP tool schemas (e.g. Claude Code) commonly set deprecated:true, making
every request with such a tool fail. Add them to
UNSUPPORTED_SCHEMA_CONSTRAINTS so cleanJSONSchemaForAntigravity removes
them recursively before the request is sent.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-29 15:12:17 +07:00
decolua
526235872a Fix OpenCode Go GLM 2026-06-29 15:00:03 +07:00
79 changed files with 8929 additions and 2490 deletions

15
.gitignore vendored
View File

@@ -1,4 +1,5 @@
# See https://help.github.com/articles/ignoring-files/ for more about ignoring files.
# dependencies
/node_modules
/.pnp
@@ -8,8 +9,10 @@
!.yarn/plugins
!.yarn/releases
!.yarn/versions
# testing
/coverage
# next.js
/.next/
/.next-cli-build/
@@ -19,22 +22,28 @@ product
# production
/build
.idea/
# misc
.DS_Store
*.pem
# debug
npm-debug.log*
yarn-debug.log*
yarn-error.log*
.pnpm-debug.log*
# env files (can opt-in for committing if needed)
.env*
!.env.example
# vercel
.vercel
# typescript
*.tsbuildinfo
next-env.d.ts
.bin/*
data/
logs/*
@@ -52,18 +61,23 @@ Thanks.md
PUBLIC.en.md
PR/*
package-lock.json
#Ignore vscode AI rules
.github/instructions/codacy.instructions.md
README1.md
deploy*.sh
ecosystem.config.*
scripts/agSniffer/*
gitbooks/*
gitbook/README.md
# Refactor backup reference (do not bundle/lint)
open-sse.old/
.graphifyignore
graphify-out/*
# Local-only working dirs (notes, vendored repos, scripts, skills)
.claude/
.docs/
@@ -72,6 +86,7 @@ graphify-out/*
.codegraph/
.PR/
.next-analyze/*
# CommandCode CLI local state (auth/taste/projects)
.commandcode/

View File

@@ -37,3 +37,7 @@ Provider-agnostic SSE engine: one OpenAI-style request → any provider (LLM cha
- `registry/index.js` is an auto-generated static import list; regenerate it (don't hand-edit) after adding a `registry/{id}.js`. REGISTRY_TEMPLATE is excluded by design.
- Special binary/protobuf formats (kiro EventStream, cursor protobuf, commandcode NDJSON) don't round-trip through OpenAI — handle in their executor.
- `rtk/` + `headroom.js` mutate the request body in-place and are **fail-open**: any error returns null and leaves the body untouched — never throw out of them. RTK skips `is_error`/`status:"error"` tool results to preserve traces.
- **HTTP 200 in-stream errors**: some upstreams signal failure INSIDE a 200 stream (AI SDK v5 `{"type":"error"}` events, error text in content). HTTP-level success checks miss these → no fallback, `Status: success` in logs. Three hook points + one config escape hatch:
1. **Translator** — never map an error event to content. Emit an OpenAI-shaped `chunk.error = { message, type }` + terminal chunk (`translator/response/commandcode-to-openai.js` is the worked example). Downstream `parseSSEToOpenAIResponse` already detects `chunk?.error`.
2. **Executor early-peek** — for streaming fallback, read the first events BEFORE returning the response; an error → non-ok Response (`executors/commandcode.js` `peekForUpstreamError`).
3. **Config escape hatch (no code)** — per-provider `streamErrorPatterns` setting (UI: provider page → Stream Error Patterns). Patterns matched against the first ~8KB of the stream and the assembled non-streaming content; see `utils/streamErrorPeek.js` + `utils/streamErrorPatterns.js`.

View File

@@ -2,6 +2,7 @@ import { HTTP_STATUS, RETRY_CONFIG, DEFAULT_RETRY_CONFIG, resolveRetryEntry, FET
import { shouldRefreshCredentials } from "../services/oauthCredentialManager.js";
import { proxyAwareFetch } from "../utils/proxyFetch.js";
import { dbg } from "../utils/debugLog.js";
import { resolveProviderTimeoutMs } from "../services/providerTimeout.js";
import { ANTHROPIC_API_VERSION, OPENAI_COMPAT_BASE, ANTHROPIC_COMPAT_BASE } from "../providers/shared.js";
import { resolveOpenAICompatibleApiType } from "../services/provider.js";
@@ -133,13 +134,14 @@ export class BaseExecutor {
// Abort if upstream doesn't return response headers within connection timeout
const connectCtrl = new AbortController();
const timeoutMs = this.config?.timeoutMs || FETCH_CONNECT_TIMEOUT_MS;
const timeoutMs = await resolveProviderTimeoutMs(this.provider, this.config?.timeoutMs, FETCH_CONNECT_TIMEOUT_MS);
const connectTimer = setTimeout(() => connectCtrl.abort(new Error("fetch connect timeout")), timeoutMs);
const mergedSignal = signal ? AbortSignal.any([signal, connectCtrl.signal]) : connectCtrl.signal;
let fetchT0 = 0;
try {
const bodyStr = JSON.stringify(transformedBody);
const fetchT0 = Date.now();
fetchT0 = Date.now();
dbg("FETCH", `${this.provider.toUpperCase()} → ${url} | body=${bodyStr.length}B | connectTimeout=${timeoutMs}ms`);
const response = await proxyAwareFetch(url, {
method: "POST",
@@ -165,6 +167,11 @@ export class BaseExecutor {
clearTimeout(connectTimer);
lastError = error;
const isConnectTimeout = connectCtrl.signal.aborted && error.name === "AbortError";
// Error diagnostic — only logs on actual upstream failure. Distinguishes
// undici connect timeout (UND_ERR_CONNECT_TIMEOUT), DNS (ENOTFOUND),
// refused (ECONNREFUSED) vs our own connectCtrl abort (AbortError).
const cause = error?.cause || {};
console.log(`[FETCH-DIAG] ${this.provider} fetch error | name=${error.name} | code=${error.code ?? cause?.code ?? "none"} | msg=${String(error.message).slice(0, 120)} | connectTimeout=${timeoutMs}ms | elapsed=${Date.now() - fetchT0}ms`);
dbg("FETCH", `${this.provider.toUpperCase()} ✖ ${error.name}: ${error.message}${isConnectTimeout ? " (connect timeout)" : ""}`);
// Connect timeout is internal — convert to retryable network error, don't propagate AbortError
if (error.name === "AbortError" && !isConnectTimeout) throw error;

View File

@@ -1,6 +1,7 @@
import { randomUUID } from "crypto";
import { BaseExecutor } from "./base.js";
import { PROVIDERS } from "../config/providers.js";
import { HTTP_STATUS } from "../config/runtimeConfig.js";
import { commandCodeToOpenAIResponse } from "../translator/response/commandcode-to-openai.js";
import { SSE_DONE } from "../utils/sseConstants.js";
@@ -14,81 +15,271 @@ import { SSE_DONE } from "../utils/sseConstants.js";
* We translate each event to an OpenAI chat.completion.chunk and emit it as SSE so
* both the streaming and non-streaming (forced SSE → JSON) downstream handlers in
* 9router can consume it without further format translation.
*
* Terminal upstream failures arrive as `{"type":"error"}` events inside the HTTP
* 200 stream, so a plain `response.ok` check cannot see them. We peek the first
* events before committing the response (see peekForUpstreamError) so a stream
* that starts with an error fails fast — the normal `!response.ok` path then
* triggers account/model fallback instead of streaming fake success content.
*/
export class CommandCodeExecutor extends BaseExecutor {
constructor() {
super("commandcode", PROVIDERS.commandcode);
}
constructor() {
super("commandcode", PROVIDERS.commandcode);
}
transformRequest(model, body, stream, credentials) {
body.stream = true;
return body;
}
transformRequest(_model, body, _stream, _credentials) {
body.stream = true;
return body;
}
buildHeaders(credentials, stream = true) {
const headers = {
"Content-Type": "application/json",
...(this.config.headers || {}),
"x-session-id": randomUUID(),
};
buildHeaders(credentials, stream = true) {
const headers = {
"Content-Type": "application/json",
...(this.config.headers || {}),
"x-session-id": randomUUID(),
};
const token = credentials?.apiKey || credentials?.accessToken;
if (token) headers["Authorization"] = `Bearer ${token}`;
const token = credentials?.apiKey || credentials?.accessToken;
if (token) headers["Authorization"] = `Bearer ${token}`;
if (stream) headers["Accept"] = "text/event-stream";
return headers;
}
if (stream) headers["Accept"] = "text/event-stream";
return headers;
}
async execute(opts) {
const result = await super.execute(opts);
if (!result?.response?.ok || !result.response.body) return result;
result.response = wrapNdjsonAsOpenAISse(result.response, opts.model);
return result;
}
async execute(opts) {
const result = await super.execute(opts);
if (!result?.response?.ok || !result.response.body) return result;
result.response = await peekForUpstreamError(result.response, opts.model, {
signal: opts.signal,
});
return result;
}
}
// How long to hold the response open while peeking the first upstream events.
// An upstream error event ("Network connection lost") is emitted at stream
// start, so the peek is fast; the bound just prevents a slow-started stream
// from being held hostage. Env: COMMANDCODE_PEEK_TIMEOUT_MS.
const PEEK_TIMEOUT_MS = (() => {
const raw = process.env.COMMANDCODE_PEEK_TIMEOUT_MS;
const n = raw ? parseInt(raw, 10) : NaN;
return Number.isFinite(n) && n > 0 ? n : 10 * 1000;
})();
// Event types that count as "the stream has started producing". Everything
// else (start, start-step, reasoning-start, text-start, ...) is metadata and
// does not end the peek.
const MEANINGFUL_EVENT_TYPES = new Set([
"text-delta",
"reasoning-delta",
"tool-input-start",
"tool-input-delta",
"tool-input-end",
"tool-call",
"finish-step",
"finish",
]);
function makeAbortError(reason) {
const error = new Error(reason?.message || reason || "Request aborted");
error.name = "AbortError";
return error;
}
function tryParseEvent(line) {
const trimmed = line.trim();
if (!trimmed) return null;
const json = trimmed.startsWith("data:") ? trimmed.slice(5).trim() : trimmed;
if (!json || json === "[DONE]") return null;
try {
return JSON.parse(json);
} catch {
return null;
}
}
function formatErrorValue(errVal) {
const errStr =
typeof errVal === "string"
? errVal
: typeof errVal?.message === "string"
? errVal.message
: JSON.stringify(errVal);
const errType =
typeof errVal === "string"
? "upstream_error"
: errVal?.type || "upstream_error";
return { message: errStr, type: errType };
}
/**
* Read the first upstream events before committing the response.
*
* - `{"type":"error"}` as the first meaningful event → return a 502 Response so
* chatCore's `!response.ok` path parses the error and triggers fallback.
* - Otherwise → re-emit the buffered bytes + the rest of the stream through the
* normal NDJSON → OpenAI SSE wrapper and return it untouched in spirit.
*
* Bounded by `timeoutMs` (default PEEK_TIMEOUT_MS): if no meaningful event
* arrives in time, or the request signal aborts, we commit whatever we have and
* let the regular stream pipeline (stall detection, abort handling) take over.
*/
export async function peekForUpstreamError(
originalResponse,
model,
{ signal = null, timeoutMs = PEEK_TIMEOUT_MS } = {},
) {
const reader = originalResponse.body.getReader();
const decoder = new TextDecoder();
const abortController = new AbortController();
const forwardAbort = () => abortController.abort(signal?.reason);
if (signal?.aborted) abortController.abort(signal?.reason);
else if (signal)
signal.addEventListener("abort", forwardAbort, { once: true });
// Raw bytes for lossless re-emission; decoded text is only used for line
// parsing / error detection. Never re-encode decoded text: TextDecoder
// holds a split multi-byte char internally and flush() would replace it
// with U+FFFD, corrupting the stream.
const rawChunks = [];
let peeked = "";
let errorEvent = null;
let committed = false;
const readWithTimeout = (ms) => {
if (abortController.signal.aborted) {
return Promise.reject(makeAbortError(abortController.signal.reason));
}
const timeoutPromise = new Promise((_, reject) => {
const t = setTimeout(() => reject(new Error("peek timeout")), ms);
t.unref?.();
});
const abortPromise = new Promise((_, reject) => {
abortController.signal.addEventListener(
"abort",
() => reject(makeAbortError(abortController.signal.reason)),
{ once: true },
);
});
return Promise.race([reader.read(), timeoutPromise, abortPromise]);
};
try {
const deadline = Date.now() + timeoutMs;
while (!errorEvent && !committed && Date.now() < deadline) {
const { done, value } = await readWithTimeout(
Math.max(deadline - Date.now(), 1),
);
if (done) break;
rawChunks.push(value);
peeked += decoder.decode(value, { stream: true });
const lines = peeked.split("\n");
// The last segment may be a partial line — only parse complete ones.
for (const line of lines.slice(0, -1)) {
const event = tryParseEvent(line);
if (!event?.type) continue;
if (event.type === "error") {
errorEvent = event;
break;
}
if (MEANINGFUL_EVENT_TYPES.has(event.type)) {
committed = true;
break;
}
}
}
} catch {
// timeout / abort / read failure during the peek → commit whatever we have;
// the downstream stream pipeline (stall detection, abort handling) takes over.
}
if (signal) signal.removeEventListener("abort", forwardAbort);
if (errorEvent) {
await reader.cancel("commandcode early error detected").catch(() => {});
const { message, type } = formatErrorValue(
errorEvent.error ?? errorEvent.message ?? "unknown",
);
return new Response(JSON.stringify({ error: { message, type } }), {
status: HTTP_STATUS.BAD_GATEWAY,
statusText: message.slice(0, 200),
headers: { "Content-Type": "application/json" },
});
}
const remaining = new ReadableStream({
start(controller) {
(async () => {
try {
// Re-emit RAW bytes (never re-encoded decoded text) so split
// multi-byte UTF-8 sequences survive the peek untouched.
for (const c of rawChunks) controller.enqueue(c);
while (true) {
const { done, value } = await reader.read();
if (done) break;
controller.enqueue(value);
}
controller.close();
} catch (err) {
controller.error(err);
}
})();
},
cancel() {
reader.cancel("commandcode stream cancelled").catch(() => {});
},
});
const combined = new Response(remaining, {
status: originalResponse.status,
statusText: originalResponse.statusText,
headers: originalResponse.headers,
});
return wrapNdjsonAsOpenAISse(combined, model);
}
function wrapNdjsonAsOpenAISse(originalResponse, model) {
const decoder = new TextDecoder();
const encoder = new TextEncoder();
let buffer = "";
const state = { model };
const decoder = new TextDecoder();
const encoder = new TextEncoder();
let buffer = "";
const state = { model };
const emitChunks = (chunks, controller) => {
if (!chunks) return;
const list = Array.isArray(chunks) ? chunks : [chunks];
for (const c of list) {
if (c == null) continue;
controller.enqueue(encoder.encode(`data: ${JSON.stringify(c)}\n\n`));
}
};
const emitChunks = (chunks, controller) => {
if (!chunks) return;
const list = Array.isArray(chunks) ? chunks : [chunks];
for (const c of list) {
if (c == null) continue;
controller.enqueue(encoder.encode(`data: ${JSON.stringify(c)}\n\n`));
}
};
const transform = new TransformStream({
transform(chunk, controller) {
buffer += decoder.decode(chunk, { stream: true });
const lines = buffer.split("\n");
buffer = lines.pop() || "";
for (const line of lines) {
const trimmed = line.trim();
if (!trimmed) continue;
// Translate AI SDK v5 NDJSON line to one or more OpenAI chunks
emitChunks(commandCodeToOpenAIResponse(trimmed, state), controller);
}
},
flush(controller) {
const trimmed = buffer.trim();
if (trimmed) {
emitChunks(commandCodeToOpenAIResponse(trimmed, state), controller);
}
controller.enqueue(encoder.encode(SSE_DONE));
},
});
const transform = new TransformStream({
transform(chunk, controller) {
buffer += decoder.decode(chunk, { stream: true });
const lines = buffer.split("\n");
buffer = lines.pop() || "";
for (const line of lines) {
const trimmed = line.trim();
if (!trimmed) continue;
// Translate AI SDK v5 NDJSON line to one or more OpenAI chunks
emitChunks(commandCodeToOpenAIResponse(trimmed, state), controller);
}
},
flush(controller) {
const trimmed = buffer.trim();
if (trimmed) {
emitChunks(commandCodeToOpenAIResponse(trimmed, state), controller);
}
controller.enqueue(encoder.encode(SSE_DONE));
},
});
const newBody = originalResponse.body.pipeThrough(transform);
return new Response(newBody, {
status: originalResponse.status,
statusText: originalResponse.statusText,
headers: originalResponse.headers,
});
const newBody = originalResponse.body.pipeThrough(transform);
return new Response(newBody, {
status: originalResponse.status,
statusText: originalResponse.statusText,
headers: originalResponse.headers,
});
}
export default CommandCodeExecutor;

View File

@@ -30,6 +30,7 @@ import { PROVIDERS } from "../config/providers.js";
import { proxyAwareFetch } from "../utils/proxyFetch.js";
import { SSE_DONE } from "../utils/sseConstants.js";
import { FETCH_CONNECT_TIMEOUT_MS } from "../config/runtimeConfig.js";
import { resolveProviderTimeoutMs } from "../services/providerTimeout.js";
import {
QODER_CHAT_URL_ENCODED,
QODER_CHAT_BASE_ALT,
@@ -536,7 +537,7 @@ export class QoderExecutor extends BaseExecutor {
};
// Abort if upstream doesn't return response headers within connect timeout.
const timeoutMs = this.config?.timeoutMs || FETCH_CONNECT_TIMEOUT_MS;
const timeoutMs = await resolveProviderTimeoutMs(this.provider, this.config?.timeoutMs, FETCH_CONNECT_TIMEOUT_MS);
const connectCtrl = new AbortController();
const connectTimer = setTimeout(() => connectCtrl.abort(new Error("fetch connect timeout")), timeoutMs);
const mergedSignal = signal ? AbortSignal.any([signal, connectCtrl.signal]) : connectCtrl.signal;

View File

@@ -372,7 +372,7 @@ export async function handleNonStreamingResponse({ providerResponse, provider, m
const totalLatency = Date.now() - requestStartTime;
saveRequestDetail(buildRequestDetail({
provider, model, connectionId,
provider, model, connectionId, apiKey,
latency: { ttft: totalLatency, total: totalLatency },
tokens: usage || { prompt_tokens: 0, completion_tokens: 0 },
request: extractRequestConfig(body, stream),

View File

@@ -58,11 +58,21 @@ export function extractUsageFromResponse(responseBody) {
return null;
}
// Mask API keys before they reach the requestDetails data blob / DB column.
// Only the prefix is kept — enough to distinguish keys without leaking them.
export function maskApiKey(key) {
if (!key || typeof key !== "string") return undefined;
const trimmed = key.trim();
if (trimmed.length <= 8) return trimmed.charAt(0) + "***";
return trimmed.slice(0, 8) + "***";
}
export function buildRequestDetail(base, overrides = {}) {
return {
provider: base.provider || "unknown",
model: base.model || "unknown",
connectionId: base.connectionId || undefined,
apiKey: maskApiKey(base.apiKey),
timestamp: new Date().toISOString(),
latency: base.latency || { ttft: 0, total: 0 },
tokens: base.tokens || { prompt_tokens: 0, completion_tokens: 0 },

View File

@@ -1,4 +1,5 @@
import { convertResponsesStreamToJson } from "../../transformer/streamToJsonConverter.js";
import { matchStreamErrorPatterns } from "../../utils/streamErrorPatterns.js";
import { createErrorResult } from "../../utils/error.js";
import { HTTP_STATUS } from "../../config/runtimeConfig.js";
import { FORMATS } from "../../translator/formats.js";
@@ -7,16 +8,17 @@ import { buildRequestDetail, extractRequestConfig, saveUsageStats, formatDoneLin
import { ROLE, RESPONSES_ITEM } from "../../translator/schema/index.js";
// Responses-API providers (e.g. codex) may emit SSE without content-type + use Responses output shape
const isResponsesProvider = (p) => PROVIDERS[p]?.format === FORMATS.OPENAI_RESPONSES;
const isResponsesProvider = (p) =>
PROVIDERS[p]?.format === FORMATS.OPENAI_RESPONSES;
import { saveRequestDetail, appendRequestLog } from "@/lib/usageDb.js";
function textFromResponsesMessageItem(item) {
if (!item?.content || !Array.isArray(item.content)) return "";
const byType = item.content.find((c) => c.type === "output_text");
if (typeof byType?.text === "string") return byType.text;
const anyText = item.content.find((c) => typeof c.text === "string");
if (typeof anyText?.text === "string") return anyText.text;
return "";
if (!item?.content || !Array.isArray(item.content)) return "";
const byType = item.content.find((c) => c.type === "output_text");
if (typeof byType?.text === "string") return byType.text;
const anyText = item.content.find((c) => typeof c.text === "string");
if (typeof anyText?.text === "string") return anyText.text;
return "";
}
/**
@@ -24,15 +26,15 @@ function textFromResponsesMessageItem(item) {
* Early message blocks often have empty output_text; the user-visible answer is usually in the last non-empty message.
*/
function pickAssistantMessageForChatCompletion(output) {
if (!Array.isArray(output)) return { msgItem: null, textContent: null };
const messages = output.filter((item) => item?.type === "message");
if (messages.length === 0) return { msgItem: null, textContent: null };
for (let i = messages.length - 1; i >= 0; i--) {
const text = textFromResponsesMessageItem(messages[i]);
if (text.length > 0) return { msgItem: messages[i], textContent: text };
}
const last = messages[messages.length - 1];
return { msgItem: last, textContent: textFromResponsesMessageItem(last) };
if (!Array.isArray(output)) return { msgItem: null, textContent: null };
const messages = output.filter((item) => item?.type === "message");
if (messages.length === 0) return { msgItem: null, textContent: null };
for (let i = messages.length - 1; i >= 0; i--) {
const text = textFromResponsesMessageItem(messages[i]);
if (text.length > 0) return { msgItem: messages[i], textContent: text };
}
const last = messages[messages.length - 1];
return { msgItem: last, textContent: textFromResponsesMessageItem(last) };
}
/**
@@ -110,250 +112,414 @@ function chatCompletionToResponses(responseBody, customToolNames = null) {
* Used when provider forces streaming but client wants non-streaming.
*/
export function parseSSEToOpenAIResponse(rawSSE, fallbackModel) {
const chunks = [];
let streamError = null;
const chunks = [];
let streamError = null;
for (const line of String(rawSSE || "").split("\n")) {
const trimmed = line.trim();
if (!trimmed.startsWith("data:")) continue;
const payload = trimmed.slice(5).trim();
if (!payload || payload === "[DONE]") continue;
try {
const chunk = JSON.parse(payload);
if (chunk?.error) streamError = chunk.error;
else chunks.push(chunk);
} catch { /* ignore malformed lines */ }
}
for (const line of String(rawSSE || "").split("\n")) {
const trimmed = line.trim();
if (!trimmed.startsWith("data:")) continue;
const payload = trimmed.slice(5).trim();
if (!payload || payload === "[DONE]") continue;
try {
const chunk = JSON.parse(payload);
if (chunk?.error) streamError = chunk.error;
else chunks.push(chunk);
} catch {
/* ignore malformed lines */
}
}
if (streamError) return { error: streamError };
if (chunks.length === 0) return null;
if (streamError) return { error: streamError };
if (chunks.length === 0) return null;
const first = chunks[0];
const contentParts = [];
const reasoningParts = [];
const toolCallMap = new Map(); // index -> { id, type, function: { name, arguments } }
let finishReason = "stop";
let usage = null;
const first = chunks[0];
const contentParts = [];
const reasoningParts = [];
const toolCallMap = new Map(); // index -> { id, type, function: { name, arguments } }
let finishReason = "stop";
let usage = null;
for (const chunk of chunks) {
const choice = chunk?.choices?.[0];
const delta = choice?.delta || {};
if (typeof delta.content === "string" && delta.content.length > 0) contentParts.push(delta.content);
if (typeof delta.reasoning_content === "string" && delta.reasoning_content.length > 0) reasoningParts.push(delta.reasoning_content);
if (choice?.finish_reason) finishReason = choice.finish_reason;
if (chunk?.usage && typeof chunk.usage === "object") usage = chunk.usage;
for (const chunk of chunks) {
const choice = chunk?.choices?.[0];
const delta = choice?.delta || {};
if (typeof delta.content === "string" && delta.content.length > 0)
contentParts.push(delta.content);
if (
typeof delta.reasoning_content === "string" &&
delta.reasoning_content.length > 0
)
reasoningParts.push(delta.reasoning_content);
if (choice?.finish_reason) finishReason = choice.finish_reason;
if (chunk?.usage && typeof chunk.usage === "object") usage = chunk.usage;
// Accumulate tool_calls from streaming deltas
if (Array.isArray(delta.tool_calls)) {
for (const tc of delta.tool_calls) {
const idx = tc.index ?? 0;
if (!toolCallMap.has(idx)) {
toolCallMap.set(idx, { id: tc.id || "", type: "function", function: { name: "", arguments: "" } });
}
const existing = toolCallMap.get(idx);
if (tc.id) existing.id = tc.id;
if (tc.function?.name) existing.function.name += tc.function.name;
if (tc.function?.arguments) existing.function.arguments += tc.function.arguments;
}
}
}
// Accumulate tool_calls from streaming deltas
if (Array.isArray(delta.tool_calls)) {
for (const tc of delta.tool_calls) {
const idx = tc.index ?? 0;
if (!toolCallMap.has(idx)) {
toolCallMap.set(idx, {
id: tc.id || "",
type: "function",
function: { name: "", arguments: "" },
});
}
const existing = toolCallMap.get(idx);
if (tc.id) existing.id = tc.id;
if (tc.function?.name) existing.function.name += tc.function.name;
if (tc.function?.arguments)
existing.function.arguments += tc.function.arguments;
}
}
}
const message = { role: "assistant", content: contentParts.join("") || (toolCallMap.size > 0 ? null : "") };
if (reasoningParts.length > 0) message.reasoning_content = reasoningParts.join("");
if (toolCallMap.size > 0) {
message.tool_calls = [...toolCallMap.entries()].sort((a, b) => a[0] - b[0]).map(([, tc]) => tc);
}
const message = {
role: "assistant",
content: contentParts.join("") || (toolCallMap.size > 0 ? null : ""),
};
if (reasoningParts.length > 0)
message.reasoning_content = reasoningParts.join("");
if (toolCallMap.size > 0) {
message.tool_calls = [...toolCallMap.entries()]
.sort((a, b) => a[0] - b[0])
.map(([, tc]) => tc);
}
const result = {
id: first.id || `chatcmpl-${Date.now()}`,
object: "chat.completion",
created: first.created || Math.floor(Date.now() / 1000),
model: first.model || fallbackModel || "unknown",
choices: [{ index: 0, message, finish_reason: finishReason }]
};
if (usage) result.usage = usage;
return result;
const result = {
id: first.id || `chatcmpl-${Date.now()}`,
object: "chat.completion",
created: first.created || Math.floor(Date.now() / 1000),
model: first.model || fallbackModel || "unknown",
choices: [{ index: 0, message, finish_reason: finishReason }],
};
if (usage) result.usage = usage;
return result;
}
/**
* Handle case: provider forced streaming but client wants JSON.
* Supports both Codex/Responses API SSE and standard Chat Completions SSE.
*/
export async function handleForcedSSEToJson({ providerResponse, sourceFormat, targetFormat, provider, model, body, stream, translatedBody, finalBody, requestStartTime, connectionId, apiKey, clientRawRequest, onRequestSuccess, customToolNames, trackDone, appendLog, reqTag, log }) {
const contentType = providerResponse.headers.get("content-type") || "";
const isSSE = contentType.includes("text/event-stream") || (contentType === "" && isResponsesProvider(provider));
if (!isSSE) return null; // not handled here
export async function handleForcedSSEToJson({
providerResponse,
sourceFormat,
targetFormat,
provider,
model,
body,
stream,
translatedBody,
finalBody,
requestStartTime,
connectionId,
apiKey,
clientRawRequest,
onRequestSuccess,
customToolNames,
trackDone,
appendLog,
reqTag,
log,
streamErrorPatterns,
}) {
const contentType = providerResponse.headers.get("content-type") || "";
const isSSE =
contentType.includes("text/event-stream") ||
(contentType === "" && isResponsesProvider(provider));
if (!isSSE) return null; // not handled here
trackDone();
trackDone();
const ctx = {
provider, model, connectionId,
request: extractRequestConfig(body, stream),
providerRequest: finalBody || translatedBody || null
};
const ctx = {
provider,
model,
connectionId,
request: extractRequestConfig(body, stream),
providerRequest: finalBody || translatedBody || null,
};
// Codex/Responses API SSE path
// Branch on the UPSTREAM format (targetFormat = format we spoke to the provider in),
// not the client format: a Responses-API client behind a chat-native forced-streaming
// provider still receives chat SSE chunks, which must go through the standard path.
const isCodexResponsesApi = isResponsesProvider(provider) || targetFormat === FORMATS.OPENAI_RESPONSES;
if (isCodexResponsesApi) {
try {
const jsonResponse = await convertResponsesStreamToJson(providerResponse.body);
if (onRequestSuccess) await onRequestSuccess();
// Codex/Responses API SSE path
// Branch on the UPSTREAM format (targetFormat = format we spoke to the provider in),
// not the client format: a Responses-API client behind a chat-native forced-streaming
// provider still receives chat SSE chunks, which must go through the standard path.
const isCodexResponsesApi =
isResponsesProvider(provider) || targetFormat === FORMATS.OPENAI_RESPONSES;
if (isCodexResponsesApi) {
try {
const jsonResponse = await convertResponsesStreamToJson(
providerResponse.body,
);
if (onRequestSuccess) await onRequestSuccess();
const usage = jsonResponse.usage || {};
appendLog({ tokens: usage, status: "200 OK" });
saveUsageStats({ provider, model, tokens: usage, connectionId, apiKey, endpoint: clientRawRequest?.endpoint, silent: true });
if (log?.line) log.line(reqTag, "📊", formatDoneLine({ usage, latency: { total: Date.now() - requestStartTime } }));
const usage = jsonResponse.usage || {};
appendLog({ tokens: usage, status: "200 OK" });
saveUsageStats({
provider,
model,
tokens: usage,
connectionId,
apiKey,
endpoint: clientRawRequest?.endpoint,
silent: true,
});
if (log?.line)
log.line(
reqTag,
"📊",
formatDoneLine({
usage,
latency: { total: Date.now() - requestStartTime },
}),
);
// Same cache-inclusive total for the recorded detail, so the DB and the
// client-facing usage can never disagree.
const inTokensForLog = (usage.input_tokens || 0)
+ (usage.cache_read_input_tokens || usage.cached_tokens || 0)
+ (usage.cache_creation_input_tokens || 0);
const { msgItem, textContent } = pickAssistantMessageForChatCompletion(jsonResponse.output);
const totalLatency = Date.now() - requestStartTime;
// Same cache-inclusive total for the recorded detail, so the DB and the
// client-facing usage can never disagree.
const inTokensForLog = (usage.input_tokens || 0)
+ (usage.cache_read_input_tokens || usage.cached_tokens || 0)
+ (usage.cache_creation_input_tokens || 0);
const { msgItem, textContent } = pickAssistantMessageForChatCompletion(
jsonResponse.output,
);
const totalLatency = Date.now() - requestStartTime;
saveRequestDetail(buildRequestDetail({
...ctx,
latency: { ttft: totalLatency, total: totalLatency },
tokens: { prompt_tokens: inTokensForLog, completion_tokens: usage.output_tokens || 0 },
response: { content: textContent, thinking: null, finish_reason: jsonResponse.status || "unknown" },
status: "success"
}, { endpoint: clientRawRequest?.endpoint || null })).catch(() => {});
saveRequestDetail(
buildRequestDetail(
{
...ctx,
apiKey,
latency: { ttft: totalLatency, total: totalLatency },
tokens: {
prompt_tokens: inTokensForLog,
completion_tokens: usage.output_tokens || 0,
},
response: {
content: textContent,
thinking: null,
finish_reason: jsonResponse.status || "unknown",
},
status: "success",
},
{ endpoint: clientRawRequest?.endpoint || null },
),
).catch(() => {});
// Client is Responses API → return as-is
if (sourceFormat === FORMATS.OPENAI_RESPONSES) {
return { success: true, response: new Response(JSON.stringify(jsonResponse), { headers: { "Content-Type": "application/json", "Access-Control-Allow-Origin": "*" } }) };
}
// Client is Responses API → return as-is
if (sourceFormat === FORMATS.OPENAI_RESPONSES) {
return {
success: true,
response: new Response(JSON.stringify(jsonResponse), {
headers: {
"Content-Type": "application/json",
"Access-Control-Allow-Origin": "*",
},
}),
};
}
// Build client-format response.
// input_tokens EXCLUDES cached tokens on cache-capable upstreams, so summing
// only input+output under-reports prompt_tokens — measured: 2012 reported
// where the real prompt was ~5344 with 5332 served from cache. Fold the cache
// counters in, and keep them visible in prompt_tokens_details so a client can
// tell a cache hit from a small prompt.
const cacheRead = usage.cache_read_input_tokens || usage.cached_tokens || 0;
const cacheCreate = usage.cache_creation_input_tokens || 0;
const inTokens = (usage.input_tokens || 0) + cacheRead + cacheCreate;
const outTokens = usage.output_tokens || 0;
const cacheDetails = (cacheRead > 0 || cacheCreate > 0)
? { prompt_tokens_details: {
...(cacheRead > 0 ? { cached_tokens: cacheRead } : {}),
...(cacheCreate > 0 ? { cache_creation_tokens: cacheCreate } : {}) } }
: {};
let finalResp;
// Build client-format response.
// input_tokens EXCLUDES cached tokens on cache-capable upstreams, so summing
// only input+output under-reports prompt_tokens — measured: 2012 reported
// where the real prompt was ~5344 with 5332 served from cache. Fold the cache
// counters in, and keep them visible in prompt_tokens_details so a client can
// tell a cache hit from a small prompt.
const cacheRead = usage.cache_read_input_tokens || usage.cached_tokens || 0;
const cacheCreate = usage.cache_creation_input_tokens || 0;
const inTokens = (usage.input_tokens || 0) + cacheRead + cacheCreate;
const outTokens = usage.output_tokens || 0;
const cacheDetails = (cacheRead > 0 || cacheCreate > 0)
? {
prompt_tokens_details: {
...(cacheRead > 0 ? { cached_tokens: cacheRead } : {}),
...(cacheCreate > 0 ? { cache_creation_tokens: cacheCreate } : {}) } }
: {};
let finalResp;
// Extract tool calls from Responses API output (function_call items)
const funcCallItems = (jsonResponse.output || []).filter(item => item.type === "function_call");
const toolCalls = funcCallItems.map((item, idx) => ({
id: item.call_id || `call_${item.name}_${Date.now()}_${idx}`,
type: "function",
function: {
name: item.name,
arguments: typeof item.arguments === "string" ? item.arguments : JSON.stringify(item.arguments || {})
}
}));
const hasToolCalls = toolCalls.length > 0;
// Extract tool calls from Responses API output (function_call items)
const funcCallItems = (jsonResponse.output || []).filter(
(item) => item.type === "function_call",
);
const toolCalls = funcCallItems.map((item, idx) => ({
id: item.call_id || `call_${item.name}_${Date.now()}_${idx}`,
type: "function",
function: {
name: item.name,
arguments:
typeof item.arguments === "string"
? item.arguments
: JSON.stringify(item.arguments || {}),
},
}));
const hasToolCalls = toolCalls.length > 0;
if (sourceFormat === FORMATS.ANTIGRAVITY || sourceFormat === FORMATS.GEMINI || sourceFormat === FORMATS.GEMINI_CLI) {
finalResp = {
response: {
candidates: [{ content: { role: "model", parts: [{ text: textContent || "" }] }, finishReason: "STOP", index: 0 }],
usageMetadata: { promptTokenCount: inTokens, candidatesTokenCount: outTokens, totalTokenCount: inTokens + outTokens },
modelVersion: model,
responseId: jsonResponse.id || `resp_${Date.now()}`
}
};
} else {
const message = { role: "assistant", content: textContent || (hasToolCalls ? null : "") };
if (hasToolCalls) message.tool_calls = toolCalls;
const responseDone = jsonResponse.status === "completed" || jsonResponse.status === "done";
const finishReason = hasToolCalls ? "tool_calls" : (responseDone ? "stop" : (jsonResponse.status || "stop"));
finalResp = {
id: jsonResponse.id || `chatcmpl-${Date.now()}`,
object: "chat.completion",
created: jsonResponse.created_at || Math.floor(Date.now() / 1000),
model: jsonResponse.model || model,
choices: [{ index: 0, message, finish_reason: finishReason }],
usage: { prompt_tokens: inTokens, completion_tokens: outTokens, total_tokens: inTokens + outTokens, ...cacheDetails }
};
}
if (
sourceFormat === FORMATS.ANTIGRAVITY ||
sourceFormat === FORMATS.GEMINI ||
sourceFormat === FORMATS.GEMINI_CLI
) {
finalResp = {
response: {
candidates: [
{
content: {
role: "model",
parts: [{ text: textContent || "" }],
},
finishReason: "STOP",
index: 0,
},
],
usageMetadata: {
promptTokenCount: inTokens,
candidatesTokenCount: outTokens,
totalTokenCount: inTokens + outTokens,
},
modelVersion: model,
responseId: jsonResponse.id || `resp_${Date.now()}`,
},
};
} else {
const message = {
role: "assistant",
content: textContent || (hasToolCalls ? null : ""),
};
if (hasToolCalls) message.tool_calls = toolCalls;
const responseDone =
jsonResponse.status === "completed" || jsonResponse.status === "done";
const finishReason = hasToolCalls
? "tool_calls"
: responseDone
? "stop"
: jsonResponse.status || "stop";
finalResp = {
id: jsonResponse.id || `chatcmpl-${Date.now()}`,
object: "chat.completion",
created: jsonResponse.created_at || Math.floor(Date.now() / 1000),
model: jsonResponse.model || model,
choices: [{ index: 0, message, finish_reason: finishReason }],
usage: {
prompt_tokens: inTokens,
completion_tokens: outTokens,
total_tokens: inTokens + outTokens,
...cacheDetails,
},
};
}
return { success: true, response: new Response(JSON.stringify(finalResp), { headers: { "Content-Type": "application/json", "Access-Control-Allow-Origin": "*" } }) };
} catch (err) {
console.error("[ChatCore] Responses API SSE→JSON failed:", err);
return createErrorResult(HTTP_STATUS.BAD_GATEWAY, "Failed to convert streaming response to JSON");
}
}
return {
success: true,
response: new Response(JSON.stringify(finalResp), {
headers: {
"Content-Type": "application/json",
"Access-Control-Allow-Origin": "*",
},
}),
};
} catch (err) {
console.error("[ChatCore] Responses API SSE→JSON failed:", err);
return createErrorResult(
HTTP_STATUS.BAD_GATEWAY,
"Failed to convert streaming response to JSON",
);
}
}
// Standard Chat Completions SSE path
try {
const sseText = await providerResponse.text();
const parsed = parseSSEToOpenAIResponse(sseText, model);
if (!parsed) return createErrorResult(HTTP_STATUS.BAD_GATEWAY, "Invalid SSE response for non-streaming request");
if (parsed.error) {
return createErrorResult(
HTTP_STATUS.BAD_GATEWAY,
parsed.error.message || "Upstream SSE stream failed"
);
}
// Standard Chat Completions SSE path
try {
const sseText = await providerResponse.text();
const parsed = parseSSEToOpenAIResponse(sseText, model);
if (!parsed)
return createErrorResult(
HTTP_STATUS.BAD_GATEWAY,
"Invalid SSE response for non-streaming request",
);
if (parsed.error) {
return createErrorResult(
HTTP_STATUS.BAD_GATEWAY,
parsed.error.message || "Upstream SSE stream failed",
);
}
if (onRequestSuccess) await onRequestSuccess();
// Config-driven in-stream error detection: the request "succeeded" at the
// HTTP level, but the content signals an upstream failure — treat it as an
// error so account/combo fallback and FAILED logging kick in.
const matchedPattern = matchStreamErrorPatterns(
streamErrorPatterns?.[provider],
parsed.choices?.[0]?.message?.content || "",
);
if (matchedPattern) {
return createErrorResult(
HTTP_STATUS.BAD_GATEWAY,
`Stream error pattern matched: ${matchedPattern}`,
);
}
const usage = parsed.usage || {};
appendLog({ tokens: usage, status: "200 OK" });
saveUsageStats({ provider, model, tokens: usage, connectionId, apiKey, endpoint: clientRawRequest?.endpoint, silent: true });
if (log?.line) log.line(reqTag, "📊", formatDoneLine({ usage, latency: { total: Date.now() - requestStartTime } }));
if (onRequestSuccess) await onRequestSuccess();
const totalLatency = Date.now() - requestStartTime;
saveRequestDetail(buildRequestDetail({
...ctx,
latency: { ttft: totalLatency, total: totalLatency },
tokens: usage,
response: {
content: parsed.choices?.[0]?.message?.content || null,
thinking: parsed.choices?.[0]?.message?.reasoning_content || null,
finish_reason: parsed.choices?.[0]?.finish_reason || "unknown"
},
status: "success"
}, { endpoint: clientRawRequest?.endpoint || null })).catch(() => {});
const usage = parsed.usage || {};
appendLog({ tokens: usage, status: "200 OK" });
saveUsageStats({
provider,
model,
tokens: usage,
connectionId,
apiKey,
endpoint: clientRawRequest?.endpoint,
silent: true,
});
if (log?.line)
log.line(
reqTag,
"📊",
formatDoneLine({
usage,
latency: { total: Date.now() - requestStartTime },
}),
);
// Re-attach usage explicitly. This handler already HAS the correct usage — it is
// the same object written to the usage DB, and for a cached Claude request that DB
// row reads cache_read_input_tokens: 11022 — yet the client was observed receiving
// no usage field at all (verified 2026-08-04 with a fingerprinted payload matched
// on both sides). Whatever drops it between assembly and serialisation, the client
// must not be left unable to account for its own token spend: a caller cannot tell
// a 90%-cached request from a cheap one without this.
if (usage && Object.keys(usage).length > 0) parsed.usage = usage;
// Re-attach usage explicitly. This handler already HAS the correct usage — it is
// the same object written to the usage DB, and for a cached Claude request that DB
// row reads cache_read_input_tokens: 11022 — yet the client was observed receiving
// no usage field at all (verified 2026-08-04 with a fingerprinted payload matched
// on both sides). Whatever drops it between assembly and serialisation, the client
// must not be left unable to account for its own token spend: a caller cannot tell
// a 90%-cached request from a cheap one without this.
if (usage && Object.keys(usage).length > 0) parsed.usage = usage;
// Strip reasoning_content only when content is non-empty.
// When content is empty (e.g. thinking models that used all tokens for reasoning),
// reasoning_content is the only useful output and must be preserved.
// Previously this was unconditional, which broke Qwen3.5, Claude extended thinking, etc.
if (parsed?.choices) {
for (const choice of parsed.choices) {
if (choice?.message?.reasoning_content && choice.message.content) {
delete choice.message.reasoning_content;
}
}
}
// Strip reasoning_content only when content is non-empty.
// When content is empty (e.g. thinking models that used all tokens for reasoning),
// reasoning_content is the only useful output and must be preserved.
// Previously this was unconditional, which broke Qwen3.5, Claude extended thinking, etc.
if (parsed?.choices) {
for (const choice of parsed.choices) {
if (choice?.message?.reasoning_content && choice.message.content) {
delete choice.message.reasoning_content;
}
}
}
// A Responses-format client (e.g. Codex) forced this provider to stream,
// but wants JSON back. parseSSEToOpenAIResponse yields a Chat Completions
// body; convert it to the Responses `output` shape so tool_calls are not
// lost on the non-streaming return path. Inlined (not imported from
// nonStreamingHandler.js) to avoid a circular import: nonStreamingHandler
// already imports parseSSEToOpenAIResponse from this module.
const finalBody = sourceFormat === FORMATS.OPENAI_RESPONSES
? chatCompletionToResponses(parsed, customToolNames)
: parsed;
// A Responses-format client (e.g. Codex) forced this provider to stream,
// but wants JSON back. parseSSEToOpenAIResponse yields a Chat Completions
// body; convert it to the Responses `output` shape so tool_calls are not
// lost on the non-streaming return path. Inlined (not imported from
// nonStreamingHandler.js) to avoid a circular import: nonStreamingHandler
// already imports parseSSEToOpenAIResponse from this module.
const finalBody = sourceFormat === FORMATS.OPENAI_RESPONSES
? chatCompletionToResponses(parsed, customToolNames)
: parsed;
return { success: true, response: new Response(JSON.stringify(finalBody), { headers: { "Content-Type": "application/json", "Access-Control-Allow-Origin": "*" } }) };
} catch (err) {
console.error("[ChatCore] Chat Completions SSE→JSON failed:", err);
return createErrorResult(HTTP_STATUS.BAD_GATEWAY, "Failed to convert streaming response to JSON");
}
return {
success: true,
response: new Response(JSON.stringify(finalBody), {
headers: {
"Content-Type": "application/json",
"Access-Control-Allow-Origin": "*",
},
}),
};
} catch (err) {
console.error("[ChatCore] Chat Completions SSE→JSON failed:", err);
return createErrorResult(
HTTP_STATUS.BAD_GATEWAY,
"Failed to convert streaming response to JSON",
);
}
}

View File

@@ -1,144 +1,343 @@
import { FORMATS } from "../../translator/formats.js";
import { needsTranslation } from "../../translator/index.js";
import { createSSETransformStreamWithLogger, createPassthroughStreamWithLogger } from "../../utils/stream.js";
import {
createSSETransformStreamWithLogger,
createPassthroughStreamWithLogger,
} from "../../utils/stream.js";
import { pipeWithDisconnect } from "../../utils/streamHandler.js";
import { PROVIDERS } from "../../config/providers.js";
import { STREAM_STALL_TIMEOUT_MS } from "../../config/runtimeConfig.js";
import { buildAbortedResponsesTerminalBytes } from "../../utils/responsesStreamHelpers.js";
import { buildRequestDetail, extractRequestConfig, saveUsageStats, formatDoneLine } from "./requestDetail.js";
import {
buildRequestDetail,
extractRequestConfig,
saveUsageStats,
formatDoneLine,
} from "./requestDetail.js";
import { streamStatusForContent } from "../../utils/streamErrorPatterns.js";
import { saveRequestDetail } from "@/lib/usageDb.js";
import { SSE_HEADERS_CORS as SSE_HEADERS } from "../../utils/sseConstants.js";
// Codex returns Responses API SSE → which client format to translate INTO, by request sourceFormat.
// Gemini-family all map to ANTIGRAVITY decoder; unknown sources fall back to OPENAI.
const CODEX_SOURCE_TO_TARGET = {
[FORMATS.OPENAI_RESPONSES]: FORMATS.OPENAI_RESPONSES,
[FORMATS.CLAUDE]: FORMATS.CLAUDE,
[FORMATS.ANTIGRAVITY]: FORMATS.ANTIGRAVITY,
[FORMATS.GEMINI]: FORMATS.ANTIGRAVITY,
[FORMATS.GEMINI_CLI]: FORMATS.ANTIGRAVITY,
[FORMATS.OPENAI_RESPONSES]: FORMATS.OPENAI_RESPONSES,
[FORMATS.CLAUDE]: FORMATS.CLAUDE,
[FORMATS.ANTIGRAVITY]: FORMATS.ANTIGRAVITY,
[FORMATS.GEMINI]: FORMATS.ANTIGRAVITY,
[FORMATS.GEMINI_CLI]: FORMATS.ANTIGRAVITY,
};
/**
* Determine which SSE transform stream to use based on provider/format.
*/
function buildTransformStream({ provider, sourceFormat, targetFormat, userAgent, reqLogger, toolNameMap, customToolNames, model, connectionId, body, onStreamComplete, apiKey }) {
const isDroidCLI = userAgent?.toLowerCase().includes("droid") || userAgent?.toLowerCase().includes("codex-cli");
// Responses-API providers (e.g. codex) emit Responses SSE → translate into client format
const isResponsesProvider = PROVIDERS[provider]?.format === FORMATS.OPENAI_RESPONSES;
const needsCodexTranslation = isResponsesProvider && targetFormat === FORMATS.OPENAI_RESPONSES && !isDroidCLI;
function buildTransformStream({
provider,
sourceFormat,
targetFormat,
userAgent,
reqLogger,
toolNameMap,
customToolNames,
model,
connectionId,
body,
onStreamComplete,
apiKey,
}) {
const isDroidCLI =
userAgent?.toLowerCase().includes("droid") ||
userAgent?.toLowerCase().includes("codex-cli");
// Responses-API providers (e.g. codex) emit Responses SSE → translate into client format
const isResponsesProvider =
PROVIDERS[provider]?.format === FORMATS.OPENAI_RESPONSES;
const needsCodexTranslation =
isResponsesProvider &&
targetFormat === FORMATS.OPENAI_RESPONSES &&
!isDroidCLI;
if (needsCodexTranslation) {
const codexTarget = CODEX_SOURCE_TO_TARGET[sourceFormat] || FORMATS.OPENAI;
return createSSETransformStreamWithLogger(FORMATS.OPENAI_RESPONSES, codexTarget, provider, reqLogger, toolNameMap, model, connectionId, body, onStreamComplete, apiKey, customToolNames);
}
if (needsCodexTranslation) {
const codexTarget = CODEX_SOURCE_TO_TARGET[sourceFormat] || FORMATS.OPENAI;
return createSSETransformStreamWithLogger(
FORMATS.OPENAI_RESPONSES,
codexTarget,
provider,
reqLogger,
toolNameMap,
model,
connectionId,
body,
onStreamComplete,
apiKey,
customToolNames,
);
}
if (needsTranslation(targetFormat, sourceFormat)) {
return createSSETransformStreamWithLogger(targetFormat, sourceFormat, provider, reqLogger, toolNameMap, model, connectionId, body, onStreamComplete, apiKey, customToolNames);
}
if (needsTranslation(targetFormat, sourceFormat)) {
return createSSETransformStreamWithLogger(
targetFormat,
sourceFormat,
provider,
reqLogger,
toolNameMap,
model,
connectionId,
body,
onStreamComplete,
apiKey,
customToolNames,
);
}
return createPassthroughStreamWithLogger(provider, reqLogger, model, connectionId, body, onStreamComplete, apiKey);
return createPassthroughStreamWithLogger(
provider,
reqLogger,
model,
connectionId,
body,
onStreamComplete,
apiKey,
);
}
/**
* Handle streaming response — pipe provider SSE through transform stream to client.
*/
export async function handleStreamingResponse({ providerResponse, provider, model, sourceFormat, targetFormat, userAgent, body, stream, translatedBody, finalBody, requestStartTime, connectionId, apiKey, clientRawRequest, onRequestSuccess, reqLogger, toolNameMap, customToolNames, streamController, onStreamComplete, streamDetailId, pxpipe, reqTag, log }) {
if (onRequestSuccess) {
Promise.resolve()
.then(onRequestSuccess)
.catch(err => {
console.error("[ChatCore] onRequestSuccess failed:", err?.message || err);
});
}
export async function handleStreamingResponse({
providerResponse,
provider,
model,
sourceFormat,
targetFormat,
userAgent,
body,
stream,
translatedBody,
finalBody,
requestStartTime,
connectionId,
apiKey,
clientRawRequest,
onRequestSuccess,
reqLogger,
toolNameMap,
customToolNames,
streamController,
onStreamComplete,
streamDetailId,
pxpipe,
reqTag,
log,
}) {
if (onRequestSuccess) {
Promise.resolve()
.then(onRequestSuccess)
.catch((err) => {
console.error(
"[ChatCore] onRequestSuccess failed:",
err?.message || err,
);
});
}
// When upstream returns HTML/text instead of SSE (e.g. Cloudflare 5xx error
// page), piping it through the SSE transform stream causes Next.js
// "failed to pipe response" and crashes the chat router. Read the body,
// pull a short human-readable message from the <title>, sanitize it, and
// return a clean JSON error instead. The message is stripped of HTML tags
// and clamped so untrusted upstream text never reaches the client verbatim
// (the UI may render error.message as HTML).
const upstreamContentType = (providerResponse.headers.get('content-type') || '').toLowerCase();
if (upstreamContentType && !upstreamContentType.includes('text/event-stream') && !upstreamContentType.includes('application/json')) {
const bodyText = await providerResponse.text().catch(() => '');
const titleMatch = bodyText.match(/<title>([^<]+)<\/title>/i);
const sanitizedTitle = (titleMatch?.[1] || '').replace(/<[^>]*>/g, '').replace(/[\r\n]+/g, ' ').trim().slice(0, 160);
const shortMsg = sanitizedTitle
|| (bodyText.length < 200 ? bodyText.replace(/<[^>]*>/g, '').trim().slice(0, 160) : `Upstream returned non-SSE response (${upstreamContentType})`);
const status = providerResponse.status || 502;
if (log?.errorLine) log.errorLine(reqTag, "✗", `BLOCKED ${status} · ${provider}/${model} · non-SSE (${upstreamContentType})\n ${shortMsg}`);
else console.warn(`[STREAM] ${provider} | ${model} | blocked pipe: ${shortMsg} [${status}]`);
streamController?.handleError?.(new Error(`upstream non-SSE: ${status}`));
return {
success: false,
response: new Response(JSON.stringify({ error: { message: `[${status}]: ${shortMsg}` } }), {
status,
headers: { 'Content-Type': 'application/json', 'Access-Control-Allow-Origin': '*' },
}),
};
}
// When upstream returns HTML/text instead of SSE (e.g. Cloudflare 5xx error
// page), piping it through the SSE transform stream causes Next.js
// "failed to pipe response" and crashes the chat router. Read the body,
// pull a short human-readable message from the <title>, sanitize it, and
// return a clean JSON error instead. The message is stripped of HTML tags
// and clamped so untrusted upstream text never reaches the client verbatim
// (the UI may render error.message as HTML).
const upstreamContentType = (
providerResponse.headers.get("content-type") || ""
).toLowerCase();
if (
upstreamContentType &&
!upstreamContentType.includes("text/event-stream") &&
!upstreamContentType.includes("application/json")
) {
const bodyText = await providerResponse.text().catch(() => "");
const titleMatch = bodyText.match(/<title>([^<]+)<\/title>/i);
const sanitizedTitle = (titleMatch?.[1] || "")
.replace(/<[^>]*>/g, "")
.replace(/[\r\n]+/g, " ")
.trim()
.slice(0, 160);
const shortMsg =
sanitizedTitle ||
(bodyText.length < 200
? bodyText
.replace(/<[^>]*>/g, "")
.trim()
.slice(0, 160)
: `Upstream returned non-SSE response (${upstreamContentType})`);
const status = providerResponse.status || 502;
if (log?.errorLine)
log.errorLine(
reqTag,
"✗",
`BLOCKED ${status} · ${provider}/${model} · non-SSE (${upstreamContentType})\n ${shortMsg}`,
);
else
console.warn(
`[STREAM] ${provider} | ${model} | blocked pipe: ${shortMsg} [${status}]`,
);
streamController?.handleError?.(new Error(`upstream non-SSE: ${status}`));
return {
success: false,
response: new Response(
JSON.stringify({ error: { message: `[${status}]: ${shortMsg}` } }),
{
status,
headers: {
"Content-Type": "application/json",
"Access-Control-Allow-Origin": "*",
},
},
),
};
}
const transformStream = buildTransformStream({ provider, sourceFormat, targetFormat, userAgent, reqLogger, toolNameMap, customToolNames, model, connectionId, body, onStreamComplete, apiKey });
const transformStream = buildTransformStream({
provider,
sourceFormat,
targetFormat,
userAgent,
reqLogger,
toolNameMap,
customToolNames,
model,
connectionId,
body,
onStreamComplete,
apiKey,
});
// Responses passthrough: synthesize response.failed + [DONE] if the stream aborts/stalls before a terminal event
const isResponsesPassthrough = sourceFormat === FORMATS.OPENAI_RESPONSES && targetFormat === FORMATS.OPENAI_RESPONSES;
const onAbortTerminal = isResponsesPassthrough ? buildAbortedResponsesTerminalBytes : null;
const stallTimeoutMs = PROVIDERS[provider]?.stallTimeoutMs || STREAM_STALL_TIMEOUT_MS;
const transformedBody = pipeWithDisconnect(providerResponse, transformStream, streamController, onAbortTerminal, stallTimeoutMs);
// Responses passthrough: synthesize response.failed + [DONE] if the stream aborts/stalls before a terminal event
const isResponsesPassthrough =
sourceFormat === FORMATS.OPENAI_RESPONSES &&
targetFormat === FORMATS.OPENAI_RESPONSES;
const onAbortTerminal = isResponsesPassthrough
? buildAbortedResponsesTerminalBytes
: null;
const stallTimeoutMs =
PROVIDERS[provider]?.stallTimeoutMs || STREAM_STALL_TIMEOUT_MS;
const transformedBody = pipeWithDisconnect(
providerResponse,
transformStream,
streamController,
onAbortTerminal,
stallTimeoutMs,
);
saveRequestDetail(buildRequestDetail({
provider, model, connectionId,
latency: { ttft: 0, total: Date.now() - requestStartTime },
tokens: { prompt_tokens: 0, completion_tokens: 0 },
request: extractRequestConfig(body, stream),
providerRequest: finalBody || translatedBody || null,
providerResponse: "[Streaming - raw response not captured]",
response: { content: "[Streaming in progress...]", thinking: null, type: "streaming" },
pxpipe,
status: "success"
}, { id: streamDetailId })).catch(err => {
console.error("[RequestDetail] Failed to save streaming request:", err.message);
});
saveRequestDetail(
buildRequestDetail(
{
provider,
model,
connectionId,
apiKey,
latency: { ttft: 0, total: Date.now() - requestStartTime },
tokens: { prompt_tokens: 0, completion_tokens: 0 },
request: extractRequestConfig(body, stream),
providerRequest: finalBody || translatedBody || null,
providerResponse: "[Streaming - raw response not captured]",
response: {
content: "[Streaming in progress...]",
thinking: null,
type: "streaming",
},
pxpipe,
status: "success",
},
{ id: streamDetailId },
),
).catch((err) => {
console.error(
"[RequestDetail] Failed to save streaming request:",
err.message,
);
});
return {
success: true,
response: new Response(transformedBody, { headers: SSE_HEADERS })
};
return {
success: true,
response: new Response(transformedBody, { headers: SSE_HEADERS }),
};
}
/**
* Build onStreamComplete callback for streaming usage tracking.
*/
export function buildOnStreamComplete({ provider, model, connectionId, apiKey, requestStartTime, body, stream, finalBody, translatedBody, clientRawRequest, pxpipe, reqTag, log }) {
const streamDetailId = `${Date.now()}-${Math.random().toString(36).slice(2, 11)}`;
export function buildOnStreamComplete({
provider,
model,
connectionId,
apiKey,
requestStartTime,
body,
stream,
finalBody,
translatedBody,
clientRawRequest,
pxpipe,
reqTag,
log,
streamErrorPatterns,
}) {
const streamDetailId = `${Date.now()}-${Math.random().toString(36).slice(2, 11)}`;
const onStreamComplete = (contentObj, usage, ttftAt) => {
const latency = {
ttft: ttftAt ? ttftAt - requestStartTime : Date.now() - requestStartTime,
total: Date.now() - requestStartTime
};
const safeContent = contentObj?.content || "[Empty streaming response]";
const safeThinking = contentObj?.thinking || null;
const onStreamComplete = (contentObj, usage, ttftAt) => {
const latency = {
ttft: ttftAt ? ttftAt - requestStartTime : Date.now() - requestStartTime,
total: Date.now() - requestStartTime,
};
const safeContent = contentObj?.content || "[Empty streaming response]";
const safeThinking = contentObj?.thinking || null;
saveRequestDetail(buildRequestDetail({
provider, model, connectionId,
latency,
tokens: usage || { prompt_tokens: 0, completion_tokens: 0 },
request: extractRequestConfig(body, stream),
providerRequest: finalBody || translatedBody || null,
providerResponse: safeContent,
response: { content: safeContent, thinking: safeThinking, type: "streaming" },
pxpipe,
status: "success"
}, { id: streamDetailId })).catch(err => {
console.error("[RequestDetail] Failed to update streaming content:", err.message);
});
saveRequestDetail(
buildRequestDetail(
{
provider,
model,
connectionId,
apiKey,
latency,
tokens: usage || { prompt_tokens: 0, completion_tokens: 0 },
request: extractRequestConfig(body, stream),
providerRequest: finalBody || translatedBody || null,
providerResponse: safeContent,
response: {
content: safeContent,
thinking: safeThinking,
type: "streaming",
},
pxpipe,
status: streamStatusForContent(
streamErrorPatterns?.[provider],
safeContent,
),
},
{ id: streamDetailId },
),
).catch((err) => {
console.error(
"[RequestDetail] Failed to update streaming content:",
err.message,
);
});
// Persist stream usage to DB (no console line; the "📊 done" line below is authoritative)
saveUsageStats({ provider, model, tokens: usage, connectionId, apiKey, endpoint: clientRawRequest?.endpoint, label: "STREAM USAGE", silent: true });
if (log?.line) log.line(reqTag, "📊", formatDoneLine({ usage, latency }));
};
// Persist stream usage to DB (no console line; the "📊 done" line below is authoritative)
saveUsageStats({
provider,
model,
tokens: usage,
connectionId,
apiKey,
endpoint: clientRawRequest?.endpoint,
label: "STREAM USAGE",
silent: true,
});
if (log?.line) log.line(reqTag, "📊", formatDoneLine({ usage, latency }));
};
return { onStreamComplete, streamDetailId };
return { onStreamComplete, streamDetailId };
}

View File

@@ -96,7 +96,7 @@ export async function handleImageGenerationCore({
let requestBody;
try {
url = adapter.buildUrl(model, credentials);
url = adapter.buildUrl(model, credentials, body);
requestBody = await adapter.buildBody(model, body);
headers = adapter.buildHeaders(credentials, requestBody, model, body);
} catch (error) {
@@ -140,7 +140,7 @@ export async function handleImageGenerationCore({
try {
const retryBody = await adapter.buildBody(model, body);
const retryHeaders = adapter.buildHeaders(credentials, retryBody, model, body);
const retryUrl = adapter.buildUrl(model, credentials);
const retryUrl = adapter.buildUrl(model, credentials, body);
providerResponse = await fetch(retryUrl, {
method: "POST",
headers: retryHeaders,

View File

@@ -12,6 +12,7 @@ import blackForestLabs from "./blackForestLabs.js";
import runwayml from "./runwayml.js";
import cloudflareAi from "./cloudflareAi.js";
import antigravity from "./antigravity.js";
import xai from "./xai.js";
const ADAPTERS = {
openai: createOpenAIAdapter("openai"),
@@ -19,7 +20,7 @@ const ADAPTERS = {
openrouter: createOpenAIAdapter("openrouter"),
recraft: createOpenAIAdapter("recraft"),
"vercel-ai-gateway": createOpenAIAdapter("vercel-ai-gateway"),
xai: createOpenAIAdapter("xai"),
xai,
gemini,
codex,
sdwebui,

View File

@@ -0,0 +1,137 @@
// xAI Grok Imagine — text-to-image + single/multi image editing
// Docs:
// https://docs.x.ai/developers/model-capabilities/images/generation
// https://docs.x.ai/developers/model-capabilities/images/editing
// https://docs.x.ai/developers/model-capabilities/images/multi-image-editing
import { sizeToAspectRatio } from "./_base.js";
import { PROVIDER_MEDIA } from "../../providers/index.js";
const IMG_CFG = PROVIDER_MEDIA["xai"]?.imageConfig || {};
const GENERATIONS_URL = IMG_CFG.baseUrl || "https://api.x.ai/v1/images/generations";
const EDITS_URL = IMG_CFG.editsUrl || "https://api.x.ai/v1/images/edits";
const ASPECT_RATIOS = new Set([
"auto",
"1:1",
"16:9",
"9:16",
"4:3",
"3:2",
"2:3",
"9:19.5",
"20:9",
]);
function hasEditInput(body) {
if (!body || typeof body !== "object") return false;
if (body.image) return true;
return Array.isArray(body.images) && body.images.some(Boolean);
}
/** Normalize client image input → xAI image ref object */
function toXaiImageRef(input) {
if (!input) return null;
if (typeof input === "object") {
// Already xAI-shaped or partial
if (input.file_id) {
return {
type: input.type || "image_url",
file_id: input.file_id,
...(input.url ? { url: input.url } : {}),
};
}
if (input.url) {
return { type: input.type || "image_url", url: input.url };
}
return null;
}
if (typeof input !== "string") return null;
const trimmed = input.trim();
if (!trimmed) return null;
// Public URL or data URI
if (/^https?:\/\//i.test(trimmed) || /^data:image\//i.test(trimmed)) {
return { type: "image_url", url: trimmed };
}
// Raw base64 → data URI
return { type: "image_url", url: `data:image/png;base64,${trimmed}` };
}
function collectImageRefs(body) {
const refs = [];
if (Array.isArray(body.images)) {
for (const item of body.images) {
const ref = toXaiImageRef(item);
if (ref) refs.push(ref);
}
}
if (body.image) {
const ref = toXaiImageRef(body.image);
if (ref) refs.push(ref);
}
// xAI multi-edit supports up to 3 source images
return refs.slice(0, 3);
}
function resolveAspectRatio(body) {
if (typeof body.aspect_ratio === "string" && body.aspect_ratio.trim()) {
const ratio = body.aspect_ratio.trim();
if (ASPECT_RATIOS.has(ratio)) return ratio;
// Pass through unknown ratio strings (upstream will validate)
return ratio;
}
// OpenAI-style size → aspect ratio (skip auto)
if (body.size && body.size !== "auto") {
return sizeToAspectRatio(body.size);
}
return undefined;
}
function resolveResolution(body) {
if (typeof body.resolution !== "string") return undefined;
const value = body.resolution.trim().toLowerCase();
if (!value || value === "auto") return undefined;
return value; // "1k" | "2k"
}
export default {
buildUrl: (_model, _credentials, body) => (hasEditInput(body) ? EDITS_URL : GENERATIONS_URL),
buildHeaders: (creds) => {
const headers = { "Content-Type": "application/json", ...(IMG_CFG.headers || {}) };
const key = creds?.apiKey || creds?.accessToken;
if (key) headers["Authorization"] = `Bearer ${key}`;
return headers;
},
buildBody: (model, body) => {
const req = {
model,
prompt: body.prompt,
};
if (body.n != null) req.n = body.n;
if (body.response_format) req.response_format = body.response_format;
const aspectRatio = resolveAspectRatio(body);
if (aspectRatio) req.aspect_ratio = aspectRatio;
const resolution = resolveResolution(body);
if (resolution) req.resolution = resolution;
const refs = collectImageRefs(body);
if (refs.length === 1) {
req.image = refs[0];
} else if (refs.length > 1) {
req.images = refs;
}
return req;
},
// xAI already returns OpenAI-compatible { created, data: [{ url | b64_json }] }
normalize: (responseBody) => responseBody,
};

View File

@@ -23,9 +23,24 @@ export default {
format: "commandcode",
forceStream: true,
headers: {
"x-command-code-version": "0.25.7",
"x-command-code-version": "1.10.0",
"x-cli-environment": "cli",
"User-Agent": "cli",
},
// Quota/billing endpoints (same alpha API the official CLI /usage calls).
// whoami resolves orgId; credits+subscription+usage/summary then report the
// 5-hour/weekly windows, plan, and period credits. See services/usage/commandcode.js.
usage: {
baseUrl: "https://api.commandcode.ai",
whoamiUrl: "/alpha/whoami",
creditsUrl: "/alpha/billing/credits",
subscriptionsUrl: "/alpha/billing/subscriptions",
summaryUrl: "/alpha/usage/summary",
},
},
features: {
usage: true,
usageApikey: true,
},
models: [
{ id: "deepseek/deepseek-v4-pro", name: "DeepSeek V4 Pro" },

View File

@@ -25,17 +25,48 @@ export default {
clientId: "b1a00492-073a-47ea-816f-4c329264a828",
tokenUrl: "https://auth.x.ai/oauth2/token",
refreshUrl: "https://auth.x.ai/oauth2/token",
// OAuth-only SuperGrok quota surfaces:
// - url: monthly API usage allotment (JSON)
// - creditsUrl: weekly SuperGrok limit (grpc-web)
// - settingsUrl: plan label (subscription_tier_display)
usage: {
url: "https://cli-chat-proxy.grok.com/v1/billing",
creditsUrl: "https://grok.com/grok_api_v2.GrokBuildBilling/GetGrokCreditsConfig",
settingsUrl: "https://cli-chat-proxy.grok.com/v1/settings",
},
},
models: [
{ id: "grok-4", name: "Grok 4" },
{ id: "grok-4-fast-reasoning", name: "Grok 4 Fast Reasoning" },
{ id: "grok-code-fast-1", name: "Grok Code Fast" },
{ id: "grok-3", name: "Grok 3" },
{ id: "grok-2-image-1212", name: "Grok 2 Image", params: ["n","response_format"], kind: "image" },
{ id: "grok-imagine-video", name: "Grok Imagine Video", params: ["duration","aspect_ratio","resolution"], kind: "video" },
{
id: "grok-imagine-image-quality",
name: "Grok Imagine Image Quality",
capabilities: ["text2img", "edit"],
params: ["n", "aspect_ratio", "resolution", "response_format", "size"],
kind: "image",
},
{
id: "grok-2-image-1212",
name: "Grok 2 Image",
capabilities: ["text2img", "edit"],
params: ["n", "aspect_ratio", "resolution", "response_format", "size"],
kind: "image",
},
{
id: "grok-imagine-video",
name: "Grok Imagine Video",
params: ["duration", "aspect_ratio", "resolution"],
kind: "video",
},
],
serviceKinds: ["llm","imageToText","webSearch","image","video"],
imageConfig: { baseUrl: "https://api.x.ai/v1/images/generations", bodyFields: ["model","prompt","n","response_format"] },
serviceKinds: ["llm", "imageToText", "webSearch", "image", "video"],
imageConfig: {
baseUrl: "https://api.x.ai/v1/images/generations",
editsUrl: "https://api.x.ai/v1/images/edits",
bodyFields: ["model", "prompt", "n", "response_format", "aspect_ratio", "resolution", "image", "images"],
},
// Async video jobs (POST returns { request_id }, GET polls until done/failed).
// Docs: https://docs.x.ai/developers/rest-api-reference/inference/videos
videoConfig: { baseUrl: "https://api.x.ai/v1/videos" },
@@ -44,4 +75,7 @@ export default {
endpoint: "https://api.x.ai/v1/responses",
pricingUrl: "https://x.ai/api#pricing",
},
features: {
usage: true,
},
};

View File

@@ -0,0 +1,56 @@
/**
* Per-provider connect timeout overrides from user settings.
* Settings are read from the DB lazily and cached with a short TTL
* so UI changes take effect without requiring a restart.
*/
let cached = {};
let cacheTs = 0;
const CACHE_TTL_MS = 10_000; // 10s — responsive enough for dashboard changes
async function refreshCache() {
const now = Date.now();
if (now - cacheTs < CACHE_TTL_MS && Object.keys(cached).length > 0) return cached;
try {
const { getSettings } = await import("@/lib/localDb");
// Return full settings so we can read providerTimeouts + globalTimeoutMs
cached = await getSettings();
cacheTs = now;
} catch {
// If DB is unavailable, keep stale cache — don't throw on hot path
}
return cached;
}
/**
* Resolve the effective connect timeout for a provider.
* Priority: per-provider override > global default timeout (settings) > registry config > env default.
* @param {string} providerId
* @param {number} configTimeoutMs - timeoutMs from the static provider registry config
* @param {number} envDefaultMs - global default from env (FETCH_CONNECT_TIMEOUT_MS)
* @returns {number} timeout in milliseconds
*/
export async function resolveProviderTimeoutMs(providerId, configTimeoutMs, envDefaultMs) {
const overrides = await refreshCache();
// 1. Per-provider override (set in provider detail page)
const providerOverride = overrides.providerTimeouts?.[providerId];
if (providerOverride?.timeoutMs && Number.isFinite(providerOverride.timeoutMs) && providerOverride.timeoutMs > 0) {
return providerOverride.timeoutMs;
}
// 2. Global default timeout (set in Profile / Settings page)
const globalDefault = overrides.defaultTimeoutMs;
if (globalDefault && Number.isFinite(globalDefault) && globalDefault > 0) {
return globalDefault;
}
// 3. Registry per-provider config
if (configTimeoutMs && Number.isFinite(configTimeoutMs) && configTimeoutMs > 0) {
return configTimeoutMs;
}
// 4. Env default
return envDefaultMs;
}

View File

@@ -11,9 +11,11 @@ export { consumeCodexRateLimitResetCredit, getCodexRateLimitResetCredits };
import { getKiroUsage } from "./usage/kiro.js";
import { getMiniMaxUsage } from "./usage/minimax.js";
import { getCodeBuddyCnUsage, getCodeBuddyIntlUsage } from "./usage/codebuddy-cn.js";
import { getXaiUsage } from "./usage/xai.js";
import { getGrokCliUsage } from "./usage/grok-cli.js";
import { getKimiUsage } from "./usage/kimi.js";
import { getDeepseekUsage } from "./usage/deepseek.js";
import { getCommandCodeUsage } from "./usage/commandcode.js";
import { resolveQoderCredentials } from "./qoderModels.js";
import {
getIflowUsage,
@@ -50,10 +52,12 @@ const USAGE_HANDLERS = {
"minimax-cn": (c) => getMiniMaxUsage(c.apiKey, c.provider, c.proxyOptions),
"vercel-ai-gateway": (c) => getVercelAiGatewayUsage(c.apiKey, c.proxyOptions),
"codebuddy-cn": (c) => getCodeBuddyCnUsage(c.accessToken, c.apiKey, c.providerSpecificData, c.proxyOptions),
xai: (c) => getXaiUsage(c.accessToken, c.proxyOptions),
"codebuddy-intl": (c) => getCodeBuddyIntlUsage(c.accessToken, c.apiKey, c.providerSpecificData, c.proxyOptions),
"grok-cli": (c) => getGrokCliUsage(c.accessToken, c.providerSpecificData, c.proxyOptions),
kimi: (c) => getKimiUsage(c.accessToken, c.apiKey, c.proxyOptions, c.providerSpecificData),
deepseek: (c) => getDeepseekUsage(c.apiKey, c.proxyOptions),
commandcode: (c) => getCommandCodeUsage(c.apiKey, c.proxyOptions),
};
export async function getUsageForProvider(connection, proxyOptions = null, options = {}) {

View File

@@ -0,0 +1,203 @@
/**
* CommandCode usage handler
*
* Mirrors the official command-code CLI /usage command: it calls the alpha API
* to surface the 5-hour + weekly usage windows, the subscription plan, and the
* credits consumed in the current billing period.
*
* GET /alpha/whoami → org.id (org-scoped billing; null for personal)
* GET /alpha/billing/credits → { credits: { monthlyCredits, purchasedCredits,
* freeCredits }, windowLimits: { fiveHour, weekly } }
* GET /alpha/billing/subscriptions → { data: { planId, currentPeriodStart, ... } }
* GET /alpha/usage/summary?since= → period token/cost totals
*
* The CLI fetches whoami first (for orgId), then credits + subscription in
* parallel, then the summary with since = currentPeriodStart. We keep the same
* order/dependencies: window limits live on credits, and the plan period start
* determines the summary window.
*/
import { proxyAwareFetch } from "../../utils/proxyFetch.js";
import { U, parseResetTime } from "./shared.js";
const USAGE = U("commandcode");
const BASE = USAGE.baseUrl || "https://api.commandcode.ai";
const WHOAMI_URL = BASE + (USAGE.whoamiUrl || "/alpha/whoami");
const CREDITS_URL = BASE + (USAGE.creditsUrl || "/alpha/billing/credits");
const SUBSCRIPTIONS_URL =
BASE + (USAGE.subscriptionsUrl || "/alpha/billing/subscriptions");
const SUMMARY_URL = BASE + (USAGE.summaryUrl || "/alpha/usage/summary");
function buildHeaders(token) {
return {
Authorization: `Bearer ${token}`,
Accept: "application/json",
};
}
/** Build a normalized quota row. `unit` is "$" — the API reports currency credits. */
function makeQuota({ used, total, resetAt, unlimited = false, unit = "$" }) {
const safeTotal = Math.max(0, Number(total) || 0);
const safeUsed = Math.max(0, Number(used) || 0);
if (unlimited || safeTotal === 0) {
return {
used: safeUsed,
total: 0,
remainingPercentage: unlimited ? 100 : 0,
resetAt: resetAt || null,
unit,
unlimited: true,
};
}
const remaining = Math.max(0, safeTotal - safeUsed);
const remainingPercentage = (remaining / safeTotal) * 100;
return {
used: safeUsed,
total: safeTotal,
remainingPercentage,
resetAt: resetAt || null,
unit,
unlimited: false,
};
}
/**
* @param {string} apiKey - commandcode API key (user_...)
* @param {object|null} proxyOptions
*/
export async function getCommandCodeUsage(apiKey, proxyOptions = null) {
if (!apiKey) {
return { message: "CommandCode credential not available." };
}
const headers = buildHeaders(apiKey);
try {
// whoami resolves the org id (billing is org-scoped; null for personal).
const whoamiRes = await proxyAwareFetch(
WHOAMI_URL,
{ method: "GET", headers },
proxyOptions,
);
if (whoamiRes.status === 401 || whoamiRes.status === 403) {
return { message: "CommandCode credential invalid or expired." };
}
if (!whoamiRes.ok) {
return { message: `CommandCode whoami API error (${whoamiRes.status}).` };
}
const whoami = await whoamiRes.json().catch(() => null);
const orgId = whoami?.org?.id ?? null;
const orgQuery = orgId ? `?orgId=${encodeURIComponent(orgId)}` : "";
const [creditsRes, subsRes] = await Promise.all([
proxyAwareFetch(
CREDITS_URL + orgQuery,
{ method: "GET", headers },
proxyOptions,
),
proxyAwareFetch(
SUBSCRIPTIONS_URL + orgQuery,
{ method: "GET", headers },
proxyOptions,
),
]);
if (
creditsRes.status === 401 ||
creditsRes.status === 403 ||
subsRes.status === 401 ||
subsRes.status === 403
) {
return { message: "CommandCode credential invalid or expired." };
}
if (!creditsRes.ok) {
return {
message: `CommandCode credits API error (${creditsRes.status}).`,
};
}
const credits = await creditsRes.json().catch(() => null);
const subs = await subsRes.json().catch(() => null);
const subData = subs?.data;
const planId = subData?.planId ?? null;
const periodStart = subData?.currentPeriodStart ?? null;
// Summary needs `since`; the CLI falls back to first-of-month when the
// subscription period start is unavailable.
const since = periodStart || firstOfMonth();
const summaryRes = await proxyAwareFetch(
`${SUMMARY_URL}?since=${encodeURIComponent(since)}`,
{ method: "GET", headers },
proxyOptions,
);
const summary = summaryRes.ok
? await summaryRes.json().catch(() => null)
: null;
const quotas = {};
const windowLimits = credits?.windowLimits || {};
const fiveHour = windowLimits.fiveHour;
if (fiveHour && Number(fiveHour.cap) > 0) {
quotas["5-hour window"] = makeQuota({
used: fiveHour.used,
total: fiveHour.cap,
resetAt: parseResetTime(fiveHour.resetAt),
});
}
const weekly = windowLimits.weekly;
if (weekly && Number(weekly.cap) > 0) {
quotas["Weekly window"] = makeQuota({
used: weekly.used,
total: weekly.cap,
resetAt: parseResetTime(weekly.resetAt),
});
}
// Monthly credits consumed this billing period (from summary when present,
// else the credits object's monthlyCredits as a fallback).
const monthlyUsed =
typeof summary?.totalCredits === "number"
? summary.totalCredits
: typeof credits?.credits?.monthlyCredits === "number"
? credits.credits.monthlyCredits
: 0;
const monthlyTotal =
typeof credits?.credits?.monthlyCredits === "number"
? credits.credits.monthlyCredits
: 0;
if (monthlyTotal > 0 || monthlyUsed > 0) {
quotas["Monthly credits"] = makeQuota({
used: monthlyUsed,
total: monthlyTotal,
resetAt: periodStart ? undefined : null,
});
}
if (Object.keys(quotas).length === 0) {
return {
plan: planId || "CommandCode",
message: "CommandCode connected, but no quota was reported.",
quotas: {},
};
}
return {
plan: planId || "CommandCode",
quotas,
periodBasis: summary?.periodBasis || "billing-period",
};
} catch (error) {
return { message: `CommandCode usage error: ${error.message}` };
}
}
function firstOfMonth() {
const now = new Date();
return new Date(now.getFullYear(), now.getMonth(), 1).toISOString();
}

View File

@@ -0,0 +1,326 @@
/**
* xAI (Grok) OAuth usage handler
*
* SuperGrok quota is split across two upstream surfaces (OAuth only):
*
* 1) Monthly / API usage allotment (JSON)
* GET https://cli-chat-proxy.grok.com/v1/billing
* {
* "config": {
* "monthlyLimit": { "val": 15000 },
* "used": { "val": 733 },
* "onDemandCap": { "val": 0 },
* "billingPeriodStart": "...",
* "billingPeriodEnd": "..."
* }
* }
*
* 2) Weekly SuperGrok limit (grpc-web protobuf)
* POST https://grok.com/grok_api_v2.GrokBuildBilling/GetGrokCreditsConfig
* Empty request frame; response message1 contains:
* usedPercent (float32), window start/end timestamps, nested windows.
* This is what the grok.com usage page labels "Weekly limit" / "Resets …".
*
* Plan label comes from cli-chat-proxy settings:
* GET https://cli-chat-proxy.grok.com/v1/settings → subscription_tier_display
*
* Note: grok.com/rest/rate-limits is a short chat-window (e.g. 2h query count)
* behind Cloudflare browser cookies — not usable with pure OAuth bearer.
*/
import { proxyAwareFetch } from "../../utils/proxyFetch.js";
import { U, parseResetTime, toFiniteNumber } from "./shared.js";
// Empty grpc-web request frame: flag(0) + length(0) + no payload.
const GRPC_WEB_EMPTY_FRAME = Buffer.from([0x00, 0x00, 0x00, 0x00, 0x00]);
function moneyVal(wrapper) {
if (wrapper == null) return null;
if (typeof wrapper === "number") return toFiniteNumber(wrapper, null);
if (typeof wrapper === "object" && wrapper.val != null) {
return toFiniteNumber(wrapper.val, null);
}
return null;
}
function authHeaders(accessToken, extra = {}) {
return {
Authorization: `Bearer ${accessToken}`,
...extra,
};
}
function readVarint(buf, offset) {
let val = 0;
let shift = 0;
let i = offset;
while (i < buf.length) {
const b = buf[i++];
val |= (b & 0x7f) << shift;
if ((b & 0x80) === 0) return { value: val >>> 0, offset: i };
shift += 7;
if (shift > 35) break;
}
return null;
}
/**
* Minimal protobuf decoder for GetGrokCreditsConfig.
* Only understands varint / fixed32 / fixed64 / length-delimited.
*/
function parseProtobufFields(buf) {
const fields = [];
let i = 0;
while (i < buf.length) {
const key = readVarint(buf, i);
if (!key) break;
i = key.offset;
const field = key.value >>> 3;
const wt = key.value & 7;
if (wt === 0) {
const v = readVarint(buf, i);
if (!v) break;
i = v.offset;
fields.push({ field, type: "varint", value: v.value });
} else if (wt === 1) {
if (i + 8 > buf.length) break;
fields.push({ field, type: "fixed64", value: buf.subarray(i, i + 8) });
i += 8;
} else if (wt === 5) {
if (i + 4 > buf.length) break;
fields.push({
field,
type: "fixed32",
value: buf.readFloatLE(i),
});
i += 4;
} else if (wt === 2) {
const ln = readVarint(buf, i);
if (!ln) break;
i = ln.offset;
if (i + ln.value > buf.length) break;
fields.push({
field,
type: "bytes",
value: buf.subarray(i, i + ln.value),
});
i += ln.value;
} else {
break;
}
}
return fields;
}
function parseTimestamp(bytes) {
if (!bytes || !bytes.length) return null;
const fields = parseProtobufFields(bytes);
const seconds = fields.find((f) => f.field === 1 && f.type === "varint")?.value;
if (!Number.isFinite(seconds) || seconds <= 0) return null;
return new Date(seconds * 1000).toISOString();
}
/**
* Parse grpc-web response bytes from GetGrokCreditsConfig.
* Returns { usedPercent, resetAt, periodStart } or null.
*/
export function parseGrokCreditsConfig(raw) {
if (!raw || !raw.length) return null;
const buf = Buffer.isBuffer(raw) ? raw : Buffer.from(raw);
// grpc-web data frame: 1-byte flag + 4-byte big-endian length + message
if (buf.length < 5) return null;
const flag = buf[0];
// Data frames have flag 0; ignore trailer frames (flag 0x80).
if (flag !== 0) return null;
const msgLen = buf.readUInt32BE(1);
if (msgLen <= 0 || 5 + msgLen > buf.length) return null;
const msg = buf.subarray(5, 5 + msgLen);
// Response is typically { 1: CreditsConfig }
const top = parseProtobufFields(msg);
const configBytes = top.find((f) => f.field === 1 && f.type === "bytes")?.value || msg;
const fields = parseProtobufFields(configBytes);
const usedPercentRaw = fields.find((f) => f.field === 1 && f.type === "fixed32")?.value;
const periodStart = parseTimestamp(fields.find((f) => f.field === 4 && f.type === "bytes")?.value);
const periodEnd = parseTimestamp(fields.find((f) => f.field === 5 && f.type === "bytes")?.value);
if (usedPercentRaw == null || !Number.isFinite(usedPercentRaw)) return null;
const usedPercent = Math.max(0, Math.min(100, usedPercentRaw));
return {
usedPercent,
remainingPercent: Math.max(0, 100 - usedPercent),
periodStart,
resetAt: periodEnd,
};
}
async function fetchBilling(accessToken, billingUrl, proxyOptions) {
const response = await proxyAwareFetch(billingUrl, {
method: "GET",
headers: authHeaders(accessToken, { Accept: "application/json" }),
}, proxyOptions);
if (response.status === 401 || response.status === 403) {
return { error: "auth", status: response.status };
}
if (!response.ok) {
return { error: "http", status: response.status };
}
const data = await response.json().catch(() => null);
if (!data || typeof data !== "object") {
return { error: "json" };
}
return { data };
}
async function fetchWeeklyCredits(accessToken, creditsUrl, proxyOptions) {
if (!creditsUrl) return null;
try {
const response = await proxyAwareFetch(creditsUrl, {
method: "POST",
headers: authHeaders(accessToken, {
"Content-Type": "application/grpc-web+proto",
"x-grpc-web": "1",
"x-user-agent": "connect-es/2.1.1",
Accept: "*/*",
Origin: "https://grok.com",
Referer: "https://grok.com/?_s=usage",
}),
body: GRPC_WEB_EMPTY_FRAME,
}, proxyOptions);
if (!response.ok) return null;
const ab = await response.arrayBuffer();
return parseGrokCreditsConfig(Buffer.from(ab));
} catch {
return null;
}
}
async function fetchPlanLabel(accessToken, settingsUrl, proxyOptions) {
if (!settingsUrl) return null;
try {
const response = await proxyAwareFetch(settingsUrl, {
method: "GET",
headers: authHeaders(accessToken, { Accept: "application/json" }),
}, proxyOptions);
if (!response.ok) return null;
const data = await response.json().catch(() => null);
const label = data?.subscription_tier_display;
return typeof label === "string" && label.trim() ? label.trim() : null;
} catch {
return null;
}
}
/**
* @param {string} accessToken - xAI OAuth access token
* @param {object|null} proxyOptions
*/
export async function getXaiUsage(accessToken, proxyOptions = null) {
if (!accessToken) {
return { message: "xAI usage unavailable: no access token. Re-authorize the connection." };
}
const cfg = U("xai") || {};
const billingUrl = cfg.url;
if (!billingUrl) {
return { message: "xAI usage endpoint is not configured." };
}
try {
const [billingResult, weekly, planLabel] = await Promise.all([
fetchBilling(accessToken, billingUrl, proxyOptions),
fetchWeeklyCredits(accessToken, cfg.creditsUrl, proxyOptions),
fetchPlanLabel(accessToken, cfg.settingsUrl, proxyOptions),
]);
if (billingResult.error === "auth") {
return { message: "xAI OAuth token expired or unauthorized. Please re-authorize." };
}
const quotas = {};
let periodStart = null;
let periodEnd = null;
let onDemandCap = 0;
if (billingResult.data) {
const config =
billingResult.data.config && typeof billingResult.data.config === "object"
? billingResult.data.config
: billingResult.data;
const monthlyLimit = moneyVal(config.monthlyLimit);
const used = moneyVal(config.used);
onDemandCap = moneyVal(config.onDemandCap) ?? 0;
periodEnd = parseResetTime(config.billingPeriodEnd);
periodStart = parseResetTime(config.billingPeriodStart);
// Absolute credit counts — do NOT put remaining credits on `remaining`
// (QuotaTable treats remaining as a 0-100 percentage; same pitfall as Qoder).
if (monthlyLimit != null && monthlyLimit > 0) {
const usedSafe = Math.max(0, used ?? 0);
quotas.api_usage = {
used: usedSafe,
total: monthlyLimit,
remainingCredits: Math.max(0, monthlyLimit - usedSafe),
unit: "credits",
resetAt: periodEnd,
unlimited: false,
};
}
if (onDemandCap > 0) {
quotas.on_demand = {
used: 0,
total: onDemandCap,
remainingCredits: onDemandCap,
unit: "credits",
resetAt: periodEnd,
unlimited: false,
};
}
}
// Weekly SuperGrok window — percentage-based like Claude/Codex windows.
if (weekly) {
quotas.weekly = {
used: weekly.usedPercent,
total: 100,
remaining: weekly.remainingPercent,
remainingPercentage: weekly.remainingPercent,
resetAt: weekly.resetAt || null,
unlimited: false,
};
if (!periodStart && weekly.periodStart) periodStart = weekly.periodStart;
}
if (Object.keys(quotas).length === 0) {
const statusHint =
billingResult.error === "http"
? ` Billing API temporarily unavailable (${billingResult.status}).`
: "";
return {
plan: planLabel || "xAI",
message: `xAI connected. No quota allotment reported for this account.${statusHint}`,
periodStart,
periodEnd,
quotas: {},
};
}
return {
plan: planLabel || "xAI",
periodStart,
periodEnd,
onDemandCap,
quotas,
};
} catch (error) {
return { message: `xAI connected. Unable to fetch billing: ${error.message}` };
}
}

View File

@@ -252,33 +252,6 @@ export function anchorClaudeCache(body) {
}
}
// 3. Drop thinking blocks whose signature is not Claude's (combo mixes models,
// so foreign signatures leak into history and Anthropic rejects them).
const thinkingEnabled = body.thinking?.type === "enabled";
if (Array.isArray(body.messages)) {
for (const msg of body.messages) {
if (msg.role !== ROLE.ASSISTANT || !Array.isArray(msg.content)) continue;
let hasToolUse = false;
let hasKeptThinking = false;
const kept = [];
for (const block of msg.content) {
if (block.type === CLAUDE_BLOCK.THINKING || block.type === CLAUDE_BLOCK.REDACTED_THINKING) {
if (isValidClaudeSignature(block.signature)) {
hasKeptThinking = true;
kept.push(block);
}
continue;
}
if (block.type === CLAUDE_BLOCK.TOOL_USE) hasToolUse = true;
kept.push(block);
}
msg.content = kept;
if (thinkingEnabled && !hasKeptThinking && hasToolUse) {
msg.content.unshift(buildThinkingPlaceholder("claude"));
}
}
}
return body;
}

View File

@@ -13,160 +13,204 @@ import { register } from "../index.js";
import { FORMATS } from "../formats.js";
import { randomUUID } from "crypto";
import { ROLE, OPENAI_BLOCK } from "../schema/index.js";
import { DEFAULT_IMAGE_MIME } from "../schema/index.js";
import { parseDataUri } from "../concerns/image.js";
import { DEFAULT_MAX_TOKENS } from "../../config/runtimeConfig.js";
function flattenText(content) {
if (content == null) return "";
if (typeof content === "string") return content;
if (Array.isArray(content)) {
const parts = [];
for (const p of content) {
if (typeof p === "string") parts.push(p);
else if (p && typeof p === "object" && typeof p.text === "string") parts.push(p.text);
}
return parts.join("\n");
}
return String(content);
if (content == null) return "";
if (typeof content === "string") return content;
if (Array.isArray(content)) {
const parts = [];
for (const p of content) {
if (typeof p === "string") parts.push(p);
else if (p && typeof p === "object" && typeof p.text === "string")
parts.push(p.text);
}
return parts.join("\n");
}
return String(content);
}
function toContentBlocks(content) {
if (content == null) return [{ type: OPENAI_BLOCK.TEXT, text: "" }];
if (typeof content === "string") return [{ type: OPENAI_BLOCK.TEXT, text: content }];
if (Array.isArray(content)) {
const blocks = [];
for (const part of content) {
if (typeof part === "string") {
blocks.push({ type: OPENAI_BLOCK.TEXT, text: part });
} else if (part && typeof part === "object") {
if (part.type === OPENAI_BLOCK.TEXT && typeof part.text === "string") {
blocks.push({ type: OPENAI_BLOCK.TEXT, text: part.text });
} else if (part.type === OPENAI_BLOCK.IMAGE_URL || part.type === OPENAI_BLOCK.IMAGE) {
blocks.push({ type: OPENAI_BLOCK.TEXT, text: "[image omitted]" });
} else if (typeof part.text === "string") {
blocks.push({ type: OPENAI_BLOCK.TEXT, text: part.text });
}
}
}
return blocks.length ? blocks : [{ type: OPENAI_BLOCK.TEXT, text: "" }];
}
return [{ type: OPENAI_BLOCK.TEXT, text: String(content) }];
if (content == null) return [{ type: OPENAI_BLOCK.TEXT, text: "" }];
if (typeof content === "string")
return [{ type: OPENAI_BLOCK.TEXT, text: content }];
if (Array.isArray(content)) {
const blocks = [];
for (const part of content) {
if (typeof part === "string") {
blocks.push({ type: OPENAI_BLOCK.TEXT, text: part });
} else if (part && typeof part === "object") {
if (part.type === OPENAI_BLOCK.TEXT && typeof part.text === "string") {
blocks.push({ type: OPENAI_BLOCK.TEXT, text: part.text });
} else if (
part.type === OPENAI_BLOCK.IMAGE_URL ||
part.type === OPENAI_BLOCK.IMAGE
) {
// CommandCode `/alpha/generate` accepts {type:"image", image:"<data URI | url>"} —
// same shape the official command-code CLI sends (verified from CLI source).
const src = part.source;
let raw = part.image_url?.url || src?.data || src?.url || "";
let parsed = parseDataUri(raw);
if (!parsed && src?.type === "base64" && src?.data) {
// Claude-style base64 source without a data-URI prefix → wrap it.
raw = `data:${src.media_type || DEFAULT_IMAGE_MIME};base64,${src.data}`;
parsed = parseDataUri(raw);
}
if (parsed) {
blocks.push({
type: "image",
image: `data:${parsed.mimeType};base64,${parsed.base64}`,
});
} else if (raw) {
blocks.push({ type: "image", image: raw });
}
} else if (typeof part.text === "string") {
blocks.push({ type: OPENAI_BLOCK.TEXT, text: part.text });
}
}
}
return blocks.length ? blocks : [{ type: OPENAI_BLOCK.TEXT, text: "" }];
}
return [{ type: OPENAI_BLOCK.TEXT, text: String(content) }];
}
function safeParseJson(s) {
if (s == null) return {};
if (typeof s !== "string") return s;
try { return JSON.parse(s); } catch { return {}; }
if (s == null) return {};
if (typeof s !== "string") return s;
try {
return JSON.parse(s);
} catch {
return {};
}
}
function convertMessages(messages = []) {
const out = [];
const systemTexts = [];
const out = [];
const systemTexts = [];
for (const m of messages) {
if (!m) continue;
const role = m.role;
for (const m of messages) {
if (!m) continue;
const role = m.role;
if (role === ROLE.SYSTEM) {
const t = flattenText(m.content);
if (t) systemTexts.push(t);
continue;
}
if (role === ROLE.SYSTEM) {
const t = flattenText(m.content);
if (t) systemTexts.push(t);
continue;
}
if (role === ROLE.TOOL) {
const value = typeof m.content === "string" ? m.content : flattenText(m.content);
out.push({
role: ROLE.TOOL,
content: [{
type: "tool-result",
toolCallId: m.tool_call_id || "",
toolName: m.name || "",
output: { type: "text", value },
}],
});
continue;
}
if (role === ROLE.TOOL) {
const value =
typeof m.content === "string" ? m.content : flattenText(m.content);
out.push({
role: ROLE.TOOL,
content: [
{
type: "tool-result",
toolCallId: m.tool_call_id || "",
toolName: m.name || "",
output: { type: "text", value },
},
],
});
continue;
}
if (role === ROLE.ASSISTANT) {
const blocks = [];
const text = flattenText(m.content);
if (text) blocks.push({ type: OPENAI_BLOCK.TEXT, text });
if (Array.isArray(m.tool_calls)) {
for (const tc of m.tool_calls) {
const fn = tc.function || {};
blocks.push({
type: "tool-call",
toolCallId: tc.id || "",
toolName: fn.name || "",
input: safeParseJson(fn.arguments),
});
}
}
out.push({ role: ROLE.ASSISTANT, content: blocks.length ? blocks : [{ type: OPENAI_BLOCK.TEXT, text: "" }] });
continue;
}
if (role === ROLE.ASSISTANT) {
const blocks = [];
const text = flattenText(m.content);
if (text) blocks.push({ type: OPENAI_BLOCK.TEXT, text });
if (Array.isArray(m.tool_calls)) {
for (const tc of m.tool_calls) {
const fn = tc.function || {};
blocks.push({
type: "tool-call",
toolCallId: tc.id || "",
toolName: fn.name || "",
input: safeParseJson(fn.arguments),
});
}
}
out.push({
role: ROLE.ASSISTANT,
content: blocks.length
? blocks
: [{ type: OPENAI_BLOCK.TEXT, text: "" }],
});
continue;
}
out.push({ role: ROLE.USER, content: toContentBlocks(m.content) });
}
out.push({ role: ROLE.USER, content: toContentBlocks(m.content) });
}
return { messages: out, system: systemTexts.join("\n\n") };
return { messages: out, system: systemTexts.join("\n\n") };
}
function convertTools(tools) {
if (!Array.isArray(tools) || tools.length === 0) return undefined;
const result = [];
for (const t of tools) {
if (!t) continue;
if (t.type === OPENAI_BLOCK.FUNCTION && t.function) {
result.push({
name: t.function.name,
description: t.function.description,
input_schema: t.function.parameters || { type: "object" },
});
} else if (t.name && (t.input_schema || t.parameters)) {
result.push({
name: t.name,
description: t.description,
input_schema: t.input_schema || t.parameters,
});
}
}
return result.length ? result : undefined;
if (!Array.isArray(tools) || tools.length === 0) return undefined;
const result = [];
for (const t of tools) {
if (!t) continue;
if (t.type === OPENAI_BLOCK.FUNCTION && t.function) {
result.push({
name: t.function.name,
description: t.function.description,
input_schema: t.function.parameters || { type: "object" },
});
} else if (t.name && (t.input_schema || t.parameters)) {
result.push({
name: t.name,
description: t.description,
input_schema: t.input_schema || t.parameters,
});
}
}
return result.length ? result : undefined;
}
export function openaiToCommandCodeRequest(model, body, stream /* , credentials */) {
const { messages, system } = convertMessages(body.messages);
const params = {
model,
messages,
stream: stream !== false,
max_tokens: body.max_tokens ?? body.max_output_tokens ?? DEFAULT_MAX_TOKENS,
temperature: body.temperature ?? 0.3,
};
export function openaiToCommandCodeRequest(
model,
body,
stream /* , credentials */,
) {
const { messages, system } = convertMessages(body.messages);
const params = {
model,
messages,
stream: stream !== false,
max_tokens: body.max_tokens ?? body.max_output_tokens ?? DEFAULT_MAX_TOKENS,
temperature: body.temperature ?? 0.3,
};
if (system) params.system = system;
if (system) params.system = system;
const tools = convertTools(body.tools);
if (tools) params.tools = tools;
if (body.top_p != null) params.top_p = body.top_p;
const tools = convertTools(body.tools);
if (tools) params.tools = tools;
if (body.top_p != null) params.top_p = body.top_p;
const today = new Date().toISOString().slice(0, 10);
const today = new Date().toISOString().slice(0, 10);
return {
threadId: randomUUID(),
memory: "",
config: {
workingDir: process.cwd(),
date: today,
environment: process.platform,
structure: [],
isGitRepo: false,
currentBranch: "",
mainBranch: "",
gitStatus: "",
recentCommits: [],
},
params,
};
// environment format mirrors the official command-code CLI (getEnvironmentInfo):
// `${platform}-${arch}, Node.js ${version}` (e.g. "darwin-arm64, Node.js v24.16.0").
const environment = `${process.platform}-${process.arch}, Node.js ${process.version}`;
return {
threadId: randomUUID(),
memory: "",
config: {
workingDir: process.cwd(),
date: today,
environment,
structure: [],
isGitRepo: false,
currentBranch: "",
mainBranch: "",
gitStatus: "",
recentCommits: [],
},
params,
};
}
register(FORMATS.OPENAI, FORMATS.COMMANDCODE, openaiToCommandCodeRequest, null);

View File

@@ -25,160 +25,198 @@ import { fallbackToolCallId } from "../concerns/toolCall.js";
import { toOpenAIFinish } from "../concerns/finishReason.js";
function ensureState(state, model) {
if (!state.responseId) {
state.responseId = `chatcmpl-${Date.now()}`;
state.created = Math.floor(Date.now() / 1000);
state.model = state.model || model || "commandcode";
state.chunkIndex = 0;
state.toolIndex = 0;
state.toolIndexById = new Map();
state.openTools = new Set();
state.openText = false;
state.finishReason = null;
state.usage = null;
}
if (!state.responseId) {
state.responseId = `chatcmpl-${Date.now()}`;
state.created = Math.floor(Date.now() / 1000);
state.model = state.model || model || "commandcode";
state.chunkIndex = 0;
state.toolIndex = 0;
state.toolIndexById = new Map();
state.openTools = new Set();
state.openText = false;
state.finishReason = null;
state.usage = null;
}
}
function makeChunk(state, delta, finishReason = null) {
return buildChunk(
{ id: state.responseId, created: state.created, model: state.model },
delta,
finishReason
);
return buildChunk(
{ id: state.responseId, created: state.created, model: state.model },
delta,
finishReason,
);
}
const mapFinishReason = (reason) => toOpenAIFinish(reason, "commandcode");
export function commandCodeToOpenAIResponse(chunk, state) {
if (!chunk) return null;
if (!chunk) return null;
// Already-OpenAI chunk: pass through
if (chunk && typeof chunk === "object" && chunk.object === "chat.completion.chunk") {
return chunk;
}
// Already-OpenAI chunk: pass through
if (
chunk &&
typeof chunk === "object" &&
chunk.object === "chat.completion.chunk"
) {
return chunk;
}
// Parse string lines coming out of upstream
let event = chunk;
if (typeof chunk === "string") {
const line = chunk.trim();
if (!line) return null;
// Tolerate raw "data: {...}" framing if the upstream wrapper inserts it
const json = line.startsWith("data:") ? line.slice(5).trim() : line;
if (!json || json === "[DONE]") return null;
try {
event = JSON.parse(json);
} catch {
return null;
}
}
// Parse string lines coming out of upstream
let event = chunk;
if (typeof chunk === "string") {
const line = chunk.trim();
if (!line) return null;
// Tolerate raw "data: {...}" framing if the upstream wrapper inserts it
const json = line.startsWith("data:") ? line.slice(5).trim() : line;
if (!json || json === "[DONE]") return null;
try {
event = JSON.parse(json);
} catch {
return null;
}
}
if (!event || typeof event !== "object" || !event.type) return null;
if (!event || typeof event !== "object" || !event.type) return null;
ensureState(state, event.model);
const out = [];
ensureState(state, event.model);
const out = [];
switch (event.type) {
case "text-delta": {
const text = event.text || event.delta || "";
if (!text) break;
const delta = state.chunkIndex === 0 ? { role: ROLE.ASSISTANT, content: text } : { content: text };
state.chunkIndex++;
state.openText = true;
out.push(makeChunk(state, delta));
break;
}
case "reasoning-delta": {
const text = event.text || "";
if (!text) break;
// Map reasoning to OpenAI "reasoning_content" field (used by deepseek-reasoner-style clients).
const delta = reasoningDelta(text, state.chunkIndex === 0);
state.chunkIndex++;
out.push(makeChunk(state, delta));
break;
}
case "tool-input-start": {
const id = event.id || event.toolCallId || fallbackToolCallId(state.toolIndex);
let idx = state.toolIndexById.get(id);
if (idx == null) {
idx = state.toolIndex++;
state.toolIndexById.set(id, idx);
}
state.openTools.add(id);
const delta = {
...(state.chunkIndex === 0 ? { role: ROLE.ASSISTANT } : {}),
tool_calls: [{
index: idx,
id,
type: OPENAI_BLOCK.FUNCTION,
function: { name: event.toolName || "", arguments: "" },
}],
};
state.chunkIndex++;
out.push(makeChunk(state, delta));
break;
}
case "tool-input-delta": {
const id = event.id || event.toolCallId;
const idx = state.toolIndexById.get(id);
if (idx == null) break;
const delta = {
tool_calls: [{
index: idx,
function: { arguments: event.delta || event.inputTextDelta || "" },
}],
};
out.push(makeChunk(state, delta));
break;
}
case "tool-call": {
// Final consolidated tool call — only emit if we never saw tool-input-* deltas.
const id = event.toolCallId;
if (state.toolIndexById.has(id)) break;
const idx = state.toolIndex++;
state.toolIndexById.set(id, idx);
const argsStr = typeof event.input === "string" ? event.input : JSON.stringify(event.input ?? {});
const delta = {
...(state.chunkIndex === 0 ? { role: ROLE.ASSISTANT } : {}),
tool_calls: [{
index: idx,
id,
type: OPENAI_BLOCK.FUNCTION,
function: { name: event.toolName || "", arguments: argsStr },
}],
};
state.chunkIndex++;
out.push(makeChunk(state, delta));
break;
}
case "finish-step": {
state.finishReason = mapFinishReason(event.finishReason);
if (event.usage) state.usage = event.usage;
break;
}
case "finish": {
const finishReason = state.finishReason || mapFinishReason(event.finishReason || "stop");
const finalChunk = makeChunk(state, {}, finishReason);
const totalUsage = event.totalUsage || state.usage;
const usage = toOpenAIUsage(totalUsage, "commandcode");
if (usage) finalChunk.usage = usage;
out.push(finalChunk);
break;
}
case "error": {
state.finishReason = OPENAI_FINISH.STOP;
const errVal = event.error ?? event.message ?? "unknown";
const errStr = typeof errVal === "string" ? errVal : JSON.stringify(errVal);
out.push(makeChunk(state, { content: `\n\n[CommandCode error: ${errStr}]` }));
out.push(makeChunk(state, {}, OPENAI_FINISH.STOP));
break;
}
// Silently ignore: start, start-step, reasoning-start, reasoning-end, text-start, text-end,
// provider-metadata, message-metadata, etc. They carry no client-visible content.
default:
break;
}
switch (event.type) {
case "text-delta": {
const text = event.text || event.delta || "";
if (!text) break;
const delta =
state.chunkIndex === 0
? { role: ROLE.ASSISTANT, content: text }
: { content: text };
state.chunkIndex++;
state.openText = true;
out.push(makeChunk(state, delta));
break;
}
case "reasoning-delta": {
const text = event.text || "";
if (!text) break;
// Map reasoning to OpenAI "reasoning_content" field (used by deepseek-reasoner-style clients).
const delta = reasoningDelta(text, state.chunkIndex === 0);
state.chunkIndex++;
out.push(makeChunk(state, delta));
break;
}
case "tool-input-start": {
const id =
event.id || event.toolCallId || fallbackToolCallId(state.toolIndex);
let idx = state.toolIndexById.get(id);
if (idx == null) {
idx = state.toolIndex++;
state.toolIndexById.set(id, idx);
}
state.openTools.add(id);
const delta = {
...(state.chunkIndex === 0 ? { role: ROLE.ASSISTANT } : {}),
tool_calls: [
{
index: idx,
id,
type: OPENAI_BLOCK.FUNCTION,
function: { name: event.toolName || "", arguments: "" },
},
],
};
state.chunkIndex++;
out.push(makeChunk(state, delta));
break;
}
case "tool-input-delta": {
const id = event.id || event.toolCallId;
const idx = state.toolIndexById.get(id);
if (idx == null) break;
const delta = {
tool_calls: [
{
index: idx,
function: { arguments: event.delta || event.inputTextDelta || "" },
},
],
};
out.push(makeChunk(state, delta));
break;
}
case "tool-call": {
// Final consolidated tool call — only emit if we never saw tool-input-* deltas.
const id = event.toolCallId;
if (state.toolIndexById.has(id)) break;
const idx = state.toolIndex++;
state.toolIndexById.set(id, idx);
const argsStr =
typeof event.input === "string"
? event.input
: JSON.stringify(event.input ?? {});
const delta = {
...(state.chunkIndex === 0 ? { role: ROLE.ASSISTANT } : {}),
tool_calls: [
{
index: idx,
id,
type: OPENAI_BLOCK.FUNCTION,
function: { name: event.toolName || "", arguments: argsStr },
},
],
};
state.chunkIndex++;
out.push(makeChunk(state, delta));
break;
}
case "finish-step": {
state.finishReason = mapFinishReason(event.finishReason);
if (event.usage) state.usage = event.usage;
break;
}
case "finish": {
const finishReason =
state.finishReason || mapFinishReason(event.finishReason || "stop");
const finalChunk = makeChunk(state, {}, finishReason);
const totalUsage = event.totalUsage || state.usage;
const usage = toOpenAIUsage(totalUsage, "commandcode");
if (usage) finalChunk.usage = usage;
out.push(finalChunk);
break;
}
case "error": {
// Terminal upstream failure (AI SDK v5 error event) — NOT content. Emit an
// OpenAI-shaped error chunk (chunk.error) so downstream — parseSSEToOpenAIResponse
// for non-streaming, OpenAI SDK clients for streaming — treats the request as
// failed instead of surfacing fake success content like "[CommandCode error: ...]".
state.finishReason = OPENAI_FINISH.STOP;
const errVal = event.error ?? event.message ?? "unknown";
const errStr =
typeof errVal === "string"
? errVal
: typeof errVal?.message === "string"
? errVal.message
: JSON.stringify(errVal);
const errType =
typeof errVal === "string"
? "upstream_error"
: errVal?.type || "upstream_error";
const errChunk = makeChunk(state, {});
errChunk.error = { message: errStr, type: errType };
out.push(errChunk);
out.push(makeChunk(state, {}, OPENAI_FINISH.STOP));
break;
}
// Silently ignore: start, start-step, reasoning-start, reasoning-end, text-start, text-end,
// provider-metadata, message-metadata, etc. They carry no client-visible content.
default:
break;
}
return out.length ? out : null;
return out.length ? out : null;
}
register(FORMATS.COMMANDCODE, FORMATS.OPENAI, null, commandCodeToOpenAIResponse);
register(
FORMATS.COMMANDCODE,
FORMATS.OPENAI,
null,
commandCodeToOpenAIResponse,
);

View File

@@ -6,95 +6,12 @@ const originalFetch = globalThis.fetch;
const proxyDispatchers = new Map();
// ─── TLS fingerprinting via got-scraping (browser-like JA3) ───────────────
// Disabled: not in use. Kept commented for future re-enable.
// Restore the original block to re-enable per-host JA3 spoofing.
// Disabled: not in use.
/*
let _gotScraping = null;
let _gotScrapingChecked = false;
const _gotScrapingLoggedHosts = new Set();
async function getGotScraping() {
if (_gotScrapingChecked) return _gotScraping;
_gotScrapingChecked = true;
try {
const mod = await import("got-scraping");
_gotScraping = typeof mod.gotScraping === "function" ? mod.gotScraping : null;
if (_gotScraping) dbg("TLS", "got-scraping loaded (browser-like JA3 enabled)");
} catch (e) {
console.warn(`[ProxyFetch] got-scraping unavailable, falling back to native fetch: ${e.message}`);
_gotScraping = null;
}
return _gotScraping;
}
async function gotScrapingFetch(url, options) {
const gs = await getGotScraping();
if (!gs) return null;
const method = (options.method || "GET").toUpperCase();
const headersInit = options.headers || {};
const headers = headersInit instanceof Headers
? Object.fromEntries(headersInit.entries())
: { ...headersInit };
return new Promise((resolve, reject) => {
let settled = false;
const stream = gs.stream({
url,
method,
headers,
body: method === "GET" || method === "HEAD" ? undefined : options.body,
throwHttpErrors: false,
retry: { limit: 0 },
timeout: { request: undefined },
followRedirect: false,
decompress: true,
});
if (options.signal) {
const onAbort = () => { try { stream.destroy(new Error("aborted")); } catch { } };
if (options.signal.aborted) onAbort();
else options.signal.addEventListener("abort", onAbort, { once: true });
}
stream.once("response", (res) => {
if (settled) return;
settled = true;
const resHeaders = new Headers();
for (const [k, v] of Object.entries(res.headers || {})) {
if (Array.isArray(v)) v.forEach((x) => resHeaders.append(k, String(x)));
else if (v != null) resHeaders.set(k, String(v));
}
const body = Readable.toWeb(stream);
resolve(new Response(body, { status: res.statusCode, statusText: res.statusMessage || "", headers: resHeaders }));
});
stream.once("error", (err) => {
if (settled) return;
settled = true;
reject(err);
});
});
}
async function tryGotScrapingFetch(url, options) {
try {
const res = await gotScrapingFetch(url, options);
if (res) {
try {
const host = new URL(typeof url === "string" ? url : url.toString()).hostname;
if (!_gotScrapingLoggedHosts.has(host)) {
_gotScrapingLoggedHosts.add(host);
dbg("TLS", `using got-scraping for ${host}`);
}
} catch { }
}
return res;
} catch (e) {
console.warn(`[ProxyFetch] got-scraping request failed, fallback to native fetch: ${e.message}`);
return null;
}
}
async function getGotScraping() { return null; }
async function tryGotScrapingFetch() { return null; }
*/
// DNS cache — use Map to avoid prototype pollution via malformed hostnames
@@ -349,7 +266,6 @@ export async function proxyAwareFetch(url, options = {}, proxyOptions = null) {
}
// got-scraping disabled — use native fetch directly
// (Re-enable per-host by wrapping with tryGotScrapingFetch when needed)
return originalFetch(url, options);
}

View File

@@ -0,0 +1,47 @@
/**
* Config-driven in-stream error pattern matching.
* See docs/2026-08-04-stream-error-patterns-design.md.
*
* Settings shape: streamErrorPatterns: { [provider]: string[] }
* Pattern syntax per entry:
* - "plain text" → case-insensitive substring match
* - "/regex/flags" → RegExp match (flags after the last "/")
*/
export function parsePatterns(patterns) {
if (!Array.isArray(patterns)) return [];
const out = [];
for (const entry of patterns) {
if (typeof entry !== "string" || !entry.trim()) continue;
const raw = entry.trim();
const m = raw.match(/^\/(.*)\/([a-z]*)$/s);
if (m) {
try {
out.push({ regex: new RegExp(m[1], m[2]), raw });
} catch {
// invalid regex → skip; config errors must never break requests
}
continue;
}
out.push({ text: raw.toLowerCase(), raw });
}
return out;
}
export function matchStreamErrorPatterns(patterns, text) {
if (!Array.isArray(patterns) || patterns.length === 0) return null;
if (typeof text !== "string" || !text) return null;
const lower = text.toLowerCase();
for (const p of parsePatterns(patterns)) {
if (p.text && lower.includes(p.text)) return p.raw;
if (p.regex) {
p.regex.lastIndex = 0; // stateless even with /g
if (p.regex.test(text)) return p.raw;
}
}
return null;
}
export function streamStatusForContent(patterns, content) {
return matchStreamErrorPatterns(patterns, content) ? "error" : "success";
}

View File

@@ -0,0 +1,136 @@
import { matchStreamErrorPatterns } from "./streamErrorPatterns.js";
import { HTTP_STATUS } from "../config/runtimeConfig.js";
const DEFAULT_TIMEOUT_MS = (() => {
const raw = process.env.STREAM_ERROR_PEEK_TIMEOUT_MS;
const n = raw ? parseInt(raw, 10) : NaN;
return Number.isFinite(n) && n > 0 ? n : 3000;
})();
const DEFAULT_MAX_BYTES = (() => {
const raw = process.env.STREAM_ERROR_PEEK_MAX_BYTES;
const n = raw ? parseInt(raw, 10) : NaN;
return Number.isFinite(n) && n > 0 ? n : 8192;
})();
function makeAbortError(reason) {
const error = new Error(reason?.message || reason || "Request aborted");
error.name = "AbortError";
return error;
}
/**
* Read the first bytes of a 200 response and reject it with a 502 when a
* configured stream-error pattern matches — BEFORE any byte reaches the
* client, so account/model fallback can still kick in for streaming too.
* On no-match/timeout/abort the response is re-emitted (buffered bytes +
* rest of stream) unchanged in spirit. Fail-open: never throws.
*/
export async function maybeRejectEarlyStreamError(
response,
patterns,
{
signal = null,
timeoutMs = DEFAULT_TIMEOUT_MS,
maxBytes = DEFAULT_MAX_BYTES,
} = {},
) {
if (!Array.isArray(patterns) || patterns.length === 0) return response;
const reader = response.body.getReader();
const decoder = new TextDecoder();
const abortController = new AbortController();
const forwardAbort = () => abortController.abort(signal?.reason);
if (signal?.aborted) abortController.abort(signal?.reason);
else if (signal)
signal.addEventListener("abort", forwardAbort, { once: true });
// Raw bytes for lossless re-emission; decoded text ONLY for pattern matching.
// Never re-encode decoded text: TextDecoder holds a split multi-byte char
// internally and flush() would replace it with U+FFFD, corrupting the stream.
const rawChunks = [];
let peekedText = "";
let total = 0;
const readWithTimeout = (ms) => {
if (abortController.signal.aborted)
return Promise.reject(makeAbortError(abortController.signal.reason));
const timeoutPromise = new Promise((_, reject) => {
const t = setTimeout(
() => reject(new Error("stream error peek timeout")),
ms,
);
t.unref?.();
});
const abortPromise = new Promise((_, reject) => {
abortController.signal.addEventListener(
"abort",
() => reject(makeAbortError(abortController.signal.reason)),
{ once: true },
);
});
return Promise.race([reader.read(), timeoutPromise, abortPromise]);
};
try {
const deadline = Date.now() + timeoutMs;
while (total < maxBytes && Date.now() < deadline) {
const { done, value } = await readWithTimeout(
Math.max(deadline - Date.now(), 1),
);
if (done) break;
rawChunks.push(value);
peekedText += decoder.decode(value, { stream: true });
total += value.byteLength;
const matched = matchStreamErrorPatterns(patterns, peekedText);
if (matched) {
await reader.cancel("stream error pattern matched").catch(() => {});
return new Response(
JSON.stringify({
error: {
message: `Stream error pattern matched: ${matched}`,
type: "upstream_error",
},
}),
{
status: HTTP_STATUS.BAD_GATEWAY,
statusText: String(matched).slice(0, 200),
headers: { "Content-Type": "application/json" },
},
);
}
}
} catch {
// timeout / abort / read failure → commit; downstream stall/abort handling takes over.
}
if (signal) signal.removeEventListener("abort", forwardAbort);
const remaining = new ReadableStream({
start(controller) {
(async () => {
try {
// Re-emit RAW bytes (never re-encoded decoded text) so split
// multi-byte UTF-8 sequences survive the peek untouched.
for (const c of rawChunks) controller.enqueue(c);
while (true) {
const { done, value } = await reader.read();
if (done) break;
controller.enqueue(value);
}
controller.close();
} catch (err) {
controller.error(err);
}
})();
},
cancel() {
reader.cancel("stream cancelled during peek commit").catch(() => {});
},
});
return new Response(remaining, {
status: response.status,
statusText: response.statusText,
headers: response.headers,
});
}

View File

@@ -397,6 +397,7 @@ export function estimateUsage(body, contentLength, targetFormat = FORMATS.OPENAI
targetFormat
);
}
/**
* Log usage with cache info (green color)
*/

View File

@@ -75,6 +75,7 @@ Common fields above work everywhere. These add/override:
| Provider | Extra/changed fields | Notes |
|---|---|---|
| `openai`, `minimax`, `openrouter`, `recraft` | `quality`, `style`, `response_format` | Standard OpenAI shape |
| `xai` (Grok Imagine) | `aspect_ratio`, `resolution`, `image`, `images[]` | Generate → `/images/generations`; edit/multi-edit → `/images/edits` (auto when `image`/`images` present). `size` maps to `aspect_ratio`. Up to 3 source images. |
| `gemini` (nano-banana) | — | Only `prompt`; ignores `size`/`n` |
| `codex` (gpt-5.4-image) | `image`, `images[]`, `image_detail`, `output_format`, `background` | SSE stream; **ChatGPT Plus/Pro required** |
| `huggingface` | — | Only `prompt`; returns single image |

View File

@@ -2,7 +2,7 @@
import { useState, useEffect, useRef, useCallback } from "react";
import PropTypes from "prop-types";
import { Card, Button, Input, Modal, CardSkeleton, Toggle, ConfirmModal } from "@/shared/components";
import { Card, Button, Input, Modal, CardSkeleton, Toggle, ConfirmModal, ApiExplorerModal } from "@/shared/components";
import { useCopyToClipboard } from "@/shared/hooks/useCopyToClipboard";
import {
TUNNEL_BENEFITS,
@@ -21,6 +21,12 @@ export default function APIPageClient({ machineId }) {
const [keys, setKeys] = useState([]);
const [loading, setLoading] = useState(true);
const [showAddModal, setShowAddModal] = useState(false);
const [showApiExplorer, setShowApiExplorer] = useState(false);
const [showImportModal, setShowImportModal] = useState(false);
const [importKeyValue, setImportKeyValue] = useState("");
const [importKeyName, setImportKeyName] = useState("");
const [importing, setImporting] = useState(false);
const [importError, setImportError] = useState(null);
const [newKeyName, setNewKeyName] = useState("");
const [createdKey, setCreatedKey] = useState(null);
const [confirmState, setConfirmState] = useState(null);
@@ -720,10 +726,20 @@ export default function APIPageClient({ machineId }) {
<div className="flex flex-col gap-8">
{/* Endpoint Card */}
<Card>
<h2 className="text-lg font-semibold mb-4 flex items-center gap-2">
<span className="material-symbols-outlined text-primary">api</span>
API Endpoint
</h2>
<div className="flex items-center justify-between gap-3 mb-4">
<h2 className="text-lg font-semibold flex items-center gap-2">
<span className="material-symbols-outlined text-primary">api</span>
API Endpoint
</h2>
<Button
size="sm"
variant="secondary"
icon="science"
onClick={() => setShowApiExplorer(true)}
>
API Explorer
</Button>
</div>
{/* Endpoint rows */}
<div className="flex flex-col gap-2">
@@ -970,9 +986,14 @@ export default function APIPageClient({ machineId }) {
<span className="material-symbols-outlined text-primary">vpn_key</span>
API Keys
</h2>
<Button icon="add" onClick={() => setShowAddModal(true)}>
Create Key
</Button>
<div className="flex gap-2">
<Button icon="add" onClick={() => setShowAddModal(true)}>
Create Key
</Button>
<Button icon="file_open" variant="secondary" onClick={() => setShowImportModal(true)}>
Import Key
</Button>
</div>
</div>
<div className="flex items-center justify-between pb-4 mb-4 border-b border-border">
@@ -1110,6 +1131,100 @@ export default function APIPageClient({ machineId }) {
</div>
</Modal>
{/* Import Key Modal */}
<Modal
isOpen={showImportModal}
title="Import Existing API Key"
onClose={() => {
setShowImportModal(false);
setImportKeyValue("");
setImportKeyName("");
setImportError(null);
}}
>
<div className="flex flex-col gap-4">
<div className="bg-surface-2 border border-border-subtle rounded-lg p-3">
<p className="text-sm text-text-muted">
Paste an existing API key to add it to this instance.
Useful for transferring keys from another 9Router instance or adding externally generated keys.
</p>
</div>
{importError && (
<div className="flex items-center gap-2 px-3 py-2 rounded-lg border border-red-300 dark:border-red-800 bg-red-500/10 text-sm text-red-600 dark:text-red-400">
<span className="material-symbols-outlined text-[16px]">error</span>
{importError}
</div>
)}
<Input
label="API Key"
value={importKeyValue}
onChange={(e) => {
setImportKeyValue(e.target.value);
setImportError(null);
}}
placeholder="Paste your API key here"
className="font-mono"
/>
<Input
label="Key Name (optional)"
value={importKeyName}
onChange={(e) => setImportKeyName(e.target.value)}
placeholder="Imported Key"
/>
<div className="flex gap-2">
<Button
onClick={async () => {
if (!importKeyValue.trim()) return;
setImporting(true);
setImportError(null);
try {
const res = await fetch("/api/keys/import", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
key: importKeyValue.trim(),
name: importKeyName.trim() || null,
}),
});
const data = await res.json();
if (res.ok) {
await fetchData();
setShowImportModal(false);
setImportKeyValue("");
setImportKeyName("");
} else {
setImportError(data.error || "Failed to import key");
}
} catch (error) {
setImportError("Network error. Please try again.");
} finally {
setImporting(false);
}
}}
fullWidth
disabled={!importKeyValue.trim() || importing}
>
{importing ? "Importing..." : "Import"}
</Button>
<Button
onClick={() => {
setShowImportModal(false);
setImportKeyValue("");
setImportKeyName("");
setImportError(null);
}}
variant="ghost"
fullWidth
disabled={importing}
>
Cancel
</Button>
</div>
</div>
</Modal>
{/* Created Key Modal */}
<Modal
isOpen={!!createdKey}
@@ -1300,6 +1415,12 @@ export default function APIPageClient({ machineId }) {
message={confirmState?.message}
variant="danger"
/>
{/* API Explorer — list + test all public AI endpoints */}
<ApiExplorerModal
isOpen={showApiExplorer}
onClose={() => setShowApiExplorer(false)}
/>
</div>
);
}

View File

@@ -43,6 +43,8 @@ export const KIND_EXAMPLE_CONFIG = {
extraFields: [
{ key: "n", label: "n", type: "number", default: 1, min: 1, max: 4 },
{ key: "size", label: "Size", type: "select", default: "auto", options: ["auto", "1024x1024", "1024x1536", "1536x1024", "1024x1792", "1792x1024"] },
{ key: "aspect_ratio", label: "Aspect", type: "select", default: "", options: ["", "auto", "1:1", "16:9", "9:16", "4:3", "3:2", "2:3", "9:19.5", "20:9"] },
{ key: "resolution", label: "Resolution", type: "select", default: "", options: ["", "1k", "2k"] },
{ key: "quality", label: "Quality", type: "select", default: "auto", options: ["auto", "low", "medium", "high", "standard", "hd"] },
{ key: "background", label: "Background", type: "select", default: "auto", options: ["auto", "transparent", "opaque"] },
{ key: "style", label: "Style", type: "select", default: "", options: ["", "vivid", "natural"] },

View File

@@ -286,6 +286,25 @@ export default function ProfilePage() {
}
};
const handleGlobalTimeoutChange = async (e) => {
const raw = e.target.value.replace(/[^0-9]/g, "");
const numTimeout = parseInt(raw, 10);
const patchValue = (raw !== "" && Number.isFinite(numTimeout) && numTimeout > 0) ? numTimeout : null;
try {
const res = await fetch("/api/settings", {
method: "PATCH",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ defaultTimeoutMs: patchValue }),
});
if (res.ok) {
setSettings(prev => ({ ...prev, defaultTimeoutMs: patchValue }));
}
} catch (err) {
console.error("Failed to update default timeout:", err);
}
};
const updateStickyLimit = async (limit) => {
const numLimit = parseInt(limit);
if (isNaN(numLimit) || numLimit < 1) return;
@@ -643,21 +662,6 @@ export default function ProfilePage() {
}
};
const updateLogLevel = async (logLevel) => {
try {
const res = await fetch("/api/settings", {
method: "PATCH",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ logLevel }),
});
if (res.ok) {
setSettings(prev => ({ ...prev, logLevel }));
}
} catch (err) {
console.error("Failed to update logLevel:", err);
}
};
const reloadSettings = async () => {
try {
const res = await fetch("/api/settings");
@@ -1535,6 +1539,43 @@ export default function ProfilePage() {
</div>
</Card>
{/* Default Timeout — global default for all providers */}
<Card>
<div className="flex items-center gap-3 mb-4">
<div className="p-2 rounded-lg bg-amber-500/10 text-amber-500 shrink-0">
<span className="material-symbols-outlined text-[20px]">timer</span>
</div>
<h3 className="text-base sm:text-lg font-semibold">Default Connect Timeout</h3>
</div>
<div className="flex flex-col gap-4">
<div className="flex items-start sm:items-center justify-between gap-4">
<div className="flex-1 min-w-0">
<p className="font-medium text-sm sm:text-base">All Providers</p>
<p className="text-xs sm:text-sm text-text-muted">
Timeout for upstream connect (applies globally unless overridden per provider). Set to 0 or leave empty for system default (60s).
</p>
</div>
<div className="flex items-center gap-1.5">
<Input
type="text"
inputMode="numeric"
placeholder="60000"
value={settings.defaultTimeoutMs != null ? String(settings.defaultTimeoutMs) : ""}
onChange={handleGlobalTimeoutChange}
disabled={loading}
className="w-20 text-center"
/>
<span className="text-xs text-text-muted shrink-0">ms</span>
</div>
</div>
<p className="text-xs text-text-muted italic">
{settings.defaultTimeoutMs
? `All providers will wait up to ${settings.defaultTimeoutMs}ms for a connection.`
: "Using system default (60s) — configure per-provider timeout on each provider's detail page for fine-grained control."}
</p>
</div>
</Card>
{/* Network */}
<Card>
<div className="flex items-center gap-3 mb-4">
@@ -1630,39 +1671,6 @@ export default function ProfilePage() {
</div>
</Card>
{/* Logging Settings */}
<Card>
<div className="flex items-center gap-3 mb-4">
<div className="p-2 rounded-lg bg-slate-500/10 text-slate-500 shrink-0">
<span className="material-symbols-outlined text-[20px]">terminal</span>
</div>
<h3 className="text-base sm:text-lg font-semibold">Logging</h3>
</div>
<div className="flex items-start sm:items-center justify-between gap-4">
<div className="flex-1 min-w-0">
<p className="font-medium text-sm sm:text-base">Log level</p>
<p className="text-xs sm:text-sm text-text-muted">
Controls how much the server prints to console. In production,
set to <span className="font-medium">Error</span> to only show
important errors, hiding the per-request INFO lines (▶ POST,
📊 DONE, [COMBO], [CHAT]). Applied immediately, no restart
needed.
</p>
</div>
<select
value={settings.logLevel || "info"}
onChange={(e) => updateLogLevel(e.target.value)}
disabled={loading}
className="shrink-0 rounded-md border border-border bg-background px-2 py-1.5 text-sm focus:border-primary focus:outline-none"
>
<option value="debug">Debug</option>
<option value="info">Info</option>
<option value="warn">Warn</option>
<option value="error">Error</option>
</select>
</div>
</Card>
{/* Account actions */}
<div className="flex flex-col sm:flex-row gap-2">
<Button

View File

@@ -3,14 +3,44 @@
import { useState, useEffect, useRef } from "react";
import { getStatusVariant as getConnectionStatusVariant } from "@/shared/utils/connectionStatus";
import PropTypes from "prop-types";
import { Badge, Toggle, Tooltip } from "@/shared/components";
import { Badge, Toggle, Tooltip, Modal, Button } from "@/shared/components";
import { useCopyToClipboard } from "@/shared/hooks/useCopyToClipboard";
import CooldownTimer from "./CooldownTimer";
export default function ConnectionRow({ connection, proxyPools, isOAuth, isFirst, isLast, onMoveUp, onMoveDown, onToggleActive, onUpdateProxy, onEdit, onDelete, oneByOneStatus = null, autoPing = null }) {
export default function ConnectionRow({
connection,
proxyPools,
isOAuth,
isFirst,
isLast,
onMoveUp,
onMoveDown,
onToggleActive,
onUpdateProxy,
onEdit,
onDelete,
oneByOneStatus = null,
autoPing = null,
testModels = [],
providerAlias = null,
}) {
const [showProxyDropdown, setShowProxyDropdown] = useState(false);
const [updatingProxy, setUpdatingProxy] = useState(false);
const [showKeyModal, setShowKeyModal] = useState(false);
const [revealedKey, setRevealedKey] = useState("");
const [loadingKey, setLoadingKey] = useState(false);
const [keyError, setKeyError] = useState(null);
const [showTestModelModal, setShowTestModelModal] = useState(false);
const [selectedTestModelId, setSelectedTestModelId] = useState("");
const [manualTestModelId, setManualTestModelId] = useState("");
const [testingModel, setTestingModel] = useState(false);
const [testModelResult, setTestModelResult] = useState(null);
const { copied, copy } = useCopyToClipboard();
const proxyDropdownRef = useRef(null);
const canTestModel = !!providerAlias && (connection.isActive !== false);
const hasCatalogModels = Array.isArray(testModels) && testModels.length > 0;
const proxyPoolMap = new Map((proxyPools || []).map((pool) => [pool.id, pool]));
const boundProxyPoolId = connection.providerSpecificData?.proxyPoolId || null;
const boundProxyPool = boundProxyPoolId ? proxyPoolMap.get(boundProxyPoolId) : null;
@@ -69,6 +99,58 @@ export default function ConnectionRow({ connection, proxyPools, isOAuth, isFirst
}
};
const openTestModelModal = () => {
setTestModelResult(null);
setManualTestModelId("");
setSelectedTestModelId(hasCatalogModels ? (testModels[0]?.id || "") : "");
setShowTestModelModal(true);
};
const resolveTestModelId = () => {
if (hasCatalogModels) return selectedTestModelId?.trim() || "";
return manualTestModelId?.trim() || "";
};
const handleTestModel = async () => {
if (!canTestModel || testingModel) return;
const modelId = resolveTestModelId();
if (!modelId) {
setTestModelResult({ ok: false, error: "Select or enter a model id" });
return;
}
setTestingModel(true);
setTestModelResult(null);
try {
const res = await fetch("/api/models/test", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
model: `${providerAlias}/${modelId}`,
connectionId: connection.id,
}),
});
const data = await res.json().catch(() => ({}));
if (!res.ok) {
setTestModelResult({
ok: false,
error: data.error || `HTTP ${res.status}`,
latencyMs: data.latencyMs,
});
return;
}
setTestModelResult({
ok: !!data.ok,
error: data.error || null,
latencyMs: data.latencyMs,
status: data.status,
});
} catch (error) {
setTestModelResult({ ok: false, error: error.message || "Network error" });
} finally {
setTestingModel(false);
}
};
const rowAuthType = connection.authType || (isOAuth ? "oauth" : "apikey");
const isOAuthConnection = rowAuthType === "oauth";
const isCookieConnection = rowAuthType === "cookie";
@@ -257,6 +339,46 @@ export default function ConnectionRow({ connection, proxyPools, isOAuth, isFirst
</button>
</Tooltip>
)}
{connection.authType === "apikey" && (
<button
onClick={async () => {
setLoadingKey(true);
setKeyError(null);
try {
const res = await fetch(`/api/providers/${connection.id}/api-key`);
const data = await res.json();
if (res.ok) {
setRevealedKey(data.apiKey || "");
setShowKeyModal(true);
} else {
setKeyError(data.error || "Failed to fetch key");
setShowKeyModal(true);
}
} catch {
setKeyError("Network error");
setShowKeyModal(true);
} finally {
setLoadingKey(false);
}
}}
className="flex flex-col items-center rounded px-2 py-1 text-text-muted hover:bg-black/5 hover:text-primary dark:hover:bg-white/5"
>
<span className={`material-symbols-outlined text-[18px] ${loadingKey ? "animate-spin" : ""}`}>
{loadingKey ? "progress_activity" : "key"}
</span>
<span className="text-[10px] leading-tight">Key</span>
</button>
)}
{canTestModel && (
<button
onClick={openTestModelModal}
className="flex flex-col items-center rounded px-2 py-1 text-text-muted hover:bg-black/5 hover:text-primary dark:hover:bg-white/5"
title="Test a model using only this account"
>
<span className="material-symbols-outlined text-[18px]">science</span>
<span className="text-[10px] leading-tight">Test</span>
</button>
)}
<button onClick={onEdit} className="flex flex-col items-center rounded px-2 py-1 text-text-muted hover:bg-black/5 hover:text-primary dark:hover:bg-white/5">
<span className="material-symbols-outlined text-[18px]">edit</span>
<span className="text-[10px] leading-tight">Edit</span>
@@ -273,6 +395,153 @@ export default function ConnectionRow({ connection, proxyPools, isOAuth, isFirst
title={(connection.isActive ?? true) ? "Disable connection" : "Enable connection"}
/>
</div>
{/* Show Key Modal */}
<Modal
isOpen={showKeyModal}
title="API Key"
onClose={() => {
setShowKeyModal(false);
setRevealedKey("");
setKeyError(null);
}}
>
<div className="flex flex-col gap-4">
{keyError ? (
<div className="flex items-center gap-2 px-3 py-2 rounded-lg border border-red-300 dark:border-red-800 bg-red-500/10 text-sm text-red-600 dark:text-red-400">
<span className="material-symbols-outlined text-[16px]">error</span>
{keyError}
</div>
) : (
<>
<p className="text-xs text-yellow-600 dark:text-yellow-400 bg-yellow-50 dark:bg-yellow-900/20 border border-yellow-200 dark:border-yellow-800 rounded-lg p-3">
<span className="material-symbols-outlined text-[14px] align-middle mr-1">warning</span>
This key provides full access to your endpoint. Keep it secure.
</p>
<div className="flex gap-2">
<input
type="text"
readOnly
value={revealedKey}
className="flex-1 px-3 py-2 text-sm font-mono bg-surface-2 rounded-lg border border-border focus:outline-none"
onClick={(e) => e.target.select()}
/>
<button
onClick={() => copy(revealedKey, connection.id)}
className="flex items-center gap-1 px-3 py-2 rounded-lg border border-border text-sm text-text-muted hover:text-primary hover:border-primary/40 transition-colors"
>
<span className="material-symbols-outlined text-[16px]">{copied === connection.id ? "check" : "content_copy"}</span>
{copied === connection.id ? "Copied!" : "Copy"}
</button>
</div>
</>
)}
<Button onClick={() => { setShowKeyModal(false); setRevealedKey(""); setKeyError(null); }} fullWidth>
Close
</Button>
</div>
</Modal>
{/* Test Model Modal — pins this connection via x-connection-id */}
<Modal
isOpen={showTestModelModal}
title="Test Model"
onClose={() => {
if (testingModel) return;
setShowTestModelModal(false);
setTestModelResult(null);
}}
>
<div className="flex flex-col gap-4">
<p className="text-xs text-text-muted">
Call one model using only this account:
<span className="ml-1 font-medium text-text-main">{displayName}</span>
</p>
{hasCatalogModels ? (
<div className="flex flex-col gap-1.5">
<label className="text-xs font-medium text-text-muted">Model</label>
<select
value={selectedTestModelId}
onChange={(e) => {
setSelectedTestModelId(e.target.value);
setTestModelResult(null);
}}
className="w-full rounded-lg border border-border bg-background px-3 py-2 text-sm focus:border-primary focus:outline-none"
disabled={testingModel}
>
{testModels.map((model) => (
<option key={model.id} value={model.id}>
{model.name && model.name !== model.id ? `${model.name} (${model.id})` : model.id}
</option>
))}
</select>
</div>
) : (
<div className="flex flex-col gap-1.5">
<label className="text-xs font-medium text-text-muted">Model ID</label>
<input
type="text"
value={manualTestModelId}
onChange={(e) => {
setManualTestModelId(e.target.value);
setTestModelResult(null);
}}
placeholder="e.g. gpt-4o-mini"
className="w-full rounded-lg border border-border bg-background px-3 py-2 text-sm font-mono focus:border-primary focus:outline-none"
disabled={testingModel}
/>
</div>
)}
{testModelResult && (
<div
className={`flex items-start gap-2 rounded-lg border px-3 py-2 text-sm ${
testModelResult.ok
? "border-green-300 bg-green-500/10 text-green-700 dark:border-green-800 dark:text-green-400"
: "border-red-300 bg-red-500/10 text-red-600 dark:border-red-800 dark:text-red-400"
}`}
>
<span className="material-symbols-outlined shrink-0 text-[16px]">
{testModelResult.ok ? "check_circle" : "error"}
</span>
<div className="min-w-0 break-words">
<p className="font-medium">
{testModelResult.ok ? "Model reachable" : "Test failed"}
{typeof testModelResult.latencyMs === "number" ? ` · ${testModelResult.latencyMs}ms` : ""}
</p>
{testModelResult.error && (
<p className="mt-0.5 text-xs opacity-90">{testModelResult.error}</p>
)}
</div>
</div>
)}
<div className="flex gap-2">
<Button
onClick={handleTestModel}
disabled={!resolveTestModelId()}
loading={testingModel}
icon="science"
fullWidth
>
{testingModel ? "Testing..." : "Run Test"}
</Button>
<Button
variant="secondary"
onClick={() => {
if (testingModel) return;
setShowTestModelModal(false);
setTestModelResult(null);
}}
disabled={testingModel}
fullWidth
>
Close
</Button>
</div>
</div>
</Modal>
</div>
);
}
@@ -315,4 +584,9 @@ ConnectionRow.propTypes = {
onToggle: PropTypes.func,
provider: PropTypes.string,
}),
testModels: PropTypes.arrayOf(PropTypes.shape({
id: PropTypes.string.isRequired,
name: PropTypes.string,
})),
providerAlias: PropTypes.string,
};

View File

@@ -65,6 +65,7 @@ export default function ProviderDetailPage() {
const [providerStrategy, setProviderStrategy] = useState(null);
const [providerStickyLimit, setProviderStickyLimit] = useState("");
const [providerNoAuthEnabled, setProviderNoAuthEnabled] = useState(true);
const [providerTimeout, setProviderTimeout] = useState("");
const [thinkingMode, setThinkingMode] = useState("auto");
const [autoPing, setAutoPing] = useState({ enabled: false, connections: {} });
const [suggestedModels, setSuggestedModels] = useState([]);
@@ -323,6 +324,9 @@ export default function ProviderDetailPage() {
setProviderStrategy(override.fallbackStrategy || null);
setProviderStickyLimit(override.stickyRoundRobinLimit != null ? String(override.stickyRoundRobinLimit) : "1");
setProviderNoAuthEnabled(override.enabled !== false);
// Load per-provider connect timeout
const timeoutCfg = (settingsData.providerTimeouts || {})[providerId] || {};
setProviderTimeout(timeoutCfg.timeoutMs != null ? String(timeoutCfg.timeoutMs) : "");
// Load per-provider thinking config
const thinkingCfg = (settingsData.providerThinking || {})[providerId] || {};
setThinkingMode(thinkingCfg.mode || "auto");
@@ -459,6 +463,38 @@ export default function ProviderDetailPage() {
saveThinkingConfig(mode);
};
const saveProviderTimeout = async (ms) => {
try {
const settingsRes = await fetch("/api/settings", { cache: "no-store" });
const settingsData = settingsRes.ok ? await settingsRes.json() : {};
const current = settingsData.providerTimeouts || {};
const updated = { ...current };
if (!ms || ms === "") {
delete updated[providerId];
} else {
const timeoutMs = parseInt(ms, 10);
if (Number.isFinite(timeoutMs) && timeoutMs > 0) {
updated[providerId] = { timeoutMs };
} else {
delete updated[providerId];
}
}
await fetch("/api/settings", {
method: "PATCH",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ providerTimeouts: updated }),
});
} catch (error) {
console.log("Error saving provider timeout:", error);
}
};
const handleTimeoutChange = (value) => {
const cleaned = value.replace(/[^0-9]/g, "");
setProviderTimeout(cleaned);
saveProviderTimeout(cleaned);
};
const saveAutoPing = async (next) => {
const autoPingSettingsKey = AUTO_PING_SETTINGS_KEYS[providerId];
if (!autoPingSettingsKey) return;
@@ -1678,6 +1714,21 @@ export default function ProviderDetailPage() {
)}
</>
)}
{/* Connect Timeout */}
<div className="flex flex-wrap items-center gap-2">
<span className="text-xs text-text-muted font-medium">Connect Timeout</span>
<div className="flex items-center gap-1.5">
<input
type="text"
inputMode="numeric"
value={providerTimeout}
onChange={(e) => handleTimeoutChange(e.target.value)}
placeholder="default"
className="w-20 px-2 py-1 text-xs border border-border rounded-md bg-background focus:outline-none focus:border-primary"
/>
<span className="text-xs text-text-muted">ms</span>
</div>
</div>
{/* Round Robin toggle */}
<div className="flex flex-wrap items-center gap-2">
<span className="text-xs text-text-muted font-medium">Round Robin</span>

View File

@@ -147,7 +147,7 @@ export default function ProviderLimits() {
const [proxyPools, setProxyPools] = useState([]);
const [providerFilter, setProviderFilter] = useState("all");
const [providerOptions, setProviderOptions] = useState([]);
const [accountFilter, setAccountFilter] = useState("all");
const [accountFilter, setAccountFilter] = useState("active");
const [quotaSortMode, setQuotaSortMode] = useState("default");
const [quotaVisibility, setQuotaVisibility] = useState({});
const [expiringFirst, setExpiringFirst] = useState(false);

View File

@@ -114,9 +114,19 @@ export default function RequestDetailsTab() {
const [providerNameCache, setProviderNameCache] = useState(null);
const [filters, setFilters] = useState({
provider: "",
model: "",
status: "",
connectionId: "",
apiKey: "",
startDate: "",
endDate: ""
});
const [filterOptions, setFilterOptions] = useState({
models: [],
apiKeys: [],
statuses: [],
connections: []
});
const fetchProviders = useCallback(async () => {
try {
@@ -131,6 +141,21 @@ export default function RequestDetailsTab() {
}
}, []);
const fetchFilterOptions = useCallback(async () => {
try {
const res = await fetch("/api/usage/filters");
const data = await res.json();
setFilterOptions({
models: data.models || [],
apiKeys: data.apiKeys || [],
statuses: data.statuses || [],
connections: data.connections || []
});
} catch (error) {
console.error("Failed to fetch filter options:", error);
}
}, []);
const fetchDetails = useCallback(async () => {
setLoading(true);
try {
@@ -139,6 +164,10 @@ export default function RequestDetailsTab() {
pageSize: pagination.pageSize.toString()
});
if (filters.provider) params.append("provider", filters.provider);
if (filters.model) params.append("model", filters.model);
if (filters.status) params.append("status", filters.status);
if (filters.connectionId) params.append("connectionId", filters.connectionId);
if (filters.apiKey) params.append("apiKey", filters.apiKey);
if (filters.startDate) params.append("startDate", filters.startDate);
if (filters.endDate) params.append("endDate", filters.endDate);
@@ -156,7 +185,8 @@ export default function RequestDetailsTab() {
useEffect(() => {
fetchProviders();
}, [fetchProviders]);
fetchFilterOptions();
}, [fetchProviders, fetchFilterOptions]);
useEffect(() => {
fetchDetails();
@@ -176,7 +206,7 @@ export default function RequestDetailsTab() {
};
const handleClearFilters = () => {
setFilters({ provider: "", startDate: "", endDate: "" });
setFilters({ provider: "", model: "", status: "", connectionId: "", apiKey: "", startDate: "", endDate: "" });
};
return (
@@ -204,6 +234,98 @@ export default function RequestDetailsTab() {
))}
</select>
</div>
<div className="flex min-w-0 flex-col gap-2">
<label htmlFor="model-filter" className="text-sm font-medium text-text-main">Model</label>
<select
id="model-filter"
value={filters.model}
onChange={(e) => setFilters({ ...filters, model: e.target.value })}
className={cn(
"h-9 px-3 rounded-lg border border-black/10 dark:border-white/10 bg-surface",
"text-sm text-text-main focus:outline-none focus:ring-2 focus:ring-primary/20",
"w-full min-w-0 cursor-pointer"
)}
style={{ colorScheme: 'auto' }}
>
<option value="">All Models</option>
{filterOptions.models.map((model) => (
<option key={model} value={model}>
{model}
</option>
))}
</select>
</div>
<div className="flex min-w-0 flex-col gap-2">
<label htmlFor="status-filter" className="text-sm font-medium text-text-main">Status</label>
<select
id="status-filter"
value={filters.status}
onChange={(e) => setFilters({ ...filters, status: e.target.value })}
className={cn(
"h-9 px-3 rounded-lg border border-black/10 dark:border-white/10 bg-surface",
"text-sm text-text-main focus:outline-none focus:ring-2 focus:ring-primary/20",
"w-full min-w-0 cursor-pointer"
)}
style={{ colorScheme: 'auto' }}
>
<option value="">All Statuses</option>
{filterOptions.statuses.map((status) => (
<option key={status} value={status}>
{status}
</option>
))}
</select>
</div>
<div className="flex min-w-0 flex-col gap-2">
<label htmlFor="connection-filter" className="text-sm font-medium text-text-main">Account</label>
<select
id="connection-filter"
value={filters.connectionId}
onChange={(e) => setFilters({ ...filters, connectionId: e.target.value })}
className={cn(
"h-9 px-3 rounded-lg border border-black/10 dark:border-white/10 bg-surface",
"text-sm text-text-main focus:outline-none focus:ring-2 focus:ring-primary/20",
"w-full min-w-0 cursor-pointer"
)}
style={{ colorScheme: 'auto' }}
>
<option value="">All Accounts</option>
{filterOptions.connections.map((group) => (
<optgroup key={group.provider} label={group.name}>
{group.accounts.map((conn) => (
<option key={conn.id} value={conn.id}>
{conn.label}
</option>
))}
</optgroup>
))}
</select>
</div>
<div className="flex min-w-0 flex-col gap-2">
<label htmlFor="apikey-filter" className="text-sm font-medium text-text-main">API Key</label>
<select
id="apikey-filter"
value={filters.apiKey}
onChange={(e) => setFilters({ ...filters, apiKey: e.target.value })}
className={cn(
"h-9 px-3 rounded-lg border border-black/10 dark:border-white/10 bg-surface",
"text-sm text-text-main focus:outline-none focus:ring-2 focus:ring-primary/20",
"w-full min-w-0 cursor-pointer"
)}
style={{ colorScheme: 'auto' }}
>
<option value="">All API Keys</option>
{filterOptions.apiKeys.map((key) => (
<option key={key} value={key}>
{key}
</option>
))}
</select>
</div>
<div className="flex min-w-0 flex-col gap-2">
<label htmlFor="start-date-filter" className="text-sm font-medium text-text-main">Start Date</label>
@@ -238,7 +360,7 @@ export default function RequestDetailsTab() {
<Button
variant="ghost"
onClick={handleClearFilters}
disabled={!filters.provider && !filters.startDate && !filters.endDate}
disabled={!filters.provider && !filters.model && !filters.status && !filters.connectionId && !filters.apiKey && !filters.startDate && !filters.endDate}
className="w-full"
>
Clear Filters
@@ -255,6 +377,7 @@ export default function RequestDetailsTab() {
<th className="text-left p-4 text-sm font-semibold text-text-main">Timestamp</th>
<th className="text-left p-4 text-sm font-semibold text-text-main">Model</th>
<th className="text-left p-4 text-sm font-semibold text-text-main">Provider</th>
<th className="text-left p-4 text-sm font-semibold text-text-main">Status</th>
<th className="text-right p-4 text-sm font-semibold text-text-main">Input Tokens</th>
<th className="text-right p-4 text-sm font-semibold text-text-main">Cached</th>
<th className="text-right p-4 text-sm font-semibold text-text-main">Cache Creation</th>
@@ -296,6 +419,20 @@ export default function RequestDetailsTab() {
{getProviderName(detail.provider, providerNameCache)}
</span>
</td>
<td className="p-4 text-sm">
<span className={cn(
"inline-flex items-center gap-1 rounded-full px-2 py-0.5 text-xs font-medium",
detail.status === "error"
? "bg-red-500/10 text-red-600"
: "bg-green-500/10 text-green-600"
)}>
<span className={cn(
"h-1.5 w-1.5 rounded-full",
detail.status === "error" ? "bg-red-500" : "bg-green-500"
)} />
{detail.status === "error" ? "Error" : "Success"}
</span>
</td>
<td className="p-4 text-sm text-text-main text-right font-mono">
{getInputTokens(detail.tokens).toLocaleString()}
</td>

View File

@@ -105,6 +105,7 @@ export default function UsageTable({
renderDetailCells,
renderSummaryCells,
emptyMessage,
providerLabel,
}) {
const [expanded, setExpanded] = useState(new Set());
@@ -198,8 +199,13 @@ export default function UsageTable({
<span className={`material-symbols-outlined text-[18px] text-text-muted transition-transform ${expanded.has(group.groupKey) ? "rotate-90" : ""}`}>
chevron_right
</span>
<span className={`font-medium transition-colors ${group.summary.pending > 0 ? "text-primary" : ""}`}>
{group.groupKey}
<span
className={`font-medium transition-colors ${group.summary.pending > 0 ? "text-primary" : ""}`}
title={group.groupKey}
>
{tableType === "provider" && providerLabel
? providerLabel(group.groupKey)
: group.groupKey}
</span>
</div>
</td>

View File

@@ -0,0 +1,31 @@
import { NextResponse } from "next/server";
import { importApiKey, getApiKeys } from "@/lib/localDb";
export const dynamic = "force-dynamic";
// POST /api/keys/import - Import existing API key
export async function POST(request) {
try {
const body = await request.json();
const { name, key } = body;
if (!key?.trim()) {
return NextResponse.json({ error: "API key value is required" }, { status: 400 });
}
const apiKey = await importApiKey(name, key.trim());
return NextResponse.json({
key: apiKey.key,
name: apiKey.name,
id: apiKey.id,
}, { status: 201 });
} catch (error) {
const message = error.message;
if (message?.includes("already exists")) {
return NextResponse.json({ error: message }, { status: 409 });
}
console.log("Error importing key:", error);
return NextResponse.json({ error: message || "Failed to import key" }, { status: 500 });
}
}

View File

@@ -50,8 +50,15 @@ async function getInternalHeaders() {
return headers;
}
export async function pingModelByKind(model, kind, baseUrl = `http://127.0.0.1:${process.env.PORT || UPDATER_CONFIG.appPort}`) {
export async function pingModelByKind(
model,
kind,
baseUrl = `http://127.0.0.1:${process.env.PORT || UPDATER_CONFIG.appPort}`,
options = {},
) {
const headers = await getInternalHeaders();
// Pin the request to a specific provider connection when testing a single account.
if (options.connectionId) headers["x-connection-id"] = options.connectionId;
const start = Date.now();
if (kind === "embedding") {

View File

@@ -2,11 +2,12 @@ import { NextResponse } from "next/server";
import { pingModelByKind } from "./ping";
// POST /api/models/test - Ping a single model via internal completions or embeddings
// Optional body.connectionId pins the request to one provider account.
export async function POST(request) {
try {
const { model, kind } = await request.json();
const { model, kind, connectionId } = await request.json();
if (!model) return NextResponse.json({ error: "Model required" }, { status: 400 });
const result = await pingModelByKind(model, kind || "llm");
const result = await pingModelByKind(model, kind || "llm", undefined, { connectionId: connectionId || null });
return NextResponse.json(result);
} catch (err) {
return NextResponse.json({ ok: false, error: err.message }, { status: 500 });

View File

@@ -0,0 +1,26 @@
import { NextResponse } from "next/server";
import { getProviderConnectionById } from "@/models";
export const dynamic = "force-dynamic";
// GET /api/providers/[id]/api-key - Get API key for a connection
// Only returns key for apikey authType connections
export async function GET(request, { params }) {
try {
const { id } = await params;
const connection = await getProviderConnectionById(id);
if (!connection) {
return NextResponse.json({ error: "Connection not found" }, { status: 404 });
}
if (connection.authType !== "apikey" && connection.authType !== "api_key") {
return NextResponse.json({ error: "This connection does not use API key authentication" }, { status: 400 });
}
return NextResponse.json({ apiKey: connection.apiKey || "" });
} catch (error) {
console.log("Error fetching API key:", error);
return NextResponse.json({ error: "Failed to fetch API key" }, { status: 500 });
}
}

View File

@@ -140,6 +140,12 @@ export async function PUT(request, { params }) {
...(providerSpecificData || {}),
};
// null sentinel = explicit delete for sensitive/optional PSD keys
for (const key of Object.keys(updateData.providerSpecificData)) {
if (updateData.providerSpecificData[key] === null) {
delete updateData.providerSpecificData[key];
}
}
if (proxyConfig.hasAnyProxyField) {
updateData.providerSpecificData.connectionProxyEnabled = proxyConfig.connectionProxyEnabled;
updateData.providerSpecificData.connectionProxyUrl = proxyConfig.connectionProxyUrl;
@@ -163,6 +169,10 @@ export async function PUT(request, { params }) {
delete result.accessToken;
delete result.refreshToken;
delete result.idToken;
if (result.providerSpecificData) {
const psd = { ...result.providerSpecificData };
result.providerSpecificData = psd;
}
return NextResponse.json({ connection: result });
} catch (error) {

View File

@@ -43,15 +43,17 @@ export async function POST(request, { params }) {
// Warm up with first model to trigger token refresh (if needed) before parallel calls.
// This prevents race condition where multiple requests concurrently refresh the same token.
// Always pin to this connection so multi-account providers test the intended account only.
const pinOpts = { connectionId: id };
const [first, ...rest] = models;
const firstKind = first.kind || first.type || "llm";
const firstResult = await pingModelByKind(`${alias}/${first.id}`, firstKind, baseUrl);
const firstResult = await pingModelByKind(`${alias}/${first.id}`, firstKind, baseUrl, pinOpts);
const results = [{ modelId: first.id, name: first.name || first.id, ...firstResult }];
if (rest.length > 0) {
const restResults = await Promise.all(
rest.map(async (model) => {
const result = await pingModelByKind(`${alias}/${model.id}`, model.kind || model.type || "llm", baseUrl);
const result = await pingModelByKind(`${alias}/${model.id}`, model.kind || model.type || "llm", baseUrl, pinOpts);
return { modelId: model.id, name: model.name || model.id, ...result };
})
);

View File

@@ -44,8 +44,12 @@ function sanitize(c) {
}
function isUsageEligible(connection) {
return USAGE_SUPPORTED_PROVIDERS.includes(connection.provider) && (
connection.authType === "oauth" || USAGE_APIKEY_PROVIDERS.includes(connection.provider)
if (!USAGE_SUPPORTED_PROVIDERS.includes(connection.provider)) return false;
// OAuth + apikey/cookie providers that expose a usage API (cookie used by grok-web).
return (
connection.authType === "oauth" ||
connection.authType === "cookie" ||
USAGE_APIKEY_PROVIDERS.includes(connection.provider)
);
}

View File

@@ -66,6 +66,7 @@ export async function GET() {
const name = isCompatible
? (c.name || nodeNameMap[c.provider] || c.providerSpecificData?.nodeName || c.provider)
: c.name;
const psd = c.providerSpecificData ? { ...c.providerSpecificData } : undefined;
return {
...c,
name,
@@ -73,6 +74,7 @@ export async function GET() {
accessToken: undefined,
refreshToken: undefined,
idToken: undefined,
providerSpecificData: psd,
};
});

View File

@@ -132,14 +132,15 @@ export async function GET(request, { params }) {
return Response.json({ error: "Connection not found" }, { status: 404 });
}
// Allow OAuth connections, plus whitelisted apikey providers (glm/minimax/kiro/...)
// Allow OAuth connections, plus whitelisted apikey/cookie providers (glm/minimax/kiro/grok-web/...)
// Kiro's headless api-key flow persists authType "api_key" (underscore) while
// generic apikey providers persist "apikey" — accept both spellings here.
// generic apikey providers persist "apikey". Web cookie providers (grok-web) use "cookie".
const isOAuth = connection.authType === "oauth";
const isApikeyAuth =
connection.authType === "apikey" || connection.authType === "api_key";
const isCookieAuth = connection.authType === "cookie";
const isApikeyEligible =
isApikeyAuth && USAGE_APIKEY_PROVIDERS.includes(connection.provider);
(isApikeyAuth || isCookieAuth) && USAGE_APIKEY_PROVIDERS.includes(connection.provider);
if (!isOAuth && !isApikeyEligible) {
return Response.json({ message: "Usage not available for this connection" });

View File

@@ -0,0 +1,66 @@
import { NextResponse } from "next/server";
import {
getDistinctModels,
getDistinctApiKeys,
getDistinctStatuses,
} from "@/lib/requestDetailsDb";
import { getProviderConnections, getProviderNodes } from "@/lib/localDb";
import { AI_PROVIDERS, getProviderByAlias } from "@/shared/constants/providers";
import { isOpenAICompatibleProvider, isAnthropicCompatibleProvider } from "@/shared/constants/providers";
/**
* GET /api/usage/filters
* Returns option lists for the Usage → Details filter bar:
* models: distinct model ids seen in requestDetails
* apiKeys: distinct (masked) api keys seen in requestDetails
* statuses: distinct status values (success/error)
* connections: provider connections grouped by provider (for <optgroup>)
*/
export async function GET() {
try {
const [models, apiKeys, statuses, connections, nodes] = await Promise.all([
getDistinctModels(),
getDistinctApiKeys(),
getDistinctStatuses(),
getProviderConnections(),
getProviderNodes(),
]);
const nodeNameMap = {};
for (const node of nodes || []) {
if (node.id && node.name) nodeNameMap[node.id] = node.name;
}
const byProvider = {};
for (const c of connections) {
const isCompatible =
isOpenAICompatibleProvider(c.provider) ||
isAnthropicCompatibleProvider(c.provider);
const baseName = isCompatible
? c.name || nodeNameMap[c.provider] || c.providerSpecificData?.nodeName
: c.name;
const providerName =
getProviderByAlias(c.provider)?.name || AI_PROVIDERS[c.provider]?.name || c.provider || "unknown";
const key = c.provider || "unknown";
if (!byProvider[key]) {
byProvider[key] = { provider: key, name: providerName, accounts: [] };
}
byProvider[key].accounts.push({
id: c.id,
label: baseName || providerName,
});
}
const connectionsList = Object.values(byProvider).sort((a, b) =>
a.name.localeCompare(b.name)
);
return NextResponse.json({ models, apiKeys, statuses, connections: connectionsList });
} catch (error) {
console.error("[API] Failed to get usage filters:", error);
return NextResponse.json(
{ error: "Failed to fetch usage filters" },
{ status: 500 }
);
}
}

View File

@@ -17,6 +17,7 @@ export async function GET(request) {
const model = searchParams.get("model");
const connectionId = searchParams.get("connectionId");
const status = searchParams.get("status");
const apiKey = searchParams.get("apiKey");
const startDate = searchParams.get("startDate");
const endDate = searchParams.get("endDate");
@@ -43,6 +44,7 @@ export async function GET(request) {
if (model) filter.model = model;
if (connectionId) filter.connectionId = connectionId;
if (status) filter.status = status;
if (apiKey) filter.apiKey = apiKey;
if (startDate) filter.startDate = startDate;
if (endDate) filter.endDate = endDate;

View File

@@ -29,7 +29,7 @@ export {
// API keys
export {
getApiKeys, getApiKeyById, createApiKey, updateApiKey, deleteApiKey, validateApiKey,
getApiKeys, getApiKeyById, createApiKey, updateApiKey, importApiKey, deleteApiKey, validateApiKey,
} from "./repos/apiKeysRepo.js";
// Combos
@@ -65,6 +65,7 @@ export {
// Request details
export {
saveRequestDetail, getRequestDetails, getRequestDetailById, getDistinctProviders,
getDistinctModels, getDistinctApiKeys, getDistinctStatuses,
} from "./repos/requestDetailsRepo.js";
// Export/import full DB

View File

@@ -0,0 +1,14 @@
// Adds an indexed apiKey column to requestDetails so the Details tab can
// filter by (masked) API key without parsing the JSON data blob.
export default {
version: 2,
name: "request-details-apikey",
up(db) {
// Idempotent: fresh DBs already get the column from TABLES schema.
const existing = db.all(`PRAGMA table_info(requestDetails)`) || [];
if (!existing.some((c) => c.name === "apiKey")) {
db.exec(`ALTER TABLE requestDetails ADD COLUMN apiKey TEXT`);
}
db.exec(`CREATE INDEX IF NOT EXISTS idx_rd_apikey ON requestDetails(apiKey)`);
},
};

View File

@@ -2,8 +2,9 @@
// Each migration: { version: number, name: string, up(db): void }
// Versions MUST be unique and monotonically increasing.
import m001 from "./001-initial.js";
import m002 from "./002-request-details-apikey.js";
export const MIGRATIONS = [m001].sort((a, b) => a.version - b.version);
export const MIGRATIONS = [m001, m002].sort((a, b) => a.version - b.version);
export function latestVersion() {
return MIGRATIONS.length ? MIGRATIONS[MIGRATIONS.length - 1].version : 0;

View File

@@ -61,6 +61,31 @@ export async function updateApiKey(id, data) {
return result;
}
export async function importApiKey(name, keyValue) {
if (!keyValue?.trim()) throw new Error("Key value is required");
const db = await getAdapter();
// Check for duplicates
const existing = db.get(`SELECT id FROM apiKeys WHERE key = ?`, [keyValue.trim()]);
if (existing) {
throw new Error("This API key already exists in the system");
}
const apiKey = {
id: uuidv4(),
name: name?.trim() || "Imported Key",
key: keyValue.trim(),
machineId: null,
isActive: true,
createdAt: new Date().toISOString(),
};
db.run(
`INSERT INTO apiKeys(id, key, name, machineId, isActive, createdAt) VALUES(?, ?, ?, ?, ?, ?)`,
[apiKey.id, apiKey.key, apiKey.name, apiKey.machineId, 1, apiKey.createdAt]
);
return apiKey;
}
export async function deleteApiKey(id) {
const db = await getAdapter();
const res = db.run(`DELETE FROM apiKeys WHERE id = ?`, [id]);

View File

@@ -3,71 +3,92 @@ import { getAdapter } from "../driver.js";
import { parseJson, stringifyJson } from "../helpers/jsonCol.js";
function rowToCombo(row) {
if (!row) return null;
return {
id: row.id,
name: row.name,
kind: row.kind,
models: parseJson(row.models, []),
createdAt: row.createdAt,
updatedAt: row.updatedAt,
};
if (!row) return null;
return {
id: row.id,
name: row.name,
kind: row.kind,
models: parseJson(row.models, []),
enabled: row.enabled === 1 || row.enabled === true,
createdAt: row.createdAt,
updatedAt: row.updatedAt,
};
}
export async function getCombos() {
const db = await getAdapter();
const rows = db.all(`SELECT * FROM combos ORDER BY createdAt ASC`);
return rows.map(rowToCombo);
const db = await getAdapter();
const rows = db.all(`SELECT * FROM combos ORDER BY createdAt ASC`);
return rows.map(rowToCombo);
}
export async function getComboById(id) {
const db = await getAdapter();
const row = db.get(`SELECT * FROM combos WHERE id = ?`, [id]);
return rowToCombo(row);
const db = await getAdapter();
const row = db.get(`SELECT * FROM combos WHERE id = ?`, [id]);
return rowToCombo(row);
}
export async function getComboByName(name) {
const db = await getAdapter();
const row = db.get(`SELECT * FROM combos WHERE name = ?`, [name]);
return rowToCombo(row);
const db = await getAdapter();
const row = db.get(`SELECT * FROM combos WHERE name = ?`, [name]);
return rowToCombo(row);
}
export async function createCombo(data) {
const db = await getAdapter();
const now = new Date().toISOString();
const combo = {
id: uuidv4(),
name: data.name,
kind: data.kind || null,
models: data.models || [],
createdAt: now,
updatedAt: now,
};
db.run(
`INSERT INTO combos(id, name, kind, models, createdAt, updatedAt) VALUES(?, ?, ?, ?, ?, ?)`,
[combo.id, combo.name, combo.kind, stringifyJson(combo.models), combo.createdAt, combo.updatedAt]
);
return combo;
const db = await getAdapter();
const now = new Date().toISOString();
const combo = {
id: uuidv4(),
name: data.name,
kind: data.kind || null,
models: data.models || [],
enabled: data.enabled !== false,
createdAt: now,
updatedAt: now,
};
db.run(
`INSERT INTO combos(id, name, kind, models, enabled, createdAt, updatedAt) VALUES(?, ?, ?, ?, ?, ?, ?)`,
[
combo.id,
combo.name,
combo.kind,
stringifyJson(combo.models),
combo.enabled !== false ? 1 : 0,
combo.createdAt,
combo.updatedAt,
],
);
return combo;
}
export async function updateCombo(id, data) {
const db = await getAdapter();
let result = null;
db.transaction(() => {
const row = db.get(`SELECT * FROM combos WHERE id = ?`, [id]);
if (!row) return;
const merged = { ...rowToCombo(row), ...data, updatedAt: new Date().toISOString() };
db.run(
`UPDATE combos SET name = ?, kind = ?, models = ?, updatedAt = ? WHERE id = ?`,
[merged.name, merged.kind, stringifyJson(merged.models || []), merged.updatedAt, id]
);
result = merged;
});
return result;
const db = await getAdapter();
let result = null;
db.transaction(() => {
const row = db.get(`SELECT * FROM combos WHERE id = ?`, [id]);
if (!row) return;
const merged = {
...rowToCombo(row),
...data,
updatedAt: new Date().toISOString(),
};
db.run(
`UPDATE combos SET name = ?, kind = ?, models = ?, enabled = ?, updatedAt = ? WHERE id = ?`,
[
merged.name,
merged.kind,
stringifyJson(merged.models || []),
merged.enabled ? 1 : 0,
merged.updatedAt,
id,
],
);
result = merged;
});
return result;
}
export async function deleteCombo(id) {
const db = await getAdapter();
const res = db.run(`DELETE FROM combos WHERE id = ?`, [id]);
return (res?.changes ?? 0) > 0;
const db = await getAdapter();
const res = db.run(`DELETE FROM combos WHERE id = ?`, [id]);
return (res?.changes ?? 0) > 0;
}

View File

@@ -109,6 +109,7 @@ async function flushToDatabase() {
connectionId: item.connectionId || null,
timestamp: item.timestamp,
status: item.status || null,
apiKey: item.apiKey || null,
latency: item.latency || {},
tokens: item.tokens || {},
request: truncateField(item.request, config.maxJsonSize),
@@ -119,8 +120,8 @@ async function flushToDatabase() {
};
db.run(
`INSERT INTO requestDetails(id, timestamp, provider, model, connectionId, status, data) VALUES(?, ?, ?, ?, ?, ?, ?) ON CONFLICT(id) DO UPDATE SET timestamp = excluded.timestamp, provider = excluded.provider, model = excluded.model, connectionId = excluded.connectionId, status = excluded.status, data = excluded.data`,
[record.id, record.timestamp, record.provider, record.model, record.connectionId, record.status, stringifyJson(record)]
`INSERT INTO requestDetails(id, timestamp, provider, model, connectionId, status, apiKey, data) VALUES(?, ?, ?, ?, ?, ?, ?, ?) ON CONFLICT(id) DO UPDATE SET timestamp = excluded.timestamp, provider = excluded.provider, model = excluded.model, connectionId = excluded.connectionId, status = excluded.status, apiKey = excluded.apiKey, data = excluded.data`,
[record.id, record.timestamp, record.provider, record.model, record.connectionId, record.status, record.apiKey, stringifyJson(record)]
);
}
@@ -168,6 +169,7 @@ export async function getRequestDetails(filter = {}) {
if (filter.model) { conds.push("model = ?"); params.push(filter.model); }
if (filter.connectionId) { conds.push("connectionId = ?"); params.push(filter.connectionId); }
if (filter.status) { conds.push("status = ?"); params.push(filter.status); }
if (filter.apiKey) { conds.push("apiKey = ?"); params.push(filter.apiKey); }
if (filter.startDate) { conds.push("timestamp >= ?"); params.push(new Date(filter.startDate).toISOString()); }
if (filter.endDate) { conds.push("timestamp <= ?"); params.push(new Date(filter.endDate).toISOString()); }
@@ -198,6 +200,24 @@ export async function getDistinctProviders() {
return rows.map((r) => r.provider);
}
export async function getDistinctModels() {
const db = await getAdapter();
const rows = db.all(`SELECT DISTINCT model FROM requestDetails WHERE model IS NOT NULL AND model != '' ORDER BY model ASC`);
return rows.map((r) => r.model);
}
export async function getDistinctApiKeys() {
const db = await getAdapter();
const rows = db.all(`SELECT DISTINCT apiKey FROM requestDetails WHERE apiKey IS NOT NULL AND apiKey != '' ORDER BY apiKey ASC`);
return rows.map((r) => r.apiKey);
}
export async function getDistinctStatuses() {
const db = await getAdapter();
const rows = db.all(`SELECT DISTINCT status FROM requestDetails WHERE status IS NOT NULL AND status != '' ORDER BY status ASC`);
return rows.map((r) => r.status);
}
export async function getRequestDetailById(id) {
const db = await getAdapter();
const row = db.get(`SELECT data FROM requestDetails WHERE id = ?`, [id]);

View File

@@ -13,6 +13,8 @@ const DEFAULT_SETTINGS = {
tailscaleUrl: "",
stickyRoundRobinLimit: 3,
providerStrategies: {},
providerTimeouts: {},
defaultTimeoutMs: null,
quotaVisibility: {},
comboStrategy: "fallback",
comboStickyRoundRobinLimit: 1,

File diff suppressed because it is too large Load Diff

View File

@@ -3,7 +3,7 @@
// pre-change safety backup in migrate.js: when the stored version is lower,
// one lightweight DB backup is taken before applying schema changes. Forgetting
// to bump only skips that backup — it does NOT break the additive auto-sync.
export const SCHEMA_VERSION = 1;
export const SCHEMA_VERSION = 2;
export const PRAGMA_SQL = `
PRAGMA journal_mode = WAL;
@@ -19,143 +19,146 @@ PRAGMA busy_timeout = 5000;
// auto-add missing tables/columns/indexes after versioned migrations.
// For destructive changes (drop/rename/type-change), write a migration file.
export const TABLES = {
_meta: {
columns: {
key: "TEXT PRIMARY KEY",
value: "TEXT NOT NULL",
},
},
settings: {
columns: {
id: "INTEGER PRIMARY KEY CHECK (id = 1)",
data: "TEXT NOT NULL",
},
},
providerConnections: {
columns: {
id: "TEXT PRIMARY KEY",
provider: "TEXT NOT NULL",
authType: "TEXT NOT NULL",
name: "TEXT",
email: "TEXT",
priority: "INTEGER",
isActive: "INTEGER DEFAULT 1",
data: "TEXT NOT NULL",
createdAt: "TEXT NOT NULL",
updatedAt: "TEXT NOT NULL",
},
indexes: [
"CREATE INDEX IF NOT EXISTS idx_pc_provider ON providerConnections(provider)",
"CREATE INDEX IF NOT EXISTS idx_pc_provider_active ON providerConnections(provider, isActive)",
"CREATE INDEX IF NOT EXISTS idx_pc_priority ON providerConnections(provider, priority)",
],
},
providerNodes: {
columns: {
id: "TEXT PRIMARY KEY",
type: "TEXT",
name: "TEXT",
data: "TEXT NOT NULL",
createdAt: "TEXT NOT NULL",
updatedAt: "TEXT NOT NULL",
},
indexes: ["CREATE INDEX IF NOT EXISTS idx_pn_type ON providerNodes(type)"],
},
proxyPools: {
columns: {
id: "TEXT PRIMARY KEY",
isActive: "INTEGER DEFAULT 1",
testStatus: "TEXT",
data: "TEXT NOT NULL",
createdAt: "TEXT NOT NULL",
updatedAt: "TEXT NOT NULL",
},
indexes: [
"CREATE INDEX IF NOT EXISTS idx_pp_active ON proxyPools(isActive)",
"CREATE INDEX IF NOT EXISTS idx_pp_status ON proxyPools(testStatus)",
],
},
apiKeys: {
columns: {
id: "TEXT PRIMARY KEY",
key: "TEXT UNIQUE NOT NULL",
name: "TEXT",
machineId: "TEXT",
isActive: "INTEGER DEFAULT 1",
createdAt: "TEXT NOT NULL",
},
indexes: ["CREATE INDEX IF NOT EXISTS idx_ak_key ON apiKeys(key)"],
},
combos: {
columns: {
id: "TEXT PRIMARY KEY",
name: "TEXT UNIQUE NOT NULL",
kind: "TEXT",
models: "TEXT NOT NULL",
createdAt: "TEXT NOT NULL",
updatedAt: "TEXT NOT NULL",
},
indexes: ["CREATE INDEX IF NOT EXISTS idx_combo_name ON combos(name)"],
},
kv: {
columns: {
scope: "TEXT NOT NULL",
key: "TEXT NOT NULL",
value: "TEXT NOT NULL",
},
primaryKey: "PRIMARY KEY (scope, key)",
indexes: ["CREATE INDEX IF NOT EXISTS idx_kv_scope ON kv(scope)"],
},
usageHistory: {
columns: {
id: "INTEGER PRIMARY KEY AUTOINCREMENT",
timestamp: "TEXT NOT NULL",
provider: "TEXT",
model: "TEXT",
connectionId: "TEXT",
apiKey: "TEXT",
endpoint: "TEXT",
promptTokens: "INTEGER DEFAULT 0",
completionTokens: "INTEGER DEFAULT 0",
cost: "REAL DEFAULT 0",
status: "TEXT",
tokens: "TEXT",
meta: "TEXT",
},
indexes: [
"CREATE INDEX IF NOT EXISTS idx_uh_ts ON usageHistory(timestamp DESC)",
"CREATE INDEX IF NOT EXISTS idx_uh_provider ON usageHistory(provider)",
"CREATE INDEX IF NOT EXISTS idx_uh_model ON usageHistory(model)",
"CREATE INDEX IF NOT EXISTS idx_uh_conn ON usageHistory(connectionId)",
],
},
usageDaily: {
columns: {
dateKey: "TEXT PRIMARY KEY",
data: "TEXT NOT NULL",
},
},
requestDetails: {
columns: {
id: "TEXT PRIMARY KEY",
timestamp: "TEXT NOT NULL",
provider: "TEXT",
model: "TEXT",
connectionId: "TEXT",
status: "TEXT",
data: "TEXT NOT NULL",
},
indexes: [
"CREATE INDEX IF NOT EXISTS idx_rd_ts ON requestDetails(timestamp DESC)",
"CREATE INDEX IF NOT EXISTS idx_rd_provider ON requestDetails(provider)",
"CREATE INDEX IF NOT EXISTS idx_rd_model ON requestDetails(model)",
"CREATE INDEX IF NOT EXISTS idx_rd_conn ON requestDetails(connectionId)",
],
},
_meta: {
columns: {
key: "TEXT PRIMARY KEY",
value: "TEXT NOT NULL",
},
},
settings: {
columns: {
id: "INTEGER PRIMARY KEY CHECK (id = 1)",
data: "TEXT NOT NULL",
},
},
providerConnections: {
columns: {
id: "TEXT PRIMARY KEY",
provider: "TEXT NOT NULL",
authType: "TEXT NOT NULL",
name: "TEXT",
email: "TEXT",
priority: "INTEGER",
isActive: "INTEGER DEFAULT 1",
data: "TEXT NOT NULL",
createdAt: "TEXT NOT NULL",
updatedAt: "TEXT NOT NULL",
},
indexes: [
"CREATE INDEX IF NOT EXISTS idx_pc_provider ON providerConnections(provider)",
"CREATE INDEX IF NOT EXISTS idx_pc_provider_active ON providerConnections(provider, isActive)",
"CREATE INDEX IF NOT EXISTS idx_pc_priority ON providerConnections(provider, priority)",
],
},
providerNodes: {
columns: {
id: "TEXT PRIMARY KEY",
type: "TEXT",
name: "TEXT",
data: "TEXT NOT NULL",
createdAt: "TEXT NOT NULL",
updatedAt: "TEXT NOT NULL",
},
indexes: ["CREATE INDEX IF NOT EXISTS idx_pn_type ON providerNodes(type)"],
},
proxyPools: {
columns: {
id: "TEXT PRIMARY KEY",
isActive: "INTEGER DEFAULT 1",
testStatus: "TEXT",
data: "TEXT NOT NULL",
createdAt: "TEXT NOT NULL",
updatedAt: "TEXT NOT NULL",
},
indexes: [
"CREATE INDEX IF NOT EXISTS idx_pp_active ON proxyPools(isActive)",
"CREATE INDEX IF NOT EXISTS idx_pp_status ON proxyPools(testStatus)",
],
},
apiKeys: {
columns: {
id: "TEXT PRIMARY KEY",
key: "TEXT UNIQUE NOT NULL",
name: "TEXT",
machineId: "TEXT",
isActive: "INTEGER DEFAULT 1",
createdAt: "TEXT NOT NULL",
},
indexes: ["CREATE INDEX IF NOT EXISTS idx_ak_key ON apiKeys(key)"],
},
combos: {
columns: {
id: "TEXT PRIMARY KEY",
name: "TEXT UNIQUE NOT NULL",
kind: "TEXT",
models: "TEXT NOT NULL",
enabled: "INTEGER NOT NULL DEFAULT 1",
createdAt: "TEXT NOT NULL",
updatedAt: "TEXT NOT NULL",
},
indexes: ["CREATE INDEX IF NOT EXISTS idx_combo_name ON combos(name)"],
},
kv: {
columns: {
scope: "TEXT NOT NULL",
key: "TEXT NOT NULL",
value: "TEXT NOT NULL",
},
primaryKey: "PRIMARY KEY (scope, key)",
indexes: ["CREATE INDEX IF NOT EXISTS idx_kv_scope ON kv(scope)"],
},
usageHistory: {
columns: {
id: "INTEGER PRIMARY KEY AUTOINCREMENT",
timestamp: "TEXT NOT NULL",
provider: "TEXT",
model: "TEXT",
connectionId: "TEXT",
apiKey: "TEXT",
endpoint: "TEXT",
promptTokens: "INTEGER DEFAULT 0",
completionTokens: "INTEGER DEFAULT 0",
cost: "REAL DEFAULT 0",
status: "TEXT",
tokens: "TEXT",
meta: "TEXT",
},
indexes: [
"CREATE INDEX IF NOT EXISTS idx_uh_ts ON usageHistory(timestamp DESC)",
"CREATE INDEX IF NOT EXISTS idx_uh_provider ON usageHistory(provider)",
"CREATE INDEX IF NOT EXISTS idx_uh_model ON usageHistory(model)",
"CREATE INDEX IF NOT EXISTS idx_uh_conn ON usageHistory(connectionId)",
],
},
usageDaily: {
columns: {
dateKey: "TEXT PRIMARY KEY",
data: "TEXT NOT NULL",
},
},
requestDetails: {
columns: {
id: "TEXT PRIMARY KEY",
timestamp: "TEXT NOT NULL",
provider: "TEXT",
model: "TEXT",
connectionId: "TEXT",
status: "TEXT",
apiKey: "TEXT",
data: "TEXT NOT NULL",
},
indexes: [
"CREATE INDEX IF NOT EXISTS idx_rd_ts ON requestDetails(timestamp DESC)",
"CREATE INDEX IF NOT EXISTS idx_rd_provider ON requestDetails(provider)",
"CREATE INDEX IF NOT EXISTS idx_rd_model ON requestDetails(model)",
"CREATE INDEX IF NOT EXISTS idx_rd_conn ON requestDetails(connectionId)",
"CREATE INDEX IF NOT EXISTS idx_rd_apikey ON requestDetails(apiKey)",
],
},
};
export function buildCreateTableSql(name, def) {
const cols = Object.entries(def.columns).map(([k, v]) => `${k} ${v}`);
if (def.primaryKey) cols.push(def.primaryKey);
return `CREATE TABLE IF NOT EXISTS ${name} (${cols.join(", ")})`;
const cols = Object.entries(def.columns).map(([k, v]) => `${k} ${v}`);
if (def.primaryKey) cols.push(def.primaryKey);
return `CREATE TABLE IF NOT EXISTS ${name} (${cols.join(", ")})`;
}

View File

@@ -10,7 +10,7 @@ export {
createProviderNode, updateProviderNode, deleteProviderNode,
getProxyPools, getProxyPoolById,
createProxyPool, updateProxyPool, deleteProxyPool,
getApiKeys, getApiKeyById, createApiKey, updateApiKey, deleteApiKey, validateApiKey,
getApiKeys, getApiKeyById, createApiKey, updateApiKey, importApiKey, deleteApiKey, validateApiKey,
getCombos, getComboById, getComboByName,
createCombo, updateCombo, deleteCombo,
getModelAliases, setModelAlias, deleteModelAlias,

View File

@@ -1,4 +1,5 @@
// Shim → re-export from new SQLite-based DB layer (src/lib/db/)
export {
saveRequestDetail, getRequestDetails, getRequestDetailById, getDistinctProviders,
getDistinctModels, getDistinctApiKeys, getDistinctStatuses,
} from "@/lib/db/index.js";

View File

@@ -4,4 +4,5 @@ export {
saveRequestUsage, getUsageHistory, getUsageStats, getChartData,
appendRequestLog, getRecentLogs,
saveRequestDetail, getRequestDetails, getRequestDetailById,
getDistinctModels, getDistinctApiKeys, getDistinctStatuses,
} from "@/lib/db/index.js";

View File

@@ -32,6 +32,7 @@ export {
setMitmAliasAll,
getApiKeys,
createApiKey,
importApiKey,
deleteApiKey,
validateApiKey,
isCloudEnabled,

File diff suppressed because it is too large Load Diff

View File

@@ -171,7 +171,7 @@ export default function EditConnectionModal({ isOpen, connection, proxyPools, on
if (providerRegions && region) {
updates.providerSpecificData = buildRegionSpecificData();
}
await onSave(updates);
} finally {
setSaving(false);
@@ -202,6 +202,8 @@ export default function EditConnectionModal({ isOpen, connection, proxyPools, on
onChange={(e) => setFormData({ ...formData, priority: Number.parseInt(e.target.value, 10) || 1 })}
/>
{!isOAuth && (
<>
<div className="flex gap-2">

View File

@@ -21,7 +21,7 @@ export default function Modal({
md: "max-w-md",
lg: "max-w-lg",
xl: "max-w-xl",
full: "max-w-4xl",
full: "max-w-6xl",
};
useEffect(() => {
@@ -99,7 +99,10 @@ export default function Modal({
)}
{/* Body */}
<div className="p-6 max-h-[calc(85vh-100px)] overflow-y-auto custom-scrollbar">{children}</div>
<div className={cn(
"p-6 overflow-y-auto custom-scrollbar",
size === "full" ? "max-h-[calc(92vh-100px)]" : "max-h-[calc(85vh-100px)]"
)}>{children}</div>
{/* Footer */}
{footer && (

File diff suppressed because it is too large Load Diff

View File

@@ -37,6 +37,7 @@ export { default as SegmentedControl } from "./SegmentedControl";
export { default as Tooltip } from "./Tooltip";
export { default as ProviderInfoCard } from "./ProviderInfoCard";
export { default as CapacityBadges } from "./CapacityBadges";
export { default as ApiExplorerModal } from "./ApiExplorerModal";
// Layouts
export * from "./layouts";

View File

@@ -0,0 +1,871 @@
/**
* Public AI API catalog for the dashboard API Explorer.
* Single source of truth for endpoint docs + form fields used by ApiExplorerModal.
*/
/** @typedef {"string"|"number"|"boolean"|"select"|"textarea"|"json"|"file"|"messages"} FieldType */
/**
* @typedef {Object} ApiField
* @property {string} key
* @property {string} label
* @property {FieldType} type
* @property {boolean} [required]
* @property {string} [description]
* @property {*} [default]
* @property {string[]} [options]
* @property {number} [min]
* @property {number} [max]
* @property {number} [step]
* @property {string} [placeholder]
* @property {string} [accept] - file accept attr
* @property {"body"|"query"|"header"} [location] - default body
* @property {boolean} [advanced] - hide behind "Advanced"
*/
/**
* @typedef {Object} ApiEndpoint
* @property {string} id
* @property {string} method
* @property {string} path
* @property {string} label
* @property {string} description
* @property {string} icon
* @property {string} category
* @property {"json"|"multipart"|"none"} contentType
* @property {boolean} [needsModel]
* @property {string} [modelsKind] - /v1/models/{kind} query
* @property {ApiField[]} fields
* @property {string} [defaultResponse]
* @property {boolean} [supportsStream]
*/
/** @type {ApiEndpoint[]} */
export const API_CATALOG = [
{
id: "models",
method: "GET",
path: "/v1/models",
label: "List Models",
description: "List chat/LLM models (default). Combos appear with owned_by: combo.",
icon: "list_alt",
category: "Discovery",
contentType: "none",
fields: [],
defaultResponse: `{\n "object": "list",\n "data": [\n { "id": "openai/gpt-4o", "object": "model", "owned_by": "openai" }\n ]\n}`,
},
{
id: "models-kind",
method: "GET",
path: "/v1/models/{kind}",
label: "List Models by Kind",
description: "List models for a capability kind: image, tts, stt, embedding, web, etc.",
icon: "category",
category: "Discovery",
contentType: "none",
fields: [
{
key: "kind",
label: "Kind",
type: "select",
required: true,
location: "path",
default: "image",
options: ["image", "tts", "stt", "embedding", "web", "image-to-text"],
description: "Path segment: /v1/models/{kind}. Chat models use GET /v1/models (no kind).",
},
],
defaultResponse: `{\n "object": "list",\n "data": [{ "id": "openai/dall-e-3", "object": "model", "owned_by": "openai" }]\n}`,
},
{
id: "models-info",
method: "GET",
path: "/v1/models/info",
label: "Model Info",
description: "Per-model metadata: params, options, contextWindow, capabilities.",
icon: "info",
category: "Discovery",
contentType: "none",
fields: [
{
key: "id",
label: "Model ID",
type: "string",
required: true,
location: "query",
placeholder: "openai/gpt-4o",
description: "Query: ?id=provider/model",
},
],
defaultResponse: `{\n "id": "openai/gpt-4o",\n "kind": "llm",\n "endpoint": "/v1/chat/completions",\n "contextWindow": 128000\n}`,
},
{
id: "chat-completions",
method: "POST",
path: "/v1/chat/completions",
label: "Chat Completions",
description: "OpenAI-compatible chat/code generation with streaming + auto-fallback.",
icon: "chat",
category: "Chat",
contentType: "json",
needsModel: true,
modelsKind: "llm",
supportsStream: true,
fields: [
{
key: "messages",
label: "Messages",
type: "messages",
required: true,
default: [{ role: "user", content: "Hello! Say hi in one sentence." }],
description: "Array of { role, content }",
},
{
key: "stream",
label: "Stream",
type: "boolean",
default: false,
description: "SSE streaming response",
},
{
key: "temperature",
label: "Temperature",
type: "number",
default: "",
min: 0,
max: 2,
step: 0.1,
advanced: true,
},
{
key: "top_p",
label: "Top P",
type: "number",
default: "",
min: 0,
max: 1,
step: 0.05,
advanced: true,
},
{
key: "max_tokens",
label: "Max Tokens",
type: "number",
default: "",
min: 1,
advanced: true,
},
{
key: "max_completion_tokens",
label: "Max Completion Tokens",
type: "number",
default: "",
min: 1,
advanced: true,
},
{
key: "presence_penalty",
label: "Presence Penalty",
type: "number",
default: "",
min: -2,
max: 2,
step: 0.1,
advanced: true,
},
{
key: "frequency_penalty",
label: "Frequency Penalty",
type: "number",
default: "",
min: -2,
max: 2,
step: 0.1,
advanced: true,
},
{
key: "seed",
label: "Seed",
type: "number",
default: "",
advanced: true,
},
{
key: "stop",
label: "Stop",
type: "json",
default: "",
placeholder: '["\\n"] or "stop"',
advanced: true,
description: "String or array of stop sequences",
},
{
key: "tools",
label: "Tools",
type: "json",
default: "",
placeholder: "[{ type: \"function\", function: {...} }]",
advanced: true,
},
{
key: "tool_choice",
label: "Tool Choice",
type: "json",
default: "",
placeholder: '"auto" | "none" | { type, function }',
advanced: true,
},
{
key: "response_format",
label: "Response Format",
type: "json",
default: "",
placeholder: '{ "type": "json_object" }',
advanced: true,
},
{
key: "reasoning_effort",
label: "Reasoning Effort",
type: "select",
default: "",
options: ["", "low", "medium", "high"],
advanced: true,
},
{
key: "thinking",
label: "Thinking",
type: "json",
default: "",
placeholder: '{ "type": "enabled", "budget_tokens": 1024 }',
advanced: true,
},
{
key: "n",
label: "n",
type: "number",
default: "",
min: 1,
max: 8,
advanced: true,
},
{
key: "user",
label: "User",
type: "string",
default: "",
advanced: true,
},
],
defaultResponse: `{\n "id": "chatcmpl-...",\n "object": "chat.completion",\n "choices": [{ "message": { "role": "assistant", "content": "..." }, "finish_reason": "stop" }],\n "usage": { "prompt_tokens": 8, "completion_tokens": 2, "total_tokens": 10 }\n}`,
},
{
id: "messages",
method: "POST",
path: "/v1/messages",
label: "Messages (Anthropic)",
description: "Anthropic Messages API format. Same models, Claude-compatible tools/clients.",
icon: "forum",
category: "Chat",
contentType: "json",
needsModel: true,
modelsKind: "llm",
supportsStream: true,
fields: [
{
key: "messages",
label: "Messages",
type: "messages",
required: true,
default: [{ role: "user", content: "Hello! Say hi in one sentence." }],
},
{
key: "max_tokens",
label: "Max Tokens",
type: "number",
required: true,
default: 1024,
min: 1,
},
{
key: "stream",
label: "Stream",
type: "boolean",
default: false,
},
{
key: "system",
label: "System",
type: "textarea",
default: "",
placeholder: "You are a helpful assistant.",
advanced: true,
},
{
key: "temperature",
label: "Temperature",
type: "number",
default: "",
min: 0,
max: 1,
step: 0.1,
advanced: true,
},
{
key: "top_p",
label: "Top P",
type: "number",
default: "",
min: 0,
max: 1,
step: 0.05,
advanced: true,
},
{
key: "top_k",
label: "Top K",
type: "number",
default: "",
min: 0,
advanced: true,
},
{
key: "stop_sequences",
label: "Stop Sequences",
type: "json",
default: "",
placeholder: '["\\n\\nHuman:"]',
advanced: true,
},
{
key: "tools",
label: "Tools",
type: "json",
default: "",
advanced: true,
},
{
key: "tool_choice",
label: "Tool Choice",
type: "json",
default: "",
advanced: true,
},
{
key: "thinking",
label: "Thinking",
type: "json",
default: "",
placeholder: '{ "type": "enabled", "budget_tokens": 1024 }',
advanced: true,
},
{
key: "anthropic-version",
label: "anthropic-version",
type: "string",
location: "header",
default: "2023-06-01",
advanced: true,
},
],
defaultResponse: `{\n "id": "msg_...",\n "type": "message",\n "role": "assistant",\n "content": [{ "type": "text", "text": "Hello!" }],\n "stop_reason": "end_turn",\n "usage": { "input_tokens": 8, "output_tokens": 2 }\n}`,
},
{
id: "responses",
method: "POST",
path: "/v1/responses",
label: "Responses API",
description: "OpenAI Responses format (Codex / OpenClaw). input can be string or items[].",
icon: "reply",
category: "Chat",
contentType: "json",
needsModel: true,
modelsKind: "llm",
supportsStream: true,
fields: [
{
key: "input",
label: "Input",
type: "textarea",
required: true,
default: "Hello! Say hi in one sentence.",
description: "String or JSON array of input items",
},
{
key: "stream",
label: "Stream",
type: "boolean",
default: false,
},
{
key: "instructions",
label: "Instructions",
type: "textarea",
default: "",
advanced: true,
},
{
key: "temperature",
label: "Temperature",
type: "number",
default: "",
min: 0,
max: 2,
step: 0.1,
advanced: true,
},
{
key: "max_output_tokens",
label: "Max Output Tokens",
type: "number",
default: "",
min: 1,
advanced: true,
},
{
key: "tools",
label: "Tools",
type: "json",
default: "",
advanced: true,
},
{
key: "reasoning",
label: "Reasoning",
type: "json",
default: "",
placeholder: '{ "effort": "medium" }',
advanced: true,
},
],
defaultResponse: `{\n "id": "resp_...",\n "object": "response",\n "status": "completed",\n "output": [{ "type": "message", "content": [{ "type": "output_text", "text": "..." }] }]\n}`,
},
{
id: "count-tokens",
method: "POST",
path: "/v1/messages/count_tokens",
label: "Count Tokens",
description: "Estimate token count for Anthropic-style messages (local estimate).",
icon: "tag",
category: "Chat",
contentType: "json",
needsModel: true,
modelsKind: "llm",
fields: [
{
key: "messages",
label: "Messages",
type: "messages",
required: true,
default: [{ role: "user", content: "Hello world" }],
},
{
key: "system",
label: "System",
type: "textarea",
default: "",
advanced: true,
},
],
defaultResponse: `{\n "input_tokens": 12\n}`,
},
{
id: "embeddings",
method: "POST",
path: "/v1/embeddings",
label: "Embeddings",
description: "Vector embeddings for RAG / semantic search.",
icon: "scatter_plot",
category: "Embeddings",
contentType: "json",
needsModel: true,
modelsKind: "embedding",
fields: [
{
key: "input",
label: "Input",
type: "textarea",
required: true,
default: "Hello world",
description: "String or JSON array of strings",
},
{
key: "encoding_format",
label: "Encoding",
type: "select",
default: "float",
options: ["float", "base64"],
},
{
key: "dimensions",
label: "Dimensions",
type: "number",
default: "",
min: 1,
advanced: true,
description: "OpenAI text-embedding-3-* only",
},
],
defaultResponse: `{\n "object": "list",\n "data": [{ "object": "embedding", "index": 0, "embedding": [0.01, -0.02, "..."] }],\n "usage": { "prompt_tokens": 2, "total_tokens": 2 }\n}`,
},
{
id: "images-generations",
method: "POST",
path: "/v1/images/generations",
label: "Image Generation",
description: "Text-to-image (DALL·E, Imagen, FLUX, Codex, MiniMax…). Supports binary output.",
icon: "brush",
category: "Media",
contentType: "json",
needsModel: true,
modelsKind: "image",
fields: [
{
key: "prompt",
label: "Prompt",
type: "textarea",
required: true,
default: "A cute cat wearing a hat, watercolor style",
},
{
key: "n",
label: "n",
type: "number",
default: 1,
min: 1,
max: 4,
},
{
key: "size",
label: "Size",
type: "select",
default: "auto",
options: ["auto", "1024x1024", "1024x1536", "1536x1024", "1024x1792", "1792x1024"],
},
{
key: "aspect_ratio",
label: "Aspect Ratio",
type: "select",
default: "",
options: ["", "auto", "1:1", "16:9", "9:16", "4:3", "3:2", "2:3", "9:19.5", "20:9"],
advanced: true,
},
{
key: "resolution",
label: "Resolution",
type: "select",
default: "",
options: ["", "1k", "2k"],
advanced: true,
},
{
key: "quality",
label: "Quality",
type: "select",
default: "auto",
options: ["auto", "low", "medium", "high", "standard", "hd"],
advanced: true,
},
{
key: "background",
label: "Background",
type: "select",
default: "auto",
options: ["auto", "transparent", "opaque"],
advanced: true,
},
{
key: "style",
label: "Style",
type: "select",
default: "",
options: ["", "vivid", "natural"],
advanced: true,
},
{
key: "response_format",
label: "Response Format",
type: "select",
default: "url",
options: ["url", "b64_json", "binary"],
description: "binary uses ?response_format=binary (raw image bytes)",
},
{
key: "output_format",
label: "Codec",
type: "select",
default: "png",
options: ["png", "jpeg", "webp"],
advanced: true,
},
{
key: "image",
label: "Reference Image",
type: "image",
default: "",
placeholder: "https://... or upload an image",
advanced: true,
description: "Edit / img2img when provider supports it. Paste URL or upload file (stored as data URL).",
accept: "image/*",
},
],
defaultResponse: `{\n "created": 1735000000,\n "data": [{ "url": "https://..." }]\n}`,
},
{
id: "audio-speech",
method: "POST",
path: "/v1/audio/speech",
label: "Text to Speech",
description: "OpenAI / ElevenLabs / Edge / Google / Deepgram voices → audio bytes.",
icon: "record_voice_over",
category: "Media",
contentType: "json",
needsModel: true,
modelsKind: "tts",
fields: [
{
key: "input",
label: "Text",
type: "textarea",
required: true,
default: "Hello from 9Router!",
},
{
key: "response_format",
label: "Response Format",
type: "select",
location: "query",
default: "mp3",
options: ["mp3", "json"],
description: "mp3 = raw audio; json = { audio: base64, format }",
},
{
key: "language",
label: "Language Hint",
type: "string",
default: "",
placeholder: "en, vi, ...",
advanced: true,
description: "Optional language hint (e.g. Gemini)",
},
{
key: "voice",
label: "Voice",
type: "string",
default: "",
placeholder: "alloy (provider-dependent)",
advanced: true,
},
{
key: "speed",
label: "Speed",
type: "number",
default: "",
min: 0.25,
max: 4,
step: 0.05,
advanced: true,
},
],
defaultResponse: "(binary audio/mp3)",
},
{
id: "audio-voices",
method: "GET",
path: "/v1/audio/voices",
label: "List Voices",
description: "List TTS voices for elevenlabs, edge-tts, deepgram, inworld, local-device.",
icon: "voice_selection",
category: "Media",
contentType: "none",
fields: [
{
key: "provider",
label: "Provider",
type: "select",
required: true,
location: "query",
default: "edge-tts",
options: ["edge-tts", "elevenlabs", "deepgram", "inworld", "local-device"],
},
{
key: "lang",
label: "Language",
type: "string",
location: "query",
default: "",
placeholder: "vi, en, ...",
advanced: true,
},
],
defaultResponse: `{\n "object": "list",\n "data": [{ "id": "...", "model": "edge-tts/vi-VN-HoaiMyNeural" }]\n}`,
},
{
id: "audio-transcriptions",
method: "POST",
path: "/v1/audio/transcriptions",
label: "Speech to Text",
description: "Whisper-compatible multipart transcription.",
icon: "mic",
category: "Media",
contentType: "multipart",
needsModel: true,
modelsKind: "stt",
fields: [
{
key: "file",
label: "Audio File",
type: "file",
required: true,
accept: "audio/*,.mp3,.wav,.m4a,.webm,.ogg,.flac",
},
{
key: "language",
label: "Language",
type: "string",
default: "",
placeholder: "en, vi, ...",
},
{
key: "prompt",
label: "Prompt",
type: "textarea",
default: "",
advanced: true,
},
{
key: "response_format",
label: "Response Format",
type: "select",
default: "json",
options: ["json", "text", "verbose_json", "srt", "vtt"],
},
{
key: "temperature",
label: "Temperature",
type: "number",
default: "",
min: 0,
max: 1,
step: 0.1,
advanced: true,
},
],
defaultResponse: `{\n "text": "..."\n}`,
},
{
id: "search",
method: "POST",
path: "/v1/search",
label: "Web Search",
description: "Tavily / Exa / Brave / Serper / SearXNG / Google PSE / You.com…",
icon: "search",
category: "Web",
contentType: "json",
needsModel: true,
modelsKind: "web",
fields: [
{
key: "query",
label: "Query",
type: "textarea",
required: true,
default: "What is the latest news about AI?",
},
{
key: "max_results",
label: "Max Results",
type: "number",
default: 5,
min: 1,
max: 100,
},
{
key: "search_type",
label: "Search Type",
type: "select",
default: "web",
options: ["web", "news"],
},
{
key: "country",
label: "Country",
type: "string",
default: "",
advanced: true,
},
{
key: "language",
label: "Language",
type: "string",
default: "",
advanced: true,
},
{
key: "time_range",
label: "Time Range",
type: "string",
default: "",
placeholder: "day, week, month, year",
advanced: true,
},
{
key: "domain_filter",
label: "Domain Filter",
type: "string",
default: "",
placeholder: "example.com",
advanced: true,
},
],
defaultResponse: `{\n "provider": "tavily",\n "query": "...",\n "results": [{ "title": "...", "url": "...", "snippet": "..." }]\n}`,
},
{
id: "web-fetch",
method: "POST",
path: "/v1/web/fetch",
label: "Web Fetch",
description: "URL → markdown / text / HTML via Firecrawl, Jina, Tavily, Exa.",
icon: "language",
category: "Web",
contentType: "json",
needsModel: true,
modelsKind: "web",
fields: [
{
key: "url",
label: "URL",
type: "string",
required: true,
default: "https://9router.com",
placeholder: "https://example.com",
},
{
key: "format",
label: "Format",
type: "select",
default: "markdown",
options: ["markdown", "text", "html"],
},
{
key: "max_characters",
label: "Max Characters",
type: "number",
default: 0,
min: 0,
advanced: true,
description: "0 = no truncate",
},
],
defaultResponse: `{\n "provider": "jina-reader",\n "url": "...",\n "title": "...",\n "content": { "format": "markdown", "text": "..." }\n}`,
},
];
export const API_CATEGORIES = [...new Set(API_CATALOG.map((e) => e.category))];
export function getApiEndpointById(id) {
return API_CATALOG.find((e) => e.id === id) || null;
}
export function getApiEndpointsByCategory(category) {
return API_CATALOG.filter((e) => e.category === category);
}

View File

@@ -89,13 +89,16 @@ export async function handleEmbeddings(request) {
log.info("ROUTING", `Provider: ${provider}, Model: ${model}`);
}
// Optional pin to a specific connection (dashboard test / client override)
const preferredConnectionId = request?.headers?.get("x-connection-id") || null;
// Credential + fallback loop (mirrors handleChat)
const excludeConnectionIds = new Set();
let lastError = null;
let lastStatus = null;
while (true) {
const credentials = await getProviderCredentials(provider, excludeConnectionIds, model);
const credentials = await getProviderCredentials(provider, excludeConnectionIds, model, { preferredConnectionId });
// All accounts unavailable
if (!credentials || credentials.allRateLimited) {
@@ -153,6 +156,10 @@ export async function handleEmbeddings(request) {
const { shouldFallback } = await markAccountUnavailable(credentials.connectionId, result.status, result.error, provider, model);
if (shouldFallback) {
if (preferredConnectionId) {
log.warn("AUTH", `Pinned account ${credentials.connectionName} unavailable (${result.status}), no fallback`);
return result.response;
}
log.warn("AUTH", `Account ${credentials.connectionName} unavailable (${result.status}), trying fallback`);
excludeConnectionIds.add(credentials.connectionId);
lastError = result.error;

View File

@@ -128,15 +128,21 @@ export async function getProviderCredentials(provider, excludeConnectionIds = nu
const strategy = providerOverride.fallbackStrategy || settings.fallbackStrategy || "fill-first";
let connection;
// Pin to preferred connection if specified and available
// Strict pin: only use the requested connection (no strategy fallback).
// Allows model-locked accounts so explicit tests can still hit that account.
if (preferredConnectionId) {
connection = availableConnections.find((c) => c.id === preferredConnectionId);
if (connection) {
log.info("AUTH", `${provider} | pinned to ${connection.id?.slice(0, 8)} (${connection.name || connection.email || "unnamed"})`);
connection = connections.find((c) => c.id === preferredConnectionId);
if (!connection || excludeSet.has(connection.id)) {
log.warn(
"AUTH",
`${provider} | preferred connection ${preferredConnectionId.slice(0, 8)} not found/active or excluded`,
);
return null;
}
}
if (connection) {
// skip strategy
log.info(
"AUTH",
`${provider} | pinned to ${connection.id?.slice(0, 8)} (${connection.name || connection.email || "unnamed"})`,
);
} else if (strategy === "round-robin") {
const stickyLimit = providerOverride.stickyRoundRobinLimit || settings.stickyRoundRobinLimit || 3;

View File

@@ -1,81 +1,103 @@
// Re-export from open-sse with localDb integration
import { getModelAliases, getComboByName, getProviderNodes } from "@/lib/localDb";
import { parseModel as parseModelCore, resolveModelAliasFromMap, getModelInfoCore } from "open-sse/services/model.js";
import {
getModelAliases,
getComboByName,
getProviderNodes,
} from "@/lib/localDb";
import {
parseModel as parseModelCore,
resolveModelAliasFromMap,
getModelInfoCore,
} from "open-sse/services/model.js";
import REGISTRY from "open-sse/providers/registry/index.js";
// Local provider alias overrides (HMR-friendly, applied on top of open-sse map)
const LOCAL_PROVIDER_ALIASES = {
xmtp: "xiaomi-tokenplan",
"xiaomi-tokenplan": "xiaomi-tokenplan",
xmtp: "xiaomi-tokenplan",
"xiaomi-tokenplan": "xiaomi-tokenplan",
};
const RESERVED_PROVIDER_PREFIXES = new Set(Object.keys(LOCAL_PROVIDER_ALIASES));
for (const entry of REGISTRY) {
RESERVED_PROVIDER_PREFIXES.add(entry.id);
if (entry.alias) RESERVED_PROVIDER_PREFIXES.add(entry.alias);
for (const alias of entry.aliases || []) RESERVED_PROVIDER_PREFIXES.add(alias);
RESERVED_PROVIDER_PREFIXES.add(entry.id);
if (entry.alias) RESERVED_PROVIDER_PREFIXES.add(entry.alias);
for (const alias of entry.aliases || [])
RESERVED_PROVIDER_PREFIXES.add(alias);
}
export function parseModel(modelStr) {
const parsed = parseModelCore(modelStr);
if (parsed?.providerAlias && LOCAL_PROVIDER_ALIASES[parsed.providerAlias]) {
return { ...parsed, provider: LOCAL_PROVIDER_ALIASES[parsed.providerAlias] };
}
return parsed;
const parsed = parseModelCore(modelStr);
if (parsed?.providerAlias && LOCAL_PROVIDER_ALIASES[parsed.providerAlias]) {
return {
...parsed,
provider: LOCAL_PROVIDER_ALIASES[parsed.providerAlias],
};
}
return parsed;
}
/**
* Resolve model alias from localDb
*/
export async function resolveModelAlias(alias) {
const aliases = await getModelAliases();
return resolveModelAliasFromMap(alias, aliases);
const aliases = await getModelAliases();
return resolveModelAliasFromMap(alias, aliases);
}
/**
* Get full model info (parse or resolve)
*/
export async function getModelInfo(modelStr) {
const parsed = parseModel(modelStr);
const parsed = parseModel(modelStr);
if (!parsed.isAlias) {
// Provider-node prefixes are user-defined. They must not override built-in
// provider ids/aliases such as `cf`, `cloudflare-ai`, `openai`, or `hf`.
if (!RESERVED_PROVIDER_PREFIXES.has(parsed.providerAlias)) {
const openaiNodes = await getProviderNodes({ type: "openai-compatible" });
const matchedOpenAI = openaiNodes.find((node) => node.prefix === parsed.providerAlias);
if (matchedOpenAI) {
return { provider: matchedOpenAI.id, model: parsed.model };
}
if (!parsed.isAlias) {
// Provider-node prefixes are user-defined. They must not override built-in
// provider ids/aliases such as `cf`, `cloudflare-ai`, `openai`, or `hf`.
if (!RESERVED_PROVIDER_PREFIXES.has(parsed.providerAlias)) {
const openaiNodes = await getProviderNodes({ type: "openai-compatible" });
const matchedOpenAI = openaiNodes.find(
(node) => node.prefix === parsed.providerAlias,
);
if (matchedOpenAI) {
return { provider: matchedOpenAI.id, model: parsed.model };
}
const anthropicNodes = await getProviderNodes({ type: "anthropic-compatible" });
const matchedAnthropic = anthropicNodes.find((node) => node.prefix === parsed.providerAlias);
if (matchedAnthropic) {
return { provider: matchedAnthropic.id, model: parsed.model };
}
const anthropicNodes = await getProviderNodes({
type: "anthropic-compatible",
});
const matchedAnthropic = anthropicNodes.find(
(node) => node.prefix === parsed.providerAlias,
);
if (matchedAnthropic) {
return { provider: matchedAnthropic.id, model: parsed.model };
}
const embeddingNodes = await getProviderNodes({ type: "custom-embedding" });
const matchedEmbedding = embeddingNodes.find((node) => node.prefix === parsed.providerAlias);
if (matchedEmbedding) {
return { provider: matchedEmbedding.id, model: parsed.model };
}
}
return {
provider: parsed.provider,
model: parsed.model
};
}
const embeddingNodes = await getProviderNodes({
type: "custom-embedding",
});
const matchedEmbedding = embeddingNodes.find(
(node) => node.prefix === parsed.providerAlias,
);
if (matchedEmbedding) {
return { provider: matchedEmbedding.id, model: parsed.model };
}
}
return {
provider: parsed.provider,
model: parsed.model,
};
}
// Check if this is a combo name before resolving as alias
// This prevents combo names from being incorrectly routed to providers
const combo = await getComboByName(parsed.model);
if (combo) {
// Return null provider to signal this should be handled as combo
// The caller (handleChat) will detect this and handle it as combo
return { provider: null, model: parsed.model };
}
// Check if this is a combo name before resolving as alias
// This prevents combo names from being incorrectly routed to providers
const combo = await getComboByName(parsed.model);
if (combo && combo.enabled !== false) {
// Return null provider to signal this should be handled as combo
// The caller (handleChat) will detect this and handle it as combo
return { provider: null, model: parsed.model };
}
return getModelInfoCore(modelStr, getModelAliases);
return getModelInfoCore(modelStr, getModelAliases);
}
/**
@@ -83,12 +105,17 @@ export async function getModelInfo(modelStr) {
* @returns {Promise<string[]|null>} Array of models or null if not a combo
*/
export async function getComboModels(modelStr) {
// Only check if it's not in provider/model format
if (modelStr.includes("/")) return null;
// Only check if it's not in provider/model format
if (modelStr.includes("/")) return null;
const combo = await getComboByName(modelStr);
if (combo && combo.models && combo.models.length > 0) {
return combo.models;
}
return null;
const combo = await getComboByName(modelStr);
if (
combo &&
combo.enabled !== false &&
combo.models &&
combo.models.length > 0
) {
return combo.models;
}
return null;
}

View File

@@ -7,11 +7,11 @@ const LOG_LEVELS = {
ERROR: 3
};
const LEVEL = LOG_LEVELS[process.env.LOG_LEVEL?.toUpperCase?.()] ?? LOG_LEVELS.INFO;
// Runtime log level. Defaults from LOG_LEVEL env, but can be changed at runtime
// via setLogLevel (dashboard Settings → Logging). WARN/ERROR hide the noisy
// INFO request lines (▶ POST / 📊 DONE / [COMBO] / [CHAT] ...) in production.
let LEVEL = LOG_LEVELS[process.env.LOG_LEVEL?.toUpperCase?.()] ?? LOG_LEVELS.INFO;
export function setLogLevel(level) {
const normalized = String(level || "").toUpperCase();
if (Object.prototype.hasOwnProperty.call(LOG_LEVELS, normalized)) {
@@ -49,6 +49,7 @@ export function tagForSession(seed) {
}
// Print one correlated line: [time] tag symbol message
// Visible at INFO and below (hidden by WARN/ERROR log levels in production).
export function line(tag, symbol, message) {
if (LEVEL > LOG_LEVELS.INFO) return;
console.log(`[${formatTime()}] ${tag} ${symbol} ${message}`);
@@ -117,12 +118,14 @@ export function request(method, path, extra) {
}
export function response(status, duration, extra) {
if (LEVEL > LOG_LEVELS.INFO) return;
const icon = status < 400 ? "📤" : "💥";
const dataStr = extra ? ` ${formatData(extra)}` : "";
console.log(`[${formatTime()}] ${icon} ${status} (${duration}ms)${dataStr}`);
}
export function stream(event, data) {
if (LEVEL > LOG_LEVELS.INFO) return;
const dataStr = data ? ` ${formatData(data)}` : "";
console.log(`[${formatTime()}] 🌊 [STREAM] ${event}${dataStr}`);
}

41
test-commandcode-vision.sh Executable file
View File

@@ -0,0 +1,41 @@
#!/usr/bin/env bash
# Test vision support for the commandcode provider in 9router.
#
# Usage:
# ./test-commandcode-vision.sh [base_url] [model]
# BASE_URL=http://localhost:20128 ./test-commandcode-vision.sh
#
# Defaults: base_url=http://localhost:20128, model=commandcode/Qwen/Qwen3.6-Max-Preview
set -euo pipefail
BASE_URL="${1:-${BASE_URL:-http://localhost:20128}}"
MODEL="${2:-${MODEL:-commandcode/Qwen/Qwen3.6-Max-Preview}}"
# 1x1 red pixel PNG (base64) — enough to prove the image reaches the model.
PNG_B64="iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAYAAAAfFcSJAAAADUlEQVR42mP8z8BQDwAEhQGAhKmMIQAAAABJRU5ErkJggg=="
echo "==> POST $BASE_URL/v1/chat/completions"
echo "==> model: $MODEL"
echo
curl -sS -N "$BASE_URL/v1/chat/completions" \
-H "Content-Type: application/json" \
-d "$(
cat <<JSON
{
"model": "$MODEL",
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": "What color is this image? Reply with just the color name."},
{"type": "image_url", "image_url": {"url": "data:image/png;base64,$PNG_B64"}}
]
}
],
"stream": false
}
JSON
)"
echo

View File

@@ -4,73 +4,133 @@ import "./registerAll.js";
import { translateRequest } from "../../open-sse/translator/index.js";
import { FORMATS } from "../../open-sse/translator/formats.js";
const O2G = (body) => translateRequest(FORMATS.OPENAI, FORMATS.GEMINI, "m", body, true, null, "gemini");
const O2C = (body) => translateRequest(FORMATS.OPENAI, FORMATS.CURSOR, "m", body, true, null, "cursor");
const O2CC = (body) => translateRequest(FORMATS.OPENAI, FORMATS.COMMANDCODE, "m", body, true, null, "commandcode");
const O2G = (body) =>
translateRequest(
FORMATS.OPENAI,
FORMATS.GEMINI,
"m",
body,
true,
null,
"gemini",
);
const O2C = (body) =>
translateRequest(
FORMATS.OPENAI,
FORMATS.CURSOR,
"m",
body,
true,
null,
"cursor",
);
const O2CC = (body) =>
translateRequest(
FORMATS.OPENAI,
FORMATS.COMMANDCODE,
"m",
body,
true,
null,
"commandcode",
);
describe("OpenAI → Gemini", () => {
// openai-to-gemini.js:92-96 — each system message overwrites systemInstruction → only last kept
// KNOWN BUG
it.fails("multiple system messages are all kept", () => {
const out = O2G({
messages: [
{ role: "system", content: "RULE_ONE" },
{ role: "system", content: "RULE_TWO" },
{ role: "user", content: "hi" },
],
});
expect(JSON.stringify(out.systemInstruction), "earlier system lost").toContain("RULE_ONE");
});
// openai-to-gemini.js:92-96 — each system message overwrites systemInstruction → only last kept
// KNOWN BUG
it.fails("multiple system messages are all kept", () => {
const out = O2G({
messages: [
{ role: "system", content: "RULE_ONE" },
{ role: "system", content: "RULE_TWO" },
{ role: "user", content: "hi" },
],
});
expect(
JSON.stringify(out.systemInstruction),
"earlier system lost",
).toContain("RULE_ONE");
});
});
describe("OpenAI → Cursor", () => {
// openai-to-cursor.js:12-24 — image content fully dropped (text only)
// KNOWN BUG
it.fails("image content is preserved", () => {
const out = O2C({
messages: [{ role: "user", content: [
{ type: "text", text: "look" },
{ type: "image_url", image_url: { url: "data:image/png;base64,AAAA" } },
] }],
});
expect(JSON.stringify(out), "image dropped").toContain("AAAA");
});
// openai-to-cursor.js:12-24 — image content fully dropped (text only)
// KNOWN BUG
it.fails("image content is preserved", () => {
const out = O2C({
messages: [
{
role: "user",
content: [
{ type: "text", text: "look" },
{
type: "image_url",
image_url: { url: "data:image/png;base64,AAAA" },
},
],
},
],
});
expect(JSON.stringify(out), "image dropped").toContain("AAAA");
});
// openai-to-cursor.js:179 — max_tokens hardcoded to 32000
// KNOWN BUG
it.fails("respects client max_tokens", () => {
const out = O2C({ max_tokens: 200, messages: [{ role: "user", content: "hi" }] });
expect(out.max_tokens).toBe(200);
});
// openai-to-cursor.js:179 — max_tokens hardcoded to 32000
// KNOWN BUG
it.fails("respects client max_tokens", () => {
const out = O2C({
max_tokens: 200,
messages: [{ role: "user", content: "hi" }],
});
expect(out.max_tokens).toBe(200);
});
});
describe("OpenAI → CommandCode", () => {
// openai-to-commandcode.js:53-57 — safeParseJson returns {} on bad JSON (args silently lost)
// KNOWN BUG
it.fails("malformed tool arguments are not silently emptied", () => {
const out = O2CC({
messages: [
{ role: "user", content: "go" },
{ role: "assistant", content: "", tool_calls: [
{ id: "c1", type: "function", function: { name: "f", arguments: "{bad" } },
] },
{ role: "tool", tool_call_id: "c1", content: "r" },
],
});
const asst = out.params.messages.find((m) => m.role === "assistant");
const call = asst.content.find((b) => b.type === "tool-call");
expect(Object.keys(call.input).length, "arguments silently dropped to {}").toBeGreaterThan(0);
});
// openai-to-commandcode.js:53-57 — safeParseJson returns {} on bad JSON (args silently lost)
// KNOWN BUG
it.fails("malformed tool arguments are not silently emptied", () => {
const out = O2CC({
messages: [
{ role: "user", content: "go" },
{
role: "assistant",
content: "",
tool_calls: [
{
id: "c1",
type: "function",
function: { name: "f", arguments: "{bad" },
},
],
},
{ role: "tool", tool_call_id: "c1", content: "r" },
],
});
const asst = out.params.messages.find((m) => m.role === "assistant");
const call = asst.content.find((b) => b.type === "tool-call");
expect(
Object.keys(call.input).length,
"arguments silently dropped to {}",
).toBeGreaterThan(0);
});
// openai-to-commandcode.js:41-42 — image becomes "[image omitted]"
// KNOWN BUG
it.fails("image content is preserved", () => {
const out = O2CC({
messages: [{ role: "user", content: [
{ type: "text", text: "look" },
{ type: "image_url", image_url: { url: "data:image/png;base64,BBBB" } },
] }],
});
expect(JSON.stringify(out), "image omitted").toContain("BBBB");
});
// openai-to-commandcode.js — image blocks now map to {type:"image", image:"data:..."}
// FIXED: was "[image omitted]", now preserved as data URI
it("image content is preserved", () => {
const out = O2CC({
messages: [
{
role: "user",
content: [
{ type: "text", text: "look" },
{
type: "image_url",
image_url: { url: "data:image/png;base64,BBBB" },
},
],
},
],
});
expect(JSON.stringify(out), "image omitted").toContain("BBBB");
});
});

View File

@@ -0,0 +1,126 @@
/**
* Unit tests for the CommandCode executor early-error peek.
*
* The upstream emits AI SDK v5 NDJSON over an HTTP 200 stream, so a terminal
* `{"type":"error"}` event is invisible to the normal `response.ok` success
* check. `peekForUpstreamError` reads the first events before committing the
* response: an error event → non-ok Response (fallback can kick in); otherwise
* the buffered bytes are re-emitted and streaming proceeds as before.
*/
import { describe, it, expect } from "vitest";
import { peekForUpstreamError } from "../../open-sse/executors/commandcode.js";
const encoder = new TextEncoder();
function ndjsonResponse(lines) {
const body = new ReadableStream({
start(controller) {
for (const line of lines) controller.enqueue(encoder.encode(line + "\n"));
controller.close();
},
});
return new Response(body, {
status: 200,
headers: { "content-type": "text/event-stream" },
});
}
describe("commandcode executor — early-error peek", () => {
it("returns 502 when the first meaningful event is an error", async () => {
const res = await peekForUpstreamError(
ndjsonResponse([
'{"type":"error","error":{"type":"server_error","message":"Network connection lost."}}',
]),
"model",
);
expect(res.status).toBe(502);
const body = await res.json();
expect(body.error.message).toBe("Network connection lost.");
expect(body.error.type).toBe("server_error");
});
it("detects an error event even when metadata events arrive first", async () => {
const res = await peekForUpstreamError(
ndjsonResponse([
'{"type":"start"}',
'{"type":"start-step"}',
'{"type":"error","error":{"type":"server_error","message":"Network connection lost."}}',
]),
"model",
);
expect(res.status).toBe(502);
const body = await res.json();
expect(body.error.message).toContain("Network connection lost");
});
it("commits and streams normally when the first meaningful event is content", async () => {
const res = await peekForUpstreamError(
ndjsonResponse([
'{"type":"start"}',
'{"type":"text-delta","text":"hi there"}',
'{"type":"finish"}',
]),
"model",
);
expect(res.status).toBe(200);
const text = await res.text();
expect(text).toContain('"content":"hi there"');
expect(text).not.toContain("[CommandCode error:");
});
it("commits when the stream ends without any event", async () => {
const res = await peekForUpstreamError(ndjsonResponse([]), "model");
expect(res.status).toBe(200);
await res.body.cancel();
});
it("commits (does not hang) when no event arrives before the peek timeout", async () => {
const stalled = new Response(new ReadableStream({ start() {} }), {
status: 200,
headers: { "content-type": "text/event-stream" },
});
const res = await peekForUpstreamError(stalled, "model", { timeoutMs: 50 });
expect(res.status).toBe(200);
await res.body.cancel();
});
it("does not hang when the request signal aborts during the peek", async () => {
const controller = new AbortController();
const stalled = new Response(new ReadableStream({ start() {} }), {
status: 200,
headers: { "content-type": "text/event-stream" },
});
setTimeout(() => controller.abort(new Error("client gone")), 10);
const res = await peekForUpstreamError(stalled, "model", {
signal: controller.signal,
timeoutMs: 2000,
});
expect(res.status).toBe(200);
await res.body.cancel();
});
it("re-emits raw bytes so a multi-byte char split across the peek boundary survives", async () => {
// chunk1 ends mid-é (0xC3); the peek commits on the first complete line
// (text-delta "hi") while the decoder still holds 0xC3. Re-emission must
// use RAW bytes — re-encoding decoded text would replace 0xC3 with U+FFFD.
const first = Buffer.from(
'{"type":"text-delta","text":"hi"}\n{"type":"text-delta","text":"caf',
);
const chunk1 = new Uint8Array([...first, 0xc3]);
const rest = new Uint8Array([0xa9, 0x22, 0x7d, 0x0a]); // é"}\n
const body = new ReadableStream({
start(c) {
c.enqueue(chunk1);
c.enqueue(rest);
c.close();
},
});
const res = await peekForUpstreamError(
new Response(body, { status: 200 }),
"m",
);
const text = await res.text();
expect(text).toContain("café");
expect(text).not.toContain("\uFFFD");
});
});

View File

@@ -12,116 +12,152 @@ import { describe, it, expect } from "vitest";
import { commandCodeToOpenAIResponse } from "../../open-sse/translator/response/commandcode-to-openai.js";
function feed(events) {
const state = {};
const all = [];
for (const e of events) {
const out = commandCodeToOpenAIResponse(JSON.stringify(e), state);
if (out) for (const c of out) all.push(c);
}
return { state, chunks: all };
const state = {};
const all = [];
for (const e of events) {
const out = commandCodeToOpenAIResponse(JSON.stringify(e), state);
if (out) for (const c of out) all.push(c);
}
return { state, chunks: all };
}
describe("commandcode-to-openai — text-delta", () => {
it("emits assistant role on first delta then content-only", () => {
const { chunks } = feed([
{ type: "text-delta", text: "Hello" },
{ type: "text-delta", text: " world" },
]);
expect(chunks[0].choices[0].delta.role).toBe("assistant");
expect(chunks[0].choices[0].delta.content).toBe("Hello");
expect(chunks[1].choices[0].delta.role).toBeUndefined();
expect(chunks[1].choices[0].delta.content).toBe(" world");
});
it("emits assistant role on first delta then content-only", () => {
const { chunks } = feed([
{ type: "text-delta", text: "Hello" },
{ type: "text-delta", text: " world" },
]);
expect(chunks[0].choices[0].delta.role).toBe("assistant");
expect(chunks[0].choices[0].delta.content).toBe("Hello");
expect(chunks[1].choices[0].delta.role).toBeUndefined();
expect(chunks[1].choices[0].delta.content).toBe(" world");
});
});
describe("commandcode-to-openai — reasoning-delta", () => {
it("maps reasoning-delta to reasoning_content delta", () => {
const { chunks } = feed([
{ type: "reasoning-delta", text: "thinking..." },
]);
expect(chunks[0].choices[0].delta.reasoning_content).toBe("thinking...");
});
it("maps reasoning-delta to reasoning_content delta", () => {
const { chunks } = feed([{ type: "reasoning-delta", text: "thinking..." }]);
expect(chunks[0].choices[0].delta.reasoning_content).toBe("thinking...");
});
});
describe("commandcode-to-openai — tool-input-* with id field (live schema)", () => {
it("registers tool index using event.id (NOT toolCallId)", () => {
const { chunks } = feed([
{ type: "tool-input-start", id: "call_X", toolName: "Bash" },
{ type: "tool-input-delta", id: "call_X", delta: "{\"cmd" },
{ type: "tool-input-delta", id: "call_X", delta: "\":\"ls\"}" },
]);
it("registers tool index using event.id (NOT toolCallId)", () => {
const { chunks } = feed([
{ type: "tool-input-start", id: "call_X", toolName: "Bash" },
{ type: "tool-input-delta", id: "call_X", delta: '{"cmd' },
{ type: "tool-input-delta", id: "call_X", delta: '":"ls"}' },
]);
// First chunk emits tool_calls with id
const startChunk = chunks[0].choices[0].delta.tool_calls[0];
expect(startChunk.id).toBe("call_X");
expect(startChunk.function.name).toBe("Bash");
// First chunk emits tool_calls with id
const startChunk = chunks[0].choices[0].delta.tool_calls[0];
expect(startChunk.id).toBe("call_X");
expect(startChunk.function.name).toBe("Bash");
// Subsequent deltas accumulate arguments
expect(chunks[1].choices[0].delta.tool_calls[0].function.arguments).toBe("{\"cmd");
expect(chunks[2].choices[0].delta.tool_calls[0].function.arguments).toBe("\":\"ls\"}");
});
// Subsequent deltas accumulate arguments
expect(chunks[1].choices[0].delta.tool_calls[0].function.arguments).toBe(
'{"cmd',
);
expect(chunks[2].choices[0].delta.tool_calls[0].function.arguments).toBe(
'":"ls"}',
);
});
it("ignores tool-input-delta when id is unknown (no prior start)", () => {
const { chunks } = feed([
{ type: "tool-input-delta", id: "unknown", delta: "x" },
]);
expect(chunks.length).toBe(0);
});
it("ignores tool-input-delta when id is unknown (no prior start)", () => {
const { chunks } = feed([
{ type: "tool-input-delta", id: "unknown", delta: "x" },
]);
expect(chunks.length).toBe(0);
});
});
describe("commandcode-to-openai — final tool-call event", () => {
it("does NOT re-emit tool_calls when tool-input-* deltas already fired", () => {
const { chunks } = feed([
{ type: "tool-input-start", id: "call_Y", toolName: "Write" },
{ type: "tool-input-delta", id: "call_Y", delta: "{\"file\":\"a\"}" },
{ type: "tool-call", toolCallId: "call_Y", toolName: "Write", input: { file: "a" } },
]);
// Should be exactly 2 chunks (start + delta), no duplicate from final tool-call
expect(chunks.length).toBe(2);
});
it("does NOT re-emit tool_calls when tool-input-* deltas already fired", () => {
const { chunks } = feed([
{ type: "tool-input-start", id: "call_Y", toolName: "Write" },
{ type: "tool-input-delta", id: "call_Y", delta: '{"file":"a"}' },
{
type: "tool-call",
toolCallId: "call_Y",
toolName: "Write",
input: { file: "a" },
},
]);
// Should be exactly 2 chunks (start + delta), no duplicate from final tool-call
expect(chunks.length).toBe(2);
});
it("emits a consolidated tool_calls when only the final tool-call event arrives", () => {
const { chunks } = feed([
{ type: "tool-call", toolCallId: "call_Z", toolName: "Read", input: { path: "/x" } },
]);
expect(chunks.length).toBe(1);
const tc = chunks[0].choices[0].delta.tool_calls[0];
expect(tc.id).toBe("call_Z");
expect(tc.function.name).toBe("Read");
expect(tc.function.arguments).toBe(JSON.stringify({ path: "/x" }));
});
it("emits a consolidated tool_calls when only the final tool-call event arrives", () => {
const { chunks } = feed([
{
type: "tool-call",
toolCallId: "call_Z",
toolName: "Read",
input: { path: "/x" },
},
]);
expect(chunks.length).toBe(1);
const tc = chunks[0].choices[0].delta.tool_calls[0];
expect(tc.id).toBe("call_Z");
expect(tc.function.name).toBe("Read");
expect(tc.function.arguments).toBe(JSON.stringify({ path: "/x" }));
});
});
describe("commandcode-to-openai — finish", () => {
it("emits a final chunk with finish_reason=tool_calls when finishReason is tool-calls", () => {
const { chunks } = feed([
{ type: "tool-input-start", id: "call_F", toolName: "Bash" },
{ type: "tool-input-delta", id: "call_F", delta: "{}" },
{ type: "finish-step", finishReason: "tool-calls" },
{ type: "finish" },
]);
const last = chunks[chunks.length - 1];
expect(last.choices[0].finish_reason).toBe("tool_calls");
});
it("emits a final chunk with finish_reason=tool_calls when finishReason is tool-calls", () => {
const { chunks } = feed([
{ type: "tool-input-start", id: "call_F", toolName: "Bash" },
{ type: "tool-input-delta", id: "call_F", delta: "{}" },
{ type: "finish-step", finishReason: "tool-calls" },
{ type: "finish" },
]);
const last = chunks[chunks.length - 1];
expect(last.choices[0].finish_reason).toBe("tool_calls");
});
it("includes usage on the final chunk when totalUsage provided", () => {
const { chunks } = feed([
{ type: "text-delta", text: "hi" },
{ type: "finish-step", finishReason: "stop", usage: { inputTokens: 10, outputTokens: 5, totalTokens: 15 } },
{ type: "finish", totalUsage: { inputTokens: 10, outputTokens: 5, totalTokens: 15 } },
]);
const last = chunks[chunks.length - 1];
expect(last.usage).toEqual({ prompt_tokens: 10, completion_tokens: 5, total_tokens: 15 });
});
it("includes usage on the final chunk when totalUsage provided", () => {
const { chunks } = feed([
{ type: "text-delta", text: "hi" },
{
type: "finish-step",
finishReason: "stop",
usage: { inputTokens: 10, outputTokens: 5, totalTokens: 15 },
},
{
type: "finish",
totalUsage: { inputTokens: 10, outputTokens: 5, totalTokens: 15 },
},
]);
const last = chunks[chunks.length - 1];
expect(last.usage).toEqual({
prompt_tokens: 10,
completion_tokens: 5,
total_tokens: 15,
});
});
});
describe("commandcode-to-openai — error event", () => {
it("stringifies object errors so client sees readable message", () => {
const { chunks } = feed([
{ type: "error", error: { type: "server_error", message: "Boom" } },
]);
const text = chunks[0].choices[0].delta.content;
expect(text).toContain("Boom");
expect(text).not.toContain("[object Object]");
});
it("emits an OpenAI-shaped error chunk instead of fake success content", () => {
const { chunks } = feed([
{ type: "error", error: { type: "server_error", message: "Boom" } },
]);
expect(chunks[0].error).toEqual({ message: "Boom", type: "server_error" });
expect(chunks[0].choices[0].delta.content).toBeUndefined();
expect(chunks[1].choices[0].finish_reason).toBe("stop");
expect(JSON.stringify(chunks)).not.toContain("[CommandCode error:");
});
it("keeps the stream terminal so clients do not hang waiting for more", () => {
const { chunks } = feed([
{ type: "start" },
{
type: "error",
error: { type: "server_error", message: "Network connection lost." },
},
]);
expect(chunks[0].error.message).toBe("Network connection lost.");
expect(chunks[1].choices[0].finish_reason).toBe("stop");
});
});

View File

@@ -0,0 +1,218 @@
import { describe, it, expect, vi, beforeEach } from "vitest";
vi.mock("../../open-sse/utils/proxyFetch.js", () => ({
proxyAwareFetch: vi.fn(),
}));
import { proxyAwareFetch } from "../../open-sse/utils/proxyFetch.js";
import { getUsageForProvider } from "../../open-sse/services/usage.js";
import {
USAGE_SUPPORTED_PROVIDERS,
USAGE_APIKEY_PROVIDERS,
} from "../../src/shared/constants/providers.js";
import { parseQuotaData } from "../../src/app/(dashboard)/dashboard/usage/components/ProviderLimits/utils.js";
const BASE = "https://api.commandcode.ai";
const WHOAMI_URL = `${BASE}/alpha/whoami`;
const CREDITS_URL = `${BASE}/alpha/billing/credits`;
const SUBS_URL = `${BASE}/alpha/billing/subscriptions`;
const SUMMARY_URL = `${BASE}/alpha/usage/summary`;
function jsonResponse(body, status = 200) {
return new Response(JSON.stringify(body), {
status,
headers: { "Content-Type": "application/json" },
});
}
const WHOAMI = { success: true, user: { id: "u1" }, org: null };
const CREDITS = {
credits: {
belowThreshold: false,
creditThreshold: 0,
monthlyCredits: 9.9,
purchasedCredits: 0,
freeCredits: 0,
},
windowLimits: {
limited: true,
exceeded: null,
fiveHour: { used: 0.05, cap: 3, exceeded: false, resetAt: 1785812386064 },
weekly: { used: 0.1, cap: 6, exceeded: false, resetAt: 1786379982640 },
},
};
const SUBS = {
success: true,
data: {
id: "sub_1",
status: "active",
orgId: null,
planId: "individual-go",
currentPeriodStart: "2026-08-03T16:38:16.000Z",
currentPeriodEnd: "2026-09-03T16:38:16.000Z",
},
};
const SUMMARY = {
totalCount: 61,
totalCost: 0.1,
totalCredits: 0.1,
totalMonthlyCredits: 0.1,
periodBasis: "billing-period",
};
describe("commandcode registry usage flags", () => {
it("is listed for apikey quota dashboard", () => {
expect(USAGE_SUPPORTED_PROVIDERS).toContain("commandcode");
expect(USAGE_APIKEY_PROVIDERS).toContain("commandcode");
});
});
describe("getUsageForProvider(commandcode)", () => {
beforeEach(() => {
vi.clearAllMocks();
});
it("fetches whoami → credits+subs → summary and maps windows + credits", async () => {
proxyAwareFetch
.mockResolvedValueOnce(jsonResponse(WHOAMI))
.mockResolvedValueOnce(jsonResponse(CREDITS))
.mockResolvedValueOnce(jsonResponse(SUBS))
.mockResolvedValueOnce(jsonResponse(SUMMARY));
const usage = await getUsageForProvider({
provider: "commandcode",
apiKey: "user_cc_test",
});
expect(usage.message).toBeUndefined();
expect(usage.plan).toBe("individual-go");
expect(usage.periodBasis).toBe("billing-period");
expect(proxyAwareFetch).toHaveBeenCalledTimes(4);
const [whoamiUrl, whoamiOpts] = proxyAwareFetch.mock.calls[0];
expect(whoamiUrl).toBe(WHOAMI_URL);
expect(whoamiOpts.headers.Authorization).toBe("Bearer user_cc_test");
// No org → credits/subscriptions called without orgId query
const creditsCall = proxyAwareFetch.mock.calls[1][0];
expect(creditsCall).toBe(CREDITS_URL);
// Summary uses currentPeriodStart as `since`
const summaryCall = proxyAwareFetch.mock.calls[3][0];
expect(summaryCall).toBe(
`${SUMMARY_URL}?since=${encodeURIComponent("2026-08-03T16:38:16.000Z")}`,
);
expect(usage.quotas["5-hour window"]).toMatchObject({
used: 0.05,
total: 3,
resetAt: new Date(1785812386064).toISOString(),
});
expect(usage.quotas["Weekly window"]).toMatchObject({
used: 0.1,
total: 6,
resetAt: new Date(1786379982640).toISOString(),
});
expect(usage.quotas["Monthly credits"]).toMatchObject({
used: 0.1,
total: 9.9,
});
});
it("adds orgId query when whoami returns an org", async () => {
proxyAwareFetch
.mockResolvedValueOnce(
jsonResponse({ success: true, org: { id: "org_1" } }),
)
.mockResolvedValueOnce(jsonResponse(CREDITS))
.mockResolvedValueOnce(jsonResponse(SUBS))
.mockResolvedValueOnce(jsonResponse(SUMMARY));
await getUsageForProvider({
provider: "commandcode",
apiKey: "user_cc_test",
});
expect(proxyAwareFetch.mock.calls[1][0]).toBe(`${CREDITS_URL}?orgId=org_1`);
expect(proxyAwareFetch.mock.calls[2][0]).toBe(`${SUBS_URL}?orgId=org_1`);
});
it("falls back to first-of-month since when subscription has no period start", async () => {
proxyAwareFetch
.mockResolvedValueOnce(jsonResponse(WHOAMI))
.mockResolvedValueOnce(jsonResponse(CREDITS))
.mockResolvedValueOnce(
jsonResponse({ success: true, data: { planId: "individual-go" } }),
)
.mockResolvedValueOnce(jsonResponse(SUMMARY));
await getUsageForProvider({
provider: "commandcode",
apiKey: "user_cc_test",
});
const since = new URL(proxyAwareFetch.mock.calls[3][0]).searchParams.get(
"since",
);
// firstOfMonth() is local-time based; assert the local date is the 1st.
const localDate = new Date(since);
expect(localDate.getDate()).toBe(1);
});
it("returns message on missing key / 401 / non-ok whoami", async () => {
const missing = await getUsageForProvider({ provider: "commandcode" });
expect(missing.message).toMatch(/credential/i);
expect(proxyAwareFetch).not.toHaveBeenCalled();
proxyAwareFetch.mockResolvedValueOnce(jsonResponse({ error: "no" }, 401));
const auth = await getUsageForProvider({
provider: "commandcode",
apiKey: "bad",
});
expect(auth.message).toMatch(/invalid|expired/i);
proxyAwareFetch.mockResolvedValueOnce(jsonResponse({ error: "x" }, 500));
const err = await getUsageForProvider({
provider: "commandcode",
apiKey: "bad",
});
expect(err.message).toMatch(/whoami/i);
});
});
describe("parseQuotaData(commandcode)", () => {
it("forwards remainingPercentage + unit for window/credit rows", () => {
const rows = parseQuotaData("commandcode", {
plan: "individual-go",
quotas: {
"5-hour window": {
used: 0.05,
total: 3,
remainingPercentage: 98.33,
resetAt: "2026-08-03T22:59:46.064Z",
unit: "$",
},
"Monthly credits": {
used: 0.1,
total: 9.9,
remainingPercentage: 98.99,
resetAt: null,
unit: "$",
},
},
});
expect(rows[0]).toMatchObject({
name: "5-hour window",
used: 0.05,
total: 3,
remainingPercentage: 98.33,
unit: "$",
});
expect(rows[1]).toMatchObject({
name: "Monthly credits",
used: 0.1,
total: 9.9,
remainingPercentage: 98.99,
});
});
});

View File

@@ -543,4 +543,146 @@ describe("handleImageGenerationCore", () => {
expect(result.success).toBe(true);
expect(onRequestSuccess).toHaveBeenCalledTimes(1);
});
it("generates image with xAI Imagine API", async () => {
global.fetch.mockResolvedValueOnce(
new Response(
JSON.stringify({
created: 1234567890,
data: [{ url: "https://example.com/xai-gen.png" }],
}),
{ status: 200, headers: { "Content-Type": "application/json" } }
)
);
const result = await handleImageGenerationCore({
body: {
prompt: "Mountain landscape at sunrise",
n: 2,
aspect_ratio: "16:9",
resolution: "2k",
response_format: "url",
},
modelInfo: { provider: "xai", model: "grok-imagine-image-quality" },
credentials: { apiKey: "xai-key" },
log: null,
});
expect(result.success).toBe(true);
expect(global.fetch).toHaveBeenCalledWith(
"https://api.x.ai/v1/images/generations",
expect.objectContaining({
method: "POST",
headers: expect.objectContaining({
"Content-Type": "application/json",
Authorization: "Bearer xai-key",
}),
})
);
const reqBody = JSON.parse(global.fetch.mock.calls[0][1].body);
expect(reqBody).toEqual({
model: "grok-imagine-image-quality",
prompt: "Mountain landscape at sunrise",
n: 2,
response_format: "url",
aspect_ratio: "16:9",
resolution: "2k",
});
});
it("maps OpenAI size to aspect_ratio for xAI generation", async () => {
global.fetch.mockResolvedValueOnce(
new Response(
JSON.stringify({ created: 1, data: [{ url: "https://example.com/xai.png" }] }),
{ status: 200, headers: { "Content-Type": "application/json" } }
)
);
await handleImageGenerationCore({
body: { prompt: "city", size: "1792x1024" },
modelInfo: { provider: "xai", model: "grok-imagine-image-quality" },
credentials: { accessToken: "oauth-token" },
log: null,
});
const reqBody = JSON.parse(global.fetch.mock.calls[0][1].body);
expect(reqBody.aspect_ratio).toBe("16:9");
expect(reqBody.size).toBeUndefined();
expect(global.fetch.mock.calls[0][1].headers.Authorization).toBe("Bearer oauth-token");
});
it("edits a single image via xAI /images/edits", async () => {
global.fetch.mockResolvedValueOnce(
new Response(
JSON.stringify({ created: 1, data: [{ url: "https://example.com/xai-edit.png" }] }),
{ status: 200, headers: { "Content-Type": "application/json" } }
)
);
const result = await handleImageGenerationCore({
body: {
prompt: "Render this as a pencil sketch",
image: "https://docs.x.ai/assets/api-examples/images/style-realistic.png",
},
modelInfo: { provider: "xai", model: "grok-imagine-image-quality" },
credentials: { apiKey: "xai-key" },
log: null,
});
expect(result.success).toBe(true);
expect(global.fetch).toHaveBeenCalledWith(
"https://api.x.ai/v1/images/edits",
expect.objectContaining({ method: "POST" })
);
const reqBody = JSON.parse(global.fetch.mock.calls[0][1].body);
expect(reqBody.image).toEqual({
type: "image_url",
url: "https://docs.x.ai/assets/api-examples/images/style-realistic.png",
});
expect(reqBody.images).toBeUndefined();
});
it("supports multi-image edit for xAI (up to 3 refs)", async () => {
global.fetch.mockResolvedValueOnce(
new Response(
JSON.stringify({ created: 1, data: [{ b64_json: "abc" }] }),
{ status: 200, headers: { "Content-Type": "application/json" } }
)
);
const result = await handleImageGenerationCore({
body: {
prompt: "Show all subjects sitting together on the grass",
images: [
"https://docs.x.ai/assets/api-examples/images/image-merge/woman.jpg",
{ url: "https://docs.x.ai/assets/api-examples/images/image-merge/man.jpg" },
"rawbase64payload",
"https://example.com/extra-ignored.jpg",
],
aspect_ratio: "3:2",
response_format: "b64_json",
},
modelInfo: { provider: "xai", model: "grok-imagine-image-quality" },
credentials: { apiKey: "xai-key" },
log: null,
});
expect(result.success).toBe(true);
expect(global.fetch).toHaveBeenCalledWith(
"https://api.x.ai/v1/images/edits",
expect.any(Object)
);
const reqBody = JSON.parse(global.fetch.mock.calls[0][1].body);
expect(reqBody.images).toEqual([
{ type: "image_url", url: "https://docs.x.ai/assets/api-examples/images/image-merge/woman.jpg" },
{ type: "image_url", url: "https://docs.x.ai/assets/api-examples/images/image-merge/man.jpg" },
{ type: "image_url", url: "data:image/png;base64,rawbase64payload" },
]);
expect(reqBody.image).toBeUndefined();
expect(reqBody.aspect_ratio).toBe("3:2");
expect(reqBody.response_format).toBe("b64_json");
});
});

View File

@@ -14,168 +14,341 @@ import { openaiToCommandCodeRequest } from "../../open-sse/translator/request/op
const MODEL = "moonshotai/Kimi-K2.6";
describe("openaiToCommandCodeRequest — basic envelope", () => {
it("returns the expected top-level envelope shape", () => {
const out = openaiToCommandCodeRequest(MODEL, {
messages: [{ role: "user", content: "hi" }],
}, true);
it("returns the expected top-level envelope shape", () => {
const out = openaiToCommandCodeRequest(
MODEL,
{
messages: [{ role: "user", content: "hi" }],
},
true,
);
expect(out).toHaveProperty("threadId");
expect(out).toHaveProperty("memory");
expect(out).toHaveProperty("config");
expect(out).toHaveProperty("params");
expect(out.params.model).toBe(MODEL);
expect(out.params.stream).toBe(true);
});
expect(out).toHaveProperty("threadId");
expect(out).toHaveProperty("memory");
expect(out).toHaveProperty("config");
expect(out).toHaveProperty("params");
expect(out.params.model).toBe(MODEL);
expect(out.params.stream).toBe(true);
});
});
describe("openaiToCommandCodeRequest — system handling", () => {
it("hoists system messages to params.system (string), not messages[]", () => {
const out = openaiToCommandCodeRequest(MODEL, {
messages: [
{ role: "system", content: "You are concise." },
{ role: "user", content: "hi" },
],
}, true);
it("hoists system messages to params.system (string), not messages[]", () => {
const out = openaiToCommandCodeRequest(
MODEL,
{
messages: [
{ role: "system", content: "You are concise." },
{ role: "user", content: "hi" },
],
},
true,
);
expect(typeof out.params.system).toBe("string");
expect(out.params.system).toBe("You are concise.");
const roles = out.params.messages.map((m) => m.role);
expect(roles).not.toContain("system");
});
expect(typeof out.params.system).toBe("string");
expect(out.params.system).toBe("You are concise.");
const roles = out.params.messages.map((m) => m.role);
expect(roles).not.toContain("system");
});
it("joins multiple system messages with blank line", () => {
const out = openaiToCommandCodeRequest(MODEL, {
messages: [
{ role: "system", content: "A" },
{ role: "system", content: "B" },
{ role: "user", content: "hi" },
],
}, true);
it("joins multiple system messages with blank line", () => {
const out = openaiToCommandCodeRequest(
MODEL,
{
messages: [
{ role: "system", content: "A" },
{ role: "system", content: "B" },
{ role: "user", content: "hi" },
],
},
true,
);
expect(out.params.system).toBe("A\n\nB");
});
expect(out.params.system).toBe("A\n\nB");
});
it("omits params.system when no system messages", () => {
const out = openaiToCommandCodeRequest(MODEL, {
messages: [{ role: "user", content: "hi" }],
}, true);
expect(out.params.system).toBeUndefined();
});
it("omits params.system when no system messages", () => {
const out = openaiToCommandCodeRequest(
MODEL,
{
messages: [{ role: "user", content: "hi" }],
},
true,
);
expect(out.params.system).toBeUndefined();
});
});
describe("openaiToCommandCodeRequest — vision / image blocks", () => {
const PNG =
"iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAYAAAAfFcSJAAAADUlEQVR42mP8z8BQDwAEhQGAhKmMIQAAAABJRU5ErkJggg==";
it('maps OpenAI image_url (data URI) → {type:"image", image:"data:..."}', () => {
const out = openaiToCommandCodeRequest(
MODEL,
{
messages: [
{
role: "user",
content: [
{ type: "text", text: "What color?" },
{
type: "image_url",
image_url: { url: `data:image/png;base64,${PNG}` },
},
],
},
],
},
true,
);
const blocks = out.params.messages[0].content;
expect(blocks[0]).toEqual({ type: "text", text: "What color?" });
expect(blocks[1]).toEqual({
type: "image",
image: `data:image/png;base64,${PNG}`,
});
});
it("maps Claude-style image block (source.base64) → data URI with media_type", () => {
const out = openaiToCommandCodeRequest(
MODEL,
{
messages: [
{
role: "user",
content: [
{
type: "image",
source: {
type: "base64",
media_type: "image/jpeg",
data: "AAAA",
},
},
],
},
],
},
true,
);
expect(out.params.messages[0].content[0]).toEqual({
type: "image",
image: "data:image/jpeg;base64,AAAA",
});
});
it("passes raw URL through when image_url is a remote http(s) URL", () => {
const out = openaiToCommandCodeRequest(
MODEL,
{
messages: [
{
role: "user",
content: [
{
type: "image_url",
image_url: { url: "https://example.com/a.png" },
},
],
},
],
},
true,
);
expect(out.params.messages[0].content[0]).toEqual({
type: "image",
image: "https://example.com/a.png",
});
});
it("skips image block when no usable image source", () => {
const out = openaiToCommandCodeRequest(
MODEL,
{
messages: [
{
role: "user",
content: [{ type: "image_url", image_url: { url: "" } }],
},
],
},
true,
);
const blocks = out.params.messages[0].content;
expect(blocks.every((b) => b.type !== "image")).toBe(true);
});
});
describe("openaiToCommandCodeRequest — content shape", () => {
it("MUST always emit content as Array (never string) for user", () => {
const out = openaiToCommandCodeRequest(MODEL, {
messages: [{ role: "user", content: "hello" }],
}, true);
it("MUST always emit content as Array (never string) for user", () => {
const out = openaiToCommandCodeRequest(
MODEL,
{
messages: [{ role: "user", content: "hello" }],
},
true,
);
const u = out.params.messages[0];
expect(Array.isArray(u.content)).toBe(true);
expect(u.content[0]).toEqual({ type: "text", text: "hello" });
});
const u = out.params.messages[0];
expect(Array.isArray(u.content)).toBe(true);
expect(u.content[0]).toEqual({ type: "text", text: "hello" });
});
it("MUST always emit content as Array for assistant", () => {
const out = openaiToCommandCodeRequest(MODEL, {
messages: [
{ role: "user", content: "a" },
{ role: "assistant", content: "b" },
],
}, true);
const a = out.params.messages[1];
expect(Array.isArray(a.content)).toBe(true);
expect(a.content[0]).toEqual({ type: "text", text: "b" });
});
it("MUST always emit content as Array for assistant", () => {
const out = openaiToCommandCodeRequest(
MODEL,
{
messages: [
{ role: "user", content: "a" },
{ role: "assistant", content: "b" },
],
},
true,
);
const a = out.params.messages[1];
expect(Array.isArray(a.content)).toBe(true);
expect(a.content[0]).toEqual({ type: "text", text: "b" });
});
});
describe("openaiToCommandCodeRequest — tool role / tool-result (AI SDK)", () => {
it("converts role:\"tool\" to role:\"tool\" with tool-result block; output is {type:\"text\",value}", () => {
const out = openaiToCommandCodeRequest(MODEL, {
messages: [
{ role: "user", content: "run X" },
{
role: "assistant",
content: null,
tool_calls: [
{ id: "call_1", type: "function", function: { name: "do_x", arguments: "{\"a\":1}" } },
],
},
{ role: "tool", tool_call_id: "call_1", name: "do_x", content: "RESULT_OK" },
],
}, true);
it('converts role:"tool" to role:"tool" with tool-result block; output is {type:"text",value}', () => {
const out = openaiToCommandCodeRequest(
MODEL,
{
messages: [
{ role: "user", content: "run X" },
{
role: "assistant",
content: null,
tool_calls: [
{
id: "call_1",
type: "function",
function: { name: "do_x", arguments: '{"a":1}' },
},
],
},
{
role: "tool",
tool_call_id: "call_1",
name: "do_x",
content: "RESULT_OK",
},
],
},
true,
);
const toolMsg = out.params.messages[out.params.messages.length - 1];
expect(toolMsg.role).toBe("tool");
const block = toolMsg.content[0];
expect(block.type).toBe("tool-result");
expect(block.toolCallId).toBe("call_1");
expect(block.toolName).toBe("do_x");
expect(block.output).toEqual({ type: "text", value: "RESULT_OK" });
});
const toolMsg = out.params.messages[out.params.messages.length - 1];
expect(toolMsg.role).toBe("tool");
const block = toolMsg.content[0];
expect(block.type).toBe("tool-result");
expect(block.toolCallId).toBe("call_1");
expect(block.toolName).toBe("do_x");
expect(block.output).toEqual({ type: "text", value: "RESULT_OK" });
});
});
describe("openaiToCommandCodeRequest — assistant tool_calls / tool-call", () => {
it("converts assistant.tool_calls[] into content blocks of type tool-call", () => {
const out = openaiToCommandCodeRequest(MODEL, {
messages: [
{ role: "user", content: "go" },
{
role: "assistant",
content: null,
tool_calls: [
{ id: "call_42", type: "function", function: { name: "search", arguments: "{\"q\":\"hi\"}" } },
],
},
],
}, true);
it("converts assistant.tool_calls[] into content blocks of type tool-call", () => {
const out = openaiToCommandCodeRequest(
MODEL,
{
messages: [
{ role: "user", content: "go" },
{
role: "assistant",
content: null,
tool_calls: [
{
id: "call_42",
type: "function",
function: { name: "search", arguments: '{"q":"hi"}' },
},
],
},
],
},
true,
);
const asst = out.params.messages[1];
expect(asst.role).toBe("assistant");
const tc = asst.content.find((b) => b.type === "tool-call");
expect(tc).toBeDefined();
expect(tc.toolCallId).toBe("call_42");
expect(tc.toolName).toBe("search");
expect(tc.input).toEqual({ q: "hi" });
});
const asst = out.params.messages[1];
expect(asst.role).toBe("assistant");
const tc = asst.content.find((b) => b.type === "tool-call");
expect(tc).toBeDefined();
expect(tc.toolCallId).toBe("call_42");
expect(tc.toolName).toBe("search");
expect(tc.input).toEqual({ q: "hi" });
});
});
describe("openaiToCommandCodeRequest — tools schema conversion", () => {
it("converts OpenAI {type:\"function\", function:{...}} to Anthropic plain {name, input_schema}", () => {
const out = openaiToCommandCodeRequest(MODEL, {
messages: [{ role: "user", content: "hi" }],
tools: [
{
type: "function",
function: {
name: "weather",
description: "Get weather",
parameters: { type: "object", properties: { city: { type: "string" } }, required: ["city"] },
},
},
],
}, true);
it('converts OpenAI {type:"function", function:{...}} to Anthropic plain {name, input_schema}', () => {
const out = openaiToCommandCodeRequest(
MODEL,
{
messages: [{ role: "user", content: "hi" }],
tools: [
{
type: "function",
function: {
name: "weather",
description: "Get weather",
parameters: {
type: "object",
properties: { city: { type: "string" } },
required: ["city"],
},
},
},
],
},
true,
);
const t = out.params.tools[0];
expect(t.name).toBe("weather");
expect(t.input_schema).toBeDefined();
expect(t.input_schema.type).toBe("object");
expect(t.function).toBeUndefined();
expect(t.parameters).toBeUndefined();
});
const t = out.params.tools[0];
expect(t.name).toBe("weather");
expect(t.input_schema).toBeDefined();
expect(t.input_schema.type).toBe("object");
expect(t.function).toBeUndefined();
expect(t.parameters).toBeUndefined();
});
it("preserves description on converted tool", () => {
const out = openaiToCommandCodeRequest(MODEL, {
messages: [{ role: "user", content: "hi" }],
tools: [
{ type: "function", function: { name: "ping", description: "Ping the server", parameters: { type: "object" } } },
],
}, true);
expect(out.params.tools[0].description).toBe("Ping the server");
});
it("preserves description on converted tool", () => {
const out = openaiToCommandCodeRequest(
MODEL,
{
messages: [{ role: "user", content: "hi" }],
tools: [
{
type: "function",
function: {
name: "ping",
description: "Ping the server",
parameters: { type: "object" },
},
},
],
},
true,
);
expect(out.params.tools[0].description).toBe("Ping the server");
});
it("does not include tools field when input has none", () => {
const out = openaiToCommandCodeRequest(MODEL, {
messages: [{ role: "user", content: "hi" }],
}, true);
expect(out.params.tools).toBeUndefined();
});
it("does not include tools field when input has none", () => {
const out = openaiToCommandCodeRequest(
MODEL,
{
messages: [{ role: "user", content: "hi" }],
},
true,
);
expect(out.params.tools).toBeUndefined();
});
});

View File

@@ -0,0 +1,90 @@
import { describe, it, expect } from "vitest";
import { handleForcedSSEToJson } from "../../open-sse/handlers/chatCore/sseToJsonHandler.js";
const encoder = new TextEncoder();
const sseResponse = (chunks) => {
const body = new ReadableStream({
start(c) {
for (const ch of chunks)
c.enqueue(encoder.encode(`data: ${JSON.stringify(ch)}\n\n`));
c.enqueue(encoder.encode("data: [DONE]\n\n"));
c.close();
},
});
return new Response(body, {
status: 200,
headers: { "content-type": "text/event-stream" },
});
};
const mkChunk = (content, finish = null) => ({
id: "x",
object: "chat.completion.chunk",
created: 1,
model: "m",
choices: [
{ index: 0, delta: content ? { content } : {}, finish_reason: finish },
],
});
const baseCtx = {
provider: "fakeprovider",
model: "m",
body: { stream: false },
stream: true,
translatedBody: null,
finalBody: null,
requestStartTime: Date.now(),
connectionId: "c1",
apiKey: null,
clientRawRequest: null,
onRequestSuccess: null,
pxpipe: null,
reqTag: "",
log: null,
trackDone: () => {},
appendLog: () => {},
reqLogger: null,
toolNameMap: null,
sourceFormat: "openai",
};
describe("Layer 1 — non-streaming stream error patterns", () => {
it("returns a 502 error result when content matches a configured pattern", async () => {
const result = await handleForcedSSEToJson({
...baseCtx,
streamErrorPatterns: { fakeprovider: ["Network connection lost"] },
providerResponse: sseResponse([
mkChunk("Network connection lost."),
mkChunk(null, "stop"),
]),
});
expect(result.success).toBe(false);
expect(result.status).toBe(502);
expect(result.error).toContain("Network connection lost");
});
it("succeeds when content does not match", async () => {
const result = await handleForcedSSEToJson({
...baseCtx,
streamErrorPatterns: { fakeprovider: ["Network connection lost"] },
providerResponse: sseResponse([
mkChunk("hello world"),
mkChunk(null, "stop"),
]),
});
expect(result.success).toBe(true);
});
it("ignores patterns for other providers", async () => {
const result = await handleForcedSSEToJson({
...baseCtx,
streamErrorPatterns: { otherprovider: ["hello"] },
providerResponse: sseResponse([
mkChunk("hello world"),
mkChunk(null, "stop"),
]),
});
expect(result.success).toBe(true);
});
});

View File

@@ -0,0 +1,99 @@
import { describe, it, expect } from "vitest";
import { maybeRejectEarlyStreamError } from "../../open-sse/utils/streamErrorPeek.js";
const encoder = new TextEncoder();
const sseResp = (lines) =>
new Response(
new ReadableStream({
start(c) {
for (const l of lines) c.enqueue(encoder.encode(l));
c.close();
},
}),
{ status: 200, headers: { "content-type": "text/event-stream" } },
);
describe("maybeRejectEarlyStreamError", () => {
it("returns 502 when a pattern matches early stream text", async () => {
const res = await maybeRejectEarlyStreamError(
sseResp([
'{"type":"start"}\n{"type":"error","error":{"type":"server_error","message":"Network connection lost."}}\n',
]),
["server_error"],
);
expect(res.status).toBe(502);
const body = await res.json();
expect(body.error.message).toContain("server_error");
});
it("passes the stream through unchanged when nothing matches", async () => {
const res = await maybeRejectEarlyStreamError(
sseResp([
'data: {"choices":[{"delta":{"content":"hello"},"finish_reason":null}]}\n\ndata: [DONE]\n\n',
]),
["server_error"],
);
expect(res.status).toBe(200);
const text = await res.text();
expect(text).toContain("hello");
expect(text).toContain("[DONE]");
});
it("commits on timeout without hanging", async () => {
const stalled = new Response(new ReadableStream({ start() {} }), {
status: 200,
});
const res = await maybeRejectEarlyStreamError(stalled, ["x"], {
timeoutMs: 50,
});
expect(res.status).toBe(200);
await res.body.cancel();
});
it("commits on abort without hanging", async () => {
const ctrl = new AbortController();
const stalled = new Response(new ReadableStream({ start() {} }), {
status: 200,
});
setTimeout(() => ctrl.abort(new Error("gone")), 10);
const res = await maybeRejectEarlyStreamError(stalled, ["x"], {
signal: ctrl.signal,
timeoutMs: 2000,
});
expect(res.status).toBe(200);
await res.body.cancel();
});
it("commits (passthrough) when patterns are empty", async () => {
const res = await maybeRejectEarlyStreamError(
sseResp(["data: hi\n\n"]),
[],
);
expect(res.status).toBe(200);
expect(await res.text()).toBe("data: hi\n\n");
});
it("multi-byte UTF-8 split across peek boundary round-trips losslessly", async () => {
// "café" split mid-é: the peek consumes bytes up to and including 0xC3
// (the first half of é); the re-emitted stream must contain the RAW bytes
// (never a re-encoded decoded string — TextDecoder flush would replace the
// lone 0xC3 with U+FFFD and corrupt the output).
const bytes = [
0x64, 0x61, 0x74, 0x61, 0x3a, 0x20, 0x22, 0x63, 0x61, 0x66, 0xc3,
];
const rest = new Uint8Array([0xa9, 0x22, 0x0a, 0x0a]);
const body = new ReadableStream({
start(c) {
c.enqueue(new Uint8Array(bytes));
c.enqueue(rest);
c.close();
},
});
const res = await maybeRejectEarlyStreamError(
new Response(body, { status: 200 }),
["nomatch"],
{ maxBytes: 11 },
);
expect(await res.text()).toBe('data: "café"\n\n');
});
});

View File

@@ -0,0 +1,70 @@
import { describe, it, expect } from "vitest";
import {
parsePatterns,
matchStreamErrorPatterns,
streamStatusForContent,
} from "../../open-sse/utils/streamErrorPatterns.js";
describe("streamErrorPatterns util", () => {
it("matches plain text as case-insensitive substring", () => {
expect(
matchStreamErrorPatterns(
["Network connection lost"],
"network CONNECTION LOST.",
),
).toBe("Network connection lost");
expect(
matchStreamErrorPatterns(["server_error"], "some normal content"),
).toBeNull();
});
it("matches /regex/flags", () => {
expect(
matchStreamErrorPatterns(
["/generation failed.*retry/i"],
"GENERATION FAILED. please RETRY",
),
).toBe("/generation failed.*retry/i");
expect(matchStreamErrorPatterns(["/\\d+ tokens/"], "used 123 tokens")).toBe(
"/\\d+ tokens/",
);
});
it("skips invalid regex and empty entries", () => {
expect(
matchStreamErrorPatterns(["/[unclosed/", "", " "], "anything"),
).toBeNull();
});
it("returns null for empty patterns or text", () => {
expect(matchStreamErrorPatterns([], "x")).toBeNull();
expect(matchStreamErrorPatterns(["x"], "")).toBeNull();
expect(matchStreamErrorPatterns(null, "x")).toBeNull();
expect(matchStreamErrorPatterns(undefined, "x")).toBeNull();
});
it("regex matching is stateless across calls", () => {
const pats = ["/error/i"];
expect(matchStreamErrorPatterns(pats, "ERROR")).toBe("/error/i");
expect(matchStreamErrorPatterns(pats, "ERROR")).toBe("/error/i");
});
it("parsePatterns normalizes entries", () => {
const parsed = parsePatterns(["Plain", "/re/g", "", "/bad["]);
expect(parsed.length).toBe(3);
expect(parsed[0]).toEqual({ text: "plain", raw: "Plain" });
expect(parsed[1].regex).toBeInstanceOf(RegExp);
expect(parsed[2]).toEqual({ text: "/bad[", raw: "/bad[" });
});
it("skips entries that look like regex but fail to compile", () => {
const parsed = parsePatterns(["/[unclosed/", "/ok/g"]);
expect(parsed.length).toBe(1);
expect(parsed[0].raw).toBe("/ok/g");
});
it("streamStatusForContent maps match to error status", () => {
expect(streamStatusForContent(["boom"], "a boom happened")).toBe("error");
expect(streamStatusForContent(["boom"], "all good")).toBe("success");
});
});

View File

@@ -0,0 +1,260 @@
import { describe, it, expect, vi, beforeEach } from "vitest";
vi.mock("../../open-sse/utils/proxyFetch.js", () => ({
proxyAwareFetch: vi.fn(),
}));
import { proxyAwareFetch } from "../../open-sse/utils/proxyFetch.js";
import { getUsageForProvider } from "../../open-sse/services/usage.js";
import { parseGrokCreditsConfig } from "../../open-sse/services/usage/xai.js";
import { parseQuotaData } from "@/app/(dashboard)/dashboard/usage/components/ProviderLimits/utils.js";
function billingResponse(config) {
return new Response(JSON.stringify({ config }), {
status: 200,
headers: { "Content-Type": "application/json" },
});
}
function settingsResponse(tier = "SuperGrok") {
return new Response(JSON.stringify({ subscription_tier_display: tier }), {
status: 200,
headers: { "Content-Type": "application/json" },
});
}
/**
* Build a minimal grpc-web GetGrokCreditsConfig payload matching the live
* SuperGrok response shape: usedPercent float + start/end timestamps.
*/
function buildCreditsGrpcFrame({ usedPercent = 18, startSec = 1783673158, endSec = 1784277958 }) {
// Encode protobuf Timestamp {1: seconds}
const encodeVarint = (n) => {
const out = [];
let v = n >>> 0;
while (v >= 0x80) {
out.push((v & 0x7f) | 0x80);
v >>>= 7;
}
out.push(v);
return Buffer.from(out);
};
const encodeKey = (field, wt) => encodeVarint((field << 3) | wt);
const encodeTimestamp = (seconds) => {
const body = Buffer.concat([encodeKey(1, 0), encodeVarint(seconds)]);
return body;
};
const encodeLen = (field, bytes) =>
Buffer.concat([encodeKey(field, 2), encodeVarint(bytes.length), bytes]);
const encodeFloat = (field, f) => {
const buf = Buffer.alloc(4);
buf.writeFloatLE(f, 0);
return Buffer.concat([encodeKey(field, 5), buf]);
};
const startTs = encodeTimestamp(startSec);
const endTs = encodeTimestamp(endSec);
const config = Buffer.concat([
encodeFloat(1, usedPercent),
encodeLen(4, startTs),
encodeLen(5, endTs),
]);
const msg = encodeLen(1, config);
const frame = Buffer.alloc(5 + msg.length);
frame[0] = 0;
frame.writeUInt32BE(msg.length, 1);
msg.copy(frame, 5);
return frame;
}
function creditsResponse(opts = {}) {
const frame = buildCreditsGrpcFrame(opts);
return new Response(frame, {
status: 200,
headers: { "Content-Type": "application/grpc-web+proto" },
});
}
function mockXaiHappyPath({ used = 733, limit = 15000, weeklyPercent = 18 } = {}) {
proxyAwareFetch.mockImplementation(async (url) => {
if (String(url).includes("/v1/billing")) {
return billingResponse({
monthlyLimit: { val: limit },
used: { val: used },
onDemandCap: { val: 0 },
billingPeriodStart: "2026-07-01T00:00:00+00:00",
billingPeriodEnd: "2026-08-01T00:00:00+00:00",
});
}
if (String(url).includes("GetGrokCreditsConfig")) {
return creditsResponse({ usedPercent: weeklyPercent });
}
if (String(url).includes("/v1/settings")) {
return settingsResponse("SuperGrok");
}
return new Response("{}", { status: 404 });
});
}
describe("xAI usage", () => {
beforeEach(() => {
vi.clearAllMocks();
});
it("parseGrokCreditsConfig extracts weekly used% and reset timestamp", () => {
const frame = buildCreditsGrpcFrame({
usedPercent: 18,
startSec: 1783673158,
endSec: 1784277958,
});
const parsed = parseGrokCreditsConfig(frame);
expect(parsed).toMatchObject({
usedPercent: 18,
remainingPercent: 82,
});
expect(parsed.resetAt).toBe(new Date(1784277958 * 1000).toISOString());
expect(parsed.periodStart).toBe(new Date(1783673158 * 1000).toISOString());
});
it("fetches billing + weekly credits + settings in parallel", async () => {
mockXaiHappyPath();
const usage = await getUsageForProvider({
provider: "xai",
accessToken: "tok-abc",
});
const urls = proxyAwareFetch.mock.calls.map((c) => String(c[0]));
expect(urls).toEqual(
expect.arrayContaining([
"https://cli-chat-proxy.grok.com/v1/billing",
"https://grok.com/grok_api_v2.GrokBuildBilling/GetGrokCreditsConfig",
"https://cli-chat-proxy.grok.com/v1/settings",
]),
);
expect(usage.plan).toBe("SuperGrok");
expect(usage.quotas.weekly).toMatchObject({
used: 18,
total: 100,
remaining: 82,
});
expect(usage.quotas.api_usage).toMatchObject({
used: 733,
total: 15000,
unit: "credits",
});
expect(usage.quotas.api_usage.resetAt).toBe("2026-08-01T00:00:00.000Z");
});
it("still returns weekly when billing fails but credits succeed", async () => {
proxyAwareFetch.mockImplementation(async (url) => {
if (String(url).includes("/v1/billing")) {
return new Response("nope", { status: 503 });
}
if (String(url).includes("GetGrokCreditsConfig")) {
return creditsResponse({ usedPercent: 42 });
}
if (String(url).includes("/v1/settings")) {
return settingsResponse("SuperGrok");
}
return new Response("{}", { status: 404 });
});
const usage = await getUsageForProvider({
provider: "xai",
accessToken: "tok",
});
expect(usage.plan).toBe("SuperGrok");
expect(usage.quotas.weekly).toMatchObject({ used: 42, total: 100, remaining: 58 });
expect(usage.quotas.api_usage).toBeUndefined();
});
it("includes on-demand row when onDemandCap > 0", async () => {
proxyAwareFetch.mockImplementation(async (url) => {
if (String(url).includes("/v1/billing")) {
return billingResponse({
monthlyLimit: { val: 1000 },
used: { val: 100 },
onDemandCap: { val: 500 },
billingPeriodEnd: "2026-08-01T00:00:00+00:00",
});
}
if (String(url).includes("GetGrokCreditsConfig")) {
return creditsResponse({ usedPercent: 10 });
}
if (String(url).includes("/v1/settings")) {
return settingsResponse("SuperGrok");
}
return new Response("{}", { status: 404 });
});
const usage = await getUsageForProvider({
provider: "xai",
accessToken: "tok",
});
expect(usage.quotas.on_demand).toMatchObject({
used: 0,
total: 500,
remainingCredits: 500,
});
});
it("returns auth message on 401 billing", async () => {
proxyAwareFetch.mockImplementation(async (url) => {
if (String(url).includes("/v1/billing")) {
return new Response("unauthorized", { status: 401 });
}
return new Response("{}", { status: 404 });
});
const usage = await getUsageForProvider({
provider: "xai",
accessToken: "tok",
});
expect(usage.message).toMatch(/re-authorize/i);
});
it("parseQuotaData maps weekly + api_usage labels", () => {
const rows = parseQuotaData("xai", {
quotas: {
weekly: {
used: 18,
total: 100,
remaining: 82,
remainingPercentage: 82,
resetAt: "2026-07-17T08:45:58.000Z",
},
api_usage: {
used: 733,
total: 15000,
remainingCredits: 14267,
unit: "credits",
resetAt: "2026-08-01T00:00:00.000Z",
},
},
});
expect(rows).toEqual(
expect.arrayContaining([
expect.objectContaining({
name: "Weekly limit",
used: 18,
total: 100,
remaining: 82,
}),
expect.objectContaining({
name: "Api usage",
used: 733,
total: 15000,
unit: "credits",
}),
]),
);
const api = rows.find((r) => r.name === "Api usage");
expect(api.remaining).toBeUndefined();
});
});