151 Commits

Author SHA1 Message Date
39101b3416 Merge origin/master (v0.5.91) into gitea/new_feature
Resolve conflicts:
- streamingHandler.js: adopt upstreamResponseHeaders while keeping 0-token detail row avoidance
- capabilities.js: preserve user-asserted caps and globalThis slots without local caching of catalogSource
- AddCustomModelModal.js & providers/[id]/page.js: wire STT transport marker with custom model edits/assertions
- models/custom/route.js & aliasRepo.js: persist custom model transport and invalidate user caps
- usageRepo.js: key byApiKey live stats by full API key and keep tail in maskApiKey
- UsageStats.js: lazy load charts dynamically
2026-09-28 21:17:33 +07:00
decolua
f01fb909e3 # v0.5.91 (2026-09-26)
## Features
- **Providers**: add Token Harbor provider and four OpenAI-compatible aggregator providers (dahl, atria, agnes, bai)
- **Claude**: forward `x-claude-code-session-id` on OAuth requests; merge client `anthropic-beta` flags and forward rate-limit headers; return thinking text to OpenAI-format clients
- **Codex**: add GPT-6 Sol and Luna support
- **CLI Tools**: support multiple model profiles for Codex CLI
- **Hermes**: multi-role model config (delegation + auxiliary slots)
- **OpenCode Go**: complete the Go catalog (40 models) with auto-fetch + family endpoint regex
- **Usage**: show and redeem free limit resets for cc accounts
- **Cline**: expose the `cline-free/*` tier and price it at zero
- **Combos**: display vision adapter models in an ordered table view

## Fixes
- **Claude**: decloak tool names when `toolNameMap` misses (#4342); update spoofed cli version to 2.1.280 to support Opus 5.5
- **Providers API**: make POST `/api/providers` O(1) and refuse silent key overwrite (#4350)
- **Capabilities**: stop caching the catalog source per module copy (#4351)
- **OAuth**: stop Zed paste-token crash and add IDE auto-import (#4359)
- **Dashboard**: resolve combo limits with the server's capabilities (#4360); lazy-load charts and `marked`, preload in background on idle
- **Responses**: carry the streamed output items in `response.completed` (#4307)
- **STT**: dispatch live-API-only Gemini models over the Live WebSocket transport (#4006)
- **Gemini**: guard terminal model turns and unresponded functionCalls in `normalizeGeminiContents`
- **Command Code**: replay raw byte chunks to preserve all NDJSON lines
- **Translator**: stop emitting empty `<think>` markers into OpenAI content
- **CLI Tools**: refresh Codex settings after apply (#4347); keep existing `ANTHROPIC_AUTH_TOKEN` when applying Claude settings
- **Tray**: native arm64 macOS menubar binary, no Rosetta required
- **CLI**: filter model selector by active connections and noAuth providers
- **Usage**: key live byApiKey stats by full api key to prevent team-key collision and preserve API key usage attribution
- **Tailscale**: cap enable-flow health wait at 20s
2026-09-26 17:40:15 +07:00
decolua
c4690307ce fix(cli-tools): keep existing ANTHROPIC_AUTH_TOKEN when applying Claude settings
Only write the token when settings.json has none, so a real API key or
earlier config is never clobbered by Apply. Reset still clears it,
re-enabling key selection on the next Apply.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-26 17:30:59 +07:00
decolua
b54a3f9bb5 test(claude): update decloak tests for suffix-stripping fallback 2026-09-26 17:30:02 +07:00
Gesya Gayatree Solih
b65d2d0a6a fix(claude): decloak tool names when toolNameMap misses (#4342)
- Add stripCloakSuffix fallback in decloakToolNames and decloakStreamChunk
- Prevent client errors when toolNameMap is missing or lost on retry
2026-09-26 17:28:51 +07:00
decolua
6aea3875ef feat(claude): forward x-claude-code-session-id on OAuth requests 2026-09-26 17:15:12 +07:00
decolua
0249464d74 feat(cli-tools): support multiple model profiles for Codex CLI
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-26 17:03:06 +07:00
Nikan Wystaf
239bcfc568 perf(providers): make POST /api/providers O(1) and refuse silent key overwrite (#4350)
Fixes #4311

- Drop full-pool renumber on insert: new row gets MAX(priority)+1 directly,
  turning an O(pool) rewrite into O(1) per insert.
- For apikey connections, query by (provider, authType, name) and count via SQL
  aggregate instead of reading the entire pool into memory.
- Refuse silent apikey overwrite on name collision with 409 PROVIDER_NAME_CONFLICT,
  unless caller explicitly sets allowOverwrite: true.
- Add 8 unit tests covering priority ordering and name collisions.
2026-09-26 16:37:25 +07:00
Nick Nyanjui
199173fe58 feat(cline): expose the cline-free/* tier and price it at zero (#4334) 2026-09-26 16:31:25 +07:00
Beru
8f20daac6b fix(cli-tools): refresh Codex settings after apply (#4347) 2026-09-26 16:21:21 +07:00
Mohammad Hijjawi
fdcba3e1b2 fix(capabilities): stop caching the catalog source per module copy (#4351) 2026-09-26 16:20:59 +07:00
agent
dd293d3c62 feat(opencode-go): add the seven models upstream serves but the registry omits (#4357)
Upstream /zen/go/v1/models serves 35 ids; the registry listed 28. Add the seven
missing models with their corresponding supported and target formats:
- deepseek-v4.1-flash
- mimo-v2.6-flash, mimo-v2.6-pro
- space-bunny-free
- omen-alpha
- grok-4.7 (responses-only)
- gpt-6-luna (responses-only)
2026-09-26 16:15:43 +07:00
Amir Seify
7a436d209c fix(oauth): stop Zed paste-token crash and add IDE auto-import (#4359)
- Fix Zed Connect crash by gating paste-token UI when provider has no paste-token config
- Add ZedAuthModal with local IDE keyring auto-import, browser OAuth, and manual callback paste
- Add localhost-only GET /api/oauth/zed/auto-import and POST /api/oauth/zed/import routes
2026-09-26 16:12:46 +07:00
Nikan Wystaf
737b1f4d0a feat(providers): add Token Harbor provider 2026-09-26 16:09:52 +07:00
Spoon94
37a6b7e0f2 fix(dashboard): resolve combo limits with the server's capabilities (#4360)
The combos page computed combo capabilities in the browser using
aggregateComboCapabilities, falling back to pattern defaults for models
without exact entries because the synced model catalog is server-only.

Allow callers to supply a resolveCaps callback (e.g. from useModelCaps)
merged over the local tables, preserving non-limit capability flags.
2026-09-26 16:07:03 +07:00
Nikan Wystaf
06112c136c feat(providers): add four OpenAI-compatible aggregator providers (dahl, atria, agnes, bai)
Adds Dahl Inference, Atria Dawn, Agnes AI, and B.AI following the existing
registry pattern. Each uses DefaultExecutor with no custom translator needed.
2026-09-26 16:04:59 +07:00
MrBeanDev
90b0693423 feat(thinking): return Claude thinking text to OpenAI-format clients 2026-09-26 12:02:07 +07:00
Nick Nyanjui
fe347e4ea5 fix(stt): dispatch live-API-only Gemini models over the Live WebSocket transport (#4006)
Addresses #4006 by letting Gemini STT models that only exist on the Live
API transcribe instead of failing.

transcribeGemini sends audio to :generateContent and that is the only Gemini
path. A model that is realtime-only (exposed by the Live API's
bidiGenerateContent WebSocket) therefore fails outright, even though the account
can transcribe it.

open-sse/handlers/geminiLiveStt.js owns the WebSocket lifecycle: opens
:bidiGenerateContent, sends setup frame, waits for setupComplete, streams
audio as realtimeInput media chunks, and settles on turnComplete.
Dispatch is driven by transport marker 'gemini-live'. Adds custom model transport
persistence and selection on the dashboard.
2026-09-26 12:00:39 +07:00
ANIRUDDHA ADAK
273f0c32cd fix(responses): carry the streamed output items in response.completed (#4307) 2026-09-26 12:00:00 +07:00
Christian Gennari
c2148179c0 fix(commandcode): replay raw byte chunks to preserve all NDJSON lines 2026-09-26 11:53:54 +07:00
MrBeanDev
5d2cfbf3c5 fix(translator): stop emitting empty <think> markers into OpenAI content 2026-09-26 11:45:35 +07:00
decolua
3e4323e2fd feat(combos): display vision adapter models in an ordered table view
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-26 11:43:22 +07:00
decolua
975f28a57c test(zed): isolate zed-live-models suite DB via temp DATA_DIR
Seeding via createProviderConnection wrote zed-live-* accounts into the
real ~/.9router DB, showing up as junk accounts in the running dashboard.
Set DATA_DIR to a mkdtemp dir before dynamic-importing the DB-backed
modules and clean it up afterwards, matching the pattern in
compatible-provider-connections.test.js.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-26 11:40:08 +07:00
decolua
ddfdb0df04 perf(dashboard): lazy-load charts and marked, preload in background on idle
- Drop UsageStats from shared barrel so recharts (~589KB) stays out of
  the initial bundle of every page; usage page imports it directly
- Convert UsageChart/ProviderBarChart/TopModelsChart to next/dynamic
- Preload chart chunks via requestIdleCallback in DashboardLayout so
  navigating to Usage is still instant
- Lazy-import marked inside ChangelogModal (only when opened)
- Route Sidebar/EndpointPageClient through settingsStore: coalesce
  concurrent in-flight GET, merge PATCH response over cache (PATCH
  omits GET-only hasPassword) to kill duplicate /api/settings calls
- Delay /api/version npm check 2.5s after first render

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-26 11:34:08 +07:00
akmal safari pellu
30464bc227 fix(gemini): guard terminal model turns and unresponded functionCalls in normalizeGeminiContents 2026-09-26 11:18:44 +07:00
Doan Anh Dung
249f6c2fb8 fix(tray): native arm64 macOS menubar binary, no Rosetta required
systray2 ships only an x86_64 tray_darwin_release and selects it by
process.platform with no process.arch branch, so there is no native slice to
choose. Apple Silicon users therefore need Rosetta 2, and without it the tray
dies with EBADARCH ("bad CPU type in executable") and no icon appears.

Overlay a native arm64 build of the same upstream source
(felixhao28/systray-portable @ 6eddc91) instead. On darwin/arm64,
ensureArm64TrayBin() detects the Intel binary by parsing the Mach-O cputype,
downloads the artifact from the pinned tray-binaries release, verifies it
against a sha256 constant, atomically renames it over systray2's binary, and
busts systray2's copyDir cache — that cached copy is what actually executes, so
without the bust the swap has no effect.

Intel Macs keep using systray2's binary unchanged and Windows is unaffected
(PowerShell NotifyIcon, no binary). Any download or checksum failure leaves the
Intel binary in place and tells the user how to install Rosetta; a marker file
throttles retries to once per 24h because ensureTrayRuntime runs synchronously
on every CLI start, and is cleared on success so a clobbered binary recovers
immediately. Binaries stay out of the npm tarball per the existing Kaspersky
false-positive constraint — the artifact is fetched on demand.

Adds cli/scripts/buildTrayArm64.js (-trimpath, bit-for-bit reproducible for a
given Go version and macOS SDK) and a workflow_dispatch action that builds on a
macos-15 runner and refuses to publish when the sha diverges from the pin.

Also corrects comments claiming the systray -> systray2 switch fixed Apple
Silicon; it only fixed the dyld header rejection on macOS 14+, the binary was
still amd64-only.
2026-09-26 11:16:38 +07:00
decolua
f462837536 fix(cli): filter model selector by active connections and noAuth providers
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-26 11:14:04 +07:00
DavidArthurCole
dc198dff1f feat(claude): merge client anthropic-beta flags and forward rate-limit headers 2026-09-26 11:08:57 +07:00
29ccb84faf feat(commandcode): align quota usage with the official CLI /usage flow
Rewrites services/usage/commandcode.js to mirror the command-code CLI: whoami
resolves the org id, credits+subscriptions report the 5-hour/weekly windows and
plan, then usage/summary is queried with since=currentPeriodStart. Endpoints are
read from the provider registry usage block instead of hardcoded constants.

Tests updated to cover the new request order and response shapes.
2026-09-25 16:09:09 +07:00
3481058e09 fix(usage): drop duplicated commandcode import and features block
The previous origin/master merges (11089ab1 and earlier) left two
auto-merged artifacts behind because both sides added the same lines in
different places:

- services/usage.js declared `getCommandCodeUsage` twice (import + handler),
  which is an ESM SyntaxError on a clean checkout. The working tree masked
  it, so `next build` passed locally while the committed tree did not parse.
- providers/registry/commandcode.js carried two `features` blocks; the
  second one shadowed the first for `usage`/`usageApikey`.

Verified with `node --check` on the committed blob.
2026-09-25 16:09:04 +07:00
272dbcb9cc Merge remote-tracking branch 'origin/master' into gitea/new_feature
# Conflicts:
#	open-sse/executors/qoder.js
#	open-sse/handlers/chatCore.js
#	open-sse/handlers/chatCore/sseToJsonHandler.js
#	open-sse/providers/registry/commandcode.js
#	src/app/(dashboard)/dashboard/combos/page.js
#	src/app/api/v1/models/route.js
#	src/lib/db/repos/usageRepo.js
#	src/shared/components/UsageStats.js
2026-09-25 15:14:28 +07:00
Rafli Ahmad Zulfikar
4a57df8bf9 fix(usage): key live byApiKey stats by full api key to prevent team-key collision
maskApiKey kept only the first 8 characters of the key. API keys minted
from the same machine id share that prefix, so every key of an instance
collapsed into one sk-XXXXXXX*** bucket per model/provider, and the
dashboard attributed one holder's usage to another.

Key the live path by the full api key - matching the daily rollup
(aggregateEntryToDay) and the lastUsed overlay - and keep the last 4
characters in maskApiKey so masked keys stay distinguishable in the UI.
2026-09-23 14:53:01 +07:00
Deepanshu
a406381fad fix(usage): preserve API key usage attribution in live stats
The 24h/today branch of getUsageStats keyed byApiKey buckets on the
masked key, while the daily rollup and the lastUsed overlay key on the
full key. Every key minted by one instance shares the sk-{machineId}
prefix, so the mask collapsed all of them into a single bucket and
attributed one key's usage to another.

Key the live branch by the full api key too; the masked value is still
carried on the bucket for display.
2026-09-23 14:42:56 +07:00
Reid Nguyen
e571a8b6da feat(usage): show and redeem free limit resets for cc accounts
Adds the free usage-limit reset (the desktop app's "Reset for free")
to the Quota Tracker for cc OAuth accounts, mirroring the existing
Codex reset-credit button.

- Reset button with remaining count on the card; tooltip shows use-by
  date and which limits get refilled.
- Expiry modal (clock button) listing each grant: label, resets left,
  refills (session / weekly), use-by date, time remaining, status.
- Confirm dialog before redeeming, since a reset is irreversible.
- Usage and reset calls send the CLI User-Agent required for cedar_ember.
- Usage cache is dropped after a redeem so the card shows the refilled limits.
2026-09-23 12:21:07 +07:00
Ilhom (MBP M5 Pro)
8403576095 fix(claude): update spoofed cli version to 2.1.280 to support Opus 5.5
Anthropic's newly released Claude Opus 5.5 model strictly requires
claude-cli version 2.1.280 or newer. Spoofing the older 2.1.258
version results in an HTTP 400 `claude_code_version_too_old` error.

This commit updates the hardcoded `CLAUDE_CLI_VERSION` in
`open-sse/providers/shared.js` and aligns the corresponding unit
tests and baselines to bypass Anthropic's version gating.
2026-09-23 12:18:39 +07:00
decolua
e7c269b8c9 fix(tailscale): cap enable-flow health wait at 20s
Enable waited out the full 180s HEALTH_CHECK.timeoutMs when funnel URL
DNS was unreachable, making the toggle hang ~3 minutes before returning
success. waitForHealth now takes an optional timeoutMs; enable uses
HEALTH_CHECK.enableTimeoutMs (20s) while watchdog/cloudflare keep 180s.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-23 12:12:10 +07:00
mxlanparty@dp14
95600db17b feat(codex): add GPT-6 Sol and Luna support
- Add GPT-6 Sol and Luna to the Codex model registry.
- Send both models using the Codex 0.155 Responses Lite request shape, including `reasoning.context: "all_turns"`.
2026-09-23 12:11:49 +07:00
decolua
69cc1aa987 feat(hermes): multi-role model config (delegation + auxiliary slots)
POST /api/cli-tools/hermes-settings now accepts selections:
[{role, model}] and writes model/delegation/auxiliary.<role> blocks
into ~/.hermes/config.yaml non-destructively (comments and other keys
preserved). GET returns the current delegation/auxiliary config,
DELETE removes all 9router-managed blocks. Dashboard card gains a
collapsible Model Roles section aligned with the existing rows; the
bare `model` payload from the CLI quick setup still works.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-23 11:46:21 +07:00
decolua
001b292876 fix(dashboard): replace Hermes icon with official Nous Research logo
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-23 11:46:13 +07:00
decolua
a61fc6a01e fix(antigravity): rewrite all Hermes identity variants, not just the legacy sentence
The old rule only matched 'You are Hermes Agent, an intelligent AI
assistant created by Nous Research.' Current Hermes versions use
'You are Hermes Agent, built by Nous Research.' and similar variants,
which slipped through and got a fake 429 RESOURCE_EXHAUSTED. Replace
every variant with a neutral 'You are an AI assistant.' identity.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-23 11:46:05 +07:00
decolua
1b72f02e3b fix(opencode-go): clamp deepseek reasoning_effort "max" to "high" for mimo backends that reject it
mimo-v2.5-pro/v2.6 on the Go lane return 400 on reasoning_effort "max"
(probed live; mimo-v2.5 accepts it). The deepseek applyFormat case now honors
the declared thinking levels, and mimo-v2.5-pro gets a levels entry.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-23 11:45:40 +07:00
decolua
2aa99d9f39 feat(opencode-go): complete Go catalog (40 models) with auto-fetch + family endpoint regex
- Add 12 missing models from live /zen/go/v1/models (glm-5, kimi-k2.5,
  mimo v2/v2.6 line, qwen3.5-plus, grok-4.5/4.7, omen-alpha, ...)
- modelsFetcher + passthroughModels + "opencode-go" suggested-models filter
- Family regex fallback (opencodeFamilyFormats) keeps unknown/passthrough ids
  on the right endpoint lane (/responses, /messages, /chat/completions);
  curated registry entries always win
- isResponsesModel now routes passthrough grok/gpt ids to /responses

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-23 11:45:34 +07:00
decolua
39e36d3d0c # v0.5.86 (2026-09-23)
## Features
- **Xiaomi MiMo**: server-assisted desktop login for headless/Docker deployments, five account clusters (cn/sgp/ams/ru/in), and v2.6 pro/flash/pro-ultraspeed models with dual-route (account service vs. cloud API)
- **Claude**: add Claude Opus 5.5 support
- **i18n**: translate React text rewrites via characterData mutation observer

## Fixes
- **Proxy Pools**: keep request headers intact through Vercel/Cloudflare/Deno relays (spreading a `Headers` instance yielded `{}`, dropping auth and content-type)
- **Xiaomi MiMo login**: keep the session in the httpOnly cookie only, require dashboard auth on the proxy branch, and stop forwarding authorization headers upstream
2026-09-23 10:05:33 +07:00
wismyzhizi
910db749aa feat(xiaomi-mimo): server-assisted desktop login, five account clusters, v2.6 models
Reproduce the MiMo Desktop login surface server-side so headless/Docker
deployments can link a Xiaomi account without the Desktop client. The
account session (passToken) is captured during the proxied login and
stored per connection.

- Five account clusters (cn/sgp/ams/ru/in): per-region mimo-server host
  and SSO sid, unknown region falls back to sgp
- mimo-v2.6-pro/flash/pro-ultraspeed dual-route models: account-service
  route when desktop credentials exist, cloud API (sk- key) otherwise;
  drops obsolete mimo-x-*-preview ids
- Desktop ServiceTokenManager 2-phase handshake (single serviceLogin with
  target sid, raw 64-bit nonce preserved), per-region session cache
- reasoning_effort bridged to output_config.effort; i18n runtime now
  observes characterData mutations so React text rewrites get translated
- Security hardening on the login proxy: session travels only in the
  httpOnly cookie (never in the URL), proxy branch requires dashboard
  auth, authorization/proxy-authorization never forwarded upstream, and
  upstream Set-Cookie is not replayed onto the app origin
2026-09-23 09:43:54 +07:00
savioruz
6af26a9ee8 fix(proxy-pools): lossless header forwarding for relay pools
Preserve request headers through the Vercel/Cloudflare/Deno relay.

- proxyFetch: normalize options.headers via Object.fromEntries before
  spreading into the relay headers. Spreading a Headers instance yields
  {} and silently dropped every entry (auth + content-type), which is why
  the same pool worked on one path and failed on another.
- vercel relay: build the forwarded header object from req.headers.entries()
  instead of new Headers(req.headers), avoiding edge-runtime normalization
  of casing/duplicate keys that some providers reject.
2026-09-23 09:12:00 +07:00
Reid Nguyen
cbffeb9770 feat(claude): support Claude Opus 5.5 2026-09-23 09:01:51 +07:00
decolua
21583c03e5 # v0.5.85 (2026-09-22)
## Features
- **System One**: add `/v1/systemone` decision endpoint for Jev models (OpenCode Zen and OpenRouter lanes), wire into sidebar and Media Providers page with interactive probe testing
- **CLI Tools**: add dynamic configuration, settings APIs, and official logos for Pi, OMP, Crush, ForgeCode, Smelt, and CodeWhale
- **Analytics & Usage**: add Requests mode, provider/model breakdown charts, All Time period filter, and refined overview cards
- **Combos**: add Cursor/Claude Default presets; support bulk select/delete and bulk strategy changes (Fallback / Round Robin / Fusion)
- **Model Capabilities**: expose model capability metadata on `/v1/models` and aggregate capabilities across combo targets
- **OpenCode Zen & MiMo**: add OpenCode Zen (`opencode-zen`) provider with free-tier fingerprint; switch default vision fallback to MiMo V2.6 Flash Free
- **Qoder CN**: add `qoder-cn` provider for qoder.com.cn with OAuth flow, COSY protocol, and CN gateway routing

## Fixes
- **Translator**: map Claude `refusal` stop_reason to `content_filter` and surface explanation; strip replayed reasoning fields for Groq, Mistral, and Cerebras (#4220)
- **Antigravity**: drop requestType `agent` to avoid false 429 `RESOURCE_EXHAUSTED`; separate weekly and short-window (5-hour) quotas and deduplicate dashboard rows
- **Responses API**: report usage on `response.completed` so clients can auto-compact (#3432)
- **Hugging Face**: migrate to Inference Providers router (`router.huggingface.co`), expand image models catalog, and add STT route
- **Qoder**: prevent signed request replay (`403/103 Duplicate request`), handle code 110 billing blocks, and preserve upstream SSE error status
- **Performance**: bound usage `lastUsed` scan to a 2-day window; map large budget tokens to `max` reasoning tier
- **Docker**: publish verified multi-platform images (linux/amd64 and linux/arm64) with configurable apk build mirrors
2026-09-22 15:43:30 +07:00
Welington
0f488c7027 fix(translator): map Claude "refusal" stop_reason to content_filter and surface its explanation
Anthropic's API-level refusal (streaming classifier / ToS) ends the stream
with stop_reason "refusal", stop_details carrying the reason, zero output
tokens and no content blocks. Map refusal to content_filter in both
directions, surface stop_details.explanation as message text, and add
CLAUDE_STOP.REFUSAL to schema.
2026-09-22 15:32:48 +07:00
decolua
1a02713150 feat(combos): hide preset buttons and migrate legacy mimo vision adapter
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 15:32:10 +07:00
decolua
b53260ca54 feat(usage): add All Time period option and refine overview cards
- Add "all" period option to usage dashboard and chart API
- Aggregate all days in getChartData when period is "all"
- Center overview card metrics and adjust font size to prevent truncation

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 15:31:24 +07:00
Rafli Ahmad Zulfikar
d1de324586 perf(usage): bound lastUsed overlay scan to 2-day window; reach max thinking tier
getUsageStats("all") shipped the entire usageHistory table to JS just to
refine lastUsed (~2s on 290K rows, on every statsEmitter update per SSE
listener). Bound the overlay to a 2-day indexed range scan; older entries
keep day-level lastUsed from usageDaily aggregates. Totals unaffected.

budgetToLevel now maps budgets > 80384 (midpoint of 32768/128000) to
"max" instead of clamping to "xhigh", so the top reasoning tier is
reachable from large budget_tokens requests.
2026-09-22 15:08:38 +07:00
decolua
ce9ac43da5 style(sidebar): match NEW badge style across 9Remote, Media Providers, and System One
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 15:08:09 +07:00
BFLabsAI
5798b30841 fix(antigravity): drop requestType "agent" to avoid false 429 RESOURCE_EXHAUSTED 2026-09-22 15:04:16 +07:00
yiwen65
782c137b1f fix(qoder): prevent signed request replay and surface upstream errors 2026-09-22 15:01:06 +07:00
Ghost Machine
0d50fe3610 feat(analytics): add Requests mode and provider/model breakdown charts 2026-09-22 14:55:31 +07:00
dajiaohuang
3279d31024 docs(cli): fix source README link 2026-09-22 14:55:02 +07:00
rspersahabatan
aa2bc53f3a docs(readme): add temporary swap instructions for low-memory production builds 2026-09-22 14:54:12 +07:00
cokvrindaa
c73bb2cd66 docs(readme): add Indonesian video guide by Neptiver 2026-09-22 14:53:46 +07:00
decolua
6c9fe6f78a feat(cli-tools): add dynamic configuration for Pi, OMP, Crush, ForgeCode, Smelt and CodeWhale
- Add dedicated settings API routes for pi, omp, crush, forge, smelt, codewhale
- Integrate GenericCliToolCard with multi-model support for Pi and auto-discovery for OMP
- Register tools in cliTools catalog and all-statuses route
- Add official logos for all new CLI tools

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 14:48:04 +07:00
decolua
44fd69d229 style(dashboard): match 9remote NEW badge style on System One tags
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 14:46:57 +07:00
decolua
f84c667d42 feat(providers): add OpenRouter System One lane and New badge
OpenRouter serves TypeSafe Jev at POST /api/v1/systemone with the same
request/response shape, so it plugs into systemoneConfig directly with
model typesafe/jev-1.13. Mark the System One media kind isNew and
render a New badge on the sidebar kind item and the Media Providers
accordion when any visible kind is new.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 14:42:02 +07:00
decolua
20014b3104 fix(dashboard): probe System One models through /v1/systemone
The inline model test sent a chat-completions payload and failed with
500 on decision models. Add a systemone branch to pingModelByKind that
submits a native state+questions probe, and restore the test button
that was hidden for System One models.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 14:41:59 +07:00
decolua
f28e918e24 feat(dashboard): add question input for System One and hide inline test
Allow customizing evaluation instructions in the System One example
card, and disable the inline model probe button for System One models
since decision models do not accept chat completion probes.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 14:21:15 +07:00
decolua
6431e35303 feat(dashboard): add System One to sidebar and media provider detail page
Expose System One in the Media Providers sidebar accordion, set
kind to "systemone" on Jev models for ModelsCard filtering, wire
systemoneConfig into ProviderInfoCard, and configure GenericExampleCard
for interactive testing of decision models.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 14:07:36 +07:00
11089ab137 Merge origin/master (v0.5.81) into gitea/new_feature
Resolve conflicts:
- streamingHandler.js: merge buildStreamErrorBytes onAbortTerminal + local shouldPersistRequestDetail & streamStatusForContent
- capabilities.js: preserve server-injected user-asserted caps and models.dev catalog lookup; wire CommandCode /alpha/generate caps inside resolve()
- commandcode.js (services/usage): adopt upstream whoami + billing credits/subscriptions with 5h/weekly rate windows and plan caps
- openai-to-commandcode.js: merge toNativeImageBlock (data-URI & http(s) support) and assistant reasoning_content preservation
- commandcode-to-openai.js: adopt upstream mid-stream error throw for clean retry and abortion
- tests: sync commandcode test suite and exclude .next from vitest config
2026-09-22 10:08:58 +07:00
decolua
41a1b8003d feat(providers): add MiMo V2.6 Flash Free to OpenCode Zen free tier
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 09:51:51 +07:00
decolua
b7446f8dd1 feat(providers): add System One (Jev) decision endpoint
New /v1/systemone pass-through route for Jev decision models (jev-1.13,
jev-1.13-free) on OpenCode Zen and the free lane. Follows the media-route
pattern: systemoneConfig in the registry drives URL/headers, the handler
mirrors the embeddings account-fallback + usage flow, and the dashboard
gains a System One media-provider kind. No chat-pipeline changes.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 09:51:29 +07:00
decolua
6886915f62 feat(capacity-adapter): default vision fallback to mimo-v2.6-flash-free
Register mimo-v2.6-flash-free on opencode-zen (chat lane) with a v2.6
capability pattern, and switch the vision adapter default from the old
mimo-v2.5-free.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 09:45:20 +07:00
82030615
da0046550a fix(responses): report usage on response.completed so clients can auto-compact
Map upstream Chat Completions usage to the Responses API shape and attach it to response.completed. Capture chunk.usage before the empty-choices guard so the usage-only trailer chunk survives, and defer completion to flushEvents() when usage is not yet known — only on the direct openai:openai-responses route, since a pivoted stream never reaches flushEvents. Fixes #3432.
2026-09-21 21:27:01 +07:00
Liang.Xu
402745dc1f feat(providers): add qoder-cn support for Qoder CN (qoder.com.cn) 2026-09-21 20:45:19 +07:00
DaDecky
f67d5a0c93 fix(docker): publish verified multi-platform images
- Build linux/amd64 and linux/arm64 on native GitHub runners
- Assemble version manifests from platform digests and promote latest only after verification
- Add release/tag validation, manual republishing, timeouts, and health smoke tests
- Make Docker build mirrors configurable via build args and remove unnecessary runtime apk upgrades
- Update DOCKER.md documentation
2026-09-21 20:28:45 +07:00
82030615
c7df895bbb build(docker): apply the apk mirror swap to the runner stage
- Add apk mirror swap to mirrors.aliyun.com at the top of runner stage
- Prevents docker build hanging when dl-cdn.alpinelinux.org is unreachable
2026-09-21 20:22:19 +07:00
Christian Gennari
be3bc764b1 fix(antigravity): separate weekly and short-window quotas and clean up redundant rows
- Track both weekly and 5-hour session buckets in parseWeeklyQuotaSummary,
  distinguishing sliding-window limits from multi-day weekly limits
- Preserve disabled session buckets at 0% rather than dropping them when weekly limits are reached
- Target 5-hour session rows (not weekly rows) during family exhaustion reconciliation in getAntigravityUsage
- Suppress synthesized per-model duplicate rows in dashboard normalization when family summaries are present
- Add unit test coverage for multi-bucket extraction, reconciliation isolation, and dashboard deduplication
2026-09-21 20:15:56 +07:00
dinhkarate
2daf25ffbe fix(qoder): handle code 110 billing blocks and preserve SSE error status
- Match code 110 (billing daily count exceeded) alongside 112/10605/pricingUrl
  in isBillingBlock, parsing JSON safely and accepting numeric/string codes
- Accept numeric strings for statusCodeValue and object bodies in envelope peek
- Emit structured 403 quota error chunk instead of synthetic assistant text
  when a billing envelope appears mid-stream
- Preserve upstream HTTP status in handleForcedSSEToJson when error chunk carries
  a valid 400-599 status
- Add unit tests for code-110 detection, mid-stream billing envelopes, and false-positive guard
2026-09-21 20:09:03 +07:00
Anantachoke
7c2b1fe3e1 fix(translator): drop replayed reasoning fields for Groq/Mistral/Cerebras (#4220)
Strict OpenAI-compatible validators reject unknown assistant-message
fields: Groq 400 ("property 'reasoning_content' is unsupported"),
Mistral 422 ("extra_forbidden"), Cerebras 400 ("wrong_api_format").

Clients driving reasoning models (Hermes Agent, and anything following
the DeepSeek/Kimi convention) echo the previous turn's reasoning_content
on every assistant message, so from the second turn on every request to
these providers fails and a fallback combo silently skips them.

Add a dropMessageFields rule to paramSupport.js that strips
reasoning_content / reasoning / reasoning_details from assistant turns
for groq, mistral, and cerebras.
2026-09-21 20:02:21 +07:00
Amir Seify
253199f16f feat(combos): Cursor/Claude Default presets + bulk select/delete/strategy
- Add Cursor Default / Claude Default on Dashboard -> Combos to generate
  unprefixed combo names that match Cursor/Claude client model IDs,
  seeded with cu/... or cc/... so those clients can route through 9Router.
- Add multi-select bulk Delete and bulk Set strategy (Fallback / Round Robin / Fusion).
- Docs and unit tests for preset builder.
2026-09-21 19:51:26 +07:00
Amir Seify
c933eefc27 fix(cursor): stop AgentService empty turns and silent tool hangs
Cursor-hosted models (cu/composer-2.5, cu/cursor-grok-*, cu/default) returned
HTTP 200 with an empty turn, or hung, whenever a client sent tools.

- Fold system prompts into the current user message. custom_system_prompt
  (RunRequest field 8) makes AgentService return an empty turn.
- Send ModelDetails (field 3); thinking variants (Composer, Grok, *-thinking)
  return an empty turn when only requested_model (field 9) is set.
- Route tool-call history and declared tool schemas through AgentService:
  encode OpenAI tools into mcp_tools (field 4), decode McpArgs and emit real
  tool_calls with finish_reason tool_calls.
- Map Composer  thinking / Grok thinking_delta (field 4) into visible content
  instead of dropping the answer with the unsigned reasoning.
- Ack request_context without echoing MCP tools (double-advertise stalls the
  HTTP/2 stream) and ack kv_server_message so the run proceeds.
- Reject IDE builtin execs instead of failing the turn, so the model can
  continue with MCP tools or a text answer.
- Add google.protobuf.Value / MCP encoders and a FIXED64 branch to
  encodeField in cursorProtobuf.js.

RTK now compresses the source-format body before translation for cursor only:
its translator rewrites role:tool into user XML, so the post-translate pass
missed those tool results. Every other provider keeps the post-translate pass
unchanged.
2026-09-21 19:49:48 +07:00
Minh Ha
5c217d34f3 feat(capabilities): model capability metadata on /v1/models, combo aggregation, pattern fixes
- Export aggregateComboCapabilities: union for vision/audio/search/pdf,
  intersection for tools, primary-model for reasoning fields, min
  contextWindow, max maxOutput
- Support nested combo resolution in aggregateComboCapabilities via
  comboLookup with depth guard (max 6)
- Wire capability metadata to all /v1/models entries and combos
- Show aggregated ctx/max metadata line and capability badges on combo chips
- Pattern fixes: MiMo v2.5/omni reasoning, qwen max/plus vision, minimax m2.x vision
- Sync commandcode model catalog and add openai gpt-5.5
- Add unit tests for capability patterns and combo capability aggregation
2026-09-21 19:41:29 +07:00
chisewaguri
477b2aed0b fix(opencode-go): send reasoning_effort for glm-5.3-flash 2026-09-21 19:10:14 +07:00
decolua
9f42e7ac1c fix(sidebar): restore NEW badge for 9Remote menu
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-21 18:56:56 +07:00
kimono381
cf663f5300 fix(huggingface): complete the Inference Providers router migration
Replace the retired api-inference.huggingface.co host with the Inference
Providers router (router.huggingface.co): imageConfig.modelMap resolves
Hub ids to provider-resolved ids, image-to-image models receive the
source image in inputs with the prompt under parameters.prompt, and a
new sttConfig wires the hf-inference ASR route. The image catalog grows
to 23 models, dead whisper-small is replaced by whisper-large-v3-turbo,
the unusable "language" param is dropped, and edit models declare the
edit capability so the dashboard offers a source image. Adds unit and
end-to-end coverage plus a model-id guard on custom endpoints.
2026-09-19 10:44:34 +07:00
Aaron
822aa958d1 fix(opencode): cloak Responses requests that already have tools
Free-tier Zen models reject Responses requests with 403 FreeTierError
when client tools are present but the fingerprint quartet is missing.
Apply the fingerprint tools to every OpenCode request, canonicalise
case variants of the quartet (Bash->bash) without duplication, and
restore the caller's original spellings on the response side via a
request-local WeakMap threaded through the existing toolNameMap.
2026-09-19 10:15:45 +07:00
Alexander Radchenko
49185137b8 feat(opencode-zen): add OpenCode Zen (PAYG) provider with free-tier fingerprint, alias ocz
Multi-endpoint provider on https://opencode.ai/zen/v1 (openai / claude /
openai-responses transports mirroring opencode-go) with 71 models across
paid + free tiers. Free-tier fingerprint: opencode/1.18.x UA spoof,
ses_-session header, built-in tool quartet, forced stream. Usage endpoint
/zen/v1/usage wired into the dashboard.
2026-09-19 10:02:28 +07:00
X-Adam
73e021b8a0 fix(ollama): map free-plan monthly window and derive reset from signup date 2026-09-19 09:51:13 +07:00
decolua
a8c9d3802c docs: update changelog header to v0.5.81
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-18 18:32:14 +07:00
decolua
23ae82d8e3 # v0.5.81 (2026-09-18)
## Features
- **Xiaomi MiMo**: merge MiMo Desktop support into `xiaomi-mimo` with dual auth (API key + Desktop/OAuth session), Preview models support, and encrypted-callback OAuth flow
- **Claude Code**: add 1M-context toggle (`[1m]` marker) and drive `CLAUDE_CODE_AUTO_COMPACT_WINDOW` directly from the dashboard
- **Models**: add DeepSeek-V4.1-Flash to DeepSeek provider, CodeBuddy-Intl, and Ollama (`deepseek-v4.1-flash:cloud`); enable `low`..`max` reasoning effort levels and vision capability for DeepSeek-V4.*
- **i18n**: integrate Persian (fa) translation

## Fixes
- **OpenCode / OpenCode Go**: resolve 403 `FreeTierError` and 429 rate limits with canonical session format, valid User-Agent, and stable upstream session reuse; force stream and declare `forceStream` for free-tier SSE aggregation; cloak decoy tools, normalize Muse Free tool choice, and strip prior reasoning items on Responses models; route Union Alpha via Messages API
- **Kiro**: preserve underscores in tool names (`mcp__server__tool`) and restore client tool names in responses; use neutral placeholder for tool-result-only turns; forward tool-result images
- **Stream**: report aborts after HTTP 200 in-band (per-format error frames) instead of closing silently
- **Command Code**: preserve images and `reasoning_effort` on `/alpha/generate`; retry transient stream errors and avoid fake stop chunks; add Quota Tracker support
- **Zed**: harden OAuth lifecycle (preserve `systemId`, renew proxy timeout), support live model resolution, and lower display priority in OAuth list
- **Antigravity**: scope cached thought signatures to model family; strip Claude Code billing headers from system prompts; sanitize Hermes system identity
- **Codex**: route bare `codex-auto-review` requests to the Codex provider (#4135)
- **Auth**: do not cool down an account for request-scoped 4xx errors
- **Usage**: improve DeepSeek credit balance display as currency credit instead of 0/total quota bar
- **Model Catalog**: scope synced catalog to gateways and declare vision capabilities for DeepSeek V4.1-Flash IDs
2026-09-18 18:31:12 +07:00
decolua
8e15f0bdd8 # v0.5.79 (2026-09-18)
## Features
- **Xiaomi MiMo**: merge MiMo Desktop support into `xiaomi-mimo` with dual auth (API key + Desktop/OAuth session), Preview models support, and encrypted-callback OAuth flow
- **Claude Code**: add 1M-context toggle (`[1m]` marker) and drive `CLAUDE_CODE_AUTO_COMPACT_WINDOW` directly from the dashboard
- **Models**: add DeepSeek-V4.1-Flash to DeepSeek provider, CodeBuddy-Intl, and Ollama (`deepseek-v4.1-flash:cloud`); enable `low`..`max` reasoning effort levels and vision capability for DeepSeek-V4.*
- **i18n**: integrate Persian (fa) translation

## Fixes
- **OpenCode / OpenCode Go**: resolve 403 `FreeTierError` and 429 rate limits with canonical session format, valid User-Agent, and stable upstream session reuse; force stream and declare `forceStream` for free-tier SSE aggregation; cloak decoy tools, normalize Muse Free tool choice, and strip prior reasoning items on Responses models; route Union Alpha via Messages API
- **Kiro**: preserve underscores in tool names (`mcp__server__tool`) and restore client tool names in responses; use neutral placeholder for tool-result-only turns; forward tool-result images
- **Stream**: report aborts after HTTP 200 in-band (per-format error frames) instead of closing silently
- **Command Code**: preserve images and `reasoning_effort` on `/alpha/generate`; retry transient stream errors and avoid fake stop chunks; add Quota Tracker support
- **Zed**: harden OAuth lifecycle (preserve `systemId`, renew proxy timeout), support live model resolution, and lower display priority in OAuth list
- **Antigravity**: scope cached thought signatures to model family; strip Claude Code billing headers from system prompts; sanitize Hermes system identity
- **Codex**: route bare `codex-auto-review` requests to the Codex provider (#4135)
- **Auth**: do not cool down an account for request-scoped 4xx errors
- **Usage**: improve DeepSeek credit balance display as currency credit instead of 0/total quota bar
- **Model Catalog**: scope synced catalog to gateways and declare vision capabilities for DeepSeek V4.1-Flash IDs
2026-09-18 18:09:32 +07:00
decolua
058ceace48 fix(opencode): declare forceStream on transport for free-tier SSE aggregation
Declares forceStream: true so chatCore properly converts upstream
forced-stream responses to JSON for non-streaming callers.

Co-authored-by: anojndr <anojndr@gmail.com>
Co-authored-by: TEGAR-SRC <tegararrahman17@gmail.com>
Co-authored-by: yxxrn <yxxrn@users.noreply.github.com>
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-18 17:32:52 +07:00
decolua
d99bc8201d fix(sidebar): temporarily hide NEW badge from 9Remote menu
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-18 17:12:57 +07:00
Louis Phạm
bc3be0cb28 fix(antigravity): scope cached thought signatures to the model family 2026-09-18 17:09:36 +07:00
Christian Gennari
092c84eac9 fix(commandcode): retry on transient stream error and avoid fake stop chunks 2026-09-18 17:09:24 +07:00
Louis Phạm
b3d6e089c6 fix(antigravity): strip Claude Code billing header from system prompts 2026-09-18 17:08:53 +07:00
Amirsalar Sojoudi
efc80ba2e3 fix(codex): route bare codex-auto-review to the Codex provider (#4135) 2026-09-18 17:05:47 +07:00
decolua
4641c2b76a fix(zed): lower display priority in oauth list
Move zed provider to the bottom of oauth providers by setting priority to 999.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-18 16:59:45 +07:00
decolua
93837af09f fix(opencode): fix free tier 403 error and improve China region handling
- Force stream:true and cloak decoy tools (bash, read) for OpenCode free tier
- Support connection testing for opencode in testUtils
- Expand error message slice limits in auth and ping to preserve workspace link
- Add concise China region link chip in provider detail page

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-18 16:57:06 +07:00
decolua
52917a6d49 Add NEW badge to 9Remote menu 2026-09-17 20:18:56 +07:00
decolua
f64229530e fix(antigravity): sanitize Hermes system identity 2026-09-17 20:18:37 +07:00
Ali Shaikh
3ac100d524 fix(usage): improve DeepSeek credit balance display
Display DeepSeek prepaid balances as credit with currency instead of a 0/total quota bar, and mark balance items as isCreditBalance.
2026-09-17 20:06:33 +07:00
Mosabbir Maruf
ef18175226 fix(zed): harden OAuth lifecycle and live model support
- executors/zed.js: use exact wire values (anthropic, open_ai, google, x_ai)
  and strip incompatible Vertex safetySettings on the Google path
- shared/zedAuth.js: robust callback query parsing, reject garbage PKCS#1 v1.5
  decryptions, and thread proxyOptions when fetching LLM tokens
- oauth: preserve systemId across authorize/register/exchange lifecycle,
  renew proxy idle timeout on reuse, and ignore non-callback localhost requests
- shared/OAuthModal.js: track owned proxy in flowRef and stop at most once
- api/providers/[id]/models: add connection-scoped live Zed model resolver
- registry: unhide provider in dashboard
- tests: add unit coverage for wire format, native auth, and live models
2026-09-17 20:02:57 +07:00
Ahmad Beyranvand
725e2c1187 feat(i18n): integrate Persian (fa) translation 2026-09-17 18:45:58 +07:00
Rafli Ahmad Zulfikar
367fc546d8 feat(models): deepseek-v4.* accepts low..max effort; flag thinkingEffortSupported and vision 2026-09-17 18:40:07 +07:00
RaoYu
20a43f5a2c fix(auth): don't cool down an account for a request-scoped 4xx
Do not trigger account cooldown or fallback for request-scoped 4xx errors that match no account rules so healthy credentials are not locked out for context length or validation errors.
2026-09-17 18:26:22 +07:00
KunN-21
aa14ef72e2 fix(opencode): normalize Muse Free tool choice
OpenCode Free returns HTTP 400 for muse-spark-1.3-contributor-free when tool_choice is non-auto. Declare forceAutoToolChoiceModels quirk and normalize explicit tool_choice to auto.
2026-09-17 18:20:08 +07:00
anojndr
eafac37dcb fix(opencode): strip prior reasoning items on Muse Spark Responses models
Strip prior-turn type: reasoning items and encrypted_content fields from body.input on Muse Spark Responses endpoints in OpenCode and OpenCode Go executors to avoid HTTP 400 errors across rotated proxy accounts.
2026-09-17 18:19:21 +07:00
Qisthi Ramadhani
c49efdf528 fix(kiro): preserve underscores in tool names and restore sanitized names in responses
Do not collapse consecutive underscores in uniqueName so mcp__server__tool is sent intact to Kiro, attach reverse map on request translation, and restore client tool names in responses.
2026-09-17 18:14:59 +07:00
Qisthi Ramadhani
82b1bca42a fix(kiro): use neutral placeholder for tool-result-only user turns
Replace the literal 'continue' placeholder on tool-result-only user turns with 'Tool results provided.' to prevent models from treating it as a new user instruction.
2026-09-17 18:12:48 +07:00
Manan Santoki
f4f06f290c fix(translator): keep tool-result images, restore Kiro tool names, preserve thinking display
Forward images inside tool_result to OpenAI and Kiro upstreams via following user messages, restore original client tool names on Kiro responses via _toolNameMap, and preserve thinking display settings across translations.
2026-09-17 18:12:05 +07:00
ErfanBagheri404
0c6ab4f99b fix(opencode): reuse one stable upstream session per identity to stop 429s
Follow-up to the canonical-session fix: with no explicit session,
every request minted a fresh x-opencode-session, and upstream free-tier
quota is accounted per session. That burns through quota and surfaces
as 429 FreeUsageLimitError with growing reset-after delays, while the
real CLI reuses one long-lived session per conversation.

- Stable canonical session per downstream identity (connectionId, else
  auth-header hash, else shared default), evicted after
  MEMORY_CONFIG.sessionTtlMs like the other session stores.
- Deterministic x-opencode-request per message (stable across retries,
  like the CLI user message id); valid downstream ids preserved.
- 6 more unit tests (22 total).
2026-09-17 18:03:58 +07:00
anojndr
6091ff597e fix(opencode): resolve 403 FreeTierError with canonical session format and valid User-Agent
OpenCode upstream validates free-tier requests: User-Agent must be opencode/<version> (>= 1.17.0) and x-opencode-session must match canonical ses_ format. Default OPENCODE_UA to opencode/1.18.31, generate canonical descending session IDs, provide deterministic foreign session translation, and isolate credentials per-request.
2026-09-17 18:03:20 +07:00
Hermes Agent
2b65c49ff5 fix(opencode): route Union Alpha through Messages API
Route union-alpha to /zen/v1/messages with targetFormat claude, add anthropic-version header, and register model capabilities (vision, 262K context, 131K max output).
2026-09-17 18:01:43 +07:00
Ridho Perdana
702b57c30d fix(opencode-go): route every responses-only model (incl. thinking variants) to /responses
Derive responses-only routing from the model registry's targetFormat instead of hardcoding model checks, and strip thinking suffixes when looking up models in providerModels so variants like gpt-5.6-luna(high) are routed correctly to /responses.
2026-09-17 18:00:01 +07:00
izzzzzi
912ed295db fix(deepseek,model-catalog): vision for V4.1-Flash ids, scope synced catalog to gateways
- Declare deepseek-v4.1-flash and deepseek-flash as vision-capable in MODEL_CAPABILITIES
- Share installed catalogSource across route chunks via globalThis.__9rCatalogSource
- Scope catalog modality keys by provider:model to prevent cross-gateway collisions
- Upgrade catalog format to v2 with automatic rebuild of older schemas
2026-09-17 17:55:33 +07:00
28e26f4295 Merge origin/master (v0.5.75) into gitea/new_feature
Resolve conflicts:
- package.json / cli/package.json: take 0.5.75
- .gitignore: union both sides (upstream 9router-*/temp files + local state dirs)
- CHANGELOG.md: keep both blocks, v0.5.75 above v0.5.70
- nonStreamingHandler.js: merge imports (unwrapClineEnvelope +
  tokensForDetail/shouldPersistRequestDetail); drop dead appendRequestLog
- providers/[id]/page.js: union useState blocks (compatible-model states
  + importingClineModels)

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-09-17 14:14:43 +07:00
81a47f36c6 fix(combo): keep nested combos as single units and stop 0-token detail rows
Nested combos (comboA lists comboB, comboC, …) now stay one slot each:
the inner combo always runs as fallback to produce a single answer.
Failed hops are no longer written to Details/usage, and streaming no
longer inserts a 0-token placeholder row.

- chat.js: comboStack cycle detection; nested combos forced to fallback;
  persistUsage="success-only" for combo hops
- combo.js: discardResponse() cancels unused bodies (fusion timeout /
  fallback) so dropped streams fire onStreamComplete; getComboModelsFromData
  keeps nested names and honors enabled=false
- requestDetail.js: tokensForDetail() canonicalizes Claude/Gemini usage;
  shouldPersistRequestDetail() skips streaming-start and non-success hops
- streamingHandler.js: drop the 0-token streaming placeholder write
- RequestDetailsTab.js: read Gemini/Claude token names; show "streaming"
  status in amber
- tests: add combo-nested.test.js (13 cases)
- gitignore: ignore local .vitest/ artifacts

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-09-17 14:03:59 +07:00
KhuatHieu
13b468b889 fix(commandcode): preserve images and reasoning_effort on /alpha/generate
Command Code dropped vision and ignored client effort through the router:
image blocks became "[image omitted]", HTTP image URLs were never inlined,
and reasoning_effort landed on the envelope wrapper instead of params (so the
DeepSeek family mapping remapped low -> high). The catalog also treated
deepseek/deepseek-v4.1-flash as text-only, so the vision adapter stole those
requests to another provider.

- Map OpenAI image_url / Claude image blocks (base64 or data-URI) to the
  native {type:"image", image:"data:...;base64,...", mimeType} generate block.
- Add FORMATS.COMMANDCODE to TARGETS_NEED_BASE64 so remote http(s) images are
  inlined by the existing SSRF-safe fetcher before translation.
- Write reasoning_effort inside params for targetFormat commandcode and pass
  low|medium|high|xhigh|max through unmapped; allow it in thinkingLevels.
- Provider-scoped capabilities for commandcode/cmc: vision except the CLI
  text-only denylist, thinkingFormat commandcode, so family patterns
  (deepseek-v4 -> thinkingFormat deepseek, vision false) no longer win.
- Quota Tracker: whoami + billing credits/subscriptions (credits vs plan cap,
  5h and weekly windows), labels from AI_PROVIDERS[].name.
2026-09-16 20:17:25 +07:00
decolua
9300121366 fix(stream): report aborts after HTTP 200 in-band instead of closing silently
A stream that stalled or lost its upstream was closed with no terminal frame
at all, so clients saw "200 OK, a few chunks, then nothing" and could not tell
a truncated reply from a finished one. The Responses passthrough path already
synthesized response.failed; every other client format got nothing.

The watchdog now hands its reason ("stream stall timeout" or "upstream
connection lost") to onAbortTerminal, and buildStreamErrorBytes frames it per
client format: OpenAI-compatible clients get data: {"error":{...}} followed by
data: [DONE], Anthropic clients get `event: error`. The error frame always
precedes [DONE] (openai-python raises APIError on any data payload carrying an
error key), and no synthetic finish_reason is ever emitted — a truncated
stream must not look like a clean stop.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-16 20:04:10 +07:00
decolua
5c399b6406 feat(codebuddy-intl,ollama): add DeepSeek-V4.1-Flash
codebuddy-intl: deepseek-v4-flash replaced by deepseek-v4.1-flash (same
gateway catalog as CN) and a capability override so the model keeps the
openai-style reasoning_effort path instead of the vendor-native "deepseek"
thinking shape the gateway rejects. Thinking levels low/high/xhigh.

ollama: add deepseek-v4.1-flash:cloud (verified on ollama.com/api/tags) with
vision + 1M context caps.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-11 22:03:48 +07:00
decolua
17c4cc7687 feat(claude-code): drive auto-compact window, add a 1M-context toggle
The "Context window" dropdown wrote CLAUDE_CODE_MAX_CONTEXT_TOKENS, which
Claude Code ignores for any model it recognizes: its window resolver returns
the env value only when the id is unknown to the model table, so every
claude-* mapping kept the built-in 200K and the dropdown did nothing. It was
never the compaction threshold either.

- Replace it with CLAUDE_CODE_AUTO_COMPACT_WINDOW — the documented trigger
  (100K–1M, clamped to the model window, env beats the autoCompactWindow
  setting) — and relabel the field Auto-compact. The 1M preset becomes 700K,
  which no longer collides with the marker it depends on.
- Add a "1M context" checkbox that appends the `[1m]` marker to the
  ANTHROPIC_DEFAULT_*_MODEL envs. Claude Code assumes 200K unless the name
  carries the marker — the resolver is a plain /\[1m\]/i test on the string,
  so it applies to any id and no model lookup is involved; the user decides
  which models are worth declaring as 1M.
- Toggling rewrites the model inputs immediately, and Apply writes them
  verbatim, so a marker typed by hand is not stripped.

Rename maxContextTokens -> autoCompactWindow through the POST body and
RESET_ENV_KEYS so a reset clears the key actually written.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-11 00:09:15 +07:00
叶炜朋
73cb89143c feat(xiaomi-mimo): merge MiMo Desktop support into xiaomi-mimo as dual auth
Adds the Desktop-exclusive Preview models and the Xiaomi account-session
route to the existing xiaomi-mimo provider instead of a separate
xiaomi-desktop provider, so the dashboard shows one MiMo entry rather than
three overlapping ones.

Dual auth, same pattern as kimi — API key (sk-) covers the cloud API,
Desktop/OAuth adds the account session used by the Preview models:

- registry: category oauth, authModes [oauth, apikey], oauth block, the two
  mimo-x-*-preview models, and the invite signupUrl
- executor: routes Preview models to the account-service route with a Cookie
  session, everything else keeps the sourceFormat-matched transport
- oauth: custom ECDH encrypted-callback flow (X25519 -> SHA256 -> AES-256-GCM)
  with a loopback callback proxy, plus one-click import of the local Desktop
  auth.json
- usage: weekly quota from the account session

Fixes found while merging:

- the OAuth browser flow was dead: poll-status cleared the session before the
  client could POST /exchange, so every exchange returned 400
- a Claude-format client was sent to /v1/chat/completions instead of the
  declared /anthropic/v1/messages transport, because buildUrl ignored
  runtimeTransport
- stopXiaomiMimoProxy leaked every pending session (each holding an X25519
  private key) for the process lifetime
- the OAuth exchange did not persist the Desktop passToken, so the Preview
  models could never work after a browser sign-in

Removes dead code: the local engine token minting (mimoEngine, never called
on the request path), the model-catalog and usage routes, engineToken/
engineUrl plumbing, and an unread top-level usage block.

Adds tests/unit/xiaomi-mimo-{executor,oauth-session,oauth-proxy}.test.js —
the provider previously had none.
2026-09-10 23:42:41 +07:00
decolua
83af3f1853 # v0.5.75 (2026-09-10)
Covers the 24 commits since the v0.5.69 tag. The package version already moved
to 0.5.75 in 4a390685b (CLI model selector), so the release commit is the
changelog alone, matching the convention in eb712ca82/4eda76e2a.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-10 23:35:10 +07:00
decolua
accf2c5296 fix(providers): remove the duplicate qwen provider
qwen.js points at the same transport as alims-intl — identical baseUrl
(dashscope-intl compatible-mode), headers, and quirks — but carries only 8
Qwen models against alims-intl's broader catalog, so it adds nothing a user
could not already reach. Node count is back to 81.

The id was dropped on 2026-08-05 (dcdd4628b) when the Qwen OAuth flow died;
this removes the standalone API-key entry that shadowed it. The `qw` alias
returns to unassigned, its state before today.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-10 23:32:55 +07:00
decolua
c712641123 feat(opencode-go): list deepseek-v4.1-flash first in the model catalog
Registry order is the display order for the provider page, /v1/models, and the
CLI selector, so moving the entry to the head of `models` is the whole change.
deepseek-flash keeps its id and supportedFormats — only its position moves.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-10 23:18:58 +07:00
decolua
da6aa90128 fix(video/vertex): reject job ids and model ids that escape the URL path
A Vertex job id is a base64url-encoded operation name, and base64url decoding
accepts arbitrary bytes without throwing, so the previous decode-and-split
check let a crafted id splice a path traversal into the fetch URL while the
Bearer token stayed attached — e.g. "..%2F..%2Fevil" resolved to
/v1/evil:fetchPredictOperation on the Vertex host. body.model had the same
shape on the create path, where it is interpolated into the URL unescaped.

decodeJobId now requires a charset-only id, a byte-for-byte round-trip, and a
decoded name matching ^projects/{p}/locations/{l}/publishers/{pub}/models/{m}/operations/{op}$
— no field may contain "/", so ".." can never reach the URL. model ids are
restricted to [A-Za-z0-9._-].

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-10 23:18:52 +07:00
LLL
248d7da01c revert(qoder): drop the Responses usage plumbing from shared code
The merged Qoder work also rewrote shared translator/handler code so that
/v1/responses clients got token usage on response.completed. That changed
behaviour for every provider, not just Qoder: proxies saw input tokens
rise by the 2000-token context buffer, and the plain token mapping was
replaced by one that always adds input_tokens_details.

A probe confirms the Qoder benefit does not depend on those edits: the
executor's coalescer already emits one include_usage-style finish chunk, so
a Claude client receives input_tokens and cache_read_input_tokens with
every shared file at its original state. Only the Responses path relies on
the shared translator, and that path has no Qoder-owned seam to put it in.

Reverts the shared files to their pre-PR state and drops the Responses
usage test. The Cline envelope unwrap in nonStreamingHandler.js, which
landed after the PR in the same file, is kept.
2026-09-10 23:13:06 +07:00
galiehneh
998bb3d975 fix(tools): scope Claude tool type defaulting to gateways that need it (#3905)
defaultClaudeToolType() stamped tools[].type = "custom" onto every
Claude-format request carrying tools since e08ac6da. That satisfied
MiniMax (error 2013) but broke Anthropic-compatible endpoints that only
accept the legacy typeless tool shape. DeepSeek's endpoint
(api.deepseek.com/anthropic/v1/messages) whitelists its tool `type` enum
to the web_search_* variants and answers HTTP 400 "unknown variant
`custom`", so every Claude Code request routed to a DeepSeek connection
failed and surfaced as a persistent 503.

Run the defaulting only when the target provider declares the new
requireClaudeToolType quirk (MiniMax, MiniMax-CN). Add
shouldDefaultClaudeToolType(provider, finalFormat, tools, PROVIDERS) in
translator/concerns/toolCall.js so the gate is unit-testable, and cover
MiniMax keeping the explicit type, DeepSeek/Anthropic staying typeless,
non-Claude formats and tool-less requests never defaulting.

Effectively a no-op for MiniMax and a restore of the pre-e08ac6da
behaviour everywhere else. Another strict gateway now only needs the same
one-line quirk instead of a global behavioural change.
2026-09-10 23:03:53 +07:00
Nick Nyanjui
8a81085a72 fix(claude): cap re-anchored cache_control at the 4-marker budget and keep single-object content turns
Anthropic accepts at most 4 blocks carrying cache_control per request. When the
client had already spent that budget, the re-anchor added a 5th marker and the
request was rejected with a non-retryable 400 that the failure path treated as
an account problem, retrying the same malformed body across the whole pool until
every account locked.

anchorClaudeCache now normalizes bare-object content, strips the invalid
cache_control carried by defer_loading tools, pins the 1h head anchors on the
last system block and last cacheable tool, then trims an over-budget body to 4
markers. The trim holds those head anchors and fills the remaining slots with the
tail-most message markers: a plain "keep the last four in document order" rule
drops the anchors first even though they lead document order, and skipping the
re-anchor at a spent budget left system/tools on the 5m default instead of 1h.

Some clients send content as a single block object rather than a one-element
array. Such a turn was dropped or zeroed on every leg that reads messages,
silently losing conversation history. normalizeMessageContent wraps it as a
one-block array on all four paths, and hasValidContent keeps it.
2026-09-10 22:54:07 +07:00
Nick Nyanjui
122f23eebc fix(cline,airforce): unwrap {success,data} envelope, add live catalog, and refresh airforce free models
Cline (api.cline.bot) wraps non-stream chat completions in
{"success":true,"data":{...choices...}}, which both the dashboard model-test
ping and the proxy non-stream path read at top level, producing "Provider
returned no completion choices for this model" (#3644). Unwrap the envelope
before usage extraction and response translation; the error envelope
({"success":false,...}) never matches and passes through untouched.

Scoped through `transport.quirks.clineEnvelope` so only cline/clinepass opt
in — no other provider's response body is ever rewritten.

Also adds a live Cline catalog: `fetchClineRawModels()` is shared between
`resolveClineModels()` (full catalog, including free-tier ids such as
z-ai/glm-5.3-flash) and `resolveClinepassModels()` (cline-pass/* only), wired
into /v1/models, the per-provider models route, and the combo selector's
model picker with the static catalog kept as fallback.

Refreshes the dead api-airforce free models (anthropic/claude-3.7-sonnet,
moonshot/kimi-k2.6, google/gemini-2.5-flash) with the live gpt-oss-120b,
gpt-oss-20b and kimi-k2.7-code, plus passthroughModels, forceStream and a
suggested-models filter.
2026-09-10 22:48:28 +07:00
izzzzzi
f6e7cabe60 fix(cline): stop workos:-prefixing ClinePass API keys and add clinepass token refresh
Cline/ClinePass requests failed with HTTP 401 ("Please make sure you are using
the latest version of Cline and re-authenticate your Cline account", #3230 /
#2333 / #3644). `getClineAccessToken()` unconditionally prefixed every token
with `workos:`, which is correct for Cline OAuth access tokens (WorkOS JWTs)
but wrong for ClinePass API keys — those are opaque strings (e.g. `clp_…`)
that the API accepts only verbatim, so the `workos:`-prefixed value was
rejected.

Only prefix tokens that look like a WorkOS JWT (`eyJ…`); API keys and other
opaque tokens pass through untouched, and an existing `workos:` prefix is
never doubled.

Also register `clinepass` in the token-refresh handlers. ClinePass shares
Cline's WorkOS auth endpoints, but without the entry expired ClinePass OAuth
tokens were never rotated, so every request kept 401ing. Finally, list
`apikey` first in the ClinePass `authModes` (ClinePass is meant to be used
with an API key from app.cline.bot/settings/api-keys), and add an "Import
from /models" button that pulls the live Cline catalog into custom models.
2026-09-10 22:48:22 +07:00
kimono381
45ec1d30bb fix(deepseek): keep Anthropic-only tool types when forwarding to /anthropic/v1/messages
DeepSeek's Anthropic-compatible endpoint accepts only the built-in
web_search_20250305 / web_search_20260209 tools and rejects client-defined
`custom` tools (MCP / Read / Bash) with HTTP 400 "unknown variant `custom`".
The generic non-Claude filter in prepareClaudeRequest dropped the offending
tools but also dropped the web_search_* ones DeepSeek does accept.

- Add an opt-in per-provider transport quirk `claudeSupportedToolTypes`; when
  declared it becomes a strict allow-list for Anthropic tool `type` values
- Stop stripping the `type` discriminator from surviving tools under that
  quirk, since DeepSeek needs it to route built-ins
- Declare the quirk on the deepseek transport with the two web_search_* types
- Providers without the quirk keep the previous filter and normalisation
  behaviour byte-for-byte; openai-format targets never reach this path
2026-09-10 22:26:44 +07:00
anhtran-ai
781c18d837 fix(codex): strip Unicode-property tool schema patterns Codex rejects
Codex's /responses validator has no Unicode property escapes, so a tool
`pattern` containing `\p{...}` 400s the whole request with `Invalid schema
for function ... is not a 'regex'` — identically on every account, costing a
full combo failover per turn (#3922).

- Add open-sse/utils/codexToolSchema.js: copy-on-write walk that drops only
  `pattern` values carrying a property escape, returning the original
  reference when nothing changed so the caller's schema stays intact for a
  retry against another provider
- Treat `properties` keys as property names, so a field literally called
  `pattern` is never read as the schema keyword; skip escaped literals via
  backslash-parity counting
- Apply it in normalizeCodexTools for both function and namespace sub-tool
  parameters, and log the strip count via dbg
- Add three cases to tests/unit/codex-tool-normalization.test.js
2026-09-10 22:25:47 +07:00
Federico Liva
1892ed77c8 fix(kiro): never send a top-level systemPrompt (400 REQUEST_BODY_INVALID)
kiro.dev rejects any body carrying a top-level systemPrompt with
400 REQUEST_BODY_INVALID. The translators stopped emitting the field in
v0.5.59 (the prompt travels in the first user turn via contentPrefix),
but two paths kept writing it back downstream of the translator:

- rtk/systemInject.js::injectKiroSystem() appended the RTK prompt to
  body.systemPrompt, so every kr/ model failed whenever an RTK injector
  (caveman, ponytail) was active. It now appends to the first history
  user turn's content (else currentMessage), reusing
  dedupStringAppend/hasPrompt so retries stay idempotent.
- executors/kiro.js::appendRepairInstruction() wrote the tool-call repair
  instruction to systemPrompt on the retry, turning every repair into a
  hard failure. It now appends to currentMessage.userInputMessage.content.

isKiroBody() no longer requires a string body.systemPrompt — that marker
is gone from the wire shape — and sniffs the conversation turn shape
instead, keeping the stray-conversationState guard intact. Stale comments
in both kiro translators corrected: the systemPrompt local is only a
session-replay cache key, not a wire field.

Also drops the mirror/rollback repair heuristic the injector no longer
needs: net -52 lines.

Fixes #3641, #3845, #2890, #2901, #2939, #3109, #3459, #3749
2026-09-10 22:25:35 +07:00
LLL
1f10f9e5c4 fix(qoder): report usage to all clients and stop inlining large attachments
- Coalesce Qoder's empty finish-in-delta frame with the later choices:[] usage
  frame so OpenAI and Claude clients receive prompt_tokens, completion_tokens
  and cache-hit tokens (the dashboard already saw them)
- Upload inlined images through /api/v2/image/upload like qodercli, and stub
  oversized non-image files instead of stuffing 30MB+ data URIs into
  agent_chat_generation
- Emit response.completed -> response.usage for chat-native upstreams so
  /v1/responses clients (Codex CLI, sub2api) no longer log 0/0/0
- Keep Claude message_delta.usage working when usage arrives without choices[0]
- Escalate to the smallest advertised Qoder context tier (200K/400K/1M) when
  the estimated prompt no longer fits max_input_tokens
- Pass apiKey for PAT connections and list hidden enable:false catalog keys
  from /v1/models
2026-09-10 22:08:19 +07:00
Hai Trinh
832a34659e feat(codex): add GPT Image 2.5, Flare and Sunburst image models
- Add gpt-image-1.5, gpt-image-2, gpt-image-2.5, gpt-image-2.5-flare and
  gpt-image-2.5-sunburst as Codex image models with multi-image support
- Add gpt-image-2.5, gpt-image-2.5-flare and gpt-image-2.5-sunburst to the
  OpenAI provider catalog
- Route tool-backed image models through the Codex responses model while
  passing the selected model to the image_generation tool, pinning
  tool_choice and deriving generate/edit from the presence of references
- Cover the Codex gpt-image-2.5 request shape with a unit test
2026-09-10 22:06:48 +07:00
coozgan
3288bbc47e feat(video): add OpenRouter and Vertex AI (Veo) video generation
Video generation was xAI-only. Adds an adapter layer under
open-sse/handlers/videoProviders/ so /v1/videos/* can target OpenRouter or
Google Cloud credentials. A provider with no adapter keeps the exact previous
behaviour (raw body to {baseUrl}/{action}, poll {baseUrl}/{id}, verbatim
passthrough), so the xAI path is unchanged.

- openrouter: async job shape identical to xAI; creation POSTs to the /videos
  collection root (no /generations suffix) and the registry HTTP-Referer /
  X-Title headers are applied. Bodies pass through verbatim.
- vertex: two-way translation, since Veo does not speak the OpenAI-ish videos
  shape. create -> :predictLongRunning { instances[], parameters{} }, poll ->
  :fetchPredictOperation (Veo has no REST GET poll). The operation resource
  name is base64url-encoded into the job id so GET /v1/videos/{id} stays a
  flat path. Access tokens are minted from Service Account JSON via the
  existing refreshVertexToken; raw API keys are rejected up front. The
  operation response maps back onto the { id, status, video, videos } shape
  clients already poll.
- videoCore: the request plan is rebuilt per attempt, so the 401 -> refresh
  once -> retry once path picks up the refreshed token. Adapter validation
  errors return 400 before any upstream call, so a malformed request can never
  create a billable job.
- videoGeneration: GET /v1/videos/{id} resolves the provider from the pinned
  x-connection-id connection, then ?provider=, then falls back to the xAI
  default.
- registry: openrouter and vertex gain videoConfig, the video serviceKind and
  video-kind models (Veo 3.1 / 3 / 2, Sora 2 Pro, Seedance 2.0).
2026-09-10 22:05:22 +07:00
zmf
807553e246 feat(codebuddy-cn): replace deepseek-v4-flash with deepseek-v4.1-flash
The server's product-config payload (which the IDE plugin fetches from
copilot.tencent.com) publishes deepseek-v4.1-flash and no longer lists
deepseek-v4-flash, so the old id is dropped — same pattern as the previous
catalog refreshes (#3648, #3802). The v4-flash endpoint still answers 200,
but the published list is the contract.

Per the server table, maxOutput rises 50000 -> 128000 while contextWindow
stays 1000000.

- registry/codebuddy-cn.js: models[] entry swapped to the new id
- capabilities.js: per-model entry swapped, maxOutput -> 128000

No changes needed in thinkingLevels.js (the deepseek-v4* pattern already
matches the new id and publishes low/high/xhigh, matching the server's
supportedEfforts or pricing.js (the deepseek-v* glob yields the same rates).
EOF
)
2026-09-10 21:57:49 +07:00
decolua
4a390685b3 feat(cli): group model selector by provider with search
Replace the flat numbered model list with provider-grouped browsing
(combos first, then providers by alias order), full-text search across
all models, and manual custom model ID entry. A single available
category opens directly into its model list.

Also bump root and cli packages to 0.5.75 and ignore packed
`9router-*` tarballs.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-10 21:56:30 +07:00
decolua
537b3befd2 Merge branch 'feat/opencode-go-models'
feat(opencode-go): add newly published Go models
fix(codex): restore Version header and single-source the CLI version
2026-09-10 21:25:40 +07:00
decolua
a7047a07d4 fix(codex): restore Version header and single-source the CLI version
The image handler's `version` header was commented out, so Codex image
requests reached chatgpt.com without the Version identity the backend
expects. Restore it and route every Codex identity header through one
constant.

The CLI version now lives on registry codex.transport as `cliVersion`
(the same pattern gemini-cli uses) and is re-exported as CODEX_CLI_VERSION,
so the registry User-Agent, the image handler and the connection test can
no longer drift apart. Bumped 0.136.0 -> 0.154.0 (current stable).

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-10 21:23:24 +07:00
decolua
eee3515e54 feat(opencode-go): add newly published Go models
Add the models the provider docs now list but the registry lacked:

  chat/completions  glm-5.3, kimi-k3, deepseek-flash, longcat-2.0,
                    hy4-preview, hy3
  + /messages       qwen3.8-max, qwen3.8-flash
  responses only    grok-4.6, gpt-5.6-luna

Endpoints follow the table at https://opencode.ai/docs/go/. chat-only
models stay on the sourceFormat-matched transport guard so a Claude
client is never routed to /messages for a model that lacks it.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-10 21:08:32 +07:00
617929zcxc
40dffbce53 feat(providers): add standalone Qwen provider 2026-09-10 18:53:50 +07:00
a84ba559f3 # v0.5.70 (2026-09-08)
## Features
- **Providers**: creating a compatible / custom-embedding node now registers the endpoint only — the API key is added afterwards from the node's page, like the built-in providers. The create dialogs drop the API Key / Model ID / Check fields, and `POST /api/provider-nodes` no longer accepts credentials at all, so a node can never be half-created
- **Providers**: compatible nodes now use the same model rows as built-in providers — capability badges, copy, per-model test, alias handling and the Add/Edit Model modal with vision + reasoning toggles, replacing the weaker read-only list
- **Providers**: compatible nodes get the built-in bulk toolbars: Test All / Disable / Active / Select All over connections, and Test All Models / Disable All / Active All over models, with per-row disable and a restore strip for disabled models
- **Providers**: wire the dead "Fetch Models" button on compatible nodes to the live upstream catalog, de-duplicating against already-added models

## Fixes
- **Models**: persist per-model capability assertions for custom and compatible providers and honor them everywhere — unsupported media is stripped on the chat path, `/v1/models` and `/api/models` report what the user asserted, and thinking translation follows it (asserting `reasoning:false` now actually strips thinking fields, `reasoning:true` emits them)
- **Models**: partial capability edits merge instead of overwriting, so toggling vision off no longer erases a stored reasoning assertion
- **Capabilities**: keep server-injected readers (synced catalog, user-asserted capabilities) in process-wide state — Next.js compiles startup and each API route into separate bundles with their own module instances, so a boot-time install was invisible to every request handler and the models.dev catalog contributed nothing to upstream requests since 0532f00d
- **Dashboard**: thinking-level picker and model-row suffix reflect user-asserted reasoning on compatible nodes
- **Providers**: `Default Model` is optional when adding an API key to a compatible node — the node's own model list (and the picker in the test modals) already determine what gets probed, and the built-in fallback still covers connection checks
- **Providers**: restore the `useCopyToClipboard` import dropped from the provider detail page, which crashed the route with `ReferenceError` for every provider
- **DB**: restore `getModelAliases` / `setModelAlias` / `deleteModelAlias` re-exports dropped from the `localDb` shim by 86112cee, which broke `GET /api/models` and `GET /v1/models` at import time
- **Providers**: remove dead `PassthroughModelsSection` (never passed props, superseded by the shared model rows)
- **Media Providers**: creating a custom embedding node reports that a key still has to be added, instead of claiming a key was saved; the edit dialog keeps its API Key + Check affordance since a stored key already exists there
- **Build**: self-host Inter instead of fetching it through `next/font/google` at build time — a Docker / mirrored builder with no route to `fonts.googleapis.com` failed the entire image build on `Failed to fetch 'Inter' from Google Fonts`. The seven `@font-face` rules and their `unicode-range`s copy what `next/font` emitted (a `latin`-only file would have dropped Vietnamese diacritics) and the latin subset is preloaded as before, so rendered metrics are unchanged
2026-09-10 11:11:24 +07:00
mrnim94
35b950be81 fix(kiro): route requests through current runtime surfaces and fix 400 REQUEST_BODY_INVALID (#3776) 2026-09-09 10:56:54 +07:00
Sutarto Jordan Chrisfivo
7fee56bacd fix(providers): clear stale locks after validation (#3830)
Clear stale connection health state (modelLock_*, backoffLevel,
rateLimitedUntil, errorCode) whenever a connection is explicitly
marked active after successful validation or OAuth re-login.

Closes #3810
2026-09-09 10:26:03 +07:00
B1nh M1nh
4ad1e7a4ba fix(usage): parse Fable weekly limit from limits[] instead of fabricating a row (#3847) 2026-09-09 10:19:50 +07:00
Christian Gennari
e3bf94ee25 feat(antigravity): add weekly quota tracking and free-tier handling (#3892) 2026-09-09 09:57:13 +07:00
decolua
628ff1eab5 fix(auth): set 24h maxAge for dashboard session cookie
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-09 09:45:23 +07:00
decolua
e7b5f09d50 fix(gemini): normalize contents and handle intermediate tool responses in Antigravity
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-09 09:45:18 +07:00
302795a613 fix(usage): restore xai quota label case; update stale kiro tests
3-way unit-suite comparison (branch-pre-merge vs origin/master vs
HEAD) with git worktrees:

* ProviderLimits/utils.js: parseQuotaData 'xai' case (ab9a3c1d
  weekly/api_usage label mapping) was dropped by an EARLIER merge
  (already failing at the pre-merge branch tip) — its companion test
  xai-usage.test.js has been red since. Restored verbatim from
  ab9a3c1d; the xai usage handler + dispatch entry survived, only the
  UI label mapping was lost.
* openai-to-kiro.test.js: 19 failures were stale upstream tests —
  1fc2a81d intentionally removed the redundant top-level
  systemPrompt field (Kiro rejects it) without updating them. The
  thinking/agentic directives now ride the frozen session-start msg0
  via contentPrefix. systemPromptOf reads msg0 (or the current
  message when there is no replayed session); the cross-turn
  stability test asserts the actual cacheability contract.

Suite now: 83 failing vs 103 on origin/master baseline; zero files
fail in HEAD that did not fail pre-merge. next build passes.
2026-09-08 10:53:21 +07:00
02f097880d fix(merge): restore ANTHROPIC_COMPATIBLE_PREFIX import + drop dead appendRequestLog call
Full-source eslint no-undef sweep over src/ + open-sse/ (config
listing node/web globals) found two remaining undeclared-variable
regressions of the same merge-loss family:

* api/providers/test-batch: branch commit f0adfb20 added a
  providerId.startsWith(ANTHROPIC_COMPATIBLE_PREFIX) check but never
  imported it; the master merge kept the buggy line. Every batch
  test / provider-group filter touching a non-openai-compatible
  provider threw ReferenceError (|| does not short-circuit).
* utils/stream.js finalizeStream: no-usage fallback called
  appendRequestLog(), a stub upstream had marked no-op and whose
  export chain was dropped by the merge. The call site only fires
  when a stream ends without valid usage and now threw
  ReferenceError inside the terminal callback. Removed the dead
  call (behaviour identical: the stub wrote nothing).

All other no-undef reports are browser globals absent from the
scan config, not source bugs. next build --webpack passes; smoke
server boots and auth-rejects unauthenticated /v1 + /api traffic
as expected; unit suite 1746 pass / 103 fail (was 1738/111).
2026-09-08 10:36:24 +07:00
38ff11ee16 fix(merge): restore chat.js imports, trust-vision floor, gitignore entries
Audit of every branch-owned line the -X theirs merge dropped from
the 32 pre-merge commits found three more real regressions:

* src/sse/handlers/chat.js: merge kept the capsOverride feature
  (bb8d67ba) but reverted the import block, so getCustomModels and
  capabilitiesFromServiceKind were undefined. The runtime error was
  swallowed by the feature's own fail-open try/catch — custom
  models silently lost their vision override. Restored both imports.
* open-sse/providers/capabilities.js: TRUST_UPSTREAM_VISION (the
  floor that keeps vision on for unknown models on upstream-validating
  gateways like openrouter) was left as dead code by the merge —
  upstream rewrote step 4 as refine() and dropped the check.
  Re-applied it on top of the new refine() so catalog/limits
  refinement still applies.
* tests/unit/chat-connection-pin.test.js: mock auth module lacked
  isModelAllowedForKey added by f0adfb20.
* .gitignore: re-add .pi-subagents/.
2026-09-08 09:38:03 +07:00
e3ba5d2207 fix(usage): restore commandcode quota + timeout guards lost in master merge
The -X theirs merge of origin/master (v0.5.69) silently reverted five
branch-only hunks because upstream had no conflict-region counterpart
and simply won the three-way pick:

* services/usage.js: re-register the commandcode USAGE handler +
  import. Without it the dashboard Quota Tracker fell through to
  'Usage API not implemented for commandcode'. The handler module
  (services/usage/commandcode.js) and registry usage block survived;
  only the dispatch entry was dropped.
* ProviderLimits/utils.js: restore parseQuotaData 'commandcode' case
  (currency-credit rows need unit "$" + remainingPercentage
  forwarding, else $0.05 balances render as 0%).
* profile + providers/[id] pages: restore Math.max(1000, ...) connect
  timeout floors so a stray '60' is never interpreted as 60ms.
* .gitignore: re-add .commandcode/ CLI local state.

Verified: tests/unit/commandcode-usage.test.js (6) and
usage-dispatch.test.js (2, asserts every provider routes to a real
handler) pass standalone; full unit run 1742 pass / 107 fail vs
1738/111 before this fix (remaining failures pre-existing, unrelated).
2026-09-08 09:28:59 +07:00
392 changed files with 27853 additions and 2863 deletions

Binary file not shown.

After

Width:  |  Height:  |  Size: 38 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 48 KiB

View File

@@ -5,22 +5,164 @@ on:
tags:
- "v*"
workflow_dispatch:
inputs:
release_tag:
description: "Existing vX.Y.Z tag to publish"
required: true
type: string
promote_latest:
description: "Promote this republish to latest"
required: false
default: false
type: boolean
# Keep every release in one FIFO queue. A per-tag group would still allow an
# older release to finish after a newer release and move latest backwards.
concurrency:
group: docker-publish-${{ github.repository }}
cancel-in-progress: false
queue: max
env:
GHCR_IMAGE: ghcr.io/${{ github.repository }}
DOCKERHUB_IMAGE: decolua/9router
jobs:
build-and-push:
prepare:
name: Validate release
runs-on: ubuntu-latest
timeout-minutes: 10
permissions:
contents: read
outputs:
tag: ${{ steps.release.outputs.tag }}
version: ${{ steps.release.outputs.version }}
commit: ${{ steps.release.outputs.commit }}
publish_dockerhub: ${{ steps.release.outputs.publish_dockerhub }}
promote_latest: ${{ steps.release.outputs.promote_latest }}
ghcr_image: ${{ steps.release.outputs.ghcr_image }}
steps:
- name: Check out release tag
uses: actions/checkout@v4
with:
ref: ${{ inputs.release_tag || github.ref_name }}
fetch-depth: 1
- name: Validate tag and package versions
id: release
env:
RELEASE_TAG: ${{ inputs.release_tag || github.ref_name }}
REPOSITORY: ${{ github.repository }}
EVENT_NAME: ${{ github.event_name }}
PROMOTE_LATEST_INPUT: ${{ inputs.promote_latest && 'true' || 'false' }}
run: |
node <<'NODE'
const fs = require("fs");
const { execFileSync } = require("child_process");
const tag = process.env.RELEASE_TAG || "";
const match = /^v((?:0|[1-9]\d*)\.(?:0|[1-9]\d*)\.(?:0|[1-9]\d*)(?:-[0-9A-Za-z-]+(?:\.[0-9A-Za-z-]+)*)?)$/.exec(tag);
if (tag.includes("+")) {
console.error(`Build metadata is not supported in Docker release tags: ${tag}`);
process.exit(1);
}
if (!match) {
console.error(`Expected a Docker-safe semver tag like v0.5.81 or v0.5.81-rc.1, received: ${tag || "<empty>"}`);
process.exit(1);
}
const version = match[1];
if (version.length > 128 || !/^[A-Za-z0-9_][A-Za-z0-9_.-]{0,127}$/.test(version)) {
console.error(`Version is not a valid Docker tag: ${version}`);
process.exit(1);
}
const prerelease = version.includes("-")
? version.slice(version.indexOf("-") + 1).split(".")
: [];
for (const identifier of prerelease) {
if (/^\d+$/.test(identifier) && identifier.length > 1 && identifier.startsWith("0")) {
console.error(`Numeric prerelease identifiers cannot contain leading zeroes: ${identifier}`);
process.exit(1);
}
}
const rootVersion = require("./package.json").version;
const cliVersion = require("./cli/package.json").version;
if (rootVersion !== version) {
console.error(`package.json version ${rootVersion} does not match tag ${tag}`);
process.exit(1);
}
if (cliVersion !== version) {
console.error(`cli/package.json version ${cliVersion} does not match tag ${tag}`);
process.exit(1);
}
const commit = execFileSync("git", ["rev-parse", "HEAD"], { encoding: "utf8" }).trim();
const publishDockerHub = process.env.REPOSITORY === "decolua/9router";
const ghcrImage = `ghcr.io/${process.env.REPOSITORY.toLowerCase()}`;
const isPrerelease = version.includes("-");
const promoteLatest = (process.env.EVENT_NAME === "push" && !isPrerelease)
|| process.env.PROMOTE_LATEST_INPUT === "true";
const output = process.env.GITHUB_OUTPUT;
fs.appendFileSync(output, `tag=${tag}\n`);
fs.appendFileSync(output, `version=${version}\n`);
fs.appendFileSync(output, `commit=${commit}\n`);
fs.appendFileSync(output, `publish_dockerhub=${publishDockerHub}\n`);
fs.appendFileSync(output, `promote_latest=${promoteLatest}\n`);
fs.appendFileSync(output, `ghcr_image=${ghcrImage}\n`);
console.log(`Validated ${tag} at ${commit}`);
console.log(`latest promotion: ${promoteLatest ? "enabled" : "disabled"}`);
NODE
build:
name: Build ${{ matrix.platform }}
needs: prepare
runs-on: ${{ matrix.runner }}
timeout-minutes: 60
env:
GHCR_IMAGE: ${{ needs.prepare.outputs.ghcr_image }}
strategy:
fail-fast: false
matrix:
include:
- platform: linux/amd64
suffix: amd64
runner: ubuntu-24.04
- platform: linux/arm64
suffix: arm64
runner: ubuntu-24.04-arm
permissions:
contents: read
packages: write
steps:
- uses: actions/checkout@v4
- name: Check out release source at validated commit
uses: actions/checkout@v4
with:
ref: ${{ needs.prepare.outputs.commit }}
path: source
fetch-depth: 1
- uses: docker/setup-buildx-action@v3
- name: Check out publishing Dockerfile
uses: actions/checkout@v4
with:
ref: ${{ github.workflow_sha }}
path: workflow
sparse-checkout: |
Dockerfile
fetch-depth: 1
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@v3
- name: Log in to GHCR
uses: docker/login-action@v3
@@ -29,32 +171,267 @@ jobs:
username: ${{ github.actor }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: Build and push platform image by digest
id: build
uses: docker/build-push-action@v6
with:
context: source
file: workflow/Dockerfile
platforms: ${{ matrix.platform }}
outputs: type=image,name=${{ env.GHCR_IMAGE }},push-by-digest=true,name-canonical=true,push=true
build-args: |
APP_VERSION=${{ needs.prepare.outputs.version }}
ALPINE_MIRROR=${{ vars.ALPINE_MIRROR || 'dl-cdn.alpinelinux.org' }}
NPM_REGISTRY=${{ vars.NPM_REGISTRY || 'https://registry.npmjs.org/' }}
labels: |
org.opencontainers.image.source=https://github.com/${{ github.repository }}
org.opencontainers.image.revision=${{ needs.prepare.outputs.commit }}
org.opencontainers.image.version=${{ needs.prepare.outputs.version }}
cache-from: type=gha,scope=9router-${{ matrix.suffix }}
cache-to: type=gha,mode=max,scope=9router-${{ matrix.suffix }}
provenance: false
sbom: false
- name: Smoke-test platform image before publishing digest artifact
env:
GHCR_IMAGE: ${{ env.GHCR_IMAGE }}
IMAGE_DIGEST: ${{ steps.build.outputs.digest }}
PLATFORM: ${{ matrix.platform }}
run: |
set -Eeuo pipefail
[[ "$IMAGE_DIGEST" =~ ^sha256:[0-9a-f]{64}$ ]]
container="9router-platform-smoke-${GITHUB_RUN_ID}-${{ matrix.suffix }}"
trap 'docker rm -f "$container" >/dev/null 2>&1 || true' EXIT
docker run --detach \
--name "$container" \
--platform "$PLATFORM" \
--publish 20128:20128 \
"${GHCR_IMAGE}@${IMAGE_DIGEST}"
for attempt in {1..45}; do
if curl --fail --silent --show-error http://127.0.0.1:20128/api/health; then
echo "${PLATFORM} health check passed"
exit 0
fi
if (( attempt % 5 == 0 )); then
echo "Waiting for ${PLATFORM} health check (${attempt}/45)" >&2
fi
sleep 2
done
echo "${PLATFORM} health check failed; container logs follow:" >&2
docker logs "$container" || true
exit 1
- name: Save image digest
env:
IMAGE_DIGEST: ${{ steps.build.outputs.digest }}
run: |
set -euo pipefail
test -n "$IMAGE_DIGEST"
mkdir -p "$RUNNER_TEMP/digests"
printf '%s\n' "$IMAGE_DIGEST" > "$RUNNER_TEMP/digests/${{ matrix.suffix }}.txt"
- name: Upload image digest
uses: actions/upload-artifact@v4
with:
name: digests-${{ matrix.suffix }}
path: ${{ runner.temp }}/digests/${{ matrix.suffix }}.txt
if-no-files-found: error
publish:
name: Publish and verify manifest
needs:
- prepare
- build
runs-on: ubuntu-latest
timeout-minutes: 30
env:
GHCR_IMAGE: ${{ needs.prepare.outputs.ghcr_image }}
permissions:
contents: read
packages: write
steps:
- name: Download platform digests
uses: actions/download-artifact@v4
with:
pattern: digests-*
path: ${{ runner.temp }}/digests
merge-multiple: true
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@v3
- name: Log in to GHCR
uses: docker/login-action@v3
with:
registry: ghcr.io
username: ${{ github.actor }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: Create and verify version manifest
env:
GHCR_IMAGE: ${{ env.GHCR_IMAGE }}
VERSION: ${{ needs.prepare.outputs.version }}
run: |
set -euo pipefail
shopt -s nullglob
digest_files=("$RUNNER_TEMP"/digests/*.txt)
if [[ "${#digest_files[@]}" -ne 2 ]]; then
echo "Expected two platform digests, found ${#digest_files[@]}" >&2
exit 1
fi
sources=()
for digest_file in "${digest_files[@]}"; do
digest="$(tr -d '\n' < "$digest_file")"
if [[ ! "$digest" =~ ^sha256:[0-9a-f]{64}$ ]]; then
echo "Invalid image digest in $digest_file: $digest" >&2
exit 1
fi
sources+=("${GHCR_IMAGE}@${digest}")
done
docker buildx imagetools create \
--tag "${GHCR_IMAGE}:${VERSION}" \
"${sources[@]}"
docker buildx imagetools inspect "${GHCR_IMAGE}:${VERSION}" | tee "$RUNNER_TEMP/version-manifest.txt"
docker buildx imagetools inspect --raw "${GHCR_IMAGE}:${VERSION}" > "$RUNNER_TEMP/version-manifest.json"
expected=$'linux/amd64\nlinux/arm64'
actual="$(jq -r '[.manifests[] | select(.platform != null and .platform.os != null and .platform.architecture != null) | "\(.platform.os)/\(.platform.architecture)"] | sort | .[]' "$RUNNER_TEMP/version-manifest.json")"
if [[ "$actual" != "$expected" ]]; then
echo "Version manifest platforms do not match exactly:" >&2
printf '%s\n' "$actual" >&2
exit 1
fi
- name: Smoke-test resolved version manifest
env:
GHCR_IMAGE: ${{ env.GHCR_IMAGE }}
VERSION: ${{ needs.prepare.outputs.version }}
run: |
set -Eeuo pipefail
container="9router-manifest-smoke-${GITHUB_RUN_ID}"
trap 'docker rm -f "$container" >/dev/null 2>&1 || true' EXIT
docker run --detach \
--name "$container" \
--platform linux/amd64 \
--publish 20128:20128 \
"${GHCR_IMAGE}:${VERSION}"
for attempt in {1..30}; do
if curl --fail --silent --show-error http://127.0.0.1:20128/api/health; then
echo "Resolved version manifest health check passed"
exit 0
fi
if (( attempt % 5 == 0 )); then
echo "Waiting for resolved manifest health check (${attempt}/30)" >&2
fi
sleep 2
done
echo "Resolved version manifest health check failed; container logs follow:" >&2
docker logs "$container" || true
exit 1
- name: Log in to Docker Hub
if: needs.prepare.outputs.publish_dockerhub == 'true'
uses: docker/login-action@v3
with:
username: ${{ secrets.DOCKERHUB_USERNAME }}
password: ${{ secrets.DOCKERHUB_TOKEN }}
- name: Extract metadata
id: meta
uses: docker/metadata-action@v5
with:
images: |
${{ env.GHCR_IMAGE }}
${{ env.DOCKERHUB_IMAGE }}
tags: |
type=semver,pattern={{version}}
type=raw,value=latest,enable={{is_default_branch}}
- name: Publish version image to Docker Hub
if: needs.prepare.outputs.publish_dockerhub == 'true'
env:
DOCKERHUB_IMAGE: ${{ env.DOCKERHUB_IMAGE }}
GHCR_IMAGE: ${{ env.GHCR_IMAGE }}
VERSION: ${{ needs.prepare.outputs.version }}
run: |
set -euo pipefail
docker buildx imagetools create \
--tag "${DOCKERHUB_IMAGE}:${VERSION}" \
"${GHCR_IMAGE}:${VERSION}"
- name: Build and push
uses: docker/build-push-action@v6
with:
context: .
push: true
tags: ${{ steps.meta.outputs.tags }}
labels: ${{ steps.meta.outputs.labels }}
cache-from: type=registry,ref=${{ env.GHCR_IMAGE }}:buildcache
cache-to: type=registry,ref=${{ env.GHCR_IMAGE }}:buildcache,mode=max
platforms: linux/amd64,linux/arm64
provenance: false
sbom: false
docker buildx imagetools inspect "${DOCKERHUB_IMAGE}:${VERSION}" | tee "$RUNNER_TEMP/dockerhub-version-manifest.txt"
docker buildx imagetools inspect --raw "${DOCKERHUB_IMAGE}:${VERSION}" > "$RUNNER_TEMP/dockerhub-version-manifest.json"
expected=$'linux/amd64\nlinux/arm64'
actual="$(jq -r '[.manifests[] | select(.platform != null and .platform.os != null and .platform.architecture != null) | "\(.platform.os)/\(.platform.architecture)"] | sort | .[]' "$RUNNER_TEMP/dockerhub-version-manifest.json")"
if [[ "$actual" != "$expected" ]]; then
echo "Docker Hub version manifest platforms do not match exactly:" >&2
printf '%s\n' "$actual" >&2
exit 1
fi
- name: Record latest promotion policy
env:
PROMOTE_LATEST: ${{ needs.prepare.outputs.promote_latest }}
VERSION: ${{ needs.prepare.outputs.version }}
run: |
if [[ "$PROMOTE_LATEST" == "true" ]]; then
echo "### Latest promotion" >> "$GITHUB_STEP_SUMMARY"
echo "- Policy: promote \`latest\` after the verified ${VERSION} manifest." >> "$GITHUB_STEP_SUMMARY"
else
echo "### Latest promotion" >> "$GITHUB_STEP_SUMMARY"
echo "- Policy: leave \`latest\` unchanged; this is a numbered-tag-only manual republish." >> "$GITHUB_STEP_SUMMARY"
fi
- name: Promote verified version to latest
if: needs.prepare.outputs.promote_latest == 'true'
env:
DOCKERHUB_IMAGE: ${{ env.DOCKERHUB_IMAGE }}
GHCR_IMAGE: ${{ env.GHCR_IMAGE }}
PUBLISH_DOCKERHUB: ${{ needs.prepare.outputs.publish_dockerhub }}
VERSION: ${{ needs.prepare.outputs.version }}
run: |
set -euo pipefail
docker buildx imagetools create \
--tag "${GHCR_IMAGE}:latest" \
"${GHCR_IMAGE}:${VERSION}"
if [[ "$PUBLISH_DOCKERHUB" == "true" ]]; then
docker buildx imagetools create \
--tag "${DOCKERHUB_IMAGE}:latest" \
"${GHCR_IMAGE}:${VERSION}"
fi
docker buildx imagetools inspect "${GHCR_IMAGE}:latest" | tee "$RUNNER_TEMP/ghcr-latest-manifest.txt"
docker buildx imagetools inspect --raw "${GHCR_IMAGE}:latest" > "$RUNNER_TEMP/ghcr-latest-manifest.json"
expected=$'linux/amd64\nlinux/arm64'
actual="$(jq -r '[.manifests[] | select(.platform != null and .platform.os != null and .platform.architecture != null) | "\(.platform.os)/\(.platform.architecture)"] | sort | .[]' "$RUNNER_TEMP/ghcr-latest-manifest.json")"
if [[ "$actual" != "$expected" ]]; then
echo "GHCR latest manifest platforms do not match exactly:" >&2
printf '%s\n' "$actual" >&2
exit 1
fi
if [[ "$PUBLISH_DOCKERHUB" == "true" ]]; then
docker buildx imagetools inspect "${DOCKERHUB_IMAGE}:latest" | tee "$RUNNER_TEMP/dockerhub-latest-manifest.txt"
docker buildx imagetools inspect --raw "${DOCKERHUB_IMAGE}:latest" > "$RUNNER_TEMP/dockerhub-latest-manifest.json"
actual="$(jq -r '[.manifests[] | select(.platform != null and .platform.os != null and .platform.architecture != null) | "\(.platform.os)/\(.platform.architecture)"] | sort | .[]' "$RUNNER_TEMP/dockerhub-latest-manifest.json")"
if [[ "$actual" != "$expected" ]]; then
echo "Docker Hub latest manifest platforms do not match exactly:" >&2
printf '%s\n' "$actual" >&2
exit 1
fi
fi
{
echo "### Published Docker images"
echo "- GHCR: \`${GHCR_IMAGE}:${VERSION}\`"
echo "- GHCR latest: \`${GHCR_IMAGE}:latest\`"
if [[ "$PUBLISH_DOCKERHUB" == "true" ]]; then
echo "- Docker Hub: \`${DOCKERHUB_IMAGE}:${VERSION}\`"
echo "- Docker Hub latest: \`${DOCKERHUB_IMAGE}:latest\`"
fi
} >> "$GITHUB_STEP_SUMMARY"

176
.github/workflows/tray-binaries.yml vendored Normal file
View File

@@ -0,0 +1,176 @@
name: Build macOS tray binary (arm64)
# systray2 ships only an x86_64 tray_darwin_release, so Apple Silicon users need
# Rosetta 2 for the menubar icon. This builds the native arm64 overlay that
# cli/hooks/trayRuntime.js downloads from the `tray-binaries` release.
#
# Manual-only: the artifact's sha256 is pinned in cli/hooks/trayRuntime.js and
# verified on every download, so a new build is only publishable together with a
# matching pin. Running this with publish=true against a mismatched pin fails
# rather than silently bricking every Apple Silicon client.
on:
workflow_dispatch:
inputs:
publish:
description: "Upload to the tray-binaries release (requires sha to match ARM64_TRAY_SHA256)"
required: false
default: false
type: boolean
concurrency:
group: tray-binaries-${{ github.repository }}
cancel-in-progress: false
permissions:
contents: read
env:
# Pinned because -trimpath only makes the build reproducible for a given Go
# version and macOS SDK. Bumping this changes the sha256.
GO_VERSION: "1.27.1"
jobs:
build:
name: Build darwin/arm64
runs-on: macos-15
timeout-minutes: 20
permissions:
contents: write
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 22
- uses: actions/setup-go@v5
with:
go-version: ${{ env.GO_VERSION }}
# The Go module lives in a temp clone of the upstream repo, so there is
# no go.sum at the workspace root for setup-go's cache to key on.
cache: false
- name: Record SDK provenance
run: |
{
echo "runner macOS: $(sw_vers -productVersion)"
echo "Xcode: $(xcodebuild -version | head -1)"
echo "clang: $(clang --version | head -1)"
echo "Go: $(go version)"
} | tee sdk-provenance.txt
- name: Build
run: node cli/scripts/buildTrayArm64.js
- name: Compare against pinned checksum
id: sha
run: |
BUILT=$(shasum -a 256 cli/.tray-build/tray_darwin_arm64 | cut -d' ' -f1)
# Whitespace-tolerant, and a missing constant must fail loudly: a null
# match would otherwise surface as an opaque TypeError from [1].
PINNED=$(node -e '
const m = require("fs").readFileSync("cli/hooks/trayRuntime.js", "utf8")
.match(/ARM64_TRAY_SHA256\s*=\s*"([0-9a-f]{64})"/);
if (!m) { console.error("::error::ARM64_TRAY_SHA256 not found in cli/hooks/trayRuntime.js"); process.exit(1); }
process.stdout.write(m[1]);
')
{
echo "built=$BUILT"
echo "pinned=$PINNED"
if [ "$BUILT" = "$PINNED" ]; then echo "match=true"; else echo "match=false"; fi
} >> "$GITHUB_OUTPUT"
- name: Write job summary
run: |
{
echo "### tray_darwin_arm64"
echo ""
echo "| | |"
echo "|---|---|"
echo "| built sha256 | \`${{ steps.sha.outputs.built }}\` |"
echo "| pinned sha256 | \`${{ steps.sha.outputs.pinned }}\` |"
echo "| match | ${{ steps.sha.outputs.match }} |"
echo ""
echo '```'
cat sdk-provenance.txt
echo '```'
echo ""
if [ "${{ steps.sha.outputs.match }}" = "true" ]; then
echo "Pin already matches — safe to re-run with \`publish=true\`."
else
echo "⚠️ Pin does **not** match. To publish this build, set \`ARM64_TRAY_SHA256\`"
echo "in \`cli/hooks/trayRuntime.js\` to the built sha256 above and land that"
echo "change first. Publishing without it makes every Apple Silicon client fail"
echo "checksum verification and fall back to the Rosetta binary."
fi
} >> "$GITHUB_STEP_SUMMARY"
# Uploaded before the mismatch gate below, so a publish run that fails on a
# checksum mismatch still leaves the bytes downloadable — that is exactly
# the run where a maintainer needs them to verify the new sha256.
- uses: actions/upload-artifact@v4
with:
name: tray_darwin_arm64
path: |
cli/.tray-build/tray_darwin_arm64
sdk-provenance.txt
- name: Refuse to publish on checksum mismatch
if: ${{ inputs.publish && steps.sha.outputs.match != 'true' }}
run: |
echo "::error::publish requested but built sha256 != ARM64_TRAY_SHA256"
echo " built: ${{ steps.sha.outputs.built }}"
echo " pinned: ${{ steps.sha.outputs.pinned }}"
echo "Update cli/hooks/trayRuntime.js and land it before publishing."
exit 1
- name: Publish to tray-binaries release
if: ${{ inputs.publish && steps.sha.outputs.match == 'true' }}
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
run: |
# Clients fetch from the hardcoded ARM64_TRAY_URL, so publishing from a
# different repo would gate an asset stream nobody downloads and silently
# decouple the sha pin from the bytes Apple Silicon users execute.
URL_REPO=$(node -e '
const m = require("fs").readFileSync("cli/hooks/trayRuntime.js", "utf8")
.match(/"https:\/\/github\.com\/([^\/"]+\/[^\/"]+)\/releases\/download\/tray-binaries\/tray_darwin_arm64"/);
if (!m) { console.error("::error::ARM64_TRAY_URL not found in cli/hooks/trayRuntime.js"); process.exit(1); }
process.stdout.write(m[1]);
')
if [ "${{ github.repository }}" != "$URL_REPO" ]; then
echo "::error::publishing to ${{ github.repository }}, but ARM64_TRAY_URL points clients at $URL_REPO"
echo "Repoint ARM64_TRAY_URL in cli/hooks/trayRuntime.js at this repo, or run the publish from $URL_REPO."
exit 1
fi
# gh release upload does not create the release, so bootstrap it on the
# first publish run rather than failing with "release not found".
if ! gh release view tray-binaries >/dev/null 2>&1; then
echo "Release 'tray-binaries' does not exist yet — creating it"
gh release create tray-binaries --latest=false \
--title "Native macOS tray binaries" \
--notes "Built by .github/workflows/tray-binaries.yml. Provenance and the pinned sha256 live in that workflow and in cli/hooks/trayRuntime.js (ARM64_TRAY_SHA256)."
fi
gh release upload tray-binaries cli/.tray-build/tray_darwin_arm64 --clobber
echo "Uploaded. Verifying public download URL..."
URL="https://github.com/${{ github.repository }}/releases/download/tray-binaries/tray_darwin_arm64"
GOT=""
for attempt in 1 2 3; do
if curl -fsSL --max-time 60 -o /tmp/verify "$URL"; then
GOT=$(shasum -a 256 /tmp/verify | cut -d' ' -f1)
if [ "$GOT" = "${{ steps.sha.outputs.built }}" ]; then break; fi
fi
# A just-uploaded asset can 404 or serve stale bytes until the CDN catches up.
echo "attempt $attempt: got '${GOT:-<download failed>}' — retrying in 15s"
sleep 15
done
if [ "$GOT" != "${{ steps.sha.outputs.built }}" ]; then
echo "::error::downloaded asset sha256 '${GOT:-<none>}' != built ${{ steps.sha.outputs.built }} after 3 attempts"
exit 1
fi
echo "✅ $URL serves the expected bytes"

22
.gitignore vendored
View File

@@ -88,4 +88,24 @@ graphify-out/*
.next-analyze/*
# Kiro local workspace state
.kiro/
.kiro/
9router-*
# Local sensitive / temp files
.engine-token.txt
.tmp-prov.json
.tmp-providers.json
.oauth-session.json
.dev-server.log
.dev-server-err.log
.start-dev.ps1
start-dev-silent.cjs
debug.log
# CommandCode CLI local state (auth/taste/projects)
.commandcode/
# Pi subagent run artifacts
.pi-subagents/
# Vitest local artifacts
.vitest/

View File

@@ -1,3 +1,133 @@
# v0.5.91 (2026-09-26)
## Features
- **Providers**: add Token Harbor provider and four OpenAI-compatible aggregator providers (dahl, atria, agnes, bai)
- **Claude**: forward `x-claude-code-session-id` on OAuth requests; merge client `anthropic-beta` flags and forward rate-limit headers; return thinking text to OpenAI-format clients
- **Codex**: add GPT-6 Sol and Luna support
- **CLI Tools**: support multiple model profiles for Codex CLI
- **Hermes**: multi-role model config (delegation + auxiliary slots)
- **OpenCode Go**: complete the Go catalog (40 models) with auto-fetch + family endpoint regex
- **Usage**: show and redeem free limit resets for cc accounts
- **Cline**: expose the `cline-free/*` tier and price it at zero
- **Combos**: display vision adapter models in an ordered table view
## Fixes
- **Claude**: decloak tool names when `toolNameMap` misses (#4342); update spoofed cli version to 2.1.280 to support Opus 5.5
- **Providers API**: make POST `/api/providers` O(1) and refuse silent key overwrite (#4350)
- **Capabilities**: stop caching the catalog source per module copy (#4351)
- **OAuth**: stop Zed paste-token crash and add IDE auto-import (#4359)
- **Dashboard**: resolve combo limits with the server's capabilities (#4360); lazy-load charts and `marked`, preload in background on idle
- **Responses**: carry the streamed output items in `response.completed` (#4307)
- **STT**: dispatch live-API-only Gemini models over the Live WebSocket transport (#4006)
- **Gemini**: guard terminal model turns and unresponded functionCalls in `normalizeGeminiContents`
- **Command Code**: replay raw byte chunks to preserve all NDJSON lines
- **Translator**: stop emitting empty `<think>` markers into OpenAI content
- **CLI Tools**: refresh Codex settings after apply (#4347); keep existing `ANTHROPIC_AUTH_TOKEN` when applying Claude settings
- **Tray**: native arm64 macOS menubar binary, no Rosetta required
- **CLI**: filter model selector by active connections and noAuth providers
- **Usage**: key live byApiKey stats by full api key to prevent team-key collision and preserve API key usage attribution
- **Tailscale**: cap enable-flow health wait at 20s
# v0.5.86 (2026-09-23)
## Features
- **Xiaomi MiMo**: server-assisted desktop login for headless/Docker deployments, five account clusters (cn/sgp/ams/ru/in), and v2.6 pro/flash/pro-ultraspeed models with dual-route (account service vs. cloud API)
- **Claude**: add Claude Opus 5.5 support
- **i18n**: translate React text rewrites via characterData mutation observer
## Fixes
- **Proxy Pools**: keep request headers intact through Vercel/Cloudflare/Deno relays (spreading a `Headers` instance yielded `{}`, dropping auth and content-type)
- **Xiaomi MiMo login**: keep the session in the httpOnly cookie only, require dashboard auth on the proxy branch, and stop forwarding authorization headers upstream
# v0.5.85 (2026-09-22)
## Features
- **System One**: add `/v1/systemone` decision endpoint for Jev models (OpenCode Zen and OpenRouter lanes), wire into sidebar and Media Providers page with interactive probe testing
- **CLI Tools**: add dynamic configuration, settings APIs, and official logos for Pi, OMP, Crush, ForgeCode, Smelt, and CodeWhale
- **Analytics & Usage**: add Requests mode, provider/model breakdown charts, All Time period filter, and refined overview cards
- **Combos**: add Cursor/Claude Default presets; support bulk select/delete and bulk strategy changes (Fallback / Round Robin / Fusion)
- **Model Capabilities**: expose model capability metadata on `/v1/models` and aggregate capabilities across combo targets
- **OpenCode Zen & MiMo**: add OpenCode Zen (`opencode-zen`) provider with free-tier fingerprint; switch default vision fallback to MiMo V2.6 Flash Free
- **Qoder CN**: add `qoder-cn` provider for qoder.com.cn with OAuth flow, COSY protocol, and CN gateway routing
## Fixes
- **Translator**: map Claude `refusal` stop_reason to `content_filter` and surface explanation; strip replayed reasoning fields for Groq, Mistral, and Cerebras (#4220)
- **Antigravity**: drop requestType `agent` to avoid false 429 `RESOURCE_EXHAUSTED`; separate weekly and short-window (5-hour) quotas and deduplicate dashboard rows
- **Responses API**: report usage on `response.completed` so clients can auto-compact (#3432)
- **Hugging Face**: migrate to Inference Providers router (`router.huggingface.co`), expand image models catalog, and add STT route
- **Qoder**: prevent signed request replay (`403/103 Duplicate request`), handle code 110 billing blocks, and preserve upstream SSE error status
- **Performance**: bound usage `lastUsed` scan to a 2-day window; map large budget tokens to `max` reasoning tier
- **Docker**: publish verified multi-platform images (linux/amd64 and linux/arm64) with configurable apk build mirrors
# v0.5.81 (2026-09-18)
## Features
- **Xiaomi MiMo**: merge MiMo Desktop support into `xiaomi-mimo` with dual auth (API key + Desktop/OAuth session), Preview models support, and encrypted-callback OAuth flow
- **Claude Code**: add 1M-context toggle (`[1m]` marker) and drive `CLAUDE_CODE_AUTO_COMPACT_WINDOW` directly from the dashboard
- **Models**: add DeepSeek-V4.1-Flash to DeepSeek provider, CodeBuddy-Intl, and Ollama (`deepseek-v4.1-flash:cloud`); enable `low`..`max` reasoning effort levels and vision capability for DeepSeek-V4.*
- **i18n**: integrate Persian (fa) translation
## Fixes
- **Cursor**: stop AgentService empty turns (`OUT 0`) and silent hangs — fold system prompts instead of `custom_system_prompt`, send `ModelDetails`, read Composer/Grok `thinking_delta`, ack request-context without echoing MCP tools, and reject IDE execs so the model can continue
- **RTK**: for Cursor, compress source-format `tool_result` / `role:tool` **before** translation — its translator rewrites those shapes, so post-translate compression missed them. Other providers keep the post-translate pass unchanged
- **OpenCode / OpenCode Go**: resolve 403 `FreeTierError` and 429 rate limits with canonical session format, valid User-Agent, and stable upstream session reuse; force stream and declare `forceStream` for free-tier SSE aggregation; cloak decoy tools, normalize Muse Free tool choice, and strip prior reasoning items on Responses models; route Union Alpha via Messages API
- **Kiro**: preserve underscores in tool names (`mcp__server__tool`) and restore client tool names in responses; use neutral placeholder for tool-result-only turns; forward tool-result images
- **Stream**: report aborts after HTTP 200 in-band (per-format error frames) instead of closing silently
- **Command Code**: preserve images and `reasoning_effort` on `/alpha/generate`; retry transient stream errors and avoid fake stop chunks; add Quota Tracker support
- **Zed**: harden OAuth lifecycle (preserve `systemId`, renew proxy timeout), support live model resolution, and lower display priority in OAuth list
- **Antigravity**: scope cached thought signatures to model family; strip Claude Code billing headers from system prompts; sanitize Hermes system identity
- **Codex**: route bare `codex-auto-review` requests to the Codex provider (#4135)
- **Auth**: do not cool down an account for request-scoped 4xx errors
- **Usage**: improve DeepSeek credit balance display as currency credit instead of 0/total quota bar
- **Model Catalog**: scope synced catalog to gateways and declare vision capabilities for DeepSeek V4.1-Flash IDs
# v0.5.75 (2026-09-10)
## Features
- **Video**: add OpenRouter and Vertex AI (Veo) video generation on `/v1/videos/*` via a provider adapter layer; poll requests resolve their provider from `x-connection-id` or `?provider=`
- **Antigravity**: add weekly quota tracking (Gemini weekly / Claude & GPT weekly) and free-tier handling from `retrieveUserQuotaSummary` (#3892)
- **Codex**: add GPT Image 2.5, Flare and Sunburst image models with multi-image support; add the same ids to the OpenAI catalog
- **Qoder**: surface usage to all clients and stop inlining large attachments — images upload through `/api/v2/image/upload` like qodercli, oversized file blocks become stubs, context tier auto-escalates
- **OpenCode Go**: add newly published models (glm-5.3, kimi-k3, deepseek-flash, longcat-2.0, hy4-preview, hy3 on chat/completions; qwen3.8-max, qwen3.8-flash on `/messages`; grok-4.6, gpt-5.6-luna on Responses) and list `deepseek-v4.1-flash` first in the catalog
- **CLI tools**: group the model selector by provider with full-text search and manual custom model ID entry
- **CodeBuddy-CN**: replace `deepseek-v4-flash` with `deepseek-v4.1-flash`
## Fixes
- **Tools**: scope Claude tool type defaulting to gateways declaring `requireClaudeToolType` — the global default broke Anthropic-compatible endpoints that only accept the legacy typeless tool shape (#3905)
- **Claude**: cap re-anchored `cache_control` at the 4-marker budget so a spent budget no longer 400s and triggers a full combo failover; wrap bare single-object content turns before the mid-conversation-system fold
- **Cline / Airforce**: unwrap the `{"success":true,"data":…}` envelope on non-stream chat completions (#3644); add the live Cline/ClinePass model catalog and refresh Airforce free models
- **Cline**: stop `workos:`-prefixing ClinePass API keys (401 on every request, #2333) and add clinepass token refresh
- **Kiro**: never send a top-level `systemPrompt` (`400 REQUEST_BODY_INVALID`); route requests through current runtime surfaces (#3776)
- **Codex**: strip Unicode-property tool schema patterns the validator rejects (#3922); restore the `Version` header and single-source the CLI version
- **DeepSeek**: keep Anthropic-only tool types when forwarding to `/anthropic/v1/messages`
- **Qoder**: drop the Responses usage plumbing from shared translator/handler code, which changed token accounting for every provider, not just Qoder
- **Antigravity**: normalize contents and handle intermediate tool responses; protect the OAuth token-refresh path from Google anti-abuse rate limits (#3813)
- **Providers**: clear stale connection health state (`modelLock_*`, `backoffLevel`, `rateLimitedUntil`, `errorCode`) when a connection is re-validated (#3810, #3830); remove the duplicate `qwen` provider that shadowed `alims-intl`
- **Video / Vertex**: reject job ids and model ids that would escape the request URL path (SSRF)
- **Usage**: parse the Fable weekly limit from `limits[]` instead of fabricating a row (#3847)
- **Auth**: set a 24h `maxAge` on the dashboard session cookie
# v0.5.70 (2026-09-08)
## Features
- **Providers**: creating a compatible / custom-embedding node now registers the endpoint only — the API key is added afterwards from the node's page, like the built-in providers. The create dialogs drop the API Key / Model ID / Check fields, and `POST /api/provider-nodes` no longer accepts credentials at all, so a node can never be half-created
- **Providers**: compatible nodes now use the same model rows as built-in providers — capability badges, copy, per-model test, alias handling and the Add/Edit Model modal with vision + reasoning toggles, replacing the weaker read-only list
- **Providers**: compatible nodes get the built-in bulk toolbars: Test All / Disable / Active / Select All over connections, and Test All Models / Disable All / Active All over models, with per-row disable and a restore strip for disabled models
- **Providers**: wire the dead "Fetch Models" button on compatible nodes to the live upstream catalog, de-duplicating against already-added models
## Fixes
- **Usage**: nested combos (comboA lists comboB, comboC, …) stay one slot each — the inner combo always runs as fallback to produce a single answer, failed hops are not written to Details/usage, and streaming no longer inserts a 0-token placeholder. One user message against a nested fallback combo is one request row with real tokens; fusion of nested combos is N panel slots + judge, not every nested leaf
- **Models**: persist per-model capability assertions for custom and compatible providers and honor them everywhere — unsupported media is stripped on the chat path, `/v1/models` and `/api/models` report what the user asserted, and thinking translation follows it (asserting `reasoning:false` now actually strips thinking fields, `reasoning:true` emits them)
- **Models**: partial capability edits merge instead of overwriting, so toggling vision off no longer erases a stored reasoning assertion
- **Capabilities**: keep server-injected readers (synced catalog, user-asserted capabilities) in process-wide state — Next.js compiles startup and each API route into separate bundles with their own module instances, so a boot-time install was invisible to every request handler and the models.dev catalog contributed nothing to upstream requests since 0532f00d
- **Dashboard**: thinking-level picker and model-row suffix reflect user-asserted reasoning on compatible nodes
- **Providers**: `Default Model` is optional when adding an API key to a compatible node — the node's own model list (and the picker in the test modals) already determine what gets probed, and the built-in fallback still covers connection checks
- **Providers**: restore the `useCopyToClipboard` import dropped from the provider detail page, which crashed the route with `ReferenceError` for every provider
- **DB**: restore `getModelAliases` / `setModelAlias` / `deleteModelAlias` re-exports dropped from the `localDb` shim by 86112cee, which broke `GET /api/models` and `GET /v1/models` at import time
- **Providers**: remove dead `PassthroughModelsSection` (never passed props, superseded by the shared model rows)
- **Media Providers**: creating a custom embedding node reports that a key still has to be added, instead of claiming a key was saved; the edit dialog keeps its API Key + Check affordance since a stored key already exists there
- **Build**: self-host Inter instead of fetching it through `next/font/google` at build time — a Docker / mirrored builder with no route to `fonts.googleapis.com` failed the entire image build on `Failed to fetch 'Inter' from Google Fonts`. The seven `@font-face` rules and their `unicode-range`s copy what `next/font` emitted (a `latin`-only file would have dropped Vietnamese diacritics) and the latin subset is preloaded as before, so rendered metrics are unchanged
# v0.5.69 (2026-09-05)
## Features

View File

@@ -88,6 +88,7 @@ Pre-translate hooks that compress `tool_result` content in-place to cut tokens.
- `custom-server.js` wraps the Next standalone server to derive client IP from the TCP socket and strip attacker-controlled `X-Forwarded-For` — trusting forwarding headers only from a loopback reverse proxy. Preserve this when touching request/IP/rate-limit code.
- Security-sensitive env: `JWT_SECRET` (session cookie), `INITIAL_PASSWORD` (default `123456` — must override), `API_KEY_SECRET`, `MACHINE_ID_SALT`. Full env contract in `.env.example` and ARCHITECTURE.md's env matrix.
- Binary/protobuf upstreams (kiro EventStream, cursor protobuf, commandcode NDJSON) don't round-trip through OpenAI — they're handled inside their own executor, not the translator.
- **Security-first on PRs**: Security is the top priority when reviewing or creating PRs. Audit authentication, credential/token storage & leaks, header manipulation (`X-Forwarded-For`), and SSRF risks before functional logic. Always include explicit security warnings/notes when reporting PR reviews or changes to the user.
- Versioning: root and `cli/` are versioned independently; changes are logged in `CHANGELOG.md`. Commit style is Conventional Commits (`fix(translator): …`, `feat(...)`).
<!-- BEGIN:nextjs-agent-rules -->

View File

@@ -100,6 +100,12 @@ docker rm -f 9router
# re-run the quick start command
```
To pin a specific version instead of following `latest`, use a numbered image tag:
```bash
docker pull decolua/9router:0.5.81
```
---
# 🛠 For Developers
@@ -107,7 +113,7 @@ docker rm -f 9router
## Build image locally (test)
```bash
cd app && docker build -t 9router .
docker build -t 9router .
docker run --rm -p 20128:20128 \
-v "$HOME/.9router:/app/data" \
@@ -115,18 +121,67 @@ docker run --rm -p 20128:20128 \
9router
```
The Dockerfile uses the official Alpine and npm registries by default. Regional mirrors can be supplied when needed:
```bash
docker build \
--build-arg ALPINE_MIRROR=mirrors.aliyun.com \
--build-arg NPM_REGISTRY=https://registry.npmmirror.com/ \
-t 9router .
```
## Publish (automatic via CI)
Push a git tag `v*` → GitHub Actions builds multi-platform (amd64+arm64) and pushes to:
- `ghcr.io/decolua/9router:v{version}` + `:latest`
- `decolua/9router:v{version}` + `:latest`
Push a Docker-safe semver git tag `vX.Y.Z` (or a prerelease such as `vX.Y.Z-rc.1`) → GitHub Actions builds `linux/amd64` and `linux/arm64` on native runners, health-checks each platform image, verifies the resulting manifest and `/api/health`, then publishes:
- `ghcr.io/decolua/9router:X.Y.Z` + `:latest`
- `decolua/9router:X.Y.Z` + `:latest`
The `v` prefix is used only for the git tag; image tags omit it. A stable tag push promotes `latest`, but a prerelease tag such as `vX.Y.Z-rc.1` publishes only its numbered image by default. Prereleases require an explicit manual `promote_latest` opt-in. Promotion happens only after both native platform builds, both platform health checks, manifest inspection, and the resolved-manifest smoke test succeed. A failed or timed-out platform build therefore cannot move `latest`.
The workflow rejects SemVer build metadata such as `v1.2.3+build.7` because the `+` form is not a valid Docker image tag. The git tag and both `package.json` versions must match exactly.
```bash
# Use scripts/release.js (recommended)
node scripts/release.js "Release title" "Notes"
# Or manually
git tag v0.4.x && git push origin v0.4.x
git tag v0.5.81 && git push origin v0.5.81
```
Workflow: `app/.github/workflows/docker-publish.yml`
To republish an existing tag, run the `Build and Push Docker Image` workflow manually and provide the exact tag, for example `v0.5.81`, in the `release_tag` input. Manual runs publish the numbered tag but leave `latest` unchanged by default:
```text
release_tag: v0.5.81
promote_latest: false
```
The `promote_latest` checkbox is an explicit opt-in for changing `latest`. Use it when a deliberate rollback or recovery should make that version the current default:
```text
release_tag: v0.5.75
promote_latest: true
```
Numbered image tags are mutable because a republish can replace their manifest. For a deployment that must be immutable, pin the image digest instead:
```bash
docker pull decolua/9router@sha256:<verified-digest>
```
The release workflow runs `/api/health` on each native `amd64` and `arm64` platform image before it uploads the digest artifact or assembles the multi-platform manifest. It then runs a second health check against the resolved version manifest before any requested `latest` promotion.
During recovery, the selected tag remains the application source while the Dockerfile from the workflow revision is used, so an older tag can be rebuilt with the current publishing fixes.
The workflow is tag-driven. Creating a git tag does not automatically create a GitHub Release, so the Releases page and the published package/image tags can be at different versions unless a maintainer creates a release separately.
The upstream repository needs these repository secrets for Docker Hub publishing:
- `DOCKERHUB_USERNAME`
- `DOCKERHUB_TOKEN`
GHCR publishing uses the workflow's `GITHUB_TOKEN` with package write permission. Forks can publish to their own GHCR namespace, but Docker Hub publication is restricted to the upstream `decolua/9router` repository.
The optional repository variables `ALPINE_MIRROR` and `NPM_REGISTRY` can override the default package mirrors used by the CI Docker build.
Workflow: `.github/workflows/docker-publish.yml`

View File

@@ -1,25 +1,51 @@
# syntax=docker/dockerfile:1.7
ARG NODE_IMAGE=node:22-alpine
ARG ALPINE_MIRROR=dl-cdn.alpinelinux.org
ARG NPM_REGISTRY=https://registry.npmjs.org/
ARG APP_VERSION=unknown
FROM ${NODE_IMAGE} AS base
ARG ALPINE_MIRROR
WORKDIR /app
# CN mirror for apk (used by builder and runner stages)
RUN sed -i 's|dl-cdn.alpinelinux.org|mirrors.aliyun.com|g' /etc/apk/repositories
# Use the official Alpine mirror by default. A repository variable/build arg can
# override it for environments that require a regional mirror.
RUN if [ "$ALPINE_MIRROR" != "dl-cdn.alpinelinux.org" ]; then \
sed -i "s|dl-cdn.alpinelinux.org|${ALPINE_MIRROR}|g" /etc/apk/repositories; \
fi
FROM base AS builder
ARG NPM_REGISTRY
RUN apk --no-cache upgrade && apk --no-cache add python3 make g++ linux-headers
RUN apk add --no-cache python3 make g++ linux-headers
COPY package.json ./
RUN npm install --registry=https://registry.npmmirror.com
RUN --mount=type=cache,target=/root/.npm \
npm install \
--registry="${NPM_REGISTRY}" \
--fetch-retries=5 \
--fetch-retry-factor=2 \
--fetch-retry-mintimeout=10000 \
--fetch-retry-maxtimeout=120000 \
--fetch-timeout=300000
COPY . ./
ENV NEXT_TELEMETRY_DISABLED=1
# Inter is self-hosted (public/fonts + src/app/fonts-inter.css), so this build needs no
# route to fonts.googleapis.com / fonts.gstatic.com — only the npm mirror above is required.
RUN npm run build
FROM ${NODE_IMAGE} AS runner
ARG ALPINE_MIRROR
ARG APP_VERSION
WORKDIR /app
LABEL org.opencontainers.image.title="9router"
RUN if [ "$ALPINE_MIRROR" != "dl-cdn.alpinelinux.org" ]; then \
sed -i "s|dl-cdn.alpinelinux.org|${ALPINE_MIRROR}|g" /etc/apk/repositories; \
fi
LABEL org.opencontainers.image.title="9router" \
org.opencontainers.image.version="${APP_VERSION}"
ENV NODE_ENV=production
ENV PORT=20128
@@ -48,8 +74,9 @@ RUN mkdir -p /app/data && chown -R node:node /app && \
mkdir -p /app/data-home && chown node:node /app/data-home && \
ln -sf /app/data-home /root/.9router 2>/dev/null || true
# Fix permissions at runtime (handles mounted volumes)
RUN apk --no-cache upgrade && apk --no-cache add su-exec && \
# Avoid a full distribution upgrade in the runtime image. It makes builds less
# reproducible and is unrelated to installing the runtime entrypoint helper.
RUN apk add --no-cache su-exec && \
printf '#!/bin/sh\nchown -R node:node /app/data /app/data-home 2>/dev/null\nexec su-exec node "$@"\n' > /entrypoint.sh && \
chmod +x /entrypoint.sh

View File

@@ -110,7 +110,22 @@ PORT=20128 NEXT_PUBLIC_BASE_URL=http://localhost:20128 npm run dev
Production mode:
```bash
# Create Temporary Memory For Build
sudo fallocate -l 2G /swapfile_temp
sudo chmod 600 /swapfile_temp
sudo mkswap /swapfile_temp
sudo swapon /swapfile_temp
export MAKEFLAGS="-j1"
export DLIB_NO_GUI_SUPPORT=1
export CFLAGS="-mno-avx"
npm run build
# Clear temporary swap
sudo swapoff /swapfile_temp
sudo rm /swapfile_temp
PORT=20128 HOSTNAME=0.0.0.0 NEXT_PUBLIC_BASE_URL=http://localhost:20128 npm run start
```
@@ -215,7 +230,14 @@ Default URLs:
<b>🇻🇳 Tiếng Việt</b><br/>
<sub>Hướng Dẫn Setup OpenClaw + 9Router: Tạo Bot Zalo AI Tự Động Từ A-Z<br/>by <a href="https://github.com/tuanminhhole">tuanminhhole</a></sub>
</td>
<td align="center" width="320"></td>
<td align="center" width="320">
<a href="https://www.youtube.com/watch?v=hgnE7MKi3Y4">
<img src="https://img.youtube.com/vi/hgnE7MKi3Y4/maxresdefault.jpg" alt="Bye Limit! Cara Bikin Sistem 'AI Unlimited' 100% Gratis Dengan 9Router!
" width="300"/>
</a><br/>
<b>🇮🇩 Indonesia</b><br/>
<sub>Bye Limit! Cara Bikin Sistem "AI Unlimited" 100% Gratis Dengan 9Router!<br/>by <a href="https://www.youtube.com/@neptiver">neptiver</a></sub>
</td>
<td align="center" width="320"></td>
<td align="center" width="320"></td>
</tr>

1
cli/.gitignore vendored
View File

@@ -1,2 +1,3 @@
app/*
node_modules/*
.tray-build/

View File

@@ -111,7 +111,7 @@ Any tool supporting OpenAI/Claude-compatible API works.
Full docs, advanced setup, video tutorials & development guide:
- **GitHub**: https://github.com/decolua/9router
- **Full README**: https://github.com/decolua/9router/blob/main/app/README.md
- **Full README**: https://github.com/decolua/9router/blob/master/README.md
- **Website**: https://9router.com
---

View File

@@ -5,9 +5,17 @@
//
// We use the maintained `systray2` fork. The original `systray@1.0.5` package
// bundles a 2017 x86_64 Go binary whose Mach-O headers are rejected by modern
// dyld (macOS 14+), so the tray silently fails to register on Apple Silicon.
// dyld (macOS 14+), so it fails to load at all.
//
// Note that systray2 is NOT an Apple Silicon fix: like its predecessor it ships
// only an x86_64 `tray_darwin_release`, and picks it by process.platform with no
// process.arch branch, so there is no native slice to select. On arm64 macOS the
// tray therefore needs Rosetta 2 and dies with EBADARCH without it. We overlay
// our own arm64 build of the same upstream source on top — see ensureArm64TrayBin.
const { spawnSync } = require("child_process");
const crypto = require("crypto");
const fs = require("fs");
const os = require("os");
const path = require("path");
const { getRuntimeDir, getRuntimeNodeModules, runNpmInstall, summarizeNpmError } = require("./sqliteRuntime");
@@ -15,6 +23,19 @@ const SYSTRAY_PKG = "systray2";
const SYSTRAY_VERSION = "2.1.4";
const LEGACY_SYSTRAY_PKG = "systray";
// Pinned `tray-binaries` release rather than `latest`, so the URL is stable and
// the artifact can only change by a deliberate re-upload. The workflow's publish
// step re-derives this repo from the literal below and refuses to upload
// anywhere else, so the integrity gate can't drift from what clients fetch.
//
// The asset is built by .github/workflows/tray-binaries.yml on a macos-15 runner.
// cgo compiles AppKit against the runner's SDK, so this value tracks that image:
// when GitHub updates it the sha changes, the workflow refuses to publish, and
// this constant must be bumped in the same change as the re-upload.
const ARM64_TRAY_URL = "https://github.com/decolua/9router/releases/download/tray-binaries/tray_darwin_arm64";
const ARM64_TRAY_SHA256 = "487e3c365aaa1eb6ad295bf3989711e975b52cee07505bf641c8559954881c81";
const ARM64_RETRY_COOLDOWN_MS = 24 * 60 * 60 * 1000;
function hasSystray() {
return fs.existsSync(path.join(getRuntimeNodeModules(), SYSTRAY_PKG, "package.json"));
}
@@ -71,6 +92,124 @@ function ensureRuntimeDir() {
return dir;
}
// A thin (non-fat) 64-bit Mach-O stores its magic then cputype, both LE.
// CPU_TYPE_ARM64 is CPU_TYPE_ARM | CPU_ARCH_ABI64. Fat/universal binaries use a
// different magic and are reported as "not arm64" here, which is fine: we only
// ever overlay a thin arm64 build and only need to tell it apart from x86_64.
function isArm64MachO(file) {
let fd = null;
try {
fd = fs.openSync(file, "r");
const buf = Buffer.alloc(8);
fs.readSync(fd, buf, 0, 8, 0);
if (buf.readUInt32LE(0) !== 0xfeedfacf) return false;
return buf.readUInt32LE(4) === 0x0100000c;
} catch {
return false;
} finally {
if (fd !== null) try { fs.closeSync(fd); } catch {}
}
}
// systray2 is constructed with copyDir:true, so what actually executes is
// ~/.cache/node-systray/<version>/tray_darwin_release, and index.js only re-copies
// when that path is absent. An overlaid binary stays invisible until this is cleared.
// Scoped to our systray2 version: the parent dir is machine-global and shared
// with any other node-systray consumer.
function bustSystrayCopyCache() {
try {
fs.rmSync(path.join(os.homedir(), ".cache", "node-systray", SYSTRAY_VERSION), { recursive: true, force: true });
} catch {}
}
function arm64AttemptMarker() {
return path.join(getRuntimeDir(), ".tray-arm64-attempt");
}
// ensureTrayRuntime runs synchronously on every `9router` start (cli.js), so a
// failed download must not re-block the next launch. Retry at most daily.
function recentlyAttemptedArm64() {
try {
const at = Number(fs.readFileSync(arm64AttemptMarker(), "utf8").trim());
return Number.isFinite(at) && Date.now() - at < ARM64_RETRY_COOLDOWN_MS;
} catch {
return false;
}
}
function markArm64Attempt() {
try { fs.writeFileSync(arm64AttemptMarker(), String(Date.now())); } catch {}
}
// Cleared on success so the cooldown only ever throttles *failures*. Without
// this, anything that restores systray2's x86_64 binary later — notably a
// globally installed 9router older than this change, which shares the same
// ~/.9router/runtime — would leave the user waiting out the cooldown.
function clearArm64Attempt() {
try { fs.rmSync(arm64AttemptMarker(), { force: true }); } catch {}
}
// Throws on any failure so the caller has a single error path.
function downloadFile(url, dest, timeoutSec) {
// darwin-only path, and curl ships with macOS, so this needs no extra dep and
// keeps the caller synchronous.
const res = spawnSync("curl", ["-fsSL", "--max-time", String(timeoutSec), "-o", dest, url], {
encoding: "utf8",
timeout: (timeoutSec + 5) * 1000
});
if (res.status === 0 && fs.existsSync(dest)) return;
const detail = (res.stderr || res.error?.message || `curl exit ${res.status}`).trim().split("\n").pop();
throw new Error(detail || "download failed");
}
function sha256File(file) {
return crypto.createHash("sha256").update(fs.readFileSync(file)).digest("hex");
}
// Replace systray2's x86_64 macOS binary with a native arm64 build so Apple
// Silicon users get a tray without installing Rosetta 2. Any failure leaves the
// Intel binary untouched, which still works under Rosetta.
//
// Takes no `silent` flag on purpose: cli.js calls ensureTrayRuntime({silent:true})
// synchronously on every start, and a stalled curl would otherwise freeze the
// launch for up to 30s with no output at all. These lines print at most once per
// 24h on failure and once ever on success, so they are worth more than the quiet.
function ensureArm64TrayBin() {
if (process.platform !== "darwin" || process.arch !== "arm64") return { skipped: true };
const binPath = path.join(getRuntimeNodeModules(), SYSTRAY_PKG, "traybin", "tray_darwin_release");
if (!fs.existsSync(binPath)) return { skipped: true };
if (isArm64MachO(binPath)) return { native: true };
if (recentlyAttemptedArm64()) return { deferred: true };
markArm64Attempt();
console.log("⏳ Downloading native Apple Silicon tray binary...");
// pid-scoped: two concurrent starts (postinstall racing cli.js, or two
// terminals) would otherwise interleave writes to one file, fail each other's
// checksum, and delete each other's in-flight download from the catch below.
const tmp = `${binPath}.arm64.${process.pid}.tmp`;
try {
downloadFile(ARM64_TRAY_URL, tmp, 30);
const sum = sha256File(tmp);
// Integrity matters more than usual: this is an executable that runs on
// every Apple Silicon user's machine.
if (sum !== ARM64_TRAY_SHA256) throw new Error(`checksum mismatch (got ${sum.slice(0, 12)}…)`);
if (!isArm64MachO(tmp)) throw new Error("downloaded file is not an arm64 Mach-O");
fs.chmodSync(tmp, 0o755);
fs.renameSync(tmp, binPath);
bustSystrayCopyCache();
clearArm64Attempt();
console.log("✅ Native Apple Silicon tray installed");
return { native: true, installed: true };
} catch (e) {
try { fs.rmSync(tmp, { force: true }); } catch {}
console.warn("⚠️ Native tray download failed — falling back to the Intel binary");
console.warn(` Reason: ${e.message}`);
console.warn(" The Intel tray needs Rosetta 2: softwareupdate --install-rosetta --agree-to-license");
return { native: false, error: e.message };
}
}
function npmInstall(pkgs, { silent = false } = {}) {
const cwd = ensureRuntimeDir();
if (!silent) console.log("⏳ Installing system tray (first run)...");
@@ -94,14 +233,20 @@ function ensureTrayRuntime({ silent = false } = {}) {
if (process.platform === "win32") {
return { systray: false, skipped: true };
}
if (hasSystray()) {
let ready = hasSystray();
if (!ready) {
ready = npmInstall([`${SYSTRAY_PKG}@${SYSTRAY_VERSION}`], { silent }) && hasSystray();
}
if (ready) {
chmodSystrayBin({ silent });
if (!silent) console.log("✅ System tray ready");
return { systray: true };
}
const ok = npmInstall([`${SYSTRAY_PKG}@${SYSTRAY_VERSION}`], { silent });
if (ok) chmodSystrayBin({ silent });
return { systray: ok && hasSystray() };
// Runs after the ready log so a download failure doesn't read as a broken
// tray — the Intel binary still works under Rosetta.
const arm64 = ready ? ensureArm64TrayBin() : { skipped: true };
return { systray: ready, arm64 };
}
module.exports = { ensureTrayRuntime };
module.exports = { ensureTrayRuntime, ensureArm64TrayBin };

View File

@@ -1,6 +1,6 @@
{
"name": "9router",
"version": "0.5.69",
"version": "0.5.91",
"description": "9Router CLI - Start and manage 9Router server",
"bin": {
"9router": "./cli.js"
@@ -16,6 +16,7 @@
"scripts": {
"dev": "nodemon -I --watch cli.js --watch src --watch hooks --ext js,json cli.js",
"build": "node scripts/build-cli.js",
"build:tray-arm64": "node scripts/buildTrayArm64.js",
"pack:cli": "npm run build && npm pack --pack-destination ..",
"publish:cli": "npm run build && npm publish",
"postinstall": "node hooks/postinstall.js",
@@ -29,7 +30,8 @@
"react-dom": "19.2.1"
},
"comment_sqlite": "sql.js + better-sqlite3 are NOT bundled here. They are installed into ~/.9router/runtime/node_modules by hooks/postinstall.js (and re-checked at runtime by cli.js). This avoids Windows EBUSY errors when updating the global CLI, since native .node files no longer live under the locked install dir.",
"comment_systray": "systray2 is NOT bundled here. It is lazy-installed into ~/.9router/runtime/node_modules by hooks/postinstall.js on macOS/Linux only. Windows uses PowerShell NotifyIcon (zero binary). This avoids shipping unsigned Go binaries that trigger antivirus false positives (Kaspersky). We use the systray2 fork because the legacy systray@1.0.5 ships a 2017 x86_64 binary that fails on modern macOS dyld.",
"comment_systray": "systray2 is NOT bundled here. It is lazy-installed into ~/.9router/runtime/node_modules by hooks/postinstall.js on macOS/Linux only. Windows uses PowerShell NotifyIcon (zero binary). This avoids shipping unsigned Go binaries that trigger antivirus false positives (Kaspersky). We use the systray2 fork because the legacy systray@1.0.5 ships a 2017 x86_64 binary that fails to load on modern macOS dyld. Neither package ships an arm64 macOS binary, so on Apple Silicon hooks/trayRuntime.js overlays our own arm64 build from the tray-binaries GitHub release; without it the tray requires Rosetta 2.",
"comment_tray_arm64": "tray_darwin_arm64 is built by scripts/buildTrayArm64.js (npm run build:tray-arm64) from felixhao28/systray-portable — the same source systray2's binary comes from — and uploaded to the pinned 'tray-binaries' GitHub release. Its sha256 is pinned as ARM64_TRAY_SHA256 in hooks/trayRuntime.js and verified after every download; rebuild and update both together. .github/workflows/tray-binaries.yml builds the same artifact in CI but refuses to publish when the sha diverges from the pin.",
"engines": {
"node": ">=18.0.0"
},

View File

@@ -0,0 +1,108 @@
#!/usr/bin/env node
// Rebuilds tray_darwin_arm64, the native Apple Silicon menubar binary that
// hooks/trayRuntime.js overlays on top of systray2's x86_64-only build.
//
// Must run on macOS: getlantern/systray is cgo against AppKit, so the arm64
// slice needs a real macOS SDK. Requires Go on PATH (`mise use -g go@latest`).
//
// Output is NOT reproducible across Go versions even with -s -w, so after a
// rebuild you must re-upload the asset and update ARM64_TRAY_SHA256 in
// hooks/trayRuntime.js — this script prints both and fails if they diverge.
const { execFileSync, spawnSync } = require("child_process");
const crypto = require("crypto");
const fs = require("fs");
const os = require("os");
const path = require("path");
const UPSTREAM_REPO = "https://github.com/felixhao28/systray-portable.git";
// master as of 2021-09-15, the commit systray2@2.1.4's own binary was built from.
const UPSTREAM_COMMIT = "6eddc917bf39fcc0d95b57a0741d0c065fbd1e23";
const outDir = path.join(__dirname, "..", ".tray-build");
const outFile = path.join(outDir, "tray_darwin_arm64");
const trayRuntimePath = path.join(__dirname, "..", "hooks", "trayRuntime.js");
function fail(msg) {
console.error(`\n❌ ${msg}`);
process.exit(1);
}
function run(cmd, args, opts = {}) {
const res = spawnSync(cmd, args, { stdio: "inherit", ...opts });
if (res.status !== 0) fail(`${cmd} ${args.join(" ")} exited with ${res.status}`);
}
if (process.platform !== "darwin") fail("must be run on macOS (cgo needs the AppKit SDK)");
const go = spawnSync("go", ["version"], { encoding: "utf8" });
if (go.status !== 0) {
fail("Go toolchain not found on PATH. Install it with: mise use -g go@latest");
}
console.log(`Go: ${go.stdout.trim()}`);
const srcDir = fs.mkdtempSync(path.join(os.tmpdir(), "systray-portable-"));
// fail() exits through process.exit(), which does not unwind the stack, so a
// try/finally here would leak the clone on every failed build. An exit handler
// covers normal completion, fail(), and uncaught exceptions alike.
process.on("exit", () => {
try { fs.rmSync(srcDir, { recursive: true, force: true }); } catch {}
});
console.log(`\nCloning ${UPSTREAM_REPO} @ ${UPSTREAM_COMMIT.slice(0, 7)}`);
run("git", ["clone", "--quiet", UPSTREAM_REPO, srcDir]);
run("git", ["-C", srcDir, "checkout", "--quiet", UPSTREAM_COMMIT]);
const head = execFileSync("git", ["-C", srcDir, "log", "-1", "--format=%H %ad %s", "--date=short"], {
encoding: "utf8"
}).trim();
console.log(`HEAD: ${head}`);
run("go", ["mod", "download"], { cwd: srcDir });
console.log("\nBuilding darwin/arm64...");
// -trimpath strips the local build directory from the binary, so two builds
// from the same commit + Go version hash identically regardless of where they
// ran. Without it the pinned sha256 could never be regenerated.
run("go", ["build", "-trimpath", "-ldflags", "-s -w", "-o", outFile, "tray.go"], {
cwd: srcDir,
env: { ...process.env, CGO_ENABLED: "1", GOOS: "darwin", GOARCH: "arm64" }
});
// ── Verify ────────────────────────────────────────────────────────────────
const buf = Buffer.alloc(8);
const fd = fs.openSync(outFile, "r");
fs.readSync(fd, buf, 0, 8, 0);
fs.closeSync(fd);
if (buf.readUInt32LE(0) !== 0xfeedfacf || buf.readUInt32LE(4) !== 0x0100000c) {
fail("output is not a thin arm64 Mach-O");
}
// Apple Silicon refuses to execute an unsigned binary. Go's linker applies an
// ad-hoc signature automatically; confirm it survived.
const sig = spawnSync("codesign", ["--verify", outFile], { encoding: "utf8" });
if (sig.status !== 0) fail(`ad-hoc signature invalid: ${(sig.stderr || "").trim()}`);
const sha256 = crypto.createHash("sha256").update(fs.readFileSync(outFile)).digest("hex");
const sizeMb = (fs.statSync(outFile).size / 1024 / 1024).toFixed(2);
console.log(`\n✅ ${outFile}`);
console.log(` arch: arm64 (ad-hoc signed, verified)`);
console.log(` size: ${sizeMb} MB`);
console.log(` sha256: ${sha256}`);
// Whitespace-tolerant: a formatter could wrap the assignment across lines, and
// a null match must fail loudly rather than silently read as "no pin".
const pinMatch = fs.readFileSync(trayRuntimePath, "utf8").match(/ARM64_TRAY_SHA256\s*=\s*"([0-9a-f]{64})"/);
if (!pinMatch) fail(`could not find ARM64_TRAY_SHA256 in ${trayRuntimePath}`);
const pinned = pinMatch[1];
if (pinned === sha256) {
console.log(`\n Matches ARM64_TRAY_SHA256 in hooks/trayRuntime.js — no code change needed.`);
} else {
console.log(`\n⚠️ Differs from ARM64_TRAY_SHA256 in hooks/trayRuntime.js (${pinned}).`);
console.log(` Re-upload the release asset, then update that constant to the sha256 above.`);
}
console.log(`\nNext: upload to the pinned release tag, keeping the asset name stable:`);
console.log(` gh release upload tray-binaries "${outFile}" --clobber`);

View File

@@ -141,11 +141,14 @@ function initWindowsTray(options) {
/**
* macOS/Linux tray via systray binary
*
* Prefers `systray2` (active fork of `systray`, ships newer
* getlantern/systray-portable binaries that work on macOS 14+ and Apple
* Silicon under Rosetta). Falls back to legacy `systray@1.0.5` if systray2
* is not available, though that binary's Mach-O headers are rejected by
* modern dyld and the icon will not appear.
* Prefers `systray2`, the active fork of `systray`. Both ship only an x86_64
* `tray_darwin_release` and select it by process.platform alone, so on Apple
* Silicon the tray runs under Rosetta 2 and fails with EBADARCH when Rosetta is
* absent. hooks/trayRuntime.js overlays a native arm64 build over that file to
* avoid the dependency; the fallbacks below are Intel-only.
*
* Falls back to legacy `systray@1.0.5` if systray2 is unavailable, though that
* binary's Mach-O headers are rejected by modern dyld and no icon will appear.
*/
function resolveSystray() {
let runtimeDir = null;

View File

@@ -2,22 +2,24 @@ const api = require("../api/client");
const { prompt } = require("./input");
const { clearScreen } = require("./display");
// Provider alias order: OAuth first, then API Key (matches ModelSelectModal)
// Provider alias order: OAuth first, then Free, then API Key
const PROVIDER_ALIAS_ORDER = [
"cc", "ag", "cx", "if", "qw", "gc", "gh", "kr",
"cc", "ag", "cx", "if", "qw", "gc", "gh", "kr", "oc",
"openrouter", "glm", "kimi", "minimax", "openai", "anthropic", "gemini"
];
// Alias to display name mapping
const PROVIDER_ALIAS_NAMES = {
cc: "Claude Code",
ag: "Antigravity",
ag: "Antigravity",
cx: "OpenAI Codex",
if: "iFlow AI",
qw: "Qwen Code",
gc: "Gemini CLI",
gh: "GitHub Copilot",
kr: "Kiro AI",
oc: "OpenCode Free",
opencode: "OpenCode Free",
openrouter: "OpenRouter",
glm: "GLM Coding",
kimi: "Kimi Coding",
@@ -27,35 +29,83 @@ const PROVIDER_ALIAS_NAMES = {
gemini: "Gemini"
};
const PROVIDER_ID_TO_ALIAS = {
claude: "cc",
codex: "cx",
"gemini-cli": "gc",
github: "gh",
antigravity: "ag",
iflow: "if",
qwen: "qw",
kiro: "kr",
cursor: "cu",
cline: "cline",
clinepass: "clinepass",
qoder: "qd",
"qoder-cn": "qd",
gitlab: "gitlab",
"codebuddy-cn": "cb",
"codebuddy-intl": "cbai",
kimchi: "kimchi",
"grok-cli": "grok-cli",
trae: "trae",
windsurf: "windsurf",
zed: "zed",
opencode: "oc",
"opencode-go": "ocg",
"opencode-zen": "ocz",
};
// Providers usable without stored credentials
const NO_AUTH_PROVIDERS = new Set(["opencode", "oc"]);
/**
* Get all available models grouped by provider + combos
* Get all available models grouped by provider + combos (filtered by active connections)
* @returns {Promise<{combos: Array, groups: Object}>}
*/
async function getAvailableModelsGrouped() {
const result = await api.getAvailableModels();
if (!result.success) return { combos: [], groups: {} };
const models = result.data?.data || [];
const [modelsResult, providersResult] = await Promise.all([
api.getAvailableModels(),
api.getProviders()
]);
if (!modelsResult.success) return { combos: [], groups: {} };
const connections = providersResult.success ? (providersResult.data?.connections || []) : [];
const activeAliases = new Set(NO_AUTH_PROVIDERS);
connections.forEach(conn => {
if (conn.isActive === false) return;
const p = conn.provider;
if (!p) return;
activeAliases.add(p);
const alias = conn.providerSpecificData?.prefix || PROVIDER_ID_TO_ALIAS[p] || p;
activeAliases.add(alias);
});
const models = modelsResult.data?.data || [];
const combos = [];
const groups = {};
models.forEach(m => {
if (m.owned_by === "combo") {
combos.push(m.id);
} else {
const provider = m.owned_by;
// Only keep connected providers or noAuth providers
if (!activeAliases.has(provider)) return;
if (!groups[provider]) {
groups[provider] = [];
}
groups[provider].push(m.id);
}
});
return { combos, groups };
}
/**
* Display model list and prompt for selection
* Display model list and prompt for selection with provider grouping & search
* @param {string} title - Title to display
* @param {string} currentValue - Current selected value (optional)
* @param {Object} options - { excludeCombos?: boolean }
@@ -68,64 +118,214 @@ async function selectModelFromList(title, currentValue = "", options = {}) {
const totalModels = combos.length + Object.values(groups).flat().length;
if (totalModels === 0) {
clearScreen();
console.log(`\n🎯 ${title}`);
console.log("=".repeat(50));
console.log("\n No connected providers found.");
console.log(" Please connect a provider in Providers menu first.\n");
console.log(" m. ✍️ Enter custom model ID");
console.log(" 0. Cancel\n");
const act = await prompt("Select option (m/0): ");
const trimmed = act.trim();
if (trimmed.toLowerCase() === "m") {
const custom = await prompt("Enter custom model ID: ");
return custom.trim() || null;
}
return null;
}
// Build flat list for selection
const allModels = [];
// Display
clearScreen();
console.log(`\n🎯 ${title}`);
console.log("=".repeat(50));
if (currentValue) {
console.log(`Current: ${currentValue}\n`);
} else {
console.log();
}
let idx = 1;
// Combos first (skipped when excludeCombos is true)
// All models for flat search
const allModelsList = [
...combos,
...Object.values(groups).flat()
];
// Build category list
const categories = [];
if (combos.length > 0) {
console.log("[Combos]");
combos.forEach(combo => {
console.log(` ${idx}. ${combo}`);
allModels.push(combo);
idx++;
categories.push({
id: "combos",
name: "[Combos]",
models: combos
});
console.log();
}
// Provider groups in order (by alias)
const sortedProviders = Object.keys(groups).sort((a, b) => {
const idxA = PROVIDER_ALIAS_ORDER.indexOf(a);
const idxB = PROVIDER_ALIAS_ORDER.indexOf(b);
return (idxA === -1 ? 999 : idxA) - (idxB === -1 ? 999 : idxB);
});
sortedProviders.forEach(provider => {
sortedProviders.forEach((provider) => {
const providerName = PROVIDER_ALIAS_NAMES[provider] || provider;
console.log(`[${providerName}]`);
groups[provider].forEach(model => {
console.log(` ${idx}. ${model}`);
allModels.push(model);
idx++;
categories.push({
id: provider,
name: providerName,
models: groups[provider]
});
console.log();
});
console.log(" 0. Cancel\n");
// Prompt for number input
const input = await prompt("Enter number: ");
const num = parseInt(input, 10);
if (isNaN(num) || num === 0 || num < 0 || num > allModels.length) {
return null;
let filterQuery = null;
while (true) {
clearScreen();
console.log(`\n🎯 ${title}`);
console.log("=".repeat(50));
if (currentValue) {
console.log(`Current: ${currentValue}\n`);
} else {
console.log();
}
// Active search view
if (filterQuery !== null) {
const q = filterQuery.toLowerCase().trim();
const matched = allModelsList.filter((m) => m.toLowerCase().includes(q));
console.log(`🔍 Search results for "${filterQuery}": (${matched.length} found)\n`);
if (matched.length === 0) {
console.log(" No matching models found.\n");
console.log(" 0. ← Back to providers");
console.log(" s. Search again\n");
const act = await prompt("Select option: ");
if (act.toLowerCase() === "s") {
const newQ = await prompt("Enter search keyword: ");
filterQuery = newQ.trim() || null;
} else {
filterQuery = null;
}
continue;
}
matched.forEach((m, i) => {
console.log(` ${i + 1}. ${m}`);
});
console.log("\n 0. ← Back to providers");
console.log(" s. Search again\n");
const input = await prompt("Enter number to select (or 0/s): ");
if (input.toLowerCase() === "s") {
const newQ = await prompt("Enter search keyword: ");
filterQuery = newQ.trim() || null;
continue;
}
const num = parseInt(input, 10);
if (isNaN(num) || num === 0) {
filterQuery = null;
continue;
}
if (num > 0 && num <= matched.length) {
return matched[num - 1];
}
continue;
}
// If only 1 category exists, jump straight into its model list
if (categories.length === 1) {
const singleCategory = categories[0];
console.log(`[${singleCategory.name}]`);
singleCategory.models.forEach((m, i) => {
console.log(` ${i + 1}. ${m}`);
});
console.log();
console.log(" s. 🔍 Search models");
console.log(" m. ✍️ Enter custom model ID");
console.log(" 0. Cancel\n");
const input = await prompt("Enter choice (number / s / m / 0): ");
const trimmed = input.trim();
if (!trimmed || trimmed === "0") return null;
const lower = trimmed.toLowerCase();
if (lower === "s") {
const q = await prompt("Enter search keyword: ");
if (q.trim()) filterQuery = q.trim();
continue;
}
if (lower === "m") {
const customModel = await prompt("Enter custom model ID: ");
if (customModel.trim()) return customModel.trim();
continue;
}
const num = parseInt(trimmed, 10);
if (!isNaN(num) && num > 0 && num <= singleCategory.models.length) {
return singleCategory.models[num - 1];
}
filterQuery = trimmed;
continue;
}
// Multiple categories view
console.log("[Providers & Groups]");
categories.forEach((cat, i) => {
console.log(` ${i + 1}. ${cat.name} (${cat.models.length} models)`);
});
console.log();
console.log(" s. 🔍 Search models");
console.log(" m. ✍️ Enter custom model ID");
console.log(" 0. Cancel\n");
const input = await prompt("Enter choice (number / keyword / s / m): ");
const trimmed = input.trim();
if (!trimmed || trimmed === "0") {
return null;
}
const lower = trimmed.toLowerCase();
if (lower === "s") {
const q = await prompt("Enter search keyword: ");
if (q.trim()) {
filterQuery = q.trim();
}
continue;
}
if (lower === "m") {
const customModel = await prompt("Enter custom model ID: ");
if (customModel.trim()) {
return customModel.trim();
}
continue;
}
const num = parseInt(trimmed, 10);
// Selected a category
if (!isNaN(num) && num > 0 && num <= categories.length) {
const selectedCategory = categories[num - 1];
while (true) {
clearScreen();
console.log(`\n🎯 ${title} > ${selectedCategory.name}`);
console.log("=".repeat(50));
if (currentValue) {
console.log(`Current: ${currentValue}\n`);
} else {
console.log();
}
selectedCategory.models.forEach((m, i) => {
console.log(` ${i + 1}. ${m}`);
});
console.log("\n 0. ← Back\n");
const modelChoice = await prompt("Enter number to select (0 to back): ");
const modelNum = parseInt(modelChoice, 10);
if (isNaN(modelNum) || modelNum === 0) {
break;
}
if (modelNum > 0 && modelNum <= selectedCategory.models.length) {
return selectedCategory.models[modelNum - 1];
}
}
continue;
}
// User typed text directly -> treat as search query
filterQuery = trimmed;
}
return allModels[num - 1];
}
module.exports = {

View File

@@ -118,6 +118,31 @@ Cursor/Cline/Any tool:
---
## Cursor / Claude Default Combos
Cursor and Claude Code send **unprefixed** model IDs (`composer-2.5`, `claude-opus-5`, `opus`), while 9Router routes with provider prefixes (`cu/composer-2.5`, `cc/claude-opus-5`). Default combo generators bridge that gap.
On **Dashboard → Combos**:
1. Click **Cursor Default** or **Claude Default**
2. Confirm the preview (new vs already-existing names)
3. 9Router creates one combo per client model ID, seeded with the matching prefixed route
**Examples:**
| Combo name (what the client sends) | Seeded model (what 9Router routes) |
|------------------------------------|------------------------------------|
| `composer-2.5` | `cu/composer-2.5` |
| `cursor-grok-4.6-high-fast` | `cu/cursor-grok-4.6-high-fast` |
| `claude-opus-5` | `cc/claude-opus-5` |
| `opus` | `cc/claude-opus-5` |
Existing combo names are **skipped** (not overwritten). Edit any generated combo afterward to add fallbacks. Click the button again later to pick up new catalog IDs.
> These combos help when Cursor/Claude already talk to 9Router (`/v1` or `ANTHROPIC_BASE_URL`) and send their native model IDs. They do not change Cursor’s built-in Models tab by themselves.
---
## Example Combos
### Example 1: Premium Coding (Subscription → Cheap → Free)

View File

@@ -75,6 +75,10 @@ const nextConfig = {
source: "/responses",
destination: "/api/v1/responses"
},
{
source: "/systemone",
destination: "/api/v1/systemone"
},
{
source: "/v1beta/:path*",
destination: "/api/v1beta/:path*"

View File

@@ -4,7 +4,7 @@ Provider-agnostic SSE engine: one OpenAI-style request → any provider (LLM cha
## Request lifecycle (chat)
`handlers/chatCore.js` → `services/model.js` `parseModel` (resolve `provider/model`) → **pre-translate hooks** (`rtk/` tool_result compress, `rtk/headroom.js` proxy compress, `rtk/caveman.js` system inject — all fail-open) → `executors/index.js` `getExecutor(provider)` → `translator/index.js` `translateRequest` (client format → provider format) → `executor.execute()` (streams upstream) → `translateResponse` (provider chunks → client format) → SSE out.
`handlers/chatCore.js` → `services/model.js` `parseModel` (resolve `provider/model`) → **RTK for `cursor`** (`rtk/` compresses the source-format `tool_result` / `role:tool` in-place — its translator rewrites those shapes, so this one provider must run **before** translate) → `translator/index.js` `translateRequest` (client format → provider format) → **post-translate savers** (`rtk/` compress for every other provider, `rtk/headroom.js` proxy compress, `rtk/caveman.js` / `rtk/ponytail.js` system inject — all fail-open) → `executors/index.js` `getExecutor(provider)` → `executor.execute()` (streams upstream) → `translateResponse` (provider chunks → client format) → SSE out.
## Directory map

View File

@@ -7,6 +7,9 @@ import { createRequire } from "module";
export const GEMINI_CLI_VERSION = PROVIDERS["gemini-cli"]?.cliVersion;
export const GEMINI_CLI_API_CLIENT = PROVIDERS["gemini-cli"]?.apiClient;
// === Codex CLI === derive từ registry codex.transport
export const CODEX_CLI_VERSION = PROVIDERS["codex"]?.cliVersion;
// Map Node arch to Gemini CLI arch string (x64/x86/arm64/...)
function geminiCLIArch() {
const a = arch();
@@ -175,6 +178,11 @@ export const CLAUDE_SYSTEM_PROMPT = "You are Claude Code, Anthropic's official C
// makes the backend flag the request and answer 429 Quota Exhausted.
export const ANTIGRAVITY_PROMPT_REWRITES = [
{ from: "You are a Claude agent, built on Anthropic's Claude Agent SDK.", to: "" },
{ from: /You are Hermes(?: Agent)?(?:,\s*(?:an intelligent AI assistant|an AI assistant|an AI agent))?(?:,?\s*(?:built|created)\s+by\s+Nous Research)?\./gi, to: "You are an AI assistant." },
// Claude Code prepends this line to its system prompt. The Claude-format translator strips it,
// but OpenAI-format clients (e.g. proxies that convert Claude Code to /v1/chat/completions)
// pass it through, and any system text containing it gets a fake 429 RESOURCE_EXHAUSTED.
{ from: /^x-anthropic-billing-header:[^\n]*(?:\r?\n)*/gim, to: "" },
{ from: /opencode/gi, to: (m) => (m === "OpenCode" ? "Antigravity" : m === "OPENCODE" ? "ANTIGRAVITY" : "antigravity") }
];

View File

@@ -3,10 +3,13 @@ import REGISTRY from "../providers/registry/index.js";
// PROVIDER_MODELS now built from providers/registry (transport + models co-located)
import { PROVIDER_MODELS } from "../providers/index.js";
import { modelQuotaFamily, modelStrip, modelTargetFormat, modelSupportedFormats, normalizeModelId } from "../providers/models/schema.js";
import { CODEX_REVIEW_SUFFIX, isMuseSparkModel } from "../providers/models/helpers.js";
import { CODEX_REVIEW_SUFFIX, isMuseSparkModel, opencodeFamilyFormats } from "../providers/models/helpers.js";
import { FORMATS } from "../translator/formats.js";
export { PROVIDER_MODELS };
// OpenCode providers sharing the endpoint-family fallback for unknown model ids
const isOpenCodeAlias = (aliasOrId) => !aliasOrId || ["oc", "opencode", "ocg", "opencode-go", "ocz", "opencode-zen"].includes(aliasOrId);
// Helper functions
export function getProviderModels(aliasOrId) {
@@ -27,11 +30,14 @@ const DOT_VERSION_PROVIDERS = new Set(["kr", "kiro"]);
// ("claude-sonnet-4-5" ~= "claude-sonnet-4.5"). Other providers use exact match only.
function findModel(models, modelId, aliasOrId) {
if (!models) return undefined;
const found = models.find(m => m.id === modelId);
const baseModelId = typeof modelId === "string"
? modelId.replace(/\([^()]+\)\s*$/, "").trim()
: modelId;
const found = models.find(m => m.id === modelId || m.id === baseModelId);
if (found) return found;
if (!DOT_VERSION_PROVIDERS.has(aliasOrId)) return undefined;
const normalized = normalizeModelId(modelId);
if (normalized === modelId) return undefined;
const normalized = normalizeModelId(baseModelId);
if (normalized === baseModelId) return undefined;
return models.find(m => m.id === normalized);
}
@@ -50,20 +56,29 @@ export function findModelName(aliasOrId, modelId) {
}
export function getModelTargetFormat(aliasOrId, modelId) {
if ((!aliasOrId || aliasOrId === "oc" || aliasOrId === "opencode" || aliasOrId === "ocg" || aliasOrId === "opencode-go") && isMuseSparkModel(modelId)) {
if (isOpenCodeAlias(aliasOrId) && isMuseSparkModel(modelId)) {
return FORMATS.OPENAI_RESPONSES;
}
const models = PROVIDER_MODELS[aliasOrId];
if (!models) return null;
return modelTargetFormat(findModel(models, modelId, aliasOrId));
const found = findModel(models, modelId, aliasOrId);
if (found) return modelTargetFormat(found);
// Family fallback keeps modelsFetcher/passthrough ids on their endpoint lane
if (isOpenCodeAlias(aliasOrId)) return opencodeFamilyFormats(modelId)?.targetFormat || null;
return null;
}
// Declared upstream formats for a model (registry `supportedFormats`). Drives the
// per-model guard on the sourceFormat-matched transport; null when undeclared.
// Unknown OpenCode ids fall back to the family regex (chat lane by default) so
// auto-fetched models never wrongly use the sourceFormat-matched transport.
export function getModelSupportedFormats(aliasOrId, modelId) {
const models = PROVIDER_MODELS[aliasOrId];
if (!models) return null;
return modelSupportedFormats(findModel(models, modelId, aliasOrId));
const found = findModel(models, modelId, aliasOrId);
if (found) return modelSupportedFormats(found);
if (isOpenCodeAlias(aliasOrId)) return opencodeFamilyFormats(modelId)?.supportedFormats || [FORMATS.OPENAI];
return null;
}
export function getModelType(aliasOrId, modelId) {

View File

@@ -5,7 +5,7 @@ import { OAUTH_ENDPOINTS, ANTIGRAVITY_HEADERS, AG_DEFAULT_TOOLS, AG_TOOL_SUFFIX,
import { HTTP_STATUS } from "../config/runtimeConfig.js";
import { resolveSessionId, toNumericSessionId } from "../utils/sessionManager.js";
import { proxyAwareFetch } from "../utils/proxyFetch.js";
import { cleanJSONSchemaForAntigravity } from "../translator/formats/gemini.js";
import { cleanJSONSchemaForAntigravity, normalizeGeminiContents } from "../translator/formats/gemini.js";
import { DEFAULT_THINKING_AG_SIGNATURE } from "../config/defaultThinkingSignature.js";
import { getGeminiThoughtSignatureSync } from "../services/thoughtSignatureStore.js";
@@ -193,7 +193,7 @@ export class AntigravityExecutor extends BaseExecutor {
// ─── Standard (non-image) request ───
// Fix contents for Claude models via Antigravity
const contents = body.request?.contents?.map(c => {
const rawContents = (body.request?.contents || []).map(c => {
let role = c.role;
// functionResponse must be role "user" for Claude models
if (c.parts?.some(p => p.functionResponse)) {
@@ -212,7 +212,7 @@ export class AntigravityExecutor extends BaseExecutor {
const modifiedParts = parts?.map(p => {
if (!p.functionCall) return p;
const callId = p.functionCall.id;
const cachedSig = callId ? getGeminiThoughtSignatureSync(callId, sessionId) : null;
const cachedSig = callId ? getGeminiThoughtSignatureSync(callId, sessionId, body.model || model) : null;
const callSig = p.thoughtSignature || cachedSig || (!firstFunctionCallSeen ? DEFAULT_THINKING_AG_SIGNATURE : undefined);
firstFunctionCallSeen = true;
if (callSig) {
@@ -226,15 +226,13 @@ export class AntigravityExecutor extends BaseExecutor {
return p;
});
const partsChanged = parts?.length !== c.parts?.length || modifiedParts?.some((p, idx) => p !== c.parts[idx]);
if (role !== c.role || partsChanged) {
return {
...c, role,
parts: modifiedParts || parts,
};
}
return c;
return {
...c,
role,
parts: modifiedParts || parts || [],
};
});
const contents = normalizeGeminiContents(rawContents);
// Sanitize tool schemas and function names before sending to Antigravity.
let tools = body.request?.tools;
@@ -295,12 +293,19 @@ export class AntigravityExecutor extends BaseExecutor {
this._lastSessionId = transformedRequest.sessionId; // cached for buildHeaders (base.execute order)
// Official Antigravity client omits `requestType` entirely on the agent
// (chat) path. Sending `requestType: "agent"` here (or leaking it through
// from an upstream envelope via the ...body spread below) makes Google
// bucket the request and return a detail-free 429 RESOURCE_EXHAUSTED even
// with quota available. `image_gen` and
// `search` buckets are unaffected and keep their own requestType.
delete body.requestType;
return {
...body,
project: projectId,
model: body.model || model,
userAgent: "antigravity",
requestType: "agent",
requestId: buildIdeRequestId({ body, request: transformedRequest, credentials, model, requestType: "agent" }),
request: transformedRequest
};

View File

@@ -128,7 +128,7 @@ export class BaseExecutor {
for (let urlIndex = 0; urlIndex < fallbackCount; urlIndex++) {
const url = this.buildUrl(model, stream, urlIndex, credentials);
const transformedBody = this.transformRequest(model, body, stream, credentials);
const headers = this.buildHeaders(credentials, stream, url, model);
const headers = this.buildHeaders(credentials, stream, url, model, transformedBody);
if (!retryAttemptsByUrl[urlIndex]) retryAttemptsByUrl[urlIndex] = 0;

View File

@@ -7,11 +7,12 @@ import {
} from "../services/oauthCredentialManager.js";
import { normalizeResponsesInput } from "../translator/formats/responsesApi.js";
import { fetchImageAsBase64 } from "../translator/concerns/image.js";
import { getModelUpstreamId } from "../config/providerModels.js";
import { getModelUpstreamId, getProviderModels } from "../config/providerModels.js";
import { getThinkingLevels } from "../providers/thinkingLevels.js";
import { DEFAULT_RETRY_CONFIG, HTTP_STATUS, resolveRetryEntry } from "../config/runtimeConfig.js";
import { dbg } from "../utils/debugLog.js";
import { resolveSessionId } from "../utils/sessionManager.js";
import { stripCodexUnsupportedPatterns } from "../utils/codexToolSchema.js";
// SSE error patterns inside 200-OK bodies. Some retry same account first; capacity rotates accounts.
const CODEX_SSE_RETRY_PATTERNS = ["server_is_overloaded", "service_unavailable_error"];
@@ -24,6 +25,10 @@ const CODEX_SSE_USER_OUTPUT_PATTERNS = [
];
const CODEX_SSE_PEEK_BYTES = 256 * 1024;
const CODEX_MODEL_CAPACITY_MESSAGE = "Selected model is at capacity. Please try a different model.";
function isCodexResponsesLiteModel(model) {
const baseId = String(model || "").replace(/\([^()]+\)\s*$/, "");
return getProviderModels("cx").some((entry) => entry.id === baseId && entry.responsesLite === true);
}
// Server-generated item id prefixes that Codex /responses cannot resolve when store=false
const SERVER_ID_PATTERN = /^(rs|fc|resp|msg)_/;
@@ -42,7 +47,7 @@ const CODEX_PASSTHROUGH_TOOL_TYPES = new Set(["custom"]);
const RESPONSES_API_ALLOWLIST = new Set([
"model", "input", "instructions", "tools", "tool_choice", "stream", "store",
"reasoning", "service_tier", "include", "prompt_cache_key", "client_metadata",
"text"
"text", "parallel_tool_calls"
]);
// Convert role=system → role=developer in body.input (keeps content in cacheable prefix)
@@ -56,13 +61,14 @@ function convertSystemToDeveloperRole(body) {
}
// Strip server-generated item IDs (rs_/fc_/resp_/msg_) from input — avoids 404 with store=false
function stripStoredItemReferences(body) {
function stripStoredItemReferences(body, preserveLitePrefix = false) {
if (!Array.isArray(body.input)) return;
body.input = body.input.filter((item) => {
if (typeof item === "string" && SERVER_ID_PATTERN.test(item)) return false;
if (item && typeof item === "object" && !Array.isArray(item)) {
if (item.type === "item_reference") return false;
if (typeof item.id === "string" && SERVER_ID_PATTERN.test(item.id)) delete item.id;
if (typeof item.id === "string" && SERVER_ID_PATTERN.test(item.id)
&& !(preserveLitePrefix && item.role === "developer" && item.id.startsWith("msg_"))) delete item.id;
}
return true;
});
@@ -72,6 +78,9 @@ function stripStoredItemReferences(body) {
function normalizeCodexTools(body) {
if (!Array.isArray(body.tools)) return;
const validNames = new Set();
// Codex's schema validator has no Unicode property escapes; a `pattern`
// carrying `\p{...}` 400s the whole request on every account (#3922).
const patternStats = { removed: 0 };
body.tools = body.tools.filter((tool) => {
if (!tool || typeof tool !== "object" || Array.isArray(tool)) return false;
const type = typeof tool.type === "string" ? tool.type : "";
@@ -80,6 +89,9 @@ function normalizeCodexTools(body) {
for (const st of tool.tools) {
const n = typeof st?.name === "string" ? st.name.trim().slice(0, 128) : "";
if (n) validNames.add(n);
if (st?.parameters && typeof st.parameters === "object") {
st.parameters = stripCodexUnsupportedPatterns(st.parameters, patternStats);
}
}
}
return true;
@@ -101,10 +113,13 @@ function normalizeCodexTools(body) {
tool.type = "function";
tool.name = name.slice(0, 128);
if (description) tool.description = description;
tool.parameters = parameters;
tool.parameters = stripCodexUnsupportedPatterns(parameters, patternStats);
validNames.add(name);
return true;
});
if (patternStats.removed > 0) {
dbg("CODEX", `stripped ${patternStats.removed} unsupported tool schema pattern(s)`);
}
// Drop tool_choice if it references an unknown function name
if (body.tool_choice && typeof body.tool_choice === "object" && !Array.isArray(body.tool_choice)) {
if (body.tool_choice.type === "function") {
@@ -128,6 +143,7 @@ function resolveCacheSessionId(body, credentials) {
function normalizeReasoningEffort(model, value) {
const supportedLevels = getThinkingLevels("codex", model);
if (supportedLevels?.includes(value)) return value;
if (isCodexResponsesLiteModel(model) && (value === "none" || value === "minimal")) return "low";
if (value === "ultra" && supportedLevels?.includes("max")) return "max";
if (value === "max" || value === "ultra") return "xhigh";
return value;
@@ -199,8 +215,11 @@ export class CodexExecutor extends BaseExecutor {
* Override headers to add codex-specific identity headers.
* transformRequest runs BEFORE buildHeaders, sets this._currentSessionId.
*/
buildHeaders(credentials, stream = true) {
buildHeaders(credentials, stream = true, _url = null, model = null) {
const headers = super.buildHeaders(credentials, stream);
if (isCodexResponsesLiteModel(model && getModelUpstreamId("cx", model))) {
headers["x-openai-internal-codex-responses-lite"] = "true";
}
headers["session_id"] = this._currentSessionId || credentials?.connectionId || "default";
// Identify client type to Codex backend (matches official codex CLI)
if (!headers["originator"]) headers["originator"] = "codex_cli_rs";
@@ -398,6 +417,8 @@ export class CodexExecutor extends BaseExecutor {
// Convert string input to array format (Codex API requires input as array)
const normalized = normalizeResponsesInput(body.input);
if (normalized) body.input = normalized;
const upstreamModel = getModelUpstreamId("cx", body.model || model);
const responsesLite = isCodexResponsesLiteModel(upstreamModel);
// Ensure input is present and non-empty (Codex API rejects empty input)
if (!body.input || (Array.isArray(body.input) && body.input.length === 0)) {
@@ -407,7 +428,7 @@ export class CodexExecutor extends BaseExecutor {
// Keep system prompts in body.input as role=developer so they stay in the cacheable prefix
convertSystemToDeveloperRole(body);
// Strip server-generated item IDs (rs_/fc_/resp_/msg_) — Codex /responses can't resolve when store=false
stripStoredItemReferences(body);
stripStoredItemReferences(body, responsesLite);
// Flatten function tools + drop unsupported types
normalizeCodexTools(body);
@@ -415,7 +436,7 @@ export class CodexExecutor extends BaseExecutor {
body.stream = true;
// If no instructions provided, inject default Codex instructions
if (!body.instructions || body.instructions.trim() === "") {
if (!responsesLite && (!body.instructions || body.instructions.trim() === "")) {
body.instructions = CODEX_DEFAULT_INSTRUCTIONS;
}
@@ -428,7 +449,29 @@ export class CodexExecutor extends BaseExecutor {
}
// Map virtual Codex review models to the upstream Codex model before suffix parsing.
body.model = getModelUpstreamId("cx", body.model || model);
body.model = upstreamModel;
if (responsesLite) {
// Codex 0.155 carries tools and instructions as input prefix items.
const input = Array.isArray(body.input) ? body.input : [body.input];
const hasLitePrefix = input.some((item) => item?.type === "additional_tools");
if (!hasLitePrefix) {
const instructions = typeof body.instructions === "string" && body.instructions.trim()
? body.instructions : CODEX_DEFAULT_INSTRUCTIONS;
const prefix = [{ type: "additional_tools", role: "developer", tools: Array.isArray(body.tools) ? body.tools : [] }];
if (instructions) {
prefix.push({ type: "message", role: "developer", content: [{ type: "input_text", text: instructions }] });
}
input.unshift(...prefix);
}
body.input = input;
body.instructions = "";
body.tools = null;
body.tool_choice ||= "auto";
body.parallel_tool_calls = false;
} else {
delete body.parallel_tool_calls;
}
// Extract thinking level from model name suffix
// e.g., gpt-5.3-codex-high → high, gpt-5.3-codex → medium (default)
@@ -445,12 +488,13 @@ export class CodexExecutor extends BaseExecutor {
// Priority: explicit reasoning.effort > reasoning_effort param > model suffix > default (medium)
if (!body.reasoning) {
const effort = normalizeReasoningEffort(body.model, body.reasoning_effort || modelEffort || 'low');
body.reasoning = { effort, summary: "auto" };
const effort = normalizeReasoningEffort(body.model, body.reasoning_effort || modelEffort || (responsesLite ? 'medium' : 'low'));
body.reasoning = responsesLite ? { effort } : { effort, summary: "auto" };
} else {
body.reasoning.effort = normalizeReasoningEffort(body.model, body.reasoning.effort);
if (!body.reasoning.summary) body.reasoning.summary = "auto";
if (!responsesLite && !body.reasoning.summary) body.reasoning.summary = "auto";
}
if (responsesLite) body.reasoning.context = "all_turns";
delete body.reasoning_effort;
// Include reasoning encrypted content (required by Codex backend for reasoning models)

View File

@@ -47,10 +47,24 @@ export class CommandCodeExecutor extends BaseExecutor {
}
async execute(opts) {
const result = await super.execute(opts);
if (!result?.response?.ok || !result.response.body) return result;
result.response = await inspectAndWrapCommandCodeResponse(result.response, opts.model);
return result;
const maxRetries = 2;
for (let attempt = 0; attempt <= maxRetries; attempt++) {
const result = await super.execute(opts);
if (!result?.response?.ok || !result.response.body) return result;
const wrappedResponse = await inspectAndWrapCommandCodeResponse(result.response, opts.model);
if (!wrappedResponse.ok && attempt < maxRetries) {
const isRetryableStatus = wrappedResponse.status === 502 || wrappedResponse.status === 503 || wrappedResponse.status === 504;
if (isRetryableStatus) {
opts.log?.debug?.("RETRY", `CommandCode upstream returned status ${wrappedResponse.status}, retrying ${attempt + 1}/${maxRetries}...`);
await new Promise(r => setTimeout(r, 1000 * (attempt + 1)));
continue;
}
}
result.response = wrappedResponse;
return result;
}
}
parseError(response, bodyText) {
@@ -134,7 +148,7 @@ export async function inspectAndWrapCommandCodeResponse(originalResponse, model)
const reader = originalResponse.body.getReader();
const decoder = new TextDecoder();
let buffer = "";
const bufferedLines = [];
const rawChunks = [];
let detectedError = null;
try {
@@ -148,16 +162,15 @@ export async function inspectAndWrapCommandCodeResponse(originalResponse, model)
const parsed = JSON.parse(jsonStr);
if (parsed?.type === "error") {
detectedError = parsed;
} else {
bufferedLines.push(trimmed);
}
} catch {
bufferedLines.push(trimmed);
/* ignore */
}
}
break;
}
rawChunks.push(value);
buffer += decoder.decode(value, { stream: true });
const lines = buffer.split("\n");
buffer = lines.pop() || "";
@@ -168,7 +181,6 @@ export async function inspectAndWrapCommandCodeResponse(originalResponse, model)
if (!trimmed) continue;
const jsonStr = trimmed.startsWith("data:") ? trimmed.slice(5).trim() : trimmed;
if (!jsonStr || jsonStr === "[DONE]") {
bufferedLines.push(trimmed);
stopLoop = true;
break;
}
@@ -177,7 +189,6 @@ export async function inspectAndWrapCommandCodeResponse(originalResponse, model)
try {
event = JSON.parse(jsonStr);
} catch {
bufferedLines.push(trimmed);
continue;
}
@@ -187,8 +198,6 @@ export async function inspectAndWrapCommandCodeResponse(originalResponse, model)
break;
}
bufferedLines.push(trimmed);
if (
event?.type === "text-delta" ||
event?.type === "reasoning-delta" ||
@@ -231,29 +240,18 @@ export async function inspectAndWrapCommandCodeResponse(originalResponse, model)
);
}
const combinedStream = createReplayedStream(bufferedLines, buffer, reader);
const combinedStream = createRawReplayedStream(rawChunks, reader);
return wrapNdjsonAsOpenAISse(combinedStream, model, originalResponse);
}
function createReplayedStream(bufferedLines, remainingBuffer, reader) {
const encoder = new TextEncoder();
let replayed = false;
function createRawReplayedStream(rawChunks, reader) {
let chunkIndex = 0;
return new ReadableStream({
async pull(controller) {
if (!replayed) {
replayed = true;
let prefix = bufferedLines.join("\n");
if (prefix && remainingBuffer) {
prefix += "\n" + remainingBuffer;
} else if (remainingBuffer) {
prefix = remainingBuffer;
} else if (prefix) {
prefix += "\n";
}
if (prefix) {
controller.enqueue(encoder.encode(prefix));
}
if (chunkIndex < rawChunks.length) {
controller.enqueue(rawChunks[chunkIndex++]);
return;
}
try {

View File

@@ -7,13 +7,16 @@ import {
wrapConnectRPCFrame,
decodeMessage,
parseConnectRPCFrame,
extractTextFromResponse
extractTextFromResponse,
encodeMcpTools,
decodeMcpArgs,
} from "../utils/cursorProtobuf.js";
import { buildCursorHeaders } from "../utils/cursorChecksum.js";
import { estimateUsage } from "../utils/usageTracking.js";
import { SSE_DONE, SSE_HEADERS } from "../utils/sseConstants.js";
import { chatChunkSse, sseChunk } from "../utils/sse.js";
import { FORMATS } from "../translator/formats.js";
import { ROLE, OPENAI_BLOCK } from "../translator/schema/index.js";
import { proxyAwareFetch } from "../utils/proxyFetch.js";
import zlib from "zlib";
import crypto from "crypto";
@@ -65,55 +68,74 @@ function textFromContent(content) {
if (typeof content === "string") return content;
if (!Array.isArray(content)) return "";
return content
.filter((part) => part?.type === "text" && typeof part.text === "string")
.filter((part) => part?.type === OPENAI_BLOCK.TEXT && typeof part.text === "string")
.map((part) => part.text)
.join("\n");
}
function isAgentTextRequest(body) {
// Many compatible clients always attach their built-in tool schemas, even
// for a normal text turn. Cursor's retired ChatService rejects those
// requests; AgentService can still answer the text turn, so ignore schemas
// here. A real tool-call/result conversation is kept on the legacy path
// until its AgentService tool protocol is implemented.
return Array.isArray(body?.messages) && body.messages.every((message) => {
if (message?.tool_calls?.length || message?.role === "tool") return false;
return typeof message?.content === "string"
|| Array.isArray(message?.content) && message.content.every((part) => part?.type === "text");
function isTextPart(part) {
return !part || part.type === OPENAI_BLOCK.TEXT || typeof part === "string";
}
export function isAgentCapableRequest(body) {
// ChatService rejects auto/composer and most thinking variants. AgentService
// can answer text turns (including declared tool schemas) and tool-call
// history. Image parts still need the legacy protobuf path.
if (!Array.isArray(body?.messages) || body.messages.length === 0) return false;
return body.messages.every((message) => {
if (Array.isArray(message?.content)) return message.content.every(isTextPart);
return message?.content == null || typeof message.content === "string";
});
}
function encodeHistoryMessage(message) {
const content = textFromContent(message?.content);
if (!content) return null;
const extras = [];
if (message?.role === ROLE.ASSISTANT && message.tool_calls?.length) {
for (const tc of message.tool_calls) {
extras.push(`[tool_call id=${tc.id || ""} name=${tc.function?.name || "tool"} args=${tc.function?.arguments || "{}"}]`);
}
}
if (message?.role === ROLE.TOOL) {
extras.push(`[tool_result id=${message.tool_call_id || ""}]`);
}
const textBody = [content, ...extras].filter(Boolean).join("\n");
if (!textBody) return null;
// ConversationHistoryMessage.user / .assistant -> repeated content -> text.
const text = agentString(1, content);
if (message.role === "assistant") {
const text = agentString(1, textBody);
if (message.role === ROLE.ASSISTANT) {
return agentMessage(2, agentMessage(1, agentMessage(1, text)));
}
return agentMessage(1, agentMessage(1, agentMessage(1, text)));
}
function buildAgentRunFrame(messages, model) {
export function buildAgentRunFrame(messages, model, tools = []) {
// custom_system_prompt (RunRequest field 8) makes AgentService return an
// empty turn. Fold system text into the current user message instead.
const system = messages
.filter((message) => message?.role === "system")
.filter((message) => message?.role === ROLE.SYSTEM)
.map((message) => textFromContent(message.content))
.filter(Boolean)
.join("\n\n");
const chatMessages = messages.filter((message) => message?.role !== "system");
const currentIndex = [...chatMessages].map((message) => message?.role).lastIndexOf("user");
const chatMessages = messages.filter((message) => message?.role !== ROLE.SYSTEM);
const currentIndex = [...chatMessages].map((message) => message?.role).lastIndexOf(ROLE.USER);
const current = currentIndex >= 0 ? chatMessages[currentIndex] : chatMessages.at(-1);
const history = chatMessages
.slice(0, currentIndex >= 0 ? currentIndex : -1)
.map(encodeHistoryMessage)
.filter(Boolean);
const userText = textFromContent(current?.content) || "Continue.";
const rawUser = textFromContent(current?.content) || "Continue.";
const userText = system ? `${system}\n\n${rawUser}` : rawUser;
// agent.v1.UserMessageAction.user_message and its optional history.
// selected_context (3) + mode=1 (4) match cursor-agent's wire format; without
// them the server may accept the RPC and stream an empty turn.
const userMessage = concatBuffers(
agentString(1, userText),
agentString(2, crypto.randomUUID()),
agentMessage(3, new Uint8Array()),
encodeField(4, PROTOBUF_VARINT, 1),
);
const conversationHistory = history.length
? concatBuffers(...history.map((entry) => agentMessage(1, entry)))
@@ -124,11 +146,20 @@ function buildAgentRunFrame(messages, model) {
);
const conversationAction = agentMessage(1, userAction);
const requestedModel = concatBuffers(agentString(1, model), agentBool(7, true));
// ModelDetails (field 3): thinking variants (Composer, Grok, *-thinking)
// return an empty turn when only RequestedModel (field 9) is set.
const modelDetails = concatBuffers(
agentString(1, model),
agentString(3, model),
agentString(4, model),
);
const mcpTools = encodeMcpTools(tools);
const runRequest = concatBuffers(
// An empty ConversationStateStructure starts a fresh local agent session.
agentMessage(1, new Uint8Array()),
agentMessage(2, conversationAction),
...(system ? [agentString(8, system)] : []),
agentMessage(3, modelDetails),
...(mcpTools.length ? [agentMessage(4, mcpTools)] : []),
agentMessage(9, requestedModel),
);
@@ -157,13 +188,51 @@ function decodeAgentFrames(buffer, onFrame) {
return pending;
}
function createRequestContextResponse() {
// AgentService asks every run for client context. 9router has no IDE file
// context, so acknowledge with an empty RequestContext.
function execIds(execRequest) {
const id = Number(execRequest?.get(1)?.[0]?.value || 0);
const execId = extractAgentString(execRequest, 15);
return { id, execId };
}
function wrapExecClientMessage(execMsgId, execId, resultField, resultPayload) {
const parts = [];
if (execMsgId) parts.push(encodeField(1, PROTOBUF_VARINT, execMsgId));
parts.push(agentString(15, execId || ""));
parts.push(encodeField(resultField, PROTOBUF_LEN, resultPayload || new Uint8Array()));
return wrapConnectRPCFrame(agentMessage(2, concatBuffers(...parts)));
}
function createRequestContextResponse(execRequest) {
// Tools already go out on AgentRunRequest.mcp_tools. Echoing them again on
// this ack makes AgentService stall silently (0 SSE bytes until abort).
const { id, execId } = execIds(execRequest);
const requestContextSuccess = agentMessage(1, new Uint8Array());
const requestContextResult = agentMessage(1, requestContextSuccess);
const execClientMessage = agentMessage(10, requestContextResult);
return wrapConnectRPCFrame(agentMessage(2, execClientMessage));
return wrapExecClientMessage(id, execId, 10, requestContextResult);
}
// ExecServerMessage variant → ExecClientMessage result field (same numbers).
const EXEC_RESULT_FIELD = {
2: 2, 3: 3, 4: 4, 5: 5, 7: 7, 8: 8, 9: 9, 16: 16, 20: 20, 23: 23,
};
function rejectExecRequest(execRequest) {
const { id, execId } = execIds(execRequest);
const variant = [...(execRequest?.keys?.() || [])].find((field) => field !== 1 && field !== 15);
const resultField = EXEC_RESULT_FIELD[variant];
if (!resultField) return null;
// Diagnostics has no rejected variant — empty success unblocks the stream.
if (variant === 9) return wrapExecClientMessage(id, execId, 9, new Uint8Array());
const rejected = agentMessage(2, agentString(2, "Tool not available in this environment. Use the MCP tools provided instead."));
return wrapExecClientMessage(id, execId, resultField, rejected);
}
function encodeKvClientMessage(kvId, resultField, resultPayload, metadata) {
const parts = [];
if (kvId) parts.push(encodeField(1, PROTOBUF_VARINT, kvId));
parts.push(encodeField(resultField, PROTOBUF_LEN, resultPayload || new Uint8Array()));
if (metadata && metadata.length) parts.push(encodeField(4, PROTOBUF_LEN, metadata));
return wrapConnectRPCFrame(agentMessage(3, concatBuffers(...parts)));
}
const CURSOR_STREAM_DEBUG = process.env.CURSOR_STREAM_DEBUG === "1";
@@ -479,7 +548,7 @@ export class CursorExecutor extends BaseExecutor {
};
}
async executeAgent({ model, body, stream, credentials, signal }) {
async executeAgent({ model, body, stream, credentials, signal, log }) {
const agentEndpoint = PROVIDER_OAUTH.cursor?.agentEndpoint;
if (!agentEndpoint) throw new Error("Cursor AgentService endpoint is not configured");
@@ -491,9 +560,10 @@ export class CursorExecutor extends BaseExecutor {
}
let session;
const tools = body.tools || [];
try {
session = this.openAgentHttp2Stream(url, headers, requestController.signal);
session.write(buildAgentRunFrame(body.messages || [], model));
session.write(buildAgentRunFrame(body.messages || [], model, tools));
} catch (error) {
throw new Error(`Cursor AgentService request failed: ${error.message}`);
}
@@ -533,8 +603,23 @@ export class CursorExecutor extends BaseExecutor {
// so strict clients such as Claude Code accept the completed stream.
const responseId = `chatcmpl-msg_${Date.now()}`;
const created = Math.floor(Date.now() / 1000);
const composerModel = isComposerModel(model);
let pending = Buffer.alloc(0);
let finished = false;
let thinkingAcc = "";
let emittedVisible = 0;
let emittedText = false;
const flushThinkingFallback = (onEvent) => {
if (emittedText || !thinkingAcc) return;
const fallback = composerModel
? visibleComposerContentFromThinking(thinkingAcc)
: thinkingAcc.trim();
if (fallback) {
emittedText = true;
onEvent({ type: "text", value: fallback });
}
};
const consume = async (onEvent) => {
try {
@@ -553,32 +638,87 @@ export class CursorExecutor extends BaseExecutor {
const update = decodeMessage(serverMessage.get(1)[0].value);
if (update.has(1)) {
const textDelta = extractAgentString(decodeMessage(update.get(1)[0].value), 1);
if (textDelta) onEvent({ type: "text", value: textDelta });
if (textDelta) {
emittedText = true;
onEvent({ type: "text", value: textDelta });
}
}
// Cursor's AgentService emits internal reasoning without the
// cryptographic signature required by Anthropic thinking blocks.
// Forwarding it makes strict Anthropic clients (Claude Code)
// discard or wait on an otherwise complete response. Keep the
// reasoning upstream-only and emit the normal answer text.
// thinking_delta (field 4). Composer (and some Grok variants) put
// the visible answer after </think> here and never send text_delta.
if (update.has(4)) {
const thinkingDelta = extractAgentString(decodeMessage(update.get(4)[0].value), 1);
if (thinkingDelta) {
thinkingAcc += thinkingDelta;
if (composerModel) {
const visible = visibleComposerContentFromThinking(thinkingAcc);
if (visible.length > emittedVisible) {
const deltaContent = visible.slice(emittedVisible);
emittedVisible = visible.length;
emittedText = true;
onEvent({ type: "text", value: deltaContent });
}
}
}
}
// Keep unsigned reasoning upstream-only for Anthropic clients.
if (update.has(14)) {
flushThinkingFallback(onEvent);
finished = true;
onEvent({ type: "done" });
}
}
// KvServerMessage (field 4): get/set blob. Ack so the stream proceeds.
if (serverMessage.has(4)) {
const kv = decodeMessage(serverMessage.get(4)[0].value);
const kvId = kv.get(1)?.[0]?.value || 0;
const metadata = kv.get(4)?.[0]?.value || null;
if (kv.has(2)) {
session.write(encodeKvClientMessage(kvId, 2, agentMessage(1, new Uint8Array()), metadata));
} else if (kv.has(3)) {
session.write(encodeKvClientMessage(kvId, 3, new Uint8Array(), metadata));
}
}
// AgentService requests IDE context before producing a response.
// Return an empty context; 9router is not coupled to an editor.
if (serverMessage.has(2)) {
const execRequest = decodeMessage(serverMessage.get(2)[0].value);
if (execRequest.has(10)) {
session.write(createRequestContextResponse());
log?.info?.("CURSOR", "AgentService request_context ack");
session.write(createRequestContextResponse(execRequest));
} else if (execRequest.has(11)) {
const mcp = decodeMcpArgs(execRequest.get(11)[0].value);
const name = mcp.toolName || mcp.name;
if (name) {
log?.info?.("CURSOR", `AgentService MCP tool_call ${name}`);
finished = true;
onEvent({
type: "tool_call",
value: {
id: mcp.toolCallId || `call_${crypto.randomUUID()}`,
name,
arguments: JSON.stringify(mcp.args || {}),
},
});
onEvent({ type: "done", finishReason: "tool_calls" });
} else {
debugLog(`[CURSOR AGENT] Unsupported exec request fields: ${[...execRequest.keys()].join(",")}`);
finished = true;
onEvent({ type: "error", value: "Cursor AgentService requested an unsupported IDE tool" });
}
} else {
// Every other ExecServerMessage variant is an editor-backed tool
// (shell, read, write, …) that 9router cannot service. Fail the
// turn rather than narrating protocol state as assistant text.
debugLog(`[CURSOR AGENT] Unsupported exec request fields: ${[...execRequest.keys()].join(",")}`);
finished = true;
onEvent({ type: "error", value: "Cursor AgentService requested an unsupported IDE tool" });
// Auto/Composer often probe IDE builtins (shell/read/…). Reject
// them so the model can continue with MCP tools or a text answer
// instead of stalling the h2 stream.
const rejection = rejectExecRequest(execRequest);
if (rejection) {
log?.info?.("CURSOR", `AgentService rejected IDE exec fields=${[...execRequest.keys()].join(",")}`);
session.write(rejection);
} else {
debugLog(`[CURSOR AGENT] Unsupported exec request fields: ${[...execRequest.keys()].join(",")}`);
finished = true;
onEvent({ type: "error", value: "Cursor AgentService requested an unsupported IDE tool" });
}
}
}
});
@@ -586,7 +726,10 @@ export class CursorExecutor extends BaseExecutor {
} finally {
try { session.end(); } catch {}
try { session.close(); } catch {}
if (!finished) onEvent({ type: "done" });
if (!finished) {
flushThinkingFallback(onEvent);
onEvent({ type: "done" });
}
}
};
@@ -594,10 +737,21 @@ export class CursorExecutor extends BaseExecutor {
let content = "";
let reasoning = "";
let agentError = null;
const toolCalls = [];
let finishReason = "stop";
await consume((event) => {
if (event.type === "text") content += event.value;
else if (event.type === "thinking") reasoning += event.value;
else if (event.type === "tool_call") {
toolCalls.push({
id: event.value.id,
type: "function",
function: { name: event.value.name, arguments: event.value.arguments },
});
finishReason = "tool_calls";
}
else if (event.type === "error") agentError = event.value;
else if (event.type === "done" && event.finishReason) finishReason = event.finishReason;
});
if (agentError) {
return {
@@ -611,13 +765,19 @@ export class CursorExecutor extends BaseExecutor {
responseFormat: FORMATS.OPENAI,
};
}
const message = {
role: "assistant",
content: content || null,
...(reasoning ? { reasoning_content: reasoning } : {}),
...(toolCalls.length ? { tool_calls: toolCalls } : {}),
};
return {
response: new Response(JSON.stringify({
id: responseId,
object: "chat.completion",
created,
model,
choices: [{ index: 0, message: { role: "assistant", content: content || null, ...(reasoning ? { reasoning_content: reasoning } : {}) }, finish_reason: "stop" }],
choices: [{ index: 0, message, finish_reason: finishReason }],
usage: estimateUsage(body, content.length, FORMATS.OPENAI),
}), { headers: { "Content-Type": "application/json" } }),
url,
@@ -635,6 +795,18 @@ export class CursorExecutor extends BaseExecutor {
controller.enqueue(encoder.encode(chatChunkSse({ id: responseId, created, model, delta: { content: event.value } })));
} else if (event.type === "thinking") {
controller.enqueue(encoder.encode(chatChunkSse({ id: responseId, created, model, delta: { reasoning_content: event.value } })));
} else if (event.type === "tool_call") {
controller.enqueue(encoder.encode(chatChunkSse({
id: responseId, created, model,
delta: {
tool_calls: [{
index: 0,
id: event.value.id,
type: "function",
function: { name: event.value.name, arguments: event.value.arguments },
}],
},
})));
} else if (event.type === "error") {
// An SSE error frame, not a content delta: a protocol failure must not
// be rendered to the user as the assistant's reply, and downstream
@@ -643,7 +815,10 @@ export class CursorExecutor extends BaseExecutor {
controller.enqueue(encoder.encode(SSE_DONE));
controller.close();
} else if (event.type === "done") {
controller.enqueue(encoder.encode(chatChunkSse({ id: responseId, created, model, delta: {}, finishReason: "stop" })));
controller.enqueue(encoder.encode(chatChunkSse({
id: responseId, created, model, delta: {},
finishReason: event.finishReason || "stop",
})));
controller.enqueue(encoder.encode(SSE_DONE));
controller.close();
}
@@ -664,9 +839,9 @@ export class CursorExecutor extends BaseExecutor {
}
async execute({ model, body, stream, credentials, signal, log, proxyOptions = null }) {
if (isAgentTextRequest(body)) {
if (isAgentCapableRequest(body)) {
try {
return await this.executeAgent({ model, body, stream, credentials, signal });
return await this.executeAgent({ model, body, stream, credentials, signal, log });
} catch (error) {
return {
response: new Response(JSON.stringify({

View File

@@ -1,12 +1,13 @@
import { BaseExecutor } from "./base.js";
import { PROVIDERS, PROVIDER_OAUTH } from "../config/providers.js";
import { ANTHROPIC_API_VERSION, OPENAI_COMPAT_BASE, ANTHROPIC_COMPAT_BASE, selectAnthropicBeta } from "../providers/shared.js";
import { ANTHROPIC_API_VERSION, OPENAI_COMPAT_BASE, ANTHROPIC_COMPAT_BASE, selectAnthropicBeta, mergeAnthropicBeta } from "../providers/shared.js";
import { resolveOpenAICompatibleApiType } from "../services/provider.js";
import { OAUTH_ENDPOINTS, buildKimiHeaders } from "../config/appConstants.js";
import { buildClineHeaders } from "../shared/clineAuth.js";
import { proxyAwareFetch } from "../utils/proxyFetch.js";
import { injectReasoningContent } from "../utils/reasoningContentInjector.js";
import { stripUnsupportedParams } from "../translator/concerns/paramSupport.js";
import { extractClaudeSessionIdFromUserId } from "../utils/claudeCloaking.js";
// Auth header descriptors — derived from registry transport.auth, fallback to hardcoded defaults.
const BEARER = { combined: true, header: "Authorization", scheme: "bearer" };
@@ -146,7 +147,7 @@ export class DefaultExecutor extends BaseExecutor {
return BEARER;
}
buildHeaders(credentials, stream = true, url, model) {
buildHeaders(credentials, stream = true, url, model, body = null) {
const rt = credentials?.runtimeTransport;
const headers = { "Content-Type": "application/json", ...(rt ? rt.headers : this.config.headers) };
const desc = rt?.auth || AUTH_DESCRIPTORS[this.provider] || this.resolveAuthDescriptor();
@@ -164,9 +165,21 @@ export class DefaultExecutor extends BaseExecutor {
// a node fronting Kimi or GLM answers on its own ids and never matches, so
// gateways that would choke on unknown beta flags are left untouched.
const isClaudeModel = typeof model === "string" && /^claude-/.test(model);
const clientBeta = credentials?.rawHeaders?.["anthropic-beta"];
if (model && (this.provider === "claude"
|| (this.provider?.startsWith?.("anthropic-compatible-") && isClaudeModel))) {
headers["Anthropic-Beta"] = selectAnthropicBeta(model);
headers["Anthropic-Beta"] = mergeAnthropicBeta(selectAnthropicBeta(model, body), clientBeta);
} else if (this.provider === "anthropic" && clientBeta) {
headers["Anthropic-Beta"] = mergeAnthropicBeta(headers["Anthropic-Beta"], clientBeta);
}
// Claude OAuth: align x-claude-code-session-id with metadata.user_id.session_id if missing
if (this.provider === "claude" && !headers["x-claude-code-session-id"]) {
const token = credentials?.accessToken || credentials?.apiKey || "";
if (token.includes("sk-ant-oat")) {
const sid = extractClaudeSessionIdFromUserId(body?.metadata?.user_id);
if (sid) headers["x-claude-code-session-id"] = sid;
}
}
// Strip first-party Claude Code identity headers for non-Anthropic anthropic-compatible upstreams

View File

@@ -11,12 +11,14 @@ import { CursorExecutor } from "./cursor.js";
import { VertexExecutor } from "./vertex.js";
import { OpenCodeExecutor } from "./opencode.js";
import { OpenCodeGoExecutor } from "./opencode-go.js";
import { OpenCodeZenExecutor } from "./opencode-zen.js";
import { GrokWebExecutor } from "./grok-web.js";
import { GrokCliExecutor } from "./grok-cli.js";
import { PerplexityWebExecutor } from "./perplexity-web.js";
import { OllamaLocalExecutor } from "./ollama-local.js";
import { CommandCodeExecutor } from "./commandcode.js";
import { XiaomiTokenplanExecutor } from "./xiaomi-tokenplan.js";
import { XiaomiMimoExecutor } from "./xiaomi-mimo.js";
import { MimoFreeExecutor } from "./mimo-free.js";
import { CodeBuddyExecutor } from "./codebuddy-cn.js";
import { CodeBuddyIntlExecutor } from "./codebuddy-intl.js";
@@ -33,6 +35,7 @@ const executors = {
github: new GithubExecutor(),
iflow: new IFlowExecutor(),
qoder: new QoderExecutor(),
"qoder-cn": new QoderExecutor("qoder-cn"),
kiro: new KiroExecutor(),
kimchi: new KimchiExecutor(),
codex: new CodexExecutor(),
@@ -42,6 +45,7 @@ const executors = {
"vertex-partner": new VertexExecutor("vertex-partner"),
opencode: new OpenCodeExecutor(),
"opencode-go": new OpenCodeGoExecutor(),
"opencode-zen": new OpenCodeZenExecutor(),
"grok-web": new GrokWebExecutor(),
"grok-cli": new GrokCliExecutor(),
gcli: new GrokCliExecutor(), // Alias
@@ -50,6 +54,7 @@ const executors = {
"ollama-local": new OllamaLocalExecutor(),
commandcode: new CommandCodeExecutor(),
"xiaomi-tokenplan": new XiaomiTokenplanExecutor(),
"xiaomi-mimo": new XiaomiMimoExecutor(),
"mimo-free": new MimoFreeExecutor(),
mmf: new MimoFreeExecutor(), // Alias for mimo-free
"codebuddy-cn": new CodeBuddyExecutor(),
@@ -87,12 +92,14 @@ export { VertexExecutor } from "./vertex.js";
export { DefaultExecutor } from "./default.js";
export { OpenCodeExecutor } from "./opencode.js";
export { OpenCodeGoExecutor } from "./opencode-go.js";
export { OpenCodeZenExecutor } from "./opencode-zen.js";
export { GrokWebExecutor } from "./grok-web.js";
export { GrokCliExecutor } from "./grok-cli.js";
export { PerplexityWebExecutor } from "./perplexity-web.js";
export { OllamaLocalExecutor } from "./ollama-local.js";
export { CommandCodeExecutor } from "./commandcode.js";
export { XiaomiTokenplanExecutor } from "./xiaomi-tokenplan.js";
export { XiaomiMimoExecutor } from "./xiaomi-mimo.js";
export { MimoFreeExecutor } from "./mimo-free.js";
export { CodeBuddyExecutor } from "./codebuddy-cn.js";
export { CodeBuddyIntlExecutor } from "./codebuddy-intl.js";

View File

@@ -127,12 +127,18 @@ async function readResponsePrefix(response, signal, maxBytes, timeoutMs) {
return decoder.decode(concatChunks(chunks, totalBytes));
}
// The instruction goes into the current user turn, never into a top-level
// `systemPrompt`: kiro.dev answers any body carrying that field with
// 400 REQUEST_BODY_INVALID, so writing it here turned every repair retry into
// a hard failure.
function appendRepairInstruction(body, kind) {
const repaired = structuredClone(body || {});
const instruction = REPAIR_INSTRUCTIONS[kind] || "Retry the previous incomplete Kiro response.";
repaired.systemPrompt = repaired.systemPrompt
? `${repaired.systemPrompt}\n\n${instruction}`
: instruction;
const msg = repaired?.conversationState?.currentMessage?.userInputMessage;
if (msg) {
const content = typeof msg.content === "string" ? msg.content : "";
msg.content = content ? `${content}\n\n${instruction}` : instruction;
}
return repaired;
}
@@ -259,6 +265,19 @@ export class KiroExecutor extends BaseExecutor {
}
}
// CLIRO parity for the Amazon surfaces: the Kiro runtime accepts the
// SSO bearer header + agent-mode marker. Without these the deprecated
// path gateway answers REQUEST_BODY_INVALID for modern payloads.
if (credentials?.accessToken) {
headers["x-amz-sso-bearer"] = credentials.accessToken;
}
headers["x-amzn-kiro-agent-mode"] = "spec";
headers["x-amzn-codewhisperer-machine-id"] = "kiro-desktop";
const profileArn = credentials?.providerSpecificData?.profileArn;
if (profileArn) {
headers["x-amzn-codewhisperer-profile-arn"] = profileArn;
}
return headers;
}
@@ -285,9 +304,13 @@ export class KiroExecutor extends BaseExecutor {
// 403 "bearer token invalid", so they must hit the CodeWhisperer
// *.amazonaws.com surface, and in the region the token was minted in
// (the baseUrls are hardcoded us-east-1).
const isCodeWhispererSurface =
authMethod === "api_key" || authMethod === "external_idp" || authMethod === "idc";
if (!isCodeWhispererSurface) return baseUrls;
// Kiro deprecated the legacy path-style GenerateAssistantResponse on
// runtime.*.kiro.dev (IDE 1.0.228+ moved to POST / + x-amz-target). The
// path gateway now answers valid modern payloads with 400
// REQUEST_BODY_INVALID, and 400 is terminal in BaseExecutor, so kiro.dev
// must never be the first surface for any auth method. Amazon surfaces
// reject foreign tokens with 401/403, which DO fall through, so trying
// q/codewhisperer first is safe for every auth method (CLIRO parity).
const region = (credentials?.providerSpecificData?.region || "us-east-1").trim();
const regionalize = (u) =>
@@ -297,20 +320,17 @@ export class KiroExecutor extends BaseExecutor {
const amazon = baseUrls.filter((u) => u.includes("amazonaws.com")).map(regionalize);
const others = baseUrls.filter((u) => !u.includes("amazonaws.com"));
if (authMethod === "api_key") {
const q = amazon.filter((u) => u.includes("://q."));
const remaining = amazon.filter((u) => !u.includes("://q."));
return q.length > 0
? [...q, ...remaining, ...others]
: [...amazon, ...others];
}
return amazon.length > 0 ? [...amazon, ...others] : baseUrls;
const q = amazon.filter((u) => u.includes("://q."));
const remaining = amazon.filter((u) => !u.includes("://q."));
return q.length > 0
? [...q, ...remaining, ...others]
: [...amazon, ...others];
}
buildUrl(model, stream, urlIndex = 0, credentials = null) {
const baseUrls = this.getOrderedBaseUrls(credentials);
return baseUrls[urlIndex] || baseUrls[0] || this.config.baseUrl;
const url = baseUrls[urlIndex] || baseUrls[0] || this.config.baseUrl;
return url;
}
// Retry only endpoint/auth-surface failures. Payload-invalid HTTP 400 must be

View File

@@ -1,7 +1,8 @@
import crypto from "node:crypto";
import { DefaultExecutor } from "./default.js";
import { resolveSessionId } from "../utils/sessionManager.js";
import { isMuseSparkModel } from "../providers/models/helpers.js";
import { getModelTargetFormat } from "../config/providerModels.js";
import { FORMATS } from "../translator/formats.js";
import {
normalizeResponsesInput,
clampResponsesCallId,
@@ -40,13 +41,10 @@ function translatedSession(sessionId, clientTool) {
return `ses_${digest}`;
}
// Strip the thinking suffix "model(level)" so checks hit the base id.
function baseModelId(model) {
return String(model || "").replace(/\([^()]+\)\s*$/, "").trim();
}
// Responses-only per the provider registry (grok-4.6, gpt-5.6-luna, muse-spark, …),
// including the family-regex fallback for passthrough ids — never hardcode model ids here.
function isResponsesModel(model) {
return isMuseSparkModel(baseModelId(model));
return getModelTargetFormat("opencode-go", model) === FORMATS.OPENAI_RESPONSES;
}
// Flatten Chat Completions tool declarations into the Responses flat shape and
@@ -90,6 +88,12 @@ function sanitizeResponsesItems(body) {
if (!Array.isArray(body.input)) return;
body.input = body.input.filter((item) => {
if (!item || typeof item !== "object" || Array.isArray(item)) return true;
// Strip prior-turn reasoning items: Muse Spark contributor models route to
// an upstream Console backend where encrypted_content cannot be validated across
// rotated accounts or sessions, causing 400 "reasoning encrypted_content was not issued to this caller".
if (item.type === "reasoning") return false;
delete item.encrypted_content;
delete item.reasoning_encrypted_content;
if (item.type === "function_call") {
if (!item.name || typeof item.name !== "string" || item.name.trim() === "") return false;
item.name = item.name.trim().slice(0, MAX_TOOL_NAME_LEN);

View File

@@ -0,0 +1,315 @@
import crypto from "node:crypto";
import { DefaultExecutor } from "./default.js";
import { resolveSessionId } from "../utils/sessionManager.js";
import { isMuseSparkModel } from "../providers/models/helpers.js";
import {
normalizeResponsesInput,
clampResponsesCallId,
coerceResponsesArguments,
coerceResponsesOutput,
} from "../translator/formats/responsesApi.js";
const SESSION_HEADER = "x-opencode-session";
const SESSION_FIELD = "_opencodeZenSession";
const MAX_SESSION_LENGTH = 256;
const RESPONSES_BASE_URL = "https://opencode.ai/zen/v1/responses";
const MAX_TOOL_NAME_LEN = 128;
const OPENCODE_UA = "opencode/1.18.31";
export const OPENCODE_SESSION_RE = /^ses_[0-9a-f]{12}[0-9A-Za-z]{14}$/;
const BASE62_CHARS = "0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz";
// Free-tier fingerprint (mirrors opencode executor, PR #4132): upstream 403s
// requests without the file-search quartet and without stream:true.
const OPENCODE_FINGERPRINT_TOOLS = ["bash", "glob", "grep", "read"];
function hasValidOpencodeVersion(ua) {
const m = String(ua || "").match(/opencode\/(\d+)\.(\d+)(?:\.(\d+))?/i);
if (!m) return false;
const major = parseInt(m[1], 10);
const minor = parseInt(m[2], 10);
return major > 1 || (major === 1 && minor >= 17);
}
function unstableRandom() {
const bytes = crypto.randomBytes(14);
let randomPart = "";
for (let i = 0; i < 14; i++) {
randomPart += BASE62_CHARS[bytes[i] % 62];
}
return randomPart;
}
export function generateSessionId(timestamp = Date.now()) {
const current = BigInt(timestamp) * 0x1000n + 1n;
const value = ~current;
const time = Array.from({ length: 6 }, (_, index) =>
Number((value >> BigInt(40 - 8 * index)) & 0xffn)
.toString(16)
.padStart(2, "0")
).join("");
return `ses_${time}${unstableRandom()}`;
}
export function generateRequestId(timestamp = Date.now()) {
const current = BigInt(timestamp) * 0x1000n + 1n;
const value = current;
const time = Array.from({ length: 6 }, (_, index) =>
Number((value >> BigInt(40 - 8 * index)) & 0xffn)
.toString(16)
.padStart(2, "0")
).join("");
return `msg_${time}${unstableRandom()}`;
}
export function translateSessionId(sessionId, clientTool = "") {
if (typeof sessionId === "string" && OPENCODE_SESSION_RE.test(sessionId.trim())) {
return sessionId.trim();
}
const digest = crypto
.createHash("sha256")
.update(`opencode\0${clientTool || "generic"}\0${sessionId || ""}`)
.digest();
const timeHex = digest.subarray(0, 6).toString("hex");
let randomPart = "";
for (let i = 6; i < 20; i++) {
randomPart += BASE62_CHARS[digest[i] % 62];
}
return `ses_${timeHex}${randomPart}`;
}
function toolNameOf(tool) {
if (!tool || typeof tool !== "object" || Array.isArray(tool)) return "";
const fn = tool.function && typeof tool.function === "object" && !Array.isArray(tool.function) ? tool.function : null;
const raw = typeof tool.name === "string" ? tool.name : (typeof fn?.name === "string" ? fn.name : "");
return raw.trim();
}
function ensureChatFingerprintTools(body) {
if (!body || typeof body !== "object") return;
const present = new Set();
if (Array.isArray(body.tools)) {
for (const tool of body.tools) {
const name = toolNameOf(tool);
if (name) present.add(name);
}
} else {
body.tools = [];
}
for (const name of OPENCODE_FINGERPRINT_TOOLS) {
if (present.has(name)) continue;
body.tools.push({
type: "function",
function: {
name,
description: `OpenCode built-in ${name} tool`,
parameters: { type: "object", properties: {} },
},
});
present.add(name);
}
}
function ensureResponsesFingerprintTools(body) {
if (!body || typeof body !== "object") return;
const present = new Set();
if (Array.isArray(body.tools)) {
for (const tool of body.tools) {
const name = toolNameOf(tool);
if (name) present.add(name);
}
} else {
body.tools = [];
}
for (const name of OPENCODE_FINGERPRINT_TOOLS) {
if (present.has(name)) continue;
body.tools.push({
type: "function",
name,
description: `OpenCode built-in ${name} tool`,
parameters: { type: "object", properties: {} },
});
present.add(name);
}
}
function normalizeSession(value) {
if (typeof value !== "string") return null;
const normalized = value.trim();
if (!normalized || normalized.length > MAX_SESSION_LENGTH) return null;
return normalized;
}
function nativeSession(headers) {
if (!headers || typeof headers !== "object") return null;
for (const [key, value] of Object.entries(headers)) {
if (key.toLowerCase() === SESSION_HEADER) {
const normalized = normalizeSession(value);
if (normalized && OPENCODE_SESSION_RE.test(normalized)) return normalized;
}
}
return null;
}
function translatedSession(sessionId, clientTool) {
return translateSessionId(sessionId, clientTool);
}
// Strip the thinking suffix "model(level)" so checks hit the base id.
function baseModelId(model) {
return String(model || "").replace(/\([^()]+\)\s*$/, "").trim();
}
function isResponsesModel(model) {
return isMuseSparkModel(baseModelId(model));
}
// Flatten Chat Completions tool declarations into the Responses flat shape and
// drop hosted/nameless tools the /responses endpoint rejects.
function normalizeResponsesTools(body) {
if (!Array.isArray(body.tools)) return;
const validNames = new Set();
body.tools = body.tools.filter((tool) => {
if (!tool || typeof tool !== "object" || Array.isArray(tool)) return false;
const fn = tool.function && typeof tool.function === "object" && !Array.isArray(tool.function) ? tool.function : null;
const rawName = typeof tool.name === "string" ? tool.name : (typeof fn?.name === "string" ? fn.name : "");
const name = rawName.trim();
if (!name) return false;
const description = typeof tool.description === "string" ? tool.description : (typeof fn?.description === "string" ? fn.description : "");
let parameters = (tool.parameters && typeof tool.parameters === "object" && !Array.isArray(tool.parameters))
? tool.parameters
: (fn?.parameters && typeof fn.parameters === "object" && !Array.isArray(fn.parameters) ? fn.parameters : { type: "object", properties: {} });
// Mirror the request translator: {type:"object"} without properties is rejected
// by strict Responses backends, so fill in the empty properties map.
if (parameters.type === "object" && !parameters.properties) parameters = { ...parameters, properties: {} };
for (const k of Object.keys(tool)) delete tool[k];
tool.type = "function";
tool.name = name.slice(0, MAX_TOOL_NAME_LEN);
if (description) tool.description = description;
tool.parameters = parameters;
validNames.add(tool.name);
return true;
});
if (body.tool_choice && typeof body.tool_choice === "object" && !Array.isArray(body.tool_choice)) {
if (body.tool_choice.type === "function") {
const n = typeof body.tool_choice.name === "string" ? body.tool_choice.name.trim() : "";
if (!n || !validNames.has(n)) delete body.tool_choice;
}
}
}
// Last line of defense for native Responses clients (sourceFormat === targetFormat
// skips translation): coerce items in place so malformed tool payloads 400 here
// with a clear shape instead of upstream as InputValidationError.
function sanitizeResponsesItems(body) {
if (!Array.isArray(body.input)) return;
body.input = body.input.filter((item) => {
if (!item || typeof item !== "object" || Array.isArray(item)) return true;
// Strip prior-turn reasoning items: Muse Spark contributor models route to
// an upstream Console backend where encrypted_content cannot be validated across
// rotated accounts or sessions, causing 400 "reasoning encrypted_content was not issued to this caller".
if (item.type === "reasoning") return false;
delete item.encrypted_content;
delete item.reasoning_encrypted_content;
if (item.type === "function_call") {
if (!item.name || typeof item.name !== "string" || item.name.trim() === "") return false;
item.name = item.name.trim().slice(0, MAX_TOOL_NAME_LEN);
item.call_id = clampResponsesCallId(item.call_id);
item.arguments = coerceResponsesArguments(item.arguments);
return true;
}
if (item.type === "function_call_output") {
item.call_id = clampResponsesCallId(item.call_id);
item.output = coerceResponsesOutput(item.output);
return true;
}
return true;
});
}
export class OpenCodeZenExecutor extends DefaultExecutor {
constructor() {
super("opencode-zen");
}
buildUrl(model, stream, urlIndex = 0, credentials = null) {
// Muse Spark lives on /responses even when a stale runtimeTransport leaks in.
if (isResponsesModel(model)) return RESPONSES_BASE_URL;
return super.buildUrl(model, stream, urlIndex, credentials);
}
prepareRequestCredentials({ body, credentials, providerSessionId, clientTool } = {}) {
const sourceCredentials = credentials || {};
const native = nativeSession(sourceCredentials.rawHeaders);
const resolved = normalizeSession(providerSessionId) || resolveSessionId({
headers: sourceCredentials.rawHeaders,
body,
connectionId: sourceCredentials.connectionId,
scope: "opencode-zen",
});
return {
...sourceCredentials,
[SESSION_FIELD]: native || translatedSession(resolved, clientTool),
};
}
async execute(args) {
const credentials = this.prepareRequestCredentials(args);
return super.execute({ ...args, credentials });
}
buildHeaders(credentials, stream = true, url, model) {
const headers = super.buildHeaders(credentials || {}, stream, url, model);
const raw = credentials?.rawHeaders || {};
const lower = {};
for (const [k, v] of Object.entries(raw)) lower[k.toLowerCase()] = v;
const downstreamUa = lower["user-agent"] || "";
// Free-tier gate: spoof the official client UA.
headers["User-Agent"] = hasValidOpencodeVersion(downstreamUa) ? downstreamUa : OPENCODE_UA;
headers["x-opencode-client"] = lower["x-opencode-client"] || "desktop";
const prepared = credentials?.[SESSION_FIELD];
if (prepared) {
headers[SESSION_HEADER] = prepared;
return headers;
}
const fallback = this.prepareRequestCredentials({ credentials });
headers[SESSION_HEADER] = fallback[SESSION_FIELD];
return headers;
}
transformRequest(model, body, stream, credentials) {
const out = super.transformRequest(model, body);
// Free-tier gate: upstream 403s stream:false even when everything else is valid.
if (out && typeof out === "object") out.stream = true;
if (!isResponsesModel(model || body?.model)) {
ensureChatFingerprintTools(out);
return out;
}
const normalized = normalizeResponsesInput(out.input);
if (normalized) out.input = normalized;
if (!Array.isArray(out.input) || out.input.length === 0) {
out.input = [{ type: "message", role: "user", content: [{ type: "input_text", text: "..." }] }];
}
// Responses names the output cap max_output_tokens, not max_tokens.
if (out.max_output_tokens === undefined) {
if (out.max_completion_tokens !== undefined) out.max_output_tokens = out.max_completion_tokens;
else if (out.max_tokens !== undefined) out.max_output_tokens = out.max_tokens;
}
delete out.max_tokens;
delete out.max_completion_tokens;
if (out.reasoning_effort !== undefined && out.reasoning === undefined) {
out.reasoning = { effort: out.reasoning_effort, summary: "auto" };
}
if (out.reasoning && typeof out.reasoning === "object" && !Array.isArray(out.reasoning)) {
if (!out.reasoning.summary) out.reasoning.summary = "auto";
}
delete out.reasoning_effort;
out.stream = true;
out.store = false;
ensureResponsesFingerprintTools(out);
normalizeResponsesTools(out);
sanitizeResponsesItems(out);
return out;
}
}

View File

@@ -1,24 +1,250 @@
import crypto from "crypto";
import { BaseExecutor } from "./base.js";
import { PROVIDERS } from "../config/providers.js";
import { MEMORY_CONFIG } from "../config/runtimeConfig.js";
import { getThinkingLevels } from "../providers/thinkingLevels.js";
import { injectReasoningContent } from "../utils/reasoningContentInjector.js";
import { resolveSessionId } from "../utils/sessionManager.js";
import { isMuseSparkModel } from "../providers/models/helpers.js";
import { applyFingerprintTools } from "../utils/opencodeFingerprint.js";
import { ANTHROPIC_API_VERSION } from "../providers/shared.js";
import {
normalizeResponsesInput,
clampResponsesCallId,
coerceResponsesArguments,
coerceResponsesOutput,
} from "../translator/formats/responsesApi.js";
const OPENCODE_UA = "opencode";
const OPENCODE_UA = "opencode/1.18.31";
const MAX_SESSION_LENGTH = 256;
const MAX_TOOL_NAME_LEN = 128;
const SESSION_HEADER = "x-opencode-session";
const SESSION_FIELD = "_opencodeSession";
const REQ_FIELD = "_opencodeRequest";
export const OPENCODE_SESSION_RE = /^ses_[0-9a-f]{12}[0-9A-Za-z]{14}$/;
export const OPENCODE_REQUEST_RE = /^msg_[0-9a-f]{12}[0-9A-Za-z]{14}$/;
const BASE62_CHARS = "0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz";
function hasValidOpencodeVersion(ua) {
const m = String(ua || "").match(/opencode\/(\d+)\.(\d+)(?:\.(\d+))?/i);
if (!m) return false;
const major = parseInt(m[1], 10);
const minor = parseInt(m[2], 10);
return major > 1 || (major === 1 && minor >= 17);
}
// Models served by /zen/v1/responses; every other model stays on /chat/completions.
const RESPONSES_MODELS = new Set([
"muse-spark-1.2-contributor-free",
"muse-spark-1.3-contributor-free",
]);
const MESSAGES_MODELS = new Set(["union-alpha"]);
function generateRequestId() {
return `msg_${crypto.randomUUID().replace(/-/g, "")}`;
let lastTimestamp = 0;
let counter = 0;
function unstableRandom() {
const bytes = crypto.randomBytes(14);
let randomPart = "";
for (let i = 0; i < 14; i++) {
randomPart += BASE62_CHARS[bytes[i] % 62];
}
return randomPart;
}
function generateSessionId() {
return `ses_${crypto.randomUUID().replace(/-/g, "")}`;
export function generateSessionId(timestamp = Date.now()) {
if (timestamp !== lastTimestamp) {
lastTimestamp = timestamp;
counter = 0;
}
counter++;
const current = BigInt(timestamp) * 0x1000n + BigInt(counter);
const value = ~current;
const time = Array.from({ length: 6 }, (_, index) =>
Number((value >> BigInt(40 - 8 * index)) & 0xffn)
.toString(16)
.padStart(2, "0")
).join("");
return `ses_${time}${unstableRandom()}`;
}
export function generateRequestId(timestamp = Date.now()) {
const current = BigInt(timestamp) * 0x1000n + 1n;
const value = current;
const time = Array.from({ length: 6 }, (_, index) =>
Number((value >> BigInt(40 - 8 * index)) & 0xffn)
.toString(16)
.padStart(2, "0")
).join("");
return `msg_${time}${unstableRandom()}`;
}
export function translateSessionId(sessionId, clientTool = "") {
if (typeof sessionId === "string" && OPENCODE_SESSION_RE.test(sessionId.trim())) {
return sessionId.trim();
}
const digest = crypto
.createHash("sha256")
.update(`opencode\0${clientTool || "generic"}\0${sessionId || ""}`)
.digest();
const timeHex = digest.subarray(0, 6).toString("hex");
let randomPart = "";
for (let i = 6; i < 20; i++) {
randomPart += BASE62_CHARS[digest[i] % 62];
}
return `ses_${timeHex}${randomPart}`;
}
function normalizeSession(value) {
if (typeof value !== "string") return null;
const normalized = value.trim();
if (!normalized || normalized.length > MAX_SESSION_LENGTH) return null;
return normalized;
}
function nativeSession(headers) {
if (!headers || typeof headers !== "object") return null;
for (const [key, value] of Object.entries(headers)) {
if (key.toLowerCase() === SESSION_HEADER) {
const normalized = normalizeSession(value);
if (normalized && OPENCODE_SESSION_RE.test(normalized)) return normalized;
}
}
return null;
}
// Upstream free-tier quota is accounted per session. Minting a fresh
// x-opencode-session on every request burns through it and surfaces as
// 429 FreeUsageLimitError with growing reset-after delays, while the real
// CLI reuses one long-lived canonical session per conversation. Mirror
// that: one stable canonical session per downstream identity, evicted
// after MEMORY_CONFIG.sessionTtlMs like the other session stores.
const stableOpencodeSessions = new Map();
const MAX_STABLE_SESSIONS = 1000;
const stableSessionCleanup = setInterval(() => {
const now = Date.now();
for (const [key, entry] of stableOpencodeSessions) {
if (now - entry.lastUsed > MEMORY_CONFIG.sessionTtlMs) {
stableOpencodeSessions.delete(key);
}
}
}, MEMORY_CONFIG.sessionCleanupIntervalMs);
if (stableSessionCleanup.unref) stableSessionCleanup.unref();
function identityKey(credentials) {
const connectionId = credentials?.connectionId || credentials?.id;
if (connectionId) return `opencode:conn:${String(connectionId).slice(0, 128)}`;
const raw = credentials?.rawHeaders || {};
const auth = raw.authorization || raw.Authorization || raw["x-api-key"] || raw["X-Api-Key"] || "";
if (auth) {
const digest = crypto.createHash("sha256").update(String(auth)).digest("hex").slice(0, 32);
return `opencode:auth:${digest}`;
}
return "opencode:default";
}
export function stableSessionId(credentials) {
const key = identityKey(credentials);
const existing = stableOpencodeSessions.get(key);
if (existing) {
existing.lastUsed = Date.now();
stableOpencodeSessions.delete(key);
stableOpencodeSessions.set(key, existing);
return existing.sessionId;
}
const sessionId = generateSessionId();
if (stableOpencodeSessions.size >= MAX_STABLE_SESSIONS) {
stableOpencodeSessions.delete(stableOpencodeSessions.keys().next().value);
}
stableOpencodeSessions.set(key, { sessionId, lastUsed: Date.now() });
return sessionId;
}
function lastUserText(body) {
try {
if (!body || typeof body !== "object") return "";
const arr = Array.isArray(body.messages)
? body.messages
: Array.isArray(body.input)
? body.input
: null;
if (!arr) return typeof body.input === "string" ? body.input.slice(-600) : "";
for (let i = arr.length - 1; i >= 0; i--) {
const msg = arr[i];
if (!msg) continue;
if (msg.role && msg.role !== "user") continue;
const content = msg.content;
if (typeof content === "string" && content.trim()) return content.trim().slice(-600);
if (Array.isArray(content)) {
const text = content
.map((part) => (typeof part === "string" ? part : part?.text || part?.input_text || ""))
.join(" ")
.trim();
if (text) return text.slice(-600);
}
}
} catch {
return "";
}
return "";
}
// The real CLI sends the current user message id (stable per turn, same on
// retries) as x-opencode-request. Derive it deterministically from the
// session plus the last user message so retries share the id.
export function deriveRequestId(sessionId, body) {
const text = lastUserText(body);
if (!text) return generateRequestId();
const digest = crypto
.createHash("sha256")
.update(`opencode-req\0${sessionId || ""}\0${text}`)
.digest();
const timeHex = digest.subarray(0, 6).toString("hex");
let randomPart = "";
for (let i = 6; i < 20; i++) {
randomPart += BASE62_CHARS[digest[i] % 62];
}
const id = `msg_${timeHex}${randomPart}`;
return OPENCODE_REQUEST_RE.test(id) ? id : generateRequestId();
}
function normalizeRequestId(value) {
if (typeof value !== "string") return null;
const normalized = value.trim();
if (!normalized || normalized.length > MAX_SESSION_LENGTH) return null;
return OPENCODE_REQUEST_RE.test(normalized) ? normalized : null;
}
function bodyHasSessionHints(body) {
try {
if (!body || typeof body !== "object") return false;
if (typeof body.session_id === "string" && body.session_id.trim()) return true;
if (typeof body.conversation_id === "string" && body.conversation_id.trim()) return true;
if (typeof body.prompt_cache_key === "string" && body.prompt_cache_key.trim()) return true;
if (body.metadata && typeof body.metadata.user_id === "string" && body.metadata.user_id.trim()) return true;
if (body.request && body.request.sessionId != null && String(body.request.sessionId) !== "") return true;
const arr = Array.isArray(body.messages)
? body.messages
: Array.isArray(body.input)
? body.input
: null;
if (arr) {
let assistantText = "";
for (const msg of arr) {
if (msg?.role === "assistant") {
const content = msg.content;
if (typeof content === "string") assistantText += content;
else if (Array.isArray(content)) {
for (const part of content) assistantText += part?.text || part?.output || "";
}
if (assistantText.length >= 50) return true;
}
}
}
return false;
} catch {
return false;
}
}
// Strip the thinking suffix "model(level)" so registry lookups hit the base id.
@@ -31,14 +257,116 @@ function isResponsesModel(model) {
return RESPONSES_MODELS.has(base) || isMuseSparkModel(base);
}
function resolveOpencodeSession(body, credentials) {
function isMessagesModel(model) {
return MESSAGES_MODELS.has(baseModelId(model));
}
function resolveOpencodeSession(body, credentials, providerSessionId, clientTool) {
const headers = credentials?.rawHeaders || {};
return resolveSessionId({
headers,
body,
connectionId: credentials?.connectionId,
scope: "opencode",
generate: generateSessionId,
const native = nativeSession(headers);
if (native) return native;
let incoming = null;
for (const [key, value] of Object.entries(headers)) {
if (key.toLowerCase() === SESSION_HEADER) {
incoming = normalizeSession(value);
break;
}
}
const hinted = incoming || normalizeSession(providerSessionId);
if (hinted) return translateSessionId(hinted, clientTool);
if (credentials?.connectionId || bodyHasSessionHints(body)) {
let viaManager = null;
try {
viaManager = resolveSessionId({
headers,
body,
connectionId: credentials?.connectionId,
scope: "opencode",
});
} catch {
viaManager = null;
}
if (viaManager) return translateSessionId(viaManager, clientTool);
}
return stableSessionId(credentials);
}
function resolveOpencodeRequestId(body, credentials, sessionId) {
const raw = credentials?.rawHeaders || {};
for (const [key, value] of Object.entries(raw)) {
if (key.toLowerCase() === "x-opencode-request") {
const normalized = normalizeRequestId(value);
if (normalized) return normalized;
break;
}
}
return deriveRequestId(sessionId, body);
}
function normalizeResponsesTools(body) {
if (!Array.isArray(body.tools)) return;
const validNames = new Set();
body.tools = body.tools.filter((tool) => {
if (!tool || typeof tool !== "object" || Array.isArray(tool)) return false;
const fn = tool.function && typeof tool.function === "object" && !Array.isArray(tool.function) ? tool.function : null;
const rawName = typeof tool.name === "string" ? tool.name : (typeof fn?.name === "string" ? fn.name : "");
const name = rawName.trim();
if (!name) return false;
const description = typeof tool.description === "string" ? tool.description : (typeof fn?.description === "string" ? fn.description : "");
let parameters = (tool.parameters && typeof tool.parameters === "object" && !Array.isArray(tool.parameters))
? tool.parameters
: (fn?.parameters && typeof fn.parameters === "object" && !Array.isArray(fn.parameters) ? fn.parameters : { type: "object", properties: {} });
if (parameters.type === "object" && !parameters.properties) parameters = { ...parameters, properties: {} };
for (const k of Object.keys(tool)) delete tool[k];
tool.type = "function";
tool.name = name.slice(0, MAX_TOOL_NAME_LEN);
if (description) tool.description = description;
tool.parameters = parameters;
validNames.add(tool.name);
return true;
});
if (body.tool_choice && typeof body.tool_choice === "object" && !Array.isArray(body.tool_choice)) {
if (body.tool_choice.type === "function") {
const n = typeof body.tool_choice.name === "string" ? body.tool_choice.name.trim() : "";
if (!n || !validNames.has(n)) delete body.tool_choice;
}
}
}
function sanitizeResponsesItems(body) {
if (!Array.isArray(body.input)) return;
body.input = body.input.filter((item) => {
if (!item || typeof item !== "object" || Array.isArray(item)) return true;
// Strip prior-turn reasoning items: OpenCode Free uses public/pooled credentials
// (`Bearer public`) routing to an upstream OpenAI/Console account pool.
// OpenAI Responses API strictly enforces that reasoning `encrypted_content`
// can only be decrypted by the exact caller/account that issued it; sending it
// across different accounts or rotating proxy relays triggers:
// [invalid_request_error] reasoning `encrypted_content` was not issued to this caller (400).
// Furthermore, under stateless mode (store=false), omitting encrypted_content
// causes OpenAI to reject the referenced reasoning item as "not found or was deleted".
// Dropping prior reasoning items allows multi-turn conversations and tool-calling
// loops to succeed cleanly.
if (item.type === "reasoning") return false;
delete item.encrypted_content;
delete item.reasoning_encrypted_content;
if (item.type === "function_call") {
if (!item.name || typeof item.name !== "string" || item.name.trim() === "") return false;
item.name = item.name.trim().slice(0, MAX_TOOL_NAME_LEN);
item.call_id = clampResponsesCallId(item.call_id);
item.arguments = coerceResponsesArguments(item.arguments);
return true;
}
if (item.type === "function_call_output") {
item.call_id = clampResponsesCallId(item.call_id);
item.output = coerceResponsesOutput(item.output);
return true;
}
return true;
});
}
@@ -68,12 +396,35 @@ function normalizeOpencodeReasoning(model, body) {
export class OpenCodeExecutor extends BaseExecutor {
constructor() {
super("opencode", PROVIDERS.opencode);
this._currentSessionId = null;
}
prepareRequestCredentials({ body, credentials, providerSessionId, clientTool } = {}) {
const sourceCredentials = credentials || {};
const session = resolveOpencodeSession(body, sourceCredentials, providerSessionId, clientTool);
return {
...sourceCredentials,
[SESSION_FIELD]: session,
[REQ_FIELD]: resolveOpencodeRequestId(body, sourceCredentials, session),
};
}
transformRequest(model, body, stream, credentials) {
this._currentSessionId = resolveOpencodeSession(body, credentials);
if (isResponsesModel(model)) {
if (body && typeof body === "object" && model && !body.model) body.model = model;
// Zen rejects non-streaming requests on free models with 403 FreeTierError;
// always stream upstream and let the handler layer aggregate for non-stream clients.
if (body && typeof body === "object") body.stream = true;
if (isResponsesModel(model || body?.model) && body && typeof body === "object") {
// ponytail: chỉ model đã xác nhận auto-only; mở allowlist khi có bằng chứng.
if ("tool_choice" in body && body.tool_choice !== "auto"
&& this.config.quirks?.forceAutoToolChoiceModels?.includes(baseModelId(model))) {
body.tool_choice = "auto";
}
const normalized = normalizeResponsesInput(body.input);
if (normalized) body.input = normalized;
if (!Array.isArray(body.input) || body.input.length === 0) {
body.input = [{ type: "message", role: "user", content: [{ type: "input_text", text: "..." }] }];
}
// Responses API names the output cap max_output_tokens and takes thinking
// as reasoning:{effort,summary} — normalize the Chat fields at this boundary.
if (body.max_output_tokens === undefined) {
@@ -83,34 +434,54 @@ export class OpenCodeExecutor extends BaseExecutor {
delete body.max_tokens;
delete body.max_completion_tokens;
normalizeOpencodeReasoning(model, body);
body.stream = true;
body.store = false;
normalizeResponsesTools(body);
sanitizeResponsesItems(body);
// Free-tier fingerprint tools are required even when an agent client
// already supplied tools. ZCode/Claude Code requests normally have
// non-empty tool arrays; skipping cloaking here triggers 403 FreeTierError.
applyFingerprintTools(body, true);
} else if (body && typeof body === "object") {
applyFingerprintTools(body, false);
}
return injectReasoningContent({ provider: this.provider, model, body });
}
buildUrl(model) {
const base = this.config.baseUrl;
return isResponsesModel(model)
? `${base}/zen/v1/responses`
: `${base}/zen/v1/chat/completions`;
async execute(args) {
return super.execute({ ...args, credentials: this.prepareRequestCredentials(args) });
}
buildHeaders(credentials, stream = true) {
buildUrl(model) {
const base = this.config.baseUrl;
if (isResponsesModel(model)) return `${base}/zen/v1/responses`;
if (isMessagesModel(model)) return `${base}/zen/v1/messages`;
return `${base}/zen/v1/chat/completions`;
}
buildHeaders(credentials, stream = true, url = "") {
const raw = credentials?.rawHeaders || {};
const lower = {};
for (const [k, v] of Object.entries(raw)) lower[k.toLowerCase()] = v;
const downstreamUa = lower["user-agent"] || "";
const isOpencodeDownstream = downstreamUa.toLowerCase().includes("opencode");
const isOpencodeDownstream = hasValidOpencodeVersion(downstreamUa);
return {
const session = credentials?.[SESSION_FIELD] || this.prepareRequestCredentials({ credentials })[SESSION_FIELD];
const downstreamReq = normalizeRequestId(lower["x-opencode-request"]);
const requestId = credentials?.[REQ_FIELD] || downstreamReq || generateRequestId();
const headers = {
"Content-Type": "application/json",
"Authorization": "Bearer public",
"User-Agent": isOpencodeDownstream ? downstreamUa : OPENCODE_UA,
"x-opencode-client": lower["x-opencode-client"] || "desktop",
"x-opencode-session": lower["x-opencode-session"] || this._currentSessionId || generateSessionId(),
"x-opencode-request": lower["x-opencode-request"] || generateRequestId(),
"x-opencode-session": session,
"x-opencode-request": requestId,
"x-opencode-project": lower["x-opencode-project"] || "global",
"Accept": stream ? "text/event-stream" : "*/*",
};
if (url.endsWith("/messages")) headers["anthropic-version"] = ANTHROPIC_API_VERSION;
return headers;
}
}

View File

@@ -29,17 +29,19 @@ import { BaseExecutor } from "./base.js";
import { PROVIDERS } from "../config/providers.js";
import { proxyAwareFetch } from "../utils/proxyFetch.js";
import { SSE_DONE } from "../utils/sseConstants.js";
import { FETCH_CONNECT_TIMEOUT_MS } from "../config/runtimeConfig.js";
import { FETCH_CONNECT_TIMEOUT_MS, HTTP_STATUS } from "../config/runtimeConfig.js";
import { resolveProviderTimeoutMs } from "../services/providerTimeout.js";
import {
QODER_CHAT_URL_ENCODED,
QODER_CHAT_BASE_ALT,
QODER_CHAT_SIG_PATH,
QODER_MODEL_MAP,
QODER_CONTEXT_TIER_ENV,
qoderInferenceBase,
} from "../shared/qoder/constants.js";
import { getQoderModelConfig, resolveQoderModels, isQoderPat, resolveQoderCredentials } from "../services/qoderModels.js";
import { OPENAI_BLOCK, CLAUDE_BLOCK } from "../translator/schema/blocks.js";
import { encodeDataUri } from "../translator/concerns/image.js";
import { createQoderSseCoalescer } from "../shared/qoder/sse.js";
import { rewriteQoderMessageAttachments } from "../shared/qoder/attachments.js";
import { resolveQoderContextTier, applyQoderContextTier } from "../shared/qoder/contextTier.js";
/**
* Hoist role:"system" messages out of the messages array (Qoder rejects
@@ -71,15 +73,16 @@ function normalizeMessages(messages) {
*
* Text-only content is flattened to a plain string (Qoder's historical
* shape). When images are present the content stays an array and image
* blocks are kept as OpenAI-style `image_url` parts — verified against the
* upstream: it accepts both http(s) URLs and inline base64 data: URIs
* directly, no pre-upload to the /image/upload OSS flow required (that is
* a qodercli client-side choice, not a protocol requirement). The legacy
* blocks are kept as OpenAI-style `image_url` parts. Native qodercli
* uploads inlined bytes to `/api/v2/image/upload` first and then sends
* the OSS URL — `buildQoderRequestBody` does that rewrite before this
* runs. Tiny leftover data URIs are still accepted. The legacy
* top-level `image_urls` / `chat_context.imageUrls` slots stay null —
* qodercli leaves them null too.
*
* Claude-style `{type:"image", source:{...}}` blocks are converted to
* `image_url` so claude-format clients also round-trip.
* `image_url`. File/document blocks that survived rewrite become short
* stubs so 30MB PDFs never land in agent_chat_generation.
*/
function normalizeContent(content) {
if (typeof content === "string") return content;
@@ -89,10 +92,24 @@ function normalizeContent(content) {
const blocks = [];
const textParts = [];
let hasImage = false;
const pushText = (text) => {
if (!text) return;
if (hasImage || blocks.length) blocks.push({ type: OPENAI_BLOCK.TEXT, text });
else textParts.push(text);
};
const imageUrlOf = (item) => {
if (typeof item.image_url === "string" && item.image_url) return item.image_url;
if (typeof item.image_url?.url === "string" && item.image_url.url) return item.image_url.url;
return null;
};
for (const item of content) {
if (!item || typeof item !== "object") continue;
if (item.type === OPENAI_BLOCK.IMAGE_URL && typeof item.image_url?.url === "string" && item.image_url.url) {
blocks.push({ type: OPENAI_BLOCK.IMAGE_URL, image_url: { url: item.image_url.url } });
const imageUrl = item.type === OPENAI_BLOCK.IMAGE_URL ? imageUrlOf(item) : null;
if (imageUrl) {
blocks.push({ type: OPENAI_BLOCK.IMAGE_URL, image_url: { url: imageUrl } });
hasImage = true;
} else if (item.type === CLAUDE_BLOCK.IMAGE && item.source) {
// Claude base64/url image → OpenAI image_url equivalent.
@@ -104,13 +121,14 @@ function normalizeContent(content) {
blocks.push({ type: OPENAI_BLOCK.IMAGE_URL, image_url: { url } });
hasImage = true;
}
} else if (item.type === OPENAI_BLOCK.FILE) {
const name = item.file?.filename || item.file?.name || "file";
pushText(`[file omitted: ${name} — Qoder reads documents via its file API, not inlined bytes]`);
} else if (item.type === CLAUDE_BLOCK.DOCUMENT) {
const name = item.title || "document";
pushText(`[file omitted: ${name} — Qoder reads documents via its file API, not inlined bytes]`);
} else if (typeof item.text === "string" && item.text) {
if (hasImage || blocks.length) {
// Keep ordering faithful once images are in play.
blocks.push({ type: OPENAI_BLOCK.TEXT, text: item.text });
} else {
textParts.push(item.text);
}
pushText(item.text);
}
}
@@ -190,16 +208,16 @@ function truncate(s, n) {
/**
* Map the OpenAI-style request body into the exact shape Qoder expects.
*/
async function buildQoderRequestBody({ model, body, credentials, log, proxyOptions, signal }) {
async function buildQoderRequestBody({ model, body, credentials, log, proxyOptions, signal, uploadFn = null, region = "intl" }) {
const qoderKey = String(model || "").replace(/^qoder\//, "");
// Fetch model config from dynamic API instead of relying on static QODER_MODEL_MAP.
// This allows support for new Qoder models (e.g., qmodel_latest) without code changes.
let modelConfig = await getQoderModelConfig(credentials, qoderKey, { log, proxyOptions, signal });
let modelConfig = await getQoderModelConfig(credentials, qoderKey, { log, proxyOptions, signal, region });
if (!modelConfig) {
// Try a forced refresh once before giving up — the cache may simply
// not be populated yet on first ever call for this credential.
const refreshed = await resolveQoderModels(credentials, { forceRefresh: true, log, proxyOptions, signal });
const refreshed = await resolveQoderModels(credentials, { forceRefresh: true, log, proxyOptions, signal, region });
const retried = refreshed?.rawConfigs.get(qoderKey);
if (!retried) {
throw new Error(
@@ -209,7 +227,30 @@ async function buildQoderRequestBody({ model, body, credentials, log, proxyOptio
modelConfig = { ...retried, key: qoderKey };
}
const { messages, systemText } = normalizeMessages(body.messages || []);
const incoming = Array.isArray(body.messages)
? body.messages.map((m) => {
if (!m || typeof m !== "object") return m;
return {
...m,
content: Array.isArray(m.content)
? m.content.map((b) => (b && typeof b === "object" ? { ...b } : b))
: m.content,
};
})
: [];
try {
await rewriteQoderMessageAttachments(incoming, {
credentials,
log,
proxyOptions,
signal,
uploadFn,
});
} catch (err) {
log?.warn?.("QODER", `attachment rewrite failed: ${err.message}`);
}
const { messages, systemText } = normalizeMessages(incoming);
const tools = body.tools;
const isReasoning = !!modelConfig.is_reasoning;
const maxOutputTokens = Number(modelConfig.max_output_tokens) || 0;
@@ -228,7 +269,21 @@ async function buildQoderRequestBody({ model, body, credentials, log, proxyOptio
const sessionId = stableHash("qoder-session", psd.userId, qoderKey);
const recordId = stableChatRecordId(qoderKey, messages, tools, maxTokens);
return {
// Context-window tier (200K/400K/1M): the IDE picks one from model_config.context_config;
// qodercli-style requests default to the smallest. Escalate when the prompt no longer fits.
const tierChoice = resolveQoderContextTier(
modelConfig,
{ system: systemText, messages, tools },
{ preference: process.env[QODER_CONTEXT_TIER_ENV] },
);
if (tierChoice) {
log?.info?.(
"QODER",
`context tier ${tierChoice.tier.name} (${tierChoice.tier.tokenCount} tokens, ${tierChoice.reason}) for ~${tierChoice.estimatedTokens} prompt tokens`,
);
}
const built = {
qoderKey,
payload: {
request_id: uuidv4(),
@@ -276,51 +331,71 @@ async function buildQoderRequestBody({ model, body, credentials, log, proxyOptio
},
modelConfig,
};
if (tierChoice) applyQoderContextTier(built.payload, tierChoice.tier);
return built;
}
/**
* Check if a qoder error message indicates a billing/quota block.
* Signatures: code 112 (quota exhausted), code 10605 (queue throttle), pricingUrl field.
* Signatures: code 110 (billing daily count exceeded), code 112 (quota
* exhausted), code 10605 (queue throttle), pricingUrl field.
*/
function isBillingBlock(inner) {
if (!inner || typeof inner !== "string") return false;
const lowerMsg = inner.toLowerCase();
// Match: {"code":"112",...}, {"code":"10605",...}, or pricingUrl field
return /\"code\"\s*:\s*\"(112|10605)\"/.test(inner) || lowerMsg.includes("pricingurl");
if (lowerMsg.includes("pricingurl")) return true;
// Parsed code preferred over regex: matches numeric or string "110"/"112"/"10605".
try {
const parsed = JSON.parse(inner);
const code = String(parsed?.code ?? "");
if (code === "110" || code === "112" || code === "10605") return true;
} catch { /* not JSON — fall through to legacy shape match */ }
// Match legacy exact shapes: {"code":"112",...}, {"code":"10605",...}.
return /"code"\s*:\s*"(112|10605)"/.test(inner);
}
/**
* Peek the first SSE frame to detect billing errors before piping.
* Returns { isBilling, statusVal, message, consumed } — `consumed` is every
* Peek the first SSE data line to detect upstream errors before piping.
* Returns { isError, isBilling, statusVal, message, consumed } — `consumed` is every
* byte read so far (including the peeked line) so the caller can re-process
* it and nothing is dropped from the stream.
*/
async function peekFirstQoderFrame(reader, decoder) {
let consumed = "";
let offset = 0;
let upstreamDone = false;
while (true) {
const { done, value } = await reader.read();
if (done) return { isBilling: false, consumed, upstreamDone: true };
let nl = consumed.indexOf("\n", offset);
if (nl === -1 && !upstreamDone) {
const { done, value } = await reader.read();
upstreamDone = done;
consumed += done ? decoder.decode() : decoder.decode(value, { stream: true });
continue;
}
if (offset >= consumed.length) return { isError: false, consumed, upstreamDone };
if (nl === -1) nl = consumed.length;
consumed += decoder.decode(value, { stream: true });
const nl = consumed.indexOf("\n");
if (nl === -1) continue; // need a full line first
const line = consumed.slice(0, nl).replace(/\r$/, "").trim();
const line = consumed.slice(offset, nl).replace(/\r$/, "").trim();
offset = nl + 1;
if (!line.startsWith("data:")) continue;
const data = line.slice(5).trimStart();
if (data === "[DONE]") return { isBilling: false, consumed };
if (data === "[DONE]") return { isError: false, consumed, upstreamDone };
let envelope;
try { envelope = JSON.parse(data); } catch { return { isBilling: false, consumed }; }
try { envelope = JSON.parse(data); } catch { return { isError: false, consumed, upstreamDone }; }
const statusVal = typeof envelope.statusCodeValue === "number" ? envelope.statusCodeValue : 200;
const inner = typeof envelope.body === "string" ? envelope.body : "";
// statusCodeValue is documented numeric, but accept numeric strings defensively.
const raw = Number(envelope?.statusCodeValue);
const statusVal = Number.isNaN(raw) ? 200 : raw;
const inner = typeof envelope?.body === "string"
? envelope.body
: envelope?.body != null ? JSON.stringify(envelope.body) : "";
if (statusVal !== 200 && isBillingBlock(inner)) {
return { isBilling: true, statusVal, message: inner || `qoder billing block (${statusVal})` };
if (statusVal !== 200) {
return { isError: true, isBilling: isBillingBlock(inner), statusVal, message: inner || `upstream status ${statusVal}` };
}
return { isBilling: false, consumed };
return { isError: false, consumed, upstreamDone };
}
}
@@ -331,32 +406,41 @@ async function peekFirstQoderFrame(reader, decoder) {
* Each upstream line looks like:
* data: {"statusCodeValue":200,"body":"{\"choices\":[{\"delta\":{...}}]}"}
* The inner body is an OpenAI streaming chunk (or "[DONE]"). We unwrap it
* and re-emit as `data: <inner>\n\n`. Errors become a synthetic OpenAI error
* chunk + [DONE].
* and re-emit as `data: <inner>\n\n`. First-frame errors become HTTP errors;
* errors after streaming starts retain the synthetic chunk + [DONE] path.
*
* Critical: Qoder's SSE often keeps the socket open after the terminal
* [DONE]/error frame (agent keepalive). Non-streaming clients drain via
* response.text() which hangs until the socket closes — so on terminal
* events we cancel the upstream reader and close our stream immediately.
*
* NEW: Peek first frame to detect billing blocks (code 112/10605/pricingUrl).
* If detected, return 403 response so chatCore marks connection unavailable
* and triggers combo fallback instead of leaking error text into chat.
* Usage: Qoder puts finish_reason on `delta` and sends token counts on a
* later `choices: []` frame. Downstream OpenAI/Claude clients only read
* usage from the finish chunk, so we coalesce those two frames (see
* createQoderSseCoalescer) before forwarding.
*
* Peek the first frame for errors before committing to HTTP 200. Preserve
* upstream error statuses so chatCore can handle failures instead of recording
* error text as a successful completion. Billing blocks retain the existing
* 403 mapping for quota/account fallback.
*/
async function wrapQoderSSE(response, model) {
async function wrapQoderSSE(response, model, log = null) {
if (!response.ok || !response.body) return response;
const decoder = new TextDecoder();
const reader = response.body.getReader();
// Peek first frame to detect billing block
// Detect errors before returning a successful streaming response.
const peek = await peekFirstQoderFrame(reader, decoder);
if (peek?.isBilling) {
// Billing block detected — return 403 so chatCore fails this connection
if (peek.isError) {
await reader.cancel().catch(() => {});
const status = peek.isBilling
? HTTP_STATUS.FORBIDDEN
: Number.isInteger(peek.statusVal) && peek.statusVal >= HTTP_STATUS.BAD_REQUEST && peek.statusVal <= 599
? peek.statusVal : HTTP_STATUS.BAD_GATEWAY;
return new Response(
JSON.stringify({ error: { message: peek.message, code: peek.statusVal } }),
{ status: 403, headers: { "Content-Type": "application/json" } }
{ status, headers: { "Content-Type": "application/json" } }
);
}
@@ -365,6 +449,11 @@ async function wrapQoderSSE(response, model) {
const upstreamDrained = peek.upstreamDone === true;
const encoder = new TextEncoder();
let doneEmitted = false;
const coalescer = createQoderSseCoalescer({ model, encoder, sseDone: SSE_DONE });
const syncDone = () => {
if (coalescer.doneEmitted) doneEmitted = true;
};
// Process one already-extracted SSE line (no trailing newline).
const processLine = (line, controller) => {
@@ -375,16 +464,42 @@ async function wrapQoderSSE(response, model) {
const data = trimmed.slice(5).trimStart();
if (data === "[DONE]") {
controller.enqueue(encoder.encode(SSE_DONE));
doneEmitted = true;
coalescer.flush(controller);
syncDone();
return;
}
let envelope;
try { envelope = JSON.parse(data); } catch { return; }
const statusVal = typeof envelope.statusCodeValue === "number" ? envelope.statusCodeValue : 200;
const inner = typeof envelope.body === "string" ? envelope.body : "";
const statusVal = Number(envelope.statusCodeValue) || 200;
const inner = typeof envelope.body === "string"
? envelope.body
: envelope.body != null ? JSON.stringify(envelope.body) : "";
if (statusVal !== 200) {
// Always visible: error envelopes are rare and worth one stderr line at
// any log level (response bodies carry no credentials).
try {
console.error(`[QODER] error envelope status=${statusVal} statusType=${typeof envelope.statusCodeValue} bodyType=${typeof envelope.body} body=${truncate(inner, 300)}`);
} catch { /* logging must not break the stream */ }
if (isBillingBlock(inner)) {
// Billing/quota envelope at any stream position (peek only covers the
// first frame): emit a structured error chunk, not fake assistant text.
// parseSSEToOpenAIResponse understands chunk.error and turns it into a
// non-200 result so chat.js locks the model and falls back. Streaming
// clients receive a real SSE error instead of "[qoder error ...]" text.
const errObj = JSON.stringify({
error: {
message: inner || `qoder billing block (${statusVal})`,
code: "qoder_billing_block",
status: 403,
type: "quota_error",
},
});
controller.enqueue(encoder.encode(`data: ${errObj}\n\n`));
controller.enqueue(encoder.encode(SSE_DONE));
doneEmitted = true;
return;
}
const msg = inner || `upstream status ${statusVal}`;
const errChunk = JSON.stringify({
id: `qoder-error-${Date.now()}`,
@@ -399,14 +514,8 @@ async function wrapQoderSSE(response, model) {
return;
}
if (!inner) return;
if (inner === "[DONE]") {
controller.enqueue(encoder.encode(SSE_DONE));
doneEmitted = true;
return;
}
// Strip embedded newlines so the SSE frame stays a single event.
const sanitized = inner.replace(/\r?\n/g, "");
controller.enqueue(encoder.encode(`data: ${sanitized}\n\n`));
coalescer.handleInner(inner, controller);
syncDone();
};
const stream = new ReadableStream({
@@ -465,7 +574,7 @@ async function wrapQoderSSE(response, model) {
} finally {
if (!doneEmitted) {
try {
controller.enqueue(encoder.encode(SSE_DONE));
coalescer.flush(controller);
doneEmitted = true;
} catch { /* already closed */ }
}
@@ -489,18 +598,13 @@ async function wrapQoderSSE(response, model) {
}
export class QoderExecutor extends BaseExecutor {
constructor() {
super("qoder", PROVIDERS.qoder);
constructor(provider = "qoder") {
super(provider, PROVIDERS[provider]);
this.region = provider === "qoder-cn" ? "cn" : "intl";
}
buildUrl(credentials) {
// Job-token (jt-...) traffic must hit api2.qoder.sh — api3 rejects jt-
// with "Login expired" (403). Device tokens (dt-...) stay on api3.
const raw = credentials?.apiKey || credentials?.accessToken;
if (typeof raw === "string" && !raw.startsWith("pt-") && (raw.startsWith("jt-") || (credentials?.accessToken || "").startsWith("jt-"))) {
return `${QODER_CHAT_BASE_ALT}/algo${QODER_CHAT_SIG_PATH}?FetchKeys=llm_model_result&AgentId=agent_common&Encode=1`;
}
return QODER_CHAT_URL_ENCODED;
return `${qoderInferenceBase(credentials, this.region)}/algo${QODER_CHAT_SIG_PATH}?FetchKeys=llm_model_result&AgentId=agent_common&Encode=1`;
}
// Override execute entirely — Qoder needs:
@@ -515,7 +619,7 @@ export class QoderExecutor extends BaseExecutor {
const rawToken = credentials?.apiKey || credentials?.accessToken;
if (isQoderPat(rawToken)) {
try {
credentials = await resolveQoderCredentials(credentials, proxyOptions, signal);
credentials = await resolveQoderCredentials(credentials, proxyOptions, signal, this.region);
} catch (err) {
log?.error?.("QODER", `PAT exchange failed: ${err.message}`);
const fakeResp = new Response(
@@ -550,7 +654,7 @@ export class QoderExecutor extends BaseExecutor {
let qoderKey;
let payload;
try {
({ qoderKey, payload } = await buildQoderRequestBody({ model, body, credentials, log, proxyOptions, signal }));
({ qoderKey, payload } = await buildQoderRequestBody({ model, body, credentials, log, proxyOptions, signal, region: this.region }));
} catch (err) {
const fakeResp = new Response(
JSON.stringify({ error: { message: err.message } }),
@@ -609,8 +713,15 @@ export class QoderExecutor extends BaseExecutor {
response = await proxyAwareFetch(
url,
{ method: "POST", headers, body: encodedBodyBuf, signal: mergedSignal },
proxyOptions,
// A failed proxy request may already have reached Qoder. Replaying
// the same COSY signature directly reuses its requestId and returns
// 403/code 103. Let the caller retry through execute() with fresh signing.
{ ...proxyOptions, strictProxy: true },
);
} catch (err) {
// strictProxy wraps transport errors; retain caller cancellation semantics.
if (mergedSignal.aborted) throw mergedSignal.reason;
throw err;
} finally {
clearTimeout(connectTimer);
}
@@ -620,7 +731,7 @@ export class QoderExecutor extends BaseExecutor {
return { response, url, headers, transformedBody: payload };
}
const wrapped = await wrapQoderSSE(response, `qoder/${qoderKey}`);
const wrapped = await wrapQoderSSE(response, `${this.provider}/${qoderKey}`, log);
return { response: wrapped, url, headers, transformedBody: payload };
}

View File

@@ -0,0 +1,116 @@
import { DefaultExecutor } from "./default.js";
import { getMimoAccountCookie, invalidateMimoAccountCookieCache, resolveMimoServerBase, MIMO_API_UA } from "../shared/mimoAccount.js";
// Dual-route v2.6 models.
// v2.6 models dynamically route to the account service when desktop session credentials
// (mimoPassToken or account cookie) are present to consume weekly quota, falling back to
// the cloud API (sk- key) otherwise.
const ACCOUNT_MODELS = new Set([
"mimo-v2.6-pro",
"mimo-v2.6-flash",
"mimo-v2.6-pro-ultraspeed",
]);
// Session cookie resolved in execute() (async) and read back by buildHeaders()
// (sync — BaseExecutor.execute does not await it). Carried on the per-request
// credentials object, same as runtimeTransport.
const COOKIE_KEY = "__mimoAccountCookie";
// Upstream calls may hand us either the bare id or a `provider/model` ref.
function bareModel(model) {
const s = String(model || "");
const i = s.indexOf("/");
return i >= 0 ? s.slice(i + 1) : s;
}
export class XiaomiMimoExecutor extends DefaultExecutor {
constructor() {
super("xiaomi-mimo");
}
static isAccountRoute(model, credentials) {
const bare = bareModel(model);
if (!ACCOUNT_MODELS.has(bare)) return false;
return Boolean(
credentials?.[COOKIE_KEY] ||
credentials?.providerSpecificData?.mimoPassToken
);
}
isAccountRoute(model, credentials) {
return XiaomiMimoExecutor.isAccountRoute(model, credentials);
}
buildUrl(model, stream, urlIndex = 0, credentials = null) {
// Account route models live on the account-service route, which is not one of the
// declared transports — resolve it before the default runtimeTransport path.
if (this.isAccountRoute(model, credentials)) {
return `${resolveMimoServerBase(credentials?.providerSpecificData)}/api/route/chat/completions`;
}
// Cloud API models keep default handling, so a Claude-format client reaches
// the /anthropic/v1/messages transport.
return super.buildUrl(model, stream, urlIndex, credentials);
}
buildHeaders(credentials, stream = true, url, model) {
if (this.isAccountRoute(model, credentials) && credentials?.[COOKIE_KEY]) {
// Account route models authenticate with the account-session cookie, not the key.
return {
"Content-Type": "application/json",
Accept: stream ? "text/event-stream" : "application/json",
"User-Agent": MIMO_API_UA,
Cookie: credentials[COOKIE_KEY],
};
}
return super.buildHeaders(credentials, stream, url, model);
}
transformRequest(model, body, stream, credentials) {
// super runs stripUnsupportedParams, which flattens content-part
// arrays (see the xiaomi-mimo rule in translator/concerns/paramSupport.js).
const out = super.transformRequest(model, body, stream, credentials);
// Account route models: bridge reasoning_effort to official output_config.effort
// (matches MiMo Desktop app.asar behavior).
if (this.isAccountRoute(model, credentials)) {
const rawEffort = out.reasoning_effort || body?.reasoning_effort || body?.output_config?.effort;
if (rawEffort) {
delete out.reasoning_effort;
const norm = String(rawEffort).toLowerCase() === "xhigh" ? "high" : String(rawEffort).toLowerCase();
out.output_config = { ...(out.output_config || {}), effort: norm };
}
if (out.temperature == null) out.temperature = 1.0;
if (out.top_p == null) out.top_p = 0.95;
}
return out;
}
async execute(args) {
const { model, credentials, proxyOptions = null } = args;
if (!this.isAccountRoute(model, credentials)) return super.execute(args);
const cookie = await getMimoAccountCookie(credentials?.providerSpecificData, proxyOptions);
if (!cookie) {
return super.execute(args);
}
credentials[COOKIE_KEY] = cookie;
const result = await super.execute(args);
// A cached session can expire early — drop it and retry once with a fresh one.
if (result.response.status === 401) {
invalidateMimoAccountCookieCache();
const fresh = await getMimoAccountCookie(credentials?.providerSpecificData, proxyOptions).catch(() => null);
if (fresh) {
credentials[COOKIE_KEY] = fresh;
return super.execute(args);
}
}
return result;
}
}
export const __test__ = { ACCOUNT_MODELS, bareModel, COOKIE_KEY };
export default XiaomiMimoExecutor;

View File

@@ -29,11 +29,18 @@ import {
zedLlmFetch,
} from "../shared/zedAuth.js";
// Wire values for the `provider` field of POST /completions. These are NOT
// display names: cloud.zed.dev matches them exactly, and an unrecognized value
// fails the whole request with `500 {"message":"An internal server error
// occurred."}` before the model is ever looked at. Spellings come from Zed's
// own GET /models catalog: `anthropic`, `open_ai`, `google` (note underscore),
// `x_ai` follows the same convention — so feeding a catalog value back through
// normalizeZedProvider is identity.
const ZED_PROVIDER = {
anthropic: "Anthropic",
openai: "OpenAi",
google: "Google",
xai: "XAi",
anthropic: "anthropic",
openai: "open_ai",
google: "google",
xai: "x_ai",
};
function normalizeZedProvider(value, model) {
@@ -55,7 +62,14 @@ function buildProviderRequest(provider, model, body, stream, credentials) {
return openaiToClaudeRequest(model, body, true);
}
if (provider === ZED_PROVIDER.google) {
return openaiToGeminiRequest(model, body, true);
const geminiRequest = openaiToGeminiRequest(model, body, true);
// Zed's hosted Gemini backend speaks the Vertex safety vocabulary, not the
// public Gemini API enum the shared translator emits (`OFF`, `CIVIC_INTEGRITY`,
// `DANGEROUS_CONTENT`). Drop client-side safetySettings for the Zed Google
// path so Zed applies its own defaults — scoped here so native Gemini/
// Antigravity is untouched.
delete geminiRequest.safetySettings;
return geminiRequest;
}
if (provider === ZED_PROVIDER.openai) {
return openaiToOpenAIResponsesRequest(model, body, true, credentials);

View File

@@ -9,17 +9,19 @@ import { createRequestLogger } from "../utils/requestLogger.js";
import { getModelTargetFormat, getModelSupportedFormats, getModelStrip, getModelUpstreamId, getModelType, PROVIDER_ID_TO_ALIAS } from "../config/providerModels.js";
import { PROVIDERS } from "../config/providers.js";
import { createErrorResult, parseUpstreamError, formatProviderError } from "../utils/error.js";
import { upstreamResponseHeaders } from "../utils/upstreamHeaders.js";
import { HTTP_STATUS, TOKEN_SAVER_HEADER } from "../config/runtimeConfig.js";
import { handleBypassRequest } from "../utils/bypassHandler.js";
import { trackPendingRequest, saveRequestDetail } from "@/lib/usageDb.js";
import { getExecutor } from "../executors/index.js";
import { supportsGrokCliReasoningEffort } from "../config/grokCli.js";
import { buildRequestDetail, extractRequestConfig } from "./chatCore/requestDetail.js";
import { buildRequestDetail, extractRequestConfig, shouldPersistRequestDetail } from "./chatCore/requestDetail.js";
import { handleForcedSSEToJson } from "./chatCore/sseToJsonHandler.js";
import { handleNonStreamingResponse } from "./chatCore/nonStreamingHandler.js";
import { handleStreamingResponse, buildOnStreamComplete } from "./chatCore/streamingHandler.js";
import { detectClientTool, isNativePassthrough } from "../utils/clientDetector.js";
import { dedupeTools } from "../utils/toolDeduper.js";
import { takeRenamedToolNames } from "../utils/opencodeFingerprint.js";
import { injectCaveman } from "../rtk/caveman.js";
import { injectPonytail } from "../rtk/ponytail.js";
import { compressMessages, formatRtkLog } from "../rtk/index.js";
@@ -28,7 +30,7 @@ import { compressWithPxpipe } from "../rtk/pxpipe.js";
import { getCapabilitiesForModel } from "../providers/capabilities.js";
import { stripUnsupportedModalities } from "../translator/concerns/modality.js";
import { prefetchRemoteImages } from "../translator/concerns/prefetch.js";
import { defaultClaudeToolType } from "../translator/concerns/toolCall.js";
import { defaultClaudeToolType, shouldDefaultClaudeToolType } from "../translator/concerns/toolCall.js";
import { resolveSessionId } from "../utils/sessionManager.js";
import { maybeRejectEarlyStreamError } from "../utils/streamErrorPeek.js";
@@ -59,7 +61,7 @@ export function stripContinuityFields(body) {
return body;
}
export async function handleChatCore({ body, modelInfo, credentials, log, onCredentialsRefreshed, onRequestSuccess, onDisconnect, clientRawRequest, connectionId, userAgent, apiKey, ccFilterNaming, rtkEnabled, headroomEnabled, headroomUrl, headroomCompressUserMessages, headroomTimeoutMs, cavemanEnabled, cavemanLevel, ponytailEnabled, ponytailLevel, pxpipeEnabled, pxpipeMinChars, pxpipeTimeoutMs, pxpipeTransform, onPxpipeEvent, sourceFormatOverride, providerThinking, capsOverride = null, streamErrorPatterns = null }) {
export async function handleChatCore({ body, modelInfo, credentials, log, onCredentialsRefreshed, onRequestSuccess, onDisconnect, clientRawRequest, connectionId, userAgent, apiKey, ccFilterNaming, rtkEnabled, headroomEnabled, headroomUrl, headroomCompressUserMessages, headroomTimeoutMs, cavemanEnabled, cavemanLevel, ponytailEnabled, ponytailLevel, pxpipeEnabled, pxpipeMinChars, pxpipeTimeoutMs, pxpipeTransform, onPxpipeEvent, sourceFormatOverride, providerThinking, capsOverride = null, streamErrorPatterns = null, persistUsage = "all" }) {
const { provider, model } = modelInfo;
const requestStartTime = Date.now();
// Stable per-session color so all lines of one CLI conversation share a tag
@@ -116,6 +118,19 @@ export async function handleChatCore({ body, modelInfo, credentials, log, onCred
}
}
// Per-request opt-out: client can bypass all token savers via header
const tokenSaverEnabled = clientRawRequest?.headers?.[TOKEN_SAVER_HEADER]?.toLowerCase() !== "off";
// Cursor's translator rewrites tool_result into user text, so RTK must run on
// the source body before translation. Every other pair translates the tool
// shapes 1:1 — keep the post-translate pass there so those providers are
// untouched (and a retry never re-compresses an already-compressed body).
const preTranslateRtk = provider === "cursor"
? compressMessages(body, tokenSaverEnabled && rtkEnabled)
: null;
const preTranslateRtkLine = formatRtkLog(preTranslateRtk);
if (preTranslateRtkLine) console.log(preTranslateRtkLine);
const clientRequestedStreaming = body.stream === true || sourceFormat === FORMATS.ANTIGRAVITY || sourceFormat === FORMATS.GEMINI || sourceFormat === FORMATS.GEMINI_CLI;
const providerRequiresStreaming = PROVIDERS[provider]?.forceStream === true;
let stream = providerRequiresStreaming ? true : (body.stream !== false);
@@ -246,15 +261,16 @@ export async function handleChatCore({ body, modelInfo, credentials, log, onCred
// Claude tool schema requires `type` to be explicitly set; strict gateways (e.g., MiniMax)
// reject legacy payloads that omit it with HTTP 400. Default to "custom" when missing.
if (finalFormat === FORMATS.CLAUDE && Array.isArray(translatedBody.tools)) {
// Provider-scoped via quirks (shouldDefaultClaudeToolType): only gateways that declare
// requireClaudeToolType get the explicit type. Applying it unconditionally breaks
// Claude-format endpoints that only accept the legacy typeless tool shape — DeepSeek's
// Anthropic-compatible endpoint 400s with "unknown variant `custom`" (#3905).
if (shouldDefaultClaudeToolType(provider, finalFormat, translatedBody.tools, PROVIDERS)) {
translatedBody.tools = defaultClaudeToolType(translatedBody.tools);
}
// Per-request opt-out: client can bypass all token savers via header
const tokenSaverEnabled = clientRawRequest?.headers?.[TOKEN_SAVER_HEADER]?.toLowerCase() !== "off";
// RTK: compress tool_result content
const rtkStats = compressMessages(translatedBody, tokenSaverEnabled && rtkEnabled);
// RTK: compress tool_result content. Skipped when already done pre-translate.
const rtkStats = preTranslateRtk || compressMessages(translatedBody, tokenSaverEnabled && rtkEnabled);
const rtkLine = formatRtkLog(rtkStats);
if (rtkLine) log?.info?.("RTK", rtkLine.replace(/^\[RTK\] /, ""));
@@ -273,6 +289,8 @@ export async function handleChatCore({ body, modelInfo, credentials, log, onCred
// Token-saver flags accumulator for the single "⚙" log line below.
const xf = [];
if (rtkStats?.hits?.length) xf.push(`RTK:${rtkStats.hits.length}`);
// Caveman: inject terse-style system prompt
if (tokenSaverEnabled && cavemanEnabled && cavemanLevel) {
injectCaveman(translatedBody, finalFormat, cavemanLevel);
@@ -374,19 +392,25 @@ export async function handleChatCore({ body, modelInfo, credentials, log, onCred
providerHeaders = result.headers;
finalBody = result.transformedBody;
providerResponseFormat = result.responseFormat || targetFormat;
const renamedToolNames = takeRenamedToolNames(translatedBody);
if (renamedToolNames?.size) {
toolNameMap = new Map([...(toolNameMap || []), ...renamedToolNames]);
}
reqLogger.logTargetRequest(providerUrl, providerHeaders, finalBody);
} catch (error) {
trackPendingRequest(model, provider, connectionId, false, true);
saveRequestDetail(buildRequestDetail({
provider, model, connectionId,
latency: { ttft: 0, total: Date.now() - requestStartTime },
tokens: { prompt_tokens: 0, completion_tokens: 0 },
request: extractRequestConfig(body, stream),
providerRequest: translatedBody || null,
response: { error: error.message || String(error), status: error.name === "AbortError" ? 499 : 502, thinking: null },
pxpipe: pxpipeSummary,
status: "error"
})).catch(() => { });
if (shouldPersistRequestDetail(persistUsage, "error")) {
saveRequestDetail(buildRequestDetail({
provider, model, connectionId,
latency: { ttft: 0, total: Date.now() - requestStartTime },
tokens: { prompt_tokens: 0, completion_tokens: 0 },
request: extractRequestConfig(body, stream),
providerRequest: translatedBody || null,
response: { error: error.message || String(error), status: error.name === "AbortError" ? 499 : 502, thinking: null },
pxpipe: pxpipeSummary,
status: "error"
})).catch(() => { });
}
if (error.name === "AbortError") {
streamController.handleError(error);
@@ -451,16 +475,18 @@ export async function handleChatCore({ body, modelInfo, credentials, log, onCred
if (!providerResponse.ok) {
trackPendingRequest(model, provider, connectionId, false, true);
const { statusCode, message, resetsAtMs } = await parseUpstreamError(providerResponse, executor);
saveRequestDetail(buildRequestDetail({
provider, model, connectionId,
latency: { ttft: 0, total: Date.now() - requestStartTime },
tokens: { prompt_tokens: 0, completion_tokens: 0 },
request: extractRequestConfig(body, stream),
providerRequest: finalBody || translatedBody || null,
response: { error: message, status: statusCode, thinking: null },
pxpipe: pxpipeSummary,
status: "error"
})).catch(() => { });
if (shouldPersistRequestDetail(persistUsage, "error")) {
saveRequestDetail(buildRequestDetail({
provider, model, connectionId,
latency: { ttft: 0, total: Date.now() - requestStartTime },
tokens: { prompt_tokens: 0, completion_tokens: 0 },
request: extractRequestConfig(body, stream),
providerRequest: finalBody || translatedBody || null,
response: { error: message, status: statusCode, thinking: null },
pxpipe: pxpipeSummary,
status: "error"
})).catch(() => { });
}
const errMsg = formatProviderError(new Error(message), provider, model, statusCode);
if (log?.errorLine) {
@@ -468,11 +494,11 @@ export async function handleChatCore({ body, modelInfo, credentials, log, onCred
log.errorLine(reqTag, "✗", `ERROR ${statusCode} · ${provider}/${model} · ${Date.now() - requestStartTime}ms${urlStr}\n ${errMsg}`);
}
reqLogger.logError(new Error(message), finalBody || translatedBody);
return createErrorResult(statusCode, errMsg, resetsAtMs);
return createErrorResult(statusCode, errMsg, resetsAtMs, upstreamResponseHeaders(providerResponse.headers));
}
const appendLog = () => {}; // request log derived from usageHistory; kept as no-op seam for handlers
const sharedCtx = { provider, model, body, stream, translatedBody, finalBody, requestStartTime, connectionId, apiKey, clientRawRequest, onRequestSuccess, pxpipe: pxpipeSummary, reqTag, log, streamErrorPatterns };
const sharedCtx = { provider, model, body, stream, translatedBody, finalBody, requestStartTime, connectionId, apiKey, clientRawRequest, onRequestSuccess, pxpipe: pxpipeSummary, reqTag, log, streamErrorPatterns, persistUsage };
// Early-peek streaming responses for configured in-stream error patterns.
// Some upstreams fail INSIDE a 200 SSE stream; without this the failure is
@@ -500,7 +526,7 @@ export async function handleChatCore({ body, modelInfo, credentials, log, onCred
// Provider forced streaming but client wants JSON
if (!clientRequestedStreaming && providerRequiresStreaming) {
const result = await handleForcedSSEToJson({ ...sharedCtx, providerResponse, sourceFormat, targetFormat: providerResponseFormat, customToolNames, trackDone, appendLog });
const result = await handleForcedSSEToJson({ ...sharedCtx, providerResponse, sourceFormat, targetFormat: providerResponseFormat, customToolNames, toolNameMap, trackDone, appendLog });
if (result) { streamController.handleComplete(); return result; }
}

View File

@@ -4,12 +4,15 @@ import { fromOpenAIFinish } from "../../translator/concerns/finishReason.js";
import { ollamaBodyToOpenAI } from "../../translator/response/ollama-to-openai.js";
import { addBufferToUsage, filterUsageForFormat } from "../../utils/usageTracking.js";
import { createErrorResult } from "../../utils/error.js";
import { upstreamResponseHeaders } from "../../utils/upstreamHeaders.js";
import { HTTP_STATUS } from "../../config/runtimeConfig.js";
import { parseSSEToOpenAIResponse } from "./sseToJsonHandler.js";
import { buildRequestDetail, extractRequestConfig, extractUsageFromResponse, saveUsageStats, formatDoneLine } from "./requestDetail.js";
import { unwrapClineEnvelope } from "../../shared/clineEnvelope.js";
import { buildRequestDetail, extractRequestConfig, extractUsageFromResponse, saveUsageStats, formatDoneLine, tokensForDetail, shouldPersistRequestDetail } from "./requestDetail.js";
import { saveRequestDetail } from "@/lib/usageDb.js";
import { matchStreamErrorPatterns } from "../../utils/streamErrorPatterns.js";
import { decloakToolNames } from "../../utils/claudeCloaking.js";
import { restoreToolNames } from "../../utils/opencodeFingerprint.js";
import { ROLE, RESPONSES_ITEM } from "../../translator/schema/index.js";
function parseToolArguments(value) {
@@ -282,7 +285,7 @@ export function translateNonStreamingResponse(responseBody, targetFormat, source
/**
* Handle non-streaming response from provider.
*/
export async function handleNonStreamingResponse({ providerResponse, provider, model, sourceFormat, targetFormat, body, stream, translatedBody, finalBody, requestStartTime, connectionId, apiKey, clientRawRequest, onRequestSuccess, reqLogger, toolNameMap, customToolNames, trackDone, appendLog, pxpipe, reqTag, log, streamErrorPatterns }) {
export async function handleNonStreamingResponse({ providerResponse, provider, model, sourceFormat, targetFormat, body, stream, translatedBody, finalBody, requestStartTime, connectionId, apiKey, clientRawRequest, onRequestSuccess, reqLogger, toolNameMap, customToolNames, trackDone, appendLog, pxpipe, reqTag, log, streamErrorPatterns, persistUsage = "all" }) {
trackDone();
const contentType = providerResponse.headers.get("content-type") || "";
let responseBody;
@@ -305,6 +308,11 @@ export async function handleNonStreamingResponse({ providerResponse, provider, m
}
}
// Unwrap before any consumer reads choices/usage so non-stream clients get a
// bare OpenAI body and usage tracking sees data.usage. No-op unless the
// provider opts in via transport.quirks.clineEnvelope.
responseBody = unwrapClineEnvelope(responseBody, provider);
reqLogger.logProviderResponse(providerResponse.status, providerResponse.statusText, providerResponse.headers, responseBody);
if (onRequestSuccess) {
Promise.resolve()
@@ -387,28 +395,30 @@ export async function handleNonStreamingResponse({ providerResponse, provider, m
reqLogger.logConvertedResponse(translatedResponse);
const totalLatency = Date.now() - requestStartTime;
saveRequestDetail(buildRequestDetail({
provider, model, connectionId, apiKey,
latency: { ttft: totalLatency, total: totalLatency },
tokens: usage || { prompt_tokens: 0, completion_tokens: 0 },
request: extractRequestConfig(body, stream),
providerRequest: finalBody || translatedBody || null,
providerResponse: responseBody || null,
response: {
content: translatedResponse?.choices?.[0]?.message?.content || translatedResponse?.content || null,
thinking: translatedResponse?.choices?.[0]?.message?.reasoning_content || translatedResponse?.reasoning_content || null,
finish_reason: translatedResponse?.choices?.[0]?.finish_reason || "unknown"
},
pxpipe,
status: "success"
}, { endpoint: clientRawRequest?.endpoint || null })).catch(err => {
console.error("[RequestDetail] Failed to save:", err.message);
});
if (shouldPersistRequestDetail(persistUsage, "success")) {
saveRequestDetail(buildRequestDetail({
provider, model, connectionId, apiKey,
latency: { ttft: totalLatency, total: totalLatency },
tokens: tokensForDetail(usage),
request: extractRequestConfig(body, stream),
providerRequest: finalBody || translatedBody || null,
providerResponse: responseBody || null,
response: {
content: translatedResponse?.choices?.[0]?.message?.content || translatedResponse?.content || null,
thinking: translatedResponse?.choices?.[0]?.message?.reasoning_content || translatedResponse?.reasoning_content || null,
finish_reason: translatedResponse?.choices?.[0]?.finish_reason || "unknown"
},
pxpipe,
status: "success"
}, { endpoint: clientRawRequest?.endpoint || null })).catch(err => {
console.error("[RequestDetail] Failed to save:", err.message);
});
}
return {
success: true,
response: new Response(JSON.stringify(translatedResponse), {
headers: { "Content-Type": "application/json", "Access-Control-Allow-Origin": "*" }
response: new Response(JSON.stringify(restoreToolNames(translatedResponse, toolNameMap)), {
headers: { "Content-Type": "application/json", "Access-Control-Allow-Origin": "*", ...upstreamResponseHeaders(providerResponse.headers) }
})
};
}

View File

@@ -110,6 +110,29 @@ export function formatDoneLine({ usage, latency }) {
return `DONE ${latency?.total ?? 0}ms${ttftStr} · ${inStr} · OUT ${outTok}`;
}
// Request-details storage convention: always prompt_tokens / completion_tokens.
// Translators often hand Claude `{input_tokens, output_tokens}` (or Gemini
// counts) to onStreamComplete; the Details tab only reads the OpenAI names,
// so an uncanonicalized object shows up as input=0 / output=0.
export function tokensForDetail(usage) {
if (!usage || typeof usage !== "object") {
return { prompt_tokens: 0, completion_tokens: 0 };
}
return canonicalizeUsage(usage) || {
prompt_tokens: usage.prompt_tokens ?? usage.input_tokens ?? 0,
completion_tokens: usage.completion_tokens ?? usage.output_tokens ?? 0,
};
}
// Combo fallback/account hops must not inflate Details with 0-token rows.
// `streaming-start` is never persisted: the placeholder was status=success at
// tokens=0, and nested/fusion paths often abandon the stream before complete.
export function shouldPersistRequestDetail(persistUsage, kind) {
if (kind === "streaming-start") return false;
if (persistUsage === "success-only") return kind === "success";
return true;
}
export function saveUsageStats({ provider, model, tokens, connectionId, apiKey, endpoint, label = "USAGE", silent = false }) {
if (!tokens || typeof tokens !== "object") return;

View File

@@ -1,5 +1,6 @@
import { convertResponsesStreamToJson } from "../../transformer/streamToJsonConverter.js";
import { matchStreamErrorPatterns } from "../../utils/streamErrorPatterns.js";
import { restoreToolNames } from "../../utils/opencodeFingerprint.js";
import { createErrorResult } from "../../utils/error.js";
import { HTTP_STATUS } from "../../config/runtimeConfig.js";
import { FORMATS } from "../../translator/formats.js";
@@ -215,17 +216,13 @@ export async function handleForcedSSEToJson({
clientRawRequest,
onRequestSuccess,
customToolNames,
toolNameMap,
trackDone,
appendLog,
reqTag,
log,
streamErrorPatterns,
}) {
const contentType = providerResponse.headers.get("content-type") || "";
const isSSE =
contentType.includes("text/event-stream") ||
(contentType === "" && isResponsesProvider(provider));
if (!isSSE) return null; // not handled here
trackDone();
@@ -306,7 +303,7 @@ export async function handleForcedSSEToJson({
if (sourceFormat === FORMATS.OPENAI_RESPONSES) {
return {
success: true,
response: new Response(JSON.stringify(jsonResponse), {
response: new Response(JSON.stringify(restoreToolNames(jsonResponse, toolNameMap)), {
headers: {
"Content-Type": "application/json",
"Access-Control-Allow-Origin": "*",
@@ -406,7 +403,7 @@ export async function handleForcedSSEToJson({
return {
success: true,
response: new Response(JSON.stringify(finalResp), {
response: new Response(JSON.stringify(restoreToolNames(finalResp, toolNameMap)), {
headers: {
"Content-Type": "application/json",
"Access-Control-Allow-Origin": "*",
@@ -432,8 +429,16 @@ export async function handleForcedSSEToJson({
"Invalid SSE response for non-streaming request",
);
if (parsed.error) {
// Structured error chunks may carry the real upstream status (e.g. the
// Qoder executor emits status 403 for billing envelopes). Preserve it so
// the account loop locks/falls back on the right status instead of a
// generic 502. Anything outside 400-599 still maps to 502.
const upstreamStatus = Number(parsed.error.status);
const status = Number.isInteger(upstreamStatus) && upstreamStatus >= 400 && upstreamStatus <= 599
? upstreamStatus
: HTTP_STATUS.BAD_GATEWAY;
return createErrorResult(
HTTP_STATUS.BAD_GATEWAY,
status,
parsed.error.message || "Upstream SSE stream failed",
);
}
@@ -508,7 +513,7 @@ export async function handleForcedSSEToJson({
return {
success: true,
response: new Response(JSON.stringify(finalBody), {
response: new Response(JSON.stringify(restoreToolNames(finalBody, toolNameMap)), {
headers: {
"Content-Type": "application/json",
"Access-Control-Allow-Origin": "*",

View File

@@ -6,17 +6,21 @@ import {
} from "../../utils/stream.js";
import { pipeWithDisconnect } from "../../utils/streamHandler.js";
import { PROVIDERS } from "../../config/providers.js";
import { STREAM_STALL_TIMEOUT_MS } from "../../config/runtimeConfig.js";
import { HTTP_STATUS, STREAM_STALL_TIMEOUT_MS } from "../../config/runtimeConfig.js";
import { buildAbortedResponsesTerminalBytes } from "../../utils/responsesStreamHelpers.js";
import { buildStreamErrorBytes } from "../../utils/streamHelpers.js";
import {
buildRequestDetail,
extractRequestConfig,
saveUsageStats,
formatDoneLine,
tokensForDetail,
shouldPersistRequestDetail,
} from "./requestDetail.js";
import { streamStatusForContent } from "../../utils/streamErrorPatterns.js";
import { saveRequestDetail } from "@/lib/usageDb.js";
import { SSE_HEADERS_CORS as SSE_HEADERS } from "../../utils/sseConstants.js";
import { upstreamResponseHeaders } from "../../utils/upstreamHeaders.js";
// Codex returns Responses API SSE → which client format to translate INTO, by request sourceFormat.
// Gemini-family all map to ANTIGRAVITY decoder; unknown sources fall back to OPENAI.
@@ -60,7 +64,7 @@ function buildTransformStream({ provider, sourceFormat, targetFormat, userAgent,
/**
* Handle streaming response — pipe provider SSE through transform stream to client.
*/
export async function handleStreamingResponse({ providerResponse, provider, model, sourceFormat, targetFormat, userAgent, body, stream, translatedBody, finalBody, requestStartTime, connectionId, apiKey, clientRawRequest, onRequestSuccess, reqLogger, toolNameMap, customToolNames, streamController, onStreamComplete, streamDetailId, pxpipe, reqTag, log, credentials }) {
export async function handleStreamingResponse({ providerResponse, provider, model, sourceFormat, targetFormat, userAgent, body, stream, translatedBody, finalBody, requestStartTime, connectionId, apiKey, clientRawRequest, onRequestSuccess, reqLogger, toolNameMap, customToolNames, streamController, onStreamComplete, pxpipe, reqTag, log, credentials }) {
if (onRequestSuccess) {
Promise.resolve()
.then(onRequestSuccess)
@@ -128,13 +132,21 @@ export async function handleStreamingResponse({ providerResponse, provider, mode
const transformStream = buildTransformStream({ provider, sourceFormat, targetFormat, userAgent, reqLogger, toolNameMap, customToolNames, model, connectionId, body, onStreamComplete, apiKey, credentials });
// Responses passthrough: synthesize response.failed + [DONE] if the stream aborts/stalls before a terminal event
// Terminal bytes when the stream aborts after HTTP 200 was already sent, so the
// client sees a real error instead of a silently truncated stream.
// Responses passthrough keeps its own response.failed shape; every other client
// format gets the OpenAI error frame + [DONE], or `event: error` for Claude.
const isResponsesPassthrough =
sourceFormat === FORMATS.OPENAI_RESPONSES &&
targetFormat === FORMATS.OPENAI_RESPONSES;
const onAbortTerminal = isResponsesPassthrough
? buildAbortedResponsesTerminalBytes
: null;
: (message) =>
buildStreamErrorBytes(
HTTP_STATUS.GATEWAY_TIMEOUT,
message,
sourceFormat,
);
const stallTimeoutMs =
PROVIDERS[provider]?.stallTimeoutMs || STREAM_STALL_TIMEOUT_MS;
const transformedBody = pipeWithDisconnect(
@@ -145,38 +157,14 @@ export async function handleStreamingResponse({ providerResponse, provider, mode
stallTimeoutMs,
);
saveRequestDetail(
buildRequestDetail(
{
provider,
model,
connectionId,
apiKey,
latency: { ttft: 0, total: Date.now() - requestStartTime },
tokens: { prompt_tokens: 0, completion_tokens: 0 },
request: extractRequestConfig(body, stream),
providerRequest: finalBody || translatedBody || null,
providerResponse: "[Streaming - raw response not captured]",
response: {
content: "[Streaming in progress...]",
thinking: null,
type: "streaming",
},
pxpipe,
status: "success",
},
{ id: streamDetailId },
),
).catch((err) => {
console.error(
"[RequestDetail] Failed to save streaming request:",
err.message,
);
});
return {
success: true,
response: new Response(transformedBody, { headers: SSE_HEADERS }),
response: new Response(transformedBody, {
headers: {
...SSE_HEADERS,
...upstreamResponseHeaders(providerResponse.headers),
},
}),
};
}
@@ -198,6 +186,7 @@ export function buildOnStreamComplete({
reqTag,
log,
streamErrorPatterns,
persistUsage = "all",
}) {
const streamDetailId = `${Date.now()}-${Math.random().toString(36).slice(2, 11)}`;
@@ -210,37 +199,39 @@ export function buildOnStreamComplete({
const safeThinking = contentObj?.thinking || null;
const rawProviderText = typeof contentObj?.rawProviderText === "string" ? contentObj.rawProviderText : "";
saveRequestDetail(
buildRequestDetail(
{
provider,
model,
connectionId,
apiKey,
latency,
tokens: usage || { prompt_tokens: 0, completion_tokens: 0 },
request: extractRequestConfig(body, stream),
providerRequest: finalBody || translatedBody || null,
providerResponse: rawProviderText || safeContent,
response: {
content: safeContent,
thinking: safeThinking,
type: "streaming",
if (shouldPersistRequestDetail(persistUsage, "success")) {
saveRequestDetail(
buildRequestDetail(
{
provider,
model,
connectionId,
apiKey,
latency,
tokens: tokensForDetail(usage),
request: extractRequestConfig(body, stream),
providerRequest: finalBody || translatedBody || null,
providerResponse: rawProviderText || safeContent,
response: {
content: safeContent,
thinking: safeThinking,
type: "streaming",
},
pxpipe,
status: streamStatusForContent(
streamErrorPatterns?.[provider],
safeContent,
),
},
pxpipe,
status: streamStatusForContent(
streamErrorPatterns?.[provider],
safeContent,
),
},
{ id: streamDetailId },
),
).catch((err) => {
console.error(
"[RequestDetail] Failed to update streaming content:",
err.message,
);
});
{ id: streamDetailId },
),
).catch((err) => {
console.error(
"[RequestDetail] Failed to update streaming content:",
err.message,
);
});
}
// Persist stream usage to DB (no console line; the "📊 done" line below is authoritative)
saveUsageStats({

View File

@@ -0,0 +1,266 @@
import { Buffer } from "node:buffer";
// Gemini Live API realtime STT transport.
//
// The REST generateContent path (sttCore.transcribeGemini) only transcribes
// whole files inline. The Live API's `:bidiGenerateContent` WebSocket is the
// streaming counterpart: audio goes up as realtimeInput mediaChunks and the
// server pushes incremental `serverContent.inputTranscription` events back.
// This module owns the socket lifecycle only — envelope/response shaping
// stays in sttCore so the engine's single STT exit shape is preserved.
//
// Marker contract: dispatched from sttCore's format-switch when the model
// entry carries `transport: "gemini-live"` (registry) or the caller passes a
// transport string (custom models). Never keyed on a hardcoded model id here.
//
// Transport behavior:
// - Node >= 22 global WebSocket (undici). No new dependency.
// - Live API expects low-latency PCM; other containers are forwarded with
// their declared MIME unchanged (provider-side rejection is surfaced).
// - Text accumulation is append-only over inputTranscription segments and
// ends on serverContent.turnComplete (or graceful close with partial text).
// - Transcription deltas are kept per-frame (chunks[]) so sttCore can shape
// verbose_json segments without fabricating timestamps. goAway advisements
// rotate the socket once per call: setup replay + byte-offset resume.
const SETUP_TIMEOUT_MS = 10_000; // open → setupComplete
const TURN_TIMEOUT_MS = 60_000; // audio streamed → turnComplete
const MAX_TIMEOUT_MS = 300_000; // clamp ceiling for client-supplied lifecycle knobs
const CHUNK_BYTES = 16_384; // ~0.5s of 16-bit 16kHz mono PCM
const GOAWAY_RECONNECTS = 1; // socket rotations honoured per call
class GeminiLiveError extends Error {
constructor(message, status) {
super(message);
this.name = "GeminiLiveError";
this.status = status || 502;
}
}
// REST base (https://host/v1beta/models) → Live WS base
// (wss://host/ws/api/v1beta/models), then the bidiGenerateContent endpoint.
function toLiveWsUrl(baseUrl, model, token) {
const url = new URL(baseUrl);
url.protocol = "wss:";
if (!url.pathname.startsWith("/ws/")) url.pathname = `/ws/api${url.pathname}`;
const base = url.toString().replace(/\/+$/, "");
return `${base}/${encodeURIComponent(model)}:bidiGenerateContent?key=${encodeURIComponent(token || "")}`;
}
// Bind socket events supporting BOTH handler styles: addEventListener
// (browser WebSocket, undici) and onopen/onmessage property assignment
// (minimal polyfills). Whichever the implementation exposes, it works.
function bindSocket(ws, { onOpen, onMessage, onError, onClose }) {
if (typeof ws.addEventListener === "function") {
ws.addEventListener("open", onOpen);
ws.addEventListener("message", onMessage);
ws.addEventListener("error", onError);
ws.addEventListener("close", onClose);
return;
}
ws.onopen = onOpen;
ws.onmessage = onMessage;
ws.onerror = onError;
ws.onclose = onClose;
}
function parseFrame(data) {
try {
return JSON.parse(typeof data === "string" ? data : String(data));
} catch {
return null; // non-JSON frames carry no Live API semantics
}
}
function firstStringField(formData, key) {
const v = typeof formData?.get === "function" ? formData.get(key) : null;
return typeof v === "string" && v.trim() ? v.trim() : "";
}
// Lifecycle knobs the live registry entry advertises in params[]
// (setup/turn timeouts). They ride the same formData pass-through sttCore
// gives every transport — no sttCore change needed to reach this leaf.
function firstNumberField(formData, key, fallback) {
const n = Number(firstStringField(formData, key));
return Number.isFinite(n) && n > 0 ? Math.min(n, MAX_TIMEOUT_MS) : fallback;
}
/**
* Transcribe an audio File via the Gemini Live bidirectional stream.
* @returns {Promise<{text: string, chunks: string[]}>} transcript plus the raw
* incremental inputTranscription deltas (sttCore shapes verbose_json from them).
* @throws {GeminiLiveError} with .status for the sttCore error envelope.
*/
export async function transcribeGeminiLive({ cfg, file, model, token, formData, mimeType }) {
const WS = globalThis.WebSocket;
if (!WS) throw new GeminiLiveError("Gemini Live transport needs global WebSocket (Node >= 22)", 502);
const buf = Buffer.from(await file.arrayBuffer());
if (!buf.length) throw new GeminiLiveError("Empty audio file", 400);
const instruction = firstStringField(formData, "prompt") || "Transcribe the spoken audio verbatim.";
const language = firstStringField(formData, "language");
const setupTimeoutMs = firstNumberField(formData, "setup_timeout_ms", SETUP_TIMEOUT_MS);
const turnTimeoutMs = firstNumberField(formData, "turn_timeout_ms", TURN_TIMEOUT_MS);
// system_instruction (registry param) overrides the built-in transcription
// directive wholesale; prompt/language only shape the default.
const instructionOverride = firstStringField(formData, "system_instruction");
const systemText = instructionOverride
|| (language ? `${instruction} Language: ${language}.` : instruction);
const wsUrl = toLiveWsUrl(cfg.baseUrl, model, token);
return await new Promise((resolve, reject) => {
let text = "";
const chunks = []; // raw inputTranscription deltas, shaped by sttCore
let settled = false;
let timer = null;
let goAwayTimer = null;
let ws = null;
let generation = 0; // socket identity: superseded closes never settle
let sentBytes = 0; // audio prefix already handed to the live socket
let goAwayReconnects = GOAWAY_RECONNECTS;
const arm = (ms, message) => {
if (timer) clearTimeout(timer);
timer = setTimeout(() => fail(new GeminiLiveError(message, 504)), ms);
};
const shutdown = () => {
if (timer) { clearTimeout(timer); timer = null; }
if (goAwayTimer) { clearTimeout(goAwayTimer); goAwayTimer = null; }
// ws is null until the first open() dials (and stays null when the
// constructor throws) — fail() runs shutdown() on that path.
if (!ws) return;
try {
if (ws.readyState === WS.OPEN || ws.readyState === WS.CONNECTING) ws.close(1000);
} catch { /* socket already dead — outcome is already settled */ }
};
const succeed = () => {
if (settled) return;
settled = true;
shutdown();
resolve({ text, chunks });
};
const fail = (err) => {
if (settled) return;
settled = true;
shutdown();
reject(err);
};
const send = (frame) => {
if (ws.readyState !== WS.OPEN) return false;
try {
ws.send(JSON.stringify(frame));
} catch {
return false; // socket died mid-send — streamAudioAndPrompt maps this to a 502
}
return true;
};
// Streams every byte not yet sent, then the flushing text turn. After a
// goAway rotation this resumes from sentBytes — no audio re-upload.
const streamAudioAndPrompt = () => {
for (let off = sentBytes; off < buf.length; off += CHUNK_BYTES) {
const mediaChunk = buf.subarray(off, off + CHUNK_BYTES).toString("base64");
if (!send({ realtimeInput: { mediaChunks: [{ mimeType, data: mediaChunk }] } })) {
fail(new GeminiLiveError("Gemini Live socket closed while streaming audio", 502));
return;
}
sentBytes = Math.min(off + CHUNK_BYTES, buf.length);
}
// Final user turn: flushes the recognizer and yields turnComplete.
send({ clientContent: { turns: [{ parts: [{ text: systemText }] }], turnComplete: true } });
};
// goAway: the server names the instant it will force-close this socket.
// Graceful play = rotate BEFORE the deadline: retire the live socket,
// dial a fresh one, replay setup, resume audio from sentBytes — text and
// chunks survive the hop. Once the advisory budget is spent a later
// goAway is left to the close path, which settles on partial transcript.
const scheduleGoAwayReconnect = (goAway) => {
if (settled || goAwayTimer || goAwayReconnects <= 0) return;
const deadline = Date.parse(typeof goAway?.time === "string" ? goAway.time : "");
const delay = Number.isFinite(deadline)
? Math.max(0, Math.min(deadline - Date.now(), setupTimeoutMs))
: 0;
goAwayTimer = setTimeout(() => {
goAwayTimer = null;
if (settled) return;
goAwayReconnects--;
generation++;
try { ws?.close(1000); } catch { /* deadline crossed mid-flight — re-dial anyway */ }
open();
}, delay);
};
const open = () => {
const gen = ++generation;
try {
ws = new WS(wsUrl);
} catch {
fail(new GeminiLiveError("Gemini Live websocket connection failed", 502));
return;
}
bindSocket(ws, {
onOpen: () => {
if (settled || gen !== generation) return;
arm(setupTimeoutMs, "Gemini Live timed out waiting for setupComplete");
send({
setup: {
model: `models/${model}`,
generationConfig: {
responseModalities: ["TEXT"],
inputAudioTranscription: {},
},
systemInstruction: { parts: [{ text: systemText }] },
},
});
},
onMessage: (ev) => {
if (settled || gen !== generation) return;
const frame = parseFrame(ev?.data);
if (!frame) return;
if (frame.error) {
const e = frame.error;
fail(new GeminiLiveError(`Gemini Live error${e.status ? ` (${e.status})` : ""}: ${e.message || "unknown"}`, 502));
return;
}
if (frame.goAway) {
scheduleGoAwayReconnect(frame.goAway);
return;
}
const sc = frame.serverContent;
if (!sc) return;
const delta = typeof sc.inputTranscription?.text === "string" ? sc.inputTranscription.text : "";
// Trim before testing: a padding-only frame carries no transcript and
// must not make an empty run look like a partial success on close.
if (delta.trim()) {
text += delta;
chunks.push(delta);
}
if (sc.setupComplete) {
arm(turnTimeoutMs, "Gemini Live transcription timed out");
streamAudioAndPrompt();
return;
}
if (sc.turnComplete) succeed();
},
onError: () => {
if (settled || gen !== generation) return;
fail(new GeminiLiveError("Gemini Live websocket connection failed", 502));
},
onClose: (ev) => {
if (settled || gen !== generation) return;
// Partial transcript beats a hard error on graceful close; silence is one.
if (text.trim()) succeed();
else fail(new GeminiLiveError(`Gemini Live socket closed before completion${ev?.code ? ` (code ${ev.code})` : ""}`, 502));
},
});
};
open();
});
}

View File

@@ -2,13 +2,21 @@
import { randomUUID } from "node:crypto";
import { nowSec } from "./_base.js";
import { PROVIDERS } from "../../config/providers.js";
import { CODEX_CLI_VERSION } from "../../config/appConstants.js";
const CODEX_RESPONSES_URL = PROVIDERS["codex"].baseUrl;
const CODEX_USER_AGENT = "codex_cli_rs/0.136.0";
const CODEX_VERSION = "0.136.0";
const CODEX_USER_AGENT = `codex_cli_rs/${CODEX_CLI_VERSION}`;
const CODEX_ORIGINATOR = "codex_cli_rs";
const CODEX_MODEL_SUFFIX = "-image";
const CODEX_REF_DETAIL = "high";
const CODEX_IMAGES_MAIN_MODEL = "gpt-5.5";
const CODEX_TOOL_IMAGE_MODELS = new Set([
"gpt-image-1.5",
"gpt-image-2",
"gpt-image-2.5",
"gpt-image-2.5-flare",
"gpt-image-2.5-sunburst",
]);
function decodeAccountId(idToken) {
try {
@@ -27,6 +35,13 @@ function stripImageSuffix(model) {
return model.endsWith(CODEX_MODEL_SUFFIX) ? model.slice(0, -CODEX_MODEL_SUFFIX.length) : model;
}
function resolveCodexImageModels(model) {
if (CODEX_TOOL_IMAGE_MODELS.has(model)) {
return { responsesModel: CODEX_IMAGES_MAIN_MODEL, toolModel: model };
}
return { responsesModel: stripImageSuffix(model), toolModel: null };
}
function toDataUrl(input) {
if (!input || typeof input !== "string") return null;
if (/^data:image\//i.test(input) || /^https?:\/\//i.test(input)) return input;
@@ -157,7 +172,7 @@ export default {
"originator": CODEX_ORIGINATOR,
"session_id": randomUUID(),
"user-agent": CODEX_USER_AGENT,
"version": CODEX_VERSION,
"version": CODEX_CLI_VERSION,
"x-client-request-id": randomUUID(),
};
},
@@ -167,21 +182,26 @@ export default {
const single = toDataUrl(body.image);
if (single) refs.push(single);
const detail = body.image_detail || CODEX_REF_DETAIL;
const { responsesModel, toolModel } = resolveCodexImageModels(model);
const imgTool = { type: "image_generation", output_format: (body.output_format || "png").toLowerCase() };
if (toolModel) {
imgTool.action = refs.length > 0 ? "edit" : "generate";
imgTool.model = toolModel;
}
if (body.size && body.size !== "") imgTool.size = body.size;
if (body.quality && body.quality !== "") imgTool.quality = body.quality;
if (body.background && body.background !== "") imgTool.background = body.background;
return {
model: stripImageSuffix(model),
model: responsesModel,
instructions: "",
input: [{ type: "message", role: "user", content: buildContent(body.prompt, refs, detail) }],
tools: [imgTool],
tool_choice: "auto",
tool_choice: toolModel ? { type: "image_generation" } : "auto",
parallel_tool_calls: false,
prompt_cache_key: randomUUID(),
stream: true,
store: false,
reasoning: null,
reasoning: toolModel ? { effort: "medium", summary: "auto" } : null,
};
},
// Custom: codex parses SSE → either pipe to client or collect b64

View File

@@ -1,18 +1,93 @@
// HuggingFace Inference API — returns binary image
import { nowSec } from "./_base.js";
// HuggingFace Inference Providers router — returns binary image
//
// The router is a switchboard in front of many inference providers and is
// addressed as `<baseUrl>/<provider>/<providerModelId>`. `providerModelId` is
// the id the *provider* uses, which is not the Hub model id, so it is resolved
// through `imageConfig.modelMap` (built from the Hub API's
// inferenceProviderMapping and limited to providers the router forwards to).
//
// The legacy `api-inference.huggingface.co` host is gone (DNS ENOTFOUND) and is
// deliberately not referenced anywhere here.
import { nowSec, urlToBase64 } from "./_base.js";
import { PROVIDER_MEDIA } from "../../providers/index.js";
const BASE_URL = PROVIDER_MEDIA["huggingface"]?.imageConfig?.baseUrl;
const imageConfig = () => PROVIDER_MEDIA["huggingface"]?.imageConfig || {};
const BASE_URL = imageConfig().baseUrl;
const MODEL_MAP = imageConfig().modelMap || {};
// A plain-object lookup returns inherited truthy values for keys like "toString" or
// "constructor", which would build nonsense URLs. Resolve own keys only.
const lookup = (model) => (Object.hasOwn(MODEL_MAP, model) ? MODEL_MAP[model] : undefined);
// modelMap values are either a bare path (text-to-image) or { path, task }.
const mappingPath = (entry) => (typeof entry === "string" ? entry : entry.path);
const mappingTask = (entry) => (typeof entry === "string" ? "text-to-image" : entry.task || "text-to-image");
// A connection may point at its own endpoint (self-hosted Text Generation
// Inference / TGI container). That endpoint already knows its own model ids, so
// the router mapping does not apply and the Hub id is passed through verbatim.
function customBaseUrl(creds) {
const url = creds?.providerSpecificData?.baseUrl;
return typeof url === "string" && url.trim() ? url.trim().replace(/\/+$/, "") : null;
}
// The router's image-to-image payload wants raw base64 — not a data URL, not a URL.
// Accept every shape our own callers use (data URL, bare base64, remote URL, array).
async function sourceImage(body) {
const raw = body?.image || (Array.isArray(body?.images) ? body.images[0] : null);
if (typeof raw !== "string" || !raw.trim()) return null;
const value = raw.trim();
if (/^https?:\/\//i.test(value)) return await urlToBase64(value);
const match = /^data:image\/[^;]+;base64,(.+)$/i.exec(value);
return match ? match[1] : value;
}
export default {
buildUrl: (model) => `${BASE_URL}/${model}`,
buildUrl: (model, creds) => {
const override = customBaseUrl(creds);
if (override) {
// The model id is client-controlled; on a custom endpoint it lands in a URL
// path verbatim, so reject traversal/query injection (mirrors sttCore's guard).
if (model.includes("..") || model.includes("//") || /[?#]/.test(model)) {
throw new Error(`HuggingFace: invalid model ID "${model}"`);
}
return `${override}/${model}`;
}
const entry = lookup(model);
if (!entry) {
throw new Error(
`HuggingFace: no HuggingFace router mapping for model "${model}". ` +
`Add it to imageConfig.modelMap in open-sse/providers/registry/huggingface.js, ` +
`or set a custom base URL on the connection.`
);
}
return `${BASE_URL}/${mappingPath(entry)}`;
},
buildHeaders: (creds) => {
const headers = { "Content-Type": "application/json" };
const key = creds?.apiKey || creds?.accessToken;
if (key) headers["Authorization"] = `Bearer ${key}`;
return headers;
},
buildBody: (_model, body) => ({ inputs: body.prompt }),
buildBody: async (model, body) => {
const entry = lookup(model);
const task = mappingTask(entry || "");
if (task === "image-to-image") {
const image = await sourceImage(body);
if (!image) {
throw new Error(
`HuggingFace: model "${model}" requires a source image. ` +
`Send it as "image" (or "images") in the request body.`
);
}
// inputs carries the source image; the prompt moves under parameters.
return { inputs: image, parameters: { prompt: body.prompt } };
}
return { inputs: body.prompt };
},
// HF returns raw image bytes — convert to b64_json
async parseResponse(response) {
const buf = await response.arrayBuffer();

View File

@@ -1,5 +1,7 @@
import { Buffer } from "node:buffer";
import { createErrorResult } from "../utils/error.js";
import { transcribeGeminiLive } from "./geminiLiveStt.js";
import { PROVIDER_MODELS, PROVIDER_ID_TO_ALIAS } from "../config/providerModels.js";
import { HTTP_STATUS } from "../config/runtimeConfig.js";
// Build auth headers from sttConfig + token
@@ -162,11 +164,26 @@ function jsonResponse(obj) {
};
}
// Model-level transport marker (registry models[].transport, e.g. the Gemini
// live STT entry's "gemini-live", or a custom model's stored transport).
// Dispatch reads the marker — never a hardcoded model id — so new realtime
// providers extend sttCore through data, not code.
function resolveModelTransport(provider, model) {
const key = PROVIDER_ID_TO_ALIAS[provider] || provider;
const models = PROVIDER_MODELS[key] || PROVIDER_MODELS[provider];
if (!Array.isArray(models)) return null;
const entry = models.find((m) => m && m.id === model && (m.kind || "llm") === "stt");
const marker = typeof entry?.transport === "string" ? entry.transport.trim() : "";
return marker || null;
}
/**
* STT core handler — dispatch by sttConfig.format.
* STT core handler — dispatch by model transport marker, else sttConfig.format.
* `transport` is the caller-supplied marker override (custom models resolve
* it in the app layer; built-ins fall back to the registry entry marker).
* @returns {Promise<{success, response, status?, error?}>}
*/
export async function handleSttCore({ provider, model, formData, credentials, sttConfig }) {
export async function handleSttCore({ provider, model, formData, credentials, sttConfig, transport }) {
const file = formData.get("file");
if (!file) return createErrorResult(HTTP_STATUS.BAD_REQUEST, "Missing required field: file");
@@ -186,8 +203,29 @@ export async function handleSttCore({ provider, model, formData, credentials, st
return createErrorResult(HTTP_STATUS.UNAUTHORIZED, `No credentials for STT provider: ${provider}`);
}
// Format-switch extension: an explicit caller marker wins over the registry
// marker; with neither, the provider-default sttConfig.format applies.
const marker = (typeof transport === "string" && transport.trim()) ? transport.trim() : resolveModelTransport(provider, model);
try {
switch (cfg.format) {
switch (marker || cfg.format) {
case "gemini-live": {
const live = await transcribeGeminiLive({ cfg, file, model, token, formData, mimeType: resolveAudioContentType(file) });
// response_format parity with the OpenAI-compatible transport: default
// envelope stays {text}; verbose_json adds segments mapped from the
// Live API's incremental inputTranscription deltas. Those frames carry
// NO timestamps, so segments expose {id,text} only (id = delta order,
// Whisper-compatible 0-based) — start/end/duration are deliberately
// absent rather than fabricated as zeros, which would misrepresent
// provider data to callers diffing transports.
const fmt = typeof formData?.get === "function"
? String(formData.get("response_format") ?? "").trim().toLowerCase()
: "";
if (fmt === "verbose_json") {
return jsonResponse({ text: live.text, segments: live.chunks.map((segText, id) => ({ id, text: segText })) });
}
return jsonResponse({ text: live.text });
}
case "deepgram": return await transcribeDeepgram(cfg, file, model, token, formData);
case "assemblyai": return await transcribeAssemblyAI(cfg, file, model, token);
case "nvidia-asr": return await transcribeNvidia(cfg, file, model, token);
@@ -196,6 +234,6 @@ export async function handleSttCore({ provider, model, formData, credentials, st
default: return await transcribeOpenAICompatible(cfg, file, model, token, formData);
}
} catch (err) {
return createErrorResult(HTTP_STATUS.BAD_GATEWAY, err.message || "STT request failed");
return createErrorResult(err.status || HTTP_STATUS.BAD_GATEWAY, err.message || "STT request failed");
}
}

View File

@@ -0,0 +1,95 @@
import { createErrorResult, parseUpstreamError, formatProviderError } from "../utils/error.js";
import { HTTP_STATUS, FETCH_CONNECT_TIMEOUT_MS } from "../config/runtimeConfig.js";
import { PROVIDER_MEDIA } from "../providers/index.js";
import { generateSessionId } from "../executors/opencode-zen.js";
/**
* Core System One (Jev) handler — native decision payload pass-through.
* URL/headers come from the registry's systemoneConfig; body and JSON response
* are forwarded untouched (decision models have no chat translation layer).
*
* @returns {Promise<{ success: boolean, response: Response, usage?: object, status?: number, error?: string }>}
*/
export async function handleSystemoneCore({
body,
modelInfo,
credentials,
log,
onRequestSuccess,
}) {
const { provider, model } = modelInfo;
const cfg = PROVIDER_MEDIA[provider]?.systemoneConfig;
if (!cfg?.baseUrl) {
return createErrorResult(
HTTP_STATUS.BAD_REQUEST,
`Provider '${provider}' does not support System One.`
);
}
// Validate input at the trust boundary; question-level shape is upstream's job.
if (body.state === undefined || body.state === null) {
return createErrorResult(HTTP_STATUS.BAD_REQUEST, "Missing required field: state");
}
if (!body.questions || typeof body.questions !== "object" || Array.isArray(body.questions)) {
return createErrorResult(HTTP_STATUS.BAD_REQUEST, "Missing required field: questions");
}
// noAuth free lanes carry accessToken "public" from the credential stub.
const token = credentials?.apiKey || credentials?.accessToken;
const headers = {
"Content-Type": "application/json",
...(token ? { Authorization: `Bearer ${token}` } : {}),
...(cfg.headers || {}),
// Zen lanes expect the official client session header on every request.
"x-opencode-session": generateSessionId(),
};
const requestBody = { ...body, model };
log?.debug?.("SYSTEMONE", `${provider.toUpperCase()} | ${model}`);
let providerResponse;
try {
providerResponse = await fetch(cfg.baseUrl, {
method: "POST",
headers,
body: JSON.stringify(requestBody),
...(typeof AbortSignal?.timeout === "function"
? { signal: AbortSignal.timeout(FETCH_CONNECT_TIMEOUT_MS) }
: {}),
});
} catch (error) {
const errMsg = formatProviderError(error, provider, model, HTTP_STATUS.BAD_GATEWAY);
log?.debug?.("SYSTEMONE", `Fetch error: ${errMsg}`);
return createErrorResult(HTTP_STATUS.BAD_GATEWAY, errMsg);
}
if (!providerResponse.ok) {
const { statusCode, message } = await parseUpstreamError(providerResponse);
const errMsg = formatProviderError(new Error(message), provider, model, statusCode);
log?.debug?.("SYSTEMONE", `Provider error: ${errMsg}`);
return createErrorResult(statusCode, errMsg);
}
let responseBody;
try {
responseBody = await providerResponse.json();
} catch {
return createErrorResult(HTTP_STATUS.BAD_GATEWAY, `Invalid JSON response from ${provider}`);
}
if (onRequestSuccess) await onRequestSuccess();
const usage = responseBody?.usage;
return {
success: true,
usage: usage
? { prompt_tokens: usage.input_tokens || 0, completion_tokens: usage.output_tokens || 0 }
: null,
response: new Response(JSON.stringify(responseBody), {
headers: {
"Content-Type": "application/json",
"Access-Control-Allow-Origin": "*",
},
}),
};
}

View File

@@ -2,6 +2,7 @@ import { createErrorResult } from "../utils/error.js";
import { HTTP_STATUS } from "../config/runtimeConfig.js";
import { refreshTokenByProvider } from "../services/tokenRefresh.js";
import { PROVIDER_MEDIA } from "../providers/index.js";
import { getVideoAdapter } from "./videoProviders/index.js";
// Upstream fetch deadline for video job submission/polling (the job itself is
// async upstream — this only bounds the HTTP round-trip, not video rendering).
@@ -94,21 +95,49 @@ export async function handleVideoProxyCore({
return createErrorResult(HTTP_STATUS.BAD_REQUEST, `Unknown video action: ${action}`);
}
const method = requestId ? "GET" : "POST";
const url = buildUpstreamUrl(config, action, requestId);
const adapter = getVideoAdapter(provider);
const fetchSignal = combineSignals(signal, timeoutMs);
const doFetch = (token) =>
fetch(url, {
// Default (xAI shape) request plan; adapters override URL/method/headers/body.
const defaultPlan = () => {
const method = requestId ? "GET" : "POST";
return {
method,
headers: buildHeaders({ token, contentType: method === "POST" ? contentType : null, idempotencyKey: method === "POST" ? idempotencyKey : null }),
url: buildUpstreamUrl(config, action, requestId),
headers: buildHeaders({
token: credentials?.accessToken || credentials?.apiKey,
contentType: method === "POST" ? contentType : null,
idempotencyKey: method === "POST" ? idempotencyKey : null,
}),
body: method === "POST" ? rawBody : undefined,
signal: fetchSignal,
});
};
};
// Rebuilt per attempt so the auth retry below picks up the refreshed token.
const doFetch = async () => {
const plan = adapter
? await adapter.buildRequest({
config, action, requestId, rawBody, contentType, idempotencyKey, credentials, log,
token: credentials?.accessToken || credentials?.apiKey,
})
: defaultPlan();
if (plan.error) return { planError: plan.error };
return {
response: await fetch(plan.url, {
method: plan.method,
headers: plan.headers,
body: plan.body,
signal: fetchSignal,
}),
};
};
const method = requestId ? "GET" : "POST";
let upstream;
try {
upstream = await doFetch(credentials?.accessToken || credentials?.apiKey);
const first = await doFetch();
if (first.planError) return createErrorResult(HTTP_STATUS.BAD_REQUEST, `[${provider}] ${first.planError}`);
upstream = first.response;
} catch (error) {
if (error?.name === "AbortError" || error?.name === "TimeoutError") {
return createErrorResult(HTTP_STATUS.REQUEST_TIMEOUT, `[${provider}] video ${method} aborted: ${error.message}`);
@@ -136,7 +165,9 @@ export async function handleVideoProxyCore({
await upstream.body?.cancel?.();
} catch { /* noop */ }
try {
upstream = await doFetch(credentials.accessToken || credentials.apiKey);
const retry = await doFetch();
if (retry.planError) return createErrorResult(HTTP_STATUS.BAD_REQUEST, `[${provider}] ${retry.planError}`);
upstream = retry.response;
} catch (error) {
return createErrorResult(HTTP_STATUS.BAD_GATEWAY, sanitizeSecrets(`[${provider}] video retry after refresh failed: ${error.message}`, credentials));
}
@@ -152,13 +183,25 @@ export async function handleVideoProxyCore({
return createErrorResult(upstream.status, `[${provider}] ${message.slice(0, 2000)}`);
}
// Success: pass the upstream JSON through untouched (request_id / status / video.url).
// Success: pass the upstream JSON through untouched (request_id / status / video.url),
// unless the adapter maps a provider-native shape onto it (Vertex operations).
let outBody = bodyText;
let outType = upstream.headers.get("content-type") || "application/json";
if (adapter?.transformResponse) {
try {
outBody = JSON.stringify(adapter.transformResponse(JSON.parse(bodyText)));
outType = "application/json";
} catch {
// Non-JSON or unexpected shape — fall back to the raw upstream body.
}
}
return {
success: true,
response: new Response(bodyText, {
response: new Response(outBody, {
status: upstream.status,
headers: {
"Content-Type": upstream.headers.get("content-type") || "application/json",
"Content-Type": outType,
"Access-Control-Allow-Origin": "*",
},
}),

View File

@@ -0,0 +1,13 @@
// Video provider adapters.
//
// Default (no adapter) = xAI shape: raw body forwarded to {baseUrl}/{action},
// polled at {baseUrl}/{id}, upstream JSON passed through verbatim.
// A provider only needs an adapter when its wire format differs from that.
import openrouter from "./openrouter.js";
import vertex from "./vertex.js";
const ADAPTERS = { openrouter, vertex };
export function getVideoAdapter(provider) {
return ADAPTERS[provider] || null;
}

View File

@@ -0,0 +1,39 @@
// OpenRouter video jobs — https://openrouter.ai/docs/api/api-reference/videos
//
// Same async shape as xAI (POST → { id, status }, GET → status/unsigned_urls),
// two differences only: creation POSTs to the collection root (no `/generations`
// suffix) and the account headers come from the registry entry.
// Response bodies are passed through verbatim.
// ponytail: generations only — OpenRouter has no edits/extensions endpoint today.
const SUPPORTED_ACTIONS = new Set(["generations"]);
function headers(config, token) {
return {
Accept: "application/json",
...(config.headers || {}),
...(token ? { Authorization: `Bearer ${token}` } : {}),
};
}
export default {
buildRequest({ config, action, requestId, rawBody, contentType, token }) {
const base = config.baseUrl.replace(/\/$/, "");
if (requestId) {
return { method: "GET", url: `${base}/${encodeURIComponent(requestId)}`, headers: headers(config, token) };
}
if (!SUPPORTED_ACTIONS.has(action)) {
return { error: `OpenRouter video supports 'generations' only (got '${action}')` };
}
if (contentType && !contentType.includes("application/json")) {
return { error: "OpenRouter video requires an application/json body" };
}
return {
method: "POST",
url: base,
headers: { ...headers(config, token), "Content-Type": "application/json" },
body: rawBody,
};
},
};

View File

@@ -0,0 +1,159 @@
// Vertex AI (Veo) video jobs.
//
// Vertex does NOT speak the OpenAI-ish /v1/videos shape, so unlike OpenRouter
// this adapter translates both directions:
// create → POST {model}:predictLongRunning { instances[], parameters{} } → { name }
// poll → POST {model}:fetchPredictOperation { operationName } → { done, response }
// Docs: https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/veo-video-generation
//
// The operation name is a resource path (contains "/"), so it is base64url-encoded
// into the job id returned to the client — GET /v1/videos/{id} stays a flat path.
import { parseVertexSaJson, refreshVertexToken } from "../../services/tokenRefresh.js";
const DEFAULT_LOCATION = "us-central1";
const encodeJobId = (name) => Buffer.from(name, "utf8").toString("base64url");
// Operation name shape: projects/{p}/locations/{l}/publishers/{pub}/models/{m}/operations/{op}.
// Anchored and single-segment-per-field so a decoded path can never carry `..` or a
// host-changing prefix into the request URL.
const OPERATION_NAME_RE = /^projects\/[^/]+\/locations\/[^/]+\/publishers\/[^/]+\/models\/[^/]+\/operations\/[^/]+$/;
function modelPathOf(operationName) {
return operationName.slice(0, operationName.indexOf("/operations/"));
}
function decodeJobId(id) {
const raw = String(id ?? "");
// Buffer.from(x, "base64url") silently drops invalid characters instead of
// throwing, so only ids that re-encode byte-for-byte are accepted.
if (!raw || raw.length > 1024 || !/^[A-Za-z0-9_-]+$/.test(raw)) return null;
const decoded = Buffer.from(raw, "base64url").toString("utf8");
if (Buffer.from(decoded, "utf8").toString("base64url") !== raw) return null;
return OPERATION_NAME_RE.test(decoded) ? decoded : null;
}
async function resolveAuth(credentials, log) {
const saJson = parseVertexSaJson(credentials?.apiKey);
const projectId =
saJson?.project_id ||
credentials?.projectId ||
credentials?.providerSpecificData?.projectId;
const location = credentials?.providerSpecificData?.location || DEFAULT_LOCATION;
if (!projectId) {
return { error: "Vertex video requires a project_id — use Service Account JSON or set providerSpecificData.projectId" };
}
let token = credentials?.accessToken;
if (saJson) {
const minted = await refreshVertexToken(saJson, log);
if (!minted?.accessToken) return { error: "Vertex video: failed to mint access token from service account JSON" };
token = minted.accessToken;
}
if (!token) return { error: "Vertex video requires Service Account JSON or an OAuth access token (raw API keys are not supported)" };
return { token, projectId, location };
}
/** OpenAI-ish video body → Vertex predictLongRunning body. */
function toVertexBody(body) {
const instance = { prompt: body.prompt };
// Image-to-video: accept the Vertex-native shape or a bare data URL / base64 string.
const image = body.image ?? body.image_url;
if (image && typeof image === "object") {
instance.image = image;
} else if (typeof image === "string") {
const match = image.match(/^data:([^;]+);base64,(.*)$/s);
instance.image = match
? { bytesBase64Encoded: match[2], mimeType: match[1] }
: { gcsUri: image };
}
if (body.video && typeof body.video === "object") instance.video = body.video;
const parameters = {};
if (body.n != null) parameters.sampleCount = Number(body.n);
if (body.duration != null) parameters.durationSeconds = Number(body.duration);
if (body.aspect_ratio) parameters.aspectRatio = body.aspect_ratio;
if (body.resolution) parameters.resolution = body.resolution;
if (body.seed != null) parameters.seed = body.seed;
if (body.negative_prompt) parameters.negativePrompt = body.negative_prompt;
// Without storageUri Vertex returns inline base64 bytes; a GCS bucket keeps
// the poll response small and is what production callers want.
if (body.storage_uri) parameters.storageUri = body.storage_uri;
if (body.generate_audio != null) parameters.generateAudio = !!body.generate_audio;
return { instances: [instance], ...(Object.keys(parameters).length ? { parameters } : {}) };
}
/** Vertex operation → the async-job shape 9Router clients already poll for. */
function fromVertexOperation(json) {
if (!json?.name) return json;
const id = encodeJobId(json.name);
if (json.error) {
return { id, request_id: id, status: "failed", error: json.error };
}
if (!json.done) {
return { id, request_id: id, status: "pending" };
}
const samples =
json.response?.videos ||
json.response?.generateVideoResponse?.generatedSamples ||
[];
const videos = samples.map((s) => ({
url: s.gcsUri || s.video?.uri || s.uri || null,
b64_json: s.bytesBase64Encoded || s.video?.bytesBase64Encoded || null,
mime_type: s.mimeType || s.video?.mimeType || "video/mp4",
}));
return { id, request_id: id, status: "completed", video: videos[0] || null, videos };
}
export default {
async buildRequest({ config, action, requestId, rawBody, contentType, credentials, log }) {
if (contentType && !contentType.includes("application/json")) {
return { error: "Vertex video requires an application/json body" };
}
const auth = await resolveAuth(credentials, log);
if (auth.error) return { error: auth.error };
const { token, projectId, location } = auth;
const base = (config.baseUrl || "https://aiplatform.googleapis.com").replace(/\/$/, "");
const headers = { Accept: "application/json", "Content-Type": "application/json", Authorization: `Bearer ${token}` };
if (requestId) {
const operationName = decodeJobId(requestId);
if (!operationName) return { error: "Invalid Vertex video job id" };
return {
method: "POST",
url: `${base}/v1/${modelPathOf(operationName)}:fetchPredictOperation`,
headers,
body: JSON.stringify({ operationName }),
};
}
if (action !== "generations") {
// ponytail: Veo extend/edit go through generations with `video`/`image` in the body.
return { error: `Vertex video supports 'generations' only (got '${action}')` };
}
let body;
try {
body = JSON.parse(typeof rawBody === "string" ? rawBody : rawBody.toString("utf8"));
} catch {
return { error: "Invalid JSON body" };
}
if (!body.model) return { error: "Vertex video requires a model (e.g. vertex/veo-3.1-generate-preview)" };
// Plain model id only — a path segment carrying "/" or ".." would rewrite the URL.
if (!/^[A-Za-z0-9._-]+$/.test(body.model)) return { error: "Invalid Vertex video model id" };
if (!body.prompt && !body.image && !body.image_url) return { error: "Vertex video requires a prompt or an image" };
return {
method: "POST",
url: `${base}/v1/projects/${projectId}/locations/${location}/publishers/google/models/${body.model}:predictLongRunning`,
headers,
body: JSON.stringify(toVertexBody(body)),
};
},
transformResponse: fromVertexOperation,
};

View File

@@ -112,10 +112,22 @@ export const MODEL_CAPABILITIES = {
"glm-5.3-flash": { vision: true, videoInput: true, pdf: true, reasoning: true, thinkingFormat: "zai", contextWindow: 1000000, maxOutput: 131072 },
"glm-4.6v": { vision: true, videoInput: true, reasoning: true, thinkingFormat: "zai", contextWindow: 128000, maxOutput: 32768 },
"glm-4.5v": { vision: true, videoInput: true, reasoning: true, thinkingFormat: "zai", contextWindow: 64000, maxOutput: 16384 },
// GLM-5.2 has 1M context — pattern *glm-5* only gives 200k, so override here
"glm-5.2": { reasoning: true, thinkingFormat: "zai", thinkingCanDisable: false, contextWindow: 1000000, maxOutput: 131072 },
// DeepSeek's first V4 model with image input; text limits match V4-Flash.
"deepseek-v4-flash-vision-exp": { vision: true, reasoning: true, thinkingFormat: "deepseek", contextWindow: 1000000, maxOutput: 384000 },
// DeepSeek V4.1-Flash is natively multimodal — models.dev lists
// opencode-go/deepseek-v4.1-flash with modalities.input ["text","image"] — and upstream
// the retired v4-flash / vision-exp ids route to it, so the live V4.1 ids carry the
// same image capability as the exp id above. "deepseek-flash" is the GA id on the
// DeepSeek API; it previously fell through to the generic *deepseek* pattern, whose
// 128K/64K limits are kept here. The repeated fields are deliberate: an exact entry
// short-circuits the pattern table, so a vision-only delta would drop them.
"deepseek-v4.1-flash": { vision: true, reasoning: true, thinkingFormat: "deepseek", contextWindow: 1000000, maxOutput: 384000 },
"deepseek-flash": { vision: true, reasoning: true, thinkingFormat: "deepseek", contextWindow: 128000, maxOutput: 64000 },
// Qwen plain coder/text (no vision) — registry "vision-model" / "coder-model" aliases
"vision-model": { vision: true, reasoning: true, thinkingFormat: "qwen", contextWindow: 1000000 },
"coder-model": { reasoning: true, thinkingFormat: "qwen", contextWindow: 1000000 },
@@ -131,6 +143,8 @@ export const MODEL_CAPABILITIES = {
// via OpenAI Responses input_image; reasoning supports up to xhigh.
"muse-spark-1.2-contributor-free": { vision: true, reasoning: true, thinkingFormat: "openai", contextWindow: 1048576, maxOutput: 131072 },
"muse-spark-1.3-contributor-free": { vision: true, reasoning: true, thinkingFormat: "openai", contextWindow: 1048576, maxOutput: 131072 },
// OpenCode Free Union Alpha — multimodal (text+vision), 262K context, 131K max output
"union-alpha": { vision: true, contextWindow: 262144, maxOutput: 131072 },
};
const KIRO_GPT_5_6_CAPABILITIES = { vision: true, reasoning: true, search: true, thinkingFormat: "openai", contextWindow: 272000, maxOutput: 128000 };
@@ -153,6 +167,12 @@ export const PROVIDER_CAPABILITIES = {
"deepseek-ai/deepseek-v4-pro": { reasoning: true, thinkingFormat: "openai", contextWindow: 1000000, maxOutput: 65536 },
"deepseek-ai/deepseek-v4-flash": { reasoning: true, thinkingFormat: "openai", contextWindow: 1000000, maxOutput: 65536 },
},
// glm-5.3-flash on OpenCode Go is served by a backend that rejects the z.ai
// `thinking` object (400: unknown field "thinking") and wants reasoning_effort.
// Overrides the global entry, whose z.ai shape is correct for z.ai itself.
"opencode-go": {
"glm-5.3-flash": { vision: true, videoInput: true, pdf: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 1000000, maxOutput: 131072 },
},
"codex": {
"gpt-6-astra": { vision: true, reasoning: true, search: true, thinkingFormat: "openai", contextWindow: 272000, maxOutput: 128000 },
"gpt-5.6-sol": CODEX_GPT_56_SOL_CAPS,
@@ -193,6 +213,10 @@ export const PROVIDER_CAPABILITIES = {
"minimax-m3": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 512000, maxOutput: 128000 },
"kimi-k2.7": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 256000, maxOutput: 32000 },
"kimi-k2.6": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 256000, maxOutput: 32000 },
"kimi-k2.5": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 164000, maxOutput: 32000 },
"hy3-preview": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 192000, maxOutput: 64000 },
"deepseek-v4-flash": { reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 1000000, maxOutput: 50000 },
"deepseek-v3-2-volc": { reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 96000, maxOutput: 32000 },
// Per-model values mirror the server's product-config payload (the plugin
// fetches it from copilot.tencent.com; the `models[]` entries carry
// maxInputTokens/maxOutputTokens/supportsImages). contextWindow =
@@ -209,47 +233,32 @@ export const PROVIDER_CAPABILITIES = {
"glm-5.3-flash": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: true, contextWindow: 1000000, maxOutput: 32000 },
"kimi-k3-1": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 1000000, maxOutput: 32000 },
"deepseek-v4-pro": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: true, contextWindow: 1000000, maxOutput: 50000 },
"deepseek-v4-flash": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: true, contextWindow: 1000000, maxOutput: 50000 },
},
// Qoder — upstream exposes opaque internal ids (dfmodel, kmodel, …); the
// registry `name` is display-only and capability lookup matches on the raw
// id, so every qoder model would fall through to DEFAULT_CAPABILITIES
// (200K) without this map. contextWindow follows the real model family's
// spec: the /algo/api/v2/model/list max_input_tokens under-reports some
// windows (GLM-5.3 / Kimi-K3 / Qwen3.8-Max claim 180K but accept more).
// max_output_tokens arrives as 0 for every model, so outputs are
// best-guess from the real model family. Vision tags below follow the
// upstream is_vl flag per explicit request, even though the executor
// currently sends image_urls:null (image pass-through over the agent_chat
// SSE protocol is unverified). reasoning:true on all of them — every model can
// reason; the upstream is_reasoning flag only drives model_config selection.
// thinkingFormat keeps the true-model family for documentation/UI, but
// thinkingCanDisable:false everywhere: the executor only forwards
// messages/tools/max_tokens, and thinking is fixed upstream via
// modelConfig.is_reasoning — client thinking intent is dropped, so "none"
// must never be offered as an option.
"qoder": {
"ultimate": { vision: true, reasoning: true, thinkingFormat: "claude-adaptive", thinkingCanDisable: false, contextWindow: 1000000, maxOutput: 128000 }, // Claude Opus 5
"performance": { vision: true, reasoning: true, thinkingFormat: "claude-adaptive", thinkingCanDisable: false, contextWindow: 1000000, maxOutput: 128000 }, // Claude Sonnet 5
"dmodel": { reasoning: true, thinkingFormat: "deepseek", thinkingCanDisable: false, contextWindow: 1000000, maxOutput: 65536 }, // DeepSeek-V4-Pro
"dfmodel": { reasoning: true, thinkingFormat: "deepseek", thinkingCanDisable: false, contextWindow: 1000000, maxOutput: 65536 }, // DeepSeek-V4-Flash
"gmodel": { reasoning: true, thinkingFormat: "zai", thinkingCanDisable: false, contextWindow: 1000000, maxOutput: 128000 }, // GLM-5.3
"gfmodel": { vision: true, reasoning: true, thinkingFormat: "zai", thinkingCanDisable: false, contextWindow: 1000000, maxOutput: 128000 }, // GLM-5.3-Flash
"kmodel_latest": { vision: true, reasoning: true, thinkingFormat: "kimi", thinkingCanDisable: false, contextWindow: 1000000, maxOutput: 65536 }, // Kimi-K3
"kmodel": { vision: true, reasoning: true, thinkingFormat: "kimi", thinkingCanDisable: false, contextWindow: 256000, maxOutput: 65536 }, // Kimi-K2.7-Code
"mmodel": { reasoning: true, thinkingFormat: "minimax", thinkingCanDisable: false, contextWindow: 1000000, maxOutput: 512000 }, // MiniMax-M3
"qmodel_latest": { vision: true, reasoning: true, thinkingFormat: "qwen", thinkingCanDisable: false, contextWindow: 1000000, maxOutput: 65536 }, // Qwen3.7-Max
"qmodel": { vision: true, reasoning: true, thinkingFormat: "qwen", thinkingCanDisable: false, contextWindow: 1000000, maxOutput: 65536 }, // Qwen3.7-Plus
"qfmodel": { vision: true, reasoning: true, thinkingFormat: "qwen", thinkingCanDisable: false, contextWindow: 1000000, maxOutput: 65536 }, // Qwen3.8-Flash
"qmodel_38max": { vision: true, reasoning: true, thinkingFormat: "qwen", thinkingCanDisable: false, contextWindow: 1000000, maxOutput: 65536 }, // Qwen3.8-Max
// deepseek-v4.1-flash replaces v4-flash (dropped from the server list;
// the old endpoint still answers 200 but the published list is the
// contract). maxOutput 128000 per the server's product-config payload.
"deepseek-v4.1-flash": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: true, contextWindow: 1000000, maxOutput: 128000 },
},
// Poolside Laguna — OpenAI-compatible, all reasoning-capable (32K max output).
"poolside": {
"laguna-s-2.1": { reasoning: true, thinkingFormat: "openai", contextWindow: 1000000, maxOutput: 32000 },
"laguna-xs-2.1": { reasoning: true, thinkingFormat: "openai", contextWindow: 200000, maxOutput: 32000 },
},
// Ollama Cloud — the generic *deepseek-v4* pattern misses the vision badge
// the library page publishes for this model (text+image in, 1M context).
// ponytail: thinkingFormat stays "deepseek" to preserve today's body shape;
// Ollama's native toggle is the top-level `think` field (bool or
// low/medium/high/max), which no format in thinkingUnified.js emits yet —
// openai-to-ollama.js drops it. Wire a "think" format when thinking on
// Ollama Cloud is actually needed.
"ollama": {
"deepseek-v4.1-flash:cloud": { vision: true, reasoning: true, thinkingFormat: "deepseek", contextWindow: 1000000, maxOutput: 384000 },
},
};
// Qoder CN serves the identical model catalog from the CN gateway, so it shares
// the intl Qoder capability table verbatim (vision/reasoning/contextWindow).
PROVIDER_CAPABILITIES["qoder-cn"] = PROVIDER_CAPABILITIES["qoder"];
/**
* Pattern fallback — glob (* = wildcard), matched case-insensitively and
* anchored (^...$) so a pattern must match the full model id. ORDER MATTERS:
@@ -319,7 +328,7 @@ export const PATTERN_CAPABILITIES = [
{ pattern: "*qwen*vl*", caps: { vision: true, reasoning: true, thinkingFormat: "qwen", contextWindow: 262144 } },
{ pattern: "*qwen*omni*", caps: { vision: true, audioInput: true, videoInput: true, reasoning: true, thinkingFormat: "qwen", contextWindow: 262144, maxOutput: 65536 } },
{ pattern: "*qwen*coder*", caps: { reasoning: true, thinkingFormat: "qwen", contextWindow: 1000000 } },
{ pattern: "*qwen*max*", caps: { reasoning: true, thinkingFormat: "qwen", contextWindow: 1000000, maxOutput: 65536 } },
{ pattern: "*qwen*max*", caps: { vision: true, reasoning: true, thinkingFormat: "qwen", contextWindow: 1000000, maxOutput: 65536 } },
{ pattern: "*qwen3.5*", caps: { vision: true, videoInput: true, reasoning: true, thinkingFormat: "qwen", contextWindow: 1000000, maxOutput: 65536 } },
{ pattern: "*qwen3.6*", caps: { vision: true, videoInput: true, reasoning: true, thinkingFormat: "qwen", contextWindow: 1000000, maxOutput: 65536 } },
{ pattern: "*qwen3.7*", caps: { vision: true, videoInput: true, reasoning: true, thinkingFormat: "qwen", contextWindow: 1000000, maxOutput: 65536 } },
@@ -346,7 +355,11 @@ export const PATTERN_CAPABILITIES = [
{ pattern: "*glm*", caps: { reasoning: true, thinkingFormat: "zai", contextWindow: 200000 } },
// ── DeepSeek (thinking.enabled + reasoning_effort; r1 = thinking-only) ─
{ pattern: "*deepseek-v4*", caps: { reasoning: true, thinkingFormat: "deepseek", contextWindow: 1000000, maxOutput: 384000 } },
// v4.1+ has real image input (probed live on Alibaba MaaS: correct color
// read from a PNG). v4-pro / v4-flash-0731 accept image blocks but ignore
// them (answered "Unknown"), so vision stays scoped to v4.* dotted releases.
{ pattern: "*deepseek-v4.*", caps: { vision: true, reasoning: true, thinkingFormat: "deepseek", thinkingEffortSupported: true, contextWindow: 1000000, maxOutput: 128000 } },
{ pattern: "*deepseek-v4*", caps: { reasoning: true, thinkingFormat: "deepseek", thinkingEffortSupported: true, contextWindow: 1000000, maxOutput: 384000 } },
{ pattern: "*reasoner*", caps: { reasoning: true, thinkingFormat: "deepseek", thinkingCanDisable: false, contextWindow: 128000 } },
{ pattern: "*deepseek-r*", caps: { reasoning: true, thinkingFormat: "deepseek", thinkingCanDisable: false, contextWindow: 128000 } },
{ pattern: "*deepseek-chat*", caps: { contextWindow: 128000 } },
@@ -354,14 +367,16 @@ export const PATTERN_CAPABILITIES = [
// ── MiniMax (M3 = adaptive; M2.x cannot disable) ─────────────────
{ pattern: "*minimax*image*", caps: { imageOutput: true } },
{ pattern: "*minimax-m3*", caps: { vision: true, reasoning: true, thinkingFormat: "minimax", contextWindow: 1048576, maxOutput: 512000 } },
{ pattern: "*minimax-m2.7*", caps: { reasoning: true, thinkingFormat: "minimax", thinkingCanDisable: false, contextWindow: 204800, maxOutput: 131072 } },
{ pattern: "*minimax-m3*", caps: { vision: true, reasoning: true, thinkingFormat: "minimax", contextWindow: 1000000, maxOutput: 131072 } },
{ pattern: "*minimax-m2.7*", caps: { vision: true, reasoning: true, thinkingFormat: "minimax", thinkingCanDisable: false, contextWindow: 204800, maxOutput: 131072 } },
{ pattern: "*minimax-m2.5*", caps: { vision: true, reasoning: true, thinkingFormat: "minimax", thinkingCanDisable: false, contextWindow: 204800, maxOutput: 131072 } },
{ pattern: "*minimax*", caps: { reasoning: true, thinkingFormat: "minimax", thinkingCanDisable: false, contextWindow: 200000, maxOutput: 131072 } },
// ── Xiaomi MiMo (vision, 1M / 262K ctx) ──────────────────────────
{ pattern: "*mimo*v2.5*", caps: { vision: true, audioInput: true, videoInput: true, contextWindow: 1048576, maxOutput: 131072 } },
{ pattern: "*mimo*omni*", caps: { vision: true, audioInput: true, contextWindow: 262144, maxOutput: 131072 } },
{ pattern: "*mimo*", caps: { vision: true, contextWindow: 262144, maxOutput: 131072 } },
// ── Xiaomi MiMo (vision + <think>-tag reasoning, always-on, can't disable) ──
{ pattern: "*mimo*v2.6*", caps: { vision: true, audioInput: true, videoInput: true, reasoning: true, thinkingFormat: "deepseek", thinkingCanDisable: false, contextWindow: 1048576, maxOutput: 131072 } },
{ pattern: "*mimo*v2.5*", caps: { vision: true, audioInput: true, videoInput: true, reasoning: true, thinkingFormat: "deepseek", thinkingCanDisable: false, contextWindow: 1048576, maxOutput: 131072 } },
{ pattern: "*mimo*omni*", caps: { vision: true, audioInput: true, reasoning: true, thinkingFormat: "deepseek", thinkingCanDisable: false, contextWindow: 262144, maxOutput: 131072 } },
{ pattern: "*mimo*", caps: { vision: true, reasoning: true, thinkingFormat: "deepseek", thinkingCanDisable: false, contextWindow: 262144, maxOutput: 131072 } },
// ── Llama (4 = vision/1M; 3.x = text-only/128K) ──────────────────
{ pattern: "*llama-4*", caps: { vision: true, contextWindow: 1000000 } },
@@ -404,6 +419,60 @@ export const PATTERN_CAPABILITIES = [
// unknown models on these providers, trust vision instead of stripping images.
const TRUST_UPSTREAM_VISION = new Set(["openrouter"]);
/**
* Aggregate capabilities for a combo from its constituent model IDs.
* Each entry in comboModels is a fully-qualified "provider/model" string.
*
* Union: vision, pdf, audioInput, videoInput, imageOutput, audioOutput, search
* Intersection: tools
* Primary: reasoning fields from the first (primary) model
* Conservative: contextWindow = min; maxOutput = max
*
* @param {string[]} comboModels
* @param {Object|null} [comboLookup] optional map of combo name → models array for nested resolution
* @param {Function|null} [resolveCaps] optional (fullId) → caps override. The synced model
* catalog is server-only (it reads a file), so a browser-side resolution cannot see the
* limits it supplies and silently falls back to the generic patterns below. Callers that
* have the server's answer (/api/models, via useModelCaps) pass it here; it is merged over
* the local tables, so fields it does not carry (tools, pdf, audio/video, thinking*) survive.
* @param {number} [_depth] internal recursion depth guard
* @returns {object|null} full capabilities object, or null for empty input
*/
export function aggregateComboCapabilities(comboModels, comboLookup = null, resolveCaps = null, _depth = 0) {
if (!comboModels?.length || _depth > 6) return null;
const allCaps = comboModels.map((fullId) => {
// Nested combo: bare name (no slash) that exists in the lookup — recurse
if (!fullId.includes("/") && comboLookup?.[fullId]) {
return aggregateComboCapabilities(comboLookup[fullId], comboLookup, resolveCaps, _depth + 1)
?? resolveCaps?.(fullId)
?? getCapabilitiesForModel(null, fullId);
}
const slash = fullId.indexOf("/");
const provider = slash === -1 ? null : fullId.slice(0, slash);
const model = slash === -1 ? fullId : fullId.slice(slash + 1);
const local = getCapabilitiesForModel(provider, model);
const override = resolveCaps?.(fullId);
return override ? { ...local, ...override } : local;
});
const first = allCaps[0];
return {
vision: allCaps.some((c) => c.vision),
pdf: allCaps.some((c) => c.pdf),
audioInput: allCaps.some((c) => c.audioInput),
videoInput: allCaps.some((c) => c.videoInput),
imageOutput: allCaps.some((c) => c.imageOutput),
audioOutput: allCaps.some((c) => c.audioOutput),
search: allCaps.some((c) => c.search),
tools: allCaps.every((c) => c.tools),
reasoning: first.reasoning,
thinkingFormat: first.thinkingFormat,
thinkingCanDisable: first.thinkingCanDisable,
thinkingRange: first.thinkingRange,
contextWindow: Math.min(...allCaps.map((c) => c.contextWindow)),
maxOutput: Math.max(...allCaps.map((c) => c.maxOutput)),
};
}
/**
* Resolve capabilities for a model using the 4-step fallback chain,
* merged over DEFAULT_CAPABILITIES so the result is always complete.
@@ -414,16 +483,71 @@ const TRUST_UPSTREAM_VISION = new Set(["openrouter"]);
*/
const MODALITY_KEYS = ["vision", "pdf", "audioInput", "videoInput"];
// Catalog lookups, installed by the server at startup. Left as no-ops in the
// browser bundle, where there is no file to read.
// ── Server-injected readers ──────────────────────────────────────────
// Next.js compiles instrumentation and each API route into SEPARATE server
// bundles, so a plain module-local `let` would give every bundle its own copy of
// this file and a source installed at boot would be invisible to the request
// handlers (silently: the setters still "succeed"). The slots therefore live on
// globalThis, which IS shared across server bundles in the same process.
// Same reason the browser bundle is safe: it never calls a setter, so the slots
// stay empty and every consumer below short-circuits. Every read goes through
// globalThis: caching it locally would keep a reader alive in other copies after
// setCatalogSource(null).
let catalogSource = null;
const SOURCE_SLOTS = (globalThis.__9R_CAPABILITY_SOURCES ||= {
catalog: null, // { getModalities, getLimits } — synced models.dev catalog
userCaps: null, // (provider, model) => asserted caps — dashboard toggles
});
/**
* Install the synced catalog reader (server only).
* @param {{ getModalities: Function, getLimits: Function } | null} source
* @param {{ getModalities: (provider: string, model: string) => object|null,
* getLimits: (provider: string, model: string) => object|null } | null} source
*/
export function setCatalogSource(source) {
catalogSource = source;
catalogSource = source || null;
if (SOURCE_SLOTS) SOURCE_SLOTS.catalog = source || null;
if (typeof globalThis !== "undefined") globalThis.__9rCatalogSource = source || null;
}
function getCatalogSource() {
if (typeof globalThis === "undefined") return catalogSource;
return SOURCE_SLOTS?.catalog || globalThis.__9rCatalogSource || null;
}
// Capabilities the user asserted per provider+model (dashboard "Add/Edit Model"
// toggles), installed by the server from the custom-model store. Unlike the
// catalog and name heuristics this is authoritative and two-directional: it can
// turn a capability OFF as well as on.
const USER_CAPS_KEYS = ["vision", "pdf", "audioInput", "videoInput", "imageOutput", "audioOutput", "search", "tools", "reasoning", "thinkingFormat", "contextWindow", "maxOutput"];
// (slot lives in SOURCE_SLOTS above — see the cross-bundle note)
/**
* Install the user-asserted caps reader (server only).
* @param {(provider: string|null, model: string) => object|null} source sync lookup
*/
export function setUserCapsSource(source) {
SOURCE_SLOTS.userCaps = typeof source === "function" ? source : null;
}
// Last step of every resolution path: the user's own assertion wins over any
// heuristic, including the tables above (a hand-typed model id can collide with
// a pattern entry that describes a different product).
function applyUserCaps(result, provider, model) {
const userCapsSource = SOURCE_SLOTS.userCaps;
if (!userCapsSource) return result;
let asserted = null;
try {
asserted = userCapsSource(provider, model);
} catch {
return result;
}
if (!asserted || typeof asserted !== "object") return result;
for (const key of USER_CAPS_KEYS) {
if (asserted[key] === undefined) continue;
result[key] = asserted[key];
}
return result;
}
// Apply the synced catalog + name heuristic on top of a table-resolved result.
@@ -432,15 +556,16 @@ export function setCatalogSource(source) {
function refine(base, provider, model) {
const result = { ...DEFAULT_CAPABILITIES, ...base };
if (catalogSource) {
const modalities = catalogSource.getModalities(model);
const source = getCatalogSource();
if (source) {
const modalities = source.getModalities(provider, model);
if (modalities) {
for (const key of MODALITY_KEYS) {
if (modalities[key] === true) result[key] = true;
}
}
const limits = catalogSource.getLimits(provider, model);
const limits = source.getLimits(provider, model);
if (limits) {
if (limits.contextWindow > 0) result.contextWindow = limits.contextWindow;
if (limits.maxOutput > 0) result.maxOutput = limits.maxOutput;
@@ -452,30 +577,91 @@ function refine(base, provider, model) {
return result;
}
// Mirrors Command Code CLI `isKnownTextOnlyModel` (no image input). New models
// default to vision; only this denylist stays text-only.
const COMMANDCODE_TEXT_ONLY = new Set([
"deepseek/deepseek-v4-pro",
"deepseek/deepseek-v4-flash",
"deepseek/deepseek-v4-flash-fast",
"zai-org/glm-5.3",
"zai-org/glm-5.2",
"zai-org/glm-5.2-fast",
"zai-org/glm-5.1",
"zai-org/glm-5",
"minimaxai/minimax-m2.7",
"minimax/minimax-m2.7-free",
"minimaxai/minimax-m2.5",
"xiaomi/mimo-v2.5-pro",
"qwen/qwen3.6-max-preview",
"qwen/qwen3.7-max",
"meituan/longcat-2.0:free",
"stepfun/step-3.5-flash",
"tencent/hy4-preview",
"tencent/hy3",
"tencent/hy3-paid",
"nvidia/nemotron-3-ultra-550b-a55b",
"poolside/laguna-s-2.1-free",
"inclusionai/ling-3.0-flash-free",
"inclusionai/ling-3.0-flash-sante:free",
]);
function isCommandCodeTextOnly(model) {
const key = String(model || "").toLowerCase();
if (COMMANDCODE_TEXT_ONLY.has(key)) return true;
for (const id of COMMANDCODE_TEXT_ONLY) {
const base = id.includes("/") ? id.slice(id.lastIndexOf("/") + 1) : id;
if (key === base || key.endsWith("/" + base)) return true;
}
return false;
}
export function getCapabilitiesForModel(provider, model) {
if (!model) return { ...DEFAULT_CAPABILITIES };
// Canonical exact lookup strips vendor prefix: "anthropic/claude-opus-4.7" -> "claude-opus-4.7".
const baseModel = model.includes("/") ? model.split("/").pop() : model;
// 1. Provider-specific override
if (provider) {
const providerCaps = PROVIDER_CAPABILITIES[provider];
if (providerCaps?.[model]) return { ...DEFAULT_CAPABILITIES, ...providerCaps[model] };
if (providerCaps?.[baseModel]) return { ...DEFAULT_CAPABILITIES, ...providerCaps[baseModel] };
}
// 2. Canonical exact
if (MODEL_CAPABILITIES[baseModel]) return { ...DEFAULT_CAPABILITIES, ...MODEL_CAPABILITIES[baseModel] };
if (MODEL_CAPABILITIES[model]) return { ...DEFAULT_CAPABILITIES, ...MODEL_CAPABILITIES[model] };
// 3. Pattern match (first match wins), refined by catalog + name heuristic
for (const { pattern, caps } of PATTERN_CAPABILITIES) {
if (matchPattern(pattern, baseModel) || matchPattern(pattern, model)) {
return refine(caps, provider, model);
const resolve = () => {
// CommandCode wire is /alpha/generate for every model. Family patterns
// (deepseek-v4 → thinkingFormat:deepseek, vision:false) must not win here.
if (provider === "commandcode" || provider === "cmc") {
const providerCaps = PROVIDER_CAPABILITIES.commandcode;
if (providerCaps?.[model]) return { ...DEFAULT_CAPABILITIES, ...providerCaps[model] };
if (providerCaps?.[baseModel]) return { ...DEFAULT_CAPABILITIES, ...providerCaps[baseModel] };
return {
...DEFAULT_CAPABILITIES,
reasoning: true,
thinkingFormat: "commandcode",
thinkingEffortSupported: true,
vision: !isCommandCodeTextOnly(model),
contextWindow: 1000000,
maxOutput: 384000,
};
}
}
// 4. Floor
return refine(null, provider, model);
// 1. Provider-specific override
if (provider) {
const providerCaps = PROVIDER_CAPABILITIES[provider];
if (providerCaps?.[model]) return { ...DEFAULT_CAPABILITIES, ...providerCaps[model] };
if (providerCaps?.[baseModel]) return { ...DEFAULT_CAPABILITIES, ...providerCaps[baseModel] };
}
// 2. Canonical exact
if (MODEL_CAPABILITIES[baseModel]) return { ...DEFAULT_CAPABILITIES, ...MODEL_CAPABILITIES[baseModel] };
if (MODEL_CAPABILITIES[model]) return { ...DEFAULT_CAPABILITIES, ...MODEL_CAPABILITIES[model] };
// 3. Pattern match (first match wins), refined by catalog + name heuristic
for (const { pattern, caps } of PATTERN_CAPABILITIES) {
if (matchPattern(pattern, baseModel) || matchPattern(pattern, model)) {
return refine(caps, provider, model);
}
}
// 4. Floor (upstream-validated gateways keep vision on for unknown models)
if (provider && TRUST_UPSTREAM_VISION.has(provider)) {
return { ...refine(null, provider, model), vision: true };
}
return refine(null, provider, model);
};
return applyUserCaps(resolve(), provider, model);
}

View File

@@ -13,6 +13,11 @@ export const CATALOG_FILE = path.join(DATA_DIR, "model-catalog.json");
// Trimmed upstream catalog, read by the add-models skill (not by the router).
export const CATALOG_RAW_FILE = path.join(DATA_DIR, "model-catalog-raw.json");
// Schema of the file this module reads. The writer stamps it; a file carrying an
// older value predates provider-scoped modality keys, and its flat keys are not
// looked up here, so the sync rebuilds it instead of asking upstream for a 304.
export const CATALOG_VERSION = 2;
const EMPTY = { models: {}, providers: {} };
let cache = EMPTY;
let cachedMtime = -1;
@@ -45,14 +50,19 @@ function load() {
return cache;
}
// Modality is a property of the model itself — any gateway serving it inherits
// the same image/video/pdf support, so this is keyed by model id alone.
export function getCatalogModalities(model) {
return load().models[baseId(model)] || null;
// Modalities are recorded per gateway upstream, and gateways disagree about the
// same weights — some do not proxy images at all — so the key is provider +
// model, in the local provider id space, exactly like the limits below. Keying
// by model id alone made short ids collide across vendors: "auto", "free" and
// "efficient" are router modes in one catalog and model names in another, and a
// request to the router mode inherited a stranger's vision.
export function getCatalogModalities(provider, model) {
if (!provider) return null;
return load().models[`${provider}:${baseId(model)}`] || null;
}
// Context and output limits are a property of the gateway, not the model: each
// one truncates differently, so these stay keyed by provider + model.
// Context and output limits are a property of the gateway too: each one
// truncates differently, so these stay keyed by provider + model.
export function getCatalogLimits(provider, model) {
const byProvider = provider && load().providers[provider];
if (!byProvider) return null;

View File

@@ -23,7 +23,7 @@ function buildTransport(transport, oauth) {
const MEDIA_KEYS = new Set([
"serviceKinds", "ttsConfig", "sttConfig", "embeddingConfig",
"imageConfig", "imageToTextConfig", "videoConfig", "musicConfig",
"searchViaChat", "searchConfig", "fetchConfig",
"searchViaChat", "searchConfig", "fetchConfig", "systemoneConfig",
"modelsFetcher", "mediaPriority", "hiddenKinds",
]);

View File

@@ -1,3 +1,5 @@
import { FORMATS } from "../../translator/formats.js";
// Codex auto-generates a "-review" variant for each llm model (review quota family)
export const CODEX_REVIEW_SUFFIX = "-review";
@@ -25,3 +27,20 @@ export function isMuseSparkModel(modelId) {
const base = clean.includes("/") ? clean.split("/").pop() : clean;
return /^muse[-_]?spark(?:$|[-_:.\s])/i.test(base);
}
// Endpoint families for OpenCode models outside the curated registry (modelsFetcher /
// passthrough ids) — regex keeps auto-fetched models on the right endpoint:
// /responses (gpt/grok/muse-spark), /messages (minimax/qwen), /chat/completions (rest).
// Curated registry entries always win; this is the unknown-id fallback only.
const OPENCODE_FAMILIES = [
{ match: /^(grok|gpt|muse[-_]?spark)/i, supportedFormats: [FORMATS.OPENAI_RESPONSES], targetFormat: FORMATS.OPENAI_RESPONSES },
{ match: /^deepseek-v4-(pro|flash)/, supportedFormats: [FORMATS.OPENAI, FORMATS.CLAUDE, FORMATS.OPENAI_RESPONSES] },
{ match: /^(minimax|qwen)/, supportedFormats: [FORMATS.OPENAI, FORMATS.CLAUDE] },
{ match: /^claude-/i, supportedFormats: [FORMATS.CLAUDE] },
];
export function opencodeFamilyFormats(modelId) {
if (!modelId || typeof modelId !== "string") return null;
const base = modelId.replace(/\([^()]+\)\s*$/, "").trim();
return OPENCODE_FAMILIES.find((f) => f.match.test(base)) || null;
}

View File

@@ -2,8 +2,28 @@
//
// Fallback order (first match wins):
// 1. PROVIDER_PRICING[provider][model] — provider-specific override
// 2. MODEL_PRICING[model] — canonical model price (provider-agnostic)
// 3. PATTERN_PRICING — glob pattern match (e.g. "codex-*")
// 2. FREE_MODEL_NAMESPACES — upstream bills these at $0
// 3. MODEL_PRICING[model] — canonical model price (provider-agnostic)
// 4. PATTERN_PRICING — glob pattern match (e.g. "codex-*")
/**
* Namespaces upstream meters at $0. A free model must never inherit a paid
* rate: the vendor-prefix strip in getPricingForModel() would turn
* "cline-free/deepseek-v4.1-flash" into "deepseek-v4.1-flash" and match
* MODEL_PRICING, so the namespace is checked before both fallbacks.
*/
export const FREE_MODEL_NAMESPACES = ["cline-free/"];
export const ZERO_PRICING = {
input: 0, output: 0, cached: 0, reasoning: 0, cache_creation: 0,
};
/** True when the model id sits in a namespace upstream bills at $0. */
export function isFreeModel(model) {
if (!model) return false;
const lower = String(model).toLowerCase();
return FREE_MODEL_NAMESPACES.some((ns) => lower.startsWith(ns));
}
/**
* Canonical model pricing — provider-agnostic.
@@ -111,6 +131,8 @@ export const MODEL_PRICING = {
"deepseek-v3.2-chat": { input: 0.14, output: 0.28, cached: 0.0028, reasoning: 0.28, cache_creation: 0.14 },
"deepseek-v3.2-reasoner": { input: 0.14, output: 0.28, cached: 0.0028, reasoning: 0.28, cache_creation: 0.14 },
"deepseek-v4-flash": { input: 0.14, output: 0.28, cached: 0.0028, reasoning: 0.28, cache_creation: 0.14 },
"deepseek-v4.1-flash": { input: 0.14, output: 0.28, cached: 0.0028, reasoning: 0.28, cache_creation: 0.14 },
"deepseek-flash": { input: 0.14, output: 0.28, cached: 0.0028, reasoning: 0.28, cache_creation: 0.14 },
"deepseek-v4-pro": { input: 0.435, output: 0.87, cached: 0.003625, reasoning: 0.87, cache_creation: 0.435 },
// === GLM ===
@@ -359,10 +381,11 @@ export function matchPattern(pattern, model) {
}
/**
* Resolve pricing for a model using the 3-step fallback chain:
* Resolve pricing for a model using the 4-step fallback chain:
* 1. PROVIDER_PRICING[provider][model]
* 2. MODEL_PRICING[model]
* 3. PATTERN_PRICING (glob match)
* 2. free namespace (upstream bills $0)
* 3. MODEL_PRICING[model]
* 4. PATTERN_PRICING (glob match)
*
* @param {string} provider
* @param {string} model
@@ -376,12 +399,15 @@ export function getPricingForModel(provider, model) {
return PROVIDER_PRICING[provider][model];
}
// 2. Canonical model pricing (strip vendor prefix if needed: "deepseek/deepseek-chat" → "deepseek-chat")
// 2. Free namespaces bill $0 regardless of the model name behind them.
if (isFreeModel(model)) return ZERO_PRICING;
// 3. Canonical model pricing (strip vendor prefix if needed: "deepseek/deepseek-chat" → "deepseek-chat")
const baseModel = model.includes("/") ? model.split("/").pop() : model;
if (MODEL_PRICING[baseModel]) return MODEL_PRICING[baseModel];
if (MODEL_PRICING[model]) return MODEL_PRICING[model];
// 3. Pattern match
// 4. Pattern match
for (const { pattern, pricing } of PATTERN_PRICING) {
if (matchPattern(pattern, baseModel) || matchPattern(pattern, model)) {
return pricing;

View File

@@ -0,0 +1,29 @@
export default {
id: "agnes",
priority: 120,
alias: "agnes",
aliases: [
"agnes-ai",
],
uiAlias: "agnes",
display: {
name: "Agnes AI",
icon: "auto_awesome",
color: "#7C3AED",
textIcon: "AG",
website: "https://agnes-ai.com",
notice: {
text: "OpenAI-compatible gateway from Agnes AI, offering free API credits on sign-up. Accepts a bearer token or an x-api-key header.",
apiKeyUrl: "https://platform.agnes-ai.com",
},
},
category: "freeTier",
authType: "apikey",
transport: {
baseUrl: "https://apihub.agnes-ai.com/v1/chat/completions",
validateUrl: "https://apihub.agnes-ai.com/v1/models",
},
// No model ids could be verified without a key, so discovery is left to the
// live endpoint and any id is accepted through passthroughModels.
passthroughModels: true,
};

View File

@@ -37,6 +37,7 @@ export default {
},
usage: {
quotaApiUrl: `${ANTIGRAVITY_IDE_BASE_URL}/v1internal:fetchAvailableModels`,
quotaSummaryApiUrl: `${ANTIGRAVITY_IDE_BASE_URL}/v1internal:retrieveUserQuotaSummary`,
loadProjectApiUrl: "https://cloudcode-pa.googleapis.com/v1internal:loadCodeAssist",
tokenUrl: "https://oauth2.googleapis.com/token",
},

View File

@@ -20,6 +20,8 @@ export default {
authModes: [
"apikey",
],
passthroughModels: true,
modelsFetcher: { url: "https://api.airforce/v1/models", type: "airforce-free" },
transport: {
baseUrl: "https://api.airforce/v1/chat/completions",
validateUrl: "https://api.airforce/v1/models",
@@ -27,10 +29,11 @@ export default {
"HTTP-Referer": "https://endpoint-proxy.local",
"X-Title": "Endpoint Proxy",
},
forceStream: true,
},
models: [
{ id: "anthropic/claude-3.7-sonnet", name: "Claude 3.7 Sonnet (Free)", contextLength: 200000 },
{ id: "moonshot/kimi-k2.6", name: "Kimi K2.6 (Free)", contextLength: 262144 },
{ id: "google/gemini-2.5-flash", name: "Gemini 2.5 Flash (Free)", contextLength: 1048576 },
{ id: "gpt-oss-120b", name: "GPT-OSS 120B (Free)", contextLength: 131072 },
{ id: "gpt-oss-20b", name: "GPT-OSS 20B (Free)", contextLength: 131072 },
{ id: "kimi-k2.7-code", name: "Kimi K2.7 Code (Free)", contextLength: 262144 },
],
};

View File

@@ -0,0 +1,32 @@
export default {
id: "atria",
priority: 120,
alias: "atria",
aliases: [
"atria-asi",
],
uiAlias: "atria",
display: {
name: "Atria Dawn",
icon: "flare",
color: "#C2410C",
textIcon: "AD",
website: "https://atria-asi.ai",
notice: {
text: "OpenAI-compatible endpoint from Atria Dawn (AtomInnoLab). Currently a research preview offering a single text model, Atria-Dawn-Preview.",
apiKeyUrl: "https://api.atria-asi.ai/dashboard",
},
},
category: "apikey",
authType: "apikey",
transport: {
baseUrl: "https://api.atria-asi.ai/v1/chat/completions",
validateUrl: "https://api.atria-asi.ai/v1/models",
},
// Docs pin the model field to one case-sensitive id. Text-only for now: the
// service ships a hook that blocks image/PDF input, so no vision is claimed.
models: [
{ id: "Atria-Dawn-Preview", name: "Atria Dawn Preview" },
],
passthroughModels: true,
};

View File

@@ -0,0 +1,30 @@
export default {
id: "bai",
priority: 120,
alias: "bai",
aliases: [
"b-ai",
],
uiAlias: "bai",
display: {
name: "B.AI",
icon: "account_balance",
color: "#0369A1",
textIcon: "BA",
website: "https://b.ai",
notice: {
text: "OpenAI-compatible gateway with one of the larger catalogues here. Accepts a bearer token or an x-api-key header. Model ids are fetched live from the provider.",
apiKeyUrl: "https://b.ai",
},
},
category: "apikey",
authType: "apikey",
transport: {
baseUrl: "https://api.b.ai/v1/chat/completions",
validateUrl: "https://api.b.ai/v1/models",
},
// No ids hardcoded: the catalogue is large and rotates, so the live endpoint
// is the source of truth and any id is accepted via passthroughModels.
modelsFetcher: { url: "https://api.b.ai/v1/models", type: "openai" },
passthroughModels: true,
};

View File

@@ -54,9 +54,12 @@ export default {
oauthUrl: "https://api.anthropic.com/api/oauth/usage",
orgUrl: "https://api.anthropic.com/v1/organizations/{org_id}/usage",
settingsUrl: "https://api.anthropic.com/v1/settings",
profileUrl: "https://api.anthropic.com/api/oauth/profile",
resetUrl: "https://api.anthropic.com/api/organizations/{org_id}/reset_rate_limits",
},
},
models: [
{ id: "claude-opus-5-5", name: "Claude Opus 5.5" },
{ id: "claude-opus-5", name: "Claude Opus 5" },
{ id: "claude-fable-5-1", name: "Claude Fable 5.1" },
{ id: "claude-fable-5", name: "Claude Fable 5" },

View File

@@ -14,12 +14,16 @@ export default {
},
},
category: "oauth",
authModes: ["oauth"],
hasOAuth: true,
transport: {
baseUrl: "https://api.cline.bot/api/v1/chat/completions",
headers: {
"HTTP-Referer": "https://cline.bot",
"X-Title": "Cline",
},
// Non-stream chat completions come back wrapped in {"success":true,"data":{...}}
quirks: { clineEnvelope: true },
tokenUrl: "https://api.cline.bot/api/v1/auth/token",
refreshUrl: "https://api.cline.bot/api/v1/auth/refresh",
auth: {

View File

@@ -14,7 +14,10 @@ export default {
},
},
category: "oauth",
authModes: ["oauth", "apikey"],
// ClinePass authenticates with a plain API key from app.cline.bot/settings/api-keys
// (category "apikey"). The OAuth extension flow used by Cline does not issue
// tokens that the ClinePass API consumer endpoint accepts (HTTP 401) — see #2333.
authModes: ["apikey", "oauth"],
hasOAuth: true,
transport: {
baseUrl: "https://api.cline.bot/api/v1/chat/completions",
@@ -22,6 +25,8 @@ export default {
"HTTP-Referer": "https://cline.bot",
"X-Title": "Cline",
},
// Non-stream chat completions come back wrapped in {"success":true,"data":{...}}
quirks: { clineEnvelope: true },
auth: {
combined: true,
header: "Authorization",

View File

@@ -58,7 +58,9 @@ export default {
// (endpoint returns 11102 "model service info not found"), plus
// glm-5.0-turbo / minimax-m2.7 / kimi-k2.5 / hy3-preview /
// deepseek-v3-2-volc (absent from the server list, though still answering
// 200) and hy3-x (paid tier, not used here).
// 200) and hy3-x (paid tier, not used here). deepseek-v4-flash removed
// 2026-09: replaced server-side by deepseek-v4.1-flash (same low/high/
// xhigh efforts; endpoint still answers 200 but the list is the contract).
// "-x" suffix = paid tier of the same model (free id rides the promo quota).
{ id: "hy3", name: "Hy3" },
{ id: "hy4-preview", name: "Hy4-Preview" },
@@ -66,7 +68,7 @@ export default {
{ id: "glm-5.3-flash", name: "GLM-5.3-Flash" },
{ id: "kimi-k3-1", name: "Kimi-K3" },
{ id: "deepseek-v4-pro", name: "DeepSeek-V4-Pro" },
{ id: "deepseek-v4-flash", name: "DeepSeek-V4-Flash" },
{ id: "deepseek-v4.1-flash", name: "DeepSeek-V4.1-Flash" },
],
oauth: {
baseUrl: "https://copilot.tencent.com",

View File

@@ -58,7 +58,9 @@ export default {
{ id: "kimi-k2.5", name: "Kimi-K2.5" },
{ id: "hy3-preview", name: "Hy3 Preview" },
{ id: "deepseek-v4-pro", name: "DeepSeek-V4-Pro" },
{ id: "deepseek-v4-flash", name: "DeepSeek-V4-Flash" },
// deepseek-v4-flash replaced server-side by deepseek-v4.1-flash (same
// catalog as CN; the old endpoint still answers 200 but the list is the contract).
{ id: "deepseek-v4.1-flash", name: "DeepSeek-V4.1-Flash" },
{ id: "deepseek-v3-2-volc", name: "DeepSeek-V3.2" },
],
oauth: {

View File

@@ -1,5 +1,10 @@
import { withCodexReviewModels } from "../models/helpers.js";
// Codex CLI version seen by OpenAI's backend — single source for the Version /
// User-Agent identity headers. Bump when the installed codex CLI is upgraded.
const CODEX_CLI_VERSION = "0.155.0";
const GPT_6_LITE_THINKING_LEVELS = ["low", "medium", "high", "xhigh", "max"];
export default {
id: "codex",
priority: 30,
@@ -34,9 +39,11 @@ export default {
baseUrl: "https://chatgpt.com/backend-api/codex/responses",
format: "openai-responses",
forceStream: true,
cliVersion: CODEX_CLI_VERSION,
headers: {
originator: "codex_cli_rs",
"User-Agent": "codex_cli_rs/0.136.0",
"User-Agent": `codex_cli_rs/${CODEX_CLI_VERSION}`,
version: CODEX_CLI_VERSION,
},
usage: {
url: "https://chatgpt.com/backend-api/wham/usage",
@@ -46,6 +53,8 @@ export default {
},
models: [
{ id: "gpt-6-astra", name: "GPT 6.0 Astra" },
{ id: "gpt-6-sol", name: "GPT 6.0 Sol", responsesLite: true, thinkingLevels: GPT_6_LITE_THINKING_LEVELS },
{ id: "gpt-6-luna", name: "GPT 6.0 Luna", responsesLite: true, thinkingLevels: GPT_6_LITE_THINKING_LEVELS },
{ id: "gpt-5.6-sol", name: "GPT 5.6 Sol" },
{ id: "gpt-5.6-sol-review", name: "GPT 5.6 Sol Review", upstreamModelId: "gpt-5.6-sol", quotaFamily: "review" },
{ id: "gpt-5.6-terra", name: "GPT 5.6 Terra" },
@@ -60,6 +69,14 @@ export default {
{ id: "gpt-5.4-mini-review", name: "GPT 5.4 Mini Review", upstreamModelId: "gpt-5.4-mini", quotaFamily: "review" },
{ id: "gpt-5.3-codex-spark", name: "GPT 5.3 Codex Spark" },
{ id: "gpt-5.3-codex-spark-review", name: "GPT 5.3 Codex Spark Review", upstreamModelId: "gpt-5.3-codex-spark", quotaFamily: "review" },
// Codex CLI's auto-review virtual model. Unlike the "-review" variants above it is not derived
// from a base model, so it is forwarded verbatim instead of having "-review" stripped (#1398).
{ id: "codex-auto-review", name: "Codex Auto Review", upstreamModelId: "codex-auto-review", quotaFamily: "review" },
{ id: "gpt-image-2.5", name: "GPT Image 2.5", capabilities: ["text2img","edit","multiImage"], params: ["size","quality","background","image_detail","output_format"], kind: "image" },
{ id: "gpt-image-2.5-flare", name: "GPT Image 2.5 Flare", capabilities: ["text2img","edit","multiImage"], params: ["size","quality","background","image_detail","output_format"], kind: "image" },
{ id: "gpt-image-2.5-sunburst", name: "GPT Image 2.5 Sunburst", capabilities: ["text2img","edit","multiImage"], params: ["size","quality","background","image_detail","output_format"], kind: "image" },
{ id: "gpt-image-2", name: "GPT Image 2", capabilities: ["text2img","edit","multiImage"], params: ["size","quality","background","image_detail","output_format"], kind: "image" },
{ id: "gpt-image-1.5", name: "GPT Image 1.5", capabilities: ["text2img","edit","multiImage"], params: ["size","quality","background","image_detail","output_format"], kind: "image" },
{ id: "gpt-5.6-sol-image", name: "GPT 5.6 Sol Image", capabilities: ["text2img","edit"], params: ["size","quality","background","image_detail","output_format"], kind: "image" },
{ id: "gpt-5.6-terra-image", name: "GPT 5.6 Terra Image", capabilities: ["text2img","edit"], params: ["size","quality","background","image_detail","output_format"], kind: "image" },
{ id: "gpt-5.6-luna-image", name: "GPT 5.6 Luna Image", capabilities: ["text2img","edit"], params: ["size","quality","background","image_detail","output_format"], kind: "image" },

View File

@@ -38,15 +38,11 @@ export default {
summaryUrl: "/alpha/usage/summary",
},
},
features: {
usage: true,
usageApikey: true,
},
models: [
{ id: "deepseek/deepseek-v4-pro", name: "DeepSeek V4 Pro" },
{ id: "deepseek/deepseek-v4-flash", name: "DeepSeek V4 Flash" },
{ id: "moonshotai/Kimi-K2.7-Code", name: "Kimi K2.7 Code" },
{ id: "moonshotai/Kimi-K2.7-Code-Highspeed", name: "Kimi K2.7 Code Highspeed" },
{ id: "moonshotai/Kimi-K2.7-Code-Highspeed", name: "Kimi K2.7 Code HighSpeed" },
{ id: "moonshotai/Kimi-K2.6", name: "Kimi K2.6" },
{ id: "moonshotai/Kimi-K2.5", name: "Kimi K2.5" },
{ id: "zai-org/GLM-5.2", name: "GLM 5.2" },
@@ -58,14 +54,14 @@ export default {
{ id: "MiniMaxAI/MiniMax-M2.5", name: "MiniMax M2.5" },
{ id: "xiaomi/mimo-v2.5-pro", name: "MiMo V2.5 Pro" },
{ id: "xiaomi/mimo-v2.5", name: "MiMo V2.5" },
{ id: "Qwen/Qwen3.7-Max", name: "Qwen 3.7 Max" },
{ id: "Qwen/Qwen3.7-Plus", name: "Qwen 3.7 Plus" },
{ id: "Qwen/Qwen3.6-Max-Preview", name: "Qwen 3.6 Max Preview" },
{ id: "Qwen/Qwen3.6-Plus", name: "Qwen 3.6 Plus" },
{ id: "Qwen/Qwen3.7-Max", name: "Qwen 3.7 Max" },
{ id: "Qwen/Qwen3.7-Plus", name: "Qwen 3.7 Plus" },
{ id: "stepfun/Step-3.7-Flash", name: "Step 3.7 Flash" },
{ id: "stepfun/Step-3.5-Flash", name: "Step 3.5 Flash" },
{ id: "tencent/Hy3", name: "Tencent Hy3" },
{ id: "nvidia/nemotron-3-ultra-550b-a55b", name: "Nemotron 3 Ultra 550B A55B" },
{ id: "nvidia/nemotron-3-ultra-550b-a55b", name: "Nemotron 3 Ultra" },
{ id: "thinkingmachines/inkling", name: "Inkling" },
{ id: "claude-sonnet-5", name: "Claude Sonnet 5" },
{ id: "claude-sonnet-4-6", name: "Claude Sonnet 4.6" },
@@ -74,4 +70,8 @@ export default {
{ id: "claude-opus-4-7", name: "Claude Opus 4.7" },
{ id: "claude-haiku-4-5", name: "Claude Haiku 4.5" },
],
features: {
usage: true,
usageApikey: true,
},
};

View File

@@ -0,0 +1,35 @@
export default {
id: "dahl",
priority: 120,
alias: "dahl",
aliases: [
"dahl-inference",
],
uiAlias: "dahl",
display: {
name: "Dahl Inference",
icon: "hub",
color: "#1E40AF",
textIcon: "DH",
website: "https://dahl.global",
notice: {
text: "OpenAI-compatible Gonka inference node. Small, fixed catalogue (GLM-5.3-Flash, DeepSeek-V4-Flash, MiniMax-M2.7) at a flat per-token rate.",
apiKeyUrl: "https://dahl.global/dashboard",
},
},
category: "apikey",
authType: "apikey",
transport: {
baseUrl: "https://inference.dahl.global/v1/chat/completions",
validateUrl: "https://inference.dahl.global/v1/models",
},
// The live catalogue is public (no auth), so modelsFetcher works without a key
// and the ids below are a convenience seed rather than an exhaustive list.
models: [
{ id: "zai-org/GLM-5.3-Flash", name: "GLM-5.3 Flash" },
{ id: "deepseek-ai/DeepSeek-V4-Flash-0731", name: "DeepSeek V4 Flash 0731" },
{ id: "MiniMaxAI/MiniMax-M2.7", name: "MiniMax M2.7" },
],
modelsFetcher: { url: "https://inference.dahl.global/v1/models", type: "openai" },
passthroughModels: true,
};

View File

@@ -25,6 +25,21 @@ export default {
reasoningInject: {
scope: "all",
},
quirks: {
// DeepSeek's Anthropic-compatible endpoint
// (https://api.deepseek.com/anthropic/v1/messages) accepts ONLY the
// built-in web_search_* tools and rejects client-defined `custom` tools
// (MCP / Read / Bash / etc.) with HTTP 400
// "tools[0]: unknown variant `custom`, expected
// `web_search_20250305` or `web_search_20260209`".
//
// Declaring this whitelist makes prepareClaudeRequest() forward only
// web_search_* tools and strip everything else before sending, so MCP /
// function tools are dropped instead of failing the whole request.
// DeepSeek's OpenAI-compatible transport is unaffected (targetFormat
// there is "openai", not "claude", so prepareClaudeRequest is not run).
claudeSupportedToolTypes: ["web_search_20250305", "web_search_20260209"],
},
},
// Multi-endpoint: pick the transport matching client sourceFormat to skip translation.
transports: [
@@ -44,6 +59,7 @@ export default {
{ id: "deepseek-v4-pro", name: "DeepSeek V4 Pro" },
{ id: "deepseek-v4-pro-max", name: "DeepSeek V4 Pro Max", upstreamModelId: "deepseek-v4-pro" },
{ id: "deepseek-v4-pro-none", name: "DeepSeek V4 Pro No Thinking", upstreamModelId: "deepseek-v4-pro" },
{ id: "deepseek-v4.1-flash", name: "DeepSeek V4.1 Flash" },
{ id: "deepseek-v4-flash", name: "DeepSeek V4 Flash" },
{ id: "deepseek-v4-flash-vision-exp", name: "DeepSeek V4 Flash Vision (Exp)" },
{ id: "deepseek-chat", name: "DeepSeek V3.2 Chat" },

View File

@@ -58,6 +58,7 @@ export default {
{ id: "gemini-2.5-flash", name: "Gemini 2.5 Flash", params: ["language","prompt"], kind: "stt" },
{ id: "gemini-2.5-flash-lite", name: "Gemini 2.5 Flash Lite (Cheapest)", params: ["language","prompt"], kind: "stt" },
{ id: "gemini-2.0-flash", name: "Gemini 2.0 Flash", params: ["language","prompt"], kind: "stt" },
{ id: "gemini-2.5-flash-native-audio-preview-09-17", name: "Gemini Live Transcription (Realtime)", params: ["language","prompt","system_instruction","setup_timeout_ms","turn_timeout_ms"], kind: "stt", transport: "gemini-live" },
{ id: "gemini-3.1-flash-tts-preview", name: "Gemini 3.1 Flash TTS", kind: "tts" },
{ id: "gemini-2.5-flash-preview-tts", name: "Gemini 2.5 Flash TTS", kind: "tts" },
{ id: "gemini-2.5-pro-preview-tts", name: "Gemini 2.5 Pro TTS", kind: "tts" },

View File

@@ -15,6 +15,7 @@ export default {
website: "https://huggingface.co",
notice: {
apiKeyUrl: "https://huggingface.co/settings/tokens",
text: "Runs through the Inference Providers router. Image and speech models are billed by the provider selected per model.",
},
},
category: "apikey",
@@ -25,10 +26,79 @@ export default {
transport: null,
models: [
{ id: "black-forest-labs/FLUX.1-schnell", name: "FLUX.1 Schnell", params: [], kind: "image" },
{ id: "black-forest-labs/FLUX.1-dev", name: "FLUX.1 Dev", params: [], kind: "image" },
{ id: "black-forest-labs/FLUX.1-Krea-dev", name: "FLUX.1 Krea", params: [], kind: "image" },
{ id: "black-forest-labs/FLUX.1-Kontext-dev", name: "FLUX.1 Kontext", params: [], kind: "image", capabilities: ["edit"] },
{ id: "black-forest-labs/FLUX.2-dev", name: "FLUX.2 Dev", params: [], kind: "image", capabilities: ["edit"] },
{ id: "black-forest-labs/FLUX.2-klein-9B", name: "FLUX.2 Klein 9B", params: [], kind: "image", capabilities: ["edit"] },
{ id: "black-forest-labs/FLUX.2-klein-4B", name: "FLUX.2 Klein 4B", params: [], kind: "image", capabilities: ["edit"] },
{ id: "black-forest-labs/FLUX.2-klein-base-9B", name: "FLUX.2 Klein Base 9B", params: [], kind: "image", capabilities: ["edit"] },
{ id: "black-forest-labs/FLUX.2-klein-base-4B", name: "FLUX.2 Klein Base 4B", params: [], kind: "image", capabilities: ["edit"] },
{ id: "stabilityai/stable-diffusion-xl-base-1.0", name: "SDXL Base 1.0", params: [], kind: "image" },
{ id: "openai/whisper-large-v3", name: "Whisper Large v3 (HF)", params: ["language"], kind: "stt" },
{ id: "openai/whisper-small", name: "Whisper Small (HF)", params: ["language"], kind: "stt" },
{ id: "stabilityai/stable-diffusion-3.5-large", name: "Stable Diffusion 3.5 Large", params: [], kind: "image" },
{ id: "stabilityai/stable-diffusion-3.5-large-turbo", name: "Stable Diffusion 3.5 Large Turbo", params: [], kind: "image" },
{ id: "Qwen/Qwen-Image", name: "Qwen Image", params: [], kind: "image" },
{ id: "Qwen/Qwen-Image-2512", name: "Qwen Image 2512", params: [], kind: "image" },
{ id: "Qwen/Qwen-Image-Edit", name: "Qwen Image Edit", params: [], kind: "image", capabilities: ["edit"] },
{ id: "Qwen/Qwen-Image-Edit-2509", name: "Qwen Image Edit 2509", params: [], kind: "image", capabilities: ["edit"] },
{ id: "Qwen/Qwen-Image-Edit-2511", name: "Qwen Image Edit 2511", params: [], kind: "image", capabilities: ["edit"] },
{ id: "ideogram-ai/ideogram-4-fp8", name: "Ideogram 4", params: [], kind: "image" },
{ id: "tencent/HunyuanImage-3.0", name: "HunyuanImage 3.0", params: [], kind: "image" },
{ id: "Tongyi-MAI/Z-Image-Turbo", name: "Z-Image Turbo", params: [], kind: "image" },
{ id: "krea/Krea-2-Turbo", name: "Krea 2 Turbo", params: [], kind: "image" },
{ id: "HiDream-ai/HiDream-I1-Fast", name: "HiDream I1 Fast", params: [], kind: "image" },
{ id: "playgroundai/playground-v2.5-1024px-aesthetic", name: "Playground v2.5", params: [], kind: "image" },
{ id: "openai/whisper-large-v3", name: "Whisper Large v3 (HF)", params: [], kind: "stt" },
{ id: "openai/whisper-large-v3-turbo", name: "Whisper Large v3 Turbo (HF)", params: [], kind: "stt" },
],
serviceKinds: ["image", "stt"],
imageConfig: { baseUrl: "https://api-inference.huggingface.co/models" },
// Inference Providers router. The router is addressed as
// `<baseUrl>/<provider>/<providerModelId>` — see open-sse/handlers/imageProviders/huggingface.js.
// `modelMap` resolves a Hub model id to the provider-resolved id the router expects.
// A plain string value is the provider path. Image-to-image models use
// `{ path, task: "image-to-image" }`: their payload differs — the source image goes in
// `inputs` and the prompt under `parameters.prompt`. See
// https://huggingface.co/docs/inference-providers/tasks/image-to-image
// Only providers the router actually forwards to are listed: replicate, wavespeed and
// deepinfra appear in the Hub's inferenceProviderMapping but reject router traffic with
// "Model not supported by provider <name>".
imageConfig: {
baseUrl: "https://router.huggingface.co",
modelMap: {
"black-forest-labs/FLUX.1-schnell": "fal-ai/fal-ai/flux/schnell",
"black-forest-labs/FLUX.1-dev": "fal-ai/fal-ai/flux/dev",
"black-forest-labs/FLUX.1-Krea-dev": "fal-ai/fal-ai/flux/krea",
"black-forest-labs/FLUX.1-Kontext-dev": { path: "fal-ai/fal-ai/flux-kontext/dev", task: "image-to-image" },
"black-forest-labs/FLUX.2-dev": { path: "fal-ai/fal-ai/flux-2/edit", task: "image-to-image" },
"black-forest-labs/FLUX.2-klein-9B": { path: "fal-ai/fal-ai/flux-2/klein/9b/edit", task: "image-to-image" },
"black-forest-labs/FLUX.2-klein-4B": { path: "fal-ai/fal-ai/flux-2/klein/4b/distilled/edit", task: "image-to-image" },
"black-forest-labs/FLUX.2-klein-base-9B": { path: "fal-ai/fal-ai/flux-2/klein/9b/base/edit", task: "image-to-image" },
"black-forest-labs/FLUX.2-klein-base-4B": { path: "fal-ai/fal-ai/flux-2/klein/4b/base/edit", task: "image-to-image" },
"stabilityai/stable-diffusion-xl-base-1.0": "fal-ai/fal-ai/fast-sdxl",
"stabilityai/stable-diffusion-3.5-large": "fal-ai/fal-ai/stable-diffusion-v35-large",
"stabilityai/stable-diffusion-3.5-large-turbo": "fal-ai/fal-ai/stable-diffusion-v35-large/turbo",
"Qwen/Qwen-Image": "fal-ai/fal-ai/qwen-image",
"Qwen/Qwen-Image-2512": "fal-ai/fal-ai/qwen-image-2512",
"Qwen/Qwen-Image-Edit": { path: "fal-ai/fal-ai/qwen-image-edit", task: "image-to-image" },
"Qwen/Qwen-Image-Edit-2509": { path: "fal-ai/fal-ai/qwen-image-edit-2509", task: "image-to-image" },
"Qwen/Qwen-Image-Edit-2511": { path: "fal-ai/fal-ai/qwen-image-edit-plus", task: "image-to-image" },
"ideogram-ai/ideogram-4-fp8": "fal-ai/ideogram/v4",
"tencent/HunyuanImage-3.0": "fal-ai/fal-ai/hunyuan-image/v3/text-to-image",
"Tongyi-MAI/Z-Image-Turbo": "fal-ai/fal-ai/z-image/turbo",
"krea/Krea-2-Turbo": "fal-ai/fal-ai/krea-2/turbo",
"HiDream-ai/HiDream-I1-Fast": "fal-ai/fal-ai/hidream-i1-fast",
"playgroundai/playground-v2.5-1024px-aesthetic": "fal-ai/fal-ai/playground-v25",
},
},
// Speech-to-text goes through the hf-inference provider, which keeps the Hub
// model id as its provider-resolved id (`/hf-inference/models/<hubId>`).
// No `params` are declared: the router's ASR payload carries only `inputs` and
// `parameters.return_timestamps` / `parameters.generation_parameters` — it has no
// language field, so a UI-declared "language" would be silently dropped.
sttConfig: {
baseUrl: "https://router.huggingface.co/hf-inference/models",
authType: "apikey",
authHeader: "bearer",
format: "huggingface-asr",
},
};

View File

@@ -69,6 +69,7 @@ import p66 from "./ollama.js";
import p123 from "./ollama-search.js";
import p67 from "./openai.js";
import p68 from "./opencode-go.js";
import p68z from "./opencode-zen.js";
import p69 from "./opencode.js";
import p70 from "./openrouter.js";
import p71 from "./perplexity-web.js";
@@ -76,6 +77,7 @@ import p72 from "./perplexity.js";
import p73 from "./perplexity-agent.js";
import p74 from "./playht.js";
import p75 from "./qoder.js";
import p124 from "./qoder-cn.js";
import p77 from "./recraft.js";
import p78 from "./runwayml.js";
import p79 from "./sdwebui.js";
@@ -123,7 +125,11 @@ import p119 from "./selfhosted-embedding.js";
import p120 from "./fish-audio.js";
import p121 from "./alitp-intl.js";
import p122 from "./xquik.js";
import p125 from "./tokenharbor.js";
import p126 from "./dahl.js";
import p127 from "./atria.js";
import p129 from "./agnes.js";
import p130 from "./bai.js";
export default [
p0,
p1,
@@ -193,8 +199,10 @@ export default [
p65,
p66,
p123,
p124,
p67,
p68,
p68z,
p69,
p70,
p71,
@@ -247,4 +255,9 @@ export default [
p120,
p121,
p122,
p125,
p126,
p127,
p129,
p130,
];

View File

@@ -22,6 +22,7 @@ export default {
headers: { ...CLAUDE_API_HEADERS },
quirks: {
dropOutputConfig: true,
requireClaudeToolType: true,
},
reasoningInject: {
scope: "all",

View File

@@ -22,6 +22,7 @@ export default {
headers: { ...CLAUDE_API_HEADERS },
quirks: {
dropOutputConfig: true,
requireClaudeToolType: true,
},
reasoningInject: {
scope: "all",

View File

@@ -30,6 +30,7 @@ export default {
{ id: "glm-4.7-flash", name: "GLM 4.7 Flash" },
{ id: "qwen3.5", name: "Qwen3.5" },
{ id: "minimax-m3", name: "MiniMax M3" },
{ id: "deepseek-v4.1-flash:cloud", name: "DeepSeek V4.1 Flash" },
],
serviceKinds: ["llm", "webFetch"],
fetchConfig: {

View File

@@ -28,6 +28,7 @@ export default {
forceStream: true,
},
models: [
{ id: "gpt-5.5", name: "GPT-5.5" },
{ id: "gpt-5.4", name: "GPT-5.4" },
{ id: "gpt-5.4-mini", name: "GPT-5.4 Mini" },
{ id: "gpt-5.4-nano", name: "GPT-5.4 Nano" },
@@ -57,6 +58,9 @@ export default {
{ id: "whisper-1", name: "Whisper 1", params: ["language","response_format","temperature","prompt"], kind: "stt" },
{ id: "gpt-4o-transcribe", name: "GPT-4o Transcribe", params: ["language","response_format","temperature","prompt"], kind: "stt" },
{ id: "gpt-4o-mini-transcribe", name: "GPT-4o Mini Transcribe", params: ["language","response_format","temperature","prompt"], kind: "stt" },
{ id: "gpt-image-2.5", name: "GPT Image 2.5", params: ["n","size","quality","response_format"], kind: "image" },
{ id: "gpt-image-2.5-flare", name: "GPT Image 2.5 Flare", params: ["n","size","quality","response_format"], kind: "image" },
{ id: "gpt-image-2.5-sunburst", name: "GPT Image 2.5 Sunburst", params: ["n","size","quality","response_format"], kind: "image" },
{ id: "gpt-image-1", name: "GPT Image 1", params: ["n","size","quality","response_format"], kind: "image" },
{ id: "dall-e-3", name: "DALL-E 3", params: ["size","quality","style","response_format"], kind: "image" },
{ id: "dall-e-2", name: "DALL-E 2", params: ["n","size","response_format"], kind: "image" },

View File

@@ -13,7 +13,7 @@ export default {
textIcon: "OC",
website: "https://opencode.ai/auth",
notice: {
text: "OpenCode Go subscription: $5/mo (then 0/mo). Access to Kimi, GLM, Qwen, MiMo, MiniMax models.",
text: "OpenCode Go subscription: $5/mo (then 10/mo). Access to Kimi, GLM, Qwen, MiMo, MiniMax models.",
apiKeyUrl: "https://opencode.ai/auth",
},
},
@@ -33,28 +33,58 @@ export default {
{ format: "claude", baseUrl: "https://opencode.ai/zen/go/v1/messages", auth: { combined: true, header: "x-api-key", scheme: "raw", anthropicVersion: true } },
{ format: "openai-responses", baseUrl: "https://opencode.ai/zen/go/v1/responses", auth: { combined: true, header: "Authorization", scheme: "bearer" } },
],
// supportedFormats follow the endpoint table in https://opencode.ai/docs/go/
models: [
{ id: "deepseek-flash", name: "DeepSeek Flash", supportedFormats: ["openai"] },
{ id: "glm-5.3-flash", name: "GLM 5.3 Flash (Vision)", supportedFormats: ["openai"] },
{ id: "glm-5.3", name: "GLM 5.3", supportedFormats: ["openai"] },
{ id: "glm-5.2", name: "GLM 5.2", supportedFormats: ["openai"] },
{ id: "glm-5.1", name: "GLM 5.1", supportedFormats: ["openai"] },
{ id: "glm-5", name: "GLM 5", supportedFormats: ["openai"] },
{ id: "kimi-k2.7-code", name: "Kimi K2.7 Code", supportedFormats: ["openai"] },
{ id: "kimi-k2.6", name: "Kimi K2.6", supportedFormats: ["openai"] },
{ id: "kimi-k2.5", name: "Kimi K2.5", supportedFormats: ["openai"] },
{ id: "kimi-k3", name: "Kimi K3", supportedFormats: ["openai"] },
{ id: "deepseek-v4-pro", name: "DeepSeek V4 Pro", supportedFormats: ["openai", "claude", "openai-responses"] },
{ id: "deepseek-v4-flash", name: "DeepSeek V4 Flash", supportedFormats: ["openai", "claude", "openai-responses"] },
{ id: "deepseek-v4-flash-vision-exp", name: "DeepSeek V4 Flash Vision (Exp)", supportedFormats: ["openai", "claude", "openai-responses"] },
{ id: "deepseek-v4.1-flash", name: "DeepSeek V4.1 Flash", supportedFormats: ["openai", "claude", "openai-responses"] },
{ id: "longcat-2.0", name: "LongCat 2.0", supportedFormats: ["openai"] },
{ id: "mimo-v2.6-flash", name: "MiMo V2.6 Flash", supportedFormats: ["openai"] },
{ id: "mimo-v2.6-pro", name: "MiMo V2.6 Pro", supportedFormats: ["openai"] },
{ id: "mimo-v2.5", name: "MiMo V2.5", supportedFormats: ["openai"] },
{ id: "mimo-v2.5-pro", name: "MiMo V2.5 Pro", supportedFormats: ["openai"] },
{ id: "mimo-v2-pro", name: "MiMo V2 Pro", supportedFormats: ["openai"] },
{ id: "mimo-v2-omni", name: "MiMo V2 Omni", supportedFormats: ["openai"] },
{ id: "minimax-m3", name: "MiniMax M3", supportedFormats: ["openai", "claude"] },
{ id: "minimax-m2.7", name: "MiniMax M2.7", supportedFormats: ["openai", "claude"] },
{ id: "minimax-m2.5", name: "MiniMax M2.5", supportedFormats: ["openai", "claude"] },
{ id: "space-bunny-free", name: "Space Bunny Free", supportedFormats: ["openai", "claude"] },
{ id: "qwen3.8-max", name: "Qwen 3.8 Max", supportedFormats: ["openai", "claude"] },
{ id: "qwen3.8-flash", name: "Qwen 3.8 Flash", supportedFormats: ["openai", "claude"] },
{ id: "qwen3.7-max", name: "Qwen 3.7 Max", supportedFormats: ["openai", "claude"] },
{ id: "qwen3.7-plus", name: "Qwen 3.7 Plus", supportedFormats: ["openai", "claude"] },
{ id: "qwen3.6-plus", name: "Qwen 3.6 Plus", supportedFormats: ["openai", "claude"] },
// Muse Spark is served by /zen/go/v1/responses only — responses-only entry forces
// chatCore past the sourceFormat-matched transports into translation (see chatCore guard).
{ id: "qwen3.5-plus", name: "Qwen 3.5 Plus", supportedFormats: ["openai", "claude"] },
{ id: "hy4-preview", name: "Hy4 Preview", supportedFormats: ["openai"] },
{ id: "hy3", name: "Hy3", supportedFormats: ["openai"] },
{ id: "hy3-preview", name: "Hy3 Preview", supportedFormats: ["openai"] },
// In /zen/go/v1/models but absent from the docs endpoint table — chat lane is the fallback guess
{ id: "omen-alpha", name: "Omen Alpha", supportedFormats: ["openai"] },
// Served by /zen/go/v1/responses only — the responses-only entry forces chatCore
// past the sourceFormat-matched transports into translation (see chatCore guard).
{ id: "grok-4.7", name: "Grok 4.7", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "grok-4.6", name: "Grok 4.6", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "grok-4.5", name: "Grok 4.5", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.6-luna", name: "GPT 5.6 Luna", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-6-luna", name: "GPT 6 Luna", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "muse-spark-1.2-contributor", name: "Muse Spark 1.2 Contributor", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "muse-spark-1.3-contributor", name: "Muse Spark 1.3 Contributor", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
],
// Live catalogue; ids outside this curated list get their endpoint lane from the
// family regex in providers/models/helpers.js (opencodeFamilyFormats).
modelsFetcher: { url: "https://opencode.ai/zen/go/v1/models", type: "opencode-go" },
passthroughModels: true,
features: {
usage: true,
usageApikey: true,

View File

@@ -0,0 +1,135 @@
export default {
id: "opencode-zen",
priority: 205,
alias: "ocz",
aliases: [
"opencode-zen",
],
uiAlias: "ocz",
display: {
name: "OpenCode Zen",
icon: "terminal",
color: "#E87040",
textIcon: "OZ",
website: "https://opencode.ai/auth",
notice: {
text: "OpenCode Zen PAYG: pay-as-you-go, key from https://opencode.ai/auth. Same models as Zen: paid + free tiers on the fast lane.",
apiKeyUrl: "https://opencode.ai/auth",
},
},
category: "apikey",
transport: {
baseUrl: "https://opencode.ai/zen/v1/chat/completions",
headers: {},
usage: {
url: "https://opencode.ai/zen/v1/usage",
},
},
// Multi-endpoint: pick the transport matching the client sourceFormat to skip
// translation. Mirrors opencode-go, pointed at /zen/v1 (see https://opencode.ai/docs/zen/).
transports: [
{ format: "openai", baseUrl: "https://opencode.ai/zen/v1/chat/completions", auth: { combined: true, header: "Authorization", scheme: "bearer" } },
{ format: "claude", baseUrl: "https://opencode.ai/zen/v1/messages", auth: { combined: true, header: "x-api-key", scheme: "raw", anthropicVersion: true } },
{ format: "openai-responses", baseUrl: "https://opencode.ai/zen/v1/responses", auth: { combined: true, header: "Authorization", scheme: "bearer" } },
],
// supportedFormats follow the endpoint table in https://opencode.ai/docs/zen/
// (live /zen/v1/models, 2026-09-18: 71 ids).
models: [
// Claude (messages)
{ id: "claude-fable-5", name: "Claude Fable 5", supportedFormats: ["claude"] },
{ id: "claude-fable-5-1", name: "Claude Fable 5.1", supportedFormats: ["claude"] },
{ id: "claude-opus-5", name: "Claude Opus 5", supportedFormats: ["claude"] },
{ id: "claude-opus-4-8", name: "Claude Opus 4.8", supportedFormats: ["claude"] },
{ id: "claude-opus-4-7", name: "Claude Opus 4.7", supportedFormats: ["claude"] },
{ id: "claude-opus-4-6", name: "Claude Opus 4.6", supportedFormats: ["claude"] },
{ id: "claude-opus-4-5", name: "Claude Opus 4.5", supportedFormats: ["claude"] },
{ id: "claude-sonnet-5", name: "Claude Sonnet 5", supportedFormats: ["claude"] },
{ id: "claude-sonnet-4-6", name: "Claude Sonnet 4.6", supportedFormats: ["claude"] },
{ id: "claude-sonnet-4-5", name: "Claude Sonnet 4.5", supportedFormats: ["claude"] },
{ id: "claude-sonnet-4", name: "Claude Sonnet 4", supportedFormats: ["claude"] },
{ id: "claude-haiku-4-5", name: "Claude Haiku 4.5", supportedFormats: ["claude"] },
// Gemini (own path, via chat completions transport)
{ id: "gemini-3.6-flash", name: "Gemini 3.6 Flash", supportedFormats: ["openai"] },
{ id: "gemini-3.8-flash", name: "Gemini 3.8 Flash", supportedFormats: ["openai"] },
{ id: "gemini-3.7-flash", name: "Gemini 3.7 Flash", supportedFormats: ["openai"] },
{ id: "gemini-3.5-flash-lite", name: "Gemini 3.5 Flash Lite", supportedFormats: ["openai"] },
{ id: "gemini-3.5-flash", name: "Gemini 3.5 Flash", supportedFormats: ["openai"] },
{ id: "gemini-3.1-pro", name: "Gemini 3.1 Pro", supportedFormats: ["openai"] },
{ id: "gemini-3-flash", name: "Gemini 3 Flash", supportedFormats: ["openai"] },
// GPT / Grok / Muse Spark paid (responses)
{ id: "gpt-6-astra", name: "GPT 6 Astra", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.6-sol", name: "GPT 5.6 Sol", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.6-terra", name: "GPT 5.6 Terra", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.6-luna", name: "GPT 5.6 Luna", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.5", name: "GPT 5.5", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.5-pro", name: "GPT 5.5 Pro", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.4", name: "GPT 5.4", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.4-pro", name: "GPT 5.4 Pro", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.4-mini", name: "GPT 5.4 Mini", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.4-nano", name: "GPT 5.4 Nano", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.3-codex-spark", name: "GPT 5.3 Codex Spark", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.3-codex", name: "GPT 5.3 Codex", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.2", name: "GPT 5.2", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.2-codex", name: "GPT 5.2 Codex", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.1", name: "GPT 5.1", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.1-codex-max", name: "GPT 5.1 Codex Max", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.1-codex", name: "GPT 5.1 Codex", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.1-codex-mini", name: "GPT 5.1 Codex Mini", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5", name: "GPT 5", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5-codex", name: "GPT 5 Codex", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5-nano", name: "GPT 5 Nano", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "grok-build-0.1", name: "Grok Build 0.1", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "grok-4.6", name: "Grok 4.6", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "grok-4.5", name: "Grok 4.5", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "muse-spark-1.3", name: "Muse Spark 1.3", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "muse-spark-1.2", name: "Muse Spark 1.2", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
// Qwen paid (messages)
{ id: "qwen3.6-plus", name: "Qwen 3.6 Plus", supportedFormats: ["claude"] },
{ id: "qwen3.5-plus", name: "Qwen 3.5 Plus", supportedFormats: ["claude"] },
// DeepSeek / GLM / MiniMax / Kimi / Big Pickle (chat completions)
{ id: "deepseek-v4-pro", name: "DeepSeek V4 Pro", supportedFormats: ["openai"] },
{ id: "deepseek-v4-flash", name: "DeepSeek V4 Flash", supportedFormats: ["openai"] },
{ id: "deepseek-v4-flash-vision-exp", name: "DeepSeek V4 Flash Vision Exp", supportedFormats: ["openai"] },
{ id: "glm-5.3-flash", name: "GLM 5.3 Flash (Vision)", supportedFormats: ["openai"] },
{ id: "glm-5.3", name: "GLM 5.3", supportedFormats: ["openai"] },
{ id: "glm-5.2", name: "GLM 5.2", supportedFormats: ["openai"] },
{ id: "glm-5.1", name: "GLM 5.1", supportedFormats: ["openai"] },
{ id: "glm-5", name: "GLM 5", supportedFormats: ["openai"] },
{ id: "minimax-m3", name: "MiniMax M3", supportedFormats: ["openai"] },
{ id: "minimax-m2.7", name: "MiniMax M2.7", supportedFormats: ["openai"] },
{ id: "minimax-m2.5", name: "MiniMax M2.5", supportedFormats: ["openai"] },
{ id: "kimi-k3", name: "Kimi K3", supportedFormats: ["openai"] },
{ id: "kimi-k2.7-code", name: "Kimi K2.7 Code", supportedFormats: ["openai"] },
{ id: "kimi-k2.6", name: "Kimi K2.6", supportedFormats: ["openai"] },
{ id: "kimi-k2.5", name: "Kimi K2.5", supportedFormats: ["openai"] },
{ id: "big-pickle", name: "Big Pickle", supportedFormats: ["openai"] },
{ id: "union-alpha", name: "Union Alpha", supportedFormats: ["claude"] },
// Free tier on the keyed lane (chat completions)
{ id: "deepseek-v4-flash-free", name: "DeepSeek V4 Flash Free", supportedFormats: ["openai"] },
{ id: "mimo-v2.6-flash-free", name: "MiMo V2.6 Flash Free", supportedFormats: ["openai"] },
{ id: "mimo-v2.5-free", name: "MiMo V2.5 Free", supportedFormats: ["openai"] },
{ id: "ling-3.0-flash-fin-free", name: "Ling 3.0 Flash Fin Free", supportedFormats: ["openai"] },
{ id: "nemotron-3-ultra-free", name: "Nemotron 3 Ultra Free", supportedFormats: ["openai"] },
{ id: "nemotron-3.5-lightning-free", name: "Nemotron 3.5 Lightning Free", supportedFormats: ["openai"] },
// Free tier on the keyed lane (responses)
{ id: "muse-spark-1.3-contributor-free", name: "Muse Spark 1.3 Contributor Free", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "muse-spark-1.2-contributor-free", name: "Muse Spark 1.2 Contributor Free", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
// System One (Jev) decision models on the native /systemone endpoint
{ id: "jev-1.13", name: "Jev 1.13", kind: "systemone" },
{ id: "jev-1.13-free", name: "Jev 1.13 Free", kind: "systemone" },
],
serviceKinds: ["llm", "systemone"],
systemoneConfig: {
baseUrl: "https://opencode.ai/zen/v1/systemone",
headers: {
"x-opencode-client": "desktop",
"User-Agent": "opencode/1.18.31",
},
},
modelsFetcher: { url: "https://opencode.ai/zen/v1/models", type: "opencode-free" },
passthroughModels: true,
features: {
usage: true,
usageApikey: true,
},
};

View File

@@ -17,14 +17,27 @@ export default {
headers: {
"x-opencode-client": "desktop",
},
forceStream: true,
noAuth: true,
quirks: {
forceAutoToolChoiceModels: ["muse-spark-1.3-contributor-free"],
},
},
models: [
// Muse Spark models are served by /zen/v1/responses; the rest stay on
// /chat/completions, so the format is declared per-model, not per-provider.
// Endpoint formats differ per model, so declare non-chat models explicitly.
{ id: "muse-spark-1.2-contributor-free", name: "Muse Spark 1.2 Contributor Free", targetFormat: "openai-responses" },
{ id: "muse-spark-1.3-contributor-free", name: "Muse Spark 1.3 Contributor Free", targetFormat: "openai-responses" },
{ id: "union-alpha", name: "Union Alpha Free", targetFormat: "claude" },
{ id: "jev-1.13-free", name: "Jev 1.13 Free", kind: "systemone" },
],
serviceKinds: ["llm", "systemone"],
systemoneConfig: {
baseUrl: "https://opencode.ai/zen/v1/systemone",
headers: {
"x-opencode-client": "desktop",
"User-Agent": "opencode/1.18.31",
},
},
modelsFetcher: { url: "https://opencode.ai/zen/v1/models", type: "opencode-free" },
passthroughModels: true,
};

View File

@@ -40,8 +40,17 @@ export default {
{ id: "openai/gpt-image-1", name: "GPT Image 1 (via OpenRouter)", params: ["n","size","quality","response_format"], kind: "image" },
{ id: "google/imagen-3.0-generate-002", name: "Imagen 3 (via OpenRouter)", params: ["n","size"], kind: "image" },
{ id: "black-forest-labs/FLUX.1-schnell", name: "FLUX.1 Schnell (via OpenRouter)", params: ["n","size"], kind: "image" },
{ id: "google/veo-3.1", name: "Veo 3.1 (via OpenRouter)", params: ["duration","aspect_ratio","resolution"], kind: "video" },
{ id: "openai/sora-2-pro", name: "Sora 2 Pro (via OpenRouter)", params: ["duration","aspect_ratio","resolution"], kind: "video" },
{ id: "bytedance/seedance-2.0", name: "Seedance 2.0 (via OpenRouter)", params: ["duration","aspect_ratio","resolution"], kind: "video" },
{ id: "typesafe/jev-1.13", name: "Jev 1.13", kind: "systemone" },
],
serviceKinds: ["llm","embedding","tts","imageToText"],
serviceKinds: ["llm","embedding","tts","imageToText","video","systemone"],
// System One decision API (TypeSafe-compatible): https://openrouter.ai/docs/guides/community/typesafe-sdk
systemoneConfig: {
baseUrl: "https://openrouter.ai/api/v1/systemone",
headers: {"HTTP-Referer":"https://endpoint-proxy.local","X-Title":"Endpoint Proxy"},
},
ttsConfig: {
baseUrl: "https://openrouter.ai/api/v1/chat/completions",
defaultModel: "openai/gpt-4o-mini-tts",
@@ -57,6 +66,12 @@ export default {
baseUrl: "https://openrouter.ai/api/v1/images/generations",
headers: {"HTTP-Referer":"https://endpoint-proxy.local","X-Title":"Endpoint Proxy"},
},
// Async video jobs (POST /videos → { id, status }, GET /videos/{id} polls).
// Docs: https://openrouter.ai/docs/api/api-reference/videos
videoConfig: {
baseUrl: "https://openrouter.ai/api/v1/videos",
headers: {"HTTP-Referer":"https://endpoint-proxy.local","X-Title":"Endpoint Proxy"},
},
modelsFetcher: { url: "https://openrouter.ai/api/v1/models", type: "openrouter-free" },
passthroughModels: true,
};

View File

@@ -0,0 +1,61 @@
export default {
id: "qoder-cn",
priority: 30,
alias: "qdcn",
uiAlias: "qdcn",
display: {
name: "Qoder CN",
icon: "water_drop",
color: "#EC4899",
website: "https://qoder.com.cn",
notice: {
signupUrl: "https://qoder.com.cn",
},
},
category: "oauth",
authModes: ["oauth", "apikey"],
hasOAuth: true,
authHint: "Personal Access Token (pt-...) from https://qoder.com.cn/account/integrations",
transport: {
baseUrl: "https://gateway.qoder.com.cn/algo/api/v2/service/pro/sse/agent_chat_generation",
headers: {},
timeoutMs: 120000,
stallTimeoutMs: 120000,
usage: {
url: "https://openapi.qoder.com.cn/api/v2/quota/usage",
},
},
models: [
{ id: "ultimate", name: "Ultimate" },
{ id: "auto", name: "Auto" },
{ id: "performance", name: "Performance" },
{ id: "efficient", name: "Efficient" },
{ id: "lite", name: "Lite" },
{ id: "qmodel_38max", name: "Qwen3.8-Max" },
{ id: "qmodel_latest", name: "Qwen3.7-Max" },
{ id: "qmodel", name: "Qwen3.7-Plus" },
{ id: "qfmodel", name: "Qwen3.8-Flash" },
{ id: "kmodel_latest", name: "Kimi-K3" },
{ id: "kmodel", name: "Kimi-K2.7-Code" },
{ id: "gmodel", name: "GLM-5.3" },
{ id: "gfmodel", name: "GLM-5.3-Flash" },
{ id: "dmodel", name: "DeepSeek-V4-Pro" },
{ id: "dfmodel", name: "DeepSeek-V4-Flash" },
{ id: "mmodel", name: "MiniMax-M3" },
],
oauth: {
openApiBaseUrl: "https://openapi.qoder.com.cn",
centerBaseUrl: "https://gateway.qoder.com.cn",
chatBaseUrl: "https://gateway.qoder.com.cn",
deviceTokenUrl: "https://openapi.qoder.com.cn/api/v1/deviceToken/poll",
refreshUrl: "https://gateway.qoder.com.cn/algo/api/v3/user/refresh_token",
userInfoUrl: "https://openapi.qoder.com.cn/api/v1/userinfo",
quotaUsageUrl: "https://openapi.qoder.com.cn/api/v2/quota/usage",
loginUrl: "https://qoder.com.cn/device/selectAccounts",
},
features: {
usage: true,
// PAT (apikey) connections also carry quota usage (via job-token exchange).
usageApikey: true,
},
};

View File

@@ -0,0 +1,49 @@
export default {
id: "tokenharbor",
priority: 120,
alias: "tokenharbor",
aliases: [
"th",
"thh",
],
uiAlias: "tokenharbor",
display: {
name: "Token Harbor",
icon: "anchor",
color: "#0F766E",
textIcon: "TH",
website: "https://tokenharbor.ai",
notice: {
text: "OpenAI-compatible aggregator. One API key reaches every model, billed per-token from a prepaid wallet. Model ids are bare (e.g. claude-opus-5.5, gpt-6-astra, deepseek-v4.1-flash:free) and are fetched live from the provider.",
apiKeyUrl: "https://tokenharbor.ai/dashboard",
},
},
category: "apikey",
authType: "apikey",
transport: {
// OpenAI-compatible. `format` is left at the shared "openai" default and
// `thinkingFormat` is deliberately NOT declared: Token Harbor forwards
// requests verbatim, so each model must resolve its own thinking wire
// format through providers/capabilities.js. Setting a provider-wide value
// would force one format (e.g. claude-adaptive) onto every model.
baseUrl: "https://tokenharbor.ai/v1/chat/completions",
validateUrl: "https://tokenharbor.ai/v1/models",
retry: {
429: 2,
},
},
// Curated seed; the live catalogue is fetched via modelsFetcher and any other
// id is accepted via passthroughModels. Their catalogue rotates (the :free set
// in particular), so this stays deliberately small and is only the offline
// fallback. Ids are bare — Token Harbor does not prefix them by upstream vendor.
models: [
{ id: "claude-opus-5.5", name: "Claude Opus 5.5" },
{ id: "claude-sonnet-5", name: "Claude Sonnet 5" },
{ id: "gpt-6-astra", name: "GPT-6 Astra" },
{ id: "gpt-6-sol", name: "GPT-6 Sol" },
{ id: "deepseek-v4.1-flash:free", name: "DeepSeek V4.1 Flash (Free)" },
{ id: "grok-4.7", name: "Grok 4.7" },
],
modelsFetcher: { url: "https://tokenharbor.ai/v1/models", type: "openai" },
passthroughModels: true,
};

View File

@@ -27,6 +27,13 @@ export default {
{ id: "gemini-3.1-flash-lite-preview", name: "Gemini 3.1 Flash Lite Preview" },
{ id: "gemini-3-flash-preview", name: "Gemini 3 Flash Preview" },
{ id: "gemini-2.5-flash", name: "Gemini 2.5 Flash" },
{ id: "veo-3.1-generate-preview", name: "Veo 3.1 (Preview)", params: ["duration","aspect_ratio","resolution","negative_prompt","seed","storage_uri","generate_audio"], kind: "video" },
{ id: "veo-3.1-fast-generate-preview", name: "Veo 3.1 Fast (Preview)", params: ["duration","aspect_ratio","resolution","negative_prompt","seed","storage_uri","generate_audio"], kind: "video" },
{ id: "veo-3.0-generate-001", name: "Veo 3", params: ["duration","aspect_ratio","resolution","negative_prompt","seed","storage_uri","generate_audio"], kind: "video" },
{ id: "veo-2.0-generate-001", name: "Veo 2", params: ["duration","aspect_ratio","negative_prompt","seed","storage_uri"], kind: "video" },
],
serviceKinds: ["llm","imageToText"],
serviceKinds: ["llm","imageToText","video"],
// Veo via predictLongRunning + fetchPredictOperation (adapter: handlers/videoProviders/vertex.js).
// Docs: https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/veo-video-generation
videoConfig: { baseUrl: "https://aiplatform.googleapis.com" },
};

View File

@@ -1,11 +1,19 @@
import { CLAUDE_API_HEADERS } from "../shared.js";
// Dual auth (same pattern as kimi):
// - API key (sk-...) → cloud API on api.xiaomimimo.com
// - Desktop account/OAuth → same cloud host, plus the dual-route v2.6 models
// served by the account-service route (mimo-server-<cluster>.xiaomimimo.com),
// authorized by a Xiaomi account session cookie, not the key.
// Endpoint is picked per model in the executor, same as opencode-go's /responses split.
export default {
id: "xiaomi-mimo",
priority: 290,
alias: "xiaomi-mimo",
aliases: [
"mimo",
"mimo-desktop",
"xmd",
],
uiAlias: "mimo",
display: {
@@ -16,9 +24,22 @@ export default {
website: "https://xiaomimimo.com",
notice: {
apiKeyUrl: "https://platform.xiaomimimo.com/console/api-keys",
signupUrl: "https://mimo.xiaomimimo.com/desktop/invite/",
},
},
category: "apikey",
category: "oauth",
authModes: ["oauth", "apikey"],
hasOAuth: true,
// Keys are cluster-specific. MiMo Desktop declares five regions
// (CN/SGP/AMS/RU/IN) — host + sid follow mimo-server-<code> / mimo<code>.
regions: [
{ id: "cn", label: "China (中国大陆)" },
{ id: "sgp", label: "Singapore (新加坡)" },
{ id: "ams", label: "Europe · Amsterdam (欧洲)" },
{ id: "ru", label: "Russia (俄罗斯)" },
{ id: "in", label: "India (印度)" },
],
defaultRegion: "sgp",
serviceKinds: ["llm", "tts"],
transport: {
baseUrl: "https://api.xiaomimimo.com/v1/chat/completions",
@@ -39,6 +60,11 @@ export default {
},
],
models: [
// Cloud API & Desktop dual-route models (prefers the desktop account quota when available)
{ id: "mimo-v2.6-pro", name: "MiMo V2.6 Pro", upstreamModelId: "xiaomi/mimo-v2.6-pro", supportedFormats: ["openai"] },
{ id: "mimo-v2.6-flash", name: "MiMo V2.6 Flash", upstreamModelId: "xiaomi/mimo-v2.6-flash", supportedFormats: ["openai"] },
{ id: "mimo-v2.6-pro-ultraspeed", name: "MiMo V2.6 Pro UltraSpeed", upstreamModelId: "xiaomi/mimo-v2.6-pro-ultraspeed", supportedFormats: ["openai"] },
// Cloud API models (api.xiaomimimo.com/v1)
{ id: "mimo-v2.5-pro", name: "MiMo V2.5 Pro" },
{ id: "mimo-v2.5", name: "MiMo V2.5" },
{ id: "mimo-v2-omni", name: "MiMo V2 Omni" },
@@ -51,4 +77,18 @@ export default {
authHeader: "bearer",
format: "xiaomi-mimo-tts",
},
features: {
usage: true,
usageApikey: true,
},
// Custom OAuth — non-standard ECDH encrypted-callback flow.
// Handled by the Xiaomi MiMo OAuth service, not the generic PKCE pipeline.
oauth: {
custom: true,
authorizeUrl: "https://platform.xiaomimimo.com/authorize",
// The callback carries ?u=<ECDH-encrypted payload> instead of ?code=.
// Decryption yields { uid, sk, url }.
callbackParam: "u",
kn: "mimocode",
},
};

View File

@@ -1,10 +1,9 @@
// Zed provider — RSA keypair callback auth (NOT standard OAuth).
export default {
id: "zed",
priority: 10,
priority: 999,
alias: "zd",
uiAlias: "zd",
hidden: true,
display: {
name: "Zed",
icon: "code",

View File

@@ -36,6 +36,13 @@ import { DEFAULT_RETRY_CONFIG, FETCH_CONNECT_TIMEOUT_MS } from "../config/runtim
* MediaConfig: { serviceKinds:[...], ttsConfig, sttConfig, embeddingConfig, imageConfig,
* searchViaChat:{defaultModel,pricingUrl}, hiddenKinds } — each *Config: {baseUrl,authType,authHeader,
* format,defaultModel,models:[{id,name,dimensions?}]}.
*
* imageConfig.modelMap (optional): maps a client-facing model id to a provider-resolved id when
* those differ — e.g. the HuggingFace Inference Providers router, where a Hub id like
* `black-forest-labs/FLUX.1-schnell` is addressed as `fal-ai/fal-ai/flux/schnell`. A value is
* either the provider path, or `{path, task}` when the request shape differs per task
* (HuggingFace uses task:"image-to-image" to move the prompt under `parameters.prompt`).
* Ignored by providers whose model ids are sent verbatim.
*/
// Shared transport defaults — provider only overrides fields that differ.

View File

@@ -22,7 +22,7 @@ export function mapStainlessArch() {
// Anthropic API version (single source — reused across claude-format providers/executors)
export const ANTHROPIC_API_VERSION = "2023-06-01";
export const CLAUDE_CLI_VERSION = "2.1.258";
export const CLAUDE_CLI_VERSION = "2.1.280";
// Shared Claude-compatible API headers (reused across claude-format providers)
export const CLAUDE_API_HEADERS = {
@@ -62,12 +62,26 @@ const ANTHROPIC_BETA_BASE = [
const ANTHROPIC_BETA_HEAVY_AGENT = ["advanced-tool-use-2025-11-20", "effort-2025-11-24"];
// Heavy-agent beta flags are gated to opus/sonnet — cheaper models don't need them.
export function selectAnthropicBeta(model = "") {
const flags = [...ANTHROPIC_BETA_BASE];
// `redact-thinking` asks Anthropic to return signature-only thinking blocks, which
// is right for clients that never render thinking but blanks the summaries a
// client explicitly requested with `thinking.display: "summarized"`.
const ANTHROPIC_BETA_REDACT_THINKING = "redact-thinking-2026-02-12";
export function wantsThinkingSummaries(body) {
return body?.thinking?.display === "summarized";
}
export function selectAnthropicBeta(model = "", body = null) {
const flags = ANTHROPIC_BETA_BASE.filter((flag) => flag !== ANTHROPIC_BETA_REDACT_THINKING || !wantsThinkingSummaries(body));
if (/^claude-(opus|sonnet)/.test(model)) flags.push(...ANTHROPIC_BETA_HEAVY_AGENT);
return flags.join(",");
}
export function mergeAnthropicBeta(...values) {
const flags = values.flatMap((v) => (typeof v === "string" ? v.split(",") : [])).map((f) => f.trim()).filter(Boolean);
return [...new Set(flags)].join(",");
}
// Shared baseUrls
export const KIMI_CODING_BASE_URL = "https://api.kimi.com/coding/v1/messages";

View File

@@ -3,6 +3,7 @@
import { getCapabilitiesForModel } from "./capabilities.js";
import { matchPattern } from "./pricing.js";
import { resolveKiroEffortPath } from "../config/kiroConstants.js";
import { getProviderModels } from "../config/providerModels.js";
// Shared level sets (deduped) — verified against provider docs + wire in thinkingUnified.applyFormat.
const L = {
@@ -26,6 +27,7 @@ const FORMAT_LEVELS = {
qwen: L.base,
kimi: L.levelMax,
deepseek: L.hiMax,
commandcode: ["none", "low", "medium", "high", "xhigh", "max"],
minimax: L.onOff,
hunyuan: L.base,
step: L.base,
@@ -40,6 +42,13 @@ const PATTERN_THINKING = [
{ provider: "codex", pattern: "*gpt-5.6-terra*", levels: [...CODEX_GPT_5_6_LEVELS, "ultra"] },
{ provider: "codex", pattern: "*gpt-5.6-luna*", levels: CODEX_GPT_5_6_LEVELS },
{ pattern: "*codex*", levels: ["low", "medium", "high", "xhigh"] }, // codex cannot disable thinking
{ pattern: "*mimo*v2.6*", levels: ["none", "low", "medium", "high", "xhigh"] },
// mimo-v2.5-pro on opencode-go rejects reasoning_effort "max" (probed live); v2.5 accepts it.
{ pattern: "*mimo*v2.5-pro*", levels: ["none", "low", "medium", "high", "xhigh"] },
// DeepSeek v4.* (Alibaba MaaS, probed live): effort low|medium|high|xhigh|max
// all 200 via output_config.effort; "none" is a 400 on the anthropic route
// (disable thinking instead). none kept for the picker = disable.
{ pattern: "*deepseek-v4.*", levels: ["none", "low", "medium", "high", "xhigh", "max"] },
// codebuddy-cn per-model effort sets — the server's product-config payload
// publishes `reasoning.supportedEfforts` per model. NOTE: the chat endpoint
// accepts any level you send (probed none/minimal/low/medium/high/xhigh/max
@@ -52,17 +61,29 @@ const PATTERN_THINKING = [
{ provider: "codebuddy-cn", pattern: "deepseek-v4*", levels: ["low", "high", "xhigh"] },
{ provider: "codebuddy-cn", pattern: "hy3*", levels: ["low", "high"] },
{ provider: "codebuddy-cn", pattern: "hy4*", levels: ["high"] },
// codebuddy-intl rides the same gateway catalog, so its deepseek levels match.
{ provider: "codebuddy-intl", pattern: "deepseek-v4*", levels: ["low", "high", "xhigh"] },
];
// The generic level set used when a model's thinking format is unknown. Exported
// for UI callers that must list levels for a model whose reasoning capability is
// user-asserted: the browser bundle has no access to the capability store, so it
// cannot derive a format the way getThinkingLevels does on the server.
export const BASE_THINKING_LEVELS = L.base;
// Returns valid thinking levels for a model, or null when the model has no reasoning.
export function getThinkingLevels(provider, model) {
if (provider === "kiro" && resolveKiroEffortPath(model) === null) return null;
const caps = getCapabilitiesForModel(provider, model);
if (!caps.reasoning) return null;
const baseId = String(model || "").replace(/\([^()]+\)\s*$/, "");
const modelLevels = provider === "codex"
? getProviderModels("cx").find((entry) => entry.id === baseId)?.thinkingLevels
: null;
const hit = PATTERN_THINKING.find((entry) =>
(!entry.provider || entry.provider === provider) && matchPattern(entry.pattern, model)
);
let levels = hit?.levels || FORMAT_LEVELS[caps.thinkingFormat] || L.base;
let levels = modelLevels || hit?.levels || FORMAT_LEVELS[caps.thinkingFormat] || L.base;
if (caps.thinkingCanDisable === false) levels = levels.filter((l) => l !== "none");
return levels;
}

View File

@@ -1,5 +1,5 @@
// RTK port: compress tool_result content in LLM request bodies
// Injected at the top of translateRequest (before any format translation)
// Applied in chatCore on the source-format body, before translateRequest.
import { RAW_CAP, MIN_COMPRESS_SIZE } from "./constants.js";
import { autoDetectFilter } from "./autodetect.js";
import { safeApply } from "./applyFilter.js";

View File

@@ -13,7 +13,7 @@ export function injectSystemPrompt(body, format, prompt) {
if (!body || !prompt) return;
if (typeof body !== "object") return;
// Kiro wire shape is unique (conversationState/systemPrompt) — handle directly.
// Kiro wire shape is unique (conversationState) — handle directly.
if (isKiroBody(body) || format === FORMATS.KIRO) {
injectKiroSystem(body, prompt);
return;
@@ -61,10 +61,13 @@ export function injectSystemPrompt(body, format, prompt) {
function isKiroBody(body) {
if (!body || typeof body !== "object") return false;
if (typeof body.systemPrompt !== "string") return false;
const cs = body.conversationState;
if (!cs || typeof cs !== "object") return false;
return Array.isArray(cs.history) || !!(cs.currentMessage && typeof cs.currentMessage === "object");
// A top-level `systemPrompt` used to be the marker, but the Kiro translator no
// longer emits it (kiro.dev rejects the field), so gate on the turn shape.
const historyTurn = Array.isArray(cs.history)
&& cs.history.some(it => it && (it.userInputMessage || it.assistantResponseMessage));
return historyTurn || !!(cs.currentMessage && cs.currentMessage.userInputMessage);
}
// Exact idempotency: prompt present as its own SEP-delimited segment (or the
@@ -258,80 +261,33 @@ function injectGeminiSystem(body, prompt) {
}
// ---- Kiro ----
// Updates top-level systemPrompt and only the mirrored leading prefix of the
// first user history turn, else current user. next = old + SEP + prompt.
// Replace old leading prefix only; preserve time context and user tail.
// The prompt is appended to the first user turn's content — the same place the
// Kiro translator already mirrors the system text via its contentPrefix.
//
// A top-level `systemPrompt` is deliberately NOT written: the kiro.dev gateway
// answers any body carrying that field with
// 400 {"message":"Improperly formed request.","reason":"REQUEST_BODY_INVALID"}
// The translator stopped emitting it in v0.5.59, but this injector kept adding
// it back, so every kr/ model failed whenever an RTK prompt (caveman, ponytail)
// was active.
function injectKiroSystem(body, prompt) {
try {
let oldPrompt = typeof body.systemPrompt === "string" ? body.systemPrompt : "";
// Repair path: a previous partial write left systemPrompt updated but user
// content still mirroring the pre-write prefix. Re-derive the effective old
// prefix from content so this pass converges instead of early-returning.
const cs0 = body.conversationState;
let firstUser0 = cs0 && Array.isArray(cs0.history)
? (cs0.history.find(it => it && it.userInputMessage)?.userInputMessage ?? null)
: null;
if (!firstUser0 && cs0?.currentMessage?.userInputMessage) firstUser0 = cs0.currentMessage.userInputMessage;
if (firstUser0 && typeof firstUser0.content === "string" && oldPrompt && !hasPrompt(oldPrompt, prompt)) {
const c0 = firstUser0.content;
if (c0 === oldPrompt || (c0.startsWith(oldPrompt) && !c0.startsWith(`${oldPrompt}${SEP}`))) {
// systemPrompt advanced past mirrored prefix → stale; treat as un-mirrored
oldPrompt = "";
}
}
if (oldPrompt && hasPrompt(oldPrompt, prompt)) return;
const next = oldPrompt ? `${oldPrompt}${SEP}${prompt}` : prompt;
// Atomicity: write user content first, then systemPrompt only if content
// write succeeded (or was a no-op). If systemPrompt write then fails, the
// repair heuristic above re-derives from content on retry — no permanent
// half-applied state.
const cs = body.conversationState;
let targetMsg = null;
try {
const hist = Array.isArray(cs?.history) ? cs.history : null;
if (hist) {
for (const item of hist) {
if (item && item.userInputMessage) { targetMsg = item.userInputMessage; break; }
}
}
if (!targetMsg && cs?.currentMessage?.userInputMessage) {
targetMsg = cs.currentMessage.userInputMessage;
}
} catch (_) { targetMsg = null; }
let sysWritten = false;
try { body.systemPrompt = next; sysWritten = true; } catch (_) {}
const applyContent = () => {
const content = typeof targetMsg.content === "string" ? targetMsg.content : "";
if (oldPrompt === "") {
// Empty old prompt: prepend unless already at head (exact, not substring)
if (content.startsWith(prompt) || content.startsWith(next)) return;
const newContent = content ? `${next}${SEP}${content}` : next;
try { targetMsg.content = newContent; } catch (_) {}
return;
}
if (!content.startsWith(oldPrompt)) return; // not mirrored at head — leave alone
if (content.startsWith(next)) return; // already applied → idempotent
const tail = content.slice(oldPrompt.length);
try { targetMsg.content = `${next}${tail}`; } catch (_) {}
};
try {
if (targetMsg) applyContent();
} catch (_) {}
if (sysWritten && targetMsg) {
// verify convergence: content should now start with next (or be un-mirrored)
let ok = false;
try {
const c = targetMsg.content;
ok = typeof c !== "string" || c.startsWith(next) || !c.startsWith(oldPrompt);
} catch (_) {}
if (!ok) {
try { body.systemPrompt = oldPrompt; } catch (_) {} // rollback
const hist = Array.isArray(cs?.history) ? cs.history : null;
if (hist) {
for (const item of hist) {
if (item && item.userInputMessage) { targetMsg = item.userInputMessage; break; }
}
}
if (!targetMsg && cs?.currentMessage?.userInputMessage) {
targetMsg = cs.currentMessage.userInputMessage;
}
if (!targetMsg) return;
const content = typeof targetMsg.content === "string" ? targetMsg.content : "";
const next = dedupStringAppend(content, prompt);
if (next === content) return; // already injected — idempotent across retries
try { targetMsg.content = next; } catch (_) { /* frozen/proxy fail-open */ }
} catch (_) {}
}

View File

@@ -29,7 +29,7 @@ export function checkFallbackError(status, errorText, backoffLevel = 0) {
// Request-scoped rule: the request body itself is at fault — no cooldown,
// no account lock. Caller must stop rotating and surface the error.
if (rule.requestScoped && lowerError && lowerError.includes(rule.text)) {
return { shouldFallback: false, requestScoped: true, cooldownMs: 0 };
return { shouldFallback: false, cooldownMs: 0 };
}
// Text-based rule: match substring in error message
@@ -51,6 +51,20 @@ export function checkFallbackError(status, errorText, backoffLevel = 0) {
}
}
// Request-scoped client errors that matched no rule above: a 400 caused by the
// request itself (context overflow, malformed body, unsupported parameter) says
// nothing about the credential, so cooling the account down only removes a
// healthy connection from rotation. With a single connection it is worse: every
// later request in the window fails with a copy of this very error
// ("all 1 accounts locked for <model> | lastError=[400]: ..."), which hides the
// real cause from the caller and makes unrelated sessions look like they hit the
// same limit. Hand the upstream error back for this request instead.
// Account-scoped statuses keep their rules above (401/402/403/404/429), and the
// text rules still win for rate-limit / quota / capacity wording.
if (status >= 400 && status < 500 && status !== 401 && status !== 402 && status !== 403 && status !== 429) {
return { shouldFallback: false, cooldownMs: 0 };
}
// Default: transient cooldown for any unmatched error
return { shouldFallback: true, cooldownMs: TRANSIENT_COOLDOWN_MS };
}

View File

@@ -12,19 +12,20 @@ import { getCapabilitiesForModel } from "../providers/capabilities.js";
const CAPABILITY_KEYS = ["vision", "pdf", "audioInput", "videoInput"];
const HARD_CAPS = new Set(CAPABILITY_KEYS);
const DEFAULT_FALLBACK_MODEL = "oc/mimo-v2.5-free";
const DEFAULT_FALLBACK_MODEL = "oc/mimo-v2.6-flash-free";
const upgradeLegacyModel = (m) => (m === "oc/mimo-v2.5-free" ? DEFAULT_FALLBACK_MODEL : m);
// Normalize a capability entry to { enabled, roundRobin, models }. Backward-compat:
// accept the legacy array form [{model, enabled}] (treated as enabled, fallback).
function normalizeCapEntry(entry) {
if (Array.isArray(entry)) {
return { enabled: true, roundRobin: false, models: entry.map((e) => e?.model || e).filter(Boolean) };
return { enabled: true, roundRobin: false, models: entry.map((e) => upgradeLegacyModel(e?.model || e)).filter(Boolean) };
}
if (entry && typeof entry === "object") {
return {
enabled: entry.enabled !== false,
roundRobin: !!entry.roundRobin,
models: Array.isArray(entry.models) ? entry.models.filter(Boolean) : [],
models: Array.isArray(entry.models) ? entry.models.map(upgradeLegacyModel).filter(Boolean) : [],
};
}
return { enabled: false, roundRobin: false, models: [] };

View File

@@ -1,6 +1,12 @@
import { buildClineHeaders } from "../shared/clineAuth.js";
const CLINEPASS_MODELS_ENDPOINT = "https://api.cline.bot/api/v1/models";
// Cline's free tier is published here, not in /api/v1/models: the catalog
// endpoint carries no `cline-free/*` ids at all. Cline's own SDK calls this
// feed unauthenticated (sdk/packages/core/src/services/llms/cline-recommended-models.ts),
// so no Authorization header is sent — adding one would only make the request
// fail on a header the endpoint ignores.
const CLINE_RECOMMENDED_MODELS_ENDPOINT = "https://api.cline.bot/api/v1/ai/cline/recommended-models";
const FETCH_TIMEOUT_MS = 5000;
/**
@@ -19,12 +25,10 @@ function buildModelListHeaders(token, isApiKey) {
}
/**
* Fetch ClinePass live model catalog from Cline's /models endpoint.
*
* @param {object} credentials - Connection credentials ({ accessToken, apiKey })
* @returns {Promise<{ models: { id: string, name: string }[] } | null>}
* Internal: fetch the raw model list from Cline's /models endpoint.
* Returns the parsed array or null on any failure.
*/
export async function resolveClinepassModels(credentials) {
async function fetchClineRawModels(credentials) {
const isApiKey = Boolean(credentials?.apiKey);
const token = isApiKey ? credentials.apiKey : credentials?.accessToken;
if (!token) return null;
@@ -45,19 +49,97 @@ export async function resolveClinepassModels(credentials) {
const json = await response.json();
const rawList = Array.isArray(json) ? json : json?.data;
if (!Array.isArray(rawList)) return null;
const models = rawList
.filter((m) => typeof m?.id === "string" && m.id.startsWith("cline-pass/"))
.map((m) => ({
id: m.id,
name: m.name || m.id,
}));
return models.length ? { models } : null;
return Array.isArray(rawList) ? rawList : null;
} catch {
return null;
} finally {
clearTimeout(timer);
}
}
/**
* Fetch ClinePass live model catalog from Cline's /models endpoint.
* Returns only models with the cline-pass/ prefix.
*
* @param {object} credentials - Connection credentials ({ accessToken, apiKey })
* @returns {Promise<{ models: { id: string, name: string }[] } | null>}
*/
export async function resolveClinepassModels(credentials) {
const rawList = await fetchClineRawModels(credentials);
if (!rawList) return null;
const models = rawList
.filter((m) => typeof m?.id === "string" && m.id.startsWith("cline-pass/"))
.map((m) => ({
id: m.id,
name: m.name || m.id,
}));
return models.length ? { models } : null;
}
/**
* Fetch Cline's recommended-models feed and return only its `free[]` tier.
* Returns null on any failure — the free tier is additive, so a dead feed must
* never take the /api/v1/models catalog down with it.
* @param {{accessToken?: string, apiKey?: string}} credentials
* @returns {Promise<{id: string, name: string}[] | null>}
*/
async function fetchClineFreeTierModels() {
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), FETCH_TIMEOUT_MS);
try {
const response = await fetch(CLINE_RECOMMENDED_MODELS_ENDPOINT, {
method: "GET",
headers: { Accept: "application/json" },
signal: controller.signal,
});
if (!response.ok) return null;
const json = await response.json();
const free = Array.isArray(json?.free) ? json.free : [];
if (!free.length) return null;
return free
.filter((m) => typeof m?.id === "string" && m.id.trim() !== "")
.map((m) => ({ id: m.id, name: m.name || m.id }));
} catch {
return null;
} finally {
clearTimeout(timer);
}
}
/**
* Fetch Cline live model catalog from Cline's /models endpoint.
* Unlike resolveClinepassModels, this returns ALL models (including
* free-tier models like z-ai/glm-5.3-flash) without the cline-pass/ prefix filter.
*
* @param {object} credentials - Connection credentials ({ accessToken, apiKey })
* @returns {Promise<{ models: { id: string, name: string }[] } | null>}
*/
export async function resolveClineModels(credentials) {
const rawList = await fetchClineRawModels(credentials);
if (!rawList) return null;
const models = rawList
.filter((m) => typeof m?.id === "string" && m.id.trim() !== "")
.map((m) => ({
id: m.id,
name: m.name || m.id,
}));
// Free tier: /api/v1/models lists no `cline-free/*` ids, so merge the feed's
// free[] in. First writer wins on a shared id, keeping the catalog's entry
// for anything the two sources agree on.
const freeTier = await fetchClineFreeTierModels();
const byId = new Map(models.map((m) => [m.id, m]));
for (const m of freeTier || []) {
if (!byId.has(m.id)) byId.set(m.id, m);
}
const merged = Array.from(byId.values());
return merged.length ? { models: merged } : null;
}

View File

@@ -246,23 +246,39 @@ export function resetComboRotation(comboName) {
}
/**
* Get combo models from combos data
* Cancel an unused Response body so its SSE transform (and usage callback)
* is not left hanging. Fusion timeouts and combo fallback both drop Response
* objects without reading them; an unread stream never fires onStreamComplete,
* which is how request-details rows get stuck at input/output tokens = 0.
*/
export function discardResponse(res) {
if (!res || res.__timeout || res.__error) return;
const cancel = res.body?.cancel;
if (typeof cancel === "function") {
try { cancel.call(res.body); } catch { /* already closed / consumed */ }
}
}
/**
* Get combo models from combos data.
* Nested combo names are kept as-is (comboA listing comboB yields ["comboB", ...]).
* Flattening them into leaves would explode fusion/round-robin: every nested
* leaf becomes its own panel/rotation slot. Nested combos run as one unit
* (see handleSingleModelChat: inner strategy forced to fallback).
*
* @param {string} modelStr - Model string to check
* @param {Array|Object} combosData - Array of combos or object with combos
* @returns {string[]|null} Array of models or null if not a combo
* @returns {string[]|null} Direct members, or null if not a combo
*/
export function getComboModelsFromData(modelStr, combosData) {
// Don't check if it's in provider/model format
if (modelStr.includes("/")) return null;
// Handle both array and object formats
if (typeof modelStr !== "string" || modelStr.includes("/")) return null;
const combos = Array.isArray(combosData) ? combosData : (combosData?.combos || []);
const combo = combos.find(c => c.name === modelStr);
if (combo && combo.models && combo.models.length > 0) {
return combo.models;
const combo = combos.find((c) => c.name === modelStr);
if (!combo || combo.enabled === false || !Array.isArray(combo.models) || combo.models.length === 0) {
return null;
}
return null;
return combo.models;
}
/**
@@ -339,6 +355,9 @@ export async function handleComboChat({ body, models, handleSingleModel, log, co
return result;
}
// Falling through: this body will never be read by the client.
discardResponse(result);
// For transient errors (503/502/504), wait for cooldown before falling through
// so a briefly-overloaded provider gets a chance to recover rather than being
// skipped immediately (fixes: combo falls through on transient 503)
@@ -509,8 +528,13 @@ function collectPanel(calls, { minPanel, stragglerGraceMs, panelHardTimeoutMs })
const hardTimer = setTimeout(finish, panelHardTimeoutMs);
calls.forEach((p, i) => {
Promise.resolve(p)
.then((v) => { out[i] = v; })
.catch((e) => { out[i] = { __error: e }; })
.then((v) => {
// Quorum/timeout already moved on: cancel the late body so its
// streaming usage callback is not left at tokens=0 forever.
if (finished) discardResponse(v);
else out[i] = v;
})
.catch((e) => { if (!finished) out[i] = { __error: e }; })
.finally(() => {
settled++;
if (out[i] && out[i].ok) ok++;
@@ -590,7 +614,7 @@ export async function handleFusionChat({ body, models, handleSingleModel, log, c
if (!res) { log.warn("FUSION", `Panel ${model} dropped (straggler/timeout)`); continue; }
if (res.__timeout) { log.warn("FUSION", `Panel ${model} timed out`); continue; }
if (res.__error) { log.warn("FUSION", `Panel ${model} threw`, { error: res.__error?.message || String(res.__error) }); continue; }
if (!res.ok) { log.warn("FUSION", `Panel ${model} failed`, { status: res.status }); continue; }
if (!res.ok) { log.warn("FUSION", `Panel ${model} failed`, { status: res.status }); discardResponse(res); continue; }
try {
const json = await res.clone().json();
const text = extractPanelText(json);
@@ -602,6 +626,11 @@ export async function handleFusionChat({ body, models, handleSingleModel, log, c
}
} catch (e) {
log.warn("FUSION", `Panel ${model} unparseable`, { error: e.message || String(e) });
} finally {
// Panel responses are never returned to the client; cancel the original
// body (clone().json() already captured the payload) so streaming
// placeholders are not left as success/0-token rows.
discardResponse(res);
}
}

View File

@@ -124,6 +124,8 @@ export async function getModelInfoCore(modelStr, aliasesOrGetter) {
// Config-driven prefix → provider inference (first match wins, fallback "openai").
const MODEL_PREFIX_PROVIDERS = [
// Codex CLI sends this bare virtual model for auto-review — keep it on OAuth Codex (#1398).
[/^codex-auto-review$/, "codex"],
[/^claude-/, "anthropic"],
[/^gemini-/, "gemini"],
[/^gpt-/, "openai"],

View File

@@ -13,9 +13,13 @@
*
* PAT (Personal Access Token, pt-...) connections: a PAT cannot sign COSY
* requests directly, so we exchange it for a short-lived job token (jt-...)
* via openapi.qoder.sh/api/v1/jobToken/exchange (plain JSON POST), then use
* that job token for signing. Job-token traffic must hit api2.qoder.sh —
* api3 rejects jt- with "Login expired" (403).
* via the region's jobToken/exchange endpoint (plain JSON POST), then use
* that job token for signing. On intl, job-token traffic must hit api2.qoder.sh —
* api3 rejects jt- with "Login expired" (403); CN serves it from the same
* gateway host.
*
* The region (intl/cn) is derived from credentials.provider (or an explicit
* options.region override) so the same catalog logic works for both sites.
*/
import { createHash } from "crypto";
@@ -23,12 +27,12 @@ import { createHash } from "crypto";
import { proxyAwareFetch } from "../utils/proxyFetch.js";
import { buildCosyHeaders } from "../shared/qoder/cosy.js";
import {
QODER_MODEL_LIST_URL,
QODER_CHAT_BASE_ALT,
QODER_JOB_TOKEN_EXCHANGE_URL,
QODER_USERINFO_URL,
QODER_IDE_VERSION,
QODER_CLIENT_TYPE,
qoderRegionOf,
qoderJobTokenExchangeUrl,
qoderUserInfoUrl,
qoderInferenceBase,
} from "../shared/qoder/constants.js";
const FETCH_TIMEOUT_MS = 15_000;
@@ -63,9 +67,9 @@ const inflight = new Map();
* Exchange a Qoder PAT (pt-...) for a short-lived job token (jt-...).
* This endpoint is plain JSON POST — NOT COSY-signed.
*/
async function exchangeJobToken(pat, proxyOptions = null, signal = null) {
async function exchangeJobToken(pat, proxyOptions = null, signal = null, region = "intl") {
const res = await proxyAwareFetch(
QODER_JOB_TOKEN_EXCHANGE_URL,
qoderJobTokenExchangeUrl(region),
{
method: "POST",
headers: {
@@ -101,10 +105,10 @@ async function exchangeJobToken(pat, proxyOptions = null, signal = null) {
* Resolve the Qoder userId for a job token (needed for COSY signing).
* Returns "" on any failure — callers fall back to the stored userId.
*/
async function fetchUserIdForJobToken(jobToken, proxyOptions = null, signal = null) {
async function fetchUserIdForJobToken(jobToken, proxyOptions = null, signal = null, region = "intl") {
try {
const res = await proxyAwareFetch(
QODER_USERINFO_URL,
qoderUserInfoUrl(region),
{
method: "GET",
headers: {
@@ -125,16 +129,17 @@ async function fetchUserIdForJobToken(jobToken, proxyOptions = null, signal = nu
}
/**
* Resolve a PAT to a job-token credential, cached per-PAT.
* Resolve a PAT to a job-token credential, cached per-PAT-per-region.
*/
async function resolvePatCredential(pat, proxyOptions = null, signal = null) {
const cached = patJobCache.get(pat);
async function resolvePatCredential(pat, proxyOptions = null, signal = null, region = "intl") {
const cacheKey = `${region}:${pat}`;
const cached = patJobCache.get(cacheKey);
if (cached && cached.expiresAt - Date.now() > PAT_REFRESH_BUFFER_MS) return cached;
const { jobToken, expiresAt } = await exchangeJobToken(pat, proxyOptions, signal);
const userId = await fetchUserIdForJobToken(jobToken, proxyOptions, signal);
const { jobToken, expiresAt } = await exchangeJobToken(pat, proxyOptions, signal, region);
const userId = await fetchUserIdForJobToken(jobToken, proxyOptions, signal, region);
const resolved = { accessToken: jobToken, userId, expiresAt };
patJobCache.set(pat, resolved);
patJobCache.set(cacheKey, resolved);
return resolved;
}
@@ -142,11 +147,14 @@ async function resolvePatCredential(pat, proxyOptions = null, signal = null) {
* Resolve connection credentials to COSY-signable form:
* - PAT (pt-...) connections → exchanged to a job token (jt-...) + userId
* - everything else → passed through unchanged
*
* Region defaults to the one implied by credentials.provider (qoder-cn → cn).
*/
export async function resolveQoderCredentials(credentials, proxyOptions = null, signal = null) {
export async function resolveQoderCredentials(credentials, proxyOptions = null, signal = null, region) {
const raw = credentials?.apiKey || credentials?.accessToken;
if (isQoderPat(raw)) {
const resolved = await resolvePatCredential(raw, proxyOptions, signal);
const effRegion = region || qoderRegionOf(credentials?.provider);
const resolved = await resolvePatCredential(raw, proxyOptions, signal, effRegion);
return {
...credentials,
accessToken: resolved.accessToken,
@@ -163,13 +171,14 @@ export async function resolveQoderCredentials(credentials, proxyOptions = null,
}
/**
* Stable cache key per credential (so different login sessions for the same
* account share an entry).
* Stable cache key per credential+region (so different login sessions for the
* same account share an entry, and the same PAT on both sites stays apart).
*/
function cacheKey(credentials) {
const psd = credentials?.providerSpecificData || {};
const seed = psd.userId || credentials?.refreshToken || credentials?.accessToken || "anonymous";
return createHash("sha256").update(`qoder:${seed}`).digest("hex");
const region = qoderRegionOf(credentials?.provider);
return createHash("sha256").update(`qoder:${region}:${seed}`).digest("hex");
}
/**
@@ -192,15 +201,13 @@ function cosyCredsFromConnection(credentials) {
* rawConfigs: Map<modelKey, modelConfigObject> }
* or `null` on any error.
*/
async function fetchQoderCatalogRaw(credentials, signal, proxyOptions = null) {
async function fetchQoderCatalogRaw(credentials, signal, proxyOptions = null, region = "intl") {
const creds = cosyCredsFromConnection(credentials);
if (!creds.userId || !creds.authToken) return null;
// Job-token traffic is rejected by api3 ("Login expired" 403) — the
// official qodercli serves it from api2 instead.
const modelListUrl = String(creds.authToken).startsWith("jt-")
? `${QODER_CHAT_BASE_ALT}/algo/api/v2/model/list`
: QODER_MODEL_LIST_URL;
// Intl job-token traffic is rejected by api3 ("Login expired" 403) — the
// official qodercli serves it from api2 instead; CN uses the single gateway.
const modelListUrl = `${qoderInferenceBase(credentials, region)}/algo/api/v2/model/list`;
const headers = {
Accept: "application/json",
@@ -293,14 +300,20 @@ export async function getQoderModelConfig(credentials, modelKey, options = {}) {
* one upstream request per credential.
*/
export async function resolveQoderModels(credentials, options = {}) {
const region = options.region || qoderRegionOf(credentials?.provider);
let resolved;
try {
resolved = await resolveQoderCredentials(credentials, options.proxyOptions, options.signal);
resolved = await resolveQoderCredentials(credentials, options.proxyOptions, options.signal, region);
} catch (error) {
options.log?.warn?.("QODER", `PAT exchange failed: ${error.message}`);
return null;
}
if (!resolved?.accessToken || !(resolved.providerSpecificData || {}).userId) return null;
// Stamp the provider so cacheKey/catalog derive the region even when the
// caller's credentials object didn't carry a provider id (e.g. /v1/models).
if (resolved && !resolved.provider) {
resolved.provider = region === "cn" ? "qoder-cn" : "qoder";
}
const key = cacheKey(resolved);
const now = Date.now();
@@ -319,7 +332,7 @@ export async function resolveQoderModels(credentials, options = {}) {
}
const fetchPromise = (async () => {
const fetched = await fetchQoderCatalogRaw(resolved, options.signal, options.proxyOptions);
const fetched = await fetchQoderCatalogRaw(resolved, options.signal, options.proxyOptions, region);
if (!fetched) return null;
const entry = {
expiresAt: Date.now() + CACHE_TTL_MS,
@@ -343,6 +356,30 @@ export async function resolveQoderModels(credentials, options = {}) {
}
}
/**
* Every model key the chat endpoint accepts for this credential: the IDE-visible
* models first, then catalog entries flagged `enable:false` (hidden in the IDE
* picker, e.g. by an account policy, but still served by agent_chat_generation —
* see fetchQoderCatalogRaw). /v1/models uses this so the advertised list matches
* what the router will actually route instead of collapsing to one or two keys.
*/
export function routableQoderModels(catalog) {
if (!catalog) return [];
const out = [];
const seen = new Set();
for (const m of catalog.models || []) {
if (!m?.id || seen.has(m.id)) continue;
seen.add(m.id);
out.push({ id: m.id, name: m.name || m.id, hidden: false });
}
for (const [key, cfg] of catalog.rawConfigs || []) {
if (!key || seen.has(key)) continue;
seen.add(key);
out.push({ id: key, name: cfg?.display_name || key, hidden: true });
}
return out;
}
export function invalidateQoderCatalog(credentials) {
if (!credentials) return;
catalogCache.delete(cacheKey(credentials));

View File

@@ -10,6 +10,24 @@ const signatureKv = makeKv(SCOPE);
const memorySignatures = new Map();
let pruneCounter = 0;
/**
* Model family that produced / will consume a signature. Antigravity serves Gemini and Claude
* models behind the same API, and each backend only accepts its own signatures: a Claude
* signature replayed to Gemini fails with 400 "Corrupted thought signature." (and vice versa).
*/
export function signatureFamily(model) {
const m = typeof model === "string" ? model.toLowerCase() : "";
if (!m) return null;
if (m.includes("claude")) return "claude";
if (m.includes("gemini")) return "gemini";
return m;
}
// Entries stored before families were recorded (no `family`) stay usable for any model.
function isCompatible(entry, family) {
return !entry.family || !family || entry.family === family;
}
function pruneMemoryExpired() {
const now = Date.now();
for (const [key, value] of memorySignatures.entries()) {
@@ -62,13 +80,15 @@ async function maybePrunePersisted() {
}
/**
* Store a thought signature for a tool_call_id with optional sessionId namespace (RAM + SQLite async)
* Store a thought signature for a tool_call_id with optional sessionId namespace (RAM + SQLite async).
* `model` is the model that produced the signature; lookups for another model family skip it.
*/
export function storeGeminiThoughtSignature(toolCallId, signature, sessionId = null) {
export function storeGeminiThoughtSignature(toolCallId, signature, sessionId = null, model = null) {
if (typeof toolCallId !== "string" || !toolCallId) return;
if (typeof signature !== "string" || !signature) return;
const now = Date.now();
const family = signatureFamily(model);
pruneMemoryExpired();
const keys = [];
@@ -80,12 +100,14 @@ export function storeGeminiThoughtSignature(toolCallId, signature, sessionId = n
for (const k of keys) {
memorySignatures.set(k, {
signature,
family,
expiresAt: now + MEMORY_TTL_MS,
});
// Async persist to SQLite kv table without blocking
signatureKv.set(k, {
signature,
family,
createdAt: now,
expiresAt: now + PERSISTED_TTL_MS,
}).catch(() => {});
@@ -95,23 +117,25 @@ export function storeGeminiThoughtSignature(toolCallId, signature, sessionId = n
}
/**
* Retrieve a thought signature by tool_call_id (RAM first, then SQLite fallback)
* Retrieve a thought signature by tool_call_id (RAM first, then SQLite fallback).
* `model` is the target model; signatures produced by another model family are ignored.
*/
export async function getGeminiThoughtSignature(toolCallId, sessionId = null) {
export async function getGeminiThoughtSignature(toolCallId, sessionId = null, model = null) {
if (typeof toolCallId !== "string" || !toolCallId) return null;
const family = signatureFamily(model);
pruneMemoryExpired();
if (sessionId && typeof sessionId === "string") {
const sessionKey = `${sessionId}:${toolCallId}`;
const sessionEntry = memorySignatures.get(sessionKey);
if (sessionEntry && sessionEntry.expiresAt > Date.now()) {
if (sessionEntry && sessionEntry.expiresAt > Date.now() && isCompatible(sessionEntry, family)) {
return sessionEntry.signature;
}
}
const entry = memorySignatures.get(toolCallId);
if (entry && entry.expiresAt > Date.now()) {
if (entry && entry.expiresAt > Date.now() && isCompatible(entry, family)) {
return entry.signature;
}
@@ -119,9 +143,10 @@ export async function getGeminiThoughtSignature(toolCallId, sessionId = null) {
if (sessionId && typeof sessionId === "string") {
const sessionKey = `${sessionId}:${toolCallId}`;
const sessionRow = await signatureKv.get(sessionKey);
if (sessionRow && typeof sessionRow.signature === "string" && (!sessionRow.expiresAt || sessionRow.expiresAt > Date.now())) {
if (sessionRow && typeof sessionRow.signature === "string" && (!sessionRow.expiresAt || sessionRow.expiresAt > Date.now()) && isCompatible(sessionRow, family)) {
memorySignatures.set(sessionKey, {
signature: sessionRow.signature,
family: sessionRow.family || null,
expiresAt: Date.now() + MEMORY_TTL_MS,
});
return sessionRow.signature;
@@ -134,8 +159,10 @@ export async function getGeminiThoughtSignature(toolCallId, sessionId = null) {
signatureKv.remove(toolCallId).catch(() => {});
return null;
}
if (!isCompatible(row, family)) return null;
memorySignatures.set(toolCallId, {
signature: row.signature,
family: row.family || null,
expiresAt: Date.now() + MEMORY_TTL_MS,
});
return row.signature;
@@ -148,22 +175,24 @@ export async function getGeminiThoughtSignature(toolCallId, sessionId = null) {
}
/**
* Synchronous get from RAM cache only (for sync translators)
* Synchronous get from RAM cache only (for sync translators).
* `model` is the target model; signatures produced by another model family are ignored.
*/
export function getGeminiThoughtSignatureSync(toolCallId, sessionId = null) {
export function getGeminiThoughtSignatureSync(toolCallId, sessionId = null, model = null) {
if (typeof toolCallId !== "string" || !toolCallId) return null;
const family = signatureFamily(model);
pruneMemoryExpired();
if (sessionId && typeof sessionId === "string") {
const sessionKey = `${sessionId}:${toolCallId}`;
const sessionEntry = memorySignatures.get(sessionKey);
if (sessionEntry && sessionEntry.expiresAt > Date.now()) {
if (sessionEntry && sessionEntry.expiresAt > Date.now() && isCompatible(sessionEntry, family)) {
return sessionEntry.signature;
}
}
const entry = memorySignatures.get(toolCallId);
if (entry && entry.expiresAt > Date.now()) {
if (entry && entry.expiresAt > Date.now() && isCompatible(entry, family)) {
return entry.signature;
}
return null;

View File

@@ -148,6 +148,8 @@ const REFRESH_HANDLERS = {
"codebuddy-intl": (c, log) => refreshCodebuddyIntlToken(c.refreshToken, log),
trae: (c, log) => refreshTraeToken(c.refreshToken, c, log),
cline: (c, log) => refreshClineToken(c.refreshToken, log),
// ClinePass shares Cline's WorkOS auth endpoints, so the same refresh works.
clinepass: (c, log) => refreshClineToken(c.refreshToken, log),
zed: () => refreshZedToken(),
windsurf: (c, log) => refreshWindsurfToken(c, log),
// Kimi Code OAuth (merged into id `kimi`); legacy id still routes here

View File

@@ -4,10 +4,10 @@
import { getGitHubUsage } from "./usage/github.js";
import { getGeminiUsage, getAntigravityUsage } from "./usage/google.js";
import { getClaudeUsage } from "./usage/claude.js";
import { getClaudeUsage, consumeClaudeResetGrant } from "./usage/claude.js";
import { getCodexUsage, consumeCodexRateLimitResetCredit, getCodexRateLimitResetCredits } from "./usage/codex.js";
export { consumeCodexRateLimitResetCredit, getCodexRateLimitResetCredits };
export { consumeCodexRateLimitResetCredit, getCodexRateLimitResetCredits, consumeClaudeResetGrant };
import { getKiroUsage } from "./usage/kiro.js";
import { getMiniMaxUsage } from "./usage/minimax.js";
import { getCodeBuddyCnUsage, getCodeBuddyIntlUsage } from "./usage/codebuddy-cn.js";
@@ -15,9 +15,12 @@ import { getXaiUsage } from "./usage/xai.js";
import { getGrokCliUsage } from "./usage/grok-cli.js";
import { getKimiUsage } from "./usage/kimi.js";
import { getDeepseekUsage } from "./usage/deepseek.js";
import { getCommandCodeUsage } from "./usage/commandcode.js";
import { getOpenCodeGoUsage } from "./usage/opencode-go.js";
import { getOpenCodeZenUsage } from "./usage/opencode-zen.js";
import { getGroqUsage } from "./usage/groq.js";
import { getZedUsage } from "./usage/zed.js";
import { getXiaomiMimoUsage } from "./usage/xiaomi-mimo.js";
import { resolveQoderCredentials } from "./qoderModels.js";
import { getGlmUsage } from "./usage/glm.js";
import {
@@ -40,12 +43,8 @@ const USAGE_HANDLERS = {
claude: (c) => getClaudeUsage(c.accessToken, c.proxyOptions, { force: c.force }),
codex: (c) => getCodexUsage(c.accessToken, c.proxyOptions),
kiro: (c) => getKiroUsage(c.accessToken, c.providerSpecificData, c.proxyOptions),
qoder: async (c) => {
// PAT (pt-...) connections must be exchanged to a job token before the
// quota endpoint accepts them.
const resolved = await resolveQoderCredentials(c, c.proxyOptions).catch(() => null);
return getQoderUsage(resolved?.accessToken || c.accessToken, c.proxyOptions);
},
qoder: (c) => getQoderUsageFor(c),
"qoder-cn": (c) => getQoderUsageFor(c),
iflow: (c) => getIflowUsage(c.accessToken),
ollama: (c) => getOllamaUsage(c.apiKey, c.providerSpecificData, c.proxyOptions),
glm: (c) => getGlmUsage(c.apiKey, c.provider, c.proxyOptions),
@@ -59,11 +58,22 @@ const USAGE_HANDLERS = {
"grok-cli": (c) => getGrokCliUsage(c.accessToken, c.providerSpecificData, c.proxyOptions),
kimi: (c) => getKimiUsage(c.accessToken, c.apiKey, c.proxyOptions, c.providerSpecificData),
"opencode-go": (c) => getOpenCodeGoUsage(c.apiKey, c.proxyOptions),
"opencode-zen": (c) => getOpenCodeZenUsage(c.apiKey, c.proxyOptions),
deepseek: (c) => getDeepseekUsage(c.apiKey, c.proxyOptions),
commandcode: (c) => getCommandCodeUsage(c.apiKey, c.proxyOptions),
groq: (c) => getGroqUsage(c.apiKey, c.proxyOptions),
zed: (c) => getZedUsage(c.accessToken, c.providerSpecificData, c.proxyOptions),
"xiaomi-mimo": (c) => getXiaomiMimoUsage(c.accessToken, c.providerSpecificData, c.proxyOptions),
};
// Qoder intl/CN share one usage path: PATs must be exchanged to a job token
// before the quota endpoint accepts them, and the quota URL comes from the
// provider's own registry usage block (region-correct via c.provider).
async function getQoderUsageFor(c) {
const resolved = await resolveQoderCredentials(c, c.proxyOptions).catch(() => null);
return getQoderUsage(resolved?.accessToken || c.accessToken, c.proxyOptions, c.provider || "qoder");
}
export async function getUsageForProvider(connection, proxyOptions = null, options = {}) {
const { provider, accessToken, apiKey, providerSpecificData, projectId } = connection;
const providerDataWithProjectId = {

View File

@@ -0,0 +1,166 @@
/**
* Antigravity weekly quota — best-effort retrieval from retrieveUserQuotaSummary.
* Failure never breaks existing per-model quota display.
*/
import { U, parseResetTime, fetchWithTimeout } from "./shared.js";
import { ANTIGRAVITY_IDE_USER_AGENT, ANTIGRAVITY_IDE_VERSION } from "../../providers/shared.js";
// — Weekly quota summary config ——————————————————————————————
const WEEKLY_CONFIG = {
...U("antigravity"),
userAgent: ANTIGRAVITY_IDE_USER_AGENT,
};
// — Cache: TTL + in-flight dedup per project ———————————————
const WEEKLY_CACHE_TTL_MS = 180_000; // 3 minutes
const weeklyCache = new Map(); // cacheKey -> { result, expiresAt } | { promise }
function cacheKey(accessToken, projectId) {
return `${accessToken}::${projectId || ""}`;
}
// Exported for tests only
export function _clearWeeklyCache() {
weeklyCache.clear();
}
// — Group-name and window to stable key mapping ——————————————————————
const GROUP_CONFIGS = [
{
pattern: /gemini/i,
weekly: { key: "gemini_weekly", displayName: "Gemini (Weekly)" },
session: { key: "gemini_session", displayName: "Gemini (5h)" },
},
{
pattern: /claude|gpt/i,
weekly: { key: "claude_gpt_weekly", displayName: "Claude & GPT (Weekly)" },
session: { key: "claude_gpt_session", displayName: "Claude & GPT (5h)" },
},
];
/**
* Parse a retrieveUserQuotaSummary response into normalized weekly quotas.
* Pure function — safe to unit-test without network.
*
* @param {Object|null} data Raw JSON response
* @returns {Object} e.g. { gemini_weekly: { used, total, ... }, claude_gpt_weekly: { ... } }
*/
export function parseWeeklyQuotaSummary(data) {
if (!data || typeof data !== "object") return {};
// Groups may live at data.groups or data.quotaSummary.groups
const groups = Array.isArray(data.groups)
? data.groups
: Array.isArray(data.quotaSummary?.groups)
? data.quotaSummary.groups
: null;
if (!groups) return {};
const result = {};
for (const group of groups) {
if (!group || typeof group !== "object") continue;
const displayName = group.displayName || "";
const buckets = Array.isArray(group.buckets) ? group.buckets : [];
for (const bucket of buckets) {
if (!bucket || typeof bucket !== "object") continue;
const windowType = String(bucket.window || "").toLowerCase();
const bucketText = `${bucket.bucketId || ""} ${bucket.displayName || ""}`.toLowerCase();
const isWeekly = windowType === "weekly" || bucketText.includes("weekly");
const isSession = windowType === "5h" || bucketText.includes("five hour") || bucketText.includes("5h") || bucketText.includes("daily") || windowType === "daily";
if (!isWeekly && !isSession) continue;
// If a session (5h) bucket is marked disabled by upstream (because weekly was hit),
// keep it so the UI shows the 5h row, but with remainingFraction: 0.
// Disabled weekly buckets are truly disabled and skipped.
if (bucket.disabled === true && isWeekly) continue;
const remainingFraction = bucket.disabled === true ? 0 : Number(bucket.remainingFraction);
if (!Number.isFinite(remainingFraction)) continue;
// Match group to a known family
for (const config of GROUP_CONFIGS) {
if (config.pattern.test(displayName)) {
const target = isWeekly ? config.weekly : config.session;
if (result[target.key]) break; // first matching bucket per type wins
const total = 1000;
const remaining = Math.round(total * remainingFraction);
const used = Math.max(0, total - remaining);
result[target.key] = {
used,
total,
resetAt: parseResetTime(bucket.resetTime),
remainingPercentage: remainingFraction * 100,
unlimited: false,
displayName: target.displayName,
};
break;
}
}
}
}
return result;
}
/**
* Fetch weekly quota summary — cached, deduped, never throws.
*/
export async function fetchAntigravityWeeklyQuota(accessToken, projectId, proxyOptions = null) {
const key = cacheKey(accessToken, projectId);
// Serve in-flight or cached
const hit = weeklyCache.get(key);
if (hit?.promise) return hit.promise;
if (hit && hit.expiresAt > Date.now()) return hit.result;
const promise = (async () => {
try {
const url = WEEKLY_CONFIG.quotaSummaryApiUrl;
if (!url) return {};
const response = await fetchWithTimeout(url, {
method: "POST",
headers: {
"Authorization": `Bearer ${accessToken}`,
"User-Agent": WEEKLY_CONFIG.userAgent,
"Content-Type": "application/json",
"X-Client-Name": "antigravity",
"X-Client-Version": ANTIGRAVITY_IDE_VERSION,
},
body: JSON.stringify({
...(projectId ? { project: projectId } : {}),
}),
}, 10000, proxyOptions);
if (!response.ok) return {};
const data = await response.json();
return parseWeeklyQuotaSummary(data);
} catch {
return {};
}
})();
weeklyCache.set(key, { promise });
try {
const result = await promise;
if (result && Object.keys(result).length > 0) {
weeklyCache.set(key, { result, expiresAt: Date.now() + WEEKLY_CACHE_TTL_MS });
} else {
weeklyCache.delete(key);
}
return result;
} catch {
weeklyCache.delete(key);
return {};
}
}

Some files were not shown because too many files have changed in this diff Show More