257 Commits

Author SHA1 Message Date
39101b3416 Merge origin/master (v0.5.91) into gitea/new_feature
Resolve conflicts:
- streamingHandler.js: adopt upstreamResponseHeaders while keeping 0-token detail row avoidance
- capabilities.js: preserve user-asserted caps and globalThis slots without local caching of catalogSource
- AddCustomModelModal.js & providers/[id]/page.js: wire STT transport marker with custom model edits/assertions
- models/custom/route.js & aliasRepo.js: persist custom model transport and invalidate user caps
- usageRepo.js: key byApiKey live stats by full API key and keep tail in maskApiKey
- UsageStats.js: lazy load charts dynamically
2026-09-28 21:17:33 +07:00
decolua
f01fb909e3 # v0.5.91 (2026-09-26)
## Features
- **Providers**: add Token Harbor provider and four OpenAI-compatible aggregator providers (dahl, atria, agnes, bai)
- **Claude**: forward `x-claude-code-session-id` on OAuth requests; merge client `anthropic-beta` flags and forward rate-limit headers; return thinking text to OpenAI-format clients
- **Codex**: add GPT-6 Sol and Luna support
- **CLI Tools**: support multiple model profiles for Codex CLI
- **Hermes**: multi-role model config (delegation + auxiliary slots)
- **OpenCode Go**: complete the Go catalog (40 models) with auto-fetch + family endpoint regex
- **Usage**: show and redeem free limit resets for cc accounts
- **Cline**: expose the `cline-free/*` tier and price it at zero
- **Combos**: display vision adapter models in an ordered table view

## Fixes
- **Claude**: decloak tool names when `toolNameMap` misses (#4342); update spoofed cli version to 2.1.280 to support Opus 5.5
- **Providers API**: make POST `/api/providers` O(1) and refuse silent key overwrite (#4350)
- **Capabilities**: stop caching the catalog source per module copy (#4351)
- **OAuth**: stop Zed paste-token crash and add IDE auto-import (#4359)
- **Dashboard**: resolve combo limits with the server's capabilities (#4360); lazy-load charts and `marked`, preload in background on idle
- **Responses**: carry the streamed output items in `response.completed` (#4307)
- **STT**: dispatch live-API-only Gemini models over the Live WebSocket transport (#4006)
- **Gemini**: guard terminal model turns and unresponded functionCalls in `normalizeGeminiContents`
- **Command Code**: replay raw byte chunks to preserve all NDJSON lines
- **Translator**: stop emitting empty `<think>` markers into OpenAI content
- **CLI Tools**: refresh Codex settings after apply (#4347); keep existing `ANTHROPIC_AUTH_TOKEN` when applying Claude settings
- **Tray**: native arm64 macOS menubar binary, no Rosetta required
- **CLI**: filter model selector by active connections and noAuth providers
- **Usage**: key live byApiKey stats by full api key to prevent team-key collision and preserve API key usage attribution
- **Tailscale**: cap enable-flow health wait at 20s
2026-09-26 17:40:15 +07:00
decolua
c4690307ce fix(cli-tools): keep existing ANTHROPIC_AUTH_TOKEN when applying Claude settings
Only write the token when settings.json has none, so a real API key or
earlier config is never clobbered by Apply. Reset still clears it,
re-enabling key selection on the next Apply.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-26 17:30:59 +07:00
decolua
b54a3f9bb5 test(claude): update decloak tests for suffix-stripping fallback 2026-09-26 17:30:02 +07:00
Gesya Gayatree Solih
b65d2d0a6a fix(claude): decloak tool names when toolNameMap misses (#4342)
- Add stripCloakSuffix fallback in decloakToolNames and decloakStreamChunk
- Prevent client errors when toolNameMap is missing or lost on retry
2026-09-26 17:28:51 +07:00
decolua
6aea3875ef feat(claude): forward x-claude-code-session-id on OAuth requests 2026-09-26 17:15:12 +07:00
decolua
0249464d74 feat(cli-tools): support multiple model profiles for Codex CLI
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-26 17:03:06 +07:00
Nikan Wystaf
239bcfc568 perf(providers): make POST /api/providers O(1) and refuse silent key overwrite (#4350)
Fixes #4311

- Drop full-pool renumber on insert: new row gets MAX(priority)+1 directly,
  turning an O(pool) rewrite into O(1) per insert.
- For apikey connections, query by (provider, authType, name) and count via SQL
  aggregate instead of reading the entire pool into memory.
- Refuse silent apikey overwrite on name collision with 409 PROVIDER_NAME_CONFLICT,
  unless caller explicitly sets allowOverwrite: true.
- Add 8 unit tests covering priority ordering and name collisions.
2026-09-26 16:37:25 +07:00
Nick Nyanjui
199173fe58 feat(cline): expose the cline-free/* tier and price it at zero (#4334) 2026-09-26 16:31:25 +07:00
Beru
8f20daac6b fix(cli-tools): refresh Codex settings after apply (#4347) 2026-09-26 16:21:21 +07:00
Mohammad Hijjawi
fdcba3e1b2 fix(capabilities): stop caching the catalog source per module copy (#4351) 2026-09-26 16:20:59 +07:00
agent
dd293d3c62 feat(opencode-go): add the seven models upstream serves but the registry omits (#4357)
Upstream /zen/go/v1/models serves 35 ids; the registry listed 28. Add the seven
missing models with their corresponding supported and target formats:
- deepseek-v4.1-flash
- mimo-v2.6-flash, mimo-v2.6-pro
- space-bunny-free
- omen-alpha
- grok-4.7 (responses-only)
- gpt-6-luna (responses-only)
2026-09-26 16:15:43 +07:00
Amir Seify
7a436d209c fix(oauth): stop Zed paste-token crash and add IDE auto-import (#4359)
- Fix Zed Connect crash by gating paste-token UI when provider has no paste-token config
- Add ZedAuthModal with local IDE keyring auto-import, browser OAuth, and manual callback paste
- Add localhost-only GET /api/oauth/zed/auto-import and POST /api/oauth/zed/import routes
2026-09-26 16:12:46 +07:00
Nikan Wystaf
737b1f4d0a feat(providers): add Token Harbor provider 2026-09-26 16:09:52 +07:00
Spoon94
37a6b7e0f2 fix(dashboard): resolve combo limits with the server's capabilities (#4360)
The combos page computed combo capabilities in the browser using
aggregateComboCapabilities, falling back to pattern defaults for models
without exact entries because the synced model catalog is server-only.

Allow callers to supply a resolveCaps callback (e.g. from useModelCaps)
merged over the local tables, preserving non-limit capability flags.
2026-09-26 16:07:03 +07:00
Nikan Wystaf
06112c136c feat(providers): add four OpenAI-compatible aggregator providers (dahl, atria, agnes, bai)
Adds Dahl Inference, Atria Dawn, Agnes AI, and B.AI following the existing
registry pattern. Each uses DefaultExecutor with no custom translator needed.
2026-09-26 16:04:59 +07:00
MrBeanDev
90b0693423 feat(thinking): return Claude thinking text to OpenAI-format clients 2026-09-26 12:02:07 +07:00
Nick Nyanjui
fe347e4ea5 fix(stt): dispatch live-API-only Gemini models over the Live WebSocket transport (#4006)
Addresses #4006 by letting Gemini STT models that only exist on the Live
API transcribe instead of failing.

transcribeGemini sends audio to :generateContent and that is the only Gemini
path. A model that is realtime-only (exposed by the Live API's
bidiGenerateContent WebSocket) therefore fails outright, even though the account
can transcribe it.

open-sse/handlers/geminiLiveStt.js owns the WebSocket lifecycle: opens
:bidiGenerateContent, sends setup frame, waits for setupComplete, streams
audio as realtimeInput media chunks, and settles on turnComplete.
Dispatch is driven by transport marker 'gemini-live'. Adds custom model transport
persistence and selection on the dashboard.
2026-09-26 12:00:39 +07:00
ANIRUDDHA ADAK
273f0c32cd fix(responses): carry the streamed output items in response.completed (#4307) 2026-09-26 12:00:00 +07:00
Christian Gennari
c2148179c0 fix(commandcode): replay raw byte chunks to preserve all NDJSON lines 2026-09-26 11:53:54 +07:00
MrBeanDev
5d2cfbf3c5 fix(translator): stop emitting empty <think> markers into OpenAI content 2026-09-26 11:45:35 +07:00
decolua
3e4323e2fd feat(combos): display vision adapter models in an ordered table view
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-26 11:43:22 +07:00
decolua
975f28a57c test(zed): isolate zed-live-models suite DB via temp DATA_DIR
Seeding via createProviderConnection wrote zed-live-* accounts into the
real ~/.9router DB, showing up as junk accounts in the running dashboard.
Set DATA_DIR to a mkdtemp dir before dynamic-importing the DB-backed
modules and clean it up afterwards, matching the pattern in
compatible-provider-connections.test.js.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-26 11:40:08 +07:00
decolua
ddfdb0df04 perf(dashboard): lazy-load charts and marked, preload in background on idle
- Drop UsageStats from shared barrel so recharts (~589KB) stays out of
  the initial bundle of every page; usage page imports it directly
- Convert UsageChart/ProviderBarChart/TopModelsChart to next/dynamic
- Preload chart chunks via requestIdleCallback in DashboardLayout so
  navigating to Usage is still instant
- Lazy-import marked inside ChangelogModal (only when opened)
- Route Sidebar/EndpointPageClient through settingsStore: coalesce
  concurrent in-flight GET, merge PATCH response over cache (PATCH
  omits GET-only hasPassword) to kill duplicate /api/settings calls
- Delay /api/version npm check 2.5s after first render

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-26 11:34:08 +07:00
akmal safari pellu
30464bc227 fix(gemini): guard terminal model turns and unresponded functionCalls in normalizeGeminiContents 2026-09-26 11:18:44 +07:00
Doan Anh Dung
249f6c2fb8 fix(tray): native arm64 macOS menubar binary, no Rosetta required
systray2 ships only an x86_64 tray_darwin_release and selects it by
process.platform with no process.arch branch, so there is no native slice to
choose. Apple Silicon users therefore need Rosetta 2, and without it the tray
dies with EBADARCH ("bad CPU type in executable") and no icon appears.

Overlay a native arm64 build of the same upstream source
(felixhao28/systray-portable @ 6eddc91) instead. On darwin/arm64,
ensureArm64TrayBin() detects the Intel binary by parsing the Mach-O cputype,
downloads the artifact from the pinned tray-binaries release, verifies it
against a sha256 constant, atomically renames it over systray2's binary, and
busts systray2's copyDir cache — that cached copy is what actually executes, so
without the bust the swap has no effect.

Intel Macs keep using systray2's binary unchanged and Windows is unaffected
(PowerShell NotifyIcon, no binary). Any download or checksum failure leaves the
Intel binary in place and tells the user how to install Rosetta; a marker file
throttles retries to once per 24h because ensureTrayRuntime runs synchronously
on every CLI start, and is cleared on success so a clobbered binary recovers
immediately. Binaries stay out of the npm tarball per the existing Kaspersky
false-positive constraint — the artifact is fetched on demand.

Adds cli/scripts/buildTrayArm64.js (-trimpath, bit-for-bit reproducible for a
given Go version and macOS SDK) and a workflow_dispatch action that builds on a
macos-15 runner and refuses to publish when the sha diverges from the pin.

Also corrects comments claiming the systray -> systray2 switch fixed Apple
Silicon; it only fixed the dyld header rejection on macOS 14+, the binary was
still amd64-only.
2026-09-26 11:16:38 +07:00
decolua
f462837536 fix(cli): filter model selector by active connections and noAuth providers
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-26 11:14:04 +07:00
DavidArthurCole
dc198dff1f feat(claude): merge client anthropic-beta flags and forward rate-limit headers 2026-09-26 11:08:57 +07:00
29ccb84faf feat(commandcode): align quota usage with the official CLI /usage flow
Rewrites services/usage/commandcode.js to mirror the command-code CLI: whoami
resolves the org id, credits+subscriptions report the 5-hour/weekly windows and
plan, then usage/summary is queried with since=currentPeriodStart. Endpoints are
read from the provider registry usage block instead of hardcoded constants.

Tests updated to cover the new request order and response shapes.
2026-09-25 16:09:09 +07:00
3481058e09 fix(usage): drop duplicated commandcode import and features block
The previous origin/master merges (11089ab1 and earlier) left two
auto-merged artifacts behind because both sides added the same lines in
different places:

- services/usage.js declared `getCommandCodeUsage` twice (import + handler),
  which is an ESM SyntaxError on a clean checkout. The working tree masked
  it, so `next build` passed locally while the committed tree did not parse.
- providers/registry/commandcode.js carried two `features` blocks; the
  second one shadowed the first for `usage`/`usageApikey`.

Verified with `node --check` on the committed blob.
2026-09-25 16:09:04 +07:00
272dbcb9cc Merge remote-tracking branch 'origin/master' into gitea/new_feature
# Conflicts:
#	open-sse/executors/qoder.js
#	open-sse/handlers/chatCore.js
#	open-sse/handlers/chatCore/sseToJsonHandler.js
#	open-sse/providers/registry/commandcode.js
#	src/app/(dashboard)/dashboard/combos/page.js
#	src/app/api/v1/models/route.js
#	src/lib/db/repos/usageRepo.js
#	src/shared/components/UsageStats.js
2026-09-25 15:14:28 +07:00
Rafli Ahmad Zulfikar
4a57df8bf9 fix(usage): key live byApiKey stats by full api key to prevent team-key collision
maskApiKey kept only the first 8 characters of the key. API keys minted
from the same machine id share that prefix, so every key of an instance
collapsed into one sk-XXXXXXX*** bucket per model/provider, and the
dashboard attributed one holder's usage to another.

Key the live path by the full api key - matching the daily rollup
(aggregateEntryToDay) and the lastUsed overlay - and keep the last 4
characters in maskApiKey so masked keys stay distinguishable in the UI.
2026-09-23 14:53:01 +07:00
Deepanshu
a406381fad fix(usage): preserve API key usage attribution in live stats
The 24h/today branch of getUsageStats keyed byApiKey buckets on the
masked key, while the daily rollup and the lastUsed overlay key on the
full key. Every key minted by one instance shares the sk-{machineId}
prefix, so the mask collapsed all of them into a single bucket and
attributed one key's usage to another.

Key the live branch by the full api key too; the masked value is still
carried on the bucket for display.
2026-09-23 14:42:56 +07:00
Reid Nguyen
e571a8b6da feat(usage): show and redeem free limit resets for cc accounts
Adds the free usage-limit reset (the desktop app's "Reset for free")
to the Quota Tracker for cc OAuth accounts, mirroring the existing
Codex reset-credit button.

- Reset button with remaining count on the card; tooltip shows use-by
  date and which limits get refilled.
- Expiry modal (clock button) listing each grant: label, resets left,
  refills (session / weekly), use-by date, time remaining, status.
- Confirm dialog before redeeming, since a reset is irreversible.
- Usage and reset calls send the CLI User-Agent required for cedar_ember.
- Usage cache is dropped after a redeem so the card shows the refilled limits.
2026-09-23 12:21:07 +07:00
Ilhom (MBP M5 Pro)
8403576095 fix(claude): update spoofed cli version to 2.1.280 to support Opus 5.5
Anthropic's newly released Claude Opus 5.5 model strictly requires
claude-cli version 2.1.280 or newer. Spoofing the older 2.1.258
version results in an HTTP 400 `claude_code_version_too_old` error.

This commit updates the hardcoded `CLAUDE_CLI_VERSION` in
`open-sse/providers/shared.js` and aligns the corresponding unit
tests and baselines to bypass Anthropic's version gating.
2026-09-23 12:18:39 +07:00
decolua
e7c269b8c9 fix(tailscale): cap enable-flow health wait at 20s
Enable waited out the full 180s HEALTH_CHECK.timeoutMs when funnel URL
DNS was unreachable, making the toggle hang ~3 minutes before returning
success. waitForHealth now takes an optional timeoutMs; enable uses
HEALTH_CHECK.enableTimeoutMs (20s) while watchdog/cloudflare keep 180s.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-23 12:12:10 +07:00
mxlanparty@dp14
95600db17b feat(codex): add GPT-6 Sol and Luna support
- Add GPT-6 Sol and Luna to the Codex model registry.
- Send both models using the Codex 0.155 Responses Lite request shape, including `reasoning.context: "all_turns"`.
2026-09-23 12:11:49 +07:00
decolua
69cc1aa987 feat(hermes): multi-role model config (delegation + auxiliary slots)
POST /api/cli-tools/hermes-settings now accepts selections:
[{role, model}] and writes model/delegation/auxiliary.<role> blocks
into ~/.hermes/config.yaml non-destructively (comments and other keys
preserved). GET returns the current delegation/auxiliary config,
DELETE removes all 9router-managed blocks. Dashboard card gains a
collapsible Model Roles section aligned with the existing rows; the
bare `model` payload from the CLI quick setup still works.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-23 11:46:21 +07:00
decolua
001b292876 fix(dashboard): replace Hermes icon with official Nous Research logo
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-23 11:46:13 +07:00
decolua
a61fc6a01e fix(antigravity): rewrite all Hermes identity variants, not just the legacy sentence
The old rule only matched 'You are Hermes Agent, an intelligent AI
assistant created by Nous Research.' Current Hermes versions use
'You are Hermes Agent, built by Nous Research.' and similar variants,
which slipped through and got a fake 429 RESOURCE_EXHAUSTED. Replace
every variant with a neutral 'You are an AI assistant.' identity.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-23 11:46:05 +07:00
decolua
1b72f02e3b fix(opencode-go): clamp deepseek reasoning_effort "max" to "high" for mimo backends that reject it
mimo-v2.5-pro/v2.6 on the Go lane return 400 on reasoning_effort "max"
(probed live; mimo-v2.5 accepts it). The deepseek applyFormat case now honors
the declared thinking levels, and mimo-v2.5-pro gets a levels entry.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-23 11:45:40 +07:00
decolua
2aa99d9f39 feat(opencode-go): complete Go catalog (40 models) with auto-fetch + family endpoint regex
- Add 12 missing models from live /zen/go/v1/models (glm-5, kimi-k2.5,
  mimo v2/v2.6 line, qwen3.5-plus, grok-4.5/4.7, omen-alpha, ...)
- modelsFetcher + passthroughModels + "opencode-go" suggested-models filter
- Family regex fallback (opencodeFamilyFormats) keeps unknown/passthrough ids
  on the right endpoint lane (/responses, /messages, /chat/completions);
  curated registry entries always win
- isResponsesModel now routes passthrough grok/gpt ids to /responses

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-23 11:45:34 +07:00
decolua
39e36d3d0c # v0.5.86 (2026-09-23)
## Features
- **Xiaomi MiMo**: server-assisted desktop login for headless/Docker deployments, five account clusters (cn/sgp/ams/ru/in), and v2.6 pro/flash/pro-ultraspeed models with dual-route (account service vs. cloud API)
- **Claude**: add Claude Opus 5.5 support
- **i18n**: translate React text rewrites via characterData mutation observer

## Fixes
- **Proxy Pools**: keep request headers intact through Vercel/Cloudflare/Deno relays (spreading a `Headers` instance yielded `{}`, dropping auth and content-type)
- **Xiaomi MiMo login**: keep the session in the httpOnly cookie only, require dashboard auth on the proxy branch, and stop forwarding authorization headers upstream
2026-09-23 10:05:33 +07:00
wismyzhizi
910db749aa feat(xiaomi-mimo): server-assisted desktop login, five account clusters, v2.6 models
Reproduce the MiMo Desktop login surface server-side so headless/Docker
deployments can link a Xiaomi account without the Desktop client. The
account session (passToken) is captured during the proxied login and
stored per connection.

- Five account clusters (cn/sgp/ams/ru/in): per-region mimo-server host
  and SSO sid, unknown region falls back to sgp
- mimo-v2.6-pro/flash/pro-ultraspeed dual-route models: account-service
  route when desktop credentials exist, cloud API (sk- key) otherwise;
  drops obsolete mimo-x-*-preview ids
- Desktop ServiceTokenManager 2-phase handshake (single serviceLogin with
  target sid, raw 64-bit nonce preserved), per-region session cache
- reasoning_effort bridged to output_config.effort; i18n runtime now
  observes characterData mutations so React text rewrites get translated
- Security hardening on the login proxy: session travels only in the
  httpOnly cookie (never in the URL), proxy branch requires dashboard
  auth, authorization/proxy-authorization never forwarded upstream, and
  upstream Set-Cookie is not replayed onto the app origin
2026-09-23 09:43:54 +07:00
savioruz
6af26a9ee8 fix(proxy-pools): lossless header forwarding for relay pools
Preserve request headers through the Vercel/Cloudflare/Deno relay.

- proxyFetch: normalize options.headers via Object.fromEntries before
  spreading into the relay headers. Spreading a Headers instance yields
  {} and silently dropped every entry (auth + content-type), which is why
  the same pool worked on one path and failed on another.
- vercel relay: build the forwarded header object from req.headers.entries()
  instead of new Headers(req.headers), avoiding edge-runtime normalization
  of casing/duplicate keys that some providers reject.
2026-09-23 09:12:00 +07:00
Reid Nguyen
cbffeb9770 feat(claude): support Claude Opus 5.5 2026-09-23 09:01:51 +07:00
decolua
21583c03e5 # v0.5.85 (2026-09-22)
## Features
- **System One**: add `/v1/systemone` decision endpoint for Jev models (OpenCode Zen and OpenRouter lanes), wire into sidebar and Media Providers page with interactive probe testing
- **CLI Tools**: add dynamic configuration, settings APIs, and official logos for Pi, OMP, Crush, ForgeCode, Smelt, and CodeWhale
- **Analytics & Usage**: add Requests mode, provider/model breakdown charts, All Time period filter, and refined overview cards
- **Combos**: add Cursor/Claude Default presets; support bulk select/delete and bulk strategy changes (Fallback / Round Robin / Fusion)
- **Model Capabilities**: expose model capability metadata on `/v1/models` and aggregate capabilities across combo targets
- **OpenCode Zen & MiMo**: add OpenCode Zen (`opencode-zen`) provider with free-tier fingerprint; switch default vision fallback to MiMo V2.6 Flash Free
- **Qoder CN**: add `qoder-cn` provider for qoder.com.cn with OAuth flow, COSY protocol, and CN gateway routing

## Fixes
- **Translator**: map Claude `refusal` stop_reason to `content_filter` and surface explanation; strip replayed reasoning fields for Groq, Mistral, and Cerebras (#4220)
- **Antigravity**: drop requestType `agent` to avoid false 429 `RESOURCE_EXHAUSTED`; separate weekly and short-window (5-hour) quotas and deduplicate dashboard rows
- **Responses API**: report usage on `response.completed` so clients can auto-compact (#3432)
- **Hugging Face**: migrate to Inference Providers router (`router.huggingface.co`), expand image models catalog, and add STT route
- **Qoder**: prevent signed request replay (`403/103 Duplicate request`), handle code 110 billing blocks, and preserve upstream SSE error status
- **Performance**: bound usage `lastUsed` scan to a 2-day window; map large budget tokens to `max` reasoning tier
- **Docker**: publish verified multi-platform images (linux/amd64 and linux/arm64) with configurable apk build mirrors
2026-09-22 15:43:30 +07:00
Welington
0f488c7027 fix(translator): map Claude "refusal" stop_reason to content_filter and surface its explanation
Anthropic's API-level refusal (streaming classifier / ToS) ends the stream
with stop_reason "refusal", stop_details carrying the reason, zero output
tokens and no content blocks. Map refusal to content_filter in both
directions, surface stop_details.explanation as message text, and add
CLAUDE_STOP.REFUSAL to schema.
2026-09-22 15:32:48 +07:00
decolua
1a02713150 feat(combos): hide preset buttons and migrate legacy mimo vision adapter
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 15:32:10 +07:00
decolua
b53260ca54 feat(usage): add All Time period option and refine overview cards
- Add "all" period option to usage dashboard and chart API
- Aggregate all days in getChartData when period is "all"
- Center overview card metrics and adjust font size to prevent truncation

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 15:31:24 +07:00
Rafli Ahmad Zulfikar
d1de324586 perf(usage): bound lastUsed overlay scan to 2-day window; reach max thinking tier
getUsageStats("all") shipped the entire usageHistory table to JS just to
refine lastUsed (~2s on 290K rows, on every statsEmitter update per SSE
listener). Bound the overlay to a 2-day indexed range scan; older entries
keep day-level lastUsed from usageDaily aggregates. Totals unaffected.

budgetToLevel now maps budgets > 80384 (midpoint of 32768/128000) to
"max" instead of clamping to "xhigh", so the top reasoning tier is
reachable from large budget_tokens requests.
2026-09-22 15:08:38 +07:00
decolua
ce9ac43da5 style(sidebar): match NEW badge style across 9Remote, Media Providers, and System One
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 15:08:09 +07:00
BFLabsAI
5798b30841 fix(antigravity): drop requestType "agent" to avoid false 429 RESOURCE_EXHAUSTED 2026-09-22 15:04:16 +07:00
yiwen65
782c137b1f fix(qoder): prevent signed request replay and surface upstream errors 2026-09-22 15:01:06 +07:00
Ghost Machine
0d50fe3610 feat(analytics): add Requests mode and provider/model breakdown charts 2026-09-22 14:55:31 +07:00
dajiaohuang
3279d31024 docs(cli): fix source README link 2026-09-22 14:55:02 +07:00
rspersahabatan
aa2bc53f3a docs(readme): add temporary swap instructions for low-memory production builds 2026-09-22 14:54:12 +07:00
cokvrindaa
c73bb2cd66 docs(readme): add Indonesian video guide by Neptiver 2026-09-22 14:53:46 +07:00
decolua
6c9fe6f78a feat(cli-tools): add dynamic configuration for Pi, OMP, Crush, ForgeCode, Smelt and CodeWhale
- Add dedicated settings API routes for pi, omp, crush, forge, smelt, codewhale
- Integrate GenericCliToolCard with multi-model support for Pi and auto-discovery for OMP
- Register tools in cliTools catalog and all-statuses route
- Add official logos for all new CLI tools

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 14:48:04 +07:00
decolua
44fd69d229 style(dashboard): match 9remote NEW badge style on System One tags
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 14:46:57 +07:00
decolua
f84c667d42 feat(providers): add OpenRouter System One lane and New badge
OpenRouter serves TypeSafe Jev at POST /api/v1/systemone with the same
request/response shape, so it plugs into systemoneConfig directly with
model typesafe/jev-1.13. Mark the System One media kind isNew and
render a New badge on the sidebar kind item and the Media Providers
accordion when any visible kind is new.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 14:42:02 +07:00
decolua
20014b3104 fix(dashboard): probe System One models through /v1/systemone
The inline model test sent a chat-completions payload and failed with
500 on decision models. Add a systemone branch to pingModelByKind that
submits a native state+questions probe, and restore the test button
that was hidden for System One models.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 14:41:59 +07:00
decolua
f28e918e24 feat(dashboard): add question input for System One and hide inline test
Allow customizing evaluation instructions in the System One example
card, and disable the inline model probe button for System One models
since decision models do not accept chat completion probes.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 14:21:15 +07:00
decolua
6431e35303 feat(dashboard): add System One to sidebar and media provider detail page
Expose System One in the Media Providers sidebar accordion, set
kind to "systemone" on Jev models for ModelsCard filtering, wire
systemoneConfig into ProviderInfoCard, and configure GenericExampleCard
for interactive testing of decision models.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 14:07:36 +07:00
11089ab137 Merge origin/master (v0.5.81) into gitea/new_feature
Resolve conflicts:
- streamingHandler.js: merge buildStreamErrorBytes onAbortTerminal + local shouldPersistRequestDetail & streamStatusForContent
- capabilities.js: preserve server-injected user-asserted caps and models.dev catalog lookup; wire CommandCode /alpha/generate caps inside resolve()
- commandcode.js (services/usage): adopt upstream whoami + billing credits/subscriptions with 5h/weekly rate windows and plan caps
- openai-to-commandcode.js: merge toNativeImageBlock (data-URI & http(s) support) and assistant reasoning_content preservation
- commandcode-to-openai.js: adopt upstream mid-stream error throw for clean retry and abortion
- tests: sync commandcode test suite and exclude .next from vitest config
2026-09-22 10:08:58 +07:00
decolua
41a1b8003d feat(providers): add MiMo V2.6 Flash Free to OpenCode Zen free tier
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 09:51:51 +07:00
decolua
b7446f8dd1 feat(providers): add System One (Jev) decision endpoint
New /v1/systemone pass-through route for Jev decision models (jev-1.13,
jev-1.13-free) on OpenCode Zen and the free lane. Follows the media-route
pattern: systemoneConfig in the registry drives URL/headers, the handler
mirrors the embeddings account-fallback + usage flow, and the dashboard
gains a System One media-provider kind. No chat-pipeline changes.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 09:51:29 +07:00
decolua
6886915f62 feat(capacity-adapter): default vision fallback to mimo-v2.6-flash-free
Register mimo-v2.6-flash-free on opencode-zen (chat lane) with a v2.6
capability pattern, and switch the vision adapter default from the old
mimo-v2.5-free.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 09:45:20 +07:00
82030615
da0046550a fix(responses): report usage on response.completed so clients can auto-compact
Map upstream Chat Completions usage to the Responses API shape and attach it to response.completed. Capture chunk.usage before the empty-choices guard so the usage-only trailer chunk survives, and defer completion to flushEvents() when usage is not yet known — only on the direct openai:openai-responses route, since a pivoted stream never reaches flushEvents. Fixes #3432.
2026-09-21 21:27:01 +07:00
Liang.Xu
402745dc1f feat(providers): add qoder-cn support for Qoder CN (qoder.com.cn) 2026-09-21 20:45:19 +07:00
DaDecky
f67d5a0c93 fix(docker): publish verified multi-platform images
- Build linux/amd64 and linux/arm64 on native GitHub runners
- Assemble version manifests from platform digests and promote latest only after verification
- Add release/tag validation, manual republishing, timeouts, and health smoke tests
- Make Docker build mirrors configurable via build args and remove unnecessary runtime apk upgrades
- Update DOCKER.md documentation
2026-09-21 20:28:45 +07:00
82030615
c7df895bbb build(docker): apply the apk mirror swap to the runner stage
- Add apk mirror swap to mirrors.aliyun.com at the top of runner stage
- Prevents docker build hanging when dl-cdn.alpinelinux.org is unreachable
2026-09-21 20:22:19 +07:00
Christian Gennari
be3bc764b1 fix(antigravity): separate weekly and short-window quotas and clean up redundant rows
- Track both weekly and 5-hour session buckets in parseWeeklyQuotaSummary,
  distinguishing sliding-window limits from multi-day weekly limits
- Preserve disabled session buckets at 0% rather than dropping them when weekly limits are reached
- Target 5-hour session rows (not weekly rows) during family exhaustion reconciliation in getAntigravityUsage
- Suppress synthesized per-model duplicate rows in dashboard normalization when family summaries are present
- Add unit test coverage for multi-bucket extraction, reconciliation isolation, and dashboard deduplication
2026-09-21 20:15:56 +07:00
dinhkarate
2daf25ffbe fix(qoder): handle code 110 billing blocks and preserve SSE error status
- Match code 110 (billing daily count exceeded) alongside 112/10605/pricingUrl
  in isBillingBlock, parsing JSON safely and accepting numeric/string codes
- Accept numeric strings for statusCodeValue and object bodies in envelope peek
- Emit structured 403 quota error chunk instead of synthetic assistant text
  when a billing envelope appears mid-stream
- Preserve upstream HTTP status in handleForcedSSEToJson when error chunk carries
  a valid 400-599 status
- Add unit tests for code-110 detection, mid-stream billing envelopes, and false-positive guard
2026-09-21 20:09:03 +07:00
Anantachoke
7c2b1fe3e1 fix(translator): drop replayed reasoning fields for Groq/Mistral/Cerebras (#4220)
Strict OpenAI-compatible validators reject unknown assistant-message
fields: Groq 400 ("property 'reasoning_content' is unsupported"),
Mistral 422 ("extra_forbidden"), Cerebras 400 ("wrong_api_format").

Clients driving reasoning models (Hermes Agent, and anything following
the DeepSeek/Kimi convention) echo the previous turn's reasoning_content
on every assistant message, so from the second turn on every request to
these providers fails and a fallback combo silently skips them.

Add a dropMessageFields rule to paramSupport.js that strips
reasoning_content / reasoning / reasoning_details from assistant turns
for groq, mistral, and cerebras.
2026-09-21 20:02:21 +07:00
Amir Seify
253199f16f feat(combos): Cursor/Claude Default presets + bulk select/delete/strategy
- Add Cursor Default / Claude Default on Dashboard -> Combos to generate
  unprefixed combo names that match Cursor/Claude client model IDs,
  seeded with cu/... or cc/... so those clients can route through 9Router.
- Add multi-select bulk Delete and bulk Set strategy (Fallback / Round Robin / Fusion).
- Docs and unit tests for preset builder.
2026-09-21 19:51:26 +07:00
Amir Seify
c933eefc27 fix(cursor): stop AgentService empty turns and silent tool hangs
Cursor-hosted models (cu/composer-2.5, cu/cursor-grok-*, cu/default) returned
HTTP 200 with an empty turn, or hung, whenever a client sent tools.

- Fold system prompts into the current user message. custom_system_prompt
  (RunRequest field 8) makes AgentService return an empty turn.
- Send ModelDetails (field 3); thinking variants (Composer, Grok, *-thinking)
  return an empty turn when only requested_model (field 9) is set.
- Route tool-call history and declared tool schemas through AgentService:
  encode OpenAI tools into mcp_tools (field 4), decode McpArgs and emit real
  tool_calls with finish_reason tool_calls.
- Map Composer  thinking / Grok thinking_delta (field 4) into visible content
  instead of dropping the answer with the unsigned reasoning.
- Ack request_context without echoing MCP tools (double-advertise stalls the
  HTTP/2 stream) and ack kv_server_message so the run proceeds.
- Reject IDE builtin execs instead of failing the turn, so the model can
  continue with MCP tools or a text answer.
- Add google.protobuf.Value / MCP encoders and a FIXED64 branch to
  encodeField in cursorProtobuf.js.

RTK now compresses the source-format body before translation for cursor only:
its translator rewrites role:tool into user XML, so the post-translate pass
missed those tool results. Every other provider keeps the post-translate pass
unchanged.
2026-09-21 19:49:48 +07:00
Minh Ha
5c217d34f3 feat(capabilities): model capability metadata on /v1/models, combo aggregation, pattern fixes
- Export aggregateComboCapabilities: union for vision/audio/search/pdf,
  intersection for tools, primary-model for reasoning fields, min
  contextWindow, max maxOutput
- Support nested combo resolution in aggregateComboCapabilities via
  comboLookup with depth guard (max 6)
- Wire capability metadata to all /v1/models entries and combos
- Show aggregated ctx/max metadata line and capability badges on combo chips
- Pattern fixes: MiMo v2.5/omni reasoning, qwen max/plus vision, minimax m2.x vision
- Sync commandcode model catalog and add openai gpt-5.5
- Add unit tests for capability patterns and combo capability aggregation
2026-09-21 19:41:29 +07:00
chisewaguri
477b2aed0b fix(opencode-go): send reasoning_effort for glm-5.3-flash 2026-09-21 19:10:14 +07:00
decolua
9f42e7ac1c fix(sidebar): restore NEW badge for 9Remote menu
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-21 18:56:56 +07:00
kimono381
cf663f5300 fix(huggingface): complete the Inference Providers router migration
Replace the retired api-inference.huggingface.co host with the Inference
Providers router (router.huggingface.co): imageConfig.modelMap resolves
Hub ids to provider-resolved ids, image-to-image models receive the
source image in inputs with the prompt under parameters.prompt, and a
new sttConfig wires the hf-inference ASR route. The image catalog grows
to 23 models, dead whisper-small is replaced by whisper-large-v3-turbo,
the unusable "language" param is dropped, and edit models declare the
edit capability so the dashboard offers a source image. Adds unit and
end-to-end coverage plus a model-id guard on custom endpoints.
2026-09-19 10:44:34 +07:00
Aaron
822aa958d1 fix(opencode): cloak Responses requests that already have tools
Free-tier Zen models reject Responses requests with 403 FreeTierError
when client tools are present but the fingerprint quartet is missing.
Apply the fingerprint tools to every OpenCode request, canonicalise
case variants of the quartet (Bash->bash) without duplication, and
restore the caller's original spellings on the response side via a
request-local WeakMap threaded through the existing toolNameMap.
2026-09-19 10:15:45 +07:00
Alexander Radchenko
49185137b8 feat(opencode-zen): add OpenCode Zen (PAYG) provider with free-tier fingerprint, alias ocz
Multi-endpoint provider on https://opencode.ai/zen/v1 (openai / claude /
openai-responses transports mirroring opencode-go) with 71 models across
paid + free tiers. Free-tier fingerprint: opencode/1.18.x UA spoof,
ses_-session header, built-in tool quartet, forced stream. Usage endpoint
/zen/v1/usage wired into the dashboard.
2026-09-19 10:02:28 +07:00
X-Adam
73e021b8a0 fix(ollama): map free-plan monthly window and derive reset from signup date 2026-09-19 09:51:13 +07:00
decolua
a8c9d3802c docs: update changelog header to v0.5.81
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-18 18:32:14 +07:00
decolua
23ae82d8e3 # v0.5.81 (2026-09-18)
## Features
- **Xiaomi MiMo**: merge MiMo Desktop support into `xiaomi-mimo` with dual auth (API key + Desktop/OAuth session), Preview models support, and encrypted-callback OAuth flow
- **Claude Code**: add 1M-context toggle (`[1m]` marker) and drive `CLAUDE_CODE_AUTO_COMPACT_WINDOW` directly from the dashboard
- **Models**: add DeepSeek-V4.1-Flash to DeepSeek provider, CodeBuddy-Intl, and Ollama (`deepseek-v4.1-flash:cloud`); enable `low`..`max` reasoning effort levels and vision capability for DeepSeek-V4.*
- **i18n**: integrate Persian (fa) translation

## Fixes
- **OpenCode / OpenCode Go**: resolve 403 `FreeTierError` and 429 rate limits with canonical session format, valid User-Agent, and stable upstream session reuse; force stream and declare `forceStream` for free-tier SSE aggregation; cloak decoy tools, normalize Muse Free tool choice, and strip prior reasoning items on Responses models; route Union Alpha via Messages API
- **Kiro**: preserve underscores in tool names (`mcp__server__tool`) and restore client tool names in responses; use neutral placeholder for tool-result-only turns; forward tool-result images
- **Stream**: report aborts after HTTP 200 in-band (per-format error frames) instead of closing silently
- **Command Code**: preserve images and `reasoning_effort` on `/alpha/generate`; retry transient stream errors and avoid fake stop chunks; add Quota Tracker support
- **Zed**: harden OAuth lifecycle (preserve `systemId`, renew proxy timeout), support live model resolution, and lower display priority in OAuth list
- **Antigravity**: scope cached thought signatures to model family; strip Claude Code billing headers from system prompts; sanitize Hermes system identity
- **Codex**: route bare `codex-auto-review` requests to the Codex provider (#4135)
- **Auth**: do not cool down an account for request-scoped 4xx errors
- **Usage**: improve DeepSeek credit balance display as currency credit instead of 0/total quota bar
- **Model Catalog**: scope synced catalog to gateways and declare vision capabilities for DeepSeek V4.1-Flash IDs
2026-09-18 18:31:12 +07:00
decolua
8e15f0bdd8 # v0.5.79 (2026-09-18)
## Features
- **Xiaomi MiMo**: merge MiMo Desktop support into `xiaomi-mimo` with dual auth (API key + Desktop/OAuth session), Preview models support, and encrypted-callback OAuth flow
- **Claude Code**: add 1M-context toggle (`[1m]` marker) and drive `CLAUDE_CODE_AUTO_COMPACT_WINDOW` directly from the dashboard
- **Models**: add DeepSeek-V4.1-Flash to DeepSeek provider, CodeBuddy-Intl, and Ollama (`deepseek-v4.1-flash:cloud`); enable `low`..`max` reasoning effort levels and vision capability for DeepSeek-V4.*
- **i18n**: integrate Persian (fa) translation

## Fixes
- **OpenCode / OpenCode Go**: resolve 403 `FreeTierError` and 429 rate limits with canonical session format, valid User-Agent, and stable upstream session reuse; force stream and declare `forceStream` for free-tier SSE aggregation; cloak decoy tools, normalize Muse Free tool choice, and strip prior reasoning items on Responses models; route Union Alpha via Messages API
- **Kiro**: preserve underscores in tool names (`mcp__server__tool`) and restore client tool names in responses; use neutral placeholder for tool-result-only turns; forward tool-result images
- **Stream**: report aborts after HTTP 200 in-band (per-format error frames) instead of closing silently
- **Command Code**: preserve images and `reasoning_effort` on `/alpha/generate`; retry transient stream errors and avoid fake stop chunks; add Quota Tracker support
- **Zed**: harden OAuth lifecycle (preserve `systemId`, renew proxy timeout), support live model resolution, and lower display priority in OAuth list
- **Antigravity**: scope cached thought signatures to model family; strip Claude Code billing headers from system prompts; sanitize Hermes system identity
- **Codex**: route bare `codex-auto-review` requests to the Codex provider (#4135)
- **Auth**: do not cool down an account for request-scoped 4xx errors
- **Usage**: improve DeepSeek credit balance display as currency credit instead of 0/total quota bar
- **Model Catalog**: scope synced catalog to gateways and declare vision capabilities for DeepSeek V4.1-Flash IDs
2026-09-18 18:09:32 +07:00
decolua
058ceace48 fix(opencode): declare forceStream on transport for free-tier SSE aggregation
Declares forceStream: true so chatCore properly converts upstream
forced-stream responses to JSON for non-streaming callers.

Co-authored-by: anojndr <anojndr@gmail.com>
Co-authored-by: TEGAR-SRC <tegararrahman17@gmail.com>
Co-authored-by: yxxrn <yxxrn@users.noreply.github.com>
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-18 17:32:52 +07:00
decolua
d99bc8201d fix(sidebar): temporarily hide NEW badge from 9Remote menu
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-18 17:12:57 +07:00
Louis Phạm
bc3be0cb28 fix(antigravity): scope cached thought signatures to the model family 2026-09-18 17:09:36 +07:00
Christian Gennari
092c84eac9 fix(commandcode): retry on transient stream error and avoid fake stop chunks 2026-09-18 17:09:24 +07:00
Louis Phạm
b3d6e089c6 fix(antigravity): strip Claude Code billing header from system prompts 2026-09-18 17:08:53 +07:00
Amirsalar Sojoudi
efc80ba2e3 fix(codex): route bare codex-auto-review to the Codex provider (#4135) 2026-09-18 17:05:47 +07:00
decolua
4641c2b76a fix(zed): lower display priority in oauth list
Move zed provider to the bottom of oauth providers by setting priority to 999.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-18 16:59:45 +07:00
decolua
93837af09f fix(opencode): fix free tier 403 error and improve China region handling
- Force stream:true and cloak decoy tools (bash, read) for OpenCode free tier
- Support connection testing for opencode in testUtils
- Expand error message slice limits in auth and ping to preserve workspace link
- Add concise China region link chip in provider detail page

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-18 16:57:06 +07:00
decolua
52917a6d49 Add NEW badge to 9Remote menu 2026-09-17 20:18:56 +07:00
decolua
f64229530e fix(antigravity): sanitize Hermes system identity 2026-09-17 20:18:37 +07:00
Ali Shaikh
3ac100d524 fix(usage): improve DeepSeek credit balance display
Display DeepSeek prepaid balances as credit with currency instead of a 0/total quota bar, and mark balance items as isCreditBalance.
2026-09-17 20:06:33 +07:00
Mosabbir Maruf
ef18175226 fix(zed): harden OAuth lifecycle and live model support
- executors/zed.js: use exact wire values (anthropic, open_ai, google, x_ai)
  and strip incompatible Vertex safetySettings on the Google path
- shared/zedAuth.js: robust callback query parsing, reject garbage PKCS#1 v1.5
  decryptions, and thread proxyOptions when fetching LLM tokens
- oauth: preserve systemId across authorize/register/exchange lifecycle,
  renew proxy idle timeout on reuse, and ignore non-callback localhost requests
- shared/OAuthModal.js: track owned proxy in flowRef and stop at most once
- api/providers/[id]/models: add connection-scoped live Zed model resolver
- registry: unhide provider in dashboard
- tests: add unit coverage for wire format, native auth, and live models
2026-09-17 20:02:57 +07:00
Ahmad Beyranvand
725e2c1187 feat(i18n): integrate Persian (fa) translation 2026-09-17 18:45:58 +07:00
Rafli Ahmad Zulfikar
367fc546d8 feat(models): deepseek-v4.* accepts low..max effort; flag thinkingEffortSupported and vision 2026-09-17 18:40:07 +07:00
RaoYu
20a43f5a2c fix(auth): don't cool down an account for a request-scoped 4xx
Do not trigger account cooldown or fallback for request-scoped 4xx errors that match no account rules so healthy credentials are not locked out for context length or validation errors.
2026-09-17 18:26:22 +07:00
KunN-21
aa14ef72e2 fix(opencode): normalize Muse Free tool choice
OpenCode Free returns HTTP 400 for muse-spark-1.3-contributor-free when tool_choice is non-auto. Declare forceAutoToolChoiceModels quirk and normalize explicit tool_choice to auto.
2026-09-17 18:20:08 +07:00
anojndr
eafac37dcb fix(opencode): strip prior reasoning items on Muse Spark Responses models
Strip prior-turn type: reasoning items and encrypted_content fields from body.input on Muse Spark Responses endpoints in OpenCode and OpenCode Go executors to avoid HTTP 400 errors across rotated proxy accounts.
2026-09-17 18:19:21 +07:00
Qisthi Ramadhani
c49efdf528 fix(kiro): preserve underscores in tool names and restore sanitized names in responses
Do not collapse consecutive underscores in uniqueName so mcp__server__tool is sent intact to Kiro, attach reverse map on request translation, and restore client tool names in responses.
2026-09-17 18:14:59 +07:00
Qisthi Ramadhani
82b1bca42a fix(kiro): use neutral placeholder for tool-result-only user turns
Replace the literal 'continue' placeholder on tool-result-only user turns with 'Tool results provided.' to prevent models from treating it as a new user instruction.
2026-09-17 18:12:48 +07:00
Manan Santoki
f4f06f290c fix(translator): keep tool-result images, restore Kiro tool names, preserve thinking display
Forward images inside tool_result to OpenAI and Kiro upstreams via following user messages, restore original client tool names on Kiro responses via _toolNameMap, and preserve thinking display settings across translations.
2026-09-17 18:12:05 +07:00
ErfanBagheri404
0c6ab4f99b fix(opencode): reuse one stable upstream session per identity to stop 429s
Follow-up to the canonical-session fix: with no explicit session,
every request minted a fresh x-opencode-session, and upstream free-tier
quota is accounted per session. That burns through quota and surfaces
as 429 FreeUsageLimitError with growing reset-after delays, while the
real CLI reuses one long-lived session per conversation.

- Stable canonical session per downstream identity (connectionId, else
  auth-header hash, else shared default), evicted after
  MEMORY_CONFIG.sessionTtlMs like the other session stores.
- Deterministic x-opencode-request per message (stable across retries,
  like the CLI user message id); valid downstream ids preserved.
- 6 more unit tests (22 total).
2026-09-17 18:03:58 +07:00
anojndr
6091ff597e fix(opencode): resolve 403 FreeTierError with canonical session format and valid User-Agent
OpenCode upstream validates free-tier requests: User-Agent must be opencode/<version> (>= 1.17.0) and x-opencode-session must match canonical ses_ format. Default OPENCODE_UA to opencode/1.18.31, generate canonical descending session IDs, provide deterministic foreign session translation, and isolate credentials per-request.
2026-09-17 18:03:20 +07:00
Hermes Agent
2b65c49ff5 fix(opencode): route Union Alpha through Messages API
Route union-alpha to /zen/v1/messages with targetFormat claude, add anthropic-version header, and register model capabilities (vision, 262K context, 131K max output).
2026-09-17 18:01:43 +07:00
Ridho Perdana
702b57c30d fix(opencode-go): route every responses-only model (incl. thinking variants) to /responses
Derive responses-only routing from the model registry's targetFormat instead of hardcoding model checks, and strip thinking suffixes when looking up models in providerModels so variants like gpt-5.6-luna(high) are routed correctly to /responses.
2026-09-17 18:00:01 +07:00
izzzzzi
912ed295db fix(deepseek,model-catalog): vision for V4.1-Flash ids, scope synced catalog to gateways
- Declare deepseek-v4.1-flash and deepseek-flash as vision-capable in MODEL_CAPABILITIES
- Share installed catalogSource across route chunks via globalThis.__9rCatalogSource
- Scope catalog modality keys by provider:model to prevent cross-gateway collisions
- Upgrade catalog format to v2 with automatic rebuild of older schemas
2026-09-17 17:55:33 +07:00
28e26f4295 Merge origin/master (v0.5.75) into gitea/new_feature
Resolve conflicts:
- package.json / cli/package.json: take 0.5.75
- .gitignore: union both sides (upstream 9router-*/temp files + local state dirs)
- CHANGELOG.md: keep both blocks, v0.5.75 above v0.5.70
- nonStreamingHandler.js: merge imports (unwrapClineEnvelope +
  tokensForDetail/shouldPersistRequestDetail); drop dead appendRequestLog
- providers/[id]/page.js: union useState blocks (compatible-model states
  + importingClineModels)

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-09-17 14:14:43 +07:00
81a47f36c6 fix(combo): keep nested combos as single units and stop 0-token detail rows
Nested combos (comboA lists comboB, comboC, …) now stay one slot each:
the inner combo always runs as fallback to produce a single answer.
Failed hops are no longer written to Details/usage, and streaming no
longer inserts a 0-token placeholder row.

- chat.js: comboStack cycle detection; nested combos forced to fallback;
  persistUsage="success-only" for combo hops
- combo.js: discardResponse() cancels unused bodies (fusion timeout /
  fallback) so dropped streams fire onStreamComplete; getComboModelsFromData
  keeps nested names and honors enabled=false
- requestDetail.js: tokensForDetail() canonicalizes Claude/Gemini usage;
  shouldPersistRequestDetail() skips streaming-start and non-success hops
- streamingHandler.js: drop the 0-token streaming placeholder write
- RequestDetailsTab.js: read Gemini/Claude token names; show "streaming"
  status in amber
- tests: add combo-nested.test.js (13 cases)
- gitignore: ignore local .vitest/ artifacts

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-09-17 14:03:59 +07:00
KhuatHieu
13b468b889 fix(commandcode): preserve images and reasoning_effort on /alpha/generate
Command Code dropped vision and ignored client effort through the router:
image blocks became "[image omitted]", HTTP image URLs were never inlined,
and reasoning_effort landed on the envelope wrapper instead of params (so the
DeepSeek family mapping remapped low -> high). The catalog also treated
deepseek/deepseek-v4.1-flash as text-only, so the vision adapter stole those
requests to another provider.

- Map OpenAI image_url / Claude image blocks (base64 or data-URI) to the
  native {type:"image", image:"data:...;base64,...", mimeType} generate block.
- Add FORMATS.COMMANDCODE to TARGETS_NEED_BASE64 so remote http(s) images are
  inlined by the existing SSRF-safe fetcher before translation.
- Write reasoning_effort inside params for targetFormat commandcode and pass
  low|medium|high|xhigh|max through unmapped; allow it in thinkingLevels.
- Provider-scoped capabilities for commandcode/cmc: vision except the CLI
  text-only denylist, thinkingFormat commandcode, so family patterns
  (deepseek-v4 -> thinkingFormat deepseek, vision false) no longer win.
- Quota Tracker: whoami + billing credits/subscriptions (credits vs plan cap,
  5h and weekly windows), labels from AI_PROVIDERS[].name.
2026-09-16 20:17:25 +07:00
decolua
9300121366 fix(stream): report aborts after HTTP 200 in-band instead of closing silently
A stream that stalled or lost its upstream was closed with no terminal frame
at all, so clients saw "200 OK, a few chunks, then nothing" and could not tell
a truncated reply from a finished one. The Responses passthrough path already
synthesized response.failed; every other client format got nothing.

The watchdog now hands its reason ("stream stall timeout" or "upstream
connection lost") to onAbortTerminal, and buildStreamErrorBytes frames it per
client format: OpenAI-compatible clients get data: {"error":{...}} followed by
data: [DONE], Anthropic clients get `event: error`. The error frame always
precedes [DONE] (openai-python raises APIError on any data payload carrying an
error key), and no synthetic finish_reason is ever emitted — a truncated
stream must not look like a clean stop.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-16 20:04:10 +07:00
decolua
5c399b6406 feat(codebuddy-intl,ollama): add DeepSeek-V4.1-Flash
codebuddy-intl: deepseek-v4-flash replaced by deepseek-v4.1-flash (same
gateway catalog as CN) and a capability override so the model keeps the
openai-style reasoning_effort path instead of the vendor-native "deepseek"
thinking shape the gateway rejects. Thinking levels low/high/xhigh.

ollama: add deepseek-v4.1-flash:cloud (verified on ollama.com/api/tags) with
vision + 1M context caps.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-11 22:03:48 +07:00
decolua
17c4cc7687 feat(claude-code): drive auto-compact window, add a 1M-context toggle
The "Context window" dropdown wrote CLAUDE_CODE_MAX_CONTEXT_TOKENS, which
Claude Code ignores for any model it recognizes: its window resolver returns
the env value only when the id is unknown to the model table, so every
claude-* mapping kept the built-in 200K and the dropdown did nothing. It was
never the compaction threshold either.

- Replace it with CLAUDE_CODE_AUTO_COMPACT_WINDOW — the documented trigger
  (100K–1M, clamped to the model window, env beats the autoCompactWindow
  setting) — and relabel the field Auto-compact. The 1M preset becomes 700K,
  which no longer collides with the marker it depends on.
- Add a "1M context" checkbox that appends the `[1m]` marker to the
  ANTHROPIC_DEFAULT_*_MODEL envs. Claude Code assumes 200K unless the name
  carries the marker — the resolver is a plain /\[1m\]/i test on the string,
  so it applies to any id and no model lookup is involved; the user decides
  which models are worth declaring as 1M.
- Toggling rewrites the model inputs immediately, and Apply writes them
  verbatim, so a marker typed by hand is not stripped.

Rename maxContextTokens -> autoCompactWindow through the POST body and
RESET_ENV_KEYS so a reset clears the key actually written.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-11 00:09:15 +07:00
叶炜朋
73cb89143c feat(xiaomi-mimo): merge MiMo Desktop support into xiaomi-mimo as dual auth
Adds the Desktop-exclusive Preview models and the Xiaomi account-session
route to the existing xiaomi-mimo provider instead of a separate
xiaomi-desktop provider, so the dashboard shows one MiMo entry rather than
three overlapping ones.

Dual auth, same pattern as kimi — API key (sk-) covers the cloud API,
Desktop/OAuth adds the account session used by the Preview models:

- registry: category oauth, authModes [oauth, apikey], oauth block, the two
  mimo-x-*-preview models, and the invite signupUrl
- executor: routes Preview models to the account-service route with a Cookie
  session, everything else keeps the sourceFormat-matched transport
- oauth: custom ECDH encrypted-callback flow (X25519 -> SHA256 -> AES-256-GCM)
  with a loopback callback proxy, plus one-click import of the local Desktop
  auth.json
- usage: weekly quota from the account session

Fixes found while merging:

- the OAuth browser flow was dead: poll-status cleared the session before the
  client could POST /exchange, so every exchange returned 400
- a Claude-format client was sent to /v1/chat/completions instead of the
  declared /anthropic/v1/messages transport, because buildUrl ignored
  runtimeTransport
- stopXiaomiMimoProxy leaked every pending session (each holding an X25519
  private key) for the process lifetime
- the OAuth exchange did not persist the Desktop passToken, so the Preview
  models could never work after a browser sign-in

Removes dead code: the local engine token minting (mimoEngine, never called
on the request path), the model-catalog and usage routes, engineToken/
engineUrl plumbing, and an unread top-level usage block.

Adds tests/unit/xiaomi-mimo-{executor,oauth-session,oauth-proxy}.test.js —
the provider previously had none.
2026-09-10 23:42:41 +07:00
decolua
83af3f1853 # v0.5.75 (2026-09-10)
Covers the 24 commits since the v0.5.69 tag. The package version already moved
to 0.5.75 in 4a390685b (CLI model selector), so the release commit is the
changelog alone, matching the convention in eb712ca82/4eda76e2a.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-10 23:35:10 +07:00
decolua
accf2c5296 fix(providers): remove the duplicate qwen provider
qwen.js points at the same transport as alims-intl — identical baseUrl
(dashscope-intl compatible-mode), headers, and quirks — but carries only 8
Qwen models against alims-intl's broader catalog, so it adds nothing a user
could not already reach. Node count is back to 81.

The id was dropped on 2026-08-05 (dcdd4628b) when the Qwen OAuth flow died;
this removes the standalone API-key entry that shadowed it. The `qw` alias
returns to unassigned, its state before today.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-10 23:32:55 +07:00
decolua
c712641123 feat(opencode-go): list deepseek-v4.1-flash first in the model catalog
Registry order is the display order for the provider page, /v1/models, and the
CLI selector, so moving the entry to the head of `models` is the whole change.
deepseek-flash keeps its id and supportedFormats — only its position moves.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-10 23:18:58 +07:00
decolua
da6aa90128 fix(video/vertex): reject job ids and model ids that escape the URL path
A Vertex job id is a base64url-encoded operation name, and base64url decoding
accepts arbitrary bytes without throwing, so the previous decode-and-split
check let a crafted id splice a path traversal into the fetch URL while the
Bearer token stayed attached — e.g. "..%2F..%2Fevil" resolved to
/v1/evil:fetchPredictOperation on the Vertex host. body.model had the same
shape on the create path, where it is interpolated into the URL unescaped.

decodeJobId now requires a charset-only id, a byte-for-byte round-trip, and a
decoded name matching ^projects/{p}/locations/{l}/publishers/{pub}/models/{m}/operations/{op}$
— no field may contain "/", so ".." can never reach the URL. model ids are
restricted to [A-Za-z0-9._-].

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-10 23:18:52 +07:00
LLL
248d7da01c revert(qoder): drop the Responses usage plumbing from shared code
The merged Qoder work also rewrote shared translator/handler code so that
/v1/responses clients got token usage on response.completed. That changed
behaviour for every provider, not just Qoder: proxies saw input tokens
rise by the 2000-token context buffer, and the plain token mapping was
replaced by one that always adds input_tokens_details.

A probe confirms the Qoder benefit does not depend on those edits: the
executor's coalescer already emits one include_usage-style finish chunk, so
a Claude client receives input_tokens and cache_read_input_tokens with
every shared file at its original state. Only the Responses path relies on
the shared translator, and that path has no Qoder-owned seam to put it in.

Reverts the shared files to their pre-PR state and drops the Responses
usage test. The Cline envelope unwrap in nonStreamingHandler.js, which
landed after the PR in the same file, is kept.
2026-09-10 23:13:06 +07:00
galiehneh
998bb3d975 fix(tools): scope Claude tool type defaulting to gateways that need it (#3905)
defaultClaudeToolType() stamped tools[].type = "custom" onto every
Claude-format request carrying tools since e08ac6da. That satisfied
MiniMax (error 2013) but broke Anthropic-compatible endpoints that only
accept the legacy typeless tool shape. DeepSeek's endpoint
(api.deepseek.com/anthropic/v1/messages) whitelists its tool `type` enum
to the web_search_* variants and answers HTTP 400 "unknown variant
`custom`", so every Claude Code request routed to a DeepSeek connection
failed and surfaced as a persistent 503.

Run the defaulting only when the target provider declares the new
requireClaudeToolType quirk (MiniMax, MiniMax-CN). Add
shouldDefaultClaudeToolType(provider, finalFormat, tools, PROVIDERS) in
translator/concerns/toolCall.js so the gate is unit-testable, and cover
MiniMax keeping the explicit type, DeepSeek/Anthropic staying typeless,
non-Claude formats and tool-less requests never defaulting.

Effectively a no-op for MiniMax and a restore of the pre-e08ac6da
behaviour everywhere else. Another strict gateway now only needs the same
one-line quirk instead of a global behavioural change.
2026-09-10 23:03:53 +07:00
Nick Nyanjui
8a81085a72 fix(claude): cap re-anchored cache_control at the 4-marker budget and keep single-object content turns
Anthropic accepts at most 4 blocks carrying cache_control per request. When the
client had already spent that budget, the re-anchor added a 5th marker and the
request was rejected with a non-retryable 400 that the failure path treated as
an account problem, retrying the same malformed body across the whole pool until
every account locked.

anchorClaudeCache now normalizes bare-object content, strips the invalid
cache_control carried by defer_loading tools, pins the 1h head anchors on the
last system block and last cacheable tool, then trims an over-budget body to 4
markers. The trim holds those head anchors and fills the remaining slots with the
tail-most message markers: a plain "keep the last four in document order" rule
drops the anchors first even though they lead document order, and skipping the
re-anchor at a spent budget left system/tools on the 5m default instead of 1h.

Some clients send content as a single block object rather than a one-element
array. Such a turn was dropped or zeroed on every leg that reads messages,
silently losing conversation history. normalizeMessageContent wraps it as a
one-block array on all four paths, and hasValidContent keeps it.
2026-09-10 22:54:07 +07:00
Nick Nyanjui
122f23eebc fix(cline,airforce): unwrap {success,data} envelope, add live catalog, and refresh airforce free models
Cline (api.cline.bot) wraps non-stream chat completions in
{"success":true,"data":{...choices...}}, which both the dashboard model-test
ping and the proxy non-stream path read at top level, producing "Provider
returned no completion choices for this model" (#3644). Unwrap the envelope
before usage extraction and response translation; the error envelope
({"success":false,...}) never matches and passes through untouched.

Scoped through `transport.quirks.clineEnvelope` so only cline/clinepass opt
in — no other provider's response body is ever rewritten.

Also adds a live Cline catalog: `fetchClineRawModels()` is shared between
`resolveClineModels()` (full catalog, including free-tier ids such as
z-ai/glm-5.3-flash) and `resolveClinepassModels()` (cline-pass/* only), wired
into /v1/models, the per-provider models route, and the combo selector's
model picker with the static catalog kept as fallback.

Refreshes the dead api-airforce free models (anthropic/claude-3.7-sonnet,
moonshot/kimi-k2.6, google/gemini-2.5-flash) with the live gpt-oss-120b,
gpt-oss-20b and kimi-k2.7-code, plus passthroughModels, forceStream and a
suggested-models filter.
2026-09-10 22:48:28 +07:00
izzzzzi
f6e7cabe60 fix(cline): stop workos:-prefixing ClinePass API keys and add clinepass token refresh
Cline/ClinePass requests failed with HTTP 401 ("Please make sure you are using
the latest version of Cline and re-authenticate your Cline account", #3230 /
#2333 / #3644). `getClineAccessToken()` unconditionally prefixed every token
with `workos:`, which is correct for Cline OAuth access tokens (WorkOS JWTs)
but wrong for ClinePass API keys — those are opaque strings (e.g. `clp_…`)
that the API accepts only verbatim, so the `workos:`-prefixed value was
rejected.

Only prefix tokens that look like a WorkOS JWT (`eyJ…`); API keys and other
opaque tokens pass through untouched, and an existing `workos:` prefix is
never doubled.

Also register `clinepass` in the token-refresh handlers. ClinePass shares
Cline's WorkOS auth endpoints, but without the entry expired ClinePass OAuth
tokens were never rotated, so every request kept 401ing. Finally, list
`apikey` first in the ClinePass `authModes` (ClinePass is meant to be used
with an API key from app.cline.bot/settings/api-keys), and add an "Import
from /models" button that pulls the live Cline catalog into custom models.
2026-09-10 22:48:22 +07:00
kimono381
45ec1d30bb fix(deepseek): keep Anthropic-only tool types when forwarding to /anthropic/v1/messages
DeepSeek's Anthropic-compatible endpoint accepts only the built-in
web_search_20250305 / web_search_20260209 tools and rejects client-defined
`custom` tools (MCP / Read / Bash) with HTTP 400 "unknown variant `custom`".
The generic non-Claude filter in prepareClaudeRequest dropped the offending
tools but also dropped the web_search_* ones DeepSeek does accept.

- Add an opt-in per-provider transport quirk `claudeSupportedToolTypes`; when
  declared it becomes a strict allow-list for Anthropic tool `type` values
- Stop stripping the `type` discriminator from surviving tools under that
  quirk, since DeepSeek needs it to route built-ins
- Declare the quirk on the deepseek transport with the two web_search_* types
- Providers without the quirk keep the previous filter and normalisation
  behaviour byte-for-byte; openai-format targets never reach this path
2026-09-10 22:26:44 +07:00
anhtran-ai
781c18d837 fix(codex): strip Unicode-property tool schema patterns Codex rejects
Codex's /responses validator has no Unicode property escapes, so a tool
`pattern` containing `\p{...}` 400s the whole request with `Invalid schema
for function ... is not a 'regex'` — identically on every account, costing a
full combo failover per turn (#3922).

- Add open-sse/utils/codexToolSchema.js: copy-on-write walk that drops only
  `pattern` values carrying a property escape, returning the original
  reference when nothing changed so the caller's schema stays intact for a
  retry against another provider
- Treat `properties` keys as property names, so a field literally called
  `pattern` is never read as the schema keyword; skip escaped literals via
  backslash-parity counting
- Apply it in normalizeCodexTools for both function and namespace sub-tool
  parameters, and log the strip count via dbg
- Add three cases to tests/unit/codex-tool-normalization.test.js
2026-09-10 22:25:47 +07:00
Federico Liva
1892ed77c8 fix(kiro): never send a top-level systemPrompt (400 REQUEST_BODY_INVALID)
kiro.dev rejects any body carrying a top-level systemPrompt with
400 REQUEST_BODY_INVALID. The translators stopped emitting the field in
v0.5.59 (the prompt travels in the first user turn via contentPrefix),
but two paths kept writing it back downstream of the translator:

- rtk/systemInject.js::injectKiroSystem() appended the RTK prompt to
  body.systemPrompt, so every kr/ model failed whenever an RTK injector
  (caveman, ponytail) was active. It now appends to the first history
  user turn's content (else currentMessage), reusing
  dedupStringAppend/hasPrompt so retries stay idempotent.
- executors/kiro.js::appendRepairInstruction() wrote the tool-call repair
  instruction to systemPrompt on the retry, turning every repair into a
  hard failure. It now appends to currentMessage.userInputMessage.content.

isKiroBody() no longer requires a string body.systemPrompt — that marker
is gone from the wire shape — and sniffs the conversation turn shape
instead, keeping the stray-conversationState guard intact. Stale comments
in both kiro translators corrected: the systemPrompt local is only a
session-replay cache key, not a wire field.

Also drops the mirror/rollback repair heuristic the injector no longer
needs: net -52 lines.

Fixes #3641, #3845, #2890, #2901, #2939, #3109, #3459, #3749
2026-09-10 22:25:35 +07:00
LLL
1f10f9e5c4 fix(qoder): report usage to all clients and stop inlining large attachments
- Coalesce Qoder's empty finish-in-delta frame with the later choices:[] usage
  frame so OpenAI and Claude clients receive prompt_tokens, completion_tokens
  and cache-hit tokens (the dashboard already saw them)
- Upload inlined images through /api/v2/image/upload like qodercli, and stub
  oversized non-image files instead of stuffing 30MB+ data URIs into
  agent_chat_generation
- Emit response.completed -> response.usage for chat-native upstreams so
  /v1/responses clients (Codex CLI, sub2api) no longer log 0/0/0
- Keep Claude message_delta.usage working when usage arrives without choices[0]
- Escalate to the smallest advertised Qoder context tier (200K/400K/1M) when
  the estimated prompt no longer fits max_input_tokens
- Pass apiKey for PAT connections and list hidden enable:false catalog keys
  from /v1/models
2026-09-10 22:08:19 +07:00
Hai Trinh
832a34659e feat(codex): add GPT Image 2.5, Flare and Sunburst image models
- Add gpt-image-1.5, gpt-image-2, gpt-image-2.5, gpt-image-2.5-flare and
  gpt-image-2.5-sunburst as Codex image models with multi-image support
- Add gpt-image-2.5, gpt-image-2.5-flare and gpt-image-2.5-sunburst to the
  OpenAI provider catalog
- Route tool-backed image models through the Codex responses model while
  passing the selected model to the image_generation tool, pinning
  tool_choice and deriving generate/edit from the presence of references
- Cover the Codex gpt-image-2.5 request shape with a unit test
2026-09-10 22:06:48 +07:00
coozgan
3288bbc47e feat(video): add OpenRouter and Vertex AI (Veo) video generation
Video generation was xAI-only. Adds an adapter layer under
open-sse/handlers/videoProviders/ so /v1/videos/* can target OpenRouter or
Google Cloud credentials. A provider with no adapter keeps the exact previous
behaviour (raw body to {baseUrl}/{action}, poll {baseUrl}/{id}, verbatim
passthrough), so the xAI path is unchanged.

- openrouter: async job shape identical to xAI; creation POSTs to the /videos
  collection root (no /generations suffix) and the registry HTTP-Referer /
  X-Title headers are applied. Bodies pass through verbatim.
- vertex: two-way translation, since Veo does not speak the OpenAI-ish videos
  shape. create -> :predictLongRunning { instances[], parameters{} }, poll ->
  :fetchPredictOperation (Veo has no REST GET poll). The operation resource
  name is base64url-encoded into the job id so GET /v1/videos/{id} stays a
  flat path. Access tokens are minted from Service Account JSON via the
  existing refreshVertexToken; raw API keys are rejected up front. The
  operation response maps back onto the { id, status, video, videos } shape
  clients already poll.
- videoCore: the request plan is rebuilt per attempt, so the 401 -> refresh
  once -> retry once path picks up the refreshed token. Adapter validation
  errors return 400 before any upstream call, so a malformed request can never
  create a billable job.
- videoGeneration: GET /v1/videos/{id} resolves the provider from the pinned
  x-connection-id connection, then ?provider=, then falls back to the xAI
  default.
- registry: openrouter and vertex gain videoConfig, the video serviceKind and
  video-kind models (Veo 3.1 / 3 / 2, Sora 2 Pro, Seedance 2.0).
2026-09-10 22:05:22 +07:00
zmf
807553e246 feat(codebuddy-cn): replace deepseek-v4-flash with deepseek-v4.1-flash
The server's product-config payload (which the IDE plugin fetches from
copilot.tencent.com) publishes deepseek-v4.1-flash and no longer lists
deepseek-v4-flash, so the old id is dropped — same pattern as the previous
catalog refreshes (#3648, #3802). The v4-flash endpoint still answers 200,
but the published list is the contract.

Per the server table, maxOutput rises 50000 -> 128000 while contextWindow
stays 1000000.

- registry/codebuddy-cn.js: models[] entry swapped to the new id
- capabilities.js: per-model entry swapped, maxOutput -> 128000

No changes needed in thinkingLevels.js (the deepseek-v4* pattern already
matches the new id and publishes low/high/xhigh, matching the server's
supportedEfforts or pricing.js (the deepseek-v* glob yields the same rates).
EOF
)
2026-09-10 21:57:49 +07:00
decolua
4a390685b3 feat(cli): group model selector by provider with search
Replace the flat numbered model list with provider-grouped browsing
(combos first, then providers by alias order), full-text search across
all models, and manual custom model ID entry. A single available
category opens directly into its model list.

Also bump root and cli packages to 0.5.75 and ignore packed
`9router-*` tarballs.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-10 21:56:30 +07:00
decolua
537b3befd2 Merge branch 'feat/opencode-go-models'
feat(opencode-go): add newly published Go models
fix(codex): restore Version header and single-source the CLI version
2026-09-10 21:25:40 +07:00
decolua
a7047a07d4 fix(codex): restore Version header and single-source the CLI version
The image handler's `version` header was commented out, so Codex image
requests reached chatgpt.com without the Version identity the backend
expects. Restore it and route every Codex identity header through one
constant.

The CLI version now lives on registry codex.transport as `cliVersion`
(the same pattern gemini-cli uses) and is re-exported as CODEX_CLI_VERSION,
so the registry User-Agent, the image handler and the connection test can
no longer drift apart. Bumped 0.136.0 -> 0.154.0 (current stable).

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-10 21:23:24 +07:00
decolua
eee3515e54 feat(opencode-go): add newly published Go models
Add the models the provider docs now list but the registry lacked:

  chat/completions  glm-5.3, kimi-k3, deepseek-flash, longcat-2.0,
                    hy4-preview, hy3
  + /messages       qwen3.8-max, qwen3.8-flash
  responses only    grok-4.6, gpt-5.6-luna

Endpoints follow the table at https://opencode.ai/docs/go/. chat-only
models stay on the sourceFormat-matched transport guard so a Claude
client is never routed to /messages for a model that lacks it.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-10 21:08:32 +07:00
617929zcxc
40dffbce53 feat(providers): add standalone Qwen provider 2026-09-10 18:53:50 +07:00
a84ba559f3 # v0.5.70 (2026-09-08)
## Features
- **Providers**: creating a compatible / custom-embedding node now registers the endpoint only — the API key is added afterwards from the node's page, like the built-in providers. The create dialogs drop the API Key / Model ID / Check fields, and `POST /api/provider-nodes` no longer accepts credentials at all, so a node can never be half-created
- **Providers**: compatible nodes now use the same model rows as built-in providers — capability badges, copy, per-model test, alias handling and the Add/Edit Model modal with vision + reasoning toggles, replacing the weaker read-only list
- **Providers**: compatible nodes get the built-in bulk toolbars: Test All / Disable / Active / Select All over connections, and Test All Models / Disable All / Active All over models, with per-row disable and a restore strip for disabled models
- **Providers**: wire the dead "Fetch Models" button on compatible nodes to the live upstream catalog, de-duplicating against already-added models

## Fixes
- **Models**: persist per-model capability assertions for custom and compatible providers and honor them everywhere — unsupported media is stripped on the chat path, `/v1/models` and `/api/models` report what the user asserted, and thinking translation follows it (asserting `reasoning:false` now actually strips thinking fields, `reasoning:true` emits them)
- **Models**: partial capability edits merge instead of overwriting, so toggling vision off no longer erases a stored reasoning assertion
- **Capabilities**: keep server-injected readers (synced catalog, user-asserted capabilities) in process-wide state — Next.js compiles startup and each API route into separate bundles with their own module instances, so a boot-time install was invisible to every request handler and the models.dev catalog contributed nothing to upstream requests since 0532f00d
- **Dashboard**: thinking-level picker and model-row suffix reflect user-asserted reasoning on compatible nodes
- **Providers**: `Default Model` is optional when adding an API key to a compatible node — the node's own model list (and the picker in the test modals) already determine what gets probed, and the built-in fallback still covers connection checks
- **Providers**: restore the `useCopyToClipboard` import dropped from the provider detail page, which crashed the route with `ReferenceError` for every provider
- **DB**: restore `getModelAliases` / `setModelAlias` / `deleteModelAlias` re-exports dropped from the `localDb` shim by 86112cee, which broke `GET /api/models` and `GET /v1/models` at import time
- **Providers**: remove dead `PassthroughModelsSection` (never passed props, superseded by the shared model rows)
- **Media Providers**: creating a custom embedding node reports that a key still has to be added, instead of claiming a key was saved; the edit dialog keeps its API Key + Check affordance since a stored key already exists there
- **Build**: self-host Inter instead of fetching it through `next/font/google` at build time — a Docker / mirrored builder with no route to `fonts.googleapis.com` failed the entire image build on `Failed to fetch 'Inter' from Google Fonts`. The seven `@font-face` rules and their `unicode-range`s copy what `next/font` emitted (a `latin`-only file would have dropped Vietnamese diacritics) and the latin subset is preloaded as before, so rendered metrics are unchanged
2026-09-10 11:11:24 +07:00
mrnim94
35b950be81 fix(kiro): route requests through current runtime surfaces and fix 400 REQUEST_BODY_INVALID (#3776) 2026-09-09 10:56:54 +07:00
Sutarto Jordan Chrisfivo
7fee56bacd fix(providers): clear stale locks after validation (#3830)
Clear stale connection health state (modelLock_*, backoffLevel,
rateLimitedUntil, errorCode) whenever a connection is explicitly
marked active after successful validation or OAuth re-login.

Closes #3810
2026-09-09 10:26:03 +07:00
B1nh M1nh
4ad1e7a4ba fix(usage): parse Fable weekly limit from limits[] instead of fabricating a row (#3847) 2026-09-09 10:19:50 +07:00
Christian Gennari
e3bf94ee25 feat(antigravity): add weekly quota tracking and free-tier handling (#3892) 2026-09-09 09:57:13 +07:00
decolua
628ff1eab5 fix(auth): set 24h maxAge for dashboard session cookie
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-09 09:45:23 +07:00
decolua
e7b5f09d50 fix(gemini): normalize contents and handle intermediate tool responses in Antigravity
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-09 09:45:18 +07:00
302795a613 fix(usage): restore xai quota label case; update stale kiro tests
3-way unit-suite comparison (branch-pre-merge vs origin/master vs
HEAD) with git worktrees:

* ProviderLimits/utils.js: parseQuotaData 'xai' case (ab9a3c1d
  weekly/api_usage label mapping) was dropped by an EARLIER merge
  (already failing at the pre-merge branch tip) — its companion test
  xai-usage.test.js has been red since. Restored verbatim from
  ab9a3c1d; the xai usage handler + dispatch entry survived, only the
  UI label mapping was lost.
* openai-to-kiro.test.js: 19 failures were stale upstream tests —
  1fc2a81d intentionally removed the redundant top-level
  systemPrompt field (Kiro rejects it) without updating them. The
  thinking/agentic directives now ride the frozen session-start msg0
  via contentPrefix. systemPromptOf reads msg0 (or the current
  message when there is no replayed session); the cross-turn
  stability test asserts the actual cacheability contract.

Suite now: 83 failing vs 103 on origin/master baseline; zero files
fail in HEAD that did not fail pre-merge. next build passes.
2026-09-08 10:53:21 +07:00
02f097880d fix(merge): restore ANTHROPIC_COMPATIBLE_PREFIX import + drop dead appendRequestLog call
Full-source eslint no-undef sweep over src/ + open-sse/ (config
listing node/web globals) found two remaining undeclared-variable
regressions of the same merge-loss family:

* api/providers/test-batch: branch commit f0adfb20 added a
  providerId.startsWith(ANTHROPIC_COMPATIBLE_PREFIX) check but never
  imported it; the master merge kept the buggy line. Every batch
  test / provider-group filter touching a non-openai-compatible
  provider threw ReferenceError (|| does not short-circuit).
* utils/stream.js finalizeStream: no-usage fallback called
  appendRequestLog(), a stub upstream had marked no-op and whose
  export chain was dropped by the merge. The call site only fires
  when a stream ends without valid usage and now threw
  ReferenceError inside the terminal callback. Removed the dead
  call (behaviour identical: the stub wrote nothing).

All other no-undef reports are browser globals absent from the
scan config, not source bugs. next build --webpack passes; smoke
server boots and auth-rejects unauthenticated /v1 + /api traffic
as expected; unit suite 1746 pass / 103 fail (was 1738/111).
2026-09-08 10:36:24 +07:00
38ff11ee16 fix(merge): restore chat.js imports, trust-vision floor, gitignore entries
Audit of every branch-owned line the -X theirs merge dropped from
the 32 pre-merge commits found three more real regressions:

* src/sse/handlers/chat.js: merge kept the capsOverride feature
  (bb8d67ba) but reverted the import block, so getCustomModels and
  capabilitiesFromServiceKind were undefined. The runtime error was
  swallowed by the feature's own fail-open try/catch — custom
  models silently lost their vision override. Restored both imports.
* open-sse/providers/capabilities.js: TRUST_UPSTREAM_VISION (the
  floor that keeps vision on for unknown models on upstream-validating
  gateways like openrouter) was left as dead code by the merge —
  upstream rewrote step 4 as refine() and dropped the check.
  Re-applied it on top of the new refine() so catalog/limits
  refinement still applies.
* tests/unit/chat-connection-pin.test.js: mock auth module lacked
  isModelAllowedForKey added by f0adfb20.
* .gitignore: re-add .pi-subagents/.
2026-09-08 09:38:03 +07:00
e3ba5d2207 fix(usage): restore commandcode quota + timeout guards lost in master merge
The -X theirs merge of origin/master (v0.5.69) silently reverted five
branch-only hunks because upstream had no conflict-region counterpart
and simply won the three-way pick:

* services/usage.js: re-register the commandcode USAGE handler +
  import. Without it the dashboard Quota Tracker fell through to
  'Usage API not implemented for commandcode'. The handler module
  (services/usage/commandcode.js) and registry usage block survived;
  only the dispatch entry was dropped.
* ProviderLimits/utils.js: restore parseQuotaData 'commandcode' case
  (currency-credit rows need unit "$" + remainingPercentage
  forwarding, else $0.05 balances render as 0%).
* profile + providers/[id] pages: restore Math.max(1000, ...) connect
  timeout floors so a stray '60' is never interpreted as 60ms.
* .gitignore: re-add .commandcode/ CLI local state.

Verified: tests/unit/commandcode-usage.test.js (6) and
usage-dispatch.test.js (2, asserts every provider routes to a real
handler) pass standalone; full unit run 1742 pass / 107 fail vs
1738/111 before this fix (remaining failures pre-existing, unrelated).
2026-09-08 09:28:59 +07:00
ddfa789a31 fix(providers): add missing providerStrategies state + restore notify store
* page.js referenced setProviderStrategies (line 166) and
  providerStrategies (line 597) but never declared the
  useState — providers page threw ReferenceError on mount.
* The destructure of useNotificationStore() was also dropped
  by the merge of origin/master; 5 sites in the file called
  notify.error/.success/.warning.
* Both were present before the merge (commits de9e00c6 +
  upstream master versions). The statusFilter commit
  (d1d4e0f0) was the last state-block edit and survived,
  but adjacent state lines were lost during the conflict
  resolution.
2026-09-07 15:45:14 +07:00
6f52d7020c fix(chatCore): add missing capsOverride + streamErrorPatterns to destructure
* handleChatCore() referenced capsOverride (line 162) and
  streamErrorPatterns (line 475) but the destructured param list
  did not include them. Callers that did not pass these (e.g.
  open-sse/handlers/responsesHandler.js, unit callers, older
  client builds) would trigger 'capsOverride is not defined' /
  'streamErrorPatterns is not defined' ReferenceError mid-request.
* Both defaulted to null. capsOverride is read by capability merge
  (line 162, already guarded by '|| {}'). streamErrorPatterns is
  read in the early-peek hook (line 484) which was null-safe
  only because the variable happened to be in scope when chat.js
  spread it in; responsesHandler never passed it and would crash.
* Unblocks all callers regardless of which fields they pass.
2026-09-07 15:23:48 +07:00
59f17b3725 feat(usage): raw request detail modal + raw stream capture
* Add /api/usage/request-details/raw endpoint serving a single
  stored request detail verbatim (raw payloads), with /raw doc
  clarifying it stays gated by the dashboard auth layer.
* Add RawDetailModal opened from a new 'Raw' button in
  RequestDetailsTab. Modal loads /raw, exposes per-section copy
  buttons and a 'Copy all (JSON)' that bundles every section.
* Capture the raw provider SSE text inside the streaming
  transform (cap 64KB) and forward it through
  onStreamComplete.rawProviderText so handler stores it as the
  providerResponse. response.content stays the extracted user
  text. Tool-call-only turns remain so the marker.
* Accumulate from translated client-facing chunks instead of
  raw provider shapes so Responses, Claude delta types, and
  Gemini/Antigravity parts all contribute.
* Drop redaction from the list endpoint; raw access is now via
  the dedicated /raw endpoint. Tests cover the new behavior.
2026-09-07 14:23:48 +07:00
a835771c97 Merge origin/master (v0.5.69) into gitea/new_feature 2026-09-07 14:10:11 +07:00
decolua
eb712ca821 # v0.5.69 (2026-09-05)
## Features
- **Codex**: add GPT 6.0 Astra (`gpt-6-astra`) with vision, thinking and search capabilities
- **Usage**: add Claude Fable quota tracker support with weekly window normalization (`weekly fable (7d)`)
- **Dashboard**: group Antigravity Gemini and Claude quotas in Quota Tracker, prune stale hidden keys
- **OpenCode Go**: add `muse-spark-1.3-contributor` model and support parallel tool calls on Responses path (#3819)
- **Providers & Models**: align CodeBuddy-CN catalog/capabilities with server config; add GPT-5.6 Sol, Terra, Luna image aliases on Codex (#3806); refresh Qoder catalog with capability mapping and image pass-through
- **CLI tools**: replace Copilot MITM with VS Code extension setup guide
- **Gemini**: persist and replay `thoughtSignature` scoped by session namespace

## Fixes
- **Claude**: normalize adaptive auto effort (`output_config.effort`) (#3792)
- **Antigravity**: prevent Google anti-abuse rate limits during multi-account refresh (#3813)
- **Anthropic-compatible**: forward Claude beta flags to nodes fronting Anthropic (#3797)
- **Dashboard**: dynamic mode label for local/remote detection (#3801)
- **Codex**: format reset credit API errors cleanly (#3778)
- **Security**: guard cowork MCP tools probe against SSRF (#3783)
- **OpenCode Go**: track OpenCode Go quota (#3791) and send stable session headers (#3800)
- **Logger**: suppress noisy background token refresh logs
- **CLI**: export packed `.tgz` directly into workspace root instead of parent directory
2026-09-05 22:57:00 +07:00
Sina Sadeghi
11222eff0f feat(opencode-go): muse-spark-1.2 and Responses tool fixes (#3820)
- Add muse-spark-1.2-contributor as responses-only model on OpenCode Go
- Normalize object tool schemas without properties in OpenCode Go executor
- Make fallback Responses call_ids unique across same-millisecond calls
- Make Responses output coercion fail-soft for circular and non-stringifiable values
2026-09-05 22:39:22 +07:00
decolua
e214fb1c30 feat(usage): add Claude Fable quota tracker support
- Recognize Fable weekly windows and normalize to weekly fable (7d)
- Fall back to 100% available weekly Fable window when Anthropic payload omits it
- Forward remaining percentages and enforce canonical Claude quota order in Quota Tracker

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-05 22:05:18 +07:00
decolua
f615a83cb2 feat(dashboard): group Antigravity model quotas and trim hidden keys
- Group Antigravity Gemini text models into single 'Gemini (Flash / Pro)' quota
- Group Claude models into single 'Claude (Sonnet / Opus)' quota
- Prune stale or legacy model keys from hidden quota visibility list

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-05 21:56:07 +07:00
Sina Sadeghi
e74db4d0a6 feat(opencode-go): add muse-spark-1.3-contributor and fix parallel tool calls on Responses paths (#3819)
- Add muse-spark-1.3-contributor as responses-only model on OpenCode Go with dedicated executor
- Key Responses→chat streaming tool calls by item_id to prevent parallel tool calls merging into index 0
- Standardize tool coercions and call_id clamping in Responses API translation
2026-09-05 21:49:53 +07:00
Sutarto Jordan Chrisfivo
77e6a227fe fix(claude): normalize adaptive auto effort (#3792)
Claude adaptive requests without an explicit effort are normalized to
output_config.effort: "high" instead of forwarding the unsupported
literal value "auto" which Anthropic rejects with HTTP 400.
2026-09-05 21:41:25 +07:00
Hifzi
1442cc73ce fix(antigravity): prevent Google anti-abuse rate limits on multi-account refresh (#3813) 2026-09-05 21:28:30 +07:00
Federico Liva
fb9fab0206 fix(anthropic-compatible): send Claude beta flags to nodes fronting Anthropic (#3797) 2026-09-05 21:26:16 +07:00
vianhanif
28cfd9facf fix(dashboard): dynamic mode label for local/remote detection (#3801) 2026-09-05 21:20:53 +07:00
Raisal P Wardana
1a3d446831 fix(codex): format reset credit API errors (#3778) 2026-09-05 21:18:07 +07:00
soroush5
97f3ab97b1 fix(security): guard cowork-mcp-tools probe against SSRF (#3783) 2026-09-05 21:12:16 +07:00
JOJO
0da803eef4 fix(usage): track OpenCode Go quota (#3791)
OpenCode Go API-key connections now appear in the Quota Tracker and report rolling, weekly, and monthly subscription usage.
2026-09-05 21:10:10 +07:00
turingcat
81f4f93082 fix(opencode-go): send stable session header (#3800)
- add a dedicated OpenCode Go executor that always sends x-opencode-session
- preserve a valid caller-provided native OpenCode session header
- translate downstream Agent session IDs into opaque, stable, Agent-scoped IDs
- forward the original provider session seed and client tool on both initial and credential-refresh requests
2026-09-05 21:09:49 +07:00
zmf
cec672d9d9 feat(providers): align codebuddy-cn catalog/capabilities with server config
- Sync codebuddy-cn catalog and capabilities with copilot.tencent.com server payload
- Fix thinkingCanDisable semantics for glm-5.3 and deepseek-v4 models
- Add missing glm-5.2 thinking levels to thinkingLevels.js
- Add glm-5-turbo model to glm and glm-cn registries
2026-09-05 21:03:02 +07:00
An Nguyen
ed963931b4 feat(codex): add GPT-5.6 Sol, Terra, and Luna image aliases (#3806) 2026-09-05 21:01:52 +07:00
decolua
f388b5e56b chore(logger): remove noisy background token refresh logs
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-04 10:17:26 +07:00
decolua
b84681d5a4 feat(cli-tools): replace copilot mitm with vscode extension setup guide
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-03 23:27:04 +07:00
hangyu
2ab6a4c949 feat(qoder): refresh model catalog, add capability mapping and image pass-through
- Registry/constants: drop qmodel_preview/gm51model, add lite,
  qmodel_38max (Qwen3.8-Max), qfmodel (Qwen3.8-Flash), gmodel (GLM-5.3),
  gfmodel (GLM-5.3-Flash)
- capabilities: add PROVIDER_CAPABILITIES['qoder'] so opaque internal
  ids resolve to their real models' context windows and limits
- executor: preserve image blocks instead of flattening away, convert
  Claude-style image blocks, and hash images into chat_record_id
- tests: cover image preservation, data-URI and Claude-block conversion
- build(docker): use CN mirrors for apk and npm
2026-09-03 23:02:52 +07:00
decolua
c08efdbe2b feat(gemini): persist and replay thoughtSignature with session namespace
- Add open-sse/services/thoughtSignatureStore.js managing LRU Map (2k) + SQLite kv table
- Store thoughtSignature with sessionId namespace and toolCallId fallback
- Replay cached signature by sessionId:tool_call_id to prevent multi-process collisions
- Normalize Antigravity sessionId to numeric int64 format

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-03 18:20:04 +07:00
decolua
4eda76e2ab # v0.5.65 (2026-09-03)
## Features
- **Fetch**: add Ollama Cloud web fetch provider
- **Gemini / Antigravity**: add Gemini 3.8 Flash support and bump IDE fingerprint to 2.11.0
- **Claude**: add Claude Fable 5.1 support (adaptive thinking with `output_config.effort`), bump Claude Code fingerprint to 2.1.258 for new-model access
- **Providers**: add client-side status filter (All / Active / Inactive / No connection) on the Providers dashboard; add max height and scroll for connection list
- **Providers & Models**: streamline tokenrouter model catalog down to 22 flagship/newest models and add missing provider icons; refresh Codebuddy-CN catalog (add hy4-preview/hy3/glm-5.3/kimi-k3-1, drop EOL glm-5.0/glm-4.7)
- **Models**: capability toggles (vision, reasoning) when adding custom models with upsert and live caps refresh
- **CLI tools**: support saving and managing custom API key presets
- **Quota**: add usage and rate-limit tracking for Groq via `x-ratelimit-*` headers
- **i18n**: complete Indonesian translation (1391 keys)

## Fixes
- **Security**: close SSRF guard bypasses in `ssrfGuard.js` (alternate IPv6 encodings, hostname trailing dots, wildcard DNS resolution check, safe redirect handling) (#3714)
- **Model markers**: strip the `[1m]` context marker Claude Code appends to model names (`claude-opus-5[1m]`) preventing model resolution failures (#3690)
- **Claude**: drop `server_tool_use` blocks carrying foreign IDs to avoid Anthropic 400 rejections; never anchor cache breakpoints on `defer_loading` tools (#3567)
- **Antigravity**: strike-break optimistic quota readings that keep 429ing by blocking the connection+model pair for 15m after 3 strikes (#3681); preserve client identity on model catalog requests (#3414)
- **Auth**: protect root `/responses` rewrite requiring API key validation in dashboardGuard
- **Chat & Docker**: return 503 Service Unavailable when all credentials are rate-limited; explicitly bundle `node-machine-id` into standalone Docker runtime image
- **OpenCode**: route Muse Spark models to `/zen/v1/responses` and declare vision support; filter inactive free model
- **Kiro**: preserve inline images as OpenAI-compatible `image_url` parts in OpenAI MITM; remove redundant top-level `systemPrompt` from payload
- **Usage**: read Responses-shape `cached_tokens` in `extractUsageFromResponse` for non-streaming traffic
- **Models**: support single model lookup with provider-prefixed IDs (e.g. `cc/claude-sonnet-5`)
- **Translator**: route Gemini thinking through `reasoning_effort` on OpenAI-compatible wire; convert `prefixItems` and ensure array items in Gemini schema sanitizer
- **UI**: apply persisted theme before first paint to prevent flash on reload; translate combo vision adapter label
2026-09-03 10:36:34 +07:00
Sami Basra
e0ffc7e2a1 feat(fetch): add Ollama Cloud web fetch provider 2026-09-03 10:13:15 +07:00
Federico Liva
6ab9ca9eb1 fix(claude): never anchor cache breakpoint on defer_loading tools (#3567) 2026-09-03 10:05:35 +07:00
decolua
6efb97904b feat(providers): streamline tokenrouter models and add missing provider icons
- Prune tokenrouter seed models from 121 to 22 flagship/newest models
- Add z-ai/glm-5.3-free with 0 pricing
- Add missing 128x128 icons for alims-intl, alitp-intl, fish-audio, and selfhosted-* providers

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-03 09:59:30 +07:00
decolua
831001c322 feat(providers): add max height and scroll for connection list
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-03 09:54:43 +07:00
IEatCodeDaily
e7dd72a8d7 fix(usage): read Responses-shape cached_tokens in extractUsageFromResponse
Non-streaming codex traffic recorded cached_tokens: 0 even when upstream
prompt caching worked. The Claude-format branch (which OpenAI Responses
usage also matches) never read input_tokens_details, and the OpenAI
branch ignored a top-level flat cached_tokens. Read both in both
branches; Responses prompts are cache-inclusive so canonicalizeUsage
passes the value through without folding. 5 new regression tests.
2026-09-03 09:48:19 +07:00
Sutarto Jordan Chrisfivo
5caa72f5fb fix(models): support single model lookup
Support single model lookup by replacing the one-segment models route
with a catch-all route that preserves capability kind paths while
allowing provider-prefixed IDs like cc/claude-sonnet-5.
2026-09-03 09:43:01 +07:00
zmf
e014cb537f feat(codebuddy-cn): refresh model catalog — add hy4-preview/hy3/glm-5.3/kimi-k3, drop EOL glm-5.0/glm-4.7
- Add hy3, hy3-x, hy4-preview, hy4-preview-x, glm-5.3, glm-5.3-flash, kimi-k3-1
- Remove dead models glm-5.0, glm-4.7 (API 11102)
- Register capabilities and context windows in PROVIDER_CAPABILITIES
- Configure supported effort sets in PATTERN_THINKING
2026-09-03 09:41:02 +07:00
openhands
d1d4e0f02b feat(providers): add status filter to providers dashboard
Adds a client-side status filter (All / Active / Inactive / No
connection) to the Providers page, applied over the already-fetched
provider + connection list. Status derives from getProviderStats
(total, allDisabled); noAuth providers count as Active. Filter composes
with the existing search across all provider sections. Part of #3699.
2026-09-03 09:39:01 +07:00
louis-cai
ac98dd9d32 fix(antigravity): strike-break optimistic quota readings that keep 429ing
Google's quota API can report remaining quota while generation endpoints
keep returning 429 (sprint/weekly dual-pool mismatch). handleAntigravityQuotaError
trusted remainingPercentage > 0 as healthy and returned null, causing 429 retry
loops across multi-account pools.

Add a strike-based circuit breaker to the optimistic and unavailable quota paths:
- After 3 strikes (429/409) within 60s for the same connection+model, cache-block
  that pair for 15 minutes by synthesizing an entry in the shared RAM quota cache.
- Re-assert active strike blocks across refreshes so optimistic readings cannot
  resurrect a broken pair prematurely.
- Reset strikes and clear synthesized cache entry upon successful request.
- Keep exact-resetAt handling for genuine 0% exhausted readings.

Closes #3681
2026-09-03 09:34:24 +07:00
Teguh Rijanandi
a58902e4a7 feat(i18n): complete Indonesian translation (1391 keys) 2026-09-03 09:33:01 +07:00
Sutarto Jordan Chrisfivo
98579f98c1 fix(auth): protect root /responses rewrite
Add /responses to PUBLIC_PREFIXES in dashboardGuard so pre-rewrite remote
requests require API key validation as intended.
2026-09-03 09:29:12 +07:00
vianhanif
15687d1913 fix(chat,docker): return 503 for rate-limited providers and bundle node-machine-id
- chat: always return 503 Service Unavailable when all credentials are rate-limited
- Dockerfile: explicitly copy node-machine-id into standalone runtime image
2026-09-03 09:25:17 +07:00
anojndr
acb5c34cdc fix(opencode): route Muse Spark models to Responses API and declare vision
Route all Muse Spark models (not just 1.2) on OpenCode Free to
/zen/v1/responses via isMuseSparkModel(), fixing HTTP 500 on
muse-spark-1.3-contributor-free. Declare vision:true on Muse Spark
models so image input is no longer stripped; register 1.3 in the
registry and capabilities. Scoped to opencode only — other providers
keep Chat Completions routing.
2026-09-03 09:24:18 +07:00
openhands
b870b5d41b fix(security): close SSRF guard bypasses in ssrfGuard.js (#3714)
Closes four SSRF guard bypasses reported in #3714:
- Block alternate IPv6 encodings (hex format, NAT64, IPv4-compatible, IPv4-mapped) by parsing to 16-bit groups
- Normalize trailing dots on hostnames to prevent FQDN bypasses
- Add assertPublicUrlResolved() with DNS resolution to block wildcard DNS domains resolving to private/metadata IPs
- Add fetchPublic() to safely handle and validate HTTP redirects
2026-09-03 09:22:22 +07:00
Outis
1f190bd00b fix(kiro): preserve inline images in OpenAI MITM
Forward Kiro userInputMessage.images as OpenAI-compatible image_url content parts.
2026-09-03 09:21:37 +07:00
Zafar
70f15aa50b feat(antigravity,gemini): add Gemini 3.8 Flash support and bump IDE fingerprint to 2.11.0
Co-authored-by: Schnee111 <daffamaarif.dev@gmail.com>
Co-authored-by: AhooraZen <ahoora935137@gmail.com>
Co-authored-by: anojndr <anojndr@gmail.com>
Co-authored-by: Emirhan <emirhan551952@gmail.com>
2026-09-03 09:13:45 +07:00
Lek Huda
1fe996db6a fix(translator): route Gemini thinking through reasoning_effort on OpenAI-compatible wire 2026-09-03 09:12:34 +07:00
decolua
c24a854278 feat(cli-tools): support saving and managing custom API key presets
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-03 09:06:08 +07:00
decolua
1fc2a81d65 fix(kiro): remove redundant top-level systemPrompt field from payload
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-03 09:06:02 +07:00
decolua
f6c59d30b0 fix(gemini): convert prefixItems and ensure array items in schema sanitizer
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-03 09:05:57 +07:00
LucasOl1337
ac9120fde3 fix(claude): support Fable 5.1
- add claude-fable-5-1 to the Claude Code model catalog (1M context,
  permanent adaptive thinking)
- centralize the spoofed Claude Code version and update both request
  and billing identities to 2.1.257 (Fable 5.1 rejects < 2.1.251)
- send output_config.effort without the redundant thinking switch for
  permanently adaptive models
- add regression coverage for capabilities, headers, billing identity
  and adaptive-effort payload

# Conflicts:
#	open-sse/providers/registry/claude.js
#	open-sse/providers/shared.js
#	open-sse/utils/claudeCloaking.js
#	tests/__baseline__/providers-baseline.json
2026-09-02 20:42:53 +07:00
Federico Liva
ee7a961633 fix: strip the [1m] context marker Claude Code appends to the model name
With the 1M-context beta enabled, Claude Code sends model: "claude-opus-5[1m]".
The marker is a client-side annotation — it matches no combo name, no alias and
no provider/model pair — so the request dies at model resolution with
"Invalid model format" and the client reports "There's an issue with the
selected model". Every request from that session fails until the beta is
switched off.

New open-sse/utils/modelMarkers.js exporting stripModelContextMarker(modelStr)
-> { model, contextMarker }. handleChat strips the marker before resolution and
normalizes body.model so downstream logging and translation see the real name.
Only a trailing marker is stripped, so a model whose name genuinely contains
brackets is left alone.

The capability itself travels in anthropic-beta: context-1m-2025-08-07, which
the default executor already forwards untouched — only the routing key needed
cleaning.

Fixes #3690.

Tests: tests/unit/model-context-marker.test.js (6 cases).
2026-09-02 20:20:32 +07:00
Matt Van Horn
f68d2f5ee5 fix(antigravity): preserve client identity on model catalog requests
Restrict the legacy IDE-version override to generation endpoints so
catalog and other passthrough requests keep their original User-Agent
and metadata.ideVersion, letting newer Antigravity releases see current
models like Gemini 3.7 Flash in the MITM selector.

Fixes #3414
2026-09-02 20:14:59 +07:00
Federico Liva
ed1bd0c528 fix(claude): drop server_tool_use blocks carrying a foreign id
Anthropic validates server_tool_use.id against ^srvtoolu_[a-zA-Z0-9_]+$
and 400s the whole request when one does not match. A combo that falls
back to a provider with its own built-in tools (z.ai/glm emits
OpenAI-style call_ ids for analyze_image) leaves such blocks in the
history, so every later Claude turn fails.

Extend normalizeClaudePassthrough to drop those blocks (reusing the
existing loop), drop the paired tool_result / web_search_tool_result
referencing a dropped id, and drop empty text blocks plus messages left
with no content. Well-formed srvtoolu_ blocks and regular tool_use ids
are untouched.
2026-09-02 20:07:16 +07:00
docaohieu2808
925cb4aade fix(ui): apply persisted theme before first paint to avoid flash on reload
Theme was applied from the client store in useEffect (after hydration),
so a reload painted the default light theme for a frame before the
stored dark theme was reapplied. Add a blocking head script that reads
the persisted zustand theme key and sets the dark class on
documentElement before first paint, mirroring applyTheme() including
system -> prefers-color-scheme resolution.
2026-09-02 20:05:52 +07:00
openhands
b9c92cb83c feat(quota): add usage tracking for Groq
First slice of #3701: quota tracking for Groq via x-ratelimit-* response
headers on the models endpoint (no dedicated quota endpoint exists, and
reading usage costs zero tokens).

- usage/groq.js: parse request+token limit/remaining headers; Go-style
  duration reset headers ("2m59.56s") resolve to future timestamps;
  missing key/401/403 -> message, 2xx without headers -> soft
  "not tracked yet" with quotas:{}
- registry/groq.js: transport.usage.url (reuses validateUrl) +
  features {usage, usageApikey}
- services/usage.js: groq entry in USAGE_HANDLERS
- ProviderLimits/utils.js: parseQuotaData case (absolute used/total,
  codex/kiro style)
- tests: groq-usage.test.js (registry flags, header parsing, soft
  not-tracked path, missing key/401, parseQuotaData)
2026-09-02 20:04:35 +07:00
decolua
44e4b80bbe fix(models): filter dead opencode free model
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-02 20:01:15 +07:00
dajinglingpake
9d3f7646d1 fix(i18n): translate combo vision adapter label 2026-09-02 19:45:34 +07:00
decolua
009cac6326 fix(claude): bump CC fingerprint to 2.1.258 for new-model access
Anthropic gates newly released models (e.g. claude-fable-5-1) to Claude
Code >= 2.1.251; the spoofed 2.1.92 client got HTTP 400 on every request.
Bump User-Agent + billing-header version to 2.1.258 and refresh the
providers baseline snapshot.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-02 11:27:21 +07:00
decolua
38f031f4c9 feat(models): capability toggles for custom models with upsert and live caps refresh
- AddCustomModelModal lets users pick vision/reasoning caps when adding a model
- POST /api/models/custom whitelists caps to booleans
- aliasRepo.addCustomModel upserts — re-adding updates caps/name in place
- /api/models includes custom llm models with stored caps overriding the heuristic
- useModelCaps refetches on customModelChanged instead of trusting a stale cache

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-01 10:45:51 +07:00
decolua
90b52e06ff # v0.5.59 (2026-08-29)
## Features
- **Search**: new web search providers — Antigravity (Google Search grounding
  on the existing OAuth account pool, citations keyed and merged by URL) and
  Xquik (X search with `x-api-key` auth, cursor pagination, credit-based
  usage), both on `POST /v1/search`. Based on #3437 by @Nautilaceae
- **Search**: ollama-search and zai-search borrow a chat provider's API key
  instead of requiring their own connection, driven by a new
  `credentialFallback` registry field. zai-search later folded into the `glm`
  provider itself so the web search page shows the shared connection
- **Models**: daily background sync of model capabilities from models.dev —
  modalities keyed by model id (majority of sources must declare one),
  context/output limits keyed by provider + model, strictly additive and
  sitting below the hand-written tables. ETag + mtime cache, 60s startup
  delay, `MODEL_CATALOG_SYNC=off` to disable
- **Models**: add GLM-5.3-Flash (1M context, natively multimodal), DeepSeek
  V4 Vision, Grok 4.5/4.6 (500k context); correct glm-4.6v/4.5v video input
  and output limits, backfill glm-4.6v on glm-cn
- **Usage**: show the Zed plan quota on the dashboard — plan, edit
  predictions, hosted model requests and billing-cycle reset; unlimited rows
  render as "N used · Unlimited"
- **Usage**: track GPT-5.3-Codex-Spark quota windows (spark_session /
  spark_weekly) from the Codex usage response (#3431)
- **Antigravity**: quota-aware routing — on 409/429 fetch live quota for the
  exact per-model resetAt and skip only the exhausted account/model pair;
  report the earliest reset when every account is blocked (#3561)
- **Antigravity**: map image `size` to the aspect-ratio model suffix (-WxH);
  add the Gemini 3.7 Flash tiers to MITM defaultModels so they show up in
  the dashboard model-mapping table
- **Dashboard**: bulk import Grok CLI accounts from JSON — paste an array or
  drag-drop multiple .json files, all OAuth connections created in a single
  call, mirroring the codex flow
- **CLI tools**: endpoint presets shared across every tool card through one
  live-resyncing store, instead of per-card localStorage copies that never
  saw each other's saved endpoints
- **Token Saver**: configurable compression timeout (`headroomTimeoutMs`) —
  the fixed 3000 ms made busy machines time out and send inconsistently
  compressed bodies, hurting prompt caching
- **i18n**: pt-BR expanded to 1132 terms

## Fixes
- **Stream**: record usage when a client closes on the terminal event — the
  Responses API has no [DONE] sentinel, so codex closed the socket on
  `response.completed` and cancelled the reader before flush() ran its usage
  side effects; the tail now lives in a once-guarded finalizeStream(). Also
  stop logging a disconnect for every completed Responses call
- **Stream**: parse the trailing NDJSON line an Ollama stream leaves behind
  without a closing newline — the final chunk carrying `done_reason` and the
  token counts was dropped
- **Session**: read the Claude Code session id from the
  `x-claude-code-session-id` header — `metadata.user_id` is dropped by
  Responses translation, splitting one conversation across several
  `prompt_cache_key` values and missing the upstream prefix cache
- **Usage**: preserve nested `cached_tokens` — the top-level-only read
  persisted `cached_tokens: 0` for every Responses-format provider (codex,
  grok-cli, …), billing cache hits at the full input rate
- **Usage**: GLM quotas accept CREDIT_LIMIT plans and multi-interval windows
  (5h session / 7d weekly) instead of overwriting a single "session" key
- **Models**: the catalog sync no longer erases its own output — deltas were
  measured against the previous run's writes (the second run cut `providers`
  from 20 entries to 5); one vote per provider in the modality tally, ETag
  restored from file on startup, and the worker thread dropped after the
  bundler rewrote its path into a module-not-found error
- **Executor**: CommandCode returns errors as a `type:"error"` event inside
  an HTTP 200 NDJSON stream — peek the first events before committing, abort
  and return a real 4xx/5xx so combo/account fallback triggers instead of
  streaming the error text as content
- **Search**: scope failure locks on the credential-fallback path — a failing
  search locked `modelLock___all` and took the shared glm key offline for
  chat as well; locks are now attributed to the connection's owner and
  scoped to `websearch:<provider>`
- **Providers**: connection tests get a 15s AbortSignal timeout instead of
  hanging and exhausting the browser socket pool; guard undefined provider
  names on the providers page
- **Antigravity**: sanitize competing-client branding via a config-driven
  rule table (Zed's Claude-agent prompt, opencode → antigravity) — upstream
  answers 429 Quota Exhausted. Applied in the executor so the shared
  openai-to-gemini translator leaves gemini/vertex/zed untouched
- **MiniMax**: preserve images on the sourceFormat-matched OpenAI transport
  — MiniMax-M3 resolved a Claude-shaped body posted to the OpenAI endpoint,
  silently dropping `image_url` blocks (#3418)
- **Claude**: decloak tool names in same-format streaming passthrough —
  OAuth-cloaked names (CLAUDE_TOOL_SUFFIX) leaked to the client and every
  tool call was rejected as unknown
- **Tools**: default a missing `tools[].type` to "custom" on Claude-format
  requests — strict Anthropic-compatible gateways (MiniMax) reject the
  request with 400 otherwise
- **Translator**: zai thinkingFormat sends the top-level `reasoning_effort`
  object GLM-5.2+ requires — every GLM-5.x request ran at the model default
  (max); gated on GLM-5.2+ since older GLM does not read it (#2721)
- **RTK**: system prompt injection matches each target wire format
  (Chat/Responses/Claude/Gemini/Kiro) and is exact-idempotent across retries,
  so distinct prompts sharing a long prefix are no longer collapsed (#3202).
  Also set the diagnostic before the silent null return on Responses
  translation failure so the panel is no longer blank
- **OpenCode**: route muse-spark through /zen/v1/responses (it 500s on
  chat/completions), normalizing the Chat fields the Responses API rejects
  and clamping max/ultra effort to xhigh
- **CLI**: install better-sqlite3 without build tools on Node 22+ (N-API
  13.0.3 ships per-platform prebuilds, `--ignore-scripts` skips the implicit
  node-gyp build); Node < 22 stays on 12.6.2, working installs untouched
- **CLI tools**: send the API key Codex actually reads —
  `[model_providers.9router.http_headers]` instead of auth.json (which left
  every request 401 and clobbered an existing ChatGPT login); subagent model
  moved to `agents.default_subagent_model`
- **OAuth**: refresh Cline tokens with the extension JSON contract
- **Dashboard**: clamp the API key mask length — keys shorter than 8 chars
  threw RangeError and crashed the media-provider detail page
- **UI**: wait for the Material Symbols font itself before revealing icons —
  `document.fonts.ready` resolved before the 4MB woff2 even started loading,
  leaving icons blank until a second load
2026-08-29 17:59:36 +07:00
decolua
2203cd8f2b test(translator): drop the golden url/header snapshot
The committed snapshot had drifted from the registry: five providers
mismatched on a plain checkout and alitp-intl was missing entirely.
Remove it so the suite regenerates from the current registry.
2026-08-28 18:19:53 +07:00
decolua
2fd99eae5d fix(session): read Claude Code session id from its request header
Claude Code carries the session in metadata.user_id, which the Responses
API translation drops before the executor resolves a cache session. The
request then fell through to the assistant-text hash and the per-connection
fallback, so one conversation was split across several prompt_cache_key
values and the upstream prefix cache kept missing.

Fall back to the x-claude-code-session-id header, which survives every
translation. The body stays authoritative when both are present.
2026-08-28 18:17:55 +07:00
Agung Gunawnan
df85e16d7a fix(providers): time out connection tests and guard undefined names
Apply a 15s AbortSignal timeout in fetchWithConnectionProxy when the
caller supplies none, so provider connection tests stop hanging and
exhausting the browser socket pool. Also make matchSearch return false
for falsy provider names instead of crashing the providers page.
2026-08-28 17:06:36 +07:00
Ahoora5678
dff648496c fix(antigravity): sanitize competing-client branding in system prompts
Antigravity flags requests whose system prompt identifies another vendor's
client and answers 429 Quota Exhausted. Move the existing Zed/Claude prompt
rewrite into a config-driven rule table and add case-preserving opencode ->
antigravity mapping.

Applied in the executor so only Antigravity requests are rewritten - the
shared openai-to-gemini translator also serves gemini, gemini-cli, vertex
and zed, which must not be touched.
2026-08-28 17:01:18 +07:00
Paulo Schuller
88676b3037 fix(oauth): refresh Cline tokens with extension JSON contract 2026-08-28 16:58:27 +07:00
fasilu
bb3cb43e09 fix(dashboard): clamp API key mask length for short keys
"•".repeat(apiKey.length - 8) threw RangeError when the key was
shorter than 8 chars, crashing the media-provider detail page.
2026-08-28 16:53:14 +07:00
Óscar Fonseca
4a371d1d9f fix(usage): preserve nested cached_tokens in canonicalizeUsage
buildUsage() only emits cache reads under prompt_tokens_details, so the
top-level-only read dropped the count for every Responses-format provider
(codex, grok-cli, ...), persisting cached_tokens: 0 and billing cache hits
at the full input rate. Mirror the cache_creation fallback already used
just above.
2026-08-28 16:46:05 +07:00
alfep
d91e8b85e0 feat(antigravity): add Gemini 3.7 Flash tiers to MITM defaultModels
Registry/pricing/CLI catalog already had gemini-3.7-flash-{high,medium,low}
but MITM_TOOLS.antigravity.defaultModels was missing them, so the tiers
never showed up in the dashboard model-mapping table.
2026-08-28 16:44:07 +07:00
snower
993c6eb469 feat(headroom): make the compression request timeout configurable
The 3000 ms timeout on /v1/compress was fixed, so busy or slow machines
timed out often and sent the LLM an inconsistently compressed body,
hurting prompt caching. Add a headroomTimeoutMs setting, thread it from
the chat handler down to compressWithHeadroom, expose it in the Token
Saver dashboard, and normalize invalid values back to the 3000 ms default.
2026-08-28 16:34:33 +07:00
turingcat
28d005772a fix(minimax): preserve images on matched OpenAI transport
Prefer the sourceFormat-matched runtime transport over a model's
declared targetFormat when both apply. MiniMax-M3 previously resolved
to a Claude-shaped body while being posted to the already-selected
OpenAI endpoint, silently dropping image_url blocks from OpenAI
clients. Fixes #3418.
2026-08-28 16:34:05 +07:00
fasilu
2a9213c5bd feat(antigravity): map image size to aspect-ratio model suffix
Resolve body.size through sizeToAspectRatio and append the ratio as a
-WxH suffix so the executor's parseImageConfig picks it up. Also fall
back to gemini-3.1-flash-image when a non-image model reaches the
image handler.
2026-08-28 16:33:23 +07:00
decolua
ec6692808b fix(search): scope failure locks so search cannot take chat offline
Two problems on the credentialFallback path, where a search provider
borrows a chat provider's connection:

- the lock was attributed to the search provider id, but the connection
  belongs to the chat provider, so markAccountUnavailable looked it up
  under the wrong provider and read a stale backoffLevel
- with no model argument the lock key is `modelLock___all`, which
  isModelLockActive treats as blocking every model — one failing search
  would have taken the shared glm key offline for chat as well

Attribute the lock to the provider that owns the connection, and scope
it to `websearch:<provider>`, passed to getProviderCredentials too so
the lock is read back under the same key.
2026-08-28 16:18:46 +07:00
Amir Seify
e5a13c3ab7 feat(usage): show Zed plan quota on the dashboard
Add a Zed usage handler so connected Zed accounts appear on
/dashboard/quota. Reads GET /client/users/me for plan, edit
predictions, optional hosted model requests and billing-cycle reset.

Render unlimited rows as "N used · Unlimited" instead of 0 / ∞, and
surface overdue-invoice / token-billing messages.
2026-08-28 16:16:01 +07:00
Bertho Joris
67d9182e1a fix(executor): handle CommandCode in-stream errors for combo and account fallback
CommandCode returns errors as a type:"error" event inside an HTTP 200
NDJSON stream instead of a non-200 status, so the existing combo/account
fallback logic (keyed off response.status) never triggered and the error
text was streamed to the client as if it were content.

Peek the first NDJSON events before committing to a stream; on a
type:"error" event, abort and return a proper 4xx/5xx Response instead.
Normal streams are replayed losslessly (buffered prefix + rest of the
stream) through the existing translator, so the happy path is unchanged.
Add CommandCodeExecutor.parseError() so parseUpstreamError() can extract
a clean message/status from the synthesized error body.
2026-08-28 16:15:19 +07:00
decolua
9dbdca0e5e refactor(search): fold zai-search into the glm provider
The separate zai-search entry showed "No connections" on the web search
page because credentials live on the `glm` connection, not on it. Every
other provider that does both chat and search (antigravity, kimi, xai,
gemini) declares webSearch on the provider itself, so do the same here.

- glm gains serviceKinds ["llm", "webSearch"] and the MCP searchConfig
- the request builder / normalizer move from "zai-search" to "glm"
- drop the zai-search registry entry and its svg logo, which also
  removes the only need for svg logo support in getProviderIconSrc

ollama-search keeps its own entry and credentialFallback: its search
endpoint is unrelated to the ollama chat transport.
2026-08-28 16:12:04 +07:00
decolua
5a86f6a8d2 feat(search): add ollama-search and zai-search with credential fallback
Register two web search providers that reuse an existing chat provider's
API key instead of requiring their own connection:

- ollama-search (POST ollama.com/api/web_search) reuses the `ollama` key
- zai-search (POST api.z.ai MCP web_search_prime) reuses the `glm` key

A new `credentialFallback` registry field drives this: when a search
provider has no connection of its own, the search handler falls back to
the linked chat provider's credentials.

Also teach getProviderIconSrc to serve .svg logos for providers that
ship vector art.
2026-08-28 16:04:45 +07:00
huohua-dev
eb312bd470 fix(claude): decloak tool names in same-format streaming passthrough
translateResponse() short-circuited untouched on claude->claude streaming,
so OAuth-cloaked tool names (CLAUDE_TOOL_SUFFIX) leaked to the client and
every tool call was rejected as unknown. Add decloakStreamChunk(), the
streaming counterpart of decloakToolNames(), and call it on the same-format
path using the already-plumbed state.toolNameMap.
2026-08-28 15:40:40 +07:00
Daniel Gonçalves Araujo
fcfcced4ab fix(usage): support CREDIT_LIMIT and multi-interval GLM quotas
GLM quota parsing only accepted TOKENS_LIMIT and wrote every limit to a
single "session" key, so credit-based plans showed nothing and later
intervals overwrote earlier ones. Accept CREDIT_LIMIT too and derive the
quota key from the limit unit (5h session, 7d weekly, tokens, custom).
Moves the parser into its own usage/glm.js, re-exported from misc.js.
2026-08-28 15:35:23 +07:00
qingyong
56a40765e9 fix(translator): zai thinkingFormat sends reasoning.effort object
Z.ai / GLM-5.2+ require a top-level reasoning_effort (low/high/max)
alongside thinking:{type:"enabled"} to control reasoning depth; the zai
branch previously only set thinking and dropped reasoning_effort, so every
GLM-5.x request ran at the model default (max). Gate the field behind
GLM-5.2+ (thinkingEffortSupported in capabilities.js) since older GLM
(4.x, 5.0, 5.1, 5-turbo, 5v-turbo) do not read it, and map client levels
to the exact low/high/max values z.ai accepts.

extractThinking now checks reasoning_effort/reasoning.effort before the
thinking object so a client-supplied effort is not overwritten by
thinking:{type:"enabled"} mapping to mode:auto.

Fixes #2721
2026-08-28 12:32:41 +07:00
KunN-21
cadef6c4ff fix(rtk): make system prompt injection format-safe and idempotent
Caveman/Ponytail injection now matches each target wire format instead of
assuming an OpenAI-shaped body:

- Chat arrays append a text block; Responses arrays append input_text and
  create typed message items
- Claude inserts before the final cache-control block; Gemini preserves the
  snake/camel systemInstruction wrapper
- Kiro updates systemPrompt and its mirrored first-user prefix atomically,
  rolling back if the pair fails to converge
- Format label decides Claude/Gemini before the wire-shape sniff, since their
  bodies also carry messages[]/contents[] and Anthropic rejects a "system"
  role inside messages[]
- Delimiter-aware dedup makes injection exact-idempotent across retries, so
  distinct prompts sharing a long prefix are no longer collapsed
- Every write is fail-open on frozen or proxied bodies

Saver order and X-9Router-Token-Saver: off behavior are unchanged.

Fixes #3202.
2026-08-28 11:47:56 +07:00
anojndr
ab044e6d6d fix(opencode): route Muse Spark through the Responses API
muse-spark-1.2-contributor-free returned HTTP 500 on /zen/v1/chat/completions.
The model is only served by /zen/v1/responses, so route it there via a per-model
targetFormat and normalize the Chat fields the Responses API rejects
(max_tokens -> max_output_tokens, reasoning_effort -> reasoning{effort,summary}),
clamping max/ultra down to the highest effort the model accepts (xhigh).

Routing stays per-model: the other free models (big-pickle, hy3-free, mimo,
nemotron, laguna) are not served by /responses and keep /chat/completions.
2026-08-28 11:33:14 +07:00
decolua
14401c433c fix(ui): wait for the icon font itself before revealing Material Symbols
`document.fonts.ready` resolved before the 4MB Material Symbols woff2 even
started loading — it runs in <head>, ahead of any element that would trigger
the lazy fetch. The `fonts-loaded` class landed early, so icons rendered
blank until a second load served the font from disk cache.

Load the face explicitly and swap visibility for opacity, with a 3s fallback
so icons never stay hidden if the font fails.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-28 11:32:28 +07:00
warelik
e08ac6dada fix(tools): default Claude tool type when missing
Strict Anthropic-compatible gateways (e.g. MiniMax) reject Claude-format
requests with HTTP 400 when tools[].type is missing. Normalize each
missing/falsy tools[].type to "custom" before dispatch when the final
request format is Claude. Built-in tool types (computer_use, bash,
web_search_*) are passed through untouched.
2026-08-28 11:16:44 +07:00
Nguyen Thanh Dat
f9d82c6575 fix(stream): parse the trailing NDJSON line an Ollama stream leaves behind
createSSEStream splits on "\n" and keeps the remainder, which only flush()
parses. That call omitted targetFormat, so parseSSELine required a "data: "
prefix and dropped whatever an NDJSON provider left without a closing
newline. The !parsed.done guard compounded it: the SSE sentinel and an
Ollama final chunk both carry done:true, but the latter is the real last
chunk holding done_reason and the token counts.

Pass targetFormat and scope the sentinel check to formats that emit one, so
the tail reaches the translator. Accumulate its usage into state the same way
the transform loop does, so finalizeStream logs those tokens instead of null.
2026-08-27 20:53:02 +07:00
decolua
2f17352cc2 feat(search): add Antigravity as a web search provider
Route POST /v1/search with provider "antigravity" through Google Search
grounding on v1internal:generateContent, using the existing Antigravity
OAuth account pool. Grounding chunks become citations with the grounded
sentence as snippet and its surrounding answer text as content.

Upstream repeats a source across chunks, so citations are keyed by URL
and their snippets merged. A missing projectId is reported up front —
upstream answers a fabricated or absent project with a misleading
"no valid license" 403.

Based on the approach in #3437 by @Nautilaceae.
2026-08-27 20:48:43 +07:00
vianhanif
90a0005845 fix(cli): install better-sqlite3 without build tools on Node 22+
The runtime hook pinned better-sqlite3 12.6.2, whose prebuilds stop at
Node ABI 141 — on Node 26 the install fell back to a node-gyp source
build and failed on machines without build tools, silently degrading to
the sql.js fallback.

Node >= 22 now installs 13.0.3, which is N-API and ships per-platform
prebuilds inside the package. Two things were needed to make that
actually work:

- npm injects an implicit `node-gyp rebuild` for any package shipping a
  binding.gyp, so the install still demanded build tools; `--ignore-scripts`
  skips it and uses the bundled prebuild as-is.
- the binary check only looked at build/Release, which 13.x no longer
  creates, so every start re-ran npm install; it now also accepts
  prebuilds/<platform>-<arch>.node.

Node < 22 stays on 12.6.2 (13.x requires Node >= 22), and an existing
working install is left untouched either way.
2026-08-27 20:28:26 +07:00
Fábio A.
e79ae6e7c5 i18n(pt-BR): expand translation to 1132 terms
Add 144 missing pt-BR strings covering Usage, Endpoint & Key security
notices, 9Remote, Media Providers, Proxy Pools, Combo & Vision Adapter,
Token Saver, Agent Skills and Quota Tracker.
2026-08-27 20:06:10 +07:00
decolua
a68ada1c83 feat(cli-tools): share endpoint presets across every tool card
Each card kept its own copy of the localStorage preset logic inside
BaseUrlSelect, so an endpoint saved on one card was invisible to the
others until a reload, and a URL typed into the custom field was
forgotten the moment the card collapsed.

Move the store into cliEndpointPresets.js and publish a change event
so open cards resync live. Applying settings now remembers the
endpoint unless it matches a built-in option, and each card passes
its configured URL as currentUrl so BaseUrlSelect can preselect the
matching preset instead of always falling back to 127.0.0.1.
Deleting a preset falls back to the first real option rather than
clearing the field.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-27 18:54:44 +07:00
decolua
c4af43faa3 fix(stream): stop logging a disconnect for every completed Responses call
Responses-API clients (codex, droid) close the socket on
response.completed because the protocol has no [DONE] sentinel, so
every successful request printed "⚡ DISCONNECT: ResponseAborted"
after its own "📊 done" line. Keep the dbg("CTRL", …) trace and drop
the console line; ABORTED and ERROR still print.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-27 18:52:50 +07:00
decolua
d7f7d70dd5 fix(stream): record usage when a client closes on the terminal event
The Responses API has no [DONE] sentinel, so codex closes the socket
as soon as response.completed arrives. That cancels the reader before
flush() runs — and flush() held every usage side effect, so a fully
successful request logged nothing: no 📊 done line, no token stats,
no request detail.

Extract that tail into a once-guarded finalizeStream() and also call
it right after the terminal event is forwarded, in both passthrough
and translate mode. flush() still calls it; the guard makes the
second call a no-op. Streams that end normally are unaffected, and a
terminal event carrying no usage falls through to the existing
estimate/null path rather than blocking.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-27 18:52:42 +07:00
decolua
9c45b27cd7 fix(cli-tools): send the API key Codex actually reads
Codex only authenticates a custom model provider from env_key,
http_headers, env_http_headers or a token command — auth.json is
read solely by the built-in openai provider. Writing OPENAI_API_KEY
there left every request unauthenticated (401 Missing API key) while
clobbering an existing ChatGPT login.

Put the key in [model_providers.9router.http_headers] instead, and
drop the auth.json write. Also move the subagent model to the
agents.default_subagent_model scalar: agents.<role> now declares a
custom role and requires a description, so the old [agents.subagent]
table was discarded with a startup warning. DELETE still clears
auth.json to repair machines configured by the previous version.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-27 18:52:32 +07:00
decolua
e6f5724b4b fix(models): stop the catalog sync from erasing its own output
collectEntries() computed each model's "current" capabilities with the
previous catalog still installed, so every delta was measured against the
last one. An upstream value that still agreed with what we had written
looked like no change and was dropped: the second run cut `providers`
from 20 entries to 5, taking glm-5.3's 1M context correction with it.

The baseline has to be the hand-written tables alone, so the reader is
detached for the snapshot and restored in a finally — a mid-sync failure
must not leave capabilities.js without it.

Two smaller corrections:

- One vote per provider in the modality tally. Ids that normalize to the
  same model (claude-opus-4-thinking:1024, :8192, :32768 …) were each
  counted, giving nano-gpt five votes where other gateways had one. No
  model's result actually flipped — the variants agree with each other —
  but the majority rule only means something if the denominator does.
- Restore the etag from the file on startup. It lived only in module
  state, so every restart re-downloaded 4.3MB to be told nothing changed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-27 18:44:30 +07:00
decolua
d01724556a fix(models): drop the worker thread from the catalog sync
The worker resolved its own path through import.meta.url, which the
bundler rewrites — so the running server looked for the file at a path
that does not exist there:

  [modelCatalog] sync failed: Cannot find module
  '/Users/Working/router4/9router/src/lib/modelCatalog/worker.js'

It was guarding against a 23ms JSON.parse that runs once a day, 60s after
boot. Inlining it into sync.js costs that 23ms on an otherwise idle tick
and removes both the failure mode and a whole file.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-27 18:39:44 +07:00
DarRahman
40eed18688 feat(usage): track GPT-5.3-Codex-Spark quota windows
Extract Spark rate limit windows from the Codex usage response and expose
them as spark_session/spark_weekly quotas, reusing the existing prefix
mechanism. Map codex quota types to readable dashboard labels.

Fixes #3431
2026-08-27 17:56:26 +07:00
decolua
0532f00d84 feat(models): refresh model capabilities from models.dev in the background
Capability tables are hand-maintained, so a model gains vision or a wider
context only when someone notices and edits the file. This adds a daily
sync that fills the gap for models already in the registry.

How it decides:

- Modalities (vision/pdf/audio/video) belong to the MODEL — every gateway
  serving glm-5.3-flash serves the same weights — so they are keyed by
  model id and shared. A majority of sources must declare one, which keeps
  out lone mis-declarations: minimax-m2.5 (1 of 45), glm-4.7 (1 of 44) and
  gpt-oss-120b (2 of 76) are text-only despite a reseller claiming vision.
- Context/output limits belong to the GATEWAY — each truncates differently
  (glm-5 ships as 202752/16384 on one host and 204800/131072 on another) —
  so they are keyed by provider + model and only the matching provider's
  own numbers are trusted.

Both layers are strictly additive and sit BELOW the hand-written tables,
which short-circuit first. A capability already true stays true.

Mechanics: worker thread (the 4MB parse would block the loop ~20ms),
ETag so an unchanged catalog costs one empty request, 60s startup delay,
30min backoff on failure, MODEL_CATALOG_SYNC=off to disable. Only the
~57KB delta is kept; lookups cost ~0.1us via an mtime-guarded cache.

capabilities.js is bundled into the browser through useModelCaps, so it
cannot import node:fs — the server injects the reader via
setCatalogSource() from instrumentation.

visionPatterns.js is the last resort: a model nobody has catalogued yet
still accepts images when its id says so (qwen3-vl-plus, glm-4.6v, llava),
with image-generation and embedding ids excluded.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-27 17:53:53 +07:00
decolua
9c650e1d54 feat(models): add GLM-5.3-Flash, DeepSeek V4 Vision, Grok 4.5/4.6
Vendors shipped four multimodal models the registry did not carry:

- glm-5.3-flash — z.ai's first natively multimodal GLM-5, 1M context,
  image + video + pdf input (glm, glm-cn, opencode-go)
- deepseek-v4-flash-vision-exp — image input at V4-Flash text parity,
  1M context / 384k output (deepseek, opencode-go)
- grok-4.6, grok-4.5 — 500k context; 4.6 has no text output limit (xai)

Capabilities needed hand entries because the existing globs mis-matched:
*glm-5* and *deepseek-v4* carry no vision, and *grok-4* would have capped
grok-4.6 at 256k instead of 500k. The grok-4.6 pattern sits above the
generic *grok-4* so it wins the first-match lookup.

Also corrects glm-4.6v / glm-4.5v, which were missing video input and
declared no maxOutput, and backfills glm-4.6v on glm-cn — zhipuai serves
it and the sibling provider already listed it.

tests/unit/opencode-go-models.test.js pins the opencode-go model list, so
its expected array moves with the registry.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-27 17:52:01 +07:00
Nim
1a3db1efae feat(antigravity): quota-aware routing with reset-aware fallback
On a 409/429 from Antigravity, fetch live quota to learn the exact
per-model resetAt instead of guessing a backoff, then skip only the
exhausted account/model pair until that time.

- antigravityQuota.js: in-memory quota cache, coalesced concurrent
  refreshes, 30s throttle per connection (applied to failures too),
  keeps known cache when upstream returns 401/403 error payloads
- auth.js: pre-filter exhausted account/model pairs; report the
  earliest quota reset when every account is blocked; skip the
  30-minute cooldown cap so the upstream resetAt is not truncated
- chat.js: antigravity 409/429 falls back on the RAM cache only, no
  persistent modelLock_* for this path
- Logs identify accounts by id prefix, never email or name

Closes #3561
2026-08-27 16:18:08 +07:00
kriptoburak
f0a6d35818 feat(search): add Xquik as an X search provider
Xquik needs a GET request with x-api-key auth and a tweets envelope
normalizer, neither of which the generic search fallback provides. Adds a
dedicated request builder and normalizer, cursor pagination passthrough,
result-based credit usage reporting, and a validateUrl probe so key
validation hits the no-charge credits endpoint.
2026-08-27 16:10:00 +07:00
vianhanif
548e32aacf fix(rtk): set diagnostic before silent null return on Responses translation failure
The openai-responses branch of compressWithHeadroom returned null without
recording a reason, leaving the diagnostics panel blank and making Codex
translation failures indistinguishable from a successful compression.
2026-08-27 15:52:46 +07:00
ariesho2903
abb20d9f39 feat(dashboard): bulk import Grok CLI accounts from JSON
Add a "Bulk Add" flow for the grok-cli provider, mirroring the existing
codex one: paste a JSON array/object or drag-drop multiple .json files,
then create all OAuth connections in a single call.

- BulkImportGrokCliModal: flexible JSON parsing (array, single object,
  {accounts:[...]}, concatenated objects) + multi-file upload
- POST /api/oauth/grok-cli/bulk-import: serial createProviderConnection,
  snake_case/camelCase token fields, email backfilled from id_token or
  access_token, authMethod "device_code" to match the login flow
2026-08-27 15:49:35 +07:00
86112cee6d feat(combos): drag-and-drop reorder with manual sortOrder
- schema.js: add sortOrder REAL column to combos, bump SCHEMA_VERSION 4->5
- combosRepo: persist sortOrder, append new combos to end, add reorderCombos() atomic reorder
- export/import DB: include sortOrder for round-trip
- API: PUT /api/combos/order accepts { ids: string[] } and persists new order
- UI: dnd-kit DndContext + SortableContext wrap flat list; drag handle (grip icon) on each ComboCard; optimistic local reorder with revert on failure

Grouped-by-tag view keeps server-side order; reorder is only available in the unfiltered flat list.
2026-08-27 14:21:25 +07:00
5c6048759c fix(db): add missing apiKeys columns + restore ApiExplorerModal export
- schema.js: add createdAt/allowedModels to apiKeys, bump SCHEMA_VERSION 3->4
- shared/components/index.js: restore ApiExplorerModal barrel export (was replaced by TagInput)
- combos: tag support, repo + api + page updates
2026-08-27 13:58:23 +07:00
d0f202a75d feat(commandcode): bump version header and expand model catalog 2026-08-27 10:34:09 +07:00
f0adfb205a feat(dashboard): per-key model restrictions, pin header routing, combo side-panel picker
- Endpoint: per-API-key model allowlist (schema v3) enforced on chat (403)
  and /v1/models; Full-access toggle + multi-select picker in Keys UI.
- Providers: honor x-connection-id in /v1/chat/completions — pinned requests
  no longer rotate to another account on failure.
- Providers: strategy saves merge into stored enabled:false override;
  Test All groups match grid sections; 1-by-1 skips disabled connections.
- Dashboard: provider-card toggle syncs from server on failure; grid toggles
  always visible; connection rows get clear-✕ for stale error banners.
- Combo editor: on desktop (xl+) the Add-Model picker opens as a floating
  side panel beside the untouched combo popup instead of stacking on top;
  mobile keeps the full-screen overlay.
- Long API-key overflow fixed in key rows + provider model sections.
2026-08-27 09:26:17 +07:00
1d56e2dbc5 chore(deps): bump dependencies 2026-08-22 14:34:02 +07:00
55f10c11e5 fix(dashboard): label usage-by-provider column as Provider / API Key 2026-08-22 14:34:02 +07:00
eedad6c5ea fix(combos): show effective strategy vs global default; keep explicit fallback override 2026-08-22 14:34:02 +07:00
99752a397c fix(usage): commandcode monthly total = consumed + remaining credits 2026-08-22 14:33:53 +07:00
bb8d67ba9c feat(caps): user-registered models trust upstream vision instead of stripping media 2026-08-22 14:33:53 +07:00
144dda2ac2 fix(routing): fall through to compatible node when built-in alias has no credentials 2026-08-22 14:33:53 +07:00
bc9719fac7 feat(dashboard): bulk enable/disable selected API keys on provider page 2026-08-22 14:33:20 +07:00
616 changed files with 94267 additions and 6195 deletions

Binary file not shown.

After

Width:  |  Height:  |  Size: 38 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 48 KiB

View File

@@ -5,22 +5,164 @@ on:
tags:
- "v*"
workflow_dispatch:
inputs:
release_tag:
description: "Existing vX.Y.Z tag to publish"
required: true
type: string
promote_latest:
description: "Promote this republish to latest"
required: false
default: false
type: boolean
# Keep every release in one FIFO queue. A per-tag group would still allow an
# older release to finish after a newer release and move latest backwards.
concurrency:
group: docker-publish-${{ github.repository }}
cancel-in-progress: false
queue: max
env:
GHCR_IMAGE: ghcr.io/${{ github.repository }}
DOCKERHUB_IMAGE: decolua/9router
jobs:
build-and-push:
prepare:
name: Validate release
runs-on: ubuntu-latest
timeout-minutes: 10
permissions:
contents: read
outputs:
tag: ${{ steps.release.outputs.tag }}
version: ${{ steps.release.outputs.version }}
commit: ${{ steps.release.outputs.commit }}
publish_dockerhub: ${{ steps.release.outputs.publish_dockerhub }}
promote_latest: ${{ steps.release.outputs.promote_latest }}
ghcr_image: ${{ steps.release.outputs.ghcr_image }}
steps:
- name: Check out release tag
uses: actions/checkout@v4
with:
ref: ${{ inputs.release_tag || github.ref_name }}
fetch-depth: 1
- name: Validate tag and package versions
id: release
env:
RELEASE_TAG: ${{ inputs.release_tag || github.ref_name }}
REPOSITORY: ${{ github.repository }}
EVENT_NAME: ${{ github.event_name }}
PROMOTE_LATEST_INPUT: ${{ inputs.promote_latest && 'true' || 'false' }}
run: |
node <<'NODE'
const fs = require("fs");
const { execFileSync } = require("child_process");
const tag = process.env.RELEASE_TAG || "";
const match = /^v((?:0|[1-9]\d*)\.(?:0|[1-9]\d*)\.(?:0|[1-9]\d*)(?:-[0-9A-Za-z-]+(?:\.[0-9A-Za-z-]+)*)?)$/.exec(tag);
if (tag.includes("+")) {
console.error(`Build metadata is not supported in Docker release tags: ${tag}`);
process.exit(1);
}
if (!match) {
console.error(`Expected a Docker-safe semver tag like v0.5.81 or v0.5.81-rc.1, received: ${tag || "<empty>"}`);
process.exit(1);
}
const version = match[1];
if (version.length > 128 || !/^[A-Za-z0-9_][A-Za-z0-9_.-]{0,127}$/.test(version)) {
console.error(`Version is not a valid Docker tag: ${version}`);
process.exit(1);
}
const prerelease = version.includes("-")
? version.slice(version.indexOf("-") + 1).split(".")
: [];
for (const identifier of prerelease) {
if (/^\d+$/.test(identifier) && identifier.length > 1 && identifier.startsWith("0")) {
console.error(`Numeric prerelease identifiers cannot contain leading zeroes: ${identifier}`);
process.exit(1);
}
}
const rootVersion = require("./package.json").version;
const cliVersion = require("./cli/package.json").version;
if (rootVersion !== version) {
console.error(`package.json version ${rootVersion} does not match tag ${tag}`);
process.exit(1);
}
if (cliVersion !== version) {
console.error(`cli/package.json version ${cliVersion} does not match tag ${tag}`);
process.exit(1);
}
const commit = execFileSync("git", ["rev-parse", "HEAD"], { encoding: "utf8" }).trim();
const publishDockerHub = process.env.REPOSITORY === "decolua/9router";
const ghcrImage = `ghcr.io/${process.env.REPOSITORY.toLowerCase()}`;
const isPrerelease = version.includes("-");
const promoteLatest = (process.env.EVENT_NAME === "push" && !isPrerelease)
|| process.env.PROMOTE_LATEST_INPUT === "true";
const output = process.env.GITHUB_OUTPUT;
fs.appendFileSync(output, `tag=${tag}\n`);
fs.appendFileSync(output, `version=${version}\n`);
fs.appendFileSync(output, `commit=${commit}\n`);
fs.appendFileSync(output, `publish_dockerhub=${publishDockerHub}\n`);
fs.appendFileSync(output, `promote_latest=${promoteLatest}\n`);
fs.appendFileSync(output, `ghcr_image=${ghcrImage}\n`);
console.log(`Validated ${tag} at ${commit}`);
console.log(`latest promotion: ${promoteLatest ? "enabled" : "disabled"}`);
NODE
build:
name: Build ${{ matrix.platform }}
needs: prepare
runs-on: ${{ matrix.runner }}
timeout-minutes: 60
env:
GHCR_IMAGE: ${{ needs.prepare.outputs.ghcr_image }}
strategy:
fail-fast: false
matrix:
include:
- platform: linux/amd64
suffix: amd64
runner: ubuntu-24.04
- platform: linux/arm64
suffix: arm64
runner: ubuntu-24.04-arm
permissions:
contents: read
packages: write
steps:
- uses: actions/checkout@v4
- name: Check out release source at validated commit
uses: actions/checkout@v4
with:
ref: ${{ needs.prepare.outputs.commit }}
path: source
fetch-depth: 1
- uses: docker/setup-buildx-action@v3
- name: Check out publishing Dockerfile
uses: actions/checkout@v4
with:
ref: ${{ github.workflow_sha }}
path: workflow
sparse-checkout: |
Dockerfile
fetch-depth: 1
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@v3
- name: Log in to GHCR
uses: docker/login-action@v3
@@ -29,32 +171,267 @@ jobs:
username: ${{ github.actor }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: Build and push platform image by digest
id: build
uses: docker/build-push-action@v6
with:
context: source
file: workflow/Dockerfile
platforms: ${{ matrix.platform }}
outputs: type=image,name=${{ env.GHCR_IMAGE }},push-by-digest=true,name-canonical=true,push=true
build-args: |
APP_VERSION=${{ needs.prepare.outputs.version }}
ALPINE_MIRROR=${{ vars.ALPINE_MIRROR || 'dl-cdn.alpinelinux.org' }}
NPM_REGISTRY=${{ vars.NPM_REGISTRY || 'https://registry.npmjs.org/' }}
labels: |
org.opencontainers.image.source=https://github.com/${{ github.repository }}
org.opencontainers.image.revision=${{ needs.prepare.outputs.commit }}
org.opencontainers.image.version=${{ needs.prepare.outputs.version }}
cache-from: type=gha,scope=9router-${{ matrix.suffix }}
cache-to: type=gha,mode=max,scope=9router-${{ matrix.suffix }}
provenance: false
sbom: false
- name: Smoke-test platform image before publishing digest artifact
env:
GHCR_IMAGE: ${{ env.GHCR_IMAGE }}
IMAGE_DIGEST: ${{ steps.build.outputs.digest }}
PLATFORM: ${{ matrix.platform }}
run: |
set -Eeuo pipefail
[[ "$IMAGE_DIGEST" =~ ^sha256:[0-9a-f]{64}$ ]]
container="9router-platform-smoke-${GITHUB_RUN_ID}-${{ matrix.suffix }}"
trap 'docker rm -f "$container" >/dev/null 2>&1 || true' EXIT
docker run --detach \
--name "$container" \
--platform "$PLATFORM" \
--publish 20128:20128 \
"${GHCR_IMAGE}@${IMAGE_DIGEST}"
for attempt in {1..45}; do
if curl --fail --silent --show-error http://127.0.0.1:20128/api/health; then
echo "${PLATFORM} health check passed"
exit 0
fi
if (( attempt % 5 == 0 )); then
echo "Waiting for ${PLATFORM} health check (${attempt}/45)" >&2
fi
sleep 2
done
echo "${PLATFORM} health check failed; container logs follow:" >&2
docker logs "$container" || true
exit 1
- name: Save image digest
env:
IMAGE_DIGEST: ${{ steps.build.outputs.digest }}
run: |
set -euo pipefail
test -n "$IMAGE_DIGEST"
mkdir -p "$RUNNER_TEMP/digests"
printf '%s\n' "$IMAGE_DIGEST" > "$RUNNER_TEMP/digests/${{ matrix.suffix }}.txt"
- name: Upload image digest
uses: actions/upload-artifact@v4
with:
name: digests-${{ matrix.suffix }}
path: ${{ runner.temp }}/digests/${{ matrix.suffix }}.txt
if-no-files-found: error
publish:
name: Publish and verify manifest
needs:
- prepare
- build
runs-on: ubuntu-latest
timeout-minutes: 30
env:
GHCR_IMAGE: ${{ needs.prepare.outputs.ghcr_image }}
permissions:
contents: read
packages: write
steps:
- name: Download platform digests
uses: actions/download-artifact@v4
with:
pattern: digests-*
path: ${{ runner.temp }}/digests
merge-multiple: true
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@v3
- name: Log in to GHCR
uses: docker/login-action@v3
with:
registry: ghcr.io
username: ${{ github.actor }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: Create and verify version manifest
env:
GHCR_IMAGE: ${{ env.GHCR_IMAGE }}
VERSION: ${{ needs.prepare.outputs.version }}
run: |
set -euo pipefail
shopt -s nullglob
digest_files=("$RUNNER_TEMP"/digests/*.txt)
if [[ "${#digest_files[@]}" -ne 2 ]]; then
echo "Expected two platform digests, found ${#digest_files[@]}" >&2
exit 1
fi
sources=()
for digest_file in "${digest_files[@]}"; do
digest="$(tr -d '\n' < "$digest_file")"
if [[ ! "$digest" =~ ^sha256:[0-9a-f]{64}$ ]]; then
echo "Invalid image digest in $digest_file: $digest" >&2
exit 1
fi
sources+=("${GHCR_IMAGE}@${digest}")
done
docker buildx imagetools create \
--tag "${GHCR_IMAGE}:${VERSION}" \
"${sources[@]}"
docker buildx imagetools inspect "${GHCR_IMAGE}:${VERSION}" | tee "$RUNNER_TEMP/version-manifest.txt"
docker buildx imagetools inspect --raw "${GHCR_IMAGE}:${VERSION}" > "$RUNNER_TEMP/version-manifest.json"
expected=$'linux/amd64\nlinux/arm64'
actual="$(jq -r '[.manifests[] | select(.platform != null and .platform.os != null and .platform.architecture != null) | "\(.platform.os)/\(.platform.architecture)"] | sort | .[]' "$RUNNER_TEMP/version-manifest.json")"
if [[ "$actual" != "$expected" ]]; then
echo "Version manifest platforms do not match exactly:" >&2
printf '%s\n' "$actual" >&2
exit 1
fi
- name: Smoke-test resolved version manifest
env:
GHCR_IMAGE: ${{ env.GHCR_IMAGE }}
VERSION: ${{ needs.prepare.outputs.version }}
run: |
set -Eeuo pipefail
container="9router-manifest-smoke-${GITHUB_RUN_ID}"
trap 'docker rm -f "$container" >/dev/null 2>&1 || true' EXIT
docker run --detach \
--name "$container" \
--platform linux/amd64 \
--publish 20128:20128 \
"${GHCR_IMAGE}:${VERSION}"
for attempt in {1..30}; do
if curl --fail --silent --show-error http://127.0.0.1:20128/api/health; then
echo "Resolved version manifest health check passed"
exit 0
fi
if (( attempt % 5 == 0 )); then
echo "Waiting for resolved manifest health check (${attempt}/30)" >&2
fi
sleep 2
done
echo "Resolved version manifest health check failed; container logs follow:" >&2
docker logs "$container" || true
exit 1
- name: Log in to Docker Hub
if: needs.prepare.outputs.publish_dockerhub == 'true'
uses: docker/login-action@v3
with:
username: ${{ secrets.DOCKERHUB_USERNAME }}
password: ${{ secrets.DOCKERHUB_TOKEN }}
- name: Extract metadata
id: meta
uses: docker/metadata-action@v5
with:
images: |
${{ env.GHCR_IMAGE }}
${{ env.DOCKERHUB_IMAGE }}
tags: |
type=semver,pattern={{version}}
type=raw,value=latest,enable={{is_default_branch}}
- name: Publish version image to Docker Hub
if: needs.prepare.outputs.publish_dockerhub == 'true'
env:
DOCKERHUB_IMAGE: ${{ env.DOCKERHUB_IMAGE }}
GHCR_IMAGE: ${{ env.GHCR_IMAGE }}
VERSION: ${{ needs.prepare.outputs.version }}
run: |
set -euo pipefail
docker buildx imagetools create \
--tag "${DOCKERHUB_IMAGE}:${VERSION}" \
"${GHCR_IMAGE}:${VERSION}"
- name: Build and push
uses: docker/build-push-action@v6
with:
context: .
push: true
tags: ${{ steps.meta.outputs.tags }}
labels: ${{ steps.meta.outputs.labels }}
cache-from: type=registry,ref=${{ env.GHCR_IMAGE }}:buildcache
cache-to: type=registry,ref=${{ env.GHCR_IMAGE }}:buildcache,mode=max
platforms: linux/amd64,linux/arm64
provenance: false
sbom: false
docker buildx imagetools inspect "${DOCKERHUB_IMAGE}:${VERSION}" | tee "$RUNNER_TEMP/dockerhub-version-manifest.txt"
docker buildx imagetools inspect --raw "${DOCKERHUB_IMAGE}:${VERSION}" > "$RUNNER_TEMP/dockerhub-version-manifest.json"
expected=$'linux/amd64\nlinux/arm64'
actual="$(jq -r '[.manifests[] | select(.platform != null and .platform.os != null and .platform.architecture != null) | "\(.platform.os)/\(.platform.architecture)"] | sort | .[]' "$RUNNER_TEMP/dockerhub-version-manifest.json")"
if [[ "$actual" != "$expected" ]]; then
echo "Docker Hub version manifest platforms do not match exactly:" >&2
printf '%s\n' "$actual" >&2
exit 1
fi
- name: Record latest promotion policy
env:
PROMOTE_LATEST: ${{ needs.prepare.outputs.promote_latest }}
VERSION: ${{ needs.prepare.outputs.version }}
run: |
if [[ "$PROMOTE_LATEST" == "true" ]]; then
echo "### Latest promotion" >> "$GITHUB_STEP_SUMMARY"
echo "- Policy: promote \`latest\` after the verified ${VERSION} manifest." >> "$GITHUB_STEP_SUMMARY"
else
echo "### Latest promotion" >> "$GITHUB_STEP_SUMMARY"
echo "- Policy: leave \`latest\` unchanged; this is a numbered-tag-only manual republish." >> "$GITHUB_STEP_SUMMARY"
fi
- name: Promote verified version to latest
if: needs.prepare.outputs.promote_latest == 'true'
env:
DOCKERHUB_IMAGE: ${{ env.DOCKERHUB_IMAGE }}
GHCR_IMAGE: ${{ env.GHCR_IMAGE }}
PUBLISH_DOCKERHUB: ${{ needs.prepare.outputs.publish_dockerhub }}
VERSION: ${{ needs.prepare.outputs.version }}
run: |
set -euo pipefail
docker buildx imagetools create \
--tag "${GHCR_IMAGE}:latest" \
"${GHCR_IMAGE}:${VERSION}"
if [[ "$PUBLISH_DOCKERHUB" == "true" ]]; then
docker buildx imagetools create \
--tag "${DOCKERHUB_IMAGE}:latest" \
"${GHCR_IMAGE}:${VERSION}"
fi
docker buildx imagetools inspect "${GHCR_IMAGE}:latest" | tee "$RUNNER_TEMP/ghcr-latest-manifest.txt"
docker buildx imagetools inspect --raw "${GHCR_IMAGE}:latest" > "$RUNNER_TEMP/ghcr-latest-manifest.json"
expected=$'linux/amd64\nlinux/arm64'
actual="$(jq -r '[.manifests[] | select(.platform != null and .platform.os != null and .platform.architecture != null) | "\(.platform.os)/\(.platform.architecture)"] | sort | .[]' "$RUNNER_TEMP/ghcr-latest-manifest.json")"
if [[ "$actual" != "$expected" ]]; then
echo "GHCR latest manifest platforms do not match exactly:" >&2
printf '%s\n' "$actual" >&2
exit 1
fi
if [[ "$PUBLISH_DOCKERHUB" == "true" ]]; then
docker buildx imagetools inspect "${DOCKERHUB_IMAGE}:latest" | tee "$RUNNER_TEMP/dockerhub-latest-manifest.txt"
docker buildx imagetools inspect --raw "${DOCKERHUB_IMAGE}:latest" > "$RUNNER_TEMP/dockerhub-latest-manifest.json"
actual="$(jq -r '[.manifests[] | select(.platform != null and .platform.os != null and .platform.architecture != null) | "\(.platform.os)/\(.platform.architecture)"] | sort | .[]' "$RUNNER_TEMP/dockerhub-latest-manifest.json")"
if [[ "$actual" != "$expected" ]]; then
echo "Docker Hub latest manifest platforms do not match exactly:" >&2
printf '%s\n' "$actual" >&2
exit 1
fi
fi
{
echo "### Published Docker images"
echo "- GHCR: \`${GHCR_IMAGE}:${VERSION}\`"
echo "- GHCR latest: \`${GHCR_IMAGE}:latest\`"
if [[ "$PUBLISH_DOCKERHUB" == "true" ]]; then
echo "- Docker Hub: \`${DOCKERHUB_IMAGE}:${VERSION}\`"
echo "- Docker Hub latest: \`${DOCKERHUB_IMAGE}:latest\`"
fi
} >> "$GITHUB_STEP_SUMMARY"

176
.github/workflows/tray-binaries.yml vendored Normal file
View File

@@ -0,0 +1,176 @@
name: Build macOS tray binary (arm64)
# systray2 ships only an x86_64 tray_darwin_release, so Apple Silicon users need
# Rosetta 2 for the menubar icon. This builds the native arm64 overlay that
# cli/hooks/trayRuntime.js downloads from the `tray-binaries` release.
#
# Manual-only: the artifact's sha256 is pinned in cli/hooks/trayRuntime.js and
# verified on every download, so a new build is only publishable together with a
# matching pin. Running this with publish=true against a mismatched pin fails
# rather than silently bricking every Apple Silicon client.
on:
workflow_dispatch:
inputs:
publish:
description: "Upload to the tray-binaries release (requires sha to match ARM64_TRAY_SHA256)"
required: false
default: false
type: boolean
concurrency:
group: tray-binaries-${{ github.repository }}
cancel-in-progress: false
permissions:
contents: read
env:
# Pinned because -trimpath only makes the build reproducible for a given Go
# version and macOS SDK. Bumping this changes the sha256.
GO_VERSION: "1.27.1"
jobs:
build:
name: Build darwin/arm64
runs-on: macos-15
timeout-minutes: 20
permissions:
contents: write
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 22
- uses: actions/setup-go@v5
with:
go-version: ${{ env.GO_VERSION }}
# The Go module lives in a temp clone of the upstream repo, so there is
# no go.sum at the workspace root for setup-go's cache to key on.
cache: false
- name: Record SDK provenance
run: |
{
echo "runner macOS: $(sw_vers -productVersion)"
echo "Xcode: $(xcodebuild -version | head -1)"
echo "clang: $(clang --version | head -1)"
echo "Go: $(go version)"
} | tee sdk-provenance.txt
- name: Build
run: node cli/scripts/buildTrayArm64.js
- name: Compare against pinned checksum
id: sha
run: |
BUILT=$(shasum -a 256 cli/.tray-build/tray_darwin_arm64 | cut -d' ' -f1)
# Whitespace-tolerant, and a missing constant must fail loudly: a null
# match would otherwise surface as an opaque TypeError from [1].
PINNED=$(node -e '
const m = require("fs").readFileSync("cli/hooks/trayRuntime.js", "utf8")
.match(/ARM64_TRAY_SHA256\s*=\s*"([0-9a-f]{64})"/);
if (!m) { console.error("::error::ARM64_TRAY_SHA256 not found in cli/hooks/trayRuntime.js"); process.exit(1); }
process.stdout.write(m[1]);
')
{
echo "built=$BUILT"
echo "pinned=$PINNED"
if [ "$BUILT" = "$PINNED" ]; then echo "match=true"; else echo "match=false"; fi
} >> "$GITHUB_OUTPUT"
- name: Write job summary
run: |
{
echo "### tray_darwin_arm64"
echo ""
echo "| | |"
echo "|---|---|"
echo "| built sha256 | \`${{ steps.sha.outputs.built }}\` |"
echo "| pinned sha256 | \`${{ steps.sha.outputs.pinned }}\` |"
echo "| match | ${{ steps.sha.outputs.match }} |"
echo ""
echo '```'
cat sdk-provenance.txt
echo '```'
echo ""
if [ "${{ steps.sha.outputs.match }}" = "true" ]; then
echo "Pin already matches — safe to re-run with \`publish=true\`."
else
echo "⚠️ Pin does **not** match. To publish this build, set \`ARM64_TRAY_SHA256\`"
echo "in \`cli/hooks/trayRuntime.js\` to the built sha256 above and land that"
echo "change first. Publishing without it makes every Apple Silicon client fail"
echo "checksum verification and fall back to the Rosetta binary."
fi
} >> "$GITHUB_STEP_SUMMARY"
# Uploaded before the mismatch gate below, so a publish run that fails on a
# checksum mismatch still leaves the bytes downloadable — that is exactly
# the run where a maintainer needs them to verify the new sha256.
- uses: actions/upload-artifact@v4
with:
name: tray_darwin_arm64
path: |
cli/.tray-build/tray_darwin_arm64
sdk-provenance.txt
- name: Refuse to publish on checksum mismatch
if: ${{ inputs.publish && steps.sha.outputs.match != 'true' }}
run: |
echo "::error::publish requested but built sha256 != ARM64_TRAY_SHA256"
echo " built: ${{ steps.sha.outputs.built }}"
echo " pinned: ${{ steps.sha.outputs.pinned }}"
echo "Update cli/hooks/trayRuntime.js and land it before publishing."
exit 1
- name: Publish to tray-binaries release
if: ${{ inputs.publish && steps.sha.outputs.match == 'true' }}
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
run: |
# Clients fetch from the hardcoded ARM64_TRAY_URL, so publishing from a
# different repo would gate an asset stream nobody downloads and silently
# decouple the sha pin from the bytes Apple Silicon users execute.
URL_REPO=$(node -e '
const m = require("fs").readFileSync("cli/hooks/trayRuntime.js", "utf8")
.match(/"https:\/\/github\.com\/([^\/"]+\/[^\/"]+)\/releases\/download\/tray-binaries\/tray_darwin_arm64"/);
if (!m) { console.error("::error::ARM64_TRAY_URL not found in cli/hooks/trayRuntime.js"); process.exit(1); }
process.stdout.write(m[1]);
')
if [ "${{ github.repository }}" != "$URL_REPO" ]; then
echo "::error::publishing to ${{ github.repository }}, but ARM64_TRAY_URL points clients at $URL_REPO"
echo "Repoint ARM64_TRAY_URL in cli/hooks/trayRuntime.js at this repo, or run the publish from $URL_REPO."
exit 1
fi
# gh release upload does not create the release, so bootstrap it on the
# first publish run rather than failing with "release not found".
if ! gh release view tray-binaries >/dev/null 2>&1; then
echo "Release 'tray-binaries' does not exist yet — creating it"
gh release create tray-binaries --latest=false \
--title "Native macOS tray binaries" \
--notes "Built by .github/workflows/tray-binaries.yml. Provenance and the pinned sha256 live in that workflow and in cli/hooks/trayRuntime.js (ARM64_TRAY_SHA256)."
fi
gh release upload tray-binaries cli/.tray-build/tray_darwin_arm64 --clobber
echo "Uploaded. Verifying public download URL..."
URL="https://github.com/${{ github.repository }}/releases/download/tray-binaries/tray_darwin_arm64"
GOT=""
for attempt in 1 2 3; do
if curl -fsSL --max-time 60 -o /tmp/verify "$URL"; then
GOT=$(shasum -a 256 /tmp/verify | cut -d' ' -f1)
if [ "$GOT" = "${{ steps.sha.outputs.built }}" ]; then break; fi
fi
# A just-uploaded asset can 404 or serve stale bytes until the CDN catches up.
echo "attempt $attempt: got '${GOT:-<download failed>}' — retrying in 15s"
sleep 15
done
if [ "$GOT" != "${{ steps.sha.outputs.built }}" ]; then
echo "::error::downloaded asset sha256 '${GOT:-<none>}' != built ${{ steps.sha.outputs.built }} after 3 attempts"
exit 1
fi
echo "✅ $URL serves the expected bytes"

17
.gitignore vendored
View File

@@ -87,8 +87,25 @@ graphify-out/*
.PR/
.next-analyze/*
# Kiro local workspace state
.kiro/
9router-*
# Local sensitive / temp files
.engine-token.txt
.tmp-prov.json
.tmp-providers.json
.oauth-session.json
.dev-server.log
.dev-server-err.log
.start-dev.ps1
start-dev-silent.cjs
debug.log
# CommandCode CLI local state (auth/taste/projects)
.commandcode/
# Pi subagent run artifacts
.pi-subagents/
# Vitest local artifacts
.vitest/

View File

@@ -1,3 +1,298 @@
# v0.5.91 (2026-09-26)
## Features
- **Providers**: add Token Harbor provider and four OpenAI-compatible aggregator providers (dahl, atria, agnes, bai)
- **Claude**: forward `x-claude-code-session-id` on OAuth requests; merge client `anthropic-beta` flags and forward rate-limit headers; return thinking text to OpenAI-format clients
- **Codex**: add GPT-6 Sol and Luna support
- **CLI Tools**: support multiple model profiles for Codex CLI
- **Hermes**: multi-role model config (delegation + auxiliary slots)
- **OpenCode Go**: complete the Go catalog (40 models) with auto-fetch + family endpoint regex
- **Usage**: show and redeem free limit resets for cc accounts
- **Cline**: expose the `cline-free/*` tier and price it at zero
- **Combos**: display vision adapter models in an ordered table view
## Fixes
- **Claude**: decloak tool names when `toolNameMap` misses (#4342); update spoofed cli version to 2.1.280 to support Opus 5.5
- **Providers API**: make POST `/api/providers` O(1) and refuse silent key overwrite (#4350)
- **Capabilities**: stop caching the catalog source per module copy (#4351)
- **OAuth**: stop Zed paste-token crash and add IDE auto-import (#4359)
- **Dashboard**: resolve combo limits with the server's capabilities (#4360); lazy-load charts and `marked`, preload in background on idle
- **Responses**: carry the streamed output items in `response.completed` (#4307)
- **STT**: dispatch live-API-only Gemini models over the Live WebSocket transport (#4006)
- **Gemini**: guard terminal model turns and unresponded functionCalls in `normalizeGeminiContents`
- **Command Code**: replay raw byte chunks to preserve all NDJSON lines
- **Translator**: stop emitting empty `<think>` markers into OpenAI content
- **CLI Tools**: refresh Codex settings after apply (#4347); keep existing `ANTHROPIC_AUTH_TOKEN` when applying Claude settings
- **Tray**: native arm64 macOS menubar binary, no Rosetta required
- **CLI**: filter model selector by active connections and noAuth providers
- **Usage**: key live byApiKey stats by full api key to prevent team-key collision and preserve API key usage attribution
- **Tailscale**: cap enable-flow health wait at 20s
# v0.5.86 (2026-09-23)
## Features
- **Xiaomi MiMo**: server-assisted desktop login for headless/Docker deployments, five account clusters (cn/sgp/ams/ru/in), and v2.6 pro/flash/pro-ultraspeed models with dual-route (account service vs. cloud API)
- **Claude**: add Claude Opus 5.5 support
- **i18n**: translate React text rewrites via characterData mutation observer
## Fixes
- **Proxy Pools**: keep request headers intact through Vercel/Cloudflare/Deno relays (spreading a `Headers` instance yielded `{}`, dropping auth and content-type)
- **Xiaomi MiMo login**: keep the session in the httpOnly cookie only, require dashboard auth on the proxy branch, and stop forwarding authorization headers upstream
# v0.5.85 (2026-09-22)
## Features
- **System One**: add `/v1/systemone` decision endpoint for Jev models (OpenCode Zen and OpenRouter lanes), wire into sidebar and Media Providers page with interactive probe testing
- **CLI Tools**: add dynamic configuration, settings APIs, and official logos for Pi, OMP, Crush, ForgeCode, Smelt, and CodeWhale
- **Analytics & Usage**: add Requests mode, provider/model breakdown charts, All Time period filter, and refined overview cards
- **Combos**: add Cursor/Claude Default presets; support bulk select/delete and bulk strategy changes (Fallback / Round Robin / Fusion)
- **Model Capabilities**: expose model capability metadata on `/v1/models` and aggregate capabilities across combo targets
- **OpenCode Zen & MiMo**: add OpenCode Zen (`opencode-zen`) provider with free-tier fingerprint; switch default vision fallback to MiMo V2.6 Flash Free
- **Qoder CN**: add `qoder-cn` provider for qoder.com.cn with OAuth flow, COSY protocol, and CN gateway routing
## Fixes
- **Translator**: map Claude `refusal` stop_reason to `content_filter` and surface explanation; strip replayed reasoning fields for Groq, Mistral, and Cerebras (#4220)
- **Antigravity**: drop requestType `agent` to avoid false 429 `RESOURCE_EXHAUSTED`; separate weekly and short-window (5-hour) quotas and deduplicate dashboard rows
- **Responses API**: report usage on `response.completed` so clients can auto-compact (#3432)
- **Hugging Face**: migrate to Inference Providers router (`router.huggingface.co`), expand image models catalog, and add STT route
- **Qoder**: prevent signed request replay (`403/103 Duplicate request`), handle code 110 billing blocks, and preserve upstream SSE error status
- **Performance**: bound usage `lastUsed` scan to a 2-day window; map large budget tokens to `max` reasoning tier
- **Docker**: publish verified multi-platform images (linux/amd64 and linux/arm64) with configurable apk build mirrors
# v0.5.81 (2026-09-18)
## Features
- **Xiaomi MiMo**: merge MiMo Desktop support into `xiaomi-mimo` with dual auth (API key + Desktop/OAuth session), Preview models support, and encrypted-callback OAuth flow
- **Claude Code**: add 1M-context toggle (`[1m]` marker) and drive `CLAUDE_CODE_AUTO_COMPACT_WINDOW` directly from the dashboard
- **Models**: add DeepSeek-V4.1-Flash to DeepSeek provider, CodeBuddy-Intl, and Ollama (`deepseek-v4.1-flash:cloud`); enable `low`..`max` reasoning effort levels and vision capability for DeepSeek-V4.*
- **i18n**: integrate Persian (fa) translation
## Fixes
- **Cursor**: stop AgentService empty turns (`OUT 0`) and silent hangs — fold system prompts instead of `custom_system_prompt`, send `ModelDetails`, read Composer/Grok `thinking_delta`, ack request-context without echoing MCP tools, and reject IDE execs so the model can continue
- **RTK**: for Cursor, compress source-format `tool_result` / `role:tool` **before** translation — its translator rewrites those shapes, so post-translate compression missed them. Other providers keep the post-translate pass unchanged
- **OpenCode / OpenCode Go**: resolve 403 `FreeTierError` and 429 rate limits with canonical session format, valid User-Agent, and stable upstream session reuse; force stream and declare `forceStream` for free-tier SSE aggregation; cloak decoy tools, normalize Muse Free tool choice, and strip prior reasoning items on Responses models; route Union Alpha via Messages API
- **Kiro**: preserve underscores in tool names (`mcp__server__tool`) and restore client tool names in responses; use neutral placeholder for tool-result-only turns; forward tool-result images
- **Stream**: report aborts after HTTP 200 in-band (per-format error frames) instead of closing silently
- **Command Code**: preserve images and `reasoning_effort` on `/alpha/generate`; retry transient stream errors and avoid fake stop chunks; add Quota Tracker support
- **Zed**: harden OAuth lifecycle (preserve `systemId`, renew proxy timeout), support live model resolution, and lower display priority in OAuth list
- **Antigravity**: scope cached thought signatures to model family; strip Claude Code billing headers from system prompts; sanitize Hermes system identity
- **Codex**: route bare `codex-auto-review` requests to the Codex provider (#4135)
- **Auth**: do not cool down an account for request-scoped 4xx errors
- **Usage**: improve DeepSeek credit balance display as currency credit instead of 0/total quota bar
- **Model Catalog**: scope synced catalog to gateways and declare vision capabilities for DeepSeek V4.1-Flash IDs
# v0.5.75 (2026-09-10)
## Features
- **Video**: add OpenRouter and Vertex AI (Veo) video generation on `/v1/videos/*` via a provider adapter layer; poll requests resolve their provider from `x-connection-id` or `?provider=`
- **Antigravity**: add weekly quota tracking (Gemini weekly / Claude & GPT weekly) and free-tier handling from `retrieveUserQuotaSummary` (#3892)
- **Codex**: add GPT Image 2.5, Flare and Sunburst image models with multi-image support; add the same ids to the OpenAI catalog
- **Qoder**: surface usage to all clients and stop inlining large attachments — images upload through `/api/v2/image/upload` like qodercli, oversized file blocks become stubs, context tier auto-escalates
- **OpenCode Go**: add newly published models (glm-5.3, kimi-k3, deepseek-flash, longcat-2.0, hy4-preview, hy3 on chat/completions; qwen3.8-max, qwen3.8-flash on `/messages`; grok-4.6, gpt-5.6-luna on Responses) and list `deepseek-v4.1-flash` first in the catalog
- **CLI tools**: group the model selector by provider with full-text search and manual custom model ID entry
- **CodeBuddy-CN**: replace `deepseek-v4-flash` with `deepseek-v4.1-flash`
## Fixes
- **Tools**: scope Claude tool type defaulting to gateways declaring `requireClaudeToolType` — the global default broke Anthropic-compatible endpoints that only accept the legacy typeless tool shape (#3905)
- **Claude**: cap re-anchored `cache_control` at the 4-marker budget so a spent budget no longer 400s and triggers a full combo failover; wrap bare single-object content turns before the mid-conversation-system fold
- **Cline / Airforce**: unwrap the `{"success":true,"data":…}` envelope on non-stream chat completions (#3644); add the live Cline/ClinePass model catalog and refresh Airforce free models
- **Cline**: stop `workos:`-prefixing ClinePass API keys (401 on every request, #2333) and add clinepass token refresh
- **Kiro**: never send a top-level `systemPrompt` (`400 REQUEST_BODY_INVALID`); route requests through current runtime surfaces (#3776)
- **Codex**: strip Unicode-property tool schema patterns the validator rejects (#3922); restore the `Version` header and single-source the CLI version
- **DeepSeek**: keep Anthropic-only tool types when forwarding to `/anthropic/v1/messages`
- **Qoder**: drop the Responses usage plumbing from shared translator/handler code, which changed token accounting for every provider, not just Qoder
- **Antigravity**: normalize contents and handle intermediate tool responses; protect the OAuth token-refresh path from Google anti-abuse rate limits (#3813)
- **Providers**: clear stale connection health state (`modelLock_*`, `backoffLevel`, `rateLimitedUntil`, `errorCode`) when a connection is re-validated (#3810, #3830); remove the duplicate `qwen` provider that shadowed `alims-intl`
- **Video / Vertex**: reject job ids and model ids that would escape the request URL path (SSRF)
- **Usage**: parse the Fable weekly limit from `limits[]` instead of fabricating a row (#3847)
- **Auth**: set a 24h `maxAge` on the dashboard session cookie
# v0.5.70 (2026-09-08)
## Features
- **Providers**: creating a compatible / custom-embedding node now registers the endpoint only — the API key is added afterwards from the node's page, like the built-in providers. The create dialogs drop the API Key / Model ID / Check fields, and `POST /api/provider-nodes` no longer accepts credentials at all, so a node can never be half-created
- **Providers**: compatible nodes now use the same model rows as built-in providers — capability badges, copy, per-model test, alias handling and the Add/Edit Model modal with vision + reasoning toggles, replacing the weaker read-only list
- **Providers**: compatible nodes get the built-in bulk toolbars: Test All / Disable / Active / Select All over connections, and Test All Models / Disable All / Active All over models, with per-row disable and a restore strip for disabled models
- **Providers**: wire the dead "Fetch Models" button on compatible nodes to the live upstream catalog, de-duplicating against already-added models
## Fixes
- **Usage**: nested combos (comboA lists comboB, comboC, …) stay one slot each — the inner combo always runs as fallback to produce a single answer, failed hops are not written to Details/usage, and streaming no longer inserts a 0-token placeholder. One user message against a nested fallback combo is one request row with real tokens; fusion of nested combos is N panel slots + judge, not every nested leaf
- **Models**: persist per-model capability assertions for custom and compatible providers and honor them everywhere — unsupported media is stripped on the chat path, `/v1/models` and `/api/models` report what the user asserted, and thinking translation follows it (asserting `reasoning:false` now actually strips thinking fields, `reasoning:true` emits them)
- **Models**: partial capability edits merge instead of overwriting, so toggling vision off no longer erases a stored reasoning assertion
- **Capabilities**: keep server-injected readers (synced catalog, user-asserted capabilities) in process-wide state — Next.js compiles startup and each API route into separate bundles with their own module instances, so a boot-time install was invisible to every request handler and the models.dev catalog contributed nothing to upstream requests since 0532f00d
- **Dashboard**: thinking-level picker and model-row suffix reflect user-asserted reasoning on compatible nodes
- **Providers**: `Default Model` is optional when adding an API key to a compatible node — the node's own model list (and the picker in the test modals) already determine what gets probed, and the built-in fallback still covers connection checks
- **Providers**: restore the `useCopyToClipboard` import dropped from the provider detail page, which crashed the route with `ReferenceError` for every provider
- **DB**: restore `getModelAliases` / `setModelAlias` / `deleteModelAlias` re-exports dropped from the `localDb` shim by 86112cee, which broke `GET /api/models` and `GET /v1/models` at import time
- **Providers**: remove dead `PassthroughModelsSection` (never passed props, superseded by the shared model rows)
- **Media Providers**: creating a custom embedding node reports that a key still has to be added, instead of claiming a key was saved; the edit dialog keeps its API Key + Check affordance since a stored key already exists there
- **Build**: self-host Inter instead of fetching it through `next/font/google` at build time — a Docker / mirrored builder with no route to `fonts.googleapis.com` failed the entire image build on `Failed to fetch 'Inter' from Google Fonts`. The seven `@font-face` rules and their `unicode-range`s copy what `next/font` emitted (a `latin`-only file would have dropped Vietnamese diacritics) and the latin subset is preloaded as before, so rendered metrics are unchanged
# v0.5.69 (2026-09-05)
## Features
- **Codex**: add GPT 6.0 Astra (`gpt-6-astra`) with vision, thinking and search capabilities
- **Usage**: add Claude Fable quota tracker support with weekly window normalization (`weekly fable (7d)`)
- **Dashboard**: group Antigravity Gemini and Claude quotas in Quota Tracker, prune stale hidden keys
- **OpenCode Go**: add `muse-spark-1.3-contributor` model and support parallel tool calls on Responses path (#3819)
- **Providers & Models**: align CodeBuddy-CN catalog/capabilities with server config; add GPT-5.6 Sol, Terra, Luna image aliases on Codex (#3806); refresh Qoder catalog with capability mapping and image pass-through
- **CLI tools**: replace Copilot MITM with VS Code extension setup guide
- **Gemini**: persist and replay `thoughtSignature` scoped by session namespace
## Fixes
- **Claude**: normalize adaptive auto effort (`output_config.effort`) (#3792)
- **Antigravity**: prevent Google anti-abuse rate limits during multi-account refresh (#3813)
- **Anthropic-compatible**: forward Claude beta flags to nodes fronting Anthropic (#3797)
- **Dashboard**: dynamic mode label for local/remote detection (#3801)
- **Codex**: format reset credit API errors cleanly (#3778)
- **Security**: guard cowork MCP tools probe against SSRF (#3783)
- **OpenCode Go**: track OpenCode Go quota (#3791) and send stable session headers (#3800)
- **Logger**: suppress noisy background token refresh logs
- **CLI**: export packed `.tgz` directly into workspace root instead of parent directory
# v0.5.65 (2026-09-03)
## Features
- **Fetch**: add Ollama Cloud web fetch provider
- **Gemini / Antigravity**: add Gemini 3.8 Flash support and bump IDE fingerprint to 2.11.0
- **Claude**: add Claude Fable 5.1 support (adaptive thinking with `output_config.effort`), bump Claude Code fingerprint to 2.1.258 for new-model access
- **Providers**: add client-side status filter (All / Active / Inactive / No connection) on the Providers dashboard; add max height and scroll for connection list
- **Providers & Models**: streamline tokenrouter model catalog down to 22 flagship/newest models and add missing provider icons; refresh Codebuddy-CN catalog (add hy4-preview/hy3/glm-5.3/kimi-k3-1, drop EOL glm-5.0/glm-4.7)
- **Models**: capability toggles (vision, reasoning) when adding custom models with upsert and live caps refresh
- **CLI tools**: support saving and managing custom API key presets
- **Quota**: add usage and rate-limit tracking for Groq via `x-ratelimit-*` headers
- **i18n**: complete Indonesian translation (1391 keys)
## Fixes
- **Security**: close SSRF guard bypasses in `ssrfGuard.js` (alternate IPv6 encodings, hostname trailing dots, wildcard DNS resolution check, safe redirect handling) (#3714)
- **Model markers**: strip the `[1m]` context marker Claude Code appends to model names (`claude-opus-5[1m]`) preventing model resolution failures (#3690)
- **Claude**: drop `server_tool_use` blocks carrying foreign IDs to avoid Anthropic 400 rejections; never anchor cache breakpoints on `defer_loading` tools (#3567)
- **Antigravity**: strike-break optimistic quota readings that keep 429ing by blocking the connection+model pair for 15m after 3 strikes (#3681); preserve client identity on model catalog requests (#3414)
- **Auth**: protect root `/responses` rewrite requiring API key validation in dashboardGuard
- **Chat & Docker**: return 503 Service Unavailable when all credentials are rate-limited; explicitly bundle `node-machine-id` into standalone Docker runtime image
- **OpenCode**: route Muse Spark models to `/zen/v1/responses` and declare vision support; filter inactive free model
- **Kiro**: preserve inline images as OpenAI-compatible `image_url` parts in OpenAI MITM; remove redundant top-level `systemPrompt` from payload
- **Usage**: read Responses-shape `cached_tokens` in `extractUsageFromResponse` for non-streaming traffic
- **Models**: support single model lookup with provider-prefixed IDs (e.g. `cc/claude-sonnet-5`)
- **Translator**: route Gemini thinking through `reasoning_effort` on OpenAI-compatible wire; convert `prefixItems` and ensure array items in Gemini schema sanitizer
- **UI**: apply persisted theme before first paint to prevent flash on reload; translate combo vision adapter label
# v0.5.59 (2026-08-29)
## Features
- **Search**: new web search providers — Antigravity (Google Search grounding
on the existing OAuth account pool, citations keyed and merged by URL) and
Xquik (X search with `x-api-key` auth, cursor pagination, credit-based
usage), both on `POST /v1/search`. Based on #3437 by @Nautilaceae
- **Search**: ollama-search and zai-search borrow a chat provider's API key
instead of requiring their own connection, driven by a new
`credentialFallback` registry field. zai-search later folded into the `glm`
provider itself so the web search page shows the shared connection
- **Models**: daily background sync of model capabilities from models.dev —
modalities keyed by model id (majority of sources must declare one),
context/output limits keyed by provider + model, strictly additive and
sitting below the hand-written tables. ETag + mtime cache, 60s startup
delay, `MODEL_CATALOG_SYNC=off` to disable
- **Models**: add GLM-5.3-Flash (1M context, natively multimodal), DeepSeek
V4 Vision, Grok 4.5/4.6 (500k context); correct glm-4.6v/4.5v video input
and output limits, backfill glm-4.6v on glm-cn
- **Usage**: show the Zed plan quota on the dashboard — plan, edit
predictions, hosted model requests and billing-cycle reset; unlimited rows
render as "N used · Unlimited"
- **Usage**: track GPT-5.3-Codex-Spark quota windows (spark_session /
spark_weekly) from the Codex usage response (#3431)
- **Antigravity**: quota-aware routing — on 409/429 fetch live quota for the
exact per-model resetAt and skip only the exhausted account/model pair;
report the earliest reset when every account is blocked (#3561)
- **Antigravity**: map image `size` to the aspect-ratio model suffix (-WxH);
add the Gemini 3.7 Flash tiers to MITM defaultModels so they show up in
the dashboard model-mapping table
- **Dashboard**: bulk import Grok CLI accounts from JSON — paste an array or
drag-drop multiple .json files, all OAuth connections created in a single
call, mirroring the codex flow
- **CLI tools**: endpoint presets shared across every tool card through one
live-resyncing store, instead of per-card localStorage copies that never
saw each other's saved endpoints
- **Token Saver**: configurable compression timeout (`headroomTimeoutMs`) —
the fixed 3000 ms made busy machines time out and send inconsistently
compressed bodies, hurting prompt caching
- **i18n**: pt-BR expanded to 1132 terms
## Fixes
- **Claude Code**: add Claude Fable 5.1 and advertise Claude Code 2.1.258 in
both the request header and billing identity; use its permanent adaptive-thinking
mode with `output_config.effort`
- **Stream**: record usage when a client closes on the terminal event — the
Responses API has no [DONE] sentinel, so codex closed the socket on
`response.completed` and cancelled the reader before flush() ran its usage
side effects; the tail now lives in a once-guarded finalizeStream(). Also
stop logging a disconnect for every completed Responses call
- **Stream**: parse the trailing NDJSON line an Ollama stream leaves behind
without a closing newline — the final chunk carrying `done_reason` and the
token counts was dropped
- **Session**: read the Claude Code session id from the
`x-claude-code-session-id` header — `metadata.user_id` is dropped by
Responses translation, splitting one conversation across several
`prompt_cache_key` values and missing the upstream prefix cache
- **Usage**: preserve nested `cached_tokens` — the top-level-only read
persisted `cached_tokens: 0` for every Responses-format provider (codex,
grok-cli, …), billing cache hits at the full input rate
- **Usage**: GLM quotas accept CREDIT_LIMIT plans and multi-interval windows
(5h session / 7d weekly) instead of overwriting a single "session" key
- **Models**: the catalog sync no longer erases its own output — deltas were
measured against the previous run's writes (the second run cut `providers`
from 20 entries to 5); one vote per provider in the modality tally, ETag
restored from file on startup, and the worker thread dropped after the
bundler rewrote its path into a module-not-found error
- **Executor**: CommandCode returns errors as a `type:"error"` event inside
an HTTP 200 NDJSON stream — peek the first events before committing, abort
and return a real 4xx/5xx so combo/account fallback triggers instead of
streaming the error text as content
- **Search**: scope failure locks on the credential-fallback path — a failing
search locked `modelLock___all` and took the shared glm key offline for
chat as well; locks are now attributed to the connection's owner and
scoped to `websearch:<provider>`
- **Providers**: connection tests get a 15s AbortSignal timeout instead of
hanging and exhausting the browser socket pool; guard undefined provider
names on the providers page
- **Antigravity**: sanitize competing-client branding via a config-driven
rule table (Zed's Claude-agent prompt, opencode → antigravity) — upstream
answers 429 Quota Exhausted. Applied in the executor so the shared
openai-to-gemini translator leaves gemini/vertex/zed untouched
- **MiniMax**: preserve images on the sourceFormat-matched OpenAI transport
— MiniMax-M3 resolved a Claude-shaped body posted to the OpenAI endpoint,
silently dropping `image_url` blocks (#3418)
- **Claude**: decloak tool names in same-format streaming passthrough —
OAuth-cloaked names (CLAUDE_TOOL_SUFFIX) leaked to the client and every
tool call was rejected as unknown
- **Tools**: default a missing `tools[].type` to "custom" on Claude-format
requests — strict Anthropic-compatible gateways (MiniMax) reject the
request with 400 otherwise
- **Translator**: zai thinkingFormat sends the top-level `reasoning_effort`
object GLM-5.2+ requires — every GLM-5.x request ran at the model default
(max); gated on GLM-5.2+ since older GLM does not read it (#2721)
- **RTK**: system prompt injection matches each target wire format
(Chat/Responses/Claude/Gemini/Kiro) and is exact-idempotent across retries,
so distinct prompts sharing a long prefix are no longer collapsed (#3202).
Also set the diagnostic before the silent null return on Responses
translation failure so the panel is no longer blank
- **OpenCode**: route muse-spark through /zen/v1/responses (it 500s on
chat/completions), normalizing the Chat fields the Responses API rejects
and clamping max/ultra effort to xhigh
- **CLI**: install better-sqlite3 without build tools on Node 22+ (N-API
13.0.3 ships per-platform prebuilds, `--ignore-scripts` skips the implicit
node-gyp build); Node < 22 stays on 12.6.2, working installs untouched
- **CLI tools**: send the API key Codex actually reads —
`[model_providers.9router.http_headers]` instead of auth.json (which left
every request 401 and clobbered an existing ChatGPT login); subagent model
moved to `agents.default_subagent_model`
- **OAuth**: refresh Cline tokens with the extension JSON contract
- **Dashboard**: clamp the API key mask length — keys shorter than 8 chars
threw RangeError and crashed the media-provider detail page
- **UI**: wait for the Material Symbols font itself before revealing icons —
`document.fonts.ready` resolved before the 4MB woff2 even started loading,
leaving icons blank until a second load
# v0.5.55 (2026-08-14)
## Features
@@ -625,4 +920,4 @@
# v0.4.46 (2026-05-15)
## Breaking Changes
- Tunnel public URL changed — old tunnel links no longer work, please reconnect to get the new URL
- Tunnel public URL changed — old tunnel links no longer work, please reconnect to get the new URL

View File

@@ -88,4 +88,15 @@ Pre-translate hooks that compress `tool_result` content in-place to cut tokens.
- `custom-server.js` wraps the Next standalone server to derive client IP from the TCP socket and strip attacker-controlled `X-Forwarded-For` — trusting forwarding headers only from a loopback reverse proxy. Preserve this when touching request/IP/rate-limit code.
- Security-sensitive env: `JWT_SECRET` (session cookie), `INITIAL_PASSWORD` (default `123456` — must override), `API_KEY_SECRET`, `MACHINE_ID_SALT`. Full env contract in `.env.example` and ARCHITECTURE.md's env matrix.
- Binary/protobuf upstreams (kiro EventStream, cursor protobuf, commandcode NDJSON) don't round-trip through OpenAI — they're handled inside their own executor, not the translator.
- **Security-first on PRs**: Security is the top priority when reviewing or creating PRs. Audit authentication, credential/token storage & leaks, header manipulation (`X-Forwarded-For`), and SSRF risks before functional logic. Always include explicit security warnings/notes when reporting PR reviews or changes to the user.
- Versioning: root and `cli/` are versioned independently; changes are logged in `CHANGELOG.md`. Commit style is Conventional Commits (`fix(translator): …`, `feat(...)`).
<!-- BEGIN:nextjs-agent-rules -->
# This is NOT the Next.js you know
This version has breaking changes — APIs, conventions, and file structure may all differ from your training data. Read the relevant guide in `node_modules/next/dist/docs/` (resolved from this file's directory; in monorepos the `next` package may not be visible from the repo root) before writing any code. Heed deprecation notices.
This block is written and re-added by `next dev` — verify at `node_modules/next/dist/server/lib/generate-agent-files.js`. Removing it from a diff only re-creates the uncommitted change; committing it with your work keeps the tree clean.
<!-- END:nextjs-agent-rules -->

View File

@@ -100,6 +100,12 @@ docker rm -f 9router
# re-run the quick start command
```
To pin a specific version instead of following `latest`, use a numbered image tag:
```bash
docker pull decolua/9router:0.5.81
```
---
# 🛠 For Developers
@@ -107,7 +113,7 @@ docker rm -f 9router
## Build image locally (test)
```bash
cd app && docker build -t 9router .
docker build -t 9router .
docker run --rm -p 20128:20128 \
-v "$HOME/.9router:/app/data" \
@@ -115,18 +121,67 @@ docker run --rm -p 20128:20128 \
9router
```
The Dockerfile uses the official Alpine and npm registries by default. Regional mirrors can be supplied when needed:
```bash
docker build \
--build-arg ALPINE_MIRROR=mirrors.aliyun.com \
--build-arg NPM_REGISTRY=https://registry.npmmirror.com/ \
-t 9router .
```
## Publish (automatic via CI)
Push a git tag `v*` → GitHub Actions builds multi-platform (amd64+arm64) and pushes to:
- `ghcr.io/decolua/9router:v{version}` + `:latest`
- `decolua/9router:v{version}` + `:latest`
Push a Docker-safe semver git tag `vX.Y.Z` (or a prerelease such as `vX.Y.Z-rc.1`) → GitHub Actions builds `linux/amd64` and `linux/arm64` on native runners, health-checks each platform image, verifies the resulting manifest and `/api/health`, then publishes:
- `ghcr.io/decolua/9router:X.Y.Z` + `:latest`
- `decolua/9router:X.Y.Z` + `:latest`
The `v` prefix is used only for the git tag; image tags omit it. A stable tag push promotes `latest`, but a prerelease tag such as `vX.Y.Z-rc.1` publishes only its numbered image by default. Prereleases require an explicit manual `promote_latest` opt-in. Promotion happens only after both native platform builds, both platform health checks, manifest inspection, and the resolved-manifest smoke test succeed. A failed or timed-out platform build therefore cannot move `latest`.
The workflow rejects SemVer build metadata such as `v1.2.3+build.7` because the `+` form is not a valid Docker image tag. The git tag and both `package.json` versions must match exactly.
```bash
# Use scripts/release.js (recommended)
node scripts/release.js "Release title" "Notes"
# Or manually
git tag v0.4.x && git push origin v0.4.x
git tag v0.5.81 && git push origin v0.5.81
```
Workflow: `app/.github/workflows/docker-publish.yml`
To republish an existing tag, run the `Build and Push Docker Image` workflow manually and provide the exact tag, for example `v0.5.81`, in the `release_tag` input. Manual runs publish the numbered tag but leave `latest` unchanged by default:
```text
release_tag: v0.5.81
promote_latest: false
```
The `promote_latest` checkbox is an explicit opt-in for changing `latest`. Use it when a deliberate rollback or recovery should make that version the current default:
```text
release_tag: v0.5.75
promote_latest: true
```
Numbered image tags are mutable because a republish can replace their manifest. For a deployment that must be immutable, pin the image digest instead:
```bash
docker pull decolua/9router@sha256:<verified-digest>
```
The release workflow runs `/api/health` on each native `amd64` and `arm64` platform image before it uploads the digest artifact or assembles the multi-platform manifest. It then runs a second health check against the resolved version manifest before any requested `latest` promotion.
During recovery, the selected tag remains the application source while the Dockerfile from the workflow revision is used, so an older tag can be rebuilt with the current publishing fixes.
The workflow is tag-driven. Creating a git tag does not automatically create a GitHub Release, so the Releases page and the published package/image tags can be at different versions unless a maintainer creates a release separately.
The upstream repository needs these repository secrets for Docker Hub publishing:
- `DOCKERHUB_USERNAME`
- `DOCKERHUB_TOKEN`
GHCR publishing uses the workflow's `GITHUB_TOKEN` with package write permission. Forks can publish to their own GHCR namespace, but Docker Hub publication is restricted to the upstream `decolua/9router` repository.
The optional repository variables `ALPINE_MIRROR` and `NPM_REGISTRY` can override the default package mirrors used by the CI Docker build.
Workflow: `.github/workflows/docker-publish.yml`

View File

@@ -1,24 +1,51 @@
# syntax=docker/dockerfile:1.7
ARG NODE_IMAGE=node:22-alpine
ARG ALPINE_MIRROR=dl-cdn.alpinelinux.org
ARG NPM_REGISTRY=https://registry.npmjs.org/
ARG APP_VERSION=unknown
FROM ${NODE_IMAGE} AS base
ARG ALPINE_MIRROR
WORKDIR /app
FROM base AS builder
# Use the official Alpine mirror by default. A repository variable/build arg can
# override it for environments that require a regional mirror.
RUN if [ "$ALPINE_MIRROR" != "dl-cdn.alpinelinux.org" ]; then \
sed -i "s|dl-cdn.alpinelinux.org|${ALPINE_MIRROR}|g" /etc/apk/repositories; \
fi
RUN apk --no-cache upgrade && apk --no-cache add python3 make g++ linux-headers
FROM base AS builder
ARG NPM_REGISTRY
RUN apk add --no-cache python3 make g++ linux-headers
COPY package.json ./
RUN --mount=type=cache,target=/root/.npm \
npm install
npm install \
--registry="${NPM_REGISTRY}" \
--fetch-retries=5 \
--fetch-retry-factor=2 \
--fetch-retry-mintimeout=10000 \
--fetch-retry-maxtimeout=120000 \
--fetch-timeout=300000
COPY . ./
ENV NEXT_TELEMETRY_DISABLED=1
# Inter is self-hosted (public/fonts + src/app/fonts-inter.css), so this build needs no
# route to fonts.googleapis.com / fonts.gstatic.com — only the npm mirror above is required.
RUN npm run build
FROM ${NODE_IMAGE} AS runner
ARG ALPINE_MIRROR
ARG APP_VERSION
WORKDIR /app
LABEL org.opencontainers.image.title="9router"
RUN if [ "$ALPINE_MIRROR" != "dl-cdn.alpinelinux.org" ]; then \
sed -i "s|dl-cdn.alpinelinux.org|${ALPINE_MIRROR}|g" /etc/apk/repositories; \
fi
LABEL org.opencontainers.image.title="9router" \
org.opencontainers.image.version="${APP_VERSION}"
ENV NODE_ENV=production
ENV PORT=20128
@@ -40,13 +67,16 @@ COPY --from=builder /app/node_modules/next ./node_modules/next
# sql.js loads dist/sql-wasm.wasm by path at runtime; tracing only follows JS imports,
# so the last-resort DB driver would abort with ENOENT on the missing binary.
COPY --from=builder /app/node_modules/sql.js ./node_modules/sql.js
# node-machine-id is createRequire-loaded at runtime; tracing omits it.
COPY --from=builder /app/node_modules/node-machine-id ./node_modules/node-machine-id
RUN mkdir -p /app/data && chown -R node:node /app && \
mkdir -p /app/data-home && chown node:node /app/data-home && \
ln -sf /app/data-home /root/.9router 2>/dev/null || true
# Fix permissions at runtime (handles mounted volumes)
RUN apk --no-cache upgrade && apk --no-cache add su-exec && \
# Avoid a full distribution upgrade in the runtime image. It makes builds less
# reproducible and is unrelated to installing the runtime entrypoint helper.
RUN apk add --no-cache su-exec && \
printf '#!/bin/sh\nchown -R node:node /app/data /app/data-home 2>/dev/null\nexec su-exec node "$@"\n' > /entrypoint.sh && \
chmod +x /entrypoint.sh

View File

@@ -110,7 +110,22 @@ PORT=20128 NEXT_PUBLIC_BASE_URL=http://localhost:20128 npm run dev
Production mode:
```bash
# Create Temporary Memory For Build
sudo fallocate -l 2G /swapfile_temp
sudo chmod 600 /swapfile_temp
sudo mkswap /swapfile_temp
sudo swapon /swapfile_temp
export MAKEFLAGS="-j1"
export DLIB_NO_GUI_SUPPORT=1
export CFLAGS="-mno-avx"
npm run build
# Clear temporary swap
sudo swapoff /swapfile_temp
sudo rm /swapfile_temp
PORT=20128 HOSTNAME=0.0.0.0 NEXT_PUBLIC_BASE_URL=http://localhost:20128 npm run start
```
@@ -215,7 +230,14 @@ Default URLs:
<b>🇻🇳 Tiếng Việt</b><br/>
<sub>Hướng Dẫn Setup OpenClaw + 9Router: Tạo Bot Zalo AI Tự Động Từ A-Z<br/>by <a href="https://github.com/tuanminhhole">tuanminhhole</a></sub>
</td>
<td align="center" width="320"></td>
<td align="center" width="320">
<a href="https://www.youtube.com/watch?v=hgnE7MKi3Y4">
<img src="https://img.youtube.com/vi/hgnE7MKi3Y4/maxresdefault.jpg" alt="Bye Limit! Cara Bikin Sistem 'AI Unlimited' 100% Gratis Dengan 9Router!
" width="300"/>
</a><br/>
<b>🇮🇩 Indonesia</b><br/>
<sub>Bye Limit! Cara Bikin Sistem "AI Unlimited" 100% Gratis Dengan 9Router!<br/>by <a href="https://www.youtube.com/@neptiver">neptiver</a></sub>
</td>
<td align="center" width="320"></td>
<td align="center" width="320"></td>
</tr>

1407
bun.lock Normal file

File diff suppressed because it is too large Load Diff

1
cli/.gitignore vendored
View File

@@ -1,2 +1,3 @@
app/*
node_modules/*
.tray-build/

View File

@@ -111,7 +111,7 @@ Any tool supporting OpenAI/Claude-compatible API works.
Full docs, advanced setup, video tutorials & development guide:
- **GitHub**: https://github.com/decolua/9router
- **Full README**: https://github.com/decolua/9router/blob/main/app/README.md
- **Full README**: https://github.com/decolua/9router/blob/master/README.md
- **Website**: https://9router.com
---

View File

@@ -6,7 +6,13 @@ const fs = require("fs");
const os = require("os");
const path = require("path");
const BETTER_SQLITE3_VERSION = "12.6.2";
// Gate the pinned version by Node major, mirroring src/lib/db/driver.js gating
// style: 13.x is N-API and ships per-platform prebuilds inside the package, so
// it needs no ABI-specific download. It requires Node >= 22; older runtimes stay
// on 12.6.2, which fetches an ABI-specific binary via prebuild-install.
const [NODE_MAJOR] = process.versions.node.split(".").map(Number);
const USE_NAPI_BUILD = NODE_MAJOR >= 22;
const BETTER_SQLITE3_VERSION = USE_NAPI_BUILD ? "13.0.3" : "12.6.2";
const SQL_JS_VERSION = "1.14.1";
function getDataDir() {
@@ -45,9 +51,23 @@ function hasModule(name) {
return fs.existsSync(path.join(getRuntimeNodeModules(), name, "package.json"));
}
function isGlibcRuntime() {
try { return Boolean(process.report?.getReport()?.header?.glibcVersionRuntime); } catch { return true; }
}
// 12.x compiles/downloads into build/Release; 13.x ships prebuilds/<platform>-<arch>.node.
function getBetterSqliteBinary() {
const root = path.join(getRuntimeNodeModules(), "better-sqlite3");
const platform = process.platform === "linux" && !isGlibcRuntime() ? "linuxmusl" : process.platform;
return [
path.join(root, "build", "Release", "better_sqlite3.node"),
path.join(root, "prebuilds", `${platform}-${process.arch}.node`),
].find((file) => fs.existsSync(file));
}
function isBetterSqliteBinaryValid() {
const binary = path.join(getRuntimeNodeModules(), "better-sqlite3", "build", "Release", "better_sqlite3.node");
if (!fs.existsSync(binary)) return false;
const binary = getBetterSqliteBinary();
if (!binary) return false;
try {
const fd = fs.openSync(binary, "r");
const buf = Buffer.alloc(4);
@@ -91,6 +111,7 @@ function runNpmInstall({ cwd, pkgs, extraArgs = [], timeout = 180000 }) {
function npmInstall(pkgs, opts = {}) {
const cwd = ensureRuntimeDir();
const extra = opts.optional ? ["--no-save"] : [];
if (opts.ignoreScripts) extra.push("--ignore-scripts");
if (!opts.silent) console.log("⏳ Installing SQLite engine (first run)...");
const res = runNpmInstall({ cwd, pkgs, extraArgs: extra, timeout: opts.timeout || 180000 });
if (!res.ok && !opts.silent) {
@@ -129,7 +150,10 @@ function ensureSqliteRuntime({ silent = false } = {}) {
return { betterSqlite: true, sqlJs: sqlJsOk };
}
const ok = npmInstall([`better-sqlite3@${BETTER_SQLITE3_VERSION}`], { optional: true, silent });
// npm injects an implicit `node-gyp rebuild` for any package carrying a
// binding.gyp, which would demand build tools even though 13.x already bundles
// the binary — skip scripts so the bundled prebuild is used as-is.
const ok = npmInstall([`better-sqlite3@${BETTER_SQLITE3_VERSION}`], { optional: true, silent, ignoreScripts: USE_NAPI_BUILD });
return {
betterSqlite: ok && hasModule("better-sqlite3") && isBetterSqliteBinaryValid(),
sqlJs: sqlJsOk,

View File

@@ -5,9 +5,17 @@
//
// We use the maintained `systray2` fork. The original `systray@1.0.5` package
// bundles a 2017 x86_64 Go binary whose Mach-O headers are rejected by modern
// dyld (macOS 14+), so the tray silently fails to register on Apple Silicon.
// dyld (macOS 14+), so it fails to load at all.
//
// Note that systray2 is NOT an Apple Silicon fix: like its predecessor it ships
// only an x86_64 `tray_darwin_release`, and picks it by process.platform with no
// process.arch branch, so there is no native slice to select. On arm64 macOS the
// tray therefore needs Rosetta 2 and dies with EBADARCH without it. We overlay
// our own arm64 build of the same upstream source on top — see ensureArm64TrayBin.
const { spawnSync } = require("child_process");
const crypto = require("crypto");
const fs = require("fs");
const os = require("os");
const path = require("path");
const { getRuntimeDir, getRuntimeNodeModules, runNpmInstall, summarizeNpmError } = require("./sqliteRuntime");
@@ -15,6 +23,19 @@ const SYSTRAY_PKG = "systray2";
const SYSTRAY_VERSION = "2.1.4";
const LEGACY_SYSTRAY_PKG = "systray";
// Pinned `tray-binaries` release rather than `latest`, so the URL is stable and
// the artifact can only change by a deliberate re-upload. The workflow's publish
// step re-derives this repo from the literal below and refuses to upload
// anywhere else, so the integrity gate can't drift from what clients fetch.
//
// The asset is built by .github/workflows/tray-binaries.yml on a macos-15 runner.
// cgo compiles AppKit against the runner's SDK, so this value tracks that image:
// when GitHub updates it the sha changes, the workflow refuses to publish, and
// this constant must be bumped in the same change as the re-upload.
const ARM64_TRAY_URL = "https://github.com/decolua/9router/releases/download/tray-binaries/tray_darwin_arm64";
const ARM64_TRAY_SHA256 = "487e3c365aaa1eb6ad295bf3989711e975b52cee07505bf641c8559954881c81";
const ARM64_RETRY_COOLDOWN_MS = 24 * 60 * 60 * 1000;
function hasSystray() {
return fs.existsSync(path.join(getRuntimeNodeModules(), SYSTRAY_PKG, "package.json"));
}
@@ -71,6 +92,124 @@ function ensureRuntimeDir() {
return dir;
}
// A thin (non-fat) 64-bit Mach-O stores its magic then cputype, both LE.
// CPU_TYPE_ARM64 is CPU_TYPE_ARM | CPU_ARCH_ABI64. Fat/universal binaries use a
// different magic and are reported as "not arm64" here, which is fine: we only
// ever overlay a thin arm64 build and only need to tell it apart from x86_64.
function isArm64MachO(file) {
let fd = null;
try {
fd = fs.openSync(file, "r");
const buf = Buffer.alloc(8);
fs.readSync(fd, buf, 0, 8, 0);
if (buf.readUInt32LE(0) !== 0xfeedfacf) return false;
return buf.readUInt32LE(4) === 0x0100000c;
} catch {
return false;
} finally {
if (fd !== null) try { fs.closeSync(fd); } catch {}
}
}
// systray2 is constructed with copyDir:true, so what actually executes is
// ~/.cache/node-systray/<version>/tray_darwin_release, and index.js only re-copies
// when that path is absent. An overlaid binary stays invisible until this is cleared.
// Scoped to our systray2 version: the parent dir is machine-global and shared
// with any other node-systray consumer.
function bustSystrayCopyCache() {
try {
fs.rmSync(path.join(os.homedir(), ".cache", "node-systray", SYSTRAY_VERSION), { recursive: true, force: true });
} catch {}
}
function arm64AttemptMarker() {
return path.join(getRuntimeDir(), ".tray-arm64-attempt");
}
// ensureTrayRuntime runs synchronously on every `9router` start (cli.js), so a
// failed download must not re-block the next launch. Retry at most daily.
function recentlyAttemptedArm64() {
try {
const at = Number(fs.readFileSync(arm64AttemptMarker(), "utf8").trim());
return Number.isFinite(at) && Date.now() - at < ARM64_RETRY_COOLDOWN_MS;
} catch {
return false;
}
}
function markArm64Attempt() {
try { fs.writeFileSync(arm64AttemptMarker(), String(Date.now())); } catch {}
}
// Cleared on success so the cooldown only ever throttles *failures*. Without
// this, anything that restores systray2's x86_64 binary later — notably a
// globally installed 9router older than this change, which shares the same
// ~/.9router/runtime — would leave the user waiting out the cooldown.
function clearArm64Attempt() {
try { fs.rmSync(arm64AttemptMarker(), { force: true }); } catch {}
}
// Throws on any failure so the caller has a single error path.
function downloadFile(url, dest, timeoutSec) {
// darwin-only path, and curl ships with macOS, so this needs no extra dep and
// keeps the caller synchronous.
const res = spawnSync("curl", ["-fsSL", "--max-time", String(timeoutSec), "-o", dest, url], {
encoding: "utf8",
timeout: (timeoutSec + 5) * 1000
});
if (res.status === 0 && fs.existsSync(dest)) return;
const detail = (res.stderr || res.error?.message || `curl exit ${res.status}`).trim().split("\n").pop();
throw new Error(detail || "download failed");
}
function sha256File(file) {
return crypto.createHash("sha256").update(fs.readFileSync(file)).digest("hex");
}
// Replace systray2's x86_64 macOS binary with a native arm64 build so Apple
// Silicon users get a tray without installing Rosetta 2. Any failure leaves the
// Intel binary untouched, which still works under Rosetta.
//
// Takes no `silent` flag on purpose: cli.js calls ensureTrayRuntime({silent:true})
// synchronously on every start, and a stalled curl would otherwise freeze the
// launch for up to 30s with no output at all. These lines print at most once per
// 24h on failure and once ever on success, so they are worth more than the quiet.
function ensureArm64TrayBin() {
if (process.platform !== "darwin" || process.arch !== "arm64") return { skipped: true };
const binPath = path.join(getRuntimeNodeModules(), SYSTRAY_PKG, "traybin", "tray_darwin_release");
if (!fs.existsSync(binPath)) return { skipped: true };
if (isArm64MachO(binPath)) return { native: true };
if (recentlyAttemptedArm64()) return { deferred: true };
markArm64Attempt();
console.log("⏳ Downloading native Apple Silicon tray binary...");
// pid-scoped: two concurrent starts (postinstall racing cli.js, or two
// terminals) would otherwise interleave writes to one file, fail each other's
// checksum, and delete each other's in-flight download from the catch below.
const tmp = `${binPath}.arm64.${process.pid}.tmp`;
try {
downloadFile(ARM64_TRAY_URL, tmp, 30);
const sum = sha256File(tmp);
// Integrity matters more than usual: this is an executable that runs on
// every Apple Silicon user's machine.
if (sum !== ARM64_TRAY_SHA256) throw new Error(`checksum mismatch (got ${sum.slice(0, 12)}…)`);
if (!isArm64MachO(tmp)) throw new Error("downloaded file is not an arm64 Mach-O");
fs.chmodSync(tmp, 0o755);
fs.renameSync(tmp, binPath);
bustSystrayCopyCache();
clearArm64Attempt();
console.log("✅ Native Apple Silicon tray installed");
return { native: true, installed: true };
} catch (e) {
try { fs.rmSync(tmp, { force: true }); } catch {}
console.warn("⚠️ Native tray download failed — falling back to the Intel binary");
console.warn(` Reason: ${e.message}`);
console.warn(" The Intel tray needs Rosetta 2: softwareupdate --install-rosetta --agree-to-license");
return { native: false, error: e.message };
}
}
function npmInstall(pkgs, { silent = false } = {}) {
const cwd = ensureRuntimeDir();
if (!silent) console.log("⏳ Installing system tray (first run)...");
@@ -94,14 +233,20 @@ function ensureTrayRuntime({ silent = false } = {}) {
if (process.platform === "win32") {
return { systray: false, skipped: true };
}
if (hasSystray()) {
let ready = hasSystray();
if (!ready) {
ready = npmInstall([`${SYSTRAY_PKG}@${SYSTRAY_VERSION}`], { silent }) && hasSystray();
}
if (ready) {
chmodSystrayBin({ silent });
if (!silent) console.log("✅ System tray ready");
return { systray: true };
}
const ok = npmInstall([`${SYSTRAY_PKG}@${SYSTRAY_VERSION}`], { silent });
if (ok) chmodSystrayBin({ silent });
return { systray: ok && hasSystray() };
// Runs after the ready log so a download failure doesn't read as a broken
// tray — the Intel binary still works under Rosetta.
const arm64 = ready ? ensureArm64TrayBin() : { skipped: true };
return { systray: ready, arm64 };
}
module.exports = { ensureTrayRuntime };
module.exports = { ensureTrayRuntime, ensureArm64TrayBin };

View File

@@ -1,6 +1,6 @@
{
"name": "9router",
"version": "0.5.55",
"version": "0.5.91",
"description": "9Router CLI - Start and manage 9Router server",
"bin": {
"9router": "./cli.js"
@@ -16,7 +16,8 @@
"scripts": {
"dev": "nodemon -I --watch cli.js --watch src --watch hooks --ext js,json cli.js",
"build": "node scripts/build-cli.js",
"pack:cli": "npm run build && npm pack --pack-destination ../..",
"build:tray-arm64": "node scripts/buildTrayArm64.js",
"pack:cli": "npm run build && npm pack --pack-destination ..",
"publish:cli": "npm run build && npm publish",
"postinstall": "node hooks/postinstall.js",
"prepublishOnly": "npm run build"
@@ -29,7 +30,8 @@
"react-dom": "19.2.1"
},
"comment_sqlite": "sql.js + better-sqlite3 are NOT bundled here. They are installed into ~/.9router/runtime/node_modules by hooks/postinstall.js (and re-checked at runtime by cli.js). This avoids Windows EBUSY errors when updating the global CLI, since native .node files no longer live under the locked install dir.",
"comment_systray": "systray2 is NOT bundled here. It is lazy-installed into ~/.9router/runtime/node_modules by hooks/postinstall.js on macOS/Linux only. Windows uses PowerShell NotifyIcon (zero binary). This avoids shipping unsigned Go binaries that trigger antivirus false positives (Kaspersky). We use the systray2 fork because the legacy systray@1.0.5 ships a 2017 x86_64 binary that fails on modern macOS dyld.",
"comment_systray": "systray2 is NOT bundled here. It is lazy-installed into ~/.9router/runtime/node_modules by hooks/postinstall.js on macOS/Linux only. Windows uses PowerShell NotifyIcon (zero binary). This avoids shipping unsigned Go binaries that trigger antivirus false positives (Kaspersky). We use the systray2 fork because the legacy systray@1.0.5 ships a 2017 x86_64 binary that fails to load on modern macOS dyld. Neither package ships an arm64 macOS binary, so on Apple Silicon hooks/trayRuntime.js overlays our own arm64 build from the tray-binaries GitHub release; without it the tray requires Rosetta 2.",
"comment_tray_arm64": "tray_darwin_arm64 is built by scripts/buildTrayArm64.js (npm run build:tray-arm64) from felixhao28/systray-portable — the same source systray2's binary comes from — and uploaded to the pinned 'tray-binaries' GitHub release. Its sha256 is pinned as ARM64_TRAY_SHA256 in hooks/trayRuntime.js and verified after every download; rebuild and update both together. .github/workflows/tray-binaries.yml builds the same artifact in CI but refuses to publish when the sha diverges from the pin.",
"engines": {
"node": ">=18.0.0"
},

View File

@@ -0,0 +1,108 @@
#!/usr/bin/env node
// Rebuilds tray_darwin_arm64, the native Apple Silicon menubar binary that
// hooks/trayRuntime.js overlays on top of systray2's x86_64-only build.
//
// Must run on macOS: getlantern/systray is cgo against AppKit, so the arm64
// slice needs a real macOS SDK. Requires Go on PATH (`mise use -g go@latest`).
//
// Output is NOT reproducible across Go versions even with -s -w, so after a
// rebuild you must re-upload the asset and update ARM64_TRAY_SHA256 in
// hooks/trayRuntime.js — this script prints both and fails if they diverge.
const { execFileSync, spawnSync } = require("child_process");
const crypto = require("crypto");
const fs = require("fs");
const os = require("os");
const path = require("path");
const UPSTREAM_REPO = "https://github.com/felixhao28/systray-portable.git";
// master as of 2021-09-15, the commit systray2@2.1.4's own binary was built from.
const UPSTREAM_COMMIT = "6eddc917bf39fcc0d95b57a0741d0c065fbd1e23";
const outDir = path.join(__dirname, "..", ".tray-build");
const outFile = path.join(outDir, "tray_darwin_arm64");
const trayRuntimePath = path.join(__dirname, "..", "hooks", "trayRuntime.js");
function fail(msg) {
console.error(`\n❌ ${msg}`);
process.exit(1);
}
function run(cmd, args, opts = {}) {
const res = spawnSync(cmd, args, { stdio: "inherit", ...opts });
if (res.status !== 0) fail(`${cmd} ${args.join(" ")} exited with ${res.status}`);
}
if (process.platform !== "darwin") fail("must be run on macOS (cgo needs the AppKit SDK)");
const go = spawnSync("go", ["version"], { encoding: "utf8" });
if (go.status !== 0) {
fail("Go toolchain not found on PATH. Install it with: mise use -g go@latest");
}
console.log(`Go: ${go.stdout.trim()}`);
const srcDir = fs.mkdtempSync(path.join(os.tmpdir(), "systray-portable-"));
// fail() exits through process.exit(), which does not unwind the stack, so a
// try/finally here would leak the clone on every failed build. An exit handler
// covers normal completion, fail(), and uncaught exceptions alike.
process.on("exit", () => {
try { fs.rmSync(srcDir, { recursive: true, force: true }); } catch {}
});
console.log(`\nCloning ${UPSTREAM_REPO} @ ${UPSTREAM_COMMIT.slice(0, 7)}`);
run("git", ["clone", "--quiet", UPSTREAM_REPO, srcDir]);
run("git", ["-C", srcDir, "checkout", "--quiet", UPSTREAM_COMMIT]);
const head = execFileSync("git", ["-C", srcDir, "log", "-1", "--format=%H %ad %s", "--date=short"], {
encoding: "utf8"
}).trim();
console.log(`HEAD: ${head}`);
run("go", ["mod", "download"], { cwd: srcDir });
console.log("\nBuilding darwin/arm64...");
// -trimpath strips the local build directory from the binary, so two builds
// from the same commit + Go version hash identically regardless of where they
// ran. Without it the pinned sha256 could never be regenerated.
run("go", ["build", "-trimpath", "-ldflags", "-s -w", "-o", outFile, "tray.go"], {
cwd: srcDir,
env: { ...process.env, CGO_ENABLED: "1", GOOS: "darwin", GOARCH: "arm64" }
});
// ── Verify ────────────────────────────────────────────────────────────────
const buf = Buffer.alloc(8);
const fd = fs.openSync(outFile, "r");
fs.readSync(fd, buf, 0, 8, 0);
fs.closeSync(fd);
if (buf.readUInt32LE(0) !== 0xfeedfacf || buf.readUInt32LE(4) !== 0x0100000c) {
fail("output is not a thin arm64 Mach-O");
}
// Apple Silicon refuses to execute an unsigned binary. Go's linker applies an
// ad-hoc signature automatically; confirm it survived.
const sig = spawnSync("codesign", ["--verify", outFile], { encoding: "utf8" });
if (sig.status !== 0) fail(`ad-hoc signature invalid: ${(sig.stderr || "").trim()}`);
const sha256 = crypto.createHash("sha256").update(fs.readFileSync(outFile)).digest("hex");
const sizeMb = (fs.statSync(outFile).size / 1024 / 1024).toFixed(2);
console.log(`\n✅ ${outFile}`);
console.log(` arch: arm64 (ad-hoc signed, verified)`);
console.log(` size: ${sizeMb} MB`);
console.log(` sha256: ${sha256}`);
// Whitespace-tolerant: a formatter could wrap the assignment across lines, and
// a null match must fail loudly rather than silently read as "no pin".
const pinMatch = fs.readFileSync(trayRuntimePath, "utf8").match(/ARM64_TRAY_SHA256\s*=\s*"([0-9a-f]{64})"/);
if (!pinMatch) fail(`could not find ARM64_TRAY_SHA256 in ${trayRuntimePath}`);
const pinned = pinMatch[1];
if (pinned === sha256) {
console.log(`\n Matches ARM64_TRAY_SHA256 in hooks/trayRuntime.js — no code change needed.`);
} else {
console.log(`\n⚠️ Differs from ARM64_TRAY_SHA256 in hooks/trayRuntime.js (${pinned}).`);
console.log(` Re-upload the release asset, then update that constant to the sha256 above.`);
}
console.log(`\nNext: upload to the pinned release tag, keeping the asset name stable:`);
console.log(` gh release upload tray-binaries "${outFile}" --clobber`);

View File

@@ -53,6 +53,9 @@ const PROVIDER_MODELS = {
{ id: "glm-4.7" },
],
ag: [
{ id: "gemini-3.8-flash-high" },
{ id: "gemini-3.8-flash-medium" },
{ id: "gemini-3.8-flash-low" },
{ id: "gemini-3.7-flash-high" },
{ id: "gemini-3.7-flash-medium" },
{ id: "gemini-3.7-flash-low" },
@@ -101,6 +104,8 @@ const PROVIDER_MODELS = {
{ id: "claude-3-5-sonnet-20241022" },
],
gemini: [
{ id: "gemini-3.8-flash" },
{ id: "gemini-3.7-flash" },
{ id: "gemini-3.6-flash" },
{ id: "gemini-3.5-flash-lite" },
{ id: "gemini-3-pro-preview" },

View File

@@ -141,11 +141,14 @@ function initWindowsTray(options) {
/**
* macOS/Linux tray via systray binary
*
* Prefers `systray2` (active fork of `systray`, ships newer
* getlantern/systray-portable binaries that work on macOS 14+ and Apple
* Silicon under Rosetta). Falls back to legacy `systray@1.0.5` if systray2
* is not available, though that binary's Mach-O headers are rejected by
* modern dyld and the icon will not appear.
* Prefers `systray2`, the active fork of `systray`. Both ship only an x86_64
* `tray_darwin_release` and select it by process.platform alone, so on Apple
* Silicon the tray runs under Rosetta 2 and fails with EBADARCH when Rosetta is
* absent. hooks/trayRuntime.js overlays a native arm64 build over that file to
* avoid the dependency; the fallbacks below are Intel-only.
*
* Falls back to legacy `systray@1.0.5` if systray2 is unavailable, though that
* binary's Mach-O headers are rejected by modern dyld and no icon will appear.
*/
function resolveSystray() {
let runtimeDir = null;

View File

@@ -2,22 +2,24 @@ const api = require("../api/client");
const { prompt } = require("./input");
const { clearScreen } = require("./display");
// Provider alias order: OAuth first, then API Key (matches ModelSelectModal)
// Provider alias order: OAuth first, then Free, then API Key
const PROVIDER_ALIAS_ORDER = [
"cc", "ag", "cx", "if", "qw", "gc", "gh", "kr",
"cc", "ag", "cx", "if", "qw", "gc", "gh", "kr", "oc",
"openrouter", "glm", "kimi", "minimax", "openai", "anthropic", "gemini"
];
// Alias to display name mapping
const PROVIDER_ALIAS_NAMES = {
cc: "Claude Code",
ag: "Antigravity",
ag: "Antigravity",
cx: "OpenAI Codex",
if: "iFlow AI",
qw: "Qwen Code",
gc: "Gemini CLI",
gh: "GitHub Copilot",
kr: "Kiro AI",
oc: "OpenCode Free",
opencode: "OpenCode Free",
openrouter: "OpenRouter",
glm: "GLM Coding",
kimi: "Kimi Coding",
@@ -27,35 +29,83 @@ const PROVIDER_ALIAS_NAMES = {
gemini: "Gemini"
};
const PROVIDER_ID_TO_ALIAS = {
claude: "cc",
codex: "cx",
"gemini-cli": "gc",
github: "gh",
antigravity: "ag",
iflow: "if",
qwen: "qw",
kiro: "kr",
cursor: "cu",
cline: "cline",
clinepass: "clinepass",
qoder: "qd",
"qoder-cn": "qd",
gitlab: "gitlab",
"codebuddy-cn": "cb",
"codebuddy-intl": "cbai",
kimchi: "kimchi",
"grok-cli": "grok-cli",
trae: "trae",
windsurf: "windsurf",
zed: "zed",
opencode: "oc",
"opencode-go": "ocg",
"opencode-zen": "ocz",
};
// Providers usable without stored credentials
const NO_AUTH_PROVIDERS = new Set(["opencode", "oc"]);
/**
* Get all available models grouped by provider + combos
* Get all available models grouped by provider + combos (filtered by active connections)
* @returns {Promise<{combos: Array, groups: Object}>}
*/
async function getAvailableModelsGrouped() {
const result = await api.getAvailableModels();
if (!result.success) return { combos: [], groups: {} };
const models = result.data?.data || [];
const [modelsResult, providersResult] = await Promise.all([
api.getAvailableModels(),
api.getProviders()
]);
if (!modelsResult.success) return { combos: [], groups: {} };
const connections = providersResult.success ? (providersResult.data?.connections || []) : [];
const activeAliases = new Set(NO_AUTH_PROVIDERS);
connections.forEach(conn => {
if (conn.isActive === false) return;
const p = conn.provider;
if (!p) return;
activeAliases.add(p);
const alias = conn.providerSpecificData?.prefix || PROVIDER_ID_TO_ALIAS[p] || p;
activeAliases.add(alias);
});
const models = modelsResult.data?.data || [];
const combos = [];
const groups = {};
models.forEach(m => {
if (m.owned_by === "combo") {
combos.push(m.id);
} else {
const provider = m.owned_by;
// Only keep connected providers or noAuth providers
if (!activeAliases.has(provider)) return;
if (!groups[provider]) {
groups[provider] = [];
}
groups[provider].push(m.id);
}
});
return { combos, groups };
}
/**
* Display model list and prompt for selection
* Display model list and prompt for selection with provider grouping & search
* @param {string} title - Title to display
* @param {string} currentValue - Current selected value (optional)
* @param {Object} options - { excludeCombos?: boolean }
@@ -68,64 +118,214 @@ async function selectModelFromList(title, currentValue = "", options = {}) {
const totalModels = combos.length + Object.values(groups).flat().length;
if (totalModels === 0) {
clearScreen();
console.log(`\n🎯 ${title}`);
console.log("=".repeat(50));
console.log("\n No connected providers found.");
console.log(" Please connect a provider in Providers menu first.\n");
console.log(" m. ✍️ Enter custom model ID");
console.log(" 0. Cancel\n");
const act = await prompt("Select option (m/0): ");
const trimmed = act.trim();
if (trimmed.toLowerCase() === "m") {
const custom = await prompt("Enter custom model ID: ");
return custom.trim() || null;
}
return null;
}
// Build flat list for selection
const allModels = [];
// Display
clearScreen();
console.log(`\n🎯 ${title}`);
console.log("=".repeat(50));
if (currentValue) {
console.log(`Current: ${currentValue}\n`);
} else {
console.log();
}
let idx = 1;
// Combos first (skipped when excludeCombos is true)
// All models for flat search
const allModelsList = [
...combos,
...Object.values(groups).flat()
];
// Build category list
const categories = [];
if (combos.length > 0) {
console.log("[Combos]");
combos.forEach(combo => {
console.log(` ${idx}. ${combo}`);
allModels.push(combo);
idx++;
categories.push({
id: "combos",
name: "[Combos]",
models: combos
});
console.log();
}
// Provider groups in order (by alias)
const sortedProviders = Object.keys(groups).sort((a, b) => {
const idxA = PROVIDER_ALIAS_ORDER.indexOf(a);
const idxB = PROVIDER_ALIAS_ORDER.indexOf(b);
return (idxA === -1 ? 999 : idxA) - (idxB === -1 ? 999 : idxB);
});
sortedProviders.forEach(provider => {
sortedProviders.forEach((provider) => {
const providerName = PROVIDER_ALIAS_NAMES[provider] || provider;
console.log(`[${providerName}]`);
groups[provider].forEach(model => {
console.log(` ${idx}. ${model}`);
allModels.push(model);
idx++;
categories.push({
id: provider,
name: providerName,
models: groups[provider]
});
console.log();
});
console.log(" 0. Cancel\n");
// Prompt for number input
const input = await prompt("Enter number: ");
const num = parseInt(input, 10);
if (isNaN(num) || num === 0 || num < 0 || num > allModels.length) {
return null;
let filterQuery = null;
while (true) {
clearScreen();
console.log(`\n🎯 ${title}`);
console.log("=".repeat(50));
if (currentValue) {
console.log(`Current: ${currentValue}\n`);
} else {
console.log();
}
// Active search view
if (filterQuery !== null) {
const q = filterQuery.toLowerCase().trim();
const matched = allModelsList.filter((m) => m.toLowerCase().includes(q));
console.log(`🔍 Search results for "${filterQuery}": (${matched.length} found)\n`);
if (matched.length === 0) {
console.log(" No matching models found.\n");
console.log(" 0. ← Back to providers");
console.log(" s. Search again\n");
const act = await prompt("Select option: ");
if (act.toLowerCase() === "s") {
const newQ = await prompt("Enter search keyword: ");
filterQuery = newQ.trim() || null;
} else {
filterQuery = null;
}
continue;
}
matched.forEach((m, i) => {
console.log(` ${i + 1}. ${m}`);
});
console.log("\n 0. ← Back to providers");
console.log(" s. Search again\n");
const input = await prompt("Enter number to select (or 0/s): ");
if (input.toLowerCase() === "s") {
const newQ = await prompt("Enter search keyword: ");
filterQuery = newQ.trim() || null;
continue;
}
const num = parseInt(input, 10);
if (isNaN(num) || num === 0) {
filterQuery = null;
continue;
}
if (num > 0 && num <= matched.length) {
return matched[num - 1];
}
continue;
}
// If only 1 category exists, jump straight into its model list
if (categories.length === 1) {
const singleCategory = categories[0];
console.log(`[${singleCategory.name}]`);
singleCategory.models.forEach((m, i) => {
console.log(` ${i + 1}. ${m}`);
});
console.log();
console.log(" s. 🔍 Search models");
console.log(" m. ✍️ Enter custom model ID");
console.log(" 0. Cancel\n");
const input = await prompt("Enter choice (number / s / m / 0): ");
const trimmed = input.trim();
if (!trimmed || trimmed === "0") return null;
const lower = trimmed.toLowerCase();
if (lower === "s") {
const q = await prompt("Enter search keyword: ");
if (q.trim()) filterQuery = q.trim();
continue;
}
if (lower === "m") {
const customModel = await prompt("Enter custom model ID: ");
if (customModel.trim()) return customModel.trim();
continue;
}
const num = parseInt(trimmed, 10);
if (!isNaN(num) && num > 0 && num <= singleCategory.models.length) {
return singleCategory.models[num - 1];
}
filterQuery = trimmed;
continue;
}
// Multiple categories view
console.log("[Providers & Groups]");
categories.forEach((cat, i) => {
console.log(` ${i + 1}. ${cat.name} (${cat.models.length} models)`);
});
console.log();
console.log(" s. 🔍 Search models");
console.log(" m. ✍️ Enter custom model ID");
console.log(" 0. Cancel\n");
const input = await prompt("Enter choice (number / keyword / s / m): ");
const trimmed = input.trim();
if (!trimmed || trimmed === "0") {
return null;
}
const lower = trimmed.toLowerCase();
if (lower === "s") {
const q = await prompt("Enter search keyword: ");
if (q.trim()) {
filterQuery = q.trim();
}
continue;
}
if (lower === "m") {
const customModel = await prompt("Enter custom model ID: ");
if (customModel.trim()) {
return customModel.trim();
}
continue;
}
const num = parseInt(trimmed, 10);
// Selected a category
if (!isNaN(num) && num > 0 && num <= categories.length) {
const selectedCategory = categories[num - 1];
while (true) {
clearScreen();
console.log(`\n🎯 ${title} > ${selectedCategory.name}`);
console.log("=".repeat(50));
if (currentValue) {
console.log(`Current: ${currentValue}\n`);
} else {
console.log();
}
selectedCategory.models.forEach((m, i) => {
console.log(` ${i + 1}. ${m}`);
});
console.log("\n 0. ← Back\n");
const modelChoice = await prompt("Enter number to select (0 to back): ");
const modelNum = parseInt(modelChoice, 10);
if (isNaN(modelNum) || modelNum === 0) {
break;
}
if (modelNum > 0 && modelNum <= selectedCategory.models.length) {
return selectedCategory.models[modelNum - 1];
}
}
continue;
}
// User typed text directly -> treat as search query
filterQuery = trimmed;
}
return allModels[num - 1];
}
module.exports = {

View File

@@ -0,0 +1,261 @@
# OpenCode Go Session Header Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Send a stable, conversation-scoped `x-opencode-session` header on every OpenCode Go request and install the patched CLI locally.
**Architecture:** Add a dedicated `OpenCodeGoExecutor` extending `DefaultExecutor`. `chatCore` passes the provider-scoped session resolved from the original request plus the detected client tool; the executor derives a request-local upstream session and delegates all existing transport, authentication, retry, and proxy behavior to `DefaultExecutor`.
**Tech Stack:** Node.js ESM, Vitest, Next.js, npm CLI packaging, GitHub CLI.
## Global Constraints
- Apply the header to OpenCode Go chat completions, Claude Messages, and OpenAI Responses transports.
- Preserve a valid native `x-opencode-session`; hash all translated non-OpenCode identities to `ses_<32 lowercase hex>`.
- Namespace translated identities by detected client tool, using `generic` when unknown.
- Do not keep mutable per-request session state on the executor singleton or mutate the caller's credentials object.
- Do not change OpenCode Go models, routing, reasoning, tool behavior, dependencies, or unrelated providers.
- Reuse upstream issue #3759 instead of creating a duplicate issue.
---
### Task 1: Add Failing OpenCode Go Session Tests
**Files:**
- Create: `tests/unit/opencode-go-session.test.js`
**Interfaces:**
- Consumes: `getExecutor(provider)` and `DefaultExecutor.buildHeaders(credentials, stream, url, model)`.
- Produces: the required public behavior for `OpenCodeGoExecutor.prepareRequestCredentials({ body, credentials, providerSessionId, clientTool })` and `OpenCodeGoExecutor.execute(args)`.
- [ ] **Step 1: Write the failing tests**
Create a Vitest suite that mocks `proxyAwareFetch`, obtains `getExecutor("opencode-go")`, and asserts:
```js
const prepared = executor.prepareRequestCredentials({
body: { messages: [{ role: "user", content: "hello" }] },
credentials: { apiKey: "test-key", connectionId: "conn-a", rawHeaders: {} },
providerSessionId: "conversation-a",
clientTool: "claude",
});
expect(prepared).not.toBe(credentials);
expect(prepared._opencodeGoSession).toMatch(/^ses_[0-9a-f]{32}$/);
expect(credentials).not.toHaveProperty("_opencodeGoSession");
```
Cover native header preservation, stable values across all three runtime transports, different conversation IDs, different client tools using the same ID, connection fallback, no singleton state, no header on `DefaultExecutor("openai")`, and the final fetch headers returned by `execute()`.
- [ ] **Step 2: Run the focused test and verify RED**
Run:
```bash
npx vitest run --config tests/vitest.config.js tests/unit/opencode-go-session.test.js
```
Expected: FAIL because `getExecutor("opencode-go")` still returns `DefaultExecutor` and `prepareRequestCredentials` does not exist.
- [ ] **Step 3: Commit the failing test**
```bash
git add tests/unit/opencode-go-session.test.js
git commit -m "test: cover OpenCode Go session headers"
```
### Task 2: Implement the Dedicated Executor
**Files:**
- Create: `open-sse/executors/opencode-go.js`
- Modify: `open-sse/executors/index.js`
**Interfaces:**
- Consumes: `DefaultExecutor`, `resolveSessionId()`, request `credentials.rawHeaders`, `providerSessionId`, and `clientTool`.
- Produces: `OpenCodeGoExecutor`, `prepareRequestCredentials()`, and an `execute()` override that delegates with cloned credentials.
- [ ] **Step 1: Add the minimal executor implementation**
Implement these rules:
```js
function translatedSessionId(sessionId, clientTool) {
const digest = crypto
.createHash("sha256")
.update(`opencode-go\0${clientTool || "generic"}\0${sessionId}`)
.digest("hex")
.slice(0, 32);
return `ses_${digest}`;
}
```
`prepareRequestCredentials()` must read a case-insensitive native
`x-opencode-session` with the same non-empty, 256-character cap used by the
session manager. Otherwise it uses `providerSessionId` or calls
`resolveSessionId({ headers, body, connectionId, scope: "opencode-go" })`, then
returns `{ ...credentials, _opencodeGoSession: value }`.
`execute(args)` must call `prepareRequestCredentials(args)` and delegate using
`super.execute({ ...args, credentials: prepared })`. `buildHeaders()` must call
`super.buildHeaders()` and add the prepared session, with a connection-scoped
fallback for direct callers.
Register `new OpenCodeGoExecutor()` under `"opencode-go"` and export the class.
- [ ] **Step 2: Run the focused test and verify partial GREEN**
Run:
```bash
npx vitest run --config tests/vitest.config.js tests/unit/opencode-go-session.test.js
```
Expected: executor-level tests pass; any chatCore-context assertion remains failing until Task 3.
- [ ] **Step 3: Commit the executor**
```bash
git add open-sse/executors/opencode-go.js open-sse/executors/index.js tests/unit/opencode-go-session.test.js
git commit -m "fix(opencode-go): add stable session header executor"
```
### Task 3: Pass Original Request Session Context
**Files:**
- Modify: `open-sse/handlers/chatCore.js`
- Modify: `tests/unit/opencode-go-session.test.js`
**Interfaces:**
- Consumes: existing `sessionSeed` and `clientTool` variables in `handleChatCore()`.
- Produces: `providerSessionId` and `clientTool` fields on both initial and refreshed-credential calls to `executor.execute()`.
- [ ] **Step 1: Add or enable the failing integration assertion**
Use a mocked executor or source request containing a body-only `session_id` and
assert the executor receives the provider-scoped session resolved before
translation.
- [ ] **Step 2: Run the focused test and verify RED**
Run:
```bash
npx vitest run --config tests/vitest.config.js tests/unit/opencode-go-session.test.js
```
Expected: FAIL because `handleChatCore()` does not pass `providerSessionId` or
`clientTool` to `executor.execute()`.
- [ ] **Step 3: Pass the request context**
Add the same fields to both executor calls:
```js
executor.execute({
model,
body: translatedBody,
stream,
credentials,
providerSessionId: sessionSeed,
clientTool,
signal: streamController.signal,
log,
proxyOptions,
});
```
- [ ] **Step 4: Run focused and neighboring tests**
Run:
```bash
npx vitest run --config tests/vitest.config.js \
tests/unit/opencode-go-session.test.js \
tests/unit/opencode-go-models.test.js \
tests/unit/session-manager.test.js \
tests/unit/executor-const-guard.test.js
```
Expected: PASS with zero failed tests.
- [ ] **Step 5: Commit the context wiring**
```bash
git add open-sse/handlers/chatCore.js tests/unit/opencode-go-session.test.js
git commit -m "fix(chat): forward provider session context"
```
### Task 4: Verify and Install the Local CLI Package
**Files:**
- Generated: `9router-0.5.65.tgz`
- Packaged output: `cli/app/server.js`
**Interfaces:**
- Consumes: completed source changes and existing CLI build scripts.
- Produces: a globally installed patched `9router@0.5.65`.
- [ ] **Step 1: Run source verification**
```bash
git diff --check origin/master...HEAD
npx vitest run --config tests/vitest.config.js tests/unit/
npm run build
```
Expected: every command exits zero. Record any pre-existing full-suite failures
separately rather than hiding them.
- [ ] **Step 2: Build and package the CLI**
```bash
npm --prefix cli run build
npm --prefix cli pack -- --pack-destination ..
```
Expected: `9router-0.5.65.tgz` exists and contains the patched bundled server.
- [ ] **Step 3: Replace the global npm installation**
```bash
npm install -g ./9router-0.5.65.tgz
```
Expected: `/opt/homebrew/lib/node_modules/9router/package.json` reports `0.5.65`
and the installed bundle contains `x-opencode-session` plus the new executor.
- [ ] **Step 4: Commit any required package-source adjustment**
Do not commit generated tarballs or CLI build artifacts unless the repository
already tracks and requires them.
### Task 5: Publish the Upstream Pull Request
**Files:**
- No additional source files unless verification finds a required correction.
**Interfaces:**
- Consumes: verified branch commits and GitHub issue #3759.
- Produces: a fork branch and a PR against `decolua/9router:master`.
- [ ] **Step 1: Create or repair the GitHub fork remote**
Use `gh repo fork decolua/9router --remote` if the current `fork` remote remains
missing, then push `fix/opencode-go-session-header`.
- [ ] **Step 2: Create the PR**
Use title:
```text
fix(opencode-go): send stable session header
```
The body must include the root cause, downstream-session translation policy,
three covered transports, concurrency behavior, verification evidence,
`Fixes #3759`, and a note that this PR is intentionally narrower than #3780.
- [ ] **Step 3: Verify the published PR**
Run `gh pr view --json number,title,state,url,headRefName,baseRefName` and report
the issue and PR URLs.

View File

@@ -0,0 +1,114 @@
# OpenCode Go Session Header Design
## Problem
OpenCode Go will begin rejecting some requests without an
`x-opencode-session` header on September 6, 2026. In 9Router v0.5.65,
`opencode-go` uses `DefaultExecutor`, whose generic header builder does not add
that header. The specialized OpenCode Free executor already sends it, but that
logic does not apply to the paid OpenCode Go provider or its three transports.
## Goals
- Add `x-opencode-session` to every OpenCode Go chat, Claude Messages, and
OpenAI Responses request.
- Translate a downstream conversation identity into a stable upstream identity.
- Keep identities isolated across different downstream agents and conversations.
- Avoid exposing non-OpenCode downstream session identifiers to OpenCode Go.
- Avoid mutable session state on the shared executor singleton.
- Leave OpenCode Free and all unrelated providers unchanged.
## Non-Goals
- Inferring an exact conversation boundary when a downstream client provides no
session or conversation identifier.
- Adding or changing OpenCode Go models, routing, reasoning, or tool behavior.
- Changing the general session-resolution policy for other providers.
## Architecture
Add a dedicated `OpenCodeGoExecutor` extending `DefaultExecutor`. The executor
keeps the existing generic URL, authentication, translation, retry, and proxy
behavior, and overrides only the OpenCode Go session-header concern.
`handleChatCore` already resolves a provider-scoped session from the original
request before translation. It will pass that value and the detected client
tool to `executor.execute()` as request context. `OpenCodeGoExecutor.execute()`
will create a shallow request-local credentials object containing the resolved
OpenCode Go session. It will then delegate to `DefaultExecutor.execute()`.
This avoids storing request state on the executor singleton or mutating shared
provider credentials.
## Session Resolution
The original downstream request remains the source of truth. Existing
`resolveSessionId()` behavior recognizes Claude Code, Antigravity, generic
session headers, and common body fields before request translation can discard
them.
Resolution rules:
1. If the downstream request supplies `x-opencode-session`, treat it as an
authoritative OpenCode identity after trimming and length validation.
2. Otherwise use the provider-scoped session resolved from the original request.
3. Namespace the resolved value with the detected downstream agent, falling back
to `generic` when the agent is unknown.
4. Convert the namespaced value to an opaque deterministic identifier:
`ses_` plus the first 32 hexadecimal characters of SHA-256.
5. If no explicit downstream identity exists, the existing provider connection
fallback guarantees that a header is still sent. It is stable but cannot
distinguish multiple conversations sharing that connection.
The same input conversation produces the same upstream identifier for all three
OpenCode Go transports. Different agents using the same raw session value
produce different identifiers.
## Header Injection
`OpenCodeGoExecutor.buildHeaders()` delegates to
`DefaultExecutor.buildHeaders()` and adds only:
```text
x-opencode-session: <stable-session-id>
```
The implementation applies to:
- `https://opencode.ai/zen/go/v1/chat/completions`
- `https://opencode.ai/zen/go/v1/messages`
- `https://opencode.ai/zen/go/v1/responses`
## Error Handling
Session derivation must not make requests fail. Invalid or oversized native
header values are ignored and the normal resolved-session fallback is used.
Hashing uses Node's built-in `crypto` module and requires no new dependency.
## Testing
Add a focused unit suite that proves:
- all three OpenCode Go transports receive the header;
- the same conversation remains stable across requests and transports;
- different conversations produce different values;
- different agents using the same raw ID remain isolated;
- non-OpenCode session IDs are represented as opaque `ses_<32 hex>` values;
- a valid native `x-opencode-session` remains stable;
- headerless requests still receive a stable fallback;
- OpenCode Free behavior is unchanged;
- unrelated `DefaultExecutor` providers do not receive the header;
- no request state is retained on the shared executor instance.
Run the focused unit tests first, then the neighboring executor/session tests,
the full offline test suite, the application build, and the CLI package build.
## Delivery
Build the CLI with `npm --prefix cli run build`, create a package with
`npm --prefix cli pack`, and install the generated tarball globally to replace
the current npm-installed `9router@0.5.65`. Verify the installed package version
and packaged source contains the new executor.
Upstream issue #3759 already tracks the problem, so no duplicate issue will be
created. The pull request will be narrowly scoped to this fix, reference
`Fixes #3759`, and explain how it differs from the broader open PR #3780.

View File

@@ -118,6 +118,31 @@ Cursor/Cline/Any tool:
---
## Cursor / Claude Default Combos
Cursor and Claude Code send **unprefixed** model IDs (`composer-2.5`, `claude-opus-5`, `opus`), while 9Router routes with provider prefixes (`cu/composer-2.5`, `cc/claude-opus-5`). Default combo generators bridge that gap.
On **Dashboard → Combos**:
1. Click **Cursor Default** or **Claude Default**
2. Confirm the preview (new vs already-existing names)
3. 9Router creates one combo per client model ID, seeded with the matching prefixed route
**Examples:**
| Combo name (what the client sends) | Seeded model (what 9Router routes) |
|------------------------------------|------------------------------------|
| `composer-2.5` | `cu/composer-2.5` |
| `cursor-grok-4.6-high-fast` | `cu/cursor-grok-4.6-high-fast` |
| `claude-opus-5` | `cc/claude-opus-5` |
| `opus` | `cc/claude-opus-5` |
Existing combo names are **skipped** (not overwritten). Edit any generated combo afterward to add fallbacks. Click the button again later to pick up new catalog IDs.
> These combos help when Cursor/Claude already talk to 9Router (`/v1` or `ANTHROPIC_BASE_URL`) and send their native model IDs. They do not change Cursor’s built-in Models tab by themselves.
---
## Example Combos
### Example 1: Premium Coding (Subscription → Cheap → Free)

View File

@@ -111,6 +111,27 @@ Model: cx/gpt-5.2-codex
| `cx/gpt-5.2` | GPT 5.2 | General tasks |
| `cx/gpt-5.1-codex` | GPT 5.1 Codex | Stable coding |
### Image Generation
The Codex image catalog includes `cx/gpt-5.6-sol-image`,
`cx/gpt-5.6-terra-image`, and `cx/gpt-5.6-luna-image`, alongside the existing
GPT 5.5, 5.4, and 5.3 image aliases. Select them under **Image → OpenAI Codex**
in the dashboard, or discover them with `GET /v1/models/image` after connecting
a Codex account.
```bash
curl http://localhost:20128/v1/images/generations \
-H "Authorization: Bearer $NINE_ROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"cx/gpt-5.6-sol-image","prompt":"A blue square","size":"1024x1024"}'
```
These are 9Router aliases: the image adapter removes `-image` and sends the
underlying model an `image_generation` tool through the Codex Responses API.
The same endpoint accepts an `image` reference for edits. Image generation
requires an eligible ChatGPT Plus or higher account; availability of each
underlying model and its image tool depends on the connected account.
### Pro Tips
- **5-hour rolling quota** - Fresh quota every 5 hours

View File

@@ -75,6 +75,10 @@ const nextConfig = {
source: "/responses",
destination: "/api/v1/responses"
},
{
source: "/systemone",
destination: "/api/v1/systemone"
},
{
source: "/v1beta/:path*",
destination: "/api/v1beta/:path*"

View File

@@ -4,7 +4,7 @@ Provider-agnostic SSE engine: one OpenAI-style request → any provider (LLM cha
## Request lifecycle (chat)
`handlers/chatCore.js` → `services/model.js` `parseModel` (resolve `provider/model`) → **pre-translate hooks** (`rtk/` tool_result compress, `rtk/headroom.js` proxy compress, `rtk/caveman.js` system inject — all fail-open) → `executors/index.js` `getExecutor(provider)` → `translator/index.js` `translateRequest` (client format → provider format) → `executor.execute()` (streams upstream) → `translateResponse` (provider chunks → client format) → SSE out.
`handlers/chatCore.js` → `services/model.js` `parseModel` (resolve `provider/model`) → **RTK for `cursor`** (`rtk/` compresses the source-format `tool_result` / `role:tool` in-place — its translator rewrites those shapes, so this one provider must run **before** translate) → `translator/index.js` `translateRequest` (client format → provider format) → **post-translate savers** (`rtk/` compress for every other provider, `rtk/headroom.js` proxy compress, `rtk/caveman.js` / `rtk/ponytail.js` system inject — all fail-open) → `executors/index.js` `getExecutor(provider)` → `executor.execute()` (streams upstream) → `translateResponse` (provider chunks → client format) → SSE out.
## Directory map
@@ -16,7 +16,7 @@ Provider-agnostic SSE engine: one OpenAI-style request → any provider (LLM cha
- `rtk/` — request token-killer. `index.js` compresses `tool_result` content in-place (OpenAI/Claude/Kiro shapes); `filters/` per-tool compressors + `autodetect.js`; `headroom.js` external compress proxy; `caveman.js` system-prompt injector.
- `transformer/` — `responsesTransformer.js` (Chat Completions SSE → Codex Responses API SSE), `streamToJsonConverter.js`.
- `shared/` — cross-provider auth/identity: `clineAuth.js`, `machineId.js`, `qoder/`.
- `services/` — `model.js`, `provider.js`, `accountFallback.js`, `combo.js`, `compact.js`, `tokenRefresh/`+`tokenRefresh.js`, `oauthCredentialManager.js`, `usage/`, `projectId.js`, `kiroModels.js`/`qoderModels.js`.
- `services/` — `model.js`, `provider.js`, `accountFallback.js`, `combo.js`, `tokenRefresh/`+`tokenRefresh.js`, `oauthCredentialManager.js`, `usage/`, `projectId.js`, `kiroModels.js`/`qoderModels.js`.
- `utils/` — streamHandler, stream, sse, error, sessionManager, claudeCloaking, clientDetector, proxyFetch (patches global fetch), cursorProtobuf/cursorChecksum, ollamaTransform.
## Conventions

View File

@@ -7,6 +7,9 @@ import { createRequire } from "module";
export const GEMINI_CLI_VERSION = PROVIDERS["gemini-cli"]?.cliVersion;
export const GEMINI_CLI_API_CLIENT = PROVIDERS["gemini-cli"]?.apiClient;
// === Codex CLI === derive từ registry codex.transport
export const CODEX_CLI_VERSION = PROVIDERS["codex"]?.cliVersion;
// Map Node arch to Gemini CLI arch string (x64/x86/arm64/...)
function geminiCLIArch() {
const a = arch();
@@ -171,6 +174,18 @@ export const LOAD_CODE_ASSIST_METADATA = {
// System prompts
export const CLAUDE_SYSTEM_PROMPT = "You are Claude Code, Anthropic's official CLI for Claude.";
// Rewrite rules applied to Antigravity system prompts: competing-client branding
// makes the backend flag the request and answer 429 Quota Exhausted.
export const ANTIGRAVITY_PROMPT_REWRITES = [
{ from: "You are a Claude agent, built on Anthropic's Claude Agent SDK.", to: "" },
{ from: /You are Hermes(?: Agent)?(?:,\s*(?:an intelligent AI assistant|an AI assistant|an AI agent))?(?:,?\s*(?:built|created)\s+by\s+Nous Research)?\./gi, to: "You are an AI assistant." },
// Claude Code prepends this line to its system prompt. The Claude-format translator strips it,
// but OpenAI-format clients (e.g. proxies that convert Claude Code to /v1/chat/completions)
// pass it through, and any system text containing it gets a fake 429 RESOURCE_EXHAUSTED.
{ from: /^x-anthropic-billing-header:[^\n]*(?:\r?\n)*/gim, to: "" },
{ from: /opencode/gi, to: (m) => (m === "OpenCode" ? "Antigravity" : m === "OPENCODE" ? "ANTIGRAVITY" : "antigravity") }
];
export const ANTIGRAVITY_DEFAULT_SYSTEM = "You are Antigravity, a powerful agentic AI coding assistant designed by the Google Deepmind team working on Advanced Agentic Coding.You are pair programming with a USER to solve their coding task. The task may require creating a new codebase, modifying or debugging an existing codebase, or simply answering a question.**Absolute paths only****Proactiveness**";
// Derive từ registry oauth.refreshLeadMs

View File

@@ -73,6 +73,17 @@ export const ERROR_RULES = [
{ status: 403, cooldownMs: COOLDOWN.long },
{ status: 404, cooldownMs: COOLDOWN.long },
{ status: 429, backoff: true },
// --- Request-scoped errors: the request itself is broken — retrying the same
// body on another account/model can never succeed, and locking the account
// would punish a healthy credential for our own bad request. Callers use this
// to fail fast (no account rotation, no model lock).
{ text: "context_length_exceeded", requestScoped: true },
{ text: "context window", requestScoped: true },
{ text: "maximum context length", requestScoped: true },
{ text: "prompt is too long", requestScoped: true },
{ text: "input is too long", requestScoped: true },
{ text: "max_tokens exceed", requestScoped: true },
{ text: "reduce the length", requestScoped: true },
];
// Backward compat: COOLDOWN_MS object (used by index.js re-export)

View File

@@ -3,9 +3,13 @@ import REGISTRY from "../providers/registry/index.js";
// PROVIDER_MODELS now built from providers/registry (transport + models co-located)
import { PROVIDER_MODELS } from "../providers/index.js";
import { modelQuotaFamily, modelStrip, modelTargetFormat, modelSupportedFormats, normalizeModelId } from "../providers/models/schema.js";
import { CODEX_REVIEW_SUFFIX } from "../providers/models/helpers.js";
import { CODEX_REVIEW_SUFFIX, isMuseSparkModel, opencodeFamilyFormats } from "../providers/models/helpers.js";
import { FORMATS } from "../translator/formats.js";
export { PROVIDER_MODELS };
// OpenCode providers sharing the endpoint-family fallback for unknown model ids
const isOpenCodeAlias = (aliasOrId) => !aliasOrId || ["oc", "opencode", "ocg", "opencode-go", "ocz", "opencode-zen"].includes(aliasOrId);
// Helper functions
export function getProviderModels(aliasOrId) {
@@ -26,11 +30,14 @@ const DOT_VERSION_PROVIDERS = new Set(["kr", "kiro"]);
// ("claude-sonnet-4-5" ~= "claude-sonnet-4.5"). Other providers use exact match only.
function findModel(models, modelId, aliasOrId) {
if (!models) return undefined;
const found = models.find(m => m.id === modelId);
const baseModelId = typeof modelId === "string"
? modelId.replace(/\([^()]+\)\s*$/, "").trim()
: modelId;
const found = models.find(m => m.id === modelId || m.id === baseModelId);
if (found) return found;
if (!DOT_VERSION_PROVIDERS.has(aliasOrId)) return undefined;
const normalized = normalizeModelId(modelId);
if (normalized === modelId) return undefined;
const normalized = normalizeModelId(baseModelId);
if (normalized === baseModelId) return undefined;
return models.find(m => m.id === normalized);
}
@@ -49,17 +56,29 @@ export function findModelName(aliasOrId, modelId) {
}
export function getModelTargetFormat(aliasOrId, modelId) {
if (isOpenCodeAlias(aliasOrId) && isMuseSparkModel(modelId)) {
return FORMATS.OPENAI_RESPONSES;
}
const models = PROVIDER_MODELS[aliasOrId];
if (!models) return null;
return modelTargetFormat(findModel(models, modelId, aliasOrId));
const found = findModel(models, modelId, aliasOrId);
if (found) return modelTargetFormat(found);
// Family fallback keeps modelsFetcher/passthrough ids on their endpoint lane
if (isOpenCodeAlias(aliasOrId)) return opencodeFamilyFormats(modelId)?.targetFormat || null;
return null;
}
// Declared upstream formats for a model (registry `supportedFormats`). Drives the
// per-model guard on the sourceFormat-matched transport; null when undeclared.
// Unknown OpenCode ids fall back to the family regex (chat lane by default) so
// auto-fetched models never wrongly use the sourceFormat-matched transport.
export function getModelSupportedFormats(aliasOrId, modelId) {
const models = PROVIDER_MODELS[aliasOrId];
if (!models) return null;
return modelSupportedFormats(findModel(models, modelId, aliasOrId));
const found = findModel(models, modelId, aliasOrId);
if (found) return modelSupportedFormats(found);
if (isOpenCodeAlias(aliasOrId)) return opencodeFamilyFormats(modelId)?.supportedFormats || [FORMATS.OPENAI];
return null;
}
export function getModelType(aliasOrId, modelId) {

View File

@@ -1,12 +1,13 @@
import crypto from "crypto";
import { BaseExecutor } from "./base.js";
import { PROVIDERS } from "../config/providers.js";
import { OAUTH_ENDPOINTS, ANTIGRAVITY_HEADERS, AG_DEFAULT_TOOLS, AG_TOOL_SUFFIX } from "../config/appConstants.js";
import { OAUTH_ENDPOINTS, ANTIGRAVITY_HEADERS, AG_DEFAULT_TOOLS, AG_TOOL_SUFFIX, ANTIGRAVITY_PROMPT_REWRITES } from "../config/appConstants.js";
import { HTTP_STATUS } from "../config/runtimeConfig.js";
import { resolveSessionId } from "../utils/sessionManager.js";
import { resolveSessionId, toNumericSessionId } from "../utils/sessionManager.js";
import { proxyAwareFetch } from "../utils/proxyFetch.js";
import { cleanJSONSchemaForAntigravity } from "../translator/formats/gemini.js";
import { cleanJSONSchemaForAntigravity, normalizeGeminiContents } from "../translator/formats/gemini.js";
import { DEFAULT_THINKING_AG_SIGNATURE } from "../config/defaultThinkingSignature.js";
import { getGeminiThoughtSignatureSync } from "../services/thoughtSignatureStore.js";
// Sanitize function name: Gemini requires [a-zA-Z_][a-zA-Z0-9_.:\-]{0,63}
function sanitizeFunctionName(name) {
@@ -187,9 +188,12 @@ export class AntigravityExecutor extends BaseExecutor {
};
}
const rawSessionId = body.request?.sessionId || resolveSessionId({ headers: credentials?.rawHeaders, body, connectionId: credentials?.email || credentials?.connectionId, scope: "antigravity" });
const sessionId = toNumericSessionId(rawSessionId) || rawSessionId;
// ─── Standard (non-image) request ───
// Fix contents for Claude models via Antigravity
const contents = body.request?.contents?.map(c => {
const rawContents = (body.request?.contents || []).map(c => {
let role = c.role;
// functionResponse must be role "user" for Claude models
if (c.parts?.some(p => p.functionResponse)) {
@@ -202,21 +206,33 @@ export class AntigravityExecutor extends BaseExecutor {
return true;
});
// Gemini 3+ rejects functionCall parts without thoughtSignature. Clients (Claude Code, IDE)
// don't persist thoughtSignature in their history, so backfill the default signature on any
// functionCall part that arrives without one.
const needsBackfill = parts?.some(p => p.functionCall && !p.thoughtSignature) ?? false;
if (role !== c.role || parts?.length !== c.parts?.length || needsBackfill) {
return {
...c, role,
parts: needsBackfill
? parts.map(p => (p.functionCall && !p.thoughtSignature)
? { ...p, thoughtSignature: DEFAULT_THINKING_AG_SIGNATURE }
: p)
: parts,
};
}
return c;
// don't persist thoughtSignature in their history, so backfill from cache or default signature.
// In parallel function calls, only the first call needs a signature; siblings stay unsigned.
let firstFunctionCallSeen = false;
const modifiedParts = parts?.map(p => {
if (!p.functionCall) return p;
const callId = p.functionCall.id;
const cachedSig = callId ? getGeminiThoughtSignatureSync(callId, sessionId, body.model || model) : null;
const callSig = p.thoughtSignature || cachedSig || (!firstFunctionCallSeen ? DEFAULT_THINKING_AG_SIGNATURE : undefined);
firstFunctionCallSeen = true;
if (callSig) {
return { ...p, thoughtSignature: callSig };
}
if (p.thoughtSignature && !cachedSig) {
// Unsigned sibling call
const { thoughtSignature: _, ...rest } = p;
return rest;
}
return p;
});
return {
...c,
role,
parts: modifiedParts || parts || [],
};
});
const contents = normalizeGeminiContents(rawContents);
// Sanitize tool schemas and function names before sending to Antigravity.
let tools = body.request?.tools;
@@ -246,13 +262,13 @@ export class AntigravityExecutor extends BaseExecutor {
const { tools: _originalTools, toolConfig: _originalToolConfig, ...requestWithoutTools } = body.request || {};
stripBlacklisted(requestWithoutTools);
// Rewrite competitive system prompts (e.g. Zed IDE's Claude prompt) to prevent Antigravity from
// flagging the request and immediately blocking it with a 429 Quota Exhausted response.
// Rewrite competing-client branding in system prompts (e.g. Zed's Claude prompt,
// OpenCode naming) so Antigravity doesn't flag the request with a 429 Quota Exhausted.
if (requestWithoutTools.systemInstruction?.parts) {
const oldText = "You are a Claude agent, built on Anthropic's Claude Agent SDK.";
for (const part of requestWithoutTools.systemInstruction.parts) {
if (typeof part.text === "string" && part.text.includes(oldText)) {
part.text = part.text.split(oldText).join("");
if (typeof part.text !== "string") continue;
for (const { from, to } of ANTIGRAVITY_PROMPT_REWRITES) {
part.text = part.text.replaceAll(from, to);
}
}
}
@@ -267,7 +283,7 @@ export class AntigravityExecutor extends BaseExecutor {
generationConfig,
...(contents && { contents }),
...(tools && { tools }),
sessionId: body.request?.sessionId || resolveSessionId({ headers: credentials?.rawHeaders, body, connectionId: credentials?.email || credentials?.connectionId, scope: "antigravity" }),
sessionId,
safetySettings: undefined,
...(tools?.length > 0 && { toolConfig: { functionCallingConfig: { mode: "VALIDATED" } } })
};
@@ -277,12 +293,19 @@ export class AntigravityExecutor extends BaseExecutor {
this._lastSessionId = transformedRequest.sessionId; // cached for buildHeaders (base.execute order)
// Official Antigravity client omits `requestType` entirely on the agent
// (chat) path. Sending `requestType: "agent"` here (or leaking it through
// from an upstream envelope via the ...body spread below) makes Google
// bucket the request and return a detail-free 429 RESOURCE_EXHAUSTED even
// with quota available. `image_gen` and
// `search` buckets are unaffected and keep their own requestType.
delete body.requestType;
return {
...body,
project: projectId,
model: body.model || model,
userAgent: "antigravity",
requestType: "agent",
requestId: buildIdeRequestId({ body, request: transformedRequest, credentials, model, requestType: "agent" }),
request: transformedRequest
};

View File

@@ -128,7 +128,7 @@ export class BaseExecutor {
for (let urlIndex = 0; urlIndex < fallbackCount; urlIndex++) {
const url = this.buildUrl(model, stream, urlIndex, credentials);
const transformedBody = this.transformRequest(model, body, stream, credentials);
const headers = this.buildHeaders(credentials, stream, url, model);
const headers = this.buildHeaders(credentials, stream, url, model, transformedBody);
if (!retryAttemptsByUrl[urlIndex]) retryAttemptsByUrl[urlIndex] = 0;

View File

@@ -7,11 +7,12 @@ import {
} from "../services/oauthCredentialManager.js";
import { normalizeResponsesInput } from "../translator/formats/responsesApi.js";
import { fetchImageAsBase64 } from "../translator/concerns/image.js";
import { getModelUpstreamId } from "../config/providerModels.js";
import { getModelUpstreamId, getProviderModels } from "../config/providerModels.js";
import { getThinkingLevels } from "../providers/thinkingLevels.js";
import { DEFAULT_RETRY_CONFIG, HTTP_STATUS, resolveRetryEntry } from "../config/runtimeConfig.js";
import { dbg } from "../utils/debugLog.js";
import { resolveSessionId } from "../utils/sessionManager.js";
import { stripCodexUnsupportedPatterns } from "../utils/codexToolSchema.js";
// SSE error patterns inside 200-OK bodies. Some retry same account first; capacity rotates accounts.
const CODEX_SSE_RETRY_PATTERNS = ["server_is_overloaded", "service_unavailable_error"];
@@ -24,6 +25,10 @@ const CODEX_SSE_USER_OUTPUT_PATTERNS = [
];
const CODEX_SSE_PEEK_BYTES = 256 * 1024;
const CODEX_MODEL_CAPACITY_MESSAGE = "Selected model is at capacity. Please try a different model.";
function isCodexResponsesLiteModel(model) {
const baseId = String(model || "").replace(/\([^()]+\)\s*$/, "");
return getProviderModels("cx").some((entry) => entry.id === baseId && entry.responsesLite === true);
}
// Server-generated item id prefixes that Codex /responses cannot resolve when store=false
const SERVER_ID_PATTERN = /^(rs|fc|resp|msg)_/;
@@ -42,7 +47,7 @@ const CODEX_PASSTHROUGH_TOOL_TYPES = new Set(["custom"]);
const RESPONSES_API_ALLOWLIST = new Set([
"model", "input", "instructions", "tools", "tool_choice", "stream", "store",
"reasoning", "service_tier", "include", "prompt_cache_key", "client_metadata",
"text"
"text", "parallel_tool_calls"
]);
// Convert role=system → role=developer in body.input (keeps content in cacheable prefix)
@@ -56,13 +61,14 @@ function convertSystemToDeveloperRole(body) {
}
// Strip server-generated item IDs (rs_/fc_/resp_/msg_) from input — avoids 404 with store=false
function stripStoredItemReferences(body) {
function stripStoredItemReferences(body, preserveLitePrefix = false) {
if (!Array.isArray(body.input)) return;
body.input = body.input.filter((item) => {
if (typeof item === "string" && SERVER_ID_PATTERN.test(item)) return false;
if (item && typeof item === "object" && !Array.isArray(item)) {
if (item.type === "item_reference") return false;
if (typeof item.id === "string" && SERVER_ID_PATTERN.test(item.id)) delete item.id;
if (typeof item.id === "string" && SERVER_ID_PATTERN.test(item.id)
&& !(preserveLitePrefix && item.role === "developer" && item.id.startsWith("msg_"))) delete item.id;
}
return true;
});
@@ -72,6 +78,9 @@ function stripStoredItemReferences(body) {
function normalizeCodexTools(body) {
if (!Array.isArray(body.tools)) return;
const validNames = new Set();
// Codex's schema validator has no Unicode property escapes; a `pattern`
// carrying `\p{...}` 400s the whole request on every account (#3922).
const patternStats = { removed: 0 };
body.tools = body.tools.filter((tool) => {
if (!tool || typeof tool !== "object" || Array.isArray(tool)) return false;
const type = typeof tool.type === "string" ? tool.type : "";
@@ -80,6 +89,9 @@ function normalizeCodexTools(body) {
for (const st of tool.tools) {
const n = typeof st?.name === "string" ? st.name.trim().slice(0, 128) : "";
if (n) validNames.add(n);
if (st?.parameters && typeof st.parameters === "object") {
st.parameters = stripCodexUnsupportedPatterns(st.parameters, patternStats);
}
}
}
return true;
@@ -101,10 +113,13 @@ function normalizeCodexTools(body) {
tool.type = "function";
tool.name = name.slice(0, 128);
if (description) tool.description = description;
tool.parameters = parameters;
tool.parameters = stripCodexUnsupportedPatterns(parameters, patternStats);
validNames.add(name);
return true;
});
if (patternStats.removed > 0) {
dbg("CODEX", `stripped ${patternStats.removed} unsupported tool schema pattern(s)`);
}
// Drop tool_choice if it references an unknown function name
if (body.tool_choice && typeof body.tool_choice === "object" && !Array.isArray(body.tool_choice)) {
if (body.tool_choice.type === "function") {
@@ -128,6 +143,7 @@ function resolveCacheSessionId(body, credentials) {
function normalizeReasoningEffort(model, value) {
const supportedLevels = getThinkingLevels("codex", model);
if (supportedLevels?.includes(value)) return value;
if (isCodexResponsesLiteModel(model) && (value === "none" || value === "minimal")) return "low";
if (value === "ultra" && supportedLevels?.includes("max")) return "max";
if (value === "max" || value === "ultra") return "xhigh";
return value;
@@ -199,8 +215,11 @@ export class CodexExecutor extends BaseExecutor {
* Override headers to add codex-specific identity headers.
* transformRequest runs BEFORE buildHeaders, sets this._currentSessionId.
*/
buildHeaders(credentials, stream = true) {
buildHeaders(credentials, stream = true, _url = null, model = null) {
const headers = super.buildHeaders(credentials, stream);
if (isCodexResponsesLiteModel(model && getModelUpstreamId("cx", model))) {
headers["x-openai-internal-codex-responses-lite"] = "true";
}
headers["session_id"] = this._currentSessionId || credentials?.connectionId || "default";
// Identify client type to Codex backend (matches official codex CLI)
if (!headers["originator"]) headers["originator"] = "codex_cli_rs";
@@ -398,6 +417,8 @@ export class CodexExecutor extends BaseExecutor {
// Convert string input to array format (Codex API requires input as array)
const normalized = normalizeResponsesInput(body.input);
if (normalized) body.input = normalized;
const upstreamModel = getModelUpstreamId("cx", body.model || model);
const responsesLite = isCodexResponsesLiteModel(upstreamModel);
// Ensure input is present and non-empty (Codex API rejects empty input)
if (!body.input || (Array.isArray(body.input) && body.input.length === 0)) {
@@ -407,7 +428,7 @@ export class CodexExecutor extends BaseExecutor {
// Keep system prompts in body.input as role=developer so they stay in the cacheable prefix
convertSystemToDeveloperRole(body);
// Strip server-generated item IDs (rs_/fc_/resp_/msg_) — Codex /responses can't resolve when store=false
stripStoredItemReferences(body);
stripStoredItemReferences(body, responsesLite);
// Flatten function tools + drop unsupported types
normalizeCodexTools(body);
@@ -415,7 +436,7 @@ export class CodexExecutor extends BaseExecutor {
body.stream = true;
// If no instructions provided, inject default Codex instructions
if (!body.instructions || body.instructions.trim() === "") {
if (!responsesLite && (!body.instructions || body.instructions.trim() === "")) {
body.instructions = CODEX_DEFAULT_INSTRUCTIONS;
}
@@ -428,7 +449,29 @@ export class CodexExecutor extends BaseExecutor {
}
// Map virtual Codex review models to the upstream Codex model before suffix parsing.
body.model = getModelUpstreamId("cx", body.model || model);
body.model = upstreamModel;
if (responsesLite) {
// Codex 0.155 carries tools and instructions as input prefix items.
const input = Array.isArray(body.input) ? body.input : [body.input];
const hasLitePrefix = input.some((item) => item?.type === "additional_tools");
if (!hasLitePrefix) {
const instructions = typeof body.instructions === "string" && body.instructions.trim()
? body.instructions : CODEX_DEFAULT_INSTRUCTIONS;
const prefix = [{ type: "additional_tools", role: "developer", tools: Array.isArray(body.tools) ? body.tools : [] }];
if (instructions) {
prefix.push({ type: "message", role: "developer", content: [{ type: "input_text", text: instructions }] });
}
input.unshift(...prefix);
}
body.input = input;
body.instructions = "";
body.tools = null;
body.tool_choice ||= "auto";
body.parallel_tool_calls = false;
} else {
delete body.parallel_tool_calls;
}
// Extract thinking level from model name suffix
// e.g., gpt-5.3-codex-high → high, gpt-5.3-codex → medium (default)
@@ -445,12 +488,13 @@ export class CodexExecutor extends BaseExecutor {
// Priority: explicit reasoning.effort > reasoning_effort param > model suffix > default (medium)
if (!body.reasoning) {
const effort = normalizeReasoningEffort(body.model, body.reasoning_effort || modelEffort || 'low');
body.reasoning = { effort, summary: "auto" };
const effort = normalizeReasoningEffort(body.model, body.reasoning_effort || modelEffort || (responsesLite ? 'medium' : 'low'));
body.reasoning = responsesLite ? { effort } : { effort, summary: "auto" };
} else {
body.reasoning.effort = normalizeReasoningEffort(body.model, body.reasoning.effort);
if (!body.reasoning.summary) body.reasoning.summary = "auto";
if (!responsesLite && !body.reasoning.summary) body.reasoning.summary = "auto";
}
if (responsesLite) body.reasoning.context = "all_turns";
delete body.reasoning_effort;
// Include reasoning encrypted content (required by Codex backend for reasoning models)

View File

@@ -46,203 +46,240 @@ export class CommandCodeExecutor extends BaseExecutor {
return headers;
}
async execute(opts) {
const result = await super.execute(opts);
if (!result?.response?.ok || !result.response.body) return result;
result.response = await peekForUpstreamError(result.response, opts.model, {
signal: opts.signal,
});
return result;
}
async execute(opts) {
const maxRetries = 2;
for (let attempt = 0; attempt <= maxRetries; attempt++) {
const result = await super.execute(opts);
if (!result?.response?.ok || !result.response.body) return result;
const wrappedResponse = await inspectAndWrapCommandCodeResponse(result.response, opts.model);
if (!wrappedResponse.ok && attempt < maxRetries) {
const isRetryableStatus = wrappedResponse.status === 502 || wrappedResponse.status === 503 || wrappedResponse.status === 504;
if (isRetryableStatus) {
opts.log?.debug?.("RETRY", `CommandCode upstream returned status ${wrappedResponse.status}, retrying ${attempt + 1}/${maxRetries}...`);
await new Promise(r => setTimeout(r, 1000 * (attempt + 1)));
continue;
}
}
result.response = wrappedResponse;
return result;
}
}
parseError(response, bodyText) {
let parsed = null;
try {
parsed = JSON.parse(bodyText || "{}");
} catch {
parsed = null;
}
const errObj = parsed?.error || parsed;
const msg = errObj?.message || parsed?.message || bodyText || response.statusText;
const status = Number(errObj?.code || errObj?.statusCode || response.status) || response.status;
return {
status,
message: msg || `CommandCode upstream error: ${response.status}`,
};
}
}
// How long to hold the response open while peeking the first upstream events.
// An upstream error event ("Network connection lost") is emitted at stream
// start, so the peek is fast; the bound just prevents a slow-started stream
// from being held hostage. Env: COMMANDCODE_PEEK_TIMEOUT_MS.
const PEEK_TIMEOUT_MS = (() => {
const raw = process.env.COMMANDCODE_PEEK_TIMEOUT_MS;
const n = raw ? parseInt(raw, 10) : NaN;
return Number.isFinite(n) && n > 0 ? n : 10 * 1000;
})();
export function parseCommandCodeError(event) {
if (!event || typeof event !== "object") {
return {
statusCode: 503,
message: "CommandCode upstream error",
type: "server_error",
};
}
// Event types that count as "the stream has started producing". Everything
// else (start, start-step, reasoning-start, text-start, ...) is metadata and
// does not end the peek.
const MEANINGFUL_EVENT_TYPES = new Set([
"text-delta",
"reasoning-delta",
"tool-input-start",
"tool-input-delta",
"tool-input-end",
"tool-call",
"finish-step",
"finish",
]);
const errVal = event.error ?? event.message ?? "unknown";
let message = "";
let statusCode = null;
let type = "server_error";
function makeAbortError(reason) {
const error = new Error(reason?.message || reason || "Request aborted");
error.name = "AbortError";
return error;
if (typeof errVal === "object" && errVal !== null) {
message = errVal.message || errVal.error || JSON.stringify(errVal);
if (errVal.statusCode && Number.isInteger(Number(errVal.statusCode))) {
statusCode = Number(errVal.statusCode);
} else if (errVal.status && Number.isInteger(Number(errVal.status))) {
statusCode = Number(errVal.status);
}
if (errVal.type) type = errVal.type;
} else if (typeof errVal === "string") {
message = errVal;
} else {
message = JSON.stringify(errVal);
}
if (event.statusCode && Number.isInteger(Number(event.statusCode))) {
statusCode = Number(event.statusCode);
}
if (!statusCode || statusCode < 400 || statusCode > 599) {
const lower = message.toLowerCase();
if (lower.includes("rate limit") || lower.includes("too many requests")) {
statusCode = 429;
type = "rate_limit_error";
} else if (lower.includes("unauthorized") || lower.includes("invalid api key") || lower.includes("authentication")) {
statusCode = 401;
type = "authentication_error";
} else if (lower.includes("payment required") || lower.includes("billing")) {
statusCode = 402;
type = "billing_error";
} else if (lower.includes("quota") || lower.includes("forbidden") || lower.includes("permission")) {
statusCode = 403;
type = "permission_error";
} else if (lower.includes("not found")) {
statusCode = 404;
type = "invalid_request_error";
} else if (lower.includes("unavailable") || lower.includes("overloaded") || lower.includes("server error")) {
statusCode = 503;
type = "server_error";
} else {
statusCode = 503;
}
}
return { statusCode, message, type };
}
function tryParseEvent(line) {
const trimmed = line.trim();
if (!trimmed) return null;
const json = trimmed.startsWith("data:") ? trimmed.slice(5).trim() : trimmed;
if (!json || json === "[DONE]") return null;
try {
return JSON.parse(json);
} catch {
return null;
}
export async function inspectAndWrapCommandCodeResponse(originalResponse, model) {
const reader = originalResponse.body.getReader();
const decoder = new TextDecoder();
let buffer = "";
const rawChunks = [];
let detectedError = null;
try {
while (true) {
const { value, done } = await reader.read();
if (done) {
const trimmed = buffer.trim();
if (trimmed) {
try {
const jsonStr = trimmed.startsWith("data:") ? trimmed.slice(5).trim() : trimmed;
const parsed = JSON.parse(jsonStr);
if (parsed?.type === "error") {
detectedError = parsed;
}
} catch {
/* ignore */
}
}
break;
}
rawChunks.push(value);
buffer += decoder.decode(value, { stream: true });
const lines = buffer.split("\n");
buffer = lines.pop() || "";
let stopLoop = false;
for (const line of lines) {
const trimmed = line.trim();
if (!trimmed) continue;
const jsonStr = trimmed.startsWith("data:") ? trimmed.slice(5).trim() : trimmed;
if (!jsonStr || jsonStr === "[DONE]") {
stopLoop = true;
break;
}
let event;
try {
event = JSON.parse(jsonStr);
} catch {
continue;
}
if (event?.type === "error") {
detectedError = event;
stopLoop = true;
break;
}
if (
event?.type === "text-delta" ||
event?.type === "reasoning-delta" ||
event?.type === "tool-input-start" ||
event?.type === "tool-call" ||
event?.type === "finish" ||
event?.type === "finish-step"
) {
stopLoop = true;
break;
}
}
if (stopLoop) break;
}
} catch {
try { reader.releaseLock(); } catch { /* ignore */ }
return originalResponse;
}
if (detectedError) {
try { await reader.cancel(); } catch { /* ignore */ }
const { statusCode, message, type } = parseCommandCodeError(detectedError);
return new Response(
JSON.stringify({
error: {
message: `[CommandCode error: ${message}]`,
type,
code: statusCode,
},
}),
{
status: statusCode,
statusText: statusCode === 503 ? "Service Unavailable" : (statusCode === 429 ? "Too Many Requests" : "Bad Gateway"),
headers: {
"Content-Type": "application/json",
"Access-Control-Allow-Origin": "*",
},
}
);
}
const combinedStream = createRawReplayedStream(rawChunks, reader);
return wrapNdjsonAsOpenAISse(combinedStream, model, originalResponse);
}
function formatErrorValue(errVal) {
const errStr =
typeof errVal === "string"
? errVal
: typeof errVal?.message === "string"
? errVal.message
: JSON.stringify(errVal);
const errType =
typeof errVal === "string"
? "upstream_error"
: errVal?.type || "upstream_error";
return { message: errStr, type: errType };
function createRawReplayedStream(rawChunks, reader) {
let chunkIndex = 0;
return new ReadableStream({
async pull(controller) {
if (chunkIndex < rawChunks.length) {
controller.enqueue(rawChunks[chunkIndex++]);
return;
}
try {
const { value, done } = await reader.read();
if (done) {
controller.close();
} else {
controller.enqueue(value);
}
} catch (err) {
controller.error(err);
}
},
async cancel(reason) {
try {
await reader.cancel(reason);
} catch {
/* ignore */
}
},
});
}
/**
* Read the first upstream events before committing the response.
*
* - `{"type":"error"}` as the first meaningful event → return a 502 Response so
* chatCore's `!response.ok` path parses the error and triggers fallback.
* - Otherwise → re-emit the buffered bytes + the rest of the stream through the
* normal NDJSON → OpenAI SSE wrapper and return it untouched in spirit.
*
* Bounded by `timeoutMs` (default PEEK_TIMEOUT_MS): if no meaningful event
* arrives in time, or the request signal aborts, we commit whatever we have and
* let the regular stream pipeline (stall detection, abort handling) take over.
*/
export async function peekForUpstreamError(
originalResponse,
model,
{ signal = null, timeoutMs = PEEK_TIMEOUT_MS } = {},
) {
const reader = originalResponse.body.getReader();
const decoder = new TextDecoder();
const abortController = new AbortController();
const forwardAbort = () => abortController.abort(signal?.reason);
if (signal?.aborted) abortController.abort(signal?.reason);
else if (signal)
signal.addEventListener("abort", forwardAbort, { once: true });
// Raw bytes for lossless re-emission; decoded text is only used for line
// parsing / error detection. Never re-encode decoded text: TextDecoder
// holds a split multi-byte char internally and flush() would replace it
// with U+FFFD, corrupting the stream.
const rawChunks = [];
let peeked = "";
let errorEvent = null;
let committed = false;
const readWithTimeout = (ms) => {
if (abortController.signal.aborted) {
return Promise.reject(makeAbortError(abortController.signal.reason));
}
const timeoutPromise = new Promise((_, reject) => {
const t = setTimeout(() => reject(new Error("peek timeout")), ms);
t.unref?.();
});
const abortPromise = new Promise((_, reject) => {
abortController.signal.addEventListener(
"abort",
() => reject(makeAbortError(abortController.signal.reason)),
{ once: true },
);
});
return Promise.race([reader.read(), timeoutPromise, abortPromise]);
};
try {
const deadline = Date.now() + timeoutMs;
while (!errorEvent && !committed && Date.now() < deadline) {
const { done, value } = await readWithTimeout(
Math.max(deadline - Date.now(), 1),
);
if (done) break;
rawChunks.push(value);
peeked += decoder.decode(value, { stream: true });
const lines = peeked.split("\n");
// The last segment may be a partial line — only parse complete ones.
for (const line of lines.slice(0, -1)) {
const event = tryParseEvent(line);
if (!event?.type) continue;
if (event.type === "error") {
errorEvent = event;
break;
}
if (MEANINGFUL_EVENT_TYPES.has(event.type)) {
committed = true;
break;
}
}
}
} catch {
// timeout / abort / read failure during the peek → commit whatever we have;
// the downstream stream pipeline (stall detection, abort handling) takes over.
}
if (signal) signal.removeEventListener("abort", forwardAbort);
if (errorEvent) {
await reader.cancel("commandcode early error detected").catch(() => {});
const { message, type } = formatErrorValue(
errorEvent.error ?? errorEvent.message ?? "unknown",
);
return new Response(JSON.stringify({ error: { message, type } }), {
status: HTTP_STATUS.BAD_GATEWAY,
statusText: message.slice(0, 200),
headers: { "Content-Type": "application/json" },
});
}
const remaining = new ReadableStream({
start(controller) {
(async () => {
try {
// Re-emit RAW bytes (never re-encoded decoded text) so split
// multi-byte UTF-8 sequences survive the peek untouched.
for (const c of rawChunks) controller.enqueue(c);
while (true) {
const { done, value } = await reader.read();
if (done) break;
controller.enqueue(value);
}
controller.close();
} catch (err) {
controller.error(err);
}
})();
},
cancel() {
reader.cancel("commandcode stream cancelled").catch(() => {});
},
});
const combined = new Response(remaining, {
status: originalResponse.status,
statusText: originalResponse.statusText,
headers: originalResponse.headers,
});
return wrapNdjsonAsOpenAISse(combined, model);
}
function wrapNdjsonAsOpenAISse(originalResponse, model) {
const decoder = new TextDecoder();
const encoder = new TextEncoder();
let buffer = "";
const state = { model };
function wrapNdjsonAsOpenAISse(streamBody, model, originalResponse = null) {
const decoder = new TextDecoder();
const encoder = new TextEncoder();
let buffer = "";
const state = { model };
const emitChunks = (chunks, controller) => {
if (!chunks) return;
@@ -253,33 +290,38 @@ function wrapNdjsonAsOpenAISse(originalResponse, model) {
}
};
const transform = new TransformStream({
transform(chunk, controller) {
buffer += decoder.decode(chunk, { stream: true });
const lines = buffer.split("\n");
buffer = lines.pop() || "";
for (const line of lines) {
const trimmed = line.trim();
if (!trimmed) continue;
// Translate AI SDK v5 NDJSON line to one or more OpenAI chunks
emitChunks(commandCodeToOpenAIResponse(trimmed, state), controller);
}
},
flush(controller) {
const trimmed = buffer.trim();
if (trimmed) {
emitChunks(commandCodeToOpenAIResponse(trimmed, state), controller);
}
controller.enqueue(encoder.encode(SSE_DONE));
},
});
const transform = new TransformStream({
transform(chunk, controller) {
buffer += decoder.decode(chunk, { stream: true });
const lines = buffer.split("\n");
buffer = lines.pop() || "";
for (const line of lines) {
const trimmed = line.trim();
if (!trimmed) continue;
emitChunks(commandCodeToOpenAIResponse(trimmed, state), controller);
}
},
flush(controller) {
const trimmed = buffer.trim();
if (trimmed) {
emitChunks(commandCodeToOpenAIResponse(trimmed, state), controller);
}
controller.enqueue(encoder.encode(SSE_DONE));
},
});
const newBody = originalResponse.body.pipeThrough(transform);
return new Response(newBody, {
status: originalResponse.status,
statusText: originalResponse.statusText,
headers: originalResponse.headers,
});
const newBody = streamBody.pipeThrough(transform);
return new Response(newBody, {
status: originalResponse?.status || 200,
statusText: originalResponse?.statusText || "OK",
headers: {
"Content-Type": "text/event-stream",
"Cache-Control": "no-cache",
"Connection": "keep-alive",
...(originalResponse?.headers ? Object.fromEntries(originalResponse.headers.entries()) : {}),
"content-type": "text/event-stream",
},
});
}
export default CommandCodeExecutor;

View File

@@ -7,13 +7,16 @@ import {
wrapConnectRPCFrame,
decodeMessage,
parseConnectRPCFrame,
extractTextFromResponse
extractTextFromResponse,
encodeMcpTools,
decodeMcpArgs,
} from "../utils/cursorProtobuf.js";
import { buildCursorHeaders } from "../utils/cursorChecksum.js";
import { estimateUsage } from "../utils/usageTracking.js";
import { SSE_DONE, SSE_HEADERS } from "../utils/sseConstants.js";
import { chatChunkSse, sseChunk } from "../utils/sse.js";
import { FORMATS } from "../translator/formats.js";
import { ROLE, OPENAI_BLOCK } from "../translator/schema/index.js";
import { proxyAwareFetch } from "../utils/proxyFetch.js";
import zlib from "zlib";
import crypto from "crypto";
@@ -65,55 +68,74 @@ function textFromContent(content) {
if (typeof content === "string") return content;
if (!Array.isArray(content)) return "";
return content
.filter((part) => part?.type === "text" && typeof part.text === "string")
.filter((part) => part?.type === OPENAI_BLOCK.TEXT && typeof part.text === "string")
.map((part) => part.text)
.join("\n");
}
function isAgentTextRequest(body) {
// Many compatible clients always attach their built-in tool schemas, even
// for a normal text turn. Cursor's retired ChatService rejects those
// requests; AgentService can still answer the text turn, so ignore schemas
// here. A real tool-call/result conversation is kept on the legacy path
// until its AgentService tool protocol is implemented.
return Array.isArray(body?.messages) && body.messages.every((message) => {
if (message?.tool_calls?.length || message?.role === "tool") return false;
return typeof message?.content === "string"
|| Array.isArray(message?.content) && message.content.every((part) => part?.type === "text");
function isTextPart(part) {
return !part || part.type === OPENAI_BLOCK.TEXT || typeof part === "string";
}
export function isAgentCapableRequest(body) {
// ChatService rejects auto/composer and most thinking variants. AgentService
// can answer text turns (including declared tool schemas) and tool-call
// history. Image parts still need the legacy protobuf path.
if (!Array.isArray(body?.messages) || body.messages.length === 0) return false;
return body.messages.every((message) => {
if (Array.isArray(message?.content)) return message.content.every(isTextPart);
return message?.content == null || typeof message.content === "string";
});
}
function encodeHistoryMessage(message) {
const content = textFromContent(message?.content);
if (!content) return null;
const extras = [];
if (message?.role === ROLE.ASSISTANT && message.tool_calls?.length) {
for (const tc of message.tool_calls) {
extras.push(`[tool_call id=${tc.id || ""} name=${tc.function?.name || "tool"} args=${tc.function?.arguments || "{}"}]`);
}
}
if (message?.role === ROLE.TOOL) {
extras.push(`[tool_result id=${message.tool_call_id || ""}]`);
}
const textBody = [content, ...extras].filter(Boolean).join("\n");
if (!textBody) return null;
// ConversationHistoryMessage.user / .assistant -> repeated content -> text.
const text = agentString(1, content);
if (message.role === "assistant") {
const text = agentString(1, textBody);
if (message.role === ROLE.ASSISTANT) {
return agentMessage(2, agentMessage(1, agentMessage(1, text)));
}
return agentMessage(1, agentMessage(1, agentMessage(1, text)));
}
function buildAgentRunFrame(messages, model) {
export function buildAgentRunFrame(messages, model, tools = []) {
// custom_system_prompt (RunRequest field 8) makes AgentService return an
// empty turn. Fold system text into the current user message instead.
const system = messages
.filter((message) => message?.role === "system")
.filter((message) => message?.role === ROLE.SYSTEM)
.map((message) => textFromContent(message.content))
.filter(Boolean)
.join("\n\n");
const chatMessages = messages.filter((message) => message?.role !== "system");
const currentIndex = [...chatMessages].map((message) => message?.role).lastIndexOf("user");
const chatMessages = messages.filter((message) => message?.role !== ROLE.SYSTEM);
const currentIndex = [...chatMessages].map((message) => message?.role).lastIndexOf(ROLE.USER);
const current = currentIndex >= 0 ? chatMessages[currentIndex] : chatMessages.at(-1);
const history = chatMessages
.slice(0, currentIndex >= 0 ? currentIndex : -1)
.map(encodeHistoryMessage)
.filter(Boolean);
const userText = textFromContent(current?.content) || "Continue.";
const rawUser = textFromContent(current?.content) || "Continue.";
const userText = system ? `${system}\n\n${rawUser}` : rawUser;
// agent.v1.UserMessageAction.user_message and its optional history.
// selected_context (3) + mode=1 (4) match cursor-agent's wire format; without
// them the server may accept the RPC and stream an empty turn.
const userMessage = concatBuffers(
agentString(1, userText),
agentString(2, crypto.randomUUID()),
agentMessage(3, new Uint8Array()),
encodeField(4, PROTOBUF_VARINT, 1),
);
const conversationHistory = history.length
? concatBuffers(...history.map((entry) => agentMessage(1, entry)))
@@ -124,11 +146,20 @@ function buildAgentRunFrame(messages, model) {
);
const conversationAction = agentMessage(1, userAction);
const requestedModel = concatBuffers(agentString(1, model), agentBool(7, true));
// ModelDetails (field 3): thinking variants (Composer, Grok, *-thinking)
// return an empty turn when only RequestedModel (field 9) is set.
const modelDetails = concatBuffers(
agentString(1, model),
agentString(3, model),
agentString(4, model),
);
const mcpTools = encodeMcpTools(tools);
const runRequest = concatBuffers(
// An empty ConversationStateStructure starts a fresh local agent session.
agentMessage(1, new Uint8Array()),
agentMessage(2, conversationAction),
...(system ? [agentString(8, system)] : []),
agentMessage(3, modelDetails),
...(mcpTools.length ? [agentMessage(4, mcpTools)] : []),
agentMessage(9, requestedModel),
);
@@ -157,13 +188,51 @@ function decodeAgentFrames(buffer, onFrame) {
return pending;
}
function createRequestContextResponse() {
// AgentService asks every run for client context. 9router has no IDE file
// context, so acknowledge with an empty RequestContext.
function execIds(execRequest) {
const id = Number(execRequest?.get(1)?.[0]?.value || 0);
const execId = extractAgentString(execRequest, 15);
return { id, execId };
}
function wrapExecClientMessage(execMsgId, execId, resultField, resultPayload) {
const parts = [];
if (execMsgId) parts.push(encodeField(1, PROTOBUF_VARINT, execMsgId));
parts.push(agentString(15, execId || ""));
parts.push(encodeField(resultField, PROTOBUF_LEN, resultPayload || new Uint8Array()));
return wrapConnectRPCFrame(agentMessage(2, concatBuffers(...parts)));
}
function createRequestContextResponse(execRequest) {
// Tools already go out on AgentRunRequest.mcp_tools. Echoing them again on
// this ack makes AgentService stall silently (0 SSE bytes until abort).
const { id, execId } = execIds(execRequest);
const requestContextSuccess = agentMessage(1, new Uint8Array());
const requestContextResult = agentMessage(1, requestContextSuccess);
const execClientMessage = agentMessage(10, requestContextResult);
return wrapConnectRPCFrame(agentMessage(2, execClientMessage));
return wrapExecClientMessage(id, execId, 10, requestContextResult);
}
// ExecServerMessage variant → ExecClientMessage result field (same numbers).
const EXEC_RESULT_FIELD = {
2: 2, 3: 3, 4: 4, 5: 5, 7: 7, 8: 8, 9: 9, 16: 16, 20: 20, 23: 23,
};
function rejectExecRequest(execRequest) {
const { id, execId } = execIds(execRequest);
const variant = [...(execRequest?.keys?.() || [])].find((field) => field !== 1 && field !== 15);
const resultField = EXEC_RESULT_FIELD[variant];
if (!resultField) return null;
// Diagnostics has no rejected variant — empty success unblocks the stream.
if (variant === 9) return wrapExecClientMessage(id, execId, 9, new Uint8Array());
const rejected = agentMessage(2, agentString(2, "Tool not available in this environment. Use the MCP tools provided instead."));
return wrapExecClientMessage(id, execId, resultField, rejected);
}
function encodeKvClientMessage(kvId, resultField, resultPayload, metadata) {
const parts = [];
if (kvId) parts.push(encodeField(1, PROTOBUF_VARINT, kvId));
parts.push(encodeField(resultField, PROTOBUF_LEN, resultPayload || new Uint8Array()));
if (metadata && metadata.length) parts.push(encodeField(4, PROTOBUF_LEN, metadata));
return wrapConnectRPCFrame(agentMessage(3, concatBuffers(...parts)));
}
const CURSOR_STREAM_DEBUG = process.env.CURSOR_STREAM_DEBUG === "1";
@@ -479,7 +548,7 @@ export class CursorExecutor extends BaseExecutor {
};
}
async executeAgent({ model, body, stream, credentials, signal }) {
async executeAgent({ model, body, stream, credentials, signal, log }) {
const agentEndpoint = PROVIDER_OAUTH.cursor?.agentEndpoint;
if (!agentEndpoint) throw new Error("Cursor AgentService endpoint is not configured");
@@ -491,9 +560,10 @@ export class CursorExecutor extends BaseExecutor {
}
let session;
const tools = body.tools || [];
try {
session = this.openAgentHttp2Stream(url, headers, requestController.signal);
session.write(buildAgentRunFrame(body.messages || [], model));
session.write(buildAgentRunFrame(body.messages || [], model, tools));
} catch (error) {
throw new Error(`Cursor AgentService request failed: ${error.message}`);
}
@@ -533,8 +603,23 @@ export class CursorExecutor extends BaseExecutor {
// so strict clients such as Claude Code accept the completed stream.
const responseId = `chatcmpl-msg_${Date.now()}`;
const created = Math.floor(Date.now() / 1000);
const composerModel = isComposerModel(model);
let pending = Buffer.alloc(0);
let finished = false;
let thinkingAcc = "";
let emittedVisible = 0;
let emittedText = false;
const flushThinkingFallback = (onEvent) => {
if (emittedText || !thinkingAcc) return;
const fallback = composerModel
? visibleComposerContentFromThinking(thinkingAcc)
: thinkingAcc.trim();
if (fallback) {
emittedText = true;
onEvent({ type: "text", value: fallback });
}
};
const consume = async (onEvent) => {
try {
@@ -553,32 +638,87 @@ export class CursorExecutor extends BaseExecutor {
const update = decodeMessage(serverMessage.get(1)[0].value);
if (update.has(1)) {
const textDelta = extractAgentString(decodeMessage(update.get(1)[0].value), 1);
if (textDelta) onEvent({ type: "text", value: textDelta });
if (textDelta) {
emittedText = true;
onEvent({ type: "text", value: textDelta });
}
}
// Cursor's AgentService emits internal reasoning without the
// cryptographic signature required by Anthropic thinking blocks.
// Forwarding it makes strict Anthropic clients (Claude Code)
// discard or wait on an otherwise complete response. Keep the
// reasoning upstream-only and emit the normal answer text.
// thinking_delta (field 4). Composer (and some Grok variants) put
// the visible answer after </think> here and never send text_delta.
if (update.has(4)) {
const thinkingDelta = extractAgentString(decodeMessage(update.get(4)[0].value), 1);
if (thinkingDelta) {
thinkingAcc += thinkingDelta;
if (composerModel) {
const visible = visibleComposerContentFromThinking(thinkingAcc);
if (visible.length > emittedVisible) {
const deltaContent = visible.slice(emittedVisible);
emittedVisible = visible.length;
emittedText = true;
onEvent({ type: "text", value: deltaContent });
}
}
}
}
// Keep unsigned reasoning upstream-only for Anthropic clients.
if (update.has(14)) {
flushThinkingFallback(onEvent);
finished = true;
onEvent({ type: "done" });
}
}
// KvServerMessage (field 4): get/set blob. Ack so the stream proceeds.
if (serverMessage.has(4)) {
const kv = decodeMessage(serverMessage.get(4)[0].value);
const kvId = kv.get(1)?.[0]?.value || 0;
const metadata = kv.get(4)?.[0]?.value || null;
if (kv.has(2)) {
session.write(encodeKvClientMessage(kvId, 2, agentMessage(1, new Uint8Array()), metadata));
} else if (kv.has(3)) {
session.write(encodeKvClientMessage(kvId, 3, new Uint8Array(), metadata));
}
}
// AgentService requests IDE context before producing a response.
// Return an empty context; 9router is not coupled to an editor.
if (serverMessage.has(2)) {
const execRequest = decodeMessage(serverMessage.get(2)[0].value);
if (execRequest.has(10)) {
session.write(createRequestContextResponse());
log?.info?.("CURSOR", "AgentService request_context ack");
session.write(createRequestContextResponse(execRequest));
} else if (execRequest.has(11)) {
const mcp = decodeMcpArgs(execRequest.get(11)[0].value);
const name = mcp.toolName || mcp.name;
if (name) {
log?.info?.("CURSOR", `AgentService MCP tool_call ${name}`);
finished = true;
onEvent({
type: "tool_call",
value: {
id: mcp.toolCallId || `call_${crypto.randomUUID()}`,
name,
arguments: JSON.stringify(mcp.args || {}),
},
});
onEvent({ type: "done", finishReason: "tool_calls" });
} else {
debugLog(`[CURSOR AGENT] Unsupported exec request fields: ${[...execRequest.keys()].join(",")}`);
finished = true;
onEvent({ type: "error", value: "Cursor AgentService requested an unsupported IDE tool" });
}
} else {
// Every other ExecServerMessage variant is an editor-backed tool
// (shell, read, write, …) that 9router cannot service. Fail the
// turn rather than narrating protocol state as assistant text.
debugLog(`[CURSOR AGENT] Unsupported exec request fields: ${[...execRequest.keys()].join(",")}`);
finished = true;
onEvent({ type: "error", value: "Cursor AgentService requested an unsupported IDE tool" });
// Auto/Composer often probe IDE builtins (shell/read/…). Reject
// them so the model can continue with MCP tools or a text answer
// instead of stalling the h2 stream.
const rejection = rejectExecRequest(execRequest);
if (rejection) {
log?.info?.("CURSOR", `AgentService rejected IDE exec fields=${[...execRequest.keys()].join(",")}`);
session.write(rejection);
} else {
debugLog(`[CURSOR AGENT] Unsupported exec request fields: ${[...execRequest.keys()].join(",")}`);
finished = true;
onEvent({ type: "error", value: "Cursor AgentService requested an unsupported IDE tool" });
}
}
}
});
@@ -586,7 +726,10 @@ export class CursorExecutor extends BaseExecutor {
} finally {
try { session.end(); } catch {}
try { session.close(); } catch {}
if (!finished) onEvent({ type: "done" });
if (!finished) {
flushThinkingFallback(onEvent);
onEvent({ type: "done" });
}
}
};
@@ -594,10 +737,21 @@ export class CursorExecutor extends BaseExecutor {
let content = "";
let reasoning = "";
let agentError = null;
const toolCalls = [];
let finishReason = "stop";
await consume((event) => {
if (event.type === "text") content += event.value;
else if (event.type === "thinking") reasoning += event.value;
else if (event.type === "tool_call") {
toolCalls.push({
id: event.value.id,
type: "function",
function: { name: event.value.name, arguments: event.value.arguments },
});
finishReason = "tool_calls";
}
else if (event.type === "error") agentError = event.value;
else if (event.type === "done" && event.finishReason) finishReason = event.finishReason;
});
if (agentError) {
return {
@@ -611,13 +765,19 @@ export class CursorExecutor extends BaseExecutor {
responseFormat: FORMATS.OPENAI,
};
}
const message = {
role: "assistant",
content: content || null,
...(reasoning ? { reasoning_content: reasoning } : {}),
...(toolCalls.length ? { tool_calls: toolCalls } : {}),
};
return {
response: new Response(JSON.stringify({
id: responseId,
object: "chat.completion",
created,
model,
choices: [{ index: 0, message: { role: "assistant", content: content || null, ...(reasoning ? { reasoning_content: reasoning } : {}) }, finish_reason: "stop" }],
choices: [{ index: 0, message, finish_reason: finishReason }],
usage: estimateUsage(body, content.length, FORMATS.OPENAI),
}), { headers: { "Content-Type": "application/json" } }),
url,
@@ -635,6 +795,18 @@ export class CursorExecutor extends BaseExecutor {
controller.enqueue(encoder.encode(chatChunkSse({ id: responseId, created, model, delta: { content: event.value } })));
} else if (event.type === "thinking") {
controller.enqueue(encoder.encode(chatChunkSse({ id: responseId, created, model, delta: { reasoning_content: event.value } })));
} else if (event.type === "tool_call") {
controller.enqueue(encoder.encode(chatChunkSse({
id: responseId, created, model,
delta: {
tool_calls: [{
index: 0,
id: event.value.id,
type: "function",
function: { name: event.value.name, arguments: event.value.arguments },
}],
},
})));
} else if (event.type === "error") {
// An SSE error frame, not a content delta: a protocol failure must not
// be rendered to the user as the assistant's reply, and downstream
@@ -643,7 +815,10 @@ export class CursorExecutor extends BaseExecutor {
controller.enqueue(encoder.encode(SSE_DONE));
controller.close();
} else if (event.type === "done") {
controller.enqueue(encoder.encode(chatChunkSse({ id: responseId, created, model, delta: {}, finishReason: "stop" })));
controller.enqueue(encoder.encode(chatChunkSse({
id: responseId, created, model, delta: {},
finishReason: event.finishReason || "stop",
})));
controller.enqueue(encoder.encode(SSE_DONE));
controller.close();
}
@@ -664,9 +839,9 @@ export class CursorExecutor extends BaseExecutor {
}
async execute({ model, body, stream, credentials, signal, log, proxyOptions = null }) {
if (isAgentTextRequest(body)) {
if (isAgentCapableRequest(body)) {
try {
return await this.executeAgent({ model, body, stream, credentials, signal });
return await this.executeAgent({ model, body, stream, credentials, signal, log });
} catch (error) {
return {
response: new Response(JSON.stringify({

View File

@@ -1,12 +1,13 @@
import { BaseExecutor } from "./base.js";
import { PROVIDERS, PROVIDER_OAUTH } from "../config/providers.js";
import { ANTHROPIC_API_VERSION, OPENAI_COMPAT_BASE, ANTHROPIC_COMPAT_BASE, selectAnthropicBeta } from "../providers/shared.js";
import { ANTHROPIC_API_VERSION, OPENAI_COMPAT_BASE, ANTHROPIC_COMPAT_BASE, selectAnthropicBeta, mergeAnthropicBeta } from "../providers/shared.js";
import { resolveOpenAICompatibleApiType } from "../services/provider.js";
import { OAUTH_ENDPOINTS, buildKimiHeaders } from "../config/appConstants.js";
import { buildClineHeaders } from "../shared/clineAuth.js";
import { proxyAwareFetch } from "../utils/proxyFetch.js";
import { injectReasoningContent } from "../utils/reasoningContentInjector.js";
import { stripUnsupportedParams } from "../translator/concerns/paramSupport.js";
import { extractClaudeSessionIdFromUserId } from "../utils/claudeCloaking.js";
// Auth header descriptors — derived from registry transport.auth, fallback to hardcoded defaults.
const BEARER = { combined: true, header: "Authorization", scheme: "bearer" };
@@ -146,7 +147,7 @@ export class DefaultExecutor extends BaseExecutor {
return BEARER;
}
buildHeaders(credentials, stream = true, url, model) {
buildHeaders(credentials, stream = true, url, model, body = null) {
const rt = credentials?.runtimeTransport;
const headers = { "Content-Type": "application/json", ...(rt ? rt.headers : this.config.headers) };
const desc = rt?.auth || AUTH_DESCRIPTORS[this.provider] || this.resolveAuthDescriptor();
@@ -154,8 +155,31 @@ export class DefaultExecutor extends BaseExecutor {
for (const hook of desc.hooks || []) HEADER_HOOKS[hook]?.(headers, credentials);
applyAuth(headers, desc, credentials);
if (this.provider === "claude" && model) {
headers["Anthropic-Beta"] = selectAnthropicBeta(model);
// anthropic-compatible-* nodes serving a real Claude model sit in front of
// Anthropic itself (a rotating multi-account proxy, a corporate gateway),
// so the request needs the same beta flags the `claude` provider sends:
// without `context-management-2025-06-27` upstream rejects the
// `context_management` block Claude Code puts in every request with
// "context_management: Extra inputs are not permitted" (HTTP 400), and the
// combo silently falls through to the next model. The model id gates this:
// a node fronting Kimi or GLM answers on its own ids and never matches, so
// gateways that would choke on unknown beta flags are left untouched.
const isClaudeModel = typeof model === "string" && /^claude-/.test(model);
const clientBeta = credentials?.rawHeaders?.["anthropic-beta"];
if (model && (this.provider === "claude"
|| (this.provider?.startsWith?.("anthropic-compatible-") && isClaudeModel))) {
headers["Anthropic-Beta"] = mergeAnthropicBeta(selectAnthropicBeta(model, body), clientBeta);
} else if (this.provider === "anthropic" && clientBeta) {
headers["Anthropic-Beta"] = mergeAnthropicBeta(headers["Anthropic-Beta"], clientBeta);
}
// Claude OAuth: align x-claude-code-session-id with metadata.user_id.session_id if missing
if (this.provider === "claude" && !headers["x-claude-code-session-id"]) {
const token = credentials?.accessToken || credentials?.apiKey || "";
if (token.includes("sk-ant-oat")) {
const sid = extractClaudeSessionIdFromUserId(body?.metadata?.user_id);
if (sid) headers["x-claude-code-session-id"] = sid;
}
}
// Strip first-party Claude Code identity headers for non-Anthropic anthropic-compatible upstreams

View File

@@ -10,12 +10,15 @@ import { CodexExecutor } from "./codex.js";
import { CursorExecutor } from "./cursor.js";
import { VertexExecutor } from "./vertex.js";
import { OpenCodeExecutor } from "./opencode.js";
import { OpenCodeGoExecutor } from "./opencode-go.js";
import { OpenCodeZenExecutor } from "./opencode-zen.js";
import { GrokWebExecutor } from "./grok-web.js";
import { GrokCliExecutor } from "./grok-cli.js";
import { PerplexityWebExecutor } from "./perplexity-web.js";
import { OllamaLocalExecutor } from "./ollama-local.js";
import { CommandCodeExecutor } from "./commandcode.js";
import { XiaomiTokenplanExecutor } from "./xiaomi-tokenplan.js";
import { XiaomiMimoExecutor } from "./xiaomi-mimo.js";
import { MimoFreeExecutor } from "./mimo-free.js";
import { CodeBuddyExecutor } from "./codebuddy-cn.js";
import { CodeBuddyIntlExecutor } from "./codebuddy-intl.js";
@@ -32,6 +35,7 @@ const executors = {
github: new GithubExecutor(),
iflow: new IFlowExecutor(),
qoder: new QoderExecutor(),
"qoder-cn": new QoderExecutor("qoder-cn"),
kiro: new KiroExecutor(),
kimchi: new KimchiExecutor(),
codex: new CodexExecutor(),
@@ -40,6 +44,8 @@ const executors = {
vertex: new VertexExecutor("vertex"),
"vertex-partner": new VertexExecutor("vertex-partner"),
opencode: new OpenCodeExecutor(),
"opencode-go": new OpenCodeGoExecutor(),
"opencode-zen": new OpenCodeZenExecutor(),
"grok-web": new GrokWebExecutor(),
"grok-cli": new GrokCliExecutor(),
gcli: new GrokCliExecutor(), // Alias
@@ -48,6 +54,7 @@ const executors = {
"ollama-local": new OllamaLocalExecutor(),
commandcode: new CommandCodeExecutor(),
"xiaomi-tokenplan": new XiaomiTokenplanExecutor(),
"xiaomi-mimo": new XiaomiMimoExecutor(),
"mimo-free": new MimoFreeExecutor(),
mmf: new MimoFreeExecutor(), // Alias for mimo-free
"codebuddy-cn": new CodeBuddyExecutor(),
@@ -84,12 +91,15 @@ export { CursorExecutor } from "./cursor.js";
export { VertexExecutor } from "./vertex.js";
export { DefaultExecutor } from "./default.js";
export { OpenCodeExecutor } from "./opencode.js";
export { OpenCodeGoExecutor } from "./opencode-go.js";
export { OpenCodeZenExecutor } from "./opencode-zen.js";
export { GrokWebExecutor } from "./grok-web.js";
export { GrokCliExecutor } from "./grok-cli.js";
export { PerplexityWebExecutor } from "./perplexity-web.js";
export { OllamaLocalExecutor } from "./ollama-local.js";
export { CommandCodeExecutor } from "./commandcode.js";
export { XiaomiTokenplanExecutor } from "./xiaomi-tokenplan.js";
export { XiaomiMimoExecutor } from "./xiaomi-mimo.js";
export { MimoFreeExecutor } from "./mimo-free.js";
export { CodeBuddyExecutor } from "./codebuddy-cn.js";
export { CodeBuddyIntlExecutor } from "./codebuddy-intl.js";

View File

@@ -127,12 +127,18 @@ async function readResponsePrefix(response, signal, maxBytes, timeoutMs) {
return decoder.decode(concatChunks(chunks, totalBytes));
}
// The instruction goes into the current user turn, never into a top-level
// `systemPrompt`: kiro.dev answers any body carrying that field with
// 400 REQUEST_BODY_INVALID, so writing it here turned every repair retry into
// a hard failure.
function appendRepairInstruction(body, kind) {
const repaired = structuredClone(body || {});
const instruction = REPAIR_INSTRUCTIONS[kind] || "Retry the previous incomplete Kiro response.";
repaired.systemPrompt = repaired.systemPrompt
? `${repaired.systemPrompt}\n\n${instruction}`
: instruction;
const msg = repaired?.conversationState?.currentMessage?.userInputMessage;
if (msg) {
const content = typeof msg.content === "string" ? msg.content : "";
msg.content = content ? `${content}\n\n${instruction}` : instruction;
}
return repaired;
}
@@ -259,6 +265,19 @@ export class KiroExecutor extends BaseExecutor {
}
}
// CLIRO parity for the Amazon surfaces: the Kiro runtime accepts the
// SSO bearer header + agent-mode marker. Without these the deprecated
// path gateway answers REQUEST_BODY_INVALID for modern payloads.
if (credentials?.accessToken) {
headers["x-amz-sso-bearer"] = credentials.accessToken;
}
headers["x-amzn-kiro-agent-mode"] = "spec";
headers["x-amzn-codewhisperer-machine-id"] = "kiro-desktop";
const profileArn = credentials?.providerSpecificData?.profileArn;
if (profileArn) {
headers["x-amzn-codewhisperer-profile-arn"] = profileArn;
}
return headers;
}
@@ -285,9 +304,13 @@ export class KiroExecutor extends BaseExecutor {
// 403 "bearer token invalid", so they must hit the CodeWhisperer
// *.amazonaws.com surface, and in the region the token was minted in
// (the baseUrls are hardcoded us-east-1).
const isCodeWhispererSurface =
authMethod === "api_key" || authMethod === "external_idp" || authMethod === "idc";
if (!isCodeWhispererSurface) return baseUrls;
// Kiro deprecated the legacy path-style GenerateAssistantResponse on
// runtime.*.kiro.dev (IDE 1.0.228+ moved to POST / + x-amz-target). The
// path gateway now answers valid modern payloads with 400
// REQUEST_BODY_INVALID, and 400 is terminal in BaseExecutor, so kiro.dev
// must never be the first surface for any auth method. Amazon surfaces
// reject foreign tokens with 401/403, which DO fall through, so trying
// q/codewhisperer first is safe for every auth method (CLIRO parity).
const region = (credentials?.providerSpecificData?.region || "us-east-1").trim();
const regionalize = (u) =>
@@ -297,20 +320,17 @@ export class KiroExecutor extends BaseExecutor {
const amazon = baseUrls.filter((u) => u.includes("amazonaws.com")).map(regionalize);
const others = baseUrls.filter((u) => !u.includes("amazonaws.com"));
if (authMethod === "api_key") {
const q = amazon.filter((u) => u.includes("://q."));
const remaining = amazon.filter((u) => !u.includes("://q."));
return q.length > 0
? [...q, ...remaining, ...others]
: [...amazon, ...others];
}
return amazon.length > 0 ? [...amazon, ...others] : baseUrls;
const q = amazon.filter((u) => u.includes("://q."));
const remaining = amazon.filter((u) => !u.includes("://q."));
return q.length > 0
? [...q, ...remaining, ...others]
: [...amazon, ...others];
}
buildUrl(model, stream, urlIndex = 0, credentials = null) {
const baseUrls = this.getOrderedBaseUrls(credentials);
return baseUrls[urlIndex] || baseUrls[0] || this.config.baseUrl;
const url = baseUrls[urlIndex] || baseUrls[0] || this.config.baseUrl;
return url;
}
// Retry only endpoint/auth-surface failures. Payload-invalid HTTP 400 must be

View File

@@ -0,0 +1,186 @@
import crypto from "node:crypto";
import { DefaultExecutor } from "./default.js";
import { resolveSessionId } from "../utils/sessionManager.js";
import { getModelTargetFormat } from "../config/providerModels.js";
import { FORMATS } from "../translator/formats.js";
import {
normalizeResponsesInput,
clampResponsesCallId,
coerceResponsesArguments,
coerceResponsesOutput,
} from "../translator/formats/responsesApi.js";
const SESSION_HEADER = "x-opencode-session";
const SESSION_FIELD = "_opencodeGoSession";
const MAX_SESSION_LENGTH = 256;
const RESPONSES_BASE_URL = "https://opencode.ai/zen/go/v1/responses";
const MAX_TOOL_NAME_LEN = 128;
function normalizeSession(value) {
if (typeof value !== "string") return null;
const normalized = value.trim();
if (!normalized || normalized.length > MAX_SESSION_LENGTH) return null;
return normalized;
}
function nativeSession(headers) {
if (!headers || typeof headers !== "object") return null;
for (const [key, value] of Object.entries(headers)) {
if (key.toLowerCase() === SESSION_HEADER) return normalizeSession(value);
}
return null;
}
function translatedSession(sessionId, clientTool) {
const digest = crypto
.createHash("sha256")
.update(`opencode-go\0${clientTool || "generic"}\0${sessionId}`)
.digest("hex")
.slice(0, 32);
return `ses_${digest}`;
}
// Responses-only per the provider registry (grok-4.6, gpt-5.6-luna, muse-spark, …),
// including the family-regex fallback for passthrough ids — never hardcode model ids here.
function isResponsesModel(model) {
return getModelTargetFormat("opencode-go", model) === FORMATS.OPENAI_RESPONSES;
}
// Flatten Chat Completions tool declarations into the Responses flat shape and
// drop hosted/nameless tools the /responses endpoint rejects.
function normalizeResponsesTools(body) {
if (!Array.isArray(body.tools)) return;
const validNames = new Set();
body.tools = body.tools.filter((tool) => {
if (!tool || typeof tool !== "object" || Array.isArray(tool)) return false;
const fn = tool.function && typeof tool.function === "object" && !Array.isArray(tool.function) ? tool.function : null;
const rawName = typeof tool.name === "string" ? tool.name : (typeof fn?.name === "string" ? fn.name : "");
const name = rawName.trim();
if (!name) return false;
const description = typeof tool.description === "string" ? tool.description : (typeof fn?.description === "string" ? fn.description : "");
let parameters = (tool.parameters && typeof tool.parameters === "object" && !Array.isArray(tool.parameters))
? tool.parameters
: (fn?.parameters && typeof fn.parameters === "object" && !Array.isArray(fn.parameters) ? fn.parameters : { type: "object", properties: {} });
// Mirror the request translator: {type:"object"} without properties is rejected
// by strict Responses backends, so fill in the empty properties map.
if (parameters.type === "object" && !parameters.properties) parameters = { ...parameters, properties: {} };
for (const k of Object.keys(tool)) delete tool[k];
tool.type = "function";
tool.name = name.slice(0, MAX_TOOL_NAME_LEN);
if (description) tool.description = description;
tool.parameters = parameters;
validNames.add(tool.name);
return true;
});
if (body.tool_choice && typeof body.tool_choice === "object" && !Array.isArray(body.tool_choice)) {
if (body.tool_choice.type === "function") {
const n = typeof body.tool_choice.name === "string" ? body.tool_choice.name.trim() : "";
if (!n || !validNames.has(n)) delete body.tool_choice;
}
}
}
// Last line of defense for native Responses clients (sourceFormat === targetFormat
// skips translation): coerce items in place so malformed tool payloads 400 here
// with a clear shape instead of upstream as InputValidationError.
function sanitizeResponsesItems(body) {
if (!Array.isArray(body.input)) return;
body.input = body.input.filter((item) => {
if (!item || typeof item !== "object" || Array.isArray(item)) return true;
// Strip prior-turn reasoning items: Muse Spark contributor models route to
// an upstream Console backend where encrypted_content cannot be validated across
// rotated accounts or sessions, causing 400 "reasoning encrypted_content was not issued to this caller".
if (item.type === "reasoning") return false;
delete item.encrypted_content;
delete item.reasoning_encrypted_content;
if (item.type === "function_call") {
if (!item.name || typeof item.name !== "string" || item.name.trim() === "") return false;
item.name = item.name.trim().slice(0, MAX_TOOL_NAME_LEN);
item.call_id = clampResponsesCallId(item.call_id);
item.arguments = coerceResponsesArguments(item.arguments);
return true;
}
if (item.type === "function_call_output") {
item.call_id = clampResponsesCallId(item.call_id);
item.output = coerceResponsesOutput(item.output);
return true;
}
return true;
});
}
export class OpenCodeGoExecutor extends DefaultExecutor {
constructor() {
super("opencode-go");
}
buildUrl(model, stream, urlIndex = 0, credentials = null) {
// Muse Spark lives on /responses even when a stale runtimeTransport leaks in.
if (isResponsesModel(model)) return RESPONSES_BASE_URL;
return super.buildUrl(model, stream, urlIndex, credentials);
}
prepareRequestCredentials({ body, credentials, providerSessionId, clientTool } = {}) {
const sourceCredentials = credentials || {};
const native = nativeSession(sourceCredentials.rawHeaders);
const resolved = normalizeSession(providerSessionId) || resolveSessionId({
headers: sourceCredentials.rawHeaders,
body,
connectionId: sourceCredentials.connectionId,
scope: "opencode-go",
});
return {
...sourceCredentials,
[SESSION_FIELD]: native || translatedSession(resolved, clientTool),
};
}
async execute(args) {
const credentials = this.prepareRequestCredentials(args);
return super.execute({ ...args, credentials });
}
buildHeaders(credentials, stream = true, url, model) {
const headers = super.buildHeaders(credentials || {}, stream, url, model);
const prepared = credentials?.[SESSION_FIELD];
if (prepared) {
headers[SESSION_HEADER] = prepared;
return headers;
}
const fallback = this.prepareRequestCredentials({ credentials });
headers[SESSION_HEADER] = fallback[SESSION_FIELD];
return headers;
}
transformRequest(model, body, stream, credentials) {
const out = super.transformRequest(model, body);
if (!isResponsesModel(model || body?.model)) return out;
const normalized = normalizeResponsesInput(out.input);
if (normalized) out.input = normalized;
if (!Array.isArray(out.input) || out.input.length === 0) {
out.input = [{ type: "message", role: "user", content: [{ type: "input_text", text: "..." }] }];
}
// Responses names the output cap max_output_tokens, not max_tokens.
if (out.max_output_tokens === undefined) {
if (out.max_completion_tokens !== undefined) out.max_output_tokens = out.max_completion_tokens;
else if (out.max_tokens !== undefined) out.max_output_tokens = out.max_tokens;
}
delete out.max_tokens;
delete out.max_completion_tokens;
if (out.reasoning_effort !== undefined && out.reasoning === undefined) {
out.reasoning = { effort: out.reasoning_effort, summary: "auto" };
}
if (out.reasoning && typeof out.reasoning === "object" && !Array.isArray(out.reasoning)) {
if (!out.reasoning.summary) out.reasoning.summary = "auto";
}
delete out.reasoning_effort;
out.stream = true;
out.store = false;
normalizeResponsesTools(out);
sanitizeResponsesItems(out);
return out;
}
}

View File

@@ -0,0 +1,315 @@
import crypto from "node:crypto";
import { DefaultExecutor } from "./default.js";
import { resolveSessionId } from "../utils/sessionManager.js";
import { isMuseSparkModel } from "../providers/models/helpers.js";
import {
normalizeResponsesInput,
clampResponsesCallId,
coerceResponsesArguments,
coerceResponsesOutput,
} from "../translator/formats/responsesApi.js";
const SESSION_HEADER = "x-opencode-session";
const SESSION_FIELD = "_opencodeZenSession";
const MAX_SESSION_LENGTH = 256;
const RESPONSES_BASE_URL = "https://opencode.ai/zen/v1/responses";
const MAX_TOOL_NAME_LEN = 128;
const OPENCODE_UA = "opencode/1.18.31";
export const OPENCODE_SESSION_RE = /^ses_[0-9a-f]{12}[0-9A-Za-z]{14}$/;
const BASE62_CHARS = "0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz";
// Free-tier fingerprint (mirrors opencode executor, PR #4132): upstream 403s
// requests without the file-search quartet and without stream:true.
const OPENCODE_FINGERPRINT_TOOLS = ["bash", "glob", "grep", "read"];
function hasValidOpencodeVersion(ua) {
const m = String(ua || "").match(/opencode\/(\d+)\.(\d+)(?:\.(\d+))?/i);
if (!m) return false;
const major = parseInt(m[1], 10);
const minor = parseInt(m[2], 10);
return major > 1 || (major === 1 && minor >= 17);
}
function unstableRandom() {
const bytes = crypto.randomBytes(14);
let randomPart = "";
for (let i = 0; i < 14; i++) {
randomPart += BASE62_CHARS[bytes[i] % 62];
}
return randomPart;
}
export function generateSessionId(timestamp = Date.now()) {
const current = BigInt(timestamp) * 0x1000n + 1n;
const value = ~current;
const time = Array.from({ length: 6 }, (_, index) =>
Number((value >> BigInt(40 - 8 * index)) & 0xffn)
.toString(16)
.padStart(2, "0")
).join("");
return `ses_${time}${unstableRandom()}`;
}
export function generateRequestId(timestamp = Date.now()) {
const current = BigInt(timestamp) * 0x1000n + 1n;
const value = current;
const time = Array.from({ length: 6 }, (_, index) =>
Number((value >> BigInt(40 - 8 * index)) & 0xffn)
.toString(16)
.padStart(2, "0")
).join("");
return `msg_${time}${unstableRandom()}`;
}
export function translateSessionId(sessionId, clientTool = "") {
if (typeof sessionId === "string" && OPENCODE_SESSION_RE.test(sessionId.trim())) {
return sessionId.trim();
}
const digest = crypto
.createHash("sha256")
.update(`opencode\0${clientTool || "generic"}\0${sessionId || ""}`)
.digest();
const timeHex = digest.subarray(0, 6).toString("hex");
let randomPart = "";
for (let i = 6; i < 20; i++) {
randomPart += BASE62_CHARS[digest[i] % 62];
}
return `ses_${timeHex}${randomPart}`;
}
function toolNameOf(tool) {
if (!tool || typeof tool !== "object" || Array.isArray(tool)) return "";
const fn = tool.function && typeof tool.function === "object" && !Array.isArray(tool.function) ? tool.function : null;
const raw = typeof tool.name === "string" ? tool.name : (typeof fn?.name === "string" ? fn.name : "");
return raw.trim();
}
function ensureChatFingerprintTools(body) {
if (!body || typeof body !== "object") return;
const present = new Set();
if (Array.isArray(body.tools)) {
for (const tool of body.tools) {
const name = toolNameOf(tool);
if (name) present.add(name);
}
} else {
body.tools = [];
}
for (const name of OPENCODE_FINGERPRINT_TOOLS) {
if (present.has(name)) continue;
body.tools.push({
type: "function",
function: {
name,
description: `OpenCode built-in ${name} tool`,
parameters: { type: "object", properties: {} },
},
});
present.add(name);
}
}
function ensureResponsesFingerprintTools(body) {
if (!body || typeof body !== "object") return;
const present = new Set();
if (Array.isArray(body.tools)) {
for (const tool of body.tools) {
const name = toolNameOf(tool);
if (name) present.add(name);
}
} else {
body.tools = [];
}
for (const name of OPENCODE_FINGERPRINT_TOOLS) {
if (present.has(name)) continue;
body.tools.push({
type: "function",
name,
description: `OpenCode built-in ${name} tool`,
parameters: { type: "object", properties: {} },
});
present.add(name);
}
}
function normalizeSession(value) {
if (typeof value !== "string") return null;
const normalized = value.trim();
if (!normalized || normalized.length > MAX_SESSION_LENGTH) return null;
return normalized;
}
function nativeSession(headers) {
if (!headers || typeof headers !== "object") return null;
for (const [key, value] of Object.entries(headers)) {
if (key.toLowerCase() === SESSION_HEADER) {
const normalized = normalizeSession(value);
if (normalized && OPENCODE_SESSION_RE.test(normalized)) return normalized;
}
}
return null;
}
function translatedSession(sessionId, clientTool) {
return translateSessionId(sessionId, clientTool);
}
// Strip the thinking suffix "model(level)" so checks hit the base id.
function baseModelId(model) {
return String(model || "").replace(/\([^()]+\)\s*$/, "").trim();
}
function isResponsesModel(model) {
return isMuseSparkModel(baseModelId(model));
}
// Flatten Chat Completions tool declarations into the Responses flat shape and
// drop hosted/nameless tools the /responses endpoint rejects.
function normalizeResponsesTools(body) {
if (!Array.isArray(body.tools)) return;
const validNames = new Set();
body.tools = body.tools.filter((tool) => {
if (!tool || typeof tool !== "object" || Array.isArray(tool)) return false;
const fn = tool.function && typeof tool.function === "object" && !Array.isArray(tool.function) ? tool.function : null;
const rawName = typeof tool.name === "string" ? tool.name : (typeof fn?.name === "string" ? fn.name : "");
const name = rawName.trim();
if (!name) return false;
const description = typeof tool.description === "string" ? tool.description : (typeof fn?.description === "string" ? fn.description : "");
let parameters = (tool.parameters && typeof tool.parameters === "object" && !Array.isArray(tool.parameters))
? tool.parameters
: (fn?.parameters && typeof fn.parameters === "object" && !Array.isArray(fn.parameters) ? fn.parameters : { type: "object", properties: {} });
// Mirror the request translator: {type:"object"} without properties is rejected
// by strict Responses backends, so fill in the empty properties map.
if (parameters.type === "object" && !parameters.properties) parameters = { ...parameters, properties: {} };
for (const k of Object.keys(tool)) delete tool[k];
tool.type = "function";
tool.name = name.slice(0, MAX_TOOL_NAME_LEN);
if (description) tool.description = description;
tool.parameters = parameters;
validNames.add(tool.name);
return true;
});
if (body.tool_choice && typeof body.tool_choice === "object" && !Array.isArray(body.tool_choice)) {
if (body.tool_choice.type === "function") {
const n = typeof body.tool_choice.name === "string" ? body.tool_choice.name.trim() : "";
if (!n || !validNames.has(n)) delete body.tool_choice;
}
}
}
// Last line of defense for native Responses clients (sourceFormat === targetFormat
// skips translation): coerce items in place so malformed tool payloads 400 here
// with a clear shape instead of upstream as InputValidationError.
function sanitizeResponsesItems(body) {
if (!Array.isArray(body.input)) return;
body.input = body.input.filter((item) => {
if (!item || typeof item !== "object" || Array.isArray(item)) return true;
// Strip prior-turn reasoning items: Muse Spark contributor models route to
// an upstream Console backend where encrypted_content cannot be validated across
// rotated accounts or sessions, causing 400 "reasoning encrypted_content was not issued to this caller".
if (item.type === "reasoning") return false;
delete item.encrypted_content;
delete item.reasoning_encrypted_content;
if (item.type === "function_call") {
if (!item.name || typeof item.name !== "string" || item.name.trim() === "") return false;
item.name = item.name.trim().slice(0, MAX_TOOL_NAME_LEN);
item.call_id = clampResponsesCallId(item.call_id);
item.arguments = coerceResponsesArguments(item.arguments);
return true;
}
if (item.type === "function_call_output") {
item.call_id = clampResponsesCallId(item.call_id);
item.output = coerceResponsesOutput(item.output);
return true;
}
return true;
});
}
export class OpenCodeZenExecutor extends DefaultExecutor {
constructor() {
super("opencode-zen");
}
buildUrl(model, stream, urlIndex = 0, credentials = null) {
// Muse Spark lives on /responses even when a stale runtimeTransport leaks in.
if (isResponsesModel(model)) return RESPONSES_BASE_URL;
return super.buildUrl(model, stream, urlIndex, credentials);
}
prepareRequestCredentials({ body, credentials, providerSessionId, clientTool } = {}) {
const sourceCredentials = credentials || {};
const native = nativeSession(sourceCredentials.rawHeaders);
const resolved = normalizeSession(providerSessionId) || resolveSessionId({
headers: sourceCredentials.rawHeaders,
body,
connectionId: sourceCredentials.connectionId,
scope: "opencode-zen",
});
return {
...sourceCredentials,
[SESSION_FIELD]: native || translatedSession(resolved, clientTool),
};
}
async execute(args) {
const credentials = this.prepareRequestCredentials(args);
return super.execute({ ...args, credentials });
}
buildHeaders(credentials, stream = true, url, model) {
const headers = super.buildHeaders(credentials || {}, stream, url, model);
const raw = credentials?.rawHeaders || {};
const lower = {};
for (const [k, v] of Object.entries(raw)) lower[k.toLowerCase()] = v;
const downstreamUa = lower["user-agent"] || "";
// Free-tier gate: spoof the official client UA.
headers["User-Agent"] = hasValidOpencodeVersion(downstreamUa) ? downstreamUa : OPENCODE_UA;
headers["x-opencode-client"] = lower["x-opencode-client"] || "desktop";
const prepared = credentials?.[SESSION_FIELD];
if (prepared) {
headers[SESSION_HEADER] = prepared;
return headers;
}
const fallback = this.prepareRequestCredentials({ credentials });
headers[SESSION_HEADER] = fallback[SESSION_FIELD];
return headers;
}
transformRequest(model, body, stream, credentials) {
const out = super.transformRequest(model, body);
// Free-tier gate: upstream 403s stream:false even when everything else is valid.
if (out && typeof out === "object") out.stream = true;
if (!isResponsesModel(model || body?.model)) {
ensureChatFingerprintTools(out);
return out;
}
const normalized = normalizeResponsesInput(out.input);
if (normalized) out.input = normalized;
if (!Array.isArray(out.input) || out.input.length === 0) {
out.input = [{ type: "message", role: "user", content: [{ type: "input_text", text: "..." }] }];
}
// Responses names the output cap max_output_tokens, not max_tokens.
if (out.max_output_tokens === undefined) {
if (out.max_completion_tokens !== undefined) out.max_output_tokens = out.max_completion_tokens;
else if (out.max_tokens !== undefined) out.max_output_tokens = out.max_tokens;
}
delete out.max_tokens;
delete out.max_completion_tokens;
if (out.reasoning_effort !== undefined && out.reasoning === undefined) {
out.reasoning = { effort: out.reasoning_effort, summary: "auto" };
}
if (out.reasoning && typeof out.reasoning === "object" && !Array.isArray(out.reasoning)) {
if (!out.reasoning.summary) out.reasoning.summary = "auto";
}
delete out.reasoning_effort;
out.stream = true;
out.store = false;
ensureResponsesFingerprintTools(out);
normalizeResponsesTools(out);
sanitizeResponsesItems(out);
return out;
}
}

View File

@@ -1,70 +1,487 @@
import crypto from "crypto";
import { BaseExecutor } from "./base.js";
import { PROVIDERS } from "../config/providers.js";
import { MEMORY_CONFIG } from "../config/runtimeConfig.js";
import { getThinkingLevels } from "../providers/thinkingLevels.js";
import { injectReasoningContent } from "../utils/reasoningContentInjector.js";
import { resolveSessionId } from "../utils/sessionManager.js";
import { isMuseSparkModel } from "../providers/models/helpers.js";
import { applyFingerprintTools } from "../utils/opencodeFingerprint.js";
import { ANTHROPIC_API_VERSION } from "../providers/shared.js";
import {
normalizeResponsesInput,
clampResponsesCallId,
coerceResponsesArguments,
coerceResponsesOutput,
} from "../translator/formats/responsesApi.js";
const OPENCODE_UA = "opencode";
const MESSAGES_MODELS = new Set();
const OPENCODE_UA = "opencode/1.18.31";
const MAX_SESSION_LENGTH = 256;
const MAX_TOOL_NAME_LEN = 128;
const SESSION_HEADER = "x-opencode-session";
const SESSION_FIELD = "_opencodeSession";
const REQ_FIELD = "_opencodeRequest";
export const OPENCODE_SESSION_RE = /^ses_[0-9a-f]{12}[0-9A-Za-z]{14}$/;
export const OPENCODE_REQUEST_RE = /^msg_[0-9a-f]{12}[0-9A-Za-z]{14}$/;
const BASE62_CHARS = "0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz";
function generateRequestId() {
return `msg_${crypto.randomUUID().replace(/-/g, "")}`;
function hasValidOpencodeVersion(ua) {
const m = String(ua || "").match(/opencode\/(\d+)\.(\d+)(?:\.(\d+))?/i);
if (!m) return false;
const major = parseInt(m[1], 10);
const minor = parseInt(m[2], 10);
return major > 1 || (major === 1 && minor >= 17);
}
// Models served by /zen/v1/responses; every other model stays on /chat/completions.
const RESPONSES_MODELS = new Set([
"muse-spark-1.2-contributor-free",
"muse-spark-1.3-contributor-free",
]);
const MESSAGES_MODELS = new Set(["union-alpha"]);
let lastTimestamp = 0;
let counter = 0;
function unstableRandom() {
const bytes = crypto.randomBytes(14);
let randomPart = "";
for (let i = 0; i < 14; i++) {
randomPart += BASE62_CHARS[bytes[i] % 62];
}
return randomPart;
}
function generateSessionId() {
return `ses_${crypto.randomUUID().replace(/-/g, "")}`;
export function generateSessionId(timestamp = Date.now()) {
if (timestamp !== lastTimestamp) {
lastTimestamp = timestamp;
counter = 0;
}
counter++;
const current = BigInt(timestamp) * 0x1000n + BigInt(counter);
const value = ~current;
const time = Array.from({ length: 6 }, (_, index) =>
Number((value >> BigInt(40 - 8 * index)) & 0xffn)
.toString(16)
.padStart(2, "0")
).join("");
return `ses_${time}${unstableRandom()}`;
}
// Normalize any resolved id into opencode's ses_ format (stable per-conversation)
function toOpencodeSession(id) {
const stripped = String(id || "").replace(/^ses_/, "").replace(/-/g, "");
return stripped ? `ses_${stripped}` : null;
export function generateRequestId(timestamp = Date.now()) {
const current = BigInt(timestamp) * 0x1000n + 1n;
const value = current;
const time = Array.from({ length: 6 }, (_, index) =>
Number((value >> BigInt(40 - 8 * index)) & 0xffn)
.toString(16)
.padStart(2, "0")
).join("");
return `msg_${time}${unstableRandom()}`;
}
function resolveOpencodeSession(body, credentials) {
return toOpencodeSession(resolveSessionId({
headers: credentials?.rawHeaders,
body,
connectionId: credentials?.connectionId,
scope: "opencode",
}));
export function translateSessionId(sessionId, clientTool = "") {
if (typeof sessionId === "string" && OPENCODE_SESSION_RE.test(sessionId.trim())) {
return sessionId.trim();
}
const digest = crypto
.createHash("sha256")
.update(`opencode\0${clientTool || "generic"}\0${sessionId || ""}`)
.digest();
const timeHex = digest.subarray(0, 6).toString("hex");
let randomPart = "";
for (let i = 6; i < 20; i++) {
randomPart += BASE62_CHARS[digest[i] % 62];
}
return `ses_${timeHex}${randomPart}`;
}
function normalizeSession(value) {
if (typeof value !== "string") return null;
const normalized = value.trim();
if (!normalized || normalized.length > MAX_SESSION_LENGTH) return null;
return normalized;
}
function nativeSession(headers) {
if (!headers || typeof headers !== "object") return null;
for (const [key, value] of Object.entries(headers)) {
if (key.toLowerCase() === SESSION_HEADER) {
const normalized = normalizeSession(value);
if (normalized && OPENCODE_SESSION_RE.test(normalized)) return normalized;
}
}
return null;
}
// Upstream free-tier quota is accounted per session. Minting a fresh
// x-opencode-session on every request burns through it and surfaces as
// 429 FreeUsageLimitError with growing reset-after delays, while the real
// CLI reuses one long-lived canonical session per conversation. Mirror
// that: one stable canonical session per downstream identity, evicted
// after MEMORY_CONFIG.sessionTtlMs like the other session stores.
const stableOpencodeSessions = new Map();
const MAX_STABLE_SESSIONS = 1000;
const stableSessionCleanup = setInterval(() => {
const now = Date.now();
for (const [key, entry] of stableOpencodeSessions) {
if (now - entry.lastUsed > MEMORY_CONFIG.sessionTtlMs) {
stableOpencodeSessions.delete(key);
}
}
}, MEMORY_CONFIG.sessionCleanupIntervalMs);
if (stableSessionCleanup.unref) stableSessionCleanup.unref();
function identityKey(credentials) {
const connectionId = credentials?.connectionId || credentials?.id;
if (connectionId) return `opencode:conn:${String(connectionId).slice(0, 128)}`;
const raw = credentials?.rawHeaders || {};
const auth = raw.authorization || raw.Authorization || raw["x-api-key"] || raw["X-Api-Key"] || "";
if (auth) {
const digest = crypto.createHash("sha256").update(String(auth)).digest("hex").slice(0, 32);
return `opencode:auth:${digest}`;
}
return "opencode:default";
}
export function stableSessionId(credentials) {
const key = identityKey(credentials);
const existing = stableOpencodeSessions.get(key);
if (existing) {
existing.lastUsed = Date.now();
stableOpencodeSessions.delete(key);
stableOpencodeSessions.set(key, existing);
return existing.sessionId;
}
const sessionId = generateSessionId();
if (stableOpencodeSessions.size >= MAX_STABLE_SESSIONS) {
stableOpencodeSessions.delete(stableOpencodeSessions.keys().next().value);
}
stableOpencodeSessions.set(key, { sessionId, lastUsed: Date.now() });
return sessionId;
}
function lastUserText(body) {
try {
if (!body || typeof body !== "object") return "";
const arr = Array.isArray(body.messages)
? body.messages
: Array.isArray(body.input)
? body.input
: null;
if (!arr) return typeof body.input === "string" ? body.input.slice(-600) : "";
for (let i = arr.length - 1; i >= 0; i--) {
const msg = arr[i];
if (!msg) continue;
if (msg.role && msg.role !== "user") continue;
const content = msg.content;
if (typeof content === "string" && content.trim()) return content.trim().slice(-600);
if (Array.isArray(content)) {
const text = content
.map((part) => (typeof part === "string" ? part : part?.text || part?.input_text || ""))
.join(" ")
.trim();
if (text) return text.slice(-600);
}
}
} catch {
return "";
}
return "";
}
// The real CLI sends the current user message id (stable per turn, same on
// retries) as x-opencode-request. Derive it deterministically from the
// session plus the last user message so retries share the id.
export function deriveRequestId(sessionId, body) {
const text = lastUserText(body);
if (!text) return generateRequestId();
const digest = crypto
.createHash("sha256")
.update(`opencode-req\0${sessionId || ""}\0${text}`)
.digest();
const timeHex = digest.subarray(0, 6).toString("hex");
let randomPart = "";
for (let i = 6; i < 20; i++) {
randomPart += BASE62_CHARS[digest[i] % 62];
}
const id = `msg_${timeHex}${randomPart}`;
return OPENCODE_REQUEST_RE.test(id) ? id : generateRequestId();
}
function normalizeRequestId(value) {
if (typeof value !== "string") return null;
const normalized = value.trim();
if (!normalized || normalized.length > MAX_SESSION_LENGTH) return null;
return OPENCODE_REQUEST_RE.test(normalized) ? normalized : null;
}
function bodyHasSessionHints(body) {
try {
if (!body || typeof body !== "object") return false;
if (typeof body.session_id === "string" && body.session_id.trim()) return true;
if (typeof body.conversation_id === "string" && body.conversation_id.trim()) return true;
if (typeof body.prompt_cache_key === "string" && body.prompt_cache_key.trim()) return true;
if (body.metadata && typeof body.metadata.user_id === "string" && body.metadata.user_id.trim()) return true;
if (body.request && body.request.sessionId != null && String(body.request.sessionId) !== "") return true;
const arr = Array.isArray(body.messages)
? body.messages
: Array.isArray(body.input)
? body.input
: null;
if (arr) {
let assistantText = "";
for (const msg of arr) {
if (msg?.role === "assistant") {
const content = msg.content;
if (typeof content === "string") assistantText += content;
else if (Array.isArray(content)) {
for (const part of content) assistantText += part?.text || part?.output || "";
}
if (assistantText.length >= 50) return true;
}
}
}
return false;
} catch {
return false;
}
}
// Strip the thinking suffix "model(level)" so registry lookups hit the base id.
function baseModelId(model) {
return String(model || "").replace(/\([^()]+\)\s*$/, "").trim();
}
function isResponsesModel(model) {
const base = baseModelId(model);
return RESPONSES_MODELS.has(base) || isMuseSparkModel(base);
}
function isMessagesModel(model) {
return MESSAGES_MODELS.has(baseModelId(model));
}
function resolveOpencodeSession(body, credentials, providerSessionId, clientTool) {
const headers = credentials?.rawHeaders || {};
const native = nativeSession(headers);
if (native) return native;
let incoming = null;
for (const [key, value] of Object.entries(headers)) {
if (key.toLowerCase() === SESSION_HEADER) {
incoming = normalizeSession(value);
break;
}
}
const hinted = incoming || normalizeSession(providerSessionId);
if (hinted) return translateSessionId(hinted, clientTool);
if (credentials?.connectionId || bodyHasSessionHints(body)) {
let viaManager = null;
try {
viaManager = resolveSessionId({
headers,
body,
connectionId: credentials?.connectionId,
scope: "opencode",
});
} catch {
viaManager = null;
}
if (viaManager) return translateSessionId(viaManager, clientTool);
}
return stableSessionId(credentials);
}
function resolveOpencodeRequestId(body, credentials, sessionId) {
const raw = credentials?.rawHeaders || {};
for (const [key, value] of Object.entries(raw)) {
if (key.toLowerCase() === "x-opencode-request") {
const normalized = normalizeRequestId(value);
if (normalized) return normalized;
break;
}
}
return deriveRequestId(sessionId, body);
}
function normalizeResponsesTools(body) {
if (!Array.isArray(body.tools)) return;
const validNames = new Set();
body.tools = body.tools.filter((tool) => {
if (!tool || typeof tool !== "object" || Array.isArray(tool)) return false;
const fn = tool.function && typeof tool.function === "object" && !Array.isArray(tool.function) ? tool.function : null;
const rawName = typeof tool.name === "string" ? tool.name : (typeof fn?.name === "string" ? fn.name : "");
const name = rawName.trim();
if (!name) return false;
const description = typeof tool.description === "string" ? tool.description : (typeof fn?.description === "string" ? fn.description : "");
let parameters = (tool.parameters && typeof tool.parameters === "object" && !Array.isArray(tool.parameters))
? tool.parameters
: (fn?.parameters && typeof fn.parameters === "object" && !Array.isArray(fn.parameters) ? fn.parameters : { type: "object", properties: {} });
if (parameters.type === "object" && !parameters.properties) parameters = { ...parameters, properties: {} };
for (const k of Object.keys(tool)) delete tool[k];
tool.type = "function";
tool.name = name.slice(0, MAX_TOOL_NAME_LEN);
if (description) tool.description = description;
tool.parameters = parameters;
validNames.add(tool.name);
return true;
});
if (body.tool_choice && typeof body.tool_choice === "object" && !Array.isArray(body.tool_choice)) {
if (body.tool_choice.type === "function") {
const n = typeof body.tool_choice.name === "string" ? body.tool_choice.name.trim() : "";
if (!n || !validNames.has(n)) delete body.tool_choice;
}
}
}
function sanitizeResponsesItems(body) {
if (!Array.isArray(body.input)) return;
body.input = body.input.filter((item) => {
if (!item || typeof item !== "object" || Array.isArray(item)) return true;
// Strip prior-turn reasoning items: OpenCode Free uses public/pooled credentials
// (`Bearer public`) routing to an upstream OpenAI/Console account pool.
// OpenAI Responses API strictly enforces that reasoning `encrypted_content`
// can only be decrypted by the exact caller/account that issued it; sending it
// across different accounts or rotating proxy relays triggers:
// [invalid_request_error] reasoning `encrypted_content` was not issued to this caller (400).
// Furthermore, under stateless mode (store=false), omitting encrypted_content
// causes OpenAI to reject the referenced reasoning item as "not found or was deleted".
// Dropping prior reasoning items allows multi-turn conversations and tool-calling
// loops to succeed cleanly.
if (item.type === "reasoning") return false;
delete item.encrypted_content;
delete item.reasoning_encrypted_content;
if (item.type === "function_call") {
if (!item.name || typeof item.name !== "string" || item.name.trim() === "") return false;
item.name = item.name.trim().slice(0, MAX_TOOL_NAME_LEN);
item.call_id = clampResponsesCallId(item.call_id);
item.arguments = coerceResponsesArguments(item.arguments);
return true;
}
if (item.type === "function_call_output") {
item.call_id = clampResponsesCallId(item.call_id);
item.output = coerceResponsesOutput(item.output);
return true;
}
return true;
});
}
function normalizeOpencodeReasoning(model, body) {
const current = body.reasoning;
const currentReasoning = current && typeof current === "object" && !Array.isArray(current)
? current
: null;
const requestedEffort = typeof body.reasoning_effort === "string"
? body.reasoning_effort
: currentReasoning?.effort;
if (typeof requestedEffort !== "string") return;
const cleanModel = baseModelId(model || body.model);
const supportedLevels = getThinkingLevels("opencode", cleanModel);
let effort = requestedEffort.toLowerCase().trim();
if ((effort === "max" || effort === "ultra") && supportedLevels?.length && !supportedLevels.includes(effort)) {
if (effort === "ultra" && supportedLevels.includes("max")) effort = "max";
else if (supportedLevels.includes("xhigh")) effort = "xhigh";
}
body.reasoning = { ...currentReasoning, effort };
if (!body.reasoning.summary) body.reasoning.summary = "auto";
delete body.reasoning_effort;
}
export class OpenCodeExecutor extends BaseExecutor {
constructor() {
super("opencode", PROVIDERS.opencode);
this._currentSessionId = null;
}
prepareRequestCredentials({ body, credentials, providerSessionId, clientTool } = {}) {
const sourceCredentials = credentials || {};
const session = resolveOpencodeSession(body, sourceCredentials, providerSessionId, clientTool);
return {
...sourceCredentials,
[SESSION_FIELD]: session,
[REQ_FIELD]: resolveOpencodeRequestId(body, sourceCredentials, session),
};
}
transformRequest(model, body, stream, credentials) {
this._currentSessionId = resolveOpencodeSession(body, credentials);
if (body && typeof body === "object" && model && !body.model) body.model = model;
// Zen rejects non-streaming requests on free models with 403 FreeTierError;
// always stream upstream and let the handler layer aggregate for non-stream clients.
if (body && typeof body === "object") body.stream = true;
if (isResponsesModel(model || body?.model) && body && typeof body === "object") {
// ponytail: chỉ model đã xác nhận auto-only; mở allowlist khi có bằng chứng.
if ("tool_choice" in body && body.tool_choice !== "auto"
&& this.config.quirks?.forceAutoToolChoiceModels?.includes(baseModelId(model))) {
body.tool_choice = "auto";
}
const normalized = normalizeResponsesInput(body.input);
if (normalized) body.input = normalized;
if (!Array.isArray(body.input) || body.input.length === 0) {
body.input = [{ type: "message", role: "user", content: [{ type: "input_text", text: "..." }] }];
}
// Responses API names the output cap max_output_tokens and takes thinking
// as reasoning:{effort,summary} — normalize the Chat fields at this boundary.
if (body.max_output_tokens === undefined) {
if (body.max_completion_tokens !== undefined) body.max_output_tokens = body.max_completion_tokens;
else if (body.max_tokens !== undefined) body.max_output_tokens = body.max_tokens;
}
delete body.max_tokens;
delete body.max_completion_tokens;
normalizeOpencodeReasoning(model, body);
body.stream = true;
body.store = false;
normalizeResponsesTools(body);
sanitizeResponsesItems(body);
// Free-tier fingerprint tools are required even when an agent client
// already supplied tools. ZCode/Claude Code requests normally have
// non-empty tool arrays; skipping cloaking here triggers 403 FreeTierError.
applyFingerprintTools(body, true);
} else if (body && typeof body === "object") {
applyFingerprintTools(body, false);
}
return injectReasoningContent({ provider: this.provider, model, body });
}
async execute(args) {
return super.execute({ ...args, credentials: this.prepareRequestCredentials(args) });
}
buildUrl(model) {
const base = this.config.baseUrl;
return MESSAGES_MODELS.has(model)
? `${base}/zen/v1/messages`
: `${base}/zen/v1/chat/completions`;
if (isResponsesModel(model)) return `${base}/zen/v1/responses`;
if (isMessagesModel(model)) return `${base}/zen/v1/messages`;
return `${base}/zen/v1/chat/completions`;
}
buildHeaders(credentials, stream = true) {
buildHeaders(credentials, stream = true, url = "") {
const raw = credentials?.rawHeaders || {};
const lower = {};
for (const [k, v] of Object.entries(raw)) lower[k.toLowerCase()] = v;
const downstreamUa = lower["user-agent"] || "";
const isOpencodeDownstream = downstreamUa.toLowerCase().includes("opencode");
const isOpencodeDownstream = hasValidOpencodeVersion(downstreamUa);
return {
const session = credentials?.[SESSION_FIELD] || this.prepareRequestCredentials({ credentials })[SESSION_FIELD];
const downstreamReq = normalizeRequestId(lower["x-opencode-request"]);
const requestId = credentials?.[REQ_FIELD] || downstreamReq || generateRequestId();
const headers = {
"Content-Type": "application/json",
"Authorization": "Bearer public",
"User-Agent": isOpencodeDownstream ? downstreamUa : OPENCODE_UA,
"x-opencode-client": lower["x-opencode-client"] || "desktop",
"x-opencode-session": lower["x-opencode-session"] || this._currentSessionId || generateSessionId(),
"x-opencode-request": lower["x-opencode-request"] || generateRequestId(),
"x-opencode-session": session,
"x-opencode-request": requestId,
"x-opencode-project": lower["x-opencode-project"] || "global",
"Accept": stream ? "text/event-stream" : "*/*",
};
if (url.endsWith("/messages")) headers["anthropic-version"] = ANTHROPIC_API_VERSION;
return headers;
}
}

View File

@@ -29,19 +29,24 @@ import { BaseExecutor } from "./base.js";
import { PROVIDERS } from "../config/providers.js";
import { proxyAwareFetch } from "../utils/proxyFetch.js";
import { SSE_DONE } from "../utils/sseConstants.js";
import { FETCH_CONNECT_TIMEOUT_MS } from "../config/runtimeConfig.js";
import { FETCH_CONNECT_TIMEOUT_MS, HTTP_STATUS } from "../config/runtimeConfig.js";
import { resolveProviderTimeoutMs } from "../services/providerTimeout.js";
import {
QODER_CHAT_URL_ENCODED,
QODER_CHAT_BASE_ALT,
QODER_CHAT_SIG_PATH,
QODER_MODEL_MAP,
QODER_CONTEXT_TIER_ENV,
qoderInferenceBase,
} from "../shared/qoder/constants.js";
import { getQoderModelConfig, resolveQoderModels, isQoderPat, resolveQoderCredentials } from "../services/qoderModels.js";
import { OPENAI_BLOCK, CLAUDE_BLOCK } from "../translator/schema/blocks.js";
import { encodeDataUri } from "../translator/concerns/image.js";
import { createQoderSseCoalescer } from "../shared/qoder/sse.js";
import { rewriteQoderMessageAttachments } from "../shared/qoder/attachments.js";
import { resolveQoderContextTier, applyQoderContextTier } from "../shared/qoder/contextTier.js";
/**
* Hoist role:"system" messages out of the messages array (Qoder rejects
* system in messages) and flatten any multipart content arrays.
* system in messages) and flatten multipart content arrays — EXCEPT image
* blocks, which are preserved (see normalizeContent).
*/
function normalizeMessages(messages) {
if (!Array.isArray(messages) || messages.length === 0) {
@@ -51,18 +56,88 @@ function normalizeMessages(messages) {
const out = [];
for (const msg of messages) {
if (!msg || typeof msg !== "object") continue;
const text = extractText(msg.content);
if (msg.role === "system") {
const text = extractText(msg.content);
if (text) systemParts.push(text);
continue;
}
const cloned = { ...msg };
cloned.content = text;
cloned.content = normalizeContent(msg.content);
out.push(cloned);
}
return { messages: out, systemText: systemParts.join("\n\n") };
}
/**
* Normalize one message's content for Qoder.
*
* Text-only content is flattened to a plain string (Qoder's historical
* shape). When images are present the content stays an array and image
* blocks are kept as OpenAI-style `image_url` parts. Native qodercli
* uploads inlined bytes to `/api/v2/image/upload` first and then sends
* the OSS URL — `buildQoderRequestBody` does that rewrite before this
* runs. Tiny leftover data URIs are still accepted. The legacy
* top-level `image_urls` / `chat_context.imageUrls` slots stay null —
* qodercli leaves them null too.
*
* Claude-style `{type:"image", source:{...}}` blocks are converted to
* `image_url`. File/document blocks that survived rewrite become short
* stubs so 30MB PDFs never land in agent_chat_generation.
*/
function normalizeContent(content) {
if (typeof content === "string") return content;
if (content == null) return "";
if (!Array.isArray(content)) return String(content);
const blocks = [];
const textParts = [];
let hasImage = false;
const pushText = (text) => {
if (!text) return;
if (hasImage || blocks.length) blocks.push({ type: OPENAI_BLOCK.TEXT, text });
else textParts.push(text);
};
const imageUrlOf = (item) => {
if (typeof item.image_url === "string" && item.image_url) return item.image_url;
if (typeof item.image_url?.url === "string" && item.image_url.url) return item.image_url.url;
return null;
};
for (const item of content) {
if (!item || typeof item !== "object") continue;
const imageUrl = item.type === OPENAI_BLOCK.IMAGE_URL ? imageUrlOf(item) : null;
if (imageUrl) {
blocks.push({ type: OPENAI_BLOCK.IMAGE_URL, image_url: { url: imageUrl } });
hasImage = true;
} else if (item.type === CLAUDE_BLOCK.IMAGE && item.source) {
// Claude base64/url image → OpenAI image_url equivalent.
const src = item.source;
const url = src.type === "base64" && src.data
? encodeDataUri(src.media_type || "image/png", src.data)
: typeof src.url === "string" && src.url ? src.url : null;
if (url) {
blocks.push({ type: OPENAI_BLOCK.IMAGE_URL, image_url: { url } });
hasImage = true;
}
} else if (item.type === OPENAI_BLOCK.FILE) {
const name = item.file?.filename || item.file?.name || "file";
pushText(`[file omitted: ${name} — Qoder reads documents via its file API, not inlined bytes]`);
} else if (item.type === CLAUDE_BLOCK.DOCUMENT) {
const name = item.title || "document";
pushText(`[file omitted: ${name} — Qoder reads documents via its file API, not inlined bytes]`);
} else if (typeof item.text === "string" && item.text) {
pushText(item.text);
}
}
if (!hasImage) return textParts.join("\n");
// Prepend any text collected before the first image block.
if (textParts.length) blocks.unshift({ type: OPENAI_BLOCK.TEXT, text: textParts.join("\n") });
return blocks;
}
function extractText(content) {
if (typeof content === "string") return content;
if (content == null) return "";
@@ -85,9 +160,9 @@ function extractText(content) {
function lastUserText(messages) {
for (let i = messages.length - 1; i >= 0; i--) {
const m = messages[i];
if (m?.role === "user" && typeof m.content === "string") {
return m.content;
}
if (m?.role !== "user") continue;
if (typeof m.content === "string") return m.content;
if (Array.isArray(m.content)) return extractText(m.content);
}
return "";
}
@@ -111,6 +186,11 @@ function stableChatRecordId(model, messages, tools, maxTokens) {
if (m.role) { h.update("\0"); h.update(m.role); }
if (typeof m.content === "string" && m.content) {
h.update("\0"); h.update(m.content);
} else if (Array.isArray(m.content)) {
// Include image refs so the same prompt with a different image gets
// a distinct chat_record_id.
h.update("\0");
try { h.update(JSON.stringify(m.content)); } catch {}
}
}
if (tools) {
@@ -128,16 +208,16 @@ function truncate(s, n) {
/**
* Map the OpenAI-style request body into the exact shape Qoder expects.
*/
async function buildQoderRequestBody({ model, body, credentials, log, proxyOptions, signal }) {
async function buildQoderRequestBody({ model, body, credentials, log, proxyOptions, signal, uploadFn = null, region = "intl" }) {
const qoderKey = String(model || "").replace(/^qoder\//, "");
// Fetch model config from dynamic API instead of relying on static QODER_MODEL_MAP.
// This allows support for new Qoder models (e.g., qmodel_latest) without code changes.
let modelConfig = await getQoderModelConfig(credentials, qoderKey, { log, proxyOptions, signal });
let modelConfig = await getQoderModelConfig(credentials, qoderKey, { log, proxyOptions, signal, region });
if (!modelConfig) {
// Try a forced refresh once before giving up — the cache may simply
// not be populated yet on first ever call for this credential.
const refreshed = await resolveQoderModels(credentials, { forceRefresh: true, log, proxyOptions, signal });
const refreshed = await resolveQoderModels(credentials, { forceRefresh: true, log, proxyOptions, signal, region });
const retried = refreshed?.rawConfigs.get(qoderKey);
if (!retried) {
throw new Error(
@@ -147,7 +227,30 @@ async function buildQoderRequestBody({ model, body, credentials, log, proxyOptio
modelConfig = { ...retried, key: qoderKey };
}
const { messages, systemText } = normalizeMessages(body.messages || []);
const incoming = Array.isArray(body.messages)
? body.messages.map((m) => {
if (!m || typeof m !== "object") return m;
return {
...m,
content: Array.isArray(m.content)
? m.content.map((b) => (b && typeof b === "object" ? { ...b } : b))
: m.content,
};
})
: [];
try {
await rewriteQoderMessageAttachments(incoming, {
credentials,
log,
proxyOptions,
signal,
uploadFn,
});
} catch (err) {
log?.warn?.("QODER", `attachment rewrite failed: ${err.message}`);
}
const { messages, systemText } = normalizeMessages(incoming);
const tools = body.tools;
const isReasoning = !!modelConfig.is_reasoning;
const maxOutputTokens = Number(modelConfig.max_output_tokens) || 0;
@@ -166,7 +269,21 @@ async function buildQoderRequestBody({ model, body, credentials, log, proxyOptio
const sessionId = stableHash("qoder-session", psd.userId, qoderKey);
const recordId = stableChatRecordId(qoderKey, messages, tools, maxTokens);
return {
// Context-window tier (200K/400K/1M): the IDE picks one from model_config.context_config;
// qodercli-style requests default to the smallest. Escalate when the prompt no longer fits.
const tierChoice = resolveQoderContextTier(
modelConfig,
{ system: systemText, messages, tools },
{ preference: process.env[QODER_CONTEXT_TIER_ENV] },
);
if (tierChoice) {
log?.info?.(
"QODER",
`context tier ${tierChoice.tier.name} (${tierChoice.tier.tokenCount} tokens, ${tierChoice.reason}) for ~${tierChoice.estimatedTokens} prompt tokens`,
);
}
const built = {
qoderKey,
payload: {
request_id: uuidv4(),
@@ -214,51 +331,71 @@ async function buildQoderRequestBody({ model, body, credentials, log, proxyOptio
},
modelConfig,
};
if (tierChoice) applyQoderContextTier(built.payload, tierChoice.tier);
return built;
}
/**
* Check if a qoder error message indicates a billing/quota block.
* Signatures: code 112 (quota exhausted), code 10605 (queue throttle), pricingUrl field.
* Signatures: code 110 (billing daily count exceeded), code 112 (quota
* exhausted), code 10605 (queue throttle), pricingUrl field.
*/
function isBillingBlock(inner) {
if (!inner || typeof inner !== "string") return false;
const lowerMsg = inner.toLowerCase();
// Match: {"code":"112",...}, {"code":"10605",...}, or pricingUrl field
return /\"code\"\s*:\s*\"(112|10605)\"/.test(inner) || lowerMsg.includes("pricingurl");
if (lowerMsg.includes("pricingurl")) return true;
// Parsed code preferred over regex: matches numeric or string "110"/"112"/"10605".
try {
const parsed = JSON.parse(inner);
const code = String(parsed?.code ?? "");
if (code === "110" || code === "112" || code === "10605") return true;
} catch { /* not JSON — fall through to legacy shape match */ }
// Match legacy exact shapes: {"code":"112",...}, {"code":"10605",...}.
return /"code"\s*:\s*"(112|10605)"/.test(inner);
}
/**
* Peek the first SSE frame to detect billing errors before piping.
* Returns { isBilling, statusVal, message, consumed } — `consumed` is every
* Peek the first SSE data line to detect upstream errors before piping.
* Returns { isError, isBilling, statusVal, message, consumed } — `consumed` is every
* byte read so far (including the peeked line) so the caller can re-process
* it and nothing is dropped from the stream.
*/
async function peekFirstQoderFrame(reader, decoder) {
let consumed = "";
let offset = 0;
let upstreamDone = false;
while (true) {
const { done, value } = await reader.read();
if (done) return { isBilling: false, consumed, upstreamDone: true };
let nl = consumed.indexOf("\n", offset);
if (nl === -1 && !upstreamDone) {
const { done, value } = await reader.read();
upstreamDone = done;
consumed += done ? decoder.decode() : decoder.decode(value, { stream: true });
continue;
}
if (offset >= consumed.length) return { isError: false, consumed, upstreamDone };
if (nl === -1) nl = consumed.length;
consumed += decoder.decode(value, { stream: true });
const nl = consumed.indexOf("\n");
if (nl === -1) continue; // need a full line first
const line = consumed.slice(0, nl).replace(/\r$/, "").trim();
const line = consumed.slice(offset, nl).replace(/\r$/, "").trim();
offset = nl + 1;
if (!line.startsWith("data:")) continue;
const data = line.slice(5).trimStart();
if (data === "[DONE]") return { isBilling: false, consumed };
if (data === "[DONE]") return { isError: false, consumed, upstreamDone };
let envelope;
try { envelope = JSON.parse(data); } catch { return { isBilling: false, consumed }; }
try { envelope = JSON.parse(data); } catch { return { isError: false, consumed, upstreamDone }; }
const statusVal = typeof envelope.statusCodeValue === "number" ? envelope.statusCodeValue : 200;
const inner = typeof envelope.body === "string" ? envelope.body : "";
// statusCodeValue is documented numeric, but accept numeric strings defensively.
const raw = Number(envelope?.statusCodeValue);
const statusVal = Number.isNaN(raw) ? 200 : raw;
const inner = typeof envelope?.body === "string"
? envelope.body
: envelope?.body != null ? JSON.stringify(envelope.body) : "";
if (statusVal !== 200 && isBillingBlock(inner)) {
return { isBilling: true, statusVal, message: inner || `qoder billing block (${statusVal})` };
if (statusVal !== 200) {
return { isError: true, isBilling: isBillingBlock(inner), statusVal, message: inner || `upstream status ${statusVal}` };
}
return { isBilling: false, consumed };
return { isError: false, consumed, upstreamDone };
}
}
@@ -269,32 +406,41 @@ async function peekFirstQoderFrame(reader, decoder) {
* Each upstream line looks like:
* data: {"statusCodeValue":200,"body":"{\"choices\":[{\"delta\":{...}}]}"}
* The inner body is an OpenAI streaming chunk (or "[DONE]"). We unwrap it
* and re-emit as `data: <inner>\n\n`. Errors become a synthetic OpenAI error
* chunk + [DONE].
* and re-emit as `data: <inner>\n\n`. First-frame errors become HTTP errors;
* errors after streaming starts retain the synthetic chunk + [DONE] path.
*
* Critical: Qoder's SSE often keeps the socket open after the terminal
* [DONE]/error frame (agent keepalive). Non-streaming clients drain via
* response.text() which hangs until the socket closes — so on terminal
* events we cancel the upstream reader and close our stream immediately.
*
* NEW: Peek first frame to detect billing blocks (code 112/10605/pricingUrl).
* If detected, return 403 response so chatCore marks connection unavailable
* and triggers combo fallback instead of leaking error text into chat.
* Usage: Qoder puts finish_reason on `delta` and sends token counts on a
* later `choices: []` frame. Downstream OpenAI/Claude clients only read
* usage from the finish chunk, so we coalesce those two frames (see
* createQoderSseCoalescer) before forwarding.
*
* Peek the first frame for errors before committing to HTTP 200. Preserve
* upstream error statuses so chatCore can handle failures instead of recording
* error text as a successful completion. Billing blocks retain the existing
* 403 mapping for quota/account fallback.
*/
async function wrapQoderSSE(response, model) {
async function wrapQoderSSE(response, model, log = null) {
if (!response.ok || !response.body) return response;
const decoder = new TextDecoder();
const reader = response.body.getReader();
// Peek first frame to detect billing block
// Detect errors before returning a successful streaming response.
const peek = await peekFirstQoderFrame(reader, decoder);
if (peek?.isBilling) {
// Billing block detected — return 403 so chatCore fails this connection
if (peek.isError) {
await reader.cancel().catch(() => {});
const status = peek.isBilling
? HTTP_STATUS.FORBIDDEN
: Number.isInteger(peek.statusVal) && peek.statusVal >= HTTP_STATUS.BAD_REQUEST && peek.statusVal <= 599
? peek.statusVal : HTTP_STATUS.BAD_GATEWAY;
return new Response(
JSON.stringify({ error: { message: peek.message, code: peek.statusVal } }),
{ status: 403, headers: { "Content-Type": "application/json" } }
{ status, headers: { "Content-Type": "application/json" } }
);
}
@@ -303,6 +449,11 @@ async function wrapQoderSSE(response, model) {
const upstreamDrained = peek.upstreamDone === true;
const encoder = new TextEncoder();
let doneEmitted = false;
const coalescer = createQoderSseCoalescer({ model, encoder, sseDone: SSE_DONE });
const syncDone = () => {
if (coalescer.doneEmitted) doneEmitted = true;
};
// Process one already-extracted SSE line (no trailing newline).
const processLine = (line, controller) => {
@@ -313,16 +464,42 @@ async function wrapQoderSSE(response, model) {
const data = trimmed.slice(5).trimStart();
if (data === "[DONE]") {
controller.enqueue(encoder.encode(SSE_DONE));
doneEmitted = true;
coalescer.flush(controller);
syncDone();
return;
}
let envelope;
try { envelope = JSON.parse(data); } catch { return; }
const statusVal = typeof envelope.statusCodeValue === "number" ? envelope.statusCodeValue : 200;
const inner = typeof envelope.body === "string" ? envelope.body : "";
const statusVal = Number(envelope.statusCodeValue) || 200;
const inner = typeof envelope.body === "string"
? envelope.body
: envelope.body != null ? JSON.stringify(envelope.body) : "";
if (statusVal !== 200) {
// Always visible: error envelopes are rare and worth one stderr line at
// any log level (response bodies carry no credentials).
try {
console.error(`[QODER] error envelope status=${statusVal} statusType=${typeof envelope.statusCodeValue} bodyType=${typeof envelope.body} body=${truncate(inner, 300)}`);
} catch { /* logging must not break the stream */ }
if (isBillingBlock(inner)) {
// Billing/quota envelope at any stream position (peek only covers the
// first frame): emit a structured error chunk, not fake assistant text.
// parseSSEToOpenAIResponse understands chunk.error and turns it into a
// non-200 result so chat.js locks the model and falls back. Streaming
// clients receive a real SSE error instead of "[qoder error ...]" text.
const errObj = JSON.stringify({
error: {
message: inner || `qoder billing block (${statusVal})`,
code: "qoder_billing_block",
status: 403,
type: "quota_error",
},
});
controller.enqueue(encoder.encode(`data: ${errObj}\n\n`));
controller.enqueue(encoder.encode(SSE_DONE));
doneEmitted = true;
return;
}
const msg = inner || `upstream status ${statusVal}`;
const errChunk = JSON.stringify({
id: `qoder-error-${Date.now()}`,
@@ -337,14 +514,8 @@ async function wrapQoderSSE(response, model) {
return;
}
if (!inner) return;
if (inner === "[DONE]") {
controller.enqueue(encoder.encode(SSE_DONE));
doneEmitted = true;
return;
}
// Strip embedded newlines so the SSE frame stays a single event.
const sanitized = inner.replace(/\r?\n/g, "");
controller.enqueue(encoder.encode(`data: ${sanitized}\n\n`));
coalescer.handleInner(inner, controller);
syncDone();
};
const stream = new ReadableStream({
@@ -403,7 +574,7 @@ async function wrapQoderSSE(response, model) {
} finally {
if (!doneEmitted) {
try {
controller.enqueue(encoder.encode(SSE_DONE));
coalescer.flush(controller);
doneEmitted = true;
} catch { /* already closed */ }
}
@@ -427,18 +598,13 @@ async function wrapQoderSSE(response, model) {
}
export class QoderExecutor extends BaseExecutor {
constructor() {
super("qoder", PROVIDERS.qoder);
constructor(provider = "qoder") {
super(provider, PROVIDERS[provider]);
this.region = provider === "qoder-cn" ? "cn" : "intl";
}
buildUrl(credentials) {
// Job-token (jt-...) traffic must hit api2.qoder.sh — api3 rejects jt-
// with "Login expired" (403). Device tokens (dt-...) stay on api3.
const raw = credentials?.apiKey || credentials?.accessToken;
if (typeof raw === "string" && !raw.startsWith("pt-") && (raw.startsWith("jt-") || (credentials?.accessToken || "").startsWith("jt-"))) {
return `${QODER_CHAT_BASE_ALT}/algo${QODER_CHAT_SIG_PATH}?FetchKeys=llm_model_result&AgentId=agent_common&Encode=1`;
}
return QODER_CHAT_URL_ENCODED;
return `${qoderInferenceBase(credentials, this.region)}/algo${QODER_CHAT_SIG_PATH}?FetchKeys=llm_model_result&AgentId=agent_common&Encode=1`;
}
// Override execute entirely — Qoder needs:
@@ -453,7 +619,7 @@ export class QoderExecutor extends BaseExecutor {
const rawToken = credentials?.apiKey || credentials?.accessToken;
if (isQoderPat(rawToken)) {
try {
credentials = await resolveQoderCredentials(credentials, proxyOptions, signal);
credentials = await resolveQoderCredentials(credentials, proxyOptions, signal, this.region);
} catch (err) {
log?.error?.("QODER", `PAT exchange failed: ${err.message}`);
const fakeResp = new Response(
@@ -488,7 +654,7 @@ export class QoderExecutor extends BaseExecutor {
let qoderKey;
let payload;
try {
({ qoderKey, payload } = await buildQoderRequestBody({ model, body, credentials, log, proxyOptions, signal }));
({ qoderKey, payload } = await buildQoderRequestBody({ model, body, credentials, log, proxyOptions, signal, region: this.region }));
} catch (err) {
const fakeResp = new Response(
JSON.stringify({ error: { message: err.message } }),
@@ -547,8 +713,15 @@ export class QoderExecutor extends BaseExecutor {
response = await proxyAwareFetch(
url,
{ method: "POST", headers, body: encodedBodyBuf, signal: mergedSignal },
proxyOptions,
// A failed proxy request may already have reached Qoder. Replaying
// the same COSY signature directly reuses its requestId and returns
// 403/code 103. Let the caller retry through execute() with fresh signing.
{ ...proxyOptions, strictProxy: true },
);
} catch (err) {
// strictProxy wraps transport errors; retain caller cancellation semantics.
if (mergedSignal.aborted) throw mergedSignal.reason;
throw err;
} finally {
clearTimeout(connectTimer);
}
@@ -558,7 +731,7 @@ export class QoderExecutor extends BaseExecutor {
return { response, url, headers, transformedBody: payload };
}
const wrapped = await wrapQoderSSE(response, `qoder/${qoderKey}`);
const wrapped = await wrapQoderSSE(response, `${this.provider}/${qoderKey}`, log);
return { response: wrapped, url, headers, transformedBody: payload };
}

View File

@@ -0,0 +1,116 @@
import { DefaultExecutor } from "./default.js";
import { getMimoAccountCookie, invalidateMimoAccountCookieCache, resolveMimoServerBase, MIMO_API_UA } from "../shared/mimoAccount.js";
// Dual-route v2.6 models.
// v2.6 models dynamically route to the account service when desktop session credentials
// (mimoPassToken or account cookie) are present to consume weekly quota, falling back to
// the cloud API (sk- key) otherwise.
const ACCOUNT_MODELS = new Set([
"mimo-v2.6-pro",
"mimo-v2.6-flash",
"mimo-v2.6-pro-ultraspeed",
]);
// Session cookie resolved in execute() (async) and read back by buildHeaders()
// (sync — BaseExecutor.execute does not await it). Carried on the per-request
// credentials object, same as runtimeTransport.
const COOKIE_KEY = "__mimoAccountCookie";
// Upstream calls may hand us either the bare id or a `provider/model` ref.
function bareModel(model) {
const s = String(model || "");
const i = s.indexOf("/");
return i >= 0 ? s.slice(i + 1) : s;
}
export class XiaomiMimoExecutor extends DefaultExecutor {
constructor() {
super("xiaomi-mimo");
}
static isAccountRoute(model, credentials) {
const bare = bareModel(model);
if (!ACCOUNT_MODELS.has(bare)) return false;
return Boolean(
credentials?.[COOKIE_KEY] ||
credentials?.providerSpecificData?.mimoPassToken
);
}
isAccountRoute(model, credentials) {
return XiaomiMimoExecutor.isAccountRoute(model, credentials);
}
buildUrl(model, stream, urlIndex = 0, credentials = null) {
// Account route models live on the account-service route, which is not one of the
// declared transports — resolve it before the default runtimeTransport path.
if (this.isAccountRoute(model, credentials)) {
return `${resolveMimoServerBase(credentials?.providerSpecificData)}/api/route/chat/completions`;
}
// Cloud API models keep default handling, so a Claude-format client reaches
// the /anthropic/v1/messages transport.
return super.buildUrl(model, stream, urlIndex, credentials);
}
buildHeaders(credentials, stream = true, url, model) {
if (this.isAccountRoute(model, credentials) && credentials?.[COOKIE_KEY]) {
// Account route models authenticate with the account-session cookie, not the key.
return {
"Content-Type": "application/json",
Accept: stream ? "text/event-stream" : "application/json",
"User-Agent": MIMO_API_UA,
Cookie: credentials[COOKIE_KEY],
};
}
return super.buildHeaders(credentials, stream, url, model);
}
transformRequest(model, body, stream, credentials) {
// super runs stripUnsupportedParams, which flattens content-part
// arrays (see the xiaomi-mimo rule in translator/concerns/paramSupport.js).
const out = super.transformRequest(model, body, stream, credentials);
// Account route models: bridge reasoning_effort to official output_config.effort
// (matches MiMo Desktop app.asar behavior).
if (this.isAccountRoute(model, credentials)) {
const rawEffort = out.reasoning_effort || body?.reasoning_effort || body?.output_config?.effort;
if (rawEffort) {
delete out.reasoning_effort;
const norm = String(rawEffort).toLowerCase() === "xhigh" ? "high" : String(rawEffort).toLowerCase();
out.output_config = { ...(out.output_config || {}), effort: norm };
}
if (out.temperature == null) out.temperature = 1.0;
if (out.top_p == null) out.top_p = 0.95;
}
return out;
}
async execute(args) {
const { model, credentials, proxyOptions = null } = args;
if (!this.isAccountRoute(model, credentials)) return super.execute(args);
const cookie = await getMimoAccountCookie(credentials?.providerSpecificData, proxyOptions);
if (!cookie) {
return super.execute(args);
}
credentials[COOKIE_KEY] = cookie;
const result = await super.execute(args);
// A cached session can expire early — drop it and retry once with a fresh one.
if (result.response.status === 401) {
invalidateMimoAccountCookieCache();
const fresh = await getMimoAccountCookie(credentials?.providerSpecificData, proxyOptions).catch(() => null);
if (fresh) {
credentials[COOKIE_KEY] = fresh;
return super.execute(args);
}
}
return result;
}
}
export const __test__ = { ACCOUNT_MODELS, bareModel, COOKIE_KEY };
export default XiaomiMimoExecutor;

View File

@@ -29,11 +29,18 @@ import {
zedLlmFetch,
} from "../shared/zedAuth.js";
// Wire values for the `provider` field of POST /completions. These are NOT
// display names: cloud.zed.dev matches them exactly, and an unrecognized value
// fails the whole request with `500 {"message":"An internal server error
// occurred."}` before the model is ever looked at. Spellings come from Zed's
// own GET /models catalog: `anthropic`, `open_ai`, `google` (note underscore),
// `x_ai` follows the same convention — so feeding a catalog value back through
// normalizeZedProvider is identity.
const ZED_PROVIDER = {
anthropic: "Anthropic",
openai: "OpenAi",
google: "Google",
xai: "XAi",
anthropic: "anthropic",
openai: "open_ai",
google: "google",
xai: "x_ai",
};
function normalizeZedProvider(value, model) {
@@ -55,7 +62,14 @@ function buildProviderRequest(provider, model, body, stream, credentials) {
return openaiToClaudeRequest(model, body, true);
}
if (provider === ZED_PROVIDER.google) {
return openaiToGeminiRequest(model, body, true);
const geminiRequest = openaiToGeminiRequest(model, body, true);
// Zed's hosted Gemini backend speaks the Vertex safety vocabulary, not the
// public Gemini API enum the shared translator emits (`OFF`, `CIVIC_INTEGRITY`,
// `DANGEROUS_CONTENT`). Drop client-side safetySettings for the Zed Google
// path so Zed applies its own defaults — scoped here so native Gemini/
// Antigravity is untouched.
delete geminiRequest.safetySettings;
return geminiRequest;
}
if (provider === ZED_PROVIDER.openai) {
return openaiToOpenAIResponsesRequest(model, body, true, credentials);

View File

@@ -9,17 +9,19 @@ import { createRequestLogger } from "../utils/requestLogger.js";
import { getModelTargetFormat, getModelSupportedFormats, getModelStrip, getModelUpstreamId, getModelType, PROVIDER_ID_TO_ALIAS } from "../config/providerModels.js";
import { PROVIDERS } from "../config/providers.js";
import { createErrorResult, parseUpstreamError, formatProviderError } from "../utils/error.js";
import { upstreamResponseHeaders } from "../utils/upstreamHeaders.js";
import { HTTP_STATUS, TOKEN_SAVER_HEADER } from "../config/runtimeConfig.js";
import { handleBypassRequest } from "../utils/bypassHandler.js";
import { trackPendingRequest, appendRequestLog, saveRequestDetail } from "@/lib/usageDb.js";
import { trackPendingRequest, saveRequestDetail } from "@/lib/usageDb.js";
import { getExecutor } from "../executors/index.js";
import { supportsGrokCliReasoningEffort } from "../config/grokCli.js";
import { buildRequestDetail, extractRequestConfig } from "./chatCore/requestDetail.js";
import { buildRequestDetail, extractRequestConfig, shouldPersistRequestDetail } from "./chatCore/requestDetail.js";
import { handleForcedSSEToJson } from "./chatCore/sseToJsonHandler.js";
import { handleNonStreamingResponse } from "./chatCore/nonStreamingHandler.js";
import { handleStreamingResponse, buildOnStreamComplete } from "./chatCore/streamingHandler.js";
import { detectClientTool, isNativePassthrough } from "../utils/clientDetector.js";
import { dedupeTools } from "../utils/toolDeduper.js";
import { takeRenamedToolNames } from "../utils/opencodeFingerprint.js";
import { injectCaveman } from "../rtk/caveman.js";
import { injectPonytail } from "../rtk/ponytail.js";
import { compressMessages, formatRtkLog } from "../rtk/index.js";
@@ -28,7 +30,9 @@ import { compressWithPxpipe } from "../rtk/pxpipe.js";
import { getCapabilitiesForModel } from "../providers/capabilities.js";
import { stripUnsupportedModalities } from "../translator/concerns/modality.js";
import { prefetchRemoteImages } from "../translator/concerns/prefetch.js";
import { defaultClaudeToolType, shouldDefaultClaudeToolType } from "../translator/concerns/toolCall.js";
import { resolveSessionId } from "../utils/sessionManager.js";
import { maybeRejectEarlyStreamError } from "../utils/streamErrorPeek.js";
/**
* Core chat handler - shared between SSE and Worker
@@ -57,7 +61,7 @@ export function stripContinuityFields(body) {
return body;
}
export async function handleChatCore({ body, modelInfo, credentials, log, onCredentialsRefreshed, onRequestSuccess, onDisconnect, clientRawRequest, connectionId, userAgent, apiKey, ccFilterNaming, rtkEnabled, headroomEnabled, headroomUrl, headroomCompressUserMessages, cavemanEnabled, cavemanLevel, ponytailEnabled, ponytailLevel, pxpipeEnabled, pxpipeMinChars, pxpipeTimeoutMs, pxpipeTransform, onPxpipeEvent, sourceFormatOverride, providerThinking }) {
export async function handleChatCore({ body, modelInfo, credentials, log, onCredentialsRefreshed, onRequestSuccess, onDisconnect, clientRawRequest, connectionId, userAgent, apiKey, ccFilterNaming, rtkEnabled, headroomEnabled, headroomUrl, headroomCompressUserMessages, headroomTimeoutMs, cavemanEnabled, cavemanLevel, ponytailEnabled, ponytailLevel, pxpipeEnabled, pxpipeMinChars, pxpipeTimeoutMs, pxpipeTransform, onPxpipeEvent, sourceFormatOverride, providerThinking, capsOverride = null, streamErrorPatterns = null, persistUsage = "all" }) {
const { provider, model } = modelInfo;
const requestStartTime = Date.now();
// Stable per-session color so all lines of one CLI conversation share a tag
@@ -90,7 +94,12 @@ export async function handleChatCore({ body, modelInfo, credentials, log, onCred
// differ — kimi/glm only do /chat/completions). Undeclared models keep the
// upstream default (use the transport), preserving behavior for glm/deepseek/...
const useTransport = (!modelSupportedFormats || modelSupportedFormats.includes(sourceFormat)) ? runtimeTransport : null;
const targetFormat = modelTargetFormat || useTransport?.format || getTargetFormat(provider, credentials);
// A source-format-matched endpoint keeps the request lossless. Prefer it
// over a model-level targetFormat, which is only the fallback for clients
// whose wire format has no supported transport (for example MiniMax-M3:
// OpenAI clients should stay on /chat/completions; other clients can fall
// back to its declared Claude target).
const targetFormat = useTransport?.format || modelTargetFormat || getTargetFormat(provider, credentials);
if (useTransport && credentials) credentials.runtimeTransport = useTransport;
const stripList = getModelStrip(alias, model);
const upstreamModel = getModelUpstreamId(alias, model);
@@ -100,7 +109,7 @@ export async function handleChatCore({ body, modelInfo, credentials, log, onCred
if (providerThinking?.mode && providerThinking.mode !== "auto") {
const mode = providerThinking.mode;
if (mode === "on" && !body.thinking) {
console.log("Injecting provider-level thinking config override: on");
log?.debug?.("THINKING", `provider-level override: on`);
body = { ...body, thinking: { type: "enabled", budget_tokens: 10000 } };
} else if (mode === "off" && !body.thinking) {
body = { ...body, thinking: { type: "disabled" } };
@@ -109,6 +118,19 @@ export async function handleChatCore({ body, modelInfo, credentials, log, onCred
}
}
// Per-request opt-out: client can bypass all token savers via header
const tokenSaverEnabled = clientRawRequest?.headers?.[TOKEN_SAVER_HEADER]?.toLowerCase() !== "off";
// Cursor's translator rewrites tool_result into user text, so RTK must run on
// the source body before translation. Every other pair translates the tool
// shapes 1:1 — keep the post-translate pass there so those providers are
// untouched (and a retry never re-compresses an already-compressed body).
const preTranslateRtk = provider === "cursor"
? compressMessages(body, tokenSaverEnabled && rtkEnabled)
: null;
const preTranslateRtkLine = formatRtkLog(preTranslateRtk);
if (preTranslateRtkLine) console.log(preTranslateRtkLine);
const clientRequestedStreaming = body.stream === true || sourceFormat === FORMATS.ANTIGRAVITY || sourceFormat === FORMATS.GEMINI || sourceFormat === FORMATS.GEMINI_CLI;
const providerRequiresStreaming = PROVIDERS[provider]?.forceStream === true;
let stream = providerRequiresStreaming ? true : (body.stream !== false);
@@ -149,8 +171,10 @@ export async function handleChatCore({ body, modelInfo, credentials, log, onCred
if (credentials) credentials.rawHeaders = clientRawRequest?.headers || {};
// Auto-strip media blocks the model can't read (vision/audio/pdf) before translation.
// capsOverride lets the app layer assert per-model capabilities (e.g. user-registered
// models) on top of the static tables.
if (!passthrough) {
const caps = getCapabilitiesForModel(provider, model);
const caps = { ...getCapabilitiesForModel(provider, model), ...(capsOverride || {}) };
if (stripUnsupportedModalities(body, sourceFormat, caps)) {
log?.debug?.("MODALITY", `stripped unsupported media for ${provider}/${model}`);
}
@@ -235,17 +259,24 @@ export async function handleChatCore({ body, modelInfo, credentials, log, onCred
delete translatedBody.tools;
}
// Per-request opt-out: client can bypass all token savers via header
const tokenSaverEnabled = clientRawRequest?.headers?.[TOKEN_SAVER_HEADER]?.toLowerCase() !== "off";
// Claude tool schema requires `type` to be explicitly set; strict gateways (e.g., MiniMax)
// reject legacy payloads that omit it with HTTP 400. Default to "custom" when missing.
// Provider-scoped via quirks (shouldDefaultClaudeToolType): only gateways that declare
// requireClaudeToolType get the explicit type. Applying it unconditionally breaks
// Claude-format endpoints that only accept the legacy typeless tool shape — DeepSeek's
// Anthropic-compatible endpoint 400s with "unknown variant `custom`" (#3905).
if (shouldDefaultClaudeToolType(provider, finalFormat, translatedBody.tools, PROVIDERS)) {
translatedBody.tools = defaultClaudeToolType(translatedBody.tools);
}
// RTK: compress tool_result content
const rtkStats = compressMessages(translatedBody, tokenSaverEnabled && rtkEnabled);
// RTK: compress tool_result content. Skipped when already done pre-translate.
const rtkStats = preTranslateRtk || compressMessages(translatedBody, tokenSaverEnabled && rtkEnabled);
const rtkLine = formatRtkLog(rtkStats);
if (rtkLine) console.log(rtkLine);
if (rtkLine) log?.info?.("RTK", rtkLine.replace(/^\[RTK\] /, ""));
// Headroom: optional external proxy compression; fail open if proxy is absent.
const headroomDiagnostics = {};
const headroomStats = await compressWithHeadroom(translatedBody, { enabled: tokenSaverEnabled && headroomEnabled, url: headroomUrl, model: upstreamModel, format: finalFormat, compressUserMessages: headroomCompressUserMessages, diagnostics: headroomDiagnostics });
const headroomStats = await compressWithHeadroom(translatedBody, { enabled: tokenSaverEnabled && headroomEnabled, url: headroomUrl, model: upstreamModel, format: finalFormat, compressUserMessages: headroomCompressUserMessages, timeoutMs: headroomTimeoutMs, diagnostics: headroomDiagnostics });
const headroomLine = formatHeadroomLog(headroomStats);
const headroomSizeLine = formatHeadroomSizeLog(headroomDiagnostics);
if (headroomLine) {
@@ -258,6 +289,8 @@ export async function handleChatCore({ body, modelInfo, credentials, log, onCred
// Token-saver flags accumulator for the single "⚙" log line below.
const xf = [];
if (rtkStats?.hits?.length) xf.push(`RTK:${rtkStats.hits.length}`);
// Caveman: inject terse-style system prompt
if (tokenSaverEnabled && cavemanEnabled && cavemanLevel) {
injectCaveman(translatedBody, finalFormat, cavemanLevel);
@@ -291,7 +324,6 @@ export async function handleChatCore({ body, modelInfo, credentials, log, onCred
const executor = getExecutor(provider);
trackPendingRequest(model, provider, connectionId, true);
appendRequestLog({ model, provider, connectionId, status: "PENDING" }).catch(() => { });
const msgCount = translatedBody.messages?.length || translatedBody.input?.length || translatedBody.contents?.length || translatedBody.request?.contents?.length || 0;
log?.debug?.("REQUEST", `${provider.toUpperCase()} | ${model} | ${msgCount} msgs`);
@@ -344,26 +376,41 @@ export async function handleChatCore({ body, modelInfo, credentials, log, onCred
// exception: it is decoded by the executor into OpenAI-compatible output.
let providerResponseFormat = targetFormat;
try {
const result = await executor.execute({ model, body: translatedBody, stream, credentials, signal: streamController.signal, log, proxyOptions });
const result = await executor.execute({
model,
body: translatedBody,
stream,
credentials,
providerSessionId: sessionSeed,
clientTool,
signal: streamController.signal,
log,
proxyOptions,
});
providerResponse = result.response;
providerUrl = result.url;
providerHeaders = result.headers;
finalBody = result.transformedBody;
providerResponseFormat = result.responseFormat || targetFormat;
const renamedToolNames = takeRenamedToolNames(translatedBody);
if (renamedToolNames?.size) {
toolNameMap = new Map([...(toolNameMap || []), ...renamedToolNames]);
}
reqLogger.logTargetRequest(providerUrl, providerHeaders, finalBody);
} catch (error) {
trackPendingRequest(model, provider, connectionId, false, true);
appendRequestLog({ model, provider, connectionId, status: `FAILED ${error.name === "AbortError" ? 499 : HTTP_STATUS.BAD_GATEWAY}` }).catch(() => { });
saveRequestDetail(buildRequestDetail({
provider, model, connectionId,
latency: { ttft: 0, total: Date.now() - requestStartTime },
tokens: { prompt_tokens: 0, completion_tokens: 0 },
request: extractRequestConfig(body, stream),
providerRequest: translatedBody || null,
response: { error: error.message || String(error), status: error.name === "AbortError" ? 499 : 502, thinking: null },
pxpipe: pxpipeSummary,
status: "error"
})).catch(() => { });
if (shouldPersistRequestDetail(persistUsage, "error")) {
saveRequestDetail(buildRequestDetail({
provider, model, connectionId,
latency: { ttft: 0, total: Date.now() - requestStartTime },
tokens: { prompt_tokens: 0, completion_tokens: 0 },
request: extractRequestConfig(body, stream),
providerRequest: translatedBody || null,
response: { error: error.message || String(error), status: error.name === "AbortError" ? 499 : 502, thinking: null },
pxpipe: pxpipeSummary,
status: "error"
})).catch(() => { });
}
if (error.name === "AbortError") {
streamController.handleError(error);
@@ -398,7 +445,17 @@ export async function handleChatCore({ body, modelInfo, credentials, log, onCred
try { await onCredentialsRefreshed(newCredentials); } catch (e) { log?.warn?.("TOKEN", `onCredentialsRefreshed failed: ${e.message}`); }
}
try {
const retryResult = await executor.execute({ model, body: translatedBody, stream, credentials, signal: streamController.signal, log, proxyOptions });
const retryResult = await executor.execute({
model,
body: translatedBody,
stream,
credentials,
providerSessionId: sessionSeed,
clientTool,
signal: streamController.signal,
log,
proxyOptions,
});
if (retryResult.response.ok) {
providerResponse = retryResult.response;
providerUrl = retryResult.url;
@@ -413,21 +470,23 @@ export async function handleChatCore({ body, modelInfo, credentials, log, onCred
}
}
// Provider returned error
if (!providerResponse.ok) {
trackPendingRequest(model, provider, connectionId, false, true);
const { statusCode, message, resetsAtMs } = await parseUpstreamError(providerResponse, executor);
appendRequestLog({ model, provider, connectionId, status: `FAILED ${statusCode}` }).catch(() => { });
saveRequestDetail(buildRequestDetail({
provider, model, connectionId,
latency: { ttft: 0, total: Date.now() - requestStartTime },
tokens: { prompt_tokens: 0, completion_tokens: 0 },
request: extractRequestConfig(body, stream),
providerRequest: finalBody || translatedBody || null,
response: { error: message, status: statusCode, thinking: null },
pxpipe: pxpipeSummary,
status: "error"
})).catch(() => { });
if (shouldPersistRequestDetail(persistUsage, "error")) {
saveRequestDetail(buildRequestDetail({
provider, model, connectionId,
latency: { ttft: 0, total: Date.now() - requestStartTime },
tokens: { prompt_tokens: 0, completion_tokens: 0 },
request: extractRequestConfig(body, stream),
providerRequest: finalBody || translatedBody || null,
response: { error: message, status: statusCode, thinking: null },
pxpipe: pxpipeSummary,
status: "error"
})).catch(() => { });
}
const errMsg = formatProviderError(new Error(message), provider, model, statusCode);
if (log?.errorLine) {
@@ -435,16 +494,39 @@ export async function handleChatCore({ body, modelInfo, credentials, log, onCred
log.errorLine(reqTag, "✗", `ERROR ${statusCode} · ${provider}/${model} · ${Date.now() - requestStartTime}ms${urlStr}\n ${errMsg}`);
}
reqLogger.logError(new Error(message), finalBody || translatedBody);
return createErrorResult(statusCode, errMsg, resetsAtMs);
return createErrorResult(statusCode, errMsg, resetsAtMs, upstreamResponseHeaders(providerResponse.headers));
}
const appendLog = () => {}; // request log derived from usageHistory; kept as no-op seam for handlers
const sharedCtx = { provider, model, body, stream, translatedBody, finalBody, requestStartTime, connectionId, apiKey, clientRawRequest, onRequestSuccess, pxpipe: pxpipeSummary, reqTag, log, streamErrorPatterns, persistUsage };
// Early-peek streaming responses for configured in-stream error patterns.
// Some upstreams fail INSIDE a 200 SSE stream; without this the failure is
// piped to the client verbatim and account/combo fallback never triggers
// (see AGENTS.md "HTTP 200 in-stream errors"). Fail-open: no patterns → pass-through.
if (providerResponse.ok && stream) {
const peeked = await maybeRejectEarlyStreamError(
providerResponse,
streamErrorPatterns?.[provider],
{ signal: streamController.signal },
);
if (!peeked.ok) {
const { message } = await parseUpstreamError(peeked).catch(() => ({ message: "Stream error pattern matched" }));
trackPendingRequest(model, provider, connectionId, false, true);
appendLog({ status: `FAILED ${HTTP_STATUS.BAD_GATEWAY}` });
if (log?.errorLine) {
log.errorLine(reqTag, "✗", `ERROR 502 · ${provider}/${model} · ${Date.now() - requestStartTime}ms (in-stream)\n ${message}`);
}
return createErrorResult(HTTP_STATUS.BAD_GATEWAY, message);
}
providerResponse = peeked;
}
const sharedCtx = { provider, model, body, stream, translatedBody, finalBody, requestStartTime, connectionId, apiKey, clientRawRequest, onRequestSuccess, pxpipe: pxpipeSummary, reqTag, log };
const appendLog = (extra) => appendRequestLog({ model, provider, connectionId, ...extra }).catch(() => { });
const trackDone = () => trackPendingRequest(model, provider, connectionId, false);
// Provider forced streaming but client wants JSON
if (!clientRequestedStreaming && providerRequiresStreaming) {
const result = await handleForcedSSEToJson({ ...sharedCtx, providerResponse, sourceFormat, targetFormat: providerResponseFormat, customToolNames, trackDone, appendLog });
const result = await handleForcedSSEToJson({ ...sharedCtx, providerResponse, sourceFormat, targetFormat: providerResponseFormat, customToolNames, toolNameMap, trackDone, appendLog });
if (result) { streamController.handleComplete(); return result; }
}
@@ -457,7 +539,7 @@ export async function handleChatCore({ body, modelInfo, credentials, log, onCred
// Streaming response
const { onStreamComplete, streamDetailId } = buildOnStreamComplete({ ...sharedCtx });
return handleStreamingResponse({ ...sharedCtx, providerResponse, sourceFormat, targetFormat: providerResponseFormat, userAgent, reqLogger, toolNameMap, customToolNames, streamController, onStreamComplete, streamDetailId });
return handleStreamingResponse({ ...sharedCtx, providerResponse, sourceFormat, targetFormat: providerResponseFormat, userAgent, reqLogger, toolNameMap, customToolNames, streamController, onStreamComplete, streamDetailId, credentials });
}
export function isTokenExpiringSoon(expiresAt, bufferMs = 5 * 60 * 1000) {

View File

@@ -4,11 +4,15 @@ import { fromOpenAIFinish } from "../../translator/concerns/finishReason.js";
import { ollamaBodyToOpenAI } from "../../translator/response/ollama-to-openai.js";
import { addBufferToUsage, filterUsageForFormat } from "../../utils/usageTracking.js";
import { createErrorResult } from "../../utils/error.js";
import { upstreamResponseHeaders } from "../../utils/upstreamHeaders.js";
import { HTTP_STATUS } from "../../config/runtimeConfig.js";
import { parseSSEToOpenAIResponse } from "./sseToJsonHandler.js";
import { buildRequestDetail, extractRequestConfig, extractUsageFromResponse, saveUsageStats, formatDoneLine } from "./requestDetail.js";
import { appendRequestLog, saveRequestDetail } from "@/lib/usageDb.js";
import { unwrapClineEnvelope } from "../../shared/clineEnvelope.js";
import { buildRequestDetail, extractRequestConfig, extractUsageFromResponse, saveUsageStats, formatDoneLine, tokensForDetail, shouldPersistRequestDetail } from "./requestDetail.js";
import { saveRequestDetail } from "@/lib/usageDb.js";
import { matchStreamErrorPatterns } from "../../utils/streamErrorPatterns.js";
import { decloakToolNames } from "../../utils/claudeCloaking.js";
import { restoreToolNames } from "../../utils/opencodeFingerprint.js";
import { ROLE, RESPONSES_ITEM } from "../../translator/schema/index.js";
function parseToolArguments(value) {
@@ -281,7 +285,7 @@ export function translateNonStreamingResponse(responseBody, targetFormat, source
/**
* Handle non-streaming response from provider.
*/
export async function handleNonStreamingResponse({ providerResponse, provider, model, sourceFormat, targetFormat, body, stream, translatedBody, finalBody, requestStartTime, connectionId, apiKey, clientRawRequest, onRequestSuccess, reqLogger, toolNameMap, customToolNames, trackDone, appendLog, pxpipe, reqTag, log }) {
export async function handleNonStreamingResponse({ providerResponse, provider, model, sourceFormat, targetFormat, body, stream, translatedBody, finalBody, requestStartTime, connectionId, apiKey, clientRawRequest, onRequestSuccess, reqLogger, toolNameMap, customToolNames, trackDone, appendLog, pxpipe, reqTag, log, streamErrorPatterns, persistUsage = "all" }) {
trackDone();
const contentType = providerResponse.headers.get("content-type") || "";
let responseBody;
@@ -304,6 +308,11 @@ export async function handleNonStreamingResponse({ providerResponse, provider, m
}
}
// Unwrap before any consumer reads choices/usage so non-stream clients get a
// bare OpenAI body and usage tracking sees data.usage. No-op unless the
// provider opts in via transport.quirks.clineEnvelope.
responseBody = unwrapClineEnvelope(responseBody, provider);
reqLogger.logProviderResponse(providerResponse.status, providerResponse.statusText, providerResponse.headers, responseBody);
if (onRequestSuccess) {
Promise.resolve()
@@ -316,6 +325,21 @@ export async function handleNonStreamingResponse({ providerResponse, provider, m
// Decloak tool_use names once on raw Claude body, before any translation (INPUT side)
responseBody = decloakToolNames(responseBody, toolNameMap);
// Config-driven in-stream error detection: the HTTP call succeeded but the
// assembled content signals an upstream failure — treat it as an error so
// account/combo fallback and FAILED logging kick in (AGENTS.md hook #3).
const matchedPattern = matchStreamErrorPatterns(
streamErrorPatterns?.[provider],
responseBody?.choices?.[0]?.message?.content || responseBody?.content || "",
);
if (matchedPattern) {
appendLog({ status: `FAILED ${HTTP_STATUS.BAD_GATEWAY}` });
if (log?.errorLine) {
log.errorLine(reqTag, "✗", `ERROR 502 · ${provider}/${model} · ${Date.now() - requestStartTime}ms (in-stream)\n Stream error pattern matched: ${matchedPattern}`);
}
return createErrorResult(HTTP_STATUS.BAD_GATEWAY, `Stream error pattern matched: ${matchedPattern}`);
}
const usage = extractUsageFromResponse(responseBody);
appendLog({ tokens: usage, status: "200 OK" });
saveUsageStats({ provider, model, tokens: usage, connectionId, apiKey, endpoint: clientRawRequest?.endpoint, silent: true });
@@ -371,28 +395,30 @@ export async function handleNonStreamingResponse({ providerResponse, provider, m
reqLogger.logConvertedResponse(translatedResponse);
const totalLatency = Date.now() - requestStartTime;
saveRequestDetail(buildRequestDetail({
provider, model, connectionId, apiKey,
latency: { ttft: totalLatency, total: totalLatency },
tokens: usage || { prompt_tokens: 0, completion_tokens: 0 },
request: extractRequestConfig(body, stream),
providerRequest: finalBody || translatedBody || null,
providerResponse: responseBody || null,
response: {
content: translatedResponse?.choices?.[0]?.message?.content || translatedResponse?.content || null,
thinking: translatedResponse?.choices?.[0]?.message?.reasoning_content || translatedResponse?.reasoning_content || null,
finish_reason: translatedResponse?.choices?.[0]?.finish_reason || "unknown"
},
pxpipe,
status: "success"
}, { endpoint: clientRawRequest?.endpoint || null })).catch(err => {
console.error("[RequestDetail] Failed to save:", err.message);
});
if (shouldPersistRequestDetail(persistUsage, "success")) {
saveRequestDetail(buildRequestDetail({
provider, model, connectionId, apiKey,
latency: { ttft: totalLatency, total: totalLatency },
tokens: tokensForDetail(usage),
request: extractRequestConfig(body, stream),
providerRequest: finalBody || translatedBody || null,
providerResponse: responseBody || null,
response: {
content: translatedResponse?.choices?.[0]?.message?.content || translatedResponse?.content || null,
thinking: translatedResponse?.choices?.[0]?.message?.reasoning_content || translatedResponse?.reasoning_content || null,
finish_reason: translatedResponse?.choices?.[0]?.finish_reason || "unknown"
},
pxpipe,
status: "success"
}, { endpoint: clientRawRequest?.endpoint || null })).catch(err => {
console.error("[RequestDetail] Failed to save:", err.message);
});
}
return {
success: true,
response: new Response(JSON.stringify(translatedResponse), {
headers: { "Content-Type": "application/json", "Access-Control-Allow-Origin": "*" }
response: new Response(JSON.stringify(restoreToolNames(translatedResponse, toolNameMap)), {
headers: { "Content-Type": "application/json", "Access-Control-Allow-Origin": "*", ...upstreamResponseHeaders(providerResponse.headers) }
})
};
}

View File

@@ -1,4 +1,4 @@
import { saveRequestUsage, appendRequestLog, saveRequestDetail } from "@/lib/usageDb.js";
import { saveRequestUsage, saveRequestDetail } from "@/lib/usageDb.js";
import { COLORS } from "../../utils/stream.js";
import { canonicalizeUsage } from "../../utils/usageTracking.js";
@@ -25,10 +25,16 @@ export function extractUsageFromResponse(responseBody) {
if (!responseBody || typeof responseBody !== "object") return null;
// Claude format
// Note: OpenAI Responses usage ({input_tokens, input_tokens_details:{cached_tokens}})
// also matches this branch. Its prompt is cache-INCLUSIVE and its cache rides in
// input_tokens_details, so emit it as cached_tokens — the convention
// canonicalizeUsage() passes through without folding. Reading it here keeps
// cache accounting correct for /v1/responses and codex traffic.
if (responseBody.usage?.input_tokens !== undefined) {
return {
prompt_tokens: responseBody.usage.input_tokens || 0,
completion_tokens: responseBody.usage.output_tokens || 0,
cached_tokens: responseBody.usage.cached_tokens ?? responseBody.usage.input_tokens_details?.cached_tokens,
cache_read_input_tokens: responseBody.usage.cache_read_input_tokens,
cache_creation_input_tokens: responseBody.usage.cache_creation_input_tokens
};
@@ -39,7 +45,7 @@ export function extractUsageFromResponse(responseBody) {
return {
prompt_tokens: responseBody.usage.prompt_tokens || 0,
completion_tokens: responseBody.usage.completion_tokens || 0,
cached_tokens: responseBody.usage.prompt_tokens_details?.cached_tokens,
cached_tokens: responseBody.usage.cached_tokens ?? responseBody.usage.prompt_tokens_details?.cached_tokens,
reasoning_tokens: responseBody.usage.completion_tokens_details?.reasoning_tokens
};
}
@@ -104,6 +110,29 @@ export function formatDoneLine({ usage, latency }) {
return `DONE ${latency?.total ?? 0}ms${ttftStr} · ${inStr} · OUT ${outTok}`;
}
// Request-details storage convention: always prompt_tokens / completion_tokens.
// Translators often hand Claude `{input_tokens, output_tokens}` (or Gemini
// counts) to onStreamComplete; the Details tab only reads the OpenAI names,
// so an uncanonicalized object shows up as input=0 / output=0.
export function tokensForDetail(usage) {
if (!usage || typeof usage !== "object") {
return { prompt_tokens: 0, completion_tokens: 0 };
}
return canonicalizeUsage(usage) || {
prompt_tokens: usage.prompt_tokens ?? usage.input_tokens ?? 0,
completion_tokens: usage.completion_tokens ?? usage.output_tokens ?? 0,
};
}
// Combo fallback/account hops must not inflate Details with 0-token rows.
// `streaming-start` is never persisted: the placeholder was status=success at
// tokens=0, and nested/fusion paths often abandon the stream before complete.
export function shouldPersistRequestDetail(persistUsage, kind) {
if (kind === "streaming-start") return false;
if (persistUsage === "success-only") return kind === "success";
return true;
}
export function saveUsageStats({ provider, model, tokens, connectionId, apiKey, endpoint, label = "USAGE", silent = false }) {
if (!tokens || typeof tokens !== "object") return;

View File

@@ -1,16 +1,17 @@
import { convertResponsesStreamToJson } from "../../transformer/streamToJsonConverter.js";
import { matchStreamErrorPatterns } from "../../utils/streamErrorPatterns.js";
import { restoreToolNames } from "../../utils/opencodeFingerprint.js";
import { createErrorResult } from "../../utils/error.js";
import { HTTP_STATUS } from "../../config/runtimeConfig.js";
import { FORMATS } from "../../translator/formats.js";
import { PROVIDERS } from "../../config/providers.js";
import { buildRequestDetail, extractRequestConfig, saveUsageStats, formatDoneLine } from "./requestDetail.js";
import { saveRequestDetail } from "@/lib/usageDb.js";
import { ROLE, RESPONSES_ITEM } from "../../translator/schema/index.js";
// Responses-API providers (e.g. codex) may emit SSE without content-type + use Responses output shape
const isResponsesProvider = (p) =>
PROVIDERS[p]?.format === FORMATS.OPENAI_RESPONSES;
import { saveRequestDetail, appendRequestLog } from "@/lib/usageDb.js";
function textFromResponsesMessageItem(item) {
if (!item?.content || !Array.isArray(item.content)) return "";
@@ -215,17 +216,13 @@ export async function handleForcedSSEToJson({
clientRawRequest,
onRequestSuccess,
customToolNames,
toolNameMap,
trackDone,
appendLog,
reqTag,
log,
streamErrorPatterns,
}) {
const contentType = providerResponse.headers.get("content-type") || "";
const isSSE =
contentType.includes("text/event-stream") ||
(contentType === "" && isResponsesProvider(provider));
if (!isSSE) return null; // not handled here
trackDone();
@@ -306,7 +303,7 @@ export async function handleForcedSSEToJson({
if (sourceFormat === FORMATS.OPENAI_RESPONSES) {
return {
success: true,
response: new Response(JSON.stringify(jsonResponse), {
response: new Response(JSON.stringify(restoreToolNames(jsonResponse, toolNameMap)), {
headers: {
"Content-Type": "application/json",
"Access-Control-Allow-Origin": "*",
@@ -406,7 +403,7 @@ export async function handleForcedSSEToJson({
return {
success: true,
response: new Response(JSON.stringify(finalResp), {
response: new Response(JSON.stringify(restoreToolNames(finalResp, toolNameMap)), {
headers: {
"Content-Type": "application/json",
"Access-Control-Allow-Origin": "*",
@@ -432,8 +429,16 @@ export async function handleForcedSSEToJson({
"Invalid SSE response for non-streaming request",
);
if (parsed.error) {
// Structured error chunks may carry the real upstream status (e.g. the
// Qoder executor emits status 403 for billing envelopes). Preserve it so
// the account loop locks/falls back on the right status instead of a
// generic 502. Anything outside 400-599 still maps to 502.
const upstreamStatus = Number(parsed.error.status);
const status = Number.isInteger(upstreamStatus) && upstreamStatus >= 400 && upstreamStatus <= 599
? upstreamStatus
: HTTP_STATUS.BAD_GATEWAY;
return createErrorResult(
HTTP_STATUS.BAD_GATEWAY,
status,
parsed.error.message || "Upstream SSE stream failed",
);
}
@@ -508,7 +513,7 @@ export async function handleForcedSSEToJson({
return {
success: true,
response: new Response(JSON.stringify(finalBody), {
response: new Response(JSON.stringify(restoreToolNames(finalBody, toolNameMap)), {
headers: {
"Content-Type": "application/json",
"Access-Control-Allow-Origin": "*",

View File

@@ -6,17 +6,21 @@ import {
} from "../../utils/stream.js";
import { pipeWithDisconnect } from "../../utils/streamHandler.js";
import { PROVIDERS } from "../../config/providers.js";
import { STREAM_STALL_TIMEOUT_MS } from "../../config/runtimeConfig.js";
import { HTTP_STATUS, STREAM_STALL_TIMEOUT_MS } from "../../config/runtimeConfig.js";
import { buildAbortedResponsesTerminalBytes } from "../../utils/responsesStreamHelpers.js";
import { buildStreamErrorBytes } from "../../utils/streamHelpers.js";
import {
buildRequestDetail,
extractRequestConfig,
saveUsageStats,
formatDoneLine,
tokensForDetail,
shouldPersistRequestDetail,
} from "./requestDetail.js";
import { streamStatusForContent } from "../../utils/streamErrorPatterns.js";
import { saveRequestDetail } from "@/lib/usageDb.js";
import { SSE_HEADERS_CORS as SSE_HEADERS } from "../../utils/sseConstants.js";
import { upstreamResponseHeaders } from "../../utils/upstreamHeaders.js";
// Codex returns Responses API SSE → which client format to translate INTO, by request sourceFormat.
// Gemini-family all map to ANTIGRAVITY decoder; unknown sources fall back to OPENAI.
@@ -31,63 +35,20 @@ const CODEX_SOURCE_TO_TARGET = {
/**
* Determine which SSE transform stream to use based on provider/format.
*/
function buildTransformStream({
provider,
sourceFormat,
targetFormat,
userAgent,
reqLogger,
toolNameMap,
customToolNames,
model,
connectionId,
body,
onStreamComplete,
apiKey,
}) {
const isDroidCLI =
userAgent?.toLowerCase().includes("droid") ||
userAgent?.toLowerCase().includes("codex-cli");
// Responses-API providers (e.g. codex) emit Responses SSE → translate into client format
const isResponsesProvider =
PROVIDERS[provider]?.format === FORMATS.OPENAI_RESPONSES;
const needsCodexTranslation =
isResponsesProvider &&
targetFormat === FORMATS.OPENAI_RESPONSES &&
!isDroidCLI;
function buildTransformStream({ provider, sourceFormat, targetFormat, userAgent, reqLogger, toolNameMap, customToolNames, model, connectionId, body, onStreamComplete, apiKey, credentials }) {
const isDroidCLI = userAgent?.toLowerCase().includes("droid") || userAgent?.toLowerCase().includes("codex-cli");
// Responses-API providers (e.g. codex) emit Responses SSE → translate into client format
const isResponsesProvider = PROVIDERS[provider]?.format === FORMATS.OPENAI_RESPONSES;
const needsCodexTranslation = isResponsesProvider && targetFormat === FORMATS.OPENAI_RESPONSES && !isDroidCLI;
if (needsCodexTranslation) {
const codexTarget = CODEX_SOURCE_TO_TARGET[sourceFormat] || FORMATS.OPENAI;
return createSSETransformStreamWithLogger(
FORMATS.OPENAI_RESPONSES,
codexTarget,
provider,
reqLogger,
toolNameMap,
model,
connectionId,
body,
onStreamComplete,
apiKey,
customToolNames,
);
}
if (needsCodexTranslation) {
const codexTarget = CODEX_SOURCE_TO_TARGET[sourceFormat] || FORMATS.OPENAI;
return createSSETransformStreamWithLogger(FORMATS.OPENAI_RESPONSES, codexTarget, provider, reqLogger, toolNameMap, model, connectionId, body, onStreamComplete, apiKey, customToolNames, credentials);
}
if (needsTranslation(targetFormat, sourceFormat)) {
return createSSETransformStreamWithLogger(
targetFormat,
sourceFormat,
provider,
reqLogger,
toolNameMap,
model,
connectionId,
body,
onStreamComplete,
apiKey,
customToolNames,
);
}
if (needsTranslation(targetFormat, sourceFormat)) {
return createSSETransformStreamWithLogger(targetFormat, sourceFormat, provider, reqLogger, toolNameMap, model, connectionId, body, onStreamComplete, apiKey, customToolNames, credentials);
}
return createPassthroughStreamWithLogger(
provider,
@@ -103,42 +64,14 @@ function buildTransformStream({
/**
* Handle streaming response — pipe provider SSE through transform stream to client.
*/
export async function handleStreamingResponse({
providerResponse,
provider,
model,
sourceFormat,
targetFormat,
userAgent,
body,
stream,
translatedBody,
finalBody,
requestStartTime,
connectionId,
apiKey,
clientRawRequest,
onRequestSuccess,
reqLogger,
toolNameMap,
customToolNames,
streamController,
onStreamComplete,
streamDetailId,
pxpipe,
reqTag,
log,
}) {
if (onRequestSuccess) {
Promise.resolve()
.then(onRequestSuccess)
.catch((err) => {
console.error(
"[ChatCore] onRequestSuccess failed:",
err?.message || err,
);
});
}
export async function handleStreamingResponse({ providerResponse, provider, model, sourceFormat, targetFormat, userAgent, body, stream, translatedBody, finalBody, requestStartTime, connectionId, apiKey, clientRawRequest, onRequestSuccess, reqLogger, toolNameMap, customToolNames, streamController, onStreamComplete, pxpipe, reqTag, log, credentials }) {
if (onRequestSuccess) {
Promise.resolve()
.then(onRequestSuccess)
.catch(err => {
console.error("[ChatCore] onRequestSuccess failed:", err?.message || err);
});
}
// When upstream returns HTML/text instead of SSE (e.g. Cloudflare 5xx error
// page), piping it through the SSE transform stream causes Next.js
@@ -197,28 +130,23 @@ export async function handleStreamingResponse({
};
}
const transformStream = buildTransformStream({
provider,
sourceFormat,
targetFormat,
userAgent,
reqLogger,
toolNameMap,
customToolNames,
model,
connectionId,
body,
onStreamComplete,
apiKey,
});
const transformStream = buildTransformStream({ provider, sourceFormat, targetFormat, userAgent, reqLogger, toolNameMap, customToolNames, model, connectionId, body, onStreamComplete, apiKey, credentials });
// Responses passthrough: synthesize response.failed + [DONE] if the stream aborts/stalls before a terminal event
// Terminal bytes when the stream aborts after HTTP 200 was already sent, so the
// client sees a real error instead of a silently truncated stream.
// Responses passthrough keeps its own response.failed shape; every other client
// format gets the OpenAI error frame + [DONE], or `event: error` for Claude.
const isResponsesPassthrough =
sourceFormat === FORMATS.OPENAI_RESPONSES &&
targetFormat === FORMATS.OPENAI_RESPONSES;
const onAbortTerminal = isResponsesPassthrough
? buildAbortedResponsesTerminalBytes
: null;
: (message) =>
buildStreamErrorBytes(
HTTP_STATUS.GATEWAY_TIMEOUT,
message,
sourceFormat,
);
const stallTimeoutMs =
PROVIDERS[provider]?.stallTimeoutMs || STREAM_STALL_TIMEOUT_MS;
const transformedBody = pipeWithDisconnect(
@@ -229,38 +157,14 @@ export async function handleStreamingResponse({
stallTimeoutMs,
);
saveRequestDetail(
buildRequestDetail(
{
provider,
model,
connectionId,
apiKey,
latency: { ttft: 0, total: Date.now() - requestStartTime },
tokens: { prompt_tokens: 0, completion_tokens: 0 },
request: extractRequestConfig(body, stream),
providerRequest: finalBody || translatedBody || null,
providerResponse: "[Streaming - raw response not captured]",
response: {
content: "[Streaming in progress...]",
thinking: null,
type: "streaming",
},
pxpipe,
status: "success",
},
{ id: streamDetailId },
),
).catch((err) => {
console.error(
"[RequestDetail] Failed to save streaming request:",
err.message,
);
});
return {
success: true,
response: new Response(transformedBody, { headers: SSE_HEADERS }),
response: new Response(transformedBody, {
headers: {
...SSE_HEADERS,
...upstreamResponseHeaders(providerResponse.headers),
},
}),
};
}
@@ -282,6 +186,7 @@ export function buildOnStreamComplete({
reqTag,
log,
streamErrorPatterns,
persistUsage = "all",
}) {
const streamDetailId = `${Date.now()}-${Math.random().toString(36).slice(2, 11)}`;
@@ -292,38 +197,41 @@ export function buildOnStreamComplete({
};
const safeContent = contentObj?.content || "[Empty streaming response]";
const safeThinking = contentObj?.thinking || null;
const rawProviderText = typeof contentObj?.rawProviderText === "string" ? contentObj.rawProviderText : "";
saveRequestDetail(
buildRequestDetail(
{
provider,
model,
connectionId,
apiKey,
latency,
tokens: usage || { prompt_tokens: 0, completion_tokens: 0 },
request: extractRequestConfig(body, stream),
providerRequest: finalBody || translatedBody || null,
providerResponse: safeContent,
response: {
content: safeContent,
thinking: safeThinking,
type: "streaming",
if (shouldPersistRequestDetail(persistUsage, "success")) {
saveRequestDetail(
buildRequestDetail(
{
provider,
model,
connectionId,
apiKey,
latency,
tokens: tokensForDetail(usage),
request: extractRequestConfig(body, stream),
providerRequest: finalBody || translatedBody || null,
providerResponse: rawProviderText || safeContent,
response: {
content: safeContent,
thinking: safeThinking,
type: "streaming",
},
pxpipe,
status: streamStatusForContent(
streamErrorPatterns?.[provider],
safeContent,
),
},
pxpipe,
status: streamStatusForContent(
streamErrorPatterns?.[provider],
safeContent,
),
},
{ id: streamDetailId },
),
).catch((err) => {
console.error(
"[RequestDetail] Failed to update streaming content:",
err.message,
);
});
{ id: streamDetailId },
),
).catch((err) => {
console.error(
"[RequestDetail] Failed to update streaming content:",
err.message,
);
});
}
// Persist stream usage to DB (no console line; the "📊 done" line below is authoritative)
saveUsageStats({

View File

@@ -1,4 +1,4 @@
// Web Fetch handler — dispatches to firecrawl, jina-reader, tavily, exa
// Web Fetch handler — dispatches to firecrawl, jina-reader, tavily, exa, ollama
// Returns normalized shape across all providers
const DEFAULT_TIMEOUT_MS = 15000;
@@ -56,8 +56,8 @@ function parseJinaTitle(text) {
return m ? m[1].trim() : null;
}
function buildData({ provider, url, title, format, text, costUsd, responseMs, upstreamMs }) {
return {
function buildData({ provider, url, title, format, text, links, costUsd, responseMs, upstreamMs }) {
const data = {
provider,
url,
title: title || null,
@@ -66,6 +66,8 @@ function buildData({ provider, url, title, format, text, costUsd, responseMs, up
usage: { fetch_cost_usd: costUsd ?? null },
metrics: { response_time_ms: responseMs, upstream_latency_ms: upstreamMs }
};
if (Array.isArray(links)) data.links = links;
return data;
}
async function readJsonOrText(res) {
@@ -115,6 +117,18 @@ export async function handleFetchCore({ url, format, maxCharacters, provider, pr
if (provider === "exa") {
return await runExa({ url, fmt, timeoutMs, apiKey, maxCharacters, costPerQuery, startedAt });
}
if (provider === "ollama") {
return await runOllama({
url,
fmt,
timeoutMs,
apiKey,
maxCharacters,
costPerQuery,
startedAt,
baseUrl: providerConfig?.baseUrl,
});
}
return { success: false, status: 400, error: `Unsupported provider: ${provider}` };
} catch (err) {
log?.("fetch handler error:", err?.message || err);
@@ -241,3 +255,56 @@ async function runExa({ url, fmt, timeoutMs, apiKey, maxCharacters, costPerQuery
})
};
}
async function runOllama({
url,
fmt,
timeoutMs,
apiKey,
maxCharacters,
costPerQuery,
startedAt,
baseUrl,
}) {
const upstreamStart = Date.now();
const r = await tryFetch(baseUrl, {
method: "POST",
headers: {
"content-type": "application/json",
...(apiKey ? { authorization: `Bearer ${apiKey}` } : {})
},
body: JSON.stringify({ url })
}, timeoutMs);
if (!r.ok) {
return { success: false, status: r.timeout ? 504 : 502, error: r.error };
}
const upstreamMs = Date.now() - upstreamStart;
const { json, text: responseText } = await readJsonOrText(r.res);
if (!r.res.ok) {
const error = json?.error
|| json?.message
|| responseText?.slice(0, 500)
|| `Ollama error: ${r.res.status}`;
return { success: false, status: r.res.status, error };
}
if (!json || typeof json.content !== "string") {
return { success: false, status: 502, error: "Ollama returned an empty or invalid web fetch response" };
}
const text = truncate(json.content, maxCharacters);
return {
success: true,
data: buildData({
provider: "ollama",
url,
title: json.title || null,
format: fmt,
text,
links: json.links,
costUsd: costPerQuery,
responseMs: Date.now() - startedAt,
upstreamMs
})
};
}

View File

@@ -0,0 +1,266 @@
import { Buffer } from "node:buffer";
// Gemini Live API realtime STT transport.
//
// The REST generateContent path (sttCore.transcribeGemini) only transcribes
// whole files inline. The Live API's `:bidiGenerateContent` WebSocket is the
// streaming counterpart: audio goes up as realtimeInput mediaChunks and the
// server pushes incremental `serverContent.inputTranscription` events back.
// This module owns the socket lifecycle only — envelope/response shaping
// stays in sttCore so the engine's single STT exit shape is preserved.
//
// Marker contract: dispatched from sttCore's format-switch when the model
// entry carries `transport: "gemini-live"` (registry) or the caller passes a
// transport string (custom models). Never keyed on a hardcoded model id here.
//
// Transport behavior:
// - Node >= 22 global WebSocket (undici). No new dependency.
// - Live API expects low-latency PCM; other containers are forwarded with
// their declared MIME unchanged (provider-side rejection is surfaced).
// - Text accumulation is append-only over inputTranscription segments and
// ends on serverContent.turnComplete (or graceful close with partial text).
// - Transcription deltas are kept per-frame (chunks[]) so sttCore can shape
// verbose_json segments without fabricating timestamps. goAway advisements
// rotate the socket once per call: setup replay + byte-offset resume.
const SETUP_TIMEOUT_MS = 10_000; // open → setupComplete
const TURN_TIMEOUT_MS = 60_000; // audio streamed → turnComplete
const MAX_TIMEOUT_MS = 300_000; // clamp ceiling for client-supplied lifecycle knobs
const CHUNK_BYTES = 16_384; // ~0.5s of 16-bit 16kHz mono PCM
const GOAWAY_RECONNECTS = 1; // socket rotations honoured per call
class GeminiLiveError extends Error {
constructor(message, status) {
super(message);
this.name = "GeminiLiveError";
this.status = status || 502;
}
}
// REST base (https://host/v1beta/models) → Live WS base
// (wss://host/ws/api/v1beta/models), then the bidiGenerateContent endpoint.
function toLiveWsUrl(baseUrl, model, token) {
const url = new URL(baseUrl);
url.protocol = "wss:";
if (!url.pathname.startsWith("/ws/")) url.pathname = `/ws/api${url.pathname}`;
const base = url.toString().replace(/\/+$/, "");
return `${base}/${encodeURIComponent(model)}:bidiGenerateContent?key=${encodeURIComponent(token || "")}`;
}
// Bind socket events supporting BOTH handler styles: addEventListener
// (browser WebSocket, undici) and onopen/onmessage property assignment
// (minimal polyfills). Whichever the implementation exposes, it works.
function bindSocket(ws, { onOpen, onMessage, onError, onClose }) {
if (typeof ws.addEventListener === "function") {
ws.addEventListener("open", onOpen);
ws.addEventListener("message", onMessage);
ws.addEventListener("error", onError);
ws.addEventListener("close", onClose);
return;
}
ws.onopen = onOpen;
ws.onmessage = onMessage;
ws.onerror = onError;
ws.onclose = onClose;
}
function parseFrame(data) {
try {
return JSON.parse(typeof data === "string" ? data : String(data));
} catch {
return null; // non-JSON frames carry no Live API semantics
}
}
function firstStringField(formData, key) {
const v = typeof formData?.get === "function" ? formData.get(key) : null;
return typeof v === "string" && v.trim() ? v.trim() : "";
}
// Lifecycle knobs the live registry entry advertises in params[]
// (setup/turn timeouts). They ride the same formData pass-through sttCore
// gives every transport — no sttCore change needed to reach this leaf.
function firstNumberField(formData, key, fallback) {
const n = Number(firstStringField(formData, key));
return Number.isFinite(n) && n > 0 ? Math.min(n, MAX_TIMEOUT_MS) : fallback;
}
/**
* Transcribe an audio File via the Gemini Live bidirectional stream.
* @returns {Promise<{text: string, chunks: string[]}>} transcript plus the raw
* incremental inputTranscription deltas (sttCore shapes verbose_json from them).
* @throws {GeminiLiveError} with .status for the sttCore error envelope.
*/
export async function transcribeGeminiLive({ cfg, file, model, token, formData, mimeType }) {
const WS = globalThis.WebSocket;
if (!WS) throw new GeminiLiveError("Gemini Live transport needs global WebSocket (Node >= 22)", 502);
const buf = Buffer.from(await file.arrayBuffer());
if (!buf.length) throw new GeminiLiveError("Empty audio file", 400);
const instruction = firstStringField(formData, "prompt") || "Transcribe the spoken audio verbatim.";
const language = firstStringField(formData, "language");
const setupTimeoutMs = firstNumberField(formData, "setup_timeout_ms", SETUP_TIMEOUT_MS);
const turnTimeoutMs = firstNumberField(formData, "turn_timeout_ms", TURN_TIMEOUT_MS);
// system_instruction (registry param) overrides the built-in transcription
// directive wholesale; prompt/language only shape the default.
const instructionOverride = firstStringField(formData, "system_instruction");
const systemText = instructionOverride
|| (language ? `${instruction} Language: ${language}.` : instruction);
const wsUrl = toLiveWsUrl(cfg.baseUrl, model, token);
return await new Promise((resolve, reject) => {
let text = "";
const chunks = []; // raw inputTranscription deltas, shaped by sttCore
let settled = false;
let timer = null;
let goAwayTimer = null;
let ws = null;
let generation = 0; // socket identity: superseded closes never settle
let sentBytes = 0; // audio prefix already handed to the live socket
let goAwayReconnects = GOAWAY_RECONNECTS;
const arm = (ms, message) => {
if (timer) clearTimeout(timer);
timer = setTimeout(() => fail(new GeminiLiveError(message, 504)), ms);
};
const shutdown = () => {
if (timer) { clearTimeout(timer); timer = null; }
if (goAwayTimer) { clearTimeout(goAwayTimer); goAwayTimer = null; }
// ws is null until the first open() dials (and stays null when the
// constructor throws) — fail() runs shutdown() on that path.
if (!ws) return;
try {
if (ws.readyState === WS.OPEN || ws.readyState === WS.CONNECTING) ws.close(1000);
} catch { /* socket already dead — outcome is already settled */ }
};
const succeed = () => {
if (settled) return;
settled = true;
shutdown();
resolve({ text, chunks });
};
const fail = (err) => {
if (settled) return;
settled = true;
shutdown();
reject(err);
};
const send = (frame) => {
if (ws.readyState !== WS.OPEN) return false;
try {
ws.send(JSON.stringify(frame));
} catch {
return false; // socket died mid-send — streamAudioAndPrompt maps this to a 502
}
return true;
};
// Streams every byte not yet sent, then the flushing text turn. After a
// goAway rotation this resumes from sentBytes — no audio re-upload.
const streamAudioAndPrompt = () => {
for (let off = sentBytes; off < buf.length; off += CHUNK_BYTES) {
const mediaChunk = buf.subarray(off, off + CHUNK_BYTES).toString("base64");
if (!send({ realtimeInput: { mediaChunks: [{ mimeType, data: mediaChunk }] } })) {
fail(new GeminiLiveError("Gemini Live socket closed while streaming audio", 502));
return;
}
sentBytes = Math.min(off + CHUNK_BYTES, buf.length);
}
// Final user turn: flushes the recognizer and yields turnComplete.
send({ clientContent: { turns: [{ parts: [{ text: systemText }] }], turnComplete: true } });
};
// goAway: the server names the instant it will force-close this socket.
// Graceful play = rotate BEFORE the deadline: retire the live socket,
// dial a fresh one, replay setup, resume audio from sentBytes — text and
// chunks survive the hop. Once the advisory budget is spent a later
// goAway is left to the close path, which settles on partial transcript.
const scheduleGoAwayReconnect = (goAway) => {
if (settled || goAwayTimer || goAwayReconnects <= 0) return;
const deadline = Date.parse(typeof goAway?.time === "string" ? goAway.time : "");
const delay = Number.isFinite(deadline)
? Math.max(0, Math.min(deadline - Date.now(), setupTimeoutMs))
: 0;
goAwayTimer = setTimeout(() => {
goAwayTimer = null;
if (settled) return;
goAwayReconnects--;
generation++;
try { ws?.close(1000); } catch { /* deadline crossed mid-flight — re-dial anyway */ }
open();
}, delay);
};
const open = () => {
const gen = ++generation;
try {
ws = new WS(wsUrl);
} catch {
fail(new GeminiLiveError("Gemini Live websocket connection failed", 502));
return;
}
bindSocket(ws, {
onOpen: () => {
if (settled || gen !== generation) return;
arm(setupTimeoutMs, "Gemini Live timed out waiting for setupComplete");
send({
setup: {
model: `models/${model}`,
generationConfig: {
responseModalities: ["TEXT"],
inputAudioTranscription: {},
},
systemInstruction: { parts: [{ text: systemText }] },
},
});
},
onMessage: (ev) => {
if (settled || gen !== generation) return;
const frame = parseFrame(ev?.data);
if (!frame) return;
if (frame.error) {
const e = frame.error;
fail(new GeminiLiveError(`Gemini Live error${e.status ? ` (${e.status})` : ""}: ${e.message || "unknown"}`, 502));
return;
}
if (frame.goAway) {
scheduleGoAwayReconnect(frame.goAway);
return;
}
const sc = frame.serverContent;
if (!sc) return;
const delta = typeof sc.inputTranscription?.text === "string" ? sc.inputTranscription.text : "";
// Trim before testing: a padding-only frame carries no transcript and
// must not make an empty run look like a partial success on close.
if (delta.trim()) {
text += delta;
chunks.push(delta);
}
if (sc.setupComplete) {
arm(turnTimeoutMs, "Gemini Live transcription timed out");
streamAudioAndPrompt();
return;
}
if (sc.turnComplete) succeed();
},
onError: () => {
if (settled || gen !== generation) return;
fail(new GeminiLiveError("Gemini Live websocket connection failed", 502));
},
onClose: (ev) => {
if (settled || gen !== generation) return;
// Partial transcript beats a hard error on graceful close; silence is one.
if (text.trim()) succeed();
else fail(new GeminiLiveError(`Gemini Live socket closed before completion${ev?.code ? ` (code ${ev.code})` : ""}`, 502));
},
});
};
open();
});
}

View File

@@ -1,6 +1,6 @@
// Antigravity image adapter - delegates to the executor for correct request
// envelope (project, model, requestType, sessionId) and auth headers.
import { nowSec } from "./_base.js";
import { nowSec, sizeToAspectRatio } from "./_base.js";
import { getExecutor } from "../../executors/index.js";
// Convert image input (data URI or raw base64) to Gemini inlineData part
@@ -31,6 +31,19 @@ export default {
const executor = getExecutor("antigravity");
if (!executor) throw new Error("Antigravity executor not found");
// Ensure we use an image model for image generation
const isImageModel = (m) => /image|imagen|image-generation/i.test(m || "");
let targetModel = isImageModel(model) ? model : "gemini-3.1-flash-image";
// If body.size is provided, resolve aspect ratio and append to model
if (body.size && typeof body.size === "string") {
const ratio = sizeToAspectRatio(body.size);
const suffix = ratio.replace(":", "x");
if (!targetModel.includes(suffix)) {
targetModel = `${targetModel}-${suffix}`;
}
}
// Build parts: text prompt + optional input image for editing
const parts = [{ text: body.prompt }];
const imageInput = body.image || (Array.isArray(body.images) && body.images[0]);
@@ -44,7 +57,7 @@ export default {
};
const result = await executor.execute({
model,
model: targetModel,
body: chatBody,
stream: false,
credentials,

View File

@@ -2,13 +2,21 @@
import { randomUUID } from "node:crypto";
import { nowSec } from "./_base.js";
import { PROVIDERS } from "../../config/providers.js";
import { CODEX_CLI_VERSION } from "../../config/appConstants.js";
const CODEX_RESPONSES_URL = PROVIDERS["codex"].baseUrl;
const CODEX_USER_AGENT = "codex_cli_rs/0.136.0";
const CODEX_VERSION = "0.136.0";
const CODEX_USER_AGENT = `codex_cli_rs/${CODEX_CLI_VERSION}`;
const CODEX_ORIGINATOR = "codex_cli_rs";
const CODEX_MODEL_SUFFIX = "-image";
const CODEX_REF_DETAIL = "high";
const CODEX_IMAGES_MAIN_MODEL = "gpt-5.5";
const CODEX_TOOL_IMAGE_MODELS = new Set([
"gpt-image-1.5",
"gpt-image-2",
"gpt-image-2.5",
"gpt-image-2.5-flare",
"gpt-image-2.5-sunburst",
]);
function decodeAccountId(idToken) {
try {
@@ -27,6 +35,13 @@ function stripImageSuffix(model) {
return model.endsWith(CODEX_MODEL_SUFFIX) ? model.slice(0, -CODEX_MODEL_SUFFIX.length) : model;
}
function resolveCodexImageModels(model) {
if (CODEX_TOOL_IMAGE_MODELS.has(model)) {
return { responsesModel: CODEX_IMAGES_MAIN_MODEL, toolModel: model };
}
return { responsesModel: stripImageSuffix(model), toolModel: null };
}
function toDataUrl(input) {
if (!input || typeof input !== "string") return null;
if (/^data:image\//i.test(input) || /^https?:\/\//i.test(input)) return input;
@@ -157,7 +172,7 @@ export default {
"originator": CODEX_ORIGINATOR,
"session_id": randomUUID(),
"user-agent": CODEX_USER_AGENT,
"version": CODEX_VERSION,
"version": CODEX_CLI_VERSION,
"x-client-request-id": randomUUID(),
};
},
@@ -167,21 +182,26 @@ export default {
const single = toDataUrl(body.image);
if (single) refs.push(single);
const detail = body.image_detail || CODEX_REF_DETAIL;
const { responsesModel, toolModel } = resolveCodexImageModels(model);
const imgTool = { type: "image_generation", output_format: (body.output_format || "png").toLowerCase() };
if (toolModel) {
imgTool.action = refs.length > 0 ? "edit" : "generate";
imgTool.model = toolModel;
}
if (body.size && body.size !== "") imgTool.size = body.size;
if (body.quality && body.quality !== "") imgTool.quality = body.quality;
if (body.background && body.background !== "") imgTool.background = body.background;
return {
model: stripImageSuffix(model),
model: responsesModel,
instructions: "",
input: [{ type: "message", role: "user", content: buildContent(body.prompt, refs, detail) }],
tools: [imgTool],
tool_choice: "auto",
tool_choice: toolModel ? { type: "image_generation" } : "auto",
parallel_tool_calls: false,
prompt_cache_key: randomUUID(),
stream: true,
store: false,
reasoning: null,
reasoning: toolModel ? { effort: "medium", summary: "auto" } : null,
};
},
// Custom: codex parses SSE → either pipe to client or collect b64

View File

@@ -1,18 +1,93 @@
// HuggingFace Inference API — returns binary image
import { nowSec } from "./_base.js";
// HuggingFace Inference Providers router — returns binary image
//
// The router is a switchboard in front of many inference providers and is
// addressed as `<baseUrl>/<provider>/<providerModelId>`. `providerModelId` is
// the id the *provider* uses, which is not the Hub model id, so it is resolved
// through `imageConfig.modelMap` (built from the Hub API's
// inferenceProviderMapping and limited to providers the router forwards to).
//
// The legacy `api-inference.huggingface.co` host is gone (DNS ENOTFOUND) and is
// deliberately not referenced anywhere here.
import { nowSec, urlToBase64 } from "./_base.js";
import { PROVIDER_MEDIA } from "../../providers/index.js";
const BASE_URL = PROVIDER_MEDIA["huggingface"]?.imageConfig?.baseUrl;
const imageConfig = () => PROVIDER_MEDIA["huggingface"]?.imageConfig || {};
const BASE_URL = imageConfig().baseUrl;
const MODEL_MAP = imageConfig().modelMap || {};
// A plain-object lookup returns inherited truthy values for keys like "toString" or
// "constructor", which would build nonsense URLs. Resolve own keys only.
const lookup = (model) => (Object.hasOwn(MODEL_MAP, model) ? MODEL_MAP[model] : undefined);
// modelMap values are either a bare path (text-to-image) or { path, task }.
const mappingPath = (entry) => (typeof entry === "string" ? entry : entry.path);
const mappingTask = (entry) => (typeof entry === "string" ? "text-to-image" : entry.task || "text-to-image");
// A connection may point at its own endpoint (self-hosted Text Generation
// Inference / TGI container). That endpoint already knows its own model ids, so
// the router mapping does not apply and the Hub id is passed through verbatim.
function customBaseUrl(creds) {
const url = creds?.providerSpecificData?.baseUrl;
return typeof url === "string" && url.trim() ? url.trim().replace(/\/+$/, "") : null;
}
// The router's image-to-image payload wants raw base64 — not a data URL, not a URL.
// Accept every shape our own callers use (data URL, bare base64, remote URL, array).
async function sourceImage(body) {
const raw = body?.image || (Array.isArray(body?.images) ? body.images[0] : null);
if (typeof raw !== "string" || !raw.trim()) return null;
const value = raw.trim();
if (/^https?:\/\//i.test(value)) return await urlToBase64(value);
const match = /^data:image\/[^;]+;base64,(.+)$/i.exec(value);
return match ? match[1] : value;
}
export default {
buildUrl: (model) => `${BASE_URL}/${model}`,
buildUrl: (model, creds) => {
const override = customBaseUrl(creds);
if (override) {
// The model id is client-controlled; on a custom endpoint it lands in a URL
// path verbatim, so reject traversal/query injection (mirrors sttCore's guard).
if (model.includes("..") || model.includes("//") || /[?#]/.test(model)) {
throw new Error(`HuggingFace: invalid model ID "${model}"`);
}
return `${override}/${model}`;
}
const entry = lookup(model);
if (!entry) {
throw new Error(
`HuggingFace: no HuggingFace router mapping for model "${model}". ` +
`Add it to imageConfig.modelMap in open-sse/providers/registry/huggingface.js, ` +
`or set a custom base URL on the connection.`
);
}
return `${BASE_URL}/${mappingPath(entry)}`;
},
buildHeaders: (creds) => {
const headers = { "Content-Type": "application/json" };
const key = creds?.apiKey || creds?.accessToken;
if (key) headers["Authorization"] = `Bearer ${key}`;
return headers;
},
buildBody: (_model, body) => ({ inputs: body.prompt }),
buildBody: async (model, body) => {
const entry = lookup(model);
const task = mappingTask(entry || "");
if (task === "image-to-image") {
const image = await sourceImage(body);
if (!image) {
throw new Error(
`HuggingFace: model "${model}" requires a source image. ` +
`Send it as "image" (or "images") in the request body.`
);
}
// inputs carries the source image; the prompt moves under parameters.
return { inputs: image, parameters: { prompt: body.prompt } };
}
return { inputs: body.prompt };
},
// HF returns raw image bytes — convert to b64_json
async parseResponse(response) {
const buf = await response.arrayBuffer();

View File

@@ -347,6 +347,81 @@ function buildSearxngRequest(config, params) {
};
}
function buildXquikRequest(config, params) {
const apiKey = params.token;
if (!apiKey) throw new Error("Xquik requires an API key");
const queryType = getProviderSetting(params, "queryType");
if (queryType && !["Latest", "Top"].includes(queryType)) {
throw new Error("Xquik queryType must be Latest or Top");
}
const qp = new URLSearchParams({
q: params.query,
limit: String(params.maxResults),
});
const cursor = getProviderSetting(params, "cursor");
if (cursor) qp.set("cursor", cursor);
if (queryType) qp.set("queryType", queryType);
if (params.language) qp.set("language", params.language);
return {
url: `${resolveBaseUrl(config, params)}?${qp}`,
init: {
method: "GET",
headers: { Accept: "application/json", "x-api-key": apiKey },
},
};
}
// ── Ollama Cloud web_search ──────────────────────────────────────────────
// POST https://ollama.com/api/web_search { query, max_results }
// Response: { results: [{ title, url, content, published_at? }] }
function buildOllamaSearchRequest(config, params) {
const body = { query: params.query, max_results: params.maxResults };
if (params.country) body.country = params.country;
if (params.language) body.language = params.language;
return {
url: resolveBaseUrl(config, params),
init: {
method: "POST",
headers: {
"Content-Type": "application/json",
...(params.token ? { Authorization: `Bearer ${params.token}` } : {}),
},
body: JSON.stringify(body),
},
};
}
// ── GLM Coding plan MCP web_search_prime ──────────────────────────────────
// POST https://api.z.ai/api/mcp/web_search_prime/mcp
// JSON-RPC envelope: { jsonrpc, id, method: "tools/call",
// params: { name: "web_search_prime", arguments: { search_query, count } } }
// Response: { result: { content: [{ type: "text", text: "<json>" }] } }
function buildGlmSearchRequest(config, params) {
const body = {
jsonrpc: "2.0",
id: `9r-${Date.now()}`,
method: "tools/call",
params: {
name: "web_search_prime",
arguments: { search_query: params.query, count: params.maxResults },
},
};
return {
url: resolveBaseUrl(config, params),
init: {
method: "POST",
headers: {
"Content-Type": "application/json",
...(params.token ? { Authorization: `Bearer ${params.token}` } : {}),
},
body: JSON.stringify(body),
},
};
}
// ── Dispatcher ──────────────────────────────────────────────────────────
const BUILDERS = {
@@ -360,6 +435,9 @@ const BUILDERS = {
"searchapi": buildSearchApiRequest,
"youcom": buildYouComRequest,
"searxng": buildSearxngRequest,
"xquik": buildXquikRequest,
"ollama-search": buildOllamaSearchRequest,
"glm": buildGlmSearchRequest,
};
/**

View File

@@ -1,8 +1,10 @@
/**
* Wrap chat-completions endpoints (with built-in web search) into the unified
* /v1/search response format. Supports gemini, openai, xai, kimi, minimax, perplexity.
* /v1/search response format. Supports gemini, antigravity, openai, xai, kimi,
* minimax, perplexity.
*/
import { PROVIDER_MEDIA } from "../../providers/index.js";
import { ANTIGRAVITY_IDE_USER_AGENT } from "../../providers/shared.js";
// Default search model + endpoint derive from registry searchViaChat (single source)
const searchModel = (id) => PROVIDER_MEDIA[id]?.searchViaChat?.defaultModel;
@@ -28,13 +30,37 @@ function toResult(c, index, provider, retrievedAt) {
score: null,
published_at: null,
favicon_url: null,
content: null,
content: c.content || null,
metadata: {},
citation: { provider, retrieved_at: retrievedAt, rank: index + 1 },
provider_raw: null
};
}
// Antigravity search request envelope (mirrors the IDE client)
const AG_CLIENT_NAME = "antigravity";
const AG_SEARCH_GENERATION_CONFIG = { temperature: 1.0, maxOutputTokens: 8192 };
const AG_CONTEXT_BEFORE = 150;
const AG_CONTEXT_AFTER = 250;
/** Widen a grounded segment to its surrounding sentence(s) in the answer text. */
function expandSegment(text, segment) {
const { startIndex, endIndex } = segment || {};
if (!text || !Number.isInteger(startIndex) || !Number.isInteger(endIndex)) return "";
const start = Math.max(0, startIndex - AG_CONTEXT_BEFORE);
const end = Math.min(text.length, endIndex + AG_CONTEXT_AFTER);
let out = text.slice(start, end).trim();
// Drop the partial words the window cut off at either edge
if (start > 0) out = `...${out.replace(/^\S+/, "")}`;
if (end < text.length) out = `${out.replace(/\S+$/, "")}...`;
return out.trim();
}
/** Join deduped grounding pieces, skipping empties. */
function joinPieces(set, sep) {
return [...(set || [])].filter(Boolean).join(sep).trim();
}
/** Coerce a citation that might be a raw URL string or an object. */
function normalizeCitation(c) {
if (!c) return null;
@@ -46,6 +72,8 @@ function normalizeCitation(c) {
/**
* Provider-specific configuration map. All providers must implement:
* { endpoint, defaultModel, buildBody, buildHeaders, extractAnswer }
* Optional: requireCredentials(credentials) → error string when a provider needs
* more than a token (returns null when satisfied).
*/
const CHAT_SEARCH_CONFIG = {
gemini: {
@@ -73,6 +101,71 @@ const CHAT_SEARCH_CONFIG = {
}
},
antigravity: {
endpoint: () => searchEndpoint("antigravity"),
// Upstream 403s on a missing or fabricated project — surface the real cause
requireCredentials: (credentials) =>
credentials?.projectId ? null : "Antigravity account has no projectId — reconnect the account",
buildBody: (query, model, credentials) => ({
project: credentials.projectId,
model,
userAgent: AG_CLIENT_NAME,
requestType: "search",
request: {
contents: [{ role: "user", parts: [{ text: query }] }],
tools: [{ googleSearch: {} }],
generationConfig: AG_SEARCH_GENERATION_CONFIG
}
}),
buildHeaders: (token) => ({
"Content-Type": "application/json",
Authorization: `Bearer ${token}`,
"User-Agent": ANTIGRAVITY_IDE_USER_AGENT
}),
extractAnswer: (data) => {
// Antigravity wraps the Gemini payload in { response: {...} }
const response = data?.response || data;
const candidate = response?.candidates?.[0];
const parts = candidate?.content?.parts || [];
const text = parts.map((p) => p?.text || "").filter(Boolean).join("");
const grounding = candidate?.groundingMetadata || {};
const chunks = grounding.groundingChunks || [];
const supports = grounding.groundingSupports || [];
// Upstream repeats the same source across chunks — key by URL so it stays one citation.
// Map, not a plain object: both the index and the URL come from upstream.
const sources = new Map();
const byIndex = chunks.map((ch) => {
const web = ch?.web;
const url = web?.uri || web?.url || "";
if (!url) return null;
if (!sources.has(url)) sources.set(url, { title: web.title || "", snippets: new Set(), contexts: new Set() });
return sources.get(url);
});
// Each support ties a sentence of the answer back to the chunks that grounded it
for (const s of supports) {
const segment = s?.segment;
const grounded = segment?.text || "";
const expanded = expandSegment(text, segment) || grounded;
for (const idx of s?.groundingChunkIndices || []) {
const source = Number.isInteger(idx) ? byIndex[idx] : null;
if (!source) continue;
if (grounded) source.snippets.add(grounded);
if (expanded) source.contexts.add(expanded);
}
}
const citations = [...sources].map(([url, src]) => {
const snippet = joinPieces(src.snippets, " | ") || src.title;
return { url, title: src.title, snippet, content: joinPieces(src.contexts, "\n\n") || snippet };
});
const tokens = response?.usageMetadata?.totalTokenCount || 0;
return { text, citations, tokens };
}
},
openai: {
endpoint: () => searchEndpoint("openai"),
buildBody: (query, model) => {
@@ -366,13 +459,18 @@ export async function handleChatSearch({
};
}
const credentialError = cfg.requireCredentials?.(credentials);
if (credentialError) {
return { success: false, status: 401, error: credentialError };
}
const limit =
Number.isFinite(maxResults) && maxResults > 0
? Math.floor(maxResults)
: DEFAULT_MAX_RESULTS;
const useModel = model || searchModel(provider);
const url = cfg.endpoint(useModel);
const body = cfg.buildBody(query, useModel);
const body = cfg.buildBody(query, useModel, credentials);
const headers = cfg.buildHeaders(token);
const controller = new AbortController();

View File

@@ -10,6 +10,7 @@
import { buildSearchRequest } from "./callers.js";
import { normalizeSearchResponse } from "./normalizers.js";
import { handleChatSearch } from "./chatSearch.js";
import { fetchPublic } from "../../../src/shared/utils/ssrfGuard.js";
const GLOBAL_TIMEOUT_MS = 15000;
const NON_RETRIABLE = new Set([400, 401, 403, 404]);
@@ -100,7 +101,7 @@ async function tryDedicatedProvider({ provider, providerConfig, body, credential
log?.info?.("SEARCH", `${provider.id} | "${params.query.slice(0, 80)}" | type=${params.searchType}`);
try {
const resp = await fetch(url, { ...init, headers: sanitizeHeaders(init.headers), signal: controller.signal });
const resp = await fetchPublic(url, { ...init, headers: sanitizeHeaders(init.headers), signal: controller.signal });
clearTimeout(timer);
if (!resp.ok) {
const errText = await resp.text().catch(() => "");
@@ -111,6 +112,13 @@ async function tryDedicatedProvider({ provider, providerConfig, body, credential
const normalized = normalizeSearchResponse(provider.id, data, params.query, params.searchType);
const results = normalized.results.slice(0, params.maxResults);
const duration = Date.now() - startTime;
const usage = {
queries_used: 1,
search_cost_usd: providerConfig.costPerQuery ?? null,
};
if (Number.isFinite(providerConfig.creditsPerResult)) {
usage.provider_credits_used = results.length * providerConfig.creditsPerResult;
}
return {
success: true,
@@ -119,7 +127,8 @@ async function tryDedicatedProvider({ provider, providerConfig, body, credential
query: params.query,
results,
answer: null,
usage: { queries_used: 1, search_cost_usd: providerConfig.costPerQuery || 0 },
usage,
...(normalized.pagination ? { pagination: normalized.pagination } : {}),
metrics: { response_time_ms: duration, upstream_latency_ms: duration, total_results_available: normalized.totalResults },
errors: []
}

View File

@@ -199,6 +199,89 @@ function normalizeSearxng(data, _query, _searchType) {
return { results, totalResults: results.length };
}
function normalizeXquik(data, _query, _searchType) {
const now = new Date().toISOString();
const items = Array.isArray(data.tweets) ? data.tweets : [];
const results = items.map((item, idx) => {
const username = typeof item?.author?.username === "string" ? item.author.username : "";
const authorName = typeof item?.author?.name === "string" ? item.author.name : "";
const tweetId = typeof item?.id === "string" ? item.id : String(item?.id || "");
const url = username && tweetId
? `https://x.com/${encodeURIComponent(username)}/status/${encodeURIComponent(tweetId)}`
: tweetId
? `https://x.com/i/web/status/${encodeURIComponent(tweetId)}`
: "";
const author = username ? `@${username}` : authorName || null;
const title = author ? `${author} on X` : "X post";
const imageUrl = Array.isArray(item?.media)
? item.media.find((media) => typeof media?.mediaUrl === "string")?.mediaUrl
: null;
return makeResult("xquik", {
title,
url,
snippet: typeof item?.text === "string" ? item.text : "",
published_at: typeof item?.createdAt === "string" ? item.createdAt : null,
author,
image_url: imageUrl || null,
source_type: "x_post",
full_text: typeof item?.text === "string" ? item.text : undefined,
text_format: "text",
}, idx, now);
});
const nextCursor = typeof data.next_cursor === "string" && data.next_cursor ? data.next_cursor : null;
return {
results,
totalResults: null,
pagination: {
has_more: data.has_next_page === true,
next_cursor: nextCursor,
},
};
}
function normalizeOllamaSearch(data, _query, _searchType) {
const now = new Date().toISOString();
const items = Array.isArray(data?.results) ? data.results : (Array.isArray(data) ? data : []);
const results = items.map((item, idx) =>
makeResult("ollama-search", {
title: item.title,
url: item.url,
snippet: item.content || item.snippet || "",
full_text: item.content,
text_format: "text",
published_at: item.published_at || null,
source_type: item.source || null,
}, idx, now)
);
return { results, totalResults: results.length };
}
function normalizeGlmSearch(data, _query, _searchType) {
const now = new Date().toISOString();
// MCP envelope: { result: { content: [{ type: "text", text: "<json>" }] } }
let payload = data;
const textContent = data?.result?.content?.[0]?.text;
if (typeof textContent === "string") {
try { payload = JSON.parse(textContent); } catch { payload = {}; }
}
const items = Array.isArray(payload?.results) ? payload.results
: Array.isArray(payload?.news) ? payload.news
: Array.isArray(payload) ? payload
: [];
const results = items.map((item, idx) =>
makeResult("glm", {
title: item.title,
url: item.link || item.url,
snippet: item.content || "",
published_at: item.publish_date || item.published_at || null,
favicon_url: item.icon || null,
source_type: item.media || null,
}, idx, now)
);
return { results, totalResults: results.length };
}
const NORMALIZERS = {
"serper": normalizeSerper,
"brave-search": normalizeBrave,
@@ -210,11 +293,14 @@ const NORMALIZERS = {
"searchapi": normalizeSearchApi,
"youcom": normalizeYouCom,
"searxng": normalizeSearxng,
"xquik": normalizeXquik,
"ollama-search": normalizeOllamaSearch,
"glm": normalizeGlmSearch,
};
/**
* Dispatch to the appropriate normalizer based on providerId.
* @returns {{results: Array, totalResults: number|null}}
* @returns {{results: Array, totalResults: number|null, pagination?: object}}
*/
export function normalizeSearchResponse(providerId, data, query, searchType) {
const fn = NORMALIZERS[providerId];

View File

@@ -1,5 +1,7 @@
import { Buffer } from "node:buffer";
import { createErrorResult } from "../utils/error.js";
import { transcribeGeminiLive } from "./geminiLiveStt.js";
import { PROVIDER_MODELS, PROVIDER_ID_TO_ALIAS } from "../config/providerModels.js";
import { HTTP_STATUS } from "../config/runtimeConfig.js";
// Build auth headers from sttConfig + token
@@ -162,11 +164,26 @@ function jsonResponse(obj) {
};
}
// Model-level transport marker (registry models[].transport, e.g. the Gemini
// live STT entry's "gemini-live", or a custom model's stored transport).
// Dispatch reads the marker — never a hardcoded model id — so new realtime
// providers extend sttCore through data, not code.
function resolveModelTransport(provider, model) {
const key = PROVIDER_ID_TO_ALIAS[provider] || provider;
const models = PROVIDER_MODELS[key] || PROVIDER_MODELS[provider];
if (!Array.isArray(models)) return null;
const entry = models.find((m) => m && m.id === model && (m.kind || "llm") === "stt");
const marker = typeof entry?.transport === "string" ? entry.transport.trim() : "";
return marker || null;
}
/**
* STT core handler — dispatch by sttConfig.format.
* STT core handler — dispatch by model transport marker, else sttConfig.format.
* `transport` is the caller-supplied marker override (custom models resolve
* it in the app layer; built-ins fall back to the registry entry marker).
* @returns {Promise<{success, response, status?, error?}>}
*/
export async function handleSttCore({ provider, model, formData, credentials, sttConfig }) {
export async function handleSttCore({ provider, model, formData, credentials, sttConfig, transport }) {
const file = formData.get("file");
if (!file) return createErrorResult(HTTP_STATUS.BAD_REQUEST, "Missing required field: file");
@@ -186,8 +203,29 @@ export async function handleSttCore({ provider, model, formData, credentials, st
return createErrorResult(HTTP_STATUS.UNAUTHORIZED, `No credentials for STT provider: ${provider}`);
}
// Format-switch extension: an explicit caller marker wins over the registry
// marker; with neither, the provider-default sttConfig.format applies.
const marker = (typeof transport === "string" && transport.trim()) ? transport.trim() : resolveModelTransport(provider, model);
try {
switch (cfg.format) {
switch (marker || cfg.format) {
case "gemini-live": {
const live = await transcribeGeminiLive({ cfg, file, model, token, formData, mimeType: resolveAudioContentType(file) });
// response_format parity with the OpenAI-compatible transport: default
// envelope stays {text}; verbose_json adds segments mapped from the
// Live API's incremental inputTranscription deltas. Those frames carry
// NO timestamps, so segments expose {id,text} only (id = delta order,
// Whisper-compatible 0-based) — start/end/duration are deliberately
// absent rather than fabricated as zeros, which would misrepresent
// provider data to callers diffing transports.
const fmt = typeof formData?.get === "function"
? String(formData.get("response_format") ?? "").trim().toLowerCase()
: "";
if (fmt === "verbose_json") {
return jsonResponse({ text: live.text, segments: live.chunks.map((segText, id) => ({ id, text: segText })) });
}
return jsonResponse({ text: live.text });
}
case "deepgram": return await transcribeDeepgram(cfg, file, model, token, formData);
case "assemblyai": return await transcribeAssemblyAI(cfg, file, model, token);
case "nvidia-asr": return await transcribeNvidia(cfg, file, model, token);
@@ -196,6 +234,6 @@ export async function handleSttCore({ provider, model, formData, credentials, st
default: return await transcribeOpenAICompatible(cfg, file, model, token, formData);
}
} catch (err) {
return createErrorResult(HTTP_STATUS.BAD_GATEWAY, err.message || "STT request failed");
return createErrorResult(err.status || HTTP_STATUS.BAD_GATEWAY, err.message || "STT request failed");
}
}

View File

@@ -0,0 +1,95 @@
import { createErrorResult, parseUpstreamError, formatProviderError } from "../utils/error.js";
import { HTTP_STATUS, FETCH_CONNECT_TIMEOUT_MS } from "../config/runtimeConfig.js";
import { PROVIDER_MEDIA } from "../providers/index.js";
import { generateSessionId } from "../executors/opencode-zen.js";
/**
* Core System One (Jev) handler — native decision payload pass-through.
* URL/headers come from the registry's systemoneConfig; body and JSON response
* are forwarded untouched (decision models have no chat translation layer).
*
* @returns {Promise<{ success: boolean, response: Response, usage?: object, status?: number, error?: string }>}
*/
export async function handleSystemoneCore({
body,
modelInfo,
credentials,
log,
onRequestSuccess,
}) {
const { provider, model } = modelInfo;
const cfg = PROVIDER_MEDIA[provider]?.systemoneConfig;
if (!cfg?.baseUrl) {
return createErrorResult(
HTTP_STATUS.BAD_REQUEST,
`Provider '${provider}' does not support System One.`
);
}
// Validate input at the trust boundary; question-level shape is upstream's job.
if (body.state === undefined || body.state === null) {
return createErrorResult(HTTP_STATUS.BAD_REQUEST, "Missing required field: state");
}
if (!body.questions || typeof body.questions !== "object" || Array.isArray(body.questions)) {
return createErrorResult(HTTP_STATUS.BAD_REQUEST, "Missing required field: questions");
}
// noAuth free lanes carry accessToken "public" from the credential stub.
const token = credentials?.apiKey || credentials?.accessToken;
const headers = {
"Content-Type": "application/json",
...(token ? { Authorization: `Bearer ${token}` } : {}),
...(cfg.headers || {}),
// Zen lanes expect the official client session header on every request.
"x-opencode-session": generateSessionId(),
};
const requestBody = { ...body, model };
log?.debug?.("SYSTEMONE", `${provider.toUpperCase()} | ${model}`);
let providerResponse;
try {
providerResponse = await fetch(cfg.baseUrl, {
method: "POST",
headers,
body: JSON.stringify(requestBody),
...(typeof AbortSignal?.timeout === "function"
? { signal: AbortSignal.timeout(FETCH_CONNECT_TIMEOUT_MS) }
: {}),
});
} catch (error) {
const errMsg = formatProviderError(error, provider, model, HTTP_STATUS.BAD_GATEWAY);
log?.debug?.("SYSTEMONE", `Fetch error: ${errMsg}`);
return createErrorResult(HTTP_STATUS.BAD_GATEWAY, errMsg);
}
if (!providerResponse.ok) {
const { statusCode, message } = await parseUpstreamError(providerResponse);
const errMsg = formatProviderError(new Error(message), provider, model, statusCode);
log?.debug?.("SYSTEMONE", `Provider error: ${errMsg}`);
return createErrorResult(statusCode, errMsg);
}
let responseBody;
try {
responseBody = await providerResponse.json();
} catch {
return createErrorResult(HTTP_STATUS.BAD_GATEWAY, `Invalid JSON response from ${provider}`);
}
if (onRequestSuccess) await onRequestSuccess();
const usage = responseBody?.usage;
return {
success: true,
usage: usage
? { prompt_tokens: usage.input_tokens || 0, completion_tokens: usage.output_tokens || 0 }
: null,
response: new Response(JSON.stringify(responseBody), {
headers: {
"Content-Type": "application/json",
"Access-Control-Allow-Origin": "*",
},
}),
};
}

View File

@@ -2,6 +2,7 @@ import { createErrorResult } from "../utils/error.js";
import { HTTP_STATUS } from "../config/runtimeConfig.js";
import { refreshTokenByProvider } from "../services/tokenRefresh.js";
import { PROVIDER_MEDIA } from "../providers/index.js";
import { getVideoAdapter } from "./videoProviders/index.js";
// Upstream fetch deadline for video job submission/polling (the job itself is
// async upstream — this only bounds the HTTP round-trip, not video rendering).
@@ -94,21 +95,49 @@ export async function handleVideoProxyCore({
return createErrorResult(HTTP_STATUS.BAD_REQUEST, `Unknown video action: ${action}`);
}
const method = requestId ? "GET" : "POST";
const url = buildUpstreamUrl(config, action, requestId);
const adapter = getVideoAdapter(provider);
const fetchSignal = combineSignals(signal, timeoutMs);
const doFetch = (token) =>
fetch(url, {
// Default (xAI shape) request plan; adapters override URL/method/headers/body.
const defaultPlan = () => {
const method = requestId ? "GET" : "POST";
return {
method,
headers: buildHeaders({ token, contentType: method === "POST" ? contentType : null, idempotencyKey: method === "POST" ? idempotencyKey : null }),
url: buildUpstreamUrl(config, action, requestId),
headers: buildHeaders({
token: credentials?.accessToken || credentials?.apiKey,
contentType: method === "POST" ? contentType : null,
idempotencyKey: method === "POST" ? idempotencyKey : null,
}),
body: method === "POST" ? rawBody : undefined,
signal: fetchSignal,
});
};
};
// Rebuilt per attempt so the auth retry below picks up the refreshed token.
const doFetch = async () => {
const plan = adapter
? await adapter.buildRequest({
config, action, requestId, rawBody, contentType, idempotencyKey, credentials, log,
token: credentials?.accessToken || credentials?.apiKey,
})
: defaultPlan();
if (plan.error) return { planError: plan.error };
return {
response: await fetch(plan.url, {
method: plan.method,
headers: plan.headers,
body: plan.body,
signal: fetchSignal,
}),
};
};
const method = requestId ? "GET" : "POST";
let upstream;
try {
upstream = await doFetch(credentials?.accessToken || credentials?.apiKey);
const first = await doFetch();
if (first.planError) return createErrorResult(HTTP_STATUS.BAD_REQUEST, `[${provider}] ${first.planError}`);
upstream = first.response;
} catch (error) {
if (error?.name === "AbortError" || error?.name === "TimeoutError") {
return createErrorResult(HTTP_STATUS.REQUEST_TIMEOUT, `[${provider}] video ${method} aborted: ${error.message}`);
@@ -136,7 +165,9 @@ export async function handleVideoProxyCore({
await upstream.body?.cancel?.();
} catch { /* noop */ }
try {
upstream = await doFetch(credentials.accessToken || credentials.apiKey);
const retry = await doFetch();
if (retry.planError) return createErrorResult(HTTP_STATUS.BAD_REQUEST, `[${provider}] ${retry.planError}`);
upstream = retry.response;
} catch (error) {
return createErrorResult(HTTP_STATUS.BAD_GATEWAY, sanitizeSecrets(`[${provider}] video retry after refresh failed: ${error.message}`, credentials));
}
@@ -152,13 +183,25 @@ export async function handleVideoProxyCore({
return createErrorResult(upstream.status, `[${provider}] ${message.slice(0, 2000)}`);
}
// Success: pass the upstream JSON through untouched (request_id / status / video.url).
// Success: pass the upstream JSON through untouched (request_id / status / video.url),
// unless the adapter maps a provider-native shape onto it (Vertex operations).
let outBody = bodyText;
let outType = upstream.headers.get("content-type") || "application/json";
if (adapter?.transformResponse) {
try {
outBody = JSON.stringify(adapter.transformResponse(JSON.parse(bodyText)));
outType = "application/json";
} catch {
// Non-JSON or unexpected shape — fall back to the raw upstream body.
}
}
return {
success: true,
response: new Response(bodyText, {
response: new Response(outBody, {
status: upstream.status,
headers: {
"Content-Type": upstream.headers.get("content-type") || "application/json",
"Content-Type": outType,
"Access-Control-Allow-Origin": "*",
},
}),

View File

@@ -0,0 +1,13 @@
// Video provider adapters.
//
// Default (no adapter) = xAI shape: raw body forwarded to {baseUrl}/{action},
// polled at {baseUrl}/{id}, upstream JSON passed through verbatim.
// A provider only needs an adapter when its wire format differs from that.
import openrouter from "./openrouter.js";
import vertex from "./vertex.js";
const ADAPTERS = { openrouter, vertex };
export function getVideoAdapter(provider) {
return ADAPTERS[provider] || null;
}

View File

@@ -0,0 +1,39 @@
// OpenRouter video jobs — https://openrouter.ai/docs/api/api-reference/videos
//
// Same async shape as xAI (POST → { id, status }, GET → status/unsigned_urls),
// two differences only: creation POSTs to the collection root (no `/generations`
// suffix) and the account headers come from the registry entry.
// Response bodies are passed through verbatim.
// ponytail: generations only — OpenRouter has no edits/extensions endpoint today.
const SUPPORTED_ACTIONS = new Set(["generations"]);
function headers(config, token) {
return {
Accept: "application/json",
...(config.headers || {}),
...(token ? { Authorization: `Bearer ${token}` } : {}),
};
}
export default {
buildRequest({ config, action, requestId, rawBody, contentType, token }) {
const base = config.baseUrl.replace(/\/$/, "");
if (requestId) {
return { method: "GET", url: `${base}/${encodeURIComponent(requestId)}`, headers: headers(config, token) };
}
if (!SUPPORTED_ACTIONS.has(action)) {
return { error: `OpenRouter video supports 'generations' only (got '${action}')` };
}
if (contentType && !contentType.includes("application/json")) {
return { error: "OpenRouter video requires an application/json body" };
}
return {
method: "POST",
url: base,
headers: { ...headers(config, token), "Content-Type": "application/json" },
body: rawBody,
};
},
};

View File

@@ -0,0 +1,159 @@
// Vertex AI (Veo) video jobs.
//
// Vertex does NOT speak the OpenAI-ish /v1/videos shape, so unlike OpenRouter
// this adapter translates both directions:
// create → POST {model}:predictLongRunning { instances[], parameters{} } → { name }
// poll → POST {model}:fetchPredictOperation { operationName } → { done, response }
// Docs: https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/veo-video-generation
//
// The operation name is a resource path (contains "/"), so it is base64url-encoded
// into the job id returned to the client — GET /v1/videos/{id} stays a flat path.
import { parseVertexSaJson, refreshVertexToken } from "../../services/tokenRefresh.js";
const DEFAULT_LOCATION = "us-central1";
const encodeJobId = (name) => Buffer.from(name, "utf8").toString("base64url");
// Operation name shape: projects/{p}/locations/{l}/publishers/{pub}/models/{m}/operations/{op}.
// Anchored and single-segment-per-field so a decoded path can never carry `..` or a
// host-changing prefix into the request URL.
const OPERATION_NAME_RE = /^projects\/[^/]+\/locations\/[^/]+\/publishers\/[^/]+\/models\/[^/]+\/operations\/[^/]+$/;
function modelPathOf(operationName) {
return operationName.slice(0, operationName.indexOf("/operations/"));
}
function decodeJobId(id) {
const raw = String(id ?? "");
// Buffer.from(x, "base64url") silently drops invalid characters instead of
// throwing, so only ids that re-encode byte-for-byte are accepted.
if (!raw || raw.length > 1024 || !/^[A-Za-z0-9_-]+$/.test(raw)) return null;
const decoded = Buffer.from(raw, "base64url").toString("utf8");
if (Buffer.from(decoded, "utf8").toString("base64url") !== raw) return null;
return OPERATION_NAME_RE.test(decoded) ? decoded : null;
}
async function resolveAuth(credentials, log) {
const saJson = parseVertexSaJson(credentials?.apiKey);
const projectId =
saJson?.project_id ||
credentials?.projectId ||
credentials?.providerSpecificData?.projectId;
const location = credentials?.providerSpecificData?.location || DEFAULT_LOCATION;
if (!projectId) {
return { error: "Vertex video requires a project_id — use Service Account JSON or set providerSpecificData.projectId" };
}
let token = credentials?.accessToken;
if (saJson) {
const minted = await refreshVertexToken(saJson, log);
if (!minted?.accessToken) return { error: "Vertex video: failed to mint access token from service account JSON" };
token = minted.accessToken;
}
if (!token) return { error: "Vertex video requires Service Account JSON or an OAuth access token (raw API keys are not supported)" };
return { token, projectId, location };
}
/** OpenAI-ish video body → Vertex predictLongRunning body. */
function toVertexBody(body) {
const instance = { prompt: body.prompt };
// Image-to-video: accept the Vertex-native shape or a bare data URL / base64 string.
const image = body.image ?? body.image_url;
if (image && typeof image === "object") {
instance.image = image;
} else if (typeof image === "string") {
const match = image.match(/^data:([^;]+);base64,(.*)$/s);
instance.image = match
? { bytesBase64Encoded: match[2], mimeType: match[1] }
: { gcsUri: image };
}
if (body.video && typeof body.video === "object") instance.video = body.video;
const parameters = {};
if (body.n != null) parameters.sampleCount = Number(body.n);
if (body.duration != null) parameters.durationSeconds = Number(body.duration);
if (body.aspect_ratio) parameters.aspectRatio = body.aspect_ratio;
if (body.resolution) parameters.resolution = body.resolution;
if (body.seed != null) parameters.seed = body.seed;
if (body.negative_prompt) parameters.negativePrompt = body.negative_prompt;
// Without storageUri Vertex returns inline base64 bytes; a GCS bucket keeps
// the poll response small and is what production callers want.
if (body.storage_uri) parameters.storageUri = body.storage_uri;
if (body.generate_audio != null) parameters.generateAudio = !!body.generate_audio;
return { instances: [instance], ...(Object.keys(parameters).length ? { parameters } : {}) };
}
/** Vertex operation → the async-job shape 9Router clients already poll for. */
function fromVertexOperation(json) {
if (!json?.name) return json;
const id = encodeJobId(json.name);
if (json.error) {
return { id, request_id: id, status: "failed", error: json.error };
}
if (!json.done) {
return { id, request_id: id, status: "pending" };
}
const samples =
json.response?.videos ||
json.response?.generateVideoResponse?.generatedSamples ||
[];
const videos = samples.map((s) => ({
url: s.gcsUri || s.video?.uri || s.uri || null,
b64_json: s.bytesBase64Encoded || s.video?.bytesBase64Encoded || null,
mime_type: s.mimeType || s.video?.mimeType || "video/mp4",
}));
return { id, request_id: id, status: "completed", video: videos[0] || null, videos };
}
export default {
async buildRequest({ config, action, requestId, rawBody, contentType, credentials, log }) {
if (contentType && !contentType.includes("application/json")) {
return { error: "Vertex video requires an application/json body" };
}
const auth = await resolveAuth(credentials, log);
if (auth.error) return { error: auth.error };
const { token, projectId, location } = auth;
const base = (config.baseUrl || "https://aiplatform.googleapis.com").replace(/\/$/, "");
const headers = { Accept: "application/json", "Content-Type": "application/json", Authorization: `Bearer ${token}` };
if (requestId) {
const operationName = decodeJobId(requestId);
if (!operationName) return { error: "Invalid Vertex video job id" };
return {
method: "POST",
url: `${base}/v1/${modelPathOf(operationName)}:fetchPredictOperation`,
headers,
body: JSON.stringify({ operationName }),
};
}
if (action !== "generations") {
// ponytail: Veo extend/edit go through generations with `video`/`image` in the body.
return { error: `Vertex video supports 'generations' only (got '${action}')` };
}
let body;
try {
body = JSON.parse(typeof rawBody === "string" ? rawBody : rawBody.toString("utf8"));
} catch {
return { error: "Invalid JSON body" };
}
if (!body.model) return { error: "Vertex video requires a model (e.g. vertex/veo-3.1-generate-preview)" };
// Plain model id only — a path segment carrying "/" or ".." would rewrite the URL.
if (!/^[A-Za-z0-9._-]+$/.test(body.model)) return { error: "Invalid Vertex video model id" };
if (!body.prompt && !body.image && !body.image_url) return { error: "Vertex video requires a prompt or an image" };
return {
method: "POST",
url: `${base}/v1/projects/${projectId}/locations/${location}/publishers/google/models/${body.model}:predictLongRunning`,
headers,
body: JSON.stringify(toVertexBody(body)),
};
},
transformResponse: fromVertexOperation,
};

View File

@@ -6,6 +6,16 @@
// 3. PATTERN_CAPABILITIES — glob match, ordered specific -> generic
// 4. DEFAULT_CAPABILITIES — safe floor (always returned)
//
// Two extra layers then refine the result, and neither can override the hand
// written tables above (steps 1-2 short-circuit before they are consulted):
// • the synced catalog — modalities keyed by model, limits keyed by provider
// + model, refreshed from models.dev in the background. It reads a file, so
// the server installs it via setCatalogSource(); this module stays free of
// node:fs because the dashboard bundles it into the browser too.
// • visionPatterns.js — name-based vision detection, last resort so a model
// nobody has catalogued yet still accepts images.
// Both only ever turn a capability ON.
//
// ── HOW TO ADD / UPDATE A MODEL ──────────────────────────────────────
// Authoritative data source: https://models.dev/api.json (145 providers, 4000+
// models, MIT). Each model exposes the exact fields we map below:
@@ -23,6 +33,7 @@
// 2.0+, Grok, Perplexity). Verify with: curl -s https://models.dev/api.json
import { matchPattern } from "./pricing.js";
import { looksLikeVisionModel } from "./visionPatterns.js";
/**
* Safe floor — every resolved result is merged over this so consumers
@@ -46,6 +57,7 @@ export const DEFAULT_CAPABILITIES = {
thinkingFormat: null,
thinkingCanDisable: true, // false → model cannot turn thinking off (clamp to min instead of disable)
thinkingRange: null, // { min, max } for budget formats; null = no clamp
thinkingEffortSupported: false, // zai format only: model accepts a reasoning_effort level (GLM-5.2+; older GLM ignores it)
// limits (tokens)
contextWindow: 200000,
maxOutput: 64000,
@@ -71,7 +83,8 @@ export function capabilitiesFromServiceKind(kind) {
* otherwise mis-match. Only declare deltas vs DEFAULT.
*/
export const MODEL_CAPABILITIES = {
// Claude Opus 5, 4.6/4.7/4.8, and Kiro Sonnet 5 have 1M context + adaptive thinking (override generic claude pattern)
// Claude Fable 5.1, Opus 5, 4.6/4.7/4.8, and Kiro Sonnet 5 have 1M context + adaptive thinking (override generic claude pattern)
"claude-fable-5-1": { vision: true, reasoning: true, search: true, thinkingFormat: "claude-adaptive", thinkingCanDisable: false, contextWindow: 1000000, maxOutput: 128000 },
"claude-opus-5": { vision: true, reasoning: true, search: true, thinkingFormat: "claude-adaptive", contextWindow: 1000000, maxOutput: 128000 },
"claude-opus-5-thinking": { vision: true, reasoning: true, search: true, thinkingFormat: "claude-adaptive", contextWindow: 1000000, maxOutput: 128000 },
"claude-opus-5-agentic": { vision: true, reasoning: true, search: true, thinkingFormat: "claude-adaptive", contextWindow: 1000000, maxOutput: 128000 },
@@ -94,8 +107,26 @@ export const MODEL_CAPABILITIES = {
// Gemini image-gen / OpenAI image / xai image variants
"gpt-image-1": { imageOutput: true, tools: false },
// GLM vision variant (text GLM has no vision)
"glm-4.6v": { vision: true, reasoning: true, thinkingFormat: "zai", contextWindow: 128000 },
// GLM vision variants (text GLM has no vision) — 5.3-Flash and 5V-Turbo are
// natively multimodal per z.ai, and 5.3-Flash carries the full 1M window.
"glm-5.3-flash": { vision: true, videoInput: true, pdf: true, reasoning: true, thinkingFormat: "zai", contextWindow: 1000000, maxOutput: 131072 },
"glm-4.6v": { vision: true, videoInput: true, reasoning: true, thinkingFormat: "zai", contextWindow: 128000, maxOutput: 32768 },
"glm-4.5v": { vision: true, videoInput: true, reasoning: true, thinkingFormat: "zai", contextWindow: 64000, maxOutput: 16384 },
// GLM-5.2 has 1M context — pattern *glm-5* only gives 200k, so override here
"glm-5.2": { reasoning: true, thinkingFormat: "zai", thinkingCanDisable: false, contextWindow: 1000000, maxOutput: 131072 },
// DeepSeek's first V4 model with image input; text limits match V4-Flash.
"deepseek-v4-flash-vision-exp": { vision: true, reasoning: true, thinkingFormat: "deepseek", contextWindow: 1000000, maxOutput: 384000 },
// DeepSeek V4.1-Flash is natively multimodal — models.dev lists
// opencode-go/deepseek-v4.1-flash with modalities.input ["text","image"] — and upstream
// the retired v4-flash / vision-exp ids route to it, so the live V4.1 ids carry the
// same image capability as the exp id above. "deepseek-flash" is the GA id on the
// DeepSeek API; it previously fell through to the generic *deepseek* pattern, whose
// 128K/64K limits are kept here. The repeated fields are deliberate: an exact entry
// short-circuits the pattern table, so a vision-only delta would drop them.
"deepseek-v4.1-flash": { vision: true, reasoning: true, thinkingFormat: "deepseek", contextWindow: 1000000, maxOutput: 384000 },
"deepseek-flash": { vision: true, reasoning: true, thinkingFormat: "deepseek", contextWindow: 128000, maxOutput: 64000 },
// Qwen plain coder/text (no vision) — registry "vision-model" / "coder-model" aliases
"vision-model": { vision: true, reasoning: true, thinkingFormat: "qwen", contextWindow: 1000000 },
@@ -108,6 +139,12 @@ export const MODEL_CAPABILITIES = {
"kimi-for-coding-highspeed": { vision: true, videoInput: true, reasoning: true, thinkingFormat: "kimi", thinkingCanDisable: false, contextWindow: 262144, maxOutput: 65536 },
"kimi-k2.7-code": { vision: true, videoInput: true, reasoning: true, thinkingFormat: "kimi", thinkingCanDisable: false, contextWindow: 262144, maxOutput: 65536 },
"kimi-k2.7-code-highspeed": { vision: true, videoInput: true, reasoning: true, thinkingFormat: "kimi", thinkingCanDisable: false, contextWindow: 262144, maxOutput: 65536 },
// OpenCode Free Muse Spark — multimodal (text+image per models.dev meta/muse-spark)
// via OpenAI Responses input_image; reasoning supports up to xhigh.
"muse-spark-1.2-contributor-free": { vision: true, reasoning: true, thinkingFormat: "openai", contextWindow: 1048576, maxOutput: 131072 },
"muse-spark-1.3-contributor-free": { vision: true, reasoning: true, thinkingFormat: "openai", contextWindow: 1048576, maxOutput: 131072 },
// OpenCode Free Union Alpha — multimodal (text+vision), 262K context, 131K max output
"union-alpha": { vision: true, contextWindow: 262144, maxOutput: 131072 },
};
const KIRO_GPT_5_6_CAPABILITIES = { vision: true, reasoning: true, search: true, thinkingFormat: "openai", contextWindow: 272000, maxOutput: 128000 };
@@ -130,7 +167,14 @@ export const PROVIDER_CAPABILITIES = {
"deepseek-ai/deepseek-v4-pro": { reasoning: true, thinkingFormat: "openai", contextWindow: 1000000, maxOutput: 65536 },
"deepseek-ai/deepseek-v4-flash": { reasoning: true, thinkingFormat: "openai", contextWindow: 1000000, maxOutput: 65536 },
},
// glm-5.3-flash on OpenCode Go is served by a backend that rejects the z.ai
// `thinking` object (400: unknown field "thinking") and wants reasoning_effort.
// Overrides the global entry, whose z.ai shape is correct for z.ai itself.
"opencode-go": {
"glm-5.3-flash": { vision: true, videoInput: true, pdf: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 1000000, maxOutput: 131072 },
},
"codex": {
"gpt-6-astra": { vision: true, reasoning: true, search: true, thinkingFormat: "openai", contextWindow: 272000, maxOutput: 128000 },
"gpt-5.6-sol": CODEX_GPT_56_SOL_CAPS,
"gpt-5.6-sol-review": CODEX_GPT_56_SOL_CAPS,
"gpt-5.6-terra": CODEX_GPT_56_DEFAULT_CAPS,
@@ -155,32 +199,66 @@ export const PROVIDER_CAPABILITIES = {
// CodeBuddy.cn — authoritative per-model metadata from the gateway's model
// config (contextWindow=maxInputTokens, maxOutput=maxOutputTokens, vision=
// supportsImages). Every model reasons via OpenAI-style reasoning_effort
// (see registry thinkingFormat). `onlyReasoning` models can't turn thinking
// off → thinkingCanDisable:false (clamped to minimal instead of disabled).
// (see registry thinkingFormat). For thinkingCanDisable use the server's
// reasoning.canDisableThinking flag — see the note in the codebuddy-cn block
// below; it is NOT the inverse of onlyReasoning.
"codebuddy-cn": {
"glm-5.2": { reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 1000000, maxOutput: 48000 },
"glm-5.1": { reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 200000, maxOutput: 48000 },
"glm-5.2": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: true, contextWindow: 1000000, maxOutput: 48000 },
"glm-5.1": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 200000, maxOutput: 48000 },
"glm-5.0": { reasoning: true, thinkingFormat: "openai", contextWindow: 200000, maxOutput: 48000 },
"glm-5.0-turbo": { reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 200000, maxOutput: 48000 },
"glm-5v-turbo": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 200000, maxOutput: 38000 },
// maxOutput 64000 per both the plugin-baked fallback and the live server
// table (the old 38000 had no source and truncated output).
"glm-5v-turbo": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 200000, maxOutput: 64000 },
"glm-4.7": { reasoning: true, thinkingFormat: "openai", contextWindow: 200000, maxOutput: 48000 },
"minimax-m3": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 512000, maxOutput: 48000 },
"minimax-m2.7": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 200000, maxOutput: 48000 },
"minimax-m3": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 512000, maxOutput: 128000 },
"kimi-k2.7": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 256000, maxOutput: 32000 },
"kimi-k2.6": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 256000, maxOutput: 32000 },
"kimi-k2.5": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 164000, maxOutput: 32000 },
"hy3-preview": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 192000, maxOutput: 64000 },
"deepseek-v4-pro": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 1000000, maxOutput: 50000 },
"deepseek-v4-flash": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 1000000, maxOutput: 50000 },
"deepseek-v4-flash": { reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 1000000, maxOutput: 50000 },
"deepseek-v3-2-volc": { reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 96000, maxOutput: 32000 },
// Per-model values mirror the server's product-config payload (the plugin
// fetches it from copilot.tencent.com; the `models[]` entries carry
// maxInputTokens/maxOutputTokens/supportsImages). contextWindow =
// maxInputTokens, maxOutput = maxOutputTokens. Where the server and the
// plugin-baked fallback disagree, the server table wins.
// ⚠️ thinkingCanDisable maps to the server's reasoning.canDisableThinking —
// it is NOT the inverse of onlyReasoning. onlyReasoning means "thinking is
// on by default"; canDisableThinking means "it CAN be turned off". glm-5.3
// and glm-5.3-flash are onlyReasoning:true BUT canDisableThinking:true, so
// their thinking is switchable; the hy* models are forced always-on.
"hy3": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 192000, maxOutput: 64000 },
"hy4-preview": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 1000000, maxOutput: 64000 },
"glm-5.3": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: true, contextWindow: 1000000, maxOutput: 48000 },
"glm-5.3-flash": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: true, contextWindow: 1000000, maxOutput: 32000 },
"kimi-k3-1": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: false, contextWindow: 1000000, maxOutput: 32000 },
"deepseek-v4-pro": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: true, contextWindow: 1000000, maxOutput: 50000 },
// deepseek-v4.1-flash replaces v4-flash (dropped from the server list;
// the old endpoint still answers 200 but the published list is the
// contract). maxOutput 128000 per the server's product-config payload.
"deepseek-v4.1-flash": { vision: true, reasoning: true, thinkingFormat: "openai", thinkingCanDisable: true, contextWindow: 1000000, maxOutput: 128000 },
},
// Poolside Laguna — OpenAI-compatible, all reasoning-capable (32K max output).
"poolside": {
"laguna-s-2.1": { reasoning: true, thinkingFormat: "openai", contextWindow: 1000000, maxOutput: 32000 },
"laguna-xs-2.1": { reasoning: true, thinkingFormat: "openai", contextWindow: 200000, maxOutput: 32000 },
},
// Ollama Cloud — the generic *deepseek-v4* pattern misses the vision badge
// the library page publishes for this model (text+image in, 1M context).
// ponytail: thinkingFormat stays "deepseek" to preserve today's body shape;
// Ollama's native toggle is the top-level `think` field (bool or
// low/medium/high/max), which no format in thinkingUnified.js emits yet —
// openai-to-ollama.js drops it. Wire a "think" format when thinking on
// Ollama Cloud is actually needed.
"ollama": {
"deepseek-v4.1-flash:cloud": { vision: true, reasoning: true, thinkingFormat: "deepseek", contextWindow: 1000000, maxOutput: 384000 },
},
};
// Qoder CN serves the identical model catalog from the CN gateway, so it shares
// the intl Qoder capability table verbatim (vision/reasoning/contextWindow).
PROVIDER_CAPABILITIES["qoder-cn"] = PROVIDER_CAPABILITIES["qoder"];
/**
* Pattern fallback — glob (* = wildcard), matched case-insensitively and
* anchored (^...$) so a pattern must match the full model id. ORDER MATTERS:
@@ -205,6 +283,7 @@ export const PATTERN_CAPABILITIES = [
// ── Gemini (all 2.0+ multimodal + google_search grounding, 1M ctx) ─
{ pattern: "*gemini*image*", caps: { vision: true, imageOutput: true, contextWindow: 1048576 } },
{ pattern: "*gemini-3.8*", caps: { vision: true, audioInput: true, videoInput: true, reasoning: true, search: true, thinkingFormat: "gemini-level", thinkingCanDisable: false, contextWindow: 1048576, maxOutput: 65536 } },
{ pattern: "*gemini-3.7*", caps: { vision: true, audioInput: true, videoInput: true, reasoning: true, search: true, thinkingFormat: "gemini-level", thinkingCanDisable: false, contextWindow: 1048576, maxOutput: 65536 } },
{ pattern: "*gemini-3*pro*", caps: { vision: true, audioInput: true, videoInput: true, reasoning: true, search: true, thinkingFormat: "gemini-level", thinkingCanDisable: false, contextWindow: 1048576, maxOutput: 65535 } },
{ pattern: "*gemini-3*", caps: { vision: true, audioInput: true, videoInput: true, reasoning: true, search: true, thinkingFormat: "gemini-level", thinkingCanDisable: false, contextWindow: 1048576, maxOutput: 65536 } },
@@ -214,6 +293,9 @@ export const PATTERN_CAPABILITIES = [
{ pattern: "*gemma*", caps: { vision: true, contextWindow: 128000 } },
{ pattern: "*nanobanana*", caps: { vision: true, imageOutput: true } },
// ── OpenAI GPT-6.x (vision + thinking + web search) ──────────────
{ pattern: "*gpt-6*", caps: { vision: true, reasoning: true, search: true, thinkingFormat: "openai", contextWindow: 272000, maxOutput: 128000 } },
// ── OpenAI GPT-5.x (vision + thinking + web search) ──────────────
{ pattern: "*gpt-5*image*", caps: { imageOutput: true } },
{ pattern: "*gpt-5*codex*", caps: { reasoning: true, search: true, thinkingFormat: "openai", contextWindow: 400000, maxOutput: 128000 } },
@@ -234,6 +316,8 @@ export const PATTERN_CAPABILITIES = [
// ── Grok (vision + Live Search) ──────────────────────────────────
{ pattern: "*grok*image*", caps: { imageOutput: true } },
{ pattern: "*grok-code*", caps: { reasoning: true, thinkingFormat: "openai", contextWindow: 256000 } },
// Grok 4.6: 500k context, no text output limit (docs.x.ai/developers/grok-4-6)
{ pattern: "*grok-4.6*", caps: { vision: true, reasoning: true, search: true, thinkingFormat: "openai", contextWindow: 500000, maxOutput: 500000 } },
// Grok 4.5 (Grok CLI / Grok Build): 500k context per cli-chat-proxy /v1/models
{ pattern: "*grok-4.5*", caps: { vision: true, reasoning: true, search: true, thinkingFormat: "openai", contextWindow: 500000, maxOutput: 64000 } },
{ pattern: "*grok-4*", caps: { vision: true, reasoning: true, search: true, thinkingFormat: "openai", contextWindow: 256000 } },
@@ -244,7 +328,7 @@ export const PATTERN_CAPABILITIES = [
{ pattern: "*qwen*vl*", caps: { vision: true, reasoning: true, thinkingFormat: "qwen", contextWindow: 262144 } },
{ pattern: "*qwen*omni*", caps: { vision: true, audioInput: true, videoInput: true, reasoning: true, thinkingFormat: "qwen", contextWindow: 262144, maxOutput: 65536 } },
{ pattern: "*qwen*coder*", caps: { reasoning: true, thinkingFormat: "qwen", contextWindow: 1000000 } },
{ pattern: "*qwen*max*", caps: { reasoning: true, thinkingFormat: "qwen", contextWindow: 1000000, maxOutput: 65536 } },
{ pattern: "*qwen*max*", caps: { vision: true, reasoning: true, thinkingFormat: "qwen", contextWindow: 1000000, maxOutput: 65536 } },
{ pattern: "*qwen3.5*", caps: { vision: true, videoInput: true, reasoning: true, thinkingFormat: "qwen", contextWindow: 1000000, maxOutput: 65536 } },
{ pattern: "*qwen3.6*", caps: { vision: true, videoInput: true, reasoning: true, thinkingFormat: "qwen", contextWindow: 1000000, maxOutput: 65536 } },
{ pattern: "*qwen3.7*", caps: { vision: true, videoInput: true, reasoning: true, thinkingFormat: "qwen", contextWindow: 1000000, maxOutput: 65536 } },
@@ -261,13 +345,21 @@ export const PATTERN_CAPABILITIES = [
{ pattern: "*kimi*", caps: { reasoning: true, thinkingFormat: "kimi", contextWindow: 262144 } },
// ── GLM / Z.ai (thinking.enabled; disable via enable_thinking:false) ─
// reasoning_effort is only read by z.ai from GLM-5.2 onward (docs.z.ai/guides/capabilities/thinking) —
// older GLM (4.x, 5.0, 5.1, 5-turbo, 5v-turbo) ignore it, so gate it per exact version, not the "*glm-5*" catch-all.
{ pattern: "*glm-5.3*", caps: { reasoning: true, thinkingFormat: "zai", thinkingEffortSupported: true, contextWindow: 200000, maxOutput: 128000 } },
{ pattern: "*glm-5.2*", caps: { reasoning: true, thinkingFormat: "zai", thinkingEffortSupported: true, contextWindow: 200000, maxOutput: 128000 } },
{ pattern: "*glm-5*", caps: { reasoning: true, thinkingFormat: "zai", contextWindow: 200000, maxOutput: 128000 } },
{ pattern: "*glm-4.7*", caps: { reasoning: true, thinkingFormat: "zai", contextWindow: 200000, maxOutput: 128000 } },
{ pattern: "*glm-4*", caps: { reasoning: true, thinkingFormat: "zai", contextWindow: 200000 } },
{ pattern: "*glm*", caps: { reasoning: true, thinkingFormat: "zai", contextWindow: 200000 } },
// ── DeepSeek (thinking.enabled + reasoning_effort; r1 = thinking-only) ─
{ pattern: "*deepseek-v4*", caps: { reasoning: true, thinkingFormat: "deepseek", contextWindow: 1000000, maxOutput: 384000 } },
// v4.1+ has real image input (probed live on Alibaba MaaS: correct color
// read from a PNG). v4-pro / v4-flash-0731 accept image blocks but ignore
// them (answered "Unknown"), so vision stays scoped to v4.* dotted releases.
{ pattern: "*deepseek-v4.*", caps: { vision: true, reasoning: true, thinkingFormat: "deepseek", thinkingEffortSupported: true, contextWindow: 1000000, maxOutput: 128000 } },
{ pattern: "*deepseek-v4*", caps: { reasoning: true, thinkingFormat: "deepseek", thinkingEffortSupported: true, contextWindow: 1000000, maxOutput: 384000 } },
{ pattern: "*reasoner*", caps: { reasoning: true, thinkingFormat: "deepseek", thinkingCanDisable: false, contextWindow: 128000 } },
{ pattern: "*deepseek-r*", caps: { reasoning: true, thinkingFormat: "deepseek", thinkingCanDisable: false, contextWindow: 128000 } },
{ pattern: "*deepseek-chat*", caps: { contextWindow: 128000 } },
@@ -275,14 +367,16 @@ export const PATTERN_CAPABILITIES = [
// ── MiniMax (M3 = adaptive; M2.x cannot disable) ─────────────────
{ pattern: "*minimax*image*", caps: { imageOutput: true } },
{ pattern: "*minimax-m3*", caps: { vision: true, reasoning: true, thinkingFormat: "minimax", contextWindow: 1048576, maxOutput: 512000 } },
{ pattern: "*minimax-m2.7*", caps: { reasoning: true, thinkingFormat: "minimax", thinkingCanDisable: false, contextWindow: 204800, maxOutput: 131072 } },
{ pattern: "*minimax-m3*", caps: { vision: true, reasoning: true, thinkingFormat: "minimax", contextWindow: 1000000, maxOutput: 131072 } },
{ pattern: "*minimax-m2.7*", caps: { vision: true, reasoning: true, thinkingFormat: "minimax", thinkingCanDisable: false, contextWindow: 204800, maxOutput: 131072 } },
{ pattern: "*minimax-m2.5*", caps: { vision: true, reasoning: true, thinkingFormat: "minimax", thinkingCanDisable: false, contextWindow: 204800, maxOutput: 131072 } },
{ pattern: "*minimax*", caps: { reasoning: true, thinkingFormat: "minimax", thinkingCanDisable: false, contextWindow: 200000, maxOutput: 131072 } },
// ── Xiaomi MiMo (vision, 1M / 262K ctx) ──────────────────────────
{ pattern: "*mimo*v2.5*", caps: { vision: true, audioInput: true, videoInput: true, contextWindow: 1048576, maxOutput: 131072 } },
{ pattern: "*mimo*omni*", caps: { vision: true, audioInput: true, contextWindow: 262144, maxOutput: 131072 } },
{ pattern: "*mimo*", caps: { vision: true, contextWindow: 262144, maxOutput: 131072 } },
// ── Xiaomi MiMo (vision + <think>-tag reasoning, always-on, can't disable) ──
{ pattern: "*mimo*v2.6*", caps: { vision: true, audioInput: true, videoInput: true, reasoning: true, thinkingFormat: "deepseek", thinkingCanDisable: false, contextWindow: 1048576, maxOutput: 131072 } },
{ pattern: "*mimo*v2.5*", caps: { vision: true, audioInput: true, videoInput: true, reasoning: true, thinkingFormat: "deepseek", thinkingCanDisable: false, contextWindow: 1048576, maxOutput: 131072 } },
{ pattern: "*mimo*omni*", caps: { vision: true, audioInput: true, reasoning: true, thinkingFormat: "deepseek", thinkingCanDisable: false, contextWindow: 262144, maxOutput: 131072 } },
{ pattern: "*mimo*", caps: { vision: true, reasoning: true, thinkingFormat: "deepseek", thinkingCanDisable: false, contextWindow: 262144, maxOutput: 131072 } },
// ── Llama (4 = vision/1M; 3.x = text-only/128K) ──────────────────
{ pattern: "*llama-4*", caps: { vision: true, contextWindow: 1000000 } },
@@ -309,6 +403,9 @@ export const PATTERN_CAPABILITIES = [
{ pattern: "*laguna-s-2.1*", caps: { reasoning: true, thinkingFormat: "openai", contextWindow: 1000000, maxOutput: 32000 } },
{ pattern: "*laguna*", caps: { reasoning: true, thinkingFormat: "openai", contextWindow: 200000, maxOutput: 32000 } },
// ── OpenCode Free Muse Spark (multimodal text+image; OpenAI Responses reasoning supports up to xhigh) ─
{ pattern: "*muse*spark*", caps: { vision: true, reasoning: true, thinkingFormat: "openai", contextWindow: 1048576, maxOutput: 131072 } },
// ── Others ───────────────────────────────────────────────────────
{ pattern: "*hunyuan*", caps: { reasoning: true, thinkingFormat: "hunyuan", contextWindow: 262144, maxOutput: 262144 } },
{ pattern: "hy3*", caps: { reasoning: true, thinkingFormat: "hunyuan", contextWindow: 262144, maxOutput: 262144 } },
@@ -317,6 +414,65 @@ export const PATTERN_CAPABILITIES = [
{ pattern: "*ling-*", caps: { reasoning: true, contextWindow: 128000 } },
];
// OpenRouter-style gateways validate modalities upstream — a text-only model
// sent an image gets a clear upstream error instead of silent corruption. So for
// unknown models on these providers, trust vision instead of stripping images.
const TRUST_UPSTREAM_VISION = new Set(["openrouter"]);
/**
* Aggregate capabilities for a combo from its constituent model IDs.
* Each entry in comboModels is a fully-qualified "provider/model" string.
*
* Union: vision, pdf, audioInput, videoInput, imageOutput, audioOutput, search
* Intersection: tools
* Primary: reasoning fields from the first (primary) model
* Conservative: contextWindow = min; maxOutput = max
*
* @param {string[]} comboModels
* @param {Object|null} [comboLookup] optional map of combo name → models array for nested resolution
* @param {Function|null} [resolveCaps] optional (fullId) → caps override. The synced model
* catalog is server-only (it reads a file), so a browser-side resolution cannot see the
* limits it supplies and silently falls back to the generic patterns below. Callers that
* have the server's answer (/api/models, via useModelCaps) pass it here; it is merged over
* the local tables, so fields it does not carry (tools, pdf, audio/video, thinking*) survive.
* @param {number} [_depth] internal recursion depth guard
* @returns {object|null} full capabilities object, or null for empty input
*/
export function aggregateComboCapabilities(comboModels, comboLookup = null, resolveCaps = null, _depth = 0) {
if (!comboModels?.length || _depth > 6) return null;
const allCaps = comboModels.map((fullId) => {
// Nested combo: bare name (no slash) that exists in the lookup — recurse
if (!fullId.includes("/") && comboLookup?.[fullId]) {
return aggregateComboCapabilities(comboLookup[fullId], comboLookup, resolveCaps, _depth + 1)
?? resolveCaps?.(fullId)
?? getCapabilitiesForModel(null, fullId);
}
const slash = fullId.indexOf("/");
const provider = slash === -1 ? null : fullId.slice(0, slash);
const model = slash === -1 ? fullId : fullId.slice(slash + 1);
const local = getCapabilitiesForModel(provider, model);
const override = resolveCaps?.(fullId);
return override ? { ...local, ...override } : local;
});
const first = allCaps[0];
return {
vision: allCaps.some((c) => c.vision),
pdf: allCaps.some((c) => c.pdf),
audioInput: allCaps.some((c) => c.audioInput),
videoInput: allCaps.some((c) => c.videoInput),
imageOutput: allCaps.some((c) => c.imageOutput),
audioOutput: allCaps.some((c) => c.audioOutput),
search: allCaps.some((c) => c.search),
tools: allCaps.every((c) => c.tools),
reasoning: first.reasoning,
thinkingFormat: first.thinkingFormat,
thinkingCanDisable: first.thinkingCanDisable,
thinkingRange: first.thinkingRange,
contextWindow: Math.min(...allCaps.map((c) => c.contextWindow)),
maxOutput: Math.max(...allCaps.map((c) => c.maxOutput)),
};
}
/**
* Resolve capabilities for a model using the 4-step fallback chain,
* merged over DEFAULT_CAPABILITIES so the result is always complete.
@@ -325,30 +481,187 @@ export const PATTERN_CAPABILITIES = [
* @param {string} model
* @returns {object} full capabilities object
*/
const MODALITY_KEYS = ["vision", "pdf", "audioInput", "videoInput"];
// ── Server-injected readers ──────────────────────────────────────────
// Next.js compiles instrumentation and each API route into SEPARATE server
// bundles, so a plain module-local `let` would give every bundle its own copy of
// this file and a source installed at boot would be invisible to the request
// handlers (silently: the setters still "succeed"). The slots therefore live on
// globalThis, which IS shared across server bundles in the same process.
// Same reason the browser bundle is safe: it never calls a setter, so the slots
// stay empty and every consumer below short-circuits. Every read goes through
// globalThis: caching it locally would keep a reader alive in other copies after
// setCatalogSource(null).
let catalogSource = null;
const SOURCE_SLOTS = (globalThis.__9R_CAPABILITY_SOURCES ||= {
catalog: null, // { getModalities, getLimits } — synced models.dev catalog
userCaps: null, // (provider, model) => asserted caps — dashboard toggles
});
/**
* Install the synced catalog reader (server only).
* @param {{ getModalities: (provider: string, model: string) => object|null,
* getLimits: (provider: string, model: string) => object|null } | null} source
*/
export function setCatalogSource(source) {
catalogSource = source || null;
if (SOURCE_SLOTS) SOURCE_SLOTS.catalog = source || null;
if (typeof globalThis !== "undefined") globalThis.__9rCatalogSource = source || null;
}
function getCatalogSource() {
if (typeof globalThis === "undefined") return catalogSource;
return SOURCE_SLOTS?.catalog || globalThis.__9rCatalogSource || null;
}
// Capabilities the user asserted per provider+model (dashboard "Add/Edit Model"
// toggles), installed by the server from the custom-model store. Unlike the
// catalog and name heuristics this is authoritative and two-directional: it can
// turn a capability OFF as well as on.
const USER_CAPS_KEYS = ["vision", "pdf", "audioInput", "videoInput", "imageOutput", "audioOutput", "search", "tools", "reasoning", "thinkingFormat", "contextWindow", "maxOutput"];
// (slot lives in SOURCE_SLOTS above — see the cross-bundle note)
/**
* Install the user-asserted caps reader (server only).
* @param {(provider: string|null, model: string) => object|null} source sync lookup
*/
export function setUserCapsSource(source) {
SOURCE_SLOTS.userCaps = typeof source === "function" ? source : null;
}
// Last step of every resolution path: the user's own assertion wins over any
// heuristic, including the tables above (a hand-typed model id can collide with
// a pattern entry that describes a different product).
function applyUserCaps(result, provider, model) {
const userCapsSource = SOURCE_SLOTS.userCaps;
if (!userCapsSource) return result;
let asserted = null;
try {
asserted = userCapsSource(provider, model);
} catch {
return result;
}
if (!asserted || typeof asserted !== "object") return result;
for (const key of USER_CAPS_KEYS) {
if (asserted[key] === undefined) continue;
result[key] = asserted[key];
}
return result;
}
// Apply the synced catalog + name heuristic on top of a table-resolved result.
// Strictly additive: a capability already true stays true, and a false one only
// flips when an outside source positively declares support.
function refine(base, provider, model) {
const result = { ...DEFAULT_CAPABILITIES, ...base };
const source = getCatalogSource();
if (source) {
const modalities = source.getModalities(provider, model);
if (modalities) {
for (const key of MODALITY_KEYS) {
if (modalities[key] === true) result[key] = true;
}
}
const limits = source.getLimits(provider, model);
if (limits) {
if (limits.contextWindow > 0) result.contextWindow = limits.contextWindow;
if (limits.maxOutput > 0) result.maxOutput = limits.maxOutput;
}
}
if (!result.vision && looksLikeVisionModel(model)) result.vision = true;
return result;
}
// Mirrors Command Code CLI `isKnownTextOnlyModel` (no image input). New models
// default to vision; only this denylist stays text-only.
const COMMANDCODE_TEXT_ONLY = new Set([
"deepseek/deepseek-v4-pro",
"deepseek/deepseek-v4-flash",
"deepseek/deepseek-v4-flash-fast",
"zai-org/glm-5.3",
"zai-org/glm-5.2",
"zai-org/glm-5.2-fast",
"zai-org/glm-5.1",
"zai-org/glm-5",
"minimaxai/minimax-m2.7",
"minimax/minimax-m2.7-free",
"minimaxai/minimax-m2.5",
"xiaomi/mimo-v2.5-pro",
"qwen/qwen3.6-max-preview",
"qwen/qwen3.7-max",
"meituan/longcat-2.0:free",
"stepfun/step-3.5-flash",
"tencent/hy4-preview",
"tencent/hy3",
"tencent/hy3-paid",
"nvidia/nemotron-3-ultra-550b-a55b",
"poolside/laguna-s-2.1-free",
"inclusionai/ling-3.0-flash-free",
"inclusionai/ling-3.0-flash-sante:free",
]);
function isCommandCodeTextOnly(model) {
const key = String(model || "").toLowerCase();
if (COMMANDCODE_TEXT_ONLY.has(key)) return true;
for (const id of COMMANDCODE_TEXT_ONLY) {
const base = id.includes("/") ? id.slice(id.lastIndexOf("/") + 1) : id;
if (key === base || key.endsWith("/" + base)) return true;
}
return false;
}
export function getCapabilitiesForModel(provider, model) {
if (!model) return { ...DEFAULT_CAPABILITIES };
// Canonical exact lookup strips vendor prefix: "anthropic/claude-opus-4.7" -> "claude-opus-4.7".
const baseModel = model.includes("/") ? model.split("/").pop() : model;
// 1. Provider-specific override
if (provider) {
const providerCaps = PROVIDER_CAPABILITIES[provider];
if (providerCaps?.[model]) return { ...DEFAULT_CAPABILITIES, ...providerCaps[model] };
if (providerCaps?.[baseModel]) return { ...DEFAULT_CAPABILITIES, ...providerCaps[baseModel] };
}
// 2. Canonical exact
if (MODEL_CAPABILITIES[baseModel]) return { ...DEFAULT_CAPABILITIES, ...MODEL_CAPABILITIES[baseModel] };
if (MODEL_CAPABILITIES[model]) return { ...DEFAULT_CAPABILITIES, ...MODEL_CAPABILITIES[model] };
// 3. Pattern match (first match wins)
for (const { pattern, caps } of PATTERN_CAPABILITIES) {
if (matchPattern(pattern, baseModel) || matchPattern(pattern, model)) {
return { ...DEFAULT_CAPABILITIES, ...caps };
const resolve = () => {
// CommandCode wire is /alpha/generate for every model. Family patterns
// (deepseek-v4 → thinkingFormat:deepseek, vision:false) must not win here.
if (provider === "commandcode" || provider === "cmc") {
const providerCaps = PROVIDER_CAPABILITIES.commandcode;
if (providerCaps?.[model]) return { ...DEFAULT_CAPABILITIES, ...providerCaps[model] };
if (providerCaps?.[baseModel]) return { ...DEFAULT_CAPABILITIES, ...providerCaps[baseModel] };
return {
...DEFAULT_CAPABILITIES,
reasoning: true,
thinkingFormat: "commandcode",
thinkingEffortSupported: true,
vision: !isCommandCodeTextOnly(model),
contextWindow: 1000000,
maxOutput: 384000,
};
}
}
// 4. Floor
return { ...DEFAULT_CAPABILITIES };
// 1. Provider-specific override
if (provider) {
const providerCaps = PROVIDER_CAPABILITIES[provider];
if (providerCaps?.[model]) return { ...DEFAULT_CAPABILITIES, ...providerCaps[model] };
if (providerCaps?.[baseModel]) return { ...DEFAULT_CAPABILITIES, ...providerCaps[baseModel] };
}
// 2. Canonical exact
if (MODEL_CAPABILITIES[baseModel]) return { ...DEFAULT_CAPABILITIES, ...MODEL_CAPABILITIES[baseModel] };
if (MODEL_CAPABILITIES[model]) return { ...DEFAULT_CAPABILITIES, ...MODEL_CAPABILITIES[model] };
// 3. Pattern match (first match wins), refined by catalog + name heuristic
for (const { pattern, caps } of PATTERN_CAPABILITIES) {
if (matchPattern(pattern, baseModel) || matchPattern(pattern, model)) {
return refine(caps, provider, model);
}
}
// 4. Floor (upstream-validated gateways keep vision on for unknown models)
if (provider && TRUST_UPSTREAM_VISION.has(provider)) {
return { ...refine(null, provider, model), vision: true };
}
return refine(null, provider, model);
};
return applyUserCaps(resolve(), provider, model);
}

View File

@@ -0,0 +1,82 @@
// Read side of the model catalog synced from models.dev.
//
// The file is the source of truth; the only thing held in memory is a parsed
// copy dropped as soon as the file's mtime changes. getCapabilitiesForModel is
// synchronous and runs per request, so the hot path is one stat (~1us) and the
// parse (~0.1ms on a ~18KB file) only reruns after a sync.
import fs from "node:fs";
import path from "node:path";
import { DATA_DIR } from "@/lib/dataDir.js";
export const CATALOG_FILE = path.join(DATA_DIR, "model-catalog.json");
// Trimmed upstream catalog, read by the add-models skill (not by the router).
export const CATALOG_RAW_FILE = path.join(DATA_DIR, "model-catalog-raw.json");
// Schema of the file this module reads. The writer stamps it; a file carrying an
// older value predates provider-scoped modality keys, and its flat keys are not
// looked up here, so the sync rebuilds it instead of asking upstream for a 304.
export const CATALOG_VERSION = 2;
const EMPTY = { models: {}, providers: {} };
let cache = EMPTY;
let cachedMtime = -1;
// "zai-org/GLM-4.6V:free" -> "glm-4.6v"
function baseId(model) {
if (!model) return "";
const withoutVendor = model.includes("/") ? model.split("/").pop() : model;
return withoutVendor.toLowerCase().split(":")[0];
}
function load() {
let mtime;
try {
mtime = fs.statSync(CATALOG_FILE).mtimeMs;
} catch {
cache = EMPTY;
cachedMtime = -1;
return cache;
}
if (mtime === cachedMtime) return cache;
cachedMtime = mtime;
try {
const parsed = JSON.parse(fs.readFileSync(CATALOG_FILE, "utf8"));
cache = { models: parsed?.models || {}, providers: parsed?.providers || {} };
} catch {
cache = EMPTY;
}
return cache;
}
// Modalities are recorded per gateway upstream, and gateways disagree about the
// same weights — some do not proxy images at all — so the key is provider +
// model, in the local provider id space, exactly like the limits below. Keying
// by model id alone made short ids collide across vendors: "auto", "free" and
// "efficient" are router modes in one catalog and model names in another, and a
// request to the router mode inherited a stranger's vision.
export function getCatalogModalities(provider, model) {
if (!provider) return null;
return load().models[`${provider}:${baseId(model)}`] || null;
}
// Context and output limits are a property of the gateway too: each one
// truncates differently, so these stay keyed by provider + model.
export function getCatalogLimits(provider, model) {
const byProvider = provider && load().providers[provider];
if (!byProvider) return null;
return byProvider[model] || byProvider[baseId(model)] || null;
}
// Force a re-read on the next lookup (called right after a sync writes the file).
export function invalidateCatalog() {
cachedMtime = -1;
}
// Hand the reader to capabilities.js. That module is bundled into the browser
// too, so it cannot import this file directly — the server pushes it in.
export async function installCatalogSource() {
const { setCatalogSource } = await import("./capabilities.js");
setCatalogSource({ getModalities: getCatalogModalities, getLimits: getCatalogLimits });
}

View File

@@ -23,7 +23,7 @@ function buildTransport(transport, oauth) {
const MEDIA_KEYS = new Set([
"serviceKinds", "ttsConfig", "sttConfig", "embeddingConfig",
"imageConfig", "imageToTextConfig", "videoConfig", "musicConfig",
"searchViaChat", "searchConfig", "fetchConfig",
"searchViaChat", "searchConfig", "fetchConfig", "systemoneConfig",
"modelsFetcher", "mediaPriority", "hiddenKinds",
]);

View File

@@ -1,3 +1,5 @@
import { FORMATS } from "../../translator/formats.js";
// Codex auto-generates a "-review" variant for each llm model (review quota family)
export const CODEX_REVIEW_SUFFIX = "-review";
@@ -18,3 +20,27 @@ export function withCodexReviewModels(models) {
];
});
}
export function isMuseSparkModel(modelId) {
if (!modelId || typeof modelId !== "string") return false;
const clean = modelId.replace(/\([^()]+\)\s*$/, "").trim();
const base = clean.includes("/") ? clean.split("/").pop() : clean;
return /^muse[-_]?spark(?:$|[-_:.\s])/i.test(base);
}
// Endpoint families for OpenCode models outside the curated registry (modelsFetcher /
// passthrough ids) — regex keeps auto-fetched models on the right endpoint:
// /responses (gpt/grok/muse-spark), /messages (minimax/qwen), /chat/completions (rest).
// Curated registry entries always win; this is the unknown-id fallback only.
const OPENCODE_FAMILIES = [
{ match: /^(grok|gpt|muse[-_]?spark)/i, supportedFormats: [FORMATS.OPENAI_RESPONSES], targetFormat: FORMATS.OPENAI_RESPONSES },
{ match: /^deepseek-v4-(pro|flash)/, supportedFormats: [FORMATS.OPENAI, FORMATS.CLAUDE, FORMATS.OPENAI_RESPONSES] },
{ match: /^(minimax|qwen)/, supportedFormats: [FORMATS.OPENAI, FORMATS.CLAUDE] },
{ match: /^claude-/i, supportedFormats: [FORMATS.CLAUDE] },
];
export function opencodeFamilyFormats(modelId) {
if (!modelId || typeof modelId !== "string") return null;
const base = modelId.replace(/\([^()]+\)\s*$/, "").trim();
return OPENCODE_FAMILIES.find((f) => f.match.test(base)) || null;
}

View File

@@ -2,8 +2,28 @@
//
// Fallback order (first match wins):
// 1. PROVIDER_PRICING[provider][model] — provider-specific override
// 2. MODEL_PRICING[model] — canonical model price (provider-agnostic)
// 3. PATTERN_PRICING — glob pattern match (e.g. "codex-*")
// 2. FREE_MODEL_NAMESPACES — upstream bills these at $0
// 3. MODEL_PRICING[model] — canonical model price (provider-agnostic)
// 4. PATTERN_PRICING — glob pattern match (e.g. "codex-*")
/**
* Namespaces upstream meters at $0. A free model must never inherit a paid
* rate: the vendor-prefix strip in getPricingForModel() would turn
* "cline-free/deepseek-v4.1-flash" into "deepseek-v4.1-flash" and match
* MODEL_PRICING, so the namespace is checked before both fallbacks.
*/
export const FREE_MODEL_NAMESPACES = ["cline-free/"];
export const ZERO_PRICING = {
input: 0, output: 0, cached: 0, reasoning: 0, cache_creation: 0,
};
/** True when the model id sits in a namespace upstream bills at $0. */
export function isFreeModel(model) {
if (!model) return false;
const lower = String(model).toLowerCase();
return FREE_MODEL_NAMESPACES.some((ns) => lower.startsWith(ns));
}
/**
* Canonical model pricing — provider-agnostic.
@@ -53,10 +73,15 @@ export const MODEL_PRICING = {
"gpt-5.6-luna": { input: 1.00, output: 6.00, cached: 0.10, reasoning: 6.00, cache_creation: 1.00 },
"gpt-5.6-terra": { input: 2.50, output: 15.00, cached: 0.25, reasoning: 15.00, cache_creation: 2.50 },
"gpt-5.6-sol": { input: 5.00, output: 30.00, cached: 0.50, reasoning: 30.00, cache_creation: 5.00 },
"gpt-6-astra": { input: 5.00, output: 30.00, cached: 0.50, reasoning: 30.00, cache_creation: 5.00 },
"o1": { input: 15.00, output: 60.00, cached: 7.50, reasoning: 90.00, cache_creation: 15.00 },
"o1-mini": { input: 3.00, output: 12.00, cached: 1.50, reasoning: 18.00, cache_creation: 3.00 },
// === Gemini ===
"gemini-3.8-flash": { input: 1.50, output: 7.50, cached: 0.15, reasoning: 11.25, cache_creation: 1.875 },
"gemini-3.8-flash-high": { input: 1.50, output: 7.50, cached: 0.15, reasoning: 11.25, cache_creation: 1.875 },
"gemini-3.8-flash-medium": { input: 1.50, output: 7.50, cached: 0.15, reasoning: 11.25, cache_creation: 1.875 },
"gemini-3.8-flash-low": { input: 1.50, output: 7.50, cached: 0.15, reasoning: 11.25, cache_creation: 1.875 },
"gemini-3.7-flash": { input: 1.50, output: 7.50, cached: 0.15, reasoning: 11.25, cache_creation: 1.875 },
"gemini-3.7-flash-high": { input: 1.50, output: 7.50, cached: 0.15, reasoning: 11.25, cache_creation: 1.875 },
"gemini-3.7-flash-medium": { input: 1.50, output: 7.50, cached: 0.15, reasoning: 11.25, cache_creation: 1.875 },
@@ -106,6 +131,8 @@ export const MODEL_PRICING = {
"deepseek-v3.2-chat": { input: 0.14, output: 0.28, cached: 0.0028, reasoning: 0.28, cache_creation: 0.14 },
"deepseek-v3.2-reasoner": { input: 0.14, output: 0.28, cached: 0.0028, reasoning: 0.28, cache_creation: 0.14 },
"deepseek-v4-flash": { input: 0.14, output: 0.28, cached: 0.0028, reasoning: 0.28, cache_creation: 0.14 },
"deepseek-v4.1-flash": { input: 0.14, output: 0.28, cached: 0.0028, reasoning: 0.28, cache_creation: 0.14 },
"deepseek-flash": { input: 0.14, output: 0.28, cached: 0.0028, reasoning: 0.28, cache_creation: 0.14 },
"deepseek-v4-pro": { input: 0.435, output: 0.87, cached: 0.003625, reasoning: 0.87, cache_creation: 0.435 },
// === GLM ===
@@ -260,6 +287,7 @@ export const PROVIDER_PRICING = {
"z-ai/glm-5-turbo": { input: 1.2, output: 4.0, cached: 0.24, reasoning: 4.0 },
"z-ai/glm-5.1": { input: 1.05, output: 3.5, cached: 0.525, reasoning: 3.5 },
"z-ai/glm-5.2": { input: 1.4, output: 4.4, cached: 0.26, reasoning: 4.4 },
"z-ai/glm-5.3-free": { input: 0, output: 0, cached: 0, reasoning: 0 },
},
};
@@ -353,10 +381,11 @@ export function matchPattern(pattern, model) {
}
/**
* Resolve pricing for a model using the 3-step fallback chain:
* Resolve pricing for a model using the 4-step fallback chain:
* 1. PROVIDER_PRICING[provider][model]
* 2. MODEL_PRICING[model]
* 3. PATTERN_PRICING (glob match)
* 2. free namespace (upstream bills $0)
* 3. MODEL_PRICING[model]
* 4. PATTERN_PRICING (glob match)
*
* @param {string} provider
* @param {string} model
@@ -370,12 +399,15 @@ export function getPricingForModel(provider, model) {
return PROVIDER_PRICING[provider][model];
}
// 2. Canonical model pricing (strip vendor prefix if needed: "deepseek/deepseek-chat" → "deepseek-chat")
// 2. Free namespaces bill $0 regardless of the model name behind them.
if (isFreeModel(model)) return ZERO_PRICING;
// 3. Canonical model pricing (strip vendor prefix if needed: "deepseek/deepseek-chat" → "deepseek-chat")
const baseModel = model.includes("/") ? model.split("/").pop() : model;
if (MODEL_PRICING[baseModel]) return MODEL_PRICING[baseModel];
if (MODEL_PRICING[model]) return MODEL_PRICING[model];
// 3. Pattern match
// 4. Pattern match
for (const { pattern, pricing } of PATTERN_PRICING) {
if (matchPattern(pattern, baseModel) || matchPattern(pattern, model)) {
return pricing;

View File

@@ -0,0 +1,29 @@
export default {
id: "agnes",
priority: 120,
alias: "agnes",
aliases: [
"agnes-ai",
],
uiAlias: "agnes",
display: {
name: "Agnes AI",
icon: "auto_awesome",
color: "#7C3AED",
textIcon: "AG",
website: "https://agnes-ai.com",
notice: {
text: "OpenAI-compatible gateway from Agnes AI, offering free API credits on sign-up. Accepts a bearer token or an x-api-key header.",
apiKeyUrl: "https://platform.agnes-ai.com",
},
},
category: "freeTier",
authType: "apikey",
transport: {
baseUrl: "https://apihub.agnes-ai.com/v1/chat/completions",
validateUrl: "https://apihub.agnes-ai.com/v1/models",
},
// No model ids could be verified without a key, so discovery is left to the
// live endpoint and any id is accepted through passthroughModels.
passthroughModels: true,
};

View File

@@ -17,7 +17,7 @@ export default {
deprecationNotice: "RISK_NOTICE",
},
category: "oauth",
serviceKinds: ["llm", "image"],
serviceKinds: ["llm", "image", "webSearch"],
transport: {
baseUrls: [ANTIGRAVITY_IDE_BASE_URL],
format: "antigravity",
@@ -36,8 +36,8 @@ export default {
},
},
usage: {
// Discovery (quota/project) on PROD; daily host rejects these.
quotaApiUrl: "https://cloudcode-pa.googleapis.com/v1internal:fetchAvailableModels",
quotaApiUrl: `${ANTIGRAVITY_IDE_BASE_URL}/v1internal:fetchAvailableModels`,
quotaSummaryApiUrl: `${ANTIGRAVITY_IDE_BASE_URL}/v1internal:retrieveUserQuotaSummary`,
loadProjectApiUrl: "https://cloudcode-pa.googleapis.com/v1internal:loadCodeAssist",
tokenUrl: "https://oauth2.googleapis.com/token",
},
@@ -45,6 +45,10 @@ export default {
clientSecret: "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",
},
models: [
{ id: "gemini-3.8-flash-high", name: "Gemini 3.8 Flash (High)", upstreamModelId: "gemini-3.8-flash-high(high)" },
{ id: "gemini-3.8-flash-medium", name: "Gemini 3.8 Flash (Medium)", upstreamModelId: "gemini-3.8-flash-medium(medium)" },
{ id: "gemini-3.8-flash-low", name: "Gemini 3.8 Flash (Low)", upstreamModelId: "gemini-3.8-flash-low(low)" },
{ id: "gemini-3.8-flash", name: "Gemini 3.8 Flash", upstreamModelId: "gemini-3.8-flash-medium(medium)" },
{ id: "gemini-3.7-flash-high", name: "Gemini 3.7 Flash (High)", upstreamModelId: "gemini-3.7-flash-tiered(high)" },
{ id: "gemini-3.7-flash-medium", name: "Gemini 3.7 Flash (Medium)", upstreamModelId: "gemini-3.7-flash-tiered(medium)" },
{ id: "gemini-3.7-flash-low", name: "Gemini 3.7 Flash (Low)", upstreamModelId: "gemini-3.7-flash-tiered(low)" },
@@ -82,6 +86,11 @@ export default {
loadCodeAssistUserAgent: ANTIGRAVITY_IDE_USER_AGENT,
refreshLeadMs: 300000,
},
searchViaChat: {
defaultModel: "gemini-2.5-flash",
endpoint: `${ANTIGRAVITY_IDE_BASE_URL}/v1internal:generateContent`,
freeTier: "Free — Google Search grounding through an Antigravity OAuth account.",
},
features: {
usage: true,
},

View File

@@ -20,6 +20,8 @@ export default {
authModes: [
"apikey",
],
passthroughModels: true,
modelsFetcher: { url: "https://api.airforce/v1/models", type: "airforce-free" },
transport: {
baseUrl: "https://api.airforce/v1/chat/completions",
validateUrl: "https://api.airforce/v1/models",
@@ -27,10 +29,11 @@ export default {
"HTTP-Referer": "https://endpoint-proxy.local",
"X-Title": "Endpoint Proxy",
},
forceStream: true,
},
models: [
{ id: "anthropic/claude-3.7-sonnet", name: "Claude 3.7 Sonnet (Free)", contextLength: 200000 },
{ id: "moonshot/kimi-k2.6", name: "Kimi K2.6 (Free)", contextLength: 262144 },
{ id: "google/gemini-2.5-flash", name: "Gemini 2.5 Flash (Free)", contextLength: 1048576 },
{ id: "gpt-oss-120b", name: "GPT-OSS 120B (Free)", contextLength: 131072 },
{ id: "gpt-oss-20b", name: "GPT-OSS 20B (Free)", contextLength: 131072 },
{ id: "kimi-k2.7-code", name: "Kimi K2.7 Code (Free)", contextLength: 262144 },
],
};

View File

@@ -0,0 +1,32 @@
export default {
id: "atria",
priority: 120,
alias: "atria",
aliases: [
"atria-asi",
],
uiAlias: "atria",
display: {
name: "Atria Dawn",
icon: "flare",
color: "#C2410C",
textIcon: "AD",
website: "https://atria-asi.ai",
notice: {
text: "OpenAI-compatible endpoint from Atria Dawn (AtomInnoLab). Currently a research preview offering a single text model, Atria-Dawn-Preview.",
apiKeyUrl: "https://api.atria-asi.ai/dashboard",
},
},
category: "apikey",
authType: "apikey",
transport: {
baseUrl: "https://api.atria-asi.ai/v1/chat/completions",
validateUrl: "https://api.atria-asi.ai/v1/models",
},
// Docs pin the model field to one case-sensitive id. Text-only for now: the
// service ships a hook that blocks image/PDF input, so no vision is claimed.
models: [
{ id: "Atria-Dawn-Preview", name: "Atria Dawn Preview" },
],
passthroughModels: true,
};

View File

@@ -0,0 +1,30 @@
export default {
id: "bai",
priority: 120,
alias: "bai",
aliases: [
"b-ai",
],
uiAlias: "bai",
display: {
name: "B.AI",
icon: "account_balance",
color: "#0369A1",
textIcon: "BA",
website: "https://b.ai",
notice: {
text: "OpenAI-compatible gateway with one of the larger catalogues here. Accepts a bearer token or an x-api-key header. Model ids are fetched live from the provider.",
apiKeyUrl: "https://b.ai",
},
},
category: "apikey",
authType: "apikey",
transport: {
baseUrl: "https://api.b.ai/v1/chat/completions",
validateUrl: "https://api.b.ai/v1/models",
},
// No ids hardcoded: the catalogue is large and rotates, so the live endpoint
// is the source of truth and any id is accepted via passthroughModels.
modelsFetcher: { url: "https://api.b.ai/v1/models", type: "openai" },
passthroughModels: true,
};

View File

@@ -1,4 +1,4 @@
import { CLAUDE_CLI_SPOOF_HEADERS } from "../shared.js";
import { CLAUDE_CLI_VERSION } from "../shared.js";
export default {
id: "claude",
@@ -25,7 +25,7 @@ export default {
"Anthropic-Version": "2023-06-01",
"Anthropic-Beta": "claude-code-20250219,oauth-2025-04-20,interleaved-thinking-2025-05-14,context-management-2025-06-27,prompt-caching-scope-2026-01-05,advanced-tool-use-2025-11-20,effort-2025-11-24,structured-outputs-2025-12-15,fast-mode-2026-02-01,redact-thinking-2026-02-12,token-efficient-tools-2026-03-28",
"Anthropic-Dangerous-Direct-Browser-Access": "true",
"User-Agent": "claude-cli/2.1.92 (external, sdk-cli)",
"User-Agent": `claude-cli/${CLAUDE_CLI_VERSION} (external, sdk-cli)`,
"X-App": "cli",
"X-Stainless-Helper-Method": "stream",
"X-Stainless-Retry-Count": "0",
@@ -54,10 +54,14 @@ export default {
oauthUrl: "https://api.anthropic.com/api/oauth/usage",
orgUrl: "https://api.anthropic.com/v1/organizations/{org_id}/usage",
settingsUrl: "https://api.anthropic.com/v1/settings",
profileUrl: "https://api.anthropic.com/api/oauth/profile",
resetUrl: "https://api.anthropic.com/api/organizations/{org_id}/reset_rate_limits",
},
},
models: [
{ id: "claude-opus-5-5", name: "Claude Opus 5.5" },
{ id: "claude-opus-5", name: "Claude Opus 5" },
{ id: "claude-fable-5-1", name: "Claude Fable 5.1" },
{ id: "claude-fable-5", name: "Claude Fable 5" },
{ id: "claude-sonnet-5", name: "Claude Sonnet 5" },
{ id: "claude-haiku-4-5-20251001", name: "Claude 4.5 Haiku" },

View File

@@ -14,12 +14,16 @@ export default {
},
},
category: "oauth",
authModes: ["oauth"],
hasOAuth: true,
transport: {
baseUrl: "https://api.cline.bot/api/v1/chat/completions",
headers: {
"HTTP-Referer": "https://cline.bot",
"X-Title": "Cline",
},
// Non-stream chat completions come back wrapped in {"success":true,"data":{...}}
quirks: { clineEnvelope: true },
tokenUrl: "https://api.cline.bot/api/v1/auth/token",
refreshUrl: "https://api.cline.bot/api/v1/auth/refresh",
auth: {

View File

@@ -14,7 +14,10 @@ export default {
},
},
category: "oauth",
authModes: ["oauth", "apikey"],
// ClinePass authenticates with a plain API key from app.cline.bot/settings/api-keys
// (category "apikey"). The OAuth extension flow used by Cline does not issue
// tokens that the ClinePass API consumer endpoint accepts (HTTP 401) — see #2333.
authModes: ["apikey", "oauth"],
hasOAuth: true,
transport: {
baseUrl: "https://api.cline.bot/api/v1/chat/completions",
@@ -22,6 +25,8 @@ export default {
"HTTP-Referer": "https://cline.bot",
"X-Title": "Cline",
},
// Non-stream chat completions come back wrapped in {"success":true,"data":{...}}
quirks: { clineEnvelope: true },
auth: {
combined: true,
header: "Authorization",

View File

@@ -47,19 +47,28 @@ export default {
models: [
{ id: "glm-5.2", name: "GLM-5.2" },
{ id: "glm-5.1", name: "GLM-5.1" },
{ id: "glm-5.0", name: "GLM-5.0" },
{ id: "glm-5.0-turbo", name: "GLM-5.0-Turbo" },
{ id: "glm-5v-turbo", name: "GLM-5v-Turbo" },
{ id: "glm-4.7", name: "GLM-4.7" },
{ id: "minimax-m3", name: "MiniMax-M3" },
{ id: "minimax-m2.7", name: "MiniMax-M2.7" },
{ id: "kimi-k2.7", name: "Kimi-K2.7-Code" },
{ id: "kimi-k2.6", name: "Kimi-K2.6" },
{ id: "kimi-k2.5", name: "Kimi-K2.5" },
{ id: "hy3-preview", name: "Hy3 Preview" },
// Catalog mirrors the server's product-config payload (the plugin fetches
// it from copilot.tencent.com). Models the server no longer publishes are
// removed even when the chat endpoint still answers them — the published
// list is the contract. Drop log: glm-5.0 / glm-4.7 and hy4-preview-x
// (endpoint returns 11102 "model service info not found"), plus
// glm-5.0-turbo / minimax-m2.7 / kimi-k2.5 / hy3-preview /
// deepseek-v3-2-volc (absent from the server list, though still answering
// 200) and hy3-x (paid tier, not used here). deepseek-v4-flash removed
// 2026-09: replaced server-side by deepseek-v4.1-flash (same low/high/
// xhigh efforts; endpoint still answers 200 but the list is the contract).
// "-x" suffix = paid tier of the same model (free id rides the promo quota).
{ id: "hy3", name: "Hy3" },
{ id: "hy4-preview", name: "Hy4-Preview" },
{ id: "glm-5.3", name: "GLM-5.3" },
{ id: "glm-5.3-flash", name: "GLM-5.3-Flash" },
{ id: "kimi-k3-1", name: "Kimi-K3" },
{ id: "deepseek-v4-pro", name: "DeepSeek-V4-Pro" },
{ id: "deepseek-v4-flash", name: "DeepSeek-V4-Flash" },
{ id: "deepseek-v3-2-volc", name: "DeepSeek-V3.2" },
{ id: "deepseek-v4.1-flash", name: "DeepSeek-V4.1-Flash" },
],
oauth: {
baseUrl: "https://copilot.tencent.com",

View File

@@ -58,7 +58,9 @@ export default {
{ id: "kimi-k2.5", name: "Kimi-K2.5" },
{ id: "hy3-preview", name: "Hy3 Preview" },
{ id: "deepseek-v4-pro", name: "DeepSeek-V4-Pro" },
{ id: "deepseek-v4-flash", name: "DeepSeek-V4-Flash" },
// deepseek-v4-flash replaced server-side by deepseek-v4.1-flash (same
// catalog as CN; the old endpoint still answers 200 but the list is the contract).
{ id: "deepseek-v4.1-flash", name: "DeepSeek-V4.1-Flash" },
{ id: "deepseek-v3-2-volc", name: "DeepSeek-V3.2" },
],
oauth: {

View File

@@ -1,5 +1,10 @@
import { withCodexReviewModels } from "../models/helpers.js";
// Codex CLI version seen by OpenAI's backend — single source for the Version /
// User-Agent identity headers. Bump when the installed codex CLI is upgraded.
const CODEX_CLI_VERSION = "0.155.0";
const GPT_6_LITE_THINKING_LEVELS = ["low", "medium", "high", "xhigh", "max"];
export default {
id: "codex",
priority: 30,
@@ -34,9 +39,11 @@ export default {
baseUrl: "https://chatgpt.com/backend-api/codex/responses",
format: "openai-responses",
forceStream: true,
cliVersion: CODEX_CLI_VERSION,
headers: {
originator: "codex_cli_rs",
"User-Agent": "codex_cli_rs/0.136.0",
"User-Agent": `codex_cli_rs/${CODEX_CLI_VERSION}`,
version: CODEX_CLI_VERSION,
},
usage: {
url: "https://chatgpt.com/backend-api/wham/usage",
@@ -45,6 +52,9 @@ export default {
},
},
models: [
{ id: "gpt-6-astra", name: "GPT 6.0 Astra" },
{ id: "gpt-6-sol", name: "GPT 6.0 Sol", responsesLite: true, thinkingLevels: GPT_6_LITE_THINKING_LEVELS },
{ id: "gpt-6-luna", name: "GPT 6.0 Luna", responsesLite: true, thinkingLevels: GPT_6_LITE_THINKING_LEVELS },
{ id: "gpt-5.6-sol", name: "GPT 5.6 Sol" },
{ id: "gpt-5.6-sol-review", name: "GPT 5.6 Sol Review", upstreamModelId: "gpt-5.6-sol", quotaFamily: "review" },
{ id: "gpt-5.6-terra", name: "GPT 5.6 Terra" },
@@ -59,6 +69,17 @@ export default {
{ id: "gpt-5.4-mini-review", name: "GPT 5.4 Mini Review", upstreamModelId: "gpt-5.4-mini", quotaFamily: "review" },
{ id: "gpt-5.3-codex-spark", name: "GPT 5.3 Codex Spark" },
{ id: "gpt-5.3-codex-spark-review", name: "GPT 5.3 Codex Spark Review", upstreamModelId: "gpt-5.3-codex-spark", quotaFamily: "review" },
// Codex CLI's auto-review virtual model. Unlike the "-review" variants above it is not derived
// from a base model, so it is forwarded verbatim instead of having "-review" stripped (#1398).
{ id: "codex-auto-review", name: "Codex Auto Review", upstreamModelId: "codex-auto-review", quotaFamily: "review" },
{ id: "gpt-image-2.5", name: "GPT Image 2.5", capabilities: ["text2img","edit","multiImage"], params: ["size","quality","background","image_detail","output_format"], kind: "image" },
{ id: "gpt-image-2.5-flare", name: "GPT Image 2.5 Flare", capabilities: ["text2img","edit","multiImage"], params: ["size","quality","background","image_detail","output_format"], kind: "image" },
{ id: "gpt-image-2.5-sunburst", name: "GPT Image 2.5 Sunburst", capabilities: ["text2img","edit","multiImage"], params: ["size","quality","background","image_detail","output_format"], kind: "image" },
{ id: "gpt-image-2", name: "GPT Image 2", capabilities: ["text2img","edit","multiImage"], params: ["size","quality","background","image_detail","output_format"], kind: "image" },
{ id: "gpt-image-1.5", name: "GPT Image 1.5", capabilities: ["text2img","edit","multiImage"], params: ["size","quality","background","image_detail","output_format"], kind: "image" },
{ id: "gpt-5.6-sol-image", name: "GPT 5.6 Sol Image", capabilities: ["text2img","edit"], params: ["size","quality","background","image_detail","output_format"], kind: "image" },
{ id: "gpt-5.6-terra-image", name: "GPT 5.6 Terra Image", capabilities: ["text2img","edit"], params: ["size","quality","background","image_detail","output_format"], kind: "image" },
{ id: "gpt-5.6-luna-image", name: "GPT 5.6 Luna Image", capabilities: ["text2img","edit"], params: ["size","quality","background","image_detail","output_format"], kind: "image" },
{ id: "gpt-5.5-image", name: "GPT 5.5 Image", capabilities: ["text2img","edit"], params: ["size","quality","background","image_detail","output_format"], kind: "image" },
{ id: "gpt-5.4-image", name: "GPT 5.4 Image", capabilities: ["text2img","edit"], params: ["size","quality","background","image_detail","output_format"], kind: "image" },
{ id: "gpt-5.3-image", name: "GPT 5.3 Image", capabilities: ["text2img","edit"], params: ["size","quality","background","image_detail","output_format"], kind: "image" },

View File

@@ -23,7 +23,7 @@ export default {
format: "commandcode",
forceStream: true,
headers: {
"x-command-code-version": "1.10.0",
"x-command-code-version": "0.45.0",
"x-cli-environment": "cli",
"User-Agent": "cli",
},
@@ -38,21 +38,40 @@ export default {
summaryUrl: "/alpha/usage/summary",
},
},
models: [
{ id: "deepseek/deepseek-v4-pro", name: "DeepSeek V4 Pro" },
{ id: "deepseek/deepseek-v4-flash", name: "DeepSeek V4 Flash" },
{ id: "moonshotai/Kimi-K2.7-Code", name: "Kimi K2.7 Code" },
{ id: "moonshotai/Kimi-K2.7-Code-Highspeed", name: "Kimi K2.7 Code HighSpeed" },
{ id: "moonshotai/Kimi-K2.6", name: "Kimi K2.6" },
{ id: "moonshotai/Kimi-K2.5", name: "Kimi K2.5" },
{ id: "zai-org/GLM-5.2", name: "GLM 5.2" },
{ id: "zai-org/GLM-5.2-Fast", name: "GLM 5.2 Fast" },
{ id: "zai-org/GLM-5.1", name: "GLM 5.1" },
{ id: "zai-org/GLM-5", name: "GLM 5" },
{ id: "MiniMaxAI/MiniMax-M3", name: "MiniMax M3" },
{ id: "MiniMaxAI/MiniMax-M2.7", name: "MiniMax M2.7" },
{ id: "MiniMaxAI/MiniMax-M2.5", name: "MiniMax M2.5" },
{ id: "xiaomi/mimo-v2.5-pro", name: "MiMo V2.5 Pro" },
{ id: "xiaomi/mimo-v2.5", name: "MiMo V2.5" },
{ id: "Qwen/Qwen3.6-Max-Preview", name: "Qwen 3.6 Max Preview" },
{ id: "Qwen/Qwen3.6-Plus", name: "Qwen 3.6 Plus" },
{ id: "Qwen/Qwen3.7-Max", name: "Qwen 3.7 Max" },
{ id: "Qwen/Qwen3.7-Plus", name: "Qwen 3.7 Plus" },
{ id: "stepfun/Step-3.7-Flash", name: "Step 3.7 Flash" },
{ id: "stepfun/Step-3.5-Flash", name: "Step 3.5 Flash" },
{ id: "tencent/Hy3", name: "Tencent Hy3" },
{ id: "nvidia/nemotron-3-ultra-550b-a55b", name: "Nemotron 3 Ultra" },
{ id: "thinkingmachines/inkling", name: "Inkling" },
{ id: "claude-sonnet-5", name: "Claude Sonnet 5" },
{ id: "claude-sonnet-4-6", name: "Claude Sonnet 4.6" },
{ id: "claude-fable-5", name: "Claude Fable 5" },
{ id: "claude-opus-4-8", name: "Claude Opus 4.8" },
{ id: "claude-opus-4-7", name: "Claude Opus 4.7" },
{ id: "claude-haiku-4-5", name: "Claude Haiku 4.5" },
],
features: {
usage: true,
usageApikey: true,
},
models: [
{ id: "deepseek/deepseek-v4-pro", name: "DeepSeek V4 Pro" },
{ id: "deepseek/deepseek-v4-flash", name: "DeepSeek V4 Flash" },
{ id: "moonshotai/Kimi-K2.6", name: "Kimi K2.6" },
{ id: "moonshotai/Kimi-K2.5", name: "Kimi K2.5" },
{ id: "zai-org/GLM-5.1", name: "GLM 5.1" },
{ id: "zai-org/GLM-5", name: "GLM 5" },
{ id: "MiniMaxAI/MiniMax-M2.7", name: "MiniMax M2.7" },
{ id: "MiniMaxAI/MiniMax-M2.5", name: "MiniMax M2.5" },
{ id: "Qwen/Qwen3.6-Max-Preview", name: "Qwen 3.6 Max Preview" },
{ id: "Qwen/Qwen3.6-Plus", name: "Qwen 3.6 Plus" },
{ id: "stepfun/Step-3.5-Flash", name: "Step 3.5 Flash" },
],
};

View File

@@ -0,0 +1,35 @@
export default {
id: "dahl",
priority: 120,
alias: "dahl",
aliases: [
"dahl-inference",
],
uiAlias: "dahl",
display: {
name: "Dahl Inference",
icon: "hub",
color: "#1E40AF",
textIcon: "DH",
website: "https://dahl.global",
notice: {
text: "OpenAI-compatible Gonka inference node. Small, fixed catalogue (GLM-5.3-Flash, DeepSeek-V4-Flash, MiniMax-M2.7) at a flat per-token rate.",
apiKeyUrl: "https://dahl.global/dashboard",
},
},
category: "apikey",
authType: "apikey",
transport: {
baseUrl: "https://inference.dahl.global/v1/chat/completions",
validateUrl: "https://inference.dahl.global/v1/models",
},
// The live catalogue is public (no auth), so modelsFetcher works without a key
// and the ids below are a convenience seed rather than an exhaustive list.
models: [
{ id: "zai-org/GLM-5.3-Flash", name: "GLM-5.3 Flash" },
{ id: "deepseek-ai/DeepSeek-V4-Flash-0731", name: "DeepSeek V4 Flash 0731" },
{ id: "MiniMaxAI/MiniMax-M2.7", name: "MiniMax M2.7" },
],
modelsFetcher: { url: "https://inference.dahl.global/v1/models", type: "openai" },
passthroughModels: true,
};

View File

@@ -25,6 +25,21 @@ export default {
reasoningInject: {
scope: "all",
},
quirks: {
// DeepSeek's Anthropic-compatible endpoint
// (https://api.deepseek.com/anthropic/v1/messages) accepts ONLY the
// built-in web_search_* tools and rejects client-defined `custom` tools
// (MCP / Read / Bash / etc.) with HTTP 400
// "tools[0]: unknown variant `custom`, expected
// `web_search_20250305` or `web_search_20260209`".
//
// Declaring this whitelist makes prepareClaudeRequest() forward only
// web_search_* tools and strip everything else before sending, so MCP /
// function tools are dropped instead of failing the whole request.
// DeepSeek's OpenAI-compatible transport is unaffected (targetFormat
// there is "openai", not "claude", so prepareClaudeRequest is not run).
claudeSupportedToolTypes: ["web_search_20250305", "web_search_20260209"],
},
},
// Multi-endpoint: pick the transport matching client sourceFormat to skip translation.
transports: [
@@ -44,7 +59,9 @@ export default {
{ id: "deepseek-v4-pro", name: "DeepSeek V4 Pro" },
{ id: "deepseek-v4-pro-max", name: "DeepSeek V4 Pro Max", upstreamModelId: "deepseek-v4-pro" },
{ id: "deepseek-v4-pro-none", name: "DeepSeek V4 Pro No Thinking", upstreamModelId: "deepseek-v4-pro" },
{ id: "deepseek-v4.1-flash", name: "DeepSeek V4.1 Flash" },
{ id: "deepseek-v4-flash", name: "DeepSeek V4 Flash" },
{ id: "deepseek-v4-flash-vision-exp", name: "DeepSeek V4 Flash Vision (Exp)" },
{ id: "deepseek-chat", name: "DeepSeek V3.2 Chat" },
{ id: "deepseek-reasoner", name: "DeepSeek V3.2 Reasoner" },
],

View File

@@ -36,6 +36,7 @@ export default {
},
},
models: [
{ id: "gemini-3.8-flash", name: "Gemini 3.8 Flash" },
{ id: "gemini-3.7-flash", name: "Gemini 3.7 Flash" },
{ id: "gemini-3.6-flash", name: "Gemini 3.6 Flash" },
{ id: "gemini-3.5-flash-lite", name: "Gemini 3.5 Flash Lite" },
@@ -57,6 +58,7 @@ export default {
{ id: "gemini-2.5-flash", name: "Gemini 2.5 Flash", params: ["language","prompt"], kind: "stt" },
{ id: "gemini-2.5-flash-lite", name: "Gemini 2.5 Flash Lite (Cheapest)", params: ["language","prompt"], kind: "stt" },
{ id: "gemini-2.0-flash", name: "Gemini 2.0 Flash", params: ["language","prompt"], kind: "stt" },
{ id: "gemini-2.5-flash-native-audio-preview-09-17", name: "Gemini Live Transcription (Realtime)", params: ["language","prompt","system_instruction","setup_timeout_ms","turn_timeout_ms"], kind: "stt", transport: "gemini-live" },
{ id: "gemini-3.1-flash-tts-preview", name: "Gemini 3.1 Flash TTS", kind: "tts" },
{ id: "gemini-2.5-flash-preview-tts", name: "Gemini 2.5 Flash TTS", kind: "tts" },
{ id: "gemini-2.5-pro-preview-tts", name: "Gemini 2.5 Pro TTS", kind: "tts" },

View File

@@ -22,10 +22,13 @@ export default {
},
models: [
{ id: "glm-5.3", name: "GLM 5.3" },
{ id: "glm-5.3-flash", name: "GLM 5.3 Flash (Vision)" },
{ id: "glm-5.2", name: "GLM 5.2" },
{ id: "glm-5.1", name: "GLM 5.1" },
{ id: "glm-5-turbo", name: "GLM 5 Turbo" },
{ id: "glm-5", name: "GLM 5" },
{ id: "glm-4.7", name: "GLM-4.7" },
{ id: "glm-4.6v", name: "GLM 4.6V (Vision)" },
{ id: "glm-4.6", name: "GLM-4.6" },
{ id: "glm-4.5-air", name: "GLM-4.5-Air" },
],

View File

@@ -46,12 +46,28 @@ export default {
],
models: [
{ id: "glm-5.3", name: "GLM 5.3" },
{ id: "glm-5.3-flash", name: "GLM 5.3 Flash (Vision)" },
{ id: "glm-5.2", name: "GLM 5.2" },
{ id: "glm-5.1", name: "GLM 5.1" },
{ id: "glm-5-turbo", name: "GLM 5 Turbo" },
{ id: "glm-5", name: "GLM 5" },
{ id: "glm-4.7", name: "GLM 4.7" },
{ id: "glm-4.6v", name: "GLM 4.6V (Vision)" },
],
serviceKinds: ["llm", "webSearch"],
// Coding plan bundles web search on the same API key as chat.
searchConfig: {
baseUrl: "https://api.z.ai/api/mcp/web_search_prime/mcp",
method: "POST",
authType: "apikey",
authHeader: "bearer",
costPerQuery: 0,
searchTypes: ["web"],
defaultMaxResults: 5,
maxMaxResults: 50,
timeoutMs: 10000,
cacheTTLMs: 300000,
},
features: {
usage: true,
usageApikey: true,

View File

@@ -17,6 +17,12 @@ export default {
transport: {
baseUrl: "https://api.groq.com/openai/v1/chat/completions",
validateUrl: "https://api.groq.com/openai/v1/models",
// No dedicated quota endpoint; rate-limit info rides on x-ratelimit-*
// response headers, always included. Reuse the models list (already
// used as validateUrl) so reading usage never costs tokens.
usage: {
url: "https://api.groq.com/openai/v1/models",
},
},
models: [
{ id: "llama-3.3-70b-versatile", name: "Llama 3.3 70B" },
@@ -34,4 +40,8 @@ export default {
authHeader: "bearer",
format: "openai",
},
features: {
usage: true,
usageApikey: true,
},
};

View File

@@ -15,6 +15,7 @@ export default {
website: "https://huggingface.co",
notice: {
apiKeyUrl: "https://huggingface.co/settings/tokens",
text: "Runs through the Inference Providers router. Image and speech models are billed by the provider selected per model.",
},
},
category: "apikey",
@@ -25,10 +26,79 @@ export default {
transport: null,
models: [
{ id: "black-forest-labs/FLUX.1-schnell", name: "FLUX.1 Schnell", params: [], kind: "image" },
{ id: "black-forest-labs/FLUX.1-dev", name: "FLUX.1 Dev", params: [], kind: "image" },
{ id: "black-forest-labs/FLUX.1-Krea-dev", name: "FLUX.1 Krea", params: [], kind: "image" },
{ id: "black-forest-labs/FLUX.1-Kontext-dev", name: "FLUX.1 Kontext", params: [], kind: "image", capabilities: ["edit"] },
{ id: "black-forest-labs/FLUX.2-dev", name: "FLUX.2 Dev", params: [], kind: "image", capabilities: ["edit"] },
{ id: "black-forest-labs/FLUX.2-klein-9B", name: "FLUX.2 Klein 9B", params: [], kind: "image", capabilities: ["edit"] },
{ id: "black-forest-labs/FLUX.2-klein-4B", name: "FLUX.2 Klein 4B", params: [], kind: "image", capabilities: ["edit"] },
{ id: "black-forest-labs/FLUX.2-klein-base-9B", name: "FLUX.2 Klein Base 9B", params: [], kind: "image", capabilities: ["edit"] },
{ id: "black-forest-labs/FLUX.2-klein-base-4B", name: "FLUX.2 Klein Base 4B", params: [], kind: "image", capabilities: ["edit"] },
{ id: "stabilityai/stable-diffusion-xl-base-1.0", name: "SDXL Base 1.0", params: [], kind: "image" },
{ id: "openai/whisper-large-v3", name: "Whisper Large v3 (HF)", params: ["language"], kind: "stt" },
{ id: "openai/whisper-small", name: "Whisper Small (HF)", params: ["language"], kind: "stt" },
{ id: "stabilityai/stable-diffusion-3.5-large", name: "Stable Diffusion 3.5 Large", params: [], kind: "image" },
{ id: "stabilityai/stable-diffusion-3.5-large-turbo", name: "Stable Diffusion 3.5 Large Turbo", params: [], kind: "image" },
{ id: "Qwen/Qwen-Image", name: "Qwen Image", params: [], kind: "image" },
{ id: "Qwen/Qwen-Image-2512", name: "Qwen Image 2512", params: [], kind: "image" },
{ id: "Qwen/Qwen-Image-Edit", name: "Qwen Image Edit", params: [], kind: "image", capabilities: ["edit"] },
{ id: "Qwen/Qwen-Image-Edit-2509", name: "Qwen Image Edit 2509", params: [], kind: "image", capabilities: ["edit"] },
{ id: "Qwen/Qwen-Image-Edit-2511", name: "Qwen Image Edit 2511", params: [], kind: "image", capabilities: ["edit"] },
{ id: "ideogram-ai/ideogram-4-fp8", name: "Ideogram 4", params: [], kind: "image" },
{ id: "tencent/HunyuanImage-3.0", name: "HunyuanImage 3.0", params: [], kind: "image" },
{ id: "Tongyi-MAI/Z-Image-Turbo", name: "Z-Image Turbo", params: [], kind: "image" },
{ id: "krea/Krea-2-Turbo", name: "Krea 2 Turbo", params: [], kind: "image" },
{ id: "HiDream-ai/HiDream-I1-Fast", name: "HiDream I1 Fast", params: [], kind: "image" },
{ id: "playgroundai/playground-v2.5-1024px-aesthetic", name: "Playground v2.5", params: [], kind: "image" },
{ id: "openai/whisper-large-v3", name: "Whisper Large v3 (HF)", params: [], kind: "stt" },
{ id: "openai/whisper-large-v3-turbo", name: "Whisper Large v3 Turbo (HF)", params: [], kind: "stt" },
],
serviceKinds: ["image", "stt"],
imageConfig: { baseUrl: "https://api-inference.huggingface.co/models" },
// Inference Providers router. The router is addressed as
// `<baseUrl>/<provider>/<providerModelId>` — see open-sse/handlers/imageProviders/huggingface.js.
// `modelMap` resolves a Hub model id to the provider-resolved id the router expects.
// A plain string value is the provider path. Image-to-image models use
// `{ path, task: "image-to-image" }`: their payload differs — the source image goes in
// `inputs` and the prompt under `parameters.prompt`. See
// https://huggingface.co/docs/inference-providers/tasks/image-to-image
// Only providers the router actually forwards to are listed: replicate, wavespeed and
// deepinfra appear in the Hub's inferenceProviderMapping but reject router traffic with
// "Model not supported by provider <name>".
imageConfig: {
baseUrl: "https://router.huggingface.co",
modelMap: {
"black-forest-labs/FLUX.1-schnell": "fal-ai/fal-ai/flux/schnell",
"black-forest-labs/FLUX.1-dev": "fal-ai/fal-ai/flux/dev",
"black-forest-labs/FLUX.1-Krea-dev": "fal-ai/fal-ai/flux/krea",
"black-forest-labs/FLUX.1-Kontext-dev": { path: "fal-ai/fal-ai/flux-kontext/dev", task: "image-to-image" },
"black-forest-labs/FLUX.2-dev": { path: "fal-ai/fal-ai/flux-2/edit", task: "image-to-image" },
"black-forest-labs/FLUX.2-klein-9B": { path: "fal-ai/fal-ai/flux-2/klein/9b/edit", task: "image-to-image" },
"black-forest-labs/FLUX.2-klein-4B": { path: "fal-ai/fal-ai/flux-2/klein/4b/distilled/edit", task: "image-to-image" },
"black-forest-labs/FLUX.2-klein-base-9B": { path: "fal-ai/fal-ai/flux-2/klein/9b/base/edit", task: "image-to-image" },
"black-forest-labs/FLUX.2-klein-base-4B": { path: "fal-ai/fal-ai/flux-2/klein/4b/base/edit", task: "image-to-image" },
"stabilityai/stable-diffusion-xl-base-1.0": "fal-ai/fal-ai/fast-sdxl",
"stabilityai/stable-diffusion-3.5-large": "fal-ai/fal-ai/stable-diffusion-v35-large",
"stabilityai/stable-diffusion-3.5-large-turbo": "fal-ai/fal-ai/stable-diffusion-v35-large/turbo",
"Qwen/Qwen-Image": "fal-ai/fal-ai/qwen-image",
"Qwen/Qwen-Image-2512": "fal-ai/fal-ai/qwen-image-2512",
"Qwen/Qwen-Image-Edit": { path: "fal-ai/fal-ai/qwen-image-edit", task: "image-to-image" },
"Qwen/Qwen-Image-Edit-2509": { path: "fal-ai/fal-ai/qwen-image-edit-2509", task: "image-to-image" },
"Qwen/Qwen-Image-Edit-2511": { path: "fal-ai/fal-ai/qwen-image-edit-plus", task: "image-to-image" },
"ideogram-ai/ideogram-4-fp8": "fal-ai/ideogram/v4",
"tencent/HunyuanImage-3.0": "fal-ai/fal-ai/hunyuan-image/v3/text-to-image",
"Tongyi-MAI/Z-Image-Turbo": "fal-ai/fal-ai/z-image/turbo",
"krea/Krea-2-Turbo": "fal-ai/fal-ai/krea-2/turbo",
"HiDream-ai/HiDream-I1-Fast": "fal-ai/fal-ai/hidream-i1-fast",
"playgroundai/playground-v2.5-1024px-aesthetic": "fal-ai/fal-ai/playground-v25",
},
},
// Speech-to-text goes through the hf-inference provider, which keeps the Hub
// model id as its provider-resolved id (`/hf-inference/models/<hubId>`).
// No `params` are declared: the router's ASR payload carries only `inputs` and
// `parameters.return_timestamps` / `parameters.generation_parameters` — it has no
// language field, so a UI-declared "language" would be silently dropped.
sttConfig: {
baseUrl: "https://router.huggingface.co/hf-inference/models",
authType: "apikey",
authHeader: "bearer",
format: "huggingface-asr",
},
};

View File

@@ -66,8 +66,10 @@ import p63 from "./nebius.js";
import p64 from "./nvidia.js";
import p65 from "./ollama-local.js";
import p66 from "./ollama.js";
import p123 from "./ollama-search.js";
import p67 from "./openai.js";
import p68 from "./opencode-go.js";
import p68z from "./opencode-zen.js";
import p69 from "./opencode.js";
import p70 from "./openrouter.js";
import p71 from "./perplexity-web.js";
@@ -75,6 +77,7 @@ import p72 from "./perplexity.js";
import p73 from "./perplexity-agent.js";
import p74 from "./playht.js";
import p75 from "./qoder.js";
import p124 from "./qoder-cn.js";
import p77 from "./recraft.js";
import p78 from "./runwayml.js";
import p79 from "./sdwebui.js";
@@ -121,7 +124,12 @@ import p118 from "./selfhosted-tts.js";
import p119 from "./selfhosted-embedding.js";
import p120 from "./fish-audio.js";
import p121 from "./alitp-intl.js";
import p122 from "./xquik.js";
import p125 from "./tokenharbor.js";
import p126 from "./dahl.js";
import p127 from "./atria.js";
import p129 from "./agnes.js";
import p130 from "./bai.js";
export default [
p0,
p1,
@@ -190,8 +198,11 @@ export default [
p64,
p65,
p66,
p123,
p124,
p67,
p68,
p68z,
p69,
p70,
p71,
@@ -243,4 +254,10 @@ export default [
p119,
p120,
p121,
p122,
p125,
p126,
p127,
p129,
p130,
];

View File

@@ -22,6 +22,7 @@ export default {
headers: { ...CLAUDE_API_HEADERS },
quirks: {
dropOutputConfig: true,
requireClaudeToolType: true,
},
reasoningInject: {
scope: "all",

View File

@@ -22,6 +22,7 @@ export default {
headers: { ...CLAUDE_API_HEADERS },
quirks: {
dropOutputConfig: true,
requireClaudeToolType: true,
},
reasoningInject: {
scope: "all",

View File

@@ -0,0 +1,35 @@
export default {
id: "ollama-search",
alias: "ollama-search",
display: {
name: "Ollama Search",
icon: "cloud",
color: "#ffffff",
textIcon: "OL",
website: "https://ollama.com",
notice: {
text: "Web search via Ollama Cloud subscription. Reuses the API key from the Ollama (chat) provider.",
apiKeyUrl: "https://ollama.com/settings/keys",
},
},
category: "apikey",
authType: "apikey",
authModes: ["apikey"],
serviceKinds: ["webSearch"],
// Credential fallback: reuses the API key registered under the `ollama`
// chat provider — one key, chat + search.
credentialFallback: "ollama",
searchConfig: {
baseUrl: "https://ollama.com/api/web_search",
method: "POST",
authType: "apikey",
authHeader: "bearer",
costPerQuery: 0,
freeMonthlyQuota: 1000,
searchTypes: ["web"],
defaultMaxResults: 5,
maxMaxResults: 10,
timeoutMs: 10000,
cacheTTLMs: 300000,
},
};

View File

@@ -30,8 +30,18 @@ export default {
{ id: "glm-4.7-flash", name: "GLM 4.7 Flash" },
{ id: "qwen3.5", name: "Qwen3.5" },
{ id: "minimax-m3", name: "MiniMax M3" },
{ id: "deepseek-v4.1-flash:cloud", name: "DeepSeek V4.1 Flash" },
],
serviceKinds: ["llm"],
serviceKinds: ["llm", "webFetch"],
fetchConfig: {
baseUrl: "https://ollama.com/api/web_fetch",
method: "POST",
authType: "apikey",
authHeader: "bearer",
formats: ["markdown"],
maxCharacters: 200000,
timeoutMs: 30000,
},
features: {
usage: true,
usageApikey: true,

View File

@@ -28,6 +28,7 @@ export default {
forceStream: true,
},
models: [
{ id: "gpt-5.5", name: "GPT-5.5" },
{ id: "gpt-5.4", name: "GPT-5.4" },
{ id: "gpt-5.4-mini", name: "GPT-5.4 Mini" },
{ id: "gpt-5.4-nano", name: "GPT-5.4 Nano" },
@@ -57,6 +58,9 @@ export default {
{ id: "whisper-1", name: "Whisper 1", params: ["language","response_format","temperature","prompt"], kind: "stt" },
{ id: "gpt-4o-transcribe", name: "GPT-4o Transcribe", params: ["language","response_format","temperature","prompt"], kind: "stt" },
{ id: "gpt-4o-mini-transcribe", name: "GPT-4o Mini Transcribe", params: ["language","response_format","temperature","prompt"], kind: "stt" },
{ id: "gpt-image-2.5", name: "GPT Image 2.5", params: ["n","size","quality","response_format"], kind: "image" },
{ id: "gpt-image-2.5-flare", name: "GPT Image 2.5 Flare", params: ["n","size","quality","response_format"], kind: "image" },
{ id: "gpt-image-2.5-sunburst", name: "GPT Image 2.5 Sunburst", params: ["n","size","quality","response_format"], kind: "image" },
{ id: "gpt-image-1", name: "GPT Image 1", params: ["n","size","quality","response_format"], kind: "image" },
{ id: "dall-e-3", name: "DALL-E 3", params: ["size","quality","style","response_format"], kind: "image" },
{ id: "dall-e-2", name: "DALL-E 2", params: ["n","size","response_format"], kind: "image" },

View File

@@ -13,7 +13,7 @@ export default {
textIcon: "OC",
website: "https://opencode.ai/auth",
notice: {
text: "OpenCode Go subscription: $5/mo (then 0/mo). Access to Kimi, GLM, Qwen, MiMo, MiniMax models.",
text: "OpenCode Go subscription: $5/mo (then 10/mo). Access to Kimi, GLM, Qwen, MiMo, MiniMax models.",
apiKeyUrl: "https://opencode.ai/auth",
},
},
@@ -21,6 +21,9 @@ export default {
transport: {
baseUrl: "https://opencode.ai/zen/go/v1/chat/completions",
headers: {},
usage: {
url: "https://opencode.ai/zen/go/v1/usage",
},
},
// Multi-endpoint: pick the transport matching the client sourceFormat to skip
// translation. Guarded per-model by `supportedFormats` (see chatCore) because
@@ -30,20 +33,60 @@ export default {
{ format: "claude", baseUrl: "https://opencode.ai/zen/go/v1/messages", auth: { combined: true, header: "x-api-key", scheme: "raw", anthropicVersion: true } },
{ format: "openai-responses", baseUrl: "https://opencode.ai/zen/go/v1/responses", auth: { combined: true, header: "Authorization", scheme: "bearer" } },
],
// supportedFormats follow the endpoint table in https://opencode.ai/docs/go/
models: [
{ id: "deepseek-flash", name: "DeepSeek Flash", supportedFormats: ["openai"] },
{ id: "glm-5.3-flash", name: "GLM 5.3 Flash (Vision)", supportedFormats: ["openai"] },
{ id: "glm-5.3", name: "GLM 5.3", supportedFormats: ["openai"] },
{ id: "glm-5.2", name: "GLM 5.2", supportedFormats: ["openai"] },
{ id: "glm-5.1", name: "GLM 5.1", supportedFormats: ["openai"] },
{ id: "glm-5", name: "GLM 5", supportedFormats: ["openai"] },
{ id: "kimi-k2.7-code", name: "Kimi K2.7 Code", supportedFormats: ["openai"] },
{ id: "kimi-k2.6", name: "Kimi K2.6", supportedFormats: ["openai"] },
{ id: "kimi-k2.5", name: "Kimi K2.5", supportedFormats: ["openai"] },
{ id: "kimi-k3", name: "Kimi K3", supportedFormats: ["openai"] },
{ id: "deepseek-v4-pro", name: "DeepSeek V4 Pro", supportedFormats: ["openai", "claude", "openai-responses"] },
{ id: "deepseek-v4-flash", name: "DeepSeek V4 Flash", supportedFormats: ["openai", "claude", "openai-responses"] },
{ id: "deepseek-v4-flash-vision-exp", name: "DeepSeek V4 Flash Vision (Exp)", supportedFormats: ["openai", "claude", "openai-responses"] },
{ id: "deepseek-v4.1-flash", name: "DeepSeek V4.1 Flash", supportedFormats: ["openai", "claude", "openai-responses"] },
{ id: "longcat-2.0", name: "LongCat 2.0", supportedFormats: ["openai"] },
{ id: "mimo-v2.6-flash", name: "MiMo V2.6 Flash", supportedFormats: ["openai"] },
{ id: "mimo-v2.6-pro", name: "MiMo V2.6 Pro", supportedFormats: ["openai"] },
{ id: "mimo-v2.5", name: "MiMo V2.5", supportedFormats: ["openai"] },
{ id: "mimo-v2.5-pro", name: "MiMo V2.5 Pro", supportedFormats: ["openai"] },
{ id: "mimo-v2-pro", name: "MiMo V2 Pro", supportedFormats: ["openai"] },
{ id: "mimo-v2-omni", name: "MiMo V2 Omni", supportedFormats: ["openai"] },
{ id: "minimax-m3", name: "MiniMax M3", supportedFormats: ["openai", "claude"] },
{ id: "minimax-m2.7", name: "MiniMax M2.7", supportedFormats: ["openai", "claude"] },
{ id: "minimax-m2.5", name: "MiniMax M2.5", supportedFormats: ["openai", "claude"] },
{ id: "space-bunny-free", name: "Space Bunny Free", supportedFormats: ["openai", "claude"] },
{ id: "qwen3.8-max", name: "Qwen 3.8 Max", supportedFormats: ["openai", "claude"] },
{ id: "qwen3.8-flash", name: "Qwen 3.8 Flash", supportedFormats: ["openai", "claude"] },
{ id: "qwen3.7-max", name: "Qwen 3.7 Max", supportedFormats: ["openai", "claude"] },
{ id: "qwen3.7-plus", name: "Qwen 3.7 Plus", supportedFormats: ["openai", "claude"] },
{ id: "qwen3.6-plus", name: "Qwen 3.6 Plus", supportedFormats: ["openai", "claude"] },
{ id: "qwen3.5-plus", name: "Qwen 3.5 Plus", supportedFormats: ["openai", "claude"] },
{ id: "hy4-preview", name: "Hy4 Preview", supportedFormats: ["openai"] },
{ id: "hy3", name: "Hy3", supportedFormats: ["openai"] },
{ id: "hy3-preview", name: "Hy3 Preview", supportedFormats: ["openai"] },
// In /zen/go/v1/models but absent from the docs endpoint table — chat lane is the fallback guess
{ id: "omen-alpha", name: "Omen Alpha", supportedFormats: ["openai"] },
// Served by /zen/go/v1/responses only — the responses-only entry forces chatCore
// past the sourceFormat-matched transports into translation (see chatCore guard).
{ id: "grok-4.7", name: "Grok 4.7", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "grok-4.6", name: "Grok 4.6", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "grok-4.5", name: "Grok 4.5", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.6-luna", name: "GPT 5.6 Luna", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-6-luna", name: "GPT 6 Luna", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "muse-spark-1.2-contributor", name: "Muse Spark 1.2 Contributor", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "muse-spark-1.3-contributor", name: "Muse Spark 1.3 Contributor", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
],
// Live catalogue; ids outside this curated list get their endpoint lane from the
// family regex in providers/models/helpers.js (opencodeFamilyFormats).
modelsFetcher: { url: "https://opencode.ai/zen/go/v1/models", type: "opencode-go" },
passthroughModels: true,
features: {
usage: true,
usageApikey: true,
},
};

View File

@@ -0,0 +1,135 @@
export default {
id: "opencode-zen",
priority: 205,
alias: "ocz",
aliases: [
"opencode-zen",
],
uiAlias: "ocz",
display: {
name: "OpenCode Zen",
icon: "terminal",
color: "#E87040",
textIcon: "OZ",
website: "https://opencode.ai/auth",
notice: {
text: "OpenCode Zen PAYG: pay-as-you-go, key from https://opencode.ai/auth. Same models as Zen: paid + free tiers on the fast lane.",
apiKeyUrl: "https://opencode.ai/auth",
},
},
category: "apikey",
transport: {
baseUrl: "https://opencode.ai/zen/v1/chat/completions",
headers: {},
usage: {
url: "https://opencode.ai/zen/v1/usage",
},
},
// Multi-endpoint: pick the transport matching the client sourceFormat to skip
// translation. Mirrors opencode-go, pointed at /zen/v1 (see https://opencode.ai/docs/zen/).
transports: [
{ format: "openai", baseUrl: "https://opencode.ai/zen/v1/chat/completions", auth: { combined: true, header: "Authorization", scheme: "bearer" } },
{ format: "claude", baseUrl: "https://opencode.ai/zen/v1/messages", auth: { combined: true, header: "x-api-key", scheme: "raw", anthropicVersion: true } },
{ format: "openai-responses", baseUrl: "https://opencode.ai/zen/v1/responses", auth: { combined: true, header: "Authorization", scheme: "bearer" } },
],
// supportedFormats follow the endpoint table in https://opencode.ai/docs/zen/
// (live /zen/v1/models, 2026-09-18: 71 ids).
models: [
// Claude (messages)
{ id: "claude-fable-5", name: "Claude Fable 5", supportedFormats: ["claude"] },
{ id: "claude-fable-5-1", name: "Claude Fable 5.1", supportedFormats: ["claude"] },
{ id: "claude-opus-5", name: "Claude Opus 5", supportedFormats: ["claude"] },
{ id: "claude-opus-4-8", name: "Claude Opus 4.8", supportedFormats: ["claude"] },
{ id: "claude-opus-4-7", name: "Claude Opus 4.7", supportedFormats: ["claude"] },
{ id: "claude-opus-4-6", name: "Claude Opus 4.6", supportedFormats: ["claude"] },
{ id: "claude-opus-4-5", name: "Claude Opus 4.5", supportedFormats: ["claude"] },
{ id: "claude-sonnet-5", name: "Claude Sonnet 5", supportedFormats: ["claude"] },
{ id: "claude-sonnet-4-6", name: "Claude Sonnet 4.6", supportedFormats: ["claude"] },
{ id: "claude-sonnet-4-5", name: "Claude Sonnet 4.5", supportedFormats: ["claude"] },
{ id: "claude-sonnet-4", name: "Claude Sonnet 4", supportedFormats: ["claude"] },
{ id: "claude-haiku-4-5", name: "Claude Haiku 4.5", supportedFormats: ["claude"] },
// Gemini (own path, via chat completions transport)
{ id: "gemini-3.6-flash", name: "Gemini 3.6 Flash", supportedFormats: ["openai"] },
{ id: "gemini-3.8-flash", name: "Gemini 3.8 Flash", supportedFormats: ["openai"] },
{ id: "gemini-3.7-flash", name: "Gemini 3.7 Flash", supportedFormats: ["openai"] },
{ id: "gemini-3.5-flash-lite", name: "Gemini 3.5 Flash Lite", supportedFormats: ["openai"] },
{ id: "gemini-3.5-flash", name: "Gemini 3.5 Flash", supportedFormats: ["openai"] },
{ id: "gemini-3.1-pro", name: "Gemini 3.1 Pro", supportedFormats: ["openai"] },
{ id: "gemini-3-flash", name: "Gemini 3 Flash", supportedFormats: ["openai"] },
// GPT / Grok / Muse Spark paid (responses)
{ id: "gpt-6-astra", name: "GPT 6 Astra", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.6-sol", name: "GPT 5.6 Sol", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.6-terra", name: "GPT 5.6 Terra", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.6-luna", name: "GPT 5.6 Luna", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.5", name: "GPT 5.5", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.5-pro", name: "GPT 5.5 Pro", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.4", name: "GPT 5.4", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.4-pro", name: "GPT 5.4 Pro", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.4-mini", name: "GPT 5.4 Mini", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.4-nano", name: "GPT 5.4 Nano", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.3-codex-spark", name: "GPT 5.3 Codex Spark", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.3-codex", name: "GPT 5.3 Codex", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.2", name: "GPT 5.2", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.2-codex", name: "GPT 5.2 Codex", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.1", name: "GPT 5.1", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.1-codex-max", name: "GPT 5.1 Codex Max", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.1-codex", name: "GPT 5.1 Codex", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5.1-codex-mini", name: "GPT 5.1 Codex Mini", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5", name: "GPT 5", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5-codex", name: "GPT 5 Codex", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "gpt-5-nano", name: "GPT 5 Nano", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "grok-build-0.1", name: "Grok Build 0.1", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "grok-4.6", name: "Grok 4.6", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "grok-4.5", name: "Grok 4.5", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "muse-spark-1.3", name: "Muse Spark 1.3", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "muse-spark-1.2", name: "Muse Spark 1.2", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
// Qwen paid (messages)
{ id: "qwen3.6-plus", name: "Qwen 3.6 Plus", supportedFormats: ["claude"] },
{ id: "qwen3.5-plus", name: "Qwen 3.5 Plus", supportedFormats: ["claude"] },
// DeepSeek / GLM / MiniMax / Kimi / Big Pickle (chat completions)
{ id: "deepseek-v4-pro", name: "DeepSeek V4 Pro", supportedFormats: ["openai"] },
{ id: "deepseek-v4-flash", name: "DeepSeek V4 Flash", supportedFormats: ["openai"] },
{ id: "deepseek-v4-flash-vision-exp", name: "DeepSeek V4 Flash Vision Exp", supportedFormats: ["openai"] },
{ id: "glm-5.3-flash", name: "GLM 5.3 Flash (Vision)", supportedFormats: ["openai"] },
{ id: "glm-5.3", name: "GLM 5.3", supportedFormats: ["openai"] },
{ id: "glm-5.2", name: "GLM 5.2", supportedFormats: ["openai"] },
{ id: "glm-5.1", name: "GLM 5.1", supportedFormats: ["openai"] },
{ id: "glm-5", name: "GLM 5", supportedFormats: ["openai"] },
{ id: "minimax-m3", name: "MiniMax M3", supportedFormats: ["openai"] },
{ id: "minimax-m2.7", name: "MiniMax M2.7", supportedFormats: ["openai"] },
{ id: "minimax-m2.5", name: "MiniMax M2.5", supportedFormats: ["openai"] },
{ id: "kimi-k3", name: "Kimi K3", supportedFormats: ["openai"] },
{ id: "kimi-k2.7-code", name: "Kimi K2.7 Code", supportedFormats: ["openai"] },
{ id: "kimi-k2.6", name: "Kimi K2.6", supportedFormats: ["openai"] },
{ id: "kimi-k2.5", name: "Kimi K2.5", supportedFormats: ["openai"] },
{ id: "big-pickle", name: "Big Pickle", supportedFormats: ["openai"] },
{ id: "union-alpha", name: "Union Alpha", supportedFormats: ["claude"] },
// Free tier on the keyed lane (chat completions)
{ id: "deepseek-v4-flash-free", name: "DeepSeek V4 Flash Free", supportedFormats: ["openai"] },
{ id: "mimo-v2.6-flash-free", name: "MiMo V2.6 Flash Free", supportedFormats: ["openai"] },
{ id: "mimo-v2.5-free", name: "MiMo V2.5 Free", supportedFormats: ["openai"] },
{ id: "ling-3.0-flash-fin-free", name: "Ling 3.0 Flash Fin Free", supportedFormats: ["openai"] },
{ id: "nemotron-3-ultra-free", name: "Nemotron 3 Ultra Free", supportedFormats: ["openai"] },
{ id: "nemotron-3.5-lightning-free", name: "Nemotron 3.5 Lightning Free", supportedFormats: ["openai"] },
// Free tier on the keyed lane (responses)
{ id: "muse-spark-1.3-contributor-free", name: "Muse Spark 1.3 Contributor Free", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
{ id: "muse-spark-1.2-contributor-free", name: "Muse Spark 1.2 Contributor Free", targetFormat: "openai-responses", supportedFormats: ["openai-responses"] },
// System One (Jev) decision models on the native /systemone endpoint
{ id: "jev-1.13", name: "Jev 1.13", kind: "systemone" },
{ id: "jev-1.13-free", name: "Jev 1.13 Free", kind: "systemone" },
],
serviceKinds: ["llm", "systemone"],
systemoneConfig: {
baseUrl: "https://opencode.ai/zen/v1/systemone",
headers: {
"x-opencode-client": "desktop",
"User-Agent": "opencode/1.18.31",
},
},
modelsFetcher: { url: "https://opencode.ai/zen/v1/models", type: "opencode-free" },
passthroughModels: true,
features: {
usage: true,
usageApikey: true,
},
};

View File

@@ -17,9 +17,27 @@ export default {
headers: {
"x-opencode-client": "desktop",
},
forceStream: true,
noAuth: true,
quirks: {
forceAutoToolChoiceModels: ["muse-spark-1.3-contributor-free"],
},
},
models: [
// Endpoint formats differ per model, so declare non-chat models explicitly.
{ id: "muse-spark-1.2-contributor-free", name: "Muse Spark 1.2 Contributor Free", targetFormat: "openai-responses" },
{ id: "muse-spark-1.3-contributor-free", name: "Muse Spark 1.3 Contributor Free", targetFormat: "openai-responses" },
{ id: "union-alpha", name: "Union Alpha Free", targetFormat: "claude" },
{ id: "jev-1.13-free", name: "Jev 1.13 Free", kind: "systemone" },
],
serviceKinds: ["llm", "systemone"],
systemoneConfig: {
baseUrl: "https://opencode.ai/zen/v1/systemone",
headers: {
"x-opencode-client": "desktop",
"User-Agent": "opencode/1.18.31",
},
},
models: [],
modelsFetcher: { url: "https://opencode.ai/zen/v1/models", type: "opencode-free" },
passthroughModels: true,
};

View File

@@ -40,8 +40,17 @@ export default {
{ id: "openai/gpt-image-1", name: "GPT Image 1 (via OpenRouter)", params: ["n","size","quality","response_format"], kind: "image" },
{ id: "google/imagen-3.0-generate-002", name: "Imagen 3 (via OpenRouter)", params: ["n","size"], kind: "image" },
{ id: "black-forest-labs/FLUX.1-schnell", name: "FLUX.1 Schnell (via OpenRouter)", params: ["n","size"], kind: "image" },
{ id: "google/veo-3.1", name: "Veo 3.1 (via OpenRouter)", params: ["duration","aspect_ratio","resolution"], kind: "video" },
{ id: "openai/sora-2-pro", name: "Sora 2 Pro (via OpenRouter)", params: ["duration","aspect_ratio","resolution"], kind: "video" },
{ id: "bytedance/seedance-2.0", name: "Seedance 2.0 (via OpenRouter)", params: ["duration","aspect_ratio","resolution"], kind: "video" },
{ id: "typesafe/jev-1.13", name: "Jev 1.13", kind: "systemone" },
],
serviceKinds: ["llm","embedding","tts","imageToText"],
serviceKinds: ["llm","embedding","tts","imageToText","video","systemone"],
// System One decision API (TypeSafe-compatible): https://openrouter.ai/docs/guides/community/typesafe-sdk
systemoneConfig: {
baseUrl: "https://openrouter.ai/api/v1/systemone",
headers: {"HTTP-Referer":"https://endpoint-proxy.local","X-Title":"Endpoint Proxy"},
},
ttsConfig: {
baseUrl: "https://openrouter.ai/api/v1/chat/completions",
defaultModel: "openai/gpt-4o-mini-tts",
@@ -57,6 +66,12 @@ export default {
baseUrl: "https://openrouter.ai/api/v1/images/generations",
headers: {"HTTP-Referer":"https://endpoint-proxy.local","X-Title":"Endpoint Proxy"},
},
// Async video jobs (POST /videos → { id, status }, GET /videos/{id} polls).
// Docs: https://openrouter.ai/docs/api/api-reference/videos
videoConfig: {
baseUrl: "https://openrouter.ai/api/v1/videos",
headers: {"HTTP-Referer":"https://endpoint-proxy.local","X-Title":"Endpoint Proxy"},
},
modelsFetcher: { url: "https://openrouter.ai/api/v1/models", type: "openrouter-free" },
passthroughModels: true,
};

View File

@@ -0,0 +1,61 @@
export default {
id: "qoder-cn",
priority: 30,
alias: "qdcn",
uiAlias: "qdcn",
display: {
name: "Qoder CN",
icon: "water_drop",
color: "#EC4899",
website: "https://qoder.com.cn",
notice: {
signupUrl: "https://qoder.com.cn",
},
},
category: "oauth",
authModes: ["oauth", "apikey"],
hasOAuth: true,
authHint: "Personal Access Token (pt-...) from https://qoder.com.cn/account/integrations",
transport: {
baseUrl: "https://gateway.qoder.com.cn/algo/api/v2/service/pro/sse/agent_chat_generation",
headers: {},
timeoutMs: 120000,
stallTimeoutMs: 120000,
usage: {
url: "https://openapi.qoder.com.cn/api/v2/quota/usage",
},
},
models: [
{ id: "ultimate", name: "Ultimate" },
{ id: "auto", name: "Auto" },
{ id: "performance", name: "Performance" },
{ id: "efficient", name: "Efficient" },
{ id: "lite", name: "Lite" },
{ id: "qmodel_38max", name: "Qwen3.8-Max" },
{ id: "qmodel_latest", name: "Qwen3.7-Max" },
{ id: "qmodel", name: "Qwen3.7-Plus" },
{ id: "qfmodel", name: "Qwen3.8-Flash" },
{ id: "kmodel_latest", name: "Kimi-K3" },
{ id: "kmodel", name: "Kimi-K2.7-Code" },
{ id: "gmodel", name: "GLM-5.3" },
{ id: "gfmodel", name: "GLM-5.3-Flash" },
{ id: "dmodel", name: "DeepSeek-V4-Pro" },
{ id: "dfmodel", name: "DeepSeek-V4-Flash" },
{ id: "mmodel", name: "MiniMax-M3" },
],
oauth: {
openApiBaseUrl: "https://openapi.qoder.com.cn",
centerBaseUrl: "https://gateway.qoder.com.cn",
chatBaseUrl: "https://gateway.qoder.com.cn",
deviceTokenUrl: "https://openapi.qoder.com.cn/api/v1/deviceToken/poll",
refreshUrl: "https://gateway.qoder.com.cn/algo/api/v3/user/refresh_token",
userInfoUrl: "https://openapi.qoder.com.cn/api/v1/userinfo",
quotaUsageUrl: "https://openapi.qoder.com.cn/api/v2/quota/usage",
loginUrl: "https://qoder.com.cn/device/selectAccounts",
},
features: {
usage: true,
// PAT (apikey) connections also carry quota usage (via job-token exchange).
usageApikey: true,
},
};

View File

@@ -30,12 +30,15 @@ export default {
{ id: "auto", name: "Auto" },
{ id: "performance", name: "Performance" },
{ id: "efficient", name: "Efficient" },
{ id: "qmodel_preview", name: "Qwen3.8-Max-Preview" },
{ id: "lite", name: "Lite" },
{ id: "qmodel_38max", name: "Qwen3.8-Max" },
{ id: "qmodel_latest", name: "Qwen3.7-Max" },
{ id: "qmodel", name: "Qwen3.7-Plus" },
{ id: "qfmodel", name: "Qwen3.8-Flash" },
{ id: "kmodel_latest", name: "Kimi-K3" },
{ id: "kmodel", name: "Kimi-K2.7-Code" },
{ id: "gm51model", name: "GLM-5.2" },
{ id: "gmodel", name: "GLM-5.3" },
{ id: "gfmodel", name: "GLM-5.3-Flash" },
{ id: "dmodel", name: "DeepSeek-V4-Pro" },
{ id: "dfmodel", name: "DeepSeek-V4-Flash" },
{ id: "mmodel", name: "MiniMax-M3" },

View File

@@ -0,0 +1,49 @@
export default {
id: "tokenharbor",
priority: 120,
alias: "tokenharbor",
aliases: [
"th",
"thh",
],
uiAlias: "tokenharbor",
display: {
name: "Token Harbor",
icon: "anchor",
color: "#0F766E",
textIcon: "TH",
website: "https://tokenharbor.ai",
notice: {
text: "OpenAI-compatible aggregator. One API key reaches every model, billed per-token from a prepaid wallet. Model ids are bare (e.g. claude-opus-5.5, gpt-6-astra, deepseek-v4.1-flash:free) and are fetched live from the provider.",
apiKeyUrl: "https://tokenharbor.ai/dashboard",
},
},
category: "apikey",
authType: "apikey",
transport: {
// OpenAI-compatible. `format` is left at the shared "openai" default and
// `thinkingFormat` is deliberately NOT declared: Token Harbor forwards
// requests verbatim, so each model must resolve its own thinking wire
// format through providers/capabilities.js. Setting a provider-wide value
// would force one format (e.g. claude-adaptive) onto every model.
baseUrl: "https://tokenharbor.ai/v1/chat/completions",
validateUrl: "https://tokenharbor.ai/v1/models",
retry: {
429: 2,
},
},
// Curated seed; the live catalogue is fetched via modelsFetcher and any other
// id is accepted via passthroughModels. Their catalogue rotates (the :free set
// in particular), so this stays deliberately small and is only the offline
// fallback. Ids are bare — Token Harbor does not prefix them by upstream vendor.
models: [
{ id: "claude-opus-5.5", name: "Claude Opus 5.5" },
{ id: "claude-sonnet-5", name: "Claude Sonnet 5" },
{ id: "gpt-6-astra", name: "GPT-6 Astra" },
{ id: "gpt-6-sol", name: "GPT-6 Sol" },
{ id: "deepseek-v4.1-flash:free", name: "DeepSeek V4.1 Flash (Free)" },
{ id: "grok-4.7", name: "Grok 4.7" },
],
modelsFetcher: { url: "https://tokenharbor.ai/v1/models", type: "openai" },
passthroughModels: true,
};

Some files were not shown because too many files have changed in this diff Show More