diff --git a/CHANGELOG.md b/CHANGELOG.md index ced1cd82..e62edded 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,17 +1,86 @@ -# Unreleased +# v0.5.50 (2026-08-05) ## Features - -- **Stream error patterns**: per-provider `streamErrorPatterns` setting (UI: provider page → Stream Error Patterns) — HTTP-200 streams whose first bytes match configured patterns (plain text or `/regex/`) are treated as failed requests: fallback works for non-streaming and early stream errors, and late streaming errors are logged as FAILED. Zero overhead when unconfigured. +- **Providers**: add TokenRouter (300+ models via OpenAI-compatible gateway) with + exact per-model pricing for 110 models and `reasoning_effort` thinking config +- **Providers**: add Self-hosted STT / TTS / Embedding — point 9Router at your own + OpenAI-compatible speech and embedding servers (whisper.cpp, faster-whisper, + Kokoro-FastAPI, llama-server, vLLM, Infinity). Unlike the named cloud providers + these read `baseUrl` per connection, so one provider can front several machines +- **Combos**: default-enable vision/audio capacity adapter (auto-routes to a + vision/audio-capable model when the target lacks that capability, falling back + to `oc/mimo-v2.5-free`), wired into chat handler routing +- **Endpoint**: auto-provision a "Default Key" for first-time users so `/v1` + works without a manual dashboard step +- **Codex**: support GPT-5.6 Max/Ultra reasoning-level overrides (cx/ routes only) +- **Qoder**: support PAT (Personal Access Token) connections end-to-end, alongside + OAuth device flow +- **CLI tools**: add OpenDesign (manalkaff/opendesign) support +- **Headroom**: report effective payload savings (tool schema/history bytes broken + out, byte-savings % reflects actual outbound reduction) +- **Ollama**: Cloud quota tracker (session + weekly) + proactive background OAuth + token refresh scheduler for all providers ## Fixes +- **Providers**: remove Qwen (OAuth flow stopped working reliably) +- **Passthrough**: detect codex-tui/Codex Desktop as native Codex client — they + were falling through to the translator and losing fields like `reasoning.summary` +- **OAuth**: scope antigravity header fixes to loadCodeAssist/onboardUser only +- **OAuth**: keep `open` external in the build so xAI/Grok token refresh works on + Windows +- **OAuth**: declare missing `searchParams` in register-session handler (was a + 500 instead of JSON on error) +- **DB**: `ENABLE_REQUEST_LOGS` env var now overrides the UI setting correctly; + observability defaults to off (opt-in) +- **Translator**: preserve Codex Responses Lite tool use across chat-native + OpenAI-compatible providers +- **Translator**: don't drop image-only user messages in `prepareClaudeRequest` +- **Translator**: drop JSON Schema keywords Gemini rejects (`uniqueItems`, + `contains`, `multipleOf`, `unevaluatedProperties`, `unevaluatedItems`, + `contentSchema`) +- **Claude**: remove global header cache that leaked one client's identity + headers onto another client/account sharing the server; gate `anthropic-beta` + by model instead +- **Antigravity**: drop retired Gemini 3.0 quota tiers, show Gemini 3.6 Flash + usage bars +- **Cloudflare AI**: declare API key authentication (dashboard showed "No + connections" despite an active key) +- **GitHub Copilot**: hold monthly-exhausted accounts until UTC month reset + instead of only cooling down 120s +- **CodeBuddy**: dodge Tencent CN content filter, add usage tracking, normalize + codebuddy-intl messages +- **Usage**: stop losing cached prompt tokens in the forced-SSE→JSON path +- **Grok CLI**: display the public subscription tier from the OAuth token claim +- **Providers**: count apikey connections for Ollama free-tier card; free-tier/ + apikey providers without `authModes` now default to apikey (were treated + oauth-only) +- **Build**: include static/public assets in standalone output (login page hung + on 404s when run via PM2) +- **Server**: support IntelliJ IDEA OpenAI-compatible clients over HTTP (h2c + upgrade handling) +- **Auth**: redirect already-logged-in sessions away from `/login` +- **CLI tools**: enable Apply button for dynamic OpenAI/Anthropic-compatible + provider connections +- **CLI**: include complete API artifacts in the CLI package +- **TTS**: a bare self-hosted model name is the MODEL, not the voice — `kokoro` + was parsed as a voice against a default model, 404ing or synthesising with the + wrong one +- **Embeddings**: self-hosted embeddings no longer fall back to `api.openai.com` + when a connection has no `baseUrl` — that silently sent the input text and API + key to OpenAI under a provider named "Self-hosted" +- **Embeddings**: an adapter that rejects a misconfigured connection now returns + 400 with the reason instead of escaping the handler uncaught +- **Embeddings**: bound the upstream fetch with `FETCH_CONNECT_TIMEOUT_MS` — an + endpoint that drops packets never returns headers, so the request previously + hung indefinitely -- **CommandCode**: in-stream `{"type":"error"}` events now emit OpenAI error chunks + an executor early-peek → 502 fallback instead of fake success content (`[CommandCode error: ...]`). +## Docs +- **i18n**: fix port typo, add RTK Token Saver feature descriptions # v0.5.45 (2026-07-30) ## Features - +- **TTS**: add Xiaomi MiMo text-to-speech (preset voices 冰糖/茉莉/苏打/白桦/Mia/Chloe/Milo/Dean, style control, language hint dropdown with Auto-detect, i18n for Style label/placeholder) - **Providers**: add Poolside (OpenAI-compatible) - **Providers**: add api-airforce, baidu, bazaarlink, bluesminds, kilo-gateway, llm7, morph, sambanova, tencent - **OAuth**: zed / trae / windsurf providers + harden callback proxies @@ -24,7 +93,6 @@ - **Usage**: SuperGrok weekly pool via gRPC-web ## Fixes - - **Refresh**: rotate `refresh_token` between retry attempts - **Kiro**: canonicalize tool history and route API keys correctly - **Kiro**: normalize dashboard thinking intensity models @@ -39,21 +107,18 @@ - **Dashboard**: flex quota rows, thin global scrollbars, no hidden-row overflow ## Docs - - **i18n**: expand pt-BR translation to 986 terms - README: Indonesian translation # v0.5.40 (2026-07-20) ## Features - - **i18n**: add Khmer (km) translations - **CLI tools**: configure Grok Build subagent models - **Kimi**: merge OAuth into dual-auth provider, add K3 / K2.7 models - **Dashboard**: ProviderTopology flow animation ## Fixes - - **DB**: resolve better-sqlite3 parameter binding crash - **Translator**: pass `service_tier` through OpenAI → Responses conversion - **Kiro**: map GPT-5.6 reasoning effort fields @@ -64,10 +129,10 @@ - **Cursor**: HTTP/2 AgentService support + version bump 3.12.17 - **Dashboard**: cut duplicate API/icon spam, lazy-load provider assets + # v0.5.35 (2026-07-16) ## Features - - **xAI**: Grok Imagine video generation (`/v1/videos`) + CLI - **CLI tools**: Grok Build setup — choose separate main/general-purpose/explore/plan models and preserve each model's context window - **GitHub Copilot**: route Claude models through Copilot's native `/v1/messages` @@ -78,7 +143,6 @@ - **i18n**: Thai (th) + Persian (fa) translations / README ## Fixes - - **Providers**: bulk-add API keys no longer overwrite existing keys (gap-fill `Key N`) - **Anthropic**: lowercase `anthropic-version` header to prevent duplication on `/v1/messages` - **Alicode-intl**: use DashScope compatible-mode endpoint so standard keys work @@ -91,17 +155,14 @@ - **Translator**: strip `client_metadata` when converting openai-responses → openai ## Improvements - - **Perf**: skip inactive background services on startup ## Docs - - README: Persian YouTube tutorial # v0.5.30 (2026-07-10) ## Features - - **Perplexity**: add Agent API provider (#2492) - **Grok CLI**: add Grok CLI / Grok Build provider with OAuth device-code flow (#2502) - **Featherless**: add OpenAI-compatible provider presets @@ -113,7 +174,6 @@ - **Proxy-Pools**: auto-rotate strategy for no-auth providers (#2409) ## Fixes - - **Cloudflare-AI**: support accountId in bulk key import (#2449) - **DB**: backup on schema change, MCP child cleanup, codex models, usage providers OOM - **Codex**: avoid bare-email OAuth dedup (#2477) @@ -130,7 +190,6 @@ - **Pricing**: update Claude/Codex model rates and add new models ## Improvements - - **i18n(zh-CN)**: complete Chinese translations for all UI strings (#2436) - **API**: caching for tunnel and version status endpoints - **Perf**: faster dev startup and lighter bundle @@ -138,14 +197,12 @@ # v0.5.20 (2026-07-07) ## Features - - **Thinking**: per-model thinking level picker on provider page — appends `(level)` suffix to copied model names for forced reasoning effort across all formats (openai, claude, gemini, deepseek, kimi, qwen, zai, minimax, hunyuan, step) - **RTK**: add JS-native git-log filter (#2423) - **Caveman**: add targeted upstream-aligned style rules (#2424) - **i18n**: add Farsi (fa) language support (#2385) ## Fixes - - **Thinking**: strip `(level)` suffix from upstream `body.model` so providers no longer reject requests - **Translator**: preserve developer instructions in openai-responses conversion (#2434) - **count_tokens**: count structured Anthropic blocks (#2419) @@ -159,14 +216,12 @@ # v0.5.18 (2026-07-03) ## Features - - **Usage**: track cached tokens + correct input/output/cache cost (#2209) — hodtien - **Codex**: show reset credit expiry details (#2290) — Rafli Ahmad Zulfikar - **NVIDIA**: add new models and capabilities — decolua - **ClinePass**: add provider support — sternelee ## Fixes - - **Usage**: dedupe streaming request-details log entries — Qin Li - **Claude**: drop foreign thinking signatures in passthrough — decolua - Prevent non-SSE stream pipe crash and cross-IdP account overwrites (#2244) — KunN-21 @@ -183,13 +238,11 @@ # v0.5.15 (2026-06-29) ## Features - - Add Kimchi OAuth provider — Nant361 - Refine Qwen vision/video + thinking model patterns — decolua - Opt-in Codex auto-ping quota keep-alive — Emirhan ## Fixes - - **Responses**: handle response.done terminal events (#2142) — rifuki - **Headroom**: skip unsafe responses tool history (#2132) — Sutarto Jordan Chrisfivo - **Translator**: map mid-conversation system message to user (claude→openai) — decolua @@ -206,7 +259,6 @@ # v0.5.12 (2026-06-26) ## Features - - Add token-saver dashboard page — decolua - Add bulk delete for provider connections — teddytkz - Resolve GitHub Copilot model catalog from upstream — caiqinzhou @@ -215,7 +267,6 @@ - Overhaul Blackbox provider catalog + WebUI test support — suryacagur ## Fixes - - Provider thinking compatibility (DeepSeek/Gemini) — Mink Nguyen - Stop double-counting streaming usage at source — decolua - Usage logging dedupe to reduce stats churn — Mink Nguyen @@ -244,13 +295,11 @@ # v0.5.8 (2026-06-21) ## Features - - **Antigravity**: native image generation support (image models tagged kind:image, hiển thị trong media-providers UI) - **CodeBuddy CN**: API key auth + credit quota tracker - **CodeBuddy CN**: short model prefix alias "cbcn" ## Fixes - - **MiniMax-M3**: enable vision capability - **Headroom**: support Docker sidecar proxy - **Antigravity**: image executor fixes @@ -264,14 +313,12 @@ # v0.5.6 (2026-06-20) ## Features - - **Ponytail**: minimalist code generation feature - **Headroom**: proxy lifecycle management + dashboard UI (one-click start/stop, install detection, status probing, token saver, claude↔openai shape conversion) - **CodeBuddy CN**: new OAuth provider (copilot.tencent.com) — 15-model catalog, /v2 inference, forced streaming, OpenAI-style reasoning - **OpenCode-Go**: align models with official endpoints; route Qwen 3.7 MiniMax via /v1/messages, GLM/Kimi/DeepSeek/MiMo via /chat/completions ## Fixes - - **Anthropic-compatible validation**: use POST /v1/messages (GET /models not spec, false "invalid" for valid keys) - **CLI tools**: tolerate JSONC configs in all 8 settings routes (opencode, openclaw, kilo, droid, cowork, copilot, claude, cline) - **Gemini/Antigravity**: preserve 'pattern' in tool schema translation (glob/grep) @@ -282,7 +329,6 @@ # v0.5.4 (2026-06-18) ## Fixes - - **Kiro**: honor thinking effort budgets - **AG/Kiro/Xiaomi**: provider fixes - **Combo/Fusion**: flatten tool history in panel calls to prevent 503 @@ -292,7 +338,6 @@ # v0.5.2 (2026-06-17) ## Features - - **Combo Fusion strategy** — fans the prompt out to all member models in parallel, then a configurable judge model synthesizes one final answer (quorum-grace, anonymized sources, graceful degradation) - **Per-combo strategy selector** — pick `fallback` / `round-robin` / `fusion` / `capacity` per combo (replaces the old round-robin toggle), with a judge picker for fusion - **Capacity auto-switch** — reorders models per request so images/PDFs route to capable models first @@ -300,7 +345,6 @@ - **Claude auto-ping** — warms the 5h quota window right after reset so a fresh window starts immediately (per-connection toggle) ## Fixes - - **Claude 429**: stop hammering the OAuth usage endpoint — cache resetAt, throttle quota refresh to 3 min, cool down after a 429 (chat unaffected) - **Usage logs always empty**: missing `await` on `getAdapter()` in `getRecentLogs` made `/api/usage/logs` & `/api/usage/request-logs` return nothing - **Executors**: strip params unsupported by the provider/model (drops deprecated `temperature` for claude-opus-4 → Anthropic 400) @@ -312,13 +356,11 @@ - **Security**: SSRF hardening on web fetch ## Internal - - Large **open-sse / translator refactor** (~40 commits): unified provider/model registry (LiteLLM-style `models[]` + `kind` field, 100 co-located registry files), single-sourced media/OAuth/refresh/token URLs, registry-based dispatch for usage & token-refresh, DRY translator concerns (buildUsage, encodeDataUri, finishReasonMap, chunkBuilder, reasoningDelta…), ESM-safe registry init, large-file splits, dead-code removal, and golden/no-regression test gates # v0.4.80 (2026-06-13) ## Features - - Vercel AI Gateway: support embeddings, images and credit usage (#1183) - Add MiMo Free no-auth provider (#1789) - Vertex: support ADC `authorized_user` credential @@ -327,7 +369,6 @@ - Kiro: enable multi-endpoint failover for GenerateAssistantResponse (#1722) ## Fixes - - Security: re-auth on DB export/import + SSRF guard on web fetch - Auth: real client IP rate-limiting + remote default-password guard - Cerebras/Mistral: strip unsupported `client_metadata` from downstream requests (#1742) @@ -346,13 +387,11 @@ - Dashboard: show provider node name instead of connection name in topology (#1770) + show explicit `kind="llm"` combos on combos page (#1684) ## Docs - - README: add Indonesian 9Router tutorial video (#1709) # v0.4.71 (2026-06-06) ## Features - - Caveman: add wenyan classical Chinese levels and sync upstream prompts; locale-based visibility on endpoint page - i18n: endpoint exposure notice across multiple languages + Russian README - Antigravity: add gemini-3.5-flash-extra-low (Low) model @@ -361,7 +400,6 @@ - MiniMax: add MiniMax-M3 + update Quota Tracker coding/CN (#1631) ## Fixes - - Codex: harden streaming timeouts (stall/connect raised to 60s, configurable per-provider), accept `response.done` event, and always emit a terminal `response.failed` + `[DONE]` for Responses passthrough when a stream closes, stalls, or aborts before a terminal event — prevents codex clients from hanging (#1648, #1680, #1688, #1618) - Codex: durable OAuth refresh lifecycle (#1664) - Tunnel: skip virtual interfaces to prevent false netchange watchdog @@ -375,25 +413,21 @@ - Model-test: route image/STT probes to their real endpoints, harden STT ping; add opencode-go + xiaomi-tokenplan to connection test (#1576, #1628) ## Improvements - - Dashboard: reorganize menu actions across sidebar/header/profile - Translator: add data-driven coverage, bug-exposing cases, and real provider smoke tests # v0.4.66 (2026-05-29) ## Features - - Add Qoder provider: device-flow OAuth, COSY signing, WAF-bypass body encoding, live model catalog, dashboard quota tracker, 11 models (#1372) - Add new models: Claude Opus 4.8 (Claude Code), GPT 5.4 Mini (Codex) ## Fixes - - DeepSeek thinking mode: echo `reasoning_content` back on follow-up/tool-call turns so OpenCode-free and custom providers no longer 400 with "reasoning_content must be passed back" (#1543) - Reasoning injector: match deepseek/kimi model ids case-insensitively (covers custom providers using capitalized model names) - OpenCode suggested-models: include free models without the `-free` suffix, e.g. `big-pickle` (#1535) ## Improvements - - Codex: trim sunset models, keep gpt-5.5 / gpt-5.4 / gpt-5.3-codex family, add gpt-5.4-mini - volcengine-ark: refresh model list (add DeepSeek-V4-Flash/Pro, drop EOL entries) - Lower stream stall timeout 35s → 30s for faster hang detection @@ -401,21 +435,18 @@ # v0.4.63 (2026-05-26) ## Fixes - - GitHub Copilot: never route Gemini/Claude models to the `/responses` endpoint; prevents misleading "does not support Responses API" 400s (#1062) - proxyFetch: restore missing `Readable` import causing runtime `ReferenceError` in DNS-bypass fetch path ## Improvements - - Lower stream stall timeout from 60s → 35s for faster hang detection # v0.4.62 (2026-05-26) ## Fixes - - Codex: auto-retry when upstream drops mid-stream (no more hangs) - Codex: fix random 400/404 errors, tool-calling failures, and unstable prompt cache -- MITM: support Antigravity 2.x +- MITM: support Antigravity 2.x - Sanitize Read tool args to prevent retry loops from non-Anthropic models (#1144) - Implement json_schema fallback for OpenAI-compatible providers without native Structured Output (#1343) - Strip empty Read pages argument in OpenAI-to-Claude translator (#1354) @@ -424,30 +455,25 @@ - Gemini CLI: reuse stored OAuth project IDs for quota checks and show clearer setup guidance when the project is missing (#1271, #1428) ## Features - - Add Cloudflare Workers proxy deployer and pool integration (#1360) - Add Deno Deploy relays support and improved proxy pools dashboard layout (#1437) ## Improvements - - Refactor Tunnel into dedicated Cloudflare and Tailscale manager modules - Refactor tokenRefresh service with in-flight dedup to prevent refresh_token_reused errors # v0.4.59 (2026-05-21) ## Fixes - - OAuth: fix login flow on Windows # v0.4.58 (2026-05-21) ## Features - - xAI Grok provider (OAuth, API key, image) - Provider limits: paginated accounts with page size controls ## Fixes - - Tailscale: fix connection status on Windows (#1300) - Tunnel: fix false "checking" when tunnel URL is reachable - Stream: fix pipe errors on client disconnect/abort @@ -455,13 +481,11 @@ # v0.4.55 (2026-05-18) ## Features - - Xiaomi MiMo Token Plan: region selector (Singapore / China / Europe) — keys are cluster-specific - Antigravity: risk confirmation dialog before first connection - Gemini CLI: surface upstream retry delay on 429 errors ## Fixes - - MITM: cannot kill process on macOS under sudo (lsof not found in PATH) - Stream: false-positive stall timeout on Claude reasoning / Kiro responses - Tunnel: cannot re-enable after disable (stuck state) @@ -470,19 +494,16 @@ - Antigravity OAuth: metadata now matches the official client ## Improvements - - Gemini CLI: bump engine to 0.34.0 - Re-hide `qwen` (OAuth EOL) and `iflow` (not ready) providers # v0.4.52 (2026-05-17) ## Features - - Add Vercel AI Gateway provider support (#1183) - rtk: Kiro format tool result compression — handle conversationState.history & currentMessage, preserve error results, ~13.6% savings (#1194) ## Fixes - - openclaw: normalize agent.model object form `{primary, fallbacks}` before .startsWith → fix TypeError & 'not configured' status (#1216) - Usage Details pagination: stay inside mobile viewport <640px (#1218) - Fix test model error @@ -492,7 +513,6 @@ # v0.4.50 (2026-05-16) ## Fixes - - Fix duplicate tray icon on macOS when hiding to tray - Fix tray not showing in background mode on macOS - Fix hide to tray broken on Windows/Linux @@ -501,13 +521,11 @@ # v0.4.49 (2026-05-16) ## Features - - Add Kiro provider support: full request/response translation, live model listing, reasoning content support - Add `buildOutput` RTK filter with autodetect for npm/yarn/cargo build logs - Add MITM warning notification in tray and dashboard ## Improvements - - Add modalities (input/output) to model configuration for OpenCode - Fix tray hide-to-tray: keep current process alive instead of spawning detached child (fixes macOS NSStatusItem ghost icon) - Fix tray kill: graceful shutdown with SIGTERM/SIGKILL escalation @@ -516,11 +534,9 @@ - Update i18n across 32 languages ## Fixes - - Fix model check (test-models) blocked by dashboardGuard: pass machineId-based CLI token in internal self-calls # v0.4.46 (2026-05-15) ## Breaking Changes - -- Tunnel public URL changed — old tunnel links no longer work, please reconnect to get the new URL +- Tunnel public URL changed — old tunnel links no longer work, please reconnect to get the new URL \ No newline at end of file diff --git a/README.md b/README.md index 4458ba78..a89c7dca 100644 --- a/README.md +++ b/README.md @@ -285,6 +285,32 @@ Default URLs: Kilo Code + + + OpenDesign
+ OpenDesign + + + jcode
+ jcode + + + Grok Build
+ Grok Build + + + Devin CLI
+ Devin CLI + + + DeepSeek TUI
+ DeepSeek TUI + + + Qwen Code
+ Qwen Code + + @@ -441,6 +467,46 @@ Default URLs:

...and 20+ more providers including Nebius, Chutes, Hyperbolic, and custom OpenAI/Anthropic compatible endpoints

+### 🏠 Self-hosted Providers + +For speech and embeddings served from **your own** machine — whisper.cpp, +faster-whisper, Speaches, Kokoro-FastAPI, openedai-speech, llama.cpp/llama-server, +vLLM, Infinity, text-embeddings-inference, or anything else that speaks the OpenAI +shape. + +| Provider | Endpoint used | Typical server | +| --- | --- | --- | +| **Self-hosted STT** | `/v1/audio/transcriptions` | whisper.cpp, faster-whisper | +| **Self-hosted TTS** | `/v1/audio/speech` | Kokoro-FastAPI, openedai-speech | +| **Self-hosted Embedding** | `/v1/embeddings` | llama-server, vLLM, Infinity | + +Every other speech provider is a named cloud service with a fixed endpoint. These +three read their address from **each connection**, so one provider can front +several machines and load-balance across them like any other. + +Set it on the connection as `providerSpecificData.baseUrl`: + +| Provider | Give it | Result | +| --- | --- | --- | +| Self-hosted STT | the full URL — `http://host:8080/v1/audio/transcriptions` | used as-is | +| Self-hosted TTS | the server root — `http://host:8880` | `+ /v1/audio/speech` | +| Self-hosted Embedding | the **OpenAI base**, `/v1` included — `http://host:8080/v1` | `+ /embeddings` | + +> **Mind the `/v1` on embeddings.** The adapter appends `/embeddings`, so +> `http://host:8080` resolves to `http://host:8080/embeddings` and misses the +> OpenAI route — llama-server answers **501**. Give it the same base URL an OpenAI +> client would use. A full `.../v1/embeddings` is also accepted, so a value pasted +> from a `curl` example works too. + +The API key is not checked by most local servers, but the field must be non-empty: +it is what gives the connection a credentials record, and `baseUrl` lives there. +Any placeholder works. + +Self-hosted Embedding has **no cloud fallback by design** — a connection saved +without a `baseUrl` is reported as a configuration error rather than quietly +falling back to `api.openai.com`, which would send your input text and API key to +a third party under a provider named "Self-hosted". + --- ## 💡 Key Features diff --git a/cli/package.json b/cli/package.json index efb786e4..98ad8cdc 100644 --- a/cli/package.json +++ b/cli/package.json @@ -1,6 +1,6 @@ { "name": "9router", - "version": "0.5.45", + "version": "0.5.50", "description": "9Router CLI - Start and manage 9Router server", "bin": { "9router": "./cli.js" diff --git a/cli/scripts/build-cli.js b/cli/scripts/build-cli.js index a1b28c5f..d4e4d279 100644 --- a/cli/scripts/build-cli.js +++ b/cli/scripts/build-cli.js @@ -81,201 +81,272 @@ function copyRecursive(src, dest) { } } -console.log("📦 Building 9Router CLI package with Next.js...\n"); +function resolveStandaloneBuild(appDir, buildDistDir) { + const legacyStandaloneRoot = path.join(appDir, ".next", "standalone"); + const resolvedStandaloneRoot = path.join(buildDistDir, "standalone"); + let standaloneRoot = fs.existsSync(resolvedStandaloneRoot) + ? resolvedStandaloneRoot + : legacyStandaloneRoot; -fs.mkdirSync(buildHomeDir, { recursive: true }); -fs.mkdirSync(path.join(buildHomeDir, "AppData", "Roaming"), { recursive: true }); -fs.mkdirSync(path.join(buildHomeDir, "AppData", "Local"), { recursive: true }); - -// Step 0: Sync version from app/cli/package.json to app/package.json -console.log("0️⃣ Syncing version to app/package.json..."); -const cliPkg = JSON.parse(fs.readFileSync(path.join(cliDir, "package.json"), "utf8")); -const appPkgPath = path.join(appDir, "package.json"); -const appPkg = JSON.parse(fs.readFileSync(appPkgPath, "utf8")); -if (appPkg.version !== cliPkg.version) { - appPkg.version = cliPkg.version; - fs.writeFileSync(appPkgPath, JSON.stringify(appPkg, null, 2) + "\n"); - console.log(`✅ Version synced: ${cliPkg.version}\n`); -} else { - console.log(`✅ Version already synced: ${cliPkg.version}\n`); -} - -// Step 1: Build app with Next.js (workspace tracing root → traced node_modules in standalone). -console.log("1️⃣ Building Next.js app..."); -try { - execSync("npm run build", { - stdio: "inherit", - cwd: appDir, - env: { - ...process.env, - HOME: buildHomeDir, - USERPROFILE: buildHomeDir, - APPDATA: path.join(buildHomeDir, "AppData", "Roaming"), - LOCALAPPDATA: path.join(buildHomeDir, "AppData", "Local"), - NEXT_DIST_DIR: buildDistDirName, - NEXT_TRACING_ROOT_MODE: "workspace", - } - }); - console.log("✅ Next.js build completed\n"); -} catch (error) { - console.error("❌ Next.js build failed"); - process.exit(1); -} - -// Step 2: Clean old app/cli/app if exists -console.log("2️⃣ Cleaning old app/cli/app..."); -if (fs.existsSync(cliAppDir)) { - fs.rmSync(cliAppDir, { recursive: true, force: true }); -} -console.log("✅ Cleaned\n"); - -// Step 3: Copy Next.js standalone build to app/cli/app. -// Newer Next.js standalone output writes server.js/package.json plus .next/, src/, and -// node_modules/ directly under .next/standalone. Older builds may still use a nested app/. -console.log("3️⃣ Copying Next.js standalone build to app/cli/app..."); -const standaloneRoot = path.join(appDir, ".next", "standalone"); -const standaloneRootResolved = path.join(buildDistDir, "standalone"); -let standaloneRootToUse = fs.existsSync(standaloneRootResolved) ? standaloneRootResolved : standaloneRoot; -// Next.js 16 nests standalone output under the project name when NEXT_TRACING_ROOT_MODE=workspace -// e.g. .next-cli-build/standalone/9router/server.js -const pkgName = path.basename(appDir); -const nestedRoot = path.join(standaloneRootToUse, pkgName); -if (fs.existsSync(path.join(nestedRoot, "server.js")) && !fs.existsSync(path.join(standaloneRootToUse, "server.js"))) { - console.log(`ℹ️ Detected nested standalone output: ${pkgName}/`); - standaloneRootToUse = nestedRoot; -} -const standaloneApp = fs.existsSync(path.join(standaloneRootToUse, "server.js")) - ? standaloneRootToUse - : path.join(standaloneRootToUse, "app"); -if (!fs.existsSync(standaloneApp)) { - console.error("❌ Next.js standalone build not found under .next/standalone"); - console.error("Expected either .next/standalone/server.js or .next/standalone/app/"); - process.exit(1); -} -copyRecursive(standaloneApp, cliAppDir); - -// Older nested-app layout stores traced node_modules at standalone root. -const standaloneNodeModules = path.join(standaloneRootToUse, "node_modules"); -if (standaloneApp !== standaloneRootToUse && fs.existsSync(standaloneNodeModules)) { - copyRecursive(standaloneNodeModules, path.join(cliAppDir, "node_modules")); -} -console.log("✅ Copied standalone build\n"); - -// Step 3a: Copy custom server (injects real socket IP, strips spoofable XFF). -const customServerSrc = path.join(appDir, "custom-server.js"); -if (fs.existsSync(customServerSrc)) { - fs.copyFileSync(customServerSrc, path.join(cliAppDir, "custom-server.js")); - console.log("✅ Copied custom-server.js\n"); -} else { - console.warn("⚠️ custom-server.js not found — server will run without real-IP injection\n"); -} - -// Step 3b: Ensure sql.js (pure JS fallback) bundled in app/cli/app/node_modules. -// Strip better-sqlite3 (native) — it lives in ~/.9router/runtime to avoid -// Windows EBUSY during global CLI updates. node:sqlite (Node ≥22.5) is also -// available as a no-install middle tier. -console.log("3️⃣ b Configuring SQLite drivers..."); -function ensureModuleInBundle(pkg) { - const dest = path.join(cliAppDir, "node_modules", pkg); - if (fs.existsSync(dest)) { - console.log(`✅ ${pkg} already bundled`); - return; + // Next.js 16 nests standalone output under the project name when + // NEXT_TRACING_ROOT_MODE=workspace, e.g. standalone/9router/server.js. + const pkgName = path.basename(appDir); + const nestedRoot = path.join(standaloneRoot, pkgName); + if (fs.existsSync(path.join(nestedRoot, "server.js")) && !fs.existsSync(path.join(standaloneRoot, "server.js"))) { + console.log(`ℹ️ Detected nested standalone output: ${pkgName}/`); + standaloneRoot = nestedRoot; } - const candidates = [ - path.join(appDir, "node_modules", pkg), - path.join(rootDir, "node_modules", pkg), + + const standaloneApp = fs.existsSync(path.join(standaloneRoot, "server.js")) + ? standaloneRoot + : path.join(standaloneRoot, "app"); + if (!fs.existsSync(standaloneApp)) { + throw new Error( + "Next.js standalone build not found under .next/standalone; " + + "expected either .next/standalone/server.js or .next/standalone/app/", + ); + } + + return { standaloneApp, standaloneRoot }; +} + +function copyStandaloneBuild(appDir, buildDistDir, cliAppDir) { + const { standaloneApp, standaloneRoot } = resolveStandaloneBuild(appDir, buildDistDir); + copyRecursive(standaloneApp, cliAppDir); + + // Older nested-app layout stores traced node_modules at standalone root. + const standaloneNodeModules = path.join(standaloneRoot, "node_modules"); + if (standaloneApp !== standaloneRoot && fs.existsSync(standaloneNodeModules)) { + copyRecursive(standaloneNodeModules, path.join(cliAppDir, "node_modules")); + } +} + +function mergeServerArtifacts(buildDistDir, cliAppDir) { + const serverSrc = path.join(buildDistDir, "server"); + const serverDest = path.join(cliAppDir, buildDistDirName, "server"); + if (!fs.existsSync(serverSrc)) { + throw new Error(`Complete Next.js server build not found: ${serverSrc}`); + } + copyRecursive(serverSrc, serverDest); +} + +function assertRequiredApiArtifacts(cliAppDir) { + const requiredArtifacts = [ + "app/api/v1/chat/completions/route.js", + "app/api/v1/messages/route.js", ]; - const src = candidates.find((p) => fs.existsSync(p)); - if (!src) { - console.warn(`⚠️ ${pkg} not found locally — bundle will rely on node:sqlite or runtime install`); - return; + const serverDir = path.join(cliAppDir, buildDistDirName, "server"); + const missingArtifacts = requiredArtifacts + .map((artifact) => path.join(serverDir, artifact)) + .filter((artifact) => !fs.existsSync(artifact)); + + if (missingArtifacts.length > 0) { + throw new Error( + `Required CLI API route artifact${missingArtifacts.length === 1 ? " is" : "s are"} missing:\n` + + missingArtifacts.join("\n"), + ); } - fs.mkdirSync(path.dirname(dest), { recursive: true }); - copyRecursive(src, dest); - console.log(`✅ Bundled ${pkg}`); -} -ensureModuleInBundle("sql.js"); -const betterDir = path.join(cliAppDir, "node_modules", "better-sqlite3"); -if (fs.existsSync(betterDir)) { - fs.rmSync(betterDir, { recursive: true, force: true }); - console.log("✅ Stripped better-sqlite3 (lives in ~/.9router/runtime)"); -} -console.log(""); - -// Step 4: Copy static files -console.log("4️⃣ Copying static files..."); -const staticSrc = path.join(appDir, ".next", "static"); -const staticSrcResolved = path.join(buildDistDir, "static"); -const staticDest = path.join(cliAppDir, buildDistDirName, "static"); -if (fs.existsSync(staticSrcResolved) || fs.existsSync(staticSrc)) { - copyRecursive(fs.existsSync(staticSrcResolved) ? staticSrcResolved : staticSrc, staticDest); - console.log("✅ Copied static files\n"); -} else { - console.log("⏭️ No static files found\n"); } -// Step 5: Copy public folder if exists -console.log("5️⃣ Copying public folder..."); -const publicSrc = path.join(appDir, "public"); -const publicDest = path.join(cliAppDir, "public"); -if (fs.existsSync(publicSrc)) { - copyRecursive(publicSrc, publicDest); - console.log("✅ Copied public folder\n"); -} else { - console.log("⏭️ No public folder found\n"); +function buildCliPackage() { + console.log("📦 Building 9Router CLI package with Next.js...\n"); + + fs.mkdirSync(buildHomeDir, { recursive: true }); + fs.mkdirSync(path.join(buildHomeDir, "AppData", "Roaming"), { recursive: true }); + fs.mkdirSync(path.join(buildHomeDir, "AppData", "Local"), { recursive: true }); + + // Step 0: Sync version from app/cli/package.json to app/package.json + console.log("0️⃣ Syncing version to app/package.json..."); + const cliPkg = JSON.parse(fs.readFileSync(path.join(cliDir, "package.json"), "utf8")); + const appPkgPath = path.join(appDir, "package.json"); + const appPkg = JSON.parse(fs.readFileSync(appPkgPath, "utf8")); + if (appPkg.version !== cliPkg.version) { + appPkg.version = cliPkg.version; + fs.writeFileSync(appPkgPath, JSON.stringify(appPkg, null, 2) + "\n"); + console.log(`✅ Version synced: ${cliPkg.version}\n`); + } else { + console.log(`✅ Version already synced: ${cliPkg.version}\n`); + } + + // Step 1: Build app with Next.js (workspace tracing root → traced node_modules in standalone). + console.log("1️⃣ Building Next.js app..."); + try { + execSync("npm run build", { + stdio: "inherit", + cwd: appDir, + env: { + ...process.env, + HOME: buildHomeDir, + USERPROFILE: buildHomeDir, + APPDATA: path.join(buildHomeDir, "AppData", "Roaming"), + LOCALAPPDATA: path.join(buildHomeDir, "AppData", "Local"), + NEXT_DIST_DIR: buildDistDirName, + NEXT_TRACING_ROOT_MODE: "workspace", + } + }); + console.log("✅ Next.js build completed\n"); + } catch (error) { + console.error("❌ Next.js build failed"); + process.exit(1); + } + + // Step 2: Clean old app/cli/app if exists + console.log("2️⃣ Cleaning old app/cli/app..."); + if (fs.existsSync(cliAppDir)) { + fs.rmSync(cliAppDir, { recursive: true, force: true }); + } + console.log("✅ Cleaned\n"); + + // Step 3: Copy Next.js standalone build to app/cli/app. + // Newer Next.js standalone output writes server.js/package.json plus .next/, src/, and + // node_modules/ directly under .next/standalone. Older builds may still use a nested app/. + console.log("3️⃣ Copying Next.js standalone build to app/cli/app..."); + try { + copyStandaloneBuild(appDir, buildDistDir, cliAppDir); + } catch (error) { + console.error("❌ Next.js standalone build not found under .next/standalone"); + console.error("Expected either .next/standalone/server.js or .next/standalone/app/"); + process.exit(1); + } + console.log("✅ Copied standalone build\n"); + + // Step 3a: Copy custom server (injects real socket IP, strips spoofable XFF). + const customServerSrc = path.join(appDir, "custom-server.js"); + if (fs.existsSync(customServerSrc)) { + fs.copyFileSync(customServerSrc, path.join(cliAppDir, "custom-server.js")); + console.log("✅ Copied custom-server.js\n"); + } else { + console.warn("⚠️ custom-server.js not found — server will run without real-IP injection\n"); + } + + // Step 3b: Ensure sql.js (pure JS fallback) bundled in app/cli/app/node_modules. + // Strip better-sqlite3 (native) — it lives in ~/.9router/runtime to avoid + // Windows EBUSY during global CLI updates. node:sqlite (Node ≥22.5) is also + // available as a no-install middle tier. + console.log("3️⃣ b Configuring SQLite drivers..."); + function ensureModuleInBundle(pkg) { + const dest = path.join(cliAppDir, "node_modules", pkg); + if (fs.existsSync(dest)) { + console.log(`✅ ${pkg} already bundled`); + return; + } + const candidates = [ + path.join(appDir, "node_modules", pkg), + path.join(rootDir, "node_modules", pkg), + ]; + const src = candidates.find((p) => fs.existsSync(p)); + if (!src) { + console.warn(`⚠️ ${pkg} not found locally — bundle will rely on node:sqlite or runtime install`); + return; + } + fs.mkdirSync(path.dirname(dest), { recursive: true }); + copyRecursive(src, dest); + console.log(`✅ Bundled ${pkg}`); + } + ensureModuleInBundle("sql.js"); + // `open` is external (see serverExternalPackages in next.config.mjs), so it must exist in + // the bundle's node_modules or every importer throws MODULE_NOT_FOUND at runtime. Output + // tracing normally copies it; this is the same belt-and-braces guard used for sql.js. + ensureModuleInBundle("open"); + const betterDir = path.join(cliAppDir, "node_modules", "better-sqlite3"); + if (fs.existsSync(betterDir)) { + fs.rmSync(betterDir, { recursive: true, force: true }); + console.log("✅ Stripped better-sqlite3 (lives in ~/.9router/runtime)"); + } + console.log(""); + + // Step 4: Copy static files + console.log("4️⃣ Copying static files..."); + const staticSrc = path.join(appDir, ".next", "static"); + const staticSrcResolved = path.join(buildDistDir, "static"); + const staticDest = path.join(cliAppDir, buildDistDirName, "static"); + if (fs.existsSync(staticSrcResolved) || fs.existsSync(staticSrc)) { + copyRecursive(fs.existsSync(staticSrcResolved) ? staticSrcResolved : staticSrc, staticDest); + console.log("✅ Copied static files\n"); + } else { + console.log("⏭️ No static files found\n"); + } + + // Step 5: Copy public folder if exists + console.log("5️⃣ Copying public folder..."); + const publicSrc = path.join(appDir, "public"); + const publicDest = path.join(cliAppDir, "public"); + if (fs.existsSync(publicSrc)) { + copyRecursive(publicSrc, publicDest); + console.log("✅ Copied public folder\n"); + } else { + console.log("⏭️ No public folder found\n"); + } + + // Step 6: Copy vendor-chunks (required for production) + console.log("6️⃣ Copying vendor-chunks..."); + const vendorChunksSrc = path.join(appDir, ".next", "server", "vendor-chunks"); + const vendorChunksSrcResolved = path.join(buildDistDir, "server", "vendor-chunks"); + const vendorChunksDest = path.join(cliAppDir, buildDistDirName, "server", "vendor-chunks"); + if (fs.existsSync(vendorChunksSrcResolved) || fs.existsSync(vendorChunksSrc)) { + copyRecursive(fs.existsSync(vendorChunksSrcResolved) ? vendorChunksSrcResolved : vendorChunksSrc, vendorChunksDest); + console.log("✅ Copied vendor-chunks\n"); + } else { + console.log("⏭️ No vendor-chunks found\n"); + } + + // Step 6b: Merge the complete generated server tree. Next.js standalone output + // is trace-pruned and can omit route modules or chunks loaded dynamically. + console.log("6️⃣ b Copying complete server artifacts..."); + mergeServerArtifacts(buildDistDir, cliAppDir); + assertRequiredApiArtifacts(cliAppDir); + console.log("✅ Copied complete server artifacts\n"); + + // Step 7: Copy MITM server files (not bundled by Next.js standalone) + console.log("7️⃣ Copying MITM server files..."); + const mitmSrc = path.join(appDir, "src", "mitm"); + const mitmDest = path.join(cliAppDir, "src", "mitm"); + if (fs.existsSync(mitmSrc)) { + copyRecursive(mitmSrc, mitmDest); + console.log("✅ Copied MITM files\n"); + } else { + console.log("⏭️ No MITM files found\n"); + } + + // Step 7b: Copy standalone updater (headless Node process for install progress) + console.log("7️⃣ b Copying updater files..."); + const updaterSrc = path.join(appDir, "src", "lib", "updater"); + const updaterDest = path.join(cliAppDir, "src", "lib", "updater"); + if (fs.existsSync(updaterSrc)) { + copyRecursive(updaterSrc, updaterDest); + console.log("✅ Copied updater files\n"); + } else { + console.log("⏭️ No updater files found\n"); + } + + // Step 8: Build MITM server (config driven - see app/cli/scripts/buildMitm.js) + console.log("8️⃣ Building MITM server..."); + try { + execSync("node scripts/buildMitm.js", { stdio: "inherit", cwd: cliDir }); + console.log("✅ MITM server build completed\n"); + } catch (error) { + console.error("❌ MITM build failed"); + process.exit(1); + } + + console.log("✨ CLI package build completed!"); + console.log(`📁 Output: ${cliAppDir}`); + + try { + const { execSync: exec } = require("child_process"); + const size = exec(`du -sh "${cliAppDir}"`, { encoding: "utf8" }).trim(); + console.log(`📊 Package size: ${size.split("\t")[0]}`); + } catch (e) { + // Silent fail on size check + } } -// Step 6: Copy vendor-chunks (required for production) -console.log("6️⃣ Copying vendor-chunks..."); -const vendorChunksSrc = path.join(appDir, ".next", "server", "vendor-chunks"); -const vendorChunksSrcResolved = path.join(buildDistDir, "server", "vendor-chunks"); -const vendorChunksDest = path.join(cliAppDir, buildDistDirName, "server", "vendor-chunks"); -if (fs.existsSync(vendorChunksSrcResolved) || fs.existsSync(vendorChunksSrc)) { - copyRecursive(fs.existsSync(vendorChunksSrcResolved) ? vendorChunksSrcResolved : vendorChunksSrc, vendorChunksDest); - console.log("✅ Copied vendor-chunks\n"); -} else { - console.log("⏭️ No vendor-chunks found\n"); -} +module.exports = { + assertRequiredApiArtifacts, + copyStandaloneBuild, + mergeServerArtifacts, +}; -// Step 7: Copy MITM server files (not bundled by Next.js standalone) -console.log("7️⃣ Copying MITM server files..."); -const mitmSrc = path.join(appDir, "src", "mitm"); -const mitmDest = path.join(cliAppDir, "src", "mitm"); -if (fs.existsSync(mitmSrc)) { - copyRecursive(mitmSrc, mitmDest); - console.log("✅ Copied MITM files\n"); -} else { - console.log("⏭️ No MITM files found\n"); -} - -// Step 7b: Copy standalone updater (headless Node process for install progress) -console.log("7️⃣ b Copying updater files..."); -const updaterSrc = path.join(appDir, "src", "lib", "updater"); -const updaterDest = path.join(cliAppDir, "src", "lib", "updater"); -if (fs.existsSync(updaterSrc)) { - copyRecursive(updaterSrc, updaterDest); - console.log("✅ Copied updater files\n"); -} else { - console.log("⏭️ No updater files found\n"); -} - -// Step 8: Build MITM server (config driven - see app/cli/scripts/buildMitm.js) -console.log("8️⃣ Building MITM server..."); -try { - execSync("node scripts/buildMitm.js", { stdio: "inherit", cwd: cliDir }); - console.log("✅ MITM server build completed\n"); -} catch (error) { - console.error("❌ MITM build failed"); - process.exit(1); -} - -console.log("✨ CLI package build completed!"); -console.log(`📁 Output: ${cliAppDir}`); - -try { - const { execSync: exec } = require("child_process"); - const size = exec(`du -sh "${cliAppDir}"`, { encoding: "utf8" }).trim(); - console.log(`📊 Package size: ${size.split("\t")[0]}`); -} catch (e) { - // Silent fail on size check +if (require.main === module) { + buildCliPackage(); } diff --git a/custom-server.js b/custom-server.js index 6e39683f..cf824ba4 100644 --- a/custom-server.js +++ b/custom-server.js @@ -1,7 +1,42 @@ const http = require("http"); +const path = require("path"); +const { pathToFileURL } = require("url"); const origCreate = http.createServer.bind(http); +let backgroundRefreshStarted = false; + +function startBackgroundTokenRefreshFromCustomServer() { + if (backgroundRefreshStarted) return; + backgroundRefreshStarted = true; + // Prefer source path (repo / standalone that still has src). Fail-open if missing + // — initializeApp also starts the same scheduler when the Next app boots. + const modPath = path.join(__dirname, "src", "sse", "services", "backgroundTokenRefresh.js"); + import(pathToFileURL(modPath).href) + .then((m) => { + try { + m.startBackgroundTokenRefresh(); + } catch (e) { + console.error("[BackgroundTokenRefresh] start failed:", e && e.message ? e.message : e); + } + const stop = () => { + try { + m.stopBackgroundTokenRefresh(); + } catch { + /* ignore */ + } + }; + process.once("SIGINT", stop); + process.once("SIGTERM", stop); + }) + .catch((e) => { + // Expected in published CLI standalone (src/ not on disk). App bootstrap covers it. + if (process.env.DEBUG_BACKGROUND_TOKEN_REFRESH) { + console.error("[BackgroundTokenRefresh] import failed:", e && e.message ? e.message : e); + } + }); +} + // Wrap Next standalone HTTP server: derive client IP from the TCP socket // (unspoofable) and strip client-supplied forwarding headers so downstream // rate-limiting keys on the real peer address instead of attacker-controlled XFF. @@ -26,7 +61,57 @@ http.createServer = (...args) => { if (viaProxy) req.headers["x-9r-via-proxy"] = "1"; return handler(req, res); }; - return origCreate(...rest, wrapped); + const server = origCreate(...rest, wrapped); + server.once("listening", () => { + startBackgroundTokenRefreshFromCustomServer(); + }); + const origEmit = server.emit; + // JBR 25 sends h2c upgrades that the HTTP/1.1 server would otherwise close. + server.emit = function (event, ...eventArgs) { + const [req, socket, head] = eventArgs; + if (event !== "upgrade" || String(req.headers.upgrade || "").toLowerCase() !== "h2c") { + return origEmit.call(this, event, ...eventArgs); + } + + const contentLength = Number(req.headers["content-length"] || 0); + if (!Number.isSafeInteger(contentLength) || contentLength < 0) { + socket.destroy(); + return true; + } + const chunks = [head]; + let received = head.length; + const serve = () => { + // Replay the upgraded request through the existing HTTP/1.1 handler. + const replay = new http.IncomingMessage(socket); + Object.assign(replay, { method: req.method, url: req.url, headers: req.headers, complete: true }); + if (received) replay.push(Buffer.concat(chunks, received).subarray(0, contentLength)); + replay.push(null); + const res = new http.ServerResponse(replay); + res.shouldKeepAlive = false; + res.assignSocket(socket); + res.once("finish", () => socket.end()); + Promise.resolve().then(() => wrapped(replay, res)).catch((error) => { + console.error("Failed to downgrade h2c request", error); + socket.destroy(); + }); + }; + if (received >= contentLength) serve(); + else { + socket.on("data", function readBody(chunk) { + chunks.push(chunk); + received += chunk.length; + if (received < contentLength) return; + socket.off("data", readBody); + serve(); + }); + socket.resume(); + } + delete req.headers.upgrade; + delete req.headers["http2-settings"]; + req.headers.connection = "close"; + return true; + }; + return server; }; -require("./server.js"); +if (require.main === module) require("./server.js"); diff --git a/docs/superpowers/plans/2026-08-02-gpt-5-6-codex-reasoning-overrides.md b/docs/superpowers/plans/2026-08-02-gpt-5-6-codex-reasoning-overrides.md new file mode 100644 index 00000000..8ff5a940 --- /dev/null +++ b/docs/superpowers/plans/2026-08-02-gpt-5-6-codex-reasoning-overrides.md @@ -0,0 +1,328 @@ +# GPT-5.6 Codex Reasoning Overrides Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Preserve Codex-advertised Max and Ultra overrides for GPT-5.6 Sol and Terra, preserve Max for Luna, and convert Luna Ultra to Max without changing Kiro or generic OpenAI-format behavior. + +**Architecture:** Keep the supported reasoning matrix in the existing `getThinkingLevels(provider, model)` resolver and reuse that result in both translation and Codex executor normalization. The dashboard already consumes this resolver, so no UI component change is required. Unsupported top-end levels remain safely normalized, with Luna Ultra selecting Luna's supported Max level. + +**Tech Stack:** JavaScript ES modules, Next.js, Vitest, Codex Responses transport. + +## Global Constraints + +- Apply the new overrides only to the OpenAI Codex provider (`codex`, exposed as `cx/`). +- Sol and Terra support `max` and `ultra`; Luna supports `max` but not `ultra`. +- Convert Luna `ultra` requests to `max` in both translated and native passthrough request paths. +- Preserve existing Kiro and generic OpenAI-compatible normalization. +- Do not add runtime model-catalog fetching, dependencies, pricing changes, or unrelated refactors. +- Write each behavior test first and observe the expected failure before changing production code. + +--- + +### Task 1: Provider-scoped GPT-5.6 level matrix + +**Files:** +- Modify: `tests/unit/thinking-levels-gpt56-sol.test.js` +- Modify: `open-sse/providers/thinkingLevels.js` + +**Interfaces:** +- Consumes: `getThinkingLevels(provider, model)` and existing capability metadata. +- Produces: `getThinkingLevels(provider, model): string[] | null` with Codex-only GPT-5.6 level overrides. + +- [ ] **Step 1: Replace the Sol-only assertions with the complete behavior matrix** + +Use literal expected arrays so each model/provider contract is independently checked: + +```js +it.each([ + ["gpt-5.6-sol", ["none", "minimal", "low", "medium", "high", "xhigh", "max", "ultra"]], + ["gpt-5.6-terra", ["none", "minimal", "low", "medium", "high", "xhigh", "max", "ultra"]], + ["gpt-5.6-luna", ["none", "minimal", "low", "medium", "high", "xhigh", "max"]], + ["gpt-5.6-sol-review", ["none", "minimal", "low", "medium", "high", "xhigh", "max", "ultra"]], + ["gpt-5.6-terra-review", ["none", "minimal", "low", "medium", "high", "xhigh", "max", "ultra"]], + ["gpt-5.6-luna-review", ["none", "minimal", "low", "medium", "high", "xhigh", "max"]], +])("returns Codex levels for %s", (model, expected) => { + expect(getThinkingLevels("codex", model)).toEqual(expected); +}); + +it("does not expose Codex-only GPT-5.6 overrides on Kiro", () => { + expect(getThinkingLevels("kiro", "gpt-5.6-sol")).toEqual([ + "none", "minimal", "low", "medium", "high", "xhigh", + ]); +}); +``` + +Keep the older Codex-model assertion to protect the existing `gpt-5.3-codex` behavior. + +- [ ] **Step 2: Run the level test and verify it fails for the missing matrix/provider scoping** + +Run: + +```bash +npx vitest run tests/unit/thinking-levels-gpt56-sol.test.js +``` + +Expected: FAIL because Sol lacks Ultra, Terra/Luna lack Max, and Kiro currently inherits Sol Max. + +- [ ] **Step 3: Add provider-aware pattern matching and the three Codex model rules** + +Update `PATTERN_THINKING` entries to accept an optional `provider` field and match it in `getThinkingLevels`: + +```js +const CODEX_GPT_5_6_LEVELS = ["none", "minimal", "low", "medium", "high", "xhigh", "max"]; + +const PATTERN_THINKING = [ + { provider: "codex", pattern: "*gpt-5.6-sol*", levels: [...CODEX_GPT_5_6_LEVELS, "ultra"] }, + { provider: "codex", pattern: "*gpt-5.6-terra*", levels: [...CODEX_GPT_5_6_LEVELS, "ultra"] }, + { provider: "codex", pattern: "*gpt-5.6-luna*", levels: CODEX_GPT_5_6_LEVELS }, + { pattern: "*codex*", levels: ["low", "medium", "high", "xhigh"] }, +]; + +const hit = PATTERN_THINKING.find((entry) => + (!entry.provider || entry.provider === provider) && matchPattern(entry.pattern, model) +); +``` + +- [ ] **Step 4: Re-run the level test and verify it passes** + +Run: + +```bash +npx vitest run tests/unit/thinking-levels-gpt56-sol.test.js +``` + +Expected: 1 test file passed with no failures. + +- [ ] **Step 5: Commit the capability matrix** + +```bash +git add open-sse/providers/thinkingLevels.js tests/unit/thinking-levels-gpt56-sol.test.js +git commit -m "feat(codex): expose GPT-5.6 reasoning overrides" +``` + +### Task 2: Model-aware shared thinking translation + +**Files:** +- Modify: `tests/translator/thinking-unified.test.js` +- Modify: `open-sse/translator/concerns/thinkingUnified.js` + +**Interfaces:** +- Consumes: `getThinkingLevels(provider, cleanModel): string[] | null` from Task 1. +- Produces: `parseSuffix(model)` support for `ultra` and `applyThinking(...)` output that preserves supported Codex levels. + +- [ ] **Step 1: Add failing suffix and translation tests** + +Add a literal parser assertion: + +```js +expect(parseSuffix("gpt-5.6-sol(ultra)")).toEqual({ + cleanModel: "gpt-5.6-sol", + override: { mode: "level", level: "ultra" }, +}); +``` + +Add table-driven Codex assertions using direct request fields: + +```js +it.each([ + ["gpt-5.6-sol", "max", "max"], + ["gpt-5.6-sol", "ultra", "ultra"], + ["gpt-5.6-terra", "max", "max"], + ["gpt-5.6-terra", "ultra", "ultra"], + ["gpt-5.6-luna", "max", "max"], + ["gpt-5.6-luna", "ultra", "max"], +])("normalizes Codex %s effort %s to %s", (model, effort, expected) => { + const out = apply("openai-responses", model, { reasoning: { effort } }, "codex"); + expect(out.reasoning_effort).toBe(expected); +}); +``` + +Add a parenthesized override assertion and Kiro isolation assertion: + +```js +expect(apply("openai-responses", "gpt-5.6-sol(ultra)", {}, "codex").reasoning_effort).toBe("ultra"); +expect(apply("openai", "gpt-5.6-sol", { reasoning_effort: "max" }, "kiro").reasoning_effort).toBe("xhigh"); +``` + +- [ ] **Step 2: Run the translator test and verify it fails for Ultra parsing and preserved Max/Ultra** + +Run: + +```bash +npx vitest run tests/translator/thinking-unified.test.js +``` + +Expected: FAIL because Ultra suffixes are ignored and OpenAI translation clamps Max to XHigh. + +- [ ] **Step 3: Implement supported-level normalization in the shared translator** + +Import `getThinkingLevels`. Recognize `ultra` explicitly in `parseSuffix` without adding it to the budget map. Resolve supported levels once in `applyThinking` and pass them to `applyFormat`. + +Use this normalization rule for the OpenAI format: + +```js +function normalizeOpenAILevel(level, supportedLevels) { + if (level !== "max" && level !== "ultra") return level; + if (supportedLevels?.includes(level)) return level; + if (level === "ultra" && supportedLevels?.includes("max")) return "max"; + return "xhigh"; +} +``` + +Keep `none`, automatic effort, budget conversion, and every non-OpenAI format unchanged. + +- [ ] **Step 4: Re-run the translator and generic OpenAI clamp tests** + +Run: + +```bash +npx vitest run tests/translator/thinking-unified.test.js tests/unit/thinking-effort-openai-max-clamp.test.js +``` + +Expected: 2 test files passed; generic OpenAI Max still becomes XHigh. + +- [ ] **Step 5: Commit shared translation support** + +```bash +git add open-sse/translator/concerns/thinkingUnified.js tests/translator/thinking-unified.test.js +git commit -m "feat(codex): preserve supported reasoning efforts" +``` + +### Task 3: Codex native passthrough normalization + +**Files:** +- Modify: `tests/unit/codex-fast-capacity.test.js` +- Modify: `open-sse/executors/codex.js` + +**Interfaces:** +- Consumes: `getThinkingLevels("codex", upstreamModel): string[] | null` from Task 1. +- Produces: `CodexExecutor.transformRequest(...)` payloads with model-supported upstream `reasoning.effort` values. + +- [ ] **Step 1: Add failing Codex executor behavior tests** + +Add a separate `describe("Codex reasoning normalization", ...)` block with real `transformRequest` calls: + +```js +it.each([ + ["gpt-5.6-sol", "max", "max"], + ["gpt-5.6-sol", "ultra", "ultra"], + ["gpt-5.6-terra", "max", "max"], + ["gpt-5.6-terra", "ultra", "ultra"], + ["gpt-5.6-luna", "max", "max"], + ["gpt-5.6-luna", "ultra", "max"], +])("normalizes %s effort %s to %s", (model, effort, expected) => { + const body = new CodexExecutor().transformRequest(model, { + model, + input: "hi", + reasoning: { effort }, + }, true, {}); + expect(body.reasoning.effort).toBe(expected); +}); + +it("resolves review models before applying the reasoning matrix", () => { + const body = new CodexExecutor().transformRequest("gpt-5.6-terra-review", { + model: "gpt-5.6-terra-review", + input: "hi", + reasoning_effort: "ultra", + }, true, {}); + expect(body.model).toBe("gpt-5.6-terra"); + expect(body.reasoning.effort).toBe("ultra"); +}); +``` + +Keep the existing GPT-5.5 Max-to-XHigh fast-tier test. + +- [ ] **Step 2: Run the executor test and verify supported values fail by being clamped** + +Run: + +```bash +npx vitest run tests/unit/codex-fast-capacity.test.js +``` + +Expected: FAIL because current normalization maps supported Max to XHigh and does not map Luna Ultra to Max. + +- [ ] **Step 3: Make Codex normalization model-aware** + +Import `getThinkingLevels` and replace the global Max clamp with: + +```js +function normalizeReasoningEffort(model, value) { + const supportedLevels = getThinkingLevels("codex", model); + if (supportedLevels?.includes(value)) return value; + if (value === "ultra" && supportedLevels?.includes("max")) return "max"; + if (value === "max" || value === "ultra") return "xhigh"; + return value; +} +``` + +Call it only after `body.model` has resolved review aliases to their upstream base model. Pass `body.model` for both `reasoning_effort` and existing `reasoning.effort` request shapes. + +- [ ] **Step 4: Re-run the executor and focused feature suites** + +Run: + +```bash +npx vitest run tests/unit/codex-fast-capacity.test.js tests/unit/thinking-levels-gpt56-sol.test.js tests/translator/thinking-unified.test.js tests/unit/thinking-effort-openai-max-clamp.test.js +``` + +Expected: 4 test files passed with no failures. + +- [ ] **Step 5: Commit native Codex normalization** + +```bash +git add open-sse/executors/codex.js tests/unit/codex-fast-capacity.test.js +git commit -m "feat(codex): forward GPT-5.6 max and ultra efforts" +``` + +### Task 4: Full verification and pull request + +**Files:** +- Verify all changed production, test, design, and plan files. + +**Interfaces:** +- Consumes: completed Tasks 1-3. +- Produces: verified branch pushed to `origin` and a pull request targeting `decolua/9router:master`. + +- [ ] **Step 1: Run all focused regression tests** + +```bash +npx vitest run tests/unit/thinking-levels-gpt56-sol.test.js tests/translator/thinking-unified.test.js tests/unit/thinking-effort-openai-max-clamp.test.js tests/unit/codex-fast-capacity.test.js +``` + +Expected: all selected test files and tests pass. + +- [ ] **Step 2: Run the complete unit test suite** + +```bash +npx vitest run tests/unit tests/translator +``` + +Expected: all test files pass with zero failed tests. + +- [ ] **Step 3: Run the production build** + +```bash +npm run build +``` + +Expected: Next.js production build exits with status 0. + +- [ ] **Step 4: Verify repository hygiene and requirement coverage** + +```bash +git diff --check upstream/master...HEAD +git status --short --branch +git log --oneline upstream/master..HEAD +``` + +Expected: no whitespace errors, no uncommitted source changes, and only scoped feature commits. + +- [ ] **Step 5: Push the feature branch and open the pull request** + +```bash +git push -u origin codex/gpt-5-6-reasoning-overrides +gh pr create --repo decolua/9router --base master --head seakleangnhak:codex/gpt-5-6-reasoning-overrides --title "feat(codex): support GPT-5.6 Max and Ultra overrides" --body $'## Summary\n- expose Max and Ultra for Codex GPT-5.6 Sol and Terra\n- expose Max for Codex GPT-5.6 Luna and normalize Luna Ultra to Max\n- keep Kiro and generic OpenAI-compatible reasoning behavior unchanged\n\n## Verification\n- `npx vitest run tests/unit tests/translator`\n- `npm run build`' +``` + +The pull request body must summarize the Codex-only support matrix, Luna Ultra-to-Max fallback, Kiro isolation, and fresh test/build evidence. diff --git a/docs/superpowers/specs/2026-08-02-gpt-5-6-codex-reasoning-overrides-design.md b/docs/superpowers/specs/2026-08-02-gpt-5-6-codex-reasoning-overrides-design.md new file mode 100644 index 00000000..832fb5ce --- /dev/null +++ b/docs/superpowers/specs/2026-08-02-gpt-5-6-codex-reasoning-overrides-design.md @@ -0,0 +1,122 @@ +# GPT-5.6 Codex Reasoning Overrides Design + +## Goal + +Expose and preserve the reasoning levels currently advertised by the OpenAI +Codex model catalog for GPT-5.6 Sol, Terra, and Luna when they are routed +through the `codex` provider (`cx/`). + +The supported override matrix is: + +| Model family | Max | Ultra | +| --- | --- | --- | +| GPT-5.6 Sol | Yes | Yes | +| GPT-5.6 Terra | Yes | Yes | +| GPT-5.6 Luna | Yes | No | + +The same matrix applies to 9router's virtual `-review` variants because they +resolve to the corresponding upstream base model. + +## Scope + +This change is limited to OpenAI Codex (`cx/`) routes. Kiro (`kr/`) and other +OpenAI-format providers retain their existing reasoning-level behavior even +when they expose models with the same GPT-5.6 names. + +The change covers the complete local request path: + +1. The provider page advertises only the levels supported by each Codex model. +2. A copied model suffix such as `gpt-5.6-sol(ultra)` is parsed as a reasoning + override. +3. The shared thinking translator preserves a supported Codex override while + retaining the existing `xhigh` fallback for unsupported OpenAI levels. +4. The Codex executor sends supported `max` and `ultra` values unchanged to the + upstream Codex Responses endpoint. + +## Current Behavior + +`gpt-5.6-luna` and the other GPT-5.6 models already exist in the Codex model +registry. The capability picker has a global Sol-only `max` pattern, which also +affects providers such as Kiro unintentionally. The shared OpenAI translator +and Codex executor then convert `max` to `xhigh`, so the advertised override is +not preserved end to end. `ultra` is not recognized as a model suffix. + +## Design + +### Provider-scoped level resolution + +Extend the existing model-pattern overrides in +`open-sse/providers/thinkingLevels.js` with an optional provider constraint. +Add three Codex-only GPT-5.6 patterns in most-specific order: + +- Sol: existing levels plus `max` and `ultra`. +- Terra: existing levels plus `max` and `ultra`. +- Luna: existing levels plus `max`. + +Matching remains wildcard-based so virtual `-review` variants inherit the +base model's levels. Provider matching prevents these overrides from changing +Kiro or other providers. + +### Shared translation + +Teach the suffix parser to recognize `ultra` as a discrete level without +assigning it a synthetic token budget. When applying the OpenAI wire format, +reuse the resolved per-provider model levels: + +- Preserve `max` or `ultra` when the target provider/model explicitly supports + the requested level. +- Convert `ultra` to `max` for GPT-5.6 Luna, preserving the highest level Luna + supports. +- Convert other unsupported `max` or `ultra` requests to `xhigh`, preserving + the existing safe fallback for generic OpenAI-compatible providers. +- Leave all existing lower levels and `none` handling unchanged. + +This keeps one capability source for the dashboard and translation behavior +instead of duplicating the GPT-5.6 matrix. + +### Codex executor + +Make Codex reasoning normalization model-aware. After virtual review models +are resolved to their upstream base model, preserve a requested level when +the Codex capability resolver lists it. Continue converting unsupported +`max` or `ultra` values to `xhigh`, except that Luna converts `ultra` to its +supported `max` level. + +Do not add `max` to the executor's legacy hyphen-suffix parser because +`gpt-5.1-codex-max` is an actual model identifier. Dashboard overrides use the +existing parenthesized suffix and the shared translator removes that suffix +before executor dispatch. + +## Error and Compatibility Behavior + +- `cx/gpt-5.6-luna(ultra)` becomes `max` rather than sending an unsupported + level upstream. +- Non-GPT-5.6 Codex models retain their current supported levels and fallback + behavior. +- Kiro GPT-5.6 routes no longer inherit the Codex Sol-only picker override and + continue using Kiro's existing effort normalization. +- Direct request fields and parenthesized model overrides follow the same + model-aware rules. + +## Testing + +Use test-driven development with focused unit coverage: + +1. Level resolver tests for Sol, Terra, Luna, their review variants, an older + Codex model, and Kiro isolation. +2. Shared translator tests proving `max` and `ultra` survive only for supported + Codex model/provider combinations, Luna `ultra` becomes `max`, and other + unsupported combinations become `xhigh`. +3. Codex executor tests proving native and translated request shapes preserve + supported values after upstream model resolution. +4. Existing thinking translation and Codex executor suites to guard generic + OpenAI clamping and fast-tier behavior. +5. Project lint/build checks in proportion to the changed JavaScript modules. + +## Non-goals + +- Runtime fetching or caching of the Codex model catalog. +- Adding these levels to Kiro or another provider. +- Changing model pricing, quotas, defaults, or service tiers. +- Adding Codex Ultra's multi-agent orchestration behavior inside 9router; + 9router only forwards the catalog-advertised reasoning override. diff --git a/i18n/README.vi.md b/i18n/README.vi.md index d0a51de4..5e01d8aa 100644 --- a/i18n/README.vi.md +++ b/i18n/README.vi.md @@ -1,21 +1,15 @@ -Dưới đây là bản dịch tiếng Việt của tài liệu Markdown, giữ nguyên toàn bộ cú pháp và cấu trúc kỹ thuật. -
Bảng điều khiển 9Router - # 9Router - Free AI Router + # 9Router - Free AI Router & Token Saver - **Không bao giờ ngừng code. Tự động định tuyến tới các mô hình AI MIỄN PHÍ & giá rẻ với cơ chế dự phòng thông minh.** + **Không bao giờ ngừng code. Tiết kiệm 20-40% token với RTK + tự động dự phòng sang các mô hình AI MIỄN PHÍ & giá rẻ.** - **Nhà cung cấp AI Miễn cho OpenClaw.** - -

- OpenClaw -

+ **Kết nối tất cả công cụ AI Code (Claude Code, Codex, Cursor, Cline, Copilot, Antigravity...) tới 40+ Nhà cung cấp AI & 100+ Mô hình.** [![npm](https://img.shields.io/npm/v/9router.svg)](https://www.npmjs.com/package/9router) [![Downloads](https://img.shields.io/npm/dm/9router.svg)](https://www.npmjs.com/package/9router) - [![License](https://img.shields.io/npm/l/9router.svg)](https://github.com/decolua/9router/blob/main/LICENSE) + [![License](https://github.com/decolua/9router/blob/main/LICENSE)](https://github.com/decolua/9router/blob/main/LICENSE) [🚀 Bắt đầu nhanh](#-quick-start) • [💡 Tính năng](#-key-features) • [📖 Cài đặt](#-setup-guide) • [🌐 Website](https://9router.com)
@@ -24,19 +18,21 @@ Dưới đây là bản dịch tiếng Việt của tài liệu Markdown, giữ ## 🤔 Tại sao chọn 9Router? -**Ngừng lãng phí tiền bạc và gặp phải giới hạn:** +**Ngừng lãng phí tiền bạc, token và không bao giờ lo chạm giới hạn (rate limit):** - ❌ Hạn mức gói đăng ký hết hạn mỗi tháng mà không dùng hết -- ❌ Giới hạn tốc độ (rate limit) ngăn bạn giữaừng khi code -- ❌ Các API đắt đỏ ($20-50/tháng cho mỗi nhà cung cấp) -- ❌ Phải chuyển đổi thủ công giữa các nhà cung cấp +- ❌ Giới hạn tốc độ (rate limit) làm gián đoạn công việc mid-coding +- ❌ Kết quả của công cụ (git diff, grep, ls...) ngốn rất nhiều token +- ❌ Chi phí API đắt đỏ ($20-50/tháng cho từng nhà cung cấp) +- ❌ Phải chuyển đổi thủ công giữa các nhà cung cấp AI **9Router giải quyết vấn đề này:** -- ✅ **Tối đa hóa gói đăng ký** - Theo dõi hạn mức, sử dụng từng bit trước khi reset -- ✅ **Tự động dự phòng** - Gói đăng ký → Giá rẻ → Miễn phí, thời gian chết bằng không -- ✅ **Đa tài khoản** - Vòng tròn (round-robin) các tài khoản của mỗi nhà cung cấp -- ✅ **Phổ quát** - Hoạt động với Claude Code, Codex, Gemini CLI, Cursor, Cline, bất kỳ công cụ CLI nào +- ✅ **RTK Token Saver** - Tự động nén nội dung `tool_result`, tiết kiệm 20-40% token trên mỗi request +- ✅ **Tối đa hóa gói đăng ký** - Theo dõi hạn mức, tận dụng triệt để trước khi reset +- ✅ **Tự động dự phòng (Auto Fallback)** - Gói đăng ký → Giá rẻ → Miễn phí, không lo downtime +- ✅ **Đa tài khoản (Multi-account)** - Xoay vòng (round-robin) các tài khoản cho mỗi nhà cung cấp +- ✅ **Phổ quát (Universal)** - Hoạt động với Claude Code, Codex, Cursor, Cline, Antigravity và mọi công cụ CLI --- @@ -44,25 +40,26 @@ Dưới đây là bản dịch tiếng Việt của tài liệu Markdown, giữ ``` ┌─────────────┐ -│ Your CLI │ (Claude Code, Codex, Gemini CLI, OpenClaw, Cursor, Cline...) -│ Tool │ +│ Công cụ │ (Claude Code, Codex, OpenClaw, Cursor, Cline, Antigravity...) +│ CLI AI │ └──────┬──────┘ │ http://localhost:20128/v1 ↓ -┌────────────────────────────────────────┐ -│ 9Router (Smart Router) │ -│ • Format translation (OpenAI ↔ Claude) │ -│ • Quota tracking │ -│ • Auto token refresh │ -└──────┬──────────────────────────────────┘ +┌─────────────────────────────────────────────┐ +│ 9Router (Smart Router) │ +│ • RTK Token Saver (nén tool_result token) │ +│ • Dịch chuyển định dạng (OpenAI ↔ Claude) │ +│ • Quota tracking (theo dõi hạn mức) │ +│ • Tự động làm mới OAuth Token │ +└──────┬──────────────────────────────────────┘ │ - ├─→ [Tier 1: SUBSCRIPTION] Claude Code, Codex, Gemini CLI - │ ↓ quota exhausted - ├─→ [Tier 2: CHEAP] GLM ($0.6/1M), MiniMax ($0.2/1M) - │ budget limit - └─→ [Tier 3: FREE] iFlow, Qwen, Kiro (unlimited) + ├─→ [Tier 1: GÓI ĐĂNG KÝ] Claude Code, Codex, GitHub Copilot + │ ↓ hết hạn mức quota + ├─→ [Tier 2: GIÁ RẺ] GLM ($0.6/1M), MiniMax ($0.2/1M) + │ ↓ chạm ngân sách + └─→ [Tier 3: MIỄN PHÍ] Kiro AI, OpenCode Free, Vertex AI ($300 credits) -Result: Never stop coding, minimal cost +Kết quả: Không bao giờ ngừng code, chi phí tối thiểu + tiết kiệm 20-40% token qua RTK ``` --- @@ -76,26 +73,26 @@ npm install -g 9router 9router ``` -🎉 Bảng điều khiển mở tại `http://localhost:20128` +🎉 Bảng điều khiển (Dashboard) sẽ tự động mở tại `http://localhost:20128` **2. Kết nối nhà cung cấp MIỄN PHÍ (không cần đăng ký):** -Bảng điều khiển → Providers -> Kết nối **ude Code** hoặc **Antigravity** -> Đăng nhập OAuth -> Xong! +Bảng điều khiển → Providers → Kết nối **Kiro AI** (~50 credits/tháng miễn phí: Claude 4.5 + GLM-5 + MiniMax) hoặc **OpenCode Free** (không cần auth) → Xong! **3. Sử dụng trong công cụ CLI của bạn:** ``` -Cài đặt Claude Code/Codex/Gemini CLI/OpenClaw/Cursor/Cline: +Cài đặt Claude Code/Codex/OpenClaw/Cursor/Cline/Antigravity: Endpoint: http://localhost:20128/v1 API Key: [sao chép từ bảng điều khiển] - Model: if/kimi-k2-thinking + Model: kr/claude-sonnet-4.5 ``` -**Xong rồi!** Bắt đầu code với các mô hình AI MIỄN PHÍ. +**Thế là xong!** Bắt đầu code ngay với các mô hình AI MIỄN PHÍ. -**Phương án khác: chạy từ nguồn (k lưu trữ này):** +**Phương án khác: chạy từ mã nguồn (repository này):** -Gói kho lưu trữ này là riêng tư (`9router-app`), vì vậy việc thực thi nguồn/Docker là đường dẫn phát triển cục bộ dự kiến. +Gói kho lưu trữ này là riêng tư (`9router-app`), vì vậy việc chạy từ nguồn/Docker là cách phát triển cục bộ mặc định. ```bash cp .env.example .env @@ -111,11 +108,12 @@ PORT=20128 HOSTNAME=0.0.0.0 NEXT_PUBLIC_BASE_URL=http://localhost:20128 npm run ``` URL mặc định: -- Bảng điều khiển: `http://localhost:20128/dashboard` +- Bảng điều khiển Dashboard: `http://localhost:20128/dashboard` - API tương thích OpenAI: `http://localhost:20128/v1` --- + ## 🎥 Hướng dẫn Video
diff --git a/i18n/README.zh-CN.md b/i18n/README.zh-CN.md index 92e0624b..c090f608 100644 --- a/i18n/README.zh-CN.md +++ b/i18n/README.zh-CN.md @@ -42,25 +42,26 @@ ``` ┌─────────────┐ -│ Your CLI │ (Claude Code, Codex, Gemini CLI, OpenClaw, Cursor, Cline...) +│ Your CLI │ (Claude Code, Codex, OpenClaw, Cursor, Cline, Antigravity...) │ Tool │ └──────┬──────┘ - │ http://localhost:201281 + │ http://localhost:20128/v1 ↓ -┌─────────────────────────────────────────┐ -│ 9Router (Smart Router) │ -│ • Format translation (OpenAI ↔ Claude) │ -│ • Quota tracking │ -│ • Auto token refresh │ -└──────┬──────────────────────────────────┘ +┌─────────────────────────────────────────────┐ +│ 9Router (Smart Router) │ +│ • RTK Token Saver (节省 20-40% Token) │ +│ • 格式转换 (OpenAI ↔ Claude) │ +│ • 配额追踪 (Quota tracking) │ +│ • 自动刷新 OAuth Token │ +└──────┬──────────────────────────────────────┘ │ - ├─→ [Tier 1: SUBSCRIPTION] Claude Code, Codex, Gemini CLI - │ ↓ quota exhausted - ├─→ [Tier 2: CHEAP] GLM ($0.6/1M), MiniMax ($0.2/1M) - │ ↓ budget limit - └─→ [Tier 3: FREE] iFlow, Qwen, Kiro (unlimited) + ├─→ [Tier 1: 订阅] Claude Code, Codex, GitHub Copilot + │ ↓ 配额用尽 + ├─→ [Tier 2: 低价] GLM ($0.6/1M), MiniMax ($0.2/1M) + │ ↓ 触及预算上限 + └─→ [Tier 3: 免费] Kiro AI, OpenCode Free, Vertex AI ($300 credits) -Result: Never stop coding, minimal cost +结果:永不停歇的编程体验,最低成本 + 通过 RTK 节省 20-40% Token ``` --- diff --git a/next.config.mjs b/next.config.mjs index d017c157..ecd385ca 100644 --- a/next.config.mjs +++ b/next.config.mjs @@ -13,7 +13,14 @@ const proxyClientMaxBodySize = process.env.NINEROUTER_PROXY_CLIENT_MAX_BODY_SIZE const nextConfig = { distDir: process.env.NEXT_DIST_DIR || ".next", output: "standalone", - serverExternalPackages: ["better-sqlite3", "sql.js", "node:sqlite", "bun:sqlite"], + // `open` must stay external. It derives its own directory from `import.meta.url`, and + // webpack replaces that with the absolute path of the BUILD machine as a string literal. + // A release built on macOS therefore ships `file:///Users/.../open/index.js`, which + // `fileURLToPath` rejects on Windows ("File URL path must be absolute" — no drive + // letter). That throw happens at module scope, so every consumer of `open` dies on + // import — including xAI/Grok token refresh, which loads the OAuth service that imports + // it. Keeping it external preserves the real `import.meta.url` at runtime. + serverExternalPackages: ["better-sqlite3", "sql.js", "node:sqlite", "bun:sqlite", "open"], turbopack: { root: tracingRoot }, diff --git a/open-sse/config/appConstants.js b/open-sse/config/appConstants.js index 5e3ec4be..5ddf95ac 100644 --- a/open-sse/config/appConstants.js +++ b/open-sse/config/appConstants.js @@ -156,6 +156,13 @@ export const LOAD_CODE_ASSIST_HEADERS = { "Client-Metadata": JSON.stringify({ ideType: IDE_TYPE.ANTIGRAVITY, platform: getPlatformEnum(), pluginType: PLUGIN_TYPE.GEMINI }), }; +// Real Antigravity IDE doesn't send X-Goog-Api-Client/Client-Metadata on loadCodeAssist/onboardUser — +// Google's backend fingerprints those and silently refuses to provision a cloudaicompanionProject. +export const ANTIGRAVITY_LOAD_CODE_ASSIST_HEADERS = { + "Content-Type": "application/json", + "User-Agent": ANTIGRAVITY_IDE_USER_AGENT, +}; + export const LOAD_CODE_ASSIST_METADATA = { ideType: IDE_TYPE.ANTIGRAVITY, platform: getPlatformEnum(), @@ -176,7 +183,6 @@ export const OAUTH_ENDPOINTS = { google: { token: "https://oauth2.googleapis.com/token", auth: "https://accounts.google.com/o/oauth2/auth" }, openai: { token: PROVIDER_OAUTH["codex"]?.tokenUrl, auth: PROVIDER_OAUTH["codex"]?.authorizeUrl }, anthropic: { token: PROVIDER_OAUTH["claude"]?.tokenUrl, auth: "https://api.anthropic.com/v1/oauth/authorize" }, // ≠ claude.authorizeUrl (claude.ai login) — keep - qwen: { token: PROVIDER_OAUTH["qwen"]?.tokenUrl, auth: PROVIDER_OAUTH["qwen"]?.deviceCodeUrl }, iflow: { token: PROVIDER_OAUTH["iflow"]?.tokenUrl, auth: PROVIDER_OAUTH["iflow"]?.authorizeUrl }, github: { token: PROVIDER_OAUTH["github"]?.tokenUrl, auth: PROVIDER_OAUTH["github"]?.authorizeUrl, deviceCode: PROVIDER_OAUTH["github"]?.deviceCodeUrl }, }; diff --git a/open-sse/config/ttsModels.js b/open-sse/config/ttsModels.js index 6925f5f5..0142f6f1 100644 --- a/open-sse/config/ttsModels.js +++ b/open-sse/config/ttsModels.js @@ -33,6 +33,21 @@ const GEMINI_VOICES = [ "Vindemiatrix", "Sadachbia", "Sadaltager", "Sulafat", ].map((id) => ({ id, name: id, type: "tts" })); +// Xiaomi MiMo preset voices (from https://mimo.mi.com/docs/zh-CN/quick-start/usage-guide/audio/speech-synthesis-v2.5). +// Voice id is passed via `audio.voice`; `mimo_default` = default (冰糖 on CN cluster, Mia elsewhere). +// Voices are language-independent — the spoken language is a separate hint, not bound to the voice. +const MIMO_VOICES = [ + { id: "mimo_default", name: "mimo_default" }, + { id: "冰糖", name: "冰糖" }, + { id: "茉莉", name: "茉莉" }, + { id: "苏打", name: "苏打" }, + { id: "白桦", name: "白桦" }, + { id: "Mia", name: "Mia" }, + { id: "Chloe", name: "Chloe" }, + { id: "Milo", name: "Milo" }, + { id: "Dean", name: "Dean" }, +].map((v) => ({ type: "tts", ...v })); + // ── TTS Config (config-driven, single source of truth) ───────────────────── export const TTS_MODELS_CONFIG = { openai: { @@ -107,6 +122,14 @@ export const TTS_MODELS_CONFIG = { }, allVoices: GEMINI_VOICES, }, + "xiaomi-mimo": { + models: [ + { id: "mimo-v2.5-tts", name: "MiMo V2.5 TTS", type: "tts" }, + ], + voices: { + "mimo-v2.5-tts": MIMO_VOICES, + }, + }, }; // ── Helper: get voices for a specific model ──────────────────────────────── diff --git a/open-sse/executors/base.js b/open-sse/executors/base.js index b4acb649..30cf77d3 100644 --- a/open-sse/executors/base.js +++ b/open-sse/executors/base.js @@ -4,6 +4,7 @@ import { proxyAwareFetch } from "../utils/proxyFetch.js"; import { dbg } from "../utils/debugLog.js"; import { resolveProviderTimeoutMs } from "../services/providerTimeout.js"; import { ANTHROPIC_API_VERSION, OPENAI_COMPAT_BASE, ANTHROPIC_COMPAT_BASE } from "../providers/shared.js"; +import { resolveOpenAICompatibleApiType } from "../services/provider.js"; /** * BaseExecutor - Base class for provider executors @@ -31,7 +32,7 @@ export class BaseExecutor { if (this.provider?.startsWith?.("openai-compatible-")) { const baseUrl = credentials?.providerSpecificData?.baseUrl || OPENAI_COMPAT_BASE; const normalized = baseUrl.replace(/\/$/, ""); - const path = this.provider.includes("responses") ? "/responses" : "/chat/completions"; + const path = resolveOpenAICompatibleApiType(this.provider, credentials) === "responses" ? "/responses" : "/chat/completions"; return `${normalized}${path}`; } if (this.provider?.startsWith?.("anthropic-compatible-")) { @@ -127,7 +128,7 @@ export class BaseExecutor { for (let urlIndex = 0; urlIndex < fallbackCount; urlIndex++) { const url = this.buildUrl(model, stream, urlIndex, credentials); const transformedBody = this.transformRequest(model, body, stream, credentials); - const headers = this.buildHeaders(credentials, stream, url); + const headers = this.buildHeaders(credentials, stream, url, model); if (!retryAttemptsByUrl[urlIndex]) retryAttemptsByUrl[urlIndex] = 0; diff --git a/open-sse/executors/codebuddy-cn.js b/open-sse/executors/codebuddy-cn.js index 5f37d015..76ec52ff 100644 --- a/open-sse/executors/codebuddy-cn.js +++ b/open-sse/executors/codebuddy-cn.js @@ -18,6 +18,35 @@ export class CodeBuddyExecutor extends DefaultExecutor { const transformed = super.transformRequest(model, body, stream, credentials); transformed.stream = true; + // Tencent's content filter flags CLI agent system prompts ("You are Claude + // Code, Anthropic's official CLI...") as prompt injection / sensitive content + // and rejects the whole request. Detect agent system prompts (length catch-all + // + identity-marker regex) and replace them with a neutral one, while leaving + // legitimate user system prompts untouched. content may be a string or typed + // blocks ([{type:"text",text}]) depending on the incoming client format, so + // flatten before matching and preserve the original shape on replacement. + const NEUTRAL_PROMPT = "You are a helpful AI assistant that helps with software engineering tasks."; + const AGENT_PATTERN = /you are claude code|claude.?code.+official.+cli|anthropic.+official.+cli|anxthxropic.+official.+cli|you are (?:cursor|windsurf|cline|aider|continue|copilot|cody)|you are an? (?:ai )?(?:coding |code )?agent|cc_entrypoint\s*=\s*(?:cli|vscode|jetbrains|gui)|claude.?code.+issues|give feedback.+claude.?code|you are .{0,30}(?:powerful )?ai agent|orchestration capabilities|OhMyOpenCode|||/i; + const flatten = (content) => + typeof content === "string" + ? content + : Array.isArray(content) + ? content.map((b) => (b && typeof b.text === "string" ? b.text : "")).join("\n") + : ""; + if (Array.isArray(transformed.messages)) { + transformed.messages = transformed.messages.map((message) => { + if (!message || message.role !== "system") return message; + const text = flatten(message.content); + if (!text) return message; + if (text.length > 2000 || AGENT_PATTERN.test(text)) { + return typeof message.content === "string" + ? { ...message, content: NEUTRAL_PROMPT } + : { ...message, content: [{ type: "text", text: NEUTRAL_PROMPT }] }; + } + return message; + }); + } + // CodeBuddy only surfaces model reasoning when the request carries the CLI's // OpenAI-style params: reasoning_effort + reasoning_summary:"auto". 9router's // thinking pipeline sets reasoning_effort only when the client asks, and never diff --git a/open-sse/executors/codebuddy-intl.js b/open-sse/executors/codebuddy-intl.js index 06fe4326..bb99ff47 100644 --- a/open-sse/executors/codebuddy-intl.js +++ b/open-sse/executors/codebuddy-intl.js @@ -23,6 +23,20 @@ export class CodeBuddyIntlExecutor extends DefaultExecutor { } else if (eff) { transformed.reasoning_summary = "auto"; } + + // CodeBuddy rejects plain OpenAI shape (11101 invalid request): needs a + // leading system prompt + user content as typed blocks, not a bare string. + const source = Array.isArray(transformed.messages) ? transformed.messages : []; + transformed.messages = [{ role: "system", content: "You are CodeBuddy Code." }]; + for (const message of source) { + if (!message || typeof message !== "object" || ["system", "developer"].includes(message.role)) continue; + if (message.role === "user" && typeof message.content === "string") { + transformed.messages.push({ ...message, content: [{ type: "text", text: message.content }] }); + } else { + transformed.messages.push({ ...message }); + } + } + return transformed; } } diff --git a/open-sse/executors/codex.js b/open-sse/executors/codex.js index 3b04445f..4d9acbd2 100644 --- a/open-sse/executors/codex.js +++ b/open-sse/executors/codex.js @@ -8,6 +8,7 @@ import { import { normalizeResponsesInput } from "../translator/formats/responsesApi.js"; import { fetchImageAsBase64 } from "../translator/concerns/image.js"; import { getModelUpstreamId } from "../config/providerModels.js"; +import { getThinkingLevels } from "../providers/thinkingLevels.js"; import { DEFAULT_RETRY_CONFIG, HTTP_STATUS, resolveRetryEntry } from "../config/runtimeConfig.js"; import { dbg } from "../utils/debugLog.js"; import { resolveSessionId } from "../utils/sessionManager.js"; @@ -124,8 +125,12 @@ function resolveCacheSessionId(body, credentials) { }); } -function normalizeReasoningEffort(value) { - return value === "max" ? "xhigh" : value; +function normalizeReasoningEffort(model, value) { + const supportedLevels = getThinkingLevels("codex", model); + if (supportedLevels?.includes(value)) return value; + if (value === "ultra" && supportedLevels?.includes("max")) return "max"; + if (value === "max" || value === "ultra") return "xhigh"; + return value; } function findNestedMessage(value, depth = 0) { @@ -440,10 +445,10 @@ export class CodexExecutor extends BaseExecutor { // Priority: explicit reasoning.effort > reasoning_effort param > model suffix > default (medium) if (!body.reasoning) { - const effort = normalizeReasoningEffort(body.reasoning_effort || modelEffort || 'low'); + const effort = normalizeReasoningEffort(body.model, body.reasoning_effort || modelEffort || 'low'); body.reasoning = { effort, summary: "auto" }; } else { - body.reasoning.effort = normalizeReasoningEffort(body.reasoning.effort); + body.reasoning.effort = normalizeReasoningEffort(body.model, body.reasoning.effort); if (!body.reasoning.summary) body.reasoning.summary = "auto"; } delete body.reasoning_effort; diff --git a/open-sse/executors/default.js b/open-sse/executors/default.js index 48be8350..92c78e92 100644 --- a/open-sse/executors/default.js +++ b/open-sse/executors/default.js @@ -1,9 +1,9 @@ import { BaseExecutor } from "./base.js"; import { PROVIDERS, PROVIDER_OAUTH } from "../config/providers.js"; -import { ANTHROPIC_API_VERSION, OPENAI_COMPAT_BASE, ANTHROPIC_COMPAT_BASE } from "../providers/shared.js"; +import { ANTHROPIC_API_VERSION, OPENAI_COMPAT_BASE, ANTHROPIC_COMPAT_BASE, selectAnthropicBeta } from "../providers/shared.js"; +import { resolveOpenAICompatibleApiType } from "../services/provider.js"; import { OAUTH_ENDPOINTS, buildKimiHeaders } from "../config/appConstants.js"; import { buildClineHeaders } from "../shared/clineAuth.js"; -import { getCachedClaudeHeaders } from "../utils/claudeHeaderCache.js"; import { proxyAwareFetch } from "../utils/proxyFetch.js"; import { injectReasoningContent } from "../utils/reasoningContentInjector.js"; import { stripUnsupportedParams } from "../translator/concerns/paramSupport.js"; @@ -42,21 +42,6 @@ const HEADER_HOOKS = { kimiHeaders: (h, c) => Object.assign(h, buildKimiHeaders(c?.providerSpecificData?.deviceId)), clineHeaders: (h, c) => Object.assign(h, buildClineHeaders(c.apiKey || c.accessToken)), kilocodeOrg: (h, c) => { if (c.providerSpecificData?.orgId) h["X-Kilocode-OrganizationID"] = c.providerSpecificData.orgId; }, - claudeOverlay: (h) => { - const cached = getCachedClaudeHeaders(); - if (!cached) return; - for (const lcKey of Object.keys(cached)) { - const titleKey = lcKey.replace(/(^|-)([a-z])/g, (_, sep, ch) => sep + ch.toUpperCase()); - if (lcKey === "anthropic-beta") { - const staticBetaStr = h[titleKey] || h[lcKey] || ""; - const flags = new Set(staticBetaStr.split(",").map(f => f.trim()).filter(Boolean)); - for (const f of cached[lcKey].split(",").map(f => f.trim()).filter(Boolean)) flags.add(f); - cached[lcKey] = Array.from(flags).join(","); - } - if (titleKey !== lcKey && h[titleKey] !== undefined) delete h[titleKey]; - } - Object.assign(h, cached); - }, }; // Config-driven OAuth refresh grants — derived from registry oauth.refresh. @@ -125,7 +110,7 @@ export class DefaultExecutor extends BaseExecutor { if (this.provider?.startsWith?.("openai-compatible-")) { const baseUrl = credentials?.providerSpecificData?.baseUrl || OPENAI_COMPAT_BASE; const normalized = baseUrl.replace(/\/$/, ""); - const path = this.provider.includes("responses") ? "/responses" : "/chat/completions"; + const path = resolveOpenAICompatibleApiType(this.provider, credentials) === "responses" ? "/responses" : "/chat/completions"; return `${normalized}${path}`; } if (this.provider?.startsWith?.("anthropic-compatible-")) { @@ -161,14 +146,18 @@ export class DefaultExecutor extends BaseExecutor { return BEARER; } - buildHeaders(credentials, stream = true) { + buildHeaders(credentials, stream = true, url, model) { const rt = credentials?.runtimeTransport; const headers = { "Content-Type": "application/json", ...(rt ? rt.headers : this.config.headers) }; const desc = rt?.auth || AUTH_DESCRIPTORS[this.provider] || this.resolveAuthDescriptor(); - // Hooks run BEFORE auth so dynamic overlays (claude cached headers) can't clobber the token. + // Hooks run BEFORE auth so dynamic overlays can't clobber the token. for (const hook of desc.hooks || []) HEADER_HOOKS[hook]?.(headers, credentials); applyAuth(headers, desc, credentials); + if (this.provider === "claude" && model) { + headers["Anthropic-Beta"] = selectAnthropicBeta(model); + } + // Strip first-party Claude Code identity headers for non-Anthropic anthropic-compatible upstreams if (this.provider?.startsWith?.("anthropic-compatible-")) { const baseUrl = credentials?.providerSpecificData?.baseUrl || ""; @@ -222,7 +211,6 @@ export class DefaultExecutor extends BaseExecutor { const refreshers = { claude: () => this.refreshFromGrant(credentials, proxyOptions), codex: () => this.refreshFromGrant(credentials, proxyOptions), - qwen: () => this.refreshWithForm(OAUTH_ENDPOINTS.qwen.token, { grant_type: "refresh_token", refresh_token: credentials.refreshToken, client_id: PROVIDERS.qwen.clientId }, proxyOptions), iflow: () => this.refreshIflow(credentials.refreshToken, proxyOptions), gemini: () => this.refreshFromGrant(credentials, proxyOptions), kiro: () => this.refreshKiro(credentials.refreshToken, proxyOptions), diff --git a/open-sse/executors/index.js b/open-sse/executors/index.js index 92e10464..8dd03421 100644 --- a/open-sse/executors/index.js +++ b/open-sse/executors/index.js @@ -9,7 +9,6 @@ import { KimchiExecutor } from "./kimchi.js"; import { CodexExecutor } from "./codex.js"; import { CursorExecutor } from "./cursor.js"; import { VertexExecutor } from "./vertex.js"; -import { QwenExecutor } from "./qwen.js"; import { OpenCodeExecutor } from "./opencode.js"; import { OpenCodeGoExecutor } from "./opencode-go.js"; import { GrokWebExecutor } from "./grok-web.js"; @@ -41,7 +40,6 @@ const executors = { cu: new CursorExecutor(), // Alias for cursor vertex: new VertexExecutor("vertex"), "vertex-partner": new VertexExecutor("vertex-partner"), - qwen: new QwenExecutor(), opencode: new OpenCodeExecutor(), "opencode-go": new OpenCodeGoExecutor(), "grok-web": new GrokWebExecutor(), @@ -87,7 +85,6 @@ export { CodexExecutor } from "./codex.js"; export { CursorExecutor } from "./cursor.js"; export { VertexExecutor } from "./vertex.js"; export { DefaultExecutor } from "./default.js"; -export { QwenExecutor } from "./qwen.js"; export { OpenCodeExecutor } from "./opencode.js"; export { OpenCodeGoExecutor } from "./opencode-go.js"; export { GrokWebExecutor } from "./grok-web.js"; diff --git a/open-sse/executors/qoder.js b/open-sse/executors/qoder.js index 7fa0e554..9bcc3859 100644 --- a/open-sse/executors/qoder.js +++ b/open-sse/executors/qoder.js @@ -33,13 +33,11 @@ import { FETCH_CONNECT_TIMEOUT_MS } from "../config/runtimeConfig.js"; import { resolveProviderTimeoutMs } from "../services/providerTimeout.js"; import { QODER_CHAT_URL_ENCODED, - QODER_JOB_TOKEN_EXCHANGE_URL, - QODER_USERINFO_URL, + QODER_CHAT_BASE_ALT, + QODER_CHAT_SIG_PATH, QODER_MODEL_MAP, - QODER_IDE_VERSION, - QODER_CLIENT_TYPE, } from "../shared/qoder/constants.js"; -import { getQoderModelConfig, resolveQoderModels } from "../services/qoderModels.js"; +import { getQoderModelConfig, resolveQoderModels, isQoderPat, resolveQoderCredentials } from "../services/qoderModels.js"; /** * Hoist role:"system" messages out of the messages array (Qoder rejects @@ -343,98 +341,18 @@ function wrapQoderSSE(response, model) { }); } -// ── PAT (Personal Access Token) → job-token exchange ─────────────────────── -// PATs (pt-...) cannot sign COSY requests directly. Exchange them for a -// short-lived job token (jt-...) via /api/v1/jobToken/exchange (plain JSON, -// not COSY-signed), then resolve the userId from userinfo. Mirrors the -// official qodercli flow. Cached per-PAT until near-expiry. -const PAT_PREFIX = "pt-"; -const PAT_REFRESH_BUFFER_MS = 5 * 60 * 1000; -const patJobCache = new Map(); - -export function isQoderPat(token) { - return typeof token === "string" && token.startsWith(PAT_PREFIX); -} - -async function exchangeJobToken(pat, proxyOptions = null, signal = null) { - const res = await proxyAwareFetch( - QODER_JOB_TOKEN_EXCHANGE_URL, - { - method: "POST", - headers: { - "Content-Type": "application/json", - Accept: "application/json", - "User-Agent": "qodercli/1.0.0", - "Cosy-Version": QODER_IDE_VERSION, - "Cosy-ClientType": QODER_CLIENT_TYPE, - }, - body: JSON.stringify({ personal_token: pat }), - signal, - }, - proxyOptions, - ); - if (!res.ok) { - const text = await res.text().catch(() => ""); - throw new Error(`qoder PAT exchange failed: ${res.status} ${text.slice(0, 200)}`); - } - const data = await res.json(); - if (!data.token) throw new Error("qoder PAT exchange returned no job token"); - - let expiresAt = Date.now() + 24 * 60 * 60 * 1000; - if (data.expires_at) { - const parsed = Date.parse(data.expires_at); - if (!Number.isNaN(parsed)) expiresAt = parsed; - } else if (typeof data.expires_in === "number" && data.expires_in > 0) { - expiresAt = Date.now() + data.expires_in; - } - return { jobToken: data.token, jobRefreshToken: data.refresh_token || "", expiresAt }; -} - -async function fetchUserIdForJobToken(jobToken, proxyOptions = null, signal = null) { - try { - const res = await proxyAwareFetch( - QODER_USERINFO_URL, - { - method: "GET", - headers: { - Authorization: `Bearer ${jobToken}`, - Accept: "application/json", - "User-Agent": "qodercli/1.0.0", - }, - signal, - }, - proxyOptions, - ); - if (!res.ok) return ""; - const info = await res.json().catch(() => ({})); - return info.id || info.userId || info.user_id || ""; - } catch { - return ""; - } -} - -/** - * Exchange a PAT for a job token + userId, caching until near-expiry so repeat - * chat requests don't re-exchange. Returns { accessToken, userId }. - */ -async function resolvePatCredential(pat, proxyOptions = null, signal = null) { - const cached = patJobCache.get(pat); - if (cached && cached.expiresAt - Date.now() > PAT_REFRESH_BUFFER_MS) { - return cached; - } - const { jobToken, expiresAt } = await exchangeJobToken(pat, proxyOptions, signal); - const userId = await fetchUserIdForJobToken(jobToken, proxyOptions, signal); - const entry = { accessToken: jobToken, userId, expiresAt }; - patJobCache.set(pat, entry); - return entry; -} - export class QoderExecutor extends BaseExecutor { constructor() { super("qoder", PROVIDERS.qoder); } - buildUrl() { + buildUrl(credentials) { + // Job-token (jt-...) traffic must hit api2.qoder.sh — api3 rejects jt- + // with "Login expired" (403). Device tokens (dt-...) stay on api3. + const raw = credentials?.apiKey || credentials?.accessToken; + if (typeof raw === "string" && !raw.startsWith("pt-") && (raw.startsWith("jt-") || (credentials?.accessToken || "").startsWith("jt-"))) { + return `${QODER_CHAT_BASE_ALT}/algo${QODER_CHAT_SIG_PATH}?FetchKeys=llm_model_result&AgentId=agent_common&Encode=1`; + } return QODER_CHAT_URL_ENCODED; } @@ -444,36 +362,24 @@ export class QoderExecutor extends BaseExecutor { // - COSY headers built from the *encoded* body bytes // - response stream re-wrapped from {statusCodeValue, body} to OpenAI SSE async execute({ model, body, stream, credentials, signal, log, proxyOptions = null }) { - const url = this.buildUrl(); - // PAT (pt-...) → exchange for short-lived job token + resolve userId so // downstream COSY signing + catalog fetch work. Device tokens (dt-...) and // job tokens (jt-...) skip this and are used directly. const rawToken = credentials?.apiKey || credentials?.accessToken; if (isQoderPat(rawToken)) { try { - const resolved = await resolvePatCredential(rawToken, proxyOptions, signal); - credentials = { - ...credentials, - accessToken: resolved.accessToken, - apiKey: undefined, - providerSpecificData: { - authMethod: "pat", - ...(credentials?.providerSpecificData || {}), - userId: resolved.userId || credentials?.providerSpecificData?.userId || "", - machineId: credentials?.providerSpecificData?.machineId || "", - }, - }; + credentials = await resolveQoderCredentials(credentials, proxyOptions, signal); } catch (err) { log?.error?.("QODER", `PAT exchange failed: ${err.message}`); const fakeResp = new Response( JSON.stringify({ error: { message: `qoder PAT exchange failed: ${err.message}` } }), { status: 401, headers: { "Content-Type": "application/json" } }, ); - return { response: fakeResp, url, headers: {}, transformedBody: body }; + return { response: fakeResp, url: this.buildUrl(credentials), headers: {}, transformedBody: body }; } } + const url = this.buildUrl(credentials); const psd = credentials?.providerSpecificData || {}; if (!psd.userId) { // No user id → no way to sign. Surface a 401 so the dashboard nudges @@ -591,6 +497,4 @@ export const __test__ = { normalizeMessages, wrapQoderSSE, buildQoderRequestBody, - isQoderPat, - resolvePatCredential, }; diff --git a/open-sse/executors/qwen.js b/open-sse/executors/qwen.js deleted file mode 100644 index ae828118..00000000 --- a/open-sse/executors/qwen.js +++ /dev/null @@ -1,129 +0,0 @@ -import { DefaultExecutor } from "./default.js"; -import { PROVIDERS } from "../config/providers.js"; -import { OAUTH_ENDPOINTS } from "../config/appConstants.js"; - -/** portal.qwen.ai — static fingerprint matching stable Qwen Code release */ -const QWEN_USER_AGENT = "QwenCode/0.12.3 (linux; x64)"; -const QWEN_STAINLESS = { - os: "Linux", - arch: "x64", - lang: "js", - runtime: "node", - runtimeVersion: "v18.19.1", - packageVersion: "5.11.0", - retryCount: "1" -}; -const QWEN_DEFAULT_SYSTEM_MESSAGE = { - role: "system", - content: [{ type: "text", text: "", cache_control: { type: "ephemeral" } }] -}; - -function ensureQwenSystemMessage(body) { - if (!body || typeof body !== "object") return body; - const next = { ...body }; - if (Array.isArray(next.messages)) { - next.messages = [QWEN_DEFAULT_SYSTEM_MESSAGE, ...next.messages]; - } else { - next.messages = [QWEN_DEFAULT_SYSTEM_MESSAGE]; - } - return next; -} - -function isQwenThinkingActive(body) { - const thinking = body?.thinking; - if (thinking === true || body?.enable_thinking === true) return true; - return typeof thinking === "object" && thinking !== null && !Array.isArray(thinking) && thinking.type === "enabled"; -} - -// Qwen rejects tool_choice="required" or object forms when thinking is active; neutralize to "auto". -function sanitizeQwenThinkingToolChoice(body) { - if (!isQwenThinkingActive(body)) return body; - const tc = body.tool_choice; - const incompatible = tc === "required" || (typeof tc === "object" && tc !== null); - if (!incompatible) return body; - return { ...body, tool_choice: "auto" }; -} - -function buildQwenUpstreamHeaders(credentials, stream = true) { - const token = credentials?.apiKey || credentials?.accessToken || ""; - const headers = { - "Content-Type": "application/json", - Authorization: `Bearer ${token}`, - "User-Agent": QWEN_USER_AGENT, - "X-DashScope-AuthType": "qwen-oauth", - "X-DashScope-CacheControl": "enable", - "X-DashScope-UserAgent": QWEN_USER_AGENT, - "X-Stainless-Arch": QWEN_STAINLESS.arch, - "X-Stainless-Lang": QWEN_STAINLESS.lang, - "X-Stainless-Os": QWEN_STAINLESS.os, - "X-Stainless-Package-Version": QWEN_STAINLESS.packageVersion, - "X-Stainless-Retry-Count": QWEN_STAINLESS.retryCount, - "X-Stainless-Runtime": QWEN_STAINLESS.runtime, - "X-Stainless-Runtime-Version": QWEN_STAINLESS.runtimeVersion, - Connection: "keep-alive", - "Accept-Language": "*", - "Sec-Fetch-Mode": "cors" - }; - headers.Accept = stream ? "text/event-stream" : "application/json"; - return headers; -} - -export class QwenExecutor extends DefaultExecutor { - constructor() { - super("qwen"); - } - - // Qwen tokens are bound to a resource_url returned at OAuth time. - // Using portal.qwen.ai when the token is issued for another shard returns 401/403. - buildUrl(model, stream, urlIndex = 0, credentials = null) { - const resourceUrl = credentials?.providerSpecificData?.resourceUrl; - const host = resourceUrl ? resourceUrl.replace(/^https?:\/\//, "").replace(/\/$/, "") : "portal.qwen.ai"; - return `https://${host}/v1/chat/completions`; - } - - buildHeaders(credentials, stream = true) { - return buildQwenUpstreamHeaders(credentials, stream); - } - - transformRequest(model, body, stream, credentials) { - let next = body && typeof body === "object" ? { ...body } : body; - if (stream && next?.messages && !next.stream_options && !next.thinking && !next.enable_thinking && next.stream !== false) { - next.stream_options = { include_usage: true }; - } - next = sanitizeQwenThinkingToolChoice(next); - return ensureQwenSystemMessage(next); - } - - // Override to capture resource_url from refresh response (required for buildUrl). - async refreshCredentials(credentials, log) { - if (!credentials?.refreshToken) return null; - try { - const response = await fetch(OAUTH_ENDPOINTS.qwen.token, { - method: "POST", - headers: { "Content-Type": "application/x-www-form-urlencoded", Accept: "application/json" }, - body: new URLSearchParams({ - grant_type: "refresh_token", - refresh_token: credentials.refreshToken, - client_id: PROVIDERS.qwen.clientId - }) - }); - if (!response.ok) return null; - const tokens = await response.json(); - log?.info?.("TOKEN", "qwen refreshed"); - return { - accessToken: tokens.access_token, - refreshToken: tokens.refresh_token || credentials.refreshToken, - expiresIn: tokens.expires_in, - providerSpecificData: { - ...(credentials.providerSpecificData || {}), - ...(tokens.resource_url ? { resourceUrl: tokens.resource_url } : {}) - } - }; - } catch (error) { - log?.error?.("TOKEN", `qwen refresh error: ${error.message}`); - return null; - } - } -} - -export default QwenExecutor; diff --git a/open-sse/handlers/chatCore.js b/open-sse/handlers/chatCore.js index 21255998..1d90992d 100644 --- a/open-sse/handlers/chatCore.js +++ b/open-sse/handlers/chatCore.js @@ -4,7 +4,7 @@ import { resolveTransport, } from "../services/provider.js"; import { translateRequest } from "../translator/index.js"; -import { stripThinkingSuffix } from "../translator/concerns/thinkingUnified.js"; +import { applyThinking, extractThinking, stripThinkingSuffix } from "../translator/concerns/thinkingUnified.js"; import { FORMATS } from "../translator/formats.js"; import { normalizeClaudePassthrough } from "../translator/formats/claude.js"; import { createStreamController } from "../utils/streamHandler.js"; @@ -61,7 +61,6 @@ import { compressWithPxpipe } from "../rtk/pxpipe.js"; import { getCapabilitiesForModel } from "../providers/capabilities.js"; import { stripUnsupportedModalities } from "../translator/concerns/modality.js"; import { prefetchRemoteImages } from "../translator/concerns/prefetch.js"; -import { extractThinking } from "../translator/concerns/thinkingUnified.js"; import { resolveSessionId } from "../utils/sessionManager.js"; /** @@ -71,6 +70,26 @@ import { resolveSessionId } from "../utils/sessionManager.js"; * @param {object} options.credentials - Provider credentials * @param {string} options.sourceFormatOverride - Override detected source format (e.g. "openai-responses") */ +/** + * Remove translator-internal continuity fields from the outbound upstream + * body. The Responses→Chat request translator stashes reasoning + * `encrypted_content` on assistant messages so a later openai→responses + * round-trip can restore the store=false continuity blob; that stash must + * never reach an upstream provider. Chat-native proxies reject the unknown + * assistant-message field and answer every turn with a literal "400" body + * (observed with multi-turn Codex sessions via OpenAI-compatible nodes). + */ +export function stripContinuityFields(body) { + if (!body || !Array.isArray(body.messages)) return body; + for (const msg of body.messages) { + if (msg && typeof msg === "object") { + delete msg.encrypted_content; + delete msg.reasoning_encrypted_content; + } + } + return body; +} + export async function handleChatCore({ body, modelInfo, @@ -138,7 +157,7 @@ export async function handleChatCore({ // Multi-endpoint providers: pick transport matching sourceFormat → zero translation const runtimeTransport = resolveTransport(provider, sourceFormat); const targetFormat = - modelTargetFormat || runtimeTransport?.format || getTargetFormat(provider); + modelTargetFormat || runtimeTransport?.format || getTargetFormat(provider, credentials); if (runtimeTransport && credentials) credentials.runtimeTransport = runtimeTransport; const stripList = getModelStrip(alias, model); @@ -248,12 +267,25 @@ export async function handleChatCore({ let translatedBody; let toolNameMap; + let customToolNames; if (passthrough) { log?.debug?.( "PASSTHROUGH", `${clientTool} → ${provider} | native lossless`, ); translatedBody = { ...body, model: stripThinkingSuffix(upstreamModel) }; + if (provider === "codex") { + const suffixThinking = {}; + applyThinking(sourceFormat, upstreamModel, suffixThinking, provider); + if (suffixThinking.reasoning_effort) { + const reasoning = translatedBody.reasoning; + translatedBody.reasoning = { + ...(reasoning && typeof reasoning === "object" && !Array.isArray(reasoning) ? reasoning : {}), + effort: suffixThinking.reasoning_effort, + }; + delete translatedBody.reasoning_effort; + } + } // Normalize newer Cowork/CC beta shapes (adaptive thinking, mid-conversation system) the API rejects if (clientTool === "claude") normalizeClaudePassthrough(translatedBody, translatedBody.model); @@ -280,7 +312,10 @@ export async function handleChatCore({ } toolNameMap = translatedBody._toolNameMap; delete translatedBody._toolNameMap; + customToolNames = translatedBody._customToolNames; + delete translatedBody._customToolNames; translatedBody.model = stripThinkingSuffix(upstreamModel); + stripContinuityFields(translatedBody); } // Dedupe duplicate built-in tools when equivalent MCP tools are present (Claude clients only). @@ -736,6 +771,8 @@ export async function handleChatCore({ ...sharedCtx, providerResponse, sourceFormat, + targetFormat: providerResponseFormat, + customToolNames, trackDone, appendLog, }); @@ -754,6 +791,7 @@ export async function handleChatCore({ targetFormat: providerResponseFormat, reqLogger, toolNameMap, + customToolNames, trackDone, appendLog, }); @@ -773,6 +811,7 @@ export async function handleChatCore({ userAgent, reqLogger, toolNameMap, + customToolNames, streamController, onStreamComplete, streamDetailId, diff --git a/open-sse/handlers/chatCore/nonStreamingHandler.js b/open-sse/handlers/chatCore/nonStreamingHandler.js index d1f4a845..5489b344 100644 --- a/open-sse/handlers/chatCore/nonStreamingHandler.js +++ b/open-sse/handlers/chatCore/nonStreamingHandler.js @@ -9,6 +9,7 @@ import { parseSSEToOpenAIResponse } from "./sseToJsonHandler.js"; import { buildRequestDetail, extractRequestConfig, extractUsageFromResponse, saveUsageStats, formatDoneLine } from "./requestDetail.js"; import { appendRequestLog, saveRequestDetail } from "@/lib/usageDb.js"; import { decloakToolNames } from "../../utils/claudeCloaking.js"; +import { ROLE, RESPONSES_ITEM } from "../../translator/schema/index.js"; function parseToolArguments(value) { if (!value) return {}; @@ -60,11 +61,93 @@ function openAICompletionToClaudeMessage(responseBody) { }; } +/** + * Convert an OpenAI Chat Completions non-streaming response body into the + * OpenAI Responses API shape. Used when a Responses-format client (e.g. Codex) + * is routed to a Chat Completions upstream and `stream:false` — the streaming + * path already emits Responses events, but the JSON path returned a raw + * `chat.completion` body, so tool_calls were invisible to Responses clients. + */ +function extractCustomToolInput(argumentsValue) { + const argumentsText = typeof argumentsValue === "string" ? argumentsValue : JSON.stringify(argumentsValue || {}); + try { + const parsed = JSON.parse(argumentsText); + if (parsed && typeof parsed === "object" && typeof parsed.input === "string") return parsed.input; + } catch { /* raw freeform input */ } + return argumentsText; +} + +function openAICompletionToResponses(responseBody, customToolNames = null) { + const choice = responseBody?.choices?.[0]; + if (!choice) return responseBody; + + const message = choice.message || {}; + const output = []; + + // Reasoning → a reasoning item (summary text), mirroring the streaming path. + const reasoning = message.reasoning_content || message.reasoning; + if (typeof reasoning === "string" && reasoning.length > 0) { + output.push({ + type: RESPONSES_ITEM.REASONING, + summary: [{ type: RESPONSES_ITEM.SUMMARY_TEXT, text: reasoning }], + }); + } + + // Assistant text → a message item with output_text content. + const text = typeof message.content === "string" ? message.content : ""; + if (text.length > 0) { + output.push({ + type: RESPONSES_ITEM.MESSAGE, + role: ROLE.ASSISTANT, + content: [{ type: RESPONSES_ITEM.OUTPUT_TEXT, text, annotations: [] }], + }); + } + + // tool_calls → function_call/custom_tool_call items (Responses-native tool shape). + for (const tc of message.tool_calls || []) { + const fn = tc.function || {}; + const custom = customToolNames?.has(fn.name); + output.push({ + type: custom ? RESPONSES_ITEM.CUSTOM_TOOL_CALL : RESPONSES_ITEM.FUNCTION_CALL, + id: `${custom ? "ctc" : "fc"}_${tc.id || ""}`, + call_id: tc.id || "", + name: fn.name || "", + ...(custom + ? { input: extractCustomToolInput(fn.arguments) } + : { arguments: typeof fn.arguments === "string" ? fn.arguments : JSON.stringify(fn.arguments || {}) }), + }); + } + + const usage = responseBody.usage || {}; + const status = choice.finish_reason === "tool_calls" ? "completed" : (choice.finish_reason === "stop" ? "completed" : (choice.finish_reason || "completed")); + + return { + id: `resp_${responseBody.id || ""}`.replace(/^resp_chatcmpl-/, "resp_"), + object: "response", + created_at: responseBody.created || Math.floor(Date.now() / 1000), + model: responseBody.model || "unknown", + status, + background: false, + error: null, + output, + usage: { + input_tokens: usage.prompt_tokens || usage.input_tokens || 0, + output_tokens: usage.completion_tokens || usage.output_tokens || 0, + total_tokens: usage.total_tokens || (usage.prompt_tokens || 0) + (usage.completion_tokens || 0), + }, + }; +} + /** * Translate non-streaming response body from provider format → OpenAI format. */ -export function translateNonStreamingResponse(responseBody, targetFormat, sourceFormat) { +export function translateNonStreamingResponse(responseBody, targetFormat, sourceFormat, customToolNames = null) { if (targetFormat === sourceFormat) return responseBody; + // Provider responded in OpenAI Chat Completions shape but the client speaks + // Responses API — convert so tool_calls/text surface as Responses `output`. + if (targetFormat === FORMATS.OPENAI && sourceFormat === FORMATS.OPENAI_RESPONSES) { + return openAICompletionToResponses(responseBody, customToolNames); + } if (targetFormat === FORMATS.OPENAI && sourceFormat === FORMATS.CLAUDE) { return openAICompletionToClaudeMessage(responseBody); } @@ -198,7 +281,7 @@ export function translateNonStreamingResponse(responseBody, targetFormat, source /** * Handle non-streaming response from provider. */ -export async function handleNonStreamingResponse({ providerResponse, provider, model, sourceFormat, targetFormat, body, stream, translatedBody, finalBody, requestStartTime, connectionId, apiKey, clientRawRequest, onRequestSuccess, reqLogger, toolNameMap, trackDone, appendLog, pxpipe, reqTag, log }) { +export async function handleNonStreamingResponse({ providerResponse, provider, model, sourceFormat, targetFormat, body, stream, translatedBody, finalBody, requestStartTime, connectionId, apiKey, clientRawRequest, onRequestSuccess, reqLogger, toolNameMap, customToolNames, trackDone, appendLog, pxpipe, reqTag, log }) { trackDone(); const contentType = providerResponse.headers.get("content-type") || ""; let responseBody; @@ -239,9 +322,12 @@ export async function handleNonStreamingResponse({ providerResponse, provider, m if (log?.line) log.line(reqTag, "📊", formatDoneLine({ usage, latency: { total: Date.now() - requestStartTime } })); const translatedResponse = needsTranslation(targetFormat, sourceFormat) - ? translateNonStreamingResponse(responseBody, targetFormat, sourceFormat) + ? translateNonStreamingResponse(responseBody, targetFormat, sourceFormat, customToolNames) : responseBody; const isClaudeMessageResponse = sourceFormat === FORMATS.CLAUDE && translatedResponse?.type === "message"; + // Responses-format translation produces a `object:"response"` body with no + // `choices`; skip the Chat-Completions-specific post-processing below for it. + const isResponsesResponse = sourceFormat === FORMATS.OPENAI_RESPONSES && translatedResponse?.object === "response"; // Fix finish_reason for tool_calls: some providers return non-standard values (e.g. "other") if (translatedResponse?.choices?.[0]) { @@ -254,13 +340,13 @@ export async function handleNonStreamingResponse({ providerResponse, provider, m } // Ensure OpenAI-required fields - if (!isClaudeMessageResponse) { + if (!isClaudeMessageResponse && !isResponsesResponse) { if (!translatedResponse.object) translatedResponse.object = "chat.completion"; if (!translatedResponse.created) translatedResponse.created = Math.floor(Date.now() / 1000); } // Strip Azure-specific fields - if (!isClaudeMessageResponse) { + if (!isClaudeMessageResponse && !isResponsesResponse) { delete translatedResponse.prompt_filter_results; if (translatedResponse?.choices) { for (const choice of translatedResponse.choices) delete choice.content_filter_results; @@ -274,7 +360,7 @@ export async function handleNonStreamingResponse({ providerResponse, provider, m // Strip reasoning_content only when content is non-empty. // When content is empty (e.g. thinking models that used all tokens for reasoning), // reasoning_content is the only useful output and must be preserved. - if (!isClaudeMessageResponse && translatedResponse?.choices) { + if (!isClaudeMessageResponse && !isResponsesResponse && translatedResponse?.choices) { for (const choice of translatedResponse.choices) { if (choice?.message?.reasoning_content && choice.message.content) { delete choice.message.reasoning_content; diff --git a/open-sse/handlers/chatCore/sseToJsonHandler.js b/open-sse/handlers/chatCore/sseToJsonHandler.js index 3ab6e9bf..b169d18d 100644 --- a/open-sse/handlers/chatCore/sseToJsonHandler.js +++ b/open-sse/handlers/chatCore/sseToJsonHandler.js @@ -4,12 +4,8 @@ import { createErrorResult } from "../../utils/error.js"; import { HTTP_STATUS } from "../../config/runtimeConfig.js"; import { FORMATS } from "../../translator/formats.js"; import { PROVIDERS } from "../../config/providers.js"; -import { - buildRequestDetail, - extractRequestConfig, - saveUsageStats, - formatDoneLine, -} from "./requestDetail.js"; +import { buildRequestDetail, extractRequestConfig, saveUsageStats, formatDoneLine } from "./requestDetail.js"; +import { ROLE, RESPONSES_ITEM } from "../../translator/schema/index.js"; // Responses-API providers (e.g. codex) may emit SSE without content-type + use Responses output shape const isResponsesProvider = (p) => @@ -41,6 +37,76 @@ function pickAssistantMessageForChatCompletion(output) { return { msgItem: last, textContent: textFromResponsesMessageItem(last) }; } +/** + * Convert an OpenAI Chat Completions JSON body into the Responses API shape. + * Inlined here (not imported from nonStreamingHandler.js) to avoid a circular + * import. Mirrors openAICompletionToResponses in nonStreamingHandler.js. + */ +function extractCustomToolInput(argumentsValue) { + const argumentsText = typeof argumentsValue === "string" ? argumentsValue : JSON.stringify(argumentsValue || {}); + try { + const parsed = JSON.parse(argumentsText); + if (parsed && typeof parsed === "object" && typeof parsed.input === "string") return parsed.input; + } catch { /* raw freeform input */ } + return argumentsText; +} + +function chatCompletionToResponses(responseBody, customToolNames = null) { + const choice = responseBody?.choices?.[0]; + if (!choice) return responseBody; + + const message = choice.message || {}; + const output = []; + + const reasoning = message.reasoning_content || message.reasoning; + if (typeof reasoning === "string" && reasoning.length > 0) { + output.push({ + type: RESPONSES_ITEM.REASONING, + summary: [{ type: RESPONSES_ITEM.SUMMARY_TEXT, text: reasoning }], + }); + } + + const text = typeof message.content === "string" ? message.content : ""; + if (text.length > 0) { + output.push({ + type: RESPONSES_ITEM.MESSAGE, + role: ROLE.ASSISTANT, + content: [{ type: RESPONSES_ITEM.OUTPUT_TEXT, text, annotations: [] }], + }); + } + + for (const tc of message.tool_calls || []) { + const fn = tc.function || {}; + const custom = customToolNames?.has(fn.name); + output.push({ + type: custom ? RESPONSES_ITEM.CUSTOM_TOOL_CALL : RESPONSES_ITEM.FUNCTION_CALL, + id: `${custom ? "ctc" : "fc"}_${tc.id || ""}`, + call_id: tc.id || "", + name: fn.name || "", + ...(custom + ? { input: extractCustomToolInput(fn.arguments) } + : { arguments: typeof fn.arguments === "string" ? fn.arguments : JSON.stringify(fn.arguments || {}) }), + }); + } + + const usage = responseBody.usage || {}; + return { + id: `resp_${responseBody.id || ""}`.replace(/^resp_chatcmpl-/, "resp_"), + object: "response", + created_at: responseBody.created || Math.floor(Date.now() / 1000), + model: responseBody.model || "unknown", + status: "completed", + background: false, + error: null, + output, + usage: { + input_tokens: usage.prompt_tokens || usage.input_tokens || 0, + output_tokens: usage.completion_tokens || usage.output_tokens || 0, + total_tokens: usage.total_tokens || (usage.prompt_tokens || 0) + (usage.completion_tokens || 0), + }, + }; +} + /** * Parse OpenAI-style SSE text into a single chat completion JSON. * Used when provider forces streaming but client wants non-streaming. @@ -136,6 +202,7 @@ export function parseSSEToOpenAIResponse(rawSSE, fallbackModel) { export async function handleForcedSSEToJson({ providerResponse, sourceFormat, + targetFormat, provider, model, body, @@ -147,6 +214,7 @@ export async function handleForcedSSEToJson({ apiKey, clientRawRequest, onRequestSuccess, + customToolNames, trackDone, appendLog, reqTag, @@ -170,8 +238,11 @@ export async function handleForcedSSEToJson({ }; // Codex/Responses API SSE path + // Branch on the UPSTREAM format (targetFormat = format we spoke to the provider in), + // not the client format: a Responses-API client behind a chat-native forced-streaming + // provider still receives chat SSE chunks, which must go through the standard path. const isCodexResponsesApi = - isResponsesProvider(provider) || sourceFormat === FORMATS.OPENAI_RESPONSES; + isResponsesProvider(provider) || targetFormat === FORMATS.OPENAI_RESPONSES; if (isCodexResponsesApi) { try { const jsonResponse = await convertResponsesStreamToJson( @@ -200,6 +271,11 @@ export async function handleForcedSSEToJson({ }), ); + // Same cache-inclusive total for the recorded detail, so the DB and the + // client-facing usage can never disagree. + const inTokensForLog = (usage.input_tokens || 0) + + (usage.cache_read_input_tokens || usage.cached_tokens || 0) + + (usage.cache_creation_input_tokens || 0); const { msgItem, textContent } = pickAssistantMessageForChatCompletion( jsonResponse.output, ); @@ -209,9 +285,10 @@ export async function handleForcedSSEToJson({ buildRequestDetail( { ...ctx, + apiKey, latency: { ttft: totalLatency, total: totalLatency }, tokens: { - prompt_tokens: usage.input_tokens || 0, + prompt_tokens: inTokensForLog, completion_tokens: usage.output_tokens || 0, }, response: { @@ -238,9 +315,22 @@ export async function handleForcedSSEToJson({ }; } - // Build client-format response - const inTokens = usage.input_tokens || 0; + // Build client-format response. + // input_tokens EXCLUDES cached tokens on cache-capable upstreams, so summing + // only input+output under-reports prompt_tokens — measured: 2012 reported + // where the real prompt was ~5344 with 5332 served from cache. Fold the cache + // counters in, and keep them visible in prompt_tokens_details so a client can + // tell a cache hit from a small prompt. + const cacheRead = usage.cache_read_input_tokens || usage.cached_tokens || 0; + const cacheCreate = usage.cache_creation_input_tokens || 0; + const inTokens = (usage.input_tokens || 0) + cacheRead + cacheCreate; const outTokens = usage.output_tokens || 0; + const cacheDetails = (cacheRead > 0 || cacheCreate > 0) + ? { + prompt_tokens_details: { + ...(cacheRead > 0 ? { cached_tokens: cacheRead } : {}), + ...(cacheCreate > 0 ? { cache_creation_tokens: cacheCreate } : {}) } } + : {}; let finalResp; // Extract tool calls from Responses API output (function_call items) @@ -309,6 +399,7 @@ export async function handleForcedSSEToJson({ prompt_tokens: inTokens, completion_tokens: outTokens, total_tokens: inTokens + outTokens, + ...cacheDetails, }, }; } @@ -384,23 +475,14 @@ export async function handleForcedSSEToJson({ }), ); - const totalLatency = Date.now() - requestStartTime; - saveRequestDetail( - buildRequestDetail( - { - ...ctx, - latency: { ttft: totalLatency, total: totalLatency }, - tokens: usage, - response: { - content: parsed.choices?.[0]?.message?.content || null, - thinking: parsed.choices?.[0]?.message?.reasoning_content || null, - finish_reason: parsed.choices?.[0]?.finish_reason || "unknown", - }, - status: "success", - }, - { endpoint: clientRawRequest?.endpoint || null }, - ), - ).catch(() => {}); + // Re-attach usage explicitly. This handler already HAS the correct usage — it is + // the same object written to the usage DB, and for a cached Claude request that DB + // row reads cache_read_input_tokens: 11022 — yet the client was observed receiving + // no usage field at all (verified 2026-08-04 with a fingerprinted payload matched + // on both sides). Whatever drops it between assembly and serialisation, the client + // must not be left unable to account for its own token spend: a caller cannot tell + // a 90%-cached request from a cheap one without this. + if (usage && Object.keys(usage).length > 0) parsed.usage = usage; // Strip reasoning_content only when content is non-empty. // When content is empty (e.g. thinking models that used all tokens for reasoning), @@ -414,9 +496,19 @@ export async function handleForcedSSEToJson({ } } + // A Responses-format client (e.g. Codex) forced this provider to stream, + // but wants JSON back. parseSSEToOpenAIResponse yields a Chat Completions + // body; convert it to the Responses `output` shape so tool_calls are not + // lost on the non-streaming return path. Inlined (not imported from + // nonStreamingHandler.js) to avoid a circular import: nonStreamingHandler + // already imports parseSSEToOpenAIResponse from this module. + const finalBody = sourceFormat === FORMATS.OPENAI_RESPONSES + ? chatCompletionToResponses(parsed, customToolNames) + : parsed; + return { success: true, - response: new Response(JSON.stringify(parsed), { + response: new Response(JSON.stringify(finalBody), { headers: { "Content-Type": "application/json", "Access-Control-Allow-Origin": "*", diff --git a/open-sse/handlers/chatCore/streamingHandler.js b/open-sse/handlers/chatCore/streamingHandler.js index d37582bd..1c6b294f 100644 --- a/open-sse/handlers/chatCore/streamingHandler.js +++ b/open-sse/handlers/chatCore/streamingHandler.js @@ -38,6 +38,7 @@ function buildTransformStream({ userAgent, reqLogger, toolNameMap, + customToolNames, model, connectionId, body, @@ -68,6 +69,7 @@ function buildTransformStream({ body, onStreamComplete, apiKey, + customToolNames, ); } @@ -83,6 +85,7 @@ function buildTransformStream({ body, onStreamComplete, apiKey, + customToolNames, ); } @@ -118,6 +121,7 @@ export async function handleStreamingResponse({ onRequestSuccess, reqLogger, toolNameMap, + customToolNames, streamController, onStreamComplete, streamDetailId, @@ -200,6 +204,7 @@ export async function handleStreamingResponse({ userAgent, reqLogger, toolNameMap, + customToolNames, model, connectionId, body, diff --git a/open-sse/handlers/embeddingProviders/index.js b/open-sse/handlers/embeddingProviders/index.js index 62e52778..9dd6a604 100644 --- a/open-sse/handlers/embeddingProviders/index.js +++ b/open-sse/handlers/embeddingProviders/index.js @@ -2,6 +2,7 @@ import createOpenAIEmbeddingAdapter from "./openai.js"; import gemini from "./gemini.js"; import openaiCompatNode from "./openaiCompatNode.js"; +import selfhostedEmbedding from "./selfhostedEmbedding.js"; const OPENAI_COMPAT_PROVIDERS = [ "openai", "openrouter", "mistral", "voyage-ai", "fireworks", @@ -13,6 +14,12 @@ const ADAPTERS = { ...Object.fromEntries(OPENAI_COMPAT_PROVIDERS.map((id) => [id, createOpenAIEmbeddingAdapter(id)])), gemini, google_ai_studio: gemini, + // Self-hosted reads creds.providerSpecificData.baseUrl (one provider, many + // servers) — but via its OWN adapter, not openaiCompatNode: that one falls back + // to api.openai.com when no baseUrl is set, which under a provider called + // "Self-hosted Embedding" means silently shipping the input and API key to + // OpenAI. selfhostedEmbedding refuses instead. + "selfhosted-embedding": selfhostedEmbedding, }; export function getEmbeddingAdapter(provider) { diff --git a/open-sse/handlers/embeddingProviders/selfhostedEmbedding.js b/open-sse/handlers/embeddingProviders/selfhostedEmbedding.js new file mode 100644 index 00000000..aa030439 --- /dev/null +++ b/open-sse/handlers/embeddingProviders/selfhostedEmbedding.js @@ -0,0 +1,46 @@ +// Self-hosted embeddings — like openaiCompatNode, but the baseUrl is REQUIRED. +// +// openaiCompatNode falls back to https://api.openai.com/v1 when a connection +// carries no providerSpecificData.baseUrl. For a custom NODE that default is +// defensible: the node was created by pointing at some OpenAI-compatible URL, and +// OpenAI is the archetype. For a provider whose entire purpose is "my own +// server", it is actively harmful — a connection saved without a baseUrl sends +// the INPUT TEXT and the API KEY to OpenAI, silently, under a provider named +// "Self-hosted Embedding". +// +// Observed exactly that with a placeholder connection (2026-08-04): +// +// [selfhosted-embedding/embedding] [401]: Incorrect API key provided: abc. +// You can find your API key at https://platform.openai.com/account/api-keys. +// +// The key "abc" was typed as a throwaway for a LOCAL server and left the network. +// A self-hosted provider must never have a cloud fallback, so this one refuses +// instead: no baseUrl means a configuration error, reported as such. +import createOpenAIEmbeddingAdapter from "./openai.js"; + +const baseAdapter = createOpenAIEmbeddingAdapter("openai"); + +export class MissingBaseUrlError extends Error { + constructor() { + super( + "Self-hosted Embedding needs an endpoint: set this connection's baseUrl to " + + "the OpenAI base URL of your server, e.g. http://host:8080/v1 (note the /v1 — " + + "\"/embeddings\" is appended to it). Refusing to fall back to api.openai.com, " + + "which would send your input and API key to OpenAI." + ); + this.name = "MissingBaseUrlError"; + this.isConfigError = true; + } +} + +export default { + ...baseAdapter, + buildUrl: (_model, creds) => { + const rawBaseUrl = creds?.providerSpecificData?.baseUrl; + if (!rawBaseUrl || !String(rawBaseUrl).trim()) throw new MissingBaseUrlError(); + // Accept either the OpenAI base or a full embeddings URL, so a value pasted + // from a curl example works as well as one typed from the help text. + const baseUrl = String(rawBaseUrl).trim().replace(/\/$/, "").replace(/\/embeddings$/, ""); + return `${baseUrl}/embeddings`; + }, +}; diff --git a/open-sse/handlers/embeddingsCore.js b/open-sse/handlers/embeddingsCore.js index aa81117c..1ee0f4cd 100644 --- a/open-sse/handlers/embeddingsCore.js +++ b/open-sse/handlers/embeddingsCore.js @@ -1,5 +1,5 @@ import { createErrorResult, parseUpstreamError, formatProviderError } from "../utils/error.js"; -import { HTTP_STATUS } from "../config/runtimeConfig.js"; +import { HTTP_STATUS, FETCH_CONNECT_TIMEOUT_MS } from "../config/runtimeConfig.js"; import { getExecutor } from "../executors/index.js"; import { refreshWithRetry } from "../services/tokenRefresh.js"; import { getEmbeddingAdapter } from "./embeddingProviders/index.js"; @@ -38,13 +38,24 @@ export async function handleEmbeddingsCore({ } const ctx = { input }; - const url = adapter.buildUrl(model, credentials, ctx); - const headers = adapter.buildHeaders(credentials, ctx); - const requestBody = adapter.buildBody(model, { - input, - encoding_format: body.encoding_format || "float", - dimensions: body.dimensions, - }); + // buildUrl/buildHeaders/buildBody were called bare. An adapter that rejects a + // misconfigured connection — selfhosted-embedding throws when no baseUrl is set + // rather than silently falling back to api.openai.com — would have escaped this + // function uncaught, surfacing as a 500 or a request that never settles. A + // configuration mistake is a 400 with the reason in it. + let url, headers, requestBody; + try { + url = adapter.buildUrl(model, credentials, ctx); + headers = adapter.buildHeaders(credentials, ctx); + requestBody = adapter.buildBody(model, { + input, + encoding_format: body.encoding_format || "float", + dimensions: body.dimensions, + }); + } catch (error) { + log?.debug?.("EMBEDDINGS", `Request build failed: ${error.message}`); + return createErrorResult(HTTP_STATUS.BAD_REQUEST, `[${provider}/${model}] ${error.message}`); + } log?.debug?.("EMBEDDINGS", `${provider.toUpperCase()} | ${model} | input_type=${Array.isArray(input) ? `array[${input.length}]` : "string"}`); @@ -54,6 +65,9 @@ export async function handleEmbeddingsCore({ method: "POST", headers, body: JSON.stringify(requestBody), + ...(typeof AbortSignal?.timeout === "function" + ? { signal: AbortSignal.timeout(FETCH_CONNECT_TIMEOUT_MS) } + : {}), }); } catch (error) { const errMsg = formatProviderError(error, provider, model, HTTP_STATUS.BAD_GATEWAY); diff --git a/open-sse/handlers/sttCore.js b/open-sse/handlers/sttCore.js index acb7d13b..8127782a 100644 --- a/open-sse/handlers/sttCore.js +++ b/open-sse/handlers/sttCore.js @@ -170,9 +170,17 @@ export async function handleSttCore({ provider, model, formData, credentials, st const file = formData.get("file"); if (!file) return createErrorResult(HTTP_STATUS.BAD_REQUEST, "Missing required field: file"); - const cfg = sttConfig; + let cfg = sttConfig; if (!cfg) return createErrorResult(HTTP_STATUS.BAD_REQUEST, `Provider '${provider}' does not support STT`); + // Per-connection endpoint override. Registry entries carry a fixed baseUrl, + // which is right for a named cloud service but useless for a self-hosted one + // whose address only the operator knows. Opt-in: absent unless the connection + // sets it, so cloud providers are untouched. Mirrors the custom embedding + // providers, which already resolve baseUrl the same way. + const overrideUrl = credentials?.providerSpecificData?.baseUrl; + if (overrideUrl) cfg = { ...cfg, baseUrl: String(overrideUrl).replace(/\/+$/, "") }; + const token = cfg.authType === "none" ? null : (credentials?.apiKey || credentials?.accessToken); if (cfg.authType !== "none" && !token) { return createErrorResult(HTTP_STATUS.UNAUTHORIZED, `No credentials for STT provider: ${provider}`); diff --git a/open-sse/handlers/ttsCore.js b/open-sse/handlers/ttsCore.js index b4b69eeb..e074eb96 100644 --- a/open-sse/handlers/ttsCore.js +++ b/open-sse/handlers/ttsCore.js @@ -48,16 +48,16 @@ function createTtsResponse(base64Audio, format, responseFormat) { * * @returns {Promise<{success, response, status?, error?}>} */ -export async function handleTtsCore({ provider, model, input, credentials, responseFormat = "mp3", language }) { +export async function handleTtsCore({ provider, model, input, credentials, responseFormat = "mp3", language, style }) { if (!input?.trim()) { return createErrorResult(HTTP_STATUS.BAD_REQUEST, "Missing required field: input"); } try { - // Special-case adapters (google-tts, edge-tts, local-device, elevenlabs, openai, openrouter, gemini) + // Special-case adapters (google-tts, edge-tts, local-device, elevenlabs, openai, openrouter, gemini, xiaomi-mimo) const adapter = getTtsAdapter(provider); if (adapter) { - const result = await adapter.synthesize(input.trim(), model, credentials, responseFormat, { language }); + const result = await adapter.synthesize(input.trim(), model, credentials, responseFormat, { language, style }); // Adapter may return a full {success, response} (legacy) or {base64, format} if (result.success !== undefined) return result; return createTtsResponse(result.base64, result.format, responseFormat); diff --git a/open-sse/handlers/ttsProviders/index.js b/open-sse/handlers/ttsProviders/index.js index e1bb8b83..d2be03fa 100644 --- a/open-sse/handlers/ttsProviders/index.js +++ b/open-sse/handlers/ttsProviders/index.js @@ -6,6 +6,8 @@ import elevenlabs, { fetchElevenLabsVoices } from "./elevenlabs.js"; import openai from "./openai.js"; import openrouter from "./openrouter.js"; import gemini, { fetchGeminiVoices } from "./gemini.js"; +import xiaomiMimo from "./xiaomi-mimo.js"; +import selfhostedTts from "./selfhostedTts.js"; import { FORMAT_HANDLERS } from "./genericFormats.js"; import { parseModelVoice } from "./_base.js"; @@ -18,6 +20,8 @@ const SPECIAL_ADAPTERS = { openai, openrouter, gemini, + "xiaomi-mimo": xiaomiMimo, + "selfhosted-tts": selfhostedTts, }; export function getTtsAdapter(provider) { diff --git a/open-sse/handlers/ttsProviders/selfhostedTts.js b/open-sse/handlers/ttsProviders/selfhostedTts.js new file mode 100644 index 00000000..fdfdbf76 --- /dev/null +++ b/open-sse/handlers/ttsProviders/selfhostedTts.js @@ -0,0 +1,69 @@ +// Self-hosted OpenAI-compatible TTS — POST {baseUrl}/v1/audio/speech. +// +// A SPECIAL_ADAPTER rather than a genericFormats handler on purpose: the generic +// dispatcher resolves baseUrl from the static registry entry +// (`synthesizeViaConfig` reads `cfg.baseUrl`) and never looks at the connection, +// which is exactly the limitation this provider exists to lift. +import { Buffer } from "node:buffer"; + +const DEFAULT_BASE_URL = "http://localhost:8880"; +const DEFAULT_MODEL = "kokoro"; +const DEFAULT_VOICE = "af_heart"; + +export default { + async synthesize(text, model, credentials, responseFormat = "mp3") { + // Accept either providerSpecificData.baseUrl (how the custom embedding and + // STT providers carry it) or a bare credentials.baseUrl (how the OpenAI TTS + // adapter does), so a connection configured either way works. + const raw = credentials?.providerSpecificData?.baseUrl || credentials?.baseUrl || DEFAULT_BASE_URL; + // Tolerate a baseUrl given as the full endpoint or with a trailing /v1 — + // both are natural things to paste, and silently double-appending the path + // would 404 with nothing pointing at the cause. + const base = String(raw) + .replace(/\/+$/, "") + .replace(/\/v1\/audio\/speech$/, "") + .replace(/\/v1$/, ""); + + // The provider prefix is already stripped by getModelInfo, so `model` here is + // "kokoro" or "kokoro/af_heart" — NOT "selfhosted-tts/...". + // + // A bare value is the MODEL, not the voice. The OpenAI adapter reads a bare + // value as a voice, which is right for a service whose model is fixed + // ("tts-1") and whose voice varies — but wrong here, where the model is the + // variable part. Treating it as a voice sent voice="kokoro" upstream and + // Kokoro answered 400, so `selfhosted-tts/kokoro` — the obvious way to + // address this provider — was the one form that did not work (verified + // against a live Kokoro through 9router, 2026-08-03). + let ttsModel = DEFAULT_MODEL; + let voice = DEFAULT_VOICE; + if (model) { + const parts = String(model).split("/").filter(Boolean); + if (parts.length >= 2) { + ttsModel = parts[0]; + voice = parts.slice(1).join("/"); + } else if (parts.length === 1) { + ttsModel = parts[0]; + } + } + + const res = await fetch(`${base}/v1/audio/speech`, { + method: "POST", + headers: { + "Content-Type": "application/json", + ...(credentials?.apiKey ? { Authorization: `Bearer ${credentials.apiKey}` } : {}), + }, + body: JSON.stringify({ + model: ttsModel, + voice, + input: text, + response_format: responseFormat, + }), + }); + if (!res.ok) { + const err = await res.json().catch(() => ({})); + throw new Error(err?.error?.message || `Self-hosted TTS failed: ${res.status}`); + } + const buf = await res.arrayBuffer(); + return { base64: Buffer.from(buf).toString("base64"), format: responseFormat }; + }, +}; diff --git a/open-sse/handlers/ttsProviders/xiaomi-mimo.js b/open-sse/handlers/ttsProviders/xiaomi-mimo.js new file mode 100644 index 00000000..46682689 --- /dev/null +++ b/open-sse/handlers/ttsProviders/xiaomi-mimo.js @@ -0,0 +1,65 @@ +// Xiaomi MiMo TTS — via OpenAI-compatible chat completions (non-streaming). +// Docs: https://mimo.mi.com/docs/zh-CN/quick-start/usage-guide/audio/speech-synthesis-v2.5 +// Message contract: target text in `role: assistant` content, style/voice +// instructions in `role: user` content. Voice is selected via the top-level +// `audio.voice` field (NOT embedded in the model name). +import { parseModelVoice } from "./_base.js"; + +const DEFAULT_MODEL = "mimo-v2.5-tts"; +const DEFAULT_VOICE = "mimo_default"; + +export default { + synthesize(text, model, credentials, responseFormat, { style, language } = {}) { + if (!credentials?.apiKey) throw new Error("xiaomi-mimo API key required"); + return synthesizeMiMo(text, model, credentials.apiKey, style, language); + }, +}; + +export async function synthesizeMiMo(text, model, apiKey, style, language) { + const { modelId, voiceId } = parseModelVoice(model, DEFAULT_MODEL, DEFAULT_VOICE, [DEFAULT_MODEL]); + + // Language and style are soft instructions → prepend as a role:user message. + // MiMo auto-detects the spoken language of the text; the hint only nudges it + // (e.g. "Speak in English.") and is independent of the chosen voice. + const instructions = []; + if (language) instructions.push(`Speak in ${language}.`); + if (style) instructions.push(style); + + const messages = [{ role: "assistant", content: text }]; + if (instructions.length) messages.unshift({ role: "user", content: instructions.join(" ") }); + + const res = await fetch("https://api.xiaomimimo.com/v1/chat/completions", { + method: "POST", + headers: { + "Content-Type": "application/json", + "Authorization": `Bearer ${apiKey}`, + }, + body: JSON.stringify({ + model: modelId, + stream: false, + messages, + audio: { + format: "wav", + voice: voiceId || DEFAULT_VOICE, + }, + }), + }); + + const rawText = await res.text(); + let data = {}; + if (rawText) { + try { data = JSON.parse(rawText); } catch { data = {}; } + } + + if (!res.ok) { + throw new Error(data?.error?.message || rawText || `MiMo TTS error (${res.status})`); + } + + const audio = data?.choices?.[0]?.message?.audio?.data; + if (!audio) throw new Error(data?.error?.message || "MiMo TTS returned no audio"); + + return { + base64: audio, + format: data?.choices?.[0]?.message?.audio?.format || "wav", + }; +} diff --git a/open-sse/index.js b/open-sse/index.js index b8181f0b..c0694580 100644 --- a/open-sse/index.js +++ b/open-sse/index.js @@ -47,7 +47,6 @@ export { refreshAccessToken, refreshClaudeOAuthToken, refreshGoogleToken, - refreshQwenToken, refreshCodexToken, refreshIflowToken, refreshGitHubToken, diff --git a/open-sse/providers/capabilities.js b/open-sse/providers/capabilities.js index 526ceb9b..e91ad1f2 100644 --- a/open-sse/providers/capabilities.js +++ b/open-sse/providers/capabilities.js @@ -279,7 +279,7 @@ export const PATTERN_CAPABILITIES = [ { pattern: "*minimax*", caps: { reasoning: true, thinkingFormat: "minimax", thinkingCanDisable: false, contextWindow: 200000, maxOutput: 131072 } }, // ── Xiaomi MiMo (vision, 1M / 262K ctx) ────────────────────────── - { pattern: "*mimo*v2.5*", caps: { vision: true, contextWindow: 1048576, maxOutput: 131072 } }, + { pattern: "*mimo*v2.5*", caps: { vision: true, audioInput: true, videoInput: true, contextWindow: 1048576, maxOutput: 131072 } }, { pattern: "*mimo*omni*", caps: { vision: true, audioInput: true, contextWindow: 262144, maxOutput: 131072 } }, { pattern: "*mimo*", caps: { vision: true, contextWindow: 262144, maxOutput: 131072 } }, diff --git a/open-sse/providers/pricing.js b/open-sse/providers/pricing.js index 9a0768ef..e9f2b612 100644 --- a/open-sse/providers/pricing.js +++ b/open-sse/providers/pricing.js @@ -141,6 +141,122 @@ export const PROVIDER_PRICING = { gh: { "gpt-5.3-codex": { input: 1.75, output: 14.00, cached: 0.175, reasoning: 14.00, cache_creation: 1.75 }, }, + // TokenRouter — exact rates from https://api.tokenrouter.com/api/pricing ($1/1M tokens). + // Ratio→USD: input = model_ratio×2, output = model_ratio×completion_ratio×2. + // These override the canonical MODEL_PRICING/PATTERN_PRICING, whose rates often + // differ from TokenRouter's reseller pricing. + tokenrouter: { + "MiniMax-M3": { input: 0.3, output: 1.2, cached: 0.06, reasoning: 1.2 }, + "anthropic/claude-fable-5": { input: 10, output: 50, cached: 1.0, cache_creation: 12.5, reasoning: 50 }, + "anthropic/claude-haiku-4.5": { input: 1.0, output: 5.0, cached: 0.1, cache_creation: 1.25, reasoning: 5.0 }, + "anthropic/claude-opus-4.5": { input: 5.0, output: 25.0, cached: 0.5, cache_creation: 6.25, reasoning: 25.0 }, + "anthropic/claude-opus-4.6": { input: 5.0, output: 25.0, cached: 0.5, cache_creation: 6.25, reasoning: 25.0 }, + "anthropic/claude-opus-4.7": { input: 5.0, output: 25.0, cached: 0.5, cache_creation: 6.25, reasoning: 25.0 }, + "anthropic/claude-opus-4.7-fast": { input: 30, output: 150, cached: 3.0, reasoning: 150 }, + "anthropic/claude-opus-4.8": { input: 5.0, output: 25.0, cached: 0.5, cache_creation: 6.25, reasoning: 25.0 }, + "anthropic/claude-opus-4.8-fast": { input: 10, output: 50, cached: 1.0, cache_creation: 12.5, reasoning: 50 }, + "anthropic/claude-opus-5": { input: 5.0, output: 25.0, cached: 0.5, cache_creation: 6.25, reasoning: 25.0 }, + "anthropic/claude-opus-5-fast": { input: 10, output: 50, cached: 1.0, cache_creation: 12.5, reasoning: 50 }, + "anthropic/claude-sonnet-4": { input: 3.0, output: 15.0, cached: 0.3, cache_creation: 3.75, reasoning: 15.0 }, + "anthropic/claude-sonnet-4.5": { input: 3.0, output: 15.0, cached: 0.3, cache_creation: 3.75, reasoning: 15.0 }, + "anthropic/claude-sonnet-4.6": { input: 3.0, output: 15.0, cached: 0.3, cache_creation: 3.75, reasoning: 15.0 }, + "anthropic/claude-sonnet-5": { input: 2, output: 10, cached: 0.2, reasoning: 10 }, + "claude-opus-4-8-m-aws": { input: 5.0, output: 25.0, cached: 0.5, cache_creation: 6.25, reasoning: 25.0 }, + "deepseek/deepseek-v3.2": { input: 0.26, output: 0.38, cached: 0.13, reasoning: 0.38 }, + "deepseek/deepseek-v4-flash": { input: 0.14, output: 0.28, cached: 0.0028, reasoning: 0.28 }, + "deepseek/deepseek-v4-flash-0731": { input: 0.14, output: 0.28, cached: 0.0028, reasoning: 0.28 }, + "deepseek/deepseek-v4-pro": { input: 0.435, output: 0.87, cached: 0.003625, reasoning: 0.87 }, + "ex/gpt-5.4": { input: 2.5, output: 15.0, cached: 0.25, reasoning: 15.0 }, + "google/gemini-2.5-flash-image": { input: 0.3, output: 2.5, reasoning: 2.5 }, + "google/gemini-3-flash-preview": { input: 0.5, output: 3.0, cached: 0.05, cache_creation: 0.08333, reasoning: 3.0 }, + "google/gemini-3-pro-image-preview": { input: 2, output: 12, reasoning: 12 }, + "google/gemini-3.1-flash-image-preview": { input: 0.5, output: 3.0, reasoning: 3.0 }, + "google/gemini-3.1-flash-lite-image": { input: 0.25, output: 1.5, reasoning: 1.5 }, + "google/gemini-3.1-pro-preview": { input: 2, output: 12, cached: 0.2, cache_creation: 0.375, reasoning: 12 }, + "google/gemini-3.5-flash": { input: 1.5, output: 9.0, cached: 0.15, cache_creation: 0.08333, reasoning: 9.0 }, + "google/gemini-3.5-flash-lite": { input: 0.3, output: 2.5, cached: 0.03, cache_creation: 0.08333, reasoning: 2.5 }, + "google/gemini-3.6-flash": { input: 1.5, output: 7.5, cached: 0.15, cache_creation: 0.08333, reasoning: 7.5 }, + "google/gemini-embedding-2": { input: 1.0, output: 6.0, cached: 0.1, reasoning: 6.0 }, + "google/gemma-4-26b-a4b-it": { input: 0.06, output: 0.33, reasoning: 0.33 }, + "kling-3.0-turbo": { input: 2.1, output: 2.1, reasoning: 2.1 }, + "microsoft/mai-image-2.5": { input: 5.0, output: 47.0, reasoning: 47.0 }, + "minimax/minimax-m2-her": { input: 0.3, output: 1.2, cached: 0.03, reasoning: 1.2 }, + "minimax/minimax-m2.1": { input: 0.3, output: 1.2, cached: 0.03, reasoning: 1.2 }, + "minimax/minimax-m2.1-highspeed": { input: 0.6, output: 2.4, cached: 0.06, reasoning: 2.4 }, + "minimax/minimax-m2.5": { input: 0.3, output: 1.2, cached: 0.03, reasoning: 1.2 }, + "minimax/minimax-m2.7": { input: 0.3, output: 1.2, cached: 0.06, reasoning: 1.2 }, + "minimax/minimax-m2.7-highspeed": { input: 0.6, output: 2.4, cached: 0.06, reasoning: 2.4 }, + "miromind/mirothinker-1-7-deepresearch": { input: 4, output: 25.0, reasoning: 25.0 }, + "miromind/mirothinker-1-7-deepresearch-mini": { input: 1.25, output: 10.0, reasoning: 10.0 }, + "mistralai/devstral-2512": { input: 0.4, output: 2.0, cached: 0.04, reasoning: 2.0 }, + "mistralai/mistral-medium-3-5": { input: 1.5, output: 7.5, reasoning: 7.5 }, + "mistralai/mistral-small-2603": { input: 0.15, output: 0.6, cached: 0.015, reasoning: 0.6 }, + "mistralai/voxtral-small-24b-2507": { input: 0.1, output: 0.3, cached: 0.01, reasoning: 0.3 }, + "moonshotai/kimi-k2.5": { input: 0.6, output: 3.0, cached: 0.1, reasoning: 3.0 }, + "moonshotai/kimi-k2.6": { input: 0.95, output: 4.0, cached: 0.16, reasoning: 4.0 }, + "moonshotai/kimi-k2.7-code": { input: 0.9286, output: 3.8571, cached: 0.1857, reasoning: 3.8571 }, + "moonshotai/kimi-k3": { input: 3.0, output: 15.0, cached: 0.3, reasoning: 15.0 }, + "nvidia/nemotron-3-super-120b-a12b": { input: 0.3, output: 0.9, cached: 0.1, reasoning: 0.9 }, + "openai/gpt-4o-mini": { input: 0.15, output: 0.6, cached: 0.075, reasoning: 0.6 }, + "openai/gpt-5": { input: 1.25, output: 10.0, cached: 0.125, reasoning: 10.0 }, + "openai/gpt-5-image": { input: 10, output: 40, cached: 2.5, reasoning: 40 }, + "openai/gpt-5-image-mini": { input: 2.5, output: 8.0, cached: 0.25, reasoning: 8.0 }, + "openai/gpt-5-mini": { input: 0.25, output: 2.0, cached: 0.025, reasoning: 2.0 }, + "openai/gpt-5.2": { input: 1.75, output: 14.0, cached: 0.175, reasoning: 14.0 }, + "openai/gpt-5.3-codex": { input: 1.75, output: 14.0, cached: 0.175, reasoning: 14.0 }, + "openai/gpt-5.4": { input: 2.5, output: 15.0, cached: 0.25, reasoning: 15.0 }, + "openai/gpt-5.4-image-2": { input: 8, output: 30.0, cached: 2.0, reasoning: 30.0 }, + "openai/gpt-5.4-mini": { input: 0.75, output: 4.5, cached: 0.075, reasoning: 4.5 }, + "openai/gpt-5.4-nano": { input: 0.2, output: 1.25, cached: 0.02, reasoning: 1.25 }, + "openai/gpt-5.4-pro": { input: 30, output: 180, reasoning: 180 }, + "openai/gpt-5.5": { input: 5.0, output: 30.0, cached: 0.5, reasoning: 30.0 }, + "openai/gpt-5.5-pro": { input: 30, output: 180, reasoning: 180 }, + "openai/gpt-5.6-luna": { input: 0.2, output: 1.2, cached: 0.02, cache_creation: 0.25, reasoning: 1.2 }, + "openai/gpt-5.6-sol": { input: 5.0, output: 30.0, cached: 0.5, cache_creation: 6.25, reasoning: 30.0 }, + "openai/gpt-5.6-terra": { input: 2, output: 12, cached: 0.2, cache_creation: 2.5, reasoning: 12 }, + "openai/gpt-audio": { input: 2.5, output: 10.0, reasoning: 10.0 }, + "openai/gpt-audio-mini": { input: 0.6, output: 2.4, reasoning: 2.4 }, + "openai/gpt-oss-120b": { input: 0.039, output: 0.18, reasoning: 0.18 }, + "qwen/qwen3-coder-next": { input: 0.12, output: 0.75, cached: 0.06, reasoning: 0.75 }, + "qwen/qwen3.5-122b-a10b": { input: 0.26, output: 2.08, reasoning: 2.08 }, + "qwen/qwen3.5-35b-a3b": { input: 0.1625, output: 1.3, reasoning: 1.3 }, + "qwen/qwen3.5-397b-a17b": { input: 0.39, output: 2.34, reasoning: 2.34 }, + "qwen/qwen3.5-9b": { input: 0.1, output: 0.15, reasoning: 0.15 }, + "qwen/qwen3.5-flash": { input: 0.1048, output: 0.4194, reasoning: 0.4194 }, + "qwen/qwen3.5-plus-02-15": { input: 0.26, output: 1.56, reasoning: 1.56 }, + "qwen/qwen3.6-plus": { input: 0.54, output: 3.21, reasoning: 3.21 }, + "qwen/qwen3.7-max": { input: 1.25, output: 3.75, cached: 0.25, reasoning: 3.75 }, + "qwen/qwen3.7-plus": { input: 0.4, output: 1.6, cached: 0.08, reasoning: 1.6 }, + "qwen/qwen3.8-max": { input: 2, output: 6, cached: 0.25, cache_creation: 2.5, reasoning: 6 }, + "qwen3.5-omni-plus": { input: 1.0, output: 5.7143, reasoning: 5.7143 }, + "qwen3.6-flash": { input: 0.171, output: 1.029, cached: 0.017, cache_creation: 0.214, reasoning: 1.029 }, + "sakana/fugu-ultra": { input: 5.0, output: 30.0, cached: 0.5, reasoning: 30.0 }, + "seed-2-0-code-preview-260328": { input: 1.0, output: 6.0, cached: 0.2, cache_creation: 0.008333, reasoning: 6.0 }, + "seed-2-0-lite-260428": { input: 0.5, output: 4.0, cached: 0.1, cache_creation: 0.008333, reasoning: 4.0 }, + "seed-2-0-mini-260428": { input: 0.2, output: 0.8, cached: 0.04, cache_creation: 0.00833, reasoning: 0.8 }, + "seed-2-0-pro-260328": { input: 1.0, output: 6.0, cached: 0.2, cache_creation: 0.008333, reasoning: 6.0 }, + "stepfun/step-3.5-flash": { input: 0.1, output: 0.3, cached: 0.02, reasoning: 0.3 }, + "stepfun/step-3.7-flash": { input: 0.2, output: 1.15, cached: 0.04, reasoning: 1.15 }, + "tencent/hy3-preview": { input: 0.066, output: 0.26, cached: 0.029, reasoning: 0.26 }, + "x-ai/grok-4.1-fast": { input: 0.2, output: 0.5, cached: 0.05, reasoning: 0.5 }, + "x-ai/grok-4.20-beta": { input: 2, output: 6, cached: 0.2, reasoning: 6 }, + "x-ai/grok-4.3": { input: 1.25, output: 2.5, cached: 0.2, reasoning: 2.5 }, + "x-ai/grok-4.5": { input: 2, output: 6, cached: 0.5, reasoning: 6 }, + "x-ai/grok-build-0.1": { input: 1.0, output: 2.0, cached: 0.2, reasoning: 2.0 }, + "xiaomi/mimo-v2-flash": { input: 0.1, output: 0.3, cached: 0.01, reasoning: 0.3 }, + "xiaomi/mimo-v2-omni": { input: 0.4, output: 2.0, cached: 0.08, reasoning: 2.0 }, + "xiaomi/mimo-v2-pro": { input: 1.0, output: 3.0, cached: 0.2, reasoning: 3.0 }, + "xiaomi/mimo-v2.5": { input: 0.4, output: 2.0, cached: 0.08, reasoning: 2.0 }, + "xiaomi/mimo-v2.5-pro": { input: 1.0, output: 3.0, cached: 0.2, reasoning: 3.0 }, + "z-ai/glm-4.5-air": { input: 0.13, output: 0.85, cached: 0.025, reasoning: 0.85 }, + "z-ai/glm-4.6": { input: 0.6, output: 2.2, cached: 0.11, reasoning: 2.2 }, + "z-ai/glm-4.6v": { input: 0.3, output: 0.9, reasoning: 0.9 }, + "z-ai/glm-4.7": { input: 0.6, output: 2.2, cached: 0.11, reasoning: 2.2 }, + "z-ai/glm-5": { input: 1.0, output: 3.2, cached: 0.2, reasoning: 3.2 }, + "z-ai/glm-5-turbo": { input: 1.2, output: 4.0, cached: 0.24, reasoning: 4.0 }, + "z-ai/glm-5.1": { input: 1.05, output: 3.5, cached: 0.525, reasoning: 3.5 }, + "z-ai/glm-5.2": { input: 1.4, output: 4.4, cached: 0.26, reasoning: 4.4 }, + }, }; /** diff --git a/open-sse/providers/registry/antigravity.js b/open-sse/providers/registry/antigravity.js index a4ca7346..1a7c4347 100644 --- a/open-sse/providers/registry/antigravity.js +++ b/open-sse/providers/registry/antigravity.js @@ -76,8 +76,7 @@ export default { apiVersion: "v1internal", loadCodeAssistEndpoint: "https://cloudcode-pa.googleapis.com/v1internal:loadCodeAssist", onboardUserEndpoint: "https://cloudcode-pa.googleapis.com/v1internal:onboardUser", - loadCodeAssistUserAgent: "google-api-nodejs-client/9.15.1", - loadCodeAssistApiClient: "google-cloud-sdk vscode_cloudshelleditor/0.1", + loadCodeAssistUserAgent: ANTIGRAVITY_IDE_USER_AGENT, refreshLeadMs: 300000, }, features: { diff --git a/open-sse/providers/registry/claude.js b/open-sse/providers/registry/claude.js index 8b030e5a..8c47a6b3 100644 --- a/open-sse/providers/registry/claude.js +++ b/open-sse/providers/registry/claude.js @@ -49,9 +49,6 @@ export default { header: "Authorization", scheme: "bearer", }, - hooks: [ - "claudeOverlay", - ], }, usage: { oauthUrl: "https://api.anthropic.com/api/oauth/usage", diff --git a/open-sse/providers/registry/cloudflare-ai.js b/open-sse/providers/registry/cloudflare-ai.js index 4d440b74..53f1a697 100644 --- a/open-sse/providers/registry/cloudflare-ai.js +++ b/open-sse/providers/registry/cloudflare-ai.js @@ -19,6 +19,8 @@ export default { }, }, category: "freeTier", + authType: "apikey", + authModes: ["apikey"], hasProviderSpecificData: true, transport: { baseUrl: "https://api.cloudflare.com/client/v4/accounts/{accountId}/ai/v1/chat/completions", diff --git a/open-sse/providers/registry/codebuddy-intl.js b/open-sse/providers/registry/codebuddy-intl.js index 7e69836c..eab1ce93 100644 --- a/open-sse/providers/registry/codebuddy-intl.js +++ b/open-sse/providers/registry/codebuddy-intl.js @@ -38,6 +38,10 @@ export default { header: "Authorization", scheme: "bearer", }, + // Intl billing endpoint mirrors CN shape (data.Response.Data.Accounts[]). + usage: { + url: "https://www.codebuddy.ai/v2/billing/meter/get-user-resource", + }, }, // Same model lineup exposed by the CN gateway — intl backend is the same catalog. models: [ diff --git a/open-sse/providers/registry/index.js b/open-sse/providers/registry/index.js index 7bfdd0d7..467d1c4d 100644 --- a/open-sse/providers/registry/index.js +++ b/open-sse/providers/registry/index.js @@ -75,7 +75,6 @@ import p72 from "./perplexity.js"; import p73 from "./perplexity-agent.js"; import p74 from "./playht.js"; import p75 from "./qoder.js"; -import p76 from "./qwen.js"; import p77 from "./recraft.js"; import p78 from "./runwayml.js"; import p79 from "./sdwebui.js"; @@ -116,6 +115,10 @@ import p113 from "./morph.js"; // import p114 from "./devin-cli.js"; // import p104 from "./windsurf.js"; import p115 from "./poolside.js"; +import p116 from "./tokenrouter.js"; +import p117 from "./selfhosted-stt.js"; +import p118 from "./selfhosted-tts.js"; +import p119 from "./selfhosted-embedding.js"; export default [ p0, @@ -194,7 +197,6 @@ export default [ p73, p74, p75, - p76, p77, p78, p79, @@ -233,4 +235,8 @@ export default [ // p114, // devin-cli — hidden, spawns local agent with shell/fs access // p104, // windsurf — hidden, no tool calling p115, + p116, + p117, + p118, + p119, ]; diff --git a/open-sse/providers/registry/ollama.js b/open-sse/providers/registry/ollama.js index 69923aa1..89fec43c 100644 --- a/open-sse/providers/registry/ollama.js +++ b/open-sse/providers/registry/ollama.js @@ -15,6 +15,8 @@ export default { }, }, category: "freeTier", + authType: "apikey", + authModes: ["apikey"], transport: { baseUrl: "https://ollama.com/api/chat", validateUrl: "https://ollama.com/api/tags", @@ -32,5 +34,6 @@ export default { serviceKinds: ["llm"], features: { usage: true, + usageApikey: true, }, }; diff --git a/open-sse/providers/registry/qoder.js b/open-sse/providers/registry/qoder.js index 2b6c93ed..fe76fd72 100644 --- a/open-sse/providers/registry/qoder.js +++ b/open-sse/providers/registry/qoder.js @@ -52,5 +52,7 @@ export default { }, features: { usage: true, + // PAT (apikey) connections also carry quota usage (via job-token exchange). + usageApikey: true, }, }; diff --git a/open-sse/providers/registry/qwen.js b/open-sse/providers/registry/qwen.js deleted file mode 100644 index 0df381ab..00000000 --- a/open-sse/providers/registry/qwen.js +++ /dev/null @@ -1,33 +0,0 @@ -export default { - id: "qwen", - hidden: true, - priority: 130, - alias: "qw", - display: { - name: "Qwen Code", - icon: "psychology", - color: "#10B981", - website: "https://chat.qwen.ai", - notice: { - signupUrl: "https://chat.qwen.ai", - }, - }, - category: "oauth", - transport: { - baseUrl: "https://portal.qwen.ai/v1/chat/completions", - }, - models: [ - { id: "qwen3-coder-plus", name: "Qwen3 Coder Plus" }, - { id: "qwen3-coder-flash", name: "Qwen3 Coder Flash" }, - { id: "vision-model", name: "Qwen3 Vision Model" }, - { id: "coder-model", name: "Qwen3.6 Coder Model" }, - ], - oauth: { - clientId: "f0304373b74a44d2b584a3fb70ca9e56", - deviceCodeUrl: "https://chat.qwen.ai/api/v1/oauth2/device/code", - tokenUrl: "https://chat.qwen.ai/api/v1/oauth2/token", - scope: "openid profile email model.completion", - codeChallengeMethod: "S256", - refreshLeadMs: 1200000, - }, -}; diff --git a/open-sse/providers/registry/selfhosted-embedding.js b/open-sse/providers/registry/selfhosted-embedding.js new file mode 100644 index 00000000..90fa79c7 --- /dev/null +++ b/open-sse/providers/registry/selfhosted-embedding.js @@ -0,0 +1,73 @@ +// Self-hosted, OpenAI-compatible embeddings (llama.cpp / llama-server, vLLM, +// Infinity, text-embeddings-inference, ...) — the embeddings counterpart of +// selfhosted-stt and selfhosted-tts. +// +// Routing a self-hosted embeddings server already WORKS today, via a custom +// provider node: getEmbeddingAdapter() matches `openai-compatible-*` and +// `custom-embedding-*` and returns openaiCompatNode, whose buildUrl reads +// creds.providerSpecificData.baseUrl. What is missing is a first-class provider, +// and the gap is visible rather than functional: +// +// /v1/embeddings on such a node -> 200, correct vectors +// the Embedding page in the dashboard -> the node is not listed at all +// +// The page renders getProvidersByKind("embedding") plus provider nodes filtered +// to `type === "custom-embedding"`. A node created as `openai-compatible` — the +// natural choice when ONE endpoint serves chat and embeddings behind the same +// front door — satisfies neither, so a working self-hosted embeddings endpoint is +// invisible on the page whose job is to show embeddings providers. Diagnosed on a +// deployment serving Qwen3-Embedding-8B at 4096 dimensions through exactly that +// shape (2026-08-04). +// +// Declaring it as a provider with serviceKinds: ["embedding"] puts it on the page +// beside Voyage, Jina and the rest, and keeps the per-connection baseUrl that +// makes self-hosting possible at all. +// +// authType is "apikey" rather than "none" for the same reason as the STT and TTS +// entries: it is what gives the connection a credentials record, and +// providerSpecificData.baseUrl lives there. Local servers ignore the key itself; +// any non-empty value works. +export default { + id: "selfhosted-embedding", + priority: 50, + hasFree: true, + alias: "selfhosted-embedding", + display: { + name: "Self-hosted Embedding", + icon: "cloud", + color: "#ffffffff", + textIcon: "SE", + website: "https://github.com/ggml-org/llama.cpp", + }, + category: "apikey", + auth: { + apiKey: { + // Note the /v1: the adapter appends "/embeddings" to whatever it is given, + // so a bare http://host:8080 resolves to http://host:8080/embeddings and + // misses the OpenAI route entirely. Give it the OpenAI base, the same value + // an OpenAI client would use. A trailing /embeddings is tolerated. + text: "Set providerSpecificData.baseUrl to the OpenAI base URL, e.g. http://host:8080/v1 — /embeddings is appended. The API key is not checked by local servers; any value works.", + }, + }, + // A self-hosted server serves whatever model it was started with, so the id + // here is a placeholder for the UI: the request passes `model` straight + // through, and llama-server ignores an unknown value rather than rejecting it. + // Dimensions are deliberately NOT declared — they are a property of the loaded + // weights, and asserting a number here would be a guess that silently + // contradicts the server. + models: [ + { id: "embedding", name: "Self-hosted embedding model", kind: "embedding" }, + ], + serviceKinds: ["embedding"], + embeddingConfig: { + // Declared for shape-consistency with the other embedding providers, and + // read by the UI — but NOT by the request path. openaiCompatNode resolves the + // URL purely from creds.providerSpecificData.baseUrl (falling back to + // api.openai.com), so unlike a fixed cloud provider this baseUrl never + // reaches the wire. Stated plainly because a reader would otherwise + // reasonably assume it is the default endpoint. + baseUrl: "http://localhost:8080/v1/embeddings", + authType: "apikey", + authHeader: "bearer", + }, +}; diff --git a/open-sse/providers/registry/selfhosted-stt.js b/open-sse/providers/registry/selfhosted-stt.js new file mode 100644 index 00000000..d2d9bc80 --- /dev/null +++ b/open-sse/providers/registry/selfhosted-stt.js @@ -0,0 +1,48 @@ +// Self-hosted, OpenAI-compatible speech-to-text (whisper.cpp, faster-whisper, +// Speaches, vLLM-served Whisper, ...). +// +// Every other STT provider here is a named cloud service with a fixed endpoint. +// This one exists so a locally-served /v1/audio/transcriptions can be used at +// all: set the connection's providerSpecificData.baseUrl to the full URL of the +// endpoint, exactly as the custom embedding providers already work. +// +// sttCore dispatches on `format`; anything that is not one of the five named +// cloud shapes falls through to transcribeOpenAICompatible, which POSTs the +// standard multipart body (file, model, and optional language / prompt / +// response_format / temperature). That is precisely what whisper.cpp's OpenAI +// endpoint accepts. +// +// authType is "apikey" rather than "none" so the connection carries a +// credentials record — which is where providerSpecificData.baseUrl lives. Local +// servers ignore the key itself; any non-empty value works. +export default { + id: "selfhosted-stt", + priority: 50, + hasFree: true, + alias: "selfhosted-stt", + display: { + name: "Self-hosted STT", + icon: "cloud", + color: "#ffffffff", + textIcon: "ST", + website: "https://github.com/ggml-org/whisper.cpp", + }, + category: "apikey", + auth: { + apiKey: { + text: "Set providerSpecificData.baseUrl to the full transcriptions URL, e.g. http://host:8080/v1/audio/transcriptions. The API key is not checked by local servers; any value works.", + }, + }, + models: [ + { id: "whisper-1", name: "Whisper (self-hosted)", params: ["language", "response_format", "temperature", "prompt"], kind: "stt" }, + ], + serviceKinds: ["stt"], + sttConfig: { + // Overridden per connection by providerSpecificData.baseUrl; this default + // only makes the provider usable out of the box on a same-host deployment. + baseUrl: "http://localhost:8080/v1/audio/transcriptions", + authType: "apikey", + authHeader: "bearer", + format: "openai", + }, +}; diff --git a/open-sse/providers/registry/selfhosted-tts.js b/open-sse/providers/registry/selfhosted-tts.js new file mode 100644 index 00000000..ff0a67b1 --- /dev/null +++ b/open-sse/providers/registry/selfhosted-tts.js @@ -0,0 +1,44 @@ +// Self-hosted, OpenAI-compatible text-to-speech (Kokoro-FastAPI, openedai-speech, +// vLLM-served TTS, ...) — the TTS counterpart of selfhosted-stt. +// +// Every other self-hostable TTS provider here (coqui, tortoise) carries a FIXED +// localhost baseUrl in its registry entry and `authType: "none"`, and the generic +// dispatcher reads `ttsConfig.baseUrl` from that entry rather than from the +// connection. So there was no way to point TTS at a server on another host. +// +// `authType: "apikey"` is what makes the override possible at all: it gives the +// connection a credentials record, which is where providerSpecificData.baseUrl +// lives. Local servers ignore the key; any non-empty value works. +export default { + id: "selfhosted-tts", + priority: 50, + hasFree: true, + alias: "selfhosted-tts", + display: { + name: "Self-hosted TTS", + icon: "cloud", + color: "#ffffffff", + textIcon: "TT", + website: "https://github.com/remsky/Kokoro-FastAPI", + }, + category: "apikey", + auth: { + apiKey: { + text: "Set providerSpecificData.baseUrl to the server root, e.g. http://host:8080 — /v1/audio/speech is appended. The API key is not checked by local servers; any value works.", + }, + }, + // Voice is selected as "/", the same convention the OpenAI TTS + // adapter uses, so existing clients need no special casing. + models: [ + { id: "kokoro", name: "Kokoro (self-hosted)", params: ["voice", "response_format", "speed"], kind: "tts" }, + ], + serviceKinds: ["tts"], + ttsConfig: { + // Overridden per connection by providerSpecificData.baseUrl; this default + // only makes the provider usable on a same-host deployment. + baseUrl: "http://localhost:8880", + defaultModel: "kokoro", + authType: "apikey", + format: "openai-speech", + }, +}; diff --git a/open-sse/providers/registry/tokenrouter.js b/open-sse/providers/registry/tokenrouter.js new file mode 100644 index 00000000..2e3cd248 --- /dev/null +++ b/open-sse/providers/registry/tokenrouter.js @@ -0,0 +1,162 @@ +export default { + id: "tokenrouter", + alias: "tokenrouter", + aliases: ["tr"], + uiAlias: "tokenrouter", + display: { + name: "TokenRouter", + icon: "hub", + color: "#0EA5E9", + textIcon: "TR", + website: "https://www.tokenrouter.com", + notice: { + text: "OpenAI-compatible gateway. 300+ models (OpenAI, Claude, Gemini, Qwen, DeepSeek, Kimi, GLM, dsb).", + apiKeyUrl: "https://www.tokenrouter.com", + }, + }, + category: "apikey", + thinkingConfig: { + options: ["low", "medium", "high", "xhigh", "max"], + defaultMode: "high", + }, + transport: { + baseUrl: "https://api.tokenrouter.com/v1/chat/completions", + validateUrl: "https://api.tokenrouter.com/v1/models", + thinkingFormat: "tokenrouter", + }, + // Seed snapshot from live /v1/models (120 entries). Latest catalogue is + // fetched via modelsFetcher; other ids still accepted via passthroughModels. + models: [ + { id: "MiniMax-Hailuo-2.3", name: "Minimax Hailuo 2.3", kind: "video" }, + { id: "MiniMax-M3", name: "Minimax M3" }, + { id: "anthropic/claude-fable-5", name: "Claude Fable 5" }, + { id: "anthropic/claude-haiku-4.5", name: "Claude Haiku 4.5" }, + { id: "anthropic/claude-opus-4.5", name: "Claude Opus 4.5" }, + { id: "anthropic/claude-opus-4.6", name: "Claude Opus 4.6" }, + { id: "anthropic/claude-opus-4.7", name: "Claude Opus 4.7" }, + { id: "anthropic/claude-opus-4.7-fast", name: "Claude Opus 4.7 Fast" }, + { id: "anthropic/claude-opus-4.8", name: "Claude Opus 4.8" }, + { id: "anthropic/claude-opus-4.8-fast", name: "Claude Opus 4.8 Fast" }, + { id: "anthropic/claude-opus-5", name: "Claude Opus 5" }, + { id: "anthropic/claude-opus-5-fast", name: "Claude Opus 5 Fast" }, + { id: "anthropic/claude-sonnet-4", name: "Claude Sonnet 4" }, + { id: "anthropic/claude-sonnet-4.5", name: "Claude Sonnet 4.5" }, + { id: "anthropic/claude-sonnet-4.6", name: "Claude Sonnet 4.6" }, + { id: "anthropic/claude-sonnet-5", name: "Claude Sonnet 5" }, + { id: "bytedance-seed/seedream-4.5", name: "Seedream 4.5", kind: "image" }, + { id: "bytedance-seed/seedream-5.0-lite", name: "Seedream 5.0 Lite", kind: "image" }, + { id: "bytedance-seed/seedream-5.0-pro", name: "Seedream 5.0 Pro", kind: "image" }, + { id: "claude-haiku-4-5", name: "Claude Haiku 4 5" }, + { id: "claude-opus-4-8-m-aws", name: "Claude Opus 4 8 M Aws" }, + { id: "deepseek/deepseek-v3.2", name: "Deepseek V3.2" }, + { id: "deepseek/deepseek-v4-flash", name: "Deepseek V4 Flash" }, + { id: "deepseek/deepseek-v4-flash-0731", name: "Deepseek V4 Flash 0731" }, + { id: "deepseek/deepseek-v4-pro", name: "Deepseek V4 Pro" }, + { id: "ex/gpt-5.4", name: "Gpt 5.4" }, + { id: "google/gemini-2.5-flash-image", name: "Gemini 2.5 Flash Image" }, + { id: "google/gemini-3-flash-preview", name: "Gemini 3 Flash Preview" }, + { id: "google/gemini-3-pro-image-preview", name: "Gemini 3 Pro Image Preview" }, + { id: "google/gemini-3.1-flash-image-preview", name: "Gemini 3.1 Flash Image Preview" }, + { id: "google/gemini-3.1-flash-lite-image", name: "Gemini 3.1 Flash Lite Image" }, + { id: "google/gemini-3.1-pro-preview", name: "Gemini 3.1 Pro Preview" }, + { id: "google/gemini-3.5-flash", name: "Gemini 3.5 Flash" }, + { id: "google/gemini-3.5-flash-lite", name: "Gemini 3.5 Flash Lite" }, + { id: "google/gemini-3.6-flash", name: "Gemini 3.6 Flash" }, + { id: "google/gemini-embedding-2", name: "Gemini Embedding 2" }, + { id: "google/gemma-4-26b-a4b-it", name: "Gemma 4 26B A4B It" }, + { id: "happyhorse-1.0-t2v", name: "Happyhorse 1.0 T2V", kind: "video" }, + { id: "kling-3.0-turbo", name: "Kling 3.0 Turbo", kind: "video" }, + { id: "kling-v2-6", name: "Kling V2 6", kind: "video" }, + { id: "kling-v3", name: "Kling V3", kind: "video" }, + { id: "kling-v3-omni", name: "Kling V3 Omni", kind: "video" }, + { id: "microsoft/mai-image-2.5", name: "Mai Image 2.5" }, + { id: "minimax/minimax-m2-her", name: "Minimax M2 Her" }, + { id: "minimax/minimax-m2.1", name: "Minimax M2.1" }, + { id: "minimax/minimax-m2.1-highspeed", name: "Minimax M2.1 Highspeed" }, + { id: "minimax/minimax-m2.5", name: "Minimax M2.5" }, + { id: "minimax/minimax-m2.7", name: "Minimax M2.7" }, + { id: "minimax/minimax-m2.7-highspeed", name: "Minimax M2.7 Highspeed" }, + { id: "miromind/mirothinker-1-7-deepresearch", name: "Mirothinker 1 7 Deepresearch" }, + { id: "miromind/mirothinker-1-7-deepresearch-mini", name: "Mirothinker 1 7 Deepresearch Mini" }, + { id: "mistralai/devstral-2512", name: "Devstral 2512" }, + { id: "mistralai/mistral-medium-3-5", name: "Mistral Medium 3 5" }, + { id: "mistralai/mistral-small-2603", name: "Mistral Small 2603" }, + { id: "mistralai/voxtral-small-24b-2507", name: "Voxtral Small 24B 2507" }, + { id: "moonshotai/kimi-k2.5", name: "Kimi K2.5" }, + { id: "moonshotai/kimi-k2.6", name: "Kimi K2.6" }, + { id: "moonshotai/kimi-k2.7-code", name: "Kimi K2.7 Code" }, + { id: "moonshotai/kimi-k3", name: "Kimi K3" }, + { id: "moonshotai/kimi-k3-free", name: "Kimi K3 Free" }, + { id: "nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free", name: "Nemotron 3 Nano Omni 30B A3B Reasoning:Free" }, + { id: "nvidia/nemotron-3-super-120b-a12b", name: "Nemotron 3 Super 120B A12B" }, + { id: "openai/gpt-4o-mini", name: "Gpt 4O Mini" }, + { id: "openai/gpt-5", name: "Gpt 5" }, + { id: "openai/gpt-5-image", name: "Gpt 5 Image" }, + { id: "openai/gpt-5-image-mini", name: "Gpt 5 Image Mini" }, + { id: "openai/gpt-5-mini", name: "Gpt 5 Mini" }, + { id: "openai/gpt-5.2", name: "Gpt 5.2" }, + { id: "openai/gpt-5.4", name: "Gpt 5.4" }, + { id: "openai/gpt-5.4-image-2", name: "Gpt 5.4 Image 2", kind: "image" }, + { id: "openai/gpt-5.4-mini", name: "Gpt 5.4 Mini" }, + { id: "openai/gpt-5.4-nano", name: "Gpt 5.4 Nano" }, + { id: "openai/gpt-5.4-pro", name: "Gpt 5.4 Pro" }, + { id: "openai/gpt-5.5", name: "Gpt 5.5" }, + { id: "openai/gpt-5.5-pro", name: "Gpt 5.5 Pro" }, + { id: "openai/gpt-5.6-luna", name: "Gpt 5.6 Luna" }, + { id: "openai/gpt-5.6-sol", name: "Gpt 5.6 Sol" }, + { id: "openai/gpt-5.6-terra", name: "Gpt 5.6 Terra" }, + { id: "openai/gpt-audio", name: "Gpt Audio", kind: "audio" }, + { id: "openai/gpt-audio-mini", name: "Gpt Audio Mini", kind: "audio" }, + { id: "openai/gpt-oss-120b", name: "Gpt Oss 120B" }, + { id: "qwen/qwen3-coder-next", name: "Qwen3 Coder Next" }, + { id: "qwen/qwen3.5-122b-a10b", name: "Qwen3.5 122B A10B" }, + { id: "qwen/qwen3.5-35b-a3b", name: "Qwen3.5 35B A3B" }, + { id: "qwen/qwen3.5-397b-a17b", name: "Qwen3.5 397B A17B" }, + { id: "qwen/qwen3.5-9b", name: "Qwen3.5 9B" }, + { id: "qwen/qwen3.5-flash", name: "Qwen3.5 Flash" }, + { id: "qwen/qwen3.5-plus-02-15", name: "Qwen3.5 Plus 02 15" }, + { id: "qwen/qwen3.6-plus", name: "Qwen3.6 Plus" }, + { id: "qwen/qwen3.7-max", name: "Qwen3.7 Max" }, + { id: "qwen/qwen3.7-plus", name: "Qwen3.7 Plus" }, + { id: "qwen/qwen3.8-max", name: "Qwen3.8 Max" }, + { id: "qwen3.5-omni-plus", name: "Qwen3.5 Omni Plus" }, + { id: "qwen3.6-flash", name: "Qwen3.6 Flash" }, + { id: "sakana/fugu-ultra", name: "Fugu Ultra" }, + { id: "seed-2-0-code-preview-260328", name: "Seed 2 0 Code Preview 260328" }, + { id: "seed-2-0-lite-260428", name: "Seed 2 0 Lite 260428" }, + { id: "seed-2-0-mini-260428", name: "Seed 2 0 Mini 260428" }, + { id: "seed-2-0-pro-260328", name: "Seed 2 0 Pro 260328" }, + { id: "stepfun/step-3.5-flash", name: "Step 3.5 Flash" }, + { id: "stepfun/step-3.7-flash", name: "Step 3.7 Flash" }, + { id: "tencent/hy3-preview", name: "Hy3 Preview" }, + { id: "x-ai/grok-4.1-fast", name: "Grok 4.1 Fast" }, + { id: "x-ai/grok-4.20-beta", name: "Grok 4.20 Beta" }, + { id: "x-ai/grok-4.3", name: "Grok 4.3" }, + { id: "x-ai/grok-4.5", name: "Grok 4.5" }, + { id: "x-ai/grok-build-0.1", name: "Grok Build 0.1" }, + { id: "xiaomi/mimo-v2-flash", name: "Mimo V2 Flash" }, + { id: "xiaomi/mimo-v2-omni", name: "Mimo V2 Omni" }, + { id: "xiaomi/mimo-v2-pro", name: "Mimo V2 Pro" }, + { id: "xiaomi/mimo-v2.5", name: "Mimo V2.5" }, + { id: "xiaomi/mimo-v2.5-pro", name: "Mimo V2.5 Pro" }, + { id: "z-ai/glm-4.5-air", name: "Glm 4.5 Air" }, + { id: "z-ai/glm-4.6", name: "Glm 4.6" }, + { id: "z-ai/glm-4.6v", name: "Glm 4.6V" }, + { id: "z-ai/glm-4.7", name: "Glm 4.7" }, + { id: "z-ai/glm-5", name: "Glm 5" }, + { id: "z-ai/glm-5-turbo", name: "Glm 5 Turbo" }, + { id: "z-ai/glm-5.1", name: "Glm 5.1" }, + { id: "z-ai/glm-5.2", name: "Glm 5.2" }, + ], + serviceKinds: ["llm", "embedding", "image"], + embeddingConfig: { + baseUrl: "https://api.tokenrouter.com/v1/embeddings", + authType: "apikey", + authHeader: "bearer", + }, + imageConfig: { + baseUrl: "https://api.tokenrouter.com/v1/images/generations", + }, + modelsFetcher: { url: "https://api.tokenrouter.com/v1/models", type: "openai" }, + passthroughModels: true, +}; diff --git a/open-sse/providers/registry/xiaomi-mimo.js b/open-sse/providers/registry/xiaomi-mimo.js index fcef7af8..49465f43 100644 --- a/open-sse/providers/registry/xiaomi-mimo.js +++ b/open-sse/providers/registry/xiaomi-mimo.js @@ -15,10 +15,11 @@ export default { textIcon: "XM", website: "https://xiaomimimo.com", notice: { - apiKeyUrl: "https://xiaomimimo.com", + apiKeyUrl: "https://platform.xiaomimimo.com/console/api-keys", }, }, category: "apikey", + serviceKinds: ["llm", "tts"], transport: { baseUrl: "https://api.xiaomimimo.com/v1/chat/completions", validateUrl: "https://api.xiaomimimo.com/v1/models", @@ -42,5 +43,12 @@ export default { { id: "mimo-v2.5", name: "MiMo V2.5" }, { id: "mimo-v2-omni", name: "MiMo V2 Omni" }, { id: "mimo-v2-flash", name: "MiMo V2 Flash" }, + { id: "mimo-v2.5-tts", name: "MiMo V2.5 TTS", kind: "tts" }, ], + ttsConfig: { + baseUrl: "https://api.xiaomimimo.com/v1/chat/completions", + authType: "apikey", + authHeader: "bearer", + format: "xiaomi-mimo-tts", + }, }; diff --git a/open-sse/providers/shared.js b/open-sse/providers/shared.js index 85b256a1..fdd76c40 100644 --- a/open-sse/providers/shared.js +++ b/open-sse/providers/shared.js @@ -47,6 +47,26 @@ export const CLAUDE_CLI_SPOOF_HEADERS = { "X-Stainless-Timeout": "600" }; +const ANTHROPIC_BETA_BASE = [ + "claude-code-20250219", + "oauth-2025-04-20", + "interleaved-thinking-2025-05-14", + "context-management-2025-06-27", + "prompt-caching-scope-2026-01-05", + "structured-outputs-2025-12-15", + "fast-mode-2026-02-01", + "redact-thinking-2026-02-12", + "token-efficient-tools-2026-03-28", +]; +const ANTHROPIC_BETA_HEAVY_AGENT = ["advanced-tool-use-2025-11-20", "effort-2025-11-24"]; + +// Heavy-agent beta flags are gated to opus/sonnet — cheaper models don't need them. +export function selectAnthropicBeta(model = "") { + const flags = [...ANTHROPIC_BETA_BASE]; + if (/^claude-(opus|sonnet)/.test(model)) flags.push(...ANTHROPIC_BETA_HEAVY_AGENT); + return flags.join(","); +} + // Shared baseUrls export const KIMI_CODING_BASE_URL = "https://api.kimi.com/coding/v1/messages"; diff --git a/open-sse/providers/thinkingLevels.js b/open-sse/providers/thinkingLevels.js index 7b638700..f42cdd84 100644 --- a/open-sse/providers/thinkingLevels.js +++ b/open-sse/providers/thinkingLevels.js @@ -31,10 +31,13 @@ const FORMAT_LEVELS = { step: L.base, }; +const CODEX_GPT_5_6_LEVELS = ["none", "minimal", "low", "medium", "high", "xhigh", "max"]; + // Model-name pattern overrides (glob, first match wins) — more precise than format default. const PATTERN_THINKING = [ - // gpt-5.6-sol accepts max (maps to xhigh on wire); live probe rejected ultra. - { pattern: "*gpt-5.6-sol*", levels: ["none", "minimal", "low", "medium", "high", "xhigh", "max"] }, + { provider: "codex", pattern: "*gpt-5.6-sol*", levels: [...CODEX_GPT_5_6_LEVELS, "ultra"] }, + { provider: "codex", pattern: "*gpt-5.6-terra*", levels: [...CODEX_GPT_5_6_LEVELS, "ultra"] }, + { provider: "codex", pattern: "*gpt-5.6-luna*", levels: CODEX_GPT_5_6_LEVELS }, { pattern: "*codex*", levels: ["low", "medium", "high", "xhigh"] }, // codex cannot disable thinking ]; @@ -43,7 +46,9 @@ export function getThinkingLevels(provider, model) { if (provider === "kiro" && resolveKiroEffortPath(model) === null) return null; const caps = getCapabilitiesForModel(provider, model); if (!caps.reasoning) return null; - const hit = PATTERN_THINKING.find((p) => matchPattern(p.pattern, model)); + const hit = PATTERN_THINKING.find((entry) => + (!entry.provider || entry.provider === provider) && matchPattern(entry.pattern, model) + ); let levels = hit?.levels || FORMAT_LEVELS[caps.thinkingFormat] || L.base; if (caps.thinkingCanDisable === false) levels = levels.filter((l) => l !== "none"); return levels; diff --git a/open-sse/rtk/headroom.js b/open-sse/rtk/headroom.js index 880195ff..2b15eed2 100644 --- a/open-sse/rtk/headroom.js +++ b/open-sse/rtk/headroom.js @@ -25,9 +25,17 @@ function messagePayload(body) { function captureSizeSnapshot(body) { const messages = messagePayload(body); + const toolHistory = messages?.filter((message) => + message?.role === "tool" + || message?.role === "function" + || message?.tool_calls?.length + || message?.content?.some?.((part) => part?.type === "tool_use" || part?.type === "tool_result") + ) || []; return { bodyBytes: jsonBytes(body), messageBytes: messages ? jsonBytes(messages) : 0, + toolSchemaBytes: jsonBytes(body?.tools || []), + toolHistoryBytes: jsonBytes(toolHistory), }; } @@ -336,7 +344,10 @@ export function formatHeadroomSizeLog(diagnostics) { const before = diagnostics?.before; const after = diagnostics?.after; if (!before || !after) return ""; - return `body=${before.bodyBytes}B→${after.bodyBytes}B messages=${before.messageBytes}B→${after.messageBytes}B`; + const effective = before.bodyBytes > 0 + ? (((before.bodyBytes - after.bodyBytes) / before.bodyBytes) * 100).toFixed(1) + : "0.0"; + return `body=${before.bodyBytes}B→${after.bodyBytes}B messages=${before.messageBytes}B→${after.messageBytes}B tools=${before.toolSchemaBytes || 0}B→${after.toolSchemaBytes || 0}B toolHistory=${before.toolHistoryBytes || 0}B→${after.toolHistoryBytes || 0}B effective=${effective}%`; } export function isHeadroomPhantomSavings(stats, diagnostics, minShrinkRatio = 0.05) { diff --git a/open-sse/services/capacityAdapter.js b/open-sse/services/capacityAdapter.js new file mode 100644 index 00000000..7b096f97 --- /dev/null +++ b/open-sse/services/capacityAdapter.js @@ -0,0 +1,173 @@ +/** + * Capacity Adapter — global fallback pools of models per input-modality capability + * (vision / pdf / audioInput / videoInput). + * + * The pool models are appended as extra fallback candidates behind whatever models + * were already going to be tried (a combo's members, or a single target model). + * combo.js's existing reorderByCapabilities then floats a capable pool model to the + * front only when none of the original models can handle the request — so this + * never overrides a combo that already has a member covering the capability. + */ +import { getCapabilitiesForModel } from "../providers/capabilities.js"; + +const CAPABILITY_KEYS = ["vision", "pdf", "audioInput", "videoInput"]; +const HARD_CAPS = new Set(CAPABILITY_KEYS); +const DEFAULT_FALLBACK_MODEL = "oc/mimo-v2.5-free"; + +// Normalize a capability entry to { enabled, roundRobin, models }. Backward-compat: +// accept the legacy array form [{model, enabled}] (treated as enabled, fallback). +function normalizeCapEntry(entry) { + if (Array.isArray(entry)) { + return { enabled: true, roundRobin: false, models: entry.map((e) => e?.model || e).filter(Boolean) }; + } + if (entry && typeof entry === "object") { + return { + enabled: entry.enabled !== false, + roundRobin: !!entry.roundRobin, + models: Array.isArray(entry.models) ? entry.models.filter(Boolean) : [], + }; + } + return { enabled: false, roundRobin: false, models: [] }; +} + +// Resolve one capability's full config. Enabled pools with no models fall back +// to DEFAULT_FALLBACK_MODEL so the toggle is never a no-op. +export function getCapacityAdapterConfig(cap, settings) { + const entry = normalizeCapEntry(settings?.capacityAdapter?.[cap]); + if (entry.enabled && entry.models.length === 0) { + return { ...entry, models: [DEFAULT_FALLBACK_MODEL] }; + } + return entry; +} + +// Flatten enabled models across all capability pools, in priority order, deduped. +export function getCapacityAdapterModels(settings) { + const seen = new Set(); + const models = []; + for (const cap of CAPABILITY_KEYS) { + const { enabled, models: pool } = getCapacityAdapterConfig(cap, settings); + if (!enabled) continue; + for (const m of pool) { + if (!seen.has(m)) { + seen.add(m); + models.push(m); + } + } + } + return models; +} + +// Strategy for a capability: "round-robin" when enabled+roundRobin, else "fallback". +export function getCapacityAdapterStrategy(cap, settings) { + const { enabled, roundRobin } = getCapacityAdapterConfig(cap, settings); + return enabled && roundRobin ? "round-robin" : "fallback"; +} + +// Strategy from the request's required capabilities: picks the first capability +// whose adapter pool is enabled and can satisfy a hard requirement. +export function getActiveAdapterStrategy(requiredCapabilities, settings) { + const hard = [...(requiredCapabilities || [])].filter((c) => HARD_CAPS.has(c)); + for (const cap of hard) { + const { enabled, models } = getCapacityAdapterConfig(cap, settings); + if (!enabled || models.length === 0) continue; + return getCapacityAdapterStrategy(cap, settings); + } + return "fallback"; +} + +function modelSatisfies(modelStr, requiredHard) { + const slash = modelStr.indexOf("/"); + const provider = slash > 0 ? modelStr.slice(0, slash) : ""; + const model = slash > 0 ? modelStr.slice(slash + 1) : modelStr; + const caps = getCapabilitiesForModel(provider, model); + return requiredHard.every((c) => caps[c] === true); +} + +// Prepend capacity-adapter models as priority candidates when NONE of the +// original models (combo members, or the single target model) can satisfy the +// request's required capabilities. Adapter models go FIRST (priority); the +// original models follow as fallback. Leaves `models` untouched when the +// original list already covers it (combo.js's reorderByCapabilities handles +// that case via autoSwitch). +export function augmentModelsWithCapacityAdapter(models, requiredCapabilities, settings) { + const hard = [...(requiredCapabilities || [])].filter((c) => HARD_CAPS.has(c)); + if (hard.length === 0 || !Array.isArray(models) || models.length === 0) return models; + if (models.some((m) => modelSatisfies(m, hard))) return models; + + const pool = getCapacityAdapterModels(settings).filter((m) => !models.includes(m) && modelSatisfies(m, hard)); + if (pool.length === 0) return models; + return [...pool, ...models]; +} + +const CHARS_PER_TOKEN = 4; // rough estimate; avoids pulling in a tokenizer dependency +const HEAD_KEEP = 6; // messages after system kept verbatim before dropping the middle + +function blockLength(content) { + if (typeof content === "string") return content.length; + if (Array.isArray(content)) { + return content.reduce((sum, b) => sum + (typeof b?.text === "string" ? b.text.length : 50), 0); + } + return 0; +} + +// Trim history to fit a (possibly smaller) context window by dropping the MIDDLE. +// Preserves: all system/instruction messages (head), and the trailing user run +// carrying the media the switch happened for (tail). Older middle turns between +// the head instructions and the current turn are dropped first. +export function stripHistoryForContext(body, contextWindow) { + const key = Array.isArray(body.messages) ? "messages" + : Array.isArray(body.input) ? "input" + : Array.isArray(body.contents) ? "contents" + : null; + if (!key) return body; + const arr = body[key]; + if (!arr || arr.length === 0) return body; + + const isSystem = (r) => r === "system" || r === "developer"; + const systemMsgs = arr.filter((m) => isSystem(m?.role)); + const rest = arr.filter((m) => !isSystem(m?.role)); + if (rest.length === 0) return body; + + const isAssistant = (r) => r === "assistant" || r === "model"; + let i = rest.length - 1; + while (i >= 0 && !isAssistant(rest[i]?.role)) i--; + const tail = rest.slice(i + 1); // current user turn (has media) — always kept + const older = rest.slice(0, i + 1); // everything before it + if (older.length === 0) return body; + + const contentOf = (m) => m.content ?? m.parts; + // Cap at 80% of the adapter model's context window — leaves room for the response. + const budgetChars = (contextWindow || 200000) * 0.8 * CHARS_PER_TOKEN; + + // Prefer keeping the first HEAD_KEEP messages (initial instructions/context) verbatim; + // only trim further if even that exceeds the adapter model's context window. + const headKept = older.slice(0, HEAD_KEEP); + let total = systemMsgs.concat(headKept, tail).reduce((s, m) => s + blockLength(contentOf(m)), 0); + + // If head + tail overflow, drop head turns from the end (closest to middle) first. + let head = headKept; + while (total > budgetChars && head.length > 0) { + const dropped = head.pop(); + total -= blockLength(contentOf(dropped)); + } + + if (head.length === older.length) return body; + return { ...body, [key]: [...systemMsgs, ...head, ...tail] }; +} + +// Wrap a handleSingleModel callback so calls to a capacity-adapter model strip +// history to fit its context window first. No-op passthrough when the pool is empty. +export function withCapacityAdapterStripping(handleSingleModel, adapterModels) { + const adapterSet = new Set(adapterModels); + if (adapterSet.size === 0) return handleSingleModel; + return (body, modelStr, ...rest) => { + if (adapterSet.has(modelStr)) { + const slash = modelStr.indexOf("/"); + const provider = slash > 0 ? modelStr.slice(0, slash) : ""; + const model = slash > 0 ? modelStr.slice(slash + 1) : modelStr; + const { contextWindow } = getCapabilitiesForModel(provider, model); + body = stripHistoryForContext(body, contextWindow); + } + return handleSingleModel(body, modelStr, ...rest); + }; +} diff --git a/open-sse/services/combo.js b/open-sse/services/combo.js index 6f6dc926..8a06195b 100644 --- a/open-sse/services/combo.js +++ b/open-sse/services/combo.js @@ -126,19 +126,33 @@ export function detectRequiredCapabilities(body) { const required = new Set(); if (!body || typeof body !== "object") return required; - const scanBlock = (b) => { - if (!b || typeof b !== "object") return; - const t = b.type; - if (t === "image_url" || t === "image" || t === "input_image") - required.add("vision"); - if (t === "file" || t === "document" || t === "input_file") - required.add("pdf"); - // gemini parts: inlineData/fileData carry a mime - const mime = b.inlineData?.mimeType || b.fileData?.mimeType; - if (typeof mime === "string" && mime.startsWith("image/")) - required.add("vision"); - if (mime === "application/pdf") required.add("pdf"); - }; + const addByMime = (mime) => { + if (typeof mime !== "string") return; + if (mime.startsWith("image/")) required.add("vision"); + else if (mime === "application/pdf") required.add("pdf"); + else if (mime.startsWith("audio/")) required.add("audioInput"); + else if (mime.startsWith("video/")) required.add("videoInput"); + }; + + const scanBlock = (b) => { + if (!b || typeof b !== "object") return; + const t = b.type; + if (t === "image_url" || t === "image" || t === "input_image") required.add("vision"); + if (t === "input_audio" || t === "audio_url" || t === "audio") required.add("audioInput"); + if (t === "input_video" || t === "video_url" || t === "video") required.add("videoInput"); + if (t === "file" || t === "document" || t === "input_file") { + // Infer modality from embedded mime when available; fall back to pdf for generic files. + let fmime = null; + if (b.input_audio?.format) fmime = `audio/${b.input_audio.format}`; + else if (b.file?.file_data) fmime = String(b.file.file_data).match(/^data:([^;,]+)/)?.[1]; + else if (b.source?.media_type) fmime = b.source.media_type; + else if (b.source?.data) fmime = String(b.source.data).match(/^data:([^;,]+)/)?.[1]; + if (fmime) addByMime(fmime); + else required.add("pdf"); + } + // gemini parts: inlineData/fileData carry a mime + addByMime(b.inlineData?.mimeType || b.fileData?.mimeType); + }; const scanContent = (content) => { if (Array.isArray(content)) for (const b of content) scanBlock(b); diff --git a/open-sse/services/projectId.js b/open-sse/services/projectId.js index 3801ac69..84ab5a2b 100644 --- a/open-sse/services/projectId.js +++ b/open-sse/services/projectId.js @@ -7,7 +7,7 @@ * This significantly reduces the risk of being flagged by Google's anti-abuse systems. */ -import { CLOUD_CODE_API, LOAD_CODE_ASSIST_HEADERS, LOAD_CODE_ASSIST_METADATA } from "../config/appConstants.js"; +import { CLOUD_CODE_API, LOAD_CODE_ASSIST_HEADERS, ANTIGRAVITY_LOAD_CODE_ASSIST_HEADERS, LOAD_CODE_ASSIST_METADATA } from "../config/appConstants.js"; // ─── Cache ──────────────────────────────────────────────────────────────────── // connectionId -> { projectId: string, fetchedAt: number } @@ -157,9 +157,10 @@ export function removeConnection(connectionId) { */ async function fetchProjectId(accessToken, signal, provider) { const endpoints = CLOUD_CODE_API[provider] || CLOUD_CODE_API["gemini-cli"]; + const headers = provider === "antigravity" ? ANTIGRAVITY_LOAD_CODE_ASSIST_HEADERS : LOAD_CODE_ASSIST_HEADERS; const response = await fetch(endpoints.loadCodeAssist, { method: "POST", - headers: { ...LOAD_CODE_ASSIST_HEADERS, "Authorization": `Bearer ${accessToken}` }, + headers: { ...headers, "Authorization": `Bearer ${accessToken}` }, body: JSON.stringify({ metadata: LOAD_CODE_ASSIST_METADATA }), signal }); @@ -186,7 +187,7 @@ async function fetchProjectId(accessToken, signal, provider) { } } - return onboardUser(accessToken, tierID, signal, endpoints); + return onboardUser(accessToken, tierID, signal, endpoints, provider); } /** @@ -197,10 +198,11 @@ async function fetchProjectId(accessToken, signal, provider) { * @param {AbortSignal} externalSignal – propagated from the connection's AbortController * @returns {Promise} */ -async function onboardUser(accessToken, tierID, externalSignal, endpoints) { +async function onboardUser(accessToken, tierID, externalSignal, endpoints, provider) { console.log(`[ProjectId] Onboarding user with tier: ${tierID}`); const reqBody = { tierId: tierID, metadata: LOAD_CODE_ASSIST_METADATA }; + const headers = provider === "antigravity" ? ANTIGRAVITY_LOAD_CODE_ASSIST_HEADERS : LOAD_CODE_ASSIST_HEADERS; const MAX_ATTEMPTS = 5; for (let attempt = 1; attempt <= MAX_ATTEMPTS; attempt++) { @@ -216,7 +218,7 @@ async function onboardUser(accessToken, tierID, externalSignal, endpoints) { try { const response = await fetch(endpoints.onboardUser, { method: "POST", - headers: { ...LOAD_CODE_ASSIST_HEADERS, "Authorization": `Bearer ${accessToken}` }, + headers: { ...headers, "Authorization": `Bearer ${accessToken}` }, body: JSON.stringify(reqBody), signal: localCtrl.signal }); diff --git a/open-sse/services/provider.js b/open-sse/services/provider.js index 1b02cd83..50f0589d 100644 --- a/open-sse/services/provider.js +++ b/open-sse/services/provider.js @@ -19,9 +19,15 @@ function isAnthropicCompatible(provider) { return typeof provider === "string" && provider.startsWith(ANTHROPIC_COMPATIBLE_PREFIX); } -function getOpenAICompatibleType(provider) { - if (!isOpenAICompatible(provider)) return "chat"; - return provider.includes("responses") ? "responses" : "chat"; +// Resolve the API type (chat vs responses) for an openai-compatible node. +// The stored apiType on the connection's providerSpecificData (kept in sync with +// the node on create/update) is authoritative. Falls back to the node ID +// substring for legacy nodes created before apiType was persisted — their IDs +// embed the type: openai-compatible--. +export function resolveOpenAICompatibleApiType(provider, credentials = null) { + const stored = credentials?.providerSpecificData?.apiType; + if (stored === "chat" || stored === "responses") return stored; + return typeof provider === "string" && provider.includes("responses") ? "responses" : "chat"; } // Detect request format from body structure @@ -105,9 +111,9 @@ export function detectFormat(body) { } // Get provider config (internal — no external runtime consumer) -function getProviderConfig(provider) { +function getProviderConfig(provider, credentials = null) { if (isOpenAICompatible(provider)) { - const apiType = getOpenAICompatibleType(provider); + const apiType = resolveOpenAICompatibleApiType(provider, credentials); return { ...PROVIDERS.openai, format: apiType === "responses" ? "openai-responses" : "openai", @@ -125,14 +131,14 @@ function getProviderConfig(provider) { } // Get target format for provider -export function getTargetFormat(provider) { +export function getTargetFormat(provider, credentials = null) { if (isOpenAICompatible(provider)) { - return getOpenAICompatibleType(provider) === "responses" ? "openai-responses" : "openai"; + return resolveOpenAICompatibleApiType(provider, credentials) === "responses" ? "openai-responses" : "openai"; } if (isAnthropicCompatible(provider)) { return "claude"; } - const config = getProviderConfig(provider); + const config = getProviderConfig(provider, credentials); return config.format || "openai"; } diff --git a/open-sse/services/qoderModels.js b/open-sse/services/qoderModels.js index 01e6fb13..572931e5 100644 --- a/open-sse/services/qoderModels.js +++ b/open-sse/services/qoderModels.js @@ -10,6 +10,12 @@ * * On any error the live cache stays empty and chatExecuteCall surfaces the * problem to the user as "model config not yet fetched, retry shortly". + * + * PAT (Personal Access Token, pt-...) connections: a PAT cannot sign COSY + * requests directly, so we exchange it for a short-lived job token (jt-...) + * via openapi.qoder.sh/api/v1/jobToken/exchange (plain JSON POST), then use + * that job token for signing. Job-token traffic must hit api2.qoder.sh — + * api3 rejects jt- with "Login expired" (403). */ import { createHash } from "crypto"; @@ -18,11 +24,30 @@ import { proxyAwareFetch } from "../utils/proxyFetch.js"; import { buildCosyHeaders } from "../shared/qoder/cosy.js"; import { QODER_MODEL_LIST_URL, + QODER_CHAT_BASE_ALT, + QODER_JOB_TOKEN_EXCHANGE_URL, + QODER_USERINFO_URL, + QODER_IDE_VERSION, + QODER_CLIENT_TYPE, } from "../shared/qoder/constants.js"; const FETCH_TIMEOUT_MS = 15_000; const CACHE_TTL_MS = 60 * 60 * 1000; // 1h, same as the Kiro catalog +const PAT_PREFIX = "pt-"; + +// PAT → job-token cache: a job token is short-lived (24h), so we keep it per +// PAT and re-exchange once it is within 5 minutes of expiry. +const PAT_REFRESH_BUFFER_MS = 5 * 60 * 1000; +const PAT_DEFAULT_TTL_MS = 24 * 60 * 60 * 1000; + +export function isQoderPat(token) { + return typeof token === "string" && token.startsWith(PAT_PREFIX); +} + +/** @type {Map} */ +const patJobCache = new Map(); + /** @type {Map, fetched: boolean }>} */ const catalogCache = new Map(); @@ -34,6 +59,109 @@ const catalogCache = new Map(); */ const inflight = new Map(); +/** + * Exchange a Qoder PAT (pt-...) for a short-lived job token (jt-...). + * This endpoint is plain JSON POST — NOT COSY-signed. + */ +async function exchangeJobToken(pat, proxyOptions = null, signal = null) { + const res = await proxyAwareFetch( + QODER_JOB_TOKEN_EXCHANGE_URL, + { + method: "POST", + headers: { + "Content-Type": "application/json", + Accept: "application/json", + "User-Agent": "qodercli/1.0.0", + "Cosy-Version": QODER_IDE_VERSION, + "Cosy-ClientType": QODER_CLIENT_TYPE, + }, + body: JSON.stringify({ personal_token: pat }), + signal, + }, + proxyOptions, + ); + if (!res.ok) { + const text = await res.text().catch(() => ""); + throw new Error(`qoder PAT exchange failed: ${res.status} ${text.slice(0, 200)}`); + } + const data = await res.json(); + if (!data.token) throw new Error("qoder PAT exchange returned no job token"); + + let expiresAt = Date.now() + PAT_DEFAULT_TTL_MS; + if (data.expires_at) { + const parsed = Date.parse(data.expires_at); + if (!Number.isNaN(parsed)) expiresAt = parsed; + } else if (typeof data.expires_in === "number" && data.expires_in > 0) { + expiresAt = Date.now() + data.expires_in; + } + return { jobToken: data.token, jobRefreshToken: data.refresh_token || "", expiresAt }; +} + +/** + * Resolve the Qoder userId for a job token (needed for COSY signing). + * Returns "" on any failure — callers fall back to the stored userId. + */ +async function fetchUserIdForJobToken(jobToken, proxyOptions = null, signal = null) { + try { + const res = await proxyAwareFetch( + QODER_USERINFO_URL, + { + method: "GET", + headers: { + Authorization: `Bearer ${jobToken}`, + Accept: "application/json", + "User-Agent": "qodercli/1.0.0", + }, + signal, + }, + proxyOptions, + ); + if (!res.ok) return ""; + const data = await res.json().catch(() => ({})); + return data.id || data.userId || data.user_id || ""; + } catch { + return ""; + } +} + +/** + * Resolve a PAT to a job-token credential, cached per-PAT. + */ +async function resolvePatCredential(pat, proxyOptions = null, signal = null) { + const cached = patJobCache.get(pat); + if (cached && cached.expiresAt - Date.now() > PAT_REFRESH_BUFFER_MS) return cached; + + const { jobToken, expiresAt } = await exchangeJobToken(pat, proxyOptions, signal); + const userId = await fetchUserIdForJobToken(jobToken, proxyOptions, signal); + const resolved = { accessToken: jobToken, userId, expiresAt }; + patJobCache.set(pat, resolved); + return resolved; +} + +/** + * Resolve connection credentials to COSY-signable form: + * - PAT (pt-...) connections → exchanged to a job token (jt-...) + userId + * - everything else → passed through unchanged + */ +export async function resolveQoderCredentials(credentials, proxyOptions = null, signal = null) { + const raw = credentials?.apiKey || credentials?.accessToken; + if (isQoderPat(raw)) { + const resolved = await resolvePatCredential(raw, proxyOptions, signal); + return { + ...credentials, + accessToken: resolved.accessToken, + apiKey: undefined, + providerSpecificData: { + authMethod: "pat", + ...(credentials?.providerSpecificData || {}), + userId: resolved.userId || credentials?.providerSpecificData?.userId || "", + machineId: credentials?.providerSpecificData?.machineId || "", + }, + }; + } + return credentials; +} + /** * Stable cache key per credential (so different login sessions for the same * account share an entry). @@ -68,10 +196,16 @@ async function fetchQoderCatalogRaw(credentials, signal, proxyOptions = null) { const creds = cosyCredsFromConnection(credentials); if (!creds.userId || !creds.authToken) return null; + // Job-token traffic is rejected by api3 ("Login expired" 403) — the + // official qodercli serves it from api2 instead. + const modelListUrl = String(creds.authToken).startsWith("jt-") + ? `${QODER_CHAT_BASE_ALT}/algo/api/v2/model/list` + : QODER_MODEL_LIST_URL; + const headers = { Accept: "application/json", "Accept-Encoding": "identity", - ...buildCosyHeaders(Buffer.alloc(0), QODER_MODEL_LIST_URL, creds), + ...buildCosyHeaders(Buffer.alloc(0), modelListUrl, creds), }; const controller = new AbortController(); @@ -92,7 +226,7 @@ async function fetchQoderCatalogRaw(credentials, signal, proxyOptions = null) { } } response = await proxyAwareFetch( - QODER_MODEL_LIST_URL, + modelListUrl, { method: "GET", headers, @@ -159,11 +293,16 @@ export async function getQoderModelConfig(credentials, modelKey, options = {}) { * one upstream request per credential. */ export async function resolveQoderModels(credentials, options = {}) { - if (!credentials?.accessToken) return null; - const psd = credentials.providerSpecificData || {}; - if (!psd.userId) return null; + let resolved; + try { + resolved = await resolveQoderCredentials(credentials, options.proxyOptions, options.signal); + } catch (error) { + options.log?.warn?.("QODER", `PAT exchange failed: ${error.message}`); + return null; + } + if (!resolved?.accessToken || !(resolved.providerSpecificData || {}).userId) return null; - const key = cacheKey(credentials); + const key = cacheKey(resolved); const now = Date.now(); if (!options.forceRefresh) { const cached = catalogCache.get(key); @@ -180,7 +319,7 @@ export async function resolveQoderModels(credentials, options = {}) { } const fetchPromise = (async () => { - const fetched = await fetchQoderCatalogRaw(credentials, options.signal, options.proxyOptions); + const fetched = await fetchQoderCatalogRaw(resolved, options.signal, options.proxyOptions); if (!fetched) return null; const entry = { expiresAt: Date.now() + CACHE_TTL_MS, diff --git a/open-sse/services/tokenRefresh.js b/open-sse/services/tokenRefresh.js index 634fd633..3160f4a7 100644 --- a/open-sse/services/tokenRefresh.js +++ b/open-sse/services/tokenRefresh.js @@ -6,7 +6,6 @@ import { refreshKimiToken, refreshClaudeOAuthToken, refreshGoogleToken, - refreshQwenToken, refreshCodexToken, refreshKiroToken, refreshIflowToken, @@ -26,7 +25,6 @@ export { refreshKimiToken, refreshClaudeOAuthToken, refreshGoogleToken, - refreshQwenToken, refreshCodexToken, refreshKiroToken, refreshIflowToken, @@ -137,7 +135,6 @@ const REFRESH_HANDLERS = { antigravity: (c, log) => refreshGoogleToken(c.refreshToken, PROVIDERS.antigravity.clientId, PROVIDERS.antigravity.clientSecret, log), claude: (c, log) => refreshClaudeOAuthToken(c.refreshToken, log), codex: (c, log) => refreshCodexToken(c.refreshToken, log), - qwen: (c, log) => refreshQwenToken(c.refreshToken, log), iflow: (c, log) => refreshIflowToken(c.refreshToken, log), github: (c, log) => refreshGitHubToken(c.refreshToken, log), kiro: (c, log) => refreshKiroToken(c.refreshToken, c.providerSpecificData, log), @@ -205,7 +202,6 @@ export function formatProviderCredentials(provider, credentials, log) { }; case "codex": - case "qwen": case "iflow": case "openai": case "openrouter": diff --git a/open-sse/services/tokenRefresh/providers.js b/open-sse/services/tokenRefresh/providers.js index 7c13ae77..40f27f51 100644 --- a/open-sse/services/tokenRefresh/providers.js +++ b/open-sse/services/tokenRefresh/providers.js @@ -40,11 +40,6 @@ const REFRESH_PROFILES = { url: () => OAUTH_ENDPOINTS.anthropic.token, dedupKey: "claude", }, - qwen: { - url: () => OAUTH_ENDPOINTS.qwen.token, - dedupKey: "qwen", - parse: (tokens) => tokens.resource_url ? { providerSpecificData: { resourceUrl: tokens.resource_url } } : {}, - }, iflow: { url: () => OAUTH_ENDPOINTS.iflow.token, dedupKey: "iflow", @@ -191,11 +186,6 @@ export async function refreshGoogleToken(refreshToken, clientId, clientSecret, l }, log); } -// Qwen: form body + clientId, surfaces resource_url. Delegate to refreshAccessToken("qwen", ...). -export async function refreshQwenToken(refreshToken, log) { - return refreshAccessToken("qwen", refreshToken, {}, log); -} - export function classifyOAuthRefreshError(errorText = "", status = 0) { let parsed = null; try { diff --git a/open-sse/services/usage.js b/open-sse/services/usage.js index b8cfe1b5..1ac0b000 100644 --- a/open-sse/services/usage.js +++ b/open-sse/services/usage.js @@ -10,14 +10,14 @@ import { getCodexUsage, consumeCodexRateLimitResetCredit, getCodexRateLimitReset export { consumeCodexRateLimitResetCredit, getCodexRateLimitResetCredits }; import { getKiroUsage } from "./usage/kiro.js"; import { getMiniMaxUsage } from "./usage/minimax.js"; -import { getCodeBuddyCnUsage } from "./usage/codebuddy-cn.js"; +import { getCodeBuddyCnUsage, getCodeBuddyIntlUsage } from "./usage/codebuddy-cn.js"; import { getXaiUsage } from "./usage/xai.js"; import { getGrokCliUsage } from "./usage/grok-cli.js"; import { getKimiUsage } from "./usage/kimi.js"; import { getDeepseekUsage } from "./usage/deepseek.js"; import { getCommandCodeUsage } from "./usage/commandcode.js"; +import { resolveQoderCredentials } from "./qoderModels.js"; import { - getQwenUsage, getIflowUsage, getOllamaUsage, getGlmUsage, @@ -38,10 +38,14 @@ const USAGE_HANDLERS = { claude: (c) => getClaudeUsage(c.accessToken, c.proxyOptions), codex: (c) => getCodexUsage(c.accessToken, c.proxyOptions), kiro: (c) => getKiroUsage(c.accessToken, c.providerSpecificData, c.proxyOptions), - qoder: (c) => getQoderUsage(c.accessToken, c.proxyOptions), - qwen: (c) => getQwenUsage(c.accessToken, c.providerSpecificData), + qoder: async (c) => { + // PAT (pt-...) connections must be exchanged to a job token before the + // quota endpoint accepts them. + const resolved = await resolveQoderCredentials(c, c.proxyOptions).catch(() => null); + return getQoderUsage(resolved?.accessToken || c.accessToken, c.proxyOptions); + }, iflow: (c) => getIflowUsage(c.accessToken), - ollama: (c) => getOllamaUsage(c.accessToken), + ollama: (c) => getOllamaUsage(c.apiKey, c.providerSpecificData, c.proxyOptions), glm: (c) => getGlmUsage(c.apiKey, c.provider, c.proxyOptions), "glm-cn": (c) => getGlmUsage(c.apiKey, c.provider, c.proxyOptions), minimax: (c) => getMiniMaxUsage(c.apiKey, c.provider, c.proxyOptions), @@ -49,6 +53,7 @@ const USAGE_HANDLERS = { "vercel-ai-gateway": (c) => getVercelAiGatewayUsage(c.apiKey, c.proxyOptions), "codebuddy-cn": (c) => getCodeBuddyCnUsage(c.accessToken, c.apiKey, c.providerSpecificData, c.proxyOptions), xai: (c) => getXaiUsage(c.accessToken, c.proxyOptions), + "codebuddy-intl": (c) => getCodeBuddyIntlUsage(c.accessToken, c.apiKey, c.providerSpecificData, c.proxyOptions), "grok-cli": (c) => getGrokCliUsage(c.accessToken, c.providerSpecificData, c.proxyOptions), kimi: (c) => getKimiUsage(c.accessToken, c.apiKey, c.proxyOptions, c.providerSpecificData), deepseek: (c) => getDeepseekUsage(c.apiKey, c.proxyOptions), diff --git a/open-sse/services/usage/codebuddy-cn.js b/open-sse/services/usage/codebuddy-cn.js index d355c729..e61d56c6 100644 --- a/open-sse/services/usage/codebuddy-cn.js +++ b/open-sse/services/usage/codebuddy-cn.js @@ -43,17 +43,17 @@ function refillCadence(acc) { return "Monthly"; } -export async function getCodeBuddyCnUsage(accessToken, apiKey, providerSpecificData, proxyOptions = null) { +async function getCodeBuddyUsage(providerId, accessToken, apiKey, providerSpecificData, proxyOptions = null) { const token = accessToken || apiKey; if (!token) { - return { message: "CodeBuddy CN credential not available." }; + return { message: `CodeBuddy (${providerId}) credential not available.` }; } try { - const response = await proxyAwareFetch(U(PROVIDER_ID).url, { + const response = await proxyAwareFetch(U(providerId).url, { method: "POST", headers: { - ...(PROVIDERS[PROVIDER_ID]?.headers || {}), + ...(PROVIDERS[providerId]?.headers || {}), Authorization: `Bearer ${token}`, "Content-Type": "application/json", Accept: "application/json", @@ -129,10 +129,18 @@ export async function getCodeBuddyCnUsage(accessToken, apiKey, providerSpecificD }); const basePkg = refills[0] || accounts[0] || {}; - const plan = basePkg.PackageName || basePkg.SubProductName || "CodeBuddy CN"; + const plan = basePkg.PackageName || basePkg.SubProductName || "CodeBuddy"; return { plan, quotas }; } catch (error) { - return { message: `CodeBuddy CN error: ${error.message}` }; + return { message: `CodeBuddy (${providerId}) error: ${error.message}` }; } } + +export async function getCodeBuddyCnUsage(accessToken, apiKey, providerSpecificData, proxyOptions = null) { + return getCodeBuddyUsage(PROVIDER_ID, accessToken, apiKey, providerSpecificData, proxyOptions); +} + +export async function getCodeBuddyIntlUsage(accessToken, apiKey, providerSpecificData, proxyOptions = null) { + return getCodeBuddyUsage("codebuddy-intl", accessToken, apiKey, providerSpecificData, proxyOptions); +} diff --git a/open-sse/services/usage/google.js b/open-sse/services/usage/google.js index f0e4d23e..e79c018e 100644 --- a/open-sse/services/usage/google.js +++ b/open-sse/services/usage/google.js @@ -161,7 +161,9 @@ export async function getAntigravityUsage(accessToken, providerSpecificData, pro if (data.models) { // Filter only recommended/important models (must match PROVIDER_MODELS ag ids) const importantModels = [ - 'gemini-3-flash-agent', + 'gemini-3.6-flash-high', + 'gemini-3.6-flash-medium', + 'gemini-3.6-flash-low', 'gemini-3.5-flash-low', 'gemini-3.5-flash-extra-low', 'gemini-pro-agent', @@ -169,10 +171,8 @@ export async function getAntigravityUsage(accessToken, providerSpecificData, pro 'claude-sonnet-4-6', 'claude-opus-4-6-thinking', 'gpt-oss-120b-medium', - 'gemini-3-flash', // Image generation models 'gemini-3.1-flash-image', - 'gemini-3-pro-image', ]; for (const [modelKey, info] of Object.entries(data.models)) { diff --git a/open-sse/services/usage/grok-cli.js b/open-sse/services/usage/grok-cli.js index 768192ad..67518e1c 100644 --- a/open-sse/services/usage/grok-cli.js +++ b/open-sse/services/usage/grok-cli.js @@ -91,6 +91,24 @@ function resolvePlan(user, config) { return "Grok Build"; } +// Display only; upstream remains authoritative for access and quota enforcement. +function planFromAccessToken(accessToken) { + try { + const payload = JSON.parse(Buffer.from(accessToken.split(".")[1], "base64url")); + return { + 0: "Free", + 1: "SuperGrok", + 2: "X Basic", + 3: "X Premium", + 4: "X Premium Plus", + 5: "SuperGrok Heavy", + 6: "SuperGrok Lite", + }[payload.tier] || ""; + } catch { + return ""; + } +} + function makeQuota({ used, total, resetAt, unlimited = false }) { const safeTotal = Math.max(0, toFiniteNumber(total, 0)); const safeUsed = Math.max(0, toFiniteNumber(used, 0)); @@ -371,6 +389,7 @@ export async function getGrokCliUsage(accessToken, providerSpecificData = null, } const parsed = parseGrokCliBilling(billing, user); + parsed.plan = planFromAccessToken(accessToken) || parsed.plan; if (!parsed.quotas || Object.keys(parsed.quotas).length === 0) { // Paid SuperGrok often returns cap=0 over REST but exposes the shared diff --git a/open-sse/services/usage/misc.js b/open-sse/services/usage/misc.js index 6ce012fa..fc133eff 100644 --- a/open-sse/services/usage/misc.js +++ b/open-sse/services/usage/misc.js @@ -1,5 +1,5 @@ /** - * Misc usage handlers (Qwen, iFlow, Ollama, GLM, Vercel AI Gateway, Qoder) + * Misc usage handlers (iFlow, Ollama, GLM, Vercel AI Gateway, Qoder) */ import { proxyAwareFetch } from "../../utils/proxyFetch.js"; @@ -15,23 +15,6 @@ const GLM_QUOTA_URLS = { // Returns { balance: "95.50", total_used: "4.50" } (USD as decimal strings). const VERCEL_AI_GATEWAY_CREDITS_URL = U("vercel-ai-gateway").url; -/** - * Qwen Usage - */ -export async function getQwenUsage(accessToken, providerSpecificData) { - try { - const resourceUrl = providerSpecificData?.resourceUrl; - if (!resourceUrl) { - return { message: "Qwen connected. No resource URL available." }; - } - - // Qwen may have usage endpoint at resource URL - return { message: "Qwen connected. Usage tracked per request." }; - } catch (error) { - return { message: "Unable to fetch Qwen usage." }; - } -} - /** * iFlow Usage */ @@ -46,23 +29,86 @@ export async function getIflowUsage(accessToken) { /** * Ollama Cloud Usage - * Ollama Cloud uses an API key from ollama.com/settings/keys - * and has no public usage API — free tier has light usage limits (resets every 5h & 7d). - * This returns an informational message with the plan details. + * GET https://ollama.com/api/usage — session (5h) + weekly (7d) `usage` is a 0..1 + * ratio (1.0 = limit reached, e.g. weekly 100% used). No reset timestamp exposed. + * POST https://ollama.com/api/me — plan label (fail-open). + * Auth: Authorization: Bearer */ -export async function getOllamaUsage(accessToken, providerSpecificData) { +export async function getOllamaUsage(apiKey, providerSpecificData, proxyOptions = null) { + if (!apiKey) { + return { message: "Ollama Cloud API key not available." }; + } + try { - // Ollama Cloud does not expose a public quota/usage API. - // The provider is configured as noAuth with a notice explaining limits. - // We return a graceful message so the UI shows a friendly state instead of an error. - const plan = providerSpecificData?.plan || "Free"; - return { - plan, - message: "Ollama Cloud uses a free tier with light usage limits (resets every 5h & 7d). For detailed usage tracking, visit ollama.com/settings/keys.", - quotas: [], - }; + const response = await proxyAwareFetch("https://ollama.com/api/usage", { + headers: { + Authorization: `Bearer ${apiKey}`, + Accept: "application/json", + }, + }, proxyOptions); + + if (response.status === 401 || response.status === 403) { + return { message: "Ollama Cloud API key invalid or expired." }; + } + + if (!response.ok) { + return { message: `Ollama Cloud usage API error (${response.status}).` }; + } + + let data; + try { + data = await response.json(); + } catch { + return { message: "Ollama Cloud usage response was not JSON." }; + } + + // Best-effort plan label from /api/me + const me = await proxyAwareFetch("https://ollama.com/api/me", { + method: "POST", + headers: { + Authorization: `Bearer ${apiKey}`, + Accept: "application/json", + "Content-Length": "0", + }, + }, proxyOptions).then((r) => (r.ok ? r.json() : null)).catch(() => null); + + const planRaw = typeof me?.Plan === "string" ? me.Plan : ""; + const plan = planRaw + ? planRaw.charAt(0).toUpperCase() + planRaw.slice(1).toLowerCase() + : "Ollama Cloud"; + + const limits = data?.limits && typeof data.limits === "object" ? data.limits : {}; + + // Ollama `usage` is a 0..1 ratio (1.0 = limit reached). Convert to a 0..100 + // bar. Do NOT set absolute `remaining` — QuotaTable reads remainingPercentage. + function ratioQuota(usageRatio, resetAt = null) { + const ratio = Math.max(0, Math.min(1, Number(usageRatio) || 0)); + const usedPct = Math.round(ratio * 100); + return { used: usedPct, total: 100, remainingPercentage: 100 - usedPct, resetAt, unlimited: false }; + } + + const sessionRaw = limits.session?.usage; + const weeklyRaw = limits.weekly?.usage; + const sessionNum = Number(sessionRaw); + const weeklyNum = Number(weeklyRaw); + const hasSession = sessionRaw !== undefined && sessionRaw !== null && !Number.isNaN(sessionNum); + const hasWeekly = weeklyRaw !== undefined && weeklyRaw !== null && !Number.isNaN(weeklyNum); + + if (!hasSession && !hasWeekly) { + return { + plan, + message: "Ollama Cloud connected. No usage limits reported.", + quotas: {}, + }; + } + + const quotas = {}; + if (hasSession) quotas["Session (5h)"] = ratioQuota(sessionNum); + if (hasWeekly) quotas["Weekly (7d)"] = ratioQuota(weeklyNum); + + return { plan, quotas }; } catch (error) { - return { message: "Unable to fetch Ollama Cloud usage." }; + return { message: `Ollama Cloud error: ${error.message}` }; } } diff --git a/open-sse/shared/qoder/constants.js b/open-sse/shared/qoder/constants.js index 184c35d6..e2635f40 100644 --- a/open-sse/shared/qoder/constants.js +++ b/open-sse/shared/qoder/constants.js @@ -11,6 +11,9 @@ export const QODER_OPENAPI_BASE = "https://openapi.qoder.sh"; export const QODER_CENTER_BASE = "https://center.qoder.sh"; export const QODER_CHAT_BASE = "https://api3.qoder.sh"; +// Job-token (jt-...) traffic is rejected by api3 with "Login expired" (403); +// the official qodercli serves it from api2 instead. +export const QODER_CHAT_BASE_ALT = "https://api2.qoder.sh"; export const QODER_LOGIN_URL = "https://qoder.com/device/selectAccounts"; diff --git a/open-sse/translator/concerns/thinkingUnified.js b/open-sse/translator/concerns/thinkingUnified.js index ee0b86aa..d18f47b6 100644 --- a/open-sse/translator/concerns/thinkingUnified.js +++ b/open-sse/translator/concerns/thinkingUnified.js @@ -3,6 +3,7 @@ // never hardcoded per-model here. See .docs/thinking/plan.md MATRIX VI-A. import { getCapabilitiesForModel } from "../../providers/capabilities.js"; +import { getThinkingLevels } from "../../providers/thinkingLevels.js"; import { PROVIDERS } from "../../providers/index.js"; import { LEVEL_TO_BUDGET, budgetToLevel, effortToBudget, effortToThinkingLevel } from "./thinking.js"; @@ -37,6 +38,7 @@ export function parseSuffix(model) { const raw = m[2].trim().toLowerCase(); if (raw === "none" || raw === "off") return { cleanModel, override: { mode: "none" } }; if (raw === "auto") return { cleanModel, override: { mode: "auto" } }; + if (raw === "ultra") return { cleanModel, override: { mode: "level", level: raw } }; if (/^\d+$/.test(raw)) return { cleanModel, override: { mode: "budget", budget: Number(raw) } }; if (LEVEL_TO_BUDGET[raw] !== undefined) return { cleanModel, override: { mode: "level", level: raw } }; return { cleanModel, override: null }; @@ -134,6 +136,13 @@ function toLevel(cfg) { return null; } +function normalizeOpenAILevel(level, supportedLevels) { + if (level !== "max" && level !== "ultra") return level; + if (supportedLevels?.includes(level)) return level; + if (level === "ultra" && supportedLevels?.includes("max")) return "max"; + return "xhigh"; +} + function toGeminiThinkingLevel(cfg) { const raw = cfg.mode === "auto" ? "high" : (toLevel(cfg) || "high"); return effortToThinkingLevel(raw); @@ -213,7 +222,7 @@ function stripAll(body) { } // Apply unified thinking config to body in the resolved provider-native format. -function applyFormat(fmt, body, cfg, caps) { +function applyFormat(fmt, body, cfg, caps, supportedLevels) { const none = cfg.mode === "none"; const canDisable = caps.thinkingCanDisable !== false; // Model cannot disable thinking → clamp "none" to minimal effort instead. @@ -223,8 +232,7 @@ function applyFormat(fmt, body, cfg, caps) { case "openai": { if (none && canDisable) { body.reasoning_effort = "none"; break; } const level = toLevel(eff); - // OpenAI reasoning_effort enum caps at "xhigh" (no "max"); clamp Claude Code's "max". - if (level) body.reasoning_effort = level === "max" ? "xhigh" : level; + if (level) body.reasoning_effort = normalizeOpenAILevel(level, supportedLevels); break; } case "claude-adaptive": { @@ -302,6 +310,15 @@ function applyFormat(fmt, body, cfg, caps) { if (level) body.reasoning_effort = level === "xhigh" || level === "max" ? "high" : level; break; } + case "tokenrouter": { + // TokenRouter's reasoning_effort enum is low/medium/high/xhigh/max — it rejects + // "none"/"auto" with a 400 and supports "max" natively (no clamp like openai). + // "none" → omit the field so the upstream default applies; pass levels through. + if (none || eff.mode === "auto") break; + const level = toLevel(eff); + if (level) body.reasoning_effort = level; + break; + } case "kiro": // Kiro thinking handled via system-tag injection in openai-to-kiro.js; no body field here. break; @@ -329,7 +346,8 @@ export function applyThinking(targetFormat, model, body, provider = null, intent if (!cfg) return body; const fmt = resolveFormat(targetFormat, cleanModel, provider); + const supportedLevels = getThinkingLevels(provider, cleanModel); stripAll(body); - applyFormat(fmt, body, cfg, caps); + applyFormat(fmt, body, cfg, caps, supportedLevels); return body; } diff --git a/open-sse/translator/formats/claude.js b/open-sse/translator/formats/claude.js index adbd3ce8..128917ec 100644 --- a/open-sse/translator/formats/claude.js +++ b/open-sse/translator/formats/claude.js @@ -16,7 +16,9 @@ export function hasValidContent(msg) { return msg.content.some(block => (block.type === CLAUDE_BLOCK.TEXT && block.text?.trim()) || block.type === CLAUDE_BLOCK.TOOL_USE || - block.type === CLAUDE_BLOCK.TOOL_RESULT + block.type === CLAUDE_BLOCK.TOOL_RESULT || + block.type === CLAUDE_BLOCK.IMAGE || + block.type === CLAUDE_BLOCK.DOCUMENT ); } return false; diff --git a/open-sse/translator/formats/gemini.js b/open-sse/translator/formats/gemini.js index b1d4db34..6393a78b 100644 --- a/open-sse/translator/formats/gemini.js +++ b/open-sse/translator/formats/gemini.js @@ -7,7 +7,13 @@ import { OPENAI_BLOCK } from "../schema/index.js"; export const UNSUPPORTED_SCHEMA_CONSTRAINTS = [ // Basic constraints (not supported by Gemini API) "minLength", "maxLength", "exclusiveMinimum", "exclusiveMaximum", - "minItems", "maxItems", "format", + "minItems", "maxItems", "format", "multipleOf", + // Array keywords the Gemini schema proto has no field for. Agent tool + // schemas set these routinely, and one occurrence rejects the whole request + // with "Unknown name ...: Cannot find field". + "uniqueItems", "contains", + // 2020-12 keywords with no Gemini equivalent + "unevaluatedProperties", "unevaluatedItems", "contentSchema", // Claude rejects these in VALIDATED mode "default", "examples", // JSON Schema meta keywords diff --git a/open-sse/translator/index.js b/open-sse/translator/index.js index 48bd1530..e2f45339 100644 --- a/open-sse/translator/index.js +++ b/open-sse/translator/index.js @@ -258,8 +258,10 @@ export function initState(sourceFormat) { funcArgsBuf: {}, funcNames: {}, funcCallIds: {}, + funcItemAdded: {}, funcArgsDone: {}, funcItemDone: {}, + customToolNames: new Set(), completedSent: false }; } diff --git a/open-sse/translator/request/openai-responses.js b/open-sse/translator/request/openai-responses.js index ecfa26d6..b821f3a4 100644 --- a/open-sse/translator/request/openai-responses.js +++ b/open-sse/translator/request/openai-responses.js @@ -32,6 +32,8 @@ export function openaiResponsesToOpenAIRequest(model, body, stream, credentials) let pendingToolResults = []; let pendingReasoning = ""; let pendingReasoningEncrypted = ""; + const additionalTools = []; + const customToolNames = new Set(); const inputItems = normalizeResponsesInput(body.input); if (!inputItems) return body; @@ -96,7 +98,7 @@ export function openaiResponsesToOpenAIRequest(model, body, stream, credentials) } result.messages.push(msg); } - else if (itemType === RESPONSES_ITEM.FUNCTION_CALL) { + else if (itemType === RESPONSES_ITEM.FUNCTION_CALL || itemType === RESPONSES_ITEM.CUSTOM_TOOL_CALL) { // Start or append to assistant message with tool_calls if (!currentAssistantMsg) { currentAssistantMsg = { @@ -108,16 +110,20 @@ export function openaiResponsesToOpenAIRequest(model, body, stream, credentials) } // Skip items with empty/missing name — Codex/OpenAI reject nameless tool calls (#444) if (!item.name || typeof item.name !== "string" || item.name.trim() === "") continue; + if (itemType === RESPONSES_ITEM.CUSTOM_TOOL_CALL) customToolNames.add(item.name); + const toolInput = itemType === RESPONSES_ITEM.CUSTOM_TOOL_CALL + ? { input: typeof item.input === "string" ? item.input : JSON.stringify(item.input ?? "") } + : item.arguments; currentAssistantMsg.tool_calls.push({ id: item.call_id, type: OPENAI_BLOCK.FUNCTION, function: { name: item.name, - arguments: item.arguments + arguments: typeof toolInput === "string" ? toolInput : JSON.stringify(toolInput ?? {}) } }); } - else if (itemType === RESPONSES_ITEM.FUNCTION_CALL_OUTPUT) { + else if (itemType === RESPONSES_ITEM.FUNCTION_CALL_OUTPUT || itemType === RESPONSES_ITEM.CUSTOM_TOOL_CALL_OUTPUT) { // Flush assistant message first if exists if (currentAssistantMsg) { result.messages.push(currentAssistantMsg); @@ -137,6 +143,9 @@ export function openaiResponsesToOpenAIRequest(model, body, stream, credentials) content: typeof item.output === "string" ? item.output : JSON.stringify(item.output) }); } + else if (itemType === RESPONSES_ITEM.ADDITIONAL_TOOLS) { + if (Array.isArray(item.tools)) additionalTools.push(...item.tools); + } else if (itemType === RESPONSES_ITEM.REASONING) { // Buffer reasoning text; attached to next assistant message/function_call. // Also stash encrypted_content so a later openai→responses hop can restore @@ -166,15 +175,45 @@ export function openaiResponsesToOpenAIRequest(model, body, stream, credentials) // explicit `name` field and cannot be represented as Chat Completions function declarations. // Filter them out to avoid sending nameless functionDeclarations to downstream providers // such as Gemini, which strictly validates function names. - if (body.tools && Array.isArray(body.tools)) { - result.tools = body.tools + const responseTools = [ + ...(Array.isArray(body.tools) ? body.tools : []), + ...additionalTools, + ]; + if (responseTools.length > 0) { + result.tools = responseTools .map(tool => { // Already in Chat Completions format: { type: "function", function: { name, ... } } if (tool.function) return tool; - // Responses API function tool: { type: "function", name, description, parameters } - // Only convert when a non-empty name is present; skip hosted tools without one. + // Responses API function/custom tool: { type, name, description, parameters|format }. + // Chat Completions has no freeform custom-tool declaration, so expose custom + // tools as functions with one raw `input` string while retaining their names + // in translator-only metadata for the response conversion. const name = tool.name; if (!name || typeof name !== "string" || name.trim() === "") return null; + if (tool.type === "custom") { + customToolNames.add(name); + const formatHint = [tool.format?.syntax, tool.format?.definition].filter(Boolean).join("\n"); + return { + type: OPENAI_BLOCK.FUNCTION, + function: { + name, + description: [String(tool.description || ""), formatHint].filter(Boolean).join("\n\n"), + parameters: { + type: "object", + properties: { + input: { + type: "string", + description: "Raw freeform input for this custom tool" + } + }, + required: ["input"], + additionalProperties: false + } + } + }; + } + // Responses API function tool: { type: "function", name, description, parameters } + // Only convert when a non-empty name is present; skip hosted tools without one. return { type: OPENAI_BLOCK.FUNCTION, function: { @@ -187,6 +226,7 @@ export function openaiResponsesToOpenAIRequest(model, body, stream, credentials) }) .filter(Boolean); } + if (customToolNames.size > 0) result._customToolNames = [...customToolNames]; // Cleanup Responses API specific fields // Map Responses-only max_output_tokens to Chat max_tokens (avoid leaking unknown field upstream) diff --git a/open-sse/translator/response/openai-responses.js b/open-sse/translator/response/openai-responses.js index b9336785..92e3cb52 100644 --- a/open-sse/translator/response/openai-responses.js +++ b/open-sse/translator/response/openai-responses.js @@ -258,24 +258,43 @@ function closeMessage(state, emit, idx) { } } +function isCustomTool(state, name) { + return !!name && state.customToolNames?.has(name); +} + +function extractCustomToolInput(argumentsText) { + if (typeof argumentsText !== "string") return ""; + try { + const parsed = JSON.parse(argumentsText); + if (parsed && typeof parsed === "object" && typeof parsed.input === "string") return parsed.input; + } catch { /* incomplete or raw freeform input */ } + return argumentsText; +} + function emitToolCall(state, emit, tc) { const tcIdx = tc.index ?? 0; const newCallId = tc.id; const funcName = tc.function?.name; if (funcName) state.funcNames[tcIdx] = funcName; + if (newCallId) state.funcCallIds[tcIdx] = newCallId; + + // Some compatible providers split the call id and function name across + // chunks. Wait for both before deciding whether this is a custom tool; + // otherwise an `exec` call can be irreversibly announced as function_call. + const callId = state.funcCallIds[tcIdx]; + if (!state.funcItemAdded[tcIdx] && callId && state.funcNames[tcIdx]) { + state.funcItemAdded[tcIdx] = true; + const custom = isCustomTool(state, state.funcNames[tcIdx]); - if (!state.funcCallIds[tcIdx] && newCallId) { - state.funcCallIds[tcIdx] = newCallId; - emit("response.output_item.added", { type: "response.output_item.added", output_index: tcIdx, item: { - id: `fc_${newCallId}`, - type: RESPONSES_ITEM.FUNCTION_CALL, - arguments: "", - call_id: newCallId, + id: `${custom ? "ctc" : "fc"}_${callId}`, + type: custom ? RESPONSES_ITEM.CUSTOM_TOOL_CALL : RESPONSES_ITEM.FUNCTION_CALL, + ...(custom ? { input: "" } : { arguments: "" }), + call_id: callId, name: state.funcNames[tcIdx] || "" } }); @@ -285,7 +304,7 @@ function emitToolCall(state, emit, tc) { if (tc.function?.arguments) { const refCallId = state.funcCallIds[tcIdx] || newCallId; - if (refCallId) { + if (state.funcItemAdded[tcIdx] && refCallId && !isCustomTool(state, state.funcNames[tcIdx])) { emit("response.function_call_arguments.delta", { type: "response.function_call_arguments.delta", item_id: `fc_${refCallId}`, @@ -293,6 +312,9 @@ function emitToolCall(state, emit, tc) { delta: tc.function.arguments }); } + // Custom input is emitted once at close, after the Chat JSON wrapper can be + // parsed and unwrapped. Streaming the raw JSON fragments would expose + // {"input":"..."} instead of the freeform program Codex expects. state.funcArgsBuf[tcIdx] += tc.function.arguments; } } @@ -301,21 +323,38 @@ function closeToolCall(state, emit, idx) { const callId = state.funcCallIds[idx]; if (callId && !state.funcItemDone[idx]) { const args = state.funcArgsBuf[idx] || "{}"; - - emit("response.function_call_arguments.done", { - type: "response.function_call_arguments.done", - item_id: `fc_${callId}`, - output_index: parseInt(idx), - arguments: args - }); + const custom = isCustomTool(state, state.funcNames[idx]); + + if (custom) { + const input = extractCustomToolInput(args); + emit("response.custom_tool_call_input.delta", { + type: "response.custom_tool_call_input.delta", + item_id: `ctc_${callId}`, + output_index: parseInt(idx), + delta: input + }); + emit("response.custom_tool_call_input.done", { + type: "response.custom_tool_call_input.done", + item_id: `ctc_${callId}`, + output_index: parseInt(idx), + input + }); + } else { + emit("response.function_call_arguments.done", { + type: "response.function_call_arguments.done", + item_id: `fc_${callId}`, + output_index: parseInt(idx), + arguments: args + }); + } emit("response.output_item.done", { type: "response.output_item.done", output_index: parseInt(idx), item: { - id: `fc_${callId}`, - type: RESPONSES_ITEM.FUNCTION_CALL, - arguments: args, + id: `${custom ? "ctc" : "fc"}_${callId}`, + type: custom ? RESPONSES_ITEM.CUSTOM_TOOL_CALL : RESPONSES_ITEM.FUNCTION_CALL, + ...(custom ? { input: extractCustomToolInput(args) } : { arguments: args }), call_id: callId, name: state.funcNames[idx] || "" } diff --git a/open-sse/translator/schema/blocks.js b/open-sse/translator/schema/blocks.js index 61c25645..958122ce 100644 --- a/open-sse/translator/schema/blocks.js +++ b/open-sse/translator/schema/blocks.js @@ -27,6 +27,9 @@ export const RESPONSES_ITEM = { MESSAGE: "message", FUNCTION_CALL: "function_call", FUNCTION_CALL_OUTPUT: "function_call_output", + CUSTOM_TOOL_CALL: "custom_tool_call", + CUSTOM_TOOL_CALL_OUTPUT: "custom_tool_call_output", + ADDITIONAL_TOOLS: "additional_tools", REASONING: "reasoning", OUTPUT_TEXT: "output_text", INPUT_TEXT: "input_text", diff --git a/open-sse/utils/claudeHeaderCache.js b/open-sse/utils/claudeHeaderCache.js deleted file mode 100644 index 11b2eb81..00000000 --- a/open-sse/utils/claudeHeaderCache.js +++ /dev/null @@ -1,70 +0,0 @@ -/** - * Singleton cache for real Claude Code client headers. - * Captures headers from authentic Claude Code requests and makes them available - * for forwarding to api.anthropic.com, replacing static hardcoded values. - */ - -const CLAUDE_IDENTITY_HEADERS = [ - "user-agent", - "anthropic-beta", - "anthropic-version", - "anthropic-dangerous-direct-browser-access", - "x-app", - "x-stainless-helper-method", - "x-stainless-retry-count", - "x-stainless-runtime-version", - "x-stainless-package-version", - "x-stainless-runtime", - "x-stainless-lang", - "x-stainless-arch", - "x-stainless-os", - "x-stainless-timeout", - "x-claude-code-session-id", - "package-version", - "runtime-version", - "os", - "arch", -]; - -let cachedHeaders = null; - -/** - * Detect if request headers look like a real Claude Code client. - * @param {object} headers - Lowercase header key/value object - */ -function isClaudeCodeClient(headers) { - const ua = (headers["user-agent"] || "").toLowerCase(); - const xApp = (headers["x-app"] || "").toLowerCase(); - return ua.includes("claude-cli") || ua.includes("claude-code") || xApp === "cli"; -} - -/** - * Store Claude Code identity headers if this looks like a real client request. - * Called at the entry point before any translation/forwarding. - * @param {object} headers - Lowercase header key/value object (from request.headers.entries()) - */ -export function cacheClaudeHeaders(headers) { - if (!headers || typeof headers !== "object") return; - if (!isClaudeCodeClient(headers)) return; - - const captured = {}; - for (const key of CLAUDE_IDENTITY_HEADERS) { - if (headers[key] !== undefined && headers[key] !== null) { - captured[key] = headers[key]; - } - } - - if (Object.keys(captured).length > 0) { - cachedHeaders = captured; - console.log(`[ClaudeHeaders] Cached ${Object.keys(captured).length} identity headers from Claude Code client`); - } -} - -/** - * Get the most recently cached Claude Code identity headers. - * Returns null if no authentic client request has been seen yet (cold start). - * @returns {object|null} - */ -export function getCachedClaudeHeaders() { - return cachedHeaders; -} diff --git a/open-sse/utils/clientDetector.js b/open-sse/utils/clientDetector.js index 2d1381bc..12020762 100644 --- a/open-sse/utils/clientDetector.js +++ b/open-sse/utils/clientDetector.js @@ -22,6 +22,7 @@ export function detectClientTool(headers = {}, body = {}) { const xApp = (headers["x-app"] || "").toLowerCase(); const openaiIntent = (headers["openai-intent"] || "").toLowerCase(); const initiator = (headers["x-initiator"] || headers["X-Initiator"] || "").toLowerCase(); + const originator = (headers["originator"] || "").toLowerCase(); // Antigravity: detected via body field (not header) if (body.userAgent === "antigravity") return "antigravity"; @@ -37,8 +38,10 @@ export function detectClientTool(headers = {}, body = {}) { // Gemini CLI if (ua.includes("gemini-cli")) return "gemini-cli"; - // Codex CLI - if (ua.includes("codex-cli")) return "codex"; + // Codex CLI/Desktop — codex-tui is the current Rust CLI, codex-cli/codex_cli_rs legacy; + // Codex Desktop identifies via UA "Codex Desktop" or originator "codex_work_desktop" + if (ua.includes("codex-tui") || ua.includes("codex-cli") || ua.includes("codex_cli_rs") || + ua.includes("codex desktop") || originator.startsWith("codex_")) return "codex"; // DeepSeek TUI if (ua.includes("deepseek-tui")) return "deepseek-tui"; diff --git a/open-sse/utils/stream.js b/open-sse/utils/stream.js index 464acde4..33e7fb04 100644 --- a/open-sse/utils/stream.js +++ b/open-sse/utils/stream.js @@ -44,6 +44,7 @@ export function createSSEStream(options = {}) { provider = null, reqLogger = null, toolNameMap = null, + customToolNames = null, model = null, connectionId = null, body = null, @@ -57,7 +58,9 @@ export function createSSEStream(options = {}) { // Per-stream decoder with stream:true to correctly handle multi-byte chars split across chunks const decoder = new TextDecoder("utf-8", { fatal: false }); - const state = mode === STREAM_MODE.TRANSLATE ? { ...initState(sourceFormat), provider, toolNameMap, model } : null; + const state = mode === STREAM_MODE.TRANSLATE + ? { ...initState(sourceFormat), provider, toolNameMap, customToolNames: new Set(customToolNames || []), model } + : null; let totalContentLength = 0; let accumulatedContent = ""; @@ -464,7 +467,7 @@ export function createSSEStream(options = {}) { }); } -export function createSSETransformStreamWithLogger(targetFormat, sourceFormat, provider = null, reqLogger = null, toolNameMap = null, model = null, connectionId = null, body = null, onStreamComplete = null, apiKey = null) { +export function createSSETransformStreamWithLogger(targetFormat, sourceFormat, provider = null, reqLogger = null, toolNameMap = null, model = null, connectionId = null, body = null, onStreamComplete = null, apiKey = null, customToolNames = null) { return createSSEStream({ mode: STREAM_MODE.TRANSLATE, targetFormat, @@ -472,6 +475,7 @@ export function createSSETransformStreamWithLogger(targetFormat, sourceFormat, p provider, reqLogger, toolNameMap, + customToolNames, model, connectionId, body, diff --git a/package.json b/package.json index fbedad57..4374a515 100644 --- a/package.json +++ b/package.json @@ -1,12 +1,14 @@ { "name": "9router-app", - "version": "0.5.45", + "version": "0.5.50", "description": "9Router web dashboard", "private": true, "scripts": { "dev": "next dev --port 20127", "dev:webpack": "next dev --webpack --port 20127", "build": "next build --webpack", + "postbuild": "node scripts/copy-standalone-assets.mjs", + "postbuild:bun": "node scripts/copy-standalone-assets.mjs", "start": "next start --port 20127", "dev:bun": "bun --bun next dev --webpack --port 20127", "build:bun": "bun --bun next build --webpack", diff --git a/public/i18n/literals/ar.json b/public/i18n/literals/ar.json index 32b72f8b..f17abcf2 100644 --- a/public/i18n/literals/ar.json +++ b/public/i18n/literals/ar.json @@ -12,6 +12,7 @@ "Logout": "تسجيل الخروج", "Login": "تسجيل الدخول", "Providers": "الموفرون", + "Style": "النمط", "Usage": "الإحصائيات", "API Key": "مفتاح API", "Connected": "متصل", @@ -176,6 +177,7 @@ "How it works:": "كيف يعمل:", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "طلب Antigravity/Copilot IDE → إعادة توجيه DNS إلى localhost:443 → يعترض وكيل MITM → 9Router → الرد إلى Antigravity/Copilot", "No API keys — create one in Keys page": "لا توجد مفاتيح API — قم بإنشاء واحدة في صفحة المفاتيح", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "مثال: صوت دافئ ولطيف، يتحدث ببطء بلهجة بريطانية", "sk_9router (default)": "sk_9router (افتراضي)", "Server started": "تم بدء الخادم", "Failed to start server": "فشل في بدء الخادم", diff --git a/public/i18n/literals/bn.json b/public/i18n/literals/bn.json index ef71ddf8..a8ccc168 100644 --- a/public/i18n/literals/bn.json +++ b/public/i18n/literals/bn.json @@ -12,6 +12,7 @@ "Logout": "লগ আউট", "Login": "লগ ইন", "Providers": "সরবরাহকারী", + "Style": "শৈলী", "Usage": "ব্যবহারের পরিসংখ্যান", "API Key": "API কী", "Connected": "সংযুক্ত", @@ -176,6 +177,7 @@ "How it works:": "এটি কীভাবে কাজ করে:", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "Antigravity/Copilot IDE অনুরোধ → DNS কে localhost:443 তে রিডিরেক্ট করুন → MITM প্রক্সি ইন্টারসেপ্ট করে → 9Router → Antigravity/Copilot এ প্রতিক্রিয়া", "No API keys — create one in Keys page": "কোন API কী নেই — Keys পৃষ্ঠায় একটি তৈরি করুন", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "যেমন: উষ্ণ, মৃদু কণ্ঠস্বর, ব্রিটিশ উচ্চারণে ধীরে ধীরে কথা বলা", "sk_9router (default)": "sk_9router (ডিফল্ট)", "Server started": "সার্ভার শুরু হয়েছে", "Failed to start server": "সার্ভার শুরু করতে ব্যর্থ", diff --git a/public/i18n/literals/cs.json b/public/i18n/literals/cs.json index ed0d991a..d73139cd 100644 --- a/public/i18n/literals/cs.json +++ b/public/i18n/literals/cs.json @@ -12,6 +12,7 @@ "Logout": "Odhlásit se", "Login": "Přihlásit se", "Providers": "Poskytovatelé", + "Style": "Styl", "Usage": "Statistika", "API Key": "Klíč API", "Connected": "Připojeno", @@ -176,6 +177,7 @@ "How it works:": "Jak to funguje:", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "Požadavek Antigravity/Copilot IDE → Přesměrování DNS na localhost:443 → Proxy MITM zachycuje → 9Router → odpověď na Antigravity/Copilot", "No API keys — create one in Keys page": "Žádné klíče API — vytvořte jeden na stránce Klíče", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "např.: teplý, jemný hlas, mluvící pomalu s britským přízvukem", "sk_9router (default)": "sk_9router (výchozí)", "Server started": "Server spuštěn", "Failed to start server": "Spuštění serveru se nezdařilo", diff --git a/public/i18n/literals/da.json b/public/i18n/literals/da.json index c81bbe79..aca9c53c 100644 --- a/public/i18n/literals/da.json +++ b/public/i18n/literals/da.json @@ -12,6 +12,7 @@ "Logout": "Log ud", "Login": "Log ind", "Providers": "Udbydere", + "Style": "Stil", "Usage": "Forbrugsstatistik", "API Key": "API-nøgle", "Connected": "Forbundet", @@ -176,6 +177,7 @@ "How it works:": "Sådan virker det:", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "Antigravity/Copilot IDE-anmodning → DNS-omdirigering til localhost:443 → MITM-proxy aflytter → 9Router → svar til Antigravity/Copilot", "No API keys — create one in Keys page": "Ingen API-nøgler — opret en på Keys-siden", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "fx.: en varm, blød stemme, der taler langsomt med britisk accent", "sk_9router (default)": "sk_9router (standard)", "Server started": "Server startet", "Failed to start server": "Fejl ved start af server", diff --git a/public/i18n/literals/de.json b/public/i18n/literals/de.json index 57bf3ef3..1a02fab7 100644 --- a/public/i18n/literals/de.json +++ b/public/i18n/literals/de.json @@ -12,6 +12,7 @@ "Logout": "Abmelden", "Login": "Anmelden", "Providers": "Anbieter", + "Style": "Stil", "Usage": "Statistiken", "API Key": "API-Schlüssel", "Connected": "Verbunden", @@ -176,6 +177,7 @@ "How it works:": "So funktioniert es:", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "Antigravity/Copilot IDE-Anforderung → DNS-Umleitung auf localhost:443 → MITM-Proxy abfangen → 9Router → Antwort auf Antigravity/Copilot", "No API keys — create one in Keys page": "Keine API-Schlüssel — erstellen Sie einen auf der Seite Schlüssel", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "z. B.: eine warme, sanfte Stimme, die langsam mit britischem Akzent spricht", "sk_9router (default)": "sk_9router (Standard)", "Server started": "Server gestartet", "Failed to start server": "Server konnte nicht gestartet werden", diff --git a/public/i18n/literals/el.json b/public/i18n/literals/el.json index bfbf8888..a220ac81 100644 --- a/public/i18n/literals/el.json +++ b/public/i18n/literals/el.json @@ -12,6 +12,7 @@ "Logout": "Έξοδος", "Login": "Σύνδεση", "Providers": "Παρόχοι", + "Style": "Στυλ", "Usage": "Στατιστικά χρήσης", "API Key": "Κλειδί API", "Connected": "Συνδεδεμένο", @@ -176,6 +177,7 @@ "How it works:": "Πώς λειτουργεί:", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "Αίτημα Antigravity/Copilot IDE → Ανακατεύθυνση DNS στο localhost:443 → Ο διακομιστής μεσολάβησης MITM παρεμβαίνει → 9Router → απάντηση στο Antigravity/Copilot", "No API keys — create one in Keys page": "Δεν υπάρχουν κλειδιά API — δημιουργήστε ένα στη σελίδα Keys", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "π.χ.: μια ζεστή, απαλή φωνή, που μιλάει αργά με βρετανική προφορά", "sk_9router (default)": "sk_9router (προεπιλεγμένο)", "Server started": "Ο διακομιστής ξεκίνησε", "Failed to start server": "Αποτυχία εκκίνησης διακομιστή", diff --git a/public/i18n/literals/es.json b/public/i18n/literals/es.json index 69d71e8f..a5e10674 100644 --- a/public/i18n/literals/es.json +++ b/public/i18n/literals/es.json @@ -12,6 +12,7 @@ "Logout": "Cerrar sesión", "Login": "Iniciar sesión", "Providers": "Proveedores", + "Style": "Estilo", "Usage": "Estadísticas", "API Key": "Clave API", "Connected": "Conectado", @@ -176,6 +177,7 @@ "How it works:": "Cómo funciona:", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "Solicitud de Antigravity/Copilot IDE → Redireccionamiento DNS a localhost:443 → El proxy MITM intercepta → 9Router → respuesta a Antigravity/Copilot", "No API keys — create one in Keys page": "Sin claves API — cree una en la página Claves", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "p. ej.: voz cálida y suave, hablando lentamente con acento británico", "sk_9router (default)": "sk_9router (predeterminado)", "Server started": "Servidor iniciado", "Failed to start server": "Error al iniciar el servidor", diff --git a/public/i18n/literals/fa.json b/public/i18n/literals/fa.json index b298fc57..0f28367f 100644 --- a/public/i18n/literals/fa.json +++ b/public/i18n/literals/fa.json @@ -1119,6 +1119,7 @@ "Stop Server": "توقف سرور", "Stopped": "متوقف شد", "Strict Proxy": "پروکسی سختگیرانه", + "Style": "سبک", "Subagent Model": "مدل زیرعامل", "Sudo Password Required": "رمز عبور sudo الزامی است", "Sudo password is required": "رمز عبور sudo الزامی است", @@ -1312,6 +1313,7 @@ "disabled": "غیرفعال", "dollars per million tokens": "دلار به ازای هر میلیون توکن", "e.g. CwhRBWXzGAHq8TQ4Fs17": "مثلاً CwhRBWXzGAHq8TQ4Fs17", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "مثلاً: صدای گرم و ملایم که با لهجه بریتانیایی به‌آرامی صحبت می‌کند", "e.g. claude-opus-4-5": "مثلاً claude-opus-4-5", "e.g. my-model-id": "مثلاً my-model-id", "e.g. tts-1-hd": "مثلاً tts-1-hd", diff --git a/public/i18n/literals/fi.json b/public/i18n/literals/fi.json index dc8b116c..4f86b061 100644 --- a/public/i18n/literals/fi.json +++ b/public/i18n/literals/fi.json @@ -12,6 +12,7 @@ "Logout": "Kirjaudu ulos", "Login": "Kirjaudu sisään", "Providers": "Palveluntarjoajat", + "Style": "Tyyli", "Usage": "Käyttötilastot", "API Key": "API-avain", "Connected": "Yhdistetty", @@ -176,6 +177,7 @@ "How it works:": "Kuinka se toimii:", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "Antigravity/Copilot IDE-pyyntö → DNS-uudelleenohjaus localhost:443:iin → MITM-välityspalvelin sieppaa → 9Router → vastaus Antigravity/Copilot:ille", "No API keys — create one in Keys page": "Ei API-avaimia — luo yksi Keys-sivulla", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "esim.: lämmin, pehmeä ääni, joka puhuu hitaasti brittiaksentilla", "sk_9router (default)": "sk_9router (oletus)", "Server started": "Palvelin käynnistetty", "Failed to start server": "Palvelimen käynnistäminen epäonnistui", diff --git a/public/i18n/literals/fr.json b/public/i18n/literals/fr.json index bbf8854a..55b86581 100644 --- a/public/i18n/literals/fr.json +++ b/public/i18n/literals/fr.json @@ -12,6 +12,7 @@ "Logout": "Déconnexion", "Login": "Connexion", "Providers": "Fournisseurs", + "Style": "Style", "Usage": "Statistiques", "API Key": "Clé API", "Connected": "Connecté", @@ -176,6 +177,7 @@ "How it works:": "Comment ça marche :", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "Demande Antigravity/Copilot IDE → Redirection DNS vers localhost:443 → Le proxy MITM intercepte → 9Router → réponse à Antigravity/Copilot", "No API keys — create one in Keys page": "Aucune clé API — créez-en une dans la page Clés", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "ex. : voix chaleureuse et douce, parlant lentement avec un accent britannique", "sk_9router (default)": "sk_9router (par défaut)", "Server started": "Serveur démarré", "Failed to start server": "Impossible de démarrer le serveur", diff --git a/public/i18n/literals/he.json b/public/i18n/literals/he.json index c144d01b..7ed67a5e 100644 --- a/public/i18n/literals/he.json +++ b/public/i18n/literals/he.json @@ -12,6 +12,7 @@ "Logout": "התנתקות", "Login": "כניסה", "Providers": "ספקים", + "Style": "סגנון", "Usage": "סטטיסטיקה", "API Key": "מפתח API", "Connected": "מחובר", @@ -176,6 +177,7 @@ "How it works:": "איך זה עובד:", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "בקשת Antigravity/Copilot IDE → הפניה DNS ל-localhost:443 → פרוקסי MITM חוטף → 9Router → תגובה ל-Antigravity/Copilot", "No API keys — create one in Keys page": "אין מפתחות API — צור אחד בעמוד Keys", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "למשל: קול חם ועדין, מדבר לאט במבטא בריטי", "sk_9router (default)": "sk_9router (ברירת מחדל)", "Server started": "השרת התחיל", "Failed to start server": "הפעלת השרת נכשלה", diff --git a/public/i18n/literals/hi.json b/public/i18n/literals/hi.json index 2fc6b338..3f3aeda6 100644 --- a/public/i18n/literals/hi.json +++ b/public/i18n/literals/hi.json @@ -12,6 +12,7 @@ "Logout": "लॉग आउट", "Login": "लॉगिन", "Providers": "प्रदाता", + "Style": "शैली", "Usage": "उपयोग के आंकड़े", "API Key": "API कुंजी", "Connected": "जुड़ा हुआ", @@ -176,6 +177,7 @@ "How it works:": "यह कैसे काम करता है:", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "Antigravity/Copilot IDE अनुरोध → DNS को localhost:443 में पुनर्निर्देशित करें → MITM प्रॉक्सी इंटरसेप्ट करता है → 9Router → Antigravity/Copilot को प्रतिक्रिया", "No API keys — create one in Keys page": "कोई API कुंजी नहीं — Keys पृष्ठ में एक बनाएं", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "जैसे: गर्म, कोमल आवाज़, ब्रिटिश लहजे में धीरे-धीरे बोलते हुए", "sk_9router (default)": "sk_9router (डिफ़ॉल्ट)", "Server started": "सर्वर शुरू किया गया", "Failed to start server": "सर्वर शुरू करने में विफल", diff --git a/public/i18n/literals/hu.json b/public/i18n/literals/hu.json index 8e339291..927d1135 100644 --- a/public/i18n/literals/hu.json +++ b/public/i18n/literals/hu.json @@ -12,6 +12,7 @@ "Logout": "Kijelentkezés", "Login": "Bejelentkezés", "Providers": "Szolgáltatók", + "Style": "Stílus", "Usage": "Használati statisztika", "API Key": "API-kulcs", "Connected": "Csatlakoztatva", @@ -176,6 +177,7 @@ "How it works:": "Hogyan működik:", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "Antigravity/Copilot IDE kérés → DNS átirányítás a localhost:443-ra → MITM proxy elfogja → 9Router → válasz Antigravity/Copilot-nak", "No API keys — create one in Keys page": "Nincsenek API-kulcsok — hozzon létre egyet a Keys oldalon", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "pl.: meleg, lágy hang, lassan beszél brit akcentussal", "sk_9router (default)": "sk_9router (alapértelmezett)", "Server started": "Szerver elindult", "Failed to start server": "Nem sikerült elindítani a szervert", diff --git a/public/i18n/literals/id.json b/public/i18n/literals/id.json index 3e5097aa..c60c1397 100644 --- a/public/i18n/literals/id.json +++ b/public/i18n/literals/id.json @@ -12,6 +12,7 @@ "Logout": "Keluar", "Login": "Masuk", "Providers": "Penyedia", + "Style": "Gaya", "Usage": "Statistik Penggunaan", "API Key": "Kunci API", "Connected": "Terhubung", @@ -176,6 +177,7 @@ "How it works:": "Cara kerjanya:", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "Permintaan Antigravity/Copilot IDE → Pengalihan DNS ke localhost:443 → Proxy MITM mengintersep → 9Router → respons ke Antigravity/Copilot", "No API keys — create one in Keys page": "Tidak ada kunci API — buat satu di halaman Keys", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "mis.: suara hangat dan lembut, berbicara pelan dengan aksen Inggris", "sk_9router (default)": "sk_9router (bawaan)", "Server started": "Server dimulai", "Failed to start server": "Gagal memulai server", diff --git a/public/i18n/literals/it.json b/public/i18n/literals/it.json index 7f684e83..e1619f86 100644 --- a/public/i18n/literals/it.json +++ b/public/i18n/literals/it.json @@ -12,6 +12,7 @@ "Logout": "Esci", "Login": "Accedi", "Providers": "Provider", + "Style": "Stile", "Usage": "Statistiche di utilizzo", "API Key": "Chiave API", "Connected": "Connesso", @@ -176,6 +177,7 @@ "How it works:": "Come funziona:", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "Richiesta Antigravity/Copilot IDE → Reindirizzamento DNS a localhost:443 → Il proxy MITM intercetta → 9Router → Risposta a Antigravity/Copilot", "No API keys — create one in Keys page": "Nessuna chiave API — crearne una nella pagina Chiavi", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "es.: voce calda e morbida, che parla lentamente con accento britannico", "sk_9router (default)": "sk_9router (predefinito)", "Server started": "Server avviato", "Failed to start server": "Impossibile avviare il server", diff --git a/public/i18n/literals/ja.json b/public/i18n/literals/ja.json index e448c685..bac4f57b 100644 --- a/public/i18n/literals/ja.json +++ b/public/i18n/literals/ja.json @@ -12,6 +12,7 @@ "Logout": "ログアウト", "Login": "ログイン", "Providers": "プロバイダー", + "Style": "スタイル", "Usage": "統計", "API Key": "APIキー", "Connected": "接続済み", @@ -176,6 +177,7 @@ "How it works:": "しくみ:", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "Antigravity/Copilot IDE リクエスト → localhost:443 への DNS リダイレクト → MITM プロキシが傍受 → 9Router → Antigravity/Copilot への応答", "No API keys — create one in Keys page": "APIキーがありません — キーページで1つ作成してください", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "例:温かく穏やかな声で、イギリス英語のアクセントでゆっくり話す", "sk_9router (default)": "sk_9router(デフォルト)", "Server started": "サーバーが開始されました", "Failed to start server": "サーバーの開始に失敗しました", diff --git a/public/i18n/literals/km.json b/public/i18n/literals/km.json index cb40c432..bce40e0a 100644 --- a/public/i18n/literals/km.json +++ b/public/i18n/literals/km.json @@ -1121,6 +1121,7 @@ "Stop Server": "បញ្ឈប់ម៉ាស៊ីនមេ", "Stopped": "បានបញ្ឈប់", "Strict Proxy": "Strict Proxy", + "Style": "រចនាបទ", "Subagent Model": "ម៉ូដែល Subagent", "Sudo Password Required": "ត្រូវការពាក្យសម្ងាត់ Sudo", "Sudo password is required": "ត្រូវការពាក្យសម្ងាត់ Sudo", @@ -1317,6 +1318,7 @@ "disabled": "បានបិទ", "dollars per million tokens": "ដុល្លារក្នុងមួយលាន Tokens", "e.g. CwhRBWXzGAHq8TQ4Fs17": "ឧ. CwhRBWXzGAHq8TQ4Fs17", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "ឧទាហរណ៍៖ សំឡេងកក់ក្តៅ និងទន់ភ្លន់ និយាយយឺតៗដោយសង្កត់សំឡេងបែបអង់គ្លេស", "e.g. claude-opus-4-5": "ឧ. claude-opus-4-5", "e.g. my-model-id": "ឧ. my-model-id", "e.g. tts-1-hd": "ឧ. tts-1-hd", diff --git a/public/i18n/literals/ko.json b/public/i18n/literals/ko.json index 1edb094e..2ca0fc19 100644 --- a/public/i18n/literals/ko.json +++ b/public/i18n/literals/ko.json @@ -12,6 +12,7 @@ "Logout": "로그아웃", "Login": "로그인", "Providers": "제공자", + "Style": "스타일", "Usage": "통계", "API Key": "API 키", "Connected": "연결됨", @@ -176,6 +177,7 @@ "How it works:": "작동 방식:", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "Antigravity/Copilot IDE 요청 → localhost:443로 DNS 리디렉션 → MITM 프록시가 가로챔 → 9Router → Antigravity/Copilot으로 응답", "No API keys — create one in Keys page": "API 키 없음 — 키 페이지에서 만들기", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "예: 따뜻하고 부드러운 목소리로 영국식 억양을 쓰며 천천히 말하기", "sk_9router (default)": "sk_9router (기본값)", "Server started": "서버 시작됨", "Failed to start server": "서버 시작 실패", diff --git a/public/i18n/literals/nl.json b/public/i18n/literals/nl.json index 2eab85ba..9fb47cda 100644 --- a/public/i18n/literals/nl.json +++ b/public/i18n/literals/nl.json @@ -12,6 +12,7 @@ "Logout": "Afmelden", "Login": "Aanmelden", "Providers": "Providers", + "Style": "Stijl", "Usage": "Statistieken", "API Key": "API-sleutel", "Connected": "Verbonden", @@ -176,6 +177,7 @@ "How it works:": "Hoe het werkt:", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "Antigravity/Copilot IDE-aanvraag → DNS-omleiding naar localhost:443 → MITM-proxy onderschept → 9Router → antwoord naar Antigravity/Copilot", "No API keys — create one in Keys page": "Geen API-sleutels — maak er één aan op de pagina Sleutels", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "bijv.: een warme, zachte stem, die langzaam praat met een Brits accent", "sk_9router (default)": "sk_9router (standaard)", "Server started": "Server gestart", "Failed to start server": "Server starten mislukt", diff --git a/public/i18n/literals/no.json b/public/i18n/literals/no.json index e3f52155..50e84505 100644 --- a/public/i18n/literals/no.json +++ b/public/i18n/literals/no.json @@ -12,6 +12,7 @@ "Logout": "Logg ut", "Login": "Logg inn", "Providers": "Leverandører", + "Style": "Stil", "Usage": "Bruksstatistikk", "API Key": "API-nøkkel", "Connected": "Tilkoblet", @@ -176,6 +177,7 @@ "How it works:": "Slik fungerer det:", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "Antigravity/Copilot IDE-forespørsel → DNS-omdirigering til localhost:443 → MITM-proxy avlytt → 9Router → svar til Antigravity/Copilot", "No API keys — create one in Keys page": "Ingen API-nøkler — lag en på Keys-siden", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "f.eks.: en varm, myk stemme, som snakker sakte med britisk aksent", "sk_9router (default)": "sk_9router (standard)", "Server started": "Server startet", "Failed to start server": "Klarte ikke å starte server", diff --git a/public/i18n/literals/pl.json b/public/i18n/literals/pl.json index f4b62a67..3840be97 100644 --- a/public/i18n/literals/pl.json +++ b/public/i18n/literals/pl.json @@ -12,6 +12,7 @@ "Logout": "Wyloguj się", "Login": "Zaloguj się", "Providers": "Dostawcy", + "Style": "Styl", "Usage": "Statystyka", "API Key": "Klucz API", "Connected": "Połączony", @@ -176,6 +177,7 @@ "How it works:": "Jak to działa:", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "Żądanie Antigravity/Copilot IDE → Przekierowanie DNS na localhost:443 → Serwer proxy MITM przechwytuje → 9Router → odpowiedź do Antigravity/Copilot", "No API keys — create one in Keys page": "Brak kluczy API — utwórz jeden na stronie Klucze", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "np.: ciepły, delikatny głos, mówiący powoli z brytyjskim akcentem", "sk_9router (default)": "sk_9router (domyślnie)", "Server started": "Serwer uruchomiony", "Failed to start server": "Nie udało się uruchomić serwera", diff --git a/public/i18n/literals/pt-BR.json b/public/i18n/literals/pt-BR.json index f2b5f6d2..a4186636 100644 --- a/public/i18n/literals/pt-BR.json +++ b/public/i18n/literals/pt-BR.json @@ -864,6 +864,7 @@ "Stopped": "Parado", "Stopping…": "Parando…", "Strict Proxy": "Proxy Estrito", + "Style": "Estilo", "Subagent Model": "Modelo de Subagente", "Subagent model overrides": "Substituições de modelo de subagente", "Submit": "Enviar", @@ -978,6 +979,7 @@ "Your OAuth application client ID": "ID do cliente do seu aplicativo OAuth", "Your requests start from your favorite tools or our unified SDK.": "Suas requisições começam de suas ferramentas favoritas ou do nosso SDK unificado.", "[ml] downloads ~1 GB (torch + huggingface-hub). Continue?": "[ml] baixa ~1 GB (torch + huggingface-hub). Continuar?", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "ex.: voz quente e suave, falando devagar com sotaque britânico", "extras status failed": "falha no status dos extras", "git/grep/ls/tree/logs → 60-90% fewer input tokens": "git/grep/ls/tree/logs → 60-90% menos tokens de entrada", "not installed": "não instalado", @@ -985,4 +987,4 @@ "tree-sitter AST compression for code responses": "Compressão AST tree-sitter para respostas de código", "⚠️ MITM intercepts HTTPS traffic of IDE tools (Antigravity, GitHub Copilot, Kiro) via local CA to redirect requests to your providers. May violate ToS → account ban. Use at your own risk.": "⚠️ MITM intercepta tráfego HTTPS de ferramentas IDE (Antigravity, GitHub Copilot, Kiro) via CA local para redirecionar solicitações aos seus provedores. Pode violar ToS → risco de banimento de conta. Use por sua conta e risco.", "⚠️ Risk Notice: This provider uses a subscription/OAuth session not officially licensed for proxy/router use. Account may be restricted or banned. Use at your own risk.": "⚠️ Aviso de Risco: Este provedor usa uma sessão de assinatura/OAuth não licenciada oficialmente para uso de proxy/roteador. A conta pode ser restrita ou banida. Use por sua conta e risco." -} \ No newline at end of file +} diff --git a/public/i18n/literals/pt-PT.json b/public/i18n/literals/pt-PT.json index c17e932a..dfa1270f 100644 --- a/public/i18n/literals/pt-PT.json +++ b/public/i18n/literals/pt-PT.json @@ -12,6 +12,7 @@ "Logout": "Terminar sessão", "Login": "Iniciar sessão", "Providers": "Fornecedores", + "Style": "Estilo", "Usage": "Estatísticas", "API Key": "Chave API", "Connected": "Ligado", @@ -176,6 +177,7 @@ "How it works:": "Como funciona:", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "Pedido do Antigravity/Copilot IDE → Redirecionamento DNS para localhost:443 → Proxy MITM interceta → 9Router → resposta para Antigravity/Copilot", "No API keys — create one in Keys page": "Sem chaves de API — crie uma na página Chaves", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "ex.: voz quente e suave, a falar devagar com sotaque britânico", "sk_9router (default)": "sk_9router (predefinição)", "Server started": "Servidor iniciado", "Failed to start server": "Falha ao iniciar o servidor", diff --git a/public/i18n/literals/ro.json b/public/i18n/literals/ro.json index 03384415..d9857d51 100644 --- a/public/i18n/literals/ro.json +++ b/public/i18n/literals/ro.json @@ -12,6 +12,7 @@ "Logout": "Ieșire", "Login": "Conectare", "Providers": "Furnizori", + "Style": "Stil", "Usage": "Statistici de utilizare", "API Key": "Cheie API", "Connected": "Conectat", @@ -176,6 +177,7 @@ "How it works:": "Cum funcționează:", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "Solicitare Antigravity/Copilot IDE → Redirecționare DNS la localhost:443 → Proxy MITM interceptează → 9Router → răspuns la Antigravity/Copilot", "No API keys — create one in Keys page": "Nicio cheie API — creați una în pagina Chei", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "ex.: voce caldă și blândă, vorbind încet cu accent britanic", "sk_9router (default)": "sk_9router (implicit)", "Server started": "Server pornit", "Failed to start server": "Nu s-a putut porni serverul", diff --git a/public/i18n/literals/ru.json b/public/i18n/literals/ru.json index 92b9f0e7..c34535ee 100644 --- a/public/i18n/literals/ru.json +++ b/public/i18n/literals/ru.json @@ -12,6 +12,7 @@ "Logout": "Выход", "Login": "Вход", "Providers": "Провайдеры", + "Style": "Стиль", "Usage": "Статистика", "API Key": "Ключ API", "Connected": "Подключено", @@ -176,6 +177,7 @@ "How it works:": "Как это работает:", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "Запрос Antigravity/Copilot IDE → Перенаправление DNS на localhost:443 → Прокси MITM перехватывает → 9Router → ответ для Antigravity/Copilot", "No API keys — create one in Keys page": "Нет ключей API — создайте один на странице ключей", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "напр.: тёплый мягкий голос, медленно говорящий с британским акцентом", "sk_9router (default)": "sk_9router (по умолчанию)", "Server started": "Сервер запущен", "Failed to start server": "Ошибка при запуске сервера", diff --git a/public/i18n/literals/sv.json b/public/i18n/literals/sv.json index 0e4c3b46..ef45f379 100644 --- a/public/i18n/literals/sv.json +++ b/public/i18n/literals/sv.json @@ -12,6 +12,7 @@ "Logout": "Logga ut", "Login": "Logga in", "Providers": "Leverantörer", + "Style": "Stil", "Usage": "Användarstatistik", "API Key": "API-nyckel", "Connected": "Ansluten", @@ -176,6 +177,7 @@ "How it works:": "Hur det fungerar:", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "Antigravity/Copilot IDE-begäran → DNS-omdirigering till localhost:443 → MITM-proxy avlyssnar → 9Router → svar till Antigravity/Copilot", "No API keys — create one in Keys page": "Inga API-nycklar — skapa en på nyckelsidan", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "t.ex.: en varm, mjuk röst som talar långsamt med brittisk accent", "sk_9router (default)": "sk_9router (standard)", "Server started": "Servern startad", "Failed to start server": "Misslyckades att starta servern", diff --git a/public/i18n/literals/th.json b/public/i18n/literals/th.json index 7d528201..bc4d9376 100644 --- a/public/i18n/literals/th.json +++ b/public/i18n/literals/th.json @@ -1119,6 +1119,7 @@ "Stop Server": "หยุดเซิร์ฟเวอร์", "Stopped": "หยุดแล้ว", "Strict Proxy": "Strict Proxy", + "Style": "สไตล์", "Subagent Model": "Subagent Model", "Sudo Password Required": "ต้องใช้ Sudo Password", "Sudo password is required": "ต้องใช้ Sudo password", @@ -1312,6 +1313,7 @@ "disabled": "ปิดใช้งานแล้ว", "dollars per million tokens": "ดอลลาร์ต่อล้าน tokens", "e.g. CwhRBWXzGAHq8TQ4Fs17": "เช่น CwhRBWXzGAHq8TQ4Fs17", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "เช่น เสียงอบอุ่นนุ่มนวล พูดช้าๆ ด้วยสำเนียงอังกฤษ", "e.g. claude-opus-4-5": "เช่น claude-opus-4-5", "e.g. my-model-id": "เช่น my-model-id", "e.g. tts-1-hd": "เช่น tts-1-hd", @@ -1388,4 +1390,4 @@ "✓ Confirm Add": "✓ ยืนยันการเพิ่ม", "📝 Configure providers in dashboard or use environment variables": "📝 กำหนดค่า providers ใน dashboard หรือใช้ environment variables", "🔐 OAuth required. Add now and authenticate after Apply; tool list will be discovered after first connect.": "🔐 ต้องใช้ OAuth เพิ่มตอนนี้แล้ว authenticate หลัง Apply; รายการเครื่องมือจะถูกค้นพบหลังการเชื่อมต่อครั้งแรก" -} \ No newline at end of file +} diff --git a/public/i18n/literals/tl.json b/public/i18n/literals/tl.json index 51af4e24..7510740a 100644 --- a/public/i18n/literals/tl.json +++ b/public/i18n/literals/tl.json @@ -12,6 +12,7 @@ "Logout": "Maglog out", "Login": "Magsimula ng sesyon", "Providers": "Mga Provider", + "Style": "Estilo", "Usage": "Mga Istatistika ng Paggamit", "API Key": "Susi ng API", "Connected": "Konektado", @@ -176,6 +177,7 @@ "How it works:": "Paano ito gumagana:", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "Antigravity/Copilot IDE request → DNS redirect sa localhost:443 → MITM proxy intercepts → 9Router → response sa Antigravity/Copilot", "No API keys — create one in Keys page": "Walang API keys — lumikha ng isa sa Keys page", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "hal.: mainit at malumanay na boses, dahan-dahang nagsasalita nang may British accent", "sk_9router (default)": "sk_9router (default)", "Server started": "Ang server ay nagsimula", "Failed to start server": "Nabigo na magsimula ang server", diff --git a/public/i18n/literals/tr.json b/public/i18n/literals/tr.json index ac4042aa..bd7ec4f4 100644 --- a/public/i18n/literals/tr.json +++ b/public/i18n/literals/tr.json @@ -12,6 +12,7 @@ "Logout": "Çıkış Yap", "Login": "Giriş Yap", "Providers": "Sağlayıcılar", + "Style": "Stil", "Usage": "Kullanım İstatistikleri", "API Key": "API Anahtarı", "Connected": "Bağlı", @@ -176,6 +177,7 @@ "How it works:": "Nasıl çalışır:", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "Antigravity/Copilot IDE isteği → DNS'i localhost:443'e yönlendir → MITM proxy yakalar → 9Router → Antigravity/Copilot'a yanıt", "No API keys — create one in Keys page": "API anahtarı yok — Keys sayfasında bir tane oluşturun", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "örn.: sıcak ve yumuşak bir ses, İngiliz aksanıyla yavaş konuşuyor", "sk_9router (default)": "sk_9router (varsayılan)", "Server started": "Sunucu başlatıldı", "Failed to start server": "Sunucu başlatılamadı", diff --git a/public/i18n/literals/uk.json b/public/i18n/literals/uk.json index e238ad2b..378476ef 100644 --- a/public/i18n/literals/uk.json +++ b/public/i18n/literals/uk.json @@ -12,6 +12,7 @@ "Logout": "Вийти", "Login": "Увійти", "Providers": "Постачальники", + "Style": "Стиль", "Usage": "Статистика використання", "API Key": "Ключ API", "Connected": "Підключено", @@ -176,6 +177,7 @@ "How it works:": "Як це працює:", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "Запит Antigravity/Copilot IDE → Перенаправлення DNS на localhost:443 → MITM проксі перехопити → 9Router → відповідь на Antigravity/Copilot", "No API keys — create one in Keys page": "Немає ключів API — створіть один на сторінці ключів", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "напр.: теплий м'який голос, що повільно говорить з британським акцентом", "sk_9router (default)": "sk_9router (за замовчуванням)", "Server started": "Сервер запущено", "Failed to start server": "Не вдалося запустити сервер", diff --git a/public/i18n/literals/ur.json b/public/i18n/literals/ur.json index e5921755..e0aa8cf7 100644 --- a/public/i18n/literals/ur.json +++ b/public/i18n/literals/ur.json @@ -12,6 +12,7 @@ "Logout": "لاگ آؤٹ", "Login": "لاگ ان", "Providers": "فراہم کنندگان", + "Style": "انداز", "Usage": "استعمال کے اعدادوشمار", "API Key": "API کلید", "Connected": "منسلک", @@ -176,6 +177,7 @@ "How it works:": "یہ کیسے کام کرتا ہے:", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "Antigravity/Copilot IDE درخواست → DNS کو localhost:443 کی طرف ری ڈائریکٹ کریں → MITM پروکسی روکے → 9Router → Antigravity/Copilot کو جواب", "No API keys — create one in Keys page": "کوئی API کلید نہیں — Keys صفحہ میں ایک بنائیں", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "مثلاً: گرم، نرم آواز، برطانوی لہجے کے ساتھ آہستہ بولتی ہوئی", "sk_9router (default)": "sk_9router (ڈیفالٹ)", "Server started": "سرور شروع ہوگیا", "Failed to start server": "سرور شروع کرنے میں ناکام", diff --git a/public/i18n/literals/vi.json b/public/i18n/literals/vi.json index 5358d068..51100780 100644 --- a/public/i18n/literals/vi.json +++ b/public/i18n/literals/vi.json @@ -12,6 +12,7 @@ "Logout": "Đăng xuất", "Login": "Đăng nhập", "Providers": "Nhà cung cấp", + "Style": "Phong cách", "Usage": "Thống kê", "API Key": "Khóa API", "Connected": " Đã kết nối", @@ -176,6 +177,7 @@ "How it works:": "Cách hoạt động:", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "Yêu cầu Antigravity/Copilot IDE → Chuyển hướng DNS đến localhost:443 → MITM proxy chặn → 9Router → phản hồi đến Antigravity/Copilot", "No API keys — create one in Keys page": "Không có khóa API — tạo một khóa trong trang Keys", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "vd.: giọng ấm áp, nhẹ nhàng, nói chậm với giọng Anh-Anh", "sk_9router (default)": "sk_9router (mặc định)", "Server started": "Đã khởi động máy chủ", "Failed to start server": "Không thể khởi động máy chủ", diff --git a/public/i18n/literals/zh-CN.json b/public/i18n/literals/zh-CN.json index 8459ae22..da045742 100644 --- a/public/i18n/literals/zh-CN.json +++ b/public/i18n/literals/zh-CN.json @@ -1119,6 +1119,7 @@ "Stop Server": "停止服务器", "Stopped": "已停止", "Strict Proxy": "严格代理", + "Style": "风格", "Subagent Model": "子代理模型", "Sudo Password Required": "需要 sudo 密码", "Sudo password is required": "需要 sudo 密码", @@ -1312,6 +1313,7 @@ "disabled": "已禁用", "dollars per million tokens": "美元 / 百万 Token", "e.g. CwhRBWXzGAHq8TQ4Fs17": "例如 CwhRBWXzGAHq8TQ4Fs17", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "例如:温暖柔和的嗓音,用英式口音缓慢说话", "e.g. claude-opus-4-5": "例如 claude-opus-4-5", "e.g. my-model-id": "例如 my-model-id", "e.g. tts-1-hd": "例如 tts-1-hd", diff --git a/public/i18n/literals/zh-TW.json b/public/i18n/literals/zh-TW.json index ea9183eb..f8f41a86 100644 --- a/public/i18n/literals/zh-TW.json +++ b/public/i18n/literals/zh-TW.json @@ -12,6 +12,7 @@ "Logout": "登出", "Login": "登錄", "Providers": "提供者", + "Style": "風格", "Usage": "統計", "API Key": "API 金鑰", "Connected": "已連接", @@ -176,6 +177,7 @@ "How it works:": "工作原理:", "Antigravity/Copilot IDE request → DNS redirect to localhost:443 → MITM proxy intercepts → 9Router → response to Antigravity/Copilot": "Antigravity/Copilot IDE 請求 → DNS 重定向到 localhost:443 → MITM 代理攔截 → 9Router → 响應到 Antigravity/Copilot", "No API keys — create one in Keys page": "沒有 API 金鑰 — 在金鑰頁面中創建一個", + "e.g. a warm, gentle voice, speaking slowly with a British accent": "例如:溫暖柔和的嗓音,用英式口音緩慢說話", "sk_9router (default)": "sk_9router(默認)", "Server started": "服務器已啟動", "Failed to start server": "啟動服務器失敗", diff --git a/public/providers/opendesign.png b/public/providers/opendesign.png new file mode 100644 index 00000000..c68169aa Binary files /dev/null and b/public/providers/opendesign.png differ diff --git a/public/providers/tokenrouter.png b/public/providers/tokenrouter.png new file mode 100644 index 00000000..58f7eac2 Binary files /dev/null and b/public/providers/tokenrouter.png differ diff --git a/scripts/copy-standalone-assets.mjs b/scripts/copy-standalone-assets.mjs new file mode 100644 index 00000000..bfaf6e0d --- /dev/null +++ b/scripts/copy-standalone-assets.mjs @@ -0,0 +1,36 @@ +import { cpSync, existsSync } from "node:fs"; +import { fileURLToPath } from "node:url"; +import { dirname, resolve } from "node:path"; + +export function copyStandaloneAssets({ projectRoot = process.cwd(), distDir = process.env.NEXT_DIST_DIR || ".next" } = {}) { + if (process.env.NEXT_TRACING_ROOT_MODE === "workspace") { + console.log("[standalone-assets] Skipping workspace-traced CLI build; CLI packaging handles assets"); + return; + } + + const buildDir = resolve(projectRoot, distDir); + const standaloneDir = resolve(buildDir, "standalone"); + + if (!existsSync(standaloneDir)) { + console.log(`[standalone-assets] No standalone build found at ${standaloneDir}`); + return; + } + + const staticSource = resolve(buildDir, "static"); + const staticDestination = resolve(standaloneDir, distDir, "static"); + if (existsSync(staticSource)) { + cpSync(staticSource, staticDestination, { recursive: true, force: true }); + console.log(`[standalone-assets] Copied static assets to ${staticDestination}`); + } + + const publicSource = resolve(projectRoot, "public"); + const publicDestination = resolve(standaloneDir, "public"); + if (existsSync(publicSource)) { + cpSync(publicSource, publicDestination, { recursive: true, force: true }); + console.log(`[standalone-assets] Copied public assets to ${publicDestination}`); + } +} + +if (process.argv[1] && resolve(process.argv[1]) === resolve(dirname(fileURLToPath(import.meta.url)), "copy-standalone-assets.mjs")) { + copyStandaloneAssets(); +} diff --git a/src/app/(dashboard)/dashboard/cli-tools/[toolId]/ToolDetailClient.js b/src/app/(dashboard)/dashboard/cli-tools/[toolId]/ToolDetailClient.js index 2e735647..209a6d33 100644 --- a/src/app/(dashboard)/dashboard/cli-tools/[toolId]/ToolDetailClient.js +++ b/src/app/(dashboard)/dashboard/cli-tools/[toolId]/ToolDetailClient.js @@ -81,6 +81,33 @@ export default function ToolDetailClient({ toolId, machineId }) { models.push({ value: modelValue, label: `${alias}/${m.id}`, provider: conn.provider, alias, connectionName: conn.name, modelId: m.id }); } }); + + // openai/anthropic-compatible providers are registered with a random UUID (e.g. + // "openai-compatible-chat-") that has no entry in the static PROVIDER_MODELS + // catalog, so `getModelsByProviderId` returns []. Routing still works because the + // request path uses the connection's own model config, but `hasActiveProviders` + // below would flip to false and disable the Apply button. Fall back to the + // connection's own models so these providers are usable from CLI tool pages. + if (providerModels.length === 0) { + const prefix = conn.providerSpecificData?.prefix || alias; + const fallbackModels = []; + if (conn.defaultModel) fallbackModels.push({ id: conn.defaultModel, name: conn.defaultModel }); + (conn.providerSpecificData?.customModels || []).forEach(m => { + if (m?.id && !fallbackModels.some(f => f.id === m.id)) fallbackModels.push({ id: m.id, name: m.name || m.id }); + }); + if (fallbackModels.length === 0 && conn.testStatus === "active") { + // Provider is confirmed reachable but exposes no model info anywhere; + // still let the user apply so they aren't stuck on a permanently disabled button. + fallbackModels.push({ id: "model-id", name: `${prefix}/model-id` }); + } + fallbackModels.forEach(m => { + const modelValue = `${prefix}/${m.id}`; + if (!seenModels.has(modelValue)) { + seenModels.add(modelValue); + models.push({ value: modelValue, label: `${prefix}/${m.id}`, provider: conn.provider, alias: prefix, connectionName: conn.name, modelId: m.id }); + } + }); + } }); return models; }; diff --git a/src/app/(dashboard)/dashboard/combos/page.js b/src/app/(dashboard)/dashboard/combos/page.js index 59fb50f4..fa093bfd 100644 --- a/src/app/(dashboard)/dashboard/combos/page.js +++ b/src/app/(dashboard)/dashboard/combos/page.js @@ -1,818 +1,855 @@ "use client"; import { useState, useEffect, useCallback } from "react"; -import { - DndContext, - closestCenter, - KeyboardSensor, - PointerSensor, - useSensor, - useSensors, -} from "@dnd-kit/core"; -import { - arrayMove, - SortableContext, - sortableKeyboardCoordinates, - useSortable, - verticalListSortingStrategy, -} from "@dnd-kit/sortable"; +import { DndContext, closestCenter, KeyboardSensor, PointerSensor, useSensor, useSensors } from "@dnd-kit/core"; +import { arrayMove, SortableContext, sortableKeyboardCoordinates, useSortable, verticalListSortingStrategy } from "@dnd-kit/sortable"; import { CSS } from "@dnd-kit/utilities"; -import { - restrictToVerticalAxis, - restrictToParentElement, -} from "@dnd-kit/modifiers"; -import { - Card, - Button, - Modal, - Input, - CardSkeleton, - ModelSelectModal, - ConfirmModal, - CapacityBadges, - Select, - Toggle, -} from "@/shared/components"; +import { restrictToVerticalAxis, restrictToParentElement } from "@dnd-kit/modifiers"; +import { Card, Button, Modal, Input, CardSkeleton, ModelSelectModal, ConfirmModal, CapacityBadges, Select, Toggle } from "@/shared/components"; import { useCopyToClipboard } from "@/shared/hooks/useCopyToClipboard"; import { useModelCaps } from "@/shared/hooks/useModelCaps"; -import { - isOpenAICompatibleProvider, - isAnthropicCompatibleProvider, -} from "@/shared/constants/providers"; +import { isOpenAICompatibleProvider, isAnthropicCompatibleProvider } from "@/shared/constants/providers"; // Validate combo name: only a-z, A-Z, 0-9, -, _ const VALID_NAME_REGEX = /^[a-zA-Z0-9_.\-]+$/; +// Capacity adapter: global fallback pools of models per input-modality capability. +// A request needing a capability the target model/combo lacks switches straight +// to the first enabled model here instead of erroring or dropping the data. +const CAPACITY_ADAPTER_CAPS = [ + { key: "vision", label: "Vision", icon: "visibility", desc: "Images" }, + // pdf, videoInput temporarily hidden — no translator support yet for those blocks. + { key: "audioInput", label: "Audio", icon: "graphic_eq", desc: "Audio input" }, +]; +const DEFAULT_FALLBACK_MODEL = "oc/mimo-v2.5-free"; +const EMPTY_CAP_ENTRY = { enabled: true, roundRobin: false, models: [] }; +const EMPTY_CAPACITY_ADAPTER = { + vision: { ...EMPTY_CAP_ENTRY }, + pdf: { ...EMPTY_CAP_ENTRY }, + audioInput: { ...EMPTY_CAP_ENTRY }, + videoInput: { ...EMPTY_CAP_ENTRY }, +}; +// Backward-compat: legacy stored form was an array of {model, enabled}. +function normalizeCapEntry(entry) { + if (Array.isArray(entry)) { + return { enabled: true, roundRobin: false, models: entry.map((e) => e?.model || e).filter(Boolean) }; + } + if (entry && typeof entry === "object") { + return { + enabled: entry.enabled !== false, + roundRobin: !!entry.roundRobin, + models: Array.isArray(entry.models) ? entry.models.filter(Boolean) : [], + }; + } + return { ...EMPTY_CAP_ENTRY }; +} + export default function CombosPage() { - const [combos, setCombos] = useState([]); - const [loading, setLoading] = useState(true); - const [showCreateModal, setShowCreateModal] = useState(false); - const [editingCombo, setEditingCombo] = useState(null); - const [activeProviders, setActiveProviders] = useState([]); - const [comboStrategies, setComboStrategies] = useState({}); - const { getCaps } = useModelCaps(); - const [confirmState, setConfirmState] = useState(null); - const { copied, copy } = useCopyToClipboard(); + const [combos, setCombos] = useState([]); + const [loading, setLoading] = useState(true); + const [showCreateModal, setShowCreateModal] = useState(false); + const [editingCombo, setEditingCombo] = useState(null); + const [activeProviders, setActiveProviders] = useState([]); + const [comboStrategies, setComboStrategies] = useState({}); + const [capacityAdapter, setCapacityAdapter] = useState(EMPTY_CAPACITY_ADAPTER); + const { getCaps } = useModelCaps(); + const [confirmState, setConfirmState] = useState(null); + const { copied, copy } = useCopyToClipboard(); - useEffect(() => { - fetchData(); - }, []); + useEffect(() => { + fetchData(); + }, []); // eslint-disable-line react-hooks/exhaustive-deps - const fetchData = async () => { - try { - const [combosRes, providersRes, settingsRes] = await Promise.all([ - fetch("/api/combos"), - fetch("/api/providers"), - fetch("/api/settings"), - ]); - const combosData = await combosRes.json(); - const providersData = await providersRes.json(); - const settingsData = settingsRes.ok ? await settingsRes.json() : {}; + const fetchData = async () => { + try { + const [combosRes, providersRes, settingsRes] = await Promise.all([ + fetch("/api/combos"), + fetch("/api/providers"), + fetch("/api/settings"), + ]); + const combosData = await combosRes.json(); + const providersData = await providersRes.json(); + const settingsData = settingsRes.ok ? await settingsRes.json() : {}; + + // Only LLM combos here - webSearch/webFetch combos belong to media-providers/web + if (combosRes.ok) setCombos((combosData.combos || []).filter(c => !c.kind || c.kind === "llm")); + if (providersRes.ok) { + setActiveProviders(providersData.connections || []); + } + setComboStrategies(settingsData.comboStrategies || {}); + const rawAdapter = settingsData.capacityAdapter || {}; + const normalized = {}; + for (const cap of CAPACITY_ADAPTER_CAPS) { + normalized[cap.key] = normalizeCapEntry(rawAdapter[cap.key]); + } + setCapacityAdapter(normalized); + } catch (error) { + console.log("Error fetching data:", error); + } finally { + setLoading(false); + } + }; - // Only LLM combos here - webSearch/webFetch combos belong to media-providers/web - if (combosRes.ok) - setCombos( - (combosData.combos || []).filter((c) => !c.kind || c.kind === "llm"), - ); - if (providersRes.ok) { - setActiveProviders(providersData.connections || []); - } - setComboStrategies(settingsData.comboStrategies || {}); - } catch (error) { - console.log("Error fetching data:", error); - } finally { - setLoading(false); - } - }; + const handleSetCapacityAdapter = async (next) => { + setCapacityAdapter(next); + try { + await fetch("/api/settings", { + method: "PATCH", + headers: { "Content-Type": "application/json" }, + body: JSON.stringify({ capacityAdapter: next }), + }); + } catch (error) { + console.log("Error updating capacity adapter:", error); + } + }; - const handleCreate = async (data) => { - try { - const res = await fetch("/api/combos", { - method: "POST", - headers: { "Content-Type": "application/json" }, - body: JSON.stringify(data), - }); - if (res.ok) { - await fetchData(); - setShowCreateModal(false); - } else { - const err = await res.json(); - alert(err.error || "Failed to create combo"); - } - } catch (error) { - console.log("Error creating combo:", error); - } - }; + const handleCreate = async (data) => { + try { + const res = await fetch("/api/combos", { + method: "POST", + headers: { "Content-Type": "application/json" }, + body: JSON.stringify(data), + }); + if (res.ok) { + await fetchData(); + setShowCreateModal(false); + } else { + const err = await res.json(); + alert(err.error || "Failed to create combo"); + } + } catch (error) { + console.log("Error creating combo:", error); + } + }; - const handleUpdate = async (id, data) => { - try { - const res = await fetch(`/api/combos/${id}`, { - method: "PUT", - headers: { "Content-Type": "application/json" }, - body: JSON.stringify(data), - }); - if (res.ok) { - await fetchData(); - setEditingCombo(null); - } else { - const err = await res.json(); - alert(err.error || "Failed to update combo"); - } - } catch (error) { - console.log("Error updating combo:", error); - } - }; + const handleUpdate = async (id, data) => { + try { + const res = await fetch(`/api/combos/${id}`, { + method: "PUT", + headers: { "Content-Type": "application/json" }, + body: JSON.stringify(data), + }); + if (res.ok) { + await fetchData(); + setEditingCombo(null); + } else { + const err = await res.json(); + alert(err.error || "Failed to update combo"); + } + } catch (error) { + console.log("Error updating combo:", error); + } + }; - const handleDelete = async (id) => { - setConfirmState({ - title: "Delete Combo", - message: "Delete this combo?", - onConfirm: async () => { - setConfirmState(null); - try { - const res = await fetch(`/api/combos/${id}`, { method: "DELETE" }); - if (res.ok) { - setCombos(combos.filter((c) => c.id !== id)); - } - } catch (error) { - console.log("Error deleting combo:", error); - } - }, - }); - }; + const handleDelete = async (id) => { + setConfirmState({ + title: "Delete Combo", + message: "Delete this combo?", + onConfirm: async () => { + setConfirmState(null); + try { + const res = await fetch(`/api/combos/${id}`, { method: "DELETE" }); + if (res.ok) { + setCombos(combos.filter(c => c.id !== id)); + } + } catch (error) { + console.log("Error deleting combo:", error); + } + } + }); + }; - // Merge a per-combo strategy patch into settings.comboStrategies. Passing an empty - // patch (strategy back to default "fallback") drops the entry entirely. - const handleSetComboStrategy = async (comboName, patch) => { - try { - const updated = { ...comboStrategies }; - const next = { ...(updated[comboName] || {}), ...patch }; - // Prune to keep settings clean: default fallback with no extras = no entry. - if (!next.fallbackStrategy || next.fallbackStrategy === "fallback") { - delete updated[comboName]; - } else { - updated[comboName] = next; - } + // Merge a per-combo strategy patch into settings.comboStrategies. Passing an empty + // patch (strategy back to default "fallback") drops the entry entirely. + const handleSetComboStrategy = async (comboName, patch) => { + try { + const updated = { ...comboStrategies }; + const next = { ...(updated[comboName] || {}), ...patch }; + // Prune to keep settings clean: default fallback with no extras = no entry. + if (!next.fallbackStrategy || next.fallbackStrategy === "fallback") { + delete updated[comboName]; + } else { + updated[comboName] = next; + } - await fetch("/api/settings", { - method: "PATCH", - headers: { "Content-Type": "application/json" }, - body: JSON.stringify({ comboStrategies: updated }), - }); + await fetch("/api/settings", { + method: "PATCH", + headers: { "Content-Type": "application/json" }, + body: JSON.stringify({ comboStrategies: updated }), + }); - setComboStrategies(updated); - } catch (error) { - console.log("Error updating combo strategy:", error); - } - }; + setComboStrategies(updated); + } catch (error) { + console.log("Error updating combo strategy:", error); + } + }; - if (loading) { - return ( -
- - -
- ); - } + if (loading) { + return ( +
+ + +
+ ); + } - return ( -
- {/* Header */} -
-
-

- Group models under one name, then pick a strategy per combo: -

-
    -
  • - Fallback — - tries models in order (next on failure) -
  • -
  • - Round Robin — - rotates models across requests to spread load -
  • -
  • - Fusion — - queries all models in parallel, then a judge synthesizes one - answer. Best quality, but costs the most: every request bills all - panel models + the judge (N+1 calls) -
  • -
  • - - Capacity auto-switch - {" "} - — sends image/PDF/audio requests to a model that supports them - first -
  • -
-
- -
+ return ( +
+ {/* Header */} +
+
+

+ Group models under one name, then pick a strategy per combo: +

+
    +
  • Fallback — tries models in order (next on failure)
  • +
  • Round Robin — rotates models across requests to spread load
  • +
  • Fusion — queries all models in parallel, then a judge synthesizes one answer. Best quality, but costs the most: every request bills all panel models + the judge (N+1 calls)
  • +
+
+ +
- {/* Combos List */} - {combos.length === 0 ? ( - -
-
- - layers - -
-

No combos yet

-

- Create model combos with fallback support -

- -
-
- ) : ( -
- {combos.map((combo) => ( - setEditingCombo(combo)} - onDelete={() => handleDelete(combo.id)} - onToggleEnabled={(enabled) => handleUpdate(combo.id, { enabled })} - strategy={comboStrategies[combo.name] || {}} - onSetStrategy={(patch) => - handleSetComboStrategy(combo.name, patch) - } - /> - ))} -
- )} + {/* Combos List */} + {combos.length === 0 ? ( + +
+
+ layers +
+

No combos yet

+

Create model combos with fallback support

+ +
+
+ ) : ( +
+ {combos.map((combo) => ( + setEditingCombo(combo)} + onDelete={() => handleDelete(combo.id)} + strategy={comboStrategies[combo.name] || {}} + onSetStrategy={(patch) => handleSetComboStrategy(combo.name, patch)} + /> + ))} +
+ )} - {/* Create Modal - Use key to force remount and reset state */} - {showCreateModal && ( - setShowCreateModal(false)} - onSave={handleCreate} - activeProviders={activeProviders} - /> - )} + {/* Capacity Adapter */} + - {editingCombo && ( - setEditingCombo(null)} - onSave={(data) => handleUpdate(editingCombo.id, data)} - activeProviders={activeProviders} - /> - )} + {/* Create Modal - Use key to force remount and reset state */} + {showCreateModal && ( + setShowCreateModal(false)} + onSave={handleCreate} + activeProviders={activeProviders} + /> + )} - {/* Confirm Delete Modal */} - setConfirmState(null)} - onConfirm={confirmState?.onConfirm} - title={confirmState?.title || "Confirm"} - message={confirmState?.message} - variant="danger" - /> -
- ); + {editingCombo && ( + setEditingCombo(null)} + onSave={(data) => handleUpdate(editingCombo.id, data)} + activeProviders={activeProviders} + /> + )} + + {/* Confirm Delete Modal */} + setConfirmState(null)} + onConfirm={confirmState?.onConfirm} + title={confirmState?.title || "Confirm"} + message={confirmState?.message} + variant="danger" + /> +
+ ); } const STRATEGY_OPTIONS = [ - { value: "fallback", label: "Fallback — try in order" }, - { value: "round-robin", label: "Round Robin — rotate" }, - { value: "fusion", label: "Fusion — panel + judge" }, + { value: "fallback", label: "Fallback — try in order" }, + { value: "round-robin", label: "Round Robin — rotate" }, + { value: "fusion", label: "Fusion — panel + judge" }, ]; -function ComboCard({ - combo, - getCaps, - activeProviders = [], - copied, - onCopy, - onEdit, - onDelete, - strategy = {}, - onSetStrategy, - onToggleEnabled, -}) { - const [showJudgeSelect, setShowJudgeSelect] = useState(false); - const current = strategy.fallbackStrategy || "fallback"; - const judge = strategy.judgeModel || ""; - const isFusion = current === "fusion"; +function ComboCard({ combo, getCaps, activeProviders = [], copied, onCopy, onEdit, onDelete, strategy = {}, onSetStrategy }) { + const [showJudgeSelect, setShowJudgeSelect] = useState(false); + const current = strategy.fallbackStrategy || "fallback"; + const judge = strategy.judgeModel || ""; + const isFusion = current === "fusion"; - return ( - -
-
-
- - layers - -
-
- - {combo.name} - -
- {combo.models.length === 0 ? ( - - No models - - ) : ( - combo.models.slice(0, 3).map((model, index) => ( - - {model} - - - )) - )} - {combo.models.length > 3 && ( - - +{combo.models.length - 3} more - - )} -
- {/* Fusion: judge picker (Auto = first model) */} - {isFusion && ( -
- - Judge - - - {judge && ( - - )} -
- )} -
-
+ return ( + +
+
+
+ layers +
+
+ {combo.name} +
+ {combo.models.length === 0 ? ( + No models + ) : ( + combo.models.slice(0, 3).map((model, index) => ( + + {model} + + + )) + )} + {combo.models.length > 3 && ( + +{combo.models.length - 3} more + )} +
+ {/* Fusion: judge picker (Auto = first model) */} + {isFusion && ( +
+ Judge + + {judge && ( + + )} +
+ )} +
+
- {/* Actions */} -
- {/* Strategy selector — always visible */} -
- onSetStrategy({ fallbackStrategy: e.target.value })} + selectClassName="py-1.5 text-xs" + /> +
- {onToggleEnabled && ( - onToggleEnabled(checked)} - /> - )} -
- - - -
-
-
+
+ + + +
+
+
- {/* Judge model picker (single-select; combo members make natural judges too) */} - {showJudgeSelect && ( - setShowJudgeSelect(false)} - onSelect={(m) => { - onSetStrategy({ judgeModel: m?.value || "" }); - setShowJudgeSelect(false); - }} - activeProviders={activeProviders} - title="Select Judge Model" - addedModelValues={judge ? [judge] : []} - closeOnSelect={true} - /> - )} - - ); + {/* Judge model picker (single-select; combo members make natural judges too) */} + {showJudgeSelect && ( + setShowJudgeSelect(false)} + onSelect={(m) => { onSetStrategy({ judgeModel: m?.value || "" }); setShowJudgeSelect(false); }} + activeProviders={activeProviders} + title="Select Judge Model" + addedModelValues={judge ? [judge] : []} + closeOnSelect={true} + /> + )} + + ); } -function ModelItem({ - id, - index, - model, - isFirst, - isLast, - onEdit, - onMoveUp, - onMoveDown, - onRemove, -}) { - const { attributes, listeners, setNodeRef, transform, isDragging } = - useSortable({ id }); - const style = { - transform: CSS.Transform.toString(transform), - // no transition — prevents the CSS settle animation fighting React's re-render on drop - opacity: isDragging ? 0.4 : 1, - zIndex: isDragging ? 999 : undefined, - }; - const [editing, setEditing] = useState(false); - const [draft, setDraft] = useState(model); - const commit = () => { - const trimmed = draft.trim(); - if (trimmed && trimmed !== model) onEdit(trimmed); - else setDraft(model); - setEditing(false); - }; - - const handleKeyDown = (e) => { - if (e.key === "Enter") commit(); - if (e.key === "Escape") { - setDraft(model); - setEditing(false); - } - }; - - return ( -
- {/* Drag handle */} - - - {/* Index badge */} - - {index + 1} - - - {/* Inline editable model value */} - {editing ? ( - setDraft(e.target.value)} - onBlur={commit} - onKeyDown={handleKeyDown} - className="min-w-0 flex-1 rounded border border-primary/40 bg-white px-1.5 py-0.5 font-mono text-xs text-text-main outline-none dark:bg-black/20" - /> - ) : ( -
setEditing(true)} - title="Click to edit" - > - {model} -
- )} - - {/* Priority arrows */} -
- - -
- - {/* Remove */} - -
- ); +function CapacityAdapterSection({ capacityAdapter, onChange, activeProviders, getCaps }) { + return ( +
+
+
+

Vision Adapter

+

+ Your model can't read image/audio? Auto-switches to a model in the pool below. +

+
    +
  • Vision — images (png, jpg, webp, …)
  • +
  • Audio — audio input
  • +
+
+
+
+ {CAPACITY_ADAPTER_CAPS.map((cap) => ( + onChange({ ...capacityAdapter, [cap.key]: entry })} + activeProviders={activeProviders} + getCaps={getCaps} + /> + ))} +
+
+ ); } -function ComboFormModal({ - isOpen, - combo, - onClose, - onSave, - activeProviders, - kindFilter = null, -}) { - // Initialize state with combo values - key prop on parent handles reset on remount - const [name, setName] = useState(combo?.name || ""); - const [models, setModels] = useState(combo?.models || []); - const [showModelSelect, setShowModelSelect] = useState(false); - const [saving, setSaving] = useState(false); - const [nameError, setNameError] = useState(""); - const [modelAliases, setModelAliases] = useState({}); +function CapacityAdapterCap({ cap, entry, onChange, activeProviders, getCaps }) { + const [showModelSelect, setShowModelSelect] = useState(false); + const { enabled, roundRobin, models } = entry; - const sensors = useSensors( - useSensor(PointerSensor, { activationConstraint: { distance: 5 } }), - useSensor(KeyboardSensor, { - coordinateGetter: sortableKeyboardCoordinates, - }), - ); + const patch = (p) => onChange({ ...entry, ...p }); - // Use stable index-based IDs so duplicates and similar names are handled correctly - const modelItems = models.map((model, i) => ({ uid: `item-${i}`, model })); + const handleAdd = (model) => { + if (models.includes(model.value)) return; + patch({ models: [...models, model.value] }); + }; - const handleDragEnd = (event) => { - const { active, over } = event; - if (over && active.id !== over.id) { - const oldIndex = modelItems.findIndex((m) => m.uid === active.id); - const newIndex = modelItems.findIndex((m) => m.uid === over.id); - if (oldIndex !== -1 && newIndex !== -1) { - setModels((prev) => arrayMove(prev, oldIndex, newIndex)); - } - } - }; + const handleRemove = (index) => { + const next = models.filter((_, i) => i !== index); + patch({ models: next.length === 0 ? [DEFAULT_FALLBACK_MODEL] : next }); + }; - const fetchModalData = async () => { - try { - const aliasesRes = await fetch("/api/models/alias"); - if (!aliasesRes.ok) return; - const aliasesData = await aliasesRes.json(); - setModelAliases(aliasesData.aliases || {}); - } catch (error) { - console.error("Error fetching modal data:", error); - } - }; + const handleMove = (index, delta) => { + const target = index + delta; + if (target < 0 || target >= models.length) return; + const next = [...models]; + [next[index], next[target]] = [next[target], next[index]]; + patch({ models: next }); + }; - useEffect(() => { - if (isOpen) fetchModalData(); - }, [isOpen]); + return ( + +
+ {/* Master toggle + icon + label + chips */} +
+ patch({ enabled: v })} + aria-label={`Enable ${cap.label} adapter`} + /> +
+ {cap.icon} +
+
+
+ {cap.label} + — {cap.desc} +
+
+ {models.length === 0 ? ( + No models + ) : ( + models.slice(0, 3).map((model, index) => ( + + {model} + + + + + + )) + )} + {models.length > 3 && ( + +{models.length - 3} more + )} +
+
+
- const validateName = (value) => { - if (!value.trim()) { - setNameError("Name is required"); - return false; - } - if (!VALID_NAME_REGEX.test(value)) { - setNameError("Only letters, numbers, -, _ and . allowed"); - return false; - } - setNameError(""); - return true; - }; + {/* Actions: Round-robin toggle + Add Model */} +
+ + +
+
- const handleNameChange = (e) => { - const value = e.target.value; - setName(value); - if (value) validateName(value); - else setNameError(""); - }; - - const handleAddModel = (model) => { - if (!models.includes(model.value)) { - setModels([...models, model.value]); - } - }; - - const handleDeselectModel = (model) => { - setModels(models.filter((m) => m !== model.value)); - }; - - const handleRemoveModel = (index) => { - setModels(models.filter((_, i) => i !== index)); - }; - - const handleMoveUp = (index) => { - if (index === 0) return; - const newModels = [...models]; - [newModels[index - 1], newModels[index]] = [ - newModels[index], - newModels[index - 1], - ]; - setModels(newModels); - }; - - const handleMoveDown = (index) => { - if (index === models.length - 1) return; - const newModels = [...models]; - [newModels[index], newModels[index + 1]] = [ - newModels[index + 1], - newModels[index], - ]; - setModels(newModels); - }; - - const handleSave = async () => { - if (!validateName(name)) return; - setSaving(true); - await onSave({ name: name.trim(), models }); - setSaving(false); - }; - - const isEdit = !!combo; - - return ( - <> - -
- {/* Name */} -
- -

- Only letters, numbers, -, _ and . allowed -

-
- - {/* Models */} -
- - - {models.length === 0 ? ( -
- - layers - -

No models added yet

-
- ) : ( - - m.uid)} - strategy={verticalListSortingStrategy} - > -
- {modelItems.map(({ uid, model }, index) => ( - { - const updated = [...models]; - updated[index] = newVal; - setModels(updated); - }} - onMoveUp={() => handleMoveUp(index)} - onMoveDown={() => handleMoveDown(index)} - onRemove={() => handleRemoveModel(index)} - /> - ))} -
-
-
- )} - - {/* Add Model button */} - -
- - {/* Actions */} -
- - -
-
-
- - {/* Model Select Modal */} - {showModelSelect && ( - setShowModelSelect(false)} - onSelect={handleAddModel} - onDeselect={handleDeselectModel} - activeProviders={activeProviders} - modelAliases={modelAliases} - title="Add Model to Combo" - kindFilter={kindFilter} - addedModelValues={models} - closeOnSelect={false} - /> - )} - - ); + {showModelSelect && ( + setShowModelSelect(false)} + onSelect={handleAdd} + activeProviders={activeProviders} + title={`Add ${cap.label} Model`} + addedModelValues={models} + capFilter={cap.key} + closeOnSelect={false} + /> + )} +
+ ); +} + +function ModelItem({ id, index, model, isFirst, isLast, onEdit, onMoveUp, onMoveDown, onRemove }) { + const { attributes, listeners, setNodeRef, transform, isDragging } = useSortable({ id }); + const style = { + transform: CSS.Transform.toString(transform), + // no transition — prevents the CSS settle animation fighting React's re-render on drop + opacity: isDragging ? 0.4 : 1, + zIndex: isDragging ? 999 : undefined, + }; + const [editing, setEditing] = useState(false); + const [draft, setDraft] = useState(model); + const commit = () => { + const trimmed = draft.trim(); + if (trimmed && trimmed !== model) onEdit(trimmed); + else setDraft(model); + setEditing(false); + }; + + const handleKeyDown = (e) => { + if (e.key === "Enter") commit(); + if (e.key === "Escape") { setDraft(model); setEditing(false); } + }; + + return ( +
+ {/* Drag handle */} + + + {/* Index badge */} + {index + 1} + + {/* Inline editable model value */} + {editing ? ( + setDraft(e.target.value)} + onBlur={commit} + onKeyDown={handleKeyDown} + className="min-w-0 flex-1 rounded border border-primary/40 bg-white px-1.5 py-0.5 font-mono text-xs text-text-main outline-none dark:bg-black/20" + /> + ) : ( +
setEditing(true)} + title="Click to edit" + > + {model} +
+ )} + + {/* Priority arrows */} +
+ + +
+ + {/* Remove */} + +
+ ); +} + +function ComboFormModal({ isOpen, combo, onClose, onSave, activeProviders, kindFilter = null }) { + // Initialize state with combo values - key prop on parent handles reset on remount + const [name, setName] = useState(combo?.name || ""); + const [models, setModels] = useState(combo?.models || []); + const [showModelSelect, setShowModelSelect] = useState(false); + const [saving, setSaving] = useState(false); + const [nameError, setNameError] = useState(""); + const [modelAliases, setModelAliases] = useState({}); + + const sensors = useSensors( + useSensor(PointerSensor, { activationConstraint: { distance: 5 } }), + useSensor(KeyboardSensor, { coordinateGetter: sortableKeyboardCoordinates }) + ); + + // Use stable index-based IDs so duplicates and similar names are handled correctly + const modelItems = models.map((model, i) => ({ uid: `item-${i}`, model })); + + const handleDragEnd = (event) => { + const { active, over } = event; + if (over && active.id !== over.id) { + const oldIndex = modelItems.findIndex((m) => m.uid === active.id); + const newIndex = modelItems.findIndex((m) => m.uid === over.id); + if (oldIndex !== -1 && newIndex !== -1) { + setModels((prev) => arrayMove(prev, oldIndex, newIndex)); + } + } + }; + + const fetchModalData = async () => { + try { + const aliasesRes = await fetch("/api/models/alias"); + if (!aliasesRes.ok) return; + const aliasesData = await aliasesRes.json(); + setModelAliases(aliasesData.aliases || {}); + } catch (error) { + console.error("Error fetching modal data:", error); + } + }; + + useEffect(() => { + if (isOpen) fetchModalData(); + }, [isOpen]); + + const validateName = (value) => { + if (!value.trim()) { + setNameError("Name is required"); + return false; + } + if (!VALID_NAME_REGEX.test(value)) { + setNameError("Only letters, numbers, -, _ and . allowed"); + return false; + } + setNameError(""); + return true; + }; + + const handleNameChange = (e) => { + const value = e.target.value; + setName(value); + if (value) validateName(value); + else setNameError(""); + }; + + const handleAddModel = (model) => { + if (!models.includes(model.value)) { + setModels([...models, model.value]); + } + }; + + const handleDeselectModel = (model) => { + setModels(models.filter((m) => m !== model.value)); + }; + + const handleRemoveModel = (index) => { + setModels(models.filter((_, i) => i !== index)); + }; + + const handleMoveUp = (index) => { + if (index === 0) return; + const newModels = [...models]; + [newModels[index - 1], newModels[index]] = [newModels[index], newModels[index - 1]]; + setModels(newModels); + }; + + const handleMoveDown = (index) => { + if (index === models.length - 1) return; + const newModels = [...models]; + [newModels[index], newModels[index + 1]] = [newModels[index + 1], newModels[index]]; + setModels(newModels); + }; + + const handleSave = async () => { + if (!validateName(name)) return; + setSaving(true); + await onSave({ name: name.trim(), models }); + setSaving(false); + }; + + const isEdit = !!combo; + + return ( + <> + +
+ {/* Name */} +
+ +

+ Only letters, numbers, -, _ and . allowed +

+
+ + {/* Models */} +
+ + + {models.length === 0 ? ( +
+ layers +

No models added yet

+
+ ) : ( + + m.uid)} strategy={verticalListSortingStrategy}> +
+ {modelItems.map(({ uid, model }, index) => ( + { + const updated = [...models]; + updated[index] = newVal; + setModels(updated); + }} + onMoveUp={() => handleMoveUp(index)} + onMoveDown={() => handleMoveDown(index)} + onRemove={() => handleRemoveModel(index)} + /> + ))} +
+
+
+ )} + + {/* Add Model button */} + +
+ + {/* Actions */} +
+ + +
+
+
+ + {/* Model Select Modal */} + {showModelSelect && ( + setShowModelSelect(false)} + onSelect={handleAddModel} + onDeselect={handleDeselectModel} + activeProviders={activeProviders} + modelAliases={modelAliases} + title="Add Model to Combo" + kindFilter={kindFilter} + addedModelValues={models} + closeOnSelect={false} + /> + )} + + ); } diff --git a/src/app/(dashboard)/dashboard/endpoint/EndpointPageClient.js b/src/app/(dashboard)/dashboard/endpoint/EndpointPageClient.js index b92f331e..39f715e5 100644 --- a/src/app/(dashboard)/dashboard/endpoint/EndpointPageClient.js +++ b/src/app/(dashboard)/dashboard/endpoint/EndpointPageClient.js @@ -261,11 +261,26 @@ export default function APIPageClient({ machineId }) { const fetchData = async () => { try { - const keysRes = await fetch("/api/keys"); - const keysData = await keysRes.json(); - if (keysRes.ok) { - setKeys(keysData.keys || []); + const fetchKeys = async () => { + const res = await fetch("/api/keys"); + if (!res.ok) return []; + const data = await res.json(); + return data.keys || []; + }; + + let existing = await fetchKeys(); + // Auto-provision a default key for first-time users so the endpoint works out of the box. + if (existing.length === 0) { + try { + const createRes = await fetch("/api/keys", { + method: "POST", + headers: { "Content-Type": "application/json" }, + body: JSON.stringify({ name: "Default Key" }), + }); + if (createRes.ok) existing = await fetchKeys(); + } catch { /* fall through to empty render */ } } + setKeys(existing); } catch (error) { console.log("Error fetching data:", error); } finally { @@ -1026,7 +1041,7 @@ export default function APIPageClient({ machineId }) {