Resolve conflicts:
- streamingHandler.js: adopt upstreamResponseHeaders while keeping 0-token detail row avoidance
- capabilities.js: preserve user-asserted caps and globalThis slots without local caching of catalogSource
- AddCustomModelModal.js & providers/[id]/page.js: wire STT transport marker with custom model edits/assertions
- models/custom/route.js & aliasRepo.js: persist custom model transport and invalidate user caps
- usageRepo.js: key byApiKey live stats by full API key and keep tail in maskApiKey
- UsageStats.js: lazy load charts dynamically
The combos page computed combo capabilities in the browser using
aggregateComboCapabilities, falling back to pattern defaults for models
without exact entries because the synced model catalog is server-only.
Allow callers to supply a resolveCaps callback (e.g. from useModelCaps)
merged over the local tables, preserving non-limit capability flags.
Register mimo-v2.6-flash-free on opencode-zen (chat lane) with a v2.6
capability pattern, and switch the vision adapter default from the old
mimo-v2.5-free.
Co-Authored-By: Claude Code <noreply@anthropic.com>
- Export aggregateComboCapabilities: union for vision/audio/search/pdf,
intersection for tools, primary-model for reasoning fields, min
contextWindow, max maxOutput
- Support nested combo resolution in aggregateComboCapabilities via
comboLookup with depth guard (max 6)
- Wire capability metadata to all /v1/models entries and combos
- Show aggregated ctx/max metadata line and capability badges on combo chips
- Pattern fixes: MiMo v2.5/omni reasoning, qwen max/plus vision, minimax m2.x vision
- Sync commandcode model catalog and add openai gpt-5.5
- Add unit tests for capability patterns and combo capability aggregation
Route union-alpha to /zen/v1/messages with targetFormat claude, add anthropic-version header, and register model capabilities (vision, 262K context, 131K max output).
- Declare deepseek-v4.1-flash and deepseek-flash as vision-capable in MODEL_CAPABILITIES
- Share installed catalogSource across route chunks via globalThis.__9rCatalogSource
- Scope catalog modality keys by provider:model to prevent cross-gateway collisions
- Upgrade catalog format to v2 with automatic rebuild of older schemas
Resolve conflicts:
- package.json / cli/package.json: take 0.5.75
- .gitignore: union both sides (upstream 9router-*/temp files + local state dirs)
- CHANGELOG.md: keep both blocks, v0.5.75 above v0.5.70
- nonStreamingHandler.js: merge imports (unwrapClineEnvelope +
tokensForDetail/shouldPersistRequestDetail); drop dead appendRequestLog
- providers/[id]/page.js: union useState blocks (compatible-model states
+ importingClineModels)
Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
Command Code dropped vision and ignored client effort through the router:
image blocks became "[image omitted]", HTTP image URLs were never inlined,
and reasoning_effort landed on the envelope wrapper instead of params (so the
DeepSeek family mapping remapped low -> high). The catalog also treated
deepseek/deepseek-v4.1-flash as text-only, so the vision adapter stole those
requests to another provider.
- Map OpenAI image_url / Claude image blocks (base64 or data-URI) to the
native {type:"image", image:"data:...;base64,...", mimeType} generate block.
- Add FORMATS.COMMANDCODE to TARGETS_NEED_BASE64 so remote http(s) images are
inlined by the existing SSRF-safe fetcher before translation.
- Write reasoning_effort inside params for targetFormat commandcode and pass
low|medium|high|xhigh|max through unmapped; allow it in thinkingLevels.
- Provider-scoped capabilities for commandcode/cmc: vision except the CLI
text-only denylist, thinkingFormat commandcode, so family patterns
(deepseek-v4 -> thinkingFormat deepseek, vision false) no longer win.
- Quota Tracker: whoami + billing credits/subscriptions (credits vs plan cap,
5h and weekly windows), labels from AI_PROVIDERS[].name.
codebuddy-intl: deepseek-v4-flash replaced by deepseek-v4.1-flash (same
gateway catalog as CN) and a capability override so the model keeps the
openai-style reasoning_effort path instead of the vendor-native "deepseek"
thinking shape the gateway rejects. Thinking levels low/high/xhigh.
ollama: add deepseek-v4.1-flash:cloud (verified on ollama.com/api/tags) with
vision + 1M context caps.
Co-Authored-By: Claude Code <noreply@anthropic.com>
- Coalesce Qoder's empty finish-in-delta frame with the later choices:[] usage
frame so OpenAI and Claude clients receive prompt_tokens, completion_tokens
and cache-hit tokens (the dashboard already saw them)
- Upload inlined images through /api/v2/image/upload like qodercli, and stub
oversized non-image files instead of stuffing 30MB+ data URIs into
agent_chat_generation
- Emit response.completed -> response.usage for chat-native upstreams so
/v1/responses clients (Codex CLI, sub2api) no longer log 0/0/0
- Keep Claude message_delta.usage working when usage arrives without choices[0]
- Escalate to the smallest advertised Qoder context tier (200K/400K/1M) when
the estimated prompt no longer fits max_input_tokens
- Pass apiKey for PAT connections and list hidden enable:false catalog keys
from /v1/models
The server's product-config payload (which the IDE plugin fetches from
copilot.tencent.com) publishes deepseek-v4.1-flash and no longer lists
deepseek-v4-flash, so the old id is dropped — same pattern as the previous
catalog refreshes (#3648, #3802). The v4-flash endpoint still answers 200,
but the published list is the contract.
Per the server table, maxOutput rises 50000 -> 128000 while contextWindow
stays 1000000.
- registry/codebuddy-cn.js: models[] entry swapped to the new id
- capabilities.js: per-model entry swapped, maxOutput -> 128000
No changes needed in thinkingLevels.js (the deepseek-v4* pattern already
matches the new id and publishes low/high/xhigh, matching the server's
supportedEfforts or pricing.js (the deepseek-v* glob yields the same rates).
EOF
)
## Features
- **Providers**: creating a compatible / custom-embedding node now registers the endpoint only — the API key is added afterwards from the node's page, like the built-in providers. The create dialogs drop the API Key / Model ID / Check fields, and `POST /api/provider-nodes` no longer accepts credentials at all, so a node can never be half-created
- **Providers**: compatible nodes now use the same model rows as built-in providers — capability badges, copy, per-model test, alias handling and the Add/Edit Model modal with vision + reasoning toggles, replacing the weaker read-only list
- **Providers**: compatible nodes get the built-in bulk toolbars: Test All / Disable / Active / Select All over connections, and Test All Models / Disable All / Active All over models, with per-row disable and a restore strip for disabled models
- **Providers**: wire the dead "Fetch Models" button on compatible nodes to the live upstream catalog, de-duplicating against already-added models
## Fixes
- **Models**: persist per-model capability assertions for custom and compatible providers and honor them everywhere — unsupported media is stripped on the chat path, `/v1/models` and `/api/models` report what the user asserted, and thinking translation follows it (asserting `reasoning:false` now actually strips thinking fields, `reasoning:true` emits them)
- **Models**: partial capability edits merge instead of overwriting, so toggling vision off no longer erases a stored reasoning assertion
- **Capabilities**: keep server-injected readers (synced catalog, user-asserted capabilities) in process-wide state — Next.js compiles startup and each API route into separate bundles with their own module instances, so a boot-time install was invisible to every request handler and the models.dev catalog contributed nothing to upstream requests since 0532f00d
- **Dashboard**: thinking-level picker and model-row suffix reflect user-asserted reasoning on compatible nodes
- **Providers**: `Default Model` is optional when adding an API key to a compatible node — the node's own model list (and the picker in the test modals) already determine what gets probed, and the built-in fallback still covers connection checks
- **Providers**: restore the `useCopyToClipboard` import dropped from the provider detail page, which crashed the route with `ReferenceError` for every provider
- **DB**: restore `getModelAliases` / `setModelAlias` / `deleteModelAlias` re-exports dropped from the `localDb` shim by 86112cee, which broke `GET /api/models` and `GET /v1/models` at import time
- **Providers**: remove dead `PassthroughModelsSection` (never passed props, superseded by the shared model rows)
- **Media Providers**: creating a custom embedding node reports that a key still has to be added, instead of claiming a key was saved; the edit dialog keeps its API Key + Check affordance since a stored key already exists there
- **Build**: self-host Inter instead of fetching it through `next/font/google` at build time — a Docker / mirrored builder with no route to `fonts.googleapis.com` failed the entire image build on `Failed to fetch 'Inter' from Google Fonts`. The seven `@font-face` rules and their `unicode-range`s copy what `next/font` emitted (a `latin`-only file would have dropped Vietnamese diacritics) and the latin subset is preloaded as before, so rendered metrics are unchanged
Audit of every branch-owned line the -X theirs merge dropped from
the 32 pre-merge commits found three more real regressions:
* src/sse/handlers/chat.js: merge kept the capsOverride feature
(bb8d67ba) but reverted the import block, so getCustomModels and
capabilitiesFromServiceKind were undefined. The runtime error was
swallowed by the feature's own fail-open try/catch — custom
models silently lost their vision override. Restored both imports.
* open-sse/providers/capabilities.js: TRUST_UPSTREAM_VISION (the
floor that keeps vision on for unknown models on upstream-validating
gateways like openrouter) was left as dead code by the merge —
upstream rewrote step 4 as refine() and dropped the check.
Re-applied it on top of the new refine() so catalog/limits
refinement still applies.
* tests/unit/chat-connection-pin.test.js: mock auth module lacked
isModelAllowedForKey added by f0adfb20.
* .gitignore: re-add .pi-subagents/.
- Sync codebuddy-cn catalog and capabilities with copilot.tencent.com server payload
- Fix thinkingCanDisable semantics for glm-5.3 and deepseek-v4 models
- Add missing glm-5.2 thinking levels to thinkingLevels.js
- Add glm-5-turbo model to glm and glm-cn registries
- Registry/constants: drop qmodel_preview/gm51model, add lite,
qmodel_38max (Qwen3.8-Max), qfmodel (Qwen3.8-Flash), gmodel (GLM-5.3),
gfmodel (GLM-5.3-Flash)
- capabilities: add PROVIDER_CAPABILITIES['qoder'] so opaque internal
ids resolve to their real models' context windows and limits
- executor: preserve image blocks instead of flattening away, convert
Claude-style image blocks, and hash images into chat_record_id
- tests: cover image preservation, data-URI and Claude-block conversion
- build(docker): use CN mirrors for apk and npm
Route all Muse Spark models (not just 1.2) on OpenCode Free to
/zen/v1/responses via isMuseSparkModel(), fixing HTTP 500 on
muse-spark-1.3-contributor-free. Declare vision:true on Muse Spark
models so image input is no longer stripped; register 1.3 in the
registry and capabilities. Scoped to opencode only — other providers
keep Chat Completions routing.
- add claude-fable-5-1 to the Claude Code model catalog (1M context,
permanent adaptive thinking)
- centralize the spoofed Claude Code version and update both request
and billing identities to 2.1.257 (Fable 5.1 rejects < 2.1.251)
- send output_config.effort without the redundant thinking switch for
permanently adaptive models
- add regression coverage for capabilities, headers, billing identity
and adaptive-effort payload
# Conflicts:
# open-sse/providers/registry/claude.js
# open-sse/providers/shared.js
# open-sse/utils/claudeCloaking.js
# tests/__baseline__/providers-baseline.json
Z.ai / GLM-5.2+ require a top-level reasoning_effort (low/high/max)
alongside thinking:{type:"enabled"} to control reasoning depth; the zai
branch previously only set thinking and dropped reasoning_effort, so every
GLM-5.x request ran at the model default (max). Gate the field behind
GLM-5.2+ (thinkingEffortSupported in capabilities.js) since older GLM
(4.x, 5.0, 5.1, 5-turbo, 5v-turbo) do not read it, and map client levels
to the exact low/high/max values z.ai accepts.
extractThinking now checks reasoning_effort/reasoning.effort before the
thinking object so a client-supplied effort is not overwritten by
thinking:{type:"enabled"} mapping to mode:auto.
Fixes#2721
muse-spark-1.2-contributor-free returned HTTP 500 on /zen/v1/chat/completions.
The model is only served by /zen/v1/responses, so route it there via a per-model
targetFormat and normalize the Chat fields the Responses API rejects
(max_tokens -> max_output_tokens, reasoning_effort -> reasoning{effort,summary}),
clamping max/ultra down to the highest effort the model accepts (xhigh).
Routing stays per-model: the other free models (big-pickle, hy3-free, mimo,
nemotron, laguna) are not served by /responses and keep /chat/completions.
Capability tables are hand-maintained, so a model gains vision or a wider
context only when someone notices and edits the file. This adds a daily
sync that fills the gap for models already in the registry.
How it decides:
- Modalities (vision/pdf/audio/video) belong to the MODEL — every gateway
serving glm-5.3-flash serves the same weights — so they are keyed by
model id and shared. A majority of sources must declare one, which keeps
out lone mis-declarations: minimax-m2.5 (1 of 45), glm-4.7 (1 of 44) and
gpt-oss-120b (2 of 76) are text-only despite a reseller claiming vision.
- Context/output limits belong to the GATEWAY — each truncates differently
(glm-5 ships as 202752/16384 on one host and 204800/131072 on another) —
so they are keyed by provider + model and only the matching provider's
own numbers are trusted.
Both layers are strictly additive and sit BELOW the hand-written tables,
which short-circuit first. A capability already true stays true.
Mechanics: worker thread (the 4MB parse would block the loop ~20ms),
ETag so an unchanged catalog costs one empty request, 60s startup delay,
30min backoff on failure, MODEL_CATALOG_SYNC=off to disable. Only the
~57KB delta is kept; lookups cost ~0.1us via an mtime-guarded cache.
capabilities.js is bundled into the browser through useModelCaps, so it
cannot import node:fs — the server injects the reader via
setCatalogSource() from instrumentation.
visionPatterns.js is the last resort: a model nobody has catalogued yet
still accepts images when its id says so (qwen3-vl-plus, glm-4.6v, llava),
with image-generation and embedding ids excluded.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Vendors shipped four multimodal models the registry did not carry:
- glm-5.3-flash — z.ai's first natively multimodal GLM-5, 1M context,
image + video + pdf input (glm, glm-cn, opencode-go)
- deepseek-v4-flash-vision-exp — image input at V4-Flash text parity,
1M context / 384k output (deepseek, opencode-go)
- grok-4.6, grok-4.5 — 500k context; 4.6 has no text output limit (xai)
Capabilities needed hand entries because the existing globs mis-matched:
*glm-5* and *deepseek-v4* carry no vision, and *grok-4* would have capped
grok-4.6 at 256k instead of 500k. The grok-4.6 pattern sits above the
generic *grok-4* so it wins the first-match lookup.
Also corrects glm-4.6v / glm-4.5v, which were missing video input and
declared no maxOutput, and backfills glm-4.6v on glm-cn — zhipuai serves
it and the sibling provider already listed it.
tests/unit/opencode-go-models.test.js pins the opencode-go model list, so
its expected array moves with the registry.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Add gemini-3.7-flash and its tiered high/medium/low variants to the
Antigravity and Gemini registries, with matching capabilities, pricing
and Antigravity quota tracking.
extractModel now recognises gemini-3.7-flash-tiered alongside 3.6 and
derives the version from the request, so thinkingLevel still maps to the
right tiered alias.
Closes#3286Closes#3281
- Enable vision + audioInput capacity-adapter pools by default for new
and existing users (mergeWithDefaults backward-compat)
- Fall back to oc/mimo-v2.5-free when an enabled pool has no models
configured, both in the backend resolver and the combos UI (auto
refill on removing the last model from a pool)
- Hide PDF/Video from the Vision Adapter UI (PDF never implemented,
Video lacks translator support) while keeping the settings shape
- Exclude combos from the model picker when opened from the Vision
Adapter section
- mimo-v2.5 registry entry now declares audioInput/videoInput
- Simplify combo strategy and Vision Adapter descriptions
Adds Poolside (inference.poolside.ai) as an API-key provider using the default OpenAI transport. Registers three Laguna models with reasoning capabilities (262K context, 32K max output).
Add GPT-5.6 Sol/Terra/Luna and their synthetic thinking/agentic/
thinking-agentic variants to the Kiro static catalog with the observed
272k context window and credit multipliers (2.4/1.2/0.6), register MITM
mapping slots for the new base ids, and override runtime capabilities so
the GPT-5.6 family reports the 272k window instead of the generic GPT-5
profile.
- Updated capabilities for NVIDIA models to enforce OpenAI-compatible reasoning formats.
- Added new models: MiniMax M3, GLM 5.2, DeepSeek V4 Pro, DeepSeek V4 Flash, Kimi K2.6, and Nemotron 3 Ultra to the NVIDIA registry.
This enhances the provider's functionality and aligns with OpenAI standards.
Add qwen omni (audio/video input), qwen3.5/3.6/3.7 (native vision/video),
and mark qwen coder & max as text-only reasoning models.
Co-authored-by: Cursor <cursoragent@cursor.com>
Registry exposes the dashed id claude-opus-4-7; matchPattern treats "."
as a literal, so it missed the dotted pattern and fell through to the
generic claude opus entry (200k / claude-budget). Add an exact entry so
it resolves to 1M context + adaptive thinking, plus a unit test covering
the dashed Opus ids.
Co-authored-by: Cursor <cursoragent@cursor.com>
Add 1M context capability overrides for claude-opus-4.8 and -thinking
variants, and use the resolved capability contextWindow (fallback 200k)
instead of the hardcoded 200k estimate in the Kiro executor.
Co-authored-by: Cursor <cursoragent@cursor.com>
The pattern matcher marked *minimax-m3* as vision: false, causing
9Router to strip image attachments before forwarding upstream. This
broke Claude Code / Cursor / Cline vision flows when routing through
MiniMax-M3.
Scoped vision: true to *minimax-m3* only. M2.7 and the catch-all
*minimax* pattern remain vision: false: those models are text-only
(per MiniMax docs / NVIDIA NIM model card), so forcing vision there
would send images to a model that errors instead of degrading.
Co-authored-by: Cursor <cursoragent@cursor.com>