Resolve conflicts:
- streamingHandler.js: adopt upstreamResponseHeaders while keeping 0-token detail row avoidance
- capabilities.js: preserve user-asserted caps and globalThis slots without local caching of catalogSource
- AddCustomModelModal.js & providers/[id]/page.js: wire STT transport marker with custom model edits/assertions
- models/custom/route.js & aliasRepo.js: persist custom model transport and invalidate user caps
- usageRepo.js: key byApiKey live stats by full API key and keep tail in maskApiKey
- UsageStats.js: lazy load charts dynamically
Fixes#4311
- Drop full-pool renumber on insert: new row gets MAX(priority)+1 directly,
turning an O(pool) rewrite into O(1) per insert.
- For apikey connections, query by (provider, authType, name) and count via SQL
aggregate instead of reading the entire pool into memory.
- Refuse silent apikey overwrite on name collision with 409 PROVIDER_NAME_CONFLICT,
unless caller explicitly sets allowOverwrite: true.
- Add 8 unit tests covering priority ordering and name collisions.
The combos page computed combo capabilities in the browser using
aggregateComboCapabilities, falling back to pattern defaults for models
without exact entries because the synced model catalog is server-only.
Allow callers to supply a resolveCaps callback (e.g. from useModelCaps)
merged over the local tables, preserving non-limit capability flags.
Adds Dahl Inference, Atria Dawn, Agnes AI, and B.AI following the existing
registry pattern. Each uses DefaultExecutor with no custom translator needed.
Addresses #4006 by letting Gemini STT models that only exist on the Live
API transcribe instead of failing.
transcribeGemini sends audio to :generateContent and that is the only Gemini
path. A model that is realtime-only (exposed by the Live API's
bidiGenerateContent WebSocket) therefore fails outright, even though the account
can transcribe it.
open-sse/handlers/geminiLiveStt.js owns the WebSocket lifecycle: opens
:bidiGenerateContent, sends setup frame, waits for setupComplete, streams
audio as realtimeInput media chunks, and settles on turnComplete.
Dispatch is driven by transport marker 'gemini-live'. Adds custom model transport
persistence and selection on the dashboard.
Seeding via createProviderConnection wrote zed-live-* accounts into the
real ~/.9router DB, showing up as junk accounts in the running dashboard.
Set DATA_DIR to a mkdtemp dir before dynamic-importing the DB-backed
modules and clean it up afterwards, matching the pattern in
compatible-provider-connections.test.js.
Co-Authored-By: Claude Code <noreply@anthropic.com>
Rewrites services/usage/commandcode.js to mirror the command-code CLI: whoami
resolves the org id, credits+subscriptions report the 5-hour/weekly windows and
plan, then usage/summary is queried with since=currentPeriodStart. Endpoints are
read from the provider registry usage block instead of hardcoded constants.
Tests updated to cover the new request order and response shapes.
The 24h/today branch of getUsageStats keyed byApiKey buckets on the
masked key, while the daily rollup and the lastUsed overlay key on the
full key. Every key minted by one instance shares the sk-{machineId}
prefix, so the mask collapsed all of them into a single bucket and
attributed one key's usage to another.
Key the live branch by the full api key too; the masked value is still
carried on the bucket for display.
Adds the free usage-limit reset (the desktop app's "Reset for free")
to the Quota Tracker for cc OAuth accounts, mirroring the existing
Codex reset-credit button.
- Reset button with remaining count on the card; tooltip shows use-by
date and which limits get refilled.
- Expiry modal (clock button) listing each grant: label, resets left,
refills (session / weekly), use-by date, time remaining, status.
- Confirm dialog before redeeming, since a reset is irreversible.
- Usage and reset calls send the CLI User-Agent required for cedar_ember.
- Usage cache is dropped after a redeem so the card shows the refilled limits.
Anthropic's newly released Claude Opus 5.5 model strictly requires
claude-cli version 2.1.280 or newer. Spoofing the older 2.1.258
version results in an HTTP 400 `claude_code_version_too_old` error.
This commit updates the hardcoded `CLAUDE_CLI_VERSION` in
`open-sse/providers/shared.js` and aligns the corresponding unit
tests and baselines to bypass Anthropic's version gating.
- Add GPT-6 Sol and Luna to the Codex model registry.
- Send both models using the Codex 0.155 Responses Lite request shape, including `reasoning.context: "all_turns"`.
Reproduce the MiMo Desktop login surface server-side so headless/Docker
deployments can link a Xiaomi account without the Desktop client. The
account session (passToken) is captured during the proxied login and
stored per connection.
- Five account clusters (cn/sgp/ams/ru/in): per-region mimo-server host
and SSO sid, unknown region falls back to sgp
- mimo-v2.6-pro/flash/pro-ultraspeed dual-route models: account-service
route when desktop credentials exist, cloud API (sk- key) otherwise;
drops obsolete mimo-x-*-preview ids
- Desktop ServiceTokenManager 2-phase handshake (single serviceLogin with
target sid, raw 64-bit nonce preserved), per-region session cache
- reasoning_effort bridged to output_config.effort; i18n runtime now
observes characterData mutations so React text rewrites get translated
- Security hardening on the login proxy: session travels only in the
httpOnly cookie (never in the URL), proxy branch requires dashboard
auth, authorization/proxy-authorization never forwarded upstream, and
upstream Set-Cookie is not replayed onto the app origin
Anthropic's API-level refusal (streaming classifier / ToS) ends the stream
with stop_reason "refusal", stop_details carrying the reason, zero output
tokens and no content blocks. Map refusal to content_filter in both
directions, surface stop_details.explanation as message text, and add
CLAUDE_STOP.REFUSAL to schema.
getUsageStats("all") shipped the entire usageHistory table to JS just to
refine lastUsed (~2s on 290K rows, on every statsEmitter update per SSE
listener). Bound the overlay to a 2-day indexed range scan; older entries
keep day-level lastUsed from usageDaily aggregates. Totals unaffected.
budgetToLevel now maps budgets > 80384 (midpoint of 32768/128000) to
"max" instead of clamping to "xhigh", so the top reasoning tier is
reachable from large budget_tokens requests.
Register mimo-v2.6-flash-free on opencode-zen (chat lane) with a v2.6
capability pattern, and switch the vision adapter default from the old
mimo-v2.5-free.
Co-Authored-By: Claude Code <noreply@anthropic.com>
Map upstream Chat Completions usage to the Responses API shape and attach it to response.completed. Capture chunk.usage before the empty-choices guard so the usage-only trailer chunk survives, and defer completion to flushEvents() when usage is not yet known — only on the direct openai:openai-responses route, since a pivoted stream never reaches flushEvents. Fixes#3432.
- Track both weekly and 5-hour session buckets in parseWeeklyQuotaSummary,
distinguishing sliding-window limits from multi-day weekly limits
- Preserve disabled session buckets at 0% rather than dropping them when weekly limits are reached
- Target 5-hour session rows (not weekly rows) during family exhaustion reconciliation in getAntigravityUsage
- Suppress synthesized per-model duplicate rows in dashboard normalization when family summaries are present
- Add unit test coverage for multi-bucket extraction, reconciliation isolation, and dashboard deduplication
- Match code 110 (billing daily count exceeded) alongside 112/10605/pricingUrl
in isBillingBlock, parsing JSON safely and accepting numeric/string codes
- Accept numeric strings for statusCodeValue and object bodies in envelope peek
- Emit structured 403 quota error chunk instead of synthetic assistant text
when a billing envelope appears mid-stream
- Preserve upstream HTTP status in handleForcedSSEToJson when error chunk carries
a valid 400-599 status
- Add unit tests for code-110 detection, mid-stream billing envelopes, and false-positive guard
Strict OpenAI-compatible validators reject unknown assistant-message
fields: Groq 400 ("property 'reasoning_content' is unsupported"),
Mistral 422 ("extra_forbidden"), Cerebras 400 ("wrong_api_format").
Clients driving reasoning models (Hermes Agent, and anything following
the DeepSeek/Kimi convention) echo the previous turn's reasoning_content
on every assistant message, so from the second turn on every request to
these providers fails and a fallback combo silently skips them.
Add a dropMessageFields rule to paramSupport.js that strips
reasoning_content / reasoning / reasoning_details from assistant turns
for groq, mistral, and cerebras.
- Add Cursor Default / Claude Default on Dashboard -> Combos to generate
unprefixed combo names that match Cursor/Claude client model IDs,
seeded with cu/... or cc/... so those clients can route through 9Router.
- Add multi-select bulk Delete and bulk Set strategy (Fallback / Round Robin / Fusion).
- Docs and unit tests for preset builder.
Cursor-hosted models (cu/composer-2.5, cu/cursor-grok-*, cu/default) returned
HTTP 200 with an empty turn, or hung, whenever a client sent tools.
- Fold system prompts into the current user message. custom_system_prompt
(RunRequest field 8) makes AgentService return an empty turn.
- Send ModelDetails (field 3); thinking variants (Composer, Grok, *-thinking)
return an empty turn when only requested_model (field 9) is set.
- Route tool-call history and declared tool schemas through AgentService:
encode OpenAI tools into mcp_tools (field 4), decode McpArgs and emit real
tool_calls with finish_reason tool_calls.
- Map Composer thinking / Grok thinking_delta (field 4) into visible content
instead of dropping the answer with the unsigned reasoning.
- Ack request_context without echoing MCP tools (double-advertise stalls the
HTTP/2 stream) and ack kv_server_message so the run proceeds.
- Reject IDE builtin execs instead of failing the turn, so the model can
continue with MCP tools or a text answer.
- Add google.protobuf.Value / MCP encoders and a FIXED64 branch to
encodeField in cursorProtobuf.js.
RTK now compresses the source-format body before translation for cursor only:
its translator rewrites role:tool into user XML, so the post-translate pass
missed those tool results. Every other provider keeps the post-translate pass
unchanged.
- Export aggregateComboCapabilities: union for vision/audio/search/pdf,
intersection for tools, primary-model for reasoning fields, min
contextWindow, max maxOutput
- Support nested combo resolution in aggregateComboCapabilities via
comboLookup with depth guard (max 6)
- Wire capability metadata to all /v1/models entries and combos
- Show aggregated ctx/max metadata line and capability badges on combo chips
- Pattern fixes: MiMo v2.5/omni reasoning, qwen max/plus vision, minimax m2.x vision
- Sync commandcode model catalog and add openai gpt-5.5
- Add unit tests for capability patterns and combo capability aggregation
Replace the retired api-inference.huggingface.co host with the Inference
Providers router (router.huggingface.co): imageConfig.modelMap resolves
Hub ids to provider-resolved ids, image-to-image models receive the
source image in inputs with the prompt under parameters.prompt, and a
new sttConfig wires the hf-inference ASR route. The image catalog grows
to 23 models, dead whisper-small is replaced by whisper-large-v3-turbo,
the unusable "language" param is dropped, and edit models declare the
edit capability so the dashboard offers a source image. Adds unit and
end-to-end coverage plus a model-id guard on custom endpoints.
Free-tier Zen models reject Responses requests with 403 FreeTierError
when client tools are present but the fingerprint quartet is missing.
Apply the fingerprint tools to every OpenCode request, canonicalise
case variants of the quartet (Bash->bash) without duplication, and
restore the caller's original spellings on the response side via a
request-local WeakMap threaded through the existing toolNameMap.
- Force stream:true and cloak decoy tools (bash, read) for OpenCode free tier
- Support connection testing for opencode in testUtils
- Expand error message slice limits in auth and ping to preserve workspace link
- Add concise China region link chip in provider detail page
Co-Authored-By: Claude Code <noreply@anthropic.com>