getUsageStats("all") shipped the entire usageHistory table to JS just to
refine lastUsed (~2s on 290K rows, on every statsEmitter update per SSE
listener). Bound the overlay to a 2-day indexed range scan; older entries
keep day-level lastUsed from usageDaily aggregates. Totals unaffected.
budgetToLevel now maps budgets > 80384 (midpoint of 32768/128000) to
"max" instead of clamping to "xhigh", so the top reasoning tier is
reachable from large budget_tokens requests.
- Add dedicated settings API routes for pi, omp, crush, forge, smelt, codewhale
- Integrate GenericCliToolCard with multi-model support for Pi and auto-discovery for OMP
- Register tools in cliTools catalog and all-statuses route
- Add official logos for all new CLI tools
Co-Authored-By: Claude Code <noreply@anthropic.com>
OpenRouter serves TypeSafe Jev at POST /api/v1/systemone with the same
request/response shape, so it plugs into systemoneConfig directly with
model typesafe/jev-1.13. Mark the System One media kind isNew and
render a New badge on the sidebar kind item and the Media Providers
accordion when any visible kind is new.
Co-Authored-By: Claude Code <noreply@anthropic.com>
The inline model test sent a chat-completions payload and failed with
500 on decision models. Add a systemone branch to pingModelByKind that
submits a native state+questions probe, and restore the test button
that was hidden for System One models.
Co-Authored-By: Claude Code <noreply@anthropic.com>
Allow customizing evaluation instructions in the System One example
card, and disable the inline model probe button for System One models
since decision models do not accept chat completion probes.
Co-Authored-By: Claude Code <noreply@anthropic.com>
Expose System One in the Media Providers sidebar accordion, set
kind to "systemone" on Jev models for ModelsCard filtering, wire
systemoneConfig into ProviderInfoCard, and configure GenericExampleCard
for interactive testing of decision models.
Co-Authored-By: Claude Code <noreply@anthropic.com>
New /v1/systemone pass-through route for Jev decision models (jev-1.13,
jev-1.13-free) on OpenCode Zen and the free lane. Follows the media-route
pattern: systemoneConfig in the registry drives URL/headers, the handler
mirrors the embeddings account-fallback + usage flow, and the dashboard
gains a System One media-provider kind. No chat-pipeline changes.
Co-Authored-By: Claude Code <noreply@anthropic.com>
Register mimo-v2.6-flash-free on opencode-zen (chat lane) with a v2.6
capability pattern, and switch the vision adapter default from the old
mimo-v2.5-free.
Co-Authored-By: Claude Code <noreply@anthropic.com>
Map upstream Chat Completions usage to the Responses API shape and attach it to response.completed. Capture chunk.usage before the empty-choices guard so the usage-only trailer chunk survives, and defer completion to flushEvents() when usage is not yet known — only on the direct openai:openai-responses route, since a pivoted stream never reaches flushEvents. Fixes#3432.
- Build linux/amd64 and linux/arm64 on native GitHub runners
- Assemble version manifests from platform digests and promote latest only after verification
- Add release/tag validation, manual republishing, timeouts, and health smoke tests
- Make Docker build mirrors configurable via build args and remove unnecessary runtime apk upgrades
- Update DOCKER.md documentation
- Track both weekly and 5-hour session buckets in parseWeeklyQuotaSummary,
distinguishing sliding-window limits from multi-day weekly limits
- Preserve disabled session buckets at 0% rather than dropping them when weekly limits are reached
- Target 5-hour session rows (not weekly rows) during family exhaustion reconciliation in getAntigravityUsage
- Suppress synthesized per-model duplicate rows in dashboard normalization when family summaries are present
- Add unit test coverage for multi-bucket extraction, reconciliation isolation, and dashboard deduplication
- Match code 110 (billing daily count exceeded) alongside 112/10605/pricingUrl
in isBillingBlock, parsing JSON safely and accepting numeric/string codes
- Accept numeric strings for statusCodeValue and object bodies in envelope peek
- Emit structured 403 quota error chunk instead of synthetic assistant text
when a billing envelope appears mid-stream
- Preserve upstream HTTP status in handleForcedSSEToJson when error chunk carries
a valid 400-599 status
- Add unit tests for code-110 detection, mid-stream billing envelopes, and false-positive guard
Strict OpenAI-compatible validators reject unknown assistant-message
fields: Groq 400 ("property 'reasoning_content' is unsupported"),
Mistral 422 ("extra_forbidden"), Cerebras 400 ("wrong_api_format").
Clients driving reasoning models (Hermes Agent, and anything following
the DeepSeek/Kimi convention) echo the previous turn's reasoning_content
on every assistant message, so from the second turn on every request to
these providers fails and a fallback combo silently skips them.
Add a dropMessageFields rule to paramSupport.js that strips
reasoning_content / reasoning / reasoning_details from assistant turns
for groq, mistral, and cerebras.
- Add Cursor Default / Claude Default on Dashboard -> Combos to generate
unprefixed combo names that match Cursor/Claude client model IDs,
seeded with cu/... or cc/... so those clients can route through 9Router.
- Add multi-select bulk Delete and bulk Set strategy (Fallback / Round Robin / Fusion).
- Docs and unit tests for preset builder.
Cursor-hosted models (cu/composer-2.5, cu/cursor-grok-*, cu/default) returned
HTTP 200 with an empty turn, or hung, whenever a client sent tools.
- Fold system prompts into the current user message. custom_system_prompt
(RunRequest field 8) makes AgentService return an empty turn.
- Send ModelDetails (field 3); thinking variants (Composer, Grok, *-thinking)
return an empty turn when only requested_model (field 9) is set.
- Route tool-call history and declared tool schemas through AgentService:
encode OpenAI tools into mcp_tools (field 4), decode McpArgs and emit real
tool_calls with finish_reason tool_calls.
- Map Composer thinking / Grok thinking_delta (field 4) into visible content
instead of dropping the answer with the unsigned reasoning.
- Ack request_context without echoing MCP tools (double-advertise stalls the
HTTP/2 stream) and ack kv_server_message so the run proceeds.
- Reject IDE builtin execs instead of failing the turn, so the model can
continue with MCP tools or a text answer.
- Add google.protobuf.Value / MCP encoders and a FIXED64 branch to
encodeField in cursorProtobuf.js.
RTK now compresses the source-format body before translation for cursor only:
its translator rewrites role:tool into user XML, so the post-translate pass
missed those tool results. Every other provider keeps the post-translate pass
unchanged.
- Export aggregateComboCapabilities: union for vision/audio/search/pdf,
intersection for tools, primary-model for reasoning fields, min
contextWindow, max maxOutput
- Support nested combo resolution in aggregateComboCapabilities via
comboLookup with depth guard (max 6)
- Wire capability metadata to all /v1/models entries and combos
- Show aggregated ctx/max metadata line and capability badges on combo chips
- Pattern fixes: MiMo v2.5/omni reasoning, qwen max/plus vision, minimax m2.x vision
- Sync commandcode model catalog and add openai gpt-5.5
- Add unit tests for capability patterns and combo capability aggregation
Replace the retired api-inference.huggingface.co host with the Inference
Providers router (router.huggingface.co): imageConfig.modelMap resolves
Hub ids to provider-resolved ids, image-to-image models receive the
source image in inputs with the prompt under parameters.prompt, and a
new sttConfig wires the hf-inference ASR route. The image catalog grows
to 23 models, dead whisper-small is replaced by whisper-large-v3-turbo,
the unusable "language" param is dropped, and edit models declare the
edit capability so the dashboard offers a source image. Adds unit and
end-to-end coverage plus a model-id guard on custom endpoints.
Free-tier Zen models reject Responses requests with 403 FreeTierError
when client tools are present but the fingerprint quartet is missing.
Apply the fingerprint tools to every OpenCode request, canonicalise
case variants of the quartet (Bash->bash) without duplication, and
restore the caller's original spellings on the response side via a
request-local WeakMap threaded through the existing toolNameMap.
## Features
- **Xiaomi MiMo**: merge MiMo Desktop support into `xiaomi-mimo` with dual auth (API key + Desktop/OAuth session), Preview models support, and encrypted-callback OAuth flow
- **Claude Code**: add 1M-context toggle (`[1m]` marker) and drive `CLAUDE_CODE_AUTO_COMPACT_WINDOW` directly from the dashboard
- **Models**: add DeepSeek-V4.1-Flash to DeepSeek provider, CodeBuddy-Intl, and Ollama (`deepseek-v4.1-flash:cloud`); enable `low`..`max` reasoning effort levels and vision capability for DeepSeek-V4.*
- **i18n**: integrate Persian (fa) translation
## Fixes
- **OpenCode / OpenCode Go**: resolve 403 `FreeTierError` and 429 rate limits with canonical session format, valid User-Agent, and stable upstream session reuse; force stream and declare `forceStream` for free-tier SSE aggregation; cloak decoy tools, normalize Muse Free tool choice, and strip prior reasoning items on Responses models; route Union Alpha via Messages API
- **Kiro**: preserve underscores in tool names (`mcp__server__tool`) and restore client tool names in responses; use neutral placeholder for tool-result-only turns; forward tool-result images
- **Stream**: report aborts after HTTP 200 in-band (per-format error frames) instead of closing silently
- **Command Code**: preserve images and `reasoning_effort` on `/alpha/generate`; retry transient stream errors and avoid fake stop chunks; add Quota Tracker support
- **Zed**: harden OAuth lifecycle (preserve `systemId`, renew proxy timeout), support live model resolution, and lower display priority in OAuth list
- **Antigravity**: scope cached thought signatures to model family; strip Claude Code billing headers from system prompts; sanitize Hermes system identity
- **Codex**: route bare `codex-auto-review` requests to the Codex provider (#4135)
- **Auth**: do not cool down an account for request-scoped 4xx errors
- **Usage**: improve DeepSeek credit balance display as currency credit instead of 0/total quota bar
- **Model Catalog**: scope synced catalog to gateways and declare vision capabilities for DeepSeek V4.1-Flash IDs
## Features
- **Xiaomi MiMo**: merge MiMo Desktop support into `xiaomi-mimo` with dual auth (API key + Desktop/OAuth session), Preview models support, and encrypted-callback OAuth flow
- **Claude Code**: add 1M-context toggle (`[1m]` marker) and drive `CLAUDE_CODE_AUTO_COMPACT_WINDOW` directly from the dashboard
- **Models**: add DeepSeek-V4.1-Flash to DeepSeek provider, CodeBuddy-Intl, and Ollama (`deepseek-v4.1-flash:cloud`); enable `low`..`max` reasoning effort levels and vision capability for DeepSeek-V4.*
- **i18n**: integrate Persian (fa) translation
## Fixes
- **OpenCode / OpenCode Go**: resolve 403 `FreeTierError` and 429 rate limits with canonical session format, valid User-Agent, and stable upstream session reuse; force stream and declare `forceStream` for free-tier SSE aggregation; cloak decoy tools, normalize Muse Free tool choice, and strip prior reasoning items on Responses models; route Union Alpha via Messages API
- **Kiro**: preserve underscores in tool names (`mcp__server__tool`) and restore client tool names in responses; use neutral placeholder for tool-result-only turns; forward tool-result images
- **Stream**: report aborts after HTTP 200 in-band (per-format error frames) instead of closing silently
- **Command Code**: preserve images and `reasoning_effort` on `/alpha/generate`; retry transient stream errors and avoid fake stop chunks; add Quota Tracker support
- **Zed**: harden OAuth lifecycle (preserve `systemId`, renew proxy timeout), support live model resolution, and lower display priority in OAuth list
- **Antigravity**: scope cached thought signatures to model family; strip Claude Code billing headers from system prompts; sanitize Hermes system identity
- **Codex**: route bare `codex-auto-review` requests to the Codex provider (#4135)
- **Auth**: do not cool down an account for request-scoped 4xx errors
- **Usage**: improve DeepSeek credit balance display as currency credit instead of 0/total quota bar
- **Model Catalog**: scope synced catalog to gateways and declare vision capabilities for DeepSeek V4.1-Flash IDs
- Force stream:true and cloak decoy tools (bash, read) for OpenCode free tier
- Support connection testing for opencode in testUtils
- Expand error message slice limits in auth and ping to preserve workspace link
- Add concise China region link chip in provider detail page
Co-Authored-By: Claude Code <noreply@anthropic.com>
- executors/zed.js: use exact wire values (anthropic, open_ai, google, x_ai)
and strip incompatible Vertex safetySettings on the Google path
- shared/zedAuth.js: robust callback query parsing, reject garbage PKCS#1 v1.5
decryptions, and thread proxyOptions when fetching LLM tokens
- oauth: preserve systemId across authorize/register/exchange lifecycle,
renew proxy idle timeout on reuse, and ignore non-callback localhost requests
- shared/OAuthModal.js: track owned proxy in flowRef and stop at most once
- api/providers/[id]/models: add connection-scoped live Zed model resolver
- registry: unhide provider in dashboard
- tests: add unit coverage for wire format, native auth, and live models