Commit Graph

335 Commits

Author SHA1 Message Date
39101b3416 Merge origin/master (v0.5.91) into gitea/new_feature
Resolve conflicts:
- streamingHandler.js: adopt upstreamResponseHeaders while keeping 0-token detail row avoidance
- capabilities.js: preserve user-asserted caps and globalThis slots without local caching of catalogSource
- AddCustomModelModal.js & providers/[id]/page.js: wire STT transport marker with custom model edits/assertions
- models/custom/route.js & aliasRepo.js: persist custom model transport and invalidate user caps
- usageRepo.js: key byApiKey live stats by full API key and keep tail in maskApiKey
- UsageStats.js: lazy load charts dynamically
2026-09-28 21:17:33 +07:00
decolua
b54a3f9bb5 test(claude): update decloak tests for suffix-stripping fallback 2026-09-26 17:30:02 +07:00
decolua
6aea3875ef feat(claude): forward x-claude-code-session-id on OAuth requests 2026-09-26 17:15:12 +07:00
decolua
0249464d74 feat(cli-tools): support multiple model profiles for Codex CLI
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-26 17:03:06 +07:00
Nikan Wystaf
239bcfc568 perf(providers): make POST /api/providers O(1) and refuse silent key overwrite (#4350)
Fixes #4311

- Drop full-pool renumber on insert: new row gets MAX(priority)+1 directly,
  turning an O(pool) rewrite into O(1) per insert.
- For apikey connections, query by (provider, authType, name) and count via SQL
  aggregate instead of reading the entire pool into memory.
- Refuse silent apikey overwrite on name collision with 409 PROVIDER_NAME_CONFLICT,
  unless caller explicitly sets allowOverwrite: true.
- Add 8 unit tests covering priority ordering and name collisions.
2026-09-26 16:37:25 +07:00
Nick Nyanjui
199173fe58 feat(cline): expose the cline-free/* tier and price it at zero (#4334) 2026-09-26 16:31:25 +07:00
Beru
8f20daac6b fix(cli-tools): refresh Codex settings after apply (#4347) 2026-09-26 16:21:21 +07:00
Mohammad Hijjawi
fdcba3e1b2 fix(capabilities): stop caching the catalog source per module copy (#4351) 2026-09-26 16:20:59 +07:00
agent
dd293d3c62 feat(opencode-go): add the seven models upstream serves but the registry omits (#4357)
Upstream /zen/go/v1/models serves 35 ids; the registry listed 28. Add the seven
missing models with their corresponding supported and target formats:
- deepseek-v4.1-flash
- mimo-v2.6-flash, mimo-v2.6-pro
- space-bunny-free
- omen-alpha
- grok-4.7 (responses-only)
- gpt-6-luna (responses-only)
2026-09-26 16:15:43 +07:00
Nikan Wystaf
737b1f4d0a feat(providers): add Token Harbor provider 2026-09-26 16:09:52 +07:00
Spoon94
37a6b7e0f2 fix(dashboard): resolve combo limits with the server's capabilities (#4360)
The combos page computed combo capabilities in the browser using
aggregateComboCapabilities, falling back to pattern defaults for models
without exact entries because the synced model catalog is server-only.

Allow callers to supply a resolveCaps callback (e.g. from useModelCaps)
merged over the local tables, preserving non-limit capability flags.
2026-09-26 16:07:03 +07:00
Nikan Wystaf
06112c136c feat(providers): add four OpenAI-compatible aggregator providers (dahl, atria, agnes, bai)
Adds Dahl Inference, Atria Dawn, Agnes AI, and B.AI following the existing
registry pattern. Each uses DefaultExecutor with no custom translator needed.
2026-09-26 16:04:59 +07:00
Nick Nyanjui
fe347e4ea5 fix(stt): dispatch live-API-only Gemini models over the Live WebSocket transport (#4006)
Addresses #4006 by letting Gemini STT models that only exist on the Live
API transcribe instead of failing.

transcribeGemini sends audio to :generateContent and that is the only Gemini
path. A model that is realtime-only (exposed by the Live API's
bidiGenerateContent WebSocket) therefore fails outright, even though the account
can transcribe it.

open-sse/handlers/geminiLiveStt.js owns the WebSocket lifecycle: opens
:bidiGenerateContent, sends setup frame, waits for setupComplete, streams
audio as realtimeInput media chunks, and settles on turnComplete.
Dispatch is driven by transport marker 'gemini-live'. Adds custom model transport
persistence and selection on the dashboard.
2026-09-26 12:00:39 +07:00
ANIRUDDHA ADAK
273f0c32cd fix(responses): carry the streamed output items in response.completed (#4307) 2026-09-26 12:00:00 +07:00
Christian Gennari
c2148179c0 fix(commandcode): replay raw byte chunks to preserve all NDJSON lines 2026-09-26 11:53:54 +07:00
MrBeanDev
5d2cfbf3c5 fix(translator): stop emitting empty <think> markers into OpenAI content 2026-09-26 11:45:35 +07:00
decolua
975f28a57c test(zed): isolate zed-live-models suite DB via temp DATA_DIR
Seeding via createProviderConnection wrote zed-live-* accounts into the
real ~/.9router DB, showing up as junk accounts in the running dashboard.
Set DATA_DIR to a mkdtemp dir before dynamic-importing the DB-backed
modules and clean it up afterwards, matching the pattern in
compatible-provider-connections.test.js.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-26 11:40:08 +07:00
akmal safari pellu
30464bc227 fix(gemini): guard terminal model turns and unresponded functionCalls in normalizeGeminiContents 2026-09-26 11:18:44 +07:00
DavidArthurCole
dc198dff1f feat(claude): merge client anthropic-beta flags and forward rate-limit headers 2026-09-26 11:08:57 +07:00
29ccb84faf feat(commandcode): align quota usage with the official CLI /usage flow
Rewrites services/usage/commandcode.js to mirror the command-code CLI: whoami
resolves the org id, credits+subscriptions report the 5-hour/weekly windows and
plan, then usage/summary is queried with since=currentPeriodStart. Endpoints are
read from the provider registry usage block instead of hardcoded constants.

Tests updated to cover the new request order and response shapes.
2026-09-25 16:09:09 +07:00
272dbcb9cc Merge remote-tracking branch 'origin/master' into gitea/new_feature
# Conflicts:
#	open-sse/executors/qoder.js
#	open-sse/handlers/chatCore.js
#	open-sse/handlers/chatCore/sseToJsonHandler.js
#	open-sse/providers/registry/commandcode.js
#	src/app/(dashboard)/dashboard/combos/page.js
#	src/app/api/v1/models/route.js
#	src/lib/db/repos/usageRepo.js
#	src/shared/components/UsageStats.js
2026-09-25 15:14:28 +07:00
Deepanshu
a406381fad fix(usage): preserve API key usage attribution in live stats
The 24h/today branch of getUsageStats keyed byApiKey buckets on the
masked key, while the daily rollup and the lastUsed overlay key on the
full key. Every key minted by one instance shares the sk-{machineId}
prefix, so the mask collapsed all of them into a single bucket and
attributed one key's usage to another.

Key the live branch by the full api key too; the masked value is still
carried on the bucket for display.
2026-09-23 14:42:56 +07:00
Reid Nguyen
e571a8b6da feat(usage): show and redeem free limit resets for cc accounts
Adds the free usage-limit reset (the desktop app's "Reset for free")
to the Quota Tracker for cc OAuth accounts, mirroring the existing
Codex reset-credit button.

- Reset button with remaining count on the card; tooltip shows use-by
  date and which limits get refilled.
- Expiry modal (clock button) listing each grant: label, resets left,
  refills (session / weekly), use-by date, time remaining, status.
- Confirm dialog before redeeming, since a reset is irreversible.
- Usage and reset calls send the CLI User-Agent required for cedar_ember.
- Usage cache is dropped after a redeem so the card shows the refilled limits.
2026-09-23 12:21:07 +07:00
Ilhom (MBP M5 Pro)
8403576095 fix(claude): update spoofed cli version to 2.1.280 to support Opus 5.5
Anthropic's newly released Claude Opus 5.5 model strictly requires
claude-cli version 2.1.280 or newer. Spoofing the older 2.1.258
version results in an HTTP 400 `claude_code_version_too_old` error.

This commit updates the hardcoded `CLAUDE_CLI_VERSION` in
`open-sse/providers/shared.js` and aligns the corresponding unit
tests and baselines to bypass Anthropic's version gating.
2026-09-23 12:18:39 +07:00
mxlanparty@dp14
95600db17b feat(codex): add GPT-6 Sol and Luna support
- Add GPT-6 Sol and Luna to the Codex model registry.
- Send both models using the Codex 0.155 Responses Lite request shape, including `reasoning.context: "all_turns"`.
2026-09-23 12:11:49 +07:00
decolua
2aa99d9f39 feat(opencode-go): complete Go catalog (40 models) with auto-fetch + family endpoint regex
- Add 12 missing models from live /zen/go/v1/models (glm-5, kimi-k2.5,
  mimo v2/v2.6 line, qwen3.5-plus, grok-4.5/4.7, omen-alpha, ...)
- modelsFetcher + passthroughModels + "opencode-go" suggested-models filter
- Family regex fallback (opencodeFamilyFormats) keeps unknown/passthrough ids
  on the right endpoint lane (/responses, /messages, /chat/completions);
  curated registry entries always win
- isResponsesModel now routes passthrough grok/gpt ids to /responses

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-23 11:45:34 +07:00
wismyzhizi
910db749aa feat(xiaomi-mimo): server-assisted desktop login, five account clusters, v2.6 models
Reproduce the MiMo Desktop login surface server-side so headless/Docker
deployments can link a Xiaomi account without the Desktop client. The
account session (passToken) is captured during the proxied login and
stored per connection.

- Five account clusters (cn/sgp/ams/ru/in): per-region mimo-server host
  and SSO sid, unknown region falls back to sgp
- mimo-v2.6-pro/flash/pro-ultraspeed dual-route models: account-service
  route when desktop credentials exist, cloud API (sk- key) otherwise;
  drops obsolete mimo-x-*-preview ids
- Desktop ServiceTokenManager 2-phase handshake (single serviceLogin with
  target sid, raw 64-bit nonce preserved), per-region session cache
- reasoning_effort bridged to output_config.effort; i18n runtime now
  observes characterData mutations so React text rewrites get translated
- Security hardening on the login proxy: session travels only in the
  httpOnly cookie (never in the URL), proxy branch requires dashboard
  auth, authorization/proxy-authorization never forwarded upstream, and
  upstream Set-Cookie is not replayed onto the app origin
2026-09-23 09:43:54 +07:00
Welington
0f488c7027 fix(translator): map Claude "refusal" stop_reason to content_filter and surface its explanation
Anthropic's API-level refusal (streaming classifier / ToS) ends the stream
with stop_reason "refusal", stop_details carrying the reason, zero output
tokens and no content blocks. Map refusal to content_filter in both
directions, surface stop_details.explanation as message text, and add
CLAUDE_STOP.REFUSAL to schema.
2026-09-22 15:32:48 +07:00
Rafli Ahmad Zulfikar
d1de324586 perf(usage): bound lastUsed overlay scan to 2-day window; reach max thinking tier
getUsageStats("all") shipped the entire usageHistory table to JS just to
refine lastUsed (~2s on 290K rows, on every statsEmitter update per SSE
listener). Bound the overlay to a 2-day indexed range scan; older entries
keep day-level lastUsed from usageDaily aggregates. Totals unaffected.

budgetToLevel now maps budgets > 80384 (midpoint of 32768/128000) to
"max" instead of clamping to "xhigh", so the top reasoning tier is
reachable from large budget_tokens requests.
2026-09-22 15:08:38 +07:00
yiwen65
782c137b1f fix(qoder): prevent signed request replay and surface upstream errors 2026-09-22 15:01:06 +07:00
11089ab137 Merge origin/master (v0.5.81) into gitea/new_feature
Resolve conflicts:
- streamingHandler.js: merge buildStreamErrorBytes onAbortTerminal + local shouldPersistRequestDetail & streamStatusForContent
- capabilities.js: preserve server-injected user-asserted caps and models.dev catalog lookup; wire CommandCode /alpha/generate caps inside resolve()
- commandcode.js (services/usage): adopt upstream whoami + billing credits/subscriptions with 5h/weekly rate windows and plan caps
- openai-to-commandcode.js: merge toNativeImageBlock (data-URI & http(s) support) and assistant reasoning_content preservation
- commandcode-to-openai.js: adopt upstream mid-stream error throw for clean retry and abortion
- tests: sync commandcode test suite and exclude .next from vitest config
2026-09-22 10:08:58 +07:00
decolua
6886915f62 feat(capacity-adapter): default vision fallback to mimo-v2.6-flash-free
Register mimo-v2.6-flash-free on opencode-zen (chat lane) with a v2.6
capability pattern, and switch the vision adapter default from the old
mimo-v2.5-free.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 09:45:20 +07:00
82030615
da0046550a fix(responses): report usage on response.completed so clients can auto-compact
Map upstream Chat Completions usage to the Responses API shape and attach it to response.completed. Capture chunk.usage before the empty-choices guard so the usage-only trailer chunk survives, and defer completion to flushEvents() when usage is not yet known — only on the direct openai:openai-responses route, since a pivoted stream never reaches flushEvents. Fixes #3432.
2026-09-21 21:27:01 +07:00
Liang.Xu
402745dc1f feat(providers): add qoder-cn support for Qoder CN (qoder.com.cn) 2026-09-21 20:45:19 +07:00
Christian Gennari
be3bc764b1 fix(antigravity): separate weekly and short-window quotas and clean up redundant rows
- Track both weekly and 5-hour session buckets in parseWeeklyQuotaSummary,
  distinguishing sliding-window limits from multi-day weekly limits
- Preserve disabled session buckets at 0% rather than dropping them when weekly limits are reached
- Target 5-hour session rows (not weekly rows) during family exhaustion reconciliation in getAntigravityUsage
- Suppress synthesized per-model duplicate rows in dashboard normalization when family summaries are present
- Add unit test coverage for multi-bucket extraction, reconciliation isolation, and dashboard deduplication
2026-09-21 20:15:56 +07:00
dinhkarate
2daf25ffbe fix(qoder): handle code 110 billing blocks and preserve SSE error status
- Match code 110 (billing daily count exceeded) alongside 112/10605/pricingUrl
  in isBillingBlock, parsing JSON safely and accepting numeric/string codes
- Accept numeric strings for statusCodeValue and object bodies in envelope peek
- Emit structured 403 quota error chunk instead of synthetic assistant text
  when a billing envelope appears mid-stream
- Preserve upstream HTTP status in handleForcedSSEToJson when error chunk carries
  a valid 400-599 status
- Add unit tests for code-110 detection, mid-stream billing envelopes, and false-positive guard
2026-09-21 20:09:03 +07:00
Anantachoke
7c2b1fe3e1 fix(translator): drop replayed reasoning fields for Groq/Mistral/Cerebras (#4220)
Strict OpenAI-compatible validators reject unknown assistant-message
fields: Groq 400 ("property 'reasoning_content' is unsupported"),
Mistral 422 ("extra_forbidden"), Cerebras 400 ("wrong_api_format").

Clients driving reasoning models (Hermes Agent, and anything following
the DeepSeek/Kimi convention) echo the previous turn's reasoning_content
on every assistant message, so from the second turn on every request to
these providers fails and a fallback combo silently skips them.

Add a dropMessageFields rule to paramSupport.js that strips
reasoning_content / reasoning / reasoning_details from assistant turns
for groq, mistral, and cerebras.
2026-09-21 20:02:21 +07:00
Amir Seify
253199f16f feat(combos): Cursor/Claude Default presets + bulk select/delete/strategy
- Add Cursor Default / Claude Default on Dashboard -> Combos to generate
  unprefixed combo names that match Cursor/Claude client model IDs,
  seeded with cu/... or cc/... so those clients can route through 9Router.
- Add multi-select bulk Delete and bulk Set strategy (Fallback / Round Robin / Fusion).
- Docs and unit tests for preset builder.
2026-09-21 19:51:26 +07:00
Amir Seify
c933eefc27 fix(cursor): stop AgentService empty turns and silent tool hangs
Cursor-hosted models (cu/composer-2.5, cu/cursor-grok-*, cu/default) returned
HTTP 200 with an empty turn, or hung, whenever a client sent tools.

- Fold system prompts into the current user message. custom_system_prompt
  (RunRequest field 8) makes AgentService return an empty turn.
- Send ModelDetails (field 3); thinking variants (Composer, Grok, *-thinking)
  return an empty turn when only requested_model (field 9) is set.
- Route tool-call history and declared tool schemas through AgentService:
  encode OpenAI tools into mcp_tools (field 4), decode McpArgs and emit real
  tool_calls with finish_reason tool_calls.
- Map Composer  thinking / Grok thinking_delta (field 4) into visible content
  instead of dropping the answer with the unsigned reasoning.
- Ack request_context without echoing MCP tools (double-advertise stalls the
  HTTP/2 stream) and ack kv_server_message so the run proceeds.
- Reject IDE builtin execs instead of failing the turn, so the model can
  continue with MCP tools or a text answer.
- Add google.protobuf.Value / MCP encoders and a FIXED64 branch to
  encodeField in cursorProtobuf.js.

RTK now compresses the source-format body before translation for cursor only:
its translator rewrites role:tool into user XML, so the post-translate pass
missed those tool results. Every other provider keeps the post-translate pass
unchanged.
2026-09-21 19:49:48 +07:00
Minh Ha
5c217d34f3 feat(capabilities): model capability metadata on /v1/models, combo aggregation, pattern fixes
- Export aggregateComboCapabilities: union for vision/audio/search/pdf,
  intersection for tools, primary-model for reasoning fields, min
  contextWindow, max maxOutput
- Support nested combo resolution in aggregateComboCapabilities via
  comboLookup with depth guard (max 6)
- Wire capability metadata to all /v1/models entries and combos
- Show aggregated ctx/max metadata line and capability badges on combo chips
- Pattern fixes: MiMo v2.5/omni reasoning, qwen max/plus vision, minimax m2.x vision
- Sync commandcode model catalog and add openai gpt-5.5
- Add unit tests for capability patterns and combo capability aggregation
2026-09-21 19:41:29 +07:00
kimono381
cf663f5300 fix(huggingface): complete the Inference Providers router migration
Replace the retired api-inference.huggingface.co host with the Inference
Providers router (router.huggingface.co): imageConfig.modelMap resolves
Hub ids to provider-resolved ids, image-to-image models receive the
source image in inputs with the prompt under parameters.prompt, and a
new sttConfig wires the hf-inference ASR route. The image catalog grows
to 23 models, dead whisper-small is replaced by whisper-large-v3-turbo,
the unusable "language" param is dropped, and edit models declare the
edit capability so the dashboard offers a source image. Adds unit and
end-to-end coverage plus a model-id guard on custom endpoints.
2026-09-19 10:44:34 +07:00
Aaron
822aa958d1 fix(opencode): cloak Responses requests that already have tools
Free-tier Zen models reject Responses requests with 403 FreeTierError
when client tools are present but the fingerprint quartet is missing.
Apply the fingerprint tools to every OpenCode request, canonicalise
case variants of the quartet (Bash->bash) without duplication, and
restore the caller's original spellings on the response side via a
request-local WeakMap threaded through the existing toolNameMap.
2026-09-19 10:15:45 +07:00
Alexander Radchenko
49185137b8 feat(opencode-zen): add OpenCode Zen (PAYG) provider with free-tier fingerprint, alias ocz
Multi-endpoint provider on https://opencode.ai/zen/v1 (openai / claude /
openai-responses transports mirroring opencode-go) with 71 models across
paid + free tiers. Free-tier fingerprint: opencode/1.18.x UA spoof,
ses_-session header, built-in tool quartet, forced stream. Usage endpoint
/zen/v1/usage wired into the dashboard.
2026-09-19 10:02:28 +07:00
X-Adam
73e021b8a0 fix(ollama): map free-plan monthly window and derive reset from signup date 2026-09-19 09:51:13 +07:00
decolua
058ceace48 fix(opencode): declare forceStream on transport for free-tier SSE aggregation
Declares forceStream: true so chatCore properly converts upstream
forced-stream responses to JSON for non-streaming callers.

Co-authored-by: anojndr <anojndr@gmail.com>
Co-authored-by: TEGAR-SRC <tegararrahman17@gmail.com>
Co-authored-by: yxxrn <yxxrn@users.noreply.github.com>
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-18 17:32:52 +07:00
Louis Phạm
bc3be0cb28 fix(antigravity): scope cached thought signatures to the model family 2026-09-18 17:09:36 +07:00
Christian Gennari
092c84eac9 fix(commandcode): retry on transient stream error and avoid fake stop chunks 2026-09-18 17:09:24 +07:00
Louis Phạm
b3d6e089c6 fix(antigravity): strip Claude Code billing header from system prompts 2026-09-18 17:08:53 +07:00
Amirsalar Sojoudi
efc80ba2e3 fix(codex): route bare codex-auto-review to the Codex provider (#4135) 2026-09-18 17:05:47 +07:00
decolua
93837af09f fix(opencode): fix free tier 403 error and improve China region handling
- Force stream:true and cloak decoy tools (bash, read) for OpenCode free tier
- Support connection testing for opencode in testUtils
- Expand error message slice limits in auth and ping to preserve workspace link
- Add concise China region link chip in provider detail page

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-18 16:57:06 +07:00