Commit Graph

781 Commits

Author SHA1 Message Date
39101b3416 Merge origin/master (v0.5.91) into gitea/new_feature
Resolve conflicts:
- streamingHandler.js: adopt upstreamResponseHeaders while keeping 0-token detail row avoidance
- capabilities.js: preserve user-asserted caps and globalThis slots without local caching of catalogSource
- AddCustomModelModal.js & providers/[id]/page.js: wire STT transport marker with custom model edits/assertions
- models/custom/route.js & aliasRepo.js: persist custom model transport and invalidate user caps
- usageRepo.js: key byApiKey live stats by full API key and keep tail in maskApiKey
- UsageStats.js: lazy load charts dynamically
2026-09-28 21:17:33 +07:00
Gesya Gayatree Solih
b65d2d0a6a fix(claude): decloak tool names when toolNameMap misses (#4342)
- Add stripCloakSuffix fallback in decloakToolNames and decloakStreamChunk
- Prevent client errors when toolNameMap is missing or lost on retry
2026-09-26 17:28:51 +07:00
decolua
6aea3875ef feat(claude): forward x-claude-code-session-id on OAuth requests 2026-09-26 17:15:12 +07:00
Nick Nyanjui
199173fe58 feat(cline): expose the cline-free/* tier and price it at zero (#4334) 2026-09-26 16:31:25 +07:00
Mohammad Hijjawi
fdcba3e1b2 fix(capabilities): stop caching the catalog source per module copy (#4351) 2026-09-26 16:20:59 +07:00
agent
dd293d3c62 feat(opencode-go): add the seven models upstream serves but the registry omits (#4357)
Upstream /zen/go/v1/models serves 35 ids; the registry listed 28. Add the seven
missing models with their corresponding supported and target formats:
- deepseek-v4.1-flash
- mimo-v2.6-flash, mimo-v2.6-pro
- space-bunny-free
- omen-alpha
- grok-4.7 (responses-only)
- gpt-6-luna (responses-only)
2026-09-26 16:15:43 +07:00
Nikan Wystaf
737b1f4d0a feat(providers): add Token Harbor provider 2026-09-26 16:09:52 +07:00
Spoon94
37a6b7e0f2 fix(dashboard): resolve combo limits with the server's capabilities (#4360)
The combos page computed combo capabilities in the browser using
aggregateComboCapabilities, falling back to pattern defaults for models
without exact entries because the synced model catalog is server-only.

Allow callers to supply a resolveCaps callback (e.g. from useModelCaps)
merged over the local tables, preserving non-limit capability flags.
2026-09-26 16:07:03 +07:00
Nikan Wystaf
06112c136c feat(providers): add four OpenAI-compatible aggregator providers (dahl, atria, agnes, bai)
Adds Dahl Inference, Atria Dawn, Agnes AI, and B.AI following the existing
registry pattern. Each uses DefaultExecutor with no custom translator needed.
2026-09-26 16:04:59 +07:00
MrBeanDev
90b0693423 feat(thinking): return Claude thinking text to OpenAI-format clients 2026-09-26 12:02:07 +07:00
Nick Nyanjui
fe347e4ea5 fix(stt): dispatch live-API-only Gemini models over the Live WebSocket transport (#4006)
Addresses #4006 by letting Gemini STT models that only exist on the Live
API transcribe instead of failing.

transcribeGemini sends audio to :generateContent and that is the only Gemini
path. A model that is realtime-only (exposed by the Live API's
bidiGenerateContent WebSocket) therefore fails outright, even though the account
can transcribe it.

open-sse/handlers/geminiLiveStt.js owns the WebSocket lifecycle: opens
:bidiGenerateContent, sends setup frame, waits for setupComplete, streams
audio as realtimeInput media chunks, and settles on turnComplete.
Dispatch is driven by transport marker 'gemini-live'. Adds custom model transport
persistence and selection on the dashboard.
2026-09-26 12:00:39 +07:00
ANIRUDDHA ADAK
273f0c32cd fix(responses): carry the streamed output items in response.completed (#4307) 2026-09-26 12:00:00 +07:00
Christian Gennari
c2148179c0 fix(commandcode): replay raw byte chunks to preserve all NDJSON lines 2026-09-26 11:53:54 +07:00
MrBeanDev
5d2cfbf3c5 fix(translator): stop emitting empty <think> markers into OpenAI content 2026-09-26 11:45:35 +07:00
akmal safari pellu
30464bc227 fix(gemini): guard terminal model turns and unresponded functionCalls in normalizeGeminiContents 2026-09-26 11:18:44 +07:00
DavidArthurCole
dc198dff1f feat(claude): merge client anthropic-beta flags and forward rate-limit headers 2026-09-26 11:08:57 +07:00
29ccb84faf feat(commandcode): align quota usage with the official CLI /usage flow
Rewrites services/usage/commandcode.js to mirror the command-code CLI: whoami
resolves the org id, credits+subscriptions report the 5-hour/weekly windows and
plan, then usage/summary is queried with since=currentPeriodStart. Endpoints are
read from the provider registry usage block instead of hardcoded constants.

Tests updated to cover the new request order and response shapes.
2026-09-25 16:09:09 +07:00
3481058e09 fix(usage): drop duplicated commandcode import and features block
The previous origin/master merges (11089ab1 and earlier) left two
auto-merged artifacts behind because both sides added the same lines in
different places:

- services/usage.js declared `getCommandCodeUsage` twice (import + handler),
  which is an ESM SyntaxError on a clean checkout. The working tree masked
  it, so `next build` passed locally while the committed tree did not parse.
- providers/registry/commandcode.js carried two `features` blocks; the
  second one shadowed the first for `usage`/`usageApikey`.

Verified with `node --check` on the committed blob.
2026-09-25 16:09:04 +07:00
272dbcb9cc Merge remote-tracking branch 'origin/master' into gitea/new_feature
# Conflicts:
#	open-sse/executors/qoder.js
#	open-sse/handlers/chatCore.js
#	open-sse/handlers/chatCore/sseToJsonHandler.js
#	open-sse/providers/registry/commandcode.js
#	src/app/(dashboard)/dashboard/combos/page.js
#	src/app/api/v1/models/route.js
#	src/lib/db/repos/usageRepo.js
#	src/shared/components/UsageStats.js
2026-09-25 15:14:28 +07:00
Reid Nguyen
e571a8b6da feat(usage): show and redeem free limit resets for cc accounts
Adds the free usage-limit reset (the desktop app's "Reset for free")
to the Quota Tracker for cc OAuth accounts, mirroring the existing
Codex reset-credit button.

- Reset button with remaining count on the card; tooltip shows use-by
  date and which limits get refilled.
- Expiry modal (clock button) listing each grant: label, resets left,
  refills (session / weekly), use-by date, time remaining, status.
- Confirm dialog before redeeming, since a reset is irreversible.
- Usage and reset calls send the CLI User-Agent required for cedar_ember.
- Usage cache is dropped after a redeem so the card shows the refilled limits.
2026-09-23 12:21:07 +07:00
mxlanparty@dp14
95600db17b feat(codex): add GPT-6 Sol and Luna support
- Add GPT-6 Sol and Luna to the Codex model registry.
- Send both models using the Codex 0.155 Responses Lite request shape, including `reasoning.context: "all_turns"`.
2026-09-23 12:11:49 +07:00
decolua
a61fc6a01e fix(antigravity): rewrite all Hermes identity variants, not just the legacy sentence
The old rule only matched 'You are Hermes Agent, an intelligent AI
assistant created by Nous Research.' Current Hermes versions use
'You are Hermes Agent, built by Nous Research.' and similar variants,
which slipped through and got a fake 429 RESOURCE_EXHAUSTED. Replace
every variant with a neutral 'You are an AI assistant.' identity.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-23 11:46:05 +07:00
decolua
1b72f02e3b fix(opencode-go): clamp deepseek reasoning_effort "max" to "high" for mimo backends that reject it
mimo-v2.5-pro/v2.6 on the Go lane return 400 on reasoning_effort "max"
(probed live; mimo-v2.5 accepts it). The deepseek applyFormat case now honors
the declared thinking levels, and mimo-v2.5-pro gets a levels entry.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-23 11:45:40 +07:00
decolua
2aa99d9f39 feat(opencode-go): complete Go catalog (40 models) with auto-fetch + family endpoint regex
- Add 12 missing models from live /zen/go/v1/models (glm-5, kimi-k2.5,
  mimo v2/v2.6 line, qwen3.5-plus, grok-4.5/4.7, omen-alpha, ...)
- modelsFetcher + passthroughModels + "opencode-go" suggested-models filter
- Family regex fallback (opencodeFamilyFormats) keeps unknown/passthrough ids
  on the right endpoint lane (/responses, /messages, /chat/completions);
  curated registry entries always win
- isResponsesModel now routes passthrough grok/gpt ids to /responses

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-23 11:45:34 +07:00
wismyzhizi
910db749aa feat(xiaomi-mimo): server-assisted desktop login, five account clusters, v2.6 models
Reproduce the MiMo Desktop login surface server-side so headless/Docker
deployments can link a Xiaomi account without the Desktop client. The
account session (passToken) is captured during the proxied login and
stored per connection.

- Five account clusters (cn/sgp/ams/ru/in): per-region mimo-server host
  and SSO sid, unknown region falls back to sgp
- mimo-v2.6-pro/flash/pro-ultraspeed dual-route models: account-service
  route when desktop credentials exist, cloud API (sk- key) otherwise;
  drops obsolete mimo-x-*-preview ids
- Desktop ServiceTokenManager 2-phase handshake (single serviceLogin with
  target sid, raw 64-bit nonce preserved), per-region session cache
- reasoning_effort bridged to output_config.effort; i18n runtime now
  observes characterData mutations so React text rewrites get translated
- Security hardening on the login proxy: session travels only in the
  httpOnly cookie (never in the URL), proxy branch requires dashboard
  auth, authorization/proxy-authorization never forwarded upstream, and
  upstream Set-Cookie is not replayed onto the app origin
2026-09-23 09:43:54 +07:00
savioruz
6af26a9ee8 fix(proxy-pools): lossless header forwarding for relay pools
Preserve request headers through the Vercel/Cloudflare/Deno relay.

- proxyFetch: normalize options.headers via Object.fromEntries before
  spreading into the relay headers. Spreading a Headers instance yields
  {} and silently dropped every entry (auth + content-type), which is why
  the same pool worked on one path and failed on another.
- vercel relay: build the forwarded header object from req.headers.entries()
  instead of new Headers(req.headers), avoiding edge-runtime normalization
  of casing/duplicate keys that some providers reject.
2026-09-23 09:12:00 +07:00
Reid Nguyen
cbffeb9770 feat(claude): support Claude Opus 5.5 2026-09-23 09:01:51 +07:00
Welington
0f488c7027 fix(translator): map Claude "refusal" stop_reason to content_filter and surface its explanation
Anthropic's API-level refusal (streaming classifier / ToS) ends the stream
with stop_reason "refusal", stop_details carrying the reason, zero output
tokens and no content blocks. Map refusal to content_filter in both
directions, surface stop_details.explanation as message text, and add
CLAUDE_STOP.REFUSAL to schema.
2026-09-22 15:32:48 +07:00
decolua
1a02713150 feat(combos): hide preset buttons and migrate legacy mimo vision adapter
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 15:32:10 +07:00
Rafli Ahmad Zulfikar
d1de324586 perf(usage): bound lastUsed overlay scan to 2-day window; reach max thinking tier
getUsageStats("all") shipped the entire usageHistory table to JS just to
refine lastUsed (~2s on 290K rows, on every statsEmitter update per SSE
listener). Bound the overlay to a 2-day indexed range scan; older entries
keep day-level lastUsed from usageDaily aggregates. Totals unaffected.

budgetToLevel now maps budgets > 80384 (midpoint of 32768/128000) to
"max" instead of clamping to "xhigh", so the top reasoning tier is
reachable from large budget_tokens requests.
2026-09-22 15:08:38 +07:00
BFLabsAI
5798b30841 fix(antigravity): drop requestType "agent" to avoid false 429 RESOURCE_EXHAUSTED 2026-09-22 15:04:16 +07:00
yiwen65
782c137b1f fix(qoder): prevent signed request replay and surface upstream errors 2026-09-22 15:01:06 +07:00
decolua
f84c667d42 feat(providers): add OpenRouter System One lane and New badge
OpenRouter serves TypeSafe Jev at POST /api/v1/systemone with the same
request/response shape, so it plugs into systemoneConfig directly with
model typesafe/jev-1.13. Mark the System One media kind isNew and
render a New badge on the sidebar kind item and the Media Providers
accordion when any visible kind is new.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 14:42:02 +07:00
decolua
6431e35303 feat(dashboard): add System One to sidebar and media provider detail page
Expose System One in the Media Providers sidebar accordion, set
kind to "systemone" on Jev models for ModelsCard filtering, wire
systemoneConfig into ProviderInfoCard, and configure GenericExampleCard
for interactive testing of decision models.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 14:07:36 +07:00
11089ab137 Merge origin/master (v0.5.81) into gitea/new_feature
Resolve conflicts:
- streamingHandler.js: merge buildStreamErrorBytes onAbortTerminal + local shouldPersistRequestDetail & streamStatusForContent
- capabilities.js: preserve server-injected user-asserted caps and models.dev catalog lookup; wire CommandCode /alpha/generate caps inside resolve()
- commandcode.js (services/usage): adopt upstream whoami + billing credits/subscriptions with 5h/weekly rate windows and plan caps
- openai-to-commandcode.js: merge toNativeImageBlock (data-URI & http(s) support) and assistant reasoning_content preservation
- commandcode-to-openai.js: adopt upstream mid-stream error throw for clean retry and abortion
- tests: sync commandcode test suite and exclude .next from vitest config
2026-09-22 10:08:58 +07:00
decolua
41a1b8003d feat(providers): add MiMo V2.6 Flash Free to OpenCode Zen free tier
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 09:51:51 +07:00
decolua
b7446f8dd1 feat(providers): add System One (Jev) decision endpoint
New /v1/systemone pass-through route for Jev decision models (jev-1.13,
jev-1.13-free) on OpenCode Zen and the free lane. Follows the media-route
pattern: systemoneConfig in the registry drives URL/headers, the handler
mirrors the embeddings account-fallback + usage flow, and the dashboard
gains a System One media-provider kind. No chat-pipeline changes.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 09:51:29 +07:00
decolua
6886915f62 feat(capacity-adapter): default vision fallback to mimo-v2.6-flash-free
Register mimo-v2.6-flash-free on opencode-zen (chat lane) with a v2.6
capability pattern, and switch the vision adapter default from the old
mimo-v2.5-free.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-22 09:45:20 +07:00
82030615
da0046550a fix(responses): report usage on response.completed so clients can auto-compact
Map upstream Chat Completions usage to the Responses API shape and attach it to response.completed. Capture chunk.usage before the empty-choices guard so the usage-only trailer chunk survives, and defer completion to flushEvents() when usage is not yet known — only on the direct openai:openai-responses route, since a pivoted stream never reaches flushEvents. Fixes #3432.
2026-09-21 21:27:01 +07:00
Liang.Xu
402745dc1f feat(providers): add qoder-cn support for Qoder CN (qoder.com.cn) 2026-09-21 20:45:19 +07:00
Christian Gennari
be3bc764b1 fix(antigravity): separate weekly and short-window quotas and clean up redundant rows
- Track both weekly and 5-hour session buckets in parseWeeklyQuotaSummary,
  distinguishing sliding-window limits from multi-day weekly limits
- Preserve disabled session buckets at 0% rather than dropping them when weekly limits are reached
- Target 5-hour session rows (not weekly rows) during family exhaustion reconciliation in getAntigravityUsage
- Suppress synthesized per-model duplicate rows in dashboard normalization when family summaries are present
- Add unit test coverage for multi-bucket extraction, reconciliation isolation, and dashboard deduplication
2026-09-21 20:15:56 +07:00
dinhkarate
2daf25ffbe fix(qoder): handle code 110 billing blocks and preserve SSE error status
- Match code 110 (billing daily count exceeded) alongside 112/10605/pricingUrl
  in isBillingBlock, parsing JSON safely and accepting numeric/string codes
- Accept numeric strings for statusCodeValue and object bodies in envelope peek
- Emit structured 403 quota error chunk instead of synthetic assistant text
  when a billing envelope appears mid-stream
- Preserve upstream HTTP status in handleForcedSSEToJson when error chunk carries
  a valid 400-599 status
- Add unit tests for code-110 detection, mid-stream billing envelopes, and false-positive guard
2026-09-21 20:09:03 +07:00
Anantachoke
7c2b1fe3e1 fix(translator): drop replayed reasoning fields for Groq/Mistral/Cerebras (#4220)
Strict OpenAI-compatible validators reject unknown assistant-message
fields: Groq 400 ("property 'reasoning_content' is unsupported"),
Mistral 422 ("extra_forbidden"), Cerebras 400 ("wrong_api_format").

Clients driving reasoning models (Hermes Agent, and anything following
the DeepSeek/Kimi convention) echo the previous turn's reasoning_content
on every assistant message, so from the second turn on every request to
these providers fails and a fallback combo silently skips them.

Add a dropMessageFields rule to paramSupport.js that strips
reasoning_content / reasoning / reasoning_details from assistant turns
for groq, mistral, and cerebras.
2026-09-21 20:02:21 +07:00
Amir Seify
c933eefc27 fix(cursor): stop AgentService empty turns and silent tool hangs
Cursor-hosted models (cu/composer-2.5, cu/cursor-grok-*, cu/default) returned
HTTP 200 with an empty turn, or hung, whenever a client sent tools.

- Fold system prompts into the current user message. custom_system_prompt
  (RunRequest field 8) makes AgentService return an empty turn.
- Send ModelDetails (field 3); thinking variants (Composer, Grok, *-thinking)
  return an empty turn when only requested_model (field 9) is set.
- Route tool-call history and declared tool schemas through AgentService:
  encode OpenAI tools into mcp_tools (field 4), decode McpArgs and emit real
  tool_calls with finish_reason tool_calls.
- Map Composer  thinking / Grok thinking_delta (field 4) into visible content
  instead of dropping the answer with the unsigned reasoning.
- Ack request_context without echoing MCP tools (double-advertise stalls the
  HTTP/2 stream) and ack kv_server_message so the run proceeds.
- Reject IDE builtin execs instead of failing the turn, so the model can
  continue with MCP tools or a text answer.
- Add google.protobuf.Value / MCP encoders and a FIXED64 branch to
  encodeField in cursorProtobuf.js.

RTK now compresses the source-format body before translation for cursor only:
its translator rewrites role:tool into user XML, so the post-translate pass
missed those tool results. Every other provider keeps the post-translate pass
unchanged.
2026-09-21 19:49:48 +07:00
Minh Ha
5c217d34f3 feat(capabilities): model capability metadata on /v1/models, combo aggregation, pattern fixes
- Export aggregateComboCapabilities: union for vision/audio/search/pdf,
  intersection for tools, primary-model for reasoning fields, min
  contextWindow, max maxOutput
- Support nested combo resolution in aggregateComboCapabilities via
  comboLookup with depth guard (max 6)
- Wire capability metadata to all /v1/models entries and combos
- Show aggregated ctx/max metadata line and capability badges on combo chips
- Pattern fixes: MiMo v2.5/omni reasoning, qwen max/plus vision, minimax m2.x vision
- Sync commandcode model catalog and add openai gpt-5.5
- Add unit tests for capability patterns and combo capability aggregation
2026-09-21 19:41:29 +07:00
chisewaguri
477b2aed0b fix(opencode-go): send reasoning_effort for glm-5.3-flash 2026-09-21 19:10:14 +07:00
kimono381
cf663f5300 fix(huggingface): complete the Inference Providers router migration
Replace the retired api-inference.huggingface.co host with the Inference
Providers router (router.huggingface.co): imageConfig.modelMap resolves
Hub ids to provider-resolved ids, image-to-image models receive the
source image in inputs with the prompt under parameters.prompt, and a
new sttConfig wires the hf-inference ASR route. The image catalog grows
to 23 models, dead whisper-small is replaced by whisper-large-v3-turbo,
the unusable "language" param is dropped, and edit models declare the
edit capability so the dashboard offers a source image. Adds unit and
end-to-end coverage plus a model-id guard on custom endpoints.
2026-09-19 10:44:34 +07:00
Aaron
822aa958d1 fix(opencode): cloak Responses requests that already have tools
Free-tier Zen models reject Responses requests with 403 FreeTierError
when client tools are present but the fingerprint quartet is missing.
Apply the fingerprint tools to every OpenCode request, canonicalise
case variants of the quartet (Bash->bash) without duplication, and
restore the caller's original spellings on the response side via a
request-local WeakMap threaded through the existing toolNameMap.
2026-09-19 10:15:45 +07:00
Alexander Radchenko
49185137b8 feat(opencode-zen): add OpenCode Zen (PAYG) provider with free-tier fingerprint, alias ocz
Multi-endpoint provider on https://opencode.ai/zen/v1 (openai / claude /
openai-responses transports mirroring opencode-go) with 71 models across
paid + free tiers. Free-tier fingerprint: opencode/1.18.x UA spoof,
ses_-session header, built-in tool quartet, forced stream. Usage endpoint
/zen/v1/usage wired into the dashboard.
2026-09-19 10:02:28 +07:00
X-Adam
73e021b8a0 fix(ollama): map free-plan monthly window and derive reset from signup date 2026-09-19 09:51:13 +07:00