Files
9router/open-sse/AGENTS.md
luulam f0adfb205a feat(dashboard): per-key model restrictions, pin header routing, combo side-panel picker
- Endpoint: per-API-key model allowlist (schema v3) enforced on chat (403)
  and /v1/models; Full-access toggle + multi-select picker in Keys UI.
- Providers: honor x-connection-id in /v1/chat/completions — pinned requests
  no longer rotate to another account on failure.
- Providers: strategy saves merge into stored enabled:false override;
  Test All groups match grid sections; 1-by-1 skips disabled connections.
- Dashboard: provider-card toggle syncs from server on failure; grid toggles
  always visible; connection rows get clear-✕ for stale error banners.
- Combo editor: on desktop (xl+) the Add-Model picker opens as a floating
  side panel beside the untouched combo popup instead of stacking on top;
  mobile keeps the full-screen overlay.
- Long API-key overflow fixed in key rows + provider model sections.
2026-08-27 09:26:17 +07:00

5.1 KiB

open-sse

Provider-agnostic SSE engine: one OpenAI-style request → any provider (LLM chat, image, embedding, tts, stt, search), streamed back in the client's format.

Request lifecycle (chat)

handlers/chatCore.js → services/model.js parseModel (resolve provider/model) → pre-translate hooks (rtk/ tool_result compress, rtk/headroom.js proxy compress, rtk/caveman.js system inject — all fail-open) → executors/index.js getExecutor(provider) → translator/index.js translateRequest (client format → provider format) → executor.execute() (streams upstream) → translateResponse (provider chunks → client format) → SSE out.

Directory map

  • config/ — ALL constants/config (no hardcode elsewhere). providers.js/registry/ (provider defs), providerModels.js (alias→models matrix), runtimeConfig.js (timeouts, token limits), *Constants.js.
  • translator/ — format conversion. request/<from>-to-<to>.js, response/<from>-to-<to>.js, schema/ (enums: ROLE, CLAUDE_BLOCK…), concerns/ (shared logic), formats.js+formats/ (per-format). index.js is the registry/entry.
  • executors/ — per-provider upstream call. base.js (BaseExecutor), one file per special provider, index.js map.
  • providers/ — registry build + capabilities.js + pricing.js. Entry: index.js (PROVIDERS).
  • handlers/ — per-modality cores (chat/image/embedding/tts/stt/search) + sub-provider folders. chatCore/ has the streaming/non-streaming/sse-to-json handlers.
  • rtk/ — request token-killer. index.js compresses tool_result content in-place (OpenAI/Claude/Kiro shapes); filters/ per-tool compressors + autodetect.js; headroom.js external compress proxy; caveman.js system-prompt injector.
  • transformer/ — responsesTransformer.js (Chat Completions SSE → Codex Responses API SSE), streamToJsonConverter.js.
  • shared/ — cross-provider auth/identity: clineAuth.js, machineId.js, qoder/.
  • services/ — model.js, provider.js, accountFallback.js, combo.js, tokenRefresh/+tokenRefresh.js, oauthCredentialManager.js, usage/, projectId.js, kiroModels.js/qoderModels.js.
  • utils/ — streamHandler, stream, sse, error, sessionManager, claudeCloaking, clientDetector, proxyFetch (patches global fetch), cursorProtobuf/cursorChecksum, ollamaTransform.

Conventions

  • Config-driven, DRY, camelCase. NEVER hardcode values, models, or block/role strings — use config/ + schema/ constants.
  • Translator pipeline pivots through OpenAI as the intermediate format. A translator registered on the exact source:target pair (e.g. claude:kiro) runs as a direct route, skipping the lossy double-hop.
  • Translators self-register via register(from, to, reqFn, resFn) as an import side-effect — new files MUST be imported in translator/index.js.

How to add

  • Provider: copy providers/REGISTRY_TEMPLATE.js → providers/registry/{id}.js; add models to config/providerModels.js. Generic providers need no executor (DefaultExecutor handles OpenAI-compatible APIs).
  • Executor (only for non-standard upstream): subclass BaseExecutor (override getBaseUrls/buildHeaders/buildUrl/execute), register in executors/index.js map. getExecutor falls back to DefaultExecutor when absent.
  • Translator: add request|response/<from>-to-<to>.js calling register(...), then import it in translator/index.js. Reuse schema/ + concerns/ — don't re-implement parsing.

Pitfalls

  • OpenAI bridge is lossy (thinking, non-base64 images, tool ids, is_error) — prefer a direct route for fragile pairs.
  • registry/index.js is an auto-generated static import list; regenerate it (don't hand-edit) after adding a registry/{id}.js. REGISTRY_TEMPLATE is excluded by design.
  • Special binary/protobuf formats (kiro EventStream, cursor protobuf, commandcode NDJSON) don't round-trip through OpenAI — handle in their executor.
  • rtk/ + headroom.js mutate the request body in-place and are fail-open: any error returns null and leaves the body untouched — never throw out of them. RTK skips is_error/status:"error" tool results to preserve traces.
  • HTTP 200 in-stream errors: some upstreams signal failure INSIDE a 200 stream (AI SDK v5 {"type":"error"} events, error text in content). HTTP-level success checks miss these → no fallback, Status: success in logs. Three hook points + one config escape hatch:
    1. Translator — never map an error event to content. Emit an OpenAI-shaped chunk.error = { message, type } + terminal chunk (translator/response/commandcode-to-openai.js is the worked example). Downstream parseSSEToOpenAIResponse already detects chunk?.error.
    2. Executor early-peek — for streaming fallback, read the first events BEFORE returning the response; an error → non-ok Response (executors/commandcode.js peekForUpstreamError).
    3. Config escape hatch (no code) — per-provider streamErrorPatterns setting (UI: provider page → Stream Error Patterns). Patterns matched against the first ~8KB of the stream and the assembled non-streaming content; see utils/streamErrorPeek.js + utils/streamErrorPatterns.js.