fix(claude): reconcile max_tokens vs thinking budget and lift per-model ceiling (#2381)

On the translated OpenAI->Claude path, adjustMaxTokens capped max_tokens
before applyThinking set thinking.budget_tokens, so max-effort budget
(128000) could exceed a 64k-clamped max_tokens -> Anthropic 400.
prepareClaudeRequest now reconciles after the budget is known: prefer
raising max_tokens, only shrink budget when it meets/exceeds the ceiling.

Also lift the global 64000 cap: the ceiling is now the model's real
maxOutput, so high-output models (fable/mythos, opus-4.8/sonnet-4.6) get
their full budget. adjustMaxTokens gains an optional ceiling arg (default
unchanged, callers untouched); openai-to-claude passes the model maxOutput.

Native Claude Code passthrough is unaffected.

Co-Authored-By: Claude <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
This commit is contained in:
thienpv
2026-07-05 17:32:25 +07:00
committed by decolua
parent 5041494e1c
commit 46e6c01a01
4 changed files with 129 additions and 8 deletions

View File

@@ -6,6 +6,7 @@ import { safeParseJSON } from "../concerns/json.js";
import { parseDataUri } from "../concerns/image.js";
import { extractTextContent } from "../formats/gemini.js";
import { ROLE, OPENAI_BLOCK, CLAUDE_BLOCK } from "../schema/index.js";
import { getCapabilitiesForModel } from "../../providers/capabilities.js";
// Empty prefix matches real Claude Code behavior (no tool name prefix).
// Previously "proxy_" was used but this is a detectable fingerprint difference.
@@ -15,9 +16,13 @@ const CLAUDE_OAUTH_TOOL_PREFIX = "";
export function openaiToClaudeRequest(model, body, stream) {
// Tool name mapping for Claude OAuth (capitalizedName → originalName)
const toolNameMap = new Map();
// Cap max_tokens at the model's real output ceiling (e.g. Opus 4.8 = 128000),
// not the conservative 64000 default — otherwise a high-output model is
// pre-clamped here before prepareClaudeRequest's model-aware step runs.
const modelCeiling = getCapabilitiesForModel(null, model).maxOutput || undefined;
const result = {
model: model,
max_tokens: adjustMaxTokens(body),
max_tokens: adjustMaxTokens(body, modelCeiling),
stream: stream
};