fix(cursor): stop AgentService empty turns and silent tool hangs

Cursor-hosted models (cu/composer-2.5, cu/cursor-grok-*, cu/default) returned
HTTP 200 with an empty turn, or hung, whenever a client sent tools.

- Fold system prompts into the current user message. custom_system_prompt
  (RunRequest field 8) makes AgentService return an empty turn.
- Send ModelDetails (field 3); thinking variants (Composer, Grok, *-thinking)
  return an empty turn when only requested_model (field 9) is set.
- Route tool-call history and declared tool schemas through AgentService:
  encode OpenAI tools into mcp_tools (field 4), decode McpArgs and emit real
  tool_calls with finish_reason tool_calls.
- Map Composer  thinking / Grok thinking_delta (field 4) into visible content
  instead of dropping the answer with the unsigned reasoning.
- Ack request_context without echoing MCP tools (double-advertise stalls the
  HTTP/2 stream) and ack kv_server_message so the run proceeds.
- Reject IDE builtin execs instead of failing the turn, so the model can
  continue with MCP tools or a text answer.
- Add google.protobuf.Value / MCP encoders and a FIXED64 branch to
  encodeField in cursorProtobuf.js.

RTK now compresses the source-format body before translation for cursor only:
its translator rewrites role:tool into user XML, so the post-translate pass
missed those tool results. Every other provider keeps the post-translate pass
unchanged.
This commit is contained in:
Amir Seify
2026-09-16 21:31:22 +07:00
parent 5c217d34f3
commit c933eefc27
9 changed files with 670 additions and 71 deletions

View File

@@ -116,6 +116,19 @@ export async function handleChatCore({ body, modelInfo, credentials, log, onCred
}
}
// Per-request opt-out: client can bypass all token savers via header
const tokenSaverEnabled = clientRawRequest?.headers?.[TOKEN_SAVER_HEADER]?.toLowerCase() !== "off";
// Cursor's translator rewrites tool_result into user text, so RTK must run on
// the source body before translation. Every other pair translates the tool
// shapes 1:1 — keep the post-translate pass there so those providers are
// untouched (and a retry never re-compresses an already-compressed body).
const preTranslateRtk = provider === "cursor"
? compressMessages(body, tokenSaverEnabled && rtkEnabled)
: null;
const preTranslateRtkLine = formatRtkLog(preTranslateRtk);
if (preTranslateRtkLine) console.log(preTranslateRtkLine);
const clientRequestedStreaming = body.stream === true || sourceFormat === FORMATS.ANTIGRAVITY || sourceFormat === FORMATS.GEMINI || sourceFormat === FORMATS.GEMINI_CLI;
const providerRequiresStreaming = PROVIDERS[provider]?.forceStream === true;
let stream = providerRequiresStreaming ? true : (body.stream !== false);
@@ -252,13 +265,8 @@ export async function handleChatCore({ body, modelInfo, credentials, log, onCred
translatedBody.tools = defaultClaudeToolType(translatedBody.tools);
}
// Per-request opt-out: client can bypass all token savers via header
const tokenSaverEnabled = clientRawRequest?.headers?.[TOKEN_SAVER_HEADER]?.toLowerCase() !== "off";
// RTK: compress tool_result content
const rtkStats = compressMessages(translatedBody, tokenSaverEnabled && rtkEnabled);
const rtkLine = formatRtkLog(rtkStats);
if (rtkLine) console.log(rtkLine);
// RTK: compress tool_result content. Skipped when already done pre-translate.
const rtkStats = preTranslateRtk || compressMessages(translatedBody, tokenSaverEnabled && rtkEnabled);
// Headroom: optional external proxy compression; fail open if proxy is absent.
const headroomDiagnostics = {};
@@ -275,6 +283,8 @@ export async function handleChatCore({ body, modelInfo, credentials, log, onCred
// Token-saver flags accumulator for the single "⚙" log line below.
const xf = [];
if (rtkStats?.hits?.length) xf.push(`RTK:${rtkStats.hits.length}`);
// Caveman: inject terse-style system prompt
if (tokenSaverEnabled && cavemanEnabled && cavemanLevel) {
injectCaveman(translatedBody, finalFormat, cavemanLevel);