GitHub Copilot's /chat/completions and /responses endpoints never surface prompt-cache token counts for Claude models. Route Claude models (detected by name pattern) to Copilot's Anthropic-native /v1/messages shim via a new executeWithMessagesEndpoint(), translating OpenAI-shape requests to Claude natively so cache_control gets injected and cached_tokens surface. Also fixes translateRequest()'s internal _toolNameMap being sent upstream, which made Anthropic's strict schema reject tool-call requests with a 400 — now stripped and threaded through response state. Removes the now-dead response_format Claude JSON-mode workaround.
17 KiB
17 KiB