fix(stream): parse the trailing NDJSON line an Ollama stream leaves behind

createSSEStream splits on "\n" and keeps the remainder, which only flush()
parses. That call omitted targetFormat, so parseSSELine required a "data: "
prefix and dropped whatever an NDJSON provider left without a closing
newline. The !parsed.done guard compounded it: the SSE sentinel and an
Ollama final chunk both carry done:true, but the latter is the real last
chunk holding done_reason and the token counts.

Pass targetFormat and scope the sentinel check to formats that emit one, so
the tail reaches the translator. Accumulate its usage into state the same way
the transform loop does, so finalizeStream logs those tokens instead of null.
This commit is contained in:
Nguyen Thanh Dat
2026-08-25 13:35:48 +07:00
parent 2f17352cc2
commit f9d82c6575
2 changed files with 111 additions and 2 deletions

View File

@@ -408,8 +408,21 @@ export function createSSEStream(options = {}) {
}
if (buffer.trim()) {
const parsed = parseSSELine(buffer.trim());
if (parsed && !parsed.done) {
// Same parse as the transform loop: without targetFormat this only
// accepts "data: " lines, so an NDJSON provider (Ollama) lost whatever
// arrived without its closing newline.
const parsed = parseSSELine(buffer.trim(), targetFormat);
// parseSSELine turns the SSE sentinel "data: [DONE]" into { done: true },
// which must not be translated. An Ollama chunk also carries done:true,
// but it is the real final chunk — it holds finish_reason and the token
// counts — so it has to go through.
const isDoneSentinel = parsed?.done && targetFormat !== FORMATS.OLLAMA;
if (parsed && !isDoneSentinel) {
// Same accumulation the transform loop does, so finalizeStream() can
// log a tail chunk's tokens instead of falling back to null.
const extracted = extractUsage(parsed);
if (extracted) state.usage = mergeUsage(state.usage, extracted);
const translated = translateResponse(targetFormat, sourceFormat, parsed, state);
if (translated?._openaiIntermediate) {