Rewrite LuaAdaptedProvider.ChatStream to match the maturity of
llmsproxy's streaming implementation:
HTTP layer:
- Dedicated stream HTTP client with no overall timeout (SSE must not
be cut by the 180s Chat timeout); only a 30s dial timeout
- Uses applyAdapterHeaders (supports build_headers dynamic signing
hook), matching the non-streaming Chat path
Non-200 response handling:
- New TransformError Lua hook (adapter.transform_error) for per-source
protocol knowledge in error messages
- Safe fallback truncation of raw error bodies (prevents HTML dump
leakage to clients)
SSE parsing enhancements:
- parseOpenAICompatibleStreamChunkFull: handles token usage in the
final chunk (prompt_tokens/prompt, total_tokens/total dual keys),
prompt cache detail fields, and empty-string finish_reason filtering
(sensenova sends "" on every chunk)
- Replaced old SSEScanner with bufio.Scanner (larger buffer, fewer
allocations)
Stream integrity:
- errorOnlyChunk detection: holds back the first chunk to reject
degenerate streams (e.g. zen free pool's finish_reason:"network_error"
with empty content) before any byte reaches the caller
- [DONE] dedup: adapters that already emit a terminating done chunk
with the real finish_reason don't get a second reason-less done
- Clean EOF sends a final Done:true if no done was seen
Struct changes:
- StreamChunk: added FinishReason and Usage fields for callers
- LuaAdaptedProvider: added streamClient (lazy) + streamMu
Tested: curl against llmsproxy SSE confirms reasoning_content parsing
is correct (delta.reasoning_content), usage chunk handling works, and
[DONE] termination is properly emitted.