The opencode adapters derived x-opencode-session from meta.timestamp, i.e. a
brand new session on every request. The upstream prefix cache is
session-scoped, so no request could ever hit it, and the cache fields the
endpoint does report (prompt_tokens_details.cached_tokens,
prompt_cache_hit_tokens/prompt_cache_miss_tokens) always came back 0/absent.
Measured against the live endpoint, same 6032-token prompt:
fixed session id -> 2nd call: hit 5888, miss 144
rotating session id -> every call: hit 0, miss 6032
Fix: derive the session from the source name (stable), matching how
x-opencode-project is already derived. x-opencode-request stays unique per
request — it is only a request identifier, not part of the cache key.
Applied to both opencodego and opencodezen.
Through the gateway the same prompt now reports, on the 2nd call:
details={'cached_tokens': 5888} hit=5888 miss=144 (non-streaming)
prompt_tokens_details={'cached_tokens': 5888} (streaming)
Test: TestOpenCodeSessionIsStableForCache asserts the session is stable
across requests for one source while the request id differs.
Zen (https://opencode.ai/zen/v1) and Go (https://opencode.ai/zen/go/v1)
are different services with different requirements, and one shared adapter
could not satisfy both.
The decisive difference is reasoning_content:
* OpenCode Go runs thinking models and REQUIRES the assistant turn's
reasoning_content to be echoed back. The shared adapter stripped it
(msg.reasoning_content = nil), so every replay of a thinking turn
failed with:
400 invalid_request_error: The `reasoning_content` in the thinking
mode must be passed back to the API.
Reproduced directly: the same request with reasoning_content -> 200,
without -> 400. That is why the Go tier never worked in an agent loop.
* The Zen free pool must not receive it, so it keeps stripping.
Both adapters keep the earlier fixes they share (never drop an assistant
turn carrying tool_calls; send stream_options only when streaming; role
whitelist; multimodal strip) and the opencode client fingerprint headers —
the Go endpoint additionally REQUIRES x-opencode-session, which the
adapter already sends.
config: localzen -> opencodezen, gozen -> opencodego.
Verified: all 25 gozen models answer correctly through the gateway with a
thinking + tool_call + tool_result history (was 0/25 before), streaming
included; the Zen free models still pass.
Test: TestOpenCodeGoVsZenReasoning pins the Go-keeps / Zen-strips split.