feat: surface zero cache hits — distinguish 'missed' from 'not reported'

Live testing across the zen pool showed models report
prompt_tokens_details.cached_tokens even when the hit count is 0 (e.g.
nemotron-3-ultra-free returns cached_tokens:0, audio_tokens:0,
cache_write_tokens:0). The previous >0 guard dropped those objects, so a
cache-enabled upstream looked identical to one without cache support.

- types: PromptTokensDetails.CachedTokens always emitted (drop inner
  omitempty) so clients see cached_tokens:0 explicitly; dsh reads it as
  a 0% hit instead of 'no data'
- adapters (9): forward prompt_tokens_details whenever the upstream
  provides it (presence check instead of >0)
- Req: add cache_reported flag set when usage carried cache accounting;
  WebUI shows an amber 0% tag for reported-but-missed rows and keeps
  the em-dash only for sources that never report cache data
This commit is contained in:
dev
2026-08-25 10:04:44 +08:00
parent ec89daad62
commit 21ec8f59d8
13 changed files with 40 additions and 23 deletions

View File

@ -99,7 +99,9 @@ type TokenUsage struct {
// PromptTokensDetails is the OpenAI v2 prompt_tokens_details object. Only
// CachedTokens is emitted (omitempty drops the whole object when zero).
type PromptTokensDetails struct {
CachedTokens int `json:"cached_tokens,omitempty"`
// CachedTokens is always emitted (even 0) so clients can distinguish
// "upstream reports cache, this request missed" from "no cache data".
CachedTokens int `json:"cached_tokens"`
}
// MarshalJSON emits both the legacy short keys (prompt/completion/total, used