mirror of
https://gitcode.com/JianFeeeee/ModelRouter.git
synced 2026-09-21 17:38:00 +00:00
feat: surface zero cache hits — distinguish 'missed' from 'not reported'
Live testing across the zen pool showed models report prompt_tokens_details.cached_tokens even when the hit count is 0 (e.g. nemotron-3-ultra-free returns cached_tokens:0, audio_tokens:0, cache_write_tokens:0). The previous >0 guard dropped those objects, so a cache-enabled upstream looked identical to one without cache support. - types: PromptTokensDetails.CachedTokens always emitted (drop inner omitempty) so clients see cached_tokens:0 explicitly; dsh reads it as a 0% hit instead of 'no data' - adapters (9): forward prompt_tokens_details whenever the upstream provides it (presence check instead of >0) - Req: add cache_reported flag set when usage carried cache accounting; WebUI shows an amber 0% tag for reported-but-missed rows and keeps the em-dash only for sources that never report cache data
This commit is contained in:
@ -99,7 +99,9 @@ type TokenUsage struct {
|
||||
// PromptTokensDetails is the OpenAI v2 prompt_tokens_details object. Only
|
||||
// CachedTokens is emitted (omitempty drops the whole object when zero).
|
||||
type PromptTokensDetails struct {
|
||||
CachedTokens int `json:"cached_tokens,omitempty"`
|
||||
// CachedTokens is always emitted (even 0) so clients can distinguish
|
||||
// "upstream reports cache, this request missed" from "no cache data".
|
||||
CachedTokens int `json:"cached_tokens"`
|
||||
}
|
||||
|
||||
// MarshalJSON emits both the legacy short keys (prompt/completion/total, used
|
||||
|
||||
Reference in New Issue
Block a user