fix(tokens): 流式统计改用上游真实 usage,图片不再记 token

两处 token 单位错误,均影响 per-model 配额计费:

1. 流式路径的 prompt/completion 只是「字节÷3」估算。
   pumpStream 明明收到了上游最后一帧的真实 usage,却只发给客户端、
   从不回写审计记录,于是配额按估算值扣。生产实测同一请求:
   上游 prompt=37/completion=179 → 记账 27/262,prompt 低估 1.4x、
   completion 高估 1.5x(双向失真)。同模型流式 completion 中位数
   是非流式的 4-27 倍。非流式路径本就用真实值,两路不一致。
   修法:lastUsage 非零时写回 rec.Prompt/rec.Compl,估算降为兜底
   (上游不报 usage 时仍保留原估算行为)。

2. 图片请求把「图片张数」记成 completion_tokens。
   rec.Compl = int64(len(resp.ImageData)),len 是切片长度即张数
   (生产 38 条 image 记录全是 1),且被计入 token 总量。
   图片生成无 token 概念 ⇒ 新增 Req.ImageCount 独立字段,
   Prompt/Compl 归 0;UI 记录表 image 行改显示张数(新增 i18n thImgs)。

顺带补 TestUILocaleKeyParity:此前无人校验 zh/en 键集合一致,
单边加键不会报错,只会显示原始键名。

新增 token_units_test.go(定值上游 6 项),做过变异验证:
回退修复实测复现 stream=16/173 vs chat=44/100、image completion=3。
This commit is contained in:
JianFeeeee
2026-09-28 22:19:12 +08:00
parent de7c372ad2
commit 0121d23f91
5 changed files with 314 additions and 3 deletions

View File

@ -957,6 +957,19 @@ func (g *Gateway) pumpStream(w http.ResponseWriter, rec *Req, chunks <-chan type
// Final usage chunk (OpenAI standard: empty choices + usage before [DONE]).
// Prefer the upstream's exact usage if the stream carried it; fall back to
// the gateway's estimate otherwise.
// Write the upstream's exact numbers back onto the audit record. Without
// this the streamed path kept only the per-chunk byte estimate, so the same
// request recorded ~1.4x its real completion tokens (measured: upstream 100,
// audit 145) while the non-streaming path recorded 100. Two paths, two
// different numbers for one request is a reporting bug, not a rounding one.
if lastUsage != nil {
if lastUsage.Prompt > 0 {
rec.Prompt = int64(lastUsage.Prompt)
}
if lastUsage.Completion > 0 {
rec.Compl = int64(lastUsage.Completion)
}
}
var tut *types.TokenUsage
if lastUsage != nil {
tut = lastUsage
@ -1125,7 +1138,10 @@ func (g *Gateway) handleImage(w http.ResponseWriter, r *http.Request) {
}
rec.OK = true
rec.Status = http.StatusOK
rec.Compl = int64(len(resp.ImageData))
// Image generation has no token concept. Recording len(ImageData)
// (the image COUNT) in completion_tokens mislabels image count as
// tokens and feeds it into the token totals; leave it 0.
rec.ImageCount = len(resp.ImageData)
g.writeRec(rec)
writeJSON(w, http.StatusOK, types.ImageGenResponse{
Created: time.Now().Unix(),
@ -1174,7 +1190,8 @@ func (g *Gateway) handleImage(w http.ResponseWriter, r *http.Request) {
}
rec.OK = true
rec.Status = http.StatusOK
rec.Compl = int64(len(resp.ImageData))
// Image generation has no token concept — see the AUTO path above.
rec.ImageCount = len(resp.ImageData)
g.writeRec(rec)
writeJSON(w, http.StatusOK, types.ImageGenResponse{
Created: time.Now().Unix(),