mirror of
https://gitcode.com/JianFeeeee/ModelRouter.git
synced 2026-10-06 07:27:30 +00:00
fix(tokens): 流式统计改用上游真实 usage,图片不再记 token
两处 token 单位错误,均影响 per-model 配额计费: 1. 流式路径的 prompt/completion 只是「字节÷3」估算。 pumpStream 明明收到了上游最后一帧的真实 usage,却只发给客户端、 从不回写审计记录,于是配额按估算值扣。生产实测同一请求: 上游 prompt=37/completion=179 → 记账 27/262,prompt 低估 1.4x、 completion 高估 1.5x(双向失真)。同模型流式 completion 中位数 是非流式的 4-27 倍。非流式路径本就用真实值,两路不一致。 修法:lastUsage 非零时写回 rec.Prompt/rec.Compl,估算降为兜底 (上游不报 usage 时仍保留原估算行为)。 2. 图片请求把「图片张数」记成 completion_tokens。 rec.Compl = int64(len(resp.ImageData)),len 是切片长度即张数 (生产 38 条 image 记录全是 1),且被计入 token 总量。 图片生成无 token 概念 ⇒ 新增 Req.ImageCount 独立字段, Prompt/Compl 归 0;UI 记录表 image 行改显示张数(新增 i18n thImgs)。 顺带补 TestUILocaleKeyParity:此前无人校验 zh/en 键集合一致, 单边加键不会报错,只会显示原始键名。 新增 token_units_test.go(定值上游 6 项),做过变异验证: 回退修复实测复现 stream=16/173 vs chat=44/100、image completion=3。
This commit is contained in:
@ -957,6 +957,19 @@ func (g *Gateway) pumpStream(w http.ResponseWriter, rec *Req, chunks <-chan type
|
||||
// Final usage chunk (OpenAI standard: empty choices + usage before [DONE]).
|
||||
// Prefer the upstream's exact usage if the stream carried it; fall back to
|
||||
// the gateway's estimate otherwise.
|
||||
// Write the upstream's exact numbers back onto the audit record. Without
|
||||
// this the streamed path kept only the per-chunk byte estimate, so the same
|
||||
// request recorded ~1.4x its real completion tokens (measured: upstream 100,
|
||||
// audit 145) while the non-streaming path recorded 100. Two paths, two
|
||||
// different numbers for one request is a reporting bug, not a rounding one.
|
||||
if lastUsage != nil {
|
||||
if lastUsage.Prompt > 0 {
|
||||
rec.Prompt = int64(lastUsage.Prompt)
|
||||
}
|
||||
if lastUsage.Completion > 0 {
|
||||
rec.Compl = int64(lastUsage.Completion)
|
||||
}
|
||||
}
|
||||
var tut *types.TokenUsage
|
||||
if lastUsage != nil {
|
||||
tut = lastUsage
|
||||
@ -1125,7 +1138,10 @@ func (g *Gateway) handleImage(w http.ResponseWriter, r *http.Request) {
|
||||
}
|
||||
rec.OK = true
|
||||
rec.Status = http.StatusOK
|
||||
rec.Compl = int64(len(resp.ImageData))
|
||||
// Image generation has no token concept. Recording len(ImageData)
|
||||
// (the image COUNT) in completion_tokens mislabels image count as
|
||||
// tokens and feeds it into the token totals; leave it 0.
|
||||
rec.ImageCount = len(resp.ImageData)
|
||||
g.writeRec(rec)
|
||||
writeJSON(w, http.StatusOK, types.ImageGenResponse{
|
||||
Created: time.Now().Unix(),
|
||||
@ -1174,7 +1190,8 @@ func (g *Gateway) handleImage(w http.ResponseWriter, r *http.Request) {
|
||||
}
|
||||
rec.OK = true
|
||||
rec.Status = http.StatusOK
|
||||
rec.Compl = int64(len(resp.ImageData))
|
||||
// Image generation has no token concept — see the AUTO path above.
|
||||
rec.ImageCount = len(resp.ImageData)
|
||||
g.writeRec(rec)
|
||||
writeJSON(w, http.StatusOK, types.ImageGenResponse{
|
||||
Created: time.Now().Unix(),
|
||||
|
||||
Reference in New Issue
Block a user