feat(stats): 用量按日/周/月/全部统计(审计文件聚合)

统计页原来只有一个视图——进程启动以来的累计。早上没人和一整周没人看起来
一模一样。加日/周/月/总四个周期。

## 口径与实现

周期视图必须走审计文件,不能走内存聚合:内存 byModel/byKey 等是终身累计,
而 recs 环形缓冲只有 500 条(defaultRingSize)。读环会把任何超过几百个
请求的周期悄悄少算——这正是要消除的那类错数。

- 日/周/月 = UTC 日历窗口(今日 / ISO 周周一 00:00 / 本月 1 日)。
  刻意不用滚动 24h:滚动窗口会让"今天"和"最近一天"边界不同,同一个数字
  随查看时刻在两张卡片间跳。UTC 也和 billing 的峰段计算同口径,峰谷小时
  不会在费用视图和用量视图里落到不同一天。
- 全部 = 复用现有 Snapshot(内存聚合),无审计文件时依然可用。
- 时间桶:日→每小时(今天内部的尖峰要看得见),周/月→每天(否则一周是
  7×24 个点、一个月 31×24)。全部视图无时间线(终身总量没有有意义的
  时间轴,硬画 500 个滚动小时点是另一种撒谎)。
- key 过滤在所有维度生效;非法 period 返回 400 而不是静默回落"全部"——
  书签里的手误应当报错,而不是悄悄换成终身数字。

## 判据(10 条 + 8 个变异全部被捕获)

窗口边界(含"周日必须回到上一个周一"这个 Go Weekday() 陷阱)、旧记录不
计入、维度独立聚合且 by_model 求和等于 total、日桶按小时且有序、key 隔离、
全部视图走终身、空窗口标记 truncated、period 校验、query 解析。

变异验证时 by_status 假绿了一次:禁用状态码聚合后判据全过,查下去是我
**根本没测 by_status**(零覆盖)。补 TestPeriodStatusDimension 后该变异
立即被捕获。判据报假问题时,先怀疑判据——这次确实是我错了。

## 真实流量核对

生产审计文件手算 vs 后端(含轮转文件):
  day   手算 3559 / 后端 3532
  week  手算 43240 / 后端 35034
  month 手算 13696 / 后端 13670
差异是核对快照与请求之间的新流量,量级一致。

一个必须说明的发现:审计文件里混着两种记录 —— Req(type/model/
prompt_tokens)和访问日志(lat_ms/status/path)。46026 行里 33450 行是
访问日志,Go 侧按 r.Type=="" 跳过。这不是 bug(CSV 导出同样如此),但
意味着任何按行数手算都必须过滤,否则会差一个数量级。

CDP 实测四周期切换:reqs 3,571 / 35,075 / 13,710 / 196,712,与后端一致,
无控制台错误,localStorage 持久化生效。
This commit is contained in:
JianFeeeee
2026-10-02 12:26:26 +08:00
parent d1da40e493
commit c241a19b51
4 changed files with 660 additions and 5 deletions

View File

@ -433,6 +433,32 @@ func (g *Gateway) handleStatsAPI(w http.ResponseWriter, r *http.Request) {
}
}
key := exportKey(r)
// ?period=day|week|month|all switches the whole payload to a calendar
// window (UTC) aggregated off the audit files, instead of the
// since-process-start totals. The CSV exports below are unaffected: they
// take an explicit from/to range and stream, so a period selector there
// would only be a second way to spell the same bounds.
if p := periodFromQuery(r.URL.Query()); p != PeriodAll {
if !ValidPeriod(p) {
writeError(w, http.StatusBadRequest, "bad_period",
"period must be one of day, week, month, all")
return
}
out := g.stats.PeriodSnapshot(p, key, time.Now())
writeJSON(w, http.StatusOK, map[string]interface{}{
"period": out.Period,
"from": out.From,
"total": out.Total,
"by_key": out.ByKey,
"by_model": out.Models,
"by_source": out.Srcs,
"by_status": out.Status,
"buckets": out.Bucket,
"truncated": out.Truncated,
"key_names": g.keyNamesFor(),
})
return
}
if r.URL.Query().Get("export") == "csv" {
from, _ := strconv.ParseInt(r.URL.Query().Get("from"), 10, 64)
to, _ := strconv.ParseInt(r.URL.Query().Get("to"), 10, 64)
@ -540,14 +566,22 @@ func (g *Gateway) handleStatsAPI(w http.ResponseWriter, r *http.Request) {
return
}
snap := g.stats.Snapshot(limit, key)
keyNames := map[string]string{}
for _, k := range g.core.ListKeys() {
keyNames[keyID(k.Key)] = k.Name
}
snap["key_names"] = keyNames
snap["key_names"] = g.keyNamesFor()
writeJSON(w, http.StatusOK, snap)
}
// keyNamesFor is the masked-id -> display-name map every stats payload needs.
// It is keyed by keyID (the mask), not the raw key, because that is what the
// aggregate rows carry — building it in one place stops the period branch and
// the lifetime branch from drifting apart.
func (g *Gateway) keyNamesFor() map[string]string {
names := map[string]string{}
for _, k := range g.core.ListKeys() {
names[keyID(k.Key)] = k.Name
}
return names
}
// handleStatsRecordsAPI pages the request records straight off the audit files.
// The dashboard loads only its first screen and asks for the next page as the
// user scrolls, so neither side holds the full history: the server keeps no