Commit Graph

16 Commits

Author SHA1 Message Date
c241a19b51 feat(stats): 用量按日/周/月/全部统计(审计文件聚合)
统计页原来只有一个视图——进程启动以来的累计。早上没人和一整周没人看起来
一模一样。加日/周/月/总四个周期。

## 口径与实现

周期视图必须走审计文件,不能走内存聚合:内存 byModel/byKey 等是终身累计,
而 recs 环形缓冲只有 500 条(defaultRingSize)。读环会把任何超过几百个
请求的周期悄悄少算——这正是要消除的那类错数。

- 日/周/月 = UTC 日历窗口(今日 / ISO 周周一 00:00 / 本月 1 日)。
  刻意不用滚动 24h:滚动窗口会让"今天"和"最近一天"边界不同,同一个数字
  随查看时刻在两张卡片间跳。UTC 也和 billing 的峰段计算同口径,峰谷小时
  不会在费用视图和用量视图里落到不同一天。
- 全部 = 复用现有 Snapshot(内存聚合),无审计文件时依然可用。
- 时间桶:日→每小时(今天内部的尖峰要看得见),周/月→每天(否则一周是
  7×24 个点、一个月 31×24)。全部视图无时间线(终身总量没有有意义的
  时间轴,硬画 500 个滚动小时点是另一种撒谎)。
- key 过滤在所有维度生效;非法 period 返回 400 而不是静默回落"全部"——
  书签里的手误应当报错,而不是悄悄换成终身数字。

## 判据(10 条 + 8 个变异全部被捕获)

窗口边界(含"周日必须回到上一个周一"这个 Go Weekday() 陷阱)、旧记录不
计入、维度独立聚合且 by_model 求和等于 total、日桶按小时且有序、key 隔离、
全部视图走终身、空窗口标记 truncated、period 校验、query 解析。

变异验证时 by_status 假绿了一次:禁用状态码聚合后判据全过,查下去是我
**根本没测 by_status**(零覆盖)。补 TestPeriodStatusDimension 后该变异
立即被捕获。判据报假问题时,先怀疑判据——这次确实是我错了。

## 真实流量核对

生产审计文件手算 vs 后端(含轮转文件):
  day   手算 3559 / 后端 3532
  week  手算 43240 / 后端 35034
  month 手算 13696 / 后端 13670
差异是核对快照与请求之间的新流量,量级一致。

一个必须说明的发现:审计文件里混着两种记录 —— Req(type/model/
prompt_tokens)和访问日志(lat_ms/status/path)。46026 行里 33450 行是
访问日志,Go 侧按 r.Type=="" 跳过。这不是 bug(CSV 导出同样如此),但
意味着任何按行数手算都必须过滤,否则会差一个数量级。

CDP 实测四周期切换:reqs 3,571 / 35,075 / 13,710 / 196,712,与后端一致,
无控制台错误,localStorage 持久化生效。
2026-10-02 12:26:26 +08:00
18e422a943 fix(sources): 源编辑不再清零 proxy_url / api_key_env / timeout
bf0657b 修好了 api_key,但 upsert 仍会重写整个源,于是请求无法表达的字段
一律被重置为零值。这四个字段的后果都不是"少个配置项":

  - api_key_env 丢失 ⇒ 盘上无明文密钥的源变成无凭据源,写操作返回 200,
    下一次调用上游才 401。而 README 恰恰把这个特性当作卖点在宣传。
  - proxy_url 丢失 ⇒ 一个走代理的上游变成直连(或反之),且完全无声。
  - timeout / queue_timeout 丢失 ⇒ 退回默认 120s / 60s。

触发路径不是只有脚本:WebUI 的 saveSource() 发的 payload 只含表单上的
11 个字段,而 editSource() 表单里根本没有这 4 项 ⇒ 运维在界面上改个并发数
就会静默清掉它们。

修法用「存在性」语义而不是「空即继承」:

  - 不传   → 保留已存值(部分更新的客户端要的就是这个)
  - 传了   → 覆盖,包括传空串表示清空

api_key 刻意保留它原有的「空即继承」规则,不跟着改成指针:该规则已随
v1.7.6 发布,脚本依赖它;而凭据丢失比代理丢失严重得多。两个字段的失败
模式相反,所以规则相反——这一条写进了两处注释。

另一处是差一点的:config.Source 把 Timeout/QueueTimeout 标成 json:"-",
所以 reveal 接口的结构体序列化**根本不返回它们**。表单读不到 → 输入框
恒空 → 而输入框每次都回传 → 每存一次就把 timeout 清零。等于把刚修好的
丢字段换个方向又造了一个。因此 reveal 分支现在显式返回 duration 字符串。

判据:
  - TestWebUIEditPayloadPreservesRoutingFields 用的是 WebUI 真实 payload
    的逐字节副本,并同时断言"确实改动的字段生效",否则"什么都不写"也能过
  - TestSourceEditPreservesAPIKeyEnv 单独盯 api_key_env(唯一造成凭据丢失的)
  - CanBeSet / CanBeCleared 分别盯两个方向:只有"空即继承"的实现过不了
    CanBeCleared(清空代理框会永远保留旧代理)
  - 持久化判据**重新加载 config.yaml 并按语义比对**:300s 会被重新序列化成
    5m0s,按字符串匹配是假红(我自己先踩了一次)
  - TestSourcePayloadCoversEveryEditableField 用反射卡住"这一类":新增
    Source 字段而没接到 API 上时立刻变红。反射查结构体而非 marshal 结果,
    因为指针 + omitempty 会合法地从序列化输出里消失,那正是"未提及"信号
  - 三个 UI 契约判据把 JS 侧也钉住(表单必须回传、必须从 reveal 读)

变异验证(每次都先确认 build 通过,再数红格):
  1. 去掉覆盖逻辑        → 7 个判据红
  2. 改成"空即继承"      → CanBeCleared + ClearIsScoped 红
  3. reveal 不返回 duration → TestSourceRevealExposesDurations 红
  4. 表单不回传 api_key_env → TestUIEditFormRoundTrips... 红

顺带修正 /api/v1 索引:DELETE /api/keys 的路径段写的是 {name},实际是 key
本身;PUT /api/keys/{key} 实现了却没列。

全量 + vet + race 全绿;WebUI 内联脚本过 node --check。
2026-10-01 20:36:48 +08:00
bf0657bb84 fix(sources): implement PUT and stop partial edits from clobbering api_key
Two defects on the admin source write path, both found while adding a model
to a live source by hand.

PUT /api/sources/{name} was advertised in the API index but never
implemented — handleSourcesAPI only switched on GET/POST/DELETE, so the
documented update verb answered 405 while the POST upsert behind it worked.

POST is an upsert that replaces the whole source, so a partial edit that did
not carry api_key persisted an empty or placeholder credential. The source
kept its name, base_url and models, the write returned 200, and the source
then answered 401 on the next request — long after the writing script exited
0. The WebUI had been routing around this by loading the real key through
?reveal=credentials; any script or partial update went straight into it.

- implement PUT, taking the name from the path and rejecting a body name
  that disagrees rather than silently resolving to one of them
- inherit the stored credential when api_key is omitted or sent as the
  literal "__KEEP__"; an explicit new key still rotates
- an empty api_key on a source that does not exist yet stays empty, since
  credential-less local upstreams are legitimate
- add model_ids, an additive shorthand, so "add these models" never has to
  read and echo the existing list back
- align the API index with the implementation

The model_ids merge had a first cut that dropped the existing list when the
request carried no models field; TestSourceModelIDsIsAdditive caught it.

Verified by mutation: removing PUT turns three tests red, flattening
resolveAPIKey into a pass-through turns TestSourceUpsertKeepsAPIKey red
on both subtests, and making model_ids replace instead of merge turns
TestSourceModelIDsIsAdditive red.
2026-10-01 18:15:29 +08:00
882288f67f fix(webui): send DELETE when removing keys and adapters
Deleting a gateway key from the admin UI did nothing and reported
"use GET /api/keys": delKey() called api() with an empty options object, so
fetch defaulted to GET and the request landed in the GET branch of
handleKeysAPI. delAdapter() had the identical bug and reported
"adapter code not exposed; edit in UI".

This is the third instance of the same mistake — ab20f1b fixed delSource and
delTemplate, missing these two — so it is now pinned by tests instead of by
review:

  * TestUIAPICallsDeclareMethod walks every api() call in the embedded
    index.html and fails if one passes an options object without a method
    (an AbortSignal-only read is allowed, being a deliberate GET).
  * TestUIDeleteHelpersUseDelete / TestUIMutatingHelpersUseWriteMethods pin the
    verb of each removal and write helper by name.
  * TestKeyDeleteRoundTrip covers create -> DELETE -> gone -> second DELETE is a
    clean 404, and TestCannotDeleteOwnKey keeps the lockout guard.

The 404 bodies for GET /api/keys/{key} and GET /api/adapters/{name} now name the
verb to use ("DELETE /api/keys/{key} to remove"), because that message is what a
mis-methoded client actually shows its user; "use GET /api/keys" read as though
the caller had done nothing wrong.

delSource's indentation, broken by ab20f1b, is also straightened out.
2026-08-30 09:07:05 +08:00
d42c02b15d perf(gateway): load request logs on demand instead of holding them in memory
Startup RSS on this deployment was 56 MB with a 29 MB audit log and ~10 MB
without one: LoadAudit() json-unmarshalled the ENTIRE file into the aggregates
and kept a 10000-entry ring of raw records. Two more paths had the same shape —
AuditRecords() materialized a whole export window into a []Req before sorting
it, and a dashboard poll serialized the full ring so the browser could render
300 rows of it.

The audit file is now the source of truth and memory only holds the live
window:

  * LoadAudit replays only the last auditReplayBytes (4 MB) and drops the
    truncated first line; the ring default drops 10000 -> 500, which still
    covers both of its consumers (the status page's 5-minute SourceRecent /
    SourceAverages windows and the first screen of the records table).
    replayPartial is exported so the UI can say the totals cover a window
    rather than all time. Token-quota accounting is unaffected: it reads the
    modelHour buckets, not the ring (pinned by a test).
  * AuditPage(cursor, limit, key) pages records straight off disk, reading the
    newest file backwards in 64 KB chunks and returning as soon as the page is
    full. The cursor is "<file>:<offset>" and walks into rotated .old files;
    a cursor whose file rotated away reports rotated=true so the client can
    reset instead of silently skipping records. No state is cached between
    requests and the file handle is closed before responding, so "release when
    the user leaves the page" is guaranteed by never retaining anything.
  * StreamAuditRecords(from,to,key,fn) replaces the accumulate-then-sort export
    path; the CSV handler writes rows as they are read and flushes every 1000,
    and a write error (client gone) aborts the walk. Export memory is O(1)
    regardless of the window. AuditRecords is kept as a test-only wrapper.
  * Snapshot ships one screen (firstScreenRecords=100) by default; aggregates
    are untouched.
  * Audit rotation 64 MB x 10 -> 16 MB x 16: same 256 MB total budget, but a
    smaller newest file keeps the first reverse page cheap.

New route: GET /api/stats/records?before=&limit=&key= (non-admins are pinned to
their own key by exportKey). /api/status additionally reports adapter_pools for
admins.

Measured with production's 29 MB audit copied to the test instance: startup RSS
19.0 MB (was 56 MB); scrolling 10 pages (1000 records) +0.7 MB; exporting the
full history (36441 rows / 4.4 MB CSV) +0.1 MB with no residual growth.
2026-08-30 08:05:54 +08:00
8334cffbc9 feat(templates): source templates for multi-key balancing
A template stores every source field except name and api_key, so
operators spin up N key-bearing sources from one shared skeleton
instead of duplicating the whole source block N times.

- config: SourceTemplate type + RuntimeConfig.SourceTemplates stored
  in runtime.json alongside runtime sources
- store: UpsertTemplate / ListTemplates / RemoveTemplate
- core: Templates / SaveTemplate / RemoveTemplate
- gateway: GET/POST/DELETE /api/source_templates
- webui: source list gains a Templates button opening a manager with
  per-template edit/delete; the add-source dialog gains 'from
  template' (event-delegated picker) and 'as template' (card modal)
  buttons in its header; z-index fixed so the template editor layers
  above the manager
2026-08-27 12:09:38 +08:00
dev
24609289e8 feat: record cache hit/miss per request in audit trail and WebUI
- Req: add CacheHit and CacheMiss fields (carrying upstream cache
  accounting from either prompt_tokens_details.cached_tokens or legacy
  prompt_cache_hit_tokens)
- recordChatUsage (non-streaming): copy cache fields from resp.TokenUsage
- pumpStream (streaming): write lastUsage cache fields back onto rec at
  stream end, so streaming requests carry cache data too
- CSV export: add first_byte_ms, cache_hit_tokens, cache_miss_tokens
  columns alongside the existing latency/prompt/completion
- WebUI request-records table: add a Cache column showing hit% per row
  (green/amber tag with tooltip hit/miss breakdown; em-dash when the
  upstream reported no cache data)
2026-08-25 09:36:21 +08:00
dev
f3d5ba6cea refactor(gateway): deduplicate stats CSV export paths
- extract csvHeaders() (Content-Type + Content-Disposition) shared by both
  export branches
- extract exportKey(): the identical admin/user key-filter logic existed
  twice (JSON path + keys-csv); now all three call sites share one function
- inline the nine single-use intermediate variables in the keys-csv loop
2026-08-24 22:41:48 +08:00
690f55c6f1 feat: proactive rate limiting (RPM) + 429 short cooldown for sources
- config.go: Source add RPM field (requests-per-minute cap, 0=unlimited)
- provider.go: RecordRateLimit() — 429 uses fixed 30s cooldown, not exponential
- provider.go: Throttle() — token-bucket proactive rate limiter, spaces requests
  at 60s/RPM interval, respects context cancellation
- provider.go: ReportStatus() — 429 -> RecordRateLimit, 5xx -> RecordFailure
- provider.go: Chat/ChatStream — wire Throttle after TryAcquire
- api.go: sourcePayload + RPM, buildSource passes RPM through
- ui/index.html: add RPM input field in source editor, bilingual i18n labels
- deploy.sh: backup old binary + rollback on healthcheck failure
- provider_test.go: TestModelStateRateLimitShortCooldown, TestThrottleSpacingAndCancel
- config.yaml: sensenova rpm: 12
2026-08-24 02:00:25 +08:00
20d2268546 fix: close P10 audit items — P10-1 sources API admin guard, P10-2 scope prefix strip, P10-3 runtime source timeout w/ stream-safe clients
- api.go: handleSourcesAPI now requires admin role (GET leaks upstream api_keys, POST/DELETE mutate routing)
- chat.go: hasScopeModel made a Gateway method that strips source-model/:// prefix strictly via Registry.EffectiveModel (only when the prefix names a real source serving the bare model) so dash-bearing ids like deepseek-v4-flash-free are never corrupted; +TestHasScopeModelWithSourcePrefix
- config.go: DefaultSourceTimeout/QueueTimeout/Concurrency constants shared by YAML ApplyDefaults and runtime sources
- core.go: mergedSources applies the same defaults to runtime sources (JSON never persisted timeout fields); a dead upstream can no longer hold a concurrency slot forever
- provider.go: split non-streaming client{Timeout} vs stream client{} sharing a Transport with ResponseHeaderTimeout, so long SSE bodies are not cut by client.Timeout; ChatStream uses doRawStream
- plan.md: mark P4-4/5/6 done, record P4-7/8 (tier-order, audit export, UI key view, zen upstream diagnosis)
- online verified: user key -> /api/sources 403 (GET+POST), admin 200, AUTO stream/non-stream healthy
2026-08-11 15:22:20 +08:00
854b3e2e37 fix: AUTO tier order ascending + stats/CSV export completeness + UI key-view toggle & exit affordance
- scheduler: BuildChain sorts tiers ascending so tier 1 (highest priority) is tried first; previously descending inverted the chain (P10-1..P10-5)
- stats/api: handleStatsAPI limit→20000; CSV reads full audit via new Stats.AuditRecords (includes *.old rotation); key_names mapped by keyID() masked key; non-admin filter and keys-csv use masked keys
- server: ring buffer NewStats(10000)
- ui: dash-row card scroll area moved to table container (fix overflow below card); key view toggle (re-click same key returns to global) + kpi-exit affordance; rec-exit span
- chat: chain-failure record now surfaces first failed tier; AUTO comment sync
2026-08-11 13:36:23 +08:00
63c2b13b4b fix: WebUI 密钥用量导出 CSV + 修复 Stats API key 过滤 bug
- api.go: 非 admin 用户过滤改用完整 key;keyNames 映射用完整 key;keys-csv 导出支持 key 查询参数过滤,安全类型断言
- stats.go: 新增 StatsRow 类型和 rows() 函数供导出使用
- server.go: handleStatusAPI 返回当前用户 key(已存在逻辑)
- index.html: 密钥用量卡片右上角添加导出 CSV 按钮(与请求记录一致)
2026-08-10 14:52:30 +08:00
d48b993010 feat(auto): AUTO-only priority chain with tiered slots + live source probing; CSV export w/ key names; audit persistence; fix prompt token accounting & deepseek thinking 2026-08-10 00:10:19 +08:00
3408c9cb1f feat(keys): role-based gateway keys with admin management UI and per-user model scope 2026-08-09 10:01:40 +08:00
dec03238dd feat(stats): gateway request stats/audit API + polish scratch sort animations (link chain, reset, cleanup) 2026-08-08 13:34:40 +08:00
f7f76e097d feat: ModelRouter — unified OpenAI-compatible multi-source LLM gateway
- Lua adapters per upstream (transform_request/response/stream_chunk, build_headers signing hooks)
- AUTO priority routing with per-model kind (chat/image), explicit source/model routing
- Per-source concurrency caps with queueing, exponential backoff, AUTO failover
- OpenAI-compatible API: chat completions, SSE streaming, image generations, models
- Gateway key auth, web UI for adapter/source management, runtime persistence
- e2e test running the real binary against mocked upstreams
2026-08-05 15:25:47 +08:00