docs: 补齐 per-key 配额的运维视角文档

代码回流 main 时配额小节已随行,但本轮新增的三项认知此前只存在于
commit message 与判据注释里,运维查不到:

1. **配额拒绝 vs 容量拒绝是两种东西**。配额在入口检查、不占上游槽位,
   是廉价拒绝(实测 19–21ms,429 + Retry-After);容量不足要等满
   busyWait 才 503(约 2.6s)。客户端据此可以区分「等窗口重置」与
   「等上游腾容量」——前者只需耐心,后者通常该降并发或换源。
   附实测对照表(容量 8、0.6s/请求、100 并发)。

2. **配额的开销**。每请求检查 149ns(配了配额)/ 42.6ns(未配配额,
   不碰桶)/ 37ns(admin);窗口查询按窗口长度扫描而非扫全量保留
   (1h 49ns、24h 55ns、30d 3.9us);生产形态内存 3.7MB。
   明确写出「未配配额的 key 几乎不付代价」,运维可放心多建 key。

3. **按源 pin 的桶是惰性创建的,以及它的代价**。无条件维护会让
   20 密钥 × 8 模型 × 3 源多占 18MB,所以只有真被 pin 查询过才维护;
   代价是配置 pinned 配额之前的历史用量无法事后按源拆分,首个窗口
   可能少算 —— 这是个会让排障困惑的行为,必须写出来。

同时更新 WebUI 密钥页说明:卡片配额徽标、「配额」编辑按钮、创建表单
可配预算(admin 自动禁用)、「我的密钥」页展示本 key 预算。

README.md / README_EN.md 同步。
This commit is contained in:
JianFeeeee
2026-09-27 18:46:39 +08:00
parent 652842783f
commit cc5e725226
2 changed files with 77 additions and 2 deletions

View File

@ -175,7 +175,11 @@ under the `keys` field of the runtime file (encrypted at rest):
delete the seed key.
- The WebUI **Keys page** creates/deletes keys. Each key has a role (`admin`
manages everything, `user` sees only its own key) and an optional **model
scope** (model + source + token quota + reset period).
scope** (model + source + token quota + reset period). Key cards show the
key-wide caps as a badge (e.g. `250.0K·6h` / `77×·6h`); a **Quota** button
edits the total token / request budget and its reset period, and the create
form takes a budget too (those fields disable themselves for `admin`, which
is never capped). The "My key" view shows a user its own budget.
- Clients authenticate with any authorized key's plaintext as
`Authorization: Bearer <key>`.
- Deleting a key removes it from the store immediately.
@ -222,6 +226,44 @@ keys:
the stored caps, sending `0` explicitly lifts a cap. Editing only the model
scope never clears a budget that was already set.
##### Quota rejection vs capacity rejection
A quota is checked at the entry point and occupies no upstream slot, so it is a
cheap rejection. Running out of capacity instead waits out the bounded
`busyWait` before answering 503. Measured (capacity 8, 0.6 s/request, 100
concurrent clients):
| Case | Status | Latency | Reaches upstream |
|---|---|---|---|
| Quota spent (`req_quota` 50, 100 concurrent) | `429 rate_limit_exceeded` + `Retry-After` | **19–21 ms** | no |
| Capacity exhausted (same topology, no quota) | `503 upstream_error` | ~2.6 s (`busyWait`) | overflow: no |
So a client can tell "wait for the window to roll over" apart from "wait for
upstream capacity" — the first only needs patience, the second usually means
back off concurrency or switch sources.
##### What the quota costs
Buckets are kept per (key, model, whole unix hour) for 40 days. Measured on an
AMD 7840HS:
- Quota check per request: **149 ns** (caps set) / **42.6 ns** (no caps — it
only looks up the key record and never touches a bucket) / **37 ns** (admin
key returns immediately) — all **0 allocations**. Keys without caps cost
almost nothing, so creating many of them is safe.
- Recording one request: 283 ns, 3 allocations (unchanged from before this
feature; the allocations come from the record ring buffer).
- A window query scans the window, not the whole retention: 49 ns for 1 h,
55 ns for 24 h, 3.9 µs for 30 d.
- Memory: the production shape (7 keys x 8 models x 2 sources at full 40-day
retention) costs about **3.7 MB**. The `source::model` bucket used by a
source-pinned quota is created **lazily** — it is maintained only once some
quota actually pins that source, because writing it on every record would
cost an extra 18 MB at 20 keys x 8 models x 3 sources.
The trade-off: usage recorded before a pinned quota was configured cannot be
split by source afterwards, so **a pinned quota may under-count its first
window**.
```
HTTP/1.1 429 Too Many Requests
Retry-After: 2100