diff --git a/README.md b/README.md index 985af10..274cfa6 100644 --- a/README.md +++ b/README.md @@ -216,6 +216,36 @@ keys: - `PUT /api/keys/{key}` 的配额字段是可选的:省略 = 保留原值,显式 `0` = 解除限制。 只改模型范围不会清空已配置的预算。 +##### 配额拒绝 vs 容量拒绝:两种「拒绝」含义不同 + +配额在**入口**检查,不占用任何上游槽位,因此是廉价拒绝;容量不足则要走 +`busyWait` 有界等待后才返回 503。实测(容量 8、0.6s/请求、100 并发): + +| 情况 | 状态码 | 延迟 | 是否打到上游 | +|---|---|---|---| +| 配额用尽(`req_quota` 50 → 100 并发) | `429 rate_limit_exceeded` + `Retry-After` | **19–21ms** | 否 | +| 容量打满(同一拓扑,未限配额) | `503 upstream_error` | ~2.6s(`busyWait`) | 超出部分否 | + +客户端因此可以区分「等窗口重置」与「等上游腾出容量」:前者只需等待,后者 +通常该降并发或换源。 + +##### 配额的开销 + +配额桶按 (密钥, 模型, 整点小时) 分桶保留 40 天,实测(AMD 7840HS): + +- 每请求配额检查:**149ns**(配了配额)/ **42.6ns**(未配配额,只查密钥记录, + 不碰桶)/ **37ns**(admin 密钥直接返回)—— 均 **0 分配**。 + 未配配额的密钥几乎不付代价,可放心大量创建。 +- 记录一条请求:283ns、3 分配(与引入配额前相同,分配来自 ring buffer)。 +- 窗口查询按窗口长度而非保留总量扫描:1h 窗口 49ns、24h 窗口 55ns、 + 30d 窗口 3.9μs。 +- 内存:生产形态(7 密钥 × 8 模型 × 2 源 × 满 40 天 retention)约 **3.7MB**。 + 按源 pin 的 `source::model` 桶**惰性创建**——只有当某条配额真的 pin 了 + 某个源时才维护,否则每条记录多写一份桶,在 20 密钥 × 8 模型 × 3 源下会 + 多占 18MB。 + 代价:配置 pinned 配额之前的历史用量无法事后按源拆分,所以 **pinned 配额的 + 首个窗口可能少算**。 + ```bash # 配额耗尽时客户端看到 HTTP/1.1 429 Too Many Requests @@ -389,7 +419,10 @@ Environment=MALLOC_ARENA_MAX=2 支持按时间范围导出 CSV;点击模型可生成 pin 到该模型的连接配置。 - **对话页**:流式/非流式调试。 - **密钥页**:创建/编辑网关 key,为每个 key 配模型范围(模型 + 源 + token 配额 + 周期), - 管理员管理全部 key,用户只看到自己的 key。 + 管理员管理全部 key,用户只看到自己的 key。key 卡片头部显示整钥配额徽标 + (如 `250.0K·6h` / `77×·6h`),「配额」按钮编辑总 token / 请求数与重置周期; + 创建 key 时可直接配预算(选 admin 角色时该组输入自动禁用,因为 admin 永不受限)。 + 「我的密钥」页对用户展示本 key 的预算。 - **优先级页**:拖拽积木配置 AUTO 链档位。 - **源页**:在线增删改上游源(API key 等敏感字段加密落盘)。 - **适配器页**:上传 / 删除 Lua 适配器脚本。 diff --git a/README_EN.md b/README_EN.md index 6f1af61..de5078b 100644 --- a/README_EN.md +++ b/README_EN.md @@ -175,7 +175,11 @@ under the `keys` field of the runtime file (encrypted at rest): delete the seed key. - The WebUI **Keys page** creates/deletes keys. Each key has a role (`admin` manages everything, `user` sees only its own key) and an optional **model - scope** (model + source + token quota + reset period). + scope** (model + source + token quota + reset period). Key cards show the + key-wide caps as a badge (e.g. `250.0K·6h` / `77×·6h`); a **Quota** button + edits the total token / request budget and its reset period, and the create + form takes a budget too (those fields disable themselves for `admin`, which + is never capped). The "My key" view shows a user its own budget. - Clients authenticate with any authorized key's plaintext as `Authorization: Bearer `. - Deleting a key removes it from the store immediately. @@ -222,6 +226,44 @@ keys: the stored caps, sending `0` explicitly lifts a cap. Editing only the model scope never clears a budget that was already set. +##### Quota rejection vs capacity rejection + +A quota is checked at the entry point and occupies no upstream slot, so it is a +cheap rejection. Running out of capacity instead waits out the bounded +`busyWait` before answering 503. Measured (capacity 8, 0.6 s/request, 100 +concurrent clients): + +| Case | Status | Latency | Reaches upstream | +|---|---|---|---| +| Quota spent (`req_quota` 50, 100 concurrent) | `429 rate_limit_exceeded` + `Retry-After` | **19–21 ms** | no | +| Capacity exhausted (same topology, no quota) | `503 upstream_error` | ~2.6 s (`busyWait`) | overflow: no | + +So a client can tell "wait for the window to roll over" apart from "wait for +upstream capacity" — the first only needs patience, the second usually means +back off concurrency or switch sources. + +##### What the quota costs + +Buckets are kept per (key, model, whole unix hour) for 40 days. Measured on an +AMD 7840HS: + +- Quota check per request: **149 ns** (caps set) / **42.6 ns** (no caps — it + only looks up the key record and never touches a bucket) / **37 ns** (admin + key returns immediately) — all **0 allocations**. Keys without caps cost + almost nothing, so creating many of them is safe. +- Recording one request: 283 ns, 3 allocations (unchanged from before this + feature; the allocations come from the record ring buffer). +- A window query scans the window, not the whole retention: 49 ns for 1 h, + 55 ns for 24 h, 3.9 µs for 30 d. +- Memory: the production shape (7 keys x 8 models x 2 sources at full 40-day + retention) costs about **3.7 MB**. The `source::model` bucket used by a + source-pinned quota is created **lazily** — it is maintained only once some + quota actually pins that source, because writing it on every record would + cost an extra 18 MB at 20 keys x 8 models x 3 sources. + The trade-off: usage recorded before a pinned quota was configured cannot be + split by source afterwards, so **a pinned quota may under-count its first + window**. + ``` HTTP/1.1 429 Too Many Requests Retry-After: 2100