mirror of
https://gitcode.com/JianFeeeee/ModelRouter.git
synced 2026-10-03 23:54:06 +00:00
docs: 补齐 per-key 配额的运维视角文档
代码回流 main 时配额小节已随行,但本轮新增的三项认知此前只存在于 commit message 与判据注释里,运维查不到: 1. **配额拒绝 vs 容量拒绝是两种东西**。配额在入口检查、不占上游槽位, 是廉价拒绝(实测 19–21ms,429 + Retry-After);容量不足要等满 busyWait 才 503(约 2.6s)。客户端据此可以区分「等窗口重置」与 「等上游腾容量」——前者只需耐心,后者通常该降并发或换源。 附实测对照表(容量 8、0.6s/请求、100 并发)。 2. **配额的开销**。每请求检查 149ns(配了配额)/ 42.6ns(未配配额, 不碰桶)/ 37ns(admin);窗口查询按窗口长度扫描而非扫全量保留 (1h 49ns、24h 55ns、30d 3.9us);生产形态内存 3.7MB。 明确写出「未配配额的 key 几乎不付代价」,运维可放心多建 key。 3. **按源 pin 的桶是惰性创建的,以及它的代价**。无条件维护会让 20 密钥 × 8 模型 × 3 源多占 18MB,所以只有真被 pin 查询过才维护; 代价是配置 pinned 配额之前的历史用量无法事后按源拆分,首个窗口 可能少算 —— 这是个会让排障困惑的行为,必须写出来。 同时更新 WebUI 密钥页说明:卡片配额徽标、「配额」编辑按钮、创建表单 可配预算(admin 自动禁用)、「我的密钥」页展示本 key 预算。 README.md / README_EN.md 同步。
This commit is contained in:
35
README.md
35
README.md
@ -216,6 +216,36 @@ keys:
|
|||||||
- `PUT /api/keys/{key}` 的配额字段是可选的:省略 = 保留原值,显式 `0` = 解除限制。
|
- `PUT /api/keys/{key}` 的配额字段是可选的:省略 = 保留原值,显式 `0` = 解除限制。
|
||||||
只改模型范围不会清空已配置的预算。
|
只改模型范围不会清空已配置的预算。
|
||||||
|
|
||||||
|
##### 配额拒绝 vs 容量拒绝:两种「拒绝」含义不同
|
||||||
|
|
||||||
|
配额在**入口**检查,不占用任何上游槽位,因此是廉价拒绝;容量不足则要走
|
||||||
|
`busyWait` 有界等待后才返回 503。实测(容量 8、0.6s/请求、100 并发):
|
||||||
|
|
||||||
|
| 情况 | 状态码 | 延迟 | 是否打到上游 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| 配额用尽(`req_quota` 50 → 100 并发) | `429 rate_limit_exceeded` + `Retry-After` | **19–21ms** | 否 |
|
||||||
|
| 容量打满(同一拓扑,未限配额) | `503 upstream_error` | ~2.6s(`busyWait`) | 超出部分否 |
|
||||||
|
|
||||||
|
客户端因此可以区分「等窗口重置」与「等上游腾出容量」:前者只需等待,后者
|
||||||
|
通常该降并发或换源。
|
||||||
|
|
||||||
|
##### 配额的开销
|
||||||
|
|
||||||
|
配额桶按 (密钥, 模型, 整点小时) 分桶保留 40 天,实测(AMD 7840HS):
|
||||||
|
|
||||||
|
- 每请求配额检查:**149ns**(配了配额)/ **42.6ns**(未配配额,只查密钥记录,
|
||||||
|
不碰桶)/ **37ns**(admin 密钥直接返回)—— 均 **0 分配**。
|
||||||
|
未配配额的密钥几乎不付代价,可放心大量创建。
|
||||||
|
- 记录一条请求:283ns、3 分配(与引入配额前相同,分配来自 ring buffer)。
|
||||||
|
- 窗口查询按窗口长度而非保留总量扫描:1h 窗口 49ns、24h 窗口 55ns、
|
||||||
|
30d 窗口 3.9μs。
|
||||||
|
- 内存:生产形态(7 密钥 × 8 模型 × 2 源 × 满 40 天 retention)约 **3.7MB**。
|
||||||
|
按源 pin 的 `source::model` 桶**惰性创建**——只有当某条配额真的 pin 了
|
||||||
|
某个源时才维护,否则每条记录多写一份桶,在 20 密钥 × 8 模型 × 3 源下会
|
||||||
|
多占 18MB。
|
||||||
|
代价:配置 pinned 配额之前的历史用量无法事后按源拆分,所以 **pinned 配额的
|
||||||
|
首个窗口可能少算**。
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# 配额耗尽时客户端看到
|
# 配额耗尽时客户端看到
|
||||||
HTTP/1.1 429 Too Many Requests
|
HTTP/1.1 429 Too Many Requests
|
||||||
@ -389,7 +419,10 @@ Environment=MALLOC_ARENA_MAX=2
|
|||||||
支持按时间范围导出 CSV;点击模型可生成 pin 到该模型的连接配置。
|
支持按时间范围导出 CSV;点击模型可生成 pin 到该模型的连接配置。
|
||||||
- **对话页**:流式/非流式调试。
|
- **对话页**:流式/非流式调试。
|
||||||
- **密钥页**:创建/编辑网关 key,为每个 key 配模型范围(模型 + 源 + token 配额 + 周期),
|
- **密钥页**:创建/编辑网关 key,为每个 key 配模型范围(模型 + 源 + token 配额 + 周期),
|
||||||
管理员管理全部 key,用户只看到自己的 key。
|
管理员管理全部 key,用户只看到自己的 key。key 卡片头部显示整钥配额徽标
|
||||||
|
(如 `250.0K·6h` / `77×·6h`),「配额」按钮编辑总 token / 请求数与重置周期;
|
||||||
|
创建 key 时可直接配预算(选 admin 角色时该组输入自动禁用,因为 admin 永不受限)。
|
||||||
|
「我的密钥」页对用户展示本 key 的预算。
|
||||||
- **优先级页**:拖拽积木配置 AUTO 链档位。
|
- **优先级页**:拖拽积木配置 AUTO 链档位。
|
||||||
- **源页**:在线增删改上游源(API key 等敏感字段加密落盘)。
|
- **源页**:在线增删改上游源(API key 等敏感字段加密落盘)。
|
||||||
- **适配器页**:上传 / 删除 Lua 适配器脚本。
|
- **适配器页**:上传 / 删除 Lua 适配器脚本。
|
||||||
|
|||||||
44
README_EN.md
44
README_EN.md
@ -175,7 +175,11 @@ under the `keys` field of the runtime file (encrypted at rest):
|
|||||||
delete the seed key.
|
delete the seed key.
|
||||||
- The WebUI **Keys page** creates/deletes keys. Each key has a role (`admin`
|
- The WebUI **Keys page** creates/deletes keys. Each key has a role (`admin`
|
||||||
manages everything, `user` sees only its own key) and an optional **model
|
manages everything, `user` sees only its own key) and an optional **model
|
||||||
scope** (model + source + token quota + reset period).
|
scope** (model + source + token quota + reset period). Key cards show the
|
||||||
|
key-wide caps as a badge (e.g. `250.0K·6h` / `77×·6h`); a **Quota** button
|
||||||
|
edits the total token / request budget and its reset period, and the create
|
||||||
|
form takes a budget too (those fields disable themselves for `admin`, which
|
||||||
|
is never capped). The "My key" view shows a user its own budget.
|
||||||
- Clients authenticate with any authorized key's plaintext as
|
- Clients authenticate with any authorized key's plaintext as
|
||||||
`Authorization: Bearer <key>`.
|
`Authorization: Bearer <key>`.
|
||||||
- Deleting a key removes it from the store immediately.
|
- Deleting a key removes it from the store immediately.
|
||||||
@ -222,6 +226,44 @@ keys:
|
|||||||
the stored caps, sending `0` explicitly lifts a cap. Editing only the model
|
the stored caps, sending `0` explicitly lifts a cap. Editing only the model
|
||||||
scope never clears a budget that was already set.
|
scope never clears a budget that was already set.
|
||||||
|
|
||||||
|
##### Quota rejection vs capacity rejection
|
||||||
|
|
||||||
|
A quota is checked at the entry point and occupies no upstream slot, so it is a
|
||||||
|
cheap rejection. Running out of capacity instead waits out the bounded
|
||||||
|
`busyWait` before answering 503. Measured (capacity 8, 0.6 s/request, 100
|
||||||
|
concurrent clients):
|
||||||
|
|
||||||
|
| Case | Status | Latency | Reaches upstream |
|
||||||
|
|---|---|---|---|
|
||||||
|
| Quota spent (`req_quota` 50, 100 concurrent) | `429 rate_limit_exceeded` + `Retry-After` | **19–21 ms** | no |
|
||||||
|
| Capacity exhausted (same topology, no quota) | `503 upstream_error` | ~2.6 s (`busyWait`) | overflow: no |
|
||||||
|
|
||||||
|
So a client can tell "wait for the window to roll over" apart from "wait for
|
||||||
|
upstream capacity" — the first only needs patience, the second usually means
|
||||||
|
back off concurrency or switch sources.
|
||||||
|
|
||||||
|
##### What the quota costs
|
||||||
|
|
||||||
|
Buckets are kept per (key, model, whole unix hour) for 40 days. Measured on an
|
||||||
|
AMD 7840HS:
|
||||||
|
|
||||||
|
- Quota check per request: **149 ns** (caps set) / **42.6 ns** (no caps — it
|
||||||
|
only looks up the key record and never touches a bucket) / **37 ns** (admin
|
||||||
|
key returns immediately) — all **0 allocations**. Keys without caps cost
|
||||||
|
almost nothing, so creating many of them is safe.
|
||||||
|
- Recording one request: 283 ns, 3 allocations (unchanged from before this
|
||||||
|
feature; the allocations come from the record ring buffer).
|
||||||
|
- A window query scans the window, not the whole retention: 49 ns for 1 h,
|
||||||
|
55 ns for 24 h, 3.9 µs for 30 d.
|
||||||
|
- Memory: the production shape (7 keys x 8 models x 2 sources at full 40-day
|
||||||
|
retention) costs about **3.7 MB**. The `source::model` bucket used by a
|
||||||
|
source-pinned quota is created **lazily** — it is maintained only once some
|
||||||
|
quota actually pins that source, because writing it on every record would
|
||||||
|
cost an extra 18 MB at 20 keys x 8 models x 3 sources.
|
||||||
|
The trade-off: usage recorded before a pinned quota was configured cannot be
|
||||||
|
split by source afterwards, so **a pinned quota may under-count its first
|
||||||
|
window**.
|
||||||
|
|
||||||
```
|
```
|
||||||
HTTP/1.1 429 Too Many Requests
|
HTTP/1.1 429 Too Many Requests
|
||||||
Retry-After: 2100
|
Retry-After: 2100
|
||||||
|
|||||||
Reference in New Issue
Block a user