mirror of
https://gitcode.com/JianFeeeee/ModelRouter.git
synced 2026-10-03 23:54:06 +00:00
refactor(quota): 配额改为按模型,删除整钥总配额
用户明确要求:配额应当是密钥对应的**每个模型的单独配额**,而非整体配额。 ## 语义变更 删除 GWKey.TokenQuota / ReqQuota / Period / Hours(整钥总额)。 ModelScope 新增 ReqQuota —— 请求数配额下沉到每条模型范围。 现在:每条 models[] 各自带 token 配额 + 请求数配额 + 重置周期, 彼此独立。一个模型用满只影响该模型。 ★ 为什么不保留整钥总额:它会让「把 A 模型的额度挪给 B」变成一次全局 重分配;按模型独立计费则每个模型各自可控,运维能直接看出哪个模型在吃预算。 ## 连带改动 - checkQuota 合并 key 级与 scope 级判定;checkKeyQuotaRetry 整体删除 (顺带修掉上轮遗留的双重判定:入口不再先判空再重算) - core:CreateKeyWithQuota / UpdateKeyWithQuota / ApplyQuota 全部删除, 改由 ValidateScopeQuotas 校验每条 scope 的配额 - admin key:scope 上的配额不强制(admin 的 scope 仍限制模型范围, 但不强制配额)—— 否则管理员会把自己锁在门外 - /api/v1/keys 不再回显 key 级配额字段(scope 里已含) - WebUI:删除整钥配额徽标 / 「配额」按钮 / 创建表单的配额组 / putScope 的整钥回传;模型砖块与范围编辑器新增「请求数配额」输入, 徽标显示 `1.0K 77×·1h`(未设配额显示 ∞) ## 判据 - TestOneModelsQuotaDoesNotBlockAnother 是本次核心保证。 ★ 它第一版是**假判据**:m2 从不消耗,key-wide 计数器与 m1 自己的计数器 读数恰好相同,退回 key-wide 仍通过。变异测试抓到后改为「先用 m2 花掉 远超 m1 配额的量,再验证 m1 仍可用」—— 这样两种设计才可区分。 - TestUncappedModelNeverBlocked / TestAdminKeyScopesAreNotEnforced 新增 - UI 契约判据重写:整钥配额界面必须彻底消失(13 个符号)、 scope 编辑器必须往返 req_quota、putScope 只发 scope 列表 - 错误消息点名具体模型(TestKeyAPIRejectionNamesTheModel) - 3/3 变异全被抓 实测(真实进程 + 浏览器):m2 配额 500000 连打 25 次全成功, m1 配额 1000 立即 429「token quota exceeded for "m1" (4315/1000)」, 此后 m2/m3 仍 200。UI:整钥配额元素全为 0,砖块各显配额, 编辑器预填/保存正确,零 JS 异常。
This commit is contained in:
59
README.md
59
README.md
@ -24,7 +24,7 @@
|
|||||||
### 强大的多租户调度能力
|
### 强大的多租户调度能力
|
||||||
|
|
||||||
- **多密钥多租户**:支持无限密钥,每个密钥独立角色、模型范围、Token 配额、重置周期
|
- **多密钥多租户**:支持无限密钥,每个密钥独立角色、模型范围、Token 配额、重置周期
|
||||||
- **密钥用量配额**:每把 key 单独配总 token 配额 + 请求数配额与重置周期(小时/周/月/自定义 N 小时),跨模型共享预算;耗尽返 429 + `Retry-After` 可自动恢复,admin key 永不受限
|
- **按模型配额**:每把 key 的每个模型单独配 token + 请求数配额与重置周期(小时/周/月/自定义 N 小时);一个模型用满只影响该模型,同 key 其它模型照常;耗尽返 429 + `Retry-After` 并点名模型,admin key 永不受限
|
||||||
- **AUTO 智能调度**:基于优先级档位的分级调度,同优先级源自动轮询负载均衡,故障自动毫秒级故障转移
|
- **AUTO 智能调度**:基于优先级档位的分级调度,同优先级源自动轮询负载均衡,故障自动毫秒级故障转移
|
||||||
- **自愈冷却**:冷却上限 5 分钟,过半后放行 1 个探测请求,上游/额度恢复即刻回归轮询,无需等满冷却窗口
|
- **自愈冷却**:冷却上限 5 分钟,过半后放行 1 个探测请求,上游/额度恢复即刻回归轮询,无需等满冷却窗口
|
||||||
- **Token 配额管理**:精确到模型级别的 Token 配额控制,支持小时/周/月/自定义小时周期自动重置
|
- **Token 配额管理**:精确到模型级别的 Token 配额控制,支持小时/周/月/自定义小时周期自动重置
|
||||||
@ -181,40 +181,44 @@ sources:
|
|||||||
|
|
||||||
#### 密钥用量配额
|
#### 密钥用量配额
|
||||||
|
|
||||||
每个密钥可单独限制用量与用量重置周期,两级配额同时生效:
|
**配额按模型单独设置**:每个密钥的 `models[]` 里,每一条模型范围各自带一份
|
||||||
|
token 配额与请求数配额。一个模型用满只影响该模型,同一密钥的其它模型照常工作。
|
||||||
|
|
||||||
```yaml
|
```yaml
|
||||||
keys:
|
keys:
|
||||||
- key: sk-gw-<hex>
|
- key: sk-gw-<hex>
|
||||||
role: user
|
role: user
|
||||||
name: agent-alice
|
name: agent-alice
|
||||||
# ---- 整钥配额(跳模型)----
|
|
||||||
token_quota: 5000000 # 本周期内这把 key 的总 token 预算,0 = 无限
|
|
||||||
req_quota: 20000 # 本周期内的请求次数,0 = 无限
|
|
||||||
period: nhour # "" | hour | week | month | nhour
|
|
||||||
hours: 6 # 仅 nhour:每 6 小时重置
|
|
||||||
# ---- 模型范围(可选,逐模型配额)----
|
|
||||||
models:
|
models:
|
||||||
- model: m1
|
- model: deepseek-v4-flash
|
||||||
token_quota: 1000000 # 本周期内该模型(该 key)的 token 预算
|
token_quota: 1000000 # 本周期内该 key 用这个模型的 token 预算
|
||||||
period: hour
|
req_quota: 20000 # 本周期内的请求次数
|
||||||
|
period: nhour # "" | hour | week | month | nhour
|
||||||
|
hours: 6 # 仅 nhour
|
||||||
- model: AUTO
|
- model: AUTO
|
||||||
|
token_quota: 5000000 # AUTO 也是一条独立配额
|
||||||
|
period: day # ← 这种写法会被拒绝(词表只有 hour/week/month/nhour)
|
||||||
|
- model: kimi-k3 # 未列配额 = 无限
|
||||||
```
|
```
|
||||||
|
|
||||||
|
- **没有「整钥总配额」**:这是刻意的设计。整钥总额会让「把 A 模型的额度挪给
|
||||||
|
B 模型」变成一次全局重分配;按模型独立计费则每个模型各自可控,运维可以
|
||||||
|
看出哪个模型吃掉了预算。
|
||||||
- `period` 词表:空 = 永不过期(累计总量),`hour` / `week` / `month` = 固定窗口,
|
- `period` 词表:空 = 永不过期(累计总量),`hour` / `week` / `month` = 固定窗口,
|
||||||
`nhour` + `hours` = 自定义小时数。**拼错的周期在写入时就被拒**,不会静默变成
|
`nhour` + `hours` = 自定义小时数。**拼错的周期在写入时就被拒**,不会静默变成
|
||||||
永不过期。
|
永不过期。
|
||||||
- 整钥配额跨该 key 所有模型共享一份预算;`models[]` 里的配额则是逐模型独立计数。
|
- 配额严格按密钥隔离,且**同一密钥内按模型隔离**:A 密钥用满 `m1` 不会消耗
|
||||||
两者都按 key 隔离,A key 的用量不会消耗 B key 的额度。
|
B 密钥的额度,同一密钥的 `m2` 也不受影响。
|
||||||
- 配额统计含聊天、流式、生图,跨重启从审计日志回放(保留 40 天,覆盖最长的
|
- 配额统计含聊天、流式、生图,跨重启从审计日志回放(保留 40 天,覆盖最长的
|
||||||
month 窗口)。
|
month 窗口)。
|
||||||
- 配额耗尽返回 **429 + `Retry-After`**(`rate_limit_exceeded`),客户端可等窗口
|
- 配额耗尽返回 **429 + `Retry-After`**(`rate_limit_exceeded`),消息里点名是哪个
|
||||||
重置后自动恢复;模型越权才是 403。**admin 密钥永不受配额限制**,
|
模型用满了,客户端可等窗口重置后自动恢复;模型越权才是 403。
|
||||||
避免把管理员锁在门外。
|
**admin 密钥永不受配额限制**(其 scope 上的配额也不强制),避免把管理员
|
||||||
|
锁在门外。
|
||||||
- 窗口用量按整点小时分桶统计,实际释放比配置窗口最多晚 1 小时(配额宁可晚释放
|
- 窗口用量按整点小时分桶统计,实际释放比配置窗口最多晚 1 小时(配额宁可晚释放
|
||||||
也不超发)。
|
也不超发)。
|
||||||
- `PUT /api/keys/{key}` 的配额字段是可选的:省略 = 保留原值,显式 `0` = 解除限制。
|
- `PUT /api/keys/{key}` 提交 `models` 即同时提交它们的配额(配额就是 scope 的一部分,
|
||||||
只改模型范围不会清空已配置的预算。
|
不存在会与模型列表脱节的第二份预算)。显式 `0` = 解除该模型的限制。
|
||||||
|
|
||||||
##### 配额拒绝 vs 容量拒绝:两种「拒绝」含义不同
|
##### 配额拒绝 vs 容量拒绝:两种「拒绝」含义不同
|
||||||
|
|
||||||
@ -233,12 +237,9 @@ keys:
|
|||||||
|
|
||||||
配额桶按 (密钥, 模型, 整点小时) 分桶保留 40 天,实测(AMD 7840HS):
|
配额桶按 (密钥, 模型, 整点小时) 分桶保留 40 天,实测(AMD 7840HS):
|
||||||
|
|
||||||
- 每请求配额检查:**149ns**(配了配额)/ **42.6ns**(未配配额,只查密钥记录,
|
|
||||||
不碰桶)/ **37ns**(admin 密钥直接返回)—— 均 **0 分配**。
|
|
||||||
未配配额的密钥几乎不付代价,可放心大量创建。
|
|
||||||
- 记录一条请求:283ns、3 分配(与引入配额前相同,分配来自 ring buffer)。
|
|
||||||
- 窗口查询按窗口长度而非保留总量扫描:1h 窗口 49ns、24h 窗口 55ns、
|
- 窗口查询按窗口长度而非保留总量扫描:1h 窗口 49ns、24h 窗口 55ns、
|
||||||
30d 窗口 3.9μs。
|
30d 窗口 3.9μs(此前全扫保留总量,960 桶时 5.9μs)。
|
||||||
|
- 记录一条请求:283ns、3 分配(与引入配额前相同,分配来自 ring buffer)。
|
||||||
- 内存:生产形态(7 密钥 × 8 模型 × 2 源 × 满 40 天 retention)约 **3.7MB**。
|
- 内存:生产形态(7 密钥 × 8 模型 × 2 源 × 满 40 天 retention)约 **3.7MB**。
|
||||||
按源 pin 的 `source::model` 桶**惰性创建**——只有当某条配额真的 pin 了
|
按源 pin 的 `source::model` 桶**惰性创建**——只有当某条配额真的 pin 了
|
||||||
某个源时才维护,否则每条记录多写一份桶,在 20 密钥 × 8 模型 × 3 源下会
|
某个源时才维护,否则每条记录多写一份桶,在 20 密钥 × 8 模型 × 3 源下会
|
||||||
@ -247,10 +248,10 @@ keys:
|
|||||||
首个窗口可能少算**。
|
首个窗口可能少算**。
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# 配额耗尽时客户端看到
|
# 配额耗尽时客户端看到(点名了具体模型)
|
||||||
HTTP/1.1 429 Too Many Requests
|
HTTP/1.1 429 Too Many Requests
|
||||||
Retry-After: 2100
|
Retry-After: 2100
|
||||||
{"error":{"type":"rate_limit_exceeded","message":"key token quota exceeded (5000000/5000000, resets every 6h)"}}
|
{"error":{"type":"rate_limit_exceeded","message":"token quota exceeded for \"deepseek-v4-flash\" (5000000/5000000)"}}
|
||||||
```
|
```
|
||||||
|
|
||||||
### 模型路由
|
### 模型路由
|
||||||
@ -419,10 +420,10 @@ Environment=MALLOC_ARENA_MAX=2
|
|||||||
支持按时间范围导出 CSV;点击模型可生成 pin 到该模型的连接配置。
|
支持按时间范围导出 CSV;点击模型可生成 pin 到该模型的连接配置。
|
||||||
- **对话页**:流式/非流式调试。
|
- **对话页**:流式/非流式调试。
|
||||||
- **密钥页**:创建/编辑网关 key,为每个 key 配模型范围(模型 + 源 + token 配额 + 周期),
|
- **密钥页**:创建/编辑网关 key,为每个 key 配模型范围(模型 + 源 + token 配额 + 周期),
|
||||||
管理员管理全部 key,用户只看到自己的 key。key 卡片头部显示整钥配额徽标
|
管理员管理全部 key,用户只看到自己的 key。每个模型砖块显示自己的配额徽标
|
||||||
(如 `250.0K·6h` / `77×·6h`),「配额」按钮编辑总 token / 请求数与重置周期;
|
(如 `1.0K 77×·1h`,未设配额显示 `∞`),点开可编辑该模型的 token 配额、
|
||||||
创建 key 时可直接配预算(选 admin 角色时该组输入自动禁用,因为 admin 永不受限)。
|
请求数配额与重置周期 —— 配额按模型独立生效,一个用满不影响同一 key 的其它模型。
|
||||||
「我的密钥」页对用户展示本 key 的预算。
|
「我的密钥」页对用户展示本 key 的模型范围与各自配额。
|
||||||
- **优先级页**:拖拽积木配置 AUTO 链档位。
|
- **优先级页**:拖拽积木配置 AUTO 链档位。
|
||||||
- **源页**:在线增删改上游源(API key 等敏感字段加密落盘)。
|
- **源页**:在线增删改上游源(API key 等敏感字段加密落盘)。
|
||||||
- **适配器页**:上传 / 删除 Lua 适配器脚本。
|
- **适配器页**:上传 / 删除 Lua 适配器脚本。
|
||||||
|
|||||||
73
README_EN.md
73
README_EN.md
@ -38,10 +38,11 @@ Extracted and independently evolved from the multi-source LLM adapter layer of
|
|||||||
`reasoning_content`, `tool_calls`, `usage`).
|
`reasoning_content`, `tool_calls`, `usage`).
|
||||||
- **Image generation**: `POST /v1/images/generations`, routed to models with
|
- **Image generation**: `POST /v1/images/generations`, routed to models with
|
||||||
`kind: image`.
|
`kind: image`.
|
||||||
- **Per-key usage quota**: each key carries its own token and request caps plus a
|
- **Per-model quota**: each key gives every model its own token and request caps
|
||||||
reset period (hour/week/month/custom N hours), shared across every model that
|
plus a reset period (hour/week/month/custom N hours). One model running out
|
||||||
key may use. Exhaustion answers 429 + `Retry-After` so a client resumes when
|
affects only that model — the key's other models keep working. Exhaustion
|
||||||
the window rolls over; admin keys are never capped.
|
answers 429 + `Retry-After` naming the model, so a client resumes when the
|
||||||
|
window rolls over; admin keys are never capped.
|
||||||
- **Multimodal**: `content` arrays (`image_url` etc.) pass through losslessly;
|
- **Multimodal**: `content` arrays (`image_url` etc.) pass through losslessly;
|
||||||
Anthropic/Gemini/Ollama are translated automatically.
|
Anthropic/Gemini/Ollama are translated automatically.
|
||||||
- **LuaJIT VM**: golua-binding LuaJIT; each adapter has its own VM + worker
|
- **LuaJIT VM**: golua-binding LuaJIT; each adapter has its own VM + worker
|
||||||
@ -175,56 +176,62 @@ under the `keys` field of the runtime file (encrypted at rest):
|
|||||||
delete the seed key.
|
delete the seed key.
|
||||||
- The WebUI **Keys page** creates/deletes keys. Each key has a role (`admin`
|
- The WebUI **Keys page** creates/deletes keys. Each key has a role (`admin`
|
||||||
manages everything, `user` sees only its own key) and an optional **model
|
manages everything, `user` sees only its own key) and an optional **model
|
||||||
scope** (model + source + token quota + reset period). Key cards show the
|
scope** (model + source + token quota + reset period). Each model brick shows
|
||||||
key-wide caps as a badge (e.g. `250.0K·6h` / `77×·6h`); a **Quota** button
|
its own budget badge (e.g. `1.0K 77×·1h`, `∞` when uncapped); clicking it
|
||||||
edits the total token / request budget and its reset period, and the create
|
edits that model's token quota, request quota and reset period. Quotas apply
|
||||||
form takes a budget too (those fields disable themselves for `admin`, which
|
per model, so one model running out never blocks the key's others. The
|
||||||
is never capped). The "My key" view shows a user its own budget.
|
"My key" view shows a user its model scopes and their budgets.
|
||||||
- Clients authenticate with any authorized key's plaintext as
|
- Clients authenticate with any authorized key's plaintext as
|
||||||
`Authorization: Bearer <key>`.
|
`Authorization: Bearer <key>`.
|
||||||
- Deleting a key removes it from the store immediately.
|
- Deleting a key removes it from the store immediately.
|
||||||
|
|
||||||
#### Per-key usage quota
|
#### Per-key usage quota
|
||||||
|
|
||||||
Each key can cap its own spend and reset period. Two levels apply at once:
|
**Quotas are per model.** Each key's `models[]` list gives every model its own
|
||||||
|
token budget and request budget. One model running out affects only that model
|
||||||
|
— the key's other models keep working.
|
||||||
|
|
||||||
```yaml
|
```yaml
|
||||||
keys:
|
keys:
|
||||||
- key: sk-gw-<hex>
|
- key: sk-gw-<hex>
|
||||||
role: user
|
role: user
|
||||||
name: agent-alice
|
name: agent-alice
|
||||||
# ---- key-wide (across every model) ----
|
|
||||||
token_quota: 5000000 # total token budget for this window, 0 = unlimited
|
|
||||||
req_quota: 20000 # requests per window, 0 = unlimited
|
|
||||||
period: nhour # "" | hour | week | month | nhour
|
|
||||||
hours: 6 # n-hour only: resets every 6 hours
|
|
||||||
# ---- per-model scope (optional) ----
|
|
||||||
models:
|
models:
|
||||||
- model: m1
|
- model: deepseek-v4-flash
|
||||||
token_quota: 1000000
|
token_quota: 1000000 # this key's token budget for this model
|
||||||
period: hour
|
req_quota: 20000 # requests within the window
|
||||||
|
period: nhour # "" | hour | week | month | nhour
|
||||||
|
hours: 6 # n-hour only
|
||||||
- model: AUTO
|
- model: AUTO
|
||||||
|
token_quota: 5000000 # AUTO is a quota entry like any other
|
||||||
|
period: hour
|
||||||
|
- model: kimi-k3 # no quota listed = unlimited
|
||||||
```
|
```
|
||||||
|
|
||||||
|
- **There is deliberately no key-wide total.** A key-wide cap would make
|
||||||
|
"move A's budget to B" a global reallocation; per-model budgets keep each
|
||||||
|
model independently controllable, so it stays visible which model is
|
||||||
|
actually consuming the spend.
|
||||||
- `period`: empty = never resets (lifetime total); `hour` / `week` / `month` =
|
- `period`: empty = never resets (lifetime total); `hour` / `week` / `month` =
|
||||||
fixed windows; `nhour` + `hours` = a custom hour count. **A misspelled
|
fixed windows; `nhour` + `hours` = a custom hour count. **A misspelled
|
||||||
period is rejected at write time** rather than silently becoming a
|
period is rejected at write time** rather than silently becoming a
|
||||||
never-resetting quota.
|
never-resetting quota.
|
||||||
- The key-wide cap is one budget shared by every model the key may use;
|
- Quotas are isolated per key *and* per model within a key: one key exhausting
|
||||||
quotas under `models[]` are counted per model. Both are isolated per key —
|
`m1` never draws on another key's budget, and never blocks the same key's
|
||||||
one key's traffic never drains another's budget.
|
`m2`.
|
||||||
- Usage counts chat, streaming and image requests, and survives a restart by
|
- Usage counts chat, streaming and image requests, and survives a restart by
|
||||||
replaying the audit log (40 days retained, covering the longest `month`
|
replaying the audit log (40 days retained, covering the longest `month`
|
||||||
window).
|
window).
|
||||||
- An exhausted quota returns **429 + `Retry-After`**
|
- An exhausted quota returns **429 + `Retry-After`**
|
||||||
(`rate_limit_exceeded`) so a client resumes when the window rolls over; a
|
(`rate_limit_exceeded`) and the message names the model that ran out, so a
|
||||||
model the key may not use stays 403. **Admin keys are never capped**, so a
|
client resumes when the window rolls over; a model the key may not use stays
|
||||||
cap can never lock the operator out.
|
403. **Admin keys are never capped** (quotas on their scopes are not
|
||||||
|
enforced either), so a cap can never lock the operator out.
|
||||||
- Buckets are whole unix hours, so a window frees up at most an hour late
|
- Buckets are whole unix hours, so a window frees up at most an hour late
|
||||||
(deliberately freeing late rather than overspending).
|
(deliberately freeing late rather than overspending).
|
||||||
- On `PUT /api/keys/{key}` the quota fields are optional: omitting them keeps
|
- `PUT /api/keys/{key}` submits quotas by submitting `models` — the caps are
|
||||||
the stored caps, sending `0` explicitly lifts a cap. Editing only the model
|
part of the scope, so there is no second budget that can drift out of sync
|
||||||
scope never clears a budget that was already set.
|
with the model list. An explicit `0` lifts that model's cap.
|
||||||
|
|
||||||
##### Quota rejection vs capacity rejection
|
##### Quota rejection vs capacity rejection
|
||||||
|
|
||||||
@ -247,14 +254,10 @@ back off concurrency or switch sources.
|
|||||||
Buckets are kept per (key, model, whole unix hour) for 40 days. Measured on an
|
Buckets are kept per (key, model, whole unix hour) for 40 days. Measured on an
|
||||||
AMD 7840HS:
|
AMD 7840HS:
|
||||||
|
|
||||||
- Quota check per request: **149 ns** (caps set) / **42.6 ns** (no caps — it
|
|
||||||
only looks up the key record and never touches a bucket) / **37 ns** (admin
|
|
||||||
key returns immediately) — all **0 allocations**. Keys without caps cost
|
|
||||||
almost nothing, so creating many of them is safe.
|
|
||||||
- Recording one request: 283 ns, 3 allocations (unchanged from before this
|
|
||||||
feature; the allocations come from the record ring buffer).
|
|
||||||
- A window query scans the window, not the whole retention: 49 ns for 1 h,
|
- A window query scans the window, not the whole retention: 49 ns for 1 h,
|
||||||
55 ns for 24 h, 3.9 µs for 30 d.
|
55 ns for 24 h, 3.9 µs for 30 d.
|
||||||
|
- Recording one request: 283 ns, 3 allocations (unchanged from before this
|
||||||
|
feature; the allocations come from the record ring buffer).
|
||||||
- Memory: the production shape (7 keys x 8 models x 2 sources at full 40-day
|
- Memory: the production shape (7 keys x 8 models x 2 sources at full 40-day
|
||||||
retention) costs about **3.7 MB**. The `source::model` bucket used by a
|
retention) costs about **3.7 MB**. The `source::model` bucket used by a
|
||||||
source-pinned quota is created **lazily** — it is maintained only once some
|
source-pinned quota is created **lazily** — it is maintained only once some
|
||||||
@ -267,7 +270,7 @@ AMD 7840HS:
|
|||||||
```
|
```
|
||||||
HTTP/1.1 429 Too Many Requests
|
HTTP/1.1 429 Too Many Requests
|
||||||
Retry-After: 2100
|
Retry-After: 2100
|
||||||
{"error":{"type":"rate_limit_exceeded","message":"key token quota exceeded (5000000/5000000, resets every 6h)"}}
|
{"error":{"type":"rate_limit_exceeded","message":"token quota exceeded for \"deepseek-v4-flash\" (5000000/5000000)"}}
|
||||||
```
|
```
|
||||||
|
|
||||||
### Model routing
|
### Model routing
|
||||||
|
|||||||
@ -375,21 +375,11 @@ type GWKey struct {
|
|||||||
Note string `yaml:"note,omitempty" json:"note,omitempty"`
|
Note string `yaml:"note,omitempty" json:"note,omitempty"`
|
||||||
CreatedAt int64 `yaml:"created_at,omitempty" json:"created_at,omitempty"`
|
CreatedAt int64 `yaml:"created_at,omitempty" json:"created_at,omitempty"`
|
||||||
Seed bool `yaml:"seed,omitempty" json:"seed,omitempty"` // true if migrated from config gateway_keys
|
Seed bool `yaml:"seed,omitempty" json:"seed,omitempty"` // true if migrated from config gateway_keys
|
||||||
// TokenQuota caps this key's TOTAL tokens across every model it may use.
|
|
||||||
// 0 = unlimited. Period/Hours define the reset window, exactly like
|
|
||||||
// ModelScope: "" never resets, "hour"/"week"/"month" fixed windows,
|
|
||||||
// "nhour" uses Hours.
|
|
||||||
TokenQuota int64 `yaml:"token_quota,omitempty" json:"token_quota,omitempty"`
|
|
||||||
Period string `yaml:"period,omitempty" json:"period,omitempty"`
|
|
||||||
Hours int64 `yaml:"hours,omitempty" json:"hours,omitempty"`
|
|
||||||
// ReqQuota caps the number of requests per reset window; 0 = unlimited.
|
|
||||||
// RPM covers short bursts; this covers sustained volume.
|
|
||||||
ReqQuota int64 `yaml:"req_quota,omitempty" json:"req_quota,omitempty"`
|
|
||||||
}
|
}
|
||||||
|
|
||||||
// KeyQuota is the set of key-wide caps accepted by the admin API. It is a
|
// KeyQuota is retained only to carry a scope entry's caps through the admin
|
||||||
// separate struct so a partial update can be expressed as a pointer (nil =
|
// API. Quotas are per model, never per key: there is deliberately no key-wide
|
||||||
// "leave the stored caps alone") instead of zero values meaning "clear".
|
// total, so exhausting one model's budget never blocks the others.
|
||||||
type KeyQuota struct {
|
type KeyQuota struct {
|
||||||
TokenQuota int64 `json:"token_quota"`
|
TokenQuota int64 `json:"token_quota"`
|
||||||
ReqQuota int64 `json:"req_quota"`
|
ReqQuota int64 `json:"req_quota"`
|
||||||
@ -397,14 +387,6 @@ type KeyQuota struct {
|
|||||||
Hours int64 `json:"hours"`
|
Hours int64 `json:"hours"`
|
||||||
}
|
}
|
||||||
|
|
||||||
// ApplyQuota writes the caps onto a key record.
|
|
||||||
func (k *GWKey) ApplyQuota(q KeyQuota) {
|
|
||||||
k.TokenQuota = q.TokenQuota
|
|
||||||
k.ReqQuota = q.ReqQuota
|
|
||||||
k.Period = q.Period
|
|
||||||
k.Hours = q.Hours
|
|
||||||
}
|
|
||||||
|
|
||||||
// NormalizeRole defaults an empty role to "user", so a key can never end up in
|
// NormalizeRole defaults an empty role to "user", so a key can never end up in
|
||||||
// a state where no role means "neither admin nor user".
|
// a state where no role means "neither admin nor user".
|
||||||
func NormalizeRole(role string) string {
|
func NormalizeRole(role string) string {
|
||||||
@ -453,15 +435,18 @@ func ValidatePeriod(period string, hours int64) error {
|
|||||||
return fmt.Errorf("period must be one of \"\", hour, week, month, nhour (got %q)", period)
|
return fmt.Errorf("period must be one of \"\", hour, week, month, nhour (got %q)", period)
|
||||||
}
|
}
|
||||||
|
|
||||||
// ModelScope is one allowed model for a key, or one AUTO scheduling slot,
|
// ModelScope is one allowed model for a key, or one AUTO scheduling slot.
|
||||||
// with an optional token quota and reset period. TokenQuota 0 = unlimited;
|
// Its TokenQuota and ReqQuota cap THAT entry only, independently of every
|
||||||
// Period "" = never resets; "hour"/"week"/"month" are fixed windows; "nhour"
|
// other entry on the same key: a model that runs out of budget stops being
|
||||||
// uses Hours as the window length in hours.
|
// served while the key's other models keep working. TokenQuota 0 / ReqQuota 0
|
||||||
|
// = unlimited. Period "" = never resets; "hour"/"week"/"month" are fixed
|
||||||
|
// windows; "nhour" uses Hours.
|
||||||
type ModelScope struct {
|
type ModelScope struct {
|
||||||
Model string `yaml:"model" json:"model"`
|
Model string `yaml:"model" json:"model"`
|
||||||
Source string `yaml:"source,omitempty" json:"source,omitempty"` // optional: pin to one upstream source; "" = any source
|
Source string `yaml:"source,omitempty" json:"source,omitempty"` // optional: pin to one upstream source; "" = any source
|
||||||
Tier int `yaml:"tier,omitempty" json:"tier,omitempty"`
|
Tier int `yaml:"tier,omitempty" json:"tier,omitempty"`
|
||||||
TokenQuota int64 `yaml:"token_quota" json:"token_quota"`
|
TokenQuota int64 `yaml:"token_quota" json:"token_quota"`
|
||||||
|
ReqQuota int64 `yaml:"req_quota,omitempty" json:"req_quota,omitempty"`
|
||||||
Period string `yaml:"period,omitempty" json:"period,omitempty"`
|
Period string `yaml:"period,omitempty" json:"period,omitempty"`
|
||||||
Hours int64 `yaml:"hours,omitempty" json:"hours,omitempty"`
|
Hours int64 `yaml:"hours,omitempty" json:"hours,omitempty"`
|
||||||
}
|
}
|
||||||
@ -480,6 +465,7 @@ func (m *ModelScope) UnmarshalJSON(b []byte) error {
|
|||||||
Source string `json:"source"`
|
Source string `json:"source"`
|
||||||
Tier int `json:"tier"`
|
Tier int `json:"tier"`
|
||||||
TokenQuota int64 `json:"token_quota"`
|
TokenQuota int64 `json:"token_quota"`
|
||||||
|
ReqQuota int64 `json:"req_quota"`
|
||||||
Period string `json:"period"`
|
Period string `json:"period"`
|
||||||
Hours int64 `json:"hours"`
|
Hours int64 `json:"hours"`
|
||||||
}
|
}
|
||||||
@ -490,6 +476,7 @@ func (m *ModelScope) UnmarshalJSON(b []byte) error {
|
|||||||
m.Source = o.Source
|
m.Source = o.Source
|
||||||
m.Tier = o.Tier
|
m.Tier = o.Tier
|
||||||
m.TokenQuota = o.TokenQuota
|
m.TokenQuota = o.TokenQuota
|
||||||
|
m.ReqQuota = o.ReqQuota
|
||||||
m.Period = o.Period
|
m.Period = o.Period
|
||||||
m.Hours = o.Hours
|
m.Hours = o.Hours
|
||||||
return nil
|
return nil
|
||||||
|
|||||||
@ -249,20 +249,17 @@ func (c *Core) FindKey(key string) (config.GWKey, bool) {
|
|||||||
}
|
}
|
||||||
|
|
||||||
// CreateKey builds a new random gateway key and persists it to config.yaml.
|
// CreateKey builds a new random gateway key and persists it to config.yaml.
|
||||||
|
// Quotas live on the model scope entries, so a new key's budget is whatever
|
||||||
|
// its scopes carry.
|
||||||
func (c *Core) CreateKey(name, role string, models []config.ModelScope, note string) (config.GWKey, error) {
|
func (c *Core) CreateKey(name, role string, models []config.ModelScope, note string) (config.GWKey, error) {
|
||||||
return c.CreateKeyWithQuota(name, role, models, note, config.KeyQuota{})
|
|
||||||
}
|
|
||||||
|
|
||||||
// CreateKeyWithQuota is CreateKey plus the key-wide token/request caps.
|
|
||||||
func (c *Core) CreateKeyWithQuota(name, role string, models []config.ModelScope, note string, q config.KeyQuota) (config.GWKey, error) {
|
|
||||||
c.mu.Lock()
|
c.mu.Lock()
|
||||||
defer c.mu.Unlock()
|
defer c.mu.Unlock()
|
||||||
models = cleanScopes(models)
|
models = cleanScopes(models)
|
||||||
key := make([]byte, 16)
|
if err := ValidateScopeQuotas(models); err != nil {
|
||||||
if _, err := rand.Read(key); err != nil {
|
|
||||||
return config.GWKey{}, err
|
return config.GWKey{}, err
|
||||||
}
|
}
|
||||||
if err := q.Validate(); err != nil {
|
key := make([]byte, 16)
|
||||||
|
if _, err := rand.Read(key); err != nil {
|
||||||
return config.GWKey{}, err
|
return config.GWKey{}, err
|
||||||
}
|
}
|
||||||
rec := config.GWKey{
|
rec := config.GWKey{
|
||||||
@ -274,7 +271,6 @@ func (c *Core) CreateKeyWithQuota(name, role string, models []config.ModelScope,
|
|||||||
CreatedAt: time.Now().Unix(),
|
CreatedAt: time.Now().Unix(),
|
||||||
}
|
}
|
||||||
rec.Role = config.NormalizeRole(rec.Role)
|
rec.Role = config.NormalizeRole(rec.Role)
|
||||||
rec.ApplyQuota(q)
|
|
||||||
c.cfg.Keys = append(c.cfg.Keys, rec)
|
c.cfg.Keys = append(c.cfg.Keys, rec)
|
||||||
if err := c.saveConfig(); err != nil {
|
if err := c.saveConfig(); err != nil {
|
||||||
return config.GWKey{}, err
|
return config.GWKey{}, err
|
||||||
@ -282,19 +278,13 @@ func (c *Core) CreateKeyWithQuota(name, role string, models []config.ModelScope,
|
|||||||
return rec, nil
|
return rec, nil
|
||||||
}
|
}
|
||||||
|
|
||||||
// UpdateKey mutates a key's name/role/model scope and persists it.
|
// UpdateKey mutates a key's name/role/model scope and persists it. The scope
|
||||||
|
// entries carry their own quotas, so replacing the scope replaces the budgets.
|
||||||
func (c *Core) UpdateKey(key, name, role string, models []config.ModelScope, note string) (config.GWKey, error) {
|
func (c *Core) UpdateKey(key, name, role string, models []config.ModelScope, note string) (config.GWKey, error) {
|
||||||
return c.UpdateKeyWithQuota(key, name, role, models, note, nil)
|
|
||||||
}
|
|
||||||
|
|
||||||
// UpdateKeyWithQuota is UpdateKey plus the key-wide caps. quota == nil leaves
|
|
||||||
// the existing caps untouched, so a caller that only edits the model scope
|
|
||||||
// does not silently clear a key's budget.
|
|
||||||
func (c *Core) UpdateKeyWithQuota(key, name, role string, models []config.ModelScope, note string, quota *config.KeyQuota) (config.GWKey, error) {
|
|
||||||
c.mu.Lock()
|
c.mu.Lock()
|
||||||
defer c.mu.Unlock()
|
defer c.mu.Unlock()
|
||||||
if quota != nil {
|
if models != nil {
|
||||||
if err := quota.Validate(); err != nil {
|
if err := ValidateScopeQuotas(models); err != nil {
|
||||||
return config.GWKey{}, err
|
return config.GWKey{}, err
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@ -312,9 +302,6 @@ func (c *Core) UpdateKeyWithQuota(key, name, role string, models []config.ModelS
|
|||||||
c.cfg.Keys[i].Models = cleanScopes(models)
|
c.cfg.Keys[i].Models = cleanScopes(models)
|
||||||
}
|
}
|
||||||
c.cfg.Keys[i].Note = note
|
c.cfg.Keys[i].Note = note
|
||||||
if quota != nil {
|
|
||||||
c.cfg.Keys[i].ApplyQuota(*quota)
|
|
||||||
}
|
|
||||||
if err := c.saveConfig(); err != nil {
|
if err := c.saveConfig(); err != nil {
|
||||||
return config.GWKey{}, err
|
return config.GWKey{}, err
|
||||||
}
|
}
|
||||||
@ -798,3 +785,20 @@ func (c *Core) Close() {
|
|||||||
c.vm.Stop()
|
c.vm.Stop()
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// ValidateScopeQuotas checks every scope entry's caps before they are stored.
|
||||||
|
// A typo in a period must be rejected at write time rather than silently
|
||||||
|
// becoming a never-resetting budget — the opposite of what was typed.
|
||||||
|
func ValidateScopeQuotas(entries []config.ModelScope) error {
|
||||||
|
for _, e := range entries {
|
||||||
|
if err := (config.KeyQuota{
|
||||||
|
TokenQuota: e.TokenQuota,
|
||||||
|
ReqQuota: e.ReqQuota,
|
||||||
|
Period: e.Period,
|
||||||
|
Hours: e.Hours,
|
||||||
|
}).Validate(); err != nil {
|
||||||
|
return fmt.Errorf("model %q: %w", e.Model, err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
|||||||
@ -116,12 +116,8 @@ func (g *Gateway) apiV1Routes(w http.ResponseWriter, r *http.Request) {
|
|||||||
"note": k.Note,
|
"note": k.Note,
|
||||||
"created_at": k.CreatedAt,
|
"created_at": k.CreatedAt,
|
||||||
"seed": k.Seed,
|
"seed": k.Seed,
|
||||||
// key-wide spend caps (0 = unlimited). Echoed so an agent can
|
// Quotas live on the scope entries (k.Models), echoed above;
|
||||||
// see what budget it has without parsing config.yaml.
|
// there is deliberately no key-wide total.
|
||||||
"token_quota": k.TokenQuota,
|
|
||||||
"req_quota": k.ReqQuota,
|
|
||||||
"period": k.Period,
|
|
||||||
"hours": k.Hours,
|
|
||||||
// The secret itself is never echoed. An operator that needs it
|
// The secret itself is never echoed. An operator that needs it
|
||||||
// already has it from creation time or from config.yaml.
|
// already has it from creation time or from config.yaml.
|
||||||
"key_prefix": maskKey(k.Key),
|
"key_prefix": maskKey(k.Key),
|
||||||
|
|||||||
@ -218,12 +218,15 @@ type quotaRejection struct {
|
|||||||
|
|
||||||
func (q *quotaRejection) Error() string { return q.msg }
|
func (q *quotaRejection) Error() string { return q.msg }
|
||||||
|
|
||||||
// checkQuota is checkKeyScope for callers that need the retry hint. It
|
// checkQuota validates the effective model against the key's model scope and
|
||||||
// separates the quota verdicts (429) from model-permission verdicts (403).
|
// that entry's quota. Returns nil when the request may proceed.
|
||||||
|
//
|
||||||
|
// Quotas are per scope entry, never key-wide: a model whose budget is spent
|
||||||
|
// is refused on its own while the key's other models keep working. The verdict
|
||||||
|
// carries the remaining seconds of the reset window so a spent budget answers
|
||||||
|
// 429 + Retry-After (come back when it rolls over) instead of 403 (which reads
|
||||||
|
// as "this key may never use this model" and makes clients give up).
|
||||||
func (g *Gateway) checkQuota(ctx context.Context, model string) *quotaRejection {
|
func (g *Gateway) checkQuota(ctx context.Context, model string) *quotaRejection {
|
||||||
if q := g.checkKeyQuotaRetry(ctx); q != nil {
|
|
||||||
return q
|
|
||||||
}
|
|
||||||
allow := g.allowedModels(ctx)
|
allow := g.allowedModels(ctx)
|
||||||
if allow == nil {
|
if allow == nil {
|
||||||
return nil
|
return nil
|
||||||
@ -232,49 +235,29 @@ func (g *Gateway) checkQuota(ctx context.Context, model string) *quotaRejection
|
|||||||
if sc.Model != model {
|
if sc.Model != model {
|
||||||
continue
|
continue
|
||||||
}
|
}
|
||||||
|
win := AutoPeriodSeconds(sc.Period, sc.Hours)
|
||||||
|
k := keyID(reqKey(ctx))
|
||||||
if sc.TokenQuota > 0 {
|
if sc.TokenQuota > 0 {
|
||||||
used := g.scopeTokens(ctx, sc)
|
if used := g.scopeTokens(ctx, sc); used >= sc.TokenQuota {
|
||||||
if used >= sc.TokenQuota {
|
|
||||||
return "aRejection{
|
return "aRejection{
|
||||||
msg: fmt.Sprintf("token quota exceeded for %q (%d/%d)", model, used, sc.TokenQuota),
|
msg: fmt.Sprintf("token quota exceeded for %q (%d/%d)", model, used, sc.TokenQuota),
|
||||||
retry: AutoSecondsToReset(sc.Period, sc.Hours),
|
retry: AutoSecondsToReset(sc.Period, sc.Hours),
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
if sc.ReqQuota > 0 {
|
||||||
|
if used := g.stats.KeyWindowReqs(k, win); used >= sc.ReqQuota {
|
||||||
|
return "aRejection{
|
||||||
|
msg: fmt.Sprintf("request quota exceeded for %q (%d/%d)", model, used, sc.ReqQuota),
|
||||||
|
retry: AutoSecondsToReset(sc.Period, sc.Hours),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
return nil
|
return nil
|
||||||
}
|
}
|
||||||
return "aRejection{msg: fmt.Sprintf("model %q is not allowed for this key", model)}
|
return "aRejection{msg: fmt.Sprintf("model %q is not allowed for this key", model)}
|
||||||
}
|
}
|
||||||
|
|
||||||
// checkKeyQuotaRetry enforces the key-wide caps and reports the remaining
|
|
||||||
// seconds of the reset window so the caller can answer with 429 + Retry-After.
|
|
||||||
// An admin key is never capped, and a key with no caps set is never rejected.
|
|
||||||
func (g *Gateway) checkKeyQuotaRetry(ctx context.Context) *quotaRejection {
|
|
||||||
rec, ok := g.core.FindKey(reqKey(ctx))
|
|
||||||
if !ok || rec.Role == "admin" {
|
|
||||||
return nil
|
|
||||||
}
|
|
||||||
k := keyID(reqKey(ctx))
|
|
||||||
win := AutoPeriodSeconds(rec.Period, rec.Hours)
|
|
||||||
if rec.TokenQuota > 0 {
|
|
||||||
if used := g.stats.KeyWindowTokens(k, win); used >= rec.TokenQuota {
|
|
||||||
return "aRejection{
|
|
||||||
msg: fmt.Sprintf("key token quota exceeded (%d/%d%s)", used, rec.TokenQuota, quotaWindowSuffix(rec.Period, rec.Hours)),
|
|
||||||
retry: AutoSecondsToReset(rec.Period, rec.Hours),
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
if rec.ReqQuota > 0 {
|
|
||||||
if used := g.stats.KeyWindowReqs(k, win); used >= rec.ReqQuota {
|
|
||||||
return "aRejection{
|
|
||||||
msg: fmt.Sprintf("key request quota exceeded (%d/%d%s)", used, rec.ReqQuota, quotaWindowSuffix(rec.Period, rec.Hours)),
|
|
||||||
retry: AutoSecondsToReset(rec.Period, rec.Hours),
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
return nil
|
|
||||||
}
|
|
||||||
|
|
||||||
// quotaWindowSuffix describes a quota's reset window for an error message, so
|
// quotaWindowSuffix describes a quota's reset window for an error message, so
|
||||||
// a rejected caller can tell a permanent block from one that clears in an hour.
|
// a rejected caller can tell a permanent block from one that clears in an hour.
|
||||||
func quotaWindowSuffix(period string, hours int64) string {
|
func quotaWindowSuffix(period string, hours int64) string {
|
||||||
@ -289,11 +272,10 @@ func quotaWindowSuffix(period string, hours int64) string {
|
|||||||
return fmt.Sprintf(", resets every %dh", hours)
|
return fmt.Sprintf(", resets every %dh", hours)
|
||||||
}
|
}
|
||||||
return ""
|
return ""
|
||||||
}
|
} // scopeTokens returns the tokens this key consumed on the scope entry's model
|
||||||
|
// within its reset window, isolated per key. For an AUTO entry the cap covers
|
||||||
// scopeTokens returns the tokens this key consumed within the scope entry's
|
// everything the key routed through AUTO; for a model entry it covers that
|
||||||
// reset window, isolated per key. For an AUTO entry the cap covers everything
|
// model only.
|
||||||
// the key routed through AUTO; for a model entry it covers that model only.
|
|
||||||
//
|
//
|
||||||
// It reads the per-key hourly buckets rather than the key-blind model
|
// It reads the per-key hourly buckets rather than the key-blind model
|
||||||
// buckets, so one key's usage can never exhaust another's quota.
|
// buckets, so one key's usage can never exhaust another's quota.
|
||||||
|
|||||||
@ -74,10 +74,24 @@ func keyRecord(t *testing.T, g *Gateway, secret string) config.GWKey {
|
|||||||
return config.GWKey{}
|
return config.GWKey{}
|
||||||
}
|
}
|
||||||
|
|
||||||
func TestKeyAPIStoresQuota(t *testing.T) {
|
func scopeOf(t *testing.T, k config.GWKey, model string) config.ModelScope {
|
||||||
|
t.Helper()
|
||||||
|
for _, m := range k.Models {
|
||||||
|
if m.Model == model {
|
||||||
|
return m
|
||||||
|
}
|
||||||
|
}
|
||||||
|
t.Fatalf("scope %q not found in %+v", model, k.Models)
|
||||||
|
return config.ModelScope{}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Quotas live on the scope entries, not on the key: creating a key with a
|
||||||
|
// budget means creating scopes that carry it, and they must survive a
|
||||||
|
// read-back (persisted, not just echoed).
|
||||||
|
func TestKeyAPICreatesPerModelQuota(t *testing.T) {
|
||||||
g := adminGateway(t, config.GWKey{Key: "sk-admin", Role: "admin"})
|
g := adminGateway(t, config.GWKey{Key: "sk-admin", Role: "admin"})
|
||||||
rr := adminReq(t, g, "POST", "/api/keys",
|
rr := adminReq(t, g, "POST", "/api/keys",
|
||||||
`{"name":"agent-x","role":"user","token_quota":50000,"req_quota":200,"period":"nhour","hours":6,"models":[{"model":"m1"}]}`)
|
`{"name":"agent-x","role":"user","models":[{"model":"m1","token_quota":50000,"req_quota":200,"period":"nhour","hours":6}]}`)
|
||||||
if rr.Code != 200 {
|
if rr.Code != 200 {
|
||||||
t.Fatalf("create: %d %s", rr.Code, rr.Body.String())
|
t.Fatalf("create: %d %s", rr.Code, rr.Body.String())
|
||||||
}
|
}
|
||||||
@ -87,54 +101,21 @@ func TestKeyAPIStoresQuota(t *testing.T) {
|
|||||||
if err := json.Unmarshal(rr.Body.Bytes(), &created); err != nil {
|
if err := json.Unmarshal(rr.Body.Bytes(), &created); err != nil {
|
||||||
t.Fatalf("decode: %v", err)
|
t.Fatalf("decode: %v", err)
|
||||||
}
|
}
|
||||||
if created.Key.TokenQuota != 50000 || created.Key.ReqQuota != 200 ||
|
sc := scopeOf(t, created.Key, "m1")
|
||||||
created.Key.Period != "nhour" || created.Key.Hours != 6 {
|
if sc.TokenQuota != 50000 || sc.ReqQuota != 200 || sc.Period != "nhour" || sc.Hours != 6 {
|
||||||
t.Fatalf("created key did not carry the caps: %+v", created.Key)
|
t.Fatalf("created scope did not carry the caps: %+v", sc)
|
||||||
}
|
}
|
||||||
// and it must survive a read-back (persisted, not just echoed)
|
back := scopeOf(t, keyRecord(t, g, created.Key.Key), "m1")
|
||||||
back := keyRecord(t, g, created.Key.Key)
|
if back.TokenQuota != 50000 || back.ReqQuota != 200 || back.Period != "nhour" || back.Hours != 6 {
|
||||||
if back.TokenQuota != 50000 || back.Period != "nhour" || back.Hours != 6 {
|
|
||||||
t.Errorf("read-back lost the caps: %+v", back)
|
t.Errorf("read-back lost the caps: %+v", back)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
// Editing only the model scope must not silently clear a key's budget: the
|
// Two models on one key carry independent budgets.
|
||||||
// caps are pointers precisely so "absent" is not "zero".
|
func TestKeyAPIKeepsPerModelQuotaIndependent(t *testing.T) {
|
||||||
func TestKeyAPIUpdateKeepsQuotaWhenOmitted(t *testing.T) {
|
|
||||||
g := adminGateway(t, config.GWKey{Key: "sk-admin", Role: "admin"})
|
g := adminGateway(t, config.GWKey{Key: "sk-admin", Role: "admin"})
|
||||||
rr := adminReq(t, g, "POST", "/api/keys",
|
rr := adminReq(t, g, "POST", "/api/keys",
|
||||||
`{"name":"agent-x","role":"user","token_quota":50000,"period":"day-typo-free","models":[{"model":"m1"}]}`)
|
`{"name":"agent-y","role":"user","models":[{"model":"m1","token_quota":1000,"period":"hour"},{"model":"m2","token_quota":9999,"req_quota":7,"period":"week"}]}`)
|
||||||
rr = adminReq(t, g, "POST", "/api/keys", `{"name":"y","role":"user","token_quota":50000,"period":"hour","models":[{"model":"m1"}]}`)
|
|
||||||
if rr.Code != 200 {
|
|
||||||
t.Fatalf("setup create: %d %s", rr.Code, rr.Body.String())
|
|
||||||
}
|
|
||||||
var created struct {
|
|
||||||
Key config.GWKey `json:"key"`
|
|
||||||
}
|
|
||||||
_ = json.Unmarshal(rr.Body.Bytes(), &created)
|
|
||||||
|
|
||||||
// a scope-only edit
|
|
||||||
rr = adminReq(t, g, "PUT", "/api/keys/"+created.Key.Key,
|
|
||||||
`{"name":"agent-y","models":[{"model":"m1"},{"model":"m2"}]}`)
|
|
||||||
if rr.Code != 200 {
|
|
||||||
t.Fatalf("update: %d %s", rr.Code, rr.Body.String())
|
|
||||||
}
|
|
||||||
back := keyRecord(t, g, created.Key.Key)
|
|
||||||
if back.TokenQuota != 50000 {
|
|
||||||
t.Errorf("token_quota was cleared by a scope-only edit: %d", back.TokenQuota)
|
|
||||||
}
|
|
||||||
if back.Period != "hour" {
|
|
||||||
t.Errorf("period was cleared by a scope-only edit: %q", back.Period)
|
|
||||||
}
|
|
||||||
if len(back.Models) != 2 {
|
|
||||||
t.Errorf("scope edit did not apply: %+v", back.Models)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
// Sending 0 explicitly must lift the cap, not be treated as "absent".
|
|
||||||
func TestKeyAPIUpdateZeroLiftsCap(t *testing.T) {
|
|
||||||
g := adminGateway(t, config.GWKey{Key: "sk-admin", Role: "admin"})
|
|
||||||
rr := adminReq(t, g, "POST", "/api/keys", `{"name":"z","role":"user","token_quota":1000,"period":"hour"}`)
|
|
||||||
if rr.Code != 200 {
|
if rr.Code != 200 {
|
||||||
t.Fatalf("create: %d %s", rr.Code, rr.Body.String())
|
t.Fatalf("create: %d %s", rr.Code, rr.Body.String())
|
||||||
}
|
}
|
||||||
@ -142,13 +123,36 @@ func TestKeyAPIUpdateZeroLiftsCap(t *testing.T) {
|
|||||||
Key config.GWKey `json:"key"`
|
Key config.GWKey `json:"key"`
|
||||||
}
|
}
|
||||||
_ = json.Unmarshal(rr.Body.Bytes(), &created)
|
_ = json.Unmarshal(rr.Body.Bytes(), &created)
|
||||||
|
back := keyRecord(t, g, created.Key.Key)
|
||||||
|
if a := scopeOf(t, back, "m1"); a.TokenQuota != 1000 || a.ReqQuota != 0 || a.Period != "hour" {
|
||||||
|
t.Errorf("m1 caps wrong: %+v", a)
|
||||||
|
}
|
||||||
|
if b := scopeOf(t, back, "m2"); b.TokenQuota != 9999 || b.ReqQuota != 7 || b.Period != "week" {
|
||||||
|
t.Errorf("m2 caps wrong: %+v", b)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
rr = adminReq(t, g, "PUT", "/api/keys/"+created.Key.Key, `{"token_quota":0}`)
|
// Sending 0 explicitly lifts that model's cap.
|
||||||
|
func TestKeyAPIUpdateZeroLiftsCap(t *testing.T) {
|
||||||
|
g := adminGateway(t, config.GWKey{Key: "sk-admin", Role: "admin"})
|
||||||
|
rr := adminReq(t, g, "POST", "/api/keys",
|
||||||
|
`{"name":"z","role":"user","models":[{"model":"m1","token_quota":1000,"period":"hour"}]}`)
|
||||||
|
if rr.Code != 200 {
|
||||||
|
t.Fatalf("create: %d %s", rr.Code, rr.Body.String())
|
||||||
|
}
|
||||||
|
var created struct {
|
||||||
|
Key config.GWKey `json:"key"`
|
||||||
|
}
|
||||||
|
_ = json.Unmarshal(rr.Body.Bytes(), &created)
|
||||||
|
secret := created.Key.Key
|
||||||
|
|
||||||
|
rr = adminReq(t, g, "PUT", "/api/keys/"+secret,
|
||||||
|
`{"models":[{"model":"m1","token_quota":0,"req_quota":0,"period":""}]}`)
|
||||||
if rr.Code != 200 {
|
if rr.Code != 200 {
|
||||||
t.Fatalf("lift: %d %s", rr.Code, rr.Body.String())
|
t.Fatalf("lift: %d %s", rr.Code, rr.Body.String())
|
||||||
}
|
}
|
||||||
if back := keyRecord(t, g, created.Key.Key); back.TokenQuota != 0 {
|
if back := scopeOf(t, keyRecord(t, g, secret), "m1"); back.TokenQuota != 0 || back.ReqQuota != 0 || back.Period != "" {
|
||||||
t.Errorf("token_quota = %d, want 0 (cap lifted)", back.TokenQuota)
|
t.Errorf("caps not lifted: %+v", back)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@ -156,7 +160,8 @@ func TestKeyAPIUpdateZeroLiftsCap(t *testing.T) {
|
|||||||
// quota — which is the exact opposite of what the operator typed.
|
// quota — which is the exact opposite of what the operator typed.
|
||||||
func TestKeyAPIRejectsBadPeriod(t *testing.T) {
|
func TestKeyAPIRejectsBadPeriod(t *testing.T) {
|
||||||
g := adminGateway(t, config.GWKey{Key: "sk-admin", Role: "admin"})
|
g := adminGateway(t, config.GWKey{Key: "sk-admin", Role: "admin"})
|
||||||
rr := adminReq(t, g, "POST", "/api/keys", `{"name":"bad","role":"user","token_quota":1000,"period":"houre"}`)
|
rr := adminReq(t, g, "POST", "/api/keys",
|
||||||
|
`{"name":"bad","role":"user","models":[{"model":"m1","token_quota":1000,"period":"houre"}]}`)
|
||||||
if rr.Code != http.StatusBadRequest {
|
if rr.Code != http.StatusBadRequest {
|
||||||
t.Fatalf("want 400 for a bad period, got %d %s", rr.Code, rr.Body.String())
|
t.Fatalf("want 400 for a bad period, got %d %s", rr.Code, rr.Body.String())
|
||||||
}
|
}
|
||||||
@ -167,19 +172,35 @@ func TestKeyAPIRejectsBadPeriod(t *testing.T) {
|
|||||||
|
|
||||||
func TestKeyAPIRejectsNegativeQuota(t *testing.T) {
|
func TestKeyAPIRejectsNegativeQuota(t *testing.T) {
|
||||||
g := adminGateway(t, config.GWKey{Key: "sk-admin", Role: "admin"})
|
g := adminGateway(t, config.GWKey{Key: "sk-admin", Role: "admin"})
|
||||||
rr := adminReq(t, g, "POST", "/api/keys", `{"name":"bad","role":"user","token_quota":-5}`)
|
rr := adminReq(t, g, "POST", "/api/keys",
|
||||||
|
`{"name":"bad","role":"user","models":[{"model":"m1","token_quota":-5}]}`)
|
||||||
if rr.Code != http.StatusBadRequest {
|
if rr.Code != http.StatusBadRequest {
|
||||||
t.Fatalf("want 400 for a negative quota, got %d %s", rr.Code, rr.Body.String())
|
t.Fatalf("want 400 for a negative quota, got %d %s", rr.Code, rr.Body.String())
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
// A non-admin key must not be able to set or read another key's budget.
|
// The rejection must say WHICH model is over budget, so an operator looking at
|
||||||
|
// a key with a dozen scopes can tell which one to raise.
|
||||||
|
func TestKeyAPIRejectionNamesTheModel(t *testing.T) {
|
||||||
|
g := adminGateway(t, config.GWKey{Key: "sk-admin", Role: "admin"})
|
||||||
|
rr := adminReq(t, g, "POST", "/api/keys",
|
||||||
|
`{"name":"bad","role":"user","models":[{"model":"m1","req_quota":-1}]}`)
|
||||||
|
if rr.Code != http.StatusBadRequest {
|
||||||
|
t.Fatalf("want 400, got %d", rr.Code)
|
||||||
|
}
|
||||||
|
if !strings.Contains(rr.Body.String(), "m1") {
|
||||||
|
t.Errorf("error should name the offending model: %s", rr.Body.String())
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// A non-admin key must not be able to mint keys.
|
||||||
func TestKeyAPIQuotaIsAdminOnly(t *testing.T) {
|
func TestKeyAPIQuotaIsAdminOnly(t *testing.T) {
|
||||||
g := adminGateway(t,
|
g := adminGateway(t,
|
||||||
config.GWKey{Key: "sk-admin", Role: "admin"},
|
config.GWKey{Key: "sk-admin", Role: "admin"},
|
||||||
config.GWKey{Key: "sk-u", Role: "user", TokenQuota: 10, Period: "hour"},
|
config.GWKey{Key: "sk-u", Role: "user", Models: []config.ModelScope{{Model: "m1", TokenQuota: 10, Period: "hour"}}},
|
||||||
)
|
)
|
||||||
req, _ := http.NewRequest("POST", "/api/keys", strings.NewReader(`{"name":"x","role":"admin","token_quota":0}`))
|
req, _ := http.NewRequest("POST", "/api/keys",
|
||||||
|
strings.NewReader(`{"name":"x","role":"admin","models":[{"model":"m1"}]}`))
|
||||||
req.Header.Set("Authorization", "Bearer sk-u")
|
req.Header.Set("Authorization", "Bearer sk-u")
|
||||||
req.Header.Set("Content-Type", "application/json")
|
req.Header.Set("Content-Type", "application/json")
|
||||||
rr := httptest.NewRecorder()
|
rr := httptest.NewRecorder()
|
||||||
@ -187,7 +208,7 @@ func TestKeyAPIQuotaIsAdminOnly(t *testing.T) {
|
|||||||
if rr.Code != http.StatusForbidden {
|
if rr.Code != http.StatusForbidden {
|
||||||
t.Fatalf("non-admin create: want 403, got %d %s", rr.Code, rr.Body.String())
|
t.Fatalf("non-admin create: want 403, got %d %s", rr.Code, rr.Body.String())
|
||||||
}
|
}
|
||||||
// /api/v1/keys echoes the caps but never the secret
|
// /api/v1/keys exposes the per-model caps but never a secret
|
||||||
rr = adminReq(t, g, "GET", "/api/v1/keys", "")
|
rr = adminReq(t, g, "GET", "/api/v1/keys", "")
|
||||||
if rr.Code != 200 {
|
if rr.Code != 200 {
|
||||||
t.Fatalf("GET /api/v1/keys: %d", rr.Code)
|
t.Fatalf("GET /api/v1/keys: %d", rr.Code)
|
||||||
@ -196,39 +217,10 @@ func TestKeyAPIQuotaIsAdminOnly(t *testing.T) {
|
|||||||
t.Error("/api/v1/keys leaked a key secret")
|
t.Error("/api/v1/keys leaked a key secret")
|
||||||
}
|
}
|
||||||
if !strings.Contains(rr.Body.String(), `"token_quota":10`) {
|
if !strings.Contains(rr.Body.String(), `"token_quota":10`) {
|
||||||
t.Errorf("/api/v1/keys should expose the cap: %s", rr.Body.String())
|
t.Errorf("/api/v1/keys should expose the per-model caps: %s", rr.Body.String())
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
// The AUTO scope entry must honour its reset window: usage that aged out of
|
|
||||||
// the window must not count against a per-key cap.
|
|
||||||
func TestAutoScopeQuotaHonoursWindow(t *testing.T) {
|
|
||||||
g, _ := quotaGateway(t,
|
|
||||||
config.GWKey{Key: "sk-a", Role: "user", Models: []config.ModelScope{
|
|
||||||
{Model: "AUTO", TokenQuota: 1000, Period: "hour"},
|
|
||||||
}},
|
|
||||||
config.GWKey{Key: "sk-b", Role: "user"},
|
|
||||||
)
|
|
||||||
ctx := quotaCtx(t, g, "sk-a")
|
|
||||||
sc := config.ModelScope{Model: "AUTO", TokenQuota: 1000, Period: "hour"}
|
|
||||||
|
|
||||||
// aged-out usage: 2 days old, 5M tokens — must be invisible to a 1h window
|
|
||||||
g.stats.Record(Req{Time: nowMSOffset(-48 * 3600 * 1000), Key: keyID("sk-a"),
|
|
||||||
Model: "m1", Source: "up", Prompt: 2500000, Compl: 2500000, OK: true, Status: 200})
|
|
||||||
if used := g.scopeTokens(ctx, sc); used != 0 {
|
|
||||||
t.Fatalf("AUTO scope saw %d tokens outside its 1h window; the period is being ignored", used)
|
|
||||||
}
|
|
||||||
|
|
||||||
// in-window usage counts
|
|
||||||
g.stats.Record(Req{Time: nowMSOffset(0), Key: keyID("sk-a"),
|
|
||||||
Model: "m1", Source: "up", Prompt: 400, Compl: 400, OK: true, Status: 200})
|
|
||||||
if used := g.scopeTokens(ctx, sc); used != 800 {
|
|
||||||
t.Fatalf("AUTO scope used = %d, want 800", used)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
var _ = fmt.Sprintf
|
|
||||||
|
|
||||||
// A user must be able to see their own budget: /api/keys/me is the only key
|
// A user must be able to see their own budget: /api/keys/me is the only key
|
||||||
// view a non-admin gets, so a cap missing from it is invisible to the very
|
// view a non-admin gets, so a cap missing from it is invisible to the very
|
||||||
// client it constrains.
|
// client it constrains.
|
||||||
@ -236,7 +228,7 @@ func TestKeyMeExposesOwnQuota(t *testing.T) {
|
|||||||
g := adminGateway(t,
|
g := adminGateway(t,
|
||||||
config.GWKey{Key: "sk-admin", Role: "admin"},
|
config.GWKey{Key: "sk-admin", Role: "admin"},
|
||||||
config.GWKey{Key: "sk-u", Role: "user", Name: "agent",
|
config.GWKey{Key: "sk-u", Role: "user", Name: "agent",
|
||||||
TokenQuota: 123456, ReqQuota: 42, Period: "week", Hours: 0},
|
Models: []config.ModelScope{{Model: "m1", TokenQuota: 123456, ReqQuota: 42, Period: "week"}}},
|
||||||
)
|
)
|
||||||
req, _ := http.NewRequest("GET", "/api/keys/me", nil)
|
req, _ := http.NewRequest("GET", "/api/keys/me", nil)
|
||||||
req.Header.Set("Authorization", "Bearer sk-u")
|
req.Header.Set("Authorization", "Bearer sk-u")
|
||||||
@ -252,8 +244,10 @@ func TestKeyMeExposesOwnQuota(t *testing.T) {
|
|||||||
if err := json.Unmarshal(rr.Body.Bytes(), &wrap); err != nil {
|
if err := json.Unmarshal(rr.Body.Bytes(), &wrap); err != nil {
|
||||||
t.Fatalf("decode: %v (%s)", err, rr.Body.String())
|
t.Fatalf("decode: %v (%s)", err, rr.Body.String())
|
||||||
}
|
}
|
||||||
me := wrap.Key
|
sc := scopeOf(t, wrap.Key, "m1")
|
||||||
if me.TokenQuota != 123456 || me.ReqQuota != 42 || me.Period != "week" {
|
if sc.TokenQuota != 123456 || sc.ReqQuota != 42 || sc.Period != "week" {
|
||||||
t.Errorf("own quota not visible to the key's owner: %+v", me)
|
t.Errorf("own quota not visible to the key's owner: %+v", sc)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
var _ = fmt.Sprintf
|
||||||
|
|||||||
@ -82,8 +82,8 @@ func chatAs(t *testing.T, g *Gateway, key, model string) (*httptest.ResponseReco
|
|||||||
// as "this key may never use this model".
|
// as "this key may never use this model".
|
||||||
func TestKeyTokenQuotaBlocksWithRetryAfter(t *testing.T) {
|
func TestKeyTokenQuotaBlocksWithRetryAfter(t *testing.T) {
|
||||||
g, _ := quotaGateway(t,
|
g, _ := quotaGateway(t,
|
||||||
config.GWKey{Key: "sk-a", Role: "user", TokenQuota: 8, Period: "hour",
|
config.GWKey{Key: "sk-a", Role: "user",
|
||||||
Models: []config.ModelScope{{Model: "m1"}}},
|
Models: []config.ModelScope{{Model: "m1", TokenQuota: 8, Period: "hour"}}},
|
||||||
config.GWKey{Key: "sk-b", Role: "user", Name: "b"},
|
config.GWKey{Key: "sk-b", Role: "user", Name: "b"},
|
||||||
)
|
)
|
||||||
// 4 tokens per call, budget 8 -> the third call crosses it
|
// 4 tokens per call, budget 8 -> the third call crosses it
|
||||||
@ -114,10 +114,10 @@ func TestKeyTokenQuotaBlocksWithRetryAfter(t *testing.T) {
|
|||||||
// key that is allowed the same model.
|
// key that is allowed the same model.
|
||||||
func TestKeyTokenQuotaIsIsolatedPerKey(t *testing.T) {
|
func TestKeyTokenQuotaIsIsolatedPerKey(t *testing.T) {
|
||||||
g, _ := quotaGateway(t,
|
g, _ := quotaGateway(t,
|
||||||
config.GWKey{Key: "sk-a", Role: "user", TokenQuota: 4, Period: "hour",
|
config.GWKey{Key: "sk-a", Role: "user",
|
||||||
Models: []config.ModelScope{{Model: "m1"}}},
|
Models: []config.ModelScope{{Model: "m1", TokenQuota: 4, Period: "hour"}}},
|
||||||
config.GWKey{Key: "sk-b", Role: "user", TokenQuota: 1000, Period: "hour",
|
config.GWKey{Key: "sk-b", Role: "user",
|
||||||
Models: []config.ModelScope{{Model: "m1"}}},
|
Models: []config.ModelScope{{Model: "m1", TokenQuota: 1000, Period: "hour"}}},
|
||||||
)
|
)
|
||||||
rr, _ := chatAs(t, g, "sk-a", "m1")
|
rr, _ := chatAs(t, g, "sk-a", "m1")
|
||||||
if rr.Code != 200 {
|
if rr.Code != 200 {
|
||||||
@ -162,7 +162,8 @@ func TestScopeModelTokenQuotaIsolatedPerKey(t *testing.T) {
|
|||||||
// operator out of the gateway they administer.
|
// operator out of the gateway they administer.
|
||||||
func TestAdminKeyIsNeverQuotaCapped(t *testing.T) {
|
func TestAdminKeyIsNeverQuotaCapped(t *testing.T) {
|
||||||
g, _ := quotaGateway(t,
|
g, _ := quotaGateway(t,
|
||||||
config.GWKey{Key: "sk-admin", Role: "admin", TokenQuota: 1, Period: "hour", ReqQuota: 1},
|
config.GWKey{Key: "sk-admin", Role: "admin",
|
||||||
|
Models: []config.ModelScope{{Model: "m1", TokenQuota: 1, ReqQuota: 1, Period: "hour"}}},
|
||||||
config.GWKey{Key: "sk-b", Role: "user"},
|
config.GWKey{Key: "sk-b", Role: "user"},
|
||||||
)
|
)
|
||||||
for i := 1; i <= 3; i++ {
|
for i := 1; i <= 3; i++ {
|
||||||
@ -176,8 +177,8 @@ func TestAdminKeyIsNeverQuotaCapped(t *testing.T) {
|
|||||||
// the count must still stop the key.
|
// the count must still stop the key.
|
||||||
func TestKeyRequestQuotaBlocks(t *testing.T) {
|
func TestKeyRequestQuotaBlocks(t *testing.T) {
|
||||||
g, ctrl := quotaGateway(t,
|
g, ctrl := quotaGateway(t,
|
||||||
config.GWKey{Key: "sk-a", Role: "user", ReqQuota: 2, Period: "hour",
|
config.GWKey{Key: "sk-a", Role: "user",
|
||||||
Models: []config.ModelScope{{Model: "m1"}}},
|
Models: []config.ModelScope{{Model: "m1", ReqQuota: 2, Period: "hour"}}},
|
||||||
config.GWKey{Key: "sk-b", Role: "user"},
|
config.GWKey{Key: "sk-b", Role: "user"},
|
||||||
)
|
)
|
||||||
for i := 1; i <= 2; i++ {
|
for i := 1; i <= 2; i++ {
|
||||||
@ -197,8 +198,8 @@ func TestKeyRequestQuotaBlocks(t *testing.T) {
|
|||||||
// A model outside the scope is still 403, not 429: retrying cannot help.
|
// A model outside the scope is still 403, not 429: retrying cannot help.
|
||||||
func TestModelOutsideScopeStaysForbidden(t *testing.T) {
|
func TestModelOutsideScopeStaysForbidden(t *testing.T) {
|
||||||
g, _ := quotaGateway(t,
|
g, _ := quotaGateway(t,
|
||||||
config.GWKey{Key: "sk-a", Role: "user", TokenQuota: 1000, Period: "hour",
|
config.GWKey{Key: "sk-a", Role: "user",
|
||||||
Models: []config.ModelScope{{Model: "other-model"}}},
|
Models: []config.ModelScope{{Model: "other-model", TokenQuota: 1000, Period: "hour"}}},
|
||||||
config.GWKey{Key: "sk-b", Role: "user"},
|
config.GWKey{Key: "sk-b", Role: "user"},
|
||||||
)
|
)
|
||||||
rr, code := chatAs(t, g, "sk-a", "m1")
|
rr, code := chatAs(t, g, "sk-a", "m1")
|
||||||
@ -256,8 +257,8 @@ func TestKeyQuotaWinsOverSlotQuota(t *testing.T) {
|
|||||||
up := upstream(t, &upstreamCtrl{})
|
up := upstream(t, &upstreamCtrl{})
|
||||||
defer up.Close()
|
defer up.Close()
|
||||||
g := newQuotaGW(t, up.URL,
|
g := newQuotaGW(t, up.URL,
|
||||||
config.GWKey{Key: "sk-a", Role: "user", TokenQuota: 4, Period: "hour",
|
config.GWKey{Key: "sk-a", Role: "user",
|
||||||
Models: []config.ModelScope{{Model: "AUTO"}}})
|
Models: []config.ModelScope{{Model: "AUTO", TokenQuota: 4, Period: "hour"}}})
|
||||||
ctx := quotaCtx(t, g, "sk-a")
|
ctx := quotaCtx(t, g, "sk-a")
|
||||||
|
|
||||||
// exhaust the key first
|
// exhaust the key first
|
||||||
@ -304,7 +305,10 @@ func newQuotaGW(t *testing.T, upURL string, keys ...config.GWKey) *Gateway {
|
|||||||
Keys: keys,
|
Keys: keys,
|
||||||
Sources: []config.Source{{
|
Sources: []config.Source{{
|
||||||
Name: "up", BaseURL: upURL, Adapter: "openai",
|
Name: "up", BaseURL: upURL, Adapter: "openai",
|
||||||
Models: []config.Model{{ID: "m1", Priority: 100}},
|
Models: []config.Model{
|
||||||
|
{ID: "m1", Priority: 100},
|
||||||
|
{ID: "m2", Priority: 90},
|
||||||
|
},
|
||||||
}},
|
}},
|
||||||
}
|
}
|
||||||
if err := cfg.ApplyDefaults(); err != nil {
|
if err := cfg.ApplyDefaults(); err != nil {
|
||||||
@ -325,3 +329,116 @@ func newQuotaGW(t *testing.T, upURL string, keys ...config.GWKey) *Gateway {
|
|||||||
}
|
}
|
||||||
return g
|
return g
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// The core of per-model quotas: a model that runs out of budget must stop
|
||||||
|
// being served on its own, while every other model on the SAME key keeps
|
||||||
|
// working. A key-wide total would fail this — it would block m2 because m1 was
|
||||||
|
// capped, which is exactly the coupling this design removes.
|
||||||
|
func TestOneModelsQuotaDoesNotBlockAnother(t *testing.T) {
|
||||||
|
up := upstream(t, &upstreamCtrl{})
|
||||||
|
defer up.Close()
|
||||||
|
td := t.TempDir()
|
||||||
|
cfgPath := td + "/config.yaml"
|
||||||
|
if err := os.WriteFile(cfgPath, []byte("listen: :0"), 0o644); err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
cfg := &config.Config{
|
||||||
|
Path: cfgPath, AdapterDir: filepath.Join(td, "adapters"),
|
||||||
|
RuntimeFile: filepath.Join(td, "runtime.json"),
|
||||||
|
Keys: []config.GWKey{{
|
||||||
|
Key: "sk-a", Role: "user", Name: "two-models",
|
||||||
|
Models: []config.ModelScope{
|
||||||
|
{Model: "m1", TokenQuota: 4, Period: "hour"},
|
||||||
|
{Model: "m2", TokenQuota: 1000, Period: "hour"},
|
||||||
|
},
|
||||||
|
}},
|
||||||
|
Sources: []config.Source{{
|
||||||
|
Name: "up", BaseURL: up.URL, Adapter: "openai",
|
||||||
|
Models: []config.Model{
|
||||||
|
{ID: "m1", Priority: 100},
|
||||||
|
{ID: "m2", Priority: 90},
|
||||||
|
},
|
||||||
|
}},
|
||||||
|
}
|
||||||
|
if err := cfg.ApplyDefaults(); err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
c, err := core.NewFromConfig(cfg)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("core: %v", err)
|
||||||
|
}
|
||||||
|
t.Cleanup(c.Close)
|
||||||
|
g, err := New(c, []string{"sk-a"})
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("gateway: %v", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
// Spend on m2 FIRST. Without this the two designs are
|
||||||
|
// indistinguishable: a key-wide counter and m1's own counter would both
|
||||||
|
// read 0 before m1 is used, so the test would pass either way (it did —
|
||||||
|
// see the commit that rewrote it).
|
||||||
|
for i := 1; i <= 3; i++ {
|
||||||
|
if rr, _ := chatAs(t, g, "sk-a", "m2"); rr.Code != 200 {
|
||||||
|
t.Fatalf("m2 priming call %d: want 200, got %d", i, rr.Code)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
// m1's budget is 4 and each call costs 4, so the key-wide total (m2+m1) is
|
||||||
|
// already 12 when m1 starts: a key-wide cap would refuse m1 immediately.
|
||||||
|
if rr, _ := chatAs(t, g, "sk-a", "m1"); rr.Code != 200 {
|
||||||
|
t.Fatalf("m1 must be served on its own budget (m2's spend is not its problem), got %d (%s)",
|
||||||
|
rr.Code, rr.Body.String())
|
||||||
|
}
|
||||||
|
if rr, code := chatAs(t, g, "sk-a", "m1"); rr.Code != http.StatusTooManyRequests {
|
||||||
|
t.Fatalf("m1 second call: want 429, got %d (%s)", rr.Code, code)
|
||||||
|
}
|
||||||
|
// m2 must keep working
|
||||||
|
for i := 1; i <= 3; i++ {
|
||||||
|
if rr, _ := chatAs(t, g, "sk-a", "m2"); rr.Code != 200 {
|
||||||
|
t.Fatalf("m2 call %d must be served while m1 is capped, got %d (%s)", i, rr.Code, rr.Body.String())
|
||||||
|
}
|
||||||
|
}
|
||||||
|
// and the message must name m1, not the key
|
||||||
|
rr, _ := chatAs(t, g, "sk-a", "m1")
|
||||||
|
if !strings.Contains(rr.Body.String(), "m1") {
|
||||||
|
t.Errorf("rejection should name the capped model: %s", rr.Body.String())
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// A model with no quota on it is never blocked by a sibling's cap, and an
|
||||||
|
// uncapped key is never blocked at all.
|
||||||
|
func TestUncappedModelNeverBlocked(t *testing.T) {
|
||||||
|
up := upstream(t, &upstreamCtrl{})
|
||||||
|
defer up.Close()
|
||||||
|
g := newQuotaGW(t, up.URL,
|
||||||
|
config.GWKey{Key: "sk-a", Role: "user", Models: []config.ModelScope{
|
||||||
|
{Model: "m1", TokenQuota: 1, Period: "hour"},
|
||||||
|
{Model: "m2"},
|
||||||
|
}})
|
||||||
|
if rr, _ := chatAs(t, g, "sk-a", "m1"); rr.Code != 200 {
|
||||||
|
t.Fatalf("m1 first: want 200, got %d", rr.Code)
|
||||||
|
}
|
||||||
|
if rr, _ := chatAs(t, g, "sk-a", "m1"); rr.Code != http.StatusTooManyRequests {
|
||||||
|
t.Fatalf("m1 second: want 429, got %d", rr.Code)
|
||||||
|
}
|
||||||
|
for i := 1; i <= 4; i++ {
|
||||||
|
if rr, _ := chatAs(t, g, "sk-a", "m2"); rr.Code != 200 {
|
||||||
|
t.Fatalf("m2 (uncapped) call %d: want 200, got %d", i, rr.Code)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// An admin key is never capped even when its scopes carry budgets: a cap that
|
||||||
|
// locked the operator out would be unrecoverable through the UI.
|
||||||
|
func TestAdminKeyScopesAreNotEnforced(t *testing.T) {
|
||||||
|
up := upstream(t, &upstreamCtrl{})
|
||||||
|
defer up.Close()
|
||||||
|
g := newQuotaGW(t, up.URL,
|
||||||
|
config.GWKey{Key: "sk-admin", Role: "admin", Models: []config.ModelScope{
|
||||||
|
{Model: "m1", TokenQuota: 1, ReqQuota: 1, Period: "hour"},
|
||||||
|
}})
|
||||||
|
for i := 1; i <= 4; i++ {
|
||||||
|
if rr, _ := chatAs(t, g, "sk-admin", "m1"); rr.Code != 200 {
|
||||||
|
t.Fatalf("admin call %d: want 200 (admin scopes are not enforced), got %d", i, rr.Code)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|||||||
@ -39,27 +39,17 @@ func (g *Gateway) handleKeysAPI(w http.ResponseWriter, r *http.Request) {
|
|||||||
writeJSON(w, http.StatusOK, map[string]interface{}{"keys": g.core.ListKeys()})
|
writeJSON(w, http.StatusOK, map[string]interface{}{"keys": g.core.ListKeys()})
|
||||||
case http.MethodPost:
|
case http.MethodPost:
|
||||||
var body struct {
|
var body struct {
|
||||||
Name string `json:"name"`
|
Name string `json:"name"`
|
||||||
Role string `json:"role"`
|
Role string `json:"role"`
|
||||||
Models []config.ModelScope `json:"models"`
|
Models []config.ModelScope `json:"models"`
|
||||||
Note string `json:"note"`
|
Note string `json:"note"`
|
||||||
TokenQuota *int64 `json:"token_quota"`
|
|
||||||
ReqQuota *int64 `json:"req_quota"`
|
|
||||||
Period *string `json:"period"`
|
|
||||||
Hours *int64 `json:"hours"`
|
|
||||||
}
|
}
|
||||||
if err := json.NewDecoder(r.Body).Decode(&body); err != nil {
|
if err := json.NewDecoder(r.Body).Decode(&body); err != nil {
|
||||||
writeError(w, http.StatusBadRequest, "invalid_request", "invalid json: "+err.Error())
|
writeError(w, http.StatusBadRequest, "invalid_request", "invalid json: "+err.Error())
|
||||||
return
|
return
|
||||||
}
|
}
|
||||||
body.Role = config.NormalizeRole(body.Role)
|
body.Role = config.NormalizeRole(body.Role)
|
||||||
q := config.KeyQuota{
|
rec, err := g.core.CreateKey(body.Name, body.Role, body.Models, body.Note)
|
||||||
TokenQuota: optInt64(body.TokenQuota),
|
|
||||||
ReqQuota: optInt64(body.ReqQuota),
|
|
||||||
Period: optString(body.Period),
|
|
||||||
Hours: optInt64(body.Hours),
|
|
||||||
}
|
|
||||||
rec, err := g.core.CreateKeyWithQuota(body.Name, body.Role, body.Models, body.Note, q)
|
|
||||||
if err != nil {
|
if err != nil {
|
||||||
writeError(w, http.StatusBadRequest, "key_error", err.Error())
|
writeError(w, http.StatusBadRequest, "key_error", err.Error())
|
||||||
return
|
return
|
||||||
@ -71,33 +61,19 @@ func (g *Gateway) handleKeysAPI(w http.ResponseWriter, r *http.Request) {
|
|||||||
return
|
return
|
||||||
}
|
}
|
||||||
var body struct {
|
var body struct {
|
||||||
Name string `json:"name"`
|
Name string `json:"name"`
|
||||||
Role string `json:"role"`
|
Role string `json:"role"`
|
||||||
Models []config.ModelScope `json:"models"`
|
Models []config.ModelScope `json:"models"`
|
||||||
Note string `json:"note"`
|
Note string `json:"note"`
|
||||||
TokenQuota *int64 `json:"token_quota"`
|
|
||||||
ReqQuota *int64 `json:"req_quota"`
|
|
||||||
Period *string `json:"period"`
|
|
||||||
Hours *int64 `json:"hours"`
|
|
||||||
}
|
}
|
||||||
if err := json.NewDecoder(r.Body).Decode(&body); err != nil {
|
if err := json.NewDecoder(r.Body).Decode(&body); err != nil {
|
||||||
writeError(w, http.StatusBadRequest, "invalid_request", "invalid json: "+err.Error())
|
writeError(w, http.StatusBadRequest, "invalid_request", "invalid json: "+err.Error())
|
||||||
return
|
return
|
||||||
}
|
}
|
||||||
// Quota fields are pointers so "absent" is distinguishable from
|
// Each scope entry carries its own token/request caps, so replacing the
|
||||||
// "set to 0": omitting them leaves the stored caps alone, while
|
// scope replaces the budgets with it — there is no separate key-wide
|
||||||
// sending 0 explicitly lifts a cap. Without this, editing only the
|
// quota that could drift out of sync with the models.
|
||||||
// model scope would silently clear a key's budget.
|
rec, err := g.core.UpdateKey(path, body.Name, body.Role, body.Models, body.Note)
|
||||||
var q *config.KeyQuota
|
|
||||||
if body.TokenQuota != nil || body.ReqQuota != nil || body.Period != nil || body.Hours != nil {
|
|
||||||
q = &config.KeyQuota{
|
|
||||||
TokenQuota: optInt64(body.TokenQuota),
|
|
||||||
ReqQuota: optInt64(body.ReqQuota),
|
|
||||||
Period: optString(body.Period),
|
|
||||||
Hours: optInt64(body.Hours),
|
|
||||||
}
|
|
||||||
}
|
|
||||||
rec, err := g.core.UpdateKeyWithQuota(path, body.Name, body.Role, body.Models, body.Note, q)
|
|
||||||
if err != nil {
|
if err != nil {
|
||||||
writeError(w, http.StatusBadRequest, "key_error", err.Error())
|
writeError(w, http.StatusBadRequest, "key_error", err.Error())
|
||||||
return
|
return
|
||||||
@ -127,22 +103,6 @@ func (g *Gateway) handleKeysAPI(w http.ResponseWriter, r *http.Request) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
// optInt64 dereferences an optional quota field, treating absent as 0.
|
|
||||||
func optInt64(p *int64) int64 {
|
|
||||||
if p == nil {
|
|
||||||
return 0
|
|
||||||
}
|
|
||||||
return *p
|
|
||||||
}
|
|
||||||
|
|
||||||
// optString dereferences an optional quota field, treating absent as "".
|
|
||||||
func optString(p *string) string {
|
|
||||||
if p == nil {
|
|
||||||
return ""
|
|
||||||
}
|
|
||||||
return *p
|
|
||||||
}
|
|
||||||
|
|
||||||
// handleKeyMe returns the authenticated key's own record (users see only
|
// handleKeyMe returns the authenticated key's own record (users see only
|
||||||
// themselves; admins can use this as a convenience too).
|
// themselves; admins can use this as a convenience too).
|
||||||
func (g *Gateway) handleKeyMe(w http.ResponseWriter, r *http.Request) {
|
func (g *Gateway) handleKeyMe(w http.ResponseWriter, r *http.Request) {
|
||||||
|
|||||||
@ -339,9 +339,6 @@
|
|||||||
.key-canvas{border:1px solid var(--line);border-radius:16px;padding:14px;margin-bottom:14px;background:var(--card);
|
.key-canvas{border:1px solid var(--line);border-radius:16px;padding:14px;margin-bottom:14px;background:var(--card);
|
||||||
backdrop-filter:blur(var(--glass));box-shadow:var(--sh-sm)}
|
backdrop-filter:blur(var(--glass));box-shadow:var(--sh-sm)}
|
||||||
.kc-head{display:flex;align-items:center;gap:10px;flex-wrap:wrap}
|
.kc-head{display:flex;align-items:center;gap:10px;flex-wrap:wrap}
|
||||||
/* key-wide caps sit inline in the head: they belong to the key, not to
|
|
||||||
any one model brick, and must not be draggable with one. */
|
|
||||||
.kc-caps{display:inline-flex;align-items:center;gap:6px;flex-wrap:wrap}
|
|
||||||
.kc-blocks{display:flex;flex-wrap:wrap;gap:10px;align-items:center;margin-top:12px;background:var(--card2);
|
.kc-blocks{display:flex;flex-wrap:wrap;gap:10px;align-items:center;margin-top:12px;background:var(--card2);
|
||||||
border:1px dashed var(--line);border-radius:12px;padding:14px;min-height:64px}
|
border:1px dashed var(--line);border-radius:12px;padding:14px;min-height:64px}
|
||||||
.kc-blocks.ovh{outline:2px dashed var(--primary);outline-offset:2px}
|
.kc-blocks.ovh{outline:2px dashed var(--primary);outline-offset:2px}
|
||||||
@ -821,15 +818,10 @@
|
|||||||
kAnySrc: "任意源",
|
kAnySrc: "任意源",
|
||||||
kQuotaB: "Token 配额",
|
kQuotaB: "Token 配额",
|
||||||
kQuotaHintB: "0 / 留空 = 无限",
|
kQuotaHintB: "0 / 留空 = 无限",
|
||||||
kKeyQuota: "密钥总配额",
|
kReqQuota: "请求数配额",
|
||||||
kKeyQuotaHint:
|
kReqQuotaHint: "限制周期内的请求次数。0 / 留空 = 无限。",
|
||||||
"限制这把密钥在重置周期内的总用量(跳模型)。0 / 留空 = 无限。",
|
kQuotaPerModelHint:
|
||||||
kKeyReqQuota: "请求数配额",
|
"配额按模型单独设置:创建后点「+ 添加模型」,逐个模型配 token 配额与周期。一个模型用满只影响该模型,同一密钥的其它模型照常。",
|
||||||
kKeyReqQuotaHint: "限制周期内的请求次数。0 / 留空 = 无限。",
|
|
||||||
kKeyQuotaAdmin:
|
|
||||||
"admin 密钥永不受配额限制(避免把管理员锁在门外)。",
|
|
||||||
kKeyQuotaEdit: "配额",
|
|
||||||
kKeyQuotaNone: "无限",
|
|
||||||
kPeriodB: "重置周期",
|
kPeriodB: "重置周期",
|
||||||
kPerNothing: "不限",
|
kPerNothing: "不限",
|
||||||
kPerHour: "每 小时",
|
kPerHour: "每 小时",
|
||||||
@ -1055,16 +1047,10 @@
|
|||||||
kAnySrc: "any source",
|
kAnySrc: "any source",
|
||||||
kQuotaB: "Token quota",
|
kQuotaB: "Token quota",
|
||||||
kQuotaHintB: "0 / empty = unlimited",
|
kQuotaHintB: "0 / empty = unlimited",
|
||||||
kKeyQuota: "Key-wide quota",
|
kReqQuota: "Request quota",
|
||||||
kKeyQuotaHint:
|
kReqQuotaHint: "Caps requests per window. 0 / empty = unlimited.",
|
||||||
"Caps this key's total spend per reset window, across every model it may use. 0 / empty = unlimited.",
|
kQuotaPerModelHint:
|
||||||
kKeyReqQuota: "Request quota",
|
"Quotas are per model: after creating the key, use \"Add model\" to give each model its own token budget and reset period. One model running out affects only that model; the key's other models keep working.",
|
||||||
kKeyReqQuotaHint:
|
|
||||||
"Caps requests per window. 0 / empty = unlimited.",
|
|
||||||
kKeyQuotaAdmin:
|
|
||||||
"Admin keys are never capped — a cap could lock the operator out.",
|
|
||||||
kKeyQuotaEdit: "Quota",
|
|
||||||
kKeyQuotaNone: "unlimited",
|
|
||||||
kPeriodB: "Reset period",
|
kPeriodB: "Reset period",
|
||||||
kPerNothing: "Never",
|
kPerNothing: "Never",
|
||||||
kPerHour: "Every hour",
|
kPerHour: "Every hour",
|
||||||
@ -4266,11 +4252,10 @@
|
|||||||
async function renderKeysUser(me) {
|
async function renderKeysUser(me) {
|
||||||
$("#tab-keys").innerHTML = `
|
$("#tab-keys").innerHTML = `
|
||||||
<div class="card"><h2>${t("kMeTitle")}</h2>
|
<div class="card"><h2>${t("kMeTitle")}</h2>
|
||||||
<div class="tbl-wrap"><table><tr><th>${t("kName")}</th><th>${t("kMeRole")}</th><th>${t("kKey")}</th><th>${t("kKeyQuota")}</th><th>${t("kMeModels")}</th></tr>
|
<div class="tbl-wrap"><table><tr><th>${t("kName")}</th><th>${t("kMeRole")}</th><th>${t("kKey")}</th><th>${t("kMeModels")}</th></tr>
|
||||||
<tr><td><b>${esc(me.name || "—")}</b></td><td>${roleTag(me.role)}</td>
|
<tr><td><b>${esc(me.name || "—")}</b></td><td>${roleTag(me.role)}</td>
|
||||||
<td><span class="kr-key">${esc(me.key)}</span>
|
<td><span class="kr-key">${esc(me.key)}</span>
|
||||||
<button class="ghost small" onclick="copyText('${escAttr(me.key)}')">${t("kCopy")}</button></td>
|
<button class="ghost small" onclick="copyText('${escAttr(me.key)}')">${t("kCopy")}</button></td>
|
||||||
<td>${keyCapBadges(me)}</td>
|
|
||||||
<td>${
|
<td>${
|
||||||
me.models && me.models.length
|
me.models && me.models.length
|
||||||
? me.models
|
? me.models
|
||||||
@ -4308,22 +4293,13 @@
|
|||||||
}
|
}
|
||||||
function keyCanvasHtml(k) {
|
function keyCanvasHtml(k) {
|
||||||
const scopes = k.models || [];
|
const scopes = k.models || [];
|
||||||
// Key-wide caps live on the canvas, not on a brick: they are a budget
|
|
||||||
// the whole key shares, so they must not be dragged around with one
|
|
||||||
// model. data-* carries them so a quota edit can round-trip them
|
|
||||||
// through the same PUT that saves the model scope.
|
|
||||||
const caps = `data-kquota="${k.token_quota || 0}" data-kreqquota="${k.req_quota || 0}"
|
|
||||||
data-kperiod="${escAttr(k.period || "")}" data-khours="${k.hours || 0}"`;
|
|
||||||
return `
|
return `
|
||||||
<div class="key-canvas" data-key="${escAttr(k.key)}" ${caps}>
|
<div class="key-canvas" data-key="${escAttr(k.key)}">
|
||||||
<div class="kc-head">
|
<div class="kc-head">
|
||||||
<b>${esc(k.name || "—")}</b>
|
<b>${esc(k.name || "—")}</b>
|
||||||
${roleTag(k.role)}
|
${roleTag(k.role)}
|
||||||
<span class="kr-key">${esc(maskKey(k.key))}</span>
|
<span class="kr-key">${esc(maskKey(k.key))}</span>
|
||||||
<button class="ghost small" onclick="copyText('${escAttr(k.key)}')">${t("kCopy")}</button>
|
<button class="ghost small" onclick="copyText('${escAttr(k.key)}')">${t("kCopy")}</button>
|
||||||
<span class="kc-caps" title="${escAttr(t("kKeyQuotaHint"))}">${keyCapBadges(k)}</span>
|
|
||||||
${k.role === "admin" ? "" : `<button class="ghost small" title="${escAttr(t("kKeyQuotaEdit"))}"
|
|
||||||
onclick="keyQuotaEdit('${escAttr(k.key)}')">${t("kKeyQuotaEdit")}</button>`}
|
|
||||||
<span class="grow"></span>
|
<span class="grow"></span>
|
||||||
<span class="muted">${fmtCreated(k.created_at)}</span>
|
<span class="muted">${fmtCreated(k.created_at)}</span>
|
||||||
<button class="ghost small errc" onclick="delKey('${escAttr(k.key)}','${escAttr(k.name || "")}')">${t("kDel")}</button>
|
<button class="ghost small errc" onclick="delKey('${escAttr(k.key)}','${escAttr(k.name || "")}')">${t("kDel")}</button>
|
||||||
@ -4336,39 +4312,19 @@
|
|||||||
</div>
|
</div>
|
||||||
</div>`;
|
</div>`;
|
||||||
}
|
}
|
||||||
// keyCapBadges renders the key-wide caps. A cap with no reset period is
|
|
||||||
// flagged as such, because "1M tokens, never resets" and "1M tokens per
|
|
||||||
// hour" are very different promises and the badge must not blur them.
|
|
||||||
function keyCapBadges(k) {
|
|
||||||
const out = [];
|
|
||||||
const suffix = periodText(k.period || "", k.hours || 0);
|
|
||||||
if (+k.token_quota > 0) {
|
|
||||||
out.push(
|
|
||||||
`<span class="mb-quota" title="${escAttr(t("kKeyQuotaHint"))}">${esc(fmtQuota(k.token_quota))}${esc(suffix)}</span>`,
|
|
||||||
);
|
|
||||||
}
|
|
||||||
if (+k.req_quota > 0) {
|
|
||||||
out.push(
|
|
||||||
`<span class="mb-quota" title="${escAttr(t("kKeyReqQuotaHint"))}">${esc(fmtQuota(k.req_quota))}×${esc(suffix)}</span>`,
|
|
||||||
);
|
|
||||||
}
|
|
||||||
if (!out.length) {
|
|
||||||
return `<span class="muted">${t("kKeyQuotaNone")}</span>`;
|
|
||||||
}
|
|
||||||
return out.join(" ");
|
|
||||||
}
|
|
||||||
function scopeHtml(key, m) {
|
function scopeHtml(key, m) {
|
||||||
const qt = fmtQuota(m.token_quota);
|
const qt = fmtQuota(m.token_quota);
|
||||||
const comb = scopeComb(m);
|
const comb = scopeComb(m);
|
||||||
const src = normSrc(m.source);
|
const src = normSrc(m.source);
|
||||||
const attrs = `data-key="${escAttr(key)}" data-model="${escAttr(comb)}"
|
const attrs = `data-key="${escAttr(key)}" data-model="${escAttr(comb)}"
|
||||||
data-quota="${m.token_quota || 0}" data-period="${escAttr(m.period || "")}" data-hours="${m.hours || 0}"`;
|
data-quota="${m.token_quota || 0}" data-reqquota="${m.req_quota || 0}"
|
||||||
|
data-period="${escAttr(m.period || "")}" data-hours="${m.hours || 0}"`;
|
||||||
return `<span class="mb" draggable="true" ${attrs} title="${escAttr(t("kBrickH"))}"
|
return `<span class="mb" draggable="true" ${attrs} title="${escAttr(t("kBrickH"))}"
|
||||||
onclick="scopeEdit('${escAttr(key)}','${escAttr(comb)}')"
|
onclick="scopeEdit('${escAttr(key)}','${escAttr(comb)}')"
|
||||||
oncontextmenu="scopeCtx(event,'${escAttr(key)}','${escAttr(comb)}')">
|
oncontextmenu="scopeCtx(event,'${escAttr(key)}','${escAttr(comb)}')">
|
||||||
<span class="mb-ico"><svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" width="10" height="10"><circle cx="12" cy="12" r="10"/></svg></span>
|
<span class="mb-ico"><svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" width="10" height="10"><circle cx="12" cy="12" r="10"/></svg></span>
|
||||||
<span class="mb-name">${esc(m.model)}${src ? `<em class="mb-src">${esc(src)}</em>` : ""}</span>
|
<span class="mb-name">${esc(m.model)}${src ? `<em class="mb-src">${esc(src)}</em>` : ""}</span>
|
||||||
<span class="mb-quota">${esc(quantBadge(m.token_quota, m.period, m.hours))}</span>
|
<span class="mb-quota">${esc(scopeQuotaBadge(m))}</span>
|
||||||
<span class="copy-b" title="${escAttr(t("kCopyB"))}" onclick="event.stopPropagation();scopeDup('${escAttr(key)}','${escAttr(comb)}')"><svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" width="12" height="12"><rect x="9" y="9" width="13" height="13" rx="2" ry="2"/><path d="M5 15H4a2 2 0 0 1-2-2V4a2 2 0 0 1 2-2h9a2 2 0 0 1 2 2v1"/></svg></span>
|
<span class="copy-b" title="${escAttr(t("kCopyB"))}" onclick="event.stopPropagation();scopeDup('${escAttr(key)}','${escAttr(comb)}')"><svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" width="12" height="12"><rect x="9" y="9" width="13" height="13" rx="2" ry="2"/><path d="M5 15H4a2 2 0 0 1-2-2V4a2 2 0 0 1 2-2h9a2 2 0 0 1 2 2v1"/></svg></span>
|
||||||
<span class="mb-x" title="${escAttr(t("kDelB"))}" onclick="event.stopPropagation();scopeRm('${escAttr(key)}','${escAttr(comb)}')"><svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" width="12" height="12"><line x1="18" y1="6" x2="6" y2="18"/><line x1="6" y1="6" x2="18" y2="18"/></svg></span>
|
<span class="mb-x" title="${escAttr(t("kDelB"))}" onclick="event.stopPropagation();scopeRm('${escAttr(key)}','${escAttr(comb)}')"><svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" width="12" height="12"><line x1="18" y1="6" x2="6" y2="18"/><line x1="6" y1="6" x2="18" y2="18"/></svg></span>
|
||||||
</span>`;
|
</span>`;
|
||||||
@ -4400,6 +4356,16 @@
|
|||||||
if (p === "nhour") return "·" + Math.max(1, h) + "h";
|
if (p === "nhour") return "·" + Math.max(1, h) + "h";
|
||||||
return "";
|
return "";
|
||||||
}
|
}
|
||||||
|
// scopeQuotaBadge shows a scope entry's budgets: token quota and, when
|
||||||
|
// set, the request count. They are per model — a spent budget blocks
|
||||||
|
// only that model, not the whole key.
|
||||||
|
function scopeQuotaBadge(m) {
|
||||||
|
const parts = [];
|
||||||
|
if (+m.token_quota > 0) parts.push(fmtQuota(m.token_quota));
|
||||||
|
if (+m.req_quota > 0) parts.push(fmtQuota(m.req_quota) + "\u00d7");
|
||||||
|
if (!parts.length) return "\u221e";
|
||||||
|
return parts.join(" ") + periodText(m.period, m.hours);
|
||||||
|
}
|
||||||
function quantBadge(quota, period, hours) {
|
function quantBadge(quota, period, hours) {
|
||||||
quota = +quota || 0;
|
quota = +quota || 0;
|
||||||
period = period || "";
|
period = period || "";
|
||||||
@ -4414,123 +4380,23 @@
|
|||||||
model: comb[0],
|
model: comb[0],
|
||||||
source: src || undefined,
|
source: src || undefined,
|
||||||
token_quota: parseInt(b.dataset.quota) || 0,
|
token_quota: parseInt(b.dataset.quota) || 0,
|
||||||
|
req_quota: parseInt(b.dataset.reqquota) || 0,
|
||||||
period: b.dataset.period || "",
|
period: b.dataset.period || "",
|
||||||
hours: parseInt(b.dataset.hours) || 0,
|
hours: parseInt(b.dataset.hours) || 0,
|
||||||
};
|
};
|
||||||
});
|
});
|
||||||
}
|
}
|
||||||
async function putScope(key, scopes) {
|
async function putScope(key, scopes) {
|
||||||
// The key-wide caps ride along with every scope write. The API reads
|
// Each scope entry carries its own quotas, so the whole budget travels
|
||||||
// them as pointers, so sending them back unchanged is a no-op, while
|
// with the models it applies to. There is no separate key-wide total
|
||||||
// omitting them would be indistinguishable from "clear the budget" to
|
// that could drift out of sync with the model list.
|
||||||
// a future reader. Round-tripping them here means editing a model's
|
|
||||||
// scope can never silently drop a key's quota.
|
|
||||||
const canvas = document.querySelector(
|
|
||||||
`.key-canvas[data-key="${CSS.escape(key)}"]`,
|
|
||||||
);
|
|
||||||
const body = { models: scopes };
|
const body = { models: scopes };
|
||||||
if (canvas) {
|
|
||||||
body.token_quota = parseInt(canvas.dataset.kquota) || 0;
|
|
||||||
body.req_quota = parseInt(canvas.dataset.kreqquota) || 0;
|
|
||||||
body.period = canvas.dataset.kperiod || "";
|
|
||||||
body.hours = parseInt(canvas.dataset.khours) || 0;
|
|
||||||
}
|
|
||||||
await api("/api/keys/" + encodeURIComponent(key), {
|
await api("/api/keys/" + encodeURIComponent(key), {
|
||||||
method: "PUT",
|
method: "PUT",
|
||||||
headers: { "Content-Type": "application/json" },
|
headers: { "Content-Type": "application/json" },
|
||||||
body: JSON.stringify(body),
|
body: JSON.stringify(body),
|
||||||
});
|
});
|
||||||
}
|
}
|
||||||
// keyQuotaEdit opens the key-wide budget form.
|
|
||||||
function keyQuotaEdit(key) {
|
|
||||||
const canvas = document.querySelector(
|
|
||||||
`.key-canvas[data-key="${CSS.escape(key)}"]`,
|
|
||||||
);
|
|
||||||
if (!canvas) return;
|
|
||||||
const cur = {
|
|
||||||
token_quota: parseInt(canvas.dataset.kquota) || 0,
|
|
||||||
req_quota: parseInt(canvas.dataset.kreqquota) || 0,
|
|
||||||
period: canvas.dataset.kperiod || "",
|
|
||||||
hours: parseInt(canvas.dataset.khours) || 0,
|
|
||||||
};
|
|
||||||
const wrap = document.createElement("div");
|
|
||||||
wrap.id = "modal-wrap";
|
|
||||||
wrap.style.cssText =
|
|
||||||
"position:fixed;inset:0;background:rgba(15,22,44,.45);display:flex;align-items:flex-start;justify-content:center;overflow:auto;padding:48px 20px;z-index:50";
|
|
||||||
wrap.innerHTML = `<div class="card" style="width:400px;max-width:100%"><h2>${t("kKeyQuotaEdit")}</h2>
|
|
||||||
<label>${t("kKeyQuota")} <span class="muted">${t("kKeyQuotaHint")}</span></label>
|
|
||||||
<input id="kq-tokens" type="number" min="0" step="1"
|
|
||||||
placeholder="${escAttr(t("kKeyQuotaNone"))}" value="${cur.token_quota || ""}">
|
|
||||||
<label>${t("kKeyReqQuota")} <span class="muted">${t("kKeyReqQuotaHint")}</span></label>
|
|
||||||
<input id="kq-reqs" type="number" min="0" step="1"
|
|
||||||
placeholder="${escAttr(t("kKeyQuotaNone"))}" value="${cur.req_quota || ""}">
|
|
||||||
<label>${t("kPeriodB")}</label>
|
|
||||||
<select id="kq-period">
|
|
||||||
<option value="" ${!cur.period ? "selected" : ""}>${t("kPerNothing")}</option>
|
|
||||||
<option value="hour" ${cur.period === "hour" ? "selected" : ""}>${t("kPerHour")}</option>
|
|
||||||
<option value="week" ${cur.period === "week" ? "selected" : ""}>${t("kPerWeek")}</option>
|
|
||||||
<option value="month" ${cur.period === "month" ? "selected" : ""}>${t("kPerMonth")}</option>
|
|
||||||
<option value="nhour" ${cur.period === "nhour" ? "selected" : ""}>${t("kPerHours")}</option>
|
|
||||||
</select>
|
|
||||||
<div id="kq-hours-box" style="display:none"><label>${t("kPerNHint")}</label>
|
|
||||||
<input id="kq-hours" type="number" min="1" step="1" value="${cur.hours || 24}"></div>
|
|
||||||
<p class="muted" style="font-size:12px">${t("kKeyQuotaAdmin")}</p>
|
|
||||||
<p><button onclick="keyQuotaSave('${escAttr(key)}', this)">${t("kSaveScope")}</button>
|
|
||||||
<button class="ghost" onclick="this.closest('#modal-wrap').remove()">${t("mCancel")}</button></p>
|
|
||||||
</div>`;
|
|
||||||
document.body.appendChild(wrap);
|
|
||||||
const toggle = () => {
|
|
||||||
$("#kq-hours-box").style.display =
|
|
||||||
$("#kq-period").value === "nhour" ? "block" : "none";
|
|
||||||
};
|
|
||||||
$("#kq-period").addEventListener("change", toggle);
|
|
||||||
toggle();
|
|
||||||
$("#kq-tokens").focus();
|
|
||||||
}
|
|
||||||
async function keyQuotaSave(key, btn) {
|
|
||||||
// Resolve our own dialog from the button that was clicked, so closing
|
|
||||||
// it can never remove a different #modal-wrap that happens to come
|
|
||||||
// first in the document.
|
|
||||||
const wrap = btn ? btn.closest("#modal-wrap") : null;
|
|
||||||
let tokens = parseInt($("#kq-tokens").value);
|
|
||||||
if (isNaN(tokens) || tokens < 0) tokens = 0;
|
|
||||||
let reqs = parseInt($("#kq-reqs").value);
|
|
||||||
if (isNaN(reqs) || reqs < 0) reqs = 0;
|
|
||||||
let hours = parseInt($("#kq-hours").value);
|
|
||||||
if (isNaN(hours) || hours < 1) hours = 1;
|
|
||||||
const period = $("#kq-period").value;
|
|
||||||
// Catch the "nhour picked but hours never filled in" case locally: the
|
|
||||||
// API rejects it too, but a round trip for a form-level mistake is
|
|
||||||
// needless.
|
|
||||||
if (period === "nhour" && hours < 1) {
|
|
||||||
toast(t("kPerNHint"));
|
|
||||||
return;
|
|
||||||
}
|
|
||||||
if (btn) btn.disabled = true;
|
|
||||||
try {
|
|
||||||
await api("/api/keys/" + encodeURIComponent(key), {
|
|
||||||
method: "PUT",
|
|
||||||
headers: { "Content-Type": "application/json" },
|
|
||||||
body: JSON.stringify({
|
|
||||||
token_quota: tokens,
|
|
||||||
req_quota: reqs,
|
|
||||||
period,
|
|
||||||
hours,
|
|
||||||
}),
|
|
||||||
});
|
|
||||||
// Close THIS modal, not whichever #modal-wrap comes first in the
|
|
||||||
// document: another dialog (e.g. the seed-key notice) may already be
|
|
||||||
// open, and a bare $("#modal-wrap") would remove that one and leave
|
|
||||||
// this form stranded on screen.
|
|
||||||
if (wrap) wrap.remove();
|
|
||||||
else closeTopModal();
|
|
||||||
toast(t("kSaved"));
|
|
||||||
await loadKeys();
|
|
||||||
} catch (e) {
|
|
||||||
toast(e.message);
|
|
||||||
if (btn) btn.disabled = false;
|
|
||||||
}
|
|
||||||
}
|
|
||||||
async function scopePush(key) {
|
async function scopePush(key) {
|
||||||
const canvas = document.querySelector(
|
const canvas = document.querySelector(
|
||||||
`.key-canvas[data-key="${CSS.escape(key)}"]`,
|
`.key-canvas[data-key="${CSS.escape(key)}"]`,
|
||||||
@ -4631,6 +4497,9 @@
|
|||||||
<label>${t("kQuotaB")} <span class="muted">${t("kQuotaHintB")}</span></label>
|
<label>${t("kQuotaB")} <span class="muted">${t("kQuotaHintB")}</span></label>
|
||||||
<input id="sc-quota" type="number" min="0" step="1"
|
<input id="sc-quota" type="number" min="0" step="1"
|
||||||
placeholder="${escAttr(t("kQuotaHintB"))}" value="${sc.token_quota ? sc.token_quota : ""}">
|
placeholder="${escAttr(t("kQuotaHintB"))}" value="${sc.token_quota ? sc.token_quota : ""}">
|
||||||
|
<label>${t("kReqQuota")} <span class="muted">${t("kReqQuotaHint")}</span></label>
|
||||||
|
<input id="sc-reqs" type="number" min="0" step="1"
|
||||||
|
placeholder="${escAttr(t("kReqQuotaHint"))}" value="${sc.req_quota || ""}">
|
||||||
<label>${t("kPeriodB")}</label>
|
<label>${t("kPeriodB")}</label>
|
||||||
<select id="sc-period">
|
<select id="sc-period">
|
||||||
<option value="" ${!sc.period ? "selected" : ""}>${t("kPerNothing")}</option>
|
<option value="" ${!sc.period ? "selected" : ""}>${t("kPerNothing")}</option>
|
||||||
@ -4664,6 +4533,8 @@
|
|||||||
const parts = splitCombKey(comb);
|
const parts = splitCombKey(comb);
|
||||||
let q = parseInt($("#sc-quota").value);
|
let q = parseInt($("#sc-quota").value);
|
||||||
if (isNaN(q) || q < 0) q = 0;
|
if (isNaN(q) || q < 0) q = 0;
|
||||||
|
let rq = parseInt($("#sc-reqs").value);
|
||||||
|
if (isNaN(rq) || rq < 0) rq = 0;
|
||||||
let hours = parseInt($("#sc-hours").value);
|
let hours = parseInt($("#sc-hours").value);
|
||||||
if (isNaN(hours) || hours < 1) hours = 1;
|
if (isNaN(hours) || hours < 1) hours = 1;
|
||||||
const period = $("#sc-period").value;
|
const period = $("#sc-period").value;
|
||||||
@ -4672,6 +4543,7 @@
|
|||||||
model: parts.model,
|
model: parts.model,
|
||||||
source: parts.source || undefined,
|
source: parts.source || undefined,
|
||||||
token_quota: q,
|
token_quota: q,
|
||||||
|
req_quota: rq,
|
||||||
period,
|
period,
|
||||||
hours,
|
hours,
|
||||||
};
|
};
|
||||||
@ -4832,41 +4704,13 @@
|
|||||||
<option value="user">${t("kRoleUser")}</option>
|
<option value="user">${t("kRoleUser")}</option>
|
||||||
<option value="admin">${t("kRoleAdmin")}</option>
|
<option value="admin">${t("kRoleAdmin")}</option>
|
||||||
</select>
|
</select>
|
||||||
<label>${t("kKeyQuota")} <span class="muted">${t("kKeyQuotaHint")}</span></label>
|
<p class="muted" style="font-size:12px">${t("kQuotaPerModelHint")}</p>
|
||||||
<input id="kc-tokens" type="number" min="0" step="1" placeholder="${escAttr(t("kKeyQuotaNone"))}">
|
|
||||||
<label>${t("kKeyReqQuota")} <span class="muted">${t("kKeyReqQuotaHint")}</span></label>
|
|
||||||
<input id="kc-reqs" type="number" min="0" step="1" placeholder="${escAttr(t("kKeyQuotaNone"))}">
|
|
||||||
<label>${t("kPeriodB")}</label>
|
|
||||||
<select id="kc-period">
|
|
||||||
<option value="" selected>${t("kPerNothing")}</option>
|
|
||||||
<option value="hour">${t("kPerHour")}</option>
|
|
||||||
<option value="week">${t("kPerWeek")}</option>
|
|
||||||
<option value="month">${t("kPerMonth")}</option>
|
|
||||||
<option value="nhour">${t("kPerHours")}</option>
|
|
||||||
</select>
|
|
||||||
<div id="kc-hours-box" style="display:none"><label>${t("kPerNHint")}</label>
|
|
||||||
<input id="kc-hours" type="number" min="1" step="1" value="24"></div>
|
|
||||||
<label>${t("kNote")}</label>
|
<label>${t("kNote")}</label>
|
||||||
<input id="kc-note">
|
<input id="kc-note">
|
||||||
<p class="muted" style="font-size:12px">${t("kKeyQuotaAdmin")}</p>
|
|
||||||
<p><button onclick="createKey(this)">${t("kCreateBtn")}</button>
|
<p><button onclick="createKey(this)">${t("kCreateBtn")}</button>
|
||||||
<button class="ghost" onclick="this.closest('#modal-wrap').remove()">${t("mCancel")}</button></p>
|
<button class="ghost" onclick="this.closest('#modal-wrap').remove()">${t("mCancel")}</button></p>
|
||||||
</div>`;
|
</div>`;
|
||||||
document.body.appendChild(wrap);
|
document.body.appendChild(wrap);
|
||||||
$("#kc-role").addEventListener("change", () => {
|
|
||||||
const admin = $("#kc-role").value === "admin";
|
|
||||||
// An admin key ignores its caps server-side; hiding the fields
|
|
||||||
// avoids the operator setting one and wondering why it never trips.
|
|
||||||
$("#kc-tokens").disabled = admin;
|
|
||||||
$("#kc-reqs").disabled = admin;
|
|
||||||
$("#kc-period").disabled = admin;
|
|
||||||
$("#kc-hours-box").style.display =
|
|
||||||
!admin && $("#kc-period").value === "nhour" ? "block" : "none";
|
|
||||||
});
|
|
||||||
$("#kc-period").addEventListener("change", () => {
|
|
||||||
$("#kc-hours-box").style.display =
|
|
||||||
$("#kc-period").value === "nhour" ? "block" : "none";
|
|
||||||
});
|
|
||||||
$("#kc-name").focus();
|
$("#kc-name").focus();
|
||||||
}
|
}
|
||||||
async function createKey(btn) {
|
async function createKey(btn) {
|
||||||
@ -4875,17 +4719,6 @@
|
|||||||
toast(t("kName"));
|
toast(t("kName"));
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
const role = $("#kc-role").value;
|
|
||||||
// An admin key is never capped; send the fields anyway (the server
|
|
||||||
// ignores them) rather than special-casing the request shape.
|
|
||||||
const readNum = (sel) => {
|
|
||||||
const el = $(sel);
|
|
||||||
if (el.disabled) return 0;
|
|
||||||
const n = parseInt(el.value);
|
|
||||||
return isNaN(n) || n < 0 ? 0 : n;
|
|
||||||
};
|
|
||||||
let hours = parseInt($("#kc-hours").value);
|
|
||||||
if (isNaN(hours) || hours < 1) hours = 1;
|
|
||||||
if (btn) btn.disabled = true;
|
if (btn) btn.disabled = true;
|
||||||
let j;
|
let j;
|
||||||
try {
|
try {
|
||||||
@ -4894,12 +4727,8 @@
|
|||||||
headers: { "Content-Type": "application/json" },
|
headers: { "Content-Type": "application/json" },
|
||||||
body: JSON.stringify({
|
body: JSON.stringify({
|
||||||
name,
|
name,
|
||||||
role,
|
role: $("#kc-role").value,
|
||||||
note: $("#kc-note").value.trim(),
|
note: $("#kc-note").value.trim(),
|
||||||
token_quota: readNum("#kc-tokens"),
|
|
||||||
req_quota: readNum("#kc-reqs"),
|
|
||||||
period: role === "admin" ? "" : $("#kc-period").value,
|
|
||||||
hours: role === "admin" ? 0 : hours,
|
|
||||||
}),
|
}),
|
||||||
});
|
});
|
||||||
} catch (e) {
|
} catch (e) {
|
||||||
|
|||||||
@ -13,9 +13,9 @@ import (
|
|||||||
// form the user just submitted stays on screen while an unrelated dialog
|
// form the user just submitted stays on screen while an unrelated dialog
|
||||||
// vanishes.
|
// vanishes.
|
||||||
//
|
//
|
||||||
// This is exactly the class of bug the api() contract test below was written
|
// This is the same class of bug the api() contract test pins: reviewing inline
|
||||||
// for: reviewing inline JS by eye does not catch it, and the visible symptom
|
// JS by eye does not catch it, and the symptom ("the dialog did not close")
|
||||||
// ("the dialog did not close") points away from the cause. It is pinned here.
|
// points away from the cause.
|
||||||
|
|
||||||
// modalCloseRe finds every `$(...)`-style lookup of the shared modal id.
|
// modalCloseRe finds every `$(...)`-style lookup of the shared modal id.
|
||||||
var modalCloseRe = regexp.MustCompile(`\$\("#modal-wrap"\)`)
|
var modalCloseRe = regexp.MustCompile(`\$\("#modal-wrap"\)`)
|
||||||
@ -25,10 +25,9 @@ var modalCloseRe = regexp.MustCompile(`\$\("#modal-wrap"\)`)
|
|||||||
var closestModalRe = regexp.MustCompile(`\.closest\("#modal-wrap"\)`)
|
var closestModalRe = regexp.MustCompile(`\.closest\("#modal-wrap"\)`)
|
||||||
|
|
||||||
func TestUIDialogClosesItselfNotTheFirstModal(t *testing.T) {
|
func TestUIDialogClosesItselfNotTheFirstModal(t *testing.T) {
|
||||||
src := uiSource(t)
|
|
||||||
// Strip comments first: prose that *names* the unsafe pattern (as the fix's
|
// Strip comments first: prose that *names* the unsafe pattern (as the fix's
|
||||||
// own comment does) would otherwise be flagged as a violation.
|
// own comment does) would otherwise be flagged as a violation.
|
||||||
code := stripJSComments(src)
|
code := stripJSComments(uiSource(t))
|
||||||
|
|
||||||
for _, m := range modalCloseRe.FindAllStringIndex(code, -1) {
|
for _, m := range modalCloseRe.FindAllStringIndex(code, -1) {
|
||||||
after := code[m[1]:]
|
after := code[m[1]:]
|
||||||
@ -45,15 +44,45 @@ func TestUIDialogClosesItselfNotTheFirstModal(t *testing.T) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
// TestUIDialogClosuresGoThroughSafePaths pins the rule across the whole
|
// stripJSComments removes // line comments and /* block */ comments from JS
|
||||||
// document by data flow rather than by pattern: every handler that closes a
|
// embedded in the UI document. It is deliberately simple: the document is our
|
||||||
// dialog must do it one of the two safe ways. A handler could contain a
|
// own source, and a false negative here only means the check is silent.
|
||||||
// correct .closest() and still close the wrong dialog on another path.
|
func stripJSComments(src string) string {
|
||||||
|
var out strings.Builder
|
||||||
|
lines := strings.Split(src, "\n")
|
||||||
|
inBlock := false
|
||||||
|
for _, ln := range lines {
|
||||||
|
trimmed := strings.TrimSpace(ln)
|
||||||
|
if inBlock {
|
||||||
|
if strings.Contains(ln, "*/") {
|
||||||
|
inBlock = false
|
||||||
|
}
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
if strings.HasPrefix(trimmed, "/*") {
|
||||||
|
if !strings.Contains(ln, "*/") {
|
||||||
|
inBlock = true
|
||||||
|
}
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
if i := strings.Index(ln, "//"); i >= 0 {
|
||||||
|
before := ln[:i]
|
||||||
|
if strings.Count(before, `"`)%2 == 0 && strings.Count(before, "'")%2 == 0 {
|
||||||
|
ln = before
|
||||||
|
}
|
||||||
|
}
|
||||||
|
out.WriteString(ln)
|
||||||
|
out.WriteString("\n")
|
||||||
|
}
|
||||||
|
return out.String()
|
||||||
|
}
|
||||||
|
|
||||||
|
// Every handler that closes a dialog must do it one of the two safe ways.
|
||||||
func TestUIDialogClosuresGoThroughSafePaths(t *testing.T) {
|
func TestUIDialogClosuresGoThroughSafePaths(t *testing.T) {
|
||||||
src := stripJSComments(uiSource(t))
|
src := stripJSComments(uiSource(t))
|
||||||
for _, fn := range []string{
|
for _, fn := range []string{
|
||||||
"downloadStatsCsv", "downloadKeysCsv", "saveSource", "saveTemplate",
|
"downloadStatsCsv", "downloadKeysCsv", "saveSource", "saveTemplate",
|
||||||
"scrAddFromForm", "sortScopeSave", "scopeSave", "keyQuotaSave", "createKey",
|
"scrAddFromForm", "sortScopeSave", "scopeSave", "createKey",
|
||||||
} {
|
} {
|
||||||
body, ok := jsFunctionBody(src, fn)
|
body, ok := jsFunctionBody(src, fn)
|
||||||
if !ok {
|
if !ok {
|
||||||
@ -83,12 +112,12 @@ func TestUICloseTopModalTakesTheLast(t *testing.T) {
|
|||||||
|
|
||||||
// Handlers that resolve their dialog from a button must actually receive one:
|
// Handlers that resolve their dialog from a button must actually receive one:
|
||||||
// a signature without the parameter means the .closest() silently yields null
|
// a signature without the parameter means the .closest() silently yields null
|
||||||
// and the save leaves its form stranded on screen.
|
// and the save leaves its form stranded.
|
||||||
func TestUIDialogHandlersReceiveTheirButton(t *testing.T) {
|
func TestUIDialogHandlersReceiveTheirButton(t *testing.T) {
|
||||||
src := stripJSComments(uiSource(t))
|
src := stripJSComments(uiSource(t))
|
||||||
for _, fn := range []string{
|
for _, fn := range []string{
|
||||||
"downloadStatsCsv", "saveSource", "scrAddFromForm", "sortScopeSave",
|
"downloadStatsCsv", "saveSource", "scrAddFromForm", "sortScopeSave",
|
||||||
"scopeSave", "keyQuotaSave", "createKey",
|
"scopeSave", "createKey",
|
||||||
} {
|
} {
|
||||||
body, ok := jsFunctionBody(src, fn)
|
body, ok := jsFunctionBody(src, fn)
|
||||||
if !ok {
|
if !ok {
|
||||||
@ -106,165 +135,133 @@ func TestUIDialogHandlersReceiveTheirButton(t *testing.T) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
// stripJSComments removes // line comments and /* block */ comments from JS
|
// Quotas belong to the scope entries, so putScope only has to ship the scope
|
||||||
// embedded in the UI document. It is deliberately simple (no string/regex
|
// list — the caps travel inside it. What must NOT come back is a key-wide
|
||||||
// awareness beyond skipping quoted spans on the same line): the document is
|
// total: it would be a second budget able to drift out of sync with the models
|
||||||
// our own source, and a false negative here only means the check is silent.
|
// it is supposed to cover.
|
||||||
func stripJSComments(src string) string {
|
func TestUIPutScopeShipsOnlyScopeQuotas(t *testing.T) {
|
||||||
var out strings.Builder
|
|
||||||
lines := strings.Split(src, "\n")
|
|
||||||
inBlock := false
|
|
||||||
for _, ln := range lines {
|
|
||||||
trimmed := strings.TrimSpace(ln)
|
|
||||||
if inBlock {
|
|
||||||
if strings.Contains(ln, "*/") {
|
|
||||||
inBlock = false
|
|
||||||
}
|
|
||||||
continue
|
|
||||||
}
|
|
||||||
if strings.HasPrefix(trimmed, "/*") {
|
|
||||||
if !strings.Contains(ln, "*/") {
|
|
||||||
inBlock = true
|
|
||||||
}
|
|
||||||
continue
|
|
||||||
}
|
|
||||||
if i := strings.Index(ln, "//"); i >= 0 {
|
|
||||||
// keep code before the comment when the // is not inside a string
|
|
||||||
before := ln[:i]
|
|
||||||
if strings.Count(before, `"`)%2 == 0 && strings.Count(before, "'")%2 == 0 {
|
|
||||||
ln = before
|
|
||||||
}
|
|
||||||
}
|
|
||||||
out.WriteString(ln)
|
|
||||||
out.WriteString("\n")
|
|
||||||
}
|
|
||||||
return out.String()
|
|
||||||
}
|
|
||||||
|
|
||||||
// TestUIKeyQuotaDialogsResolveOwnModal pins the dialog-closing rule for the
|
|
||||||
// two forms this change added.
|
|
||||||
func TestUIKeyQuotaDialogsResolveOwnModal(t *testing.T) {
|
|
||||||
src := uiSource(t)
|
|
||||||
for _, fn := range []string{"keyQuotaSave", "createKey"} {
|
|
||||||
body, ok := jsFunctionBody(src, fn)
|
|
||||||
if !ok {
|
|
||||||
t.Errorf("%s not found in the UI source", fn)
|
|
||||||
continue
|
|
||||||
}
|
|
||||||
if !closestModalRe.MatchString(body) {
|
|
||||||
t.Errorf("%s does not resolve its own dialog via .closest(\"#modal-wrap\");\n"+
|
|
||||||
"with another dialog open it would close that one instead and leave this form stranded", fn)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
// The quota editor must read and write the key-wide caps, and putScope must
|
|
||||||
// carry them along: the API treats the quota fields as pointers, so dropping
|
|
||||||
// them on a scope-only write is indistinguishable from "clear the budget".
|
|
||||||
func TestUIPutScopeCarriesKeyQuota(t *testing.T) {
|
|
||||||
body, ok := jsFunctionBody(uiSource(t), "putScope")
|
body, ok := jsFunctionBody(uiSource(t), "putScope")
|
||||||
if !ok {
|
if !ok {
|
||||||
t.Fatal("putScope not found")
|
t.Fatal("putScope not found")
|
||||||
}
|
}
|
||||||
// Scan the code with comments removed, or a comment that merely *names* a
|
|
||||||
// field would satisfy the check while the field is never sent.
|
|
||||||
code := stripJSComments(body)
|
code := stripJSComments(body)
|
||||||
for _, field := range []string{"token_quota", "req_quota", "period", "hours"} {
|
if !strings.Contains(code, "models:") || !strings.Contains(code, "scopes") {
|
||||||
if !strings.Contains(code, field) {
|
t.Error("putScope must ship the scope list the caps live in")
|
||||||
t.Errorf("putScope does not send %q — editing a model scope would clear the key's quota", field)
|
}
|
||||||
|
for _, gone := range []string{"kquota", "kreqquota", "kperiod", "khours"} {
|
||||||
|
if strings.Contains(code, gone) {
|
||||||
|
t.Errorf("putScope still references the removed key-wide quota field %q", gone)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
// A key's caps are rendered from the API record and shown on the canvas, so
|
// A model brick carries its own budgets, and the scope editor reads and writes
|
||||||
// the badge and the data attributes must not drift from the field names. The
|
// both of them: dropping req_quota on the round trip would silently lift a
|
||||||
// create form must send them too, or a key would only be cappable after an
|
// request cap every time someone edited a token cap.
|
||||||
// extra round of edits.
|
func TestUIScopeEditorRoundTripsBothQuotas(t *testing.T) {
|
||||||
func TestUIKeyQuotaRendersFromAPIFields(t *testing.T) {
|
src := stripJSComments(uiSource(t))
|
||||||
|
for _, fn := range []string{"scopeHtml", "readScopes", "scopeEdit", "scopeSave", "scopeQuotaBadge"} {
|
||||||
|
if _, ok := jsFunctionBody(src, fn); !ok {
|
||||||
|
t.Errorf("%s not found in the UI source", fn)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
for _, fn := range []string{"scopeHtml", "readScopes", "scopeSave", "scopeQuotaBadge"} {
|
||||||
|
body, ok := jsFunctionBody(src, fn)
|
||||||
|
if !ok {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
if !strings.Contains(body, "req_quota") && !strings.Contains(body, "reqquota") {
|
||||||
|
t.Errorf("%s does not carry req_quota — a request cap would be lost on edit", fn)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if form, ok := jsFunctionBody(src, "scopeEdit"); ok && !strings.Contains(form, "sc-reqs") {
|
||||||
|
t.Error("the scope editor has no request-quota input")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// The whole key-wide quota surface must be gone from the UI: a badge or a
|
||||||
|
// button reading a field the server no longer has would render "undefined" or
|
||||||
|
// silently do nothing.
|
||||||
|
func TestUIHasNoKeyWideQuotaSurface(t *testing.T) {
|
||||||
src := uiSource(t)
|
src := uiSource(t)
|
||||||
for _, token := range []string{
|
for _, gone := range []string{
|
||||||
"keyCapBadges", // shared renderer
|
"keyCapBadges", "keyQuotaEdit", "keyQuotaSave",
|
||||||
"kq-tokens", "kq-reqs", "kq-period", "kq-hours", // editor fields
|
"kq-tokens", "kq-reqs", "kq-period", "kq-hours",
|
||||||
"kc-tokens", "kc-reqs", "kc-period", "kc-hours", // create form fields
|
"kc-tokens", "kc-reqs", "kc-period", "kc-hours",
|
||||||
|
"data-kquota", "data-kreqquota", "data-kperiod", "data-khours",
|
||||||
} {
|
} {
|
||||||
if !strings.Contains(src, token) {
|
if strings.Contains(src, gone) {
|
||||||
t.Errorf("UI never references %q — the quota form is not wired up", token)
|
t.Errorf("UI still references the removed key-wide quota surface %q", gone)
|
||||||
}
|
|
||||||
}
|
|
||||||
// the create request must actually carry the caps
|
|
||||||
full, ok := jsFunctionBody(src, "createKey")
|
|
||||||
if !ok {
|
|
||||||
t.Fatal("createKey not found")
|
|
||||||
}
|
|
||||||
body := stripJSComments(full)
|
|
||||||
for _, field := range []string{"token_quota", "req_quota", "period"} {
|
|
||||||
if !strings.Contains(body, field) {
|
|
||||||
t.Errorf("createKey does not send %q — a new key could never be created with a budget", field)
|
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
// Existence of the strings is not enough: the badge has to READ the API
|
// Existence of the strings is not enough: the badge has to READ the scope's
|
||||||
// fields, and the editor has to read the canvas data attributes it writes.
|
// fields, and the editor has to read back what the brick writes. A field can
|
||||||
// A field can be present in the source and still never reach the screen —
|
// be present in the source and still never reach the screen — e.g. left in a
|
||||||
// e.g. left in a dead branch, or read from a name the writer never sets.
|
// dead branch, or read from a data attribute the writer never sets.
|
||||||
func TestUIKeyQuotaDataflowIsLive(t *testing.T) {
|
func TestUIKeyQuotaDataflowIsLive(t *testing.T) {
|
||||||
src := uiSource(t)
|
src := stripJSComments(uiSource(t))
|
||||||
|
|
||||||
badge, ok := jsFunctionBody(src, "keyCapBadges")
|
badge, ok := jsFunctionBody(src, "scopeQuotaBadge")
|
||||||
if !ok {
|
if !ok {
|
||||||
t.Fatal("keyCapBadges not found")
|
t.Fatal("scopeQuotaBadge not found")
|
||||||
}
|
}
|
||||||
badgeCode := stripJSComments(badge)
|
for _, field := range []string{"m.token_quota", "m.req_quota", "m.period"} {
|
||||||
for _, field := range []string{"k.token_quota", "k.req_quota", "k.period"} {
|
if !strings.Contains(badge, field) {
|
||||||
if !strings.Contains(badgeCode, field) {
|
t.Errorf("scopeQuotaBadge does not read %q — the cap would never show on the model brick", field)
|
||||||
t.Errorf("keyCapBadges does not read %q — the cap would never show on the key card", field)
|
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
// the editor must read back what keyCanvasHtml wrote
|
brick, ok := jsFunctionBody(src, "scopeHtml")
|
||||||
canvas, ok := jsFunctionBody(src, "keyCanvasHtml")
|
|
||||||
if !ok {
|
if !ok {
|
||||||
t.Fatal("keyCanvasHtml not found")
|
t.Fatal("scopeHtml not found")
|
||||||
}
|
}
|
||||||
editor, ok := jsFunctionBody(src, "keyQuotaEdit")
|
edit, ok := jsFunctionBody(src, "scopeEdit")
|
||||||
if !ok {
|
if !ok {
|
||||||
t.Fatal("keyQuotaEdit not found")
|
t.Fatal("scopeEdit not found")
|
||||||
}
|
}
|
||||||
canvasCode, editorCode := stripJSComments(canvas), stripJSComments(editor)
|
brickCode, editCode := stripJSComments(brick), stripJSComments(edit)
|
||||||
for _, ds := range []string{"kquota", "kreqquota", "kperiod", "khours"} {
|
reader := stripJSComments(mustBody(t, src, "readScopes"))
|
||||||
// written as data-<name> on the canvas
|
for _, ds := range []string{"quota", "reqquota", "period", "hours"} {
|
||||||
if !strings.Contains(canvasCode, "data-"+ds+"=") {
|
if !strings.Contains(brickCode, "data-"+ds+"=") {
|
||||||
t.Errorf("keyCanvasHtml does not write data-%s, so the editor has nothing to prefill", ds)
|
t.Errorf("scopeHtml does not write data-%s, so the editor has nothing to prefill", ds)
|
||||||
}
|
}
|
||||||
// read back as dataset.<name> by the editor
|
// the editor pre-fills through readScopes(), which is what walks the
|
||||||
if !strings.Contains(editorCode, "dataset."+ds) {
|
// bricks' data attributes — check the reader, not the form
|
||||||
t.Errorf("keyQuotaEdit does not read dataset.%s — the form would open blank and save zeros", ds)
|
if !strings.Contains(reader, "dataset."+ds) {
|
||||||
|
t.Errorf("readScopes does not read dataset.%s — editing a brick would save zeros over it", ds)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
// and the form must actually consume what readScopes produced
|
||||||
|
for _, field := range []string{"sc.token_quota", "sc.req_quota", "sc.period", "sc.hours"} {
|
||||||
|
if !strings.Contains(editCode, field) {
|
||||||
|
t.Errorf("scopeEdit does not prefill from %q", field)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func mustBody(t *testing.T, src, fn string) string {
|
||||||
|
t.Helper()
|
||||||
|
b, ok := jsFunctionBody(src, fn)
|
||||||
|
if !ok {
|
||||||
|
t.Fatalf("%s not found", fn)
|
||||||
|
}
|
||||||
|
return b
|
||||||
}
|
}
|
||||||
|
|
||||||
// The quota period vocabulary must match the server's, or the UI can offer a
|
// The quota period vocabulary must match the server's, or the UI can offer a
|
||||||
// value the API rejects.
|
// value the API rejects.
|
||||||
func TestUIQuotaPeriodsMatchServer(t *testing.T) {
|
func TestUIQuotaPeriodsMatchServer(t *testing.T) {
|
||||||
src := uiSource(t)
|
src := uiSource(t)
|
||||||
// the shared period select options, as rendered in both forms
|
|
||||||
for _, p := range []string{`value=""`, `value="hour"`, `value="week"`, `value="month"`, `value="nhour"`} {
|
for _, p := range []string{`value=""`, `value="hour"`, `value="week"`, `value="month"`, `value="nhour"`} {
|
||||||
if !strings.Contains(src, p) {
|
if !strings.Contains(src, p) {
|
||||||
t.Errorf("UI period select is missing %s", p)
|
t.Errorf("UI period select is missing %s", p)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
// the server's accepted vocabulary
|
|
||||||
for _, p := range []string{`"hour"`, `"week"`, `"month"`, `"nhour"`} {
|
for _, p := range []string{`"hour"`, `"week"`, `"month"`, `"nhour"`} {
|
||||||
if !strings.Contains(src, `if (p === `+p+`)`) && !strings.Contains(src, `=== `+p+`)`) {
|
if !strings.Contains(src, `if (p === `+p+`)`) && !strings.Contains(src, `=== `+p+`)`) {
|
||||||
t.Errorf("periodText() does not describe %s, so a badge would omit the window", p)
|
t.Errorf("periodText() does not describe %s, so a badge would omit the window", p)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
func max(a, b int) int {
|
|
||||||
if a > b {
|
|
||||||
return a
|
|
||||||
}
|
|
||||||
return b
|
|
||||||
}
|
|
||||||
|
|||||||
Reference in New Issue
Block a user