refactor(quota): 配额改为按模型,删除整钥总配额

用户明确要求:配额应当是密钥对应的**每个模型的单独配额**,而非整体配额。

## 语义变更

删除 GWKey.TokenQuota / ReqQuota / Period / Hours(整钥总额)。
ModelScope 新增 ReqQuota —— 请求数配额下沉到每条模型范围。

现在:每条 models[] 各自带 token 配额 + 请求数配额 + 重置周期,
彼此独立。一个模型用满只影响该模型。

★ 为什么不保留整钥总额:它会让「把 A 模型的额度挪给 B」变成一次全局
重分配;按模型独立计费则每个模型各自可控,运维能直接看出哪个模型在吃预算。

## 连带改动

- checkQuota 合并 key 级与 scope 级判定;checkKeyQuotaRetry 整体删除
  (顺带修掉上轮遗留的双重判定:入口不再先判空再重算)
- core:CreateKeyWithQuota / UpdateKeyWithQuota / ApplyQuota 全部删除,
  改由 ValidateScopeQuotas 校验每条 scope 的配额
- admin key:scope 上的配额不强制(admin 的 scope 仍限制模型范围,
  但不强制配额)—— 否则管理员会把自己锁在门外
- /api/v1/keys 不再回显 key 级配额字段(scope 里已含)
- WebUI:删除整钥配额徽标 / 「配额」按钮 / 创建表单的配额组 /
  putScope 的整钥回传;模型砖块与范围编辑器新增「请求数配额」输入,
  徽标显示 `1.0K 77×·1h`(未设配额显示 ∞)

## 判据

- TestOneModelsQuotaDoesNotBlockAnother 是本次核心保证。
  ★ 它第一版是**假判据**:m2 从不消耗,key-wide 计数器与 m1 自己的计数器
  读数恰好相同,退回 key-wide 仍通过。变异测试抓到后改为「先用 m2 花掉
  远超 m1 配额的量,再验证 m1 仍可用」—— 这样两种设计才可区分。
- TestUncappedModelNeverBlocked / TestAdminKeyScopesAreNotEnforced 新增
- UI 契约判据重写:整钥配额界面必须彻底消失(13 个符号)、
  scope 编辑器必须往返 req_quota、putScope 只发 scope 列表
- 错误消息点名具体模型(TestKeyAPIRejectionNamesTheModel)
- 3/3 变异全被抓

实测(真实进程 + 浏览器):m2 配额 500000 连打 25 次全成功,
m1 配额 1000 立即 429「token quota exceeded for "m1" (4315/1000)」,
此后 m2/m3 仍 200。UI:整钥配额元素全为 0,砖块各显配额,
编辑器预填/保存正确,零 JS 异常。
This commit is contained in:
JianFeeeee
2026-09-27 19:02:13 +08:00
parent 5530912d32
commit c51066f0b6
11 changed files with 515 additions and 645 deletions

View File

@ -24,7 +24,7 @@
### 强大的多租户调度能力
- **多密钥多租户**:支持无限密钥,每个密钥独立角色、模型范围、Token 配额、重置周期
- **密钥用量配额**:每把 key 单独配总 token 配额 + 请求数配额与重置周期(小时/周/月/自定义 N 小时),跨模型共享预算;耗尽返 429 + `Retry-After` 可自动恢复,admin key 永不受限
- **按模型配额**:每把 key 的每个模型单独配 token + 请求数配额与重置周期(小时/周/月/自定义 N 小时);一个模型用满只影响该模型,同 key 其它模型照常;耗尽返 429 + `Retry-After` 并点名模型,admin key 永不受限
- **AUTO 智能调度**:基于优先级档位的分级调度,同优先级源自动轮询负载均衡,故障自动毫秒级故障转移
- **自愈冷却**:冷却上限 5 分钟,过半后放行 1 个探测请求,上游/额度恢复即刻回归轮询,无需等满冷却窗口
- **Token 配额管理**:精确到模型级别的 Token 配额控制,支持小时/周/月/自定义小时周期自动重置
@ -181,40 +181,44 @@ sources:
#### 密钥用量配额
每个密钥可单独限制用量与用量重置周期,两级配额同时生效:
**配额按模型单独设置**:每个密钥的 `models[]` 里,每一条模型范围各自带一份
token 配额与请求数配额。一个模型用满只影响该模型,同一密钥的其它模型照常工作。
```yaml
keys:
- key: sk-gw-<hex>
role: user
name: agent-alice
# ---- 整钥配额(跳模型)----
token_quota: 5000000 # 本周期内这把 key 的总 token 预算,0 = 无限
req_quota: 20000 # 本周期内的请求次数,0 = 无限
period: nhour # "" | hour | week | month | nhour
hours: 6 # 仅 nhour:每 6 小时重置
# ---- 模型范围(可选,逐模型配额)----
models:
- model: m1
token_quota: 1000000 # 本周期内该模型(该 key)的 token 预算
period: hour
- model: deepseek-v4-flash
token_quota: 1000000 # 本周期内该 key 用这个模型的 token 预算
req_quota: 20000 # 本周期内的请求次数
period: nhour # "" | hour | week | month | nhour
hours: 6 # 仅 nhour
- model: AUTO
token_quota: 5000000 # AUTO 也是一条独立配额
period: day # ← 这种写法会被拒绝(词表只有 hour/week/month/nhour)
- model: kimi-k3 # 未列配额 = 无限
```
- **没有「整钥总配额」**:这是刻意的设计。整钥总额会让「把 A 模型的额度挪给
B 模型」变成一次全局重分配;按模型独立计费则每个模型各自可控,运维可以
看出哪个模型吃掉了预算。
- `period` 词表:空 = 永不过期(累计总量),`hour` / `week` / `month` = 固定窗口,
`nhour` + `hours` = 自定义小时数。**拼错的周期在写入时就被拒**,不会静默变成
永不过期。
- 整钥配额跨该 key 所有模型共享一份预算;`models[]` 里的配额则是逐模型独立计数。
两者都按 key 隔离,A key 的用量不会消耗 B key 的额度。
- 配额严格按密钥隔离,且**同一密钥内按模型隔离**:A 密钥用满 `m1` 不会消耗
B 密钥的额度,同一密钥的 `m2` 也不受影响。
- 配额统计含聊天、流式、生图,跨重启从审计日志回放(保留 40 天,覆盖最长的
month 窗口)。
- 配额耗尽返回 **429 + `Retry-After`**(`rate_limit_exceeded`),客户端可等窗口
重置后自动恢复;模型越权才是 403。**admin 密钥永不受配额限制**,
避免把管理员锁在门外。
- 配额耗尽返回 **429 + `Retry-After`**(`rate_limit_exceeded`),消息里点名是哪个
模型用满了,客户端可等窗口重置后自动恢复;模型越权才是 403。
**admin 密钥永不受配额限制**(其 scope 上的配额也不强制),避免把管理员
锁在门外。
- 窗口用量按整点小时分桶统计,实际释放比配置窗口最多晚 1 小时(配额宁可晚释放
也不超发)。
- `PUT /api/keys/{key}` 的配额字段是可选的:省略 = 保留原值,显式 `0` = 解除限制。
只改模型范围不会清空已配置的预算。
- `PUT /api/keys/{key}` 提交 `models` 即同时提交它们的配额(配额就是 scope 的一部分,
不存在会与模型列表脱节的第二份预算)。显式 `0` = 解除该模型的限制。
##### 配额拒绝 vs 容量拒绝:两种「拒绝」含义不同
@ -233,12 +237,9 @@ keys:
配额桶按 (密钥, 模型, 整点小时) 分桶保留 40 天,实测(AMD 7840HS):
- 每请求配额检查:**149ns**(配了配额)/ **42.6ns**(未配配额,只查密钥记录,
不碰桶)/ **37ns**(admin 密钥直接返回)—— 均 **0 分配**。
未配配额的密钥几乎不付代价,可放心大量创建。
- 记录一条请求:283ns、3 分配(与引入配额前相同,分配来自 ring buffer)。
- 窗口查询按窗口长度而非保留总量扫描:1h 窗口 49ns、24h 窗口 55ns、
30d 窗口 3.9μs。
30d 窗口 3.9μs(此前全扫保留总量,960 桶时 5.9μs)。
- 记录一条请求:283ns、3 分配(与引入配额前相同,分配来自 ring buffer)。
- 内存:生产形态(7 密钥 × 8 模型 × 2 源 × 满 40 天 retention)约 **3.7MB**。
按源 pin 的 `source::model` 桶**惰性创建**——只有当某条配额真的 pin 了
某个源时才维护,否则每条记录多写一份桶,在 20 密钥 × 8 模型 × 3 源下会
@ -247,10 +248,10 @@ keys:
首个窗口可能少算**。
```bash
# 配额耗尽时客户端看到
# 配额耗尽时客户端看到(点名了具体模型)
HTTP/1.1 429 Too Many Requests
Retry-After: 2100
{"error":{"type":"rate_limit_exceeded","message":"key token quota exceeded (5000000/5000000, resets every 6h)"}}
{"error":{"type":"rate_limit_exceeded","message":"token quota exceeded for \"deepseek-v4-flash\" (5000000/5000000)"}}
```
### 模型路由
@ -419,10 +420,10 @@ Environment=MALLOC_ARENA_MAX=2
支持按时间范围导出 CSV;点击模型可生成 pin 到该模型的连接配置。
- **对话页**:流式/非流式调试。
- **密钥页**:创建/编辑网关 key,为每个 key 配模型范围(模型 + 源 + token 配额 + 周期),
管理员管理全部 key,用户只看到自己的 key。key 卡片头部显示整钥配额徽标
(如 `250.0K·6h` / `77×·6h`),「配额」按钮编辑总 token / 请求数与重置周期;
创建 key 时可直接配预算(选 admin 角色时该组输入自动禁用,因为 admin 永不受限)。
「我的密钥」页对用户展示本 key 的预算。
管理员管理全部 key,用户只看到自己的 key。每个模型砖块显示自己的配额徽标
(如 `1.0K 77×·1h`,未设配额显示 `∞`),点开可编辑该模型的 token 配额、
请求数配额与重置周期 —— 配额按模型独立生效,一个用满不影响同一 key 的其它模型。
「我的密钥」页对用户展示本 key 的模型范围与各自配额。
- **优先级页**:拖拽积木配置 AUTO 链档位。
- **源页**:在线增删改上游源(API key 等敏感字段加密落盘)。
- **适配器页**:上传 / 删除 Lua 适配器脚本。

View File

@ -38,10 +38,11 @@ Extracted and independently evolved from the multi-source LLM adapter layer of
`reasoning_content`, `tool_calls`, `usage`).
- **Image generation**: `POST /v1/images/generations`, routed to models with
`kind: image`.
- **Per-key usage quota**: each key carries its own token and request caps plus a
reset period (hour/week/month/custom N hours), shared across every model that
key may use. Exhaustion answers 429 + `Retry-After` so a client resumes when
the window rolls over; admin keys are never capped.
- **Per-model quota**: each key gives every model its own token and request caps
plus a reset period (hour/week/month/custom N hours). One model running out
affects only that model — the key's other models keep working. Exhaustion
answers 429 + `Retry-After` naming the model, so a client resumes when the
window rolls over; admin keys are never capped.
- **Multimodal**: `content` arrays (`image_url` etc.) pass through losslessly;
Anthropic/Gemini/Ollama are translated automatically.
- **LuaJIT VM**: golua-binding LuaJIT; each adapter has its own VM + worker
@ -175,56 +176,62 @@ under the `keys` field of the runtime file (encrypted at rest):
delete the seed key.
- The WebUI **Keys page** creates/deletes keys. Each key has a role (`admin`
manages everything, `user` sees only its own key) and an optional **model
scope** (model + source + token quota + reset period). Key cards show the
key-wide caps as a badge (e.g. `250.0K·6h` / `77×·6h`); a **Quota** button
edits the total token / request budget and its reset period, and the create
form takes a budget too (those fields disable themselves for `admin`, which
is never capped). The "My key" view shows a user its own budget.
scope** (model + source + token quota + reset period). Each model brick shows
its own budget badge (e.g. `1.0K 77×·1h`, `∞` when uncapped); clicking it
edits that model's token quota, request quota and reset period. Quotas apply
per model, so one model running out never blocks the key's others. The
"My key" view shows a user its model scopes and their budgets.
- Clients authenticate with any authorized key's plaintext as
`Authorization: Bearer <key>`.
- Deleting a key removes it from the store immediately.
#### Per-key usage quota
Each key can cap its own spend and reset period. Two levels apply at once:
**Quotas are per model.** Each key's `models[]` list gives every model its own
token budget and request budget. One model running out affects only that model
— the key's other models keep working.
```yaml
keys:
- key: sk-gw-<hex>
role: user
name: agent-alice
# ---- key-wide (across every model) ----
token_quota: 5000000 # total token budget for this window, 0 = unlimited
req_quota: 20000 # requests per window, 0 = unlimited
period: nhour # "" | hour | week | month | nhour
hours: 6 # n-hour only: resets every 6 hours
# ---- per-model scope (optional) ----
models:
- model: m1
token_quota: 1000000
period: hour
- model: deepseek-v4-flash
token_quota: 1000000 # this key's token budget for this model
req_quota: 20000 # requests within the window
period: nhour # "" | hour | week | month | nhour
hours: 6 # n-hour only
- model: AUTO
token_quota: 5000000 # AUTO is a quota entry like any other
period: hour
- model: kimi-k3 # no quota listed = unlimited
```
- **There is deliberately no key-wide total.** A key-wide cap would make
"move A's budget to B" a global reallocation; per-model budgets keep each
model independently controllable, so it stays visible which model is
actually consuming the spend.
- `period`: empty = never resets (lifetime total); `hour` / `week` / `month` =
fixed windows; `nhour` + `hours` = a custom hour count. **A misspelled
period is rejected at write time** rather than silently becoming a
never-resetting quota.
- The key-wide cap is one budget shared by every model the key may use;
quotas under `models[]` are counted per model. Both are isolated per key —
one key's traffic never drains another's budget.
- Quotas are isolated per key *and* per model within a key: one key exhausting
`m1` never draws on another key's budget, and never blocks the same key's
`m2`.
- Usage counts chat, streaming and image requests, and survives a restart by
replaying the audit log (40 days retained, covering the longest `month`
window).
- An exhausted quota returns **429 + `Retry-After`**
(`rate_limit_exceeded`) so a client resumes when the window rolls over; a
model the key may not use stays 403. **Admin keys are never capped**, so a
cap can never lock the operator out.
(`rate_limit_exceeded`) and the message names the model that ran out, so a
client resumes when the window rolls over; a model the key may not use stays
403. **Admin keys are never capped** (quotas on their scopes are not
enforced either), so a cap can never lock the operator out.
- Buckets are whole unix hours, so a window frees up at most an hour late
(deliberately freeing late rather than overspending).
- On `PUT /api/keys/{key}` the quota fields are optional: omitting them keeps
the stored caps, sending `0` explicitly lifts a cap. Editing only the model
scope never clears a budget that was already set.
- `PUT /api/keys/{key}` submits quotas by submitting `models` — the caps are
part of the scope, so there is no second budget that can drift out of sync
with the model list. An explicit `0` lifts that model's cap.
##### Quota rejection vs capacity rejection
@ -247,14 +254,10 @@ back off concurrency or switch sources.
Buckets are kept per (key, model, whole unix hour) for 40 days. Measured on an
AMD 7840HS:
- Quota check per request: **149 ns** (caps set) / **42.6 ns** (no caps — it
only looks up the key record and never touches a bucket) / **37 ns** (admin
key returns immediately) — all **0 allocations**. Keys without caps cost
almost nothing, so creating many of them is safe.
- Recording one request: 283 ns, 3 allocations (unchanged from before this
feature; the allocations come from the record ring buffer).
- A window query scans the window, not the whole retention: 49 ns for 1 h,
55 ns for 24 h, 3.9 µs for 30 d.
- Recording one request: 283 ns, 3 allocations (unchanged from before this
feature; the allocations come from the record ring buffer).
- Memory: the production shape (7 keys x 8 models x 2 sources at full 40-day
retention) costs about **3.7 MB**. The `source::model` bucket used by a
source-pinned quota is created **lazily** — it is maintained only once some
@ -267,7 +270,7 @@ AMD 7840HS:
```
HTTP/1.1 429 Too Many Requests
Retry-After: 2100
{"error":{"type":"rate_limit_exceeded","message":"key token quota exceeded (5000000/5000000, resets every 6h)"}}
{"error":{"type":"rate_limit_exceeded","message":"token quota exceeded for \"deepseek-v4-flash\" (5000000/5000000)"}}
```
### Model routing

View File

@ -375,21 +375,11 @@ type GWKey struct {
Note string `yaml:"note,omitempty" json:"note,omitempty"`
CreatedAt int64 `yaml:"created_at,omitempty" json:"created_at,omitempty"`
Seed bool `yaml:"seed,omitempty" json:"seed,omitempty"` // true if migrated from config gateway_keys
// TokenQuota caps this key's TOTAL tokens across every model it may use.
// 0 = unlimited. Period/Hours define the reset window, exactly like
// ModelScope: "" never resets, "hour"/"week"/"month" fixed windows,
// "nhour" uses Hours.
TokenQuota int64 `yaml:"token_quota,omitempty" json:"token_quota,omitempty"`
Period string `yaml:"period,omitempty" json:"period,omitempty"`
Hours int64 `yaml:"hours,omitempty" json:"hours,omitempty"`
// ReqQuota caps the number of requests per reset window; 0 = unlimited.
// RPM covers short bursts; this covers sustained volume.
ReqQuota int64 `yaml:"req_quota,omitempty" json:"req_quota,omitempty"`
}
// KeyQuota is the set of key-wide caps accepted by the admin API. It is a
// separate struct so a partial update can be expressed as a pointer (nil =
// "leave the stored caps alone") instead of zero values meaning "clear".
// KeyQuota is retained only to carry a scope entry's caps through the admin
// API. Quotas are per model, never per key: there is deliberately no key-wide
// total, so exhausting one model's budget never blocks the others.
type KeyQuota struct {
TokenQuota int64 `json:"token_quota"`
ReqQuota int64 `json:"req_quota"`
@ -397,14 +387,6 @@ type KeyQuota struct {
Hours int64 `json:"hours"`
}
// ApplyQuota writes the caps onto a key record.
func (k *GWKey) ApplyQuota(q KeyQuota) {
k.TokenQuota = q.TokenQuota
k.ReqQuota = q.ReqQuota
k.Period = q.Period
k.Hours = q.Hours
}
// NormalizeRole defaults an empty role to "user", so a key can never end up in
// a state where no role means "neither admin nor user".
func NormalizeRole(role string) string {
@ -453,15 +435,18 @@ func ValidatePeriod(period string, hours int64) error {
return fmt.Errorf("period must be one of \"\", hour, week, month, nhour (got %q)", period)
}
// ModelScope is one allowed model for a key, or one AUTO scheduling slot,
// with an optional token quota and reset period. TokenQuota 0 = unlimited;
// Period "" = never resets; "hour"/"week"/"month" are fixed windows; "nhour"
// uses Hours as the window length in hours.
// ModelScope is one allowed model for a key, or one AUTO scheduling slot.
// Its TokenQuota and ReqQuota cap THAT entry only, independently of every
// other entry on the same key: a model that runs out of budget stops being
// served while the key's other models keep working. TokenQuota 0 / ReqQuota 0
// = unlimited. Period "" = never resets; "hour"/"week"/"month" are fixed
// windows; "nhour" uses Hours.
type ModelScope struct {
Model string `yaml:"model" json:"model"`
Source string `yaml:"source,omitempty" json:"source,omitempty"` // optional: pin to one upstream source; "" = any source
Tier int `yaml:"tier,omitempty" json:"tier,omitempty"`
TokenQuota int64 `yaml:"token_quota" json:"token_quota"`
ReqQuota int64 `yaml:"req_quota,omitempty" json:"req_quota,omitempty"`
Period string `yaml:"period,omitempty" json:"period,omitempty"`
Hours int64 `yaml:"hours,omitempty" json:"hours,omitempty"`
}
@ -480,6 +465,7 @@ func (m *ModelScope) UnmarshalJSON(b []byte) error {
Source string `json:"source"`
Tier int `json:"tier"`
TokenQuota int64 `json:"token_quota"`
ReqQuota int64 `json:"req_quota"`
Period string `json:"period"`
Hours int64 `json:"hours"`
}
@ -490,6 +476,7 @@ func (m *ModelScope) UnmarshalJSON(b []byte) error {
m.Source = o.Source
m.Tier = o.Tier
m.TokenQuota = o.TokenQuota
m.ReqQuota = o.ReqQuota
m.Period = o.Period
m.Hours = o.Hours
return nil

View File

@ -249,20 +249,17 @@ func (c *Core) FindKey(key string) (config.GWKey, bool) {
}
// CreateKey builds a new random gateway key and persists it to config.yaml.
// Quotas live on the model scope entries, so a new key's budget is whatever
// its scopes carry.
func (c *Core) CreateKey(name, role string, models []config.ModelScope, note string) (config.GWKey, error) {
return c.CreateKeyWithQuota(name, role, models, note, config.KeyQuota{})
}
// CreateKeyWithQuota is CreateKey plus the key-wide token/request caps.
func (c *Core) CreateKeyWithQuota(name, role string, models []config.ModelScope, note string, q config.KeyQuota) (config.GWKey, error) {
c.mu.Lock()
defer c.mu.Unlock()
models = cleanScopes(models)
key := make([]byte, 16)
if _, err := rand.Read(key); err != nil {
if err := ValidateScopeQuotas(models); err != nil {
return config.GWKey{}, err
}
if err := q.Validate(); err != nil {
key := make([]byte, 16)
if _, err := rand.Read(key); err != nil {
return config.GWKey{}, err
}
rec := config.GWKey{
@ -274,7 +271,6 @@ func (c *Core) CreateKeyWithQuota(name, role string, models []config.ModelScope,
CreatedAt: time.Now().Unix(),
}
rec.Role = config.NormalizeRole(rec.Role)
rec.ApplyQuota(q)
c.cfg.Keys = append(c.cfg.Keys, rec)
if err := c.saveConfig(); err != nil {
return config.GWKey{}, err
@ -282,19 +278,13 @@ func (c *Core) CreateKeyWithQuota(name, role string, models []config.ModelScope,
return rec, nil
}
// UpdateKey mutates a key's name/role/model scope and persists it.
// UpdateKey mutates a key's name/role/model scope and persists it. The scope
// entries carry their own quotas, so replacing the scope replaces the budgets.
func (c *Core) UpdateKey(key, name, role string, models []config.ModelScope, note string) (config.GWKey, error) {
return c.UpdateKeyWithQuota(key, name, role, models, note, nil)
}
// UpdateKeyWithQuota is UpdateKey plus the key-wide caps. quota == nil leaves
// the existing caps untouched, so a caller that only edits the model scope
// does not silently clear a key's budget.
func (c *Core) UpdateKeyWithQuota(key, name, role string, models []config.ModelScope, note string, quota *config.KeyQuota) (config.GWKey, error) {
c.mu.Lock()
defer c.mu.Unlock()
if quota != nil {
if err := quota.Validate(); err != nil {
if models != nil {
if err := ValidateScopeQuotas(models); err != nil {
return config.GWKey{}, err
}
}
@ -312,9 +302,6 @@ func (c *Core) UpdateKeyWithQuota(key, name, role string, models []config.ModelS
c.cfg.Keys[i].Models = cleanScopes(models)
}
c.cfg.Keys[i].Note = note
if quota != nil {
c.cfg.Keys[i].ApplyQuota(*quota)
}
if err := c.saveConfig(); err != nil {
return config.GWKey{}, err
}
@ -798,3 +785,20 @@ func (c *Core) Close() {
c.vm.Stop()
}
}
// ValidateScopeQuotas checks every scope entry's caps before they are stored.
// A typo in a period must be rejected at write time rather than silently
// becoming a never-resetting budget — the opposite of what was typed.
func ValidateScopeQuotas(entries []config.ModelScope) error {
for _, e := range entries {
if err := (config.KeyQuota{
TokenQuota: e.TokenQuota,
ReqQuota: e.ReqQuota,
Period: e.Period,
Hours: e.Hours,
}).Validate(); err != nil {
return fmt.Errorf("model %q: %w", e.Model, err)
}
}
return nil
}

View File

@ -116,12 +116,8 @@ func (g *Gateway) apiV1Routes(w http.ResponseWriter, r *http.Request) {
"note": k.Note,
"created_at": k.CreatedAt,
"seed": k.Seed,
// key-wide spend caps (0 = unlimited). Echoed so an agent can
// see what budget it has without parsing config.yaml.
"token_quota": k.TokenQuota,
"req_quota": k.ReqQuota,
"period": k.Period,
"hours": k.Hours,
// Quotas live on the scope entries (k.Models), echoed above;
// there is deliberately no key-wide total.
// The secret itself is never echoed. An operator that needs it
// already has it from creation time or from config.yaml.
"key_prefix": maskKey(k.Key),

View File

@ -218,12 +218,15 @@ type quotaRejection struct {
func (q *quotaRejection) Error() string { return q.msg }
// checkQuota is checkKeyScope for callers that need the retry hint. It
// separates the quota verdicts (429) from model-permission verdicts (403).
// checkQuota validates the effective model against the key's model scope and
// that entry's quota. Returns nil when the request may proceed.
//
// Quotas are per scope entry, never key-wide: a model whose budget is spent
// is refused on its own while the key's other models keep working. The verdict
// carries the remaining seconds of the reset window so a spent budget answers
// 429 + Retry-After (come back when it rolls over) instead of 403 (which reads
// as "this key may never use this model" and makes clients give up).
func (g *Gateway) checkQuota(ctx context.Context, model string) *quotaRejection {
if q := g.checkKeyQuotaRetry(ctx); q != nil {
return q
}
allow := g.allowedModels(ctx)
if allow == nil {
return nil
@ -232,49 +235,29 @@ func (g *Gateway) checkQuota(ctx context.Context, model string) *quotaRejection
if sc.Model != model {
continue
}
win := AutoPeriodSeconds(sc.Period, sc.Hours)
k := keyID(reqKey(ctx))
if sc.TokenQuota > 0 {
used := g.scopeTokens(ctx, sc)
if used >= sc.TokenQuota {
if used := g.scopeTokens(ctx, sc); used >= sc.TokenQuota {
return &quotaRejection{
msg: fmt.Sprintf("token quota exceeded for %q (%d/%d)", model, used, sc.TokenQuota),
retry: AutoSecondsToReset(sc.Period, sc.Hours),
}
}
}
if sc.ReqQuota > 0 {
if used := g.stats.KeyWindowReqs(k, win); used >= sc.ReqQuota {
return &quotaRejection{
msg: fmt.Sprintf("request quota exceeded for %q (%d/%d)", model, used, sc.ReqQuota),
retry: AutoSecondsToReset(sc.Period, sc.Hours),
}
}
}
return nil
}
return &quotaRejection{msg: fmt.Sprintf("model %q is not allowed for this key", model)}
}
// checkKeyQuotaRetry enforces the key-wide caps and reports the remaining
// seconds of the reset window so the caller can answer with 429 + Retry-After.
// An admin key is never capped, and a key with no caps set is never rejected.
func (g *Gateway) checkKeyQuotaRetry(ctx context.Context) *quotaRejection {
rec, ok := g.core.FindKey(reqKey(ctx))
if !ok || rec.Role == "admin" {
return nil
}
k := keyID(reqKey(ctx))
win := AutoPeriodSeconds(rec.Period, rec.Hours)
if rec.TokenQuota > 0 {
if used := g.stats.KeyWindowTokens(k, win); used >= rec.TokenQuota {
return &quotaRejection{
msg: fmt.Sprintf("key token quota exceeded (%d/%d%s)", used, rec.TokenQuota, quotaWindowSuffix(rec.Period, rec.Hours)),
retry: AutoSecondsToReset(rec.Period, rec.Hours),
}
}
}
if rec.ReqQuota > 0 {
if used := g.stats.KeyWindowReqs(k, win); used >= rec.ReqQuota {
return &quotaRejection{
msg: fmt.Sprintf("key request quota exceeded (%d/%d%s)", used, rec.ReqQuota, quotaWindowSuffix(rec.Period, rec.Hours)),
retry: AutoSecondsToReset(rec.Period, rec.Hours),
}
}
}
return nil
}
// quotaWindowSuffix describes a quota's reset window for an error message, so
// a rejected caller can tell a permanent block from one that clears in an hour.
func quotaWindowSuffix(period string, hours int64) string {
@ -289,11 +272,10 @@ func quotaWindowSuffix(period string, hours int64) string {
return fmt.Sprintf(", resets every %dh", hours)
}
return ""
}
// scopeTokens returns the tokens this key consumed within the scope entry's
// reset window, isolated per key. For an AUTO entry the cap covers everything
// the key routed through AUTO; for a model entry it covers that model only.
} // scopeTokens returns the tokens this key consumed on the scope entry's model
// within its reset window, isolated per key. For an AUTO entry the cap covers
// everything the key routed through AUTO; for a model entry it covers that
// model only.
//
// It reads the per-key hourly buckets rather than the key-blind model
// buckets, so one key's usage can never exhaust another's quota.

View File

@ -74,10 +74,24 @@ func keyRecord(t *testing.T, g *Gateway, secret string) config.GWKey {
return config.GWKey{}
}
func TestKeyAPIStoresQuota(t *testing.T) {
func scopeOf(t *testing.T, k config.GWKey, model string) config.ModelScope {
t.Helper()
for _, m := range k.Models {
if m.Model == model {
return m
}
}
t.Fatalf("scope %q not found in %+v", model, k.Models)
return config.ModelScope{}
}
// Quotas live on the scope entries, not on the key: creating a key with a
// budget means creating scopes that carry it, and they must survive a
// read-back (persisted, not just echoed).
func TestKeyAPICreatesPerModelQuota(t *testing.T) {
g := adminGateway(t, config.GWKey{Key: "sk-admin", Role: "admin"})
rr := adminReq(t, g, "POST", "/api/keys",
`{"name":"agent-x","role":"user","token_quota":50000,"req_quota":200,"period":"nhour","hours":6,"models":[{"model":"m1"}]}`)
`{"name":"agent-x","role":"user","models":[{"model":"m1","token_quota":50000,"req_quota":200,"period":"nhour","hours":6}]}`)
if rr.Code != 200 {
t.Fatalf("create: %d %s", rr.Code, rr.Body.String())
}
@ -87,54 +101,21 @@ func TestKeyAPIStoresQuota(t *testing.T) {
if err := json.Unmarshal(rr.Body.Bytes(), &created); err != nil {
t.Fatalf("decode: %v", err)
}
if created.Key.TokenQuota != 50000 || created.Key.ReqQuota != 200 ||
created.Key.Period != "nhour" || created.Key.Hours != 6 {
t.Fatalf("created key did not carry the caps: %+v", created.Key)
sc := scopeOf(t, created.Key, "m1")
if sc.TokenQuota != 50000 || sc.ReqQuota != 200 || sc.Period != "nhour" || sc.Hours != 6 {
t.Fatalf("created scope did not carry the caps: %+v", sc)
}
// and it must survive a read-back (persisted, not just echoed)
back := keyRecord(t, g, created.Key.Key)
if back.TokenQuota != 50000 || back.Period != "nhour" || back.Hours != 6 {
back := scopeOf(t, keyRecord(t, g, created.Key.Key), "m1")
if back.TokenQuota != 50000 || back.ReqQuota != 200 || back.Period != "nhour" || back.Hours != 6 {
t.Errorf("read-back lost the caps: %+v", back)
}
}
// Editing only the model scope must not silently clear a key's budget: the
// caps are pointers precisely so "absent" is not "zero".
func TestKeyAPIUpdateKeepsQuotaWhenOmitted(t *testing.T) {
// Two models on one key carry independent budgets.
func TestKeyAPIKeepsPerModelQuotaIndependent(t *testing.T) {
g := adminGateway(t, config.GWKey{Key: "sk-admin", Role: "admin"})
rr := adminReq(t, g, "POST", "/api/keys",
`{"name":"agent-x","role":"user","token_quota":50000,"period":"day-typo-free","models":[{"model":"m1"}]}`)
rr = adminReq(t, g, "POST", "/api/keys", `{"name":"y","role":"user","token_quota":50000,"period":"hour","models":[{"model":"m1"}]}`)
if rr.Code != 200 {
t.Fatalf("setup create: %d %s", rr.Code, rr.Body.String())
}
var created struct {
Key config.GWKey `json:"key"`
}
_ = json.Unmarshal(rr.Body.Bytes(), &created)
// a scope-only edit
rr = adminReq(t, g, "PUT", "/api/keys/"+created.Key.Key,
`{"name":"agent-y","models":[{"model":"m1"},{"model":"m2"}]}`)
if rr.Code != 200 {
t.Fatalf("update: %d %s", rr.Code, rr.Body.String())
}
back := keyRecord(t, g, created.Key.Key)
if back.TokenQuota != 50000 {
t.Errorf("token_quota was cleared by a scope-only edit: %d", back.TokenQuota)
}
if back.Period != "hour" {
t.Errorf("period was cleared by a scope-only edit: %q", back.Period)
}
if len(back.Models) != 2 {
t.Errorf("scope edit did not apply: %+v", back.Models)
}
}
// Sending 0 explicitly must lift the cap, not be treated as "absent".
func TestKeyAPIUpdateZeroLiftsCap(t *testing.T) {
g := adminGateway(t, config.GWKey{Key: "sk-admin", Role: "admin"})
rr := adminReq(t, g, "POST", "/api/keys", `{"name":"z","role":"user","token_quota":1000,"period":"hour"}`)
`{"name":"agent-y","role":"user","models":[{"model":"m1","token_quota":1000,"period":"hour"},{"model":"m2","token_quota":9999,"req_quota":7,"period":"week"}]}`)
if rr.Code != 200 {
t.Fatalf("create: %d %s", rr.Code, rr.Body.String())
}
@ -142,13 +123,36 @@ func TestKeyAPIUpdateZeroLiftsCap(t *testing.T) {
Key config.GWKey `json:"key"`
}
_ = json.Unmarshal(rr.Body.Bytes(), &created)
back := keyRecord(t, g, created.Key.Key)
if a := scopeOf(t, back, "m1"); a.TokenQuota != 1000 || a.ReqQuota != 0 || a.Period != "hour" {
t.Errorf("m1 caps wrong: %+v", a)
}
if b := scopeOf(t, back, "m2"); b.TokenQuota != 9999 || b.ReqQuota != 7 || b.Period != "week" {
t.Errorf("m2 caps wrong: %+v", b)
}
}
rr = adminReq(t, g, "PUT", "/api/keys/"+created.Key.Key, `{"token_quota":0}`)
// Sending 0 explicitly lifts that model's cap.
func TestKeyAPIUpdateZeroLiftsCap(t *testing.T) {
g := adminGateway(t, config.GWKey{Key: "sk-admin", Role: "admin"})
rr := adminReq(t, g, "POST", "/api/keys",
`{"name":"z","role":"user","models":[{"model":"m1","token_quota":1000,"period":"hour"}]}`)
if rr.Code != 200 {
t.Fatalf("create: %d %s", rr.Code, rr.Body.String())
}
var created struct {
Key config.GWKey `json:"key"`
}
_ = json.Unmarshal(rr.Body.Bytes(), &created)
secret := created.Key.Key
rr = adminReq(t, g, "PUT", "/api/keys/"+secret,
`{"models":[{"model":"m1","token_quota":0,"req_quota":0,"period":""}]}`)
if rr.Code != 200 {
t.Fatalf("lift: %d %s", rr.Code, rr.Body.String())
}
if back := keyRecord(t, g, created.Key.Key); back.TokenQuota != 0 {
t.Errorf("token_quota = %d, want 0 (cap lifted)", back.TokenQuota)
if back := scopeOf(t, keyRecord(t, g, secret), "m1"); back.TokenQuota != 0 || back.ReqQuota != 0 || back.Period != "" {
t.Errorf("caps not lifted: %+v", back)
}
}
@ -156,7 +160,8 @@ func TestKeyAPIUpdateZeroLiftsCap(t *testing.T) {
// quota — which is the exact opposite of what the operator typed.
func TestKeyAPIRejectsBadPeriod(t *testing.T) {
g := adminGateway(t, config.GWKey{Key: "sk-admin", Role: "admin"})
rr := adminReq(t, g, "POST", "/api/keys", `{"name":"bad","role":"user","token_quota":1000,"period":"houre"}`)
rr := adminReq(t, g, "POST", "/api/keys",
`{"name":"bad","role":"user","models":[{"model":"m1","token_quota":1000,"period":"houre"}]}`)
if rr.Code != http.StatusBadRequest {
t.Fatalf("want 400 for a bad period, got %d %s", rr.Code, rr.Body.String())
}
@ -167,19 +172,35 @@ func TestKeyAPIRejectsBadPeriod(t *testing.T) {
func TestKeyAPIRejectsNegativeQuota(t *testing.T) {
g := adminGateway(t, config.GWKey{Key: "sk-admin", Role: "admin"})
rr := adminReq(t, g, "POST", "/api/keys", `{"name":"bad","role":"user","token_quota":-5}`)
rr := adminReq(t, g, "POST", "/api/keys",
`{"name":"bad","role":"user","models":[{"model":"m1","token_quota":-5}]}`)
if rr.Code != http.StatusBadRequest {
t.Fatalf("want 400 for a negative quota, got %d %s", rr.Code, rr.Body.String())
}
}
// A non-admin key must not be able to set or read another key's budget.
// The rejection must say WHICH model is over budget, so an operator looking at
// a key with a dozen scopes can tell which one to raise.
func TestKeyAPIRejectionNamesTheModel(t *testing.T) {
g := adminGateway(t, config.GWKey{Key: "sk-admin", Role: "admin"})
rr := adminReq(t, g, "POST", "/api/keys",
`{"name":"bad","role":"user","models":[{"model":"m1","req_quota":-1}]}`)
if rr.Code != http.StatusBadRequest {
t.Fatalf("want 400, got %d", rr.Code)
}
if !strings.Contains(rr.Body.String(), "m1") {
t.Errorf("error should name the offending model: %s", rr.Body.String())
}
}
// A non-admin key must not be able to mint keys.
func TestKeyAPIQuotaIsAdminOnly(t *testing.T) {
g := adminGateway(t,
config.GWKey{Key: "sk-admin", Role: "admin"},
config.GWKey{Key: "sk-u", Role: "user", TokenQuota: 10, Period: "hour"},
config.GWKey{Key: "sk-u", Role: "user", Models: []config.ModelScope{{Model: "m1", TokenQuota: 10, Period: "hour"}}},
)
req, _ := http.NewRequest("POST", "/api/keys", strings.NewReader(`{"name":"x","role":"admin","token_quota":0}`))
req, _ := http.NewRequest("POST", "/api/keys",
strings.NewReader(`{"name":"x","role":"admin","models":[{"model":"m1"}]}`))
req.Header.Set("Authorization", "Bearer sk-u")
req.Header.Set("Content-Type", "application/json")
rr := httptest.NewRecorder()
@ -187,7 +208,7 @@ func TestKeyAPIQuotaIsAdminOnly(t *testing.T) {
if rr.Code != http.StatusForbidden {
t.Fatalf("non-admin create: want 403, got %d %s", rr.Code, rr.Body.String())
}
// /api/v1/keys echoes the caps but never the secret
// /api/v1/keys exposes the per-model caps but never a secret
rr = adminReq(t, g, "GET", "/api/v1/keys", "")
if rr.Code != 200 {
t.Fatalf("GET /api/v1/keys: %d", rr.Code)
@ -196,39 +217,10 @@ func TestKeyAPIQuotaIsAdminOnly(t *testing.T) {
t.Error("/api/v1/keys leaked a key secret")
}
if !strings.Contains(rr.Body.String(), `"token_quota":10`) {
t.Errorf("/api/v1/keys should expose the cap: %s", rr.Body.String())
t.Errorf("/api/v1/keys should expose the per-model caps: %s", rr.Body.String())
}
}
// The AUTO scope entry must honour its reset window: usage that aged out of
// the window must not count against a per-key cap.
func TestAutoScopeQuotaHonoursWindow(t *testing.T) {
g, _ := quotaGateway(t,
config.GWKey{Key: "sk-a", Role: "user", Models: []config.ModelScope{
{Model: "AUTO", TokenQuota: 1000, Period: "hour"},
}},
config.GWKey{Key: "sk-b", Role: "user"},
)
ctx := quotaCtx(t, g, "sk-a")
sc := config.ModelScope{Model: "AUTO", TokenQuota: 1000, Period: "hour"}
// aged-out usage: 2 days old, 5M tokens — must be invisible to a 1h window
g.stats.Record(Req{Time: nowMSOffset(-48 * 3600 * 1000), Key: keyID("sk-a"),
Model: "m1", Source: "up", Prompt: 2500000, Compl: 2500000, OK: true, Status: 200})
if used := g.scopeTokens(ctx, sc); used != 0 {
t.Fatalf("AUTO scope saw %d tokens outside its 1h window; the period is being ignored", used)
}
// in-window usage counts
g.stats.Record(Req{Time: nowMSOffset(0), Key: keyID("sk-a"),
Model: "m1", Source: "up", Prompt: 400, Compl: 400, OK: true, Status: 200})
if used := g.scopeTokens(ctx, sc); used != 800 {
t.Fatalf("AUTO scope used = %d, want 800", used)
}
}
var _ = fmt.Sprintf
// A user must be able to see their own budget: /api/keys/me is the only key
// view a non-admin gets, so a cap missing from it is invisible to the very
// client it constrains.
@ -236,7 +228,7 @@ func TestKeyMeExposesOwnQuota(t *testing.T) {
g := adminGateway(t,
config.GWKey{Key: "sk-admin", Role: "admin"},
config.GWKey{Key: "sk-u", Role: "user", Name: "agent",
TokenQuota: 123456, ReqQuota: 42, Period: "week", Hours: 0},
Models: []config.ModelScope{{Model: "m1", TokenQuota: 123456, ReqQuota: 42, Period: "week"}}},
)
req, _ := http.NewRequest("GET", "/api/keys/me", nil)
req.Header.Set("Authorization", "Bearer sk-u")
@ -252,8 +244,10 @@ func TestKeyMeExposesOwnQuota(t *testing.T) {
if err := json.Unmarshal(rr.Body.Bytes(), &wrap); err != nil {
t.Fatalf("decode: %v (%s)", err, rr.Body.String())
}
me := wrap.Key
if me.TokenQuota != 123456 || me.ReqQuota != 42 || me.Period != "week" {
t.Errorf("own quota not visible to the key's owner: %+v", me)
sc := scopeOf(t, wrap.Key, "m1")
if sc.TokenQuota != 123456 || sc.ReqQuota != 42 || sc.Period != "week" {
t.Errorf("own quota not visible to the key's owner: %+v", sc)
}
}
var _ = fmt.Sprintf

View File

@ -82,8 +82,8 @@ func chatAs(t *testing.T, g *Gateway, key, model string) (*httptest.ResponseReco
// as "this key may never use this model".
func TestKeyTokenQuotaBlocksWithRetryAfter(t *testing.T) {
g, _ := quotaGateway(t,
config.GWKey{Key: "sk-a", Role: "user", TokenQuota: 8, Period: "hour",
Models: []config.ModelScope{{Model: "m1"}}},
config.GWKey{Key: "sk-a", Role: "user",
Models: []config.ModelScope{{Model: "m1", TokenQuota: 8, Period: "hour"}}},
config.GWKey{Key: "sk-b", Role: "user", Name: "b"},
)
// 4 tokens per call, budget 8 -> the third call crosses it
@ -114,10 +114,10 @@ func TestKeyTokenQuotaBlocksWithRetryAfter(t *testing.T) {
// key that is allowed the same model.
func TestKeyTokenQuotaIsIsolatedPerKey(t *testing.T) {
g, _ := quotaGateway(t,
config.GWKey{Key: "sk-a", Role: "user", TokenQuota: 4, Period: "hour",
Models: []config.ModelScope{{Model: "m1"}}},
config.GWKey{Key: "sk-b", Role: "user", TokenQuota: 1000, Period: "hour",
Models: []config.ModelScope{{Model: "m1"}}},
config.GWKey{Key: "sk-a", Role: "user",
Models: []config.ModelScope{{Model: "m1", TokenQuota: 4, Period: "hour"}}},
config.GWKey{Key: "sk-b", Role: "user",
Models: []config.ModelScope{{Model: "m1", TokenQuota: 1000, Period: "hour"}}},
)
rr, _ := chatAs(t, g, "sk-a", "m1")
if rr.Code != 200 {
@ -162,7 +162,8 @@ func TestScopeModelTokenQuotaIsolatedPerKey(t *testing.T) {
// operator out of the gateway they administer.
func TestAdminKeyIsNeverQuotaCapped(t *testing.T) {
g, _ := quotaGateway(t,
config.GWKey{Key: "sk-admin", Role: "admin", TokenQuota: 1, Period: "hour", ReqQuota: 1},
config.GWKey{Key: "sk-admin", Role: "admin",
Models: []config.ModelScope{{Model: "m1", TokenQuota: 1, ReqQuota: 1, Period: "hour"}}},
config.GWKey{Key: "sk-b", Role: "user"},
)
for i := 1; i <= 3; i++ {
@ -176,8 +177,8 @@ func TestAdminKeyIsNeverQuotaCapped(t *testing.T) {
// the count must still stop the key.
func TestKeyRequestQuotaBlocks(t *testing.T) {
g, ctrl := quotaGateway(t,
config.GWKey{Key: "sk-a", Role: "user", ReqQuota: 2, Period: "hour",
Models: []config.ModelScope{{Model: "m1"}}},
config.GWKey{Key: "sk-a", Role: "user",
Models: []config.ModelScope{{Model: "m1", ReqQuota: 2, Period: "hour"}}},
config.GWKey{Key: "sk-b", Role: "user"},
)
for i := 1; i <= 2; i++ {
@ -197,8 +198,8 @@ func TestKeyRequestQuotaBlocks(t *testing.T) {
// A model outside the scope is still 403, not 429: retrying cannot help.
func TestModelOutsideScopeStaysForbidden(t *testing.T) {
g, _ := quotaGateway(t,
config.GWKey{Key: "sk-a", Role: "user", TokenQuota: 1000, Period: "hour",
Models: []config.ModelScope{{Model: "other-model"}}},
config.GWKey{Key: "sk-a", Role: "user",
Models: []config.ModelScope{{Model: "other-model", TokenQuota: 1000, Period: "hour"}}},
config.GWKey{Key: "sk-b", Role: "user"},
)
rr, code := chatAs(t, g, "sk-a", "m1")
@ -256,8 +257,8 @@ func TestKeyQuotaWinsOverSlotQuota(t *testing.T) {
up := upstream(t, &upstreamCtrl{})
defer up.Close()
g := newQuotaGW(t, up.URL,
config.GWKey{Key: "sk-a", Role: "user", TokenQuota: 4, Period: "hour",
Models: []config.ModelScope{{Model: "AUTO"}}})
config.GWKey{Key: "sk-a", Role: "user",
Models: []config.ModelScope{{Model: "AUTO", TokenQuota: 4, Period: "hour"}}})
ctx := quotaCtx(t, g, "sk-a")
// exhaust the key first
@ -304,7 +305,10 @@ func newQuotaGW(t *testing.T, upURL string, keys ...config.GWKey) *Gateway {
Keys: keys,
Sources: []config.Source{{
Name: "up", BaseURL: upURL, Adapter: "openai",
Models: []config.Model{{ID: "m1", Priority: 100}},
Models: []config.Model{
{ID: "m1", Priority: 100},
{ID: "m2", Priority: 90},
},
}},
}
if err := cfg.ApplyDefaults(); err != nil {
@ -325,3 +329,116 @@ func newQuotaGW(t *testing.T, upURL string, keys ...config.GWKey) *Gateway {
}
return g
}
// The core of per-model quotas: a model that runs out of budget must stop
// being served on its own, while every other model on the SAME key keeps
// working. A key-wide total would fail this — it would block m2 because m1 was
// capped, which is exactly the coupling this design removes.
func TestOneModelsQuotaDoesNotBlockAnother(t *testing.T) {
up := upstream(t, &upstreamCtrl{})
defer up.Close()
td := t.TempDir()
cfgPath := td + "/config.yaml"
if err := os.WriteFile(cfgPath, []byte("listen: :0"), 0o644); err != nil {
t.Fatal(err)
}
cfg := &config.Config{
Path: cfgPath, AdapterDir: filepath.Join(td, "adapters"),
RuntimeFile: filepath.Join(td, "runtime.json"),
Keys: []config.GWKey{{
Key: "sk-a", Role: "user", Name: "two-models",
Models: []config.ModelScope{
{Model: "m1", TokenQuota: 4, Period: "hour"},
{Model: "m2", TokenQuota: 1000, Period: "hour"},
},
}},
Sources: []config.Source{{
Name: "up", BaseURL: up.URL, Adapter: "openai",
Models: []config.Model{
{ID: "m1", Priority: 100},
{ID: "m2", Priority: 90},
},
}},
}
if err := cfg.ApplyDefaults(); err != nil {
t.Fatal(err)
}
c, err := core.NewFromConfig(cfg)
if err != nil {
t.Fatalf("core: %v", err)
}
t.Cleanup(c.Close)
g, err := New(c, []string{"sk-a"})
if err != nil {
t.Fatalf("gateway: %v", err)
}
// Spend on m2 FIRST. Without this the two designs are
// indistinguishable: a key-wide counter and m1's own counter would both
// read 0 before m1 is used, so the test would pass either way (it did —
// see the commit that rewrote it).
for i := 1; i <= 3; i++ {
if rr, _ := chatAs(t, g, "sk-a", "m2"); rr.Code != 200 {
t.Fatalf("m2 priming call %d: want 200, got %d", i, rr.Code)
}
}
// m1's budget is 4 and each call costs 4, so the key-wide total (m2+m1) is
// already 12 when m1 starts: a key-wide cap would refuse m1 immediately.
if rr, _ := chatAs(t, g, "sk-a", "m1"); rr.Code != 200 {
t.Fatalf("m1 must be served on its own budget (m2's spend is not its problem), got %d (%s)",
rr.Code, rr.Body.String())
}
if rr, code := chatAs(t, g, "sk-a", "m1"); rr.Code != http.StatusTooManyRequests {
t.Fatalf("m1 second call: want 429, got %d (%s)", rr.Code, code)
}
// m2 must keep working
for i := 1; i <= 3; i++ {
if rr, _ := chatAs(t, g, "sk-a", "m2"); rr.Code != 200 {
t.Fatalf("m2 call %d must be served while m1 is capped, got %d (%s)", i, rr.Code, rr.Body.String())
}
}
// and the message must name m1, not the key
rr, _ := chatAs(t, g, "sk-a", "m1")
if !strings.Contains(rr.Body.String(), "m1") {
t.Errorf("rejection should name the capped model: %s", rr.Body.String())
}
}
// A model with no quota on it is never blocked by a sibling's cap, and an
// uncapped key is never blocked at all.
func TestUncappedModelNeverBlocked(t *testing.T) {
up := upstream(t, &upstreamCtrl{})
defer up.Close()
g := newQuotaGW(t, up.URL,
config.GWKey{Key: "sk-a", Role: "user", Models: []config.ModelScope{
{Model: "m1", TokenQuota: 1, Period: "hour"},
{Model: "m2"},
}})
if rr, _ := chatAs(t, g, "sk-a", "m1"); rr.Code != 200 {
t.Fatalf("m1 first: want 200, got %d", rr.Code)
}
if rr, _ := chatAs(t, g, "sk-a", "m1"); rr.Code != http.StatusTooManyRequests {
t.Fatalf("m1 second: want 429, got %d", rr.Code)
}
for i := 1; i <= 4; i++ {
if rr, _ := chatAs(t, g, "sk-a", "m2"); rr.Code != 200 {
t.Fatalf("m2 (uncapped) call %d: want 200, got %d", i, rr.Code)
}
}
}
// An admin key is never capped even when its scopes carry budgets: a cap that
// locked the operator out would be unrecoverable through the UI.
func TestAdminKeyScopesAreNotEnforced(t *testing.T) {
up := upstream(t, &upstreamCtrl{})
defer up.Close()
g := newQuotaGW(t, up.URL,
config.GWKey{Key: "sk-admin", Role: "admin", Models: []config.ModelScope{
{Model: "m1", TokenQuota: 1, ReqQuota: 1, Period: "hour"},
}})
for i := 1; i <= 4; i++ {
if rr, _ := chatAs(t, g, "sk-admin", "m1"); rr.Code != 200 {
t.Fatalf("admin call %d: want 200 (admin scopes are not enforced), got %d", i, rr.Code)
}
}
}

View File

@ -39,27 +39,17 @@ func (g *Gateway) handleKeysAPI(w http.ResponseWriter, r *http.Request) {
writeJSON(w, http.StatusOK, map[string]interface{}{"keys": g.core.ListKeys()})
case http.MethodPost:
var body struct {
Name string `json:"name"`
Role string `json:"role"`
Models []config.ModelScope `json:"models"`
Note string `json:"note"`
TokenQuota *int64 `json:"token_quota"`
ReqQuota *int64 `json:"req_quota"`
Period *string `json:"period"`
Hours *int64 `json:"hours"`
Name string `json:"name"`
Role string `json:"role"`
Models []config.ModelScope `json:"models"`
Note string `json:"note"`
}
if err := json.NewDecoder(r.Body).Decode(&body); err != nil {
writeError(w, http.StatusBadRequest, "invalid_request", "invalid json: "+err.Error())
return
}
body.Role = config.NormalizeRole(body.Role)
q := config.KeyQuota{
TokenQuota: optInt64(body.TokenQuota),
ReqQuota: optInt64(body.ReqQuota),
Period: optString(body.Period),
Hours: optInt64(body.Hours),
}
rec, err := g.core.CreateKeyWithQuota(body.Name, body.Role, body.Models, body.Note, q)
rec, err := g.core.CreateKey(body.Name, body.Role, body.Models, body.Note)
if err != nil {
writeError(w, http.StatusBadRequest, "key_error", err.Error())
return
@ -71,33 +61,19 @@ func (g *Gateway) handleKeysAPI(w http.ResponseWriter, r *http.Request) {
return
}
var body struct {
Name string `json:"name"`
Role string `json:"role"`
Models []config.ModelScope `json:"models"`
Note string `json:"note"`
TokenQuota *int64 `json:"token_quota"`
ReqQuota *int64 `json:"req_quota"`
Period *string `json:"period"`
Hours *int64 `json:"hours"`
Name string `json:"name"`
Role string `json:"role"`
Models []config.ModelScope `json:"models"`
Note string `json:"note"`
}
if err := json.NewDecoder(r.Body).Decode(&body); err != nil {
writeError(w, http.StatusBadRequest, "invalid_request", "invalid json: "+err.Error())
return
}
// Quota fields are pointers so "absent" is distinguishable from
// "set to 0": omitting them leaves the stored caps alone, while
// sending 0 explicitly lifts a cap. Without this, editing only the
// model scope would silently clear a key's budget.
var q *config.KeyQuota
if body.TokenQuota != nil || body.ReqQuota != nil || body.Period != nil || body.Hours != nil {
q = &config.KeyQuota{
TokenQuota: optInt64(body.TokenQuota),
ReqQuota: optInt64(body.ReqQuota),
Period: optString(body.Period),
Hours: optInt64(body.Hours),
}
}
rec, err := g.core.UpdateKeyWithQuota(path, body.Name, body.Role, body.Models, body.Note, q)
// Each scope entry carries its own token/request caps, so replacing the
// scope replaces the budgets with it — there is no separate key-wide
// quota that could drift out of sync with the models.
rec, err := g.core.UpdateKey(path, body.Name, body.Role, body.Models, body.Note)
if err != nil {
writeError(w, http.StatusBadRequest, "key_error", err.Error())
return
@ -127,22 +103,6 @@ func (g *Gateway) handleKeysAPI(w http.ResponseWriter, r *http.Request) {
}
}
// optInt64 dereferences an optional quota field, treating absent as 0.
func optInt64(p *int64) int64 {
if p == nil {
return 0
}
return *p
}
// optString dereferences an optional quota field, treating absent as "".
func optString(p *string) string {
if p == nil {
return ""
}
return *p
}
// handleKeyMe returns the authenticated key's own record (users see only
// themselves; admins can use this as a convenience too).
func (g *Gateway) handleKeyMe(w http.ResponseWriter, r *http.Request) {

View File

@ -339,9 +339,6 @@
.key-canvas{border:1px solid var(--line);border-radius:16px;padding:14px;margin-bottom:14px;background:var(--card);
backdrop-filter:blur(var(--glass));box-shadow:var(--sh-sm)}
.kc-head{display:flex;align-items:center;gap:10px;flex-wrap:wrap}
/* key-wide caps sit inline in the head: they belong to the key, not to
any one model brick, and must not be draggable with one. */
.kc-caps{display:inline-flex;align-items:center;gap:6px;flex-wrap:wrap}
.kc-blocks{display:flex;flex-wrap:wrap;gap:10px;align-items:center;margin-top:12px;background:var(--card2);
border:1px dashed var(--line);border-radius:12px;padding:14px;min-height:64px}
.kc-blocks.ovh{outline:2px dashed var(--primary);outline-offset:2px}
@ -821,15 +818,10 @@
kAnySrc: "任意源",
kQuotaB: "Token 配额",
kQuotaHintB: "0 / 留空 = 无限",
kKeyQuota: "密钥总配额",
kKeyQuotaHint:
"限制这把密钥在重置周期内的总用量(跳模型)。0 / 留空 = 无限。",
kKeyReqQuota: "请求数配额",
kKeyReqQuotaHint: "限制周期内的请求次数。0 / 留空 = 无限。",
kKeyQuotaAdmin:
"admin 密钥永不受配额限制(避免把管理员锁在门外)。",
kKeyQuotaEdit: "配额",
kKeyQuotaNone: "无限",
kReqQuota: "请求数配额",
kReqQuotaHint: "限制周期内的请求次数。0 / 留空 = 无限。",
kQuotaPerModelHint:
"配额按模型单独设置:创建后点「+ 添加模型」,逐个模型配 token 配额与周期。一个模型用满只影响该模型,同一密钥的其它模型照常。",
kPeriodB: "重置周期",
kPerNothing: "不限",
kPerHour: "每 小时",
@ -1055,16 +1047,10 @@
kAnySrc: "any source",
kQuotaB: "Token quota",
kQuotaHintB: "0 / empty = unlimited",
kKeyQuota: "Key-wide quota",
kKeyQuotaHint:
"Caps this key's total spend per reset window, across every model it may use. 0 / empty = unlimited.",
kKeyReqQuota: "Request quota",
kKeyReqQuotaHint:
"Caps requests per window. 0 / empty = unlimited.",
kKeyQuotaAdmin:
"Admin keys are never capped — a cap could lock the operator out.",
kKeyQuotaEdit: "Quota",
kKeyQuotaNone: "unlimited",
kReqQuota: "Request quota",
kReqQuotaHint: "Caps requests per window. 0 / empty = unlimited.",
kQuotaPerModelHint:
"Quotas are per model: after creating the key, use \"Add model\" to give each model its own token budget and reset period. One model running out affects only that model; the key's other models keep working.",
kPeriodB: "Reset period",
kPerNothing: "Never",
kPerHour: "Every hour",
@ -4266,11 +4252,10 @@
async function renderKeysUser(me) {
$("#tab-keys").innerHTML = `
<div class="card"><h2>${t("kMeTitle")}</h2>
<div class="tbl-wrap"><table><tr><th>${t("kName")}</th><th>${t("kMeRole")}</th><th>${t("kKey")}</th><th>${t("kKeyQuota")}</th><th>${t("kMeModels")}</th></tr>
<div class="tbl-wrap"><table><tr><th>${t("kName")}</th><th>${t("kMeRole")}</th><th>${t("kKey")}</th><th>${t("kMeModels")}</th></tr>
<tr><td><b>${esc(me.name || "—")}</b></td><td>${roleTag(me.role)}</td>
<td><span class="kr-key">${esc(me.key)}</span>
<button class="ghost small" onclick="copyText('${escAttr(me.key)}')">${t("kCopy")}</button></td>
<td>${keyCapBadges(me)}</td>
<td>${
me.models && me.models.length
? me.models
@ -4308,22 +4293,13 @@
}
function keyCanvasHtml(k) {
const scopes = k.models || [];
// Key-wide caps live on the canvas, not on a brick: they are a budget
// the whole key shares, so they must not be dragged around with one
// model. data-* carries them so a quota edit can round-trip them
// through the same PUT that saves the model scope.
const caps = `data-kquota="${k.token_quota || 0}" data-kreqquota="${k.req_quota || 0}"
data-kperiod="${escAttr(k.period || "")}" data-khours="${k.hours || 0}"`;
return `
<div class="key-canvas" data-key="${escAttr(k.key)}" ${caps}>
<div class="key-canvas" data-key="${escAttr(k.key)}">
<div class="kc-head">
<b>${esc(k.name || "—")}</b>
${roleTag(k.role)}
<span class="kr-key">${esc(maskKey(k.key))}</span>
<button class="ghost small" onclick="copyText('${escAttr(k.key)}')">${t("kCopy")}</button>
<span class="kc-caps" title="${escAttr(t("kKeyQuotaHint"))}">${keyCapBadges(k)}</span>
${k.role === "admin" ? "" : `<button class="ghost small" title="${escAttr(t("kKeyQuotaEdit"))}"
onclick="keyQuotaEdit('${escAttr(k.key)}')">${t("kKeyQuotaEdit")}</button>`}
<span class="grow"></span>
<span class="muted">${fmtCreated(k.created_at)}</span>
<button class="ghost small errc" onclick="delKey('${escAttr(k.key)}','${escAttr(k.name || "")}')">${t("kDel")}</button>
@ -4336,39 +4312,19 @@
</div>
</div>`;
}
// keyCapBadges renders the key-wide caps. A cap with no reset period is
// flagged as such, because "1M tokens, never resets" and "1M tokens per
// hour" are very different promises and the badge must not blur them.
function keyCapBadges(k) {
const out = [];
const suffix = periodText(k.period || "", k.hours || 0);
if (+k.token_quota > 0) {
out.push(
`<span class="mb-quota" title="${escAttr(t("kKeyQuotaHint"))}">${esc(fmtQuota(k.token_quota))}${esc(suffix)}</span>`,
);
}
if (+k.req_quota > 0) {
out.push(
`<span class="mb-quota" title="${escAttr(t("kKeyReqQuotaHint"))}">${esc(fmtQuota(k.req_quota))}×${esc(suffix)}</span>`,
);
}
if (!out.length) {
return `<span class="muted">${t("kKeyQuotaNone")}</span>`;
}
return out.join(" ");
}
function scopeHtml(key, m) {
const qt = fmtQuota(m.token_quota);
const comb = scopeComb(m);
const src = normSrc(m.source);
const attrs = `data-key="${escAttr(key)}" data-model="${escAttr(comb)}"
data-quota="${m.token_quota || 0}" data-period="${escAttr(m.period || "")}" data-hours="${m.hours || 0}"`;
data-quota="${m.token_quota || 0}" data-reqquota="${m.req_quota || 0}"
data-period="${escAttr(m.period || "")}" data-hours="${m.hours || 0}"`;
return `<span class="mb" draggable="true" ${attrs} title="${escAttr(t("kBrickH"))}"
onclick="scopeEdit('${escAttr(key)}','${escAttr(comb)}')"
oncontextmenu="scopeCtx(event,'${escAttr(key)}','${escAttr(comb)}')">
<span class="mb-ico"><svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" width="10" height="10"><circle cx="12" cy="12" r="10"/></svg></span>
<span class="mb-name">${esc(m.model)}${src ? `<em class="mb-src">${esc(src)}</em>` : ""}</span>
<span class="mb-quota">${esc(quantBadge(m.token_quota, m.period, m.hours))}</span>
<span class="mb-quota">${esc(scopeQuotaBadge(m))}</span>
<span class="copy-b" title="${escAttr(t("kCopyB"))}" onclick="event.stopPropagation();scopeDup('${escAttr(key)}','${escAttr(comb)}')"><svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" width="12" height="12"><rect x="9" y="9" width="13" height="13" rx="2" ry="2"/><path d="M5 15H4a2 2 0 0 1-2-2V4a2 2 0 0 1 2-2h9a2 2 0 0 1 2 2v1"/></svg></span>
<span class="mb-x" title="${escAttr(t("kDelB"))}" onclick="event.stopPropagation();scopeRm('${escAttr(key)}','${escAttr(comb)}')"><svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" width="12" height="12"><line x1="18" y1="6" x2="6" y2="18"/><line x1="6" y1="6" x2="18" y2="18"/></svg></span>
</span>`;
@ -4400,6 +4356,16 @@
if (p === "nhour") return "·" + Math.max(1, h) + "h";
return "";
}
// scopeQuotaBadge shows a scope entry's budgets: token quota and, when
// set, the request count. They are per model — a spent budget blocks
// only that model, not the whole key.
function scopeQuotaBadge(m) {
const parts = [];
if (+m.token_quota > 0) parts.push(fmtQuota(m.token_quota));
if (+m.req_quota > 0) parts.push(fmtQuota(m.req_quota) + "\u00d7");
if (!parts.length) return "\u221e";
return parts.join(" ") + periodText(m.period, m.hours);
}
function quantBadge(quota, period, hours) {
quota = +quota || 0;
period = period || "";
@ -4414,123 +4380,23 @@
model: comb[0],
source: src || undefined,
token_quota: parseInt(b.dataset.quota) || 0,
req_quota: parseInt(b.dataset.reqquota) || 0,
period: b.dataset.period || "",
hours: parseInt(b.dataset.hours) || 0,
};
});
}
async function putScope(key, scopes) {
// The key-wide caps ride along with every scope write. The API reads
// them as pointers, so sending them back unchanged is a no-op, while
// omitting them would be indistinguishable from "clear the budget" to
// a future reader. Round-tripping them here means editing a model's
// scope can never silently drop a key's quota.
const canvas = document.querySelector(
`.key-canvas[data-key="${CSS.escape(key)}"]`,
);
// Each scope entry carries its own quotas, so the whole budget travels
// with the models it applies to. There is no separate key-wide total
// that could drift out of sync with the model list.
const body = { models: scopes };
if (canvas) {
body.token_quota = parseInt(canvas.dataset.kquota) || 0;
body.req_quota = parseInt(canvas.dataset.kreqquota) || 0;
body.period = canvas.dataset.kperiod || "";
body.hours = parseInt(canvas.dataset.khours) || 0;
}
await api("/api/keys/" + encodeURIComponent(key), {
method: "PUT",
headers: { "Content-Type": "application/json" },
body: JSON.stringify(body),
});
}
// keyQuotaEdit opens the key-wide budget form.
function keyQuotaEdit(key) {
const canvas = document.querySelector(
`.key-canvas[data-key="${CSS.escape(key)}"]`,
);
if (!canvas) return;
const cur = {
token_quota: parseInt(canvas.dataset.kquota) || 0,
req_quota: parseInt(canvas.dataset.kreqquota) || 0,
period: canvas.dataset.kperiod || "",
hours: parseInt(canvas.dataset.khours) || 0,
};
const wrap = document.createElement("div");
wrap.id = "modal-wrap";
wrap.style.cssText =
"position:fixed;inset:0;background:rgba(15,22,44,.45);display:flex;align-items:flex-start;justify-content:center;overflow:auto;padding:48px 20px;z-index:50";
wrap.innerHTML = `<div class="card" style="width:400px;max-width:100%"><h2>${t("kKeyQuotaEdit")}</h2>
<label>${t("kKeyQuota")} <span class="muted">${t("kKeyQuotaHint")}</span></label>
<input id="kq-tokens" type="number" min="0" step="1"
placeholder="${escAttr(t("kKeyQuotaNone"))}" value="${cur.token_quota || ""}">
<label>${t("kKeyReqQuota")} <span class="muted">${t("kKeyReqQuotaHint")}</span></label>
<input id="kq-reqs" type="number" min="0" step="1"
placeholder="${escAttr(t("kKeyQuotaNone"))}" value="${cur.req_quota || ""}">
<label>${t("kPeriodB")}</label>
<select id="kq-period">
<option value="" ${!cur.period ? "selected" : ""}>${t("kPerNothing")}</option>
<option value="hour" ${cur.period === "hour" ? "selected" : ""}>${t("kPerHour")}</option>
<option value="week" ${cur.period === "week" ? "selected" : ""}>${t("kPerWeek")}</option>
<option value="month" ${cur.period === "month" ? "selected" : ""}>${t("kPerMonth")}</option>
<option value="nhour" ${cur.period === "nhour" ? "selected" : ""}>${t("kPerHours")}</option>
</select>
<div id="kq-hours-box" style="display:none"><label>${t("kPerNHint")}</label>
<input id="kq-hours" type="number" min="1" step="1" value="${cur.hours || 24}"></div>
<p class="muted" style="font-size:12px">${t("kKeyQuotaAdmin")}</p>
<p><button onclick="keyQuotaSave('${escAttr(key)}', this)">${t("kSaveScope")}</button>
<button class="ghost" onclick="this.closest('#modal-wrap').remove()">${t("mCancel")}</button></p>
</div>`;
document.body.appendChild(wrap);
const toggle = () => {
$("#kq-hours-box").style.display =
$("#kq-period").value === "nhour" ? "block" : "none";
};
$("#kq-period").addEventListener("change", toggle);
toggle();
$("#kq-tokens").focus();
}
async function keyQuotaSave(key, btn) {
// Resolve our own dialog from the button that was clicked, so closing
// it can never remove a different #modal-wrap that happens to come
// first in the document.
const wrap = btn ? btn.closest("#modal-wrap") : null;
let tokens = parseInt($("#kq-tokens").value);
if (isNaN(tokens) || tokens < 0) tokens = 0;
let reqs = parseInt($("#kq-reqs").value);
if (isNaN(reqs) || reqs < 0) reqs = 0;
let hours = parseInt($("#kq-hours").value);
if (isNaN(hours) || hours < 1) hours = 1;
const period = $("#kq-period").value;
// Catch the "nhour picked but hours never filled in" case locally: the
// API rejects it too, but a round trip for a form-level mistake is
// needless.
if (period === "nhour" && hours < 1) {
toast(t("kPerNHint"));
return;
}
if (btn) btn.disabled = true;
try {
await api("/api/keys/" + encodeURIComponent(key), {
method: "PUT",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
token_quota: tokens,
req_quota: reqs,
period,
hours,
}),
});
// Close THIS modal, not whichever #modal-wrap comes first in the
// document: another dialog (e.g. the seed-key notice) may already be
// open, and a bare $("#modal-wrap") would remove that one and leave
// this form stranded on screen.
if (wrap) wrap.remove();
else closeTopModal();
toast(t("kSaved"));
await loadKeys();
} catch (e) {
toast(e.message);
if (btn) btn.disabled = false;
}
}
async function scopePush(key) {
const canvas = document.querySelector(
`.key-canvas[data-key="${CSS.escape(key)}"]`,
@ -4631,6 +4497,9 @@
<label>${t("kQuotaB")} <span class="muted">${t("kQuotaHintB")}</span></label>
<input id="sc-quota" type="number" min="0" step="1"
placeholder="${escAttr(t("kQuotaHintB"))}" value="${sc.token_quota ? sc.token_quota : ""}">
<label>${t("kReqQuota")} <span class="muted">${t("kReqQuotaHint")}</span></label>
<input id="sc-reqs" type="number" min="0" step="1"
placeholder="${escAttr(t("kReqQuotaHint"))}" value="${sc.req_quota || ""}">
<label>${t("kPeriodB")}</label>
<select id="sc-period">
<option value="" ${!sc.period ? "selected" : ""}>${t("kPerNothing")}</option>
@ -4664,6 +4533,8 @@
const parts = splitCombKey(comb);
let q = parseInt($("#sc-quota").value);
if (isNaN(q) || q < 0) q = 0;
let rq = parseInt($("#sc-reqs").value);
if (isNaN(rq) || rq < 0) rq = 0;
let hours = parseInt($("#sc-hours").value);
if (isNaN(hours) || hours < 1) hours = 1;
const period = $("#sc-period").value;
@ -4672,6 +4543,7 @@
model: parts.model,
source: parts.source || undefined,
token_quota: q,
req_quota: rq,
period,
hours,
};
@ -4832,41 +4704,13 @@
<option value="user">${t("kRoleUser")}</option>
<option value="admin">${t("kRoleAdmin")}</option>
</select>
<label>${t("kKeyQuota")} <span class="muted">${t("kKeyQuotaHint")}</span></label>
<input id="kc-tokens" type="number" min="0" step="1" placeholder="${escAttr(t("kKeyQuotaNone"))}">
<label>${t("kKeyReqQuota")} <span class="muted">${t("kKeyReqQuotaHint")}</span></label>
<input id="kc-reqs" type="number" min="0" step="1" placeholder="${escAttr(t("kKeyQuotaNone"))}">
<label>${t("kPeriodB")}</label>
<select id="kc-period">
<option value="" selected>${t("kPerNothing")}</option>
<option value="hour">${t("kPerHour")}</option>
<option value="week">${t("kPerWeek")}</option>
<option value="month">${t("kPerMonth")}</option>
<option value="nhour">${t("kPerHours")}</option>
</select>
<div id="kc-hours-box" style="display:none"><label>${t("kPerNHint")}</label>
<input id="kc-hours" type="number" min="1" step="1" value="24"></div>
<p class="muted" style="font-size:12px">${t("kQuotaPerModelHint")}</p>
<label>${t("kNote")}</label>
<input id="kc-note">
<p class="muted" style="font-size:12px">${t("kKeyQuotaAdmin")}</p>
<p><button onclick="createKey(this)">${t("kCreateBtn")}</button>
<button class="ghost" onclick="this.closest('#modal-wrap').remove()">${t("mCancel")}</button></p>
</div>`;
document.body.appendChild(wrap);
$("#kc-role").addEventListener("change", () => {
const admin = $("#kc-role").value === "admin";
// An admin key ignores its caps server-side; hiding the fields
// avoids the operator setting one and wondering why it never trips.
$("#kc-tokens").disabled = admin;
$("#kc-reqs").disabled = admin;
$("#kc-period").disabled = admin;
$("#kc-hours-box").style.display =
!admin && $("#kc-period").value === "nhour" ? "block" : "none";
});
$("#kc-period").addEventListener("change", () => {
$("#kc-hours-box").style.display =
$("#kc-period").value === "nhour" ? "block" : "none";
});
$("#kc-name").focus();
}
async function createKey(btn) {
@ -4875,17 +4719,6 @@
toast(t("kName"));
return;
}
const role = $("#kc-role").value;
// An admin key is never capped; send the fields anyway (the server
// ignores them) rather than special-casing the request shape.
const readNum = (sel) => {
const el = $(sel);
if (el.disabled) return 0;
const n = parseInt(el.value);
return isNaN(n) || n < 0 ? 0 : n;
};
let hours = parseInt($("#kc-hours").value);
if (isNaN(hours) || hours < 1) hours = 1;
if (btn) btn.disabled = true;
let j;
try {
@ -4894,12 +4727,8 @@
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
name,
role,
role: $("#kc-role").value,
note: $("#kc-note").value.trim(),
token_quota: readNum("#kc-tokens"),
req_quota: readNum("#kc-reqs"),
period: role === "admin" ? "" : $("#kc-period").value,
hours: role === "admin" ? 0 : hours,
}),
});
} catch (e) {

View File

@ -13,9 +13,9 @@ import (
// form the user just submitted stays on screen while an unrelated dialog
// vanishes.
//
// This is exactly the class of bug the api() contract test below was written
// for: reviewing inline JS by eye does not catch it, and the visible symptom
// ("the dialog did not close") points away from the cause. It is pinned here.
// This is the same class of bug the api() contract test pins: reviewing inline
// JS by eye does not catch it, and the symptom ("the dialog did not close")
// points away from the cause.
// modalCloseRe finds every `$(...)`-style lookup of the shared modal id.
var modalCloseRe = regexp.MustCompile(`\$\("#modal-wrap"\)`)
@ -25,10 +25,9 @@ var modalCloseRe = regexp.MustCompile(`\$\("#modal-wrap"\)`)
var closestModalRe = regexp.MustCompile(`\.closest\("#modal-wrap"\)`)
func TestUIDialogClosesItselfNotTheFirstModal(t *testing.T) {
src := uiSource(t)
// Strip comments first: prose that *names* the unsafe pattern (as the fix's
// own comment does) would otherwise be flagged as a violation.
code := stripJSComments(src)
code := stripJSComments(uiSource(t))
for _, m := range modalCloseRe.FindAllStringIndex(code, -1) {
after := code[m[1]:]
@ -45,15 +44,45 @@ func TestUIDialogClosesItselfNotTheFirstModal(t *testing.T) {
}
}
// TestUIDialogClosuresGoThroughSafePaths pins the rule across the whole
// document by data flow rather than by pattern: every handler that closes a
// dialog must do it one of the two safe ways. A handler could contain a
// correct .closest() and still close the wrong dialog on another path.
// stripJSComments removes // line comments and /* block */ comments from JS
// embedded in the UI document. It is deliberately simple: the document is our
// own source, and a false negative here only means the check is silent.
func stripJSComments(src string) string {
var out strings.Builder
lines := strings.Split(src, "\n")
inBlock := false
for _, ln := range lines {
trimmed := strings.TrimSpace(ln)
if inBlock {
if strings.Contains(ln, "*/") {
inBlock = false
}
continue
}
if strings.HasPrefix(trimmed, "/*") {
if !strings.Contains(ln, "*/") {
inBlock = true
}
continue
}
if i := strings.Index(ln, "//"); i >= 0 {
before := ln[:i]
if strings.Count(before, `"`)%2 == 0 && strings.Count(before, "'")%2 == 0 {
ln = before
}
}
out.WriteString(ln)
out.WriteString("\n")
}
return out.String()
}
// Every handler that closes a dialog must do it one of the two safe ways.
func TestUIDialogClosuresGoThroughSafePaths(t *testing.T) {
src := stripJSComments(uiSource(t))
for _, fn := range []string{
"downloadStatsCsv", "downloadKeysCsv", "saveSource", "saveTemplate",
"scrAddFromForm", "sortScopeSave", "scopeSave", "keyQuotaSave", "createKey",
"scrAddFromForm", "sortScopeSave", "scopeSave", "createKey",
} {
body, ok := jsFunctionBody(src, fn)
if !ok {
@ -83,12 +112,12 @@ func TestUICloseTopModalTakesTheLast(t *testing.T) {
// Handlers that resolve their dialog from a button must actually receive one:
// a signature without the parameter means the .closest() silently yields null
// and the save leaves its form stranded on screen.
// and the save leaves its form stranded.
func TestUIDialogHandlersReceiveTheirButton(t *testing.T) {
src := stripJSComments(uiSource(t))
for _, fn := range []string{
"downloadStatsCsv", "saveSource", "scrAddFromForm", "sortScopeSave",
"scopeSave", "keyQuotaSave", "createKey",
"scopeSave", "createKey",
} {
body, ok := jsFunctionBody(src, fn)
if !ok {
@ -106,165 +135,133 @@ func TestUIDialogHandlersReceiveTheirButton(t *testing.T) {
}
}
// stripJSComments removes // line comments and /* block */ comments from JS
// embedded in the UI document. It is deliberately simple (no string/regex
// awareness beyond skipping quoted spans on the same line): the document is
// our own source, and a false negative here only means the check is silent.
func stripJSComments(src string) string {
var out strings.Builder
lines := strings.Split(src, "\n")
inBlock := false
for _, ln := range lines {
trimmed := strings.TrimSpace(ln)
if inBlock {
if strings.Contains(ln, "*/") {
inBlock = false
}
continue
}
if strings.HasPrefix(trimmed, "/*") {
if !strings.Contains(ln, "*/") {
inBlock = true
}
continue
}
if i := strings.Index(ln, "//"); i >= 0 {
// keep code before the comment when the // is not inside a string
before := ln[:i]
if strings.Count(before, `"`)%2 == 0 && strings.Count(before, "'")%2 == 0 {
ln = before
}
}
out.WriteString(ln)
out.WriteString("\n")
}
return out.String()
}
// TestUIKeyQuotaDialogsResolveOwnModal pins the dialog-closing rule for the
// two forms this change added.
func TestUIKeyQuotaDialogsResolveOwnModal(t *testing.T) {
src := uiSource(t)
for _, fn := range []string{"keyQuotaSave", "createKey"} {
body, ok := jsFunctionBody(src, fn)
if !ok {
t.Errorf("%s not found in the UI source", fn)
continue
}
if !closestModalRe.MatchString(body) {
t.Errorf("%s does not resolve its own dialog via .closest(\"#modal-wrap\");\n"+
"with another dialog open it would close that one instead and leave this form stranded", fn)
}
}
}
// The quota editor must read and write the key-wide caps, and putScope must
// carry them along: the API treats the quota fields as pointers, so dropping
// them on a scope-only write is indistinguishable from "clear the budget".
func TestUIPutScopeCarriesKeyQuota(t *testing.T) {
// Quotas belong to the scope entries, so putScope only has to ship the scope
// list — the caps travel inside it. What must NOT come back is a key-wide
// total: it would be a second budget able to drift out of sync with the models
// it is supposed to cover.
func TestUIPutScopeShipsOnlyScopeQuotas(t *testing.T) {
body, ok := jsFunctionBody(uiSource(t), "putScope")
if !ok {
t.Fatal("putScope not found")
}
// Scan the code with comments removed, or a comment that merely *names* a
// field would satisfy the check while the field is never sent.
code := stripJSComments(body)
for _, field := range []string{"token_quota", "req_quota", "period", "hours"} {
if !strings.Contains(code, field) {
t.Errorf("putScope does not send %q — editing a model scope would clear the key's quota", field)
if !strings.Contains(code, "models:") || !strings.Contains(code, "scopes") {
t.Error("putScope must ship the scope list the caps live in")
}
for _, gone := range []string{"kquota", "kreqquota", "kperiod", "khours"} {
if strings.Contains(code, gone) {
t.Errorf("putScope still references the removed key-wide quota field %q", gone)
}
}
}
// A key's caps are rendered from the API record and shown on the canvas, so
// the badge and the data attributes must not drift from the field names. The
// create form must send them too, or a key would only be cappable after an
// extra round of edits.
func TestUIKeyQuotaRendersFromAPIFields(t *testing.T) {
// A model brick carries its own budgets, and the scope editor reads and writes
// both of them: dropping req_quota on the round trip would silently lift a
// request cap every time someone edited a token cap.
func TestUIScopeEditorRoundTripsBothQuotas(t *testing.T) {
src := stripJSComments(uiSource(t))
for _, fn := range []string{"scopeHtml", "readScopes", "scopeEdit", "scopeSave", "scopeQuotaBadge"} {
if _, ok := jsFunctionBody(src, fn); !ok {
t.Errorf("%s not found in the UI source", fn)
}
}
for _, fn := range []string{"scopeHtml", "readScopes", "scopeSave", "scopeQuotaBadge"} {
body, ok := jsFunctionBody(src, fn)
if !ok {
continue
}
if !strings.Contains(body, "req_quota") && !strings.Contains(body, "reqquota") {
t.Errorf("%s does not carry req_quota — a request cap would be lost on edit", fn)
}
}
if form, ok := jsFunctionBody(src, "scopeEdit"); ok && !strings.Contains(form, "sc-reqs") {
t.Error("the scope editor has no request-quota input")
}
}
// The whole key-wide quota surface must be gone from the UI: a badge or a
// button reading a field the server no longer has would render "undefined" or
// silently do nothing.
func TestUIHasNoKeyWideQuotaSurface(t *testing.T) {
src := uiSource(t)
for _, token := range []string{
"keyCapBadges", // shared renderer
"kq-tokens", "kq-reqs", "kq-period", "kq-hours", // editor fields
"kc-tokens", "kc-reqs", "kc-period", "kc-hours", // create form fields
for _, gone := range []string{
"keyCapBadges", "keyQuotaEdit", "keyQuotaSave",
"kq-tokens", "kq-reqs", "kq-period", "kq-hours",
"kc-tokens", "kc-reqs", "kc-period", "kc-hours",
"data-kquota", "data-kreqquota", "data-kperiod", "data-khours",
} {
if !strings.Contains(src, token) {
t.Errorf("UI never references %q — the quota form is not wired up", token)
}
}
// the create request must actually carry the caps
full, ok := jsFunctionBody(src, "createKey")
if !ok {
t.Fatal("createKey not found")
}
body := stripJSComments(full)
for _, field := range []string{"token_quota", "req_quota", "period"} {
if !strings.Contains(body, field) {
t.Errorf("createKey does not send %q — a new key could never be created with a budget", field)
if strings.Contains(src, gone) {
t.Errorf("UI still references the removed key-wide quota surface %q", gone)
}
}
}
// Existence of the strings is not enough: the badge has to READ the API
// fields, and the editor has to read the canvas data attributes it writes.
// A field can be present in the source and still never reach the screen —
// e.g. left in a dead branch, or read from a name the writer never sets.
// Existence of the strings is not enough: the badge has to READ the scope's
// fields, and the editor has to read back what the brick writes. A field can
// be present in the source and still never reach the screen — e.g. left in a
// dead branch, or read from a data attribute the writer never sets.
func TestUIKeyQuotaDataflowIsLive(t *testing.T) {
src := uiSource(t)
src := stripJSComments(uiSource(t))
badge, ok := jsFunctionBody(src, "keyCapBadges")
badge, ok := jsFunctionBody(src, "scopeQuotaBadge")
if !ok {
t.Fatal("keyCapBadges not found")
t.Fatal("scopeQuotaBadge not found")
}
badgeCode := stripJSComments(badge)
for _, field := range []string{"k.token_quota", "k.req_quota", "k.period"} {
if !strings.Contains(badgeCode, field) {
t.Errorf("keyCapBadges does not read %q — the cap would never show on the key card", field)
for _, field := range []string{"m.token_quota", "m.req_quota", "m.period"} {
if !strings.Contains(badge, field) {
t.Errorf("scopeQuotaBadge does not read %q — the cap would never show on the model brick", field)
}
}
// the editor must read back what keyCanvasHtml wrote
canvas, ok := jsFunctionBody(src, "keyCanvasHtml")
brick, ok := jsFunctionBody(src, "scopeHtml")
if !ok {
t.Fatal("keyCanvasHtml not found")
t.Fatal("scopeHtml not found")
}
editor, ok := jsFunctionBody(src, "keyQuotaEdit")
edit, ok := jsFunctionBody(src, "scopeEdit")
if !ok {
t.Fatal("keyQuotaEdit not found")
t.Fatal("scopeEdit not found")
}
canvasCode, editorCode := stripJSComments(canvas), stripJSComments(editor)
for _, ds := range []string{"kquota", "kreqquota", "kperiod", "khours"} {
// written as data-<name> on the canvas
if !strings.Contains(canvasCode, "data-"+ds+"=") {
t.Errorf("keyCanvasHtml does not write data-%s, so the editor has nothing to prefill", ds)
brickCode, editCode := stripJSComments(brick), stripJSComments(edit)
reader := stripJSComments(mustBody(t, src, "readScopes"))
for _, ds := range []string{"quota", "reqquota", "period", "hours"} {
if !strings.Contains(brickCode, "data-"+ds+"=") {
t.Errorf("scopeHtml does not write data-%s, so the editor has nothing to prefill", ds)
}
// read back as dataset.<name> by the editor
if !strings.Contains(editorCode, "dataset."+ds) {
t.Errorf("keyQuotaEdit does not read dataset.%s — the form would open blank and save zeros", ds)
// the editor pre-fills through readScopes(), which is what walks the
// bricks' data attributes — check the reader, not the form
if !strings.Contains(reader, "dataset."+ds) {
t.Errorf("readScopes does not read dataset.%s — editing a brick would save zeros over it", ds)
}
}
// and the form must actually consume what readScopes produced
for _, field := range []string{"sc.token_quota", "sc.req_quota", "sc.period", "sc.hours"} {
if !strings.Contains(editCode, field) {
t.Errorf("scopeEdit does not prefill from %q", field)
}
}
}
func mustBody(t *testing.T, src, fn string) string {
t.Helper()
b, ok := jsFunctionBody(src, fn)
if !ok {
t.Fatalf("%s not found", fn)
}
return b
}
// The quota period vocabulary must match the server's, or the UI can offer a
// value the API rejects.
func TestUIQuotaPeriodsMatchServer(t *testing.T) {
src := uiSource(t)
// the shared period select options, as rendered in both forms
for _, p := range []string{`value=""`, `value="hour"`, `value="week"`, `value="month"`, `value="nhour"`} {
if !strings.Contains(src, p) {
t.Errorf("UI period select is missing %s", p)
}
}
// the server's accepted vocabulary
for _, p := range []string{`"hour"`, `"week"`, `"month"`, `"nhour"`} {
if !strings.Contains(src, `if (p === `+p+`)`) && !strings.Contains(src, `=== `+p+`)`) {
t.Errorf("periodText() does not describe %s, so a badge would omit the window", p)
}
}
}
func max(a, b int) int {
if a > b {
return a
}
return b
}