refactor(quota): 配额改为按模型,删除整钥总配额

用户明确要求:配额应当是密钥对应的**每个模型的单独配额**,而非整体配额。

## 语义变更

删除 GWKey.TokenQuota / ReqQuota / Period / Hours(整钥总额)。
ModelScope 新增 ReqQuota —— 请求数配额下沉到每条模型范围。

现在:每条 models[] 各自带 token 配额 + 请求数配额 + 重置周期,
彼此独立。一个模型用满只影响该模型。

★ 为什么不保留整钥总额:它会让「把 A 模型的额度挪给 B」变成一次全局
重分配;按模型独立计费则每个模型各自可控,运维能直接看出哪个模型在吃预算。

## 连带改动

- checkQuota 合并 key 级与 scope 级判定;checkKeyQuotaRetry 整体删除
  (顺带修掉上轮遗留的双重判定:入口不再先判空再重算)
- core:CreateKeyWithQuota / UpdateKeyWithQuota / ApplyQuota 全部删除,
  改由 ValidateScopeQuotas 校验每条 scope 的配额
- admin key:scope 上的配额不强制(admin 的 scope 仍限制模型范围,
  但不强制配额)—— 否则管理员会把自己锁在门外
- /api/v1/keys 不再回显 key 级配额字段(scope 里已含)
- WebUI:删除整钥配额徽标 / 「配额」按钮 / 创建表单的配额组 /
  putScope 的整钥回传;模型砖块与范围编辑器新增「请求数配额」输入,
  徽标显示 `1.0K 77×·1h`(未设配额显示 ∞)

## 判据

- TestOneModelsQuotaDoesNotBlockAnother 是本次核心保证。
  ★ 它第一版是**假判据**:m2 从不消耗,key-wide 计数器与 m1 自己的计数器
  读数恰好相同,退回 key-wide 仍通过。变异测试抓到后改为「先用 m2 花掉
  远超 m1 配额的量,再验证 m1 仍可用」—— 这样两种设计才可区分。
- TestUncappedModelNeverBlocked / TestAdminKeyScopesAreNotEnforced 新增
- UI 契约判据重写:整钥配额界面必须彻底消失(13 个符号)、
  scope 编辑器必须往返 req_quota、putScope 只发 scope 列表
- 错误消息点名具体模型(TestKeyAPIRejectionNamesTheModel)
- 3/3 变异全被抓

实测(真实进程 + 浏览器):m2 配额 500000 连打 25 次全成功,
m1 配额 1000 立即 429「token quota exceeded for "m1" (4315/1000)」,
此后 m2/m3 仍 200。UI:整钥配额元素全为 0,砖块各显配额,
编辑器预填/保存正确,零 JS 异常。
This commit is contained in:
JianFeeeee
2026-09-27 19:02:13 +08:00
parent 5530912d32
commit c51066f0b6
11 changed files with 515 additions and 645 deletions

View File

@ -375,21 +375,11 @@ type GWKey struct {
Note string `yaml:"note,omitempty" json:"note,omitempty"`
CreatedAt int64 `yaml:"created_at,omitempty" json:"created_at,omitempty"`
Seed bool `yaml:"seed,omitempty" json:"seed,omitempty"` // true if migrated from config gateway_keys
// TokenQuota caps this key's TOTAL tokens across every model it may use.
// 0 = unlimited. Period/Hours define the reset window, exactly like
// ModelScope: "" never resets, "hour"/"week"/"month" fixed windows,
// "nhour" uses Hours.
TokenQuota int64 `yaml:"token_quota,omitempty" json:"token_quota,omitempty"`
Period string `yaml:"period,omitempty" json:"period,omitempty"`
Hours int64 `yaml:"hours,omitempty" json:"hours,omitempty"`
// ReqQuota caps the number of requests per reset window; 0 = unlimited.
// RPM covers short bursts; this covers sustained volume.
ReqQuota int64 `yaml:"req_quota,omitempty" json:"req_quota,omitempty"`
}
// KeyQuota is the set of key-wide caps accepted by the admin API. It is a
// separate struct so a partial update can be expressed as a pointer (nil =
// "leave the stored caps alone") instead of zero values meaning "clear".
// KeyQuota is retained only to carry a scope entry's caps through the admin
// API. Quotas are per model, never per key: there is deliberately no key-wide
// total, so exhausting one model's budget never blocks the others.
type KeyQuota struct {
TokenQuota int64 `json:"token_quota"`
ReqQuota int64 `json:"req_quota"`
@ -397,14 +387,6 @@ type KeyQuota struct {
Hours int64 `json:"hours"`
}
// ApplyQuota writes the caps onto a key record.
func (k *GWKey) ApplyQuota(q KeyQuota) {
k.TokenQuota = q.TokenQuota
k.ReqQuota = q.ReqQuota
k.Period = q.Period
k.Hours = q.Hours
}
// NormalizeRole defaults an empty role to "user", so a key can never end up in
// a state where no role means "neither admin nor user".
func NormalizeRole(role string) string {
@ -453,15 +435,18 @@ func ValidatePeriod(period string, hours int64) error {
return fmt.Errorf("period must be one of \"\", hour, week, month, nhour (got %q)", period)
}
// ModelScope is one allowed model for a key, or one AUTO scheduling slot,
// with an optional token quota and reset period. TokenQuota 0 = unlimited;
// Period "" = never resets; "hour"/"week"/"month" are fixed windows; "nhour"
// uses Hours as the window length in hours.
// ModelScope is one allowed model for a key, or one AUTO scheduling slot.
// Its TokenQuota and ReqQuota cap THAT entry only, independently of every
// other entry on the same key: a model that runs out of budget stops being
// served while the key's other models keep working. TokenQuota 0 / ReqQuota 0
// = unlimited. Period "" = never resets; "hour"/"week"/"month" are fixed
// windows; "nhour" uses Hours.
type ModelScope struct {
Model string `yaml:"model" json:"model"`
Source string `yaml:"source,omitempty" json:"source,omitempty"` // optional: pin to one upstream source; "" = any source
Tier int `yaml:"tier,omitempty" json:"tier,omitempty"`
TokenQuota int64 `yaml:"token_quota" json:"token_quota"`
ReqQuota int64 `yaml:"req_quota,omitempty" json:"req_quota,omitempty"`
Period string `yaml:"period,omitempty" json:"period,omitempty"`
Hours int64 `yaml:"hours,omitempty" json:"hours,omitempty"`
}
@ -480,6 +465,7 @@ func (m *ModelScope) UnmarshalJSON(b []byte) error {
Source string `json:"source"`
Tier int `json:"tier"`
TokenQuota int64 `json:"token_quota"`
ReqQuota int64 `json:"req_quota"`
Period string `json:"period"`
Hours int64 `json:"hours"`
}
@ -490,6 +476,7 @@ func (m *ModelScope) UnmarshalJSON(b []byte) error {
m.Source = o.Source
m.Tier = o.Tier
m.TokenQuota = o.TokenQuota
m.ReqQuota = o.ReqQuota
m.Period = o.Period
m.Hours = o.Hours
return nil