feat(plugin): Lua 插件机制 + 计费插件 + 插件文档

插件 = plugin_dir 下的单个 .lua 文件,做两件事:挂请求流水线的钩子、在启动时
贡献 WebUI 界面(整页或往现有页面追加组件)。两者独立。

## 流水线 stage(三个)
  request_start  已解析鉴权、未选源
  routed         已选定 (source, model)、未发往上游
  request_end    每请求恰好一次,带最终计量
request_end 挂在 gateway.writeRec——四条入口路径(直连/AUTO × 流式/非流式)的
唯一汇合点:既不漏(流式 token 只有流结束才知道)也不重。

## 计费插件(plugins/billing.lua,默认 seed,开箱可用)
源 / 模型 / 密钥三个维度定价。token 价优先级 keys > models > default;per_request
固定价是**叠加**的(生图模型可以既算 token 又收固定费)。单位是 USD/单 token,
即各家 provider 的公布口径。累计 total / by_source / by_model / by_key / by_day。
失败请求保留 token 费用、丢弃固定费(可经 count_failures 翻转)。
界面 = 一个独立页 + 状态页顶部一块总开销 tile。

## 一个明确的设计边界
计费插件**只报表,不执法**。网关自己的配额会计(stats.go,入口强制)才是限额
权威,插件不参与任何路由/配额决策。两套独立会计若对不上,比一套功能略少的
更糟。

## ★ 中途改掉的一个根本设计错误
最初让插件复用适配器的**弹性 worker 池**(多状态)。这对适配器是对的(它们无
状态),对插件是错的:计费插件往 plugin.state 累加,多状态意味着总量被劈成
几份;而 SetState 写价格只写进其中一个 worker,钩子恰好跑到另一个时**所有请求
按 0 计费**。改为**单状态 + 互斥锁**。代价写进文档:钩子必须短、同步、不阻塞,
卡住的钩子会卡住所有插件的钩子。
这个 bug 是测试逼出来的——先写了 SetState+Fire 的用例,数字全是 0 才挖出来。

另一个连带缺陷:只带 prices 的 PUT 会整体替换 state,把累计量清零。改为
prices/state 分离——prices 是配置、state 是历史,改价不动账。

## 撞到的三个 Lua 绑定的坑(都写进注释)
  - SetGlobal **会 pop 栈**:连着调两次,第二次从空栈取,赋成 nil
  - GetField 索引越界是 **SIGABRT 整个进程**,不是 panic,recover 救不了
  - Call(nargs, n) **不接受函数索引**,它调的是 nargs 个参数正下方那个;
    传索引会调到参数上("attempt to call a table value")
另外 GetField/SetField 用绝对索引,SetTop(0) 之后必须重取。

## 错误隔离
钩子 error() 不影响转发:捕获 → 记进 hook_errors → 跳下一个插件。适配器出错
会让源进冷却,插件出错**零惩罚**——插件是可选功能。/api/plugins 的 hook_errors
让"坏掉的插件"可见而不是静默消失。

## 界面注入
GET /api/ui-inject 一次返回所有插件的扩展(侧栏需要全部 page 才能建好)。
WebUI 在首次 render **之前** await 注入:先插 HTML 再重建 <script> 让它执行
(innerHTML/template 插入的 script 不会执行,这正是要的效果——避免脚本跑在
自己 DOM 之前)。注入失败不影响仪表盘。
browser 侧 pluginAPI 暴露 fetchState / postState / onTabShown。

## 文档
docs/plugins.md —— 快速上手、加载与热更新、三个 stage 的完整字段表、界面扩展、
状态与 HTTP API、运行时约束(单状态/异常隔离/内置函数)、计费插件的定价与
计费策略、排错表、与适配器的对比表。

## 判据(328 个测试全绿,插件相关 33 个)
  - 计费断言的是**具体金额**(0.00625 / 0.0402 / 0.0075…),不是"能加载"
  - 4 个变异都红:钩子异常不隔离 / prices 清空累计 / 忽略 key 优先级 /
    毫秒时间戳不换算
  - UI 侧 6 个判据把注入顺序、script 执行时机、pluginAPI 名称、tab 路由、
    anchor 四种形式、失败非致命全钉住
  - 鉴权:state 读任意角色、写仅 admin
This commit is contained in:
JianFeeeee
2026-10-02 00:37:29 +08:00
parent a2e1adc2d8
commit a51a6811a6
14 changed files with 3173 additions and 5 deletions

View File

@ -14,6 +14,7 @@ import (
"time"
"llmsproxy/internal/config"
"llmsproxy/internal/lua"
"llmsproxy/internal/provider"
"llmsproxy/internal/scheduler"
"llmsproxy/internal/types"
@ -414,6 +415,7 @@ func (g *Gateway) handleChat(w http.ResponseWriter, r *http.Request) {
Type: "chat",
OK: false,
}
g.fireStart(ctx, &req, "chat", "AUTO", len(req.Messages), len(req.Tools))
// quotaExhausted reports a slot whose token window has been used up;
// exhausted slots are dropped from scheduling without penalty.
quotaExhausted := func(sl *scheduler.Slot) bool {
@ -817,6 +819,7 @@ func (g *Gateway) singleChat(w http.ResponseWriter, ctx context.Context, cands [
recordChatUsage(rec, req, resp)
rec.Source = usedSrc
rec.Model = usedModel
g.fireRouted(ctx, "chat", usedSrc, usedModel, -1, false)
// Non-streaming: the whole response arrives at once, so TTFB equals
// the total latency.
rec.FirstByteMs = rec.LatMs
@ -824,7 +827,18 @@ func (g *Gateway) singleChat(w http.ResponseWriter, ctx context.Context, cands [
writeChatCompletion(w, resp, effective)
}
// writeRec records a finished request (audit + aggregates).
// writeRec records a finished request (audit + aggregates) and fires the
// plugin request_end stage.
//
// This is the ONE place every request passes through on its way out, which is
// what makes it the right hook point: the four entry points (single/stream ×
// direct/auto) all funnel here, so a plugin sees each request exactly once with
// its final accounting. Firing earlier would miss the streamed ones (their
// numbers are only known once the stream finishes), and firing in each entry
// point would mean four call sites to keep in sync.
//
// Hooks run AFTER the record is written: a plugin must not be able to delay or
// lose the audit trail, and a plugin that throws is contained by Fire.
func (g *Gateway) writeRec(rec *Req) {
if rec == nil {
return
@ -833,6 +847,81 @@ func (g *Gateway) writeRec(rec *Req) {
rec.Time = time.Now().UnixMilli()
}
g.stats.Record(*rec)
g.fireEnd(rec)
}
// fireStart dispatches the plugin request_start stage: the request has been
// parsed and authorized but no upstream slot has been chosen yet, so `source`
// is empty. A plugin that only wants volume/acceptance counts can subscribe
// here and stay out of the per-request hot path entirely.
func (g *Gateway) fireStart(ctx context.Context, req *chatRequest, kind, model string, msgs, tools int) {
ps := g.core.Plugins()
if ps == nil || ps.Count() == 0 {
return
}
ps.Fire(lua.StageRequestStart, map[string]interface{}{
"stage": string(lua.StageRequestStart),
"type": kind,
"model": model,
"key": keyID(reqKey(ctx)),
"role": reqRole(ctx),
"source": "",
"stream": req.Stream,
"messages_count": msgs,
"tools_count": tools,
"ts": time.Now().Unix(),
})
}
// fireRouted dispatches the plugin routed stage once a (source, model) slot has
// been selected. tier is the AUTO tier index, or -1 on the direct path, so a
// plugin can tell "this came from tier 1" from "this bypassed the chain".
func (g *Gateway) fireRouted(ctx context.Context, kind, source, model string, tier int, stream bool) {
ps := g.core.Plugins()
if ps == nil || ps.Count() == 0 {
return
}
ps.Fire(lua.StageRouted, map[string]interface{}{
"stage": string(lua.StageRouted),
"type": kind,
"source": source,
"model": model,
"key": keyID(reqKey(ctx)),
"tier": tier,
"stream": stream,
"ts": time.Now().Unix(),
})
}
// fireEnd dispatches the plugin request_end stage for one finished request.
func (g *Gateway) fireEnd(rec *Req) {
ps := g.core.Plugins()
if ps == nil || ps.Count() == 0 {
return
}
payload := map[string]interface{}{
"stage": string(lua.StageRequestEnd),
"type": rec.Type,
"model": rec.Model,
"source": rec.Source,
"key": rec.Key,
"ok": rec.OK,
"status": rec.Status,
"latency_ms": rec.LatMs,
"first_byte_ms": rec.FirstByteMs,
"prompt_tokens": rec.Prompt,
"completion_tokens": rec.Compl,
"cache_hit_tokens": rec.CacheHit,
"cache_miss_tokens": rec.CacheMiss,
"image_count": rec.ImageCount,
"error": rec.Err,
"time": rec.Time,
}
// The merged result is intentionally discarded: request_end is the last
// stage, so there is nobody downstream to read a plugin's additions. Plugins
// that need to publish derived numbers (the billing plugin) do it in their
// OWN state and expose them through the /api/plugins/<name>/state endpoint.
ps.Fire(lua.StageRequestEnd, payload)
}
// mergeUsage combines token usage across stream chunks additively. Some
@ -1039,6 +1128,7 @@ func (g *Gateway) streamChat(w http.ResponseWriter, ctx context.Context, cands [
// failover it differs from the first candidate). Direct streams previously
// discarded it.
rec.Source = usedSrc
g.fireRouted(ctx, "stream", usedSrc, rec.Model, -1, true)
rec.Prompt = estimatePromptTokens(req)
g.pumpStream(w, rec, chunks, effective, t0)
}
@ -1064,6 +1154,10 @@ func (g *Gateway) singleChatAuto(w http.ResponseWriter, ctx context.Context, cha
recordChatUsage(rec, req, resp)
rec.Source = usedSrc
rec.Model = usedModel
// AUTO has no single tier to report: the chain may have walked several
// before this slot served the request, so -2 means "resolved by the chain"
// and a plugin can tell that apart from the direct path's -1.
g.fireRouted(ctx, "chat", usedSrc, usedModel, -2, false)
rec.FirstByteMs = rec.LatMs
g.writeRec(rec)
writeChatCompletion(w, resp, usedModel)
@ -1091,6 +1185,7 @@ func (g *Gateway) streamChatAuto(w http.ResponseWriter, ctx context.Context, cha
rec.Model = usedModel
}
rec.Source = usedSrc
g.fireRouted(ctx, "stream", usedSrc, rec.Model, -2, true)
rec.Prompt = estimatePromptTokens(req)
g.pumpStream(w, rec, chunks, usedModel, t0)
}