mirror of
https://gitcode.com/JianFeeeee/ModelRouter.git
synced 2026-10-06 07:27:30 +00:00
feat(plugin): Lua 插件机制 + 计费插件 + 插件文档
插件 = plugin_dir 下的单个 .lua 文件,做两件事:挂请求流水线的钩子、在启动时
贡献 WebUI 界面(整页或往现有页面追加组件)。两者独立。
## 流水线 stage(三个)
request_start 已解析鉴权、未选源
routed 已选定 (source, model)、未发往上游
request_end 每请求恰好一次,带最终计量
request_end 挂在 gateway.writeRec——四条入口路径(直连/AUTO × 流式/非流式)的
唯一汇合点:既不漏(流式 token 只有流结束才知道)也不重。
## 计费插件(plugins/billing.lua,默认 seed,开箱可用)
源 / 模型 / 密钥三个维度定价。token 价优先级 keys > models > default;per_request
固定价是**叠加**的(生图模型可以既算 token 又收固定费)。单位是 USD/单 token,
即各家 provider 的公布口径。累计 total / by_source / by_model / by_key / by_day。
失败请求保留 token 费用、丢弃固定费(可经 count_failures 翻转)。
界面 = 一个独立页 + 状态页顶部一块总开销 tile。
## 一个明确的设计边界
计费插件**只报表,不执法**。网关自己的配额会计(stats.go,入口强制)才是限额
权威,插件不参与任何路由/配额决策。两套独立会计若对不上,比一套功能略少的
更糟。
## ★ 中途改掉的一个根本设计错误
最初让插件复用适配器的**弹性 worker 池**(多状态)。这对适配器是对的(它们无
状态),对插件是错的:计费插件往 plugin.state 累加,多状态意味着总量被劈成
几份;而 SetState 写价格只写进其中一个 worker,钩子恰好跑到另一个时**所有请求
按 0 计费**。改为**单状态 + 互斥锁**。代价写进文档:钩子必须短、同步、不阻塞,
卡住的钩子会卡住所有插件的钩子。
这个 bug 是测试逼出来的——先写了 SetState+Fire 的用例,数字全是 0 才挖出来。
另一个连带缺陷:只带 prices 的 PUT 会整体替换 state,把累计量清零。改为
prices/state 分离——prices 是配置、state 是历史,改价不动账。
## 撞到的三个 Lua 绑定的坑(都写进注释)
- SetGlobal **会 pop 栈**:连着调两次,第二次从空栈取,赋成 nil
- GetField 索引越界是 **SIGABRT 整个进程**,不是 panic,recover 救不了
- Call(nargs, n) **不接受函数索引**,它调的是 nargs 个参数正下方那个;
传索引会调到参数上("attempt to call a table value")
另外 GetField/SetField 用绝对索引,SetTop(0) 之后必须重取。
## 错误隔离
钩子 error() 不影响转发:捕获 → 记进 hook_errors → 跳下一个插件。适配器出错
会让源进冷却,插件出错**零惩罚**——插件是可选功能。/api/plugins 的 hook_errors
让"坏掉的插件"可见而不是静默消失。
## 界面注入
GET /api/ui-inject 一次返回所有插件的扩展(侧栏需要全部 page 才能建好)。
WebUI 在首次 render **之前** await 注入:先插 HTML 再重建 <script> 让它执行
(innerHTML/template 插入的 script 不会执行,这正是要的效果——避免脚本跑在
自己 DOM 之前)。注入失败不影响仪表盘。
browser 侧 pluginAPI 暴露 fetchState / postState / onTabShown。
## 文档
docs/plugins.md —— 快速上手、加载与热更新、三个 stage 的完整字段表、界面扩展、
状态与 HTTP API、运行时约束(单状态/异常隔离/内置函数)、计费插件的定价与
计费策略、排错表、与适配器的对比表。
## 判据(328 个测试全绿,插件相关 33 个)
- 计费断言的是**具体金额**(0.00625 / 0.0402 / 0.0075…),不是"能加载"
- 4 个变异都红:钩子异常不隔离 / prices 清空累计 / 忽略 key 优先级 /
毫秒时间戳不换算
- UI 侧 6 个判据把注入顺序、script 执行时机、pluginAPI 名称、tab 路由、
anchor 四种形式、失败非致命全钉住
- 鉴权:state 读任意角色、写仅 admin
This commit is contained in:
@ -14,6 +14,7 @@ import (
|
||||
"time"
|
||||
|
||||
"llmsproxy/internal/config"
|
||||
"llmsproxy/internal/lua"
|
||||
"llmsproxy/internal/provider"
|
||||
"llmsproxy/internal/scheduler"
|
||||
"llmsproxy/internal/types"
|
||||
@ -414,6 +415,7 @@ func (g *Gateway) handleChat(w http.ResponseWriter, r *http.Request) {
|
||||
Type: "chat",
|
||||
OK: false,
|
||||
}
|
||||
g.fireStart(ctx, &req, "chat", "AUTO", len(req.Messages), len(req.Tools))
|
||||
// quotaExhausted reports a slot whose token window has been used up;
|
||||
// exhausted slots are dropped from scheduling without penalty.
|
||||
quotaExhausted := func(sl *scheduler.Slot) bool {
|
||||
@ -817,6 +819,7 @@ func (g *Gateway) singleChat(w http.ResponseWriter, ctx context.Context, cands [
|
||||
recordChatUsage(rec, req, resp)
|
||||
rec.Source = usedSrc
|
||||
rec.Model = usedModel
|
||||
g.fireRouted(ctx, "chat", usedSrc, usedModel, -1, false)
|
||||
// Non-streaming: the whole response arrives at once, so TTFB equals
|
||||
// the total latency.
|
||||
rec.FirstByteMs = rec.LatMs
|
||||
@ -824,7 +827,18 @@ func (g *Gateway) singleChat(w http.ResponseWriter, ctx context.Context, cands [
|
||||
writeChatCompletion(w, resp, effective)
|
||||
}
|
||||
|
||||
// writeRec records a finished request (audit + aggregates).
|
||||
// writeRec records a finished request (audit + aggregates) and fires the
|
||||
// plugin request_end stage.
|
||||
//
|
||||
// This is the ONE place every request passes through on its way out, which is
|
||||
// what makes it the right hook point: the four entry points (single/stream ×
|
||||
// direct/auto) all funnel here, so a plugin sees each request exactly once with
|
||||
// its final accounting. Firing earlier would miss the streamed ones (their
|
||||
// numbers are only known once the stream finishes), and firing in each entry
|
||||
// point would mean four call sites to keep in sync.
|
||||
//
|
||||
// Hooks run AFTER the record is written: a plugin must not be able to delay or
|
||||
// lose the audit trail, and a plugin that throws is contained by Fire.
|
||||
func (g *Gateway) writeRec(rec *Req) {
|
||||
if rec == nil {
|
||||
return
|
||||
@ -833,6 +847,81 @@ func (g *Gateway) writeRec(rec *Req) {
|
||||
rec.Time = time.Now().UnixMilli()
|
||||
}
|
||||
g.stats.Record(*rec)
|
||||
g.fireEnd(rec)
|
||||
}
|
||||
|
||||
// fireStart dispatches the plugin request_start stage: the request has been
|
||||
// parsed and authorized but no upstream slot has been chosen yet, so `source`
|
||||
// is empty. A plugin that only wants volume/acceptance counts can subscribe
|
||||
// here and stay out of the per-request hot path entirely.
|
||||
func (g *Gateway) fireStart(ctx context.Context, req *chatRequest, kind, model string, msgs, tools int) {
|
||||
ps := g.core.Plugins()
|
||||
if ps == nil || ps.Count() == 0 {
|
||||
return
|
||||
}
|
||||
ps.Fire(lua.StageRequestStart, map[string]interface{}{
|
||||
"stage": string(lua.StageRequestStart),
|
||||
"type": kind,
|
||||
"model": model,
|
||||
"key": keyID(reqKey(ctx)),
|
||||
"role": reqRole(ctx),
|
||||
"source": "",
|
||||
"stream": req.Stream,
|
||||
"messages_count": msgs,
|
||||
"tools_count": tools,
|
||||
"ts": time.Now().Unix(),
|
||||
})
|
||||
}
|
||||
|
||||
// fireRouted dispatches the plugin routed stage once a (source, model) slot has
|
||||
// been selected. tier is the AUTO tier index, or -1 on the direct path, so a
|
||||
// plugin can tell "this came from tier 1" from "this bypassed the chain".
|
||||
func (g *Gateway) fireRouted(ctx context.Context, kind, source, model string, tier int, stream bool) {
|
||||
ps := g.core.Plugins()
|
||||
if ps == nil || ps.Count() == 0 {
|
||||
return
|
||||
}
|
||||
ps.Fire(lua.StageRouted, map[string]interface{}{
|
||||
"stage": string(lua.StageRouted),
|
||||
"type": kind,
|
||||
"source": source,
|
||||
"model": model,
|
||||
"key": keyID(reqKey(ctx)),
|
||||
"tier": tier,
|
||||
"stream": stream,
|
||||
"ts": time.Now().Unix(),
|
||||
})
|
||||
}
|
||||
|
||||
// fireEnd dispatches the plugin request_end stage for one finished request.
|
||||
func (g *Gateway) fireEnd(rec *Req) {
|
||||
ps := g.core.Plugins()
|
||||
if ps == nil || ps.Count() == 0 {
|
||||
return
|
||||
}
|
||||
payload := map[string]interface{}{
|
||||
"stage": string(lua.StageRequestEnd),
|
||||
"type": rec.Type,
|
||||
"model": rec.Model,
|
||||
"source": rec.Source,
|
||||
"key": rec.Key,
|
||||
"ok": rec.OK,
|
||||
"status": rec.Status,
|
||||
"latency_ms": rec.LatMs,
|
||||
"first_byte_ms": rec.FirstByteMs,
|
||||
"prompt_tokens": rec.Prompt,
|
||||
"completion_tokens": rec.Compl,
|
||||
"cache_hit_tokens": rec.CacheHit,
|
||||
"cache_miss_tokens": rec.CacheMiss,
|
||||
"image_count": rec.ImageCount,
|
||||
"error": rec.Err,
|
||||
"time": rec.Time,
|
||||
}
|
||||
// The merged result is intentionally discarded: request_end is the last
|
||||
// stage, so there is nobody downstream to read a plugin's additions. Plugins
|
||||
// that need to publish derived numbers (the billing plugin) do it in their
|
||||
// OWN state and expose them through the /api/plugins/<name>/state endpoint.
|
||||
ps.Fire(lua.StageRequestEnd, payload)
|
||||
}
|
||||
|
||||
// mergeUsage combines token usage across stream chunks additively. Some
|
||||
@ -1039,6 +1128,7 @@ func (g *Gateway) streamChat(w http.ResponseWriter, ctx context.Context, cands [
|
||||
// failover it differs from the first candidate). Direct streams previously
|
||||
// discarded it.
|
||||
rec.Source = usedSrc
|
||||
g.fireRouted(ctx, "stream", usedSrc, rec.Model, -1, true)
|
||||
rec.Prompt = estimatePromptTokens(req)
|
||||
g.pumpStream(w, rec, chunks, effective, t0)
|
||||
}
|
||||
@ -1064,6 +1154,10 @@ func (g *Gateway) singleChatAuto(w http.ResponseWriter, ctx context.Context, cha
|
||||
recordChatUsage(rec, req, resp)
|
||||
rec.Source = usedSrc
|
||||
rec.Model = usedModel
|
||||
// AUTO has no single tier to report: the chain may have walked several
|
||||
// before this slot served the request, so -2 means "resolved by the chain"
|
||||
// and a plugin can tell that apart from the direct path's -1.
|
||||
g.fireRouted(ctx, "chat", usedSrc, usedModel, -2, false)
|
||||
rec.FirstByteMs = rec.LatMs
|
||||
g.writeRec(rec)
|
||||
writeChatCompletion(w, resp, usedModel)
|
||||
@ -1091,6 +1185,7 @@ func (g *Gateway) streamChatAuto(w http.ResponseWriter, ctx context.Context, cha
|
||||
rec.Model = usedModel
|
||||
}
|
||||
rec.Source = usedSrc
|
||||
g.fireRouted(ctx, "stream", usedSrc, rec.Model, -2, true)
|
||||
rec.Prompt = estimatePromptTokens(req)
|
||||
g.pumpStream(w, rec, chunks, usedModel, t0)
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user