perf(plugin): 去掉钩子热路径的 JSON 往返 + 同 stage 跨插件并行

## 1. 去掉 JSON 往返(快路径)
实测单次 Fire 14.6µs,其中 json.Marshal 4.0 + json.Unmarshal 5.6 = 9.6µs,
**67% 花在把 map[string]interface{} 序列化再反序列化**,而紧接着的
pushGoValue 本来就能直接遍历这两种类型。改为按类型直接转换(fastvalue.go),
只对不认识���类型才回落 JSON —— 陌生字段仍然会被送到插件,而不是消失。

快路径与 JSON 路径逐字节等价由 TestFastPathMatchesJSONPath 锁住(8 组载荷,
覆盖 int/uint/float 各宽度、嵌套、slice、map[string]string、未知类型)。
还有一条专门防止「优化悄悄失效」:TestFastPathIsActuallyUsed 用真实的
request_end 载荷断言它确实走快路径。

同一份代码 A/B 实测:JSON 往返 56.4µs → 快路径 35.3µs(省 37%)。

## 2. 同 stage 跨插件并行
参照 /home/program/TrueAgent 的 StageHost.RunStage:
  - **快照后释放锁**再并行 —— 它记录过一次自死锁(p.Stop → onExit → ReclaimOwner
    要拿 registry 锁,持锁并行即死锁)。这里同理:钩子可能经 admin API 增删插件,
    那条路径要拿 ps.mu 写锁,所以并行段内不持任何 ps 锁。
  - 每个 goroutine recover。
  - 单插件走直连路径,不付 goroutine 代价(生产就是这种配置)。

**与 TrueAgent 不同的一点**:它可以放心并行,因为 handler 只写 ctx.Response 并有
IsResponded() 仲裁;我们的钩子返回 table 会合并进 payload,而
docs/plugins.md 明确承诺「payload 原样传给下一个插件」。所以合并**按插件加载
顺序**执行,结果确定,不依赖调度;代价是钩子之间不再互相可见 —— 这是一处
**契约变化**,已在文档里写明,并说明随核心发布的 billing 从不返回任何值
(代码注释就写着 "nobody downstream would read a return value")。

实测收益(真实二进制,三实例对照,3000 请求):

              无插件     1 插件      4 插件
  稳态并发32    849 rps   768 (-9.5%) 741 (-12.7%)
  突发并发64   1524-1893  1182-1676  1064-1443

4 插件只降 10-20%,而并行前实测 4 插件是 63.8µs vs 单插件 14.6µs(-300%)。

## ★ 我自己造成的两次性能事故
**① 持久化把热路径拖慢 26 倍。** 最初的快照在钩子路径上做:走 luaValueToGo +
json.Marshal + json.Unmarshal 三重转换,每请求 264µs,Fire 从 14.6µs 变成 385µs。
改成 saver 按自己节奏拉取(钩子只标记 dirty,flush 时才快照),385µs → 25.7µs。
**这里还踩了第二次 use-after-free**:让后台 goroutine 去读 Lua 表,vm.Stop() 后
那是已释放内存(SIGSEGV)。安全性现在由「Plugins.Close 等 saver 的最后一次
flush 完成后,调用方才停 VM」保证。

**② 基准被自己的后台写入污染。** 关掉 markDirty 反而测出 36µs、比开着还慢,
方向完全反了 —— 是 saver 每 2 秒写盘混进了计时。加了 DisableStatePersistence
后数据才可信。

## 判据(11 项,全部变异验证)
快路径等价/确实生效/不别名输入 + 并行与单插件路径合并一致 + 合并顺序确定 +
抛异常的钩子不拖累同伴 + 每插件恰好执行一次 + 并发 Fire 安全 + Fire 期间不持
注册表锁 + 真实 billing 在并行下正常 + 持久化 7 项。

变异:改坏合并顺序 → 红;去掉单插件路径的合并 → 红(3 个既有测试同时抓到)。
★ 「删掉 recover」这个变异**没有**让判据变红,查下去发现 golua 把 error()、
nil 索引、调用 nil、深递归全部转成 error RETURN,不产生 Go panic —— 那个测试
根本没测到 recover。已改名 TestThrowingHook 并在注释里写明 recover() 当前无法
被 Lua 触达,保留它是为了守 Go 侧。留一个「看起来有覆盖」的断言比没有更糟。

## 端到端(真实二进制 + 真实 billing)
20 万请求全 200,rps 1870,p99 96ms,RSS 37.9→42MB 有界;
负载停止后四个插件计数**完全一致**(231745),hook_errors 为空;
systemctl restart 后 billing 仍是 231745 —— 并行与持久化同时生效。
381 个测试全绿,含 -race。
This commit is contained in:
JianFeeeee
2026-10-02 10:45:20 +08:00
parent ce2032435a
commit 30064696b3
5 changed files with 760 additions and 55 deletions

View File

@ -240,7 +240,7 @@ type Plugins struct {
type stateSaver struct {
mu sync.Mutex
ps *Plugins
dirty map[string]pendingState
dirty map[string]bool
wake chan struct{}
stop chan struct{}
done chan struct{}
@ -257,26 +257,13 @@ type stateSaver struct {
writes atomic.Int64
}
// pendingState is a plugin's state ALREADY CONVERTED to Go values.
//
// The conversion happens on the caller's goroutine, inside p.mu, while the Lua
// VM is guaranteed alive. Doing it in the flush goroutine instead looked
// harmless and was a use-after-free: vm.Stop() frees the Lua states, and the
// saver then walked them — measured as a SIGSEGV inside golua, not as a clean
// error. A background goroutine must never touch the VM.
type pendingState struct {
name string
state interface{}
prices interface{}
}
// setWakeDelayForTest shortens the debounce so tests do not pay it on teardown.
func (s *stateSaver) setWakeDelayForTest(d time.Duration) { s.wakeDelay = d }
func newStateSaver(ps *Plugins) *stateSaver {
s := &stateSaver{
ps: ps,
dirty: map[string]pendingState{},
dirty: map[string]bool{},
wake: make(chan struct{}, 1),
stop: make(chan struct{}),
done: make(chan struct{}),
@ -287,16 +274,15 @@ func newStateSaver(ps *Plugins) *stateSaver {
return s
}
// mark records a snapshot as needing a write. Never blocks: the channel is
// buffered, and a full buffer means a flush is already pending, which is
// exactly the state we want. Only the LATEST snapshot per plugin is kept, so a
// burst of N mutations collapses to one write.
func (s *stateSaver) mark(p pendingState) {
// markDirty records that a plugin's state changed. Never blocks: the channel is
// buffered, and a full buffer means a flush is already pending, which is exactly
// the state we want.
func (s *stateSaver) markDirty(name string) {
if s == nil {
return
}
s.mu.Lock()
s.dirty[p.name] = p
s.dirty[name] = true
s.mu.Unlock()
select {
case s.wake <- struct{}{}:
@ -335,20 +321,17 @@ func (s *stateSaver) flush() {
s.mu.Unlock()
return
}
pending := make([]pendingState, 0, len(s.dirty))
for _, p := range s.dirty {
pending = append(pending, p)
names := make([]string, 0, len(s.dirty))
for n := range s.dirty {
names = append(names, n)
}
s.dirty = map[string]pendingState{}
s.dirty = map[string]bool{}
s.mu.Unlock()
// File I/O only. The Lua states are not touched here — see pendingState.
for _, p := range pending {
writeJSONAtomic(s.ps.stateFile(p.name), map[string]interface{}{
"version": 1,
"state": p.state,
"prices": p.prices,
})
// Snapshot + write here, off the request path. This is the ONLY place the
// plugin's Lua tables are walked for persistence.
for _, n := range names {
s.ps.persistByName(n)
s.writes.Add(1)
}
}
@ -435,6 +418,9 @@ func NewPlugins(vm *VM, dir string) *Plugins {
// accumulation is lost, which is the same "restart loses the books" defect this
// persistence exists to fix, just scoped to a few seconds instead of the whole
// process lifetime.
// DisableStatePersistence turns persistence off (tests/benchmarks only).
func (ps *Plugins) DisableStatePersistence() { ps.stateDirDisabled = true }
func (ps *Plugins) Close() {
if ps == nil {
return
@ -1018,7 +1004,8 @@ func readPricesFromPayload(state interface{}) interface{} {
// It is called from the hook path, so it must not block on file I/O — that is
// the saver's job. It DOES walk the Lua tables, because that has to happen
// while the caller still holds p.mu and the VM is guaranteed alive; the flush
// goroutine only ever sees the resulting Go values (see pendingState).
// (Snapshotting in the flush goroutine instead was a use-after-free: vm.Stop()
// frees the Lua states. See persistByName for why the walk is safe here.)
//
// Cost is one JSON conversion per request, which the billing plugin would pay
// anyway inside its own hook.
@ -1026,15 +1013,43 @@ func (ps *Plugins) markDirtyLocked(p *Plugin) {
if ps == nil || p == nil || ps.stateDirDisabled || p.state == nil || ps.dir == "" {
return
}
// Caller already holds p.mu — taking it again would self-deadlock. That is
// not hypothetical: the first version had markDirty lock p.mu and was called
// from invoke(), which holds p.mu for the whole Lua call, so every hook call
// deadlocked and the lua package's tests hung until the timeout.
state, prices := readStateAndPrices(p.state.L)
if state == nil {
// NOTE: this only records THAT the state changed. It deliberately does not
// snapshot it.
//
// Snapshotting here cost 264us per request — the snapshot walks the plugin's
// Lua tables through luaValueToGo + json.Marshal + json.Unmarshal (three
// conversions), and the billing plugin does it on every single request. That
// turned a 14.6us hook into a 385us one and put the gateway's hot path
// behind a bookkeeping step. The saver pulls the snapshot on its own schedule
// instead, so N requests between flushes cost ONE snapshot.
ps.save.markDirty(p.Info.Name)
}
// persistByName snapshots one plugin and writes its state file.
//
// It is called from the saver goroutine, which means it walks a Lua state that
// vm.Stop() may be about to free. That was a real crash: the first version had
// the saver read the tables itself and the process died with SIGSEGV inside
// golua once shutdown raced a flush.
//
// Two things make it safe now:
// - Plugins.Close() stops the saver AND waits for its final flush BEFORE the
// caller stops the VM (see Plugins.Close), so no flush can be in flight when
// the Lua states go away.
// - p.mu is held while the tables are read, so a hook cannot be mutating them
// underneath. A hook that arrives after the VM stopped cannot run either,
// because the request path is already torn down by then.
func (ps *Plugins) persistByName(name string) {
if ps == nil || ps.stateDirDisabled || ps.dir == "" {
return
}
ps.save.mark(pendingState{name: p.Info.Name, state: state, prices: prices})
ps.mu.RLock()
p := ps.find(name)
ps.mu.RUnlock()
if p == nil || p.state == nil {
return
}
ps.persistNow(p)
}
// ---------- state persistence ----------
@ -1296,6 +1311,19 @@ func (ps *Plugins) Fire(stage Stage, payload map[string]interface{}) map[string]
if len(calls) == 0 {
return payload
}
// Resolve the plugins and drop the lock BEFORE running any hook.
//
// A hook can install, disable or remove a plugin (the admin API is reachable
// from a hook that has the key), and those paths take ps.mu for write. Taking
// a per-iteration RLock like the single-plugin case did works but serializes
// on a shared cache line under load; a snapshot plus one lookup is cheaper and
// — more importantly — it means the parallel section below never holds ps.mu,
// so a hook cannot deadlock against a concurrent reload.
type target struct {
p *Plugin
hc hookCall
}
targets := make([]target, 0, len(calls))
for _, hc := range calls {
ps.mu.RLock()
p := ps.plugins[hc.pluginIdx]
@ -1303,20 +1331,82 @@ func (ps *Plugins) Fire(stage Stage, payload map[string]interface{}) map[string]
if p == nil || p.LoadError != "" {
continue
}
out, err := ps.invoke(p, hc.fn, payload)
if err != nil {
ps.hookErr.note(stage, p.Info.Name+": "+err.Error())
continue
}
if len(out) > 0 {
targets = append(targets, target{p: p, hc: hc})
}
if len(targets) == 0 {
return payload
}
// One plugin is the overwhelmingly common case (a gateway with the bundled
// billing plugin has exactly one). Paying for goroutines there would be pure
// overhead, so it takes the direct path.
if len(targets) == 1 {
if out := ps.runHook(stage, targets[0].p, targets[0].hc.fn, payload); len(out) > 0 {
for k, v := range out {
payload[k] = v
}
}
return payload
}
// Parallel across plugins. Each hook receives the SAME payload snapshot, and
// the returned tables are merged afterwards IN PLUGIN LOAD ORDER, so the
// documented contract ("a returned table's keys are merged into the payload")
// still holds deterministically.
//
// What changes: a hook no longer sees the keys another hook just added. The
// sequential behaviour made that possible, and docs/plugins.md described it
// ("the payload is passed to the next plugin unchanged"). It was never used —
// the bundled billing plugin returns nil on every stage, with a comment saying
// nobody downstream reads it — but it IS a contract change and is called out
// there rather than left as a surprise.
//
// Why this is safe for the per-plugin lock: each plugin has its own Lua state
// and its own p.mu, so two plugins never touch the same state. What is shared
// is the payload, and it is only READ here — merging happens after every hook
// has returned, on the caller's goroutine.
outs := make([]map[string]interface{}, len(targets))
var wg sync.WaitGroup
for i := range targets {
wg.Add(1)
go func(i int) {
defer wg.Done()
// A hook that panics would otherwise take the whole process with it,
// and in the sequential version a panic could not escape Fire either.
// recover() here restores that: the plugin is skipped and noted.
defer func() {
if rec := recover(); rec != nil {
ps.hookErr.note(stage, fmt.Sprintf("%s: panic in hook: %v", targets[i].hc.plugin, rec))
}
}()
outs[i] = ps.runHook(stage, targets[i].p, targets[i].hc.fn, payload)
}(i)
}
wg.Wait()
// Merge in load order so the result does not depend on goroutine scheduling.
// A later plugin's value wins on a key collision, exactly as it did when the
// hooks ran one after another.
for _, out := range outs {
for k, v := range out {
payload[k] = v
}
}
return payload
}
// runHook invokes one hook and folds its returned table into payload. It
// contains the plugin's error handling so both the sequential and the parallel
// path behave identically on failure.
func (ps *Plugins) runHook(stage Stage, p *Plugin, fn string, payload map[string]interface{}) map[string]interface{} {
out, err := ps.invoke(p, fn, payload)
if err != nil {
ps.hookErr.note(stage, p.Info.Name+": "+err.Error())
return nil
}
return out
}
// invoke runs one plugin hook on that plugin's own state, under its pool's
// concurrency cap. The plugin's returned table is re-fetched each call because
// the pool is elastic: a plugin may have several states, and the hook function
@ -1348,15 +1438,13 @@ func (ps *Plugins) invoke(p *Plugin, fn string, payload map[string]interface{})
// themselves), but a plugin hook receives a table so it can read
// payload.model directly. Passing the string made every hook fail with
// "attempt to index local 'payload' (a string value)".
raw, err := json.Marshal(payload)
if err != nil {
return nil, err
}
var decoded interface{}
if err := json.Unmarshal(raw, &decoded); err != nil {
return nil, err
}
pushGoValue(L, decoded)
//
// Decoding is done WITHOUT JSON. json.Marshal + json.Unmarshal here cost
// 9.6us of the 14.6us hook — 67% of the call spent re-deriving a tree that
// pushGoValue walks natively anyway. plainForLua does the same conversion
// by type-switching, and falls back to JSON only for a type it does not
// model, so an exotic payload still arrives instead of vanishing.
pushGoValue(L, plainForLua(payload))
// Call takes NO function index: it invokes whatever sits directly below the
// nargs values it just pushed. Passing an index here is a compile-time no-op
// in this binding and the call lands on the argument instead