mirror of
https://gitcode.com/JianFeeeee/ModelRouter.git
synced 2026-10-06 15:37:42 +00:00
feat(plugin): AUTO 调度轨迹可见(chain_step stage)
被问"还有 auto 调度相关 stage 呢?"问出来的真实缺口。
## 问题
chainDrive 只返回 (resp, src, model, err),调用方只知道**最终哪个槽位赢了**。
遍历过程中算出来又丢掉的东西——哪些档被跳过、为什么跳过、哪些槽位硬失败、
哪档全忙——一律不可见。ChainErr 里其实有这些,但**只在全部失败时**才填,
而它是 error 返回值不是记录。于是:
"tier 1 冷却所以降级到 tier 3" == "tier 1 正常接单"
对插件而言 tier 只是个常量 -2("resolved by the chain"),信息量为零。而这
恰恰是优先级链存在的全部理由,也是"我那个贵模型为什么没被用"的答案。
## 做法(scheduler 侧零新依赖)
新增 TraceEvent / TraceSink,chainDrive 多一个可选 sink 参数:
- TraceEvent 是本包的普通 struct,sink 是 func 参数 ⇒ **不新增 import**,
scheduler 仍然可独立测试
- sink 为 nil 时每次 emit 只多一次 nil 判断;没有插件的网关在 AUTO 热路径上
零开销(gateway 的 chainTraceSink 直接返回 nil)
- 事件是纯观测:scheduler 不基于它做任何分支,gateway 也不把它喂回路由/
冷却/配额
四种 kind:tier_skip / slot_fail / tier_busy / selected,selected 每次成功
遍历恰好一次且是最后一步。顺序保证所有 step 在 routed 之前。
## 暴露给插件
新增 chain_step stage(逐个步骤),并在 request_end 载荷里加三个便于做报表的
字段:chain_walk(上限 12 步,防审计记录膨胀)、degraded、tier_served。
## ★ 计费口径(我按推荐的做,已写进文档,需要你确认)
**按实际服务的模型计费**:降级到 tier 3 仍按 tier 3 的价算,轨迹只作观测。
理由与 §7.5 的边界一致——插件只报表不执法,两套口径混在一起会引出"降级该不该
多收钱"这种无法从代码判断的争议。若要改成"按本该用的档计价",需要在 models
价目里允许按 tier 定价,这我没做,因为那是个产品决策。
## 计费插件同步消费
by_tier_served / skip_reasons / degraded_reqs 三个新维度。skip_reasons 的等待
时长做了归一(`no free slot within <wait>`),否则 busy-wait 文案一变就多一行。
降级次数在 request_end 里计而不是在 chain_step 里计:一次降级的请求要走多步,
按步计会重复计数。
## 判据(346 个测试全绿,新增 15 个)
scheduler 6 个:正常路径只发一个 selected / 跳档+降级可见 / 硬失败与跳档
严格区分(不可混为一谈,否则抖动上游看起来像空闲上游)/
nil sink 安全 / 全失败时轨迹与 ChainErr 并存且不互相破坏 /
空链不发事件
gateway 1 个端到端:tier 1 全 500 → 插件收到 slot_fail(tier 1) +
selected(tier 2),request_end 的 tier_served=2 且 degraded=true
lua 2 个:降级计数与按实际模型计价 / 跳过原因归一聚合
lua 1 个:chain_step 是真 stage 且顺序正确
3 个变异都红:去掉 slot_fail(3 个判据红)/ 去掉 tier_skip(1 个)/
去掉 degraded 字段(1 个)。
This commit is contained in:
@ -293,13 +293,75 @@ func runTier(ctx context.Context, tn *TierNode, cands []candidate, base int64, r
|
||||
return tierResult{hard: hard}
|
||||
}
|
||||
|
||||
// TraceKind classifies one step of an AUTO chain walk.
|
||||
type TraceKind string
|
||||
|
||||
const (
|
||||
// TraceTierSkip: the whole tier was skipped — every slot was cooling,
|
||||
// quota-exhausted, or none was schedulable. Reason says which.
|
||||
TraceTierSkip TraceKind = "tier_skip"
|
||||
// TraceSlotFail: one slot failed hard (upstream error / bad adapter). The
|
||||
// walk continues to the next slot or tier.
|
||||
TraceSlotFail TraceKind = "slot_fail"
|
||||
// TraceTierBusy: the tier was fully busy and the bounded wait expired.
|
||||
TraceTierBusy TraceKind = "tier_busy"
|
||||
// TraceSelected: this slot served the request. Exactly one per successful
|
||||
// chain walk, and the last event emitted.
|
||||
TraceSelected TraceKind = "selected"
|
||||
)
|
||||
|
||||
// TraceEvent is one observable step of an AUTO chain walk.
|
||||
//
|
||||
// WHY THIS EXISTS: chainDrive's return value is (resp, src, model, err), so a
|
||||
// caller learns only which slot finally served the request. Everything the
|
||||
// scheduler decided on the way there — which tiers it skipped and WHY, which
|
||||
// slots hard-failed, whether a tier was merely busy — was computed and then
|
||||
// discarded. That is invisible to operators and to plugins: "tier 1 was cooling
|
||||
// so we degraded to tier 3" looked exactly like "tier 1 served it".
|
||||
//
|
||||
// The walk already accumulates this in ChainErr, but ONLY on total failure, and
|
||||
// ChainErr is an error return, not a record. Emitting a trace as it happens
|
||||
// covers the far more common case: a request that SUCCEEDED after degrading.
|
||||
//
|
||||
// Design constraints:
|
||||
// - scheduler stays dependency-free and independently testable. A TraceEvent
|
||||
// is a plain struct in this package and the sink is a func parameter, so no
|
||||
// import is added and no test has to change to observe a walk.
|
||||
// - The sink is optional (nil = emit nothing). The overhead on the hot path
|
||||
// is one nil check per event.
|
||||
// - Events are OBSERVATION ONLY. Nothing in the scheduler branches on them,
|
||||
// and the gateway does not feed them back into routing, cooldown or quota —
|
||||
// see docs/plugins.md for why accounting and enforcement are kept apart.
|
||||
type TraceEvent struct {
|
||||
Kind TraceKind
|
||||
Tier int
|
||||
Source string
|
||||
Model string
|
||||
Reason string // human-readable, for TraceTierSkip / TraceSlotFail
|
||||
Err string // the underlying error text, for TraceSlotFail
|
||||
// Attempt counts the 1-based slot attempt within the whole walk.
|
||||
Attempt int
|
||||
}
|
||||
|
||||
// TraceSink receives chain-walk events. It must not block: it is called from the
|
||||
// request path, and a slow sink slows the request.
|
||||
type TraceSink func(TraceEvent)
|
||||
|
||||
// chainDrive runs a request down the chain (plan 2.3): tiers ascending (tier
|
||||
// 1, the highest priority, first), per-tier round-robin starting at the tier
|
||||
// cursor, same-tier runs ordered by preference (negative prefs sink but stay
|
||||
// reachable). Quota-exhausted and cooling slots are filtered up front; a
|
||||
// fully busy tier is polled for a bounded time before falling through.
|
||||
// Failures are summarized in *ChainErr for the caller to map to HTTP 503.
|
||||
func (s *Scheduler) chainDrive(ctx context.Context, chain *Chain, req *types.ChatRequest, exhausted func(*Slot) bool, stream bool) (*types.UnifiedResponse, <-chan types.UnifiedChunk, string, string, error) {
|
||||
//
|
||||
// trace may be nil; when set it receives one event per observable step.
|
||||
func (s *Scheduler) chainDrive(ctx context.Context, chain *Chain, req *types.ChatRequest, exhausted func(*Slot) bool, stream bool, trace TraceSink) (*types.UnifiedResponse, <-chan types.UnifiedChunk, string, string, error) {
|
||||
emit := func(ev TraceEvent) {
|
||||
if trace != nil {
|
||||
trace(ev)
|
||||
}
|
||||
}
|
||||
attempt := 0
|
||||
if chain == nil || len(chain.Tiers) == 0 {
|
||||
return nil, nil, "", "", fmt.Errorf("no auto slot configured")
|
||||
}
|
||||
@ -309,7 +371,9 @@ func (s *Scheduler) chainDrive(ctx context.Context, chain *Chain, req *types.Cha
|
||||
// dropped unless they qualify as half-cooldown probes (appended last).
|
||||
cands := collectCands(tn.Slots, exhausted)
|
||||
if len(cands) == 0 {
|
||||
ce.Skipped = append(ce.Skipped, fmt.Sprintf("tier %d: no schedulable slot (cooling or quota exhausted)", tn.Tier))
|
||||
reason := "no schedulable slot (cooling or quota exhausted)"
|
||||
ce.Skipped = append(ce.Skipped, fmt.Sprintf("tier %d: %s", tn.Tier, reason))
|
||||
emit(TraceEvent{Kind: TraceTierSkip, Tier: tn.Tier, Reason: reason})
|
||||
continue
|
||||
}
|
||||
// No Pref sort: load balancing is done by round-robin cursor.
|
||||
@ -318,6 +382,8 @@ func (s *Scheduler) chainDrive(ctx context.Context, chain *Chain, req *types.Cha
|
||||
base := tn.NextStart()
|
||||
res := runTier(ctx, tn, cands, base, req, stream)
|
||||
if res.resp != nil || res.chunks != nil {
|
||||
attempt++
|
||||
emit(TraceEvent{Kind: TraceSelected, Tier: tn.Tier, Source: res.src, Model: res.model, Attempt: attempt})
|
||||
releaseProbes(cands)
|
||||
return res.resp, res.chunks, res.src, res.model, nil
|
||||
}
|
||||
@ -327,6 +393,13 @@ func (s *Scheduler) chainDrive(ctx context.Context, chain *Chain, req *types.Cha
|
||||
}
|
||||
if len(res.hard) > 0 {
|
||||
ce.Tiers = append(ce.Tiers, res.hard...)
|
||||
for _, h := range res.hard {
|
||||
attempt++
|
||||
emit(TraceEvent{
|
||||
Kind: TraceSlotFail, Tier: tn.Tier, Source: h.Source, Model: h.Model,
|
||||
Err: types.OneLine(h.Err.Error(), 200), Attempt: attempt,
|
||||
})
|
||||
}
|
||||
releaseProbes(cands)
|
||||
continue // hard failures: fall through to the next tier, no waiting
|
||||
}
|
||||
@ -334,10 +407,13 @@ func (s *Scheduler) chainDrive(ctx context.Context, chain *Chain, req *types.Cha
|
||||
if err := s.pollBusyTier(ctx, tn, cands, base, req, stream, &ce); err != nil {
|
||||
releaseProbes(cands)
|
||||
if r, ok := err.(*tierSuccess); ok {
|
||||
attempt++
|
||||
emit(TraceEvent{Kind: TraceSelected, Tier: tn.Tier, Source: r.res.src, Model: r.res.model, Attempt: attempt})
|
||||
return r.res.resp, r.res.chunks, r.res.src, r.res.model, nil
|
||||
}
|
||||
return nil, nil, "", "", err
|
||||
}
|
||||
emit(TraceEvent{Kind: TraceTierBusy, Tier: tn.Tier, Reason: fmt.Sprintf("no free slot within %v", busyWait)})
|
||||
releaseProbes(cands)
|
||||
}
|
||||
if len(ce.Tiers) == 0 && len(ce.Skipped) == 0 {
|
||||
@ -401,16 +477,16 @@ func (s *Scheduler) pollBusyTier(ctx context.Context, tn *TierNode, cands []cand
|
||||
// non-nil, decides slot token-quota exhaustion. Returns the response, the
|
||||
// serving source and the exact model id used; on total failure a *ChainErr
|
||||
// summarizing every tier.
|
||||
func (s *Scheduler) ChainChat(ctx context.Context, chain *Chain, req *types.ChatRequest, exhausted func(*Slot) bool) (*types.UnifiedResponse, string, string, error) {
|
||||
resp, _, src, model, err := s.chainDrive(ctx, chain, req, exhausted, false)
|
||||
func (s *Scheduler) ChainChat(ctx context.Context, chain *Chain, req *types.ChatRequest, exhausted func(*Slot) bool, trace TraceSink) (*types.UnifiedResponse, string, string, error) {
|
||||
resp, _, src, model, err := s.chainDrive(ctx, chain, req, exhausted, false, trace)
|
||||
return resp, src, model, err
|
||||
}
|
||||
|
||||
// ChainChatStream runs a streaming AUTO request down the chain. A slot is
|
||||
// abandoned only on connect failures / busy (before its first chunk); after a
|
||||
// stream starts it is pinned. Same return contract as ChainChat.
|
||||
func (s *Scheduler) ChainChatStream(ctx context.Context, chain *Chain, req *types.ChatRequest, exhausted func(*Slot) bool) (<-chan types.UnifiedChunk, string, string, error) {
|
||||
_, chunks, src, model, err := s.chainDrive(ctx, chain, req, exhausted, true)
|
||||
func (s *Scheduler) ChainChatStream(ctx context.Context, chain *Chain, req *types.ChatRequest, exhausted func(*Slot) bool, trace TraceSink) (<-chan types.UnifiedChunk, string, string, error) {
|
||||
_, chunks, src, model, err := s.chainDrive(ctx, chain, req, exhausted, true, trace)
|
||||
return chunks, src, model, err
|
||||
}
|
||||
|
||||
|
||||
Reference in New Issue
Block a user