mirror of
https://gitcode.com/JianFeeeee/ModelRouter.git
synced 2026-10-03 23:54:06 +00:00
三个问题都来自生产实测,不是代码审阅。
## 1. 缓存命中被计费却不被统计
网关确实从上游 usage 提取了 prompt_cache_hit_tokens(审计里能看到
cache_hit_tokens: 270104 / cache_reported: true,占 prompt 的 99.9%),
costFor() 也用它给缓存段定价了 —— 但**没有任何 bucket 记录它**。
结果:一个 99.88% 命中率的网关,报表显示 prompt_tokens 却看不出其中
多少是缓存读,也无从按源/模型/key 看命中率。
每个 bucket 现在多三个字段:
cache_hit_tokens 命中数(按上游上报)
cache_fresh_tokens 未命中的 prompt
cache_reported_reqs 上游确实上报了缓存数的请求数
第三个字段是刻意的:**「零命中」与「上游根本不上报」在命中总量里完全一样**,
而它们在「缓存折扣有没有生效」这个问题上含义相反。没有它就无法区分,
只能猜。
chat.go 的 payload 之前**没有** cache_reported(审计有、插件没有),
所以任何插件侧的缓存统计都只能猜 —— 已补上。
旧 state 文件的 bucket 没有这些字段:Lua 里 nil + number 会抛错,而钩子抛错
会让**该请求完全不记账**(一个统计缺口会变成静默缺口)。add() 里做了回填。
UI 增加 fresh/cache/cache% 三列 + Cache hit rate KPI;未上报的显示 n/r 而不是 0%。
## 2. Billing 页空白:render() 引用了未定义的 s
`render(st)` 里两处 KPI 写成 `s.degraded_reqs`,ReferenceError 让整个渲染
中断,所有表格停在初始的空 innerHTML。症状是「页面加载了但什么都没有」,
而 /api/plugins/billing/state 返回 200 且有真实数据 —— 载荷完全正确,
DOM 是空的。
更糟的是 refresh() 里的 `catch (e) { /* never break the page */ }` 把错误
**静默吞掉**了:网络面板一切正常,页面什么都没有。现在 catch 会
console.error(仍然不抛,装饰性组件不该拖垮宿主页,但必须留痕)。
## 3. 侧栏图标
billing 声明 icon = "💰",而原生 tab 全是内联 SVG(stroke: currentColor)。
emoji 尺寸不对、不跟随主题。
WebUI 增加 pluginIconHTML:插件图标可以是文本,也可以是内联 SVG。
**SVG 走严格白名单**(tag + 属性都是 allowlist,不是 denylist)——
插件是在运维者浏览器里跑的第三方代码,不能"信任插件";但也不能直接拒绝
SVG,因为那是唯一能和原生 tab 视觉一致的方式。
用真实 Chromium 验证 12 个用例,全部挡住,包括 foreignObject 里嵌 HTML
命名空间 <img onerror> 这个经典绕过(整体丢弃,所以 img/onerror 也没了)。
★ node 里没有 DOMParser/jsdom,所以没法在单测里跑这个过滤器 —— 用正则近似
会得到一个"测试通过但浏览器里失效"的过滤器,这比没有测试更糟。
顺带修了过滤器的两个真缺陷:输出里嵌套了空 `<svg></svg>`,且 viewBox
是从包装元素读的(永远是 null)而不是插件自己的,所以任何自定义 viewBox
的图标都会丢失。
## 判据(新增 7 项,全部变异验证)
写「注入脚本能否正常执行」这个守卫时我错了四次:
1. 静态扫「已声明的名字」→ 把 HTML 字符串里的 CSS 类名(class/div/td)
全报成未定义
2. 用 CSS 选择器解析器查样式表 → 报样式表本身坏了
3. 只挂 process 的 uncaughtException → 脚本在 IIFE 里异步跑,错误是
unhandledRejection,判据对原 bug 全绿
4. 只查「有没有抛错」→ render() 开头是 `if (!st) return`,传错字段是
**静默 no-op**:不抛、不打日志、不报错,只是页面空白
最终判据是:在 node 里用 DOM stub 真跑一遍,同时要求「无异常」且
「至少写进一个容器」,并监听 console.error。变异验证:还原 s → 红;
render 收到 undefined 字段 → 红。
表头/行列数一致性也有守卫:row() 加了缓存列而表头没加时,表格会整体错位
(cache% 落到 completion 列下)—— 渲染正常、有数据、但要仔细看才发现。
## 生产验证
重启后价目表与累计账完整保留(1.17 亿 prompt tokens)。
新请求缓存统计生效:cache_hit 947,436 / cache_fresh 888,
cache_reported_reqs 7 / 395(其余来自旧 state,正是该字段存在的意义)。
真实浏览器:表格 3 行、KPI 7 项、表头 name/cost/reqs/prompt/fresh/cache/cache%/completion、
SVG 图标 currentColor 渲染、控制台无 billing 错误。391 个测试全绿。
## 另发现一个无关 bug(未修)
首页 stats 图表抛 IndexSizeError: arc 半径为负(-2),在 ui/index.html 的
paintStats 附近。属状态页图表,不在本次范围。
639 lines
28 KiB
Lua
639 lines
28 KiB
Lua
-- billing.lua — usage accounting plugin for ModelRouter.
|
||
--
|
||
-- Computes what each request cost, from three configurable dimensions:
|
||
--
|
||
-- source a flat per-request price for an upstream source
|
||
-- model a per-token price for a model id (prompt / completion separately)
|
||
-- key an override price for one gateway key
|
||
--
|
||
-- It then keeps running totals for the whole gateway, per source, per model
|
||
-- and per key, and publishes them in `plugin.state` so the kernel can serve
|
||
-- them at GET /api/plugins/billing/state — which is what its own dashboard
|
||
-- component reads.
|
||
--
|
||
-- ACCOUNTING BOUNDARY (important, and deliberate):
|
||
-- this plugin REPORTS; it does not ENFORCE. The gateway's own quota accounting
|
||
-- (internal/gateway/stats.go, enforced at request admission) stays
|
||
-- authoritative for limits. Two independent accounting paths that disagree are
|
||
-- worse than one that is slightly less featureful, so nothing here feeds back
|
||
-- into routing or quota decisions.
|
||
--
|
||
-- PRICE CONFIGURATION
|
||
-- Prices are supplied as a Lua table assigned to `billing.prices` before the
|
||
-- plugin is loaded, OR at runtime through PUT /api/plugins/billing/state. The
|
||
-- shape is:
|
||
--
|
||
-- billing.prices = {
|
||
-- currency = "USD", -- display only, no conversion happens
|
||
-- default = { prompt = 0, completion = 0, per_request = 0 },
|
||
-- sources = {
|
||
-- ["localzen"] = { per_request = 0.0 },
|
||
-- ["trae"] = { per_request = 0.01 },
|
||
-- },
|
||
-- models = {
|
||
-- ["gpt-5.4"] = { prompt = 1.25e-6, completion = 1e-5 }, -- USD per TOKEN
|
||
-- ["kimi-k3"] = { prompt = 6e-7, completion = 2.5e-6 },
|
||
-- ["kolors"] = { per_request = 0.04 }, -- image: flat
|
||
-- },
|
||
-- keys = {
|
||
-- -- by gateway key (the same value the audit log masks to ***xxxxxx)
|
||
-- ["***a1b2c3"] = { prompt = 1.1e-6, completion = 9e-6 },
|
||
-- },
|
||
-- }
|
||
--
|
||
-- Precedence for a token price: keys > models > default. A flat per_request
|
||
-- price, when present at any level, is ADDED on top of the token cost, so an
|
||
-- image model can carry both (e.g. tokens billed plus a fixed fee).
|
||
--
|
||
-- Numbers are USD per single token, which is how providers publish prices. That
|
||
-- makes a typical entry look like 1.25e-6; the plugin multiplies by the token
|
||
-- count, so no unit conversion happens anywhere.
|
||
|
||
local plugin = {
|
||
name = "billing",
|
||
version = "1.0.0",
|
||
description = "Per-source / per-model / per-key cost accounting with a dashboard",
|
||
author = "ModelRouter",
|
||
}
|
||
|
||
-- ---------- prices ----------
|
||
|
||
-- plugin.prices can be pre-seeded by embedding this file (an operator edits the
|
||
-- table below) or replaced at runtime through the state API. It is a SEPARATE
|
||
-- field from plugin.state on purpose: PUT /api/plugins/billing/state replaces
|
||
-- `state` wholesale, and prices must not live there or a price update would
|
||
-- wipe the accumulated totals. See docs/plugins.md.
|
||
local DEFAULT_PRICES = {
|
||
currency = "USD",
|
||
default = { prompt = 0, completion = 0, per_request = 0 },
|
||
sources = {},
|
||
models = {},
|
||
keys = {},
|
||
}
|
||
plugin.prices = DEFAULT_PRICES
|
||
|
||
-- Default prompt-cache discount. 0.1 = a cache read costs a tenth of a fresh
|
||
-- token, which is what DeepSeek/Qwen/Kimi and most others charge. It can be
|
||
-- overridden per price entry (prices.models.<m>.cache_discount) or globally by
|
||
-- setting plugin.cache_discount; 1 restores flat prompt pricing.
|
||
plugin.cache_discount = 0.1
|
||
|
||
-- ---------- accumulated totals ----------
|
||
|
||
-- state is what the kernel serves at GET /api/plugins/billing/state. It holds
|
||
-- ACCUMULATED TOTALS ONLY — prices live in plugin.prices (see above), so
|
||
-- replacing state never destroys a price table and updating prices never
|
||
-- destroys history.
|
||
--
|
||
-- Structure:
|
||
-- total { cost, requests, prompt_tokens, completion_tokens }
|
||
-- by_source { <name> = { cost, requests, ...tokens } }
|
||
-- by_model { <model> = { cost, ... } }
|
||
-- by_key { <masked key id> = { cost, ... } }
|
||
-- by_day { "YYYY-MM-DD" = { cost, ... } }
|
||
-- top_sources [ {name, cost, requests}, ... ] sorted, capped
|
||
-- top_models [ ... ]
|
||
-- top_keys [ ... ]
|
||
--
|
||
-- Sorted top-N lists are maintained incrementally rather than re-sorted on
|
||
-- every request: this hook runs once per request on the hot path, so it does
|
||
-- map updates only. The sort happens when state is READ.
|
||
-- CACHE ACCOUNTING (added after production showed the gap):
|
||
-- the gateway extracts prompt_cache_hit_tokens from upstream usage and puts it
|
||
-- in the request_end payload, and costFor() already used it to price the cache
|
||
-- leg — but no bucket recorded it. So a gateway where 99.88% of prompt tokens
|
||
-- were cache reads showed a prompt_tokens number with no indication of that,
|
||
-- and there was no way to see cache hit rate per source/model/key at all.
|
||
--
|
||
-- cache_hit_tokens hits, as reported by upstream
|
||
-- cache_fresh_tokens prompt tokens that were NOT cache reads
|
||
-- cache_reported_reqs requests where upstream gave a cache number at all.
|
||
-- Kept separate from a zero: "upstream does not report cache usage" and
|
||
-- "upstream reported zero hits" look identical in a hit total, and they mean
|
||
-- opposite things when you are trying to work out whether a cache discount is
|
||
-- doing anything.
|
||
local function emptyBucket()
|
||
return {
|
||
cost = 0, requests = 0, prompt_tokens = 0, completion_tokens = 0, failures = 0,
|
||
cache_hit_tokens = 0, cache_fresh_tokens = 0, cache_reported_reqs = 0,
|
||
}
|
||
end
|
||
|
||
plugin.state = {
|
||
total = emptyBucket(),
|
||
by_source = {},
|
||
by_model = {},
|
||
by_key = {},
|
||
by_day = {},
|
||
started = os.time and 0 or 0,
|
||
}
|
||
|
||
local function bucket(tbl, k)
|
||
local b = tbl[k]
|
||
if b == nil then
|
||
b = emptyBucket()
|
||
tbl[k] = b
|
||
end
|
||
return b
|
||
end
|
||
|
||
local function add(b, cost, prompt, completion, ok, cacheHit, cacheReported)
|
||
b.cost = b.cost + cost
|
||
b.requests = b.requests + 1
|
||
b.prompt_tokens = b.prompt_tokens + prompt
|
||
b.completion_tokens = b.completion_tokens + completion
|
||
if not ok then b.failures = b.failures + 1 end
|
||
-- Backfill guards a bucket that predates these fields (a state file written
|
||
-- by an older build, or one restored from disk): nil + number is an error in
|
||
-- Lua, and a hook that throws stops accounting for that request entirely.
|
||
if b.cache_hit_tokens == nil then b.cache_hit_tokens = 0 end
|
||
if b.cache_fresh_tokens == nil then b.cache_fresh_tokens = 0 end
|
||
if b.cache_reported_reqs == nil then b.cache_reported_reqs = 0 end
|
||
b.cache_hit_tokens = b.cache_hit_tokens + (cacheHit or 0)
|
||
b.cache_fresh_tokens = b.cache_fresh_tokens + ((prompt or 0) - (cacheHit or 0))
|
||
if cacheReported then b.cache_reported_reqs = b.cache_reported_reqs + 1 end
|
||
end
|
||
|
||
-- ---------- pricing ----------
|
||
|
||
-- lookup walks keys > models > default and returns a price triple plus whether
|
||
-- a flat per_request component applies.
|
||
local function priceFor(payload)
|
||
local p = plugin.prices or DEFAULT_PRICES
|
||
local d = p.default or {}
|
||
-- Whether ANY dimension actually priced this request. A request that ends up
|
||
-- with all-zero prices is not "free", it is UNPRICED, and the two must not
|
||
-- look the same: an unpriced model silently costing 0 is the most dangerous
|
||
-- failure mode a cost plugin has, because the bill still adds up and just
|
||
-- quietly under-reports. It is counted separately and surfaced in the UI.
|
||
out = {
|
||
prompt = d.prompt or 0, completion = d.completion or 0,
|
||
per_request = 0, cache_discount = d.cache_discount, peak = d.peak,
|
||
}
|
||
|
||
-- model dimension (a token price overrides the default's token prices)
|
||
local mp = p.models and p.models[payload.model]
|
||
if mp then
|
||
out.priced = true
|
||
if mp.prompt ~= nil then out.prompt = mp.prompt end
|
||
if mp.completion ~= nil then out.completion = mp.completion end
|
||
if mp.per_request ~= nil then out.per_request = out.per_request + mp.per_request end
|
||
if mp.cache_discount ~= nil then out.cache_discount = mp.cache_discount end
|
||
if mp.peak ~= nil then out.peak = mp.peak end
|
||
end
|
||
|
||
-- source dimension: usually a flat fee, but may also carry token prices
|
||
local sp = p.sources and p.sources[payload.source]
|
||
if sp then
|
||
out.priced = true
|
||
if sp.prompt ~= nil then out.prompt = sp.prompt end
|
||
if sp.completion ~= nil then out.completion = sp.completion end
|
||
if sp.per_request ~= nil then out.per_request = out.per_request + sp.per_request end
|
||
if sp.cache_discount ~= nil then out.cache_discount = sp.cache_discount end
|
||
if sp.peak ~= nil then out.peak = sp.peak end
|
||
end
|
||
|
||
-- key dimension wins over the others (an operator pricing one customer
|
||
-- specially must be able to override both the model and the source price)
|
||
local kp = p.keys and p.keys[payload.key]
|
||
if kp then
|
||
out.priced = true
|
||
if kp.prompt ~= nil then out.prompt = kp.prompt end
|
||
if kp.completion ~= nil then out.completion = kp.completion end
|
||
if kp.per_request ~= nil then out.per_request = out.per_request + kp.per_request end
|
||
if kp.cache_discount ~= nil then out.cache_discount = kp.cache_discount end
|
||
if kp.peak ~= nil then out.peak = kp.peak end
|
||
end
|
||
return out
|
||
end
|
||
|
||
-- ===== 峰谷 / 时段定价 ================================================
|
||
--
|
||
-- 有些 provider 按 UTC 时段分价(commandcode 的 DeepSeek V4 系列就是:高峰
|
||
-- 01-04 & 06-10 UTC 工作日,价格恰好是非高峰的 2 倍)。静态价目无法表达这一点,
|
||
-- 而算错方向通常是【静默高估或低估】,不会报错——所以这里显式支持。
|
||
--
|
||
-- 配置形态(挂在任一维度的价目条目上):
|
||
--
|
||
-- "deepseek-v4.1-flash": {
|
||
-- prompt = 1.5e-7, completion = 6e-7,
|
||
-- peak = {
|
||
-- multiplier = 2, -- 高峰时单价乘以它
|
||
-- windows = [ -- UTC 星期几 = os.date 的 %w(周日=1)
|
||
-- { days = {2,3,4,5,6}, hours = {{1,2,3},{6,7,8,9}} },
|
||
-- ],
|
||
-- },
|
||
-- }
|
||
--
|
||
-- 语义:命中任一 window ⇒ 乘以 multiplier。hours 用 {起,止} 闭区间,跨零点
|
||
-- 用 {{22,24}} 表示 22:00-24:00(24 是"当天最后一刻")。
|
||
--
|
||
-- ★ 为什么用 os.date 的 ! 前缀取 UTC:provider 的费率表按 UTC 标注,而网关
|
||
-- 跑在本地时区(这台机是 Asia/Hong_Kong)。混用本地小时会让峰谷整体偏移 8
|
||
-- 小时,白天算成夜间——比不做峰谷还糟。
|
||
local function inPeakWindow(ev)
|
||
if ev == nil then return false end
|
||
local w = ev.windows
|
||
if type(w) ~= "table" or #w == 0 then return false end
|
||
local dow = tonumber(os.date("!%w")) or 0 -- 0=Sunday
|
||
local hour = tonumber(os.date("!%H")) or 0
|
||
for _, win in ipairs(w) do
|
||
local days = win.days
|
||
if type(days) == "table" then
|
||
local day_ok = false
|
||
for _, d in ipairs(days) do
|
||
if tonumber(d) == dow then day_ok = true break end
|
||
end
|
||
if not day_ok then goto continue_win end
|
||
end
|
||
local hours = win.hours
|
||
if type(hours) == "table" then
|
||
for _, h in ipairs(hours) do
|
||
local lo, hi = tonumber(h[1]), tonumber(h[2])
|
||
if lo and hi and hour >= lo and hour <= hi then return true end
|
||
end
|
||
end
|
||
::continue_win::
|
||
end
|
||
return false
|
||
end
|
||
|
||
-- applyPeak multiplies a price by the peak rule, if the request lands in a peak
|
||
-- window. It is a no-op when no rule is configured, so the common case costs one
|
||
-- nil check.
|
||
--
|
||
-- The multiplier is RECORDED, not applied to price.prompt in place. That looks
|
||
-- like a roundabout way to do it, but applying it there was a real bug: the
|
||
-- cache-read rate is DERIVED from price.prompt inside costFor, so doubling
|
||
-- price.prompt silently doubled the cache read too — compounding two separate
|
||
-- discounts. Keeping the multiplier separate lets costFor scale the fresh-prompt
|
||
-- and completion legs and leave the cache leg alone, which is what "peak rates
|
||
-- apply to the token price, cache reads are billed at their own rate" means.
|
||
local function applyPeak(price)
|
||
local pk = price.peak
|
||
if pk == nil then return price end
|
||
if not inPeakWindow(pk) then return price end
|
||
local m = tonumber(pk.multiplier) or 1
|
||
if m <= 0 then return price end
|
||
price.peak_multiplier = m
|
||
return price
|
||
end
|
||
|
||
-- costFor computes one request's price.
|
||
--
|
||
-- PROMPT CACHE: a cached prompt token is not billed like a fresh one. Almost
|
||
-- every provider sells cache reads at a steep discount (commonly 10% of the
|
||
-- fresh rate), and cache-heavy agent traffic hits long shared prefixes hard.
|
||
-- Charging the full prompt rate made a 1M-token request of which 900k were
|
||
-- cache reads come out at 10 USD instead of ~1.9 — an order of magnitude, on
|
||
-- exactly the traffic the cache exists to make cheap. The plugin therefore
|
||
-- splits the prompt count:
|
||
--
|
||
-- fresh = prompt_tokens - cache_hit_tokens -> full rate
|
||
-- cached = cache_hit_tokens -> rate * cache_discount
|
||
--
|
||
-- cache_discount defaults to 0.1 (the common 10x). It is configurable because
|
||
-- the ratio is a per-provider fact, not a constant of nature: set it to 1 to
|
||
-- keep the old flat behaviour, or 0 for providers that do not discount.
|
||
--
|
||
-- A request that reports cache_hit_tokens LARGER than prompt_tokens (a
|
||
-- misbehaving adapter, or two upstreams' numbers being mixed) is clamped: the
|
||
-- fresh count never goes negative, which would silently turn a request into
|
||
-- billable negative tokens.
|
||
local function costFor(payload, price)
|
||
price = applyPeak(price or priceFor(payload))
|
||
local prompt = tonumber(payload.prompt_tokens) or 0
|
||
local completion = tonumber(payload.completion_tokens) or 0
|
||
local cacheHit = tonumber(payload.cache_hit_tokens) or 0
|
||
if cacheHit < 0 then cacheHit = 0 end
|
||
if cacheHit > prompt then cacheHit = prompt end
|
||
|
||
local discount = tonumber(price.cache_discount)
|
||
if discount == nil then discount = plugin.cache_discount end
|
||
if discount == nil then discount = 0.1 end
|
||
if discount < 0 then discount = 0 elseif discount > 1 then discount = 1 end
|
||
|
||
-- The peak multiplier applies to the freshly-read prompt tokens and the
|
||
-- completion, but NOT to the cache read: a cache read is a separate upstream
|
||
-- rate that the off-peak figures already discount, and doubling it would
|
||
-- stack two discounts the provider never intended to stack.
|
||
local mult = tonumber(price.peak_multiplier) or 1
|
||
local fresh = prompt - cacheHit
|
||
local cost = fresh * price.prompt * mult
|
||
+ cacheHit * price.prompt * discount
|
||
+ completion * price.completion * mult
|
||
|
||
local flat = price.per_request
|
||
if not payload.ok and not plugin.count_failures then
|
||
flat = 0
|
||
end
|
||
return cost + flat
|
||
end
|
||
|
||
-- ---------- day bucket ----------
|
||
|
||
local function dayKey(epoch_seconds)
|
||
-- os.date is available in LuaJIT; fall back to a UTC-ish arithmetic stamp if
|
||
-- the host build has no os.date (keeps the plugin from erroring out on a
|
||
-- stripped runtime, which would otherwise look like a plugin failure).
|
||
if os and os.date then
|
||
return os.date("!%Y-%m-%d", epoch_seconds)
|
||
end
|
||
return tostring(math.floor(epoch_seconds / 86400))
|
||
end
|
||
|
||
-- ---------- hooks ----------
|
||
|
||
plugin.hooks = {
|
||
-- chain_step gives the per-tier walk; request_end gives the final accounting.
|
||
-- Subscribing to chain_step is OPTIONAL here: the totals are driven by
|
||
-- request_end alone, and the degradation counters below are pure observation.
|
||
-- A gateway with thousands of requests can drop this hook to save the
|
||
-- per-step Lua call without losing a single billed request.
|
||
chain_step = "on_chain_step",
|
||
request_end = "on_request_end",
|
||
}
|
||
|
||
-- Tracks how often a request had to drop below the top tier, and which tier
|
||
-- actually served it. Without this, "tier 1 was cooling" and "tier 1 served it"
|
||
-- are indistinguishable in the accounts, and a quietly degraded gateway looks
|
||
-- exactly like a healthy one.
|
||
plugin.state.degraded_reqs = 0
|
||
plugin.state.by_tier_served = {}
|
||
plugin.state.skip_reasons = {}
|
||
|
||
function plugin.on_chain_step(payload)
|
||
if payload == nil then return nil end
|
||
local s = plugin.state
|
||
if s == nil then return nil end
|
||
if s.by_tier_served == nil then s.by_tier_served = {} end
|
||
if s.skip_reasons == nil then s.skip_reasons = {} end
|
||
|
||
if payload.kind == "selected" then
|
||
local t = tostring(payload.tier or "?")
|
||
s.by_tier_served[t] = (s.by_tier_served[t] or 0) + 1
|
||
elseif payload.kind == "tier_skip" or payload.kind == "tier_busy" then
|
||
-- reason text is the ACTIONABLE part; normalise the volatile bits so the
|
||
-- same cause aggregates instead of creating a new row per request.
|
||
local r = tostring(payload.reason or payload.kind or "unknown")
|
||
r = string.gsub(r, "within [%d%.%a]+", "within <wait>")
|
||
s.skip_reasons[r] = (s.skip_reasons[r] or 0) + 1
|
||
end
|
||
return nil
|
||
end
|
||
|
||
function plugin.on_request_end(payload)
|
||
if payload == nil then return nil end
|
||
local prompt = tonumber(payload.prompt_tokens) or 0
|
||
local completion = tonumber(payload.completion_tokens) or 0
|
||
local ok = payload.ok and true or false
|
||
local price = priceFor(payload)
|
||
local cost = costFor(payload, price)
|
||
local s = plugin.state
|
||
-- Rebuild any missing container. This is reached in two real situations:
|
||
-- a fresh plugin, and an admin who PUT a partial state (e.g. only "prices"),
|
||
-- which legitimately replaces `state` with a sparse table. Checking only the
|
||
-- outer table would leave `s.total` nil and crash the hook on the next call.
|
||
if s == nil then s = {} plugin.state = s end
|
||
if s.total == nil then s.total = emptyBucket() end
|
||
if s.by_source == nil then s.by_source = {} end
|
||
if s.by_model == nil then s.by_model = {} end
|
||
if s.by_key == nil then s.by_key = {} end
|
||
if s.by_day == nil then s.by_day = {} end
|
||
if s.started == nil then s.started = payload.time or 0 end
|
||
if s.unpriced_reqs == nil then s.unpriced_reqs = 0 end
|
||
if s.unpriced_models == nil then s.unpriced_models = {} end
|
||
-- Track traffic that no price entry covered. This MUST come after the
|
||
-- container rebuild above: an earlier version referenced `s` before it was
|
||
-- declared, so on a fresh plugin the hook threw and the request recorded
|
||
-- NOTHING at all — the worst possible failure for a billing plugin, and one
|
||
-- that only showed up as "requests = 0" in a test.
|
||
if not price.priced then
|
||
s.unpriced_reqs = s.unpriced_reqs + 1
|
||
local m = payload.model or "?"
|
||
s.unpriced_models[m] = (s.unpriced_models[m] or 0) + 1
|
||
end
|
||
if s.degraded_reqs == nil then s.degraded_reqs = 0 end
|
||
-- Degradation is counted here rather than in the chain_step hook because
|
||
-- request_end sees the whole walk at once: one degraded request must count
|
||
-- once, whereas the walk may contain several skipped tiers.
|
||
if payload.degraded then s.degraded_reqs = s.degraded_reqs + 1 end
|
||
|
||
local cacheHit = tonumber(payload.cache_hit_tokens) or 0
|
||
if cacheHit < 0 then cacheHit = 0 end
|
||
if cacheHit > prompt then cacheHit = prompt end
|
||
-- cache_reported is the gateway's own signal that UPSTREAM gave a cache
|
||
-- number. Without it a source that never reports cache usage is
|
||
-- indistinguishable from one that always reports zero hits.
|
||
local cacheReported = payload.cache_reported and true or false
|
||
local C = cacheHit
|
||
local R = cacheReported
|
||
|
||
add(s.total, cost, prompt, completion, ok, C, R)
|
||
if payload.source ~= nil and payload.source ~= "" then
|
||
add(bucket(s.by_source, payload.source), cost, prompt, completion, ok, C, R)
|
||
end
|
||
if payload.model ~= nil and payload.model ~= "" then
|
||
add(bucket(s.by_model, payload.model), cost, prompt, completion, ok, C, R)
|
||
end
|
||
if payload.key ~= nil and payload.key ~= "" then
|
||
add(bucket(s.by_key, payload.key), cost, prompt, completion, ok, C, R)
|
||
end
|
||
|
||
-- Daily rollup, so the dashboard can draw a trend without the browser
|
||
-- re-deriving it. Keyed off the request's own timestamp, not os.time(), so a
|
||
-- replayed or imported record lands on the right day.
|
||
local ts = payload.time
|
||
if ts ~= nil and ts > 0 then
|
||
if ts > 1000000000000 then ts = ts / 1000 end -- kernel sends unix MILLIseconds
|
||
add(bucket(s.by_day, dayKey(ts)), cost, prompt, completion, ok, C, R)
|
||
end
|
||
return nil -- last stage: nobody downstream would read a return value
|
||
end
|
||
|
||
-- ---------- dashboard UI ----------
|
||
|
||
-- A whole page. The kernel injects this HTML and evaluates the <script> after
|
||
-- the DOM exists, and exposes `pluginAPI` for talking to the gateway.
|
||
plugin.ui = {
|
||
page = {
|
||
page_id = "billing",
|
||
title = "Billing",
|
||
icon = [==[<svg viewBox="0 0 24 24"><circle cx="12" cy="12" r="9"/><path d="M14.5 9.5a3 3 0 0 0-2.5-1.3c-1.4 0-2.4.7-2.4 1.8 0 2.6 5.2 1.4 5.2 4 0 1.1-1 1.8-2.5 1.8-1.1 0-2.1-.4-2.7-1.2"/><path d="M12 6.4v11.2"/></svg>]==],
|
||
order = 40,
|
||
mount = [==[
|
||
<div id="billing-root" style="padding:16px">
|
||
<div class="kpis" id="billing-kpis" style="display:grid;grid-template-columns:repeat(auto-fit,minmax(170px,1fr));gap:12px;margin-bottom:18px"></div>
|
||
<div style="display:grid;grid-template-columns:repeat(auto-fit,minmax(320px,1fr));gap:16px">
|
||
<div class="card" style="padding:14px">
|
||
<h3 style="margin:0 0 10px;font-size:14px">Per source</h3>
|
||
<div id="billing-by-source"></div>
|
||
</div>
|
||
<div class="card" style="padding:14px">
|
||
<h3 style="margin:0 0 10px;font-size:14px">Per model</h3>
|
||
<div id="billing-by-model"></div>
|
||
</div>
|
||
<div class="card" style="padding:14px">
|
||
<h3 style="margin:0 0 10px;font-size:14px">Per gateway key</h3>
|
||
<div id="billing-by-key"></div>
|
||
</div>
|
||
</div>
|
||
<div class="card" style="padding:14px;margin-top:16px">
|
||
<h3 style="margin:0 0 10px;font-size:14px">Daily</h3>
|
||
<div id="billing-by-day"></div>
|
||
</div>
|
||
</div>
|
||
<script>
|
||
(function () {
|
||
var ROOT = "billing";
|
||
function fmt(n) {
|
||
if (n === null || n === undefined) return "-";
|
||
n = Number(n);
|
||
if (!isFinite(n)) return "-";
|
||
if (n === 0) return "0";
|
||
if (Math.abs(n) < 0.000001) return n.toExponential(2);
|
||
return n.toFixed(Math.abs(n) < 1 ? 6 : 4);
|
||
}
|
||
function money(v, cur) { return (cur || "USD") + " " + fmt(v); }
|
||
function esc(s) {
|
||
return String(s == null ? "" : s).replace(/[&<>"]/g, function (c) {
|
||
return { "&": "&", "<": "<", ">": ">", '"': """ }[c];
|
||
});
|
||
}
|
||
// Cache hit rate, with the reporting caveat made visible.
|
||
//
|
||
// A bucket whose upstream never reports cache usage would render as "0%" from
|
||
// a 0/0 and read as "the cache is not working", when the truth is "this
|
||
// provider does not tell us". "n/r" keeps those apart.
|
||
function cacheRate(b) {
|
||
var prompt = b.prompt_tokens || 0;
|
||
var hit = b.cache_hit_tokens || 0;
|
||
if (!prompt) return "\u2014";
|
||
if (!b.cache_reported_reqs) return "n/r";
|
||
return ((hit / prompt) * 100).toFixed(1) + "%";
|
||
}
|
||
function row(name, b, cur) {
|
||
var fresh = (b.cache_fresh_tokens === undefined) ? (b.prompt_tokens || 0) : b.cache_fresh_tokens;
|
||
return "<tr><td><b>" + esc(name) + "</b></td><td>" + money(b.cost, cur) +
|
||
"</td><td>" + (b.requests || 0) + "</td><td>" + (b.prompt_tokens || 0) +
|
||
"</td><td>" + fresh +
|
||
"</td><td>" + (b.cache_hit_tokens || 0) +
|
||
"</td><td>" + esc(cacheRate(b)) +
|
||
"</td><td>" + (b.completion_tokens || 0) + "</td></tr>";
|
||
}
|
||
function tableFor(el, obj, cur, empty) {
|
||
var keys = Object.keys(obj || {});
|
||
if (!keys.length) { el.innerHTML = '<div class="muted">' + empty + "</div>"; return; }
|
||
keys.sort(function (a, b) { return (obj[b].cost || 0) - (obj[a].cost || 0); });
|
||
var h = "<table style='width:100%;border-collapse:collapse;font-size:13px'>" +
|
||
"<tr style='text-align:left;opacity:.65'><th>name</th><th>cost</th><th>reqs</th>" +
|
||
"<th>prompt</th><th>fresh</th><th>cache</th><th>cache%</th>" +
|
||
"<th>completion</th></tr>";
|
||
for (var i = 0; i < keys.length; i++) {
|
||
var k = keys[i];
|
||
h += "<tr style='border-top:1px solid rgba(120,90,150,.14)'>" + row(k, obj[k], cur) + "</tr>";
|
||
}
|
||
el.innerHTML = h + "</table>";
|
||
}
|
||
function render(st) {
|
||
if (!st) return;
|
||
var cur = (st.currency || "USD");
|
||
var t = st.total || {};
|
||
document.getElementById("billing-kpis").innerHTML = [
|
||
["Total", money(t.cost, cur)],
|
||
["Requests", t.requests || 0],
|
||
["Degraded", st.degraded_reqs || 0],
|
||
["Unpriced", st.unpriced_reqs || 0],
|
||
["Prompt tokens", t.prompt_tokens || 0],
|
||
["Completion tokens", t.completion_tokens || 0],
|
||
["Failures", t.failures || 0]
|
||
].map(function (kv) {
|
||
return "<div class='card' style='padding:12px'><div style='font-size:11px;opacity:.65'>" +
|
||
kv[0] + "</div><div style='font-size:19px;font-weight:600;margin-top:4px'>" +
|
||
esc(kv[1]) + "</div></div>";
|
||
}).join("");
|
||
tableFor(document.getElementById("billing-by-source"), st.by_source, cur, "no per-source data yet");
|
||
tableFor(document.getElementById("billing-by-model"), st.by_model, cur, "no per-model data yet");
|
||
tableFor(document.getElementById("billing-by-key"), st.by_key, cur, "no per-key data yet");
|
||
tableFor(document.getElementById("billing-by-day"), st.by_day, cur, "no daily data yet");
|
||
}
|
||
async function refresh() {
|
||
try {
|
||
var r = await fetch("/api/plugins/" + ROOT + "/state", { credentials: "same-origin" });
|
||
if (!r.ok) return;
|
||
var j = await r.json();
|
||
render(j.state);
|
||
} catch (e) {
|
||
// Swallowing this is what made the production bug invisible: render() threw
|
||
// a ReferenceError on an undefined `s`, the catch ate it, every table kept
|
||
// its empty placeholder, and the page looked fine in the network tab while
|
||
// showing nothing. Still must not THROW (the pane is decoration and must
|
||
// never break the host page) — but it must leave a trace.
|
||
if (window.console && console.error) console.error("[billing] render failed", e);
|
||
}
|
||
}
|
||
window.__billingRefresh = refresh;
|
||
refresh();
|
||
if (window.pluginAPI && pluginAPI.onTabShown) pluginAPI.onTabShown(refresh);
|
||
})();
|
||
</script>
|
||
]==],
|
||
},
|
||
-- Two elements on the EXISTING status page: a headline tile and a
|
||
-- per-source cost breakdown, so the number is visible without opening the
|
||
-- Billing tab.
|
||
elements = {
|
||
{
|
||
target = "status",
|
||
anchor = "top",
|
||
order = 5,
|
||
mount = [==[
|
||
<div class="card" id="billing-status-tile" style="padding:12px;margin-bottom:12px">
|
||
<div style="font-size:11px;opacity:.65">Total spend (billing plugin)</div>
|
||
<div id="billing-status-total" style="font-size:22px;font-weight:600;margin-top:4px">—</div>
|
||
<div id="billing-status-sub" style="font-size:12px;opacity:.65;margin-top:2px"></div>
|
||
</div>
|
||
<script>
|
||
(function () {
|
||
function fmt(n) {
|
||
n = Number(n || 0);
|
||
if (n === 0) return "0";
|
||
if (Math.abs(n) < 0.000001) return n.toExponential(2);
|
||
return n.toFixed(Math.abs(n) < 1 ? 6 : 4);
|
||
}
|
||
async function tick() {
|
||
try {
|
||
var r = await fetch("/api/plugins/billing/state", { credentials: "same-origin" });
|
||
if (!r.ok) return;
|
||
var j = await r.json();
|
||
var st = j.state;
|
||
if (!st || !st.total) return;
|
||
var cur = st.currency || "USD";
|
||
document.getElementById("billing-status-total").textContent = cur + " " + fmt(st.total.cost);
|
||
var parts = [];
|
||
var srcs = st.by_source || {};
|
||
var names = Object.keys(srcs).sort(function (a, b) {
|
||
return (srcs[b].cost || 0) - (srcs[a].cost || 0);
|
||
});
|
||
for (var i = 0; i < Math.min(3, names.length); i++) {
|
||
parts.push(names[i] + " " + fmt(srcs[names[i]].cost));
|
||
}
|
||
document.getElementById("billing-status-sub").textContent =
|
||
(st.total.requests || 0) + " requests" + (parts.length ? " · top: " + parts.join(" · ") : "");
|
||
} catch (e) {
|
||
// Same reasoning as the Billing page: decoration must never break the
|
||
// host page, but a silent catch turns a broken widget into "the plugin
|
||
// just doesn't show anything" with no way to tell why.
|
||
if (window.console && console.error) console.error("[billing] status tile refresh failed", e);
|
||
}
|
||
}
|
||
if (window.pluginAPI && pluginAPI.onTabShown) pluginAPI.onTabShown(tick);
|
||
tick();
|
||
})();
|
||
</script>
|
||
]==],
|
||
},
|
||
},
|
||
}
|
||
|
||
return plugin |