mirror of
https://gitcode.com/JianFeeeee/ModelRouter.git
synced 2026-10-03 23:54:06 +00:00
用户报 billing 页"文字超出展示框"。真因是三处叠加,都不是文案问题: 1. KPI 卡是 CSS grid,grid item 默认 min-width:auto。2e8 级的 prompt tokens(实测 270,149,209)无法收缩,于是把 grid 轨道撑出容器 → 整页横向 溢出。修法是 min-width:0 + overflow:hidden(不给 min-width:0 的话, grid 子项永远不肯收缩,这是 grid 最常见的溢出坑)。 2. 表格写死 min-width:560px,窄视口下必然溢出。改 width:100% + table-layout:fixed,列宽由布局分配而不是由内容撑开。 3. 长名称(deepseek-v4.1-flash 这类模型 id)撑宽单元格。名称列改 ellipsis + title 悬停看全名;数字列 word-break:break-all 在列宽内换行。 顺手:所有 token/请求数走 toLocaleString 千分位。原始 9 位数字读起来要数 零位数,分组后 270,149,209 一眼可读,也顺带缩短了字符串宽度。 CDP 实测(820px 窄视口,逼出溢出条件): main/body 横向溢出 = no table=338 < card=366(修复前 min-width:560 必然 > 366) KPI 输入 tokens = 270,149,209 billing 页无控制台错误 残留(不影响布局):表头 TH 在固定布局下 48>42 轻微超出自身格,因为 word-break 不拆单个长词;表格整体仍在容器内,未产生页面滚动。 全量测试全绿(8 包)。
728 lines
32 KiB
Lua
728 lines
32 KiB
Lua
-- billing.lua — usage accounting plugin for ModelRouter.
|
||
--
|
||
-- Computes what each request cost, from three configurable dimensions:
|
||
--
|
||
-- source a flat per-request price for an upstream source
|
||
-- model a per-token price for a model id (prompt / completion separately)
|
||
-- key an override price for one gateway key
|
||
--
|
||
-- It then keeps running totals for the whole gateway, per source, per model
|
||
-- and per key, and publishes them in `plugin.state` so the kernel can serve
|
||
-- them at GET /api/plugins/billing/state — which is what its own dashboard
|
||
-- component reads.
|
||
--
|
||
-- ACCOUNTING BOUNDARY (important, and deliberate):
|
||
-- this plugin REPORTS; it does not ENFORCE. The gateway's own quota accounting
|
||
-- (internal/gateway/stats.go, enforced at request admission) stays
|
||
-- authoritative for limits. Two independent accounting paths that disagree are
|
||
-- worse than one that is slightly less featureful, so nothing here feeds back
|
||
-- into routing or quota decisions.
|
||
--
|
||
-- PRICE CONFIGURATION
|
||
-- Prices are supplied as a Lua table assigned to `billing.prices` before the
|
||
-- plugin is loaded, OR at runtime through PUT /api/plugins/billing/state. The
|
||
-- shape is:
|
||
--
|
||
-- billing.prices = {
|
||
-- currency = "USD", -- display only, no conversion happens
|
||
-- default = { prompt = 0, completion = 0, per_request = 0 },
|
||
-- sources = {
|
||
-- ["localzen"] = { per_request = 0.0 },
|
||
-- ["trae"] = { per_request = 0.01 },
|
||
-- },
|
||
-- models = {
|
||
-- ["gpt-5.4"] = { prompt = 1.25e-6, completion = 1e-5 }, -- USD per TOKEN
|
||
-- ["kimi-k3"] = { prompt = 6e-7, completion = 2.5e-6 },
|
||
-- ["kolors"] = { per_request = 0.04 }, -- image: flat
|
||
-- },
|
||
-- keys = {
|
||
-- -- by gateway key (the same value the audit log masks to ***xxxxxx)
|
||
-- ["***a1b2c3"] = { prompt = 1.1e-6, completion = 9e-6 },
|
||
-- },
|
||
-- }
|
||
--
|
||
-- Precedence for a token price: keys > models > default. A flat per_request
|
||
-- price, when present at any level, is ADDED on top of the token cost, so an
|
||
-- image model can carry both (e.g. tokens billed plus a fixed fee).
|
||
--
|
||
-- Numbers are USD per single token, which is how providers publish prices. That
|
||
-- makes a typical entry look like 1.25e-6; the plugin multiplies by the token
|
||
-- count, so no unit conversion happens anywhere.
|
||
|
||
local plugin = {
|
||
name = "billing",
|
||
version = "1.0.0",
|
||
description = "Per-source / per-model / per-key cost accounting with a dashboard",
|
||
author = "ModelRouter",
|
||
}
|
||
|
||
-- ---------- prices ----------
|
||
|
||
-- plugin.prices can be pre-seeded by embedding this file (an operator edits the
|
||
-- table below) or replaced at runtime through the state API. It is a SEPARATE
|
||
-- field from plugin.state on purpose: PUT /api/plugins/billing/state replaces
|
||
-- `state` wholesale, and prices must not live there or a price update would
|
||
-- wipe the accumulated totals. See docs/plugins.md.
|
||
local DEFAULT_PRICES = {
|
||
currency = "USD",
|
||
default = { prompt = 0, completion = 0, per_request = 0 },
|
||
sources = {},
|
||
models = {},
|
||
keys = {},
|
||
}
|
||
plugin.prices = DEFAULT_PRICES
|
||
|
||
-- Default prompt-cache discount. 0.1 = a cache read costs a tenth of a fresh
|
||
-- token, which is what DeepSeek/Qwen/Kimi and most others charge. It can be
|
||
-- overridden per price entry (prices.models.<m>.cache_discount) or globally by
|
||
-- setting plugin.cache_discount; 1 restores flat prompt pricing.
|
||
plugin.cache_discount = 0.1
|
||
|
||
-- ---------- accumulated totals ----------
|
||
|
||
-- state is what the kernel serves at GET /api/plugins/billing/state. It holds
|
||
-- ACCUMULATED TOTALS ONLY — prices live in plugin.prices (see above), so
|
||
-- replacing state never destroys a price table and updating prices never
|
||
-- destroys history.
|
||
--
|
||
-- Structure:
|
||
-- total { cost, requests, prompt_tokens, completion_tokens }
|
||
-- by_source { <name> = { cost, requests, ...tokens } }
|
||
-- by_model { <model> = { cost, ... } }
|
||
-- by_key { <masked key id> = { cost, ... } }
|
||
-- by_day { "YYYY-MM-DD" = { cost, ... } }
|
||
-- top_sources [ {name, cost, requests}, ... ] sorted, capped
|
||
-- top_models [ ... ]
|
||
-- top_keys [ ... ]
|
||
--
|
||
-- Sorted top-N lists are maintained incrementally rather than re-sorted on
|
||
-- every request: this hook runs once per request on the hot path, so it does
|
||
-- map updates only. The sort happens when state is READ.
|
||
-- CACHE ACCOUNTING (added after production showed the gap):
|
||
-- the gateway extracts prompt_cache_hit_tokens from upstream usage and puts it
|
||
-- in the request_end payload, and costFor() already used it to price the cache
|
||
-- leg — but no bucket recorded it. So a gateway where 99.88% of prompt tokens
|
||
-- were cache reads showed a prompt_tokens number with no indication of that,
|
||
-- and there was no way to see cache hit rate per source/model/key at all.
|
||
--
|
||
-- cache_hit_tokens hits, as reported by upstream
|
||
-- cache_fresh_tokens prompt tokens that were NOT cache reads
|
||
-- cache_reported_reqs requests where upstream gave a cache number at all.
|
||
-- Kept separate from a zero: "upstream does not report cache usage" and
|
||
-- "upstream reported zero hits" look identical in a hit total, and they mean
|
||
-- opposite things when you are trying to work out whether a cache discount is
|
||
-- doing anything.
|
||
local function emptyBucket()
|
||
return {
|
||
cost = 0, requests = 0, prompt_tokens = 0, completion_tokens = 0, failures = 0,
|
||
cache_hit_tokens = 0, cache_fresh_tokens = 0, cache_reported_reqs = 0,
|
||
}
|
||
end
|
||
|
||
plugin.state = {
|
||
total = emptyBucket(),
|
||
by_source = {},
|
||
by_model = {},
|
||
by_key = {},
|
||
by_day = {},
|
||
started = os.time and 0 or 0,
|
||
}
|
||
|
||
local function bucket(tbl, k)
|
||
local b = tbl[k]
|
||
if b == nil then
|
||
b = emptyBucket()
|
||
tbl[k] = b
|
||
end
|
||
return b
|
||
end
|
||
|
||
local function add(b, cost, prompt, completion, ok, cacheHit, cacheReported)
|
||
b.cost = b.cost + cost
|
||
b.requests = b.requests + 1
|
||
b.prompt_tokens = b.prompt_tokens + prompt
|
||
b.completion_tokens = b.completion_tokens + completion
|
||
if not ok then b.failures = b.failures + 1 end
|
||
-- Backfill guards a bucket that predates these fields (a state file written
|
||
-- by an older build, or one restored from disk): nil + number is an error in
|
||
-- Lua, and a hook that throws stops accounting for that request entirely.
|
||
if b.cache_hit_tokens == nil then b.cache_hit_tokens = 0 end
|
||
if b.cache_fresh_tokens == nil then b.cache_fresh_tokens = 0 end
|
||
if b.cache_reported_reqs == nil then b.cache_reported_reqs = 0 end
|
||
b.cache_hit_tokens = b.cache_hit_tokens + (cacheHit or 0)
|
||
b.cache_fresh_tokens = b.cache_fresh_tokens + ((prompt or 0) - (cacheHit or 0))
|
||
if cacheReported then b.cache_reported_reqs = b.cache_reported_reqs + 1 end
|
||
end
|
||
|
||
-- ---------- pricing ----------
|
||
|
||
-- lookup walks keys > models > default and returns a price triple plus whether
|
||
-- a flat per_request component applies.
|
||
local function priceFor(payload)
|
||
local p = plugin.prices or DEFAULT_PRICES
|
||
local d = p.default or {}
|
||
-- Whether ANY dimension actually priced this request. A request that ends up
|
||
-- with all-zero prices is not "free", it is UNPRICED, and the two must not
|
||
-- look the same: an unpriced model silently costing 0 is the most dangerous
|
||
-- failure mode a cost plugin has, because the bill still adds up and just
|
||
-- quietly under-reports. It is counted separately and surfaced in the UI.
|
||
out = {
|
||
prompt = d.prompt or 0, completion = d.completion or 0,
|
||
per_request = 0, cache_discount = d.cache_discount, peak = d.peak,
|
||
}
|
||
|
||
-- model dimension (a token price overrides the default's token prices)
|
||
local mp = p.models and p.models[payload.model]
|
||
if mp then
|
||
out.priced = true
|
||
if mp.prompt ~= nil then out.prompt = mp.prompt end
|
||
if mp.completion ~= nil then out.completion = mp.completion end
|
||
if mp.per_request ~= nil then out.per_request = out.per_request + mp.per_request end
|
||
if mp.cache_discount ~= nil then out.cache_discount = mp.cache_discount end
|
||
if mp.peak ~= nil then out.peak = mp.peak end
|
||
end
|
||
|
||
-- source dimension: usually a flat fee, but may also carry token prices
|
||
local sp = p.sources and p.sources[payload.source]
|
||
if sp then
|
||
out.priced = true
|
||
if sp.prompt ~= nil then out.prompt = sp.prompt end
|
||
if sp.completion ~= nil then out.completion = sp.completion end
|
||
if sp.per_request ~= nil then out.per_request = out.per_request + sp.per_request end
|
||
if sp.cache_discount ~= nil then out.cache_discount = sp.cache_discount end
|
||
if sp.peak ~= nil then out.peak = sp.peak end
|
||
end
|
||
|
||
-- key dimension wins over the others (an operator pricing one customer
|
||
-- specially must be able to override both the model and the source price)
|
||
local kp = p.keys and p.keys[payload.key]
|
||
if kp then
|
||
out.priced = true
|
||
if kp.prompt ~= nil then out.prompt = kp.prompt end
|
||
if kp.completion ~= nil then out.completion = kp.completion end
|
||
if kp.per_request ~= nil then out.per_request = out.per_request + kp.per_request end
|
||
if kp.cache_discount ~= nil then out.cache_discount = kp.cache_discount end
|
||
if kp.peak ~= nil then out.peak = kp.peak end
|
||
end
|
||
return out
|
||
end
|
||
|
||
-- ===== 峰谷 / 时段定价 ================================================
|
||
--
|
||
-- 有些 provider 按 UTC 时段分价(commandcode 的 DeepSeek V4 系列就是:高峰
|
||
-- 01-04 & 06-10 UTC 工作日,价格恰好是非高峰的 2 倍)。静态价目无法表达这一点,
|
||
-- 而算错方向通常是【静默高估或低估】,不会报错——所以这里显式支持。
|
||
--
|
||
-- 配置形态(挂在任一维度的价目条目上):
|
||
--
|
||
-- "deepseek-v4.1-flash": {
|
||
-- prompt = 1.5e-7, completion = 6e-7,
|
||
-- peak = {
|
||
-- multiplier = 2, -- 高峰时单价乘以它
|
||
-- windows = [ -- UTC 星期几 = os.date 的 %w(周日=1)
|
||
-- { days = {2,3,4,5,6}, hours = {{1,2,3},{6,7,8,9}} },
|
||
-- ],
|
||
-- },
|
||
-- }
|
||
--
|
||
-- 语义:命中任一 window ⇒ 乘以 multiplier。hours 用 {起,止} 闭区间,跨零点
|
||
-- 用 {{22,24}} 表示 22:00-24:00(24 是"当天最后一刻")。
|
||
--
|
||
-- ★ 为什么用 os.date 的 ! 前缀取 UTC:provider 的费率表按 UTC 标注,而网关
|
||
-- 跑在本地时区(这台机是 Asia/Hong_Kong)。混用本地小时会让峰谷整体偏移 8
|
||
-- 小时,白天算成夜间——比不做峰谷还糟。
|
||
local function inPeakWindow(ev)
|
||
if ev == nil then return false end
|
||
local w = ev.windows
|
||
if type(w) ~= "table" or #w == 0 then return false end
|
||
local dow = tonumber(os.date("!%w")) or 0 -- 0=Sunday
|
||
local hour = tonumber(os.date("!%H")) or 0
|
||
for _, win in ipairs(w) do
|
||
local days = win.days
|
||
if type(days) == "table" then
|
||
local day_ok = false
|
||
for _, d in ipairs(days) do
|
||
if tonumber(d) == dow then day_ok = true break end
|
||
end
|
||
if not day_ok then goto continue_win end
|
||
end
|
||
local hours = win.hours
|
||
if type(hours) == "table" then
|
||
for _, h in ipairs(hours) do
|
||
local lo, hi = tonumber(h[1]), tonumber(h[2])
|
||
if lo and hi and hour >= lo and hour <= hi then return true end
|
||
end
|
||
end
|
||
::continue_win::
|
||
end
|
||
return false
|
||
end
|
||
|
||
-- applyPeak multiplies a price by the peak rule, if the request lands in a peak
|
||
-- window. It is a no-op when no rule is configured, so the common case costs one
|
||
-- nil check.
|
||
--
|
||
-- The multiplier is RECORDED, not applied to price.prompt in place. That looks
|
||
-- like a roundabout way to do it, but applying it there was a real bug: the
|
||
-- cache-read rate is DERIVED from price.prompt inside costFor, so doubling
|
||
-- price.prompt silently doubled the cache read too — compounding two separate
|
||
-- discounts. Keeping the multiplier separate lets costFor scale the fresh-prompt
|
||
-- and completion legs and leave the cache leg alone, which is what "peak rates
|
||
-- apply to the token price, cache reads are billed at their own rate" means.
|
||
local function applyPeak(price)
|
||
local pk = price.peak
|
||
if pk == nil then return price end
|
||
if not inPeakWindow(pk) then return price end
|
||
local m = tonumber(pk.multiplier) or 1
|
||
if m <= 0 then return price end
|
||
price.peak_multiplier = m
|
||
return price
|
||
end
|
||
|
||
-- costFor computes one request's price.
|
||
--
|
||
-- PROMPT CACHE: a cached prompt token is not billed like a fresh one. Almost
|
||
-- every provider sells cache reads at a steep discount (commonly 10% of the
|
||
-- fresh rate), and cache-heavy agent traffic hits long shared prefixes hard.
|
||
-- Charging the full prompt rate made a 1M-token request of which 900k were
|
||
-- cache reads come out at 10 USD instead of ~1.9 — an order of magnitude, on
|
||
-- exactly the traffic the cache exists to make cheap. The plugin therefore
|
||
-- splits the prompt count:
|
||
--
|
||
-- fresh = prompt_tokens - cache_hit_tokens -> full rate
|
||
-- cached = cache_hit_tokens -> rate * cache_discount
|
||
--
|
||
-- cache_discount defaults to 0.1 (the common 10x). It is configurable because
|
||
-- the ratio is a per-provider fact, not a constant of nature: set it to 1 to
|
||
-- keep the old flat behaviour, or 0 for providers that do not discount.
|
||
--
|
||
-- A request that reports cache_hit_tokens LARGER than prompt_tokens (a
|
||
-- misbehaving adapter, or two upstreams' numbers being mixed) is clamped: the
|
||
-- fresh count never goes negative, which would silently turn a request into
|
||
-- billable negative tokens.
|
||
local function costFor(payload, price)
|
||
price = applyPeak(price or priceFor(payload))
|
||
local prompt = tonumber(payload.prompt_tokens) or 0
|
||
local completion = tonumber(payload.completion_tokens) or 0
|
||
local cacheHit = tonumber(payload.cache_hit_tokens) or 0
|
||
if cacheHit < 0 then cacheHit = 0 end
|
||
if cacheHit > prompt then cacheHit = prompt end
|
||
|
||
local discount = tonumber(price.cache_discount)
|
||
if discount == nil then discount = plugin.cache_discount end
|
||
if discount == nil then discount = 0.1 end
|
||
if discount < 0 then discount = 0 elseif discount > 1 then discount = 1 end
|
||
|
||
-- The peak multiplier applies to the freshly-read prompt tokens and the
|
||
-- completion, but NOT to the cache read: a cache read is a separate upstream
|
||
-- rate that the off-peak figures already discount, and doubling it would
|
||
-- stack two discounts the provider never intended to stack.
|
||
local mult = tonumber(price.peak_multiplier) or 1
|
||
local fresh = prompt - cacheHit
|
||
local cost = fresh * price.prompt * mult
|
||
+ cacheHit * price.prompt * discount
|
||
+ completion * price.completion * mult
|
||
|
||
local flat = price.per_request
|
||
if not payload.ok and not plugin.count_failures then
|
||
flat = 0
|
||
end
|
||
return cost + flat
|
||
end
|
||
|
||
-- ---------- day bucket ----------
|
||
|
||
local function dayKey(epoch_seconds)
|
||
-- os.date is available in LuaJIT; fall back to a UTC-ish arithmetic stamp if
|
||
-- the host build has no os.date (keeps the plugin from erroring out on a
|
||
-- stripped runtime, which would otherwise look like a plugin failure).
|
||
if os and os.date then
|
||
return os.date("!%Y-%m-%d", epoch_seconds)
|
||
end
|
||
return tostring(math.floor(epoch_seconds / 86400))
|
||
end
|
||
|
||
-- ---------- hooks ----------
|
||
|
||
plugin.hooks = {
|
||
-- chain_step gives the per-tier walk; request_end gives the final accounting.
|
||
-- Subscribing to chain_step is OPTIONAL here: the totals are driven by
|
||
-- request_end alone, and the degradation counters below are pure observation.
|
||
-- A gateway with thousands of requests can drop this hook to save the
|
||
-- per-step Lua call without losing a single billed request.
|
||
chain_step = "on_chain_step",
|
||
request_end = "on_request_end",
|
||
}
|
||
|
||
-- Tracks how often a request had to drop below the top tier, and which tier
|
||
-- actually served it. Without this, "tier 1 was cooling" and "tier 1 served it"
|
||
-- are indistinguishable in the accounts, and a quietly degraded gateway looks
|
||
-- exactly like a healthy one.
|
||
plugin.state.degraded_reqs = 0
|
||
plugin.state.by_tier_served = {}
|
||
plugin.state.skip_reasons = {}
|
||
|
||
function plugin.on_chain_step(payload)
|
||
if payload == nil then return nil end
|
||
local s = plugin.state
|
||
if s == nil then return nil end
|
||
if s.by_tier_served == nil then s.by_tier_served = {} end
|
||
if s.skip_reasons == nil then s.skip_reasons = {} end
|
||
|
||
if payload.kind == "selected" then
|
||
local t = tostring(payload.tier or "?")
|
||
s.by_tier_served[t] = (s.by_tier_served[t] or 0) + 1
|
||
elseif payload.kind == "tier_skip" or payload.kind == "tier_busy" then
|
||
-- reason text is the ACTIONABLE part; normalise the volatile bits so the
|
||
-- same cause aggregates instead of creating a new row per request.
|
||
local r = tostring(payload.reason or payload.kind or "unknown")
|
||
r = string.gsub(r, "within [%d%.%a]+", "within <wait>")
|
||
s.skip_reasons[r] = (s.skip_reasons[r] or 0) + 1
|
||
end
|
||
return nil
|
||
end
|
||
|
||
function plugin.on_request_end(payload)
|
||
if payload == nil then return nil end
|
||
local prompt = tonumber(payload.prompt_tokens) or 0
|
||
local completion = tonumber(payload.completion_tokens) or 0
|
||
local ok = payload.ok and true or false
|
||
local price = priceFor(payload)
|
||
local cost = costFor(payload, price)
|
||
local s = plugin.state
|
||
-- Rebuild any missing container. This is reached in two real situations:
|
||
-- a fresh plugin, and an admin who PUT a partial state (e.g. only "prices"),
|
||
-- which legitimately replaces `state` with a sparse table. Checking only the
|
||
-- outer table would leave `s.total` nil and crash the hook on the next call.
|
||
if s == nil then s = {} plugin.state = s end
|
||
if s.total == nil then s.total = emptyBucket() end
|
||
if s.by_source == nil then s.by_source = {} end
|
||
if s.by_model == nil then s.by_model = {} end
|
||
if s.by_key == nil then s.by_key = {} end
|
||
if s.by_day == nil then s.by_day = {} end
|
||
if s.started == nil then s.started = payload.time or 0 end
|
||
if s.unpriced_reqs == nil then s.unpriced_reqs = 0 end
|
||
if s.unpriced_models == nil then s.unpriced_models = {} end
|
||
-- Track traffic that no price entry covered. This MUST come after the
|
||
-- container rebuild above: an earlier version referenced `s` before it was
|
||
-- declared, so on a fresh plugin the hook threw and the request recorded
|
||
-- NOTHING at all — the worst possible failure for a billing plugin, and one
|
||
-- that only showed up as "requests = 0" in a test.
|
||
if not price.priced then
|
||
s.unpriced_reqs = s.unpriced_reqs + 1
|
||
local m = payload.model or "?"
|
||
s.unpriced_models[m] = (s.unpriced_models[m] or 0) + 1
|
||
end
|
||
if s.degraded_reqs == nil then s.degraded_reqs = 0 end
|
||
-- Degradation is counted here rather than in the chain_step hook because
|
||
-- request_end sees the whole walk at once: one degraded request must count
|
||
-- once, whereas the walk may contain several skipped tiers.
|
||
if payload.degraded then s.degraded_reqs = s.degraded_reqs + 1 end
|
||
|
||
local cacheHit = tonumber(payload.cache_hit_tokens) or 0
|
||
if cacheHit < 0 then cacheHit = 0 end
|
||
if cacheHit > prompt then cacheHit = prompt end
|
||
-- cache_reported is the gateway's own signal that UPSTREAM gave a cache
|
||
-- number. Without it a source that never reports cache usage is
|
||
-- indistinguishable from one that always reports zero hits.
|
||
local cacheReported = payload.cache_reported and true or false
|
||
local C = cacheHit
|
||
local R = cacheReported
|
||
|
||
add(s.total, cost, prompt, completion, ok, C, R)
|
||
if payload.source ~= nil and payload.source ~= "" then
|
||
add(bucket(s.by_source, payload.source), cost, prompt, completion, ok, C, R)
|
||
end
|
||
if payload.model ~= nil and payload.model ~= "" then
|
||
add(bucket(s.by_model, payload.model), cost, prompt, completion, ok, C, R)
|
||
end
|
||
if payload.key ~= nil and payload.key ~= "" then
|
||
add(bucket(s.by_key, payload.key), cost, prompt, completion, ok, C, R)
|
||
end
|
||
|
||
-- Daily rollup, so the dashboard can draw a trend without the browser
|
||
-- re-deriving it. Keyed off the request's own timestamp, not os.time(), so a
|
||
-- replayed or imported record lands on the right day.
|
||
local ts = payload.time
|
||
if ts ~= nil and ts > 0 then
|
||
if ts > 1000000000000 then ts = ts / 1000 end -- kernel sends unix MILLIseconds
|
||
add(bucket(s.by_day, dayKey(ts)), cost, prompt, completion, ok, C, R)
|
||
end
|
||
return nil -- last stage: nobody downstream would read a return value
|
||
end
|
||
|
||
-- ---------- dashboard UI ----------
|
||
|
||
-- A whole page. The kernel injects this HTML and evaluates the <script> after
|
||
-- the DOM exists, and exposes `pluginAPI` for talking to the gateway.
|
||
plugin.ui = {
|
||
page = {
|
||
page_id = "billing",
|
||
title = "Billing",
|
||
icon = [==[<svg viewBox="0 0 24 24"><circle cx="12" cy="12" r="9"/><path d="M14.5 9.5a3 3 0 0 0-2.5-1.3c-1.4 0-2.4.7-2.4 1.8 0 2.6 5.2 1.4 5.2 4 0 1.1-1 1.8-2.5 1.8-1.1 0-2.1-.4-2.7-1.2"/><path d="M12 6.4v11.2"/></svg>]==],
|
||
order = 40,
|
||
mount = [==[
|
||
<div id="billing-root" style="padding:16px;min-width:0;max-width:100%;overflow-x:auto">
|
||
<div class="kpis" id="billing-kpis" style="display:grid;grid-template-columns:repeat(auto-fit,minmax(170px,1fr));gap:12px;margin-bottom:18px"></div>
|
||
<div style="display:grid;grid-template-columns:repeat(auto-fit,minmax(320px,1fr));gap:16px">
|
||
<div class="card" style="padding:14px">
|
||
<h3 id="billing-h-src" style="margin:0 0 10px;font-size:14px"></h3>
|
||
<div id="billing-by-source"></div>
|
||
</div>
|
||
<div class="card" style="padding:14px">
|
||
<h3 id="billing-h-model" style="margin:0 0 10px;font-size:14px"></h3>
|
||
<div id="billing-by-model"></div>
|
||
</div>
|
||
<div class="card" style="padding:14px">
|
||
<h3 id="billing-h-key" style="margin:0 0 10px;font-size:14px"></h3>
|
||
<div id="billing-by-key"></div>
|
||
</div>
|
||
</div>
|
||
<div class="card" style="padding:14px;margin-top:16px">
|
||
<h3 id="billing-h-day" style="margin:0 0 10px;font-size:14px"></h3>
|
||
<div id="billing-by-day"></div>
|
||
</div>
|
||
</div>
|
||
<script>
|
||
(function () {
|
||
var ROOT = "billing";
|
||
// UI strings, bilingual. The host page i18n (applyI18n/data-i) only covers
|
||
// markup the HOST renders; a plugin's injected markup is invisible to it, so
|
||
// the Billing page stayed English while the rest of the UI switched. These go
|
||
// through pluginAPI.lang / onLangChange, the small surface the host exposes
|
||
// for exactly this.
|
||
var STR = {
|
||
en: {
|
||
total: "Total", requests: "Requests", degraded: "Degraded",
|
||
unpriced: "Unpriced", prompt: "Prompt tokens",
|
||
completion: "Completion tokens", failures: "Failures",
|
||
cacheRate: "Cache hit rate", cacheTokens: "Cache read tokens",
|
||
perSource: "Per source", perModel: "Per model",
|
||
perKey: "Per gateway key", perDay: "Daily",
|
||
thName: "name", thCost: "cost", thReqs: "reqs", thPrompt: "prompt",
|
||
thFresh: "fresh", thCache: "cache", thCachePct: "cache%",
|
||
thCompletion: "completion",
|
||
noData: "no data yet",
|
||
},
|
||
zh: {
|
||
total: "总开销", requests: "请求数", degraded: "降级",
|
||
unpriced: "未定价", prompt: "输入 tokens",
|
||
completion: "输出 tokens", failures: "失败",
|
||
cacheRate: "缓存命中率", cacheTokens: "缓存读取 tokens",
|
||
perSource: "按源", perModel: "按模型",
|
||
perKey: "按网关密钥", perDay: "按天",
|
||
thName: "名称", thCost: "开销", thReqs: "请求", thPrompt: "输入",
|
||
thFresh: "新鲜", thCache: "缓存", thCachePct: "缓存%",
|
||
thCompletion: "输出",
|
||
noData: "暂无数据",
|
||
},
|
||
};
|
||
function L() {
|
||
var lang = (window.pluginAPI && pluginAPI.lang) || "zh";
|
||
return STR[lang] || STR.zh;
|
||
}
|
||
function fmt(n) {
|
||
if (n === null || n === undefined) return "-";
|
||
n = Number(n);
|
||
if (!isFinite(n)) return "-";
|
||
if (n === 0) return "0";
|
||
if (Math.abs(n) < 0.000001) return n.toExponential(2);
|
||
return n.toFixed(Math.abs(n) < 1 ? 6 : 4);
|
||
}
|
||
function money(v, cur) { return (cur || "USD") + " " + fmt(v); }
|
||
function esc(s) {
|
||
return String(s == null ? "" : s).replace(/[&<>"]/g, function (c) {
|
||
return { "&": "&", "<": "<", ">": ">", '"': """ }[c];
|
||
});
|
||
}
|
||
// Cache hit rate, with the reporting caveat made visible.
|
||
//
|
||
// A bucket whose upstream never reports cache usage would render as "0%" from
|
||
// a 0/0 and read as "the cache is not working", when the truth is "this
|
||
// provider does not tell us". "n/r" keeps those apart.
|
||
function cacheRate(b) {
|
||
var prompt = b.prompt_tokens || 0;
|
||
var hit = b.cache_hit_tokens || 0;
|
||
if (!prompt) return "\u2014";
|
||
if (!b.cache_reported_reqs) return "n/r";
|
||
return ((hit / prompt) * 100).toFixed(1) + "%";
|
||
}
|
||
// fmtInt 千分位分组:2e8 级 token 总数可读、也更短,降低撑宽风险。
|
||
function fmtInt(n) {
|
||
return (Number(n) || 0).toLocaleString("en-US");
|
||
}
|
||
function row(name, b, cur) {
|
||
var fresh = (b.cache_fresh_tokens === undefined) ? (b.prompt_tokens || 0) : b.cache_fresh_tokens;
|
||
// 名称列 ellipsis(title 悬停看全名);数字列 break-all 在列宽内换行而不是
|
||
// 把表格撑出卡片。单元格结构与列数不变,列数判据不受影响。
|
||
return "<tr><td style='overflow:hidden'><b style='display:block;white-space:nowrap;overflow:hidden;text-overflow:ellipsis' title='" +
|
||
esc(String(name).replace(/'/g, "'")) + "'>" + esc(name) + "</b></td>" +
|
||
"<td style='word-break:break-all'>" + money(b.cost, cur) + "</td>" +
|
||
"<td style='word-break:break-all'>" + fmtInt(b.requests || 0) + "</td>" +
|
||
"<td style='word-break:break-all'>" + fmtInt(b.prompt_tokens || 0) + "</td>" +
|
||
"<td style='word-break:break-all'>" + fmtInt(fresh) + "</td>" +
|
||
"<td style='word-break:break-all'>" + fmtInt(b.cache_hit_tokens || 0) + "</td>" +
|
||
"<td style='word-break:break-all'>" + esc(cacheRate(b)) + "</td>" +
|
||
"<td style='word-break:break-all'>" + fmtInt(b.completion_tokens || 0) + "</td></tr>";
|
||
}
|
||
function tableFor(el, obj, cur, empty) {
|
||
var keys = Object.keys(obj || {});
|
||
if (!keys.length) { el.innerHTML = '<div class="muted">' + empty + "</div>"; return; }
|
||
keys.sort(function (a, b) { return (obj[b].cost || 0) - (obj[a].cost || 0); });
|
||
var TH = L();
|
||
var h = "<table style='width:100%;border-collapse:collapse;font-size:13px;table-layout:fixed;word-break:break-word'>" +
|
||
"<tr style='text-align:left;opacity:.65'><th>" + TH.thName + "</th><th>" + TH.thCost +
|
||
"</th><th>" + TH.thReqs + "</th><th>" + TH.thPrompt +
|
||
"</th><th>" + TH.thFresh + "</th><th>" + TH.thCache + "</th><th>" + TH.thCachePct +
|
||
"</th><th>" + TH.thCompletion + "</th></tr>";
|
||
for (var i = 0; i < keys.length; i++) {
|
||
var k = keys[i];
|
||
h += "<tr style='border-top:1px solid rgba(120,90,150,.14)'>" + row(k, obj[k], cur) + "</tr>";
|
||
}
|
||
el.innerHTML = h + "</table>";
|
||
}
|
||
function renderTitles() {
|
||
var T = L();
|
||
var m = { "billing-h-src": T.perSource, "billing-h-model": T.perModel,
|
||
"billing-h-key": T.perKey, "billing-h-day": T.perDay };
|
||
for (var id in m) {
|
||
var el = document.getElementById(id);
|
||
if (el) el.textContent = m[id];
|
||
}
|
||
}
|
||
function render(st) {
|
||
if (!st) return;
|
||
renderTitles();
|
||
var T = L();
|
||
var cur = (st.currency || "USD");
|
||
var t = st.total || {};
|
||
document.getElementById("billing-kpis").innerHTML = [
|
||
[T.total, money(t.cost, cur)],
|
||
[T.requests, fmtInt(t.requests || 0)],
|
||
[T.degraded, fmtInt(st.degraded_reqs || 0)],
|
||
[T.unpriced, fmtInt(st.unpriced_reqs || 0)],
|
||
[T.prompt, fmtInt(t.prompt_tokens || 0)],
|
||
[T.completion, fmtInt(t.completion_tokens || 0)],
|
||
[T.failures, t.failures || 0],
|
||
// Cache KPIs: last session added the table columns but the KPI cards
|
||
// were left out — the edit's assert failed and the retry only re-did the
|
||
// tables. The numbers existed in state and nowhere in the UI.
|
||
[T.cacheRate, cacheRate(t)],
|
||
[T.cacheTokens, fmtInt(t.cache_hit_tokens || 0)]
|
||
].map(function (kv) {
|
||
// min-width:0:grid item 默认 min-width:auto,2e8 级长数字会把轨道撑出
|
||
// 容器造成横向溢出。标签 nowrap 截断,数值 break-all 换行。
|
||
return "<div class='card' style='padding:12px;min-width:0;overflow:hidden'>" +
|
||
"<div style='font-size:11px;opacity:.65;white-space:nowrap;overflow:hidden;text-overflow:ellipsis'>" +
|
||
esc(kv[0]) + "</div><div style='font-size:19px;font-weight:600;margin-top:4px;word-break:break-all;line-height:1.2'>" +
|
||
esc(kv[1]) + "</div></div>";
|
||
}).join("");
|
||
var TD = L();
|
||
tableFor(document.getElementById("billing-by-source"), st.by_source, cur, TD.noData);
|
||
tableFor(document.getElementById("billing-by-model"), st.by_model, cur, TD.noData);
|
||
tableFor(document.getElementById("billing-by-key"), st.by_key, cur, TD.noData);
|
||
tableFor(document.getElementById("billing-by-day"), st.by_day, cur, TD.noData);
|
||
}
|
||
async function refresh() {
|
||
try {
|
||
var r = await fetch("/api/plugins/" + ROOT + "/state", { credentials: "same-origin" });
|
||
if (!r.ok) return;
|
||
var j = await r.json();
|
||
render(j.state);
|
||
} catch (e) {
|
||
// Swallowing this is what made the production bug invisible: render() threw
|
||
// a ReferenceError on an undefined `s`, the catch ate it, every table kept
|
||
// its empty placeholder, and the page looked fine in the network tab while
|
||
// showing nothing. Still must not THROW (the pane is decoration and must
|
||
// never break the host page) — but it must leave a trace.
|
||
if (window.console && console.error) console.error("[billing] render failed", e);
|
||
}
|
||
}
|
||
window.__billingRefresh = refresh;
|
||
refresh();
|
||
if (window.pluginAPI) {
|
||
if (pluginAPI.onTabShown) pluginAPI.onTabShown(refresh);
|
||
if (pluginAPI.onLangChange) {
|
||
pluginAPI.onLangChange(function () {
|
||
renderTitles();
|
||
refresh();
|
||
});
|
||
}
|
||
}
|
||
})();
|
||
</script>
|
||
]==],
|
||
},
|
||
-- Two elements on the EXISTING status page: a headline tile and a
|
||
-- per-source cost breakdown, so the number is visible without opening the
|
||
-- Billing tab.
|
||
elements = {
|
||
{
|
||
target = "status",
|
||
anchor = "top",
|
||
order = 5,
|
||
mount = [==[
|
||
<div class="card" id="billing-status-tile" style="padding:12px;margin-bottom:12px">
|
||
<div style="font-size:11px;opacity:.65" id="billing-tile-label"></div>
|
||
<div id="billing-status-total" style="font-size:22px;font-weight:600;margin-top:4px">—</div>
|
||
<div id="billing-status-sub" style="font-size:12px;opacity:.65;margin-top:2px"></div>
|
||
</div>
|
||
<script>
|
||
(function () {
|
||
function fmt(n) {
|
||
n = Number(n || 0);
|
||
if (n === 0) return "0";
|
||
if (Math.abs(n) < 0.000001) return n.toExponential(2);
|
||
return n.toFixed(Math.abs(n) < 1 ? 6 : 4);
|
||
}
|
||
async function tick() {
|
||
try {
|
||
var r = await fetch("/api/plugins/billing/state", { credentials: "same-origin" });
|
||
if (!r.ok) return;
|
||
var j = await r.json();
|
||
var st = j.state;
|
||
if (!st || !st.total) return;
|
||
var cur = st.currency || "USD";
|
||
// The await above yields, so the host page may have rebuilt or torn down
|
||
// this element in the meantime — and it does: renderStatus assigns
|
||
// pane.innerHTML wholesale on every refresh. Assigning to a null element
|
||
// threw a TypeError that the surrounding catch logged on every repaint.
|
||
// Re-check after every await rather than assuming the DOM survived it.
|
||
var totalEl = document.getElementById("billing-status-total");
|
||
if (!totalEl) return;
|
||
totalEl.textContent = cur + " " + fmt(st.total.cost);
|
||
// The tile's label is plugin UI text, so it follows the host language via
|
||
// the same pluginAPI surface the Billing page uses.
|
||
var lab = document.getElementById("billing-tile-label");
|
||
if (lab) {
|
||
var lang = (window.pluginAPI && pluginAPI.lang) || "zh";
|
||
lab.textContent = lang === "zh" ? "总开销(billing 插件)" : "Total spend (billing plugin)";
|
||
}
|
||
var parts = [];
|
||
var srcs = st.by_source || {};
|
||
var names = Object.keys(srcs).sort(function (a, b) {
|
||
return (srcs[b].cost || 0) - (srcs[a].cost || 0);
|
||
});
|
||
for (var i = 0; i < Math.min(3, names.length); i++) {
|
||
parts.push(names[i] + " " + fmt(srcs[names[i]].cost));
|
||
}
|
||
var sub2 = document.getElementById("billing-status-sub");
|
||
if (sub2) sub2.textContent =
|
||
(st.total.requests || 0) + " requests" + (parts.length ? " · top: " + parts.join(" · ") : "");
|
||
} catch (e) {
|
||
// Same reasoning as the Billing page: decoration must never break the
|
||
// host page, but a silent catch turns a broken widget into "the plugin
|
||
// just doesn't show anything" with no way to tell why.
|
||
if (window.console && console.error) console.error("[billing] status tile refresh failed", e);
|
||
}
|
||
}
|
||
if (window.pluginAPI && pluginAPI.onTabShown) pluginAPI.onTabShown(tick);
|
||
tick();
|
||
})();
|
||
</script>
|
||
]==],
|
||
},
|
||
},
|
||
}
|
||
|
||
return plugin |