Files
ModelRouter/internal/lua/plugins/billing.lua
JianFeeeee d1da40e493 fix(billing): 页面文字溢出卡片 + 数字千分位
用户报 billing 页"文字超出展示框"。真因是三处叠加,都不是文案问题:

1. KPI 卡是 CSS grid,grid item 默认 min-width:auto。2e8 级的 prompt
   tokens(实测 270,149,209)无法收缩,于是把 grid 轨道撑出容器 → 整页横向
   溢出。修法是 min-width:0 + overflow:hidden(不给 min-width:0 的话,
   grid 子项永远不肯收缩,这是 grid 最常见的溢出坑)。
2. 表格写死 min-width:560px,窄视口下必然溢出。改 width:100% +
   table-layout:fixed,列宽由布局分配而不是由内容撑开。
3. 长名称(deepseek-v4.1-flash 这类模型 id)撑宽单元格。名称列改 ellipsis +
   title 悬停看全名;数字列 word-break:break-all 在列宽内换行。

顺手:所有 token/请求数走 toLocaleString 千分位。原始 9 位数字读起来要数
零位数,分组后 270,149,209 一眼可读,也顺带缩短了字符串宽度。

CDP 实测(820px 窄视口,逼出溢出条件):
  main/body 横向溢出 = no
  table=338 < card=366(修复前 min-width:560 必然 > 366)
  KPI 输入 tokens = 270,149,209
  billing 页无控制台错误

残留(不影响布局):表头 TH 在固定布局下 48>42 轻微超出自身格,因为
word-break 不拆单个长词;表格整体仍在容器内,未产生页面滚动。

全量测试全绿(8 包)。
2026-10-02 12:16:40 +08:00

728 lines
32 KiB
Lua
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

-- billing.lua — usage accounting plugin for ModelRouter.
--
-- Computes what each request cost, from three configurable dimensions:
--
-- source a flat per-request price for an upstream source
-- model a per-token price for a model id (prompt / completion separately)
-- key an override price for one gateway key
--
-- It then keeps running totals for the whole gateway, per source, per model
-- and per key, and publishes them in `plugin.state` so the kernel can serve
-- them at GET /api/plugins/billing/state — which is what its own dashboard
-- component reads.
--
-- ACCOUNTING BOUNDARY (important, and deliberate):
-- this plugin REPORTS; it does not ENFORCE. The gateway's own quota accounting
-- (internal/gateway/stats.go, enforced at request admission) stays
-- authoritative for limits. Two independent accounting paths that disagree are
-- worse than one that is slightly less featureful, so nothing here feeds back
-- into routing or quota decisions.
--
-- PRICE CONFIGURATION
-- Prices are supplied as a Lua table assigned to `billing.prices` before the
-- plugin is loaded, OR at runtime through PUT /api/plugins/billing/state. The
-- shape is:
--
-- billing.prices = {
-- currency = "USD", -- display only, no conversion happens
-- default = { prompt = 0, completion = 0, per_request = 0 },
-- sources = {
-- ["localzen"] = { per_request = 0.0 },
-- ["trae"] = { per_request = 0.01 },
-- },
-- models = {
-- ["gpt-5.4"] = { prompt = 1.25e-6, completion = 1e-5 }, -- USD per TOKEN
-- ["kimi-k3"] = { prompt = 6e-7, completion = 2.5e-6 },
-- ["kolors"] = { per_request = 0.04 }, -- image: flat
-- },
-- keys = {
-- -- by gateway key (the same value the audit log masks to ***xxxxxx)
-- ["***a1b2c3"] = { prompt = 1.1e-6, completion = 9e-6 },
-- },
-- }
--
-- Precedence for a token price: keys > models > default. A flat per_request
-- price, when present at any level, is ADDED on top of the token cost, so an
-- image model can carry both (e.g. tokens billed plus a fixed fee).
--
-- Numbers are USD per single token, which is how providers publish prices. That
-- makes a typical entry look like 1.25e-6; the plugin multiplies by the token
-- count, so no unit conversion happens anywhere.
local plugin = {
name = "billing",
version = "1.0.0",
description = "Per-source / per-model / per-key cost accounting with a dashboard",
author = "ModelRouter",
}
-- ---------- prices ----------
-- plugin.prices can be pre-seeded by embedding this file (an operator edits the
-- table below) or replaced at runtime through the state API. It is a SEPARATE
-- field from plugin.state on purpose: PUT /api/plugins/billing/state replaces
-- `state` wholesale, and prices must not live there or a price update would
-- wipe the accumulated totals. See docs/plugins.md.
local DEFAULT_PRICES = {
currency = "USD",
default = { prompt = 0, completion = 0, per_request = 0 },
sources = {},
models = {},
keys = {},
}
plugin.prices = DEFAULT_PRICES
-- Default prompt-cache discount. 0.1 = a cache read costs a tenth of a fresh
-- token, which is what DeepSeek/Qwen/Kimi and most others charge. It can be
-- overridden per price entry (prices.models.<m>.cache_discount) or globally by
-- setting plugin.cache_discount; 1 restores flat prompt pricing.
plugin.cache_discount = 0.1
-- ---------- accumulated totals ----------
-- state is what the kernel serves at GET /api/plugins/billing/state. It holds
-- ACCUMULATED TOTALS ONLY — prices live in plugin.prices (see above), so
-- replacing state never destroys a price table and updating prices never
-- destroys history.
--
-- Structure:
-- total { cost, requests, prompt_tokens, completion_tokens }
-- by_source { <name> = { cost, requests, ...tokens } }
-- by_model { <model> = { cost, ... } }
-- by_key { <masked key id> = { cost, ... } }
-- by_day { "YYYY-MM-DD" = { cost, ... } }
-- top_sources [ {name, cost, requests}, ... ] sorted, capped
-- top_models [ ... ]
-- top_keys [ ... ]
--
-- Sorted top-N lists are maintained incrementally rather than re-sorted on
-- every request: this hook runs once per request on the hot path, so it does
-- map updates only. The sort happens when state is READ.
-- CACHE ACCOUNTING (added after production showed the gap):
-- the gateway extracts prompt_cache_hit_tokens from upstream usage and puts it
-- in the request_end payload, and costFor() already used it to price the cache
-- leg — but no bucket recorded it. So a gateway where 99.88% of prompt tokens
-- were cache reads showed a prompt_tokens number with no indication of that,
-- and there was no way to see cache hit rate per source/model/key at all.
--
-- cache_hit_tokens hits, as reported by upstream
-- cache_fresh_tokens prompt tokens that were NOT cache reads
-- cache_reported_reqs requests where upstream gave a cache number at all.
-- Kept separate from a zero: "upstream does not report cache usage" and
-- "upstream reported zero hits" look identical in a hit total, and they mean
-- opposite things when you are trying to work out whether a cache discount is
-- doing anything.
local function emptyBucket()
return {
cost = 0, requests = 0, prompt_tokens = 0, completion_tokens = 0, failures = 0,
cache_hit_tokens = 0, cache_fresh_tokens = 0, cache_reported_reqs = 0,
}
end
plugin.state = {
total = emptyBucket(),
by_source = {},
by_model = {},
by_key = {},
by_day = {},
started = os.time and 0 or 0,
}
local function bucket(tbl, k)
local b = tbl[k]
if b == nil then
b = emptyBucket()
tbl[k] = b
end
return b
end
local function add(b, cost, prompt, completion, ok, cacheHit, cacheReported)
b.cost = b.cost + cost
b.requests = b.requests + 1
b.prompt_tokens = b.prompt_tokens + prompt
b.completion_tokens = b.completion_tokens + completion
if not ok then b.failures = b.failures + 1 end
-- Backfill guards a bucket that predates these fields (a state file written
-- by an older build, or one restored from disk): nil + number is an error in
-- Lua, and a hook that throws stops accounting for that request entirely.
if b.cache_hit_tokens == nil then b.cache_hit_tokens = 0 end
if b.cache_fresh_tokens == nil then b.cache_fresh_tokens = 0 end
if b.cache_reported_reqs == nil then b.cache_reported_reqs = 0 end
b.cache_hit_tokens = b.cache_hit_tokens + (cacheHit or 0)
b.cache_fresh_tokens = b.cache_fresh_tokens + ((prompt or 0) - (cacheHit or 0))
if cacheReported then b.cache_reported_reqs = b.cache_reported_reqs + 1 end
end
-- ---------- pricing ----------
-- lookup walks keys > models > default and returns a price triple plus whether
-- a flat per_request component applies.
local function priceFor(payload)
local p = plugin.prices or DEFAULT_PRICES
local d = p.default or {}
-- Whether ANY dimension actually priced this request. A request that ends up
-- with all-zero prices is not "free", it is UNPRICED, and the two must not
-- look the same: an unpriced model silently costing 0 is the most dangerous
-- failure mode a cost plugin has, because the bill still adds up and just
-- quietly under-reports. It is counted separately and surfaced in the UI.
out = {
prompt = d.prompt or 0, completion = d.completion or 0,
per_request = 0, cache_discount = d.cache_discount, peak = d.peak,
}
-- model dimension (a token price overrides the default's token prices)
local mp = p.models and p.models[payload.model]
if mp then
out.priced = true
if mp.prompt ~= nil then out.prompt = mp.prompt end
if mp.completion ~= nil then out.completion = mp.completion end
if mp.per_request ~= nil then out.per_request = out.per_request + mp.per_request end
if mp.cache_discount ~= nil then out.cache_discount = mp.cache_discount end
if mp.peak ~= nil then out.peak = mp.peak end
end
-- source dimension: usually a flat fee, but may also carry token prices
local sp = p.sources and p.sources[payload.source]
if sp then
out.priced = true
if sp.prompt ~= nil then out.prompt = sp.prompt end
if sp.completion ~= nil then out.completion = sp.completion end
if sp.per_request ~= nil then out.per_request = out.per_request + sp.per_request end
if sp.cache_discount ~= nil then out.cache_discount = sp.cache_discount end
if sp.peak ~= nil then out.peak = sp.peak end
end
-- key dimension wins over the others (an operator pricing one customer
-- specially must be able to override both the model and the source price)
local kp = p.keys and p.keys[payload.key]
if kp then
out.priced = true
if kp.prompt ~= nil then out.prompt = kp.prompt end
if kp.completion ~= nil then out.completion = kp.completion end
if kp.per_request ~= nil then out.per_request = out.per_request + kp.per_request end
if kp.cache_discount ~= nil then out.cache_discount = kp.cache_discount end
if kp.peak ~= nil then out.peak = kp.peak end
end
return out
end
-- ===== 峰谷 / 时段定价 ================================================
--
-- 有些 provider 按 UTC 时段分价(commandcode 的 DeepSeek V4 系列就是:高峰
-- 01-04 & 06-10 UTC 工作日,价格恰好是非高峰的 2 倍)。静态价目无法表达这一点,
-- 而算错方向通常是【静默高估或低估】,不会报错——所以这里显式支持。
--
-- 配置形态(挂在任一维度的价目条目上):
--
-- "deepseek-v4.1-flash": {
-- prompt = 1.5e-7, completion = 6e-7,
-- peak = {
-- multiplier = 2, -- 高峰时单价乘以它
-- windows = [ -- UTC 星期几 = os.date 的 %w(周日=1)
-- { days = {2,3,4,5,6}, hours = {{1,2,3},{6,7,8,9}} },
-- ],
-- },
-- }
--
-- 语义:命中任一 window ⇒ 乘以 multiplier。hours 用 {起,止} 闭区间,跨零点
-- 用 {{22,24}} 表示 22:00-24:00(24 是"当天最后一刻")。
--
-- ★ 为什么用 os.date 的 ! 前缀取 UTC:provider 的费率表按 UTC 标注,而网关
-- 跑在本地时区(这台机是 Asia/Hong_Kong)。混用本地小时会让峰谷整体偏移 8
-- 小时,白天算成夜间——比不做峰谷还糟。
local function inPeakWindow(ev)
if ev == nil then return false end
local w = ev.windows
if type(w) ~= "table" or #w == 0 then return false end
local dow = tonumber(os.date("!%w")) or 0 -- 0=Sunday
local hour = tonumber(os.date("!%H")) or 0
for _, win in ipairs(w) do
local days = win.days
if type(days) == "table" then
local day_ok = false
for _, d in ipairs(days) do
if tonumber(d) == dow then day_ok = true break end
end
if not day_ok then goto continue_win end
end
local hours = win.hours
if type(hours) == "table" then
for _, h in ipairs(hours) do
local lo, hi = tonumber(h[1]), tonumber(h[2])
if lo and hi and hour >= lo and hour <= hi then return true end
end
end
::continue_win::
end
return false
end
-- applyPeak multiplies a price by the peak rule, if the request lands in a peak
-- window. It is a no-op when no rule is configured, so the common case costs one
-- nil check.
--
-- The multiplier is RECORDED, not applied to price.prompt in place. That looks
-- like a roundabout way to do it, but applying it there was a real bug: the
-- cache-read rate is DERIVED from price.prompt inside costFor, so doubling
-- price.prompt silently doubled the cache read too — compounding two separate
-- discounts. Keeping the multiplier separate lets costFor scale the fresh-prompt
-- and completion legs and leave the cache leg alone, which is what "peak rates
-- apply to the token price, cache reads are billed at their own rate" means.
local function applyPeak(price)
local pk = price.peak
if pk == nil then return price end
if not inPeakWindow(pk) then return price end
local m = tonumber(pk.multiplier) or 1
if m <= 0 then return price end
price.peak_multiplier = m
return price
end
-- costFor computes one request's price.
--
-- PROMPT CACHE: a cached prompt token is not billed like a fresh one. Almost
-- every provider sells cache reads at a steep discount (commonly 10% of the
-- fresh rate), and cache-heavy agent traffic hits long shared prefixes hard.
-- Charging the full prompt rate made a 1M-token request of which 900k were
-- cache reads come out at 10 USD instead of ~1.9 — an order of magnitude, on
-- exactly the traffic the cache exists to make cheap. The plugin therefore
-- splits the prompt count:
--
-- fresh = prompt_tokens - cache_hit_tokens -> full rate
-- cached = cache_hit_tokens -> rate * cache_discount
--
-- cache_discount defaults to 0.1 (the common 10x). It is configurable because
-- the ratio is a per-provider fact, not a constant of nature: set it to 1 to
-- keep the old flat behaviour, or 0 for providers that do not discount.
--
-- A request that reports cache_hit_tokens LARGER than prompt_tokens (a
-- misbehaving adapter, or two upstreams' numbers being mixed) is clamped: the
-- fresh count never goes negative, which would silently turn a request into
-- billable negative tokens.
local function costFor(payload, price)
price = applyPeak(price or priceFor(payload))
local prompt = tonumber(payload.prompt_tokens) or 0
local completion = tonumber(payload.completion_tokens) or 0
local cacheHit = tonumber(payload.cache_hit_tokens) or 0
if cacheHit < 0 then cacheHit = 0 end
if cacheHit > prompt then cacheHit = prompt end
local discount = tonumber(price.cache_discount)
if discount == nil then discount = plugin.cache_discount end
if discount == nil then discount = 0.1 end
if discount < 0 then discount = 0 elseif discount > 1 then discount = 1 end
-- The peak multiplier applies to the freshly-read prompt tokens and the
-- completion, but NOT to the cache read: a cache read is a separate upstream
-- rate that the off-peak figures already discount, and doubling it would
-- stack two discounts the provider never intended to stack.
local mult = tonumber(price.peak_multiplier) or 1
local fresh = prompt - cacheHit
local cost = fresh * price.prompt * mult
+ cacheHit * price.prompt * discount
+ completion * price.completion * mult
local flat = price.per_request
if not payload.ok and not plugin.count_failures then
flat = 0
end
return cost + flat
end
-- ---------- day bucket ----------
local function dayKey(epoch_seconds)
-- os.date is available in LuaJIT; fall back to a UTC-ish arithmetic stamp if
-- the host build has no os.date (keeps the plugin from erroring out on a
-- stripped runtime, which would otherwise look like a plugin failure).
if os and os.date then
return os.date("!%Y-%m-%d", epoch_seconds)
end
return tostring(math.floor(epoch_seconds / 86400))
end
-- ---------- hooks ----------
plugin.hooks = {
-- chain_step gives the per-tier walk; request_end gives the final accounting.
-- Subscribing to chain_step is OPTIONAL here: the totals are driven by
-- request_end alone, and the degradation counters below are pure observation.
-- A gateway with thousands of requests can drop this hook to save the
-- per-step Lua call without losing a single billed request.
chain_step = "on_chain_step",
request_end = "on_request_end",
}
-- Tracks how often a request had to drop below the top tier, and which tier
-- actually served it. Without this, "tier 1 was cooling" and "tier 1 served it"
-- are indistinguishable in the accounts, and a quietly degraded gateway looks
-- exactly like a healthy one.
plugin.state.degraded_reqs = 0
plugin.state.by_tier_served = {}
plugin.state.skip_reasons = {}
function plugin.on_chain_step(payload)
if payload == nil then return nil end
local s = plugin.state
if s == nil then return nil end
if s.by_tier_served == nil then s.by_tier_served = {} end
if s.skip_reasons == nil then s.skip_reasons = {} end
if payload.kind == "selected" then
local t = tostring(payload.tier or "?")
s.by_tier_served[t] = (s.by_tier_served[t] or 0) + 1
elseif payload.kind == "tier_skip" or payload.kind == "tier_busy" then
-- reason text is the ACTIONABLE part; normalise the volatile bits so the
-- same cause aggregates instead of creating a new row per request.
local r = tostring(payload.reason or payload.kind or "unknown")
r = string.gsub(r, "within [%d%.%a]+", "within <wait>")
s.skip_reasons[r] = (s.skip_reasons[r] or 0) + 1
end
return nil
end
function plugin.on_request_end(payload)
if payload == nil then return nil end
local prompt = tonumber(payload.prompt_tokens) or 0
local completion = tonumber(payload.completion_tokens) or 0
local ok = payload.ok and true or false
local price = priceFor(payload)
local cost = costFor(payload, price)
local s = plugin.state
-- Rebuild any missing container. This is reached in two real situations:
-- a fresh plugin, and an admin who PUT a partial state (e.g. only "prices"),
-- which legitimately replaces `state` with a sparse table. Checking only the
-- outer table would leave `s.total` nil and crash the hook on the next call.
if s == nil then s = {} plugin.state = s end
if s.total == nil then s.total = emptyBucket() end
if s.by_source == nil then s.by_source = {} end
if s.by_model == nil then s.by_model = {} end
if s.by_key == nil then s.by_key = {} end
if s.by_day == nil then s.by_day = {} end
if s.started == nil then s.started = payload.time or 0 end
if s.unpriced_reqs == nil then s.unpriced_reqs = 0 end
if s.unpriced_models == nil then s.unpriced_models = {} end
-- Track traffic that no price entry covered. This MUST come after the
-- container rebuild above: an earlier version referenced `s` before it was
-- declared, so on a fresh plugin the hook threw and the request recorded
-- NOTHING at all — the worst possible failure for a billing plugin, and one
-- that only showed up as "requests = 0" in a test.
if not price.priced then
s.unpriced_reqs = s.unpriced_reqs + 1
local m = payload.model or "?"
s.unpriced_models[m] = (s.unpriced_models[m] or 0) + 1
end
if s.degraded_reqs == nil then s.degraded_reqs = 0 end
-- Degradation is counted here rather than in the chain_step hook because
-- request_end sees the whole walk at once: one degraded request must count
-- once, whereas the walk may contain several skipped tiers.
if payload.degraded then s.degraded_reqs = s.degraded_reqs + 1 end
local cacheHit = tonumber(payload.cache_hit_tokens) or 0
if cacheHit < 0 then cacheHit = 0 end
if cacheHit > prompt then cacheHit = prompt end
-- cache_reported is the gateway's own signal that UPSTREAM gave a cache
-- number. Without it a source that never reports cache usage is
-- indistinguishable from one that always reports zero hits.
local cacheReported = payload.cache_reported and true or false
local C = cacheHit
local R = cacheReported
add(s.total, cost, prompt, completion, ok, C, R)
if payload.source ~= nil and payload.source ~= "" then
add(bucket(s.by_source, payload.source), cost, prompt, completion, ok, C, R)
end
if payload.model ~= nil and payload.model ~= "" then
add(bucket(s.by_model, payload.model), cost, prompt, completion, ok, C, R)
end
if payload.key ~= nil and payload.key ~= "" then
add(bucket(s.by_key, payload.key), cost, prompt, completion, ok, C, R)
end
-- Daily rollup, so the dashboard can draw a trend without the browser
-- re-deriving it. Keyed off the request's own timestamp, not os.time(), so a
-- replayed or imported record lands on the right day.
local ts = payload.time
if ts ~= nil and ts > 0 then
if ts > 1000000000000 then ts = ts / 1000 end -- kernel sends unix MILLIseconds
add(bucket(s.by_day, dayKey(ts)), cost, prompt, completion, ok, C, R)
end
return nil -- last stage: nobody downstream would read a return value
end
-- ---------- dashboard UI ----------
-- A whole page. The kernel injects this HTML and evaluates the <script> after
-- the DOM exists, and exposes `pluginAPI` for talking to the gateway.
plugin.ui = {
page = {
page_id = "billing",
title = "Billing",
icon = [==[<svg viewBox="0 0 24 24"><circle cx="12" cy="12" r="9"/><path d="M14.5 9.5a3 3 0 0 0-2.5-1.3c-1.4 0-2.4.7-2.4 1.8 0 2.6 5.2 1.4 5.2 4 0 1.1-1 1.8-2.5 1.8-1.1 0-2.1-.4-2.7-1.2"/><path d="M12 6.4v11.2"/></svg>]==],
order = 40,
mount = [==[
<div id="billing-root" style="padding:16px;min-width:0;max-width:100%;overflow-x:auto">
<div class="kpis" id="billing-kpis" style="display:grid;grid-template-columns:repeat(auto-fit,minmax(170px,1fr));gap:12px;margin-bottom:18px"></div>
<div style="display:grid;grid-template-columns:repeat(auto-fit,minmax(320px,1fr));gap:16px">
<div class="card" style="padding:14px">
<h3 id="billing-h-src" style="margin:0 0 10px;font-size:14px"></h3>
<div id="billing-by-source"></div>
</div>
<div class="card" style="padding:14px">
<h3 id="billing-h-model" style="margin:0 0 10px;font-size:14px"></h3>
<div id="billing-by-model"></div>
</div>
<div class="card" style="padding:14px">
<h3 id="billing-h-key" style="margin:0 0 10px;font-size:14px"></h3>
<div id="billing-by-key"></div>
</div>
</div>
<div class="card" style="padding:14px;margin-top:16px">
<h3 id="billing-h-day" style="margin:0 0 10px;font-size:14px"></h3>
<div id="billing-by-day"></div>
</div>
</div>
<script>
(function () {
var ROOT = "billing";
// UI strings, bilingual. The host page i18n (applyI18n/data-i) only covers
// markup the HOST renders; a plugin's injected markup is invisible to it, so
// the Billing page stayed English while the rest of the UI switched. These go
// through pluginAPI.lang / onLangChange, the small surface the host exposes
// for exactly this.
var STR = {
en: {
total: "Total", requests: "Requests", degraded: "Degraded",
unpriced: "Unpriced", prompt: "Prompt tokens",
completion: "Completion tokens", failures: "Failures",
cacheRate: "Cache hit rate", cacheTokens: "Cache read tokens",
perSource: "Per source", perModel: "Per model",
perKey: "Per gateway key", perDay: "Daily",
thName: "name", thCost: "cost", thReqs: "reqs", thPrompt: "prompt",
thFresh: "fresh", thCache: "cache", thCachePct: "cache%",
thCompletion: "completion",
noData: "no data yet",
},
zh: {
total: "总开销", requests: "请求数", degraded: "降级",
unpriced: "未定价", prompt: "输入 tokens",
completion: "输出 tokens", failures: "失败",
cacheRate: "缓存命中率", cacheTokens: "缓存读取 tokens",
perSource: "按源", perModel: "按模型",
perKey: "按网关密钥", perDay: "按天",
thName: "名称", thCost: "开销", thReqs: "请求", thPrompt: "输入",
thFresh: "新鲜", thCache: "缓存", thCachePct: "缓存%",
thCompletion: "输出",
noData: "暂无数据",
},
};
function L() {
var lang = (window.pluginAPI && pluginAPI.lang) || "zh";
return STR[lang] || STR.zh;
}
function fmt(n) {
if (n === null || n === undefined) return "-";
n = Number(n);
if (!isFinite(n)) return "-";
if (n === 0) return "0";
if (Math.abs(n) < 0.000001) return n.toExponential(2);
return n.toFixed(Math.abs(n) < 1 ? 6 : 4);
}
function money(v, cur) { return (cur || "USD") + " " + fmt(v); }
function esc(s) {
return String(s == null ? "" : s).replace(/[&<>"]/g, function (c) {
return { "&": "&amp;", "<": "&lt;", ">": "&gt;", '"': "&quot;" }[c];
});
}
// Cache hit rate, with the reporting caveat made visible.
//
// A bucket whose upstream never reports cache usage would render as "0%" from
// a 0/0 and read as "the cache is not working", when the truth is "this
// provider does not tell us". "n/r" keeps those apart.
function cacheRate(b) {
var prompt = b.prompt_tokens || 0;
var hit = b.cache_hit_tokens || 0;
if (!prompt) return "\u2014";
if (!b.cache_reported_reqs) return "n/r";
return ((hit / prompt) * 100).toFixed(1) + "%";
}
// fmtInt 千分位分组:2e8 级 token 总数可读、也更短,降低撑宽风险。
function fmtInt(n) {
return (Number(n) || 0).toLocaleString("en-US");
}
function row(name, b, cur) {
var fresh = (b.cache_fresh_tokens === undefined) ? (b.prompt_tokens || 0) : b.cache_fresh_tokens;
// 名称列 ellipsis(title 悬停看全名);数字列 break-all 在列宽内换行而不是
// 把表格撑出卡片。单元格结构与列数不变,列数判据不受影响。
return "<tr><td style='overflow:hidden'><b style='display:block;white-space:nowrap;overflow:hidden;text-overflow:ellipsis' title='" +
esc(String(name).replace(/'/g, "&#39;")) + "'>" + esc(name) + "</b></td>" +
"<td style='word-break:break-all'>" + money(b.cost, cur) + "</td>" +
"<td style='word-break:break-all'>" + fmtInt(b.requests || 0) + "</td>" +
"<td style='word-break:break-all'>" + fmtInt(b.prompt_tokens || 0) + "</td>" +
"<td style='word-break:break-all'>" + fmtInt(fresh) + "</td>" +
"<td style='word-break:break-all'>" + fmtInt(b.cache_hit_tokens || 0) + "</td>" +
"<td style='word-break:break-all'>" + esc(cacheRate(b)) + "</td>" +
"<td style='word-break:break-all'>" + fmtInt(b.completion_tokens || 0) + "</td></tr>";
}
function tableFor(el, obj, cur, empty) {
var keys = Object.keys(obj || {});
if (!keys.length) { el.innerHTML = '<div class="muted">' + empty + "</div>"; return; }
keys.sort(function (a, b) { return (obj[b].cost || 0) - (obj[a].cost || 0); });
var TH = L();
var h = "<table style='width:100%;border-collapse:collapse;font-size:13px;table-layout:fixed;word-break:break-word'>" +
"<tr style='text-align:left;opacity:.65'><th>" + TH.thName + "</th><th>" + TH.thCost +
"</th><th>" + TH.thReqs + "</th><th>" + TH.thPrompt +
"</th><th>" + TH.thFresh + "</th><th>" + TH.thCache + "</th><th>" + TH.thCachePct +
"</th><th>" + TH.thCompletion + "</th></tr>";
for (var i = 0; i < keys.length; i++) {
var k = keys[i];
h += "<tr style='border-top:1px solid rgba(120,90,150,.14)'>" + row(k, obj[k], cur) + "</tr>";
}
el.innerHTML = h + "</table>";
}
function renderTitles() {
var T = L();
var m = { "billing-h-src": T.perSource, "billing-h-model": T.perModel,
"billing-h-key": T.perKey, "billing-h-day": T.perDay };
for (var id in m) {
var el = document.getElementById(id);
if (el) el.textContent = m[id];
}
}
function render(st) {
if (!st) return;
renderTitles();
var T = L();
var cur = (st.currency || "USD");
var t = st.total || {};
document.getElementById("billing-kpis").innerHTML = [
[T.total, money(t.cost, cur)],
[T.requests, fmtInt(t.requests || 0)],
[T.degraded, fmtInt(st.degraded_reqs || 0)],
[T.unpriced, fmtInt(st.unpriced_reqs || 0)],
[T.prompt, fmtInt(t.prompt_tokens || 0)],
[T.completion, fmtInt(t.completion_tokens || 0)],
[T.failures, t.failures || 0],
// Cache KPIs: last session added the table columns but the KPI cards
// were left out — the edit's assert failed and the retry only re-did the
// tables. The numbers existed in state and nowhere in the UI.
[T.cacheRate, cacheRate(t)],
[T.cacheTokens, fmtInt(t.cache_hit_tokens || 0)]
].map(function (kv) {
// min-width:0:grid item 默认 min-width:auto,2e8 级长数字会把轨道撑出
// 容器造成横向溢出。标签 nowrap 截断,数值 break-all 换行。
return "<div class='card' style='padding:12px;min-width:0;overflow:hidden'>" +
"<div style='font-size:11px;opacity:.65;white-space:nowrap;overflow:hidden;text-overflow:ellipsis'>" +
esc(kv[0]) + "</div><div style='font-size:19px;font-weight:600;margin-top:4px;word-break:break-all;line-height:1.2'>" +
esc(kv[1]) + "</div></div>";
}).join("");
var TD = L();
tableFor(document.getElementById("billing-by-source"), st.by_source, cur, TD.noData);
tableFor(document.getElementById("billing-by-model"), st.by_model, cur, TD.noData);
tableFor(document.getElementById("billing-by-key"), st.by_key, cur, TD.noData);
tableFor(document.getElementById("billing-by-day"), st.by_day, cur, TD.noData);
}
async function refresh() {
try {
var r = await fetch("/api/plugins/" + ROOT + "/state", { credentials: "same-origin" });
if (!r.ok) return;
var j = await r.json();
render(j.state);
} catch (e) {
// Swallowing this is what made the production bug invisible: render() threw
// a ReferenceError on an undefined `s`, the catch ate it, every table kept
// its empty placeholder, and the page looked fine in the network tab while
// showing nothing. Still must not THROW (the pane is decoration and must
// never break the host page) — but it must leave a trace.
if (window.console && console.error) console.error("[billing] render failed", e);
}
}
window.__billingRefresh = refresh;
refresh();
if (window.pluginAPI) {
if (pluginAPI.onTabShown) pluginAPI.onTabShown(refresh);
if (pluginAPI.onLangChange) {
pluginAPI.onLangChange(function () {
renderTitles();
refresh();
});
}
}
})();
</script>
]==],
},
-- Two elements on the EXISTING status page: a headline tile and a
-- per-source cost breakdown, so the number is visible without opening the
-- Billing tab.
elements = {
{
target = "status",
anchor = "top",
order = 5,
mount = [==[
<div class="card" id="billing-status-tile" style="padding:12px;margin-bottom:12px">
<div style="font-size:11px;opacity:.65" id="billing-tile-label"></div>
<div id="billing-status-total" style="font-size:22px;font-weight:600;margin-top:4px">—</div>
<div id="billing-status-sub" style="font-size:12px;opacity:.65;margin-top:2px"></div>
</div>
<script>
(function () {
function fmt(n) {
n = Number(n || 0);
if (n === 0) return "0";
if (Math.abs(n) < 0.000001) return n.toExponential(2);
return n.toFixed(Math.abs(n) < 1 ? 6 : 4);
}
async function tick() {
try {
var r = await fetch("/api/plugins/billing/state", { credentials: "same-origin" });
if (!r.ok) return;
var j = await r.json();
var st = j.state;
if (!st || !st.total) return;
var cur = st.currency || "USD";
// The await above yields, so the host page may have rebuilt or torn down
// this element in the meantime — and it does: renderStatus assigns
// pane.innerHTML wholesale on every refresh. Assigning to a null element
// threw a TypeError that the surrounding catch logged on every repaint.
// Re-check after every await rather than assuming the DOM survived it.
var totalEl = document.getElementById("billing-status-total");
if (!totalEl) return;
totalEl.textContent = cur + " " + fmt(st.total.cost);
// The tile's label is plugin UI text, so it follows the host language via
// the same pluginAPI surface the Billing page uses.
var lab = document.getElementById("billing-tile-label");
if (lab) {
var lang = (window.pluginAPI && pluginAPI.lang) || "zh";
lab.textContent = lang === "zh" ? "总开销(billing 插件)" : "Total spend (billing plugin)";
}
var parts = [];
var srcs = st.by_source || {};
var names = Object.keys(srcs).sort(function (a, b) {
return (srcs[b].cost || 0) - (srcs[a].cost || 0);
});
for (var i = 0; i < Math.min(3, names.length); i++) {
parts.push(names[i] + " " + fmt(srcs[names[i]].cost));
}
var sub2 = document.getElementById("billing-status-sub");
if (sub2) sub2.textContent =
(st.total.requests || 0) + " requests" + (parts.length ? " · top: " + parts.join(" · ") : "");
} catch (e) {
// Same reasoning as the Billing page: decoration must never break the
// host page, but a silent catch turns a broken widget into "the plugin
// just doesn't show anything" with no way to tell why.
if (window.console && console.error) console.error("[billing] status tile refresh failed", e);
}
}
if (window.pluginAPI && pluginAPI.onTabShown) pluginAPI.onTabShown(tick);
tick();
})();
</script>
]==],
},
},
}
return plugin