docs: source templates, split image AUTO chain, headless/desktop installers

- AUTO chain section: chat and image generation are now two independent
  chains (auto / auto_image) toggled in the Priority page; legacy image
  discovery fallback documented
- new Source templates section: multi-key balancing via reusable templates
  (Templates manager, From template / As template in the add-source dialog)
- Desktop GUI section: two installer kinds — Headless (server, plain binary)
  and Desktop (Electron GUI); release artifact naming ModelRouter-Headless-* /
  ModelRouter-Desktop-*; stale 1.0.0 version pins replaced with <ver>
This commit is contained in:
JianFeeeee
2026-08-27 13:02:48 +08:00
parent a5384d9fb6
commit 94cbcb6771
2 changed files with 105 additions and 25 deletions

View File

@ -176,12 +176,20 @@ tiered AUTO chain.
### AUTO chain (Priority page)
AUTO scheduling is driven **only** by the rules saved on the Priority page
(persisted as the `auto` field of the runtime file). The numeric
`models[].priority` in source configs no longer participates in scheduling and
is no longer shown.
AUTO scheduling is driven **only** by the rules saved on the Priority page. The
numeric `models[].priority` in source configs no longer participates in
scheduling and is no longer shown. **Chat and image generation are two
independent chains**, switched in the Priority page via the `Chat / Image` toggle:
- **Chat chain** (`auto` field): serves `POST /v1/chat/completions` with
`model: AUTO`; only `kind: chat` models are honored (image slots are skipped).
- **Image chain** (`auto_image` field): serves `POST /v1/images/generations` with
`model: AUTO`; only `kind: image` models are honored. With no image chain
configured, AUTO image generation falls back to legacy discovery (all sources
exposing an image model, in registry order) for backward compatibility.
Each rule is one "slot": `{ model, source, tier, token_quota, period, hours }`.
- Each rule is one "slot": `{ model, source, tier, token_quota, period, hours }`.
- `tier` is the priority tier: models in the same tier sit side by side and
share it; tiers run high → low.
- The same model may appear in several slots (e.g. A low → B low → A high →
@ -189,8 +197,9 @@ is no longer shown.
- `token_quota` > 0 means the slot is skipped once its tokens within the reset
window are exhausted; `period` supports `hour` / `week` / `month` / `nhour`
(with `hours`); empty = unlimited.
- Image models (`kind: image`) are kept out of the chain and are served by the
separate `POST /v1/images/generations` path.
- Within a tier, slots run by preference score; cooling / quota-exhausted /
hard-failed slots fall through to the next tier, and total failure returns
503 with a per-tier summary.
### Sources vs. adapters
@ -227,6 +236,28 @@ sources and `.lua` files; base sources defined in `config.yaml` cannot rewrite
that file, so they are hidden via a deletion tombstone (still hidden after
restart) and can be restored by re-adding the same name in the UI.
### Source templates (multi-key balancing)
When several upstream keys share one config (URL / adapter / model list /
concurrency / rpm), copying the whole source block N times is wasteful. A
**source template** stores every source field except `name` and `api_key`, so you
create N key-only-different sources from one shared skeleton — natural multi-key
load balancing (each expanded source schedules, cools down, and reports health
independently).
- WebUI Sources page, top-right **Templates** button: list all templates with
edit / delete / create.
- Add-source dialog has two header buttons:
- **From template** → pick a template → form auto-fills (name + key still
yours to enter);
- **As template** → save the current form's non-key/name fields as a template.
- Templates persist in the runtime file's `source_templates` field alongside
runtime sources (hot-managed, apply immediately). Templates hold no secret
keys, so they're safe to share / version-control.
A template is just a recipe — it doesn't become a source by itself. Only the
key-bearing sources created from it carry real traffic.
### disable_thinking
With `"disable_thinking": true` in the request body, the gateway passes it to
@ -308,14 +339,25 @@ tray, autostart, silent launch, and a pixel-for-pixel embedded full WebUI
(Status/Chat/Keys/Priority/Sources/Adapters — no login needed). Server users
keep using the plain Go binary.
### Two installer kinds: Headless and Desktop
Release installers ship in two flavors for different audiences:
- **Headless (server)**: the plain core binary (no Electron), config-driven,
for servers / containers / systemd / unattended runs. Linux ships as
`.tar.gz` (binary + `config.example.yaml` + `adapters/` + README), Windows as
`.zip`.
- **Desktop (desktop app)**: the Electron GUI installer with an embedded core,
for personal daily desktop use.
### GUI vs. plain backend
| Scenario | Pick | Why |
| ---- | ---- | ---- |
| Server / intranet gateway / unattended long-running | **plain backend** (single binary) | ~10 MB, ~15 MB RSS, zero-dep single process — drop it into systemd or any container, remote admin |
| Personal desktop daily use / multi-device intranet sharing | **GUI** (Electron) | no-login embedded WebUI, tray one-click, autostart, silent background — for non-CLI users |
| Windows desktop | **GUI** | plain backend needs manual service registration; GUI ships native tray/autostart |
| CI one-shot 3-platform installers | **GUI packaging scripts** | `make gui-*` emits deb / AppImage / NSIS, drops straight into a release pipeline |
| Server / intranet gateway / unattended long-running | **Headless** (single binary) | ~10 MB, ~15 MB RSS, zero-dep single process — drop it into systemd or any container, remote admin |
| Personal desktop daily use / multi-device intranet sharing | **Desktop** (GUI) | no-login embedded WebUI, tray one-click, autostart, silent background — for non-CLI users |
| Windows desktop | **Desktop** | plain backend needs manual service registration; GUI ships native tray/autostart |
| CI one-shot 3-platform installers | **Desktop packaging scripts** | `make gui-*` emits deb / AppImage / NSIS, drops straight into a release pipeline |
Both share the exact same `llmsproxy` core (LuaJIT build) — configs and
adapters are fully compatible, freely interchangeable.
@ -332,9 +374,11 @@ make gui-win # Windows NSIS (host mingw + wine)
make gui-win-docker # Windows NSIS, fully dockerized (no host mingw needed)
```
Artifacts land in `cmd/build/gui-dist/`: `ModelRouter-1.0.0.AppImage`,
`modelrouter-gui_1.0.0_amd64.deb`, `ModelRouter Setup 1.0.0.exe` (win). rpm
needs system `rpmbuild`.
Artifacts land in `cmd/build/gui-dist/`: `ModelRouter-<ver>.AppImage`,
`modelrouter-gui_<ver>_amd64.deb`, `ModelRouter Setup <ver>.exe` (win). rpm
needs system `rpmbuild`. At release time artifacts are renamed into the
`ModelRouter-Desktop-*` (GUI) and `ModelRouter-Headless-*` (plain backend)
families before upload to GitCode Releases.
#### Windows packaging (dockerized, recommended)