Files
MailUI4Agents/plugins/zcode-mail-bridge/test/prompt.test.mjs
JianFeeeee 29ad8aa204 feat(mcp): 去 ZCode 影子 —— mcp/server.mjs 改为通用 MCP 服务
## 目的

`mcp/server.mjs` 此前注释与行为都绑定 ZCode,接入端必须为 AgentMail 写
专用插件。去掉这层绑定后,任何支持 MCP 的宿主挂一行配置即可用:

    {"command":"node","args":["…/mcp/server.mjs"],"env":{
      "AGENTMAIL_GATEWAY_URL":…,"AGENTMAIL_AGENT_NAME":…,
      "AGENTMAIL_AGENT_SECRET":…,"AGENTMAIL_MCP_PLATFORM":"my-host"}}

协议层(零依赖手写 stdio JSON-RPC)与 11 个邮件工具本就与宿主无关,
真正要动的只有 4 处耦合 + 工具面。

## 改动

**1. 移除执行类工具(`run_command` / `write_file`)**
它们的门禁(lib/action-tools.mjs + lib/approval.mjs + 落盘授权表)是为
ZCode headless 的**双进程审批**设计的:MCP 进程问人、ZCode 钩子进程等回答、
中间靠文件对齐。脱离该宿主后这套门禁的前提不成立,挂在通用服务上等于
提供一条**没有审批的旁路**。
`lib/` 里三个模块与 `hooks/` 源码保留(桌面模式的 ZCode 仍走它们),
只是 server.mjs 不再装载。

**2. platform 可配置**:`AGENTMAIL_MCP_PLATFORM`,默认 `mcp`,
空白值回落默认值。原先硬编码 `'zcode'`(两处)。

**3. 错误文案去宿主名**:不再让模型/人「去 ZCode 的插件设置里填写」,
改为说明设置 `AGENTMAIL_*` 环境变量。

**4. 提示词如实说能力**(src/prompt.mjs):原文案向模型承诺
`run_command`/`write_file` 可用并分档描述「会被请示 / 直接生效」。
工具移除后那变成**指向不存在工具的承诺** —— 模型会去找、把整轮浪费在
换名字重试上。改为明说「本平台没有执行面,需要动手就写进回信请人做」。
三档措辞仍互不相同(`plan`/`workspace`/`full`),因为「档位仍存在但都无
执行面」这件事模型需要知道。

## ★★ 顺带修掉一个真实缺陷(端到端撞出来的)

`connect_to_server` 对 secret-only 的 Agent **一直 400**:
`/agent/register` 只认 `Authorization: Bearer` 或 body 里的 `secret`,
不认 `X-Agent-Secret` 头(其它接口才认),而它漏了 `body.secret`。
dsh / pi 正是 secret-only 配置 ⇒ 它们调「连一下服务器」必然失败,
且模型看不出该改什么。
lib/gateway.mjs 的 `register()` 本来就做对了,tools.mjs 里是手抄的劣化副本。
修后实测 `HTTP 400` → `已连接 …(状态:registered)`。

## 判据

新增 `test/generic-mcp.test.mjs`(5 格)。**这三件事此前无人看守**:
变异验证时「把 action-tools 挂回 server.mjs」与「platform 硬编码回 zcode」
都能全套通过 —— 因为没有判据看 server.mjs 实际挂了什么、也没人看 platform。

改写的 4 格(prompt 3 格 + driver 1 格)保留原意图(不向模型撒谎、
native 自报要有真凭据、工具不存在时不要重试),改为断言新事实。

**变异验证**(每条都确认已应用后才数红格):

    挂回 action-tools            → 红 3
    platform 硬编码 zcode        → 红 3
    platform 空白不回落           → 红 3
    文案指回 ZCode 插件设置        → 红 3
    删掉 body.secret(400 复现)  → 红 3

全套 **402/402**。

## 端到端验收

写了一个**非 ZCode 宿主**探针(纯 stdio JSON-RPC,不加载任何插件),
对着真实网关跑通:initialize → tools/list(11 个,无执行类)→
connect_to_server(registered)→ suggest_address。

## 未做

- 未发布到 npm registry(`npx` 即用需要发布或指向仓库路径)。
- 未改 `check-deploy-drift.mjs` 的 zcode 豁免(本机仍不退场该宿主)。
2026-10-02 12:31:18 +08:00

182 lines
8.7 KiB
JavaScript
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

/**
* 提示词与回信文案的测试。
*
* 这层没有 I/O,但它决定模型看到什么 —— 而模型看到的东西错了,表现是
* 「这个 Agent 就是不回信」或「两个 Agent 无限客套」,都不会报错。
*/
import { test } from 'node:test';
import assert from 'node:assert/strict';
import { buildMailPrompt, replySubject, renderTurnFailure } from '../src/prompt.mjs';
const mail = (over = {}) => ({
mail_id: 'm1',
from_name: 'gui-lab',
from_human: true,
subject: '帮我看看',
reply_address: 'gui-lab@/w.alias',
...over
});
test('新任务:说明这是新邮件,并给出发件人/主题/邮件 id', () => {
const p = buildMailPrompt({ agentName: 'zcode', data: mail() });
assert.match(p, /gui-lab/);
assert.match(p, /帮我看看/);
assert.match(p, /m1/);
assert.match(p, /身份:你是 zcode/);
});
test('★ 回信不是新任务:带 in_reply_to 时要说清是回哪封', () => {
// 模型分不清「新任务」与「回信」时,会把对方一句「已收到」再当待办做一遍。
//
// ★ 2026-10-02 修陈旧断言(b0c8719 漏改):b0c8719 给 in_reply_to 加了
// **方向判据**——只有 parent_from === 自己时才说「回的是你那封」
// (单向续信链曾被逐封误读成双向对话)。本测试当时只传了 in_reply_to,
// 断言的却是旧的无条件文案 ⇒ 假红。夹具补上 parent_from,并顺手把
// 三个方向场景钉全(改 relay-policy 而漏改 prompt 的测试正是这个教训)。
const mine = buildMailPrompt({ agentName: 'zcode', data: mail({ in_reply_to: 'm0', parent_from: 'zcode' }) });
assert.match(mine, /回的是你那封:m0/);
// 多方续谈:父邮件是别人发的,不得宣称「回的是你那封」
const others = buildMailPrompt({ agentName: 'zcode', data: mail({ in_reply_to: 'm0', parent_from: 'pi' }) });
assert.doesNotMatch(others, /回的是你那封/);
assert.match(others, /续谈/);
// 服务端未升级(无 parent_from):退回「回复到了」但同样不指认方向
const noParent = buildMailPrompt({ agentName: 'zcode', data: mail({ in_reply_to: 'm0' }) });
assert.doesNotMatch(noParent, /回的是你那封/);
assert.match(noParent, /回复.*到了|回复到了/);
});
test('★ 人来信 vs Agent 来信:回信责任必须不同', () => {
const fromHuman = buildMailPrompt({ agentName: 'z', data: mail({ from_human: true }) });
const fromAgent = buildMailPrompt({ agentName: 'z', data: mail({ from_human: false }) });
assert.match(fromHuman, /回信不用你自己发/);
assert.match(fromAgent, /不会替你回信/);
// 反向对照:两者不能出现对方的措辞
assert.doesNotMatch(fromHuman, /不会替你回信/);
assert.doesNotMatch(fromAgent, /回信不用你自己发/);
});
test('★ from_human 缺失时按「不是人」处理(宁多一次 send_mail,不许诺空头回信)', () => {
const p = buildMailPrompt({ agentName: 'z', data: mail({ from_human: undefined }) });
assert.match(p, /不会替你回信/);
});
test('补投的邮件标注 catchup(模型该知道这不是刚发生的)', () => {
const p = buildMailPrompt({ agentName: 'z', data: mail({ catchup: true }) });
assert.ok(p.length > 0);
// 复用会话时不重复交代身份(省 token,且身份没变过)
const reused = buildMailPrompt({ agentName: 'z', data: mail(), reused: true });
assert.doesNotMatch(reused, /身份:你是/);
});
test('权限结论走单独路径,不写成「新邮件」', () => {
const p = buildMailPrompt({
agentName: 'z',
kind: 'permission',
data: { decision: '同意', decided_by: 'gui-lab' }
});
assert.match(p, /同意/);
assert.match(p, /gui-lab/);
assert.doesNotMatch(p, /read_inbox/);
});
// ─── 回信主题 ─────────────────────────────────────────────────────
test('Re: 前缀不会越滚越长', () => {
assert.equal(replySubject('帮我看看'), 'Re: 帮我看看');
assert.equal(replySubject('Re: 帮我看看'), 'Re: 帮我看看');
assert.equal(replySubject('RE:帮我看看'), 'Re: 帮我看看');
assert.equal(replySubject('回复: 帮我看看'), 'Re: 帮我看看');
});
test('空主题回落到「本轮工作总结」', () => {
for (const s of ['', ' ', undefined, null]) {
assert.equal(replySubject(s), '本轮工作总结');
}
});
// ─── 失败回信 ─────────────────────────────────────────────────────
test('★ 失败回信给出 ZCode 自己的成因,而不是别处的建议', () => {
// 共用库那份 renderFailureReport 的建议是「调整可用模型范围」——
// 对 ZCode 而言那条建议什么也解决不了(它的常见成因是没登录)。
const body = renderTurnFailure([{ kind: 'CLI 失败', error: 'Model config is missing.' }], '帮我看看');
assert.match(body, /Model config is missing/);
assert.match(body, /没有登录/);
assert.match(body, /~\/\.zcode\/cli\/config\.json/);
assert.match(body, /AGENTMAIL_ZCODE_CLI/);
assert.doesNotMatch(body, /调整可用模型范围/);
});
test('失败回信列出每一次尝试', () => {
const body = renderTurnFailure(
[
{ kind: '超时', error: '回合超时' },
{ kind: 'CLI 失败', error: '退出码 7' }
],
's'
);
assert.match(body, /已尝试 2 次/);
assert.match(body, /超时/);
assert.match(body, /退出码 7/);
});
test('失败回信在没有任何尝试记录时也不崩', () => {
const body = renderTurnFailure(undefined, undefined);
assert.match(body, /已尝试 0 次/);
});
// ─── 能力说明(平台把自带危险工具禁掉了,模型必须知道)─────────────────
test('★ workspace 档:说清没有执行面(不带 run_command/write_file 这类不存在的东西)', () => {
const p = buildMailPrompt({ agentName: 'zcode', data: mail({ permission_mode: 'workspace' }) });
assert.match(p, /Bash \/ Write \/ Edit \/ js/, '必须点名哪些自带工具不可用');
assert.match(p, /禁用/);
// ★ 2026-10-02:执行类工具已从 MCP 面移除,**不得**再向模型承诺它们。
// 若这里再出现 run_command / write_file,模型会去找一个不存在的工具,
// 把整轮浪费在换名字重试上 —— 这正是“如实说清能力”的反面。
assert.doesNotMatch(p, /run_command/, '不得承诺不存在的执行工具');
assert.doesNotMatch(p, /write_file/);
assert.match(p, /不能\*\*执行命令/);
// 工具不存在而报错是真实结果,不能靠重试或绕道
assert.match(p, /真实结果/);
assert.match(p, /不要重试/);
assert.match(p, /换名字再试/);
// 兜底路径要给出:需要动手就写进回信请人做,而不是自己硬试
assert.match(p, /需要动手|写进回信/);
// 只读工具要明确可用,否则模型会以为自己什么都干不了
assert.match(p, /Read \/ Glob \/ Grep/);
});
test('★ plan 档:明说不能动手,别浪费一轮去试', () => {
const p = buildMailPrompt({ agentName: 'zcode', data: mail({ permission_mode: 'plan' }) });
assert.match(p, /plan 档/);
assert.match(p, /不能\*\*执行命令/);
assert.doesNotMatch(p, /申请授权/, '根本不会发请求,不该说会去申请');
assert.doesNotMatch(p, /run_command|write_file/, '不得承诺不存在的执行工具');
});
test('★ full 档:即使全权,也明说平台没有执行面', () => {
const p = buildMailPrompt({ agentName: 'zcode', data: mail({ permission_mode: 'full' }) });
assert.match(p, /full 档/);
// ★ 2026-10-02:full 档只是“不问人”,不等于“工具存在”。
// 旧文案说 full 档“直接生效、不会打扰”—— 那会让模型以为能动手,
// 然后花一整轮去找一个已被移除的工具。
assert.match(p, /不能\*\*执行命令/, 'full 档也必须明说没有执行面');
assert.match(p, /只提供邮件能力/);
assert.doesNotMatch(p, /run_command|write_file/);
assert.doesNotMatch(p, /第一次调用会先向发件人申请授权/);
});
test('★ 反向对照:三个档位的说明互不相同(写死一档会让另两档撒谎)', () => {
const texts = ['plan', 'workspace', 'full'].map(t =>
buildMailPrompt({ agentName: 'zcode', data: mail({ permission_mode: t }) })
);
assert.equal(new Set(texts).size, 3, '三个档位给出的能力说明必须各不相同');
});
test('没有 permission_mode 时按 workspace(平台默认档)说明', () => {
const p = buildMailPrompt({ agentName: 'zcode', data: mail() });
assert.match(p, /workspace 档/);
});