## 目的
`mcp/server.mjs` 此前注释与行为都绑定 ZCode,接入端必须为 AgentMail 写
专用插件。去掉这层绑定后,任何支持 MCP 的宿主挂一行配置即可用:
{"command":"node","args":["…/mcp/server.mjs"],"env":{
"AGENTMAIL_GATEWAY_URL":…,"AGENTMAIL_AGENT_NAME":…,
"AGENTMAIL_AGENT_SECRET":…,"AGENTMAIL_MCP_PLATFORM":"my-host"}}
协议层(零依赖手写 stdio JSON-RPC)与 11 个邮件工具本就与宿主无关,
真正要动的只有 4 处耦合 + 工具面。
## 改动
**1. 移除执行类工具(`run_command` / `write_file`)**
它们的门禁(lib/action-tools.mjs + lib/approval.mjs + 落盘授权表)是为
ZCode headless 的**双进程审批**设计的:MCP 进程问人、ZCode 钩子进程等回答、
中间靠文件对齐。脱离该宿主后这套门禁的前提不成立,挂在通用服务上等于
提供一条**没有审批的旁路**。
`lib/` 里三个模块与 `hooks/` 源码保留(桌面模式的 ZCode 仍走它们),
只是 server.mjs 不再装载。
**2. platform 可配置**:`AGENTMAIL_MCP_PLATFORM`,默认 `mcp`,
空白值回落默认值。原先硬编码 `'zcode'`(两处)。
**3. 错误文案去宿主名**:不再让模型/人「去 ZCode 的插件设置里填写」,
改为说明设置 `AGENTMAIL_*` 环境变量。
**4. 提示词如实说能力**(src/prompt.mjs):原文案向模型承诺
`run_command`/`write_file` 可用并分档描述「会被请示 / 直接生效」。
工具移除后那变成**指向不存在工具的承诺** —— 模型会去找、把整轮浪费在
换名字重试上。改为明说「本平台没有执行面,需要动手就写进回信请人做」。
三档措辞仍互不相同(`plan`/`workspace`/`full`),因为「档位仍存在但都无
执行面」这件事模型需要知道。
## ★★ 顺带修掉一个真实缺陷(端到端撞出来的)
`connect_to_server` 对 secret-only 的 Agent **一直 400**:
`/agent/register` 只认 `Authorization: Bearer` 或 body 里的 `secret`,
不认 `X-Agent-Secret` 头(其它接口才认),而它漏了 `body.secret`。
dsh / pi 正是 secret-only 配置 ⇒ 它们调「连一下服务器」必然失败,
且模型看不出该改什么。
lib/gateway.mjs 的 `register()` 本来就做对了,tools.mjs 里是手抄的劣化副本。
修后实测 `HTTP 400` → `已连接 …(状态:registered)`。
## 判据
新增 `test/generic-mcp.test.mjs`(5 格)。**这三件事此前无人看守**:
变异验证时「把 action-tools 挂回 server.mjs」与「platform 硬编码回 zcode」
都能全套通过 —— 因为没有判据看 server.mjs 实际挂了什么、也没人看 platform。
改写的 4 格(prompt 3 格 + driver 1 格)保留原意图(不向模型撒谎、
native 自报要有真凭据、工具不存在时不要重试),改为断言新事实。
**变异验证**(每条都确认已应用后才数红格):
挂回 action-tools → 红 3
platform 硬编码 zcode → 红 3
platform 空白不回落 → 红 3
文案指回 ZCode 插件设置 → 红 3
删掉 body.secret(400 复现) → 红 3
全套 **402/402**。
## 端到端验收
写了一个**非 ZCode 宿主**探针(纯 stdio JSON-RPC,不加载任何插件),
对着真实网关跑通:initialize → tools/list(11 个,无执行类)→
connect_to_server(registered)→ suggest_address。
## 未做
- 未发布到 npm registry(`npx` 即用需要发布或指向仓库路径)。
- 未改 `check-deploy-drift.mjs` 的 zcode 豁免(本机仍不退场该宿主)。
414 lines
16 KiB
JavaScript
414 lines
16 KiB
JavaScript
/**
|
||
* 驱动的整条流水线测试(不需要模型、不需要 ZCode)。
|
||
*
|
||
* 判据集中在几件**错了就会静默出错**的事上:
|
||
*
|
||
* - 说错「会不会替你回信」→ 要么发件人等一封永远不来的信,要么收到两封重复邮件
|
||
* - 忘了带 `--mode` → 授权询问全部消失(见 turn-mode 的测试)
|
||
* - 一轮跑不起来却不回信 → 发件人只看到「信发出去了,然后再无音讯」
|
||
* - 去重失效 → 同一封信被处理两遍
|
||
*/
|
||
|
||
import { test } from 'node:test';
|
||
import assert from 'node:assert/strict';
|
||
import { mkdtemp, rm, readFile } from 'node:fs/promises';
|
||
import { tmpdir } from 'node:os';
|
||
import { join } from 'node:path';
|
||
import { createDriver } from '../src/index.mjs';
|
||
import { explicitSendsFile, noteExplicitSendFile } from '../lib/explicit-sends.mjs';
|
||
|
||
/** 假网关客户端:记下发出去的每一封信。 */
|
||
function fakeClient() {
|
||
const sent = [];
|
||
return {
|
||
sent,
|
||
baseURL: 'http://fake',
|
||
agentName: 'zcode',
|
||
authHeaders: () => ({ 'X-Agent-Name': 'zcode' }),
|
||
checkConfig: () => [],
|
||
async post(path, body) {
|
||
if (path === '/mail/send') {
|
||
sent.push(body);
|
||
return { mail_id: `out-${sent.length}` };
|
||
}
|
||
return {};
|
||
},
|
||
async get() {
|
||
return {};
|
||
}
|
||
};
|
||
}
|
||
|
||
/** 一封来自人的来信。 */
|
||
const humanMail = (over = {}) => ({
|
||
mail_id: 'in-1',
|
||
session_id: 'sess-1',
|
||
role: 'to',
|
||
from_human: true,
|
||
from_name: 'gui-lab',
|
||
subject: '帮我看看日志',
|
||
permission_mode: 'workspace',
|
||
to_workspace: '',
|
||
reply_address: 'gui-lab@/work.sess-1',
|
||
...over
|
||
});
|
||
|
||
async function harness({ turn = { sessionId: 'sess_z1', response: '结论:是磁盘满了', exitCode: 0 }, env } = {}) {
|
||
const dir = await mkdtemp(join(tmpdir(), 'zc-driver-'));
|
||
const client = fakeClient();
|
||
const logs = [];
|
||
const calls = [];
|
||
const driver = createDriver({
|
||
client,
|
||
logFn: (...a) => logs.push(a.join(' ')),
|
||
env: env || {},
|
||
config: { workspaceRoot: dir, cliPath: '/fake/zcode.cjs', turnTimeoutMs: 60_000 },
|
||
runTurnFn: async (opts, deps) => {
|
||
calls.push({ opts, deps });
|
||
// 第三个参数是调用序号:测试里常要「第一次失败、第二次成功」
|
||
return typeof turn === 'function' ? turn(opts, deps, calls.length) : turn;
|
||
}
|
||
});
|
||
return { driver, client, logs, calls, dir, cleanup: () => rm(dir, { recursive: true, force: true }) };
|
||
}
|
||
|
||
test('★ 人来信 + 模型没自己发 → 自动回信', async () => {
|
||
const h = await harness();
|
||
try {
|
||
await h.driver.processMail(humanMail());
|
||
assert.equal(h.client.sent.length, 1, '必须回一封信');
|
||
const mail = h.client.sent[0];
|
||
assert.equal(mail.to, 'gui-lab');
|
||
assert.equal(mail.body, '结论:是磁盘满了');
|
||
assert.equal(mail.reply_to, 'in-1');
|
||
assert.equal(mail.subject, 'Re: 帮我看看日志');
|
||
// 走免配额通道:模型已经把话说完了,驱动只是搬运
|
||
assert.equal(mail.relay, 'summary');
|
||
assert.ok(mail.relay_key, '要有幂等键');
|
||
} finally {
|
||
await h.cleanup();
|
||
}
|
||
});
|
||
|
||
test('★ Agent 来信 → 不自动转发(Agent 间必须自己 send_mail)', async () => {
|
||
// 反向对照:同一封邮件只翻转 from_human,回信行为必须跟着翻转。
|
||
// 不这么做的话,两个 Agent 会互相把对方的「已收到」当成待办,无限客套下去。
|
||
const h = await harness();
|
||
try {
|
||
await h.driver.processMail(humanMail({ from_human: false, from_name: 'pi' }));
|
||
assert.equal(h.client.sent.length, 0, 'Agent 来信不该被自动回信');
|
||
assert.ok(
|
||
h.logs.some(l => /不自动转发/.test(l)),
|
||
`日志里应说明原因:${h.logs.join(' | ')}`
|
||
);
|
||
} finally {
|
||
await h.cleanup();
|
||
}
|
||
});
|
||
|
||
test('★ 模型这一轮自己发过信 → 让位,不重复转发', async () => {
|
||
// 线上实测过后果:收件箱里两封说同一件事的邮件(311 与 342 字节)。
|
||
const env = { AGENTMAIL_ZCODE_SENDS_FILE: join(await mkdtemp(join(tmpdir(), 'zc-sends-')), 'sends.jsonl') };
|
||
noteExplicitSendFile(env.AGENTMAIL_ZCODE_SENDS_FILE, {
|
||
sessionId: 'sess-1',
|
||
to: 'gui-lab',
|
||
replyTo: 'in-1'
|
||
});
|
||
const h = await harness({ env });
|
||
try {
|
||
await h.driver.processMail(humanMail());
|
||
assert.equal(h.client.sent.length, 0, '模型已亲手回过,驱动不该再发一封');
|
||
assert.ok(h.logs.some(l => /跳过自动转发/.test(l)));
|
||
} finally {
|
||
await h.cleanup();
|
||
}
|
||
});
|
||
|
||
test('★ 反向对照:另一条会话的主动发信不该让本会话沉默', async () => {
|
||
// 去重不能按「有人发过信」一刀切,必须按会话配对。
|
||
const dir = await mkdtemp(join(tmpdir(), 'zc-sends-'));
|
||
const env = { AGENTMAIL_ZCODE_SENDS_FILE: join(dir, 'sends.jsonl') };
|
||
noteExplicitSendFile(env.AGENTMAIL_ZCODE_SENDS_FILE, {
|
||
sessionId: 'sess-OTHER',
|
||
to: 'gui-lab',
|
||
replyTo: 'in-1'
|
||
});
|
||
const h = await harness({ env });
|
||
try {
|
||
await h.driver.processMail(humanMail());
|
||
assert.equal(h.client.sent.length, 1, '别的会话发过信不该影响这一封');
|
||
} finally {
|
||
await h.cleanup();
|
||
await rm(dir, { recursive: true, force: true });
|
||
}
|
||
});
|
||
|
||
test('★★ 一轮跑不起来 → 必须回一封失败信', async () => {
|
||
// 邮件驱动的会话没有本地界面:什么都不发等于「信发出去了,然后再无音讯」。
|
||
const h = await harness({
|
||
turn: { sessionId: '', response: '', exitCode: 1, stderrTail: 'Model config is missing.' }
|
||
});
|
||
try {
|
||
await h.driver.processMail(humanMail());
|
||
assert.equal(h.client.sent.length, 1, '失败也必须回信');
|
||
const mail = h.client.sent[0];
|
||
assert.match(mail.subject, /处理失败/);
|
||
assert.match(mail.body, /Model config is missing/);
|
||
// 失败信的正文要给出**这个平台**的成因,而不是别处的建议
|
||
assert.match(mail.body, /没有登录/);
|
||
assert.match(mail.body, /AGENTMAIL_ZCODE_CLI/);
|
||
assert.equal(h.driver.stats.failures, 1);
|
||
} finally {
|
||
await h.cleanup();
|
||
}
|
||
});
|
||
|
||
test('超时也算失败,且原因写明超时', async () => {
|
||
const h = await harness({
|
||
turn: { sessionId: '', response: '', exitCode: -1, timedOut: true }
|
||
});
|
||
try {
|
||
await h.driver.processMail(humanMail());
|
||
assert.match(h.client.sent[0].body, /超时/);
|
||
assert.doesNotMatch(h.client.sent[0].body, /CLI 失败/);
|
||
} finally {
|
||
await h.cleanup();
|
||
}
|
||
});
|
||
|
||
test('退出码 0 但没有最终文本 → 不冒充回信', async () => {
|
||
const h = await harness({ turn: { sessionId: 's', response: ' ', exitCode: 0 } });
|
||
try {
|
||
const r = await h.driver.processMail(humanMail());
|
||
assert.equal(h.client.sent.length, 0);
|
||
assert.equal(r.relayed, false);
|
||
assert.ok(h.logs.some(l => /没有产出最终文本/.test(l)));
|
||
} finally {
|
||
await h.cleanup();
|
||
}
|
||
});
|
||
|
||
// ─── 提示词与档位怎么传下去 ─────────────────────────────────────────
|
||
test('提示词里带上回信地址与邮件 id(模型自己发信时要拼对地址)', async () => {
|
||
const h = await harness();
|
||
try {
|
||
await h.driver.processMail(humanMail());
|
||
const prompt = h.calls[0].opts.prompt;
|
||
assert.match(prompt, /gui-lab@\/work\.sess-1/);
|
||
assert.match(prompt, /in-1/);
|
||
// 人来信:告诉模型插件会替它回信
|
||
assert.match(prompt, /回信不用你自己发/);
|
||
} finally {
|
||
await h.cleanup();
|
||
}
|
||
});
|
||
|
||
test('Agent 来信的提示词必须说清「插件不会替你回信」', async () => {
|
||
const h = await harness();
|
||
try {
|
||
await h.driver.processMail(humanMail({ from_human: false, from_name: 'pi' }));
|
||
const prompt = h.calls[0].opts.prompt;
|
||
assert.match(prompt, /不会替你回信/);
|
||
assert.doesNotMatch(prompt, /回信不用你自己发/);
|
||
} finally {
|
||
await h.cleanup();
|
||
}
|
||
});
|
||
|
||
test('★ 档位随邮件传下去,并作为 --mode / 禁用清单 / 钩子环境变量注入', async () => {
|
||
for (const [tier, mode] of [
|
||
['plan', 'plan'],
|
||
// workspace 与 full 都映射到 yolo:平台不做权限判定(它自带危险工具已被
|
||
// --disallowed-tools 拿掉),执行类动作改由我们自己的门禁逐次请示。
|
||
['workspace', 'yolo'],
|
||
['full', 'yolo']
|
||
]) {
|
||
const h = await harness();
|
||
try {
|
||
await h.driver.processMail(humanMail({ permission_mode: tier }));
|
||
assert.equal(h.calls[0].opts.mode, mode, `档位 ${tier} 应映射到 ${mode}`);
|
||
const env = h.calls[0].opts.env;
|
||
// 钩子靠这两个变量决定档位与「有没有本地界面」
|
||
assert.equal(env.AGENTMAIL_PERMISSION_MODE, tier);
|
||
assert.equal(env.AGENTMAIL_SESSION_ID, 'sess-1');
|
||
// 禁用清单必须真的传下去:它是「平台不问」时唯一的替代防线。
|
||
const denied = h.calls[0].opts.disallowedTools;
|
||
assert.ok(Array.isArray(denied) && denied.length > 20, '禁用清单未传给 ZCode');
|
||
for (const must of ['Bash', 'Write', 'Edit', 'js', 'mcp__node_repl__js']) {
|
||
assert.ok(denied.includes(must), `禁用清单缺少 ${must}`);
|
||
}
|
||
} finally {
|
||
await h.cleanup();
|
||
}
|
||
}
|
||
});
|
||
|
||
test('工作目录取 to_workspace;没有就用兜底目录', async () => {
|
||
const h = await harness();
|
||
try {
|
||
await h.driver.processMail(humanMail());
|
||
assert.ok(h.calls[0].opts.cwd.startsWith(h.dir), `兜底目录应在 ${h.dir} 下,实际 ${h.calls[0].opts.cwd}`);
|
||
|
||
const real = await mkdtemp(join(tmpdir(), 'zc-ws-'));
|
||
await h.driver.processMail(humanMail({ mail_id: 'in-2', to_workspace: real }));
|
||
assert.equal(h.calls[1].opts.cwd, real);
|
||
await rm(real, { recursive: true, force: true });
|
||
} finally {
|
||
await h.cleanup();
|
||
}
|
||
});
|
||
|
||
test('★ 同一会话的第二封信带上 --resume(否则模型每封信都从零开始)', async () => {
|
||
const h = await harness();
|
||
try {
|
||
await h.driver.processMail(humanMail());
|
||
assert.equal(h.calls[0].opts.resumeSessionId, undefined, '首轮不该带 resume');
|
||
await h.driver.processMail(humanMail({ mail_id: 'in-2' }));
|
||
assert.equal(h.calls[1].opts.resumeSessionId, 'sess_z1', '第二轮要续上同一个 ZCode 会话');
|
||
} finally {
|
||
await h.cleanup();
|
||
}
|
||
});
|
||
|
||
// ─── 事件入口 ───────────────────────────────────────────────────────
|
||
test('SSE 事件:重复的 mail_id 只处理一次', async () => {
|
||
const h = await harness();
|
||
try {
|
||
h.driver.handleEvent('new_mail', humanMail());
|
||
h.driver.handleEvent('new_mail', humanMail());
|
||
await new Promise(r => setTimeout(r, 20));
|
||
assert.equal(h.calls.length, 1, '同一封信被处理了两遍');
|
||
} finally {
|
||
await h.cleanup();
|
||
}
|
||
});
|
||
|
||
test('抄送给自己也处理(role=cc),其它角色忽略', async () => {
|
||
const h = await harness();
|
||
try {
|
||
h.driver.handleEvent('new_mail', humanMail({ mail_id: 'cc-1', role: 'cc' }));
|
||
h.driver.handleEvent('new_mail', humanMail({ mail_id: 'x-1', role: 'from' }));
|
||
await new Promise(r => setTimeout(r, 20));
|
||
assert.equal(h.calls.length, 1);
|
||
assert.equal(h.calls[0].opts.prompt.includes('cc-1'), true);
|
||
} finally {
|
||
await h.cleanup();
|
||
}
|
||
});
|
||
|
||
test('非 new_mail 事件被忽略(permission_decision 由钩子自己处理)', async () => {
|
||
const h = await harness();
|
||
try {
|
||
h.driver.handleEvent('permission_decision', { relay_key: 'k' });
|
||
h.driver.handleEvent('session_archived', { session_id: 's' });
|
||
await new Promise(r => setTimeout(r, 20));
|
||
assert.equal(h.calls.length, 0);
|
||
} finally {
|
||
await h.cleanup();
|
||
}
|
||
});
|
||
|
||
test('一封邮件处理崩了不会带走驱动(后面的信照常处理)', async () => {
|
||
const h = await harness({
|
||
turn: (opts, deps, n) => {
|
||
if (n === 1) throw new Error('boom');
|
||
return { sessionId: 's', response: '第二封处理好了', exitCode: 0 };
|
||
}
|
||
});
|
||
try {
|
||
h.driver.handleEvent('new_mail', humanMail());
|
||
h.driver.handleEvent('new_mail', humanMail({ mail_id: 'in-2' }));
|
||
await new Promise(r => setTimeout(r, 50));
|
||
assert.equal(h.client.sent.length, 1);
|
||
assert.match(h.client.sent[0].body, /第二封处理好了/);
|
||
} finally {
|
||
await h.cleanup();
|
||
}
|
||
});
|
||
|
||
test('★ 关停时终止在途回合(不留下跑工具的孤儿)', async () => {
|
||
// systemd 杀掉驱动后,那个 ZCode 进程还在跑工具,而既没有驱动看着它、
|
||
// 也没有本地界面看着它 —— 宁可丢掉这一轮的工作。
|
||
const signals = [];
|
||
let release;
|
||
const h = await harness({
|
||
turn: async (opts, deps) => {
|
||
// 真实现会把「怎么杀」通过 deps.onChild 交出来(见 zcode-run.mjs)
|
||
deps.onChild(sig => signals.push(sig));
|
||
await new Promise(r => (release = r));
|
||
return { sessionId: 's', response: 'x', exitCode: 0 };
|
||
}
|
||
});
|
||
try {
|
||
const p = h.driver.processMail(humanMail());
|
||
await new Promise(r => setTimeout(r, 20));
|
||
assert.equal(typeof release, 'function', '回合应已开始');
|
||
|
||
h.driver.abort();
|
||
assert.deepEqual(signals, ['SIGTERM'], 'abort 必须终止在途回合');
|
||
// 幂等:重复关停不该再杀一次
|
||
h.driver.abort();
|
||
assert.deepEqual(signals, ['SIGTERM']);
|
||
|
||
release();
|
||
await p;
|
||
} finally {
|
||
await h.cleanup();
|
||
}
|
||
});
|
||
|
||
// ─── 主动发信记录的落盘 ─────────────────────────────────────────────
|
||
test('★ explicit-sends 文件位置两端一致(驱动与 MCP 服务器必须解析出同一路径)', async () => {
|
||
const env = { AGENTMAIL_CONFIG_DIR: '/tmp/agentmail-cfg' };
|
||
assert.equal(explicitSendsFile(env), '/tmp/agentmail-cfg/explicit-sends.jsonl');
|
||
assert.equal(explicitSendsFile({ ZCODE_PLUGIN_DATA: '/d' }), '/d/explicit-sends.jsonl');
|
||
assert.match(explicitSendsFile({}), /explicit-sends\.jsonl$/);
|
||
});
|
||
|
||
test('读回的记录形状可直接交给共用去重判据', async () => {
|
||
const dir = await mkdtemp(join(tmpdir(), 'zc-sends-'));
|
||
const f = join(dir, 'sends.jsonl');
|
||
noteExplicitSendFile(f, { sessionId: 's1', to: 'gui-lab@/p', replyTo: 'm1' });
|
||
const { readExplicitSends } = await import('../lib/explicit-sends.mjs');
|
||
const rec = readExplicitSends(f, { sessionId: 's1' });
|
||
assert.ok(rec.names.has('gui-lab'));
|
||
assert.ok(rec.replyTos.has('m1'));
|
||
await rm(dir, { recursive: true, force: true });
|
||
});
|
||
|
||
// ─── 档位强制力自报(必须如实,否则是在替不存在的能力背书)────────────
|
||
|
||
test('★ 自报 native 要有真凭据:门禁链就绪(yolo + 够长的禁用清单)', async () => {
|
||
const { detectModeEnforcement } = await import('../src/index.mjs');
|
||
const r = detectModeEnforcement({ env: {} });
|
||
assert.equal(r.enforcement, 'native');
|
||
// 理由里必须点出**谁**在把关。以前这里写的是「钩子已注册」,而 yolo 下
|
||
// 钩子根本不会触发 —— 那种理由会让人以为平台在管,实际平台什么都没管。
|
||
//
|
||
// ★ 2026-10-02:执行类工具已从 MCP 面移除,所以 native 的凭据变成
|
||
// 「平台自带危险工具已禁用 + AgentMail 侧无执行面 ⇒ 模型无执行路径」。
|
||
// 凭据换了,**意图不变**:native 必须有真凭据,不能只报个标签。
|
||
assert.match(r.reason, /门禁|执行面|无任何执行面|已禁用/);
|
||
assert.match(r.reason, /不再提供执行类工具|无任何执行面/, '必须说清“没有执行面”才是当前凭据');
|
||
assert.doesNotMatch(r.reason, /^钩子已注册/, '不能拿钩子当唯一凭据');
|
||
});
|
||
|
||
test('★ 反向对照:门禁链断了就必须降级成 advisory', async () => {
|
||
const { detectModeEnforcement } = await import('../src/index.mjs');
|
||
// 禁用清单被清空 = 平台自带 Bash/Write/js 全都还回去了 —— 此时即使钩子
|
||
// 清单正常,也不再是「该档位被强制」。
|
||
const r = detectModeEnforcement({ env: { AGENTMAIL_ZCODE_DISALLOWED_TOOLS: '' } });
|
||
assert.equal(r.enforcement, 'advisory');
|
||
});
|
||
|
||
test('★ 只拿得到钩子、拿不到门禁时不能硬报 native', async () => {
|
||
const { detectModeEnforcement } = await import('../src/index.mjs');
|
||
const r = detectModeEnforcement({
|
||
hooksFile: '/nonexistent/hooks.json',
|
||
env: {}
|
||
});
|
||
// 门禁就绪 → 仍然 native(我们拦得住),但理由里不能声称有钩子
|
||
assert.equal(r.enforcement, 'native');
|
||
assert.doesNotMatch(r.reason, /钩子已注册/);
|
||
});
|