fix(llm): 参数无法解析时给出真因,不再静默丢弃整条调用

★ 上次修复误判了成因。真实根因(本次运行日志 34/34 同形):
    {"command": "…完好的长命令…", "timeout": 20s}
  command 一字节没错,只是 timeout 值少了引号 —— cmd_run 的 schema 把 timeout
  声明成 string、示例写着 "10s, 1m, 30s",模型照抄格式却忘了引号。
  finish_reason=length 出现 0 次 ⇒ 上次那条"截断"分支从不生效。

旧行为把**整个参数**丢掉,模型只看到 "command is required",看不出坏在 timeout,
只能原样重试。实测本次运行 cmd_run 失败率 35%(34 败 / 71 成),
12 分钟的任务里更是 48% 时间耗在这上面 —— 每次失败都付一次完整 LLM 往返。

三处改动:
1. repairToolArgsJSON:解析失败时先试窄修复 —— 只给"值位置上未加引号的带单位
   数字"补引号,且修完必须真能解析成功才接受。不碰合法 JSON、不动正文里的 20s、
   不会把真截断"修好"。
2. 修复仍失败时不再静默降级成空 map,改为带 __arg_error 交给模型,并按成因
   分流文案:截断→拆小参数;JSON 写坏→提醒带单位的值要加引号。
3. 统一键名 __arg_error(原 __truncated_error 只覆盖截断,语义过窄)。

同一缺陷面不止 cmd:agentcli/healthcheck/timer 都有 string 类型却以
"5m, 1h" 作示例的参数,此修复一并覆盖。

回归测试:真实日志样本修复、保守性(不碰合法/正文/截断)、
端到端(修复后 timeout 仍能被 time.ParseDuration 接受)。
This commit is contained in:
JianFeeeee
2026-09-19 16:49:16 +08:00
parent 6dca498bc8
commit 499a4f3aa1
7 changed files with 85 additions and 31 deletions

View File

@ -218,6 +218,19 @@ func truncatedArgsError(name string, rawLen, maxTokens int) string {
name, rawLen, maxTokens)
}
// malformedArgsError 把「参数 JSON 写坏了」变成模型能自己改对的一句话。
//
// 实测最常见的一种:带单位的值忘了加引号 —— `{"command": "ls", "timeout": 20s}`。
// 工具 schema 把这类参数声明为 string、示例又写成 “10s, 1m, 30s”,模型容易照抄格式。
// 旧实现丢整条参数,模型只看到 “command is required”,永远不知道坏在 timeout。
func malformedArgsError(name, raw string) string {
return fmt.Sprintf(
"工具 %s 的参数不是合法 JSON,本次调用未执行(这是参数格式问题,不是工具故障)。"+
"请重新生成完整参数并注意:所有字符串值必须带引号 —— 特别是超时/时长这种"+
"带单位的值,要写成 \"20s\" 而不是 20s。收到 %d 字节,开头是:%s",
name, len(raw), truncateStr(raw, 160))
}
// accumulateStream 消费 chunk channel,累积为完整 CompletionResponse,
// 同时发布增量事件。返回的 response 与非流式 Chat() 的返回等价。
//
@ -248,13 +261,22 @@ func accumulateStream(ctx context.Context, ch <-chan agentAPI.StreamChunk, a *Ag
} else if raw == "" {
log.Printf("[agent] stream tool_call %s (idx=%d) received NO argument fragments", acc.name, idx)
}
// 被 MaxTokens 截断时**不要**静默降级成空参数:那会让工具报
// "path is required" 这类与真因无关的错,模型据此重试只会再撞一次。
// 改成把「参数不完整」原样交给模型,并附上可执行的收缩指引。
if !argsOK && lastFinish == "length" {
log.Printf("[agent] stream tool_call %s (idx=%d) TRUNCATED by max_tokens=%d (%d bytes of args) — surfacing to model",
acc.name, idx, maxTokens, len(raw))
args = map[string]interface{}{"__truncated_error": truncatedArgsError(acc.name, len(raw), maxTokens)}
// 参数没法解析时**绝不能**静默降级成空 map:工具只能报 “xxx is required”,
// 那与真因(参数 JSON 写坏了)毫无关系,模型据此重试只会再撞一次
// (实测 2026-09-19:cmd_run 失败 34 次、某任务 48% 时间耗在这上面)。
// 分两种成因给出可执行的指引:
// - finish_reason=length → 被输出上限截断,需拆小参数
// - 其他 → JSON 写坏了(常见:带单位的值忘了引号)
if !argsOK {
if lastFinish == "length" {
log.Printf("[agent] stream tool_call %s (idx=%d) TRUNCATED by max_tokens=%d (%d bytes of args) — surfacing to model",
acc.name, idx, maxTokens, len(raw))
args = map[string]interface{}{"__arg_error": truncatedArgsError(acc.name, len(raw), maxTokens)}
} else {
log.Printf("[agent] stream tool_call %s (idx=%d) args unparseable (%d bytes) — surfacing to model instead of calling with empty args",
acc.name, idx, len(raw))
args = map[string]interface{}{"__arg_error": malformedArgsError(acc.name, raw)}
}
}
tc := agentAPI.ToolCall{
ID: acc.id,