feat(memory): 场景双通道——主动声明与被动涌现并存,且互不吞噬

按「声明式的也要支持,相当于主动被动两条路」落实。此前两者只是恰好并存,
没有边界,实测会互相吃掉(下面的坑就是)。

- Triple.Scenes []string(多值):一轮写下的记忆**两条路都挂**。
  只挂一条会丢东西——只挂声明则细粒度唤起丢失,只挂涌现则首次交互
  (场景还没长出来)没有兜底。单值 Scene 保留兼容。
- TurnScene:Primary 用于写(优先涌现场景,首次退到声明场景兜底),
  Keys 是两条路的并集,用于召回(声明+涌动的场景一起进 RecallByScene)。
- EnterSceneWithHint:主动路 EnsureScene(声明即建场景,不等第二次),
  被动路 EnterScene(指纹聚类)。写侧由 executeToolCall 把本轮场景集合
  传给 memory_commit,模型不需要知道"场景"这回事。

踩到并修掉的坑(两条路互相吞噬):
  最初让声明场景也吸收**整轮指纹**,于是 chan:qq 的相似度永远是 1.0,
  把后续所有同类轮次全部吃掉 → 被动路再也长不出更细的场面,
  实测 turn2.Emergent=true 但 Primary 仍是 chan:qq、没有 auto: 场景。
  修法:给场景加 origin(declared/emergent):
  - 被动聚类只认 origin='emergent' 的场景(声明场景不进相似度空间);
  - 声明场景的特征**只从键自身解析**(chan:qq/peer:group_1 → {chan:qq, peer:group_1}),
    白名单 kind(chan/peer/peer_group/tool/topic/part),不猜——
    「老大2026-09-04_12:27_qq私聊图片」里的 12:27 也是 kind:value 形态,
    放进特征空间就是往相似度里灌垃圾(有测试钉住)。
  - 声明路的泛化靠**层级键前缀**(chan:qq 覆盖 chan:qq/peer:x),机制各归各。
- memgc -scene-stats 增加 [declared|emergent] 与 strength/features 两栏,
  可直接观察两条路各自在长什么。

新增/改写用例:
- TestDeclaredAndEmergentBothLearn:首次交互兜底到声明场景 → 第 2 轮长出
  细粒度涌现场景且**优先用于写入** → 声明场景不进相似度空间(防止压死被动路)
  但仍走声明键取回 → 两条路都进召回集合 → 声明键特征解析与白名单。
- TestEffectiveScenes:多值+单值合并去重保序。

go build/vet 干净,go test -count=1 ./... 全绿。
This commit is contained in:
JianFeeeee
2026-09-15 09:37:29 +08:00
parent bb7e7979ae
commit 49695c38f3
12 changed files with 429 additions and 74 deletions

View File

@ -13,7 +13,7 @@ import (
"gitcode.com/JianFeeeee/HomeAgent/internal/memory/text"
)
func (a *Agent) executeToolCall(tc agentAPI.ToolCall, channel string, turnScene ...string) (ret string) {
func (a *Agent) executeToolCall(tc agentAPI.ToolCall, channel string, turnScenes ...string) (ret string) {
defer func() {
if r := recover(); r != nil {
stack := debug.Stack()
@ -31,7 +31,7 @@ func (a *Agent) executeToolCall(tc agentAPI.ToolCall, channel string, turnScene
done := make(chan string, 1)
go func() {
done <- a.executeToolCallInner(tc, channel, firstOr(turnScene))
done <- a.executeToolCallInner(tc, channel, turnScenes)
}()
select {
@ -43,21 +43,12 @@ func (a *Agent) executeToolCall(tc agentAPI.ToolCall, channel string, turnScene
}
}
// firstOr 取可选参数的首个值(工具执行路径只有调用方知道本轮场景,
// 用变参是为了不让「不关心场景」的调用点(spawn/测试)被迫传空串)。
func firstOr(v []string) string {
if len(v) == 0 {
return ""
}
return v[0]
}
func (a *Agent) executeToolCallInner(tc agentAPI.ToolCall, channel string, scene string) string {
func (a *Agent) executeToolCallInner(tc agentAPI.ToolCall, channel string, turnScenes []string) string {
switch {
case tc.Name == "persona_set":
return a.executePersonaTool(tc)
case strings.HasPrefix(tc.Name, "memory_"):
return a.executeMemoryTool(tc, scene)
return a.executeMemoryTool(tc, turnScenes)
case strings.HasPrefix(tc.Name, "social_"):
return a.executeSocialTool(tc)
case strings.HasPrefix(tc.Name, "knowledge_"):
@ -132,7 +123,7 @@ func (a *Agent) executeToolCallInner(tc agentAPI.ToolCall, channel string, scene
return fmt.Sprintf("%v", result)
}
func (a *Agent) executeMemoryTool(tc agentAPI.ToolCall, turnScene string) string {
func (a *Agent) executeMemoryTool(tc agentAPI.ToolCall, turnScenes []string) string {
g := a.graphMem()
if g == nil {
if tc.Name == "memory_document_query" {
@ -223,10 +214,14 @@ func (a *Agent) executeMemoryTool(tc agentAPI.ToolCall, turnScene string) string
// 两条路都为空则这条记忆不参与场景召回——不做猜测:猜错的场景会把
// 无关记忆钉死,之后每次进入该场面都会被注入,比漏标更难发现。
batchScene := getString(tc.Arguments, "scene")
// 没有显式声明时,落到本轮**涌现**出来的场景上:模型不需要知道场景
// 这回事,记忆也会因为「是在什么场面里写下的」而自动获得唤起入口。
// 写侧的场景是**两条路都挂**:
// 显式声明(模型在参数里点名)优先;
// 否则挂本轮解析出的场景集合——主动声明的 + 被动涌现的。
// 只挂一条会丢东西:只挂声明则细粒度唤起丢失,只挂涌现则首次交互
// (场景还没长出来)没有兜底。
var batchScenes []string
if batchScene == "" {
batchScene = turnScene
batchScenes = turnScenes
}
var triples []memory.Triple
for _, td := range triplesData {
@ -241,6 +236,9 @@ func (a *Agent) executeMemoryTool(tc agentAPI.ToolCall, turnScene string) string
if t.Scene == "" {
t.Scene = batchScene
}
if len(t.Scenes) == 0 {
t.Scenes = batchScenes
}
// 模型显式关联的媒体:结构化字段随三元组一起提交,
// 由 commitTriplesWithMedia 变成 L3 一等块并与句子建边——
// 不再把 marker 写进句子文本。