mirror of
https://gitcode.com/JianFeeeee/HomeAgent.git
synced 2026-09-21 17:38:10 +00:00
feat(memory): 媒体接入 L0/L2——digest 挂到对话事件,归档时引用随之转移
在 a822674 的 CAS 层之上把媒体真正接进记忆链路。此前 CAS 只是个孤立的
存储包,没有任何写入方。
## 媒体进入对话有两条路,两条都只把文字留给记忆
1. 用户直接发图 → processMediaInput → mediaToBlocks
ContextEvent.Input 只存 alt 文本("[从 qq 收到了 image]"),
base64 随 message 数组发给模型后就丢了。
2. 插件注入 → SetToolBlocks → process.go 的 mediaMsg
ToolResultItem.Output 只存那句 "[已将图片注入后续对话] /tmp/x.png"。
于是下一轮起,模型能看到的只剩一句路径或一句 alt。那个文件被删、被覆盖,
或者本来就是 /tmp 下的临时产物,连线索都断了。
现在两条路在同一处收口(captureBlockMedia):从 ContentBlock 的 data URL
取出字节存进 CAS,digest 挂到当轮 ContextEvent。
## 改动
internal/agent/core/mediaref.go(新)
- captureBlockMedia:ContentBlock → CAS。只处理 data URL——http(s) URL
拿不到字节就无法内容寻址,而「下载它再存」会把一次对话变成一次网络
请求(超时、鉴权、SSRF 全来了),不在本层解决。
- stage/drainMediaDigests:媒体在 process() 期间被捕获,而承载它的
ContextEvent 要等 process() 返回后才 Append——此刻还没有 owner_id,
故先缓存。与既有 pendingMedia 同一手法,同受 a.mu 保护。
- bindEventMedia:双向落地。evt.Media 让事件记得引了什么(随
context.json 持久化),media_refs 让 CAS 知道谁在引用(GC 的判断依据)。
只写一边的话,要么 GC 误删仍被引用的内容,要么孤儿永远清不掉。
- mediaSummaryForEvent:把已有描述拼成一行写进 Input。这是方案 C 的
落点——**描述文本才是持久语义记忆,blob 只是缓存**。blob 可能被容量
GC 淘汰,但描述会一直留在 L0/L2/L3 的文本里,让「那张紫蓝红三色带图」
几个月后仍可被检索。
ContextEvent 新增 ID 与 Media 两个字段,都是 omitempty:
- ID 懒生成,只有真要挂媒体时才赋值。绝大多数对话没有媒体,全量生成
会让每条事件都多一个字段进 context.json。
- 存量 context.json 读回来两字段皆空,不影响任何既有行为(有测试)。
RelevanceContext.Prune 归档时转移引用(transferMediaRefs):
**先挂到归档文档、再注销原事件引用**。顺序不能反——先销后挂会让引用
计数瞬时归零,若此刻后台 GC 正在跑就会把仍被记忆引用的内容当孤儿清掉。
为此把 Prune 内的局部类型 scored 提为包级 scoredEvent(局部类型无法
出现在方法签名上)。
media 包新增 OwnerContext/OwnerDocument/OwnerGraphSentence 常量:
owner_kind 进了主键,拼错一个字符就是一条永远对不上的孤立引用——
AddRef 不报错,DropOwner 也永远匹配不到。
## 配置
core.memory.media.enabled(默认 true)、.dir、.max_mb(2048)、
.gc_interval(6h)、.gc_min_age(1h)。
关闭后全链路静默跳过,对话行为与本特性上线前完全一致(有测试)。
mediaStore 为 nil 时同理——它是记忆增强,不是对话必需品,开不起来
只记一条 warning 不阻止启动。
## 测试(11 例)
入库与 MIME 归类、http URL 跳过、nil store 全链路 no-op、音视频混合、
stage/drain 清空语义、懒生成 ID、描述作为持久记忆、**归档转移期间内容
始终可读且 refcount 不归零**、无媒体存储时归档照常、context.json
向后兼容往返。
全仓 go build / go vet / go test 通过,SDK 冻结 diff = 0。
## 尚未接入
L3 图库的 graph_sentence owner(常量已备好,无写入方)、
描述生成的后台任务(Pending() 已就绪,尚无消费者)、
媒体 GC 的定时触发(配置项已注册,尚未接 ticker)。
This commit is contained in:
@ -174,6 +174,11 @@ func (a *Agent) processMediaInput(evt *agentIO.InputEvent) {
|
||||
|
||||
blocks, fallback := a.mediaToBlocks(evt.Payload, evt.Type, evt.Source)
|
||||
|
||||
// 用户直接发来的媒体:先落进 CAS。
|
||||
// 不存的后果是 ContextEvent.Input 只剩一句 alt 文本
|
||||
//("[从 qq 收到了 image]"),base64 随 message 数组发给模型后就丢了。
|
||||
a.stageMediaDigests(a.captureBlockMedia(blocks, "input_"+evt.Type)...)
|
||||
|
||||
stageCtx := a.stageCtxFromInput(fallback, evt.Source, "")
|
||||
stageCtx.Extra = map[string]interface{}{
|
||||
"media_blocks": blocks,
|
||||
@ -216,14 +221,22 @@ func (a *Agent) processMediaInput(evt *agentIO.InputEvent) {
|
||||
elapsed := time.Since(start)
|
||||
log.Printf("[agent] %s from %s → response (%dms, tools=%v)", evt.Type, evt.Source, elapsed.Milliseconds(), toolsUsed)
|
||||
|
||||
a.context.Append(ContextEvent{
|
||||
// 本轮捕获的媒体(用户发的 + 工具注入的)挂到这条事件上。
|
||||
// 媒体描述并进 Input:描述文本才是持久语义记忆,blob 只是缓存。
|
||||
digests := a.drainMediaDigests()
|
||||
mediaEvt := ContextEvent{
|
||||
Timestamp: time.Now(),
|
||||
Source: "agent",
|
||||
Input: fallback,
|
||||
Response: response,
|
||||
ToolsUsed: toolsUsed,
|
||||
ToolResults: toolResults,
|
||||
})
|
||||
}
|
||||
a.bindEventMedia(&mediaEvt, digests)
|
||||
if s := a.mediaSummaryForEvent(mediaEvt.Media); s != "" {
|
||||
mediaEvt.Input = mediaEvt.Input + "\n" + s
|
||||
}
|
||||
a.context.Append(mediaEvt)
|
||||
|
||||
a.emitResponse(evt, response)
|
||||
|
||||
@ -271,12 +284,12 @@ func (a *Agent) mediaToBlocks(payload map[string]interface{}, mediaType string,
|
||||
}
|
||||
if mediaType == "image" {
|
||||
blocks = append(blocks, agentAPI.ContentBlock{
|
||||
Type: "image_url",
|
||||
Type: "image_url",
|
||||
ImageURL: &agentAPI.ImageURL{URL: imgURL, Detail: "auto"},
|
||||
})
|
||||
} else if mediaType == "audio" {
|
||||
blocks = append(blocks, agentAPI.ContentBlock{
|
||||
Type: "audio_url",
|
||||
Type: "audio_url",
|
||||
AudioURL: &agentAPI.AudioURL{URL: imgURL},
|
||||
})
|
||||
}
|
||||
@ -371,14 +384,21 @@ func (a *Agent) processTextInput(evt *agentIO.InputEvent, input string) {
|
||||
elapsed := time.Since(start)
|
||||
log.Printf("[agent] input from %s → response (%dms, tools=%v)", evt.Source, elapsed.Milliseconds(), toolsUsed)
|
||||
|
||||
a.context.Append(ContextEvent{
|
||||
// 纯文本输入也可能产生媒体:模型调 multimodal_see_picture / see_video 等工具时,
|
||||
// 插件经 SetToolBlocks 注入的块已在 process() 里被捕获。
|
||||
textEvt := ContextEvent{
|
||||
Timestamp: time.Now(),
|
||||
Source: "agent",
|
||||
Input: cleanInput,
|
||||
Response: response,
|
||||
ToolsUsed: toolsUsed,
|
||||
ToolResults: toolResults,
|
||||
})
|
||||
}
|
||||
a.bindEventMedia(&textEvt, a.drainMediaDigests())
|
||||
if s := a.mediaSummaryForEvent(textEvt.Media); s != "" {
|
||||
textEvt.Input = textEvt.Input + "\n" + s
|
||||
}
|
||||
a.context.Append(textEvt)
|
||||
|
||||
a.emitResponse(evt, response)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user