Title: Zhejiang University open-sources LightMem-Ego: the first ...
URL Source: https://www.bestblogs.dev/status/2079059418168311813?amp%3Butm_medium=feed&%3Butm_campaign=resources&%3Bentry=rss_article_item
Published Time: 2026-07-20 12:22:47
Markdown Content: Skip to main contentAudio 2 
Search Ctrl+K
Change language Switch ThemeSign In
Curated Daily BriefWeekly PicksTopicsWorld Cup Special SettingsHelp CenterCollapse
Narrow Mode
Zhejiang University open-sources LightMem-Ego: the first multimodal egocentric memory system for smart glasses
Zhejiang University open-sources LightMem-Ego: the first multimodal egocentric memory system for smart glasses
 AIGCLINK@aigclink
"我钥匙刚才放哪了""上午那人跟我说的会议改几点"——你戴着智能AI 眼镜过了一整天,它能不能替你回答这些?浙大 NLP 团队开源了 LightMem-Ego,专门干这件事:把你第一视角看到、听到的一切流式存成"能被问答的记忆",可与Rokid 眼镜结合。
LightMem-Ego(浙大知识图谱组 zjunlp,Ningyu Zhang 团队):一套端到端的第一视角(egocentric)记忆系统:接一副 Rokid AI 眼镜的安卓 App + 浏览器前端 + 在线后端,把眼镜的摄像头画面和麦克风音频流式送上去,边过日子边建记忆,然后问它当下或过去的事。
它的核心设计,是把连续的视听体验分成三层记忆(按时间维度):
~当前记忆:正在发生的场景理解;
~短期记忆:最近的事件、动作、对话;
~长期记忆:沉淀下来的情节、习惯、偏好、语义事实。
回答问题时,它先按问题动态路由到对应的记忆层、检索出带时间戳的多模态证据,再据此生成答案——不是凭空答,是"翻到那一刻的画面和录音"再说。设计针对的场景很具体:找东西、回忆对话、总结一天、发现日常规律、免手操作的可穿戴助手。
放进记忆类赛道看:Agent 的"记忆"这半年一直是热点,但多数是给文本对话 agent 做长期记忆。LightMem-Ego 换了个更难的场景——多模态、流式、还要跑在算力受限的可穿戴设备上。它对标的不是"聊天机器人记住你说过啥",是"你的眼睛耳朵替你记了一整天,事后能被检索"。这也是智能眼镜这波硬件想要、但一直缺的那层软件。
Rokid 眼镜是现成的落地硬件载体,可落地的方向很清楚:记忆层做成可私有部署的 SDK,个性化 AI 眼镜/可穿戴的"记忆中间件,从而用于特定场景。
#LightMemEgo #AI记忆 #智能眼镜 #多模态 #浙大 Show More
Jul 20, 2026, 4:22 AM View on X
7 Replies
0 Retweets
10 Likes
1,867 Views  AIGCLINK @aigclink
Follow
One Sentence Summary
The Zhejiang University NLP team open-sources LightMem-Ego, building a queryable three-layer egocentric memory from streaming audiovisual data, which can be integrated with Rokid glasses.
Summary
The Zhejiang University Knowledge Graph team (Ningyu Zhang's team) has released the open-source project LightMem-Ego, an end-to-end egocentric memory system that can connect to the Rokid AI glasses Android app, web frontend, and online backend, streaming the glasses' camera frames and microphone audio to build memory on the fly. The system divides continuous audiovisual experience into three memory layers: current memory (understanding the ongoing scene), short-term memory (recent events, actions, and dialogues), and long-term memory (consolidated episodes, habits, preferences, and semantic facts). When answering a question, it dynamically routes the query to the appropriate memory layer, retrieves timestamped multimodal evidence, and generates an answer based on that evidence—effectively “flipping back to that moment’s video and audio” as a searchable memory. Use cases include finding objects, recalling conversations, summarizing the day, discovering daily routines, and providing a hands-free wearable assistant. This work fills the missing software memory layer for smart glasses hardware, proposing a privately deployable SDK as a “memory middleware” for personalized AI glasses.
AI Screening
77
Influence Score 10
Published Today
Language
Chinese
Tags
LightMem-Ego
AI Memory
Smart Glasses
Multimodal
Zhejiang University
Make your daily reading actually fit you.A daily brief built from the sources you follow. Get started free HomeDiscoverWorld CupSettings