Search Ctrl+K
Change language Switch ThemeSign In
Curated Daily BriefWeekly PicksTopics SettingsHelp CenterCollapse
Narrow Mode
SeedRealtime Proactive Interaction: From "Can Converse" to "Can Act"
SeedRealtime Proactive Interaction: From "Can Converse" to "Can Act"
 向阳乔木@vista8
除了音视频联合理解,还有一个点很打动我,就是主动交互。
传统对话模型是被动的,用户问,模型答。
SeedRealtime 引入了持续的视觉感知,让模型能在画面状态发生变化时主动开口。
官方博物馆案例:用户说"看到错金银铜虎噬鹿屏座就提醒我",然后开始逛展。
模型跟随镜头一直在"看"画面,当扫到用户提到的展品时,主动出声提醒。
任务是存在上下文里,目标出现才触发,不是每隔几秒就打扰用户。
另外就是交互节奏和流畅性都比之前版本好很多,接话停顿自然,能分清闲聊和背景噪音,抗干扰强,误触发少。
SeedRealtime下一步重点:更低延迟、更自然节奏。更主动的感知与决策,多人复杂场景稳定性。
个人最期待是:与现实世界打通,把多模态理解变成具体行动,帮用户实现查询、预定与任务办理,从“能对话”到“能行动”。
感觉这个模型会影响具身机器人的发展速度,字节还是有东西!Show More
#### 向阳乔木
@vista8 · 16h ago
昨天豆包的视频通话功能升级了,感觉比 GPT Live 还强!
可以看我录屏,真的超预期。
测试让豆包陪我一起看一个Youtube视频,画面出现个人,直接问这人是谁,竟然答对了!
画面有个一闪而过的图表,问:“刚才有个图表,你能给我解读下吗?”。
它竟然知道我指的是哪个图,并完整解释。
对话延迟低,语音自然,也没有GPT Live的老外腔(国产模型天然优势),很像真人对话。
未来场景相当多啊!K12教育、导览、机器人、实时翻译总结等。
Seed官方公众号说这个技术叫:原生音视频全双工大模型 SeedRealtime。
豆包 App 已全量上线,更新到最新版,点"打电话"图标,开摄像头就能用,不用申请内测。
下面技术解读和另一个测试。Show More
04:37
15
8
48
17.6K
Aug 6, 2026, 8:15 AM View on X
22 Replies
0 Retweets
4 Likes
4,011 Views  向阳乔木 @vista8
Follow
One Sentence Summary
The author analyzes SeedRealtime's proactive interaction capabilities, envisioning its impact on embodied AI and real-world task execution.
Summary
The author highlights SeedRealtime's ability to initiate interaction based on continuous visual perception (e.g., reminding a user when a specific museum exhibit appears). The model distinguishes between casual chat and background noise with high stability. The author looks forward to the model moving from "conversation" to "action," integrating with the real world for queries and bookings, which will accelerate the development of embodied robots.
AI Screening
82
Influence Score 15
Published Yesterday
Language
Chinese
Tags
SeedRealtime
Proactive Interaction
Embodied AI
Multimodal
ByteDance
Make your daily reading actually fit you.A daily brief built from the sources you follow. Get started free HomeDiscoverSettings