Search Ctrl+K
Change language Switch ThemeSign In
Curated Daily BriefWeekly PicksTopicsWorld Cup Special SettingsHelp CenterCollapse
Narrow Mode
JD.com Open-Sources Real-Time Video VLM: JoyAI-VL-Interaction, Enabling 'See and Speak' Proactive Interaction
JD.com Open-Sources Real-Time Video VLM: JoyAI-VL-Interaction, Enabling 'See and Speak' Proactive Interaction
 AIGCLINK@aigclink
京东最新开源的实时视频视觉语言交互模型:JoyAI-VL-Interaction,让大模型从“一问一答”走向“边看边说”
也就是说它会像人一样“在场”,持续观察视频流,自主判断什么时候说话、什么时候沉默,并实时响应关键事件
在58个真人盲评的实时流式场景中,对豆包胜率77.6%、对Gemini胜率 87.9%,监控预警场景对两个基线均100%胜率
可以用来搭建需要持续观察、自主判断、即时响应的实景AI,比如说安防监控、老人儿童看护、直播讲解、电商导购、操作指导、AI眼镜等等
核心特点是主动判断非被动回答,它会持续观察视频流,来自主做判断,不是等用户发起问题才开始处理当前画面
比如说,当设置"裁判出示红牌时提醒我",模型就会持续值守画面,在事件发生时自动预警
第二个,它会面向正在发生的视频流即时响应,画面变化时即能响应
前台+后台的分工协作设计,前台模型实时观察视频流,后台大模型/Agent接复杂推理、代码生成、工具调用的重活
后台结果返回后前台自然接回对话,形成前台实时助手+后台智能大脑的协作系统,端侧用小模型持续值守,复杂任务才调用大模型,使得成本和延迟更可控
模型+数据+训练方案+可部署系统全栈开源,各模块可替换,拿去即能用
#JoyAIVLInteraction #VLM Show More
01:04
Jun 23, 2026, 3:11 AM View on X
19 Replies
7 Retweets
41 Likes
7,613 Views  AIGCLINK @aigclink
Follow
One Sentence Summary
JD.com open-sources the real-time video vision-language model JoyAI-VL-Interaction, which continuously observes video streams and autonomously decides when to speak, achieving 77.6% and 87.9% win rates against Doubao and Gemini respectively in 58 human blind evaluations.
Summary
This tweet details the open-source real-time video vision-language interaction model JoyAI-VL-Interaction from JD.com. The model can continuously observe video streams and autonomously decide when to speak or remain silent, achieving a 'see and speak' human-like presence. In 58 human blind evaluations of real-time streaming scenarios, it achieved a 77.6% win rate against Doubao and 87.9% against Gemini, and a 100% win rate against both baselines in monitoring and alerting scenarios. The model uses a frontend + backend design: a lightweight frontend model continuously observes the video stream, while a backend large model/Agent handles complex reasoning, code generation, and other heavy tasks, keeping costs and latency manageable. Applicable scenarios include security monitoring, elderly and child care, live streaming commentary, e-commerce shopping assistance, operational guidance, and AI glasses. All technology (model, data, training scheme, deployment system) is fully open-sourced, with replaceable modules.
AI Screening
86
Influence Score 33
Published Yesterday
Language
Chinese
Tags
Real-Time Video VLM
JoyAI-VL-Interaction
JD.com Open Source
Vision Language Model
Proactive Interaction
Make your daily reading actually fit you.A daily brief built from the sources you follow. Get started free HomeDiscoverWorld CupSettings