← 回總覽

京东开源实时视频 VLM:JoyAI-VL-Interaction,实现“边看边说”主动交互

📅 2026-06-23 11:11 AIGCLINK 人工智能 4 分鐘 4329 字 評分: 86
实时视频VLM JoyAI-VL-Interaction 京东开源 视觉语言模型 主动交互
📌 一句话摘要 京东开源实时视频视觉语言模型 JoyAI-VL-Interaction,可持续观察视频流、自主判断何时说话,在 58 个真人盲评中对豆包和 Gemini 分别取得 77.6%和 87.9%胜率。 📝 详细摘要 推文详细介绍了京东开源的实时视频视觉语言交互模型 JoyAI-VL-Interaction,该模型具备持续观察视频流、自主判断何时发言/沉默的能力,实现“边看边说”的类人在场交互。在 58 个真人盲评的实时流式场景中,对豆包胜率 77.6%、对 Gemini 胜率 87.9%,监控预警场景对两者均 100%胜率。模型采用前台+后台分工设计:前台小模型实时观察视频流,后
Skip to main contentAudio 2 ![Image 1: LogoBest Blogs](https://www.bestblogs.dev/ "BestBlogs.dev")

Search Ctrl+K

Change language Switch ThemeSign In

Curated Daily BriefWeekly PicksTopicsWorld Cup Special SettingsHelp CenterCollapse

Narrow Mode

JD.com Open-Sources Real-Time Video VLM: JoyAI-VL-Interaction, Enabling 'See and Speak' Proactive Interaction

JD.com Open-Sources Real-Time Video VLM: JoyAI-VL-Interaction, Enabling 'See and Speak' Proactive Interaction

![Image 2: AIGCLINK](https://www.bestblogs.dev/en/tweets?sourceId=SOURCE_fa9efd59) AIGCLINK

@aigclink

京东最新开源的实时视频视觉语言交互模型:JoyAI-VL-Interaction,让大模型从“一问一答”走向“边看边说”

也就是说它会像人一样“在场”,持续观察视频流,自主判断什么时候说话、什么时候沉默,并实时响应关键事件

在58个真人盲评的实时流式场景中,对豆包胜率77.6%、对Gemini胜率 87.9%,监控预警场景对两个基线均100%胜率

可以用来搭建需要持续观察、自主判断、即时响应的实景AI,比如说安防监控、老人儿童看护、直播讲解、电商导购、操作指导、AI眼镜等等

核心特点是主动判断非被动回答,它会持续观察视频流,来自主做判断,不是等用户发起问题才开始处理当前画面

比如说,当设置"裁判出示红牌时提醒我",模型就会持续值守画面,在事件发生时自动预警

第二个,它会面向正在发生的视频流即时响应,画面变化时即能响应

前台+后台的分工协作设计,前台模型实时观察视频流,后台大模型/Agent接复杂推理、代码生成、工具调用的重活

后台结果返回后前台自然接回对话,形成前台实时助手+后台智能大脑的协作系统,端侧用小模型持续值守,复杂任务才调用大模型,使得成本和延迟更可控

模型+数据+训练方案+可部署系统全栈开源,各模块可替换,拿去即能用

#JoyAIVLInteraction #VLM Show More

!Image 3: Video thumbnail

01:04

Jun 23, 2026, 3:11 AM View on X

19 Replies

7 Retweets

41 Likes

7,613 Views ![Image 4: AIGCLINK](https://www.bestblogs.dev/en/tweets?sourceid=fa9efd59) AIGCLINK @aigclink

Follow

One Sentence Summary

JD.com open-sources the real-time video vision-language model JoyAI-VL-Interaction, which continuously observes video streams and autonomously decides when to speak, achieving 77.6% and 87.9% win rates against Doubao and Gemini respectively in 58 human blind evaluations.

Summary

This tweet details the open-source real-time video vision-language interaction model JoyAI-VL-Interaction from JD.com. The model can continuously observe video streams and autonomously decide when to speak or remain silent, achieving a 'see and speak' human-like presence. In 58 human blind evaluations of real-time streaming scenarios, it achieved a 77.6% win rate against Doubao and 87.9% against Gemini, and a 100% win rate against both baselines in monitoring and alerting scenarios. The model uses a frontend + backend design: a lightweight frontend model continuously observes the video stream, while a backend large model/Agent handles complex reasoning, code generation, and other heavy tasks, keeping costs and latency manageable. Applicable scenarios include security monitoring, elderly and child care, live streaming commentary, e-commerce shopping assistance, operational guidance, and AI glasses. All technology (model, data, training scheme, deployment system) is fully open-sourced, with replaceable modules.

AI Screening

86

Influence Score 33

Published Yesterday

Language

Chinese

Tags

Real-Time Video VLM

JoyAI-VL-Interaction

JD.com Open Source

Vision Language Model

Proactive Interaction

Make your daily reading actually fit you.A daily brief built from the sources you follow. Get started free HomeDiscoverWorld CupSettings

查看原文 → 發佈: 2026-06-23 11:11:19 收錄: 2026-06-23 22:00:39

🤖 問 AI

針對這篇文章提問,AI 會根據文章內容回答。按 Ctrl+Enter 送出。