Search Ctrl+K
Change language Switch ThemeSign In
Curated Daily BriefWeekly PicksTopics SettingsHelp CenterCollapse
Narrow Mode
OpenMOSS Open-Sources MOSS-VL-Realtime: An 11B Multimodal Model for Real-Time Video Streaming Interaction
OpenMOSS Open-Sources MOSS-VL-Realtime: An 11B Multimodal Model for Real-Time Video Streaming Interaction
 Berryxia.AI@berryxia
兄弟们,开源大小模型多端开花啊!
OpenMOSS 今天正式开源了 MOSS-VL-Realtime,这是一个专注于实时视频流交互的 11B 多模态模型。
不同于传统“先看完视频再回答”的离线模式,MOSS-VL-Realtime 支持在视频持续输入的过程中同步进行感知和生成。
它可以一边处理新帧,一边生成回复,并在场景变化时主动修改或打断自己的回答,也能在信息不足时保持沉默,真正实现“边看边聊”的实时交互体验。
模型亮点包括:
- 采用 Cross-Attention 架构,将视觉编码与语言推理分离
- XRoPE 实现统一的时空位置编码
- 支持文本、单图/多图、单视频/多视频以及图文交错输入(中英双语)
- 256K 超长上下文窗口
- 统一的对话模板,可同时支持离线、流式和实时交互模式
同时,新发布的 MOSS-VL-0708 Instruct 在细粒度感知、时间动作定位以及长视频理解等任务上也有明显提升。
模型已全部开源(Apache-2.0 协议),支持本地部署和实时推理。
Hugging Face:huggingface.co/OpenMOSS-Team/…
GitHub:github.com/OpenMOSS/MOSS-…
技术博客:github.com/OpenMOSS/MOSS-…(内含详细说明)
这可能是目前开源社区在实时多模态交互方向上的一次重要推进,尤其适合需要持续视频理解和动态对话的场景。Show More
#### OpenMOSS
@Open_MOSS · 16h ago
🤗 MOSS-VL-Realtime is now open source on @huggingface .
The 11B model family supports text, single and multiple images, single and multiple videos, and interleaved visual-text inputs in Chinese and English.@MosiAI_Official
Highlights:
🏗️ Cross-Attention architecture separating visual encoding from language reasoning
🧭 XRoPE for unified temporal-spatial positioning
🧩 Unified conversation templates for offline, streaming, and real-time interaction
🧠 256K-token context window
📜 Apache-2.0 license
MOSS-VL-Realtime continues processing new frames while generating a response, allowing it to revise or interrupt that response as the scene evolves—or remain silent when more evidence is needed.
Thank you @sgl_project @lmsysorg for day-0 support! 🚀
Hugginhuggingface.co/OpenMOSS-Team/…bBocrP6
Ggithub.com/OpenMOSS/MOSS-…l4QZQyk
Technicalopenmoss.github.io/MOSS-VLJFSBhxH
Join the commdiscord.gg/SmVQHGffZUaMLZ1nn
👇Show More
00:40
2
16
72
7,587
Jul 15, 2026, 12:25 AM View on X
4 Replies
0 Retweets
2 Likes
746 Views  Berryxia.AI @berryxia
One Sentence Summary
OpenMOSS open-sources the 11B model MOSS-VL-Realtime, supporting chatting while watching video, with the ability to actively modify or interrupt responses as scenes change.
Summary
The author quotes an official tweet from OpenMOSS introducing the open-sourcing of the MOSS-VL-Realtime model. The 11B multimodal model focuses on real-time video streaming interaction. It uses a Cross-Attention architecture separating visual encoding from language reasoning, XRoPE for unified temporal-spatial position encoding, and supports 256K context length. It can perceive and generate responses synchronously while the video streams in, actively modifying or interrupting its own responses when the scene changes, and staying silent when more evidence is needed. The model supports text, single/multiple images, single/multiple videos, and interleaved visual-text inputs in Chinese and English. It achieves leading performance among open-source models on streaming video understanding benchmarks, particularly in proactive alerting and timely response. The model is fully open-sourced under Apache-2.0, along with an instruction-tuned version MOSS-VL-0708 Instruct that shows improvements in fine-grained perception, temporal action localization, and long video understanding.
AI Screening
85
Influence Score 4
Published Today
Language
Chinese
Tags
Multimodal AI
Video Understanding
Open Source
Real-Time AI
MOSS
Make your daily reading actually fit you.A daily brief built from the sources you follow. Get started free HomeDiscoverSettings