@HuggingPapers
MOSS-VL: real-time video understanding that perceives while speaking OpenMOSS released MOSS-VL, an 11B open vision-language model that keeps watching live frames while answering — proactive silence and dynamic self-correction built in. https://t.co/UtHiYqJyjO