Issue 01AI 리서치
AC POST
AI 리서치 목록
arxiv2026년 9월 18일 09:00

OmniVChat: 네이티브 오디오-비주얼 대화를 위한 합성, 벤치마킹 및 학습

본 논문은 사용자와 omni model 간의 native audio-visual dialogue를 수행하는 OmniVChat 태스크를 정의합니다. 이 태스크는 별도의 text question이나 speech recognition 없이 audio와 video를 직접 동시에 입력받아 text를 출력하는 것을 목표로 합니다. 데이터 부족과 평가의 어려움을 해결하기 위해 synthesized dialogues를 활용한 training 및 evaluation 방안을 제시합니다.

We define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model. In OmniVChat, omni models directly and simultaneously receive audio and video from a user and return text. The user's query is embedded in the audio and video, without a separate text question, external captioning, or speech recognition. Direct audio-visual input reduces external latency and computation while preserving perceptual cues. However, research on OmniVChat faces two constraints: data availability and evaluation. Recordings of people using their own devices are scarce. Furthermore, a good reply often needs to account for the user's surroundings, facial expressions, and nearby objects, and such responses can be expressed in many different ways, making keyword matching unreliable for evaluating reply quality. Recent progress in agent systems and video generation makes generation for comprehension viable, which means using synthesized dialogues for training and eva