OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Paper • 2607.23855 • Published Jul 26 • 27
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Paper • 2606.07639 • Published Jun 1 • 5
MOSS Transcribe Collection A unified multimodal large language model for end-to-end speaker-attributed, time-stamped transcription. • 4 items • Updated Jul 11 • 13
OpenMOSS-Team/MOSS-Transcribe-Diarize Audio-Text-to-Text • 0.9B • Updated about 1 month ago • 237k • 405
OpenMOSS-Team/MOSS-Transcribe-preview-2B Automatic Speech Recognition • 2B • Updated Jun 26 • 1.25k • 50