Multimodal model processes audio and text, handling over 30 language and sound tasks with multi-turn dialogue, universal audio understanding, and strong benchmark results.
Cost / License
- Free
- Open Source (Apache-2.0)
Application types
Platforms
- Mac
- Windows
- Linux
- FFmpeg
- Python
- PyTorch




