Processes both audio and text inputs with a multi-task framework supporting over 30 language and sound tasks, enabling multi-turn dialogue, sound reasoning, and tool use, while excelling in benchmarks without task-specific fine-tuning or retraining.
Cost / License
- Free
- Open Source
Application types
Platforms
- Mac
- Windows
- Linux
- FFmpeg
- Python
- PyTorch

























































It was great but when they removed the voices its not great