4 ms·
OmniAudio-2.6B: Fastest Audio Language Model for Edge Deployment
- BUFU 2y agoOn a 2024 Mac Mini M4 Pro, Qwen2-Audio-7B-Instruct running on Transformers achieves an average decoding speed of 6.38 tokens/second, while OmniAudio-2.6B through Nexa SDK reaches 35.23 tokens/second in FP16 GGUF version and 66 tokens/second in Q4_K_M quantized GGUF version - delivering 5.5x to 10.3x faster performance on consumer hardware. Blogs for more details: https://nexa.ai/blogs/OmniAudio-2.6B https://nexa.ai/blogs/OmniAudio-2.6B HuggingFace Repo: https://huggingface.co/NexaAIDev/OmniAudio-2.6B https://huggingface.co/NexaAIDev/OmniAudio-2.6B Run locally: https://huggingface.co/NexaAIDev/OmniAudio-2.6B#how-to-use-on-device https://huggingface.co/NexaAIDev/OmniAudio-2.6B#how-to-use-o... Interactive Demo: https://huggingface.co/spaces/NexaAIDev/omni-audio-demo https://huggingface.co/spaces/NexaAIDev/omni-audio-demo