How to Set Up Fast Inference for Voice and Multimodal Models with SGLang-Omni
SGLang-Omni provides a specialized runtime for multi-stage inference of voice and multimodal models, handling the complex pipeline from audio encoding to speech synthesis.
Language
HomeLanguages
Sections
SGLang-Omni provides a specialized runtime for multi-stage inference of voice and multimodal models, handling the complex pipeline from audio encoding to speech synthesis.
GigaAM is an open-source acoustic model family from SberDevices that outperforms Whisper on Russian speech recognition. Here's what it can do and how to use it.
Tired of drowning in neural network courses? Deep Learning Drizzle is a curated collection of university courses and lectures from top AI labs worldwide—your antidote to fragmented YouTube tutorials.
MLX-Audio brings lightning-fast TTS, STT, and STS capabilities to Apple Silicon Macs. Built on Apple's MLX framework, it offers Whisper, Kokoro, voice cloning, and more—all running locally without cloud costs.