Discover
Audio Models
Voice, music, and sound-effect generation models, plus speech and audio editing tools.
Open Whisper is the model suitable for transcription and summarizes the content of the audio based on speaker input
Open SOTA stemming model for voice, drums, bass, guitar and more.
Open Whether you're working on memes, videos, games, or AI agents, Chatterbox brings your content to life. Use the first tts from resemble ai.
Open Generate speech from text prompts and different voices using the MiniMax Speech-02 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.
Open Analyzes your video’s pacing, mood, and timing to generate a frame-synced, licensed, commercial-use-safe soundtrack in seconds.
Open Audio reasoning variant of NVIDIA's Nemotron 3 Nano Omni. 30B A3B hybrid Transformer-Mamba MoE - accepts audio plus a prompt and returns text.
Open Change the voices in your audios with voices in ElevenLabs!
Open Generate multilingual text-to-speech audio using ElevenLabs TTS Multilingual v2.
Open Transform voices using Resemble AI's Chatterbox. Convert audio to new voices or your own samples, with expressive results and built-in perceptual watermarking.
Open Use Scribe-V2 from ElevenLabs to do blazingly fast speech to text inferences!
Open Generate high-speed text-to-speech audio using ElevenLabs TTS Turbo v2.5.
Open Analyzes a video and generates synchronized, royalty-free sound effects timed to visible actions. Returns the generated sound-effects audio track for commercial use.
Open Detect speech presence and timestamps with accuracy and speed using the ultra-lightweight Silero VAD model
Open Isolate audio tracks using ElevenLabs advanced audio isolation technology.
Open Generate sound effects using ElevenLabs advanced sound effects model.
Open Generate text from speech using ElevenLabs advanced speech-to-text model.
Open Generate speech from text prompts and different voices using the MiniMax Speech-2.8 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.
Open Generate synced sounds for any video, and return the new sound track (like MMAudio)
Open Audio separation with SAM Audio. Isolate any sound using natural language—professional-grade audio editing made simple for creators, researchers, and accessibility applications.
Open Generate high quality, realistic music with fine controls using Elevenlabs Music!