Discover

Audio Models

Voice, music, and sound-effect generation models, plus speech and audio editing tools.

139Models
Speechify
Open
SpeechifyFCAST

Whisper is the model suitable for transcription and summarizes the content of the audio based on speaker input

Demucs
Open
DemucsMeta

SOTA stemming model for voice, drums, bass, guitar and more.

Chatterbox
Open
ChatterboxResemble AI

Whether you're working on memes, videos, games, or AI agents, Chatterbox brings your content to life. Use the first tts from resemble ai.

MiniMax Speech-02 HD
Open
MiniMax Speech-02 HDMiniMax

Generate speech from text prompts and different voices using the MiniMax Speech-02 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.

V1.1 Video to Music
Open
V1.1 Video to MusicSonilo

Analyzes your video’s pacing, mood, and timing to generate a frame-synced, licensed, commercial-use-safe soundtrack in seconds.

Nemotron 3 Nano Omni Audio
Open
Nemotron 3 Nano Omni AudioNvidia

Audio reasoning variant of NVIDIA's Nemotron 3 Nano Omni. 30B A3B hybrid Transformer-Mamba MoE - accepts audio plus a prompt and returns text.

ElevenLabs Voice Changer
Open
ElevenLabs Voice ChangerElevenLabs

Change the voices in your audios with voices in ElevenLabs!

ElevenLabs TTS Multilingual v2
Open
ElevenLabs TTS Multilingual v2ElevenLabs

Generate multilingual text-to-speech audio using ElevenLabs TTS Multilingual v2.

Chatterboxhd
Open
ChatterboxhdResemble AI

Transform voices using Resemble AI's Chatterbox. Convert audio to new voices or your own samples, with expressive results and built-in perceptual watermarking.

ElevenLabs - Scribe V2
Open
ElevenLabs - Scribe V2ElevenLabs

Use Scribe-V2 from ElevenLabs to do blazingly fast speech to text inferences!

ElevenLabs TTS Turbo v2.5
Open
ElevenLabs TTS Turbo v2.5ElevenLabs

Generate high-speed text-to-speech audio using ElevenLabs TTS Turbo v2.5.

V1.1 Video to Sound Effects
Open
V1.1 Video to Sound EffectsSonilo

Analyzes a video and generates synchronized, royalty-free sound effects timed to visible actions. Returns the generated sound-effects audio track for commercial use.

Silero VAD
Open
Silero VADSilero

Detect speech presence and timestamps with accuracy and speed using the ultra-lightweight Silero VAD model

ElevenLabs Audio Isolation
Open
ElevenLabs Audio IsolationElevenLabs

Isolate audio tracks using ElevenLabs advanced audio isolation technology.

Elevenlabs Sound Effects V2
Open
Elevenlabs Sound Effects V2ElevenLabs

Generate sound effects using ElevenLabs advanced sound effects model.

ElevenLabs Scribe V1
Open
ElevenLabs Scribe V1ElevenLabs

Generate text from speech using ElevenLabs advanced speech-to-text model.

MiniMax Speech 2.8 [HD]
Open
MiniMax Speech 2.8 [HD]MiniMax

Generate speech from text prompts and different voices using the MiniMax Speech-2.8 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.

Mirelo SFX V1.5
Open
Mirelo SFX V1.5Mirelo AI

Generate synced sounds for any video, and return the new sound track (like MMAudio)

Sam Audio Text-guided
Open
Sam Audio Text-guidedOpen Source

Audio separation with SAM Audio. Isolate any sound using natural language—professional-grade audio editing made simple for creators, researchers, and accessibility applications.

Elevenlabs Music
Open
Elevenlabs MusicElevenLabs

Generate high quality, realistic music with fine controls using Elevenlabs Music!

Showing 1–20 of 139