Popular Audio & Voice AI tools

35 category leaders in Audio & Voice, selected from the full directory.

ElevenLabs

  • Voice synthesis, cloning, and conversational AI APIs.
  • Studio-quality TTS used by creators, media, and products.
  • Multilingual voices with emotion and style controls.
  • Default choice for production voice apps and agents.

Murf

  • AI voiceover studio for video, courses, and ads.
  • Natural TTS with accents, styles, and voice cloning.
  • Timeline editor for timing, music, and captions.
  • Widely used by marketers and e-learning teams.

Inworld

  • Realtime TTS, STT, and LLM serving via modular APIs.
  • Inworld TTS-2 ranks #1 on Artificial Analysis.
  • Built for low-latency consumer voice apps.
  • Former game-NPC platform turned voice AI lab.

Deepgram

  • Speech-to-text and voice AI APIs for products.
  • Real-time and batch ASR with speaker diarization.
  • Custom models and low-latency streaming endpoints.
  • Standard ASR stack for contact centers and apps.

WellSaid

  • Enterprise TTS for training, marketing, and media.
  • Studio voices with brand-safe commercial licensing.
  • Team workspaces and API for production pipelines.
  • Go-to for corporate voiceover at scale.

Resemble

  • Enterprise voice cloning and real-time TTS APIs.
  • Custom voice models with brand and compliance controls.
  • Speech-to-speech and detection tools for media teams.
  • Common stack for localized and branded voice products.

Assembly

  • Speech-to-text and audio intelligence APIs for products.
  • Transcription, summarization, and speaker labels.
  • Streaming and batch ASR used across many apps.
  • Default developer choice for speech understanding layers.

Respeecher

  • Emmy-winning AI voice cloning for film and games.
  • Credits include The Mandalorian, Cyberpunk 2077, The Brutalist.
  • Used by Lucasfilm, Sony, Blumhouse, and Digital Domain.
  • Studio-grade speech-to-speech for media production.

Vapi

  • Build and host voice AI agents over phone and web.
  • Low-latency orchestration across STT, LLM, and TTS.
  • Dashboards, evals, and tooling for production agents.
  • Go-to developer platform for conversational voice apps.

Bland

  • AI phone agents for outbound and inbound enterprise calls.
  • Self-host options with compliance-minded deployments.
  • High concurrent call capacity for large dial programs.
  • Leading platform for automated business phone voice.

Cartesia

  • Real-time voice AI API platform.
  • Known for low-latency speech synthesis.
  • Popular in agent and app voice stacks.
  • Strong developer adoption in interactive audio.

Hume

  • Expressive TTS and empathic speech-to-speech APIs.
  • Octave models tuned for emotion-aware delivery.
  • Fits agents and apps that need nuanced voice tone.
  • Strong pick when expression matters as much as clarity.

Speechmatics

  • Speech-to-text and audio intelligence APIs with broad language coverage.
  • Targets real-time and batch transcription for contact centers and products.
  • Known for accuracy on challenging audio and regulated-industry deployments.
  • Strong fit when dependable ASR and audio understanding power the core product.

Speechify

  • Text-to-speech for articles, PDFs, and ebooks.
  • Neural voices across languages with speed controls.
  • OCR and sync across mobile, web, and extensions.
  • Widely used for accessibility and faster reading.

Krisp

  • Real-time AI noise and echo cancellation for calls.
  • Works with Zoom, Meet, Teams, and major dialers.
  • Voice isolation for clearer remote meetings.
  • Standard tool for distributed teams and call centers.

Retell

  • Build, deploy, and monitor AI phone voice agents.
  • Low-latency calls with knowledge bases and batch dial.
  • Used by thousands of businesses; YC W24.
  • Pay-per-minute platform for contact-center automation.

Sesame

  • Conversational speech models with human-like presence.
  • Viral Maya/Miles voice demos and open CSM weights.
  • Research lab push on full-duplex spoken agents.
  • Defining reference for natural conversational voice.

Amazon Polly

  • AWS neural text-to-speech for apps and telephony.
  • Broad language coverage with SSML voice controls.
  • Default TTS path inside many Amazon stacks.
  • Hyperscaler speech synthesis at global scale.

Google Cloud Text-to-Speech

  • Google Cloud neural and Studio TTS voices.
  • WaveNet-class quality across many languages.
  • SSML and product-ready Cloud APIs.
  • Standard enterprise TTS on Google Cloud.

Azure AI Speech

  • Microsoft TTS, STT, and custom neural voice.
  • Speech Studio for enterprise voice programs.
  • Translation and real-time speech in Azure.
  • Default speech stack for Microsoft clouds.

Adobe Podcast

  • Adobe Enhance Speech for studio-like mic cleanup.
  • Browser tools to record and polish podcast audio.
  • Widely used free upgrade path for laptop recordings.
  • Household creative-cloud brand in AI audio.

Whisper

  • OpenAI open speech recognition trained at large scale.
  • Multilingual transcription and speech translation.
  • Default open ASR baseline across research and products.
  • MIT-licensed weights with wide hosted inference support.

OpenAI TTS

  • OpenAI text-to-speech API for apps and agents.
  • Built-in voices with streaming audio output.
  • Tight integration with the OpenAI platform stack.
  • Common default when shipping GPT-backed voice UX.

OpenAI Realtime

  • Speech-to-speech API for live voice agents.
  • Handles interruptions, tools, and turn-taking live.
  • Skips separate STT-LLM-TTS pipelines for many apps.
  • Realtime conversational voice without a classic STT stack.

Rev

  • High-accuracy speech-to-text API for developers.
  • Async and streaming ASR for captions and compliance.
  • Long track record in professional transcription.
  • Major alternative to Deepgram and AssemblyAI.

NVIDIA Riva

  • GPU-accelerated ASR and TTS speech SDK.
  • Enterprise deployment on NVIDIA infrastructure.
  • Low-latency transcription and synthesis pipelines.
  • NVIDIA's speech AI stack for production systems.

MiniMax Audio

  • MiniMax foundation models for speech and music.
  • Expressive multilingual TTS and voice cloning APIs.
  • Built for agents, media, and product voice stacks.
  • Major Asia-rooted lab expanding global audio AI.

Amazon Nova Sonic

  • AWS speech-to-speech model for live voice agents.
  • Bidirectional audio without separate STT-TTS hops.
  • Tool use and conversation on Nova infrastructure.
  • Hyperscaler native path for realtime voice AI.

Gemini Live

  • Google Gemini realtime voice and multimodal Live API.
  • Streaming speech with Gemini reasoning in the loop.
  • Built for interactive agents and live assistants.
  • Default Google path for low-latency spoken AI.

Amazon Transcribe

  • AWS speech-to-text for batch and streaming ASR.
  • Diarization, custom vocab, and language ID features.
  • Widely used in contact centers and media pipelines.
  • Default Amazon cloud transcription service.

Google Cloud Speech-to-Text

  • Google Cloud ASR with Chirp-class recognition models.
  • Realtime and batch APIs across many languages.
  • Captions, voice typing, and product STT workloads.
  • Hyperscaler counterpart to Google Cloud TTS.

Ultravox

  • Speech-native voice AI without transcription lag.
  • Open-weight models strong on Big Bench Audio.
  • Preserves paralinguistics for agent conversations.
  • Chosen by leading voice-agent product teams.

Fish Audio

  • Creator-friendly TTS with cloning and voice library.
  • Realtime and batch APIs for apps and agents.
  • Strong adoption for conversational voice products.
  • Popular open-leaning alternative to premium TTS.

Inworld TTS

  • Realtime TTS tuned for agents and games.
  • Sub-200ms latency with cloning and emotion controls.
  • Aggressive per-character pricing at volume.
  • Benchmarks among the fastest production voice APIs.

Gladia

  • Speech-to-text and audio intelligence APIs.
  • Diarization and transcription for apps and media.
  • Developer-first alternative to Assembly-class STT.
  • Used across contact centers and content pipelines.