Audio & Voice AI platforms provide tools for generating, processing, and manipulating audio content using artificial intelligence. These platforms offer capabilities like text-to-speech, voice cloning, audio enhancement, and music generation, enabling the creation of high-quality audio content for various applications.
Voice synthesis platform. Features text-to-speech generation, voice cloning, and language support. Includes voice customization and emotion control. Supports audio content creation.
No-code phone AI agent platform. Build and deploy conversational voice agents for inbound and outbound calls with natural speech and workflow automation for sales and support teams.
Pricing: Pay-per-use starting at $0.10/minute. Volume discounts available. Enterprise pricing on request.
Cloud text-to-speech tool that turns pasted scripts into MP3 voiceovers. About 30 voices across 20+ languages with tone, pause, and breathing controls. Download and drop the file into any editor.
Pricing: One-time purchase from $47. Optional Pro and Tube add-ons on site.
Advanced text-to-speech platform that transforms written content into natural-sounding audio using AI-powered voice synthesis. Features high-quality neural voices across multiple languages and accents, with customizable speaking speeds and voice options. Includes OCR technology for converting images and PDFs to audio, real-time synchronization across devices, and intelligent text highlighting. Offers seamless integration with popular platforms including websites, documents, and e-readers. Supports multiple file formats and provides dyslexia-friendly features. Particularly valuable for students, professionals, and individuals with reading difficulties seeking efficient content consumption through audio.
Pricing: Free plan available. Premium plans start at $11.99/month. Team and Enterprise solutions with custom pricing.
Advanced AI-powered noise cancellation platform that removes background noise and echo from voice calls in real-time. Features sophisticated audio processing algorithms that filter out unwanted sounds while preserving voice quality. Includes voice cancellation for privacy, HD voice quality enhancement, and acoustic echo cancellation. Offers seamless integration with all major video conferencing and communication platforms. Supports multiple audio devices and provides meeting analytics for voice quality and participation metrics. Particularly valuable for remote workers, professionals, and teams requiring clear communication in noisy environments.
Pricing: Free plan available. Pro plan starts at $8/month. Enterprise solutions with custom pricing.
Realtime voice AI lab and inference provider for consumer apps. Inworld TTS-2 ranks #1 on Artificial Analysis, alongside streaming STT and LLM serving through modular APIs aimed at cutting voice AI costs.
Pricing: Usage-based API pricing; free tier to start; enterprise plans.
Speech recognition and understanding platform. Features real-time transcription, speaker identification, and language detection. Includes custom model training and API integration. Supports voice applications.
Voice generation and text-to-speech platform. Features natural voice synthesis, custom voice cloning, and multi-language support. Includes voice library, real-time generation, and enterprise integration. Supports professional voice content creation.
Real-time, speech-native voice AI platform. Trains Ultravox model - state-of-the-art on Big Bench Audio (97% with reasoning). Features speech-native processing without transcription latency, paralinguistic signal preservation, and purpose-built inference infrastructure. Open-weight models on Hugging Face. Used by 11x and other leading voice agent companies.
Enterprise voice AI and cloning platform. Features high-fidelity voice synthesis, emotion control, and real-time voice generation. Includes voice preservation and custom model training. Supports enterprise voice applications.
Text-to-speech generator with 1000+ voices across 140+ languages, plus voice cloning and text-to-video. Browser studio for turning scripts into narrated audio.
Pricing: Free trial (1000 words); paid plans on site.
Advanced AI speech recognition and audio intelligence platform. Features real-time transcription, speaker identification, content summarization, and sentiment analysis. Includes multi-language support, custom vocabulary training, entity detection, and comprehensive API integration. Helps developers and businesses convert speech to text, analyze audio content, and extract insights from voice data at scale.
Platform specializing in voice technology and AI agents. Features include text-to-speech generation with natural-sounding voices across multiple languages, voice cloning capabilities, and AI calling agents that can handle customer interactions. Offers two main products: Waves for creating human-like AI voices and Atoms for building conversational AI agents that can take calls and automate workflows.
Pricing: Free tier available. Paid plans based on usage volume. Enterprise solutions with custom pricing.
Language technology platform specializing in automatic speech recognition (ASR), neural machine translation (NMT), and natural language processing. Offers enterprise-grade solutions for speech-to-text, text-to-text translation, generative text with large language models, and neural speech synthesis across multiple languages. Used in media production, government applications, customer engagement, and accessibility compliance.
Pricing: Custom pricing based on enterprise needs. Free demos available for specific products including ASR API, Speech Translate mobile app, and Neural Machine Translation.
Platform that automates podcast creation by converting RSS feeds, podcasts, and YouTube playlists into new episodes. Features content summarization, translation, and scheduled publishing with natural-sounding AI voices.
Pricing: Free plan available, Pro: Contact for pricing
Text-to-speech and voice cloning across 150+ languages, with a studio editor and TTS API including a low-latency Flash model. Voice design, captions, and related creative tools in one dashboard.
Pricing: Starts from $1; API usage-based; see Verbatik.
Emotionally expressive text-to-speech with 700+ AI voices. Controls for tone, pitch, speed, and emotion, plus an API for commercial, audiobook, and social voiceovers. Free tier available.
Enterprise-grade AI voice cloning and synthesis platform trusted by Lucasfilm, Sony, Blumhouse, and Digital Domain. Emmy Award winner for voice work on The Mandalorian (young Luke Skywalker). Features AI Voice Lab for film/TV production, text-to-speech API, speech-to-speech conversion, cross-language voice cloning, and Pro Tools plugin integration. Used in Cyberpunk 2077, The Brutalist, Emilia Pérez. Specializes in ethical voice recreation for entertainment, dubbing, games, and advertising.
Pricing: Enterprise pricing. Custom quotes based on project needs.
High-accuracy automated transcription platform with 97-99% accuracy in 53+ languages and translation in 54+ languages. Features AI-powered summaries, automated subtitles, topic detection, sentiment analysis, and multi-user collaboration. Used by Google, Microsoft, NBC Universal, Warner Bros, Adobe, IBM, Wall Street Journal, Stanford, Yale, and UCLA. Includes enterprise-grade security and integrations with Zoom, Adobe Premiere, and major workflow tools.
Pricing: Pay-as-you-go: $5-10/hour. Premium and Enterprise plans available.
Voice AI agent platform for enterprises. Build and deploy conversational AI voice agents in minutes. Sub-700ms latency, multilingual support, developer-friendly API. Handles customer calls, sales, support at scale. $20M Series A (Bessemer, Y Combinator). Used by Mindtickle, Luma Health, Ellipsis Health.
Pricing: Pay-per-minute. Contact for enterprise pricing.
AI voice agent platform for contact centers. Build, test, deploy, and monitor voice agents for phone calls. ~600ms latency, 3200+ businesses, profitable. Knowledge base integration, verified phone numbers, batch calling. Y Combinator W24; $4.6M seed from Alt Capital.
Conversational AI platform for enterprise phone automation. AI voice agents for calls, SMS, chat. Voice cloning from single MP3; self-hosted stack for data control. SOC2, GDPR, HIPAA compliant. 1M+ concurrent call capacity. Trusted by Samsara, Snapchat, Gallup. Forward-deployed engineers for custom agent build-out.
Real-time text-to-speech and voice API platform for developers and products. Features streaming synthesis, low-latency playback, voice cloning, and multilingual models suited to agents, apps, and interactive voice experiences.
Pricing: Usage-based API pricing; see site for current tiers.
AI voice synthesis platform with real-time and batch TTS APIs, voice cloning, and a large voice library. Targets creators and developers building conversational, entertainment, and product voice experiences.
Pricing: Usage-based and subscription options; see site for current pricing.
Audio intelligence API platform for speech-to-text, diarization, and related speech workflows. Built for developers integrating transcription and audio understanding into apps, contact centers, and media pipelines.
Pricing: Pay-as-you-go and enterprise plans; see site for current rates.
Research-backed voice AI stack combining expressive text-to-speech (Octave), real-time empathic speech-to-speech (EVI), and measurement models for prosody and emotion. Octave is built as an LLM-aware TTS system for nuanced delivery, streaming APIs, voice design from prompts, and fast cloning from short samples. Strong fit for developers building emotionally rich assistants, interactive media, and production voice products that need APIs, SDKs, and integrations with agent frameworks.
Pricing: Usage-based API pricing and plans; see site for current tiers and free trial options.
Enterprise-focused text-to-speech APIs tuned for real-time voice agents, IVR, and high-stakes phone experiences. Offers flagship Arcana models for natural, expressive speech plus Mist variants optimized for speed, pronunciation control, and streaming at scale. Supports multilingual code-switching, on-prem and VPC deployment options, and observability for production traffic. Best for teams shipping voice agents and telephony where latency, accuracy, and compliance matter.
Pricing: Usage-based API pricing; enterprise and private deployment options; see site for quotes.
Speech-to-text and audio intelligence APIs with broad language coverage, real-time and batch transcription, entity formatting, and features aimed at contact centers, media, and product teams. Emphasizes accuracy on noisy audio, flexible deployment, and developer-first integration for apps that need reliable ASR at scale. Strong fit when voice input, analytics, or compliance-grade transcription pipelines are core to the product.
Pricing: Pay-as-you-go and enterprise licensing; see site for current per-hour rates and packages.
Voice AI company focused on low-latency, human-like TTS and cloning for conversational products, with emphasis on on-device and private deployments as well as managed cloud APIs. Positions models for support lines, IVR, and assistants where natural cadence and operational cost matter. Useful for builders exploring edge-first voice, streaming synthesis, and enterprise-scale rollouts.
Pricing: Managed API and enterprise plans; open-source model paths advertised; see site for current pricing.
AI dubbing platform for creators, broadcasters, and studios. Translates video into 40+ languages with voice cloning that preserves speaker cadence, warmth, and timing. Scene-aware translation and optional lip-sync for broadcast-ready localization.
Personal AI agents with conversational voice - research-rooted speech models (CSM) plus a consumer agent preview on iOS, with intelligent eyewear on the roadmap. Voice-native personal intelligence rather than batch TTS alone.
Pricing: Research demos free; product pricing TBD.
AWS neural text-to-speech service for apps and IVR. Dozens of voices and languages, SSML controls, and brand voice customization on Amazon infrastructure.
Pricing: Pay-as-you-go; Free Tier available. See AWS pricing.
Realtime voice AI and TTS from Inworld - sub-200ms latency, instant cloning, emotion and non-verbal controls, multilingual support. Positioned for agents and games at aggressive per-character pricing.
Pricing: Usage-based; volume rates as low as ~$5 per 1M characters-see Inworld.
Multilingual speech AI platform with realtime STT, TTS, and translation APIs across 60+ languages. Sub-200ms latency and language switching for live voice applications.
OpenAI open-source speech recognition trained on large-scale weak supervision. Robust multilingual transcription and translation that became the default open ASR baseline for products and research.
Pricing: Open source (MIT); hosted inference available from many providers.
OpenAI text-to-speech API for apps and agents. Multiple built-in voices, streaming audio, and tight integration with the OpenAI platform for product and voice-agent workloads.
OpenAI speech-to-speech Realtime API for live voice agents. Low-latency audio in/out with tool calling, interruptions, and multimodal conversation without a separate STT-LLM-TTS pipeline.
Pricing: Usage-based API pricing; see OpenAI Realtime docs.
Low-latency voice AI for TTS, speech-to-text, and voice cloning. Built for real-time agents where time-to-first-audio and consistency matter as much as voice quality.
Speech-to-text API known for high accuracy transcription. Asynchronous and streaming ASR for developers building captions, compliance, and voice workflows.
NVIDIA Speech AI SDK for GPU-accelerated ASR, TTS, and related speech services. Enterprise deployment for low-latency transcription and synthesis on NVIDIA infrastructure.
Pricing: NVIDIA developer and enterprise licensing; see NVIDIA Riva.
MiniMax speech and music models for expressive multilingual TTS, voice cloning, and audio generation. API access for agents, media, and product voice experiences at foundation-model scale.
Pricing: API and platform plans; see MiniMax Audio.
Amazon Nova speech-to-speech foundation model for realtime conversational voice. Bidirectional audio with tool use on AWS without a separate STT-LLM-TTS pipeline.
Pricing: AWS usage-based; see Amazon Nova pricing.
Google Gemini Live API for low-latency realtime voice and multimodal conversation. Streaming speech in/out with Gemini reasoning for agents and interactive apps.
Pricing: Google Cloud / Gemini API usage; see Live API docs.
AWS automatic speech recognition for batch and streaming transcription. Speaker diarization, custom vocabularies, and language identification for apps and contact centers.
Pricing: Pay-per-minute AWS pricing; see Amazon Transcribe.
Google Cloud ASR with Chirp-class models for accurate multilingual transcription. Real-time and batch recognition APIs for products, captions, and voice pipelines.
Pricing: Google Cloud usage-based; free tier available.