AI Speech APIs | Deepgram & ElevenLabs | Locus
All use cases

AI Speech APIs

Transcribe audio with Deepgram and synthesize natural speech with Deepgram or ElevenLabs — pay per request with no provider subscriptions.

Why agents need speech APIs

AI agents increasingly work across both sides of audio: transcribing meetings, podcasts, and support calls into text, then turning generated responses into natural speech. Through Locus, agents can use Deepgram for URL-based transcription, text-to-speech, and text analysis, or ElevenLabs for high-quality voice generation. Each request is charged to the same USDC wallet with no separate provider account or subscription.

Deepgram speech tools through Locus

Industry-leading accuracy

Deepgram's Nova model achieves state-of-the-art accuracy on speech recognition benchmarks. It handles accents, background noise, and multi-speaker conversations with high fidelity. Supports 30+ languages and dialects.

URL-based transcription

Send a publicly accessible URL for a pre-recorded audio file and receive a transcript in one request. Live WebSocket streaming is not part of this managed tool.

Rich features

Speaker diarization (who said what), punctuation, paragraph formatting, topic detection, sentiment analysis, and entity recognition. Get structured output that's ready for LLM processing.

Predictable per-request pricing

Wrapped transcription publishes its current catalog price before the call. Deepgram text-to-speech and text analysis are available as separate operations with their own pricing basis.

ElevenLabs text-to-speech through Locus

Three production models

Choose Eleven Turbo v2.5 or Flash v2.5 for low-latency speech, or Multilingual v2 for multilingual voice generation.

MP3 or PCM output

Generate MP3 audio by default or request PCM output for downstream audio processing. Responses return base64-encoded audio through the shared API pipeline.

Wrapped and public x402 access

Authenticated wrapped calls can use a voice available to the platform account. Public x402 calls use the curated premade voice IDs documented on the ElevenLabs endpoint page.

Character-based pricing

Text-to-speech is metered by the selected model and character usage. The catalog publishes the current rate for each available operation.

Speech-to-text workflows for agents

  • Meeting summarization Agent transcribes a recorded meeting via Deepgram, then uses a current Claude or GPT model to generate action items, key decisions, and a structured summary.
  • Podcast processing Transcribe podcast episodes, extract key quotes, generate show notes, and create social media clips — all autonomously.
  • Customer support analysis Transcribe support calls, analyze sentiment, identify common issues, and generate reports on support quality and customer satisfaction trends.
  • Content repurposing Convert video content into blog posts, articles, or documentation. Agent transcribes the video, then rewrites the content for the target format using an LLM.

Combining transcription with other APIs

Speech processing is often part of a multi-API pipeline. An agent can transcribe audio with Deepgram, analyze the text with Claude, enrich it with Tavily, and synthesize a spoken response with ElevenLabs. Through Locus, every call is paid from the same USDC wallet with a single audit trail, so the complete cost of transcription, reasoning, enrichment, and voice generation stays visible.

Related resources