ToolPicker
Inworld AI (TTS)

Inworld AI (TTS)

51/100

Realtime TTS, 100+ languages, natural-language voice steering.

inworld.ai

// scorecard · 51/100

verified 1mo ago
methodology →
Agent-readiness23/35
Agent registration (api_key)6/9
Public API5/5
OpenAPI spec9/9
MCP server0/9
llms.txt3/3
Reliability & performance0/15
Actively maintained
Production-ready0/4
Updated recently
Pricing12/13
Itemized public pricing6/6
Free tier / trial4/4
Free-tier generosity2/3
Docs & DX2/12
Docs quality
Public API2/2
Quickstart / examples0/2
Security & compliance10/12
SOC 26/6
ISO 270010/2
GDPR2/2
Security / trust page2/2
Openness4/10
Open source0/6
Open / un-gated API4/4
Company maturity0/3
Company size
Years operating

// rankings · cost per standard workload

Realtime TTS-2 on-demand $25/1M chars (volume → $10)

inworld.ai/pricing

// overview

Inworld AI provides realtime text-to-speech (TTS-2), speech-to-text, and LLM routing APIs optimized for low-latency conversational voice AI. It supports voice cloning from 15s audio, text-based voice design, bracketed voice direction, and 100+ languages with cross-lingual cloning. Targeted at scalable voice agents for companions, education, wellness, and interactive media.

Best for: Realtime voice AI apps needing natural low-latency TTS, voice cloning and steering at scale

// pricing

Free tier · generosity 3/5
Includes: 70 min TTS · 100 custom voices · At-cost LLMs · Commercial license
Limits: Evaluation/prototyping only
On-Demandfree
  • 70 min TTS
  • 100 custom voices
  • At-cost LLMs
Creator$25/mo
  • $25 credits
  • Up to 33% off rates
  • 500 custom voices
Builder$100/mo
  • $100 credits
  • Up to 40% off rates
  • 3,000 custom voices
Developer$300/mo
  • $300 credits
  • Up to 47% off rates
  • 10,000 custom voices
Growth$1,500/mo
  • $1,500 credits
  • Up to 53% off rates
  • 30,000 custom voices
Enterprisecustom
  • Custom rates as low as $5/M chars
  • SLA, DPA, on-prem options

// security & compliance

SOC 2 ISO 27001 GDPR HIPAA

// features

  • Realtime TTS <130ms first-chunk latency
  • Voice cloning from 15s audio + text-based voice design
  • Bracketed instructions for tone/speed/style control
  • 100+ languages with cross-lingual cloning
  • LLM Router at 0% markup
  • SOC 2 Type II, HIPAA, GDPR, ZDR compliance