ToolPicker
Speechmatics

Speechmatics

41/100

Batch & real-time speech-to-text; high accuracy, 50+ languages.

speechmatics.com

// scorecard · 41/100

verified 1mo ago
methodology →
Agent-readiness11/35
Agent registration (api_key)6/9
Public API5/5
OpenAPI spec0/9
MCP server0/9
llms.txt0/3
Reliability & performance0/15
Actively maintained
Production-ready0/4
Updated recently
Pricing12/13
Itemized public pricing6/6
Free tier / trial4/4
Free-tier generosity2/3
Docs & DX2/12
Docs quality
Public API2/2
Quickstart / examples0/2
Security & compliance12/12
SOC 26/6
ISO 270012/2
GDPR2/2
Security / trust page2/2
Openness4/10
Open source0/6
Open / un-gated API4/4
Company maturity0/3
Company size
Years operating

// rankings · cost per standard workload

#10 in Transcription$24/ 100h audio

Batch Standard $0.24/hr x100 = $24/mo (Enhanced $40)

speechmatics.com/pricing

// overview

Speechmatics provides low-latency speech-to-text APIs supporting 56+ languages for real-time multilingual, multi-speaker conversations, plus text-to-speech capabilities. It offers deployments on-device, on-prem or cloud with high accuracy, speaker awareness and enterprise security features. Targeted at use cases including media captioning, contact centers, healthcare, legal transcription and voice agents.

Best for: Enterprises and developers needing accurate, real-time multilingual STT/TTS with flexible secure deployments

// pricing

Free tier · generosity 3/5
Includes: 3,000 min STT · 1M TTS chars
Limits: 2 concurrent real-time sessions
Free$0
  • 3,000 min STT/month
  • 1M TTS chars/month
  • 2 concurrent sessions
Profrom $0.129/hr
  • 50 concurrent sessions
  • 10 file jobs/sec
  • email support
Enterprisevolume discounts
  • unlimited scale
  • custom models
  • on-prem/SaaS
  • dedicated support

// security & compliance

SOC 2 ISO 27001 GDPR HIPAA

// features

  • Real-time STT with sub-second latency
  • 56+ languages and dialects
  • On-device/on-prem/cloud deployment options
  • Speaker diarization and awareness
  • Custom models and vocabularies
  • Text-to-speech (English, low-latency)