// scorecard · 36/100
verified 1mo ago// rankings · cost per standard workload
Dedicated GPU deployment per-minute (L4 $0.014/min), not per-search
baseten.co/pricing// overview
Baseten is an inference platform for deploying AI models in production with dedicated inference for high-scale workloads, pre-optimized Model APIs, training, and Frontier Gateway. It provides the fastest model runtimes, cross-cloud high availability, and developer workflows powered by the Baseten Inference Stack, with optimizations for Gen AI apps including image generation, transcription, text-to-speech, LLM runtimes, embeddings, and compound AI. Deployments are available in Baseten Cloud or self-hosted in your VPC, with compliance including SOC 2 Type 2, HIPAA, GDPR, and others.
Best for: High-scale production inference and Gen AI apps needing performance, autoscaling, and compliance
// pricing
- Dedicated deployments
- Model APIs
- Training
- Fast cold starts
- SOC 2 Type II and HIPAA compliant
- Email and in-app chat support
- Everything in Basic
- Priority access to high-demand GPUs
- Dedicated compute
- Higher Model API rate limits
- Dedicated support on Slack and Zoom
- Everything in Pro
- Custom SLAs
- Self-host deployments
- On-demand flex compute
- Advanced security and compliance
- Custom global regions
- Advanced RBAC with Teams
// security & compliance
// features
- Dedicated inference for open-source/custom models
- Pre-optimized Model APIs with per-token pricing
- On-demand training deployments
- Custom performance optimizations for images/transcription/TTS/LLMs/embeddings
- Baseten Cloud or self-hosted VPC deployments
- Forward deployed engineers and rapid iteration DevEx