// scorecard · 42/100
verified 1mo agoAgent-readiness11/35
Agent registration (api_key)6/9
Public API5/5
OpenAPI spec0/9
MCP server0/9
llms.txt0/3
Reliability & performance0/15
Actively maintained—
Production-ready0/4
Updated recently—
Pricing11/13
Itemized public pricing6/6
Free tier / trial4/4
Free-tier generosity1/3
Docs & DX10/12
Docs quality6/8
Public API2/2
Quickstart / examples2/2
Security & compliance6/12
SOC 26/6
ISO 270010/2
GDPR0/2
Security / trust page0/2
Openness4/10
Open source0/6
Open / un-gated API4/4
Company maturity0/3
Company size—
Years operating—
// rankings · cost per standard workload
// overview
Fireworks AI is a platform for fast inference and fine-tuning of open source generative AI models, offering serverless pay-per-token serving, dedicated on-demand deployments, and full-spectrum training options including supervised, preference, and reinforcement fine-tuning. It supports text, vision, audio, embeddings, and image models with optimized throughput/latency and OpenAI/Anthropic-compatible APIs. The service processes 30T+ tokens per day and includes a large model library with instant access to popular OSS models.
Best for: developers and enterprises needing high-performance inference, autoscaling, and fine-tuning of open models at scale
// pricing
Free tier · generosity 1/5
Includes: $1 free credits on serverless
Limits: self-serve only
Serverless Inferencepay per token
- Priority and Fast tiers
- 50% cached input
- 50% batch pricing
Fine Tuningper 1M training tokens or GPU hour
- LoRA SFT/DPO and Full Param options
- same serving price as base models
On-Demand Deploymentsper GPU hour
- H100 $7, H200 $7, B200 $10, B300 $12
// security & compliance
✓ SOC 2✗ ISO 27001✗ GDPR✓ HIPAA
// features
- Serverless inference with per-token pricing and no cold starts
- On-demand and reserved GPU deployments with autoscaling
- Supervised, preference, and reinforcement fine-tuning up to 1T+ parameters
- Large library of optimized open models (text, vision, audio, embeddings)
- Batch inference at 50% serverless rates and cached input at 50%
- Function calling, structured outputs, and vision model support