ToolPicker
Fireworks AI

Fireworks AI

42/100

Serverless embeddings (nomic-embed, etc.), per-1M tokens.

fireworks.ai

// scorecard · 42/100

verified 1mo ago
methodology →
Agent-readiness11/35
Agent registration (api_key)6/9
Public API5/5
OpenAPI spec0/9
MCP server0/9
llms.txt0/3
Reliability & performance0/15
Actively maintained
Production-ready0/4
Updated recently
Pricing11/13
Itemized public pricing6/6
Free tier / trial4/4
Free-tier generosity1/3
Docs & DX8/12
Docs quality6/8
Public API2/2
Quickstart / examples0/2
Security & compliance8/12
SOC 26/6
ISO 270010/2
GDPR0/2
Security / trust page2/2
Openness4/10
Open source0/6
Open / un-gated API4/4
Company maturity0/3
Company size
Years operating

// rankings · cost per standard workload

#1 in Embeddings$0.8/ 100M tokens

nomic-embed-text-v1.5 $0.008/1M x100 = $0.80/mo

docs.fireworks.ai

// overview

Fireworks AI is a platform for fast inference and fine-tuning of open source generative AI models, offering serverless pay-per-token serving, on-demand GPU deployments, and multiple training options including supervised, preference, and reinforcement fine-tuning. It supports 100+ models (text, vision, embeddings, audio) with optimizations for throughput/latency and OpenAI-compatible APIs. Processes 30T+ tokens daily and includes a model library and full training-to-production workflow.

Best for: Developers and teams building production inference, agentic systems, RAG, or specialized fine-tuned models at scale with open-source economics.

// pricing

Free tier · generosity 1/5
Includes: $1 free credits
Limits: self-serve only
Serverless Inferencepay per token
  • Fast/Priority tiers
  • 50% cached input
  • 50% batch
Fine Tuningper 1M training tokens
  • LoRA SFT/DPO
  • Full Param SFT/DPO
  • Reinforcement FT per GPU hour
On-Demand Deploymentsper GPU hour
  • H100 $7
  • H200 $7
  • B200 $10

// security & compliance

SOC 2 ISO 27001 GDPR HIPAA

// features

  • Serverless Inference with pay-per-token and cached/batch discounts
  • On-Demand and Reserved GPU deployments (H100/H200/B200)
  • Supervised/Preference/RL fine-tuning with LoRA and full-param options
  • Model library with instant access to latest OSS models (Deepseek, GLM, Qwen, etc.)
  • Vision, embeddings, tool calling, structured outputs, and batch API
  • Autoscaling deployments and post-training model serving