ToolPicker
Baseten Rerank

Baseten Rerank

36/100

Host reranker models (bge-reranker-v2-m3) on GPUs.

baseten.co

// scorecard · 36/100

verified 1mo ago
methodology →
Agent-readiness14/35
Agent registration (api_key)6/9
Public API5/5
OpenAPI spec0/9
MCP server0/9
llms.txt3/3
Reliability & performance0/15
Actively maintained
Production-ready0/4
Updated recently
Pricing6/13
Itemized public pricing6/6
Free tier / trial0/4
Free-tier generosity
Docs & DX2/12
Docs quality
Public API2/2
Quickstart / examples0/2
Security & compliance10/12
SOC 26/6
ISO 270010/2
GDPR2/2
Security / trust page2/2
Openness4/10
Open source0/6
Open / un-gated API4/4
Company maturity0/3
Company size
Years operating

// rankings · cost per standard workload

#11 in Rerankersnot comparable

Dedicated GPU deployment per-minute (L4 $0.014/min), not per-search

baseten.co/pricing

// overview

Baseten is an inference platform for deploying AI models in production with dedicated inference for high-scale workloads, pre-optimized Model APIs, training, and Frontier Gateway. It provides the fastest model runtimes, cross-cloud high availability, and developer workflows powered by the Baseten Inference Stack, with optimizations for Gen AI apps including image generation, transcription, text-to-speech, LLM runtimes, embeddings, and compound AI. Deployments are available in Baseten Cloud or self-hosted in your VPC, with compliance including SOC 2 Type 2, HIPAA, GDPR, and others.

Best for: High-scale production inference and Gen AI apps needing performance, autoscaling, and compliance

// pricing

Basic$0 per month pay as you go
  • Dedicated deployments
  • Model APIs
  • Training
  • Fast cold starts
  • SOC 2 Type II and HIPAA compliant
  • Email and in-app chat support
ProVolume discounts available
  • Everything in Basic
  • Priority access to high-demand GPUs
  • Dedicated compute
  • Higher Model API rate limits
  • Dedicated support on Slack and Zoom
EnterpriseVolume discounts available
  • Everything in Pro
  • Custom SLAs
  • Self-host deployments
  • On-demand flex compute
  • Advanced security and compliance
  • Custom global regions
  • Advanced RBAC with Teams

// security & compliance

SOC 2 ISO 27001 GDPR HIPAA

// features

  • Dedicated inference for open-source/custom models
  • Pre-optimized Model APIs with per-token pricing
  • On-demand training deployments
  • Custom performance optimizations for images/transcription/TTS/LLMs/embeddings
  • Baseten Cloud or self-hosted VPC deployments
  • Forward deployed engineers and rapid iteration DevEx

sources