ToolPicker
Together AI Rerank

Together AI Rerank

36/100

Serverless rerank (Salesforce LlamaRank) for RAG.

together.ai

// scorecard · 36/100

verified 1mo ago
methodology →
Agent-readiness14/35
Agent registration (api_key)6/9
Public API5/5
OpenAPI spec0/9
MCP server0/9
llms.txt3/3
Reliability & performance0/15
Actively maintained
Production-ready0/4
Updated recently
Pricing6/13
Itemized public pricing6/6
Free tier / trial0/4
Free-tier generosity0/3
Docs & DX2/12
Docs quality
Public API2/2
Quickstart / examples0/2
Security & compliance10/12
SOC 26/6
ISO 270012/2
GDPR0/2
Security / trust page2/2
Openness4/10
Open source0/6
Open / un-gated API4/4
Company maturity0/3
Company size
Years operating

// rankings · cost per standard workload

#5 in Rerankers$1/ 1k rerank searches

LlamaRank $0.10/1M tokens x10M = $1/mo (per-token)

together.ai/models/salesforce-llamarank

// overview

Together AI is a full-stack AI platform providing serverless inference, batch inference, dedicated inference, GPU clusters, fine-tuning, evaluations, and managed storage for open-source models. It offers high-performance APIs for models including MiniMax M3, Gemma 4 31B, DeepSeek V4 Pro, GLM-5.2, and Kimi K2.7 Code with per-million-token pricing. The platform emphasizes research-driven optimizations for faster inference and lower costs.

Best for: Teams starting with serverless inference for open-source models and scaling to dedicated endpoints

// pricing

Serverless Inferenceper 1M tokens
  • Chat, Vision, Embeddings, Rerank
  • Model-specific rates e.g. DeepSeek V4 Pro $1.74 input / $3.48 output
  • MiniMax M3 $0.30 input / $1.20 output

// security & compliance

SOC 2 ISO 27001 GDPR HIPAA

// features

  • Serverless Inference APIs
  • Batch Inference
  • Dedicated Model Inference
  • GPU Clusters (GB200, B200, H100 etc.)
  • Fine-Tuning and Evaluations
  • Managed Storage

sources