Together AI vs Groq vs Fireworks AI vs Replicate

Four serverless LLM-inference providers with very different free tiers — OpenAI-compatible open-model endpoints with signup credit, LPU-class low-latency inference with a free rate-limited API, fine-tuned-model hosting, and a huge catalog of community models — compared on free credit, rate limits, and latency.

Together AIGroqFireworks AIReplicate
Free tier noteServerless OpenAI-compatible endpoints for 200+ open models; $5 signup credit, then pay-per-token.LPU-hosted Llama + Mixtral with a free rate-limited API Console tier; very low latency.Serverless host for open + fine-tuned models with OpenAI-compatible API; $1 signup credit, then pay-per-token.Run thousands of community + open models via a single API; trial credit, then pay-per-second GPU billing.
Bandwidthn/an/an/an/a
Buildsn/an/an/an/a
Requests$5 free credits on signup (serverless endpoints)Free API with rate limits (e.g. 30 req/min, daily token cap)$1 free credit on signup (serverless endpoints)Free trial credit for new accounts
Custom domainsn/an/an/an/a
Paid from
Generosity ★41444
Last verified2025-09-012025-09-012025-09-012025-09-01
Visit siteTogether AI ↗Groq ↗Fireworks AI ↗Replicate ↗

Data verified 2025-09-01.