Together AI vs Groq vs Fireworks AI vs Replicate
Four serverless LLM-inference providers with very different free tiers — OpenAI-compatible open-model endpoints with signup credit, LPU-class low-latency inference with a free rate-limited API, fine-tuned-model hosting, and a huge catalog of community models — compared on free credit, rate limits, and latency.
| Together AI | Groq | Fireworks AI | Replicate | |
|---|---|---|---|---|
| Free tier note | Serverless OpenAI-compatible endpoints for 200+ open models; $5 signup credit, then pay-per-token. | LPU-hosted Llama + Mixtral with a free rate-limited API Console tier; very low latency. | Serverless host for open + fine-tuned models with OpenAI-compatible API; $1 signup credit, then pay-per-token. | Run thousands of community + open models via a single API; trial credit, then pay-per-second GPU billing. |
| Bandwidth | n/a | n/a | n/a | n/a |
| Builds | n/a | n/a | n/a | n/a |
| Requests | $5 free credits on signup (serverless endpoints) | Free API with rate limits (e.g. 30 req/min, daily token cap) | $1 free credit on signup (serverless endpoints) | Free trial credit for new accounts |
| Custom domains | n/a | n/a | n/a | n/a |
| Paid from | — | — | — | — |
| Generosity ★ | 4 | 14 | 4 | 4 |
| Last verified | 2025-09-01 | 2025-09-01 | 2025-09-01 | 2025-09-01 |
| Visit site | Together AI ↗ | Groq ↗ | Fireworks AI ↗ | Replicate ↗ |
Data verified 2025-09-01.