Best free tier for inference

Groq

★ 14

LPU-hosted Llama + Mixtral with a free rate-limited API Console tier; very low latency.

Bandwidth
n/a
Requests
Free API with rate limits (e.g. 30 req/min, daily token cap)
Paid from
Visit site ↗

Together AI

★ 4

Serverless OpenAI-compatible endpoints for 200+ open models; $5 signup credit, then pay-per-token.

Bandwidth
n/a
Requests
$5 free credits on signup (serverless endpoints)
Paid from
Visit site ↗

Fireworks AI

★ 4

Serverless host for open + fine-tuned models with OpenAI-compatible API; $1 signup credit, then pay-per-token.

Bandwidth
n/a
Requests
$1 free credit on signup (serverless endpoints)
Paid from
Visit site ↗

Replicate

★ 4

Run thousands of community + open models via a single API; trial credit, then pay-per-second GPU billing.

Bandwidth
n/a
Requests
Free trial credit for new accounts
Paid from
Visit site ↗