🔌 LLM APIs with free tier · no card · live — last verified by a probe on 2026-09-17 · groq.com · back to the whole list
Fast inference against a free plan Groq publishes as a per-model rate table
gpt-oss, qwen3.6, qwen3.8-27b
Groq states the free plan as a table rather than one quota, in RPM / RPD / TPM / TPD: 30 / 1K / 8K / 200K on openai/gpt-oss-120b, gpt-oss-20b, gpt-oss-safeguard-20b, qwen/qwen3.6-27b and qwen/qwen3.8-27b, 30 / 14.4K / 15K / 500K on the two meta-llama/llama-prompt-guard classifiers, 30 / 250 / 70K on groq/compound and compound-mini, 20 / 2K on the two whisper models and 10 / 100 / 1.2K / 3.6K on the two canopylabs/orpheus voices (read 2026-09-08). Those thirteen rows are the whole free plan, with no Llama among them — the Llama ids in the page’s API samples are not on it. Groq calls the table “a high level summary and there may be exceptions”, and points at the limits page in an account for the exact figures
https://api.groq.com/openai/v1GROQ_API_KEY — get one at https://console.groq.com/keysopenai/gpt-oss-120b, openai/gpt-oss-20b, qwen/qwen3.6-27b, qwen/qwen3.8-27bfree plan limits, qwen/qwen3.8-27b2026-09-10 — Free models changed: added qwen3.8; dropped llama-3.32026-08-17 — Free models changed: added gpt-oss, llama-3.3, qwen3.6; dropped llama-4, qwen32026-07-19 — Added to the list: Fast inference free tierGenerated from registry.yaml on 2026-09-18 and re-verified twice a week; the full list, the Atom feed and the machinery are at https://github.com/mvalentsev/awesome-free-ai-coding.