🔌 LLM APIs with free tier · card required · provisional — added recently, two weeks of probes still to pass · live — last verified by a probe on 2026-09-17 · ibm.com · back to the whole list
IBM’s watsonx.ai Runtime on its Lite plan — 300,000 tokens a month of foundation-model inference (Granite, Llama, Mistral and other hosted models) on IBM Cloud, a plan IBM’s own docs call free and never bill
The page this row is verified against names no free model, so the column stays empty; callable ids, where the row has them, are under Connect.
“300,000 tokens per month”, “20 CUH per month” of compute and “2 inference requests per second”, on “A free plan with limited capacity” — the Lite plan of watsonx.ai Runtime as its service-plans page reads on 2026-09-16, with no expiry named. The card is taken at the door and not charged: the sign-up doc says “For your IBM Cloud account, you enter your email address, personal information, and credit card information, which is used to verify your identity” and “Lite plans do not incur charges”. Which foundation models the 300,000 tokens reach is on a separate docs page the probe does not read, which is why the Free models column is empty
https://us-south.ml.cloud.ibm.com/ml/v1 (not OpenAI-shaped)IBM_WATSONX_AI_API_KEY — get one at https://cloud.ibm.com/iam/apikeys300,000 tokens per month, 2 inference requests per second2026-09-07 — Added to the list: IBM’s watsonx.ai Runtime on its Lite plan — 300,000 tokens a month of foundation-model inference (Granite, Llama, Mistral and other hosted models) on IBM Cloud, a plan IBM’s own docs call free and never billGenerated from registry.yaml on 2026-09-18 and re-verified twice a week; the full list, the Atom feed and the machinery are at https://github.com/mvalentsev/awesome-free-ai-coding.