Free AI models & providers
Every provider here has a permanent free tier or is free outright — no trial credits, no time-limited promos. Compiled by searching and scraping live provider directories, then separating what is genuinely free from what quietly requires a card, a quota, or identity verification.
01 First-party APIs
APIs run by the companies that train or fine-tune the models themselves. All are OpenAI-SDK compatible unless noted, and none require a credit card.
| Model | Context | Max out | Modality | Free limit |
|---|---|---|---|---|
| Gemini 3.7 Flash | 1M | 65K | Text · Image · Audio · Video | — |
| Gemini 3.6 Flash | 1M | 65K | Text · Image · Audio · Video | 15 rpm · 1 500 rpd |
| Gemini 3.5 Flash-Lite | 1M | 65K | Text · Image · Audio · Video | 30 rpm · 1 500 rpd |
| Gemini 2.5 Pro | 1M | 65K | Text · Image · Audio · Video | 5 rpm · 50 rpd |
| Gemma 4 31B | 256K | 32K | Text | — |
| Model | Context | Modality | Free limit |
|---|---|---|---|
| Mistral Large 3 | 256K | Multimodal | ~1 rps · 500K tpm |
| Mistral Medium 3.5 · 128B | 256K | Text · Image · Code | ~1 rps · 500K tpm |
| Mistral Small 4 | 256K | Text · Image · Code | ~1 rps · 500K tpm |
| Codestral | 128K | Code | ~1 rps · 500K tpm |
| Ministral 3 · 3B / 8B / 14B | 256K | Text · Vision | ~1 rps · 500K tpm |
Worth knowing: the Mistral free mode grants roughly $10/month in credits and is on by default — but free-mode prompts may be used to train Mistral models unless you opt out.
| Model | Context | Max out | Modality | Free limit |
|---|---|---|---|---|
| GLM-4.7-Flash | 200K | 128K | Text · reasoning | 1 concurrent |
| GLM-4.6V-Flash | 128K | 32K | Multimodal | 1 concurrent |
| GLM-4.5-Flash | 128K | 96K | Text · reasoning | 1 concurrent · retiring |
The same free models are served internationally at api.z.ai/api/paas/v4. Registration
accepts overseas numbers and the chat API does not require real-name verification.
| Model | Context | Modality | Notes |
|---|---|---|---|
| Command A+ · 218B | 128K | Text · Image | 20 rpm |
| Command A · 111B | 256K | Text | 20 rpm |
| Command A Reasoning | 256K | Text · reasoning | 20 rpm |
| Command A Vision | 128K | Text · Image | 20 rpm |
| Aya Expanse 32B · Aya Vision 32B | 128K / 16K | Text / Text · Image | 20 rpm |
non-commercial Cohere's trial key allows 1 000 calls/month for non-commercial use only. No credit card required.
| Model | Context | Modality | Free limit |
|---|---|---|---|
| aion-labs/aion-3.0 | 128K | Text · reasoning | 15 rpm · 20K tpd |
| aion-labs/aion-3.0-mini | 128K | Text · reasoning | 15 rpm · 20K tpd |
| aion-labs/aion-rp-llama-3.1-8b | 32K | Text · roleplay | 15 rpm · 20K tpd |
No credit card. Specialised for roleplay and storytelling, with a small daily token budget.
02 Inference hosts
Third-party platforms hosting open-weight models from many sources. This is where the widest free model coverage lives.
| Model | Context | Max out | Free limit |
|---|---|---|---|
| openai/gpt-oss-120b | 131K | 65K | 30 rpm · 1 000 rpd |
| openai/gpt-oss-20b | 131K | 65K | 30 rpm · 1 000 rpd |
| qwen/qwen3.6-27b | 131K | 16K | 30 rpm · 1 000 rpd |
| groq/compound · compound-mini | 131K | 8K | 30 rpm · 250 rpd |
| Model | Context | Max out | Free limit |
|---|---|---|---|
| nvidia/nemotron-3-super-120b-a12b | 1M | 262K | 40 rpm · 10 000 rpd |
| nvidia/nemotron-3-ultra-550b-a55b | 1M | 262K | 40 rpm · 10 000 rpd |
| openai/gpt-oss-120b · gpt-oss-20b | 131K | 131K | 40 rpm · 10 000 rpd |
| mistralai/mistral-nemotron | 128K | 8K | 40 rpm · 10 000 rpd |
| google/gemma-4-31b-it | 262K | 8K | 40 rpm · 10 000 rpd |
| + 92 more | — | — | text · image · video · speech |
| Model | Context | Modality | Free limit |
|---|---|---|---|
| @cf/openai/gpt-oss-120b | 128K | Text | shared 10K neurons/day |
| @cf/meta/llama-4-scout-17b-16e-instruct | 131K | Multimodal | shared |
| @cf/google/gemma-4-26b-a4b-it | 256K | Text · Vision | shared |
| @cf/zai-org/glm-4.7-flash | 131K | Text | shared |
| @cf/deepseek-ai/deepseek-r1-distill-qwen-32b | 80K | Text · reasoning | shared |
| + 72 more | — | Text · Image · Audio · Embeddings | shared |
The daily neuron budget is shared across all usage, not per model, and resets at 00:00 UTC. Exceeding it fails the request rather than billing you. Five models are excluded from the free plan.
| Model | Context | Modality | Free limit |
|---|---|---|---|
| nvidia/nemotron-3-super-120b-a12b:free | 262K | Text | 20 rpm · 50 rpd |
| openai/gpt-oss-20b:free | 131K | Text | 20 rpm · 50 rpd |
| google/gemma-4-31b-it:free | 262K | Text · Image | 20 rpm · 50 rpd |
| cohere/north-mini-code:free | 256K | Text · code | 20 rpm · 50 rpd |
| poolside/laguna-xs-2.1:free | 262K | Text · code | 20 rpm · 50 rpd |
| nvidia/nemotron-nano-12b-v2-vl:free | 128K | Text · Image | 20 rpm · 50 rpd |
| + 11 more free models | — | Text · Image | 20 rpm · 50 rpd |
Free models allow 50 requests/day each. A one-time $10 credit purchase unlocks 1 000 rpd while keeping the models free. There is also a free-models router and model-fallback chaining.
| Model | Context | Free limit |
|---|---|---|
| deepseek-v4-pro · deepseek-v4-flash | 1M | session · weekly |
| kimi-k3 | 1M | session · weekly |
| minimax-m3 | 512K | session · weekly |
| gpt-oss:120b · gpt-oss:20b | 128K · 131K | session · weekly |
| mistral-large-3:675b | 256K | session · weekly |
| qwen3.5:397b | 256K | session · weekly |
Usage is measured by weighted input plus output tokens. Session limits reset every 5 hours, weekly limits every 7 days. Exact numbers are not published.
| Model | Context | Max out | Modality |
|---|---|---|---|
| nvidia/nemotron-3-ultra-550b-a55b:free | 1M | 65K | Text |
| nvidia/nemotron-3.5-lightning:free | 1M | 65K | Text |
| stepfun/step-3.7-flash:free | 262K | 262K | Text · Vision |
| poolside/laguna-s-2.1:free · laguna-xs-2.1:free | 262K | 32K | Text · code |
| tencent/hy3:free | 262K | 128K | Text |
| nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free | 256K | 65K | Multimodal |
Reachable with no API key at all. An auto-router picks from the free pool. The catalogue lags what is actually served, so models absent from it have still been observed answering.
| Route | Modality | Free limit |
|---|---|---|
| Inference providers — Fireworks · Together · Hyperbolic · Nebius · Novita · DeepInfra | Text · Image · Audio | $0.10/month credits |
| Meta-Llama-3.1-8B-Instruct | Text · 128K | credit-metered |
| phi-4 · Qwen2.5-Coder-7B · Qwen2.5-7B | Text | credit-metered |
| + thousands of community models | Text · Image · Audio · Embeddings | credit-metered |
| Model | Context | Free limit |
|---|---|---|
| Qwen/Qwen3.5-35B-A3B | 256K | ≤500 rpd · 2 000 rpd total |
| Qwen/Qwen3.5-27B | 256K | ≤500 rpd · 2 000 rpd total |
verification Requires binding an Alibaba Cloud account and passing real-name verification.
| Model | Context | Modality | Anonymous limit |
|---|---|---|---|
| gpt-oss:20b | 128K | Text | 10 rpm · 60 req/hr |
| mistral-Nemo-Instruct-2407 | 128K | Text | 10 rpm · 60 req/hr |
| minimax-m2.7 | 180K | Text · reasoning | 10 rpm · 60 req/hr |
A free token (no card) raises limits to 2 rps, 40 rpm, 100 req/hr and 1M tokens/day. The catalogue rotates.
| Model | Context | Free limit |
|---|---|---|
| Qwen/Qwen3-8B | 128K | 1 000 rpm · 50 000 tpm |
verification Real-name verification is required. Verification supports mainland-Chinese documents; international users must contact support. Most of the 100+ catalogue is paid.
03 No key, no signup
Genuinely zero-friction: no account, no API key, no card. Useful for testing and low-volume work, with tight rate limits.
| Provider | Endpoint | Anonymous limit |
|---|---|---|
| OVHcloud AI Endpoints · FR | oai.endpoints.kepler.ai.cloud.ovh.net/v1 | 2 rpm per IP, per model |
| LLM7.io · GB | api.llm7.io/v1 | 10 rpm · 60 req/hr |
| Kilo Code · US | api.kilo.ai/api/gateway | 200 req/hr per IP |
OVHcloud serves 20+ open-weight models from EU data centres on this anonymous tier, including Qwen3.5-397B-A17B, gpt-oss-120b, Qwen3-Coder-30B-A3B, Qwen2.5-VL-72B, Mistral-Small-3.2-24B and Llama-3.3-70B. Higher limits need a key and are billed per token.
04 Open-weight models
Weights you can download, inspect and run on your own hardware — no account, no metering, no telemetry. Most of what the free tiers above host originates here.
| Family | Latest generation | Strength | Licence |
|---|---|---|---|
| Qwen · Alibaba | Qwen 3.8 · 3.5 series | Best price-to-performance for general inference | Open weights |
| DeepSeek | DeepSeek V4 | Reasoning and coding at frontier-adjacent quality | Open weights |
| GLM · Zhipu | GLM-5.3 · 4.7-Flash | Strong multilingual and agentic behaviour | Open weights |
| Kimi · Moonshot | Kimi K3 · K2.6 | Long-context and agentic coding | Open weights |
| gpt-oss · OpenAI | gpt-oss-120b · 20b | Efficient MoE, widely hosted on free tiers | Open weights |
| Gemma · Google | Gemma 4 · 26B / 31B | Strong small-model quality for local use | Open weights |
| Llama · Meta | Llama 4 · 3.x | Broad ecosystem and tooling support | Community licence |
| Mistral · FR | Mistral Large 3 · Small 4 | Efficient European models, strong multilingual | Open / research |
| Nemotron · NVIDIA | Nemotron 3 · Super / Ultra / Nano | MoE families, long context, free on NVIDIA NIM | Open weights |
| Phi · Microsoft | Phi-4 series | Small models for constrained hardware | Open weights |
Practically: open-weight means the weights are free to download — it does not always mean a fully permissive data licence. Check the licence before commercial use, particularly for Llama and some Mistral releases.
05 Image & video generation
Free and open-weight generators you can self-host or reach through the hosts above.
| Model | Type | Notes |
|---|---|---|
| FLUX.2 Dev | Image · text-to-image | Widely rated the most balanced open image model in 2026 |
| Stable Diffusion 3 / SDXL | Image · text-to-image | The longest-standing open family, huge ecosystem |
| Qwen-Image | Image · text-to-image | Strong prompt adherence and text rendering |
| Z-Image | Image | Efficient variant aimed at lower VRAM |
| Wan | Video | Open video generation |
Free hosted image generation also appears inside several providers above — Cloudflare Workers AI serves image and audio models on the shared neuron budget, and NVIDIA NIM's catalogue covers image, video and speech alongside text.
06 Free chat apps
No API involved — consumer products with permanent free tiers, useful when you need a capable assistant without wiring anything up.
| Product | Vendor | Card required |
|---|---|---|
| ChatGPT free | OpenAI | no |
| Claude free | Anthropic | no |
| Gemini free | no | |
| Perplexity free | Perplexity | no |
| GitHub Copilot free | GitHub | no |
| DeepSeek chat | DeepSeek | no |
| Le Chat free | Mistral AI | no |
| Qwen Chat | Alibaba | no |
| Kimi | Moonshot | no |
Free consumer tiers are usually rate-limited during peak hours and may apply lower model versions than the paid plan. They are not a substitute for an API when building.
07 Caveats & sources
Read before relying on any free tier
- Free tiers change without notice. Models get retired and limits get cut — Groq retired two Llama models in August 2026, and one provider here has a flash model already announced for retirement.
- "Free" often means your prompts are training data. Gemini, Mistral, OpenRouter and Kilo all publish this caveat. Do not send confidential data through a free tier.
- Verification gates are real. Two providers listed require identity verification — one of them supporting mainland-Chinese documents only.
- Regional restrictions apply. The Gemini free tier has region-specific terms around the EEA, Switzerland and the UK.
- Catalogue ≠ availability. At least one gateway serves models that do not appear in its own model list, and conversely lists models that no longer answer.
Verification method
Provider lists were gathered by web search and then scraped directly from the source pages using a
headless browser pipeline, rather than copied from search summaries. Model rows, context windows and rate
limits come from the scraped provider tables. Where a value is genuinely unpublished, this document shows
— rather than an estimate.
Sources
github.com/mnfst/awesome-free-llm-apis— permanent free tiers, provider-by-provider tables (primary source)wotai.co/blog/best-free-llm-apis— live-probed free API directory, 110 models across 16 providersopenrouter.ai/blog/tutorials/free-llm-apis-compared— free-tier comparisonkdnuggets.com·dataiku.com/blog— free LLM API overviewscomputingforgeeks.com/open-source-llm-comparison— open-weight model comparisonthundercompute.com·bentoml.com— open image-generation models