Free & unlimited AI — the complete list
Every provider offering free LLM access, across four tiers — permanent free, renewable, trial, and genuinely unlimited. Compiled by scraping provider directories directly, then verifying claims against each provider's own published limits.
01 The "unlimited" truth
You asked for free unlimited. This is the most important finding in the whole document, so it comes first.
No API provider offers truly unlimited free inference. Every hosted free tier in existence is metered — by requests per minute, requests per day, tokens per day, or a credit budget. Some market themselves as "unlimited", but on inspection they carry session quotas, 5-hour reset windows, or fair-use cutoffs.
There is exactly one genuinely unlimited free option: running the model yourself. Open-weight models carry no quota at all — no rate limit, no daily cap, no metering, no logging. The limit is your own hardware. That is why the local tier in §05 is the real answer to this request, and why every hosted tier should be treated as a quota to spend, not a well to draw from.
| Tier | What you get | Unlimited? | Providers |
|---|---|---|---|
| Permanent free | Free forever, no card, hard rate limits | No — metered | 19 |
| Renewable | Small daily allowance that resets | No — resets daily | 4 |
| Trial credits | One-time credits that expire — not free forever | No — expires | 17 |
| Self-hosted | Open weights on your own hardware | Yes — no quota at all | 9 tools |
02 Permanent free tiers — 19 providers
Free forever, no credit card, no expiring credits. Card status and rate limits are exactly as each provider publishes them.
| Provider | Card | Rate limit | Daily / monthly | Key models |
|---|---|---|---|---|
| Hetzner Inference | No | 3M in / 60K out per 60s | 500M in / 5M out per 24h | Qwen3.6 35B A3B |
| LLM7.io | No | 30 rpm anon · 120 rpm free token | up to 5M tokens/day | DeepSeek-R1 · Qwen 2.5 |
| Google AI Studio | No | 5–30 rpm by model | 9 000 rpd Flash · 25 rpd Pro | Gemini 3.1 Pro · Flash · Flash-Lite |
| Groq | No | 30 rpm | 14 400 requests/day | Qwen3.6 27B · MiniMax M2.7 · Whisper Large v3 |
| Cloudflare Workers AI | No | varies by model | 10 000 neurons/day · ~300k/month | Llama 3.1 8B · Llama 3.2 3B · Mistral 7B · Qwen 1.5 7B |
| Cohere | No | 20 rpm | 1 000 requests/month · non-commercial | Command A 111B · Command R+ · Command R7B |
| Hugging Face | No | 300 req/hour | $0.10/month credits (PRO $2) | Llama 3.2 11B Vision · Qwen 2.5 72B · Gemma 2 9B |
| Ollama Cloud | No | 1 concurrent model | session reset ~5h · weekly reset | gpt-oss 120B · Qwen3.5 · DeepSeek V4 Flash |
| Nous Portal | No | not published | $0/month free tier | Hermes 4 |
| Pollinations.ai | No | ~1 request / 15s anon | fair use | GPT-class · Mistral-class (proxied) |
| Inference.net | No | 30 rpm fair use | fair-use policy | DeepSeek-R1 · Llama 3.1 8B · Llama 3.1 70B |
| Aion Labs | No | 15 rpm | 20K tokens/day | aion-3.0 · aion-3.0-mini · aion-rp-llama-3.1-8b |
| OVH AI Endpoints | Reg. | 2 rpm anon · 400 rpm auth | 20+ EU-hosted models | Qwen3.5-397B · gpt-oss-120b · SDXL · TTS |
| Z.AI (GLM) | Reg. | ~1 req/sec | ~1 000 requests/day Flash | GLM-4.7-Flash · GLM-4.5-Flash · GLM-4.6V-Flash |
| Coze | Reg. | varies | token-based, resets daily | GPT-4o via Coze · Gemini 1.5 Pro via Coze |
| Venice.ai | Reg. | 10 rpm | limited daily usage | Llama 3.1 405B · Dolphin Mixtral · SD3 |
| Mistral La Plateforme | Phone | 1 request/second | free | Mistral Small · Nemo · 7B · Mixtral 8x7B |
| NVIDIA NIM | Phone | 40 rpm | 10 000 rpd · 100+ models | Nemotron 3 Super 120B · Ultra 550B · gpt-oss-120b |
| ModelScope | Phone | dynamic | 500 rpd model · 2 000 rpd total | Qwen3.5-35B-A3B · Qwen3.5-27B |
SiliconFlow also runs permanently free models (Qwen/Qwen3-8B at
1 000 rpm / 50 000 tpm) but gates them behind real-name identity verification — mainland-Chinese
documents only for international support.
03 Renewable credits — 4 providers
Free allowance that resets, so it keeps working indefinitely — but with a small daily ceiling rather than a permanent quota.
| Provider | Card | Rate limit | Free allowance | Key models |
|---|---|---|---|---|
| OpenRouter | No | 20 rpm | 50 requests/day · 1 000 with a one-time $10 top-up | 19 live :free models — see §06 |
| Requesty | No | 60 rpm | 200 requests/day across free models | free model pool |
| Grok (xAI) | Reg. | low on free tier | $25 one-time signup credit | Grok-2 · Grok-2 Mini · Grok-2 Vision |
| Venice.ai | Reg. | 10 rpm | limited daily usage, resets | Llama 3.1 405B · SD3 |
OpenRouter is the strongest aggregator here: a single key reaching 19 free models with independent quotas, plus a free-models router and fallback chaining. A one-time $10 raises the daily ceiling from 50 to 1 000 requests without making the models themselves paid.
04 Trial credits — 17 providers
Listed for completeness, and explicitly marked as not free forever. These are one-time credits that expire, so they fail the "permanent free" test.
| Provider | Credit | Expiry |
|---|---|---|
| Cerebrium | $30 | one-time |
| AI21 Labs | $10 — Jamba Large · Mini | 3 months |
| Upstage | $10 | 3 months |
| Friendli AI | $10 | one-time |
| SambaNova Cloud | $5 | 3 months |
| Cerebras | $5 — Llama 3.1 · Llama 4 Scout · Qwen3 32B | 30 days |
| DeepInfra | $5 | 90 days |
| Nscale | $5 | one-time |
| DeepSeek | 5M tokens | 30 days |
| Qwen (Alibaba Bailian) | 1M tokens per model | one-time |
| Scaleway Generative APIs | 1M tokens | one-time |
| Fireworks AI | $1 | one-time |
| Hyperbolic | $1 | one-time |
| Nebius Token Factory | $1 — card on file | one-time |
| Novita AI | $0.50 | one-time |
| Replicate | small trial credit | one-time |
| Together.AI | free research models — $5 deposit | — |
05 Unlimited: local & self-hosted
The only tier that is genuinely unlimited, private and free forever. Open weights running on hardware you control — no quota, no metering, no prompt logging.
| Tool | Type | Best for | Cost |
|---|---|---|---|
| Ollama | CLI + server | Fastest path to a local model; OpenAI-compatible API | Free |
| LM Studio | Desktop GUI | Browsing and chatting with models without a terminal | Free |
| llama.cpp | Inference engine | Maximum control and efficiency; runs on CPU and GPU | Free |
| GPT4All | Desktop app | Consumer hardware, private offline assistant | Free |
| Jan.ai | Desktop app | Open-source ChatGPT alternative, fully offline | Free |
| KoboldCpp | Single-file runtime | Creative writing and roleplay workloads | Free |
| llamafile | Single binary | One file that runs anywhere, no install | Free |
| Text Generation WebUI | Gradio UI | Advanced experimentation and customisation | Free |
| BentoML | Inference platform | Deploying any model as a production service | Free |
Models worth running locally
The open-weight families that make self-hosting viable — all downloadable with no account, no telemetry and no usage limit:
- Qwen — Qwen 3.8 / 3.5 series; best price-to-performance for general inference, and strong small variants
- DeepSeek — V4 family; reasoning and coding close to frontier quality
- Gemma — Gemma 4 26B / 31B; excellent quality per gigabyte for local machines
- gpt-oss — 120B and 20B mixture-of-experts; efficient and widely supported
- Llama — Llama 4 / 3.x; the broadest tooling and quantisation support
- Mistral — Large 3 / Small 4; efficient, strong multilingual
- GLM — GLM 5.3 / 4.7-Flash; strong agentic and multilingual behaviour
- Kimi — K3 / K2.6; long-context and agentic coding
- Nemotron — Nemotron 3 Super / Ultra / Nano; long context, MoE efficiency
- Phi — Phi-4 series; small models for constrained hardware
06 Live model inventory
Model lists pulled live from provider APIs during this research, rather than copied from directories — so these reflect what is actually served today.
| Model ID | Context | Daily limit |
|---|---|---|
| thinkingmachines/inkling:free | 1 048 576 | 50 rpd |
| thinkingmachines/inkling-small:free | 1 048 576 | 50 rpd |
| nvidia/nemotron-3.5-lightning:free | 1 000 000 | 50 rpd |
| nvidia/nemotron-3-ultra-550b-a55b:free | 1 000 000 | 50 rpd |
| dots-studio/dots-3-note-preview:free | 512 000 | 50 rpd |
| inclusionai/ling-3.0-flash-vl:free | 262 144 | 50 rpd |
| inclusionai/ling-3.0-flash-sante:free | 262 144 | 50 rpd |
| inclusionai/ling-3.0-flash-fin:free | 262 144 | 50 rpd |
| nex-agi/nex-n2.5-pro:free | 262 144 | 50 rpd |
| nex-agi/nex-n2.5-mini:free | 262 144 | 50 rpd |
| poolside/laguna-s-2.1:free | 262 144 | 50 rpd |
| poolside/laguna-xs-2.1:free | 262 144 | 50 rpd |
| google/gemma-4-26b-a4b-it:free | 262 144 | 50 rpd |
| google/gemma-4-31b-it:free | 262 144 | 50 rpd |
| nvidia/nemotron-3-super-120b-a12b:free | 262 144 | 50 rpd |
| cohere/north-mini-code:free | 256 000 | 50 rpd |
| nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free | 256 000 | 50 rpd |
| nvidia/nemotron-3.5-content-safety:free | 128 000 | 50 rpd |
| liquid/lfm-2.5-2.6b:free | 65 536 | 50 rpd |
| Provider | Free models | Verified by |
|---|---|---|
| OpenRouter | 19 of 445 total models are free | live API |
| NVIDIA NIM | 100+ models on the free tier | catalogue probe |
| Cloudflare Workers AI | 75+ models on the free tier | provider docs |
| Kilo Code | 11+ free models, no API key needed | provider docs |
| Hugging Face | thousands via inference providers | provider docs |
| Chutes.ai | 14 models — paid, $0.12–$3.00/M | scraped, excluded |
Chutes.ai is worth calling out because it circulates on free-API lists: scraping its catalogue showed per-token pricing on every model, so it does not belong in a free list.
07 Caveats & method
Read before relying on any of this
- Free tiers change monthly. Limits get cut and models get retired without notice. One provider's flash model is already announced for retirement while still listed as free.
- "Free" usually means prompts as training data. Gemini, Mistral, OpenRouter and Kilo all publish this. Never send confidential data through a free tier.
- Verification gates are common. Four providers here need phone or identity verification — and identity verification often supports mainland-Chinese documents only.
- Trial credits are not free. §04 is separated deliberately: those 17 providers expire.
- Fair-use is not unlimited. Where a provider says "fair use", there is an unpublished ceiling that can end access.
- Search snippets are unreliable. Several widely-cited "unlimited free" claims did not survive checking against provider documentation — which is exactly why this list was scraped from sources rather than summarised from search results.
How this was researched
Provider and model data was obtained by driving a headless browser pipeline against the source
repositories and provider catalogues, then parsing the raw content directly. Live model inventories
(OpenRouter, Chutes) were pulled from provider JSON APIs. Where a provider does not publish a
figure, this document shows — or "not published" rather than an estimate.
Sources
github.com/nejib1/Free-LLM— 41 providers, 120+ models, tiered by card requirement and credit type (primary)github.com/mnfst/awesome-free-llm-apis— permanent free tiers with per-model rate limitsopenrouter.ai/api/v1/models— live free-model inventory, fetched directlychutes.ai/app/chutes— scraped catalogue, verified as paid and excludedwotai.co/blog/best-free-llm-apis— live-probed free API directory