Deep research · All providers

Free & unlimited AI — the complete list

Every provider offering free LLM access, across four tiers — permanent free, renewable, trial, and genuinely unlimited. Compiled by scraping provider directories directly, then verifying claims against each provider's own published limits.

tiers 4 providers 40 free models 120+ self-host tools 9 compiled 2026-09-15
19Permanent free
4Renewable
17Trial only
9Unlimited (local)

01 The "unlimited" truth

You asked for free unlimited. This is the most important finding in the whole document, so it comes first.

No API provider offers truly unlimited free inference. Every hosted free tier in existence is metered — by requests per minute, requests per day, tokens per day, or a credit budget. Some market themselves as "unlimited", but on inspection they carry session quotas, 5-hour reset windows, or fair-use cutoffs.

There is exactly one genuinely unlimited free option: running the model yourself. Open-weight models carry no quota at all — no rate limit, no daily cap, no metering, no logging. The limit is your own hardware. That is why the local tier in §05 is the real answer to this request, and why every hosted tier should be treated as a quota to spend, not a well to draw from.

What "free" actually means, tier by tier
TierWhat you getUnlimited?Providers
Permanent free Free forever, no card, hard rate limits No — metered 19
Renewable Small daily allowance that resets No — resets daily 4
Trial credits One-time credits that expire — not free forever No — expires 17
Self-hosted Open weights on your own hardware Yes — no quota at all 9 tools
Bottom line For unlimited free AI, self-host. For convenience, stack several permanent free tiers — their limits are independent, so four providers give you four separate daily budgets.

02 Permanent free tiers — 19 providers

Free forever, no credit card, no expiring credits. Card status and rate limits are exactly as each provider publishes them.

All 19 permanent free providers
ProviderCardRate limitDaily / monthlyKey models
Hetzner InferenceNo 3M in / 60K out per 60s 500M in / 5M out per 24h Qwen3.6 35B A3B
LLM7.ioNo 30 rpm anon · 120 rpm free token up to 5M tokens/day DeepSeek-R1 · Qwen 2.5
Google AI StudioNo 5–30 rpm by model 9 000 rpd Flash · 25 rpd Pro Gemini 3.1 Pro · Flash · Flash-Lite
GroqNo 30 rpm 14 400 requests/day Qwen3.6 27B · MiniMax M2.7 · Whisper Large v3
Cloudflare Workers AINo varies by model 10 000 neurons/day · ~300k/month Llama 3.1 8B · Llama 3.2 3B · Mistral 7B · Qwen 1.5 7B
CohereNo 20 rpm 1 000 requests/month · non-commercial Command A 111B · Command R+ · Command R7B
Hugging FaceNo 300 req/hour $0.10/month credits (PRO $2) Llama 3.2 11B Vision · Qwen 2.5 72B · Gemma 2 9B
Ollama CloudNo 1 concurrent model session reset ~5h · weekly reset gpt-oss 120B · Qwen3.5 · DeepSeek V4 Flash
Nous PortalNo not published $0/month free tier Hermes 4
Pollinations.aiNo ~1 request / 15s anon fair use GPT-class · Mistral-class (proxied)
Inference.netNo 30 rpm fair use fair-use policy DeepSeek-R1 · Llama 3.1 8B · Llama 3.1 70B
Aion LabsNo 15 rpm 20K tokens/day aion-3.0 · aion-3.0-mini · aion-rp-llama-3.1-8b
OVH AI EndpointsReg. 2 rpm anon · 400 rpm auth 20+ EU-hosted models Qwen3.5-397B · gpt-oss-120b · SDXL · TTS
Z.AI (GLM)Reg. ~1 req/sec ~1 000 requests/day Flash GLM-4.7-Flash · GLM-4.5-Flash · GLM-4.6V-Flash
CozeReg. varies token-based, resets daily GPT-4o via Coze · Gemini 1.5 Pro via Coze
Venice.aiReg. 10 rpm limited daily usage Llama 3.1 405B · Dolphin Mixtral · SD3
Mistral La PlateformePhone 1 request/second free Mistral Small · Nemo · 7B · Mixtral 8x7B
NVIDIA NIMPhone 40 rpm 10 000 rpd · 100+ models Nemotron 3 Super 120B · Ultra 550B · gpt-oss-120b
ModelScopePhone dynamic 500 rpd model · 2 000 rpd total Qwen3.5-35B-A3B · Qwen3.5-27B

SiliconFlow also runs permanently free models (Qwen/Qwen3-8B at 1 000 rpm / 50 000 tpm) but gates them behind real-name identity verification — mainland-Chinese documents only for international support.

Most generous quotas found Hetzner Inference — 500M input tokens and 5M output tokens per 24 hours, no card, no billing system attached. LLM7.io — up to 5M tokens/day with a free token and no card. Google AI Studio — 9 000 requests/day on Flash models. These three alone exceed most people's realistic usage.

03 Renewable credits — 4 providers

Free allowance that resets, so it keeps working indefinitely — but with a small daily ceiling rather than a permanent quota.

ProviderCardRate limitFree allowanceKey models
OpenRouterNo20 rpm 50 requests/day · 1 000 with a one-time $10 top-up 19 live :free models — see §06
RequestyNo60 rpm 200 requests/day across free models free model pool
Grok (xAI)Reg.low on free tier $25 one-time signup credit Grok-2 · Grok-2 Mini · Grok-2 Vision
Venice.aiReg.10 rpm limited daily usage, resets Llama 3.1 405B · SD3

OpenRouter is the strongest aggregator here: a single key reaching 19 free models with independent quotas, plus a free-models router and fallback chaining. A one-time $10 raises the daily ceiling from 50 to 1 000 requests without making the models themselves paid.

04 Trial credits — 17 providers

Listed for completeness, and explicitly marked as not free forever. These are one-time credits that expire, so they fail the "permanent free" test.

ProviderCreditExpiry
Cerebrium$30one-time
AI21 Labs$10 — Jamba Large · Mini3 months
Upstage$103 months
Friendli AI$10one-time
SambaNova Cloud$53 months
Cerebras$5 — Llama 3.1 · Llama 4 Scout · Qwen3 32B30 days
DeepInfra$590 days
Nscale$5one-time
DeepSeek5M tokens30 days
Qwen (Alibaba Bailian)1M tokens per modelone-time
Scaleway Generative APIs1M tokensone-time
Fireworks AI$1one-time
Hyperbolic$1one-time
Nebius Token Factory$1 — card on fileone-time
Novita AI$0.50one-time
Replicatesmall trial creditone-time
Together.AIfree research models — $5 deposit
Watch out Several providers marketed as "free" belong here rather than in §02 — Cerebras, SambaNova, Together.AI and Fireworks all run one-time or expiring credits, not permanent free tiers. If a provider has an expiry date, it is not free forever.

05 Unlimited: local & self-hosted

The only tier that is genuinely unlimited, private and free forever. Open weights running on hardware you control — no quota, no metering, no prompt logging.

ToolTypeBest forCost
OllamaCLI + serverFastest path to a local model; OpenAI-compatible APIFree
LM StudioDesktop GUIBrowsing and chatting with models without a terminalFree
llama.cppInference engineMaximum control and efficiency; runs on CPU and GPUFree
GPT4AllDesktop appConsumer hardware, private offline assistantFree
Jan.aiDesktop appOpen-source ChatGPT alternative, fully offlineFree
KoboldCppSingle-file runtimeCreative writing and roleplay workloadsFree
llamafileSingle binaryOne file that runs anywhere, no installFree
Text Generation WebUIGradio UIAdvanced experimentation and customisationFree
BentoMLInference platformDeploying any model as a production serviceFree

Models worth running locally

The open-weight families that make self-hosting viable — all downloadable with no account, no telemetry and no usage limit:

  • Qwen — Qwen 3.8 / 3.5 series; best price-to-performance for general inference, and strong small variants
  • DeepSeek — V4 family; reasoning and coding close to frontier quality
  • Gemma — Gemma 4 26B / 31B; excellent quality per gigabyte for local machines
  • gpt-oss — 120B and 20B mixture-of-experts; efficient and widely supported
  • Llama — Llama 4 / 3.x; the broadest tooling and quantisation support
  • Mistral — Large 3 / Small 4; efficient, strong multilingual
  • GLM — GLM 5.3 / 4.7-Flash; strong agentic and multilingual behaviour
  • Kimi — K3 / K2.6; long-context and agentic coding
  • Nemotron — Nemotron 3 Super / Ultra / Nano; long context, MoE efficiency
  • Phi — Phi-4 series; small models for constrained hardware
Practical minimum A 7–9B model at 4-bit quantisation runs comfortably on 8 GB of unified memory; a 30B model needs roughly 20 GB. Everything in the list above has quantised builds for exactly this reason.

06 Live model inventory

Model lists pulled live from provider APIs during this research, rather than copied from directories — so these reflect what is actually served today.

OpenRouter — all 19 free models, fetched live from the public model API
Model IDContextDaily limit
thinkingmachines/inkling:free1 048 57650 rpd
thinkingmachines/inkling-small:free1 048 57650 rpd
nvidia/nemotron-3.5-lightning:free1 000 00050 rpd
nvidia/nemotron-3-ultra-550b-a55b:free1 000 00050 rpd
dots-studio/dots-3-note-preview:free512 00050 rpd
inclusionai/ling-3.0-flash-vl:free262 14450 rpd
inclusionai/ling-3.0-flash-sante:free262 14450 rpd
inclusionai/ling-3.0-flash-fin:free262 14450 rpd
nex-agi/nex-n2.5-pro:free262 14450 rpd
nex-agi/nex-n2.5-mini:free262 14450 rpd
poolside/laguna-s-2.1:free262 14450 rpd
poolside/laguna-xs-2.1:free262 14450 rpd
google/gemma-4-26b-a4b-it:free262 14450 rpd
google/gemma-4-31b-it:free262 14450 rpd
nvidia/nemotron-3-super-120b-a12b:free262 14450 rpd
cohere/north-mini-code:free256 00050 rpd
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free256 00050 rpd
nvidia/nemotron-3.5-content-safety:free128 00050 rpd
liquid/lfm-2.5-2.6b:free65 53650 rpd
Provider catalogue sizes verified during research
ProviderFree modelsVerified by
OpenRouter19 of 445 total models are freelive API
NVIDIA NIM100+ models on the free tiercatalogue probe
Cloudflare Workers AI75+ models on the free tierprovider docs
Kilo Code11+ free models, no API key neededprovider docs
Hugging Facethousands via inference providersprovider docs
Chutes.ai14 models — paid, $0.12–$3.00/Mscraped, excluded

Chutes.ai is worth calling out because it circulates on free-API lists: scraping its catalogue showed per-token pricing on every model, so it does not belong in a free list.

07 Caveats & method

Read before relying on any of this

  • Free tiers change monthly. Limits get cut and models get retired without notice. One provider's flash model is already announced for retirement while still listed as free.
  • "Free" usually means prompts as training data. Gemini, Mistral, OpenRouter and Kilo all publish this. Never send confidential data through a free tier.
  • Verification gates are common. Four providers here need phone or identity verification — and identity verification often supports mainland-Chinese documents only.
  • Trial credits are not free. §04 is separated deliberately: those 17 providers expire.
  • Fair-use is not unlimited. Where a provider says "fair use", there is an unpublished ceiling that can end access.
  • Search snippets are unreliable. Several widely-cited "unlimited free" claims did not survive checking against provider documentation — which is exactly why this list was scraped from sources rather than summarised from search results.

How this was researched

Provider and model data was obtained by driving a headless browser pipeline against the source repositories and provider catalogues, then parsing the raw content directly. Live model inventories (OpenRouter, Chutes) were pulled from provider JSON APIs. Where a provider does not publish a figure, this document shows or "not published" rather than an estimate.

Sources

  • github.com/nejib1/Free-LLM — 41 providers, 120+ models, tiered by card requirement and credit type (primary)
  • github.com/mnfst/awesome-free-llm-apis — permanent free tiers with per-model rate limits
  • openrouter.ai/api/v1/models — live free-model inventory, fetched directly
  • chutes.ai/app/chutes — scraped catalogue, verified as paid and excluded
  • wotai.co/blog/best-free-llm-apis — live-probed free API directory
Scope Only permanent, no-card, API-accessible free tiers qualify for §02. Web-UI-only free access, expiring credits and card-required tiers were deliberately separated out.