Reference · Free tiers only

Free AI models & providers

Every provider here has a permanent free tier or is free outright — no trial credits, no time-limited promos. Compiled by searching and scraping live provider directories, then separating what is genuinely free from what quietly requires a card, a quota, or identity verification.

scoped to 2026 providers 16 models listed 120+ currency USD compiled 2026-09-15
16Providers
5First-party APIs
11Inference hosts
3No key at all

01 First-party APIs

APIs run by the companies that train or fine-tune the models themselves. All are OpenAI-SDK compatible unless noted, and none require a credit card.

Google Gemini — free tier · US · generativelanguage.googleapis.com/v1beta
ModelContextMax outModalityFree limit
Gemini 3.7 Flash1M65KText · Image · Audio · Video
Gemini 3.6 Flash1M65KText · Image · Audio · Video15 rpm · 1 500 rpd
Gemini 3.5 Flash-Lite1M65KText · Image · Audio · Video30 rpm · 1 500 rpd
Gemini 2.5 Pro1M65KText · Image · Audio · Video5 rpm · 50 rpd
Gemma 4 31B256K32KText
Mistral AI — free mode · FR · api.mistral.ai/v1
ModelContextModalityFree limit
Mistral Large 3256KMultimodal~1 rps · 500K tpm
Mistral Medium 3.5 · 128B256KText · Image · Code~1 rps · 500K tpm
Mistral Small 4256KText · Image · Code~1 rps · 500K tpm
Codestral128KCode~1 rps · 500K tpm
Ministral 3 · 3B / 8B / 14B256KText · Vision~1 rps · 500K tpm

Worth knowing: the Mistral free mode grants roughly $10/month in credits and is on by default — but free-mode prompts may be used to train Mistral models unless you opt out.

Z AI (Zhipu) — permanently free models · CN · open.bigmodel.cn/api/paas/v4
ModelContextMax outModalityFree limit
GLM-4.7-Flash200K128KText · reasoning1 concurrent
GLM-4.6V-Flash128K32KMultimodal1 concurrent
GLM-4.5-Flash128K96KText · reasoning1 concurrent · retiring

The same free models are served internationally at api.z.ai/api/paas/v4. Registration accepts overseas numbers and the chat API does not require real-name verification.

Cohere — trial key · CA · api.cohere.com/v2
ModelContextModalityNotes
Command A+ · 218B128KText · Image20 rpm
Command A · 111B256KText20 rpm
Command A Reasoning256KText · reasoning20 rpm
Command A Vision128KText · Image20 rpm
Aya Expanse 32B · Aya Vision 32B128K / 16KText / Text · Image20 rpm

non-commercial Cohere's trial key allows 1 000 calls/month for non-commercial use only. No credit card required.

Aion Labs — permanent free tier · IL · api.aionlabs.ai/v1
ModelContextModalityFree limit
aion-labs/aion-3.0128KText · reasoning15 rpm · 20K tpd
aion-labs/aion-3.0-mini128KText · reasoning15 rpm · 20K tpd
aion-labs/aion-rp-llama-3.1-8b32KText · roleplay15 rpm · 20K tpd

No credit card. Specialised for roleplay and storytelling, with a small daily token budget.

02 Inference hosts

Third-party platforms hosting open-weight models from many sources. This is where the widest free model coverage lives.

Groq — free tier · US · api.groq.com/openai/v1 · ultra-fast LPU inference
ModelContextMax outFree limit
openai/gpt-oss-120b131K65K30 rpm · 1 000 rpd
openai/gpt-oss-20b131K65K30 rpm · 1 000 rpd
qwen/qwen3.6-27b131K16K30 rpm · 1 000 rpd
groq/compound · compound-mini131K8K30 rpm · 250 rpd
NVIDIA NIM — free with Developer Program · US · integrate.api.nvidia.com/v1 · 100+ models
ModelContextMax outFree limit
nvidia/nemotron-3-super-120b-a12b1M262K40 rpm · 10 000 rpd
nvidia/nemotron-3-ultra-550b-a55b1M262K40 rpm · 10 000 rpd
openai/gpt-oss-120b · gpt-oss-20b131K131K40 rpm · 10 000 rpd
mistralai/mistral-nemotron128K8K40 rpm · 10 000 rpd
google/gemma-4-31b-it262K8K40 rpm · 10 000 rpd
+ 92 moretext · image · video · speech
Cloudflare Workers AI — 10 000 neurons/day · US · api.cloudflare.com/…/ai/run · 75+ models
ModelContextModalityFree limit
@cf/openai/gpt-oss-120b128KTextshared 10K neurons/day
@cf/meta/llama-4-scout-17b-16e-instruct131KMultimodalshared
@cf/google/gemma-4-26b-a4b-it256KText · Visionshared
@cf/zai-org/glm-4.7-flash131KTextshared
@cf/deepseek-ai/deepseek-r1-distill-qwen-32b80KText · reasoningshared
+ 72 moreText · Image · Audio · Embeddingsshared

The daily neuron budget is shared across all usage, not per model, and resets at 00:00 UTC. Exceeding it fails the request rather than billing you. Five models are excluded from the free plan.

OpenRouter — 17 free models · US · openrouter.ai/api/v1
ModelContextModalityFree limit
nvidia/nemotron-3-super-120b-a12b:free262KText20 rpm · 50 rpd
openai/gpt-oss-20b:free131KText20 rpm · 50 rpd
google/gemma-4-31b-it:free262KText · Image20 rpm · 50 rpd
cohere/north-mini-code:free256KText · code20 rpm · 50 rpd
poolside/laguna-xs-2.1:free262KText · code20 rpm · 50 rpd
nvidia/nemotron-nano-12b-v2-vl:free128KText · Image20 rpm · 50 rpd
+ 11 more free modelsText · Image20 rpm · 50 rpd

Free models allow 50 requests/day each. A one-time $10 credit purchase unlocks 1 000 rpd while keeping the models free. There is also a free-models router and model-fallback chaining.

Ollama Cloud — free tier · US · ollama.com/api · OpenAI-compatible at ollama.com/v1
ModelContextFree limit
deepseek-v4-pro · deepseek-v4-flash1Msession · weekly
kimi-k31Msession · weekly
minimax-m3512Ksession · weekly
gpt-oss:120b · gpt-oss:20b128K · 131Ksession · weekly
mistral-large-3:675b256Ksession · weekly
qwen3.5:397b256Ksession · weekly

Usage is measured by weighted input plus output tokens. Session limits reset every 5 hours, weekly limits every 7 days. Exact numbers are not published.

Kilo Code — free pool, no API key · US · api.kilo.ai/api/gateway · 200 req/hour
ModelContextMax outModality
nvidia/nemotron-3-ultra-550b-a55b:free1M65KText
nvidia/nemotron-3.5-lightning:free1M65KText
stepfun/step-3.7-flash:free262K262KText · Vision
poolside/laguna-s-2.1:free · laguna-xs-2.1:free262K32KText · code
tencent/hy3:free262K128KText
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free256K65KMultimodal

Reachable with no API key at all. An auto-router picks from the free pool. The catalogue lags what is actually served, so models absent from it have still been observed answering.

Hugging Face — inference credits · US · router.huggingface.co/v1
RouteModalityFree limit
Inference providers — Fireworks · Together · Hyperbolic · Nebius · Novita · DeepInfraText · Image · Audio$0.10/month credits
Meta-Llama-3.1-8B-InstructText · 128Kcredit-metered
phi-4 · Qwen2.5-Coder-7B · Qwen2.5-7BTextcredit-metered
+ thousands of community modelsText · Image · Audio · Embeddingscredit-metered
ModelScope — free inference · CN · api-inference.modelscope.cn/v1
ModelContextFree limit
Qwen/Qwen3.5-35B-A3B256K≤500 rpd · 2 000 rpd total
Qwen/Qwen3.5-27B256K≤500 rpd · 2 000 rpd total

verification Requires binding an Alibaba Cloud account and passing real-name verification.

LLM7.io — free gateway · GB · api.llm7.io/v1
ModelContextModalityAnonymous limit
gpt-oss:20b128KText10 rpm · 60 req/hr
mistral-Nemo-Instruct-2407128KText10 rpm · 60 req/hr
minimax-m2.7180KText · reasoning10 rpm · 60 req/hr

A free token (no card) raises limits to 2 rps, 40 rpm, 100 req/hr and 1M tokens/day. The catalogue rotates.

SiliconFlow — permanently free models · CN · api.siliconflow.cn/v1
ModelContextFree limit
Qwen/Qwen3-8B128K1 000 rpm · 50 000 tpm

verification Real-name verification is required. Verification supports mainland-Chinese documents; international users must contact support. Most of the 100+ catalogue is paid.

03 No key, no signup

Genuinely zero-friction: no account, no API key, no card. Useful for testing and low-volume work, with tight rate limits.

ProviderEndpointAnonymous limit
OVHcloud AI Endpoints · FR oai.endpoints.kepler.ai.cloud.ovh.net/v1 2 rpm per IP, per model
LLM7.io · GB api.llm7.io/v1 10 rpm · 60 req/hr
Kilo Code · US api.kilo.ai/api/gateway 200 req/hr per IP

OVHcloud serves 20+ open-weight models from EU data centres on this anonymous tier, including Qwen3.5-397B-A17B, gpt-oss-120b, Qwen3-Coder-30B-A3B, Qwen2.5-VL-72B, Mistral-Small-3.2-24B and Llama-3.3-70B. Higher limits need a key and are billed per token.

Privacy note Free tiers commonly log prompts. Kilo's own documentation warns its router may send requests to providers that log prompts and outputs, and that NVIDIA's free endpoints are marked trial use — do not submit personal or confidential data. Gemini, Mistral and OpenRouter publish similar caveats.

04 Open-weight models

Weights you can download, inspect and run on your own hardware — no account, no metering, no telemetry. Most of what the free tiers above host originates here.

FamilyLatest generationStrengthLicence
Qwen · AlibabaQwen 3.8 · 3.5 seriesBest price-to-performance for general inferenceOpen weights
DeepSeekDeepSeek V4Reasoning and coding at frontier-adjacent qualityOpen weights
GLM · ZhipuGLM-5.3 · 4.7-FlashStrong multilingual and agentic behaviourOpen weights
Kimi · MoonshotKimi K3 · K2.6Long-context and agentic codingOpen weights
gpt-oss · OpenAIgpt-oss-120b · 20bEfficient MoE, widely hosted on free tiersOpen weights
Gemma · GoogleGemma 4 · 26B / 31BStrong small-model quality for local useOpen weights
Llama · MetaLlama 4 · 3.xBroad ecosystem and tooling supportCommunity licence
Mistral · FRMistral Large 3 · Small 4Efficient European models, strong multilingualOpen / research
Nemotron · NVIDIANemotron 3 · Super / Ultra / NanoMoE families, long context, free on NVIDIA NIMOpen weights
Phi · MicrosoftPhi-4 seriesSmall models for constrained hardwareOpen weights

Practically: open-weight means the weights are free to download — it does not always mean a fully permissive data licence. Check the licence before commercial use, particularly for Llama and some Mistral releases.

05 Image & video generation

Free and open-weight generators you can self-host or reach through the hosts above.

ModelTypeNotes
FLUX.2 DevImage · text-to-imageWidely rated the most balanced open image model in 2026
Stable Diffusion 3 / SDXLImage · text-to-imageThe longest-standing open family, huge ecosystem
Qwen-ImageImage · text-to-imageStrong prompt adherence and text rendering
Z-ImageImageEfficient variant aimed at lower VRAM
WanVideoOpen video generation

Free hosted image generation also appears inside several providers above — Cloudflare Workers AI serves image and audio models on the shared neuron budget, and NVIDIA NIM's catalogue covers image, video and speech alongside text.

Best route for zero cost For image and video at genuinely no cost, self-hosting open weights is the most durable option: no quota, no logging, no rate limit. The free hosted tiers are better suited to occasional or prototype use.

06 Free chat apps

No API involved — consumer products with permanent free tiers, useful when you need a capable assistant without wiring anything up.

ProductVendorCard required
ChatGPT freeOpenAIno
Claude freeAnthropicno
Gemini freeGoogleno
Perplexity freePerplexityno
GitHub Copilot freeGitHubno
DeepSeek chatDeepSeekno
Le Chat freeMistral AIno
Qwen ChatAlibabano
KimiMoonshotno

Free consumer tiers are usually rate-limited during peak hours and may apply lower model versions than the paid plan. They are not a substitute for an API when building.

07 Caveats & sources

Read before relying on any free tier

  • Free tiers change without notice. Models get retired and limits get cut — Groq retired two Llama models in August 2026, and one provider here has a flash model already announced for retirement.
  • "Free" often means your prompts are training data. Gemini, Mistral, OpenRouter and Kilo all publish this caveat. Do not send confidential data through a free tier.
  • Verification gates are real. Two providers listed require identity verification — one of them supporting mainland-Chinese documents only.
  • Regional restrictions apply. The Gemini free tier has region-specific terms around the EEA, Switzerland and the UK.
  • Catalogue ≠ availability. At least one gateway serves models that do not appear in its own model list, and conversely lists models that no longer answer.

Verification method

Provider lists were gathered by web search and then scraped directly from the source pages using a headless browser pipeline, rather than copied from search summaries. Model rows, context windows and rate limits come from the scraped provider tables. Where a value is genuinely unpublished, this document shows rather than an estimate.

Sources

  • github.com/mnfst/awesome-free-llm-apis — permanent free tiers, provider-by-provider tables (primary source)
  • wotai.co/blog/best-free-llm-apis — live-probed free API directory, 110 models across 16 providers
  • openrouter.ai/blog/tutorials/free-llm-apis-compared — free-tier comparison
  • kdnuggets.com · dataiku.com/blog — free LLM API overviews
  • computingforgeeks.com/open-source-llm-comparison — open-weight model comparison
  • thundercompute.com · bentoml.com — open image-generation models
Scope Only permanent free offerings are listed. Trial credits, time-limited promotions and free-to-start-but-paid-later products were excluded by design.