OpenAI
commercialMaker of GPT and the OpenAI API. Offers frontier models (GPT-5.5, GPT-5.4), reasoning (o4-mini), and the open-source gpt-oss famil…
31 providers in the directory — commercial labs, open-source model makers, and inference hosting platforms. Model counts and cheapest prices come from the current price board.
Maker of GPT and the OpenAI API. Offers frontier models (GPT-5.5, GPT-5.4), reasoning (o4-mini), and the open-source gpt-oss famil…
Maker of the Claude model family. Flagship Fable 5 and Opus 5, plus Sonnet and Haiku tiers. Known for safety research.
Maker of Gemini multimodal models with up to 2M context. Current flagship Gemini 3.5 Flash; also open-weights Gemma.
Maker of Grok models. Current flagship Grok 4.3 offers 1M context at $1.25/$2.50 per 1M tokens — strong value.
Chinese AI lab offering extremely cheap APIs with aggressive prompt caching. V4 Flash at $0.14/$0.28 per 1M tokens.
European AI lab. Mistral Large 3 at $0.50/$1.50 per 1M tokens, plus Codestral for code and open-weight models.
Maker of the Nova model family on AWS. Nova Micro at $0.035/$0.14 per 1M tokens is among the cheapest hosted APIs.
Enterprise-focused lab. Command A flagship at $2.50/$10 per 1M tokens with a 256K context — the same list price as the older Comma…
Maker of the open-weight Llama family. Llama 4 Scout (17B, 16 experts, 10M context) and Llama 3.3 70B are widely hosted.
Inference platform running open models on custom LPU chips for extreme speed. Up to 1000 tokens/sec, cheapest hosted prices for ma…
Inference platform hosting 200+ open-source models with simple per-token pricing. Strong coverage of Qwen, GLM, Kimi, MiniMax, and…
Inference platform serving open-weight models (Llama, Qwen, DeepSeek, Kimi, GLM, MiniMax, gpt-oss) on its own infrastructure at fi…
Search-augmented AI from Perplexity. The Sonar family (Sonar, Sonar Pro, Sonar Reasoning Pro, Sonar Deep Research) adds live web g…
NVIDIA's Nemotron models for enterprise and agentic use — Llama Nemotron Ultra 253B and Nemotron 3 Ultra, from $0.60/$3.60 per 1M …
IBM's open-weight Granite 4 family for enterprise workloads — Granite 4 H Large down to Granite 4.0 Micro at $0.017/1M input token…
Israeli AI lab behind the Jamba hybrid SSM-Transformer models — Jamba Large and Jamba Mini ($0.20/$0.40 per 1M) with a 256K-token …
Multimodal AI lab. Reka Core, Flash, and Edge span frontier to on-device inference, with Reka Edge from $0.10/1M tokens.
Voyage AI (a MongoDB company) builds high-accuracy embedding and reranking models for retrieval and RAG, billed per 1M input token…
Alibaba's Qwen models — Qwen3-Max, Qwen-Plus, and Qwen-Turbo (from $0.05/$0.20 per 1M) with context windows up to 1M tokens.
Z.AI (formerly Zhipu AI) — makers of the GLM model family. Frontier and open-weight models with competitive pricing.
Chinese AI lab behind the Kimi models. Kimi K2.5 through K2.7 Code offer a 256K-token context from $0.60/1M input tokens.
Chinese AI lab. The MiniMax-M2 series (M2 through M2.7, plus M3 at 1M context) starts at $0.30/$1.20 per 1M tokens.
Founded by Kai-Fu Lee. Yi Large is a bilingual English/Chinese model at $3/$9 per 1M tokens with a 32K context window.
Chinese AI lab. Baichuan M2-32B is an open-weight 32B model priced at $0.07/1M for both input and output tokens.
Chinese technology conglomerate. Its Hunyuan family is served via Tencent Cloud's TokenHub API; Hunyuan Hy3 is a 295B-parameter Mo…
Serverless inference platform serving 100+ open-weight models (DeepSeek, Llama, Qwen, Gemma, Mistral) at pay-per-token rates — one…
Chinese technology company whose LongCat model family is served via the LongCat API Platform. LongCat-2.0 is offered pay-as-you-go…
AI research lab founded by former OpenAI CTO Mira Murati. Its Inkling family ships as open weights on Hugging Face and is also sol…
Tokyo research company founded by David Ha, Llion Jones and Ren Ito, known for nature-inspired and evolutionary approaches to mode…
Korean AI company behind the Solar model family and a document-intelligence suite (Document Parse, Information Extract). Solar Pro…
San Francisco lab building small, code-specific models for coding agents rather than general-purpose chat — retrieval, fast merge …