Answer

What is the cheapest LLM API in 2026?

By standard list price per token, OpenAI's GPT-6 Luna is the cheapest of the major frontier-vendor APIs we checked on September 25, 2026, at 0.10 US dollars per million input tokens and 0.50 per million output. DeepSeek Flash is close behind and cheaper on cached input and during off-peak hours, and Google's Gemini API has a free tier with a data-use condition. The cheapest price per token is not always the cheapest bill, so this answer shows how caching, batch discounts, time-of-day pricing and tokenizers change the result.

Published · Updated · Evidence-linked, not search-volume ranked.

Short answer

On standard list prices checked on September 25, 2026, OpenAI's GPT-6 Luna is the cheapest API from the major model vendors we compared: 0.10 US dollars per million input tokens, 0.01 for cached input and 0.50 per million output tokens. DeepSeek's deepseek-flash model costs 0.30 input and 1.20 output at peak hours and half that off-peak (0.15 and 0.60), and its cached input is the cheapest of the group at 0.003 off-peak. Google's Gemini 3.1 Flash-Lite lists at 0.25 input and 1.50 output, and the Gemini API has a free tier, but Google says free-tier content is used to improve its products. Anthropic's cheapest current model, Claude Haiku 4.5, lists at 1 and 5. Three things change the ranking for real workloads: how much of your input is cached, whether you can use batch pricing (half price at OpenAI, Anthropic and Google), and the time of day for DeepSeek. The cheapest price per token is also not the same as the cheapest cost per finished task, so run your own tasks on the two cheapest candidates before committing.

Why this question is current

Exact query-volume data was unavailable, so RepoRadar uses these as current demand and intent signals rather than a claimed volume ranking.

  • cheapest llm api · Google Suggest · US; English · checked 2026-09-25T22:26:30Z
    Observed completions: cheapest llm api, cheapest llm api provider, cheapest llm api reddit, cheapest llm api key, cheapest llm api pricing, cheapest llm api for coding, cheapest llm api for openclaw. A formulation signal captured at this time, not a volume or ranking claim.
  • best llm for the money · Google Suggest · US; English · checked 2026-09-25T22:26:30Z
    Observed completions included best llm for the money and best llm value for money. A formulation signal, not a volume or ranking claim.
  • stories with more than 60 points, trailing 48 hours · Hacker News Algolia search_by_date · global English-language developer community · checked 2026-09-25T22:26:10Z
    Story 49830866, Best LLM for every budget, updated daily (bestmodelforyourbudget.terrydjony.com), created 2026-09-24T14:09Z, carried 180 points and 111 comments at check time. Interest signal, not search volume.

Who this helps

  • developers building high-volume features such as classification, extraction or routing
  • founders estimating the running cost of an AI product
  • hobbyists who want the lowest bill for a side project
  • teams comparing vendors before committing to one API

List prices side by side

These are standard paid-tier prices per million tokens in US dollars, taken from each vendor's official pricing page on September 25, 2026. They are the cheapest current general-purpose model from each vendor we checked, not a complete market survey.

  • OpenAI GPT-6 Luna (gpt-6-luna): 0.10 input, 0.01 cached input, 0.50 output. Prompts over 272,000 input tokens are billed at 2x input and 1.5x output. Batch and Flex are half price. The API free tier is not supported.
  • DeepSeek Flash (deepseek-flash, served by DeepSeek-V4.1-Flash): peak 0.30 input, 0.006 cached input, 1.20 output; off-peak exactly half. Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays. 1M-token context, up to 384K output.
  • Google Gemini 3.1 Flash-Lite: 0.25 input for text, image and video (0.50 for audio), 1.50 output; batch 0.125 and 0.75. Gemini 3.5 Flash-Lite lists at 0.30 and 2.50.
  • Anthropic Claude Haiku 4.5: 1 input, 5 output, 0.10 for cache reads. Batch is half price.

So which is cheapest

On plain standard pricing, GPT-6 Luna is lowest on both input and output. A common way to compare is a blended price that weights input three to one against output, the convention used by the Artificial Analysis data behind the value tracker that trended on Hacker News. By our arithmetic on the list prices above, that blend is about 0.20 for Luna, 0.26 for DeepSeek Flash off-peak, 0.53 for DeepSeek Flash at peak, 0.56 for Gemini 3.1 Flash-Lite and 2.00 for Claude Haiku 4.5.

The ranking flips in two common cases. If most of your input is a repeated prompt that hits the cache, DeepSeek's cached rate of 0.003 off-peak is about a third of Luna's 0.01. And if your jobs can wait, Luna's batch price of 0.05 and 0.25 halves again, as do Gemini and Claude batch rates.

Free is not the same as cheap

Google's Gemini API offers a free tier with free input and output tokens on several models, including Gemini 3.1 Flash-Lite and Gemini 3.8 Flash. Google's pricing page states that free-tier content is used to improve its products, while paid-tier content is not. For prototypes and public data that trade can be fine. For customer data or anything confidential, budget for the paid tier.

Watch dated promotions too. Gemini 3.8 Flash lists 0.75 input and 3.75 output through December 31, 2026, doubling to 1.50 and 7.50 from January 1, 2027. A cost estimate built on a promotional price has an expiry date.

Why the cheapest per token can cost more per task

Tokens are not the same size across vendors. Anthropic's pricing page says its Claude 4.7 and later models use a newer tokenizer that produces roughly 30 percent more tokens for the same text. Two models with identical list prices can therefore bill differently for the same prompt.

Models also use different numbers of tokens to finish the same job, especially reasoning models that bill thinking tokens as output. A cheap model that needs three attempts, or a long reasoning trace, can cost more than a pricier model that gets it right once. Price per token tells you the rate, not the bill.

How to choose

  • High-volume, simple tasks such as classification, tagging and routing: start with GPT-6 Luna or DeepSeek Flash, and check whether batch pricing fits your latency needs.
  • Workloads with a long, repeated system prompt or shared document: compare cached-input rates, where DeepSeek Flash is lowest off-peak.
  • Jobs you can schedule: DeepSeek off-peak hours and batch APIs at the other vendors each cut the bill in half.
  • Prototypes on public data: the Gemini free tier costs nothing, subject to its data-use terms.
  • Anything with confidential data: read each provider's data terms before price, and exclude any tier that uses your content for training if that matters to you.

Limits of this answer

Prices come from official vendor pricing pages as checked on September 25, 2026, and change often. We compared the cheapest current general-purpose model from four major vendors; hosted open-weight models from other providers and aggregators such as OpenRouter can be cheaper still and are not covered. Blended figures are our own arithmetic on list prices and exclude caching, batch and long-context surcharges. RepoRadar did not run cost-per-task tests for this answer, and nothing here reflects any sponsorship or affiliate relationship.

A useful next action

Take 50 real requests from your workload, run them through the two cheapest candidates for your case, and record the pass rate and the billed tokens for each. Divide total cost by the number of passing results. That cost per good answer is the number to compare, and it often differs from the list-price ranking above.

Sources checked

  • OpenAI API docs: GPT-6 Luna model page ↗ checked · vendor documentation, global

    Primary source for GPT-6 Luna pricing (0.10 input, 0.01 cached input, 0.125 cache writes, 0.50 output), the 272K-token long-prompt surcharge, Batch and Flex at 50 percent, fast mode at 2x, and free tier not supported.

  • DeepSeek API Docs: Models and Pricing ↗ checked · vendor documentation, global

    Primary source for deepseek-flash (DeepSeek-V4.1-Flash) and deepseek-v4-pro peak and off-peak prices, cached-input rates, off-peak at half of peak, the UTC peak-hour windows, 1M context and 384K max output.

  • Google: Gemini Developer API pricing ↗ checked · vendor documentation, global

    Primary source for Gemini 3.1 Flash-Lite, 3.5 Flash-Lite and 3.8 Flash prices including the December 31, 2026 promotional end date, batch pricing, the free tier, and the statement that free-tier content is used to improve Google products while paid-tier content is not.

  • Anthropic Claude Platform Docs: Pricing ↗ checked · vendor documentation, global

    Primary source for Claude Haiku 4.5 at 1 input and 5 output, 0.10 cache reads, the 50 percent Batch API discount, and the note that Claude 4.7 and later models use a tokenizer producing roughly 30 percent more tokens for the same text.

  • Best value LLM tracker (Artificial Analysis data) ↗ checked · community project, global

    Secondary source used only for the 3:1 input-to-output blended-price convention from Artificial Analysis and its exclusion of caching, batch and fast modes. Not used for any price figure.

RepoRadar separates factual source claims from analysis. Recheck vendor docs before purchase, deployment, or policy decisions.