DeepFrugal Early Preview Find the cheapest way to run any model

Cheapest provider to run DeepSeek V4.1 Flash

Compare DeepSeek V4.1 Flash API prices across every gateway and plan, with live widgets for the cheapest endpoints and the best monthly plan.

13 min read DeepFrugal
Cheapest provider to run DeepSeek V4.1 Flash

DeepSeek shipped V4.1 Flash on 2026-09-10 and moved its flagship traffic onto it: from 2026-09-14, every V4 Pro request is served by V4.1 Flash at Flash rates. The model takes text and images, holds a one-million-token context, and bills on a peak and an off-peak rate. The same weights cost different amounts on each gateway, so the endpoint you pick sets the bill.

The cheapest endpoints

Providers list this model in one of two shapes: a single flat rate, or a pair of peak and off-peak rates. The three tables below rank each shape on its own, one row per plan. They start on Effective prices; turn the toggle off for the raw listed rate. All three read the same data as the main table.

Flat-rate endpoints

Providers that publish one rate, with no time window.

Flat rate
PlanPricing /1MPrivacyPerformanceMonthly quota
InputOutputCached readLogsTrains
DevPass LiteLLM Gateway$0.05$0.2$0.003โ€”$87 of model usage
DevPass ProLLM Gateway$0.05$0.2$0.003โ€”$237 of model usage
DevPass MaxLLM Gateway$0.05$0.2$0.003โ€”$537 of model usage
Ozore BasicOzore$0.055$0.21โ€”โ€”$20 of usage credits
Ozore ProOzore$0.055$0.21โ€”โ€”$70 of usage credits
Ozore APIOzore$0.11$0.42โ€”โ€”โ€”
Portal PlusNous Portal$0.118$0.473$0.002โ€”$22 of usage credits ยท $10 rollover cap
Portal SuperNous Portal$0.118$0.473$0.002โ€”$110 of usage credits ยท $50 rollover cap
Portal UltraNous Portal$0.118$0.473$0.002โ€”$220 of usage credits ยท $100 rollover cap
OpenRouter APIOpenRouter$0.137$0.549$0.00323 tpsโ€”

Effective prices apply the gateway's service fee, and scale a subscription's listed rate by its monthly quota. A subscription's effective rate holds only if you use the full quota. The fee is included only where the gateway publishes it, and sales tax is never included โ€” confirm both with the provider.

Peak endpoints

The rate that applies inside the peak windows.

Peak rate
PlanPricing /1MPrivacyPerformanceMonthly quota
InputOutputCached readLogsTrains
OpenCode GoOpenCode$0.05$0.2$0.001โ€”$60$15
Command Code GOATCommand Code$0.05$0.2$0.001247 tps$60$40
Command Code ProCommand Code$0.086$0.343$0.002247 tps$70$50
Ollama Cloud ProOllama$0.1$0.4$0.002โ€”$60 of usage credits
Ollama Cloud MaxOllama$0.1$0.4$0.002โ€”$300 of usage credits
Command Code Max 10xCommand Code$0.2$0.8$0.004โ€”$150
Command Code Max 20xCommand Code$0.2$0.8$0.004โ€”$300
DeepSeek APIDeepSeek$0.3$1.2$0.006โ€”โ€”
LLM Gateway APILLM Gateway$0.315$1.26$0.032โ€”โ€”
OpenRouter APIOpenRouter$0.316$1.266$0.00692 tpsโ€”

Effective prices apply the gateway's service fee, and scale a subscription's listed rate by its monthly quota. A subscription's effective rate holds only if you use the full quota. The fee is included only where the gateway publishes it, and sales tax is never included โ€” confirm both with the provider.

Off-peak endpoints

The rate that applies outside the peak windows.

Off-peak rate
PlanPricing /1MPrivacyPerformanceMonthly quota
InputOutputCached readLogsTrains
OpenCode GoOpenCode$0.025$0.1$0.001โ€”$60$15
Command Code GOATCommand Code$0.025$0.1$0.001247 tps$60$40
Command Code ProCommand Code$0.043$0.171$0.001247 tps$70$50
Ollama Cloud ProOllama$0.05$0.2$0.001โ€”$60 of usage credits
Ollama Cloud MaxOllama$0.05$0.2$0.001โ€”$300 of usage credits
Command Code Max 10xCommand Code$0.1$0.4$0.002โ€”$150
Command Code Max 20xCommand Code$0.1$0.4$0.002โ€”$300
DeepSeek APIDeepSeek$0.15$0.6$0.003โ€”โ€”
LLM Gateway APILLM Gateway$0.158$0.63$0.016โ€”โ€”
OpenRouter APIOpenRouter$0.158$0.633$0.00392 tpsโ€”

Effective prices apply the gateway's service fee, and scale a subscription's listed rate by its monthly quota. A subscription's effective rate holds only if you use the full quota. The fee is included only where the gateway publishes it, and sales tax is never included โ€” confirm both with the provider.

Benchmarks

DeepFrugal tracks an Intelligence score for the models it catalogs. V4.1 Flash publishes an Intelligence score; it publishes no separate coding or agentic score yet, so those columns read โ€œโ€”โ€. Treat one score as a summary, not a verdict: DeepSeek reports the model ahead of V4 Pro on its coding and agent tasks, and behind it on some knowledge tests.

Benchmark results for DeepSeek V4.1 Flash
ModelIntelligenceCodingAgentic
DeepSeek V4.1 Flash39.5โ€”โ€”

How DeepSeek V4.1 Flash is priced

A gateway can bill the same model in more than one shape. Each shape lands on its own row, so you compare like with like:

  • Pay as you go. You pay per token on input and output.
  • Cached input. Reused prompt tokens usually bill at a lower rate.
  • Cache writes. Some providers charge a fee to store a prompt in the cache.
  • Long context. Requests above a context threshold can bill at a higher rate.
  • Peak and off-peak. A provider that varies its rate publishes a separate row for each window.
  • Subscription quota. A monthly plan covers usage up to a published quota.

DeepSeek publishes two windows for this model. Off-peak rates are half the peak rate. Peak hours run 01:00โ€“04:00 and 06:00โ€“10:00 UTC, Monday to Friday; every other hour is off-peak. A gateway that resells the model can then add a service fee; sales tax is not included and depends on your billing country. Each provider sets its own rate for the same weights.

Pay-as-you-go or subscription?

A subscription only wins above a usage threshold. Below it, you pay for capacity you never use. The threshold depends on your monthly mix โ€” how much input, cached read, cache write and output you send.

The break-even guide explains how DeepFrugal turns a plan's list price into an effective rate per token, and how AI subscription credits work explains what the quota covers. The break-even calculator runs the same maths on your own usage.

Best plan for this model

Providers split into two groups. Some publish a peak and an off-peak rate; others publish one flat rate for the model. The two finders below rank each group on the same monthly workload โ€” a mix of input, cached read, cache write and output. Both value the plans on this model's usage, not on the listed token rate alone.

Best plan with peak and off-peak rates

The first finder splits the workload across the two windows: peak usage is one fifth of off-peak usage. It ranks the plans that publish both rates.

Peak and off-peak rates ยท Experimental

Ranking every subscription plan and pay-as-you-go gateway for a monthly mix of 288M tokens. Open the article for the interactive ranking.

#PlanTypeMonthly costvs cheapestQuota used
1OpenCode GoOpenCodeSubscription$10.00โ€”43%
2Command Code GOATCommand CodeSubscription$10.00โ€”43%
3Ollama Cloud ProOllamaSubscription$20.00+ $10.0043%
4Command Code ProCommand CodeSubscription$20.00+ $10.0037%
5DeepSeek APIDeepSeekPay as you go$25.87+ $15.87โ€”
Break-even details
$0.015$0.035$0.090200M500M1B288M111M668MTotal tokens per month (M tokens/mo, log scale)Average cost per 1M tokens overflow billed at the listed rate your mix
  • Average cost per token
  • Listed rate
  • Quota limit
  • Cheaper than the cheapest
  • Quota usage break-even
  • Plan break-even / loses
  • Average at your mix
Hover the curve to read the values.
โ–พ How the graph is calculated

The winner's curve is the average cost per 1M tokens as your mix scales up. Inside the quota it is the monthly commitment divided by the tokens used, so it falls and reaches the plan's effective rate at the quota limit. Above the limit the extra tokens are billed at the listed rate, so the average climbs towards it.

The quota is one shared budget: each model's published quota is that same budget restated at its rates, allocated in the mix order shown, so a different order or blend moves the curve. A token type a model does not price still counts in the volume but adds no cost. The token axis is logarithmic, so a wide range of usage fits on one chart.

The pay-as-you-go line is the cheapest real rate per model (fee included; sales tax excluded) โ€” the same composite baseline the ranking calls Cheapest pay-as-you-go route โ€” so it can combine more than one gateway. When it equals the plan's own listed rate the chart shows the listed line only. With the no logging or training filter on, both the plan's quota rows and the pay-as-you-go baseline use only clean routes, so the chart matches the filtered ranking.

The markers: the quota usage break-even is where the average meets the list rate; the plan break-even is where the average starts to beat the cheaper of the two references; loses to the cheapest is where it stops doing so; the your mix point is your stated usage, labelled with its monthly token volume, where the horizontal guide marks the average cost of the mix per 1M, and the shaded band is where the subscription wins. Only a subscription draws a curve; a pay-as-you-go plan is a flat line at its real blended rate for the mix, with no quota limit. Select a plan in the ranking to chart it instead of the winner, or pick two to compare them side by side. Two subscriptions also draw their quota limits and listed rates, with the break-even markers on the cheaper plan.

Real rates add the service fee the gateway publishes; sales tax is not included โ€” confirm it with the provider.

Best plan with a flat rate

The second finder bills the whole workload at one flat rate. It ranks the providers that make no distinction between windows and publish one variant.

A single flat rate ยท Experimental

Ranking every subscription plan and pay-as-you-go gateway for a monthly mix of 288M tokens. Open the article for the interactive ranking.

#PlanTypeMonthly costvs cheapestQuota used
1Ozore BasicOzoreSubscription$10.00โ€”77%
2Ozore APIOzorePay as you go$15.36+ $5.36โ€”
3OpenRouter APIOpenRouterPay as you go$18.57+ $8.57โ€”
4Portal PlusNous PortalSubscription$20.00+ $10.0087%
5LLM Gateway APILLM GatewayPay as you go$24.70+ $14.70โ€”
Break-even details
$0.027$0.035$0.053500M288M188M375MTotal tokens per month (M tokens/mo, log scale)Average cost per 1M tokens overflow billed at the listed rate your mix
  • Average cost per token
  • Listed rate
  • Quota limit
  • Cheaper than the cheapest
  • Quota usage break-even
  • Plan break-even / loses
  • Average at your mix
Hover the curve to read the values.
โ–พ How the graph is calculated

The winner's curve is the average cost per 1M tokens as your mix scales up. Inside the quota it is the monthly commitment divided by the tokens used, so it falls and reaches the plan's effective rate at the quota limit. Above the limit the extra tokens are billed at the listed rate, so the average climbs towards it.

The quota is one shared budget: each model's published quota is that same budget restated at its rates, allocated in the mix order shown, so a different order or blend moves the curve. A token type a model does not price still counts in the volume but adds no cost. The token axis is logarithmic, so a wide range of usage fits on one chart.

The pay-as-you-go line is the cheapest real rate per model (fee included; sales tax excluded) โ€” the same composite baseline the ranking calls Cheapest pay-as-you-go route โ€” so it can combine more than one gateway. When it equals the plan's own listed rate the chart shows the listed line only. With the no logging or training filter on, both the plan's quota rows and the pay-as-you-go baseline use only clean routes, so the chart matches the filtered ranking.

The markers: the quota usage break-even is where the average meets the list rate; the plan break-even is where the average starts to beat the cheaper of the two references; loses to the cheapest is where it stops doing so; the your mix point is your stated usage, labelled with its monthly token volume, where the horizontal guide marks the average cost of the mix per 1M, and the shaded band is where the subscription wins. Only a subscription draws a curve; a pay-as-you-go plan is a flat line at its real blended rate for the mix, with no quota limit. Select a plan in the ranking to chart it instead of the winner, or pick two to compare them side by side. Two subscriptions also draw their quota limits and listed rates, with the break-even markers on the cheaper plan.

Real rates add the service fee the gateway publishes; sales tax is not included โ€” confirm it with the provider.

Change the usage numbers to your own to see how a ranking moves. Rank by reorders the list: cheapest, best privacy, performance, or the broadest catalog. Open How the best plan is calculated for the method.

What to check beyond price

Two endpoints with the same rate are still not equal. Check:

  • Context window. A long-context tier can cost more and cap at a lower size.
  • Throughput. Tokens per second tells you whether the endpoint keeps up with an interactive or agentic workload.
  • Logging and training. The main table carries a privacy column per row. Treat a blank as unknown, never as "no".
  • Rate limits. A subscription can throttle bursts well below its monthly quota.

Frequently asked questions

Why does the same model cost more on one gateway?

A gateway marks up the underlying provider and can add a service fee and sales tax. Each step raises the effective rate.

Does a subscription make DeepSeek V4.1 Flash cheaper?

Only above a usage threshold. Under it, pay-as-you-go is cheaper. The threshold moves with your mix, because cached reads and output bill at different rates.

How do peak hours change the bill?

Usage inside the peak windows bills at the higher rate; usage outside them bills at half. A workload that runs mostly off-peak reaches a lower effective rate than one that runs at peak, on the same plan.

Can cached reads change the ranking?

Yes. A cache-heavy workload shifts the ranking, because a cache read often costs a fraction of fresh input. Enter your own cached read and cache write volumes in the calculators to see the effect.

Summary

The cheapest way to run DeepSeek V4.1 Flash depends on your usage, not on the list price alone. Compare the endpoints first, then value your own month against the plans. Remember that the rate depends on the time of day.

Next steps

Frequently asked questions

Where is DeepSeek V4.1 Flash cheapest?

The cheapest-endpoint widget ranks every gateway by listed input rate, with a toggle for effective prices.

Is a subscription cheaper than an API for this model?

It depends on usage. The plan finder widget ranks quota plans for a monthly mix, and the break-even calculator shows where the plan wins.

What changes between peak and off-peak?

The listed rate. DeepFrugal tracks peak and off-peak as separate variants, each with its own window.

Do the shown prices include fees and tax?

The effective-prices toggle applies a gateway's service fee; sales tax is never included. Confirm it with the provider.