Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
DeepSeek Previews V4 Pro and V4 Flash With 1M-Token Context · News · Kaino
DeepSeek Previews V4 Pro and V4 Flash With 1M-Token Context
Kaino
2d agoAug 4, 2026, 12:00 AM1 views

DeepSeek Previews V4 Pro and V4 Flash With 1M-Token Context

DeepSeek has released preview versions of V4 Pro and V4 Flash, two API models with a stated 1 million-token context window. The company positions Flash as a lower-cost, faster option, with published pricing starting at $0.14 per million uncached input tokens.

llmsdeepseek

DeepSeek adds two V4 preview models

DeepSeek has announced preview releases of V4 Pro and V4 Flash, expanding its model range with two variants that support a stated context window of up to 1 million tokens.

According to DeepSeek’s V4 release documentation, both models are available through its API and are released with open weights. The company describes V4 Flash as the lower-cost variant, designed for fast and economical inference, while V4 Pro is positioned as the higher-end option.

The Associated Press also reported the release of the V4 Pro and Flash previews. AP said DeepSeek presented the models as updates with improved reasoning and agent-oriented capabilities, and reported that both have a 1 million-token context window.

Published API pricing favors Flash

DeepSeek’s pricing documentation lists V4 Flash at $0.14 per million uncached input tokens, $0.0028 per million cached input tokens, and $0.28 per million output tokens. The difference between cached and uncached input prices matters for applications that repeatedly send the same background material, such as reference documents, instructions, code repositories, or conversation context.

Artificial Analysis lists the V4 Flash 0731 model at the same published prices for input and output. Its analysis estimates a weighted blended cost of $0.06 per million tokens and assigns the model an Intelligence Index score of 50.

That blended figure should not be treated as a universal cost comparison. Actual spending depends on the ratio of inputs to outputs, the extent to which prompt caching applies, the size of requests, reasoning settings, and the way an application manages long context. List prices can therefore be useful for initial evaluation without predicting the total cost of a production deployment.

Long-context use will require practical testing

A 1 million-token context limit is intended for workloads involving large volumes of material, including long technical documents, extensive software codebases, transcripts, research collections, and multi-step tasks that retain substantial background information.

A maximum context specification alone does not establish performance on those tasks. Developers will need to assess whether a model can reliably find relevant details in long inputs, maintain accuracy across extended interactions, use tools effectively where applicable, and return results within acceptable latency limits.

DeepSeek’s release combines a very large stated context capacity with low published per-token prices for V4 Flash. For organizations considering the models, the relevant comparison will be workload-specific: reasoning quality, reliability over long inputs, caching behavior, response speed, and end-to-end costs may matter more than a single headline price.

Sources

DeepSeek’s V4 Preview Release announcement and Models & Pricing documentation provide the model availability, context-window, open-weights, positioning, and pricing details. The Associated Press independently reported the V4 Pro and Flash launch and DeepSeek’s claims around reasoning and agent-oriented improvements. Artificial Analysis published its separate price listing and benchmark index for V4 Flash 0731.

Key takeaways
  • 1

    DeepSeek adds two V4 preview models DeepSeek has announced preview releases of V4 Pro and V4 Flash , expanding its model range with two variants that support a stated context window of up to 1 million tokens.

  • 2

    According to DeepSeek’s V4 release documentation, both models are available through its API and are released with open weights.

  • 3

    The company describes V4 Flash as the lower cost variant, designed for fast and economical inference, while V4 Pro is positioned as the higher end option.

Continue reading

Latest from Kaino News

Story pulse

Freshness

2d ago

Views

1

Reading

3 min

Byline

Kainotomic Team

Utilities

Topics

llmsdeepseek

Sources

Reference material and original reporting used in this story.

DeepSeek

Published Aug 4, 2026, 12:00 AM

View source