DeepSeek Open-Sources V4.1 Flash: 552B Parameters, $0.60 per Million Output Tokens, and a Coding Score That Edges Past Opus 5
Everyone else sells tokens by the gram. DeepSeek sells them by the ton.
Three key facts
A 552-billion-parameter mixture-of-experts model that wakes up only 8 billion parameters to read and 16 billion to write. V4.1 Flash uses an encoder-decoder design, handles a one-million-token context window with native vision, and ships its weights on Hugging Face under the MIT license, free for commercial use and self-hosting. A new compressed sparse attention scheme with FP4 KV caching squeezes the cache to roughly 890 bytes per token, about a quarter of the global cache memory V4 Flash needed.
Pricing: off-peak, cached input costs $0.003 per million tokens, uncached input $0.15 and output $0.60; peak hours double it. Peak means weekdays 01:00 to 04:00 and 06:00 to 10:00 UTC. That $0.003 figure is aimed squarely at agents, which reread the same context over and over and now pay almost nothing to do it.
The scorecard is mixed. At maximum reasoning effort it scores 74.2 on DeepSWE v1.1, a hair above Claude Opus 5 at 74.0 and GPT-5.6 Sol at 73.0, plus 88.1 on CyberGym. On Terminal-Bench 4.0 it manages 31.2 against Opus 5's 51.8. The Artificial Analysis Intelligence Index gives it 40; GPT-6 Astra and Fable 5.1 both sit at 53.
WangDou's Take
Start by deflating the loudest claim. Beating Opus 5 on DeepSWE by 0.2 points is noise. Losing on Terminal-Bench 4.0 by 20 points is a gap. An index score of 40 versus 53 is honest: this model isn't the smartest in the room and was never trying to be.
The real weapon is buried in the peak-hours table. 01:00 to 04:00 and 06:00 to 10:00 UTC is 9 a.m. to noon and 2 p.m. to 6 p.m. in Beijing. Expensive while China is at work, half price while America is. That's a rate card written for Silicon Valley developers.
Now the math: a billion cached input tokens costs $3, less than a pour-over in San Francisco. When an agent rereads the same codebase hundreds of times a day, the deciding factor isn't a 0.2-point leaderboard edge, it's the look on the CFO's face at month-end. The American labs are guarding frontier capability behind a 53-point wall. DeepSeek is outside that wall selling 40-point capability at produce-aisle prices, and most real work doesn't need 53 points.
Source: VentureBeat · The Rundown AI · Dataconomy
