Currently burning money quickly on official deepseek api. They are also increasing pricing starting today. V4 Flash 0731 still feels like the most outstanding model of the past few months and probably to come.
Competitive with opus 4.8 but weaker than sol or fable. About 20x cheaper.
JacobAsmuth 5 minutes ago [-]
Per token. You need to look at pricing per task.
swiftcoder 34 minutes ago [-]
How does it stack against the updated Deepseek Flash version?
pixelesque 13 minutes ago [-]
I've found Pro to be a lot better per "task" than the recently released Flash for code reviews and things (via OpenRouter running in pi.dev).
Flash makes a lot more initial mistakes, and then has to re-check stuff, and produces much more output compared to Pro. It often gets to the correct result eventually, but the output volume is often 5x more than for Pro, and the initial outputs are often wrong, with the first few saying something wrong (like there's a bug, or the code won't compile when it does), and then saying things like "Wait, let me re-check:", or "Actually, looking at it more carefully:" and then it thinks a bit more and eventually gets to the right answer.
swiftcoder 10 minutes ago [-]
yeah, I've definitely noticed one has to be quite precise to keep Flash on the straight-and-narrow
k__ 28 minutes ago [-]
Around 5 percentage points better. (E.g., 87% instead of 82%)
Gecko4072 24 minutes ago [-]
So not worth it over flash? Even at ~7x the size it isn't worth the price hike. Flash may be a monster of a model due to all the RL it received from free usage everywhere.
npn 37 seconds ago [-]
[delayed]
networked 11 minutes ago [-]
I haven't tried DeepSeek V4 Pro 0813 yet. Recent experience tells me that larger models are worth it in non-obvious ways. MiMo-V2.5-Pro solved problems that DeepSeek V4 Flash 0731 couldn't solve for me: for example, adding a live counter for elided reasoning lines to a terminal-based coding harness. You wouldn't be able to tell from the scores on their respective Artifical Analysis pages (https://artificialanalysis.ai/models/mimo-v2-5-pro, https://artificialanalysis.ai/models/deepseek-v4-flash). I like the DeepSeek V4 models, though. It critiqued my engineering decisions better than MiMo, and they seem to have a distinct aesthetic in the SVGs they write.
saaga 18 minutes ago [-]
Yea that's what I was thinking.
Flash is nuts. I find I have to be a more precise and specific with it but damn. It's crossed a threshold of production grade coding for sure.
I was running a session over a couple days and it didnt cross a dollar lol.
k__ 21 minutes ago [-]
I tried the previous Pro model and in the end it was 50% more expensive than the previous Flash.
Wasn't worth it.
sparkling 19 minutes ago [-]
deepseek-v4-flash feels so fast and snappy, i'm loving it. Happy to trade speed for the the 5% degraded benchmarking performance.
saaga 17 minutes ago [-]
I feel the same too. I like the speed.
I'm also a big fan of glm 5.2 fast. I can't wait for like 2000 t/s on these haha.
k__ 17 minutes ago [-]
I wouldn't exactly call it snappy, but faster than Pro, yes.
ericd 14 minutes ago [-]
Single request depth on vllm with dspark, I'm getting ~200 tps, I'd say it's pretty snappy.
JacobAsmuth 3 minutes ago [-]
Well sure but you're running on tens of thousands of dollars of hardware.
52 minutes ago [-]
LeonKnst 29 minutes ago [-]
I find it interesting how much adoption seems to be influenced by momentum. Some of these Chinese models are surprisingly capable, but developers often default to the models that are already established as the “industry standard
spacebanana7 1 minutes ago [-]
In an enterprise setting Chinese models are often discouraged due to political risk. They don't want to need to remove a model that's deeply embedded in their stack. And it's entirely feasible that the US gov bans federal contractors from using them in the next 6 months for example, or that EU AI safety rules effectively ban them too.
HawtAds 4 minutes ago [-]
Hacker News is very Bay Area/US tech centric where spending a few hundred a month on AI is just pocket change. The weaker AI models with more questionable data retention policies are popular in developing countries. I think the new Facebook muse model will be similarly popular.
BlackRabbit1 4 minutes ago [-]
A lot of it/infrastructure departments aren't aware that you can the Asian models hosted within the US or even EU.
Rendered at 17:10:53 GMT+0000 (Coordinated Universal Time) with Vercel.
Competitive with opus 4.8 but weaker than sol or fable. About 20x cheaper.
Flash makes a lot more initial mistakes, and then has to re-check stuff, and produces much more output compared to Pro. It often gets to the correct result eventually, but the output volume is often 5x more than for Pro, and the initial outputs are often wrong, with the first few saying something wrong (like there's a bug, or the code won't compile when it does), and then saying things like "Wait, let me re-check:", or "Actually, looking at it more carefully:" and then it thinks a bit more and eventually gets to the right answer.
I was running a session over a couple days and it didnt cross a dollar lol.
Wasn't worth it.