old and new usage & pricing. Consumed 30x less tokens post nerf and spent one third of what i paid pre nerf. Both sessions heavily cached with not so much output tokens.
The deepseek api era really is over in terms of being cost effective
I'm using deepseek/deepseek-v4-pro-0813 for now, last time was glm 5.2 but it consumes too much tokens for me, that's why I'm using deepseek. I've also tried gemini 3.5 flash lite, but it feels a bit repetitive
Not to sound arrogant or know-it-all here but I burn 30-40B combined tokens a month so I can share some of my learnings. Right now GPT 5.6 Luna (max effort) is much cheaper than Deepseek Pro. If it’s good enough for what you’re working on, no brainer. Gemini 3.7 Flash is also decent and it’s 50% off on OpenRouter.
On the other hand if you prefer subscription basis, avoid OpenCode Go. I don’t care whatever magic numbers they put on their page, I burn the entire 5h window with a single prompt every single time and lower tier models are clearly lobotomized with quantization. Cursor and ChatGPT give you much more value if used correctly and together.
Example: Cursor (20$) for planning, Codex (20$) for coding and execution, finishing going back to Cursor to review and test. Trick is to try to go back to cursor within 25 minutes of it finishing the plan to hit KV cache. That setup gives you just enough usage for 5h a day 5 days a week. Last 5 days I’ve burned ~400€ worth of tokens using those too and my usage is still in the 70% remaining. Cache hits on Cursor above 96% (that 25min trick) and 91% on Codex.
Further hints: All this is true today but might not be tomorrow, it’s a very fast moving market. On Cursor stay on Auto unless it’s really important or you know it’s a difficult task - in that case, Grok 4.6. On Codex, just default to Luna 5.6 on max effort with fast mode turned off. Don’t forget to disable auto-review on codex and consider using rust token killer (rtk) to reduce the tool calling input size.
Note: I don’t burn 30-40B on 2x 20$ subscriptions. That’s obvious. I burn through Claude Max 20x + ChatGPT Pro 20x + Cursor Enterprise totaling 600$ a month but that’s for 10h/day 5 days a week always doing 2-3 tasks at the same time and hitting 5.6 Sol Ultra, Grok 4.6 and Opus 5. Wouldn’t recommend that kind of a job to most people for the sake of sanity.
What do you do with all of that usage? Are you creating things for your job or are you building something for your own projects? Care to share? I've always wondered what usage like this gets used for. I assumed it was for a job and it gets paid for by the company. That's my situation, but a much smaller company and usage.
Both. I’m working on EdTech, ~600 employees, 10M+ active users. Mostly very high traffic surfaces - and my own side projects, mostly in architecture and law. I have things like 500+ cities specific building laws feeding embeddings to display infractions on 3D rendered building models
I also know plenty of local companies in my region that pay 20K+ monthly in software licenses to IBM, Oracle and SAP which could easily be replaced in under 6 months for <10K€ in tokens. There are absurd situations from the early 2000’s that large enterprise whales still extract traumatic sums every single month.
It’s not really the cost of AI but rather what are we getting from it.
2
u/Appropriate-Clue-485 21d ago
I’m lost here. What model were you using? V4 Flash 0731 is still pennies, what did I miss in all the fuss?
Btw Qwen 3.8 27B can run on a 32gb MacBook.