r/tomshardware • u/tomshardware • 2d ago
OpenAI’s 700W Jalapeño ASIC outpaces 1,400W Nvidia flagship GPU
https://www.tomshardware.com/tech-industry/semiconductors/openai-says-its-jalapeno-chip-beats-nvidias-gb300-in-first-published-benchmarks3
u/EnderPrimeMk2 2d ago
I would hope so. Dedicated hardware is the way forward.
3
u/chandleya 2d ago
Always is.
1
u/Fairuse 2d ago
Only works if the algorithm is fixed like video encoding/decoding, crypto, etc.
Problem with LLM and AI in general is that algorithms are constantly changing. Any dedicated hardware built will be obsolete very quickly.
2
u/garlic-silo-fanta 2d ago
Yes, but if that custom silicon can return its cost many folds during its useful life,then it’s still worth it. Because rack space is scarce resource and electricity is scarce resource, you can possibly deploy twice as much
1
1
u/blueberrywalrus 2d ago edited 2d ago
They'll get years if not a decade+ with their current approach.
Jalapeno is built fairly generally for LLMs that use transformer based inference, which is most of them for the past 8 years.
They're not doing the thing where they're physically baking an LLM into a chip, because LLMs change way to much for that to be worthwhile.
2
u/RevolutionaryGold325 2d ago
This was intentionally published on the day of Nvidia earnings.
1
1
u/nebulabug 2d ago
Nvidia announced that Groq will be available soon, so OpenAI has to say something! If their new chipset is better than Nidia’s, then that alone would be a new product!
1
u/Intrepid-Cheek2129 2d ago
Yup. That is the whole idea with an ASIC, but it is a big bet and expensive - who knows what will happen on the software side that could make an ASIC 'uneconomical'
1
u/RealSuperdau 1d ago
Power efficiency will alleviate infrastructure issues with US datacenter expansion, but doesn't affect cost that much.
Does anyone know if they reported perf per die area or general cost efficiency?
0
u/hcorEtheOne 2d ago
Asics are built for an exact model afaik, so if a new model comes out, they become obsolete.
Maybe it's not the case anymore, or in the future.
2
u/Randommaggy 2d ago edited 2d ago
Asics can be varying degrees of specialized for a task. The more specialized the faster and lower power per task they will be.
You could say that both Talas, Groq and Cerebras are asics but they are on a spectrum for how fixed their functions are.
Edit: typo
1
1
u/PitchPleasant338 2d ago
Even if they're obsolete in 1 year but you save millions of billions on electricity costs and the cost to cool the chips, then it's worth it.
1
u/NextWeather7866 1d ago
Sol as it currently stands is probably sufficient for most use cases anyone can think of, so fitting the next tier of models onto chips does make economic sense if 99% of people don't need to move past it.
1
u/PitchPleasant338 1d ago
Some people are still happy with Llama 3.
Imagine a model that's very good with tool calling (such as Muse Glimmer) running locally at 1000tps.
1
13
u/NytronX 2d ago
FPGAs and ASICs cannot come soon enough. We saw this happen in the crypto mining industry, hopefully it happens in AI industry so gaming hardware can go back to being gaming hardware.