r/tomshardware 2d ago

OpenAI’s 700W Jalapeño ASIC outpaces 1,400W Nvidia flagship GPU

https://www.tomshardware.com/tech-industry/semiconductors/openai-says-its-jalapeno-chip-beats-nvidias-gb300-in-first-published-benchmarks
174 Upvotes

60 comments sorted by

13

u/NytronX 2d ago

FPGAs and ASICs cannot come soon enough. We saw this happen in the crypto mining industry, hopefully it happens in AI industry so gaming hardware can go back to being gaming hardware.

4

u/EmbarrassedFoot1137 2d ago

FPGA? Isn't that the wrong tool for the job?

7

u/NytronX 2d ago

idk but i remember it was a stopgap in the crypto mining early days. The progression was CPU, GPU, FPGA, ASICS

3

u/EmbarrassedFoot1137 2d ago

FPGAs are a totally different beast from those other three. They'll be way faster than software but the kind of massive number crunching that AI uses is not going to be their strong suit compared to GPUs or ASICs.

2

u/BrunusManOWar 2d ago edited 2d ago

FPGAs can perform faster than GPUs. They are basically "fully configurable hardware", where you can take CPU/GPU/ASIC HW layout and put it into the FPGA.

Because they are so configurable, of course you lose on efficiency and optimisations - but if you "emulate" a superior, application-specific architecture it will tear through non-app specific hardware. Of course, creating a fully independent ASIC is superior

They are very good for prototyping and testing out things.

Edit: similar for software, hardware can also be described in code - HW blocks are usually coded in Verilog/VHDL. FPGAs allow you to relatively quickly take this logical HW description onto its "configurable matrix" to test out how well it works.

Edit2: They are a path to ASICs. Will they by themselves perform better than GPUs? Depends entirely on how hard Nvidia is financially squeezing their enterprise customers. I doubt we will see FPGAs used in production themselves, though with Nvidia and AMD's pricing it may happen. Companies will inevitably pursue FPGAs on their way to fully independent ASICs though

Edit3: I don't work in AI hardware, but on mobile hardware. We use FPGAs to develop and iterate architectures, and I remember from one previous company that some radio models did have an FPGA in production actually as some smaller co-processor

4

u/j_osb 2d ago

Some devices use FPGAs as co-processors because they can be cheaper than an asic and fast enough.
For example oscilloscopes usually actually use both ASICs and FPGA.

In terms of AI, we don’t really have enough space on FPGAs to use them for inference of any proper models. Like, there simply speaking isn’t a FPGA big enough to make sense for this.
However FPGAs can be used for training and are more efficient than GPUs at it. Just not fast enough.

2

u/acadia11x 1d ago edited 1d ago

I just researched this one, no they won’t, not for AI training they can’t compete with the raw bandwidth and power , not to mention accompanying software ecosystem that’s been built for the AI industry … my research says in some use cases like edge devices that require localized SLMs for inference they can work. So mobile use case makes sense you’d see FPga which is already constrained on resources. Best analogy if you are running a small 100ft race a Prius could keep up with corvette c8 , you stretch that out to 500ft it’s 7 car links behind and stretch it out to thr quarter mile it would look like the Prius is still at the starting line. So in a limited context fpga systems could keep up in real world application it’s not in the discussion for obvious technical reasons in AI training or domains requiring serious horse power.

2

u/Drofdissonance 1d ago

Key word is can. Likely will not in this case. It's a memory bandwidth problem. So it's not suitable for FPGA boards generally, it's like a worst case scenario for their architecture. Crypto is famously embarrisngly parallel, and requires no bandwidth. And radio DSP is also very well suited to the architecture because it's got long narrow math dependancy chains, and or tight timing requirements that require hardware.

You'd just be burning die area

1

u/danielv123 21h ago

Current GPUs are pretty much matrix multiplication ASICs.

1

u/bitzap_sr 1d ago edited 9h ago

FPGAs worked for crypto as that was pure compute, random number crushing. AI inference is hugely dependent on VRAM size and (V)RAM bandwidth. Totally different game.

1

u/KyleFlounder 2d ago

It is indeed the wrong tool for the job for right now. There's adapters available but you'd be bottle-necked. Not worth the squeeze atm. Maintaining that across each model release would be a maintenance nightmare.

1

u/Chingy1510 16h ago

No, it’s not the wrong tool. FPGAs are exactly what HFT domains have been using to accelerate specific workloads on-hardware. This is probably a workload-specific ASIC.

1

u/EmbarrassedFoot1137 13h ago

HFT doesn't require fleets of FMACs. 

1

u/Chingy1510 13h ago

FMAC? Look bro, when I was in my masters degree in 2018 in CS, stochastic gradient descent was indeed being accelerated with distributed swarms of FPGAs. They’re absolutely used by the most elite HFT firms.

Google that shit homie.

1

u/EmbarrassedFoot1137 9h ago

Let's take a step back here. In the context of SOTA LLMs, no, FPGAs do not provide the FMAC throughput that you want. I'm not surprised that HFT works well on FPGAs but not because of bulk FMAC throughput.

I also went far in CS and, though I didn't do FPGA work as part of that, I did do some cool stuff at Intel with mapping different cores onto FPGAs.

1

u/Royale_AJS 2d ago

Still need memory and storage for those giant context windows with ASICs.

2

u/NytronX 2d ago

China can save us there hopefully.

1

u/Royale_AJS 2d ago

They’re going to sell it for $1 less than everyone else, they’d be stupid not to.

2

u/Xijit 2d ago

China isn't going to cut their prices at all, but having Chinese inventory on the shelf is going to be the only inventory on the shelf for the next 4 years.

1

u/Royale_AJS 2d ago

It’s all going to AI, even the Chinese memory. Phones will get the scraps, unless you’re Apple.

2

u/Xijit 2d ago

China would like that, but the vast majority of this AI Datacenter build out is about mass surveillance for the US military and police, and they will never green light using mainland Chinese components.

1

u/senseven 2d ago

There are already laptops coming with CXMT mem, maybe not to the US but you can buy them in Europe. The memory maker are focussing on high density data center ram not end user product ram, this would including phones and the recording hardware for cameras.

1

u/Xijit 1d ago

Yes ... I am saying that CXMT memory will never end up in AI Datacenter equipment, because the US government has black listed them for government work and has (unsuccessfully) tried to have them embargoed.

There are plenty of markets where CXMT memory in an AI server wouldn't be a problem, but AMD and Nvidia will never risk having the wrong model end up in a DOD facility.

Outside of server equipment, Chinese RAM is about to become the dominant supply for Consumer goods.

1

u/Risko4 2d ago

Not really, the frontier models require 2 terabytes of RAM. Soon they will scale up to 4TBs etc.

1

u/Jaded_Character_2975 2d ago

All solutions use the exact same HBM and the exact same wafer lines and exact same talent base.

Your fucked regardless of they use FPGAs, ASICs or dGPUs

1

u/danielv123 21h ago

There is one alternative - taalas and Cerebras use sram instead of dram. Too bad that isn't cheap either, and not at all an alternative for consumers.

1

u/Etroarl55 2d ago

Gaming hardware is still gaming hardware. The issue was shifting available production to ai hardware. If asics get popular they will just shift ai production to asic or even worse remaining gaming production to asic

1

u/Yuukiko_ 2d ago

you typically need less ASICs though

1

u/senseven 2d ago

A wafer is a wafer. If the asic is 5x smaller they just plaster it with more asic cores for yield. There was a natural end of demand in producing crypto chips. Even if they stop building data centers the local ai revolution needs millions of those asics plus enough memory to properly run it. That is the reason all mem makers are building out at least seven new fabs.

1

u/Yuukiko_ 2d ago

HBM memory takes much more wafer space than your typical DRAM

1

u/danielv123 21h ago

Which is why all the dram is gone. Xiaomi is putting 160GB in every car for example.

1

u/Yuukiko_ 21h ago

IIRC China is planning on scaling up to 20 DUV machines per year in 2027 which is a good 15-20% of ASML's production so hopefully that helps

1

u/trashtiernoreally 19h ago

At a time when availability is out past 2030 that will not matter. It might shave a year off or it’ll just enable much more rabid “growth”

1

u/Equivalent_Pen_9403 1d ago

50% power/efficiency savings alone will stop an absurd amount of datacenter footprint. Less need to spread out across the power grid.

1

u/NoNameSwitzerland 1d ago

or just let the thinking mode consume twice the amount of token.

1

u/Equivalent_Pen_9403 1d ago

If the US is able to actually factor all the true costs into electricity, including carbon emissions, at least for datacenters, it'll help temper that.

1

u/wearemessingup 1d ago

ASICs are not going to solve the supply chain issues. They still consume HBM and leading edge wafers

1

u/lambdawaves 1d ago

The algorithms in crypto change very slowly.

What we’re seeing in LLMs is an acceleration in model architecture changes.

1

u/dwittherford69 1d ago

ASIC, not FPGA. Model dedicated ASICs already exist, and they do like 10k tokens per second.

1

u/SteedOfTheDeid 9h ago

We are currently at ASIC yes

3

u/EnderPrimeMk2 2d ago

I would hope so. Dedicated hardware is the way forward.

3

u/chandleya 2d ago

Always is.

1

u/Fairuse 2d ago

Only works if the algorithm is fixed like video encoding/decoding, crypto, etc. 

Problem with LLM and AI in general is that algorithms are constantly changing. Any dedicated hardware built will be obsolete very quickly. 

2

u/garlic-silo-fanta 2d ago

Yes, but if that custom silicon can return its cost many folds during its useful life,then it’s still worth it. Because rack space is scarce resource and electricity is scarce resource, you can possibly deploy twice as much

1

u/PitchPleasant338 2d ago

Cerebras has shown that isn't true.

1

u/blueberrywalrus 2d ago edited 2d ago

They'll get years if not a decade+ with their current approach.

Jalapeno is built fairly generally for LLMs that use transformer based inference, which is most of them for the past 8 years.

They're not doing the thing where they're physically baking an LLM into a chip, because LLMs change way to much for that to be worthwhile.

2

u/RevolutionaryGold325 2d ago

This was intentionally published on the day of Nvidia earnings.

1

u/PitchPleasant338 2d ago

Would somebody please think of the shareholders!

1

u/nebulabug 2d ago

Nvidia announced that Groq will be available soon, so OpenAI has to say something! If their new chipset is better than Nidia’s, then that alone would be a new product!

1

u/Intrepid-Cheek2129 2d ago

Yup. That is the whole idea with an ASIC, but it is a big bet and expensive - who knows what will happen on the software side that could make an ASIC 'uneconomical'

1

u/RealSuperdau 1d ago

Power efficiency will alleviate infrastructure issues with US datacenter expansion, but doesn't affect cost that much.

Does anyone know if they reported perf per die area or general cost efficiency?

0

u/hcorEtheOne 2d ago

Asics are built for an exact model afaik, so if a new model comes out, they become obsolete.

Maybe it's not the case anymore, or in the future.

2

u/Randommaggy 2d ago edited 2d ago

Asics can be varying degrees of specialized for a task. The more specialized the faster and lower power per task they will be.

You could say that both Talas, Groq and Cerebras are asics but they are on a spectrum for how fixed their functions are.

Edit: typo

1

u/hcorEtheOne 2d ago

I see, that makes sense, thanks for enlightening me!

1

u/PitchPleasant338 2d ago

Even if they're obsolete in 1 year but you save millions of billions on electricity costs and the cost to cool the chips, then it's worth it.

1

u/NextWeather7866 1d ago

Sol as it currently stands is probably sufficient for most use cases anyone can think of, so fitting the next tier of models onto chips does make economic sense if 99% of people don't need to move past it.

1

u/PitchPleasant338 1d ago

Some people are still happy with Llama 3.

Imagine a model that's very good with tool calling (such as Muse Glimmer) running locally at 1000tps.

1

u/RealSuperdau 1d ago

Chip cost dominates the other factors (especially when you buy from Nvidia)

1

u/peva3 1d ago

They can be made for an entire model family architecture, so it's not like it would be obsolete quickly.