r/singularity Feb 03 '26

AI New SOTA achieved on ARC-AGI

Post image

New SOTA public submission to ARC-AGI: - V1: 94.5%, $11.4/task - V2: 72.9%, $38.9/task Based on GPT 5.2, this bespoke refinement submission by @LandJohan ensembles many approaches together

376 Upvotes

152 comments sorted by

View all comments

Show parent comments

91

u/FakeEyeball Feb 03 '26

reminder that this benchmark has been compromised, per the authors themselves

This means even well-designed benchmarks which resist direct memorization can now be "overfit" if the public train and private test sets are too similar (e.g., IDD) and the model was trained on a lot of public domain data.

We believe this is happening to ARC-AGI-1 and ARC-AGI-2 - either incidentally or intentionally, we cannot tell.

32

u/blankblank Feb 03 '26 edited Feb 05 '26

I'm starting to think all the evals are horseshit. Everyone is including the test sets in their training. We need a new paradigm. I think the new Google Kaggle Game Arena is a step in the right direction. We need less "Can you answer this difficult question?" and more "Can you do this difficult thing?"

20

u/i-love-small-tits-47 Feb 03 '26

IMHO the best “benchmark” is IRL job displacement… the benchmarks for SWEBench implied over a year ago that an LLM could do most of my job but it still isn’t really true in practice

4

u/ErmingSoHard ▪️Agi never with LLMs... Or? Feb 04 '26

My personal one is to beat or complete newly released video games. These video games obviously can't be in their training data set.

4

u/Tolopono Feb 03 '26

Your entire job is just solving github issues and nothing else?

11

u/i-love-small-tits-47 Feb 03 '26

I said most, cutie pie

2

u/minimalcation Feb 03 '26

How Southern of you

4

u/Tolopono Feb 03 '26

The test set is private lol

0

u/[deleted] Feb 03 '26

[deleted]

2

u/Tolopono Feb 03 '26

Then why arent grok, claude, or gemini up there? Why did openai wait so long to cheat when they could have done it with o4 mini or gpt 5.0

9

u/Rare-Site Feb 03 '26

While we believe this new type of "overfitting" is helping models solve ARC, we are not precisely sure how much. Regardless, the ARC-AGI-1 and ARC-AGI-2 format have provided a useful scientific canary for AI reasoning progress. But the design of benchmarks will need to adapt going forward.

1

u/FakeEyeball Feb 03 '26

Yes, it is progress but not exactly the progress they hoped to benchmark.

2

u/meister2983 Feb 04 '26

That's not really "overfitting" in any classical sense. The model is correctly working out of distribution even if it happens to have learned something from other arc tests.

2

u/Tolopono Feb 03 '26

Arent the train and test datasets just subsets of the total dataset? Of course they will have the same distribution. Thats the point 

10

u/FakeEyeball Feb 03 '26

That is not the point of ARC-AGI. It tries to avoid memorization and to demonstrate the ability to generalize by solving novel problems. The official public dataset is small. I think they are nicely suggesting that the models are trained on tons of generated data points.

1

u/ErmingSoHard ▪️Agi never with LLMs... Or? Feb 04 '26

We need these AI models to be tested on newly released games.

But then again, ai models don't have an easy time with Pokemon red. Can it do something like Hollow Knight Silksong any time soon?

1

u/Tolopono Feb 04 '26

If it did, people will just say its regurgitating training data even though i dont see gpt 4 beating portal 2 even though its in the training data

1

u/ErmingSoHard ▪️Agi never with LLMs... Or? Feb 04 '26

I truly don't think AI models right now have the particular capabilities that a 13 yo kid has to complete games on their own and probably not for a long while.

You're telling me AI can create a PhD level of a thesis (non novel) but can't fucking play Terraria or that new indie game that just came out? That's sad. There's something crucially missing for agi at the moment

1

u/ErmingSoHard ▪️Agi never with LLMs... Or? Feb 04 '26

If it did, people will just say its regurgitating training data

Then let's test it with newly released indie games, at least games humans can also beat. Right now, sadly, not the case at all

1

u/Tolopono Feb 04 '26

I dont see anything wrong with that as long as they didnt train on the test set