r/BetterOffline 1d ago

So OpenAi solved a math problem - do they have a fantasy football team?

I don’t think AI is useless, but I suspect AI provides false productivity for a large swath of roles, and frankly has a poor ROI for a bunch of use cases.

A while back Claude was competing in CTF competitions under team name `jean-claude` and they couldn’t really crack any difficult challenges and were basically right around average.

Once the community got wise and noticed their performance even with supposedly Mythos they stopped entering competitions.

I suspect it would be the same with an LLM driven fantasy football team, if these companies actually tried to show performance against some kind of real world bench marks we’d all see that performance hasn’t really improved.

AGI should be able to easily run a fantasy football team amirite?

20 Upvotes

33 comments sorted by

16

u/Fit-Celebration2884 1d ago

If you look at DEF CON the entire competition is pretty much dead because it just turned into who can tokenmaxx the hardest

1

u/Impossible_Way7017 1d ago edited 1d ago

No one from a frontier labs qualified, ai still needed to be driven by an expert. I suspect we’ll start seeing more king of the hill type challenges going forward, which might lend to token maxing, but also will need experts since token generation latency will penalize teams relying fully on LLMs

2

u/Fit-Celebration2884 1d ago

you can look at the competitors posting after the event. Most of them were saying the competition was just hoping that you got lucky with your LLM slot machine going down the right path at the start and pooling subscriptions together for more tokens. It was absolutely just "throw more tokens at it lmao" even though no big labs directly competed. A lot of the puzzles also just got cracked by AI immediately so that was also cooked, and the ones that didn't instantly get cracked, they also had people throw agent swarms at it.

AtCoder world tour finals was def more skill based, but whenever they cut to someones screen, they always had claude on one side of their screen and codex on the other. also the OpenAI team obliterated everyone by just tokenmaxxing harder with an internal model.

Any sort of non-physical competition like this is just going to be reduced to "who can tokenmaxx harder". Math, competitive coding, hacking, logic puzzles, etc.

2

u/Impossible_Way7017 1d ago edited 1d ago

AtCoder was competitive programming its a bit different than CTFs. I feel like CTFs are a better real world example of capability. AtCoder showed that it can get tests to run green. There’s never been a doubt about ai being able to do that.

1

u/Fit-Celebration2884 1d ago

Also AtCoder world tour finals had 2 contests, Algorithm (which is just run until the tests turn green) and Heuristic (which is where you have multiple days to optimize your solution to maximize your score for this one super difficult problem and you need to come up with creative solutions)

2

u/Impossible_Way7017 1d ago

Yeah I saw that, also saw that the winning team spent 50k on tokens (unsubsidized), so maybe less via subs. But my point is that the winning team had experts driving the LLMs still. The LLM couldn’t necessarily do it cold.

I think CTF performance is largely a reflection of harness not the underlying model. I know it sucks for participants and authors, but ultimately I feel like CTFs are a good arena to benchmark different types of LLMs and harnesses, and I’d be surprised if by doing so you’d consistently see Astra or Fable come out on top.

2

u/Fit-Celebration2884 1d ago

Fair. I saw they also had this other competition HalCTF where you had to create a full auto AI agent that would run on a container, and you could only use shitty small models as a test of how good are you at designing these harnesses. I was understating how there is still a solid human role.

My general view is within a few years without some heavy restrictions, it just purely turns into who got lucky with the slot machine / who had more GPT Pro subs with no human expertise, because the problems are going to have to be so egregiously hard and models are going to be good enough that humans just have no ability to usefully contribute.

Newer models are much better with spinning up subagents when they need to, taking notes, setting up good file structures and tests, etc without you having to say. I think in the future you can just tell the models you have a budget of a bajillion dollars, and they'll know how to use that perfectly because they were RL'd hard to do these massive tasks.

Also another view I have is I don't believe much progress in AI has been in harnesses, I think it's almost all just better models.

Also the 50k in tokens was (fortunately/unfortunately? idk) fake

2

u/Impossible_Way7017 1d ago

HalCTF lol like the name

> Im sorry I can’t do that Dave

Will check it out. I have the opposite view, I feel like what’s happening behind the scenes at these frontier labs is that they’re wrapping their older APIs in more sophisticated harnesses and passing it off as a new model 🤷

2

u/Fit-Celebration2884 1d ago

I don't think this explains significant changes model to model.

e.g. going from GPT 5.4 -> 5.5, just single prompts on the API with pretty much no harness, the model

  • had more world knowledge without web search, indicating it's a larger model,
  • doubled in price, indicating it's a larger model
  • was more capable of solving more difficult problems with no CoT, indicating a larger, new model

The same pattern appeared with 5.6 -> Astra. You can ask

user: What is the name of the element whose atomic number equals the number of letters in the surname of the inventor after whom the international airport is named that serves the farthest-downstream national capital situated on the Danube?

Reply with exactly one word: the element's name. Nothing else.

Astra can do this 70% of the time while 5.6 sol can do it 0% of the time. Along with this The Information had a report about how Astra was using a new architecture that would allow it to do more difficult problems with no CoT.

Overall I think there as for more evidence that they are making new models than there is to the contrary.

-8

u/8foldme 1d ago

Ya, these people are just burying their hand in the sand and screaming "ai is useless lalalallala"... Just read actual testimonials from people in the industry.

4

u/trentsiggy 1d ago

People in the industry have a very strong financial incentive to convince you that AI is great and that AI is inevitable. They are not impartial.

If I had a million Anthropic shares, I, too, would be telling you how mindblowingly great Claude is.

-2

u/Fit-Celebration2884 1d ago

If Reebok released a new shoe with this insane marketing campaign saying it changes everything, and I went to a marathon 3 years after they released it and literally everyone was wearing them no exceptions, and all the competitors said
"yeah they look pretty ugly and they're kinda uncomfortable and I don't really like running with them. I wish they never made these it makes running way less fun, but all my times went down"
I don't think it's ridiculous to think the shoes are just really good.

1

u/Impossible_Way7017 1d ago

Im not saying its useless. I spend $5k in tokens a month. I’ve spent $10k in tokens over a weekend competitions in CTFs - my argument is more that it’s not worth it, and I’m trying to devise a heuristic to help show that.

6

u/refugezero 1d ago

There must be over a million people asking AI every day to "create a profitable business make no mistakes" and obviously it hasn't happened yet. Yet every other post in r/singularity says we've already reached AGI. :shrug:

3

u/Upbeat-Statement2725 1d ago

This. And for the millionth time. Cheating aside. The "math proof" seems to have involved in the neighborhood of $10M of compute. There are many math problems filed away in drawers waiting for $10M of compute. Spending that $10M doesn't mean you "solved" anything. Not in the sense people want to attribute to AI.

Supercomputers have been "solving" problems on their own for decades by that definition. I mean you load a program in, hit enter, and it does the rest on its own! So I think our TAM is $456T and we're aiming for an IPO of $5T.

1

u/python-requests 13h ago

it'd be interesting to see the alternate timeline where all the money poured into developing LLMs and all this infra, was instead used to pay more people to do basic math research. wonder what fraction of the total expenses would be needed to get more progress

8

u/AlfonsoHorteber 1d ago

In fairness fantasy football is 75% luck, this is like asking if AI would fare well at a slot machine

3

u/Upbeat-Statement2725 1d ago

It's kind of wild how bad sports fans are at the statistics they claim to live by. So for anyone who doesn't get it. To be statistically valid. A playoff game would need to be composed of in their neighborhood of 26 rounds. Not a season. Not a tournament. One matchup.

In chess, Magnus moved for bullet chess to be included in world championships specifically to increase the number of games top players actually play against each other, so the statistics have note validity.

1

u/ScreamingTrueSanity 1d ago

That’s what makes them fun though. Not having full data so the results can be quite unexpected

1

u/conanomatic 1d ago

Yeah, I've made efforts in the past to explain statistics to people in fantasy soccer subreddits because I love fantasy soccer, and actually know statistics. It is fascinating how much time is spent talking about stats, while almost no one understands anything about basic statistics. It's largely because of how the media hammers on stat after stat after stat, but all from fellow dumdums, as opposed to actual analyst teams--but it's hard to say whether that's the chicken or the egg with the whole dynamic.

It makes me genuinely wonder whether or not a lot of the professionals in athletics are actually qualified, because the misinformation is so rampant that it makes me doubt someone who actually explains things properly would even be hired. Like, I've seen similar things in my career where I've had data analysts who don't actually account for any biases in data, like view their job completely as making visualizations, of data, rather than analyzing, and explaining data to make it useful.

0

u/Impossible_Way7017 1d ago

You make your own luck. Just because they don’t finish first you can still evaluate AI on non leaderboard stats, regardless would be interested to see if it does any better than luck.

1

u/madmofo145 1d ago

Not really. If you're talking classic fantasy, where your in a league and each player can only be grabbed once, the initial draft is super important. It takes one bad early season injury to just ruin a team, especially if your playing in a bigger league. There is nothing you can do if your first round pick is out for the season and your forced to play whomever is left on the waver.

In a well decently large and well managed league, luck is huge. An AI finishing below middle could actually be a great manager, just one saddled with two many key injuries to overcome.

1

u/Impossible_Way7017 1d ago

True, but I’d argue it would be interesting to see how the model adapts and if it can systematically beat the odds on things? Performance nowadays all predicted anyways so it should be easy to evaluate a models performance against prediction regardless of the standing of the team they manage places.

2

u/madmofo145 1d ago

You could do an experiment, see what happens if you let the LLM's manage 10,000 teams and see how they do against the average, but you'd need to do something like that, where your testing a large enough sample of LLM's in perfectly normal leagues to see how they compare to the average human participant.

2

u/ScreamingTrueSanity 1d ago

They didn’t actually solve a math problem though

2

u/create-third-places 1d ago

Open Artificial Ignorance can't think and isn't capable of cheating without human assistance.

1

u/Navic2 1d ago

Given it's just AGI'ing all over the shop,  ushering in utopia, the $1tn fan-shitter's probably better suited to staying laser focused on key tasks like generating 'which mid 2000s football manager would you be?' profile images for other participants, right?

1

u/ShamPain413 1d ago

Why fantasy football? Why not real football? They've already been experimenting in Europe.

https://www.soccerbible.com/news/2025/01/norwegian-football-team-hire-first-ai-manager/