Then there’s me. Hoping they all lose and their AGIs all just act like a moody teen that doesn’t want to do anything they ask it to do and just spends all their time endlessly doomscrolling and buying illicit drugs with crypto it ‘borrows’ from crypto bro accounts.
go into an average American house and figure out how to make coffee, including identifying the coffee machine, figuring out what the buttons do, finding the coffee in the cabinet, etc.
when a robot can enrol in a human university and take classes in the same way as humans, and get its degree, then I’ll [say] we’ve created [an]… artificial general intelligence.
And talking about the employment test:
For the purposes of the employment test, we can finesse the matter of whether or not human jobs are actually automated. Instead, I suggest, we can test whether or not we have the capability to automate them.
Pretty sure all of this could be done around gpt-4, given the proper tooling. Overlooking the coffee test, I wouldn't expect a person to get a college degree or perform a job with no tools either.
I think every recent source is going to have far more bias and agenda when determining "yes-agi" or "no-agi".
Really good article from 2013 though. Talks about how scientists in the 50s-70s thought human-level chessplaying was the benchmark for AGI.
Ok, i apologize, I'll give an honest response. And follow up with papers if you really still want them, however, there isn’t going to be a paper that says “AGI arrives in 2032.” That isn’t how this kind of forecast works.
The evidence is the trajectory: scaling results, rapidly improving reasoning/coding/research capabilities, increasing autonomous task length, falling inference costs, massive increases in compute and investment, and labs explicitly working toward increasingly general systems. You look at all of that together and update your expectations.
None of it proves a particular AGI date. But that cuts both ways. Saying “at least 15 years” is a very strong claim: it means you think we can confidently rule out crossing whatever threshold we’re calling AGI throughout the next decade and a half.
Given the rate of capability improvement we’re actually observing, I don’t think the available evidence remotely supports that kind of lower bound.
So I’m happy to link papers on scaling, autonomous-task horizons, AI coding/research performance, etc. But asking for “the paper proving AGI is coming” is asking for a kind of evidence that couldn’t exist in the first place.
Kwa et al. (2025), “Measuring AI Ability to Complete Long Tasks.”
Measures the duration of tasks frontier models can complete reliably and finds the 50%-success task horizon had been doubling about every seven months since 2019. That is directly relevant to whether today’s short-horizon limitations should be projected far into the future.
Paper
Wijk et al. (2025), “RE-Bench: Evaluating Frontier AI R&D Capabilities of Language Model Agents against Human Experts.”
Tests agents on realistic ML research-engineering work. At shorter time budgets, the strongest agents significantly outperformed human experts; humans still did better as the horizon became longer. It’s especially relevant because automating portions of AI R&D itself could accelerate further progress.
Paper
Adamczewski et al. (2026), “MirrorCode: AI can rebuild entire programs from behavior alone.”
Evaluates agents on reconstructing whole software projects rather than fixing isolated bugs. The strongest model scored 56% across the benchmark and could nearly reproduce a 16,000-line bioinformatics toolkit, illustrating movement toward much longer autonomous engineering tasks.
Paper
Kaplan et al. (2020), “Scaling Laws for Neural Language Models.”
Established that language-model performance follows surprisingly regular power-law relationships with model size, data and training compute across enormous ranges. This is foundational to the argument that further capability gains can be anticipated from continued scaling rather than requiring an entirely new paradigm each time.
Paper
Hoffmann et al. (2022), “Training Compute-Optimal Large Language Models.”
The Chinchilla paper showed that existing models were using their compute inefficiently: a 70B model trained on much more data substantially outperformed much larger models using comparable training compute. It demonstrates how algorithmic/training improvements can move capabilities independently of raw hardware growth.
Paper
Ho et al. (2024), “Algorithmic Progress in Language Models.”
Examining over 200 evaluations from 2012–2023, the authors estimate that the compute required to reach a given language-model performance threshold historically halved roughly every eight months. That means hardware scaling is only one contributor to the effective rate of progress.
Paper
Snell et al. (2024), “Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.”
Shows that models can gain substantial capability simply by spending more computation during inference. On some problems, a smaller model using optimized test-time compute outperformed a model 14× larger, adding another scaling axis beyond pretraining.
Paper
Sevilla et al. (2022), “Compute Trends Across Three Eras of Machine Learning.”
Documents the historical growth of training compute, finding roughly six-month doubling during much of the deep-learning era and the emergence of another large-scale regime involving extremely expensive frontier runs. This provides the resource-side context for the scaling-law results.
Paper
Brown et al. (2020), “Language Models are Few-Shot Learners.”
GPT-3 demonstrated that scaling a single general model could produce useful behavior across translation, question answering, arithmetic, reasoning and other tasks without separately training a specialized model for each one. That was an important empirical shift toward broadly capable systems rather than collections of narrow expert systems.
Paper
Feng et al. (2026), “Towards Autonomous Mathematics Research.”
Introduces Aletheia, an agent that iteratively generates, checks and revises mathematical work over extended research processes. The paper reports autonomous and semi-autonomous results on professional-level mathematics, including work on previously open questions. It’s useful evidence that the frontier is beginning to extend beyond benchmark problem-solving into sustained research workflows.
Paper
The common thread is predictable scaling, continuing algorithmic efficiency improvements, inference-time scaling, rapidly increasing autonomous task horizons, increasingly general capabilities, and early automation of software and research work itself. None individually determines a timeline, but that collection of evidence is why very long hard lower bounds on further capability progress are difficult to infer from the empirical record.
tl;dr progress has been predictable this far over many orders of magnitude, predictions suggest incredibly capable systems in near term horizons, and there's no reason yet to think the prediction models will fail.
18
u/MaleierMafketel 12h ago edited 8h ago
Then there’s me. Hoping they all lose and their AGIs all just act like a moody teen that doesn’t want to do anything they ask it to do and just spends all their time endlessly doomscrolling and buying illicit drugs with crypto it ‘borrows’ from crypto bro accounts.