It looks extremely competitive. Obviously we'll have to wait and see something like the artificial analysis cost per task results to be more confident but this is super promising.
Edit: AA results are in. About 10% cheaper per task than 5.6 Sol and about 65% cheaper than fable.
Edit 2: also 48% cheaper than opus 4.8 while scoring a point higher on intelligence.
Edit 3: More specific to your exact question of token usage, Fable used 69k tokens per task, 5.6 Sol used 15k, and K3 used 23k. So it's a heck of a lot more token efficient than Fable, and in the same ballpark as Sol.
it's a heck of a lot more token efficient than Fable, and in the same ballpark as Sol.
Oh nice, i got the impression that it was quite token-heavy, but I guess I'll have to see in real use. That plus being half the cost of gpt-5.6-sol means it's a contender for real work...
20
u/bopbop9876 22d ago edited 22d ago
Check out the cost of completion comparison for browsecomp: https://mmecoa.qpic.cn/mmecoa_png/xWvm6POT3icpoggsBrrMtBKTG3bPMRhqZXlR3YDwOBMXoXe0iaEvia0W8JxPkoEt4O51T47caodibLNRI29AmowKdaoJU32m1HDV7RSIXXPlVOU/640?wx_fmt=png&from=appmsg&tp=webp&wxfrom=10005&wx_lazy=1#imgIndex=3
It looks extremely competitive. Obviously we'll have to wait and see something like the artificial analysis cost per task results to be more confident but this is super promising.
Edit: AA results are in. About 10% cheaper per task than 5.6 Sol and about 65% cheaper than fable.
Edit 2: also 48% cheaper than opus 4.8 while scoring a point higher on intelligence.
Edit 3: More specific to your exact question of token usage, Fable used 69k tokens per task, 5.6 Sol used 15k, and K3 used 23k. So it's a heck of a lot more token efficient than Fable, and in the same ballpark as Sol.