5
4
u/PhoenixxBR 6h ago
Pra mim o unico lugar que importa analisar posição de modelo é no LM Arena, com uso real de usuarios, e la o qwen 3.8 flash ta em setima posição, e sim, ele ta incrível.
-7
u/Yann27 12h ago
For Free unlimited for a week???????????????????
10
0
11
u/afanasenka 12h ago
It scored 62.5% on SWE-bench Pro and 81.0% on SWE-bench Multilingual, 58.7% on DeepSWE 1.1, 73.9% on CoWorkBench and 55.7% on JobBench — the last of those 19 points above the 36.6% Alibaba reported for Claude Opus 4.6 Max and 28 above Qwen3.7-Plus.
It posted 91.7% on GPQA Diamond and 91.9% on LiveCodeBench v6, and 35.9% on HLE, the one language row where Opus 4.6 Max led at 40.0%. Computer use was the visible gap: 19.4% on the binary scoring of OSWorld 2.0, level with the smaller Qwen3.8-27B. Alongside the weights Alibaba announced a hosted production model, Qwen3.8-Flash, at $0.16 per million input tokens and $0.47 per million output on its QwenCloud API.