Benchmark leaderboards answer one question: which model scores highest on a test taken in a single sitting. They stay silent on the question that matters if you actually want to delegate work: how long can a model keep going with no human in the loop? This review measures continuous autonomous work time: the hours or days a system can run one task end to end, self-check as it goes, recover from failures, and return something verifiable.
This is deliberately a different axis from our million-token output comparison, which measured how much text a model can emit in one response and what that volume costs. Single-output capacity is a lung test; overnight autonomy is a marathon. The model must maintain state across hours, decide when a checkpoint is good enough, and stay on target with nobody watching. A model can ace the first test and fail the second.
One warning before the numbers: every result here is officially self-reported, with no third-party reproduction at the time of writing. What separates the contenders is verifiability - raw traces or re-runnable test suites outsiders can audit, versus a blog post. We grade each case on that scale.
The scoreboard at a glance
One snapshot before the details. Longest documented continuous case per model, with how each claim can be checked:
| Model | Longest documented case | Verifiability | Access and cost | Sourcing |
|---|---|---|---|---|
| Ling-3.1-flash (Ant Group) | ~17h writing a Lua to x86-64 ELF compiler from scratch: 178 of 182 tests pass (97.8%); ~20h porting a C image library to Rust at an 8.015x speedup, all 30 correctness checks pass | Official blog; test suites are re-runnable | Free trial until around Oct 13 (web chat, Ant MaaS, Vercel AI Gateway); trial context capped at 256K, 1M on paid tier | Official self-reported |
| Qwen3.8-Max (Alibaba) | 16-day autonomous build of oh-my-cli: 265 commits, 127 PRs, 151 issues; separate ~5-day, 125h paper reproduction (7,600 lines, 1,100+ actions, 33 GPU training rounds) lifting AIME24 from 49.58% to 52.29% | Full process trace public on GitHub (qwen-code-dev-bot/oh-my-cli) | Open weights 4.89TB BF16 (needs 24 GPUs BF16, 16 FP8, 8 NVFP4); API from $2/M input, $6/M output (Singapore standard) | Official self-reported, trace auditable |
| Gemini 4 Argon (Google DeepMind) | 800,000+ line migration of the Fuchsia Zircon kernel from C/C++ to Rust; libgav1: 32,000 lines of SIMD rewritten, 2.7x faster with frame-identical output; DeepSWE v1.1 77.9% | Official blog; states automated tests, simulation, and human review before deployment | Intro pricing $2/M input, $10/M output ($4/$20 after intro); but Fairwind partners only - the general public is still locked out | Official self-reported |
| Reference: DeepSeek-V4-Flash-Vision-Exp | DeepSWE v1.1 59.3%, same benchmark version as Argon's 77.9% | Public model card; weights and benchmark downloadable | MIT-licensed open weights, runnable locally | Official self-reported |
Two notes. First, the Argon and DeepSeek DeepSWE numbers sit on the same v1.1 benchmark version, so 77.9 versus 59.3 is a fair same-paper comparison; neither may be compared against other benchmark versions. Second, raw hours are not the only yardstick: a 17-hour run with a re-runnable test suite beats a longer run you cannot inspect.
Ant Ling-3.1-flash: 17 hours, 178 out of 182
Ant Group's Ling team shipped Ling-3.1-flash on October 6: 560B total parameters, 25B active per token, a 1M context ceiling, and a 512-expert MoE router selecting 8 routed experts plus 1 shared expert. It is roughly 4.5x the total scale of its predecessor, Ling-3.0-flash. The autonomy showcase came from the official blog, which documented two long runs.
The first run lasted about 17 hours: the model wrote a Lua to x86-64 ELF compiler from zero and passed 178 of 182 tests in the accompanying suite, a 97.8% pass rate. The second ran about 20 hours: porting a C image library to Rust, landing an 8.015x speedup with all 30 correctness checks passing. Both cases ship with re-runnable test sets, which puts them above pure narrative claims - though the runs themselves remain official self-reported numbers.
Access is the easiest of the three. The free trial runs until around October 13 through three doors: the official web chat, Ant's MaaS platform, and Vercel AI Gateway, which lists the free period explicitly. The catch: the trial caps context at 256K; the full 1M opens on the paid tier, and open-sourcing is planned without a date.
Qwen3.8-Max: 16 days on the record
Alibaba's Qwen3.8-Max is the heavyweight of this comparison and the only contender whose autonomy claim you can audit commit by commit. The headline case: a 16-day autonomous engineering effort on a project called oh-my-cli, which by the July 30 snapshot had accumulated 265 commits, 127 pull requests, and 151 issues. The system ran an issue state machine, a watchdog, and a CI self-healing loop: it decomposed work into tracked issues, monitored its own runs, and fixed broken builds unattended.
What elevates this from press release to evidence is that the full process trace is public on GitHub under qwen-code-dev-bot/oh-my-cli. Every commit, every issue transition, every recovery is readable. Among the three contenders, this is the strongest verifiability grade.
The second case is a paper reproduction that lasted about five days, or 125 hours: 7,600 lines of code, more than 1,100 actions, and 33 rounds of GPU training, ending with AIME24 accuracy lifted from 49.58% to 52.29%. Across four rounds the system proposed 18 of its own improvement ideas and ultimately beat the original paper's method by 2.7 points. Again self-reported, but with the methodology and runs documented for inspection.
Self-hosting is steep: the open weights ship as 213 shards totaling 4.89TB in BF16, needing roughly 24 GPUs in BF16, 16 in FP8, or 8 in NVFP4. Most readers will take the API instead: $2 per million input tokens ($0.25 cached) and $6 per million output on the Singapore standard tier, $1.65/$4.951 in Virginia-region pricing, batch jobs at half price - roughly 40% of Opus 5's input price and 24% of its output, per CGTN.
Gemini 4 Argon: 800K lines behind a locked door
Google DeepMind's Argon, announced September 30 as the first Gemini 4 family model, plays in the same league but in a different currency: scale. The flagship internal case is a migration of more than 800,000 lines of the Fuchsia Zircon kernel from C/C++ to Rust. A second case rewrote 32,000 lines of SIMD code in libgav1, coming out 2.7x faster than the original Rust port with frame-identical output. Other documented wins include a datacenter memory optimization freeing 300-plus TiB. Google says deployments followed automated tests, simulation, and human review.
On the long-horizon software engineering benchmark DeepSWE v1.1, Argon posts 77.9%, ahead of GPT-6 Astra at 74.1, Claude Opus 5.5 at 74.2, and Claude Fable 5.1 at 67.4 - a same-version comparison only. Its output ceiling is 1M tokens per response, about 16x the previous generation. On safety, Google claims the strongest resilience to indirect prompt injection of its generation, with chain-of-thought and action monitoring that can abort execution mid-run, plus sandbox isolation - relevant when a model runs unattended for days.
Honest weaknesses, from Google's own table: FrontierSWE v2 at 55.0 trails Astra's 65.5, and Terminal-bench 4.0 at 57.4 trails Opus 5.5's 66.4. The bigger problem is access. Argon is in controlled release through the Fairwind Program - 650-plus partners spanning government, critical infrastructure, and core platforms, with trusted defenders receiving a version without cyber guardrails. The general public still locked out, the paid API not open, no date given. The intro pricing of $2 per million input and $10 per million output tokens ($4/$20 afterward) stays on paper until that changes.
The reference row: what MIT-licensed open weights buy you
DeepSeek-V4-Flash-Vision-Exp anchors the table as the reference row: 59.3% on DeepSWE v1.1, the same benchmark version as Argon's 77.9%. It is not competitive on the score; what it offers instead is total verifiability: MIT-licensed open weights you can download, run locally, and benchmark yourself - the only row where the score is limited by your hardware rather than a vendor's willingness to let you in. For how this family scores on coding benchmarks specifically, see our DeepSeek V4 Flash benchmark review.
Ling vs Qwen: two opposite answers to one question
The most common real-world question - Ling or Qwen? - deserves an architecture-level answer, because the two teams made opposite bets. Both are 512-expert MoE models. The difference is where they spend their attention budget.
Ling-3.1-flash leans hard into linear attention: 7 layers of KDA (a linear kernel-delta attention) against a single layer of Gated MLA. Qwen3.8-Max keeps a larger share of full attention: 3 layers of Gated DeltaNet against 1 layer of Gated Attention. The trade-off is textbook: linear-heavy stacks get dramatically cheaper as context grows - what a 17-hour run with a huge working set needs - while more full-attention layers preserve retrieval precision over long dependencies, what a 16-day run needs to catch its own mistakes. Neither bet is wrong; they optimize for different shapes of long work. For how these two families compare on pure coding and reasoning scores, see our flagship coding comparison.
In practice, though, architecture sets the cost curve, not the reliability curve. Qwen's 16-day run worked because of scaffolding - the issue state machine, the watchdog, and the CI self-healing loop layered on top of the model. Any overnight deployment lives or dies on that outer loop as much as on the weights.
The half-open problem
Qwen3.8-Max's open weights come with asterisks that matter. The released model, Qwen3.8-2.4T-A95B, is text-only, exposes a native 262,144-token context, and forces thinking mode - setting enable_thinking to false throws an exception rather than silently complying. The API version is a different product: it adds vision, defaults to 1M context, and ships built-in tools. What you download is not what you call.
Then there is the license, and this is the red line for any commercial plan: it is a custom qwen3.8-max license with revenue sharing for large commercial deployments - not Apache 2.0. The revenue-share threshold and ratio have not been finalized; the official position is that they are still being worked out. Only the 27B sibling model carries the Apache-2.0 license.
Critics call this half-open: a text-only, forced-thinking release under an unfinished commercial license, with the flagship capabilities reserved for the paid API. Defenders make a fair counterpoint: the full 2.4T weights are downloadable at all, which several competitors - Argon most obviously - do not offer; the custom license is more permissive than a closed API for most small and mid-size commercial uses; and publishing an unfinished threshold beats silently changing terms later. Both readings fit the same facts; budget for the uncertainty if your plan depends on where the threshold lands.
How to pick your overnight worker
Match the claim to your verification appetite. If you want evidence you can audit yourself, Qwen's public oh-my-cli trace is the strongest record in this comparison, and self-hosting the MIT-licensed DeepSeek-V4-Flash-Vision-Exp removes vendor trust entirely. If you want to test long-run autonomy this week without spending anything, Ling's free window - until around October 13 - is the cheapest on-ramp, with the 256K trial cap as the trade. If you are inside the Fairwind Program, Argon's numbers justify trying it; if not, no benchmark SOTA changes the fact that you cannot buy it yet. For where the previous flagship generation stood, our Claude Sonnet 5.5 release coverage provides the baseline these three are measured against.
A closing rule of thumb: prefer cases with re-runnable tests or public traces over blog numbers, however impressive. Everything in this piece is self-reported. The real difference between the contenders is how much of the claim you can check yourself.
FAQ
Q1: How long can AI agents work autonomously as of late 2026? A1: Documented cases range from about 17 hours (Ling-3.1-flash writing a compiler from scratch) to 16 days (Qwen3.8-Max building oh-my-cli). All figures are officially self-reported; the strongest one is auditable commit by commit on GitHub. The real ceiling depends on scaffolding as much as on the model itself.
Q2: Is the Ling-3.1-flash 17-hour compiler result independently verified? A2: No. It is officially self-reported with no third-party reproduction yet. However, the vendor published a re-runnable test suite (178 of 182 tests passing), so outsiders can reproduce part of the claim.
Q3: Can the Qwen3.8-Max open weights be used commercially? A3: Yes, but under a custom qwen3.8-max license with revenue sharing for large commercial use - it is not an Apache 2.0 release. The threshold and ratio are not finalized. Only the 27B sibling model carries the Apache-2.0 license, and the open weights are text-only with forced thinking, unlike the API version.
Q4: Is Gemini 4 Argon available to the general public? A4: Not yet. It is in a controlled release through the Fairwind Program with 650-plus partners in government, critical infrastructure, and core platforms. The paid API is not open and Google has given no public date.
Q5: Ling-3.1-flash or Qwen3.8-Max for long coding tasks? A5: Ling is free to trial until around October 13 and is cheap on long context thanks to its linear-attention-heavy architecture. Qwen offers auditable long-run traces and a more capable API with vision and tools, at higher cost. A practical path: test on Ling's free window first, then compare survivors against Qwen's API.
Search Keywords
- ling 3.1 flash vs qwen3.8 max
- how long can ai agents work autonomously
- best ai for long running coding tasks
This article was drafted with AI assistance and reviewed by a human editor.