Hardcore Reviews
Hardcore Reviews

17 hours vs 16 days: which AI can actually work overnight

A continuous-autonomy comparison (October 7, 2026 basis; a different axis from the site's October 5 single-output-capacity review, which it links rather than repeats): how long can one AI work on a single task, and what does it deliver. Four rows: Ant's Ling-3.1-flash wrote a Lua-to-x86-64 compiler from scratch in about 17 hours (178/182 tests passing, self-reported) and ported a C image library to Rust in about 20 hours with an 8.015x speedup (30/30 checks); Alibaba's Qwen3.8-Max ran a 16-day autonomous build of oh-my-cli (265 commits/127 PRs/151 issues, with a public GitHub trace) plus a ~125-hour paper reproduction that beat the original method (AIME24 49.58% to 52.29%, 7,600 lines of code, 33 training rounds); Google's Argon handled the 800K+ line Zircon migration and a 2.7x libgav1 rewrite (DeepSWE v1.1 77.9%, general public still locked out); reference row DeepSeek-V4-Flash-Vision-Exp scores 59.3 on the same benchmark (MIT, self-hostable). Table columns: model / longest case / verifiability / access and cost / sourcing. An architecture section answers "Ling vs Qwen": Ling uses 7 KDA + 1 Gated MLA layers, Qwen 3 Gated DeltaNet + 1 Gated Attention, both 512-expert MoE. License red lines reported faithfully: Qwen3.8-Max open weights ship under a custom qwen3.8-max license (revenue share, threshold not finalized), not Apache 2.0 (only the 27B sibling is), and the open build is text-only with forced thinking - not the API version with vision, 1M context and tools; the "half-open" controversy is presented from both sides; 4.89TB of weights need 24 GPUs (BF16) or 16 (FP8). All numbers self-reported; DeepSWE v1.1 compared only within v1.1.

Published October 7, 202610 min read
<!-- long-autonomy-comparison-review | review | 17 hours vs 16 days: which AI can actually work overnight -->

Benchmark leaderboards answer one question: which model scores highest on a test taken in a single sitting. They stay silent on the question that matters if you actually want to delegate work: how long can a model keep going with no human in the loop? This review measures continuous autonomous work time: the hours or days a system can run one task end to end, self-check as it goes, recover from failures, and return something verifiable.

This is deliberately a different axis from our million-token output comparison, which measured how much text a model can emit in one response and what that volume costs. Single-output capacity is a lung test; overnight autonomy is a marathon. The model must maintain state across hours, decide when a checkpoint is good enough, and stay on target with nobody watching. A model can ace the first test and fail the second.

One warning before the numbers: every result here is officially self-reported, with no third-party reproduction at the time of writing. What separates the contenders is verifiability - raw traces or re-runnable test suites outsiders can audit, versus a blog post. We grade each case on that scale.

The scoreboard at a glance

One snapshot before the details. Longest documented continuous case per model, with how each claim can be checked:

ModelLongest documented caseVerifiabilityAccess and costSourcing
Ling-3.1-flash (Ant Group)~17h writing a Lua to x86-64 ELF compiler from scratch: 178 of 182 tests pass (97.8%); ~20h porting a C image library to Rust at an 8.015x speedup, all 30 correctness checks passOfficial blog; test suites are re-runnableFree trial until around Oct 13 (web chat, Ant MaaS, Vercel AI Gateway); trial context capped at 256K, 1M on paid tierOfficial self-reported
Qwen3.8-Max (Alibaba)16-day autonomous build of oh-my-cli: 265 commits, 127 PRs, 151 issues; separate ~5-day, 125h paper reproduction (7,600 lines, 1,100+ actions, 33 GPU training rounds) lifting AIME24 from 49.58% to 52.29%Full process trace public on GitHub (qwen-code-dev-bot/oh-my-cli)Open weights 4.89TB BF16 (needs 24 GPUs BF16, 16 FP8, 8 NVFP4); API from $2/M input, $6/M output (Singapore standard)Official self-reported, trace auditable
Gemini 4 Argon (Google DeepMind)800,000+ line migration of the Fuchsia Zircon kernel from C/C++ to Rust; libgav1: 32,000 lines of SIMD rewritten, 2.7x faster with frame-identical output; DeepSWE v1.1 77.9%Official blog; states automated tests, simulation, and human review before deploymentIntro pricing $2/M input, $10/M output ($4/$20 after intro); but Fairwind partners only - the general public is still locked outOfficial self-reported
Reference: DeepSeek-V4-Flash-Vision-ExpDeepSWE v1.1 59.3%, same benchmark version as Argon's 77.9%Public model card; weights and benchmark downloadableMIT-licensed open weights, runnable locallyOfficial self-reported

Two notes. First, the Argon and DeepSeek DeepSWE numbers sit on the same v1.1 benchmark version, so 77.9 versus 59.3 is a fair same-paper comparison; neither may be compared against other benchmark versions. Second, raw hours are not the only yardstick: a 17-hour run with a re-runnable test suite beats a longer run you cannot inspect.

Ant Ling-3.1-flash: 17 hours, 178 out of 182

Ant Group's Ling team shipped Ling-3.1-flash on October 6: 560B total parameters, 25B active per token, a 1M context ceiling, and a 512-expert MoE router selecting 8 routed experts plus 1 shared expert. It is roughly 4.5x the total scale of its predecessor, Ling-3.0-flash. The autonomy showcase came from the official blog, which documented two long runs.

The first run lasted about 17 hours: the model wrote a Lua to x86-64 ELF compiler from zero and passed 178 of 182 tests in the accompanying suite, a 97.8% pass rate. The second ran about 20 hours: porting a C image library to Rust, landing an 8.015x speedup with all 30 correctness checks passing. Both cases ship with re-runnable test sets, which puts them above pure narrative claims - though the runs themselves remain official self-reported numbers.

Access is the easiest of the three. The free trial runs until around October 13 through three doors: the official web chat, Ant's MaaS platform, and Vercel AI Gateway, which lists the free period explicitly. The catch: the trial caps context at 256K; the full 1M opens on the paid tier, and open-sourcing is planned without a date.

Qwen3.8-Max: 16 days on the record

Alibaba's Qwen3.8-Max is the heavyweight of this comparison and the only contender whose autonomy claim you can audit commit by commit. The headline case: a 16-day autonomous engineering effort on a project called oh-my-cli, which by the July 30 snapshot had accumulated 265 commits, 127 pull requests, and 151 issues. The system ran an issue state machine, a watchdog, and a CI self-healing loop: it decomposed work into tracked issues, monitored its own runs, and fixed broken builds unattended.

What elevates this from press release to evidence is that the full process trace is public on GitHub under qwen-code-dev-bot/oh-my-cli. Every commit, every issue transition, every recovery is readable. Among the three contenders, this is the strongest verifiability grade.

The second case is a paper reproduction that lasted about five days, or 125 hours: 7,600 lines of code, more than 1,100 actions, and 33 rounds of GPU training, ending with AIME24 accuracy lifted from 49.58% to 52.29%. Across four rounds the system proposed 18 of its own improvement ideas and ultimately beat the original paper's method by 2.7 points. Again self-reported, but with the methodology and runs documented for inspection.

Self-hosting is steep: the open weights ship as 213 shards totaling 4.89TB in BF16, needing roughly 24 GPUs in BF16, 16 in FP8, or 8 in NVFP4. Most readers will take the API instead: $2 per million input tokens ($0.25 cached) and $6 per million output on the Singapore standard tier, $1.65/$4.951 in Virginia-region pricing, batch jobs at half price - roughly 40% of Opus 5's input price and 24% of its output, per CGTN.

Gemini 4 Argon: 800K lines behind a locked door

Google DeepMind's Argon, announced September 30 as the first Gemini 4 family model, plays in the same league but in a different currency: scale. The flagship internal case is a migration of more than 800,000 lines of the Fuchsia Zircon kernel from C/C++ to Rust. A second case rewrote 32,000 lines of SIMD code in libgav1, coming out 2.7x faster than the original Rust port with frame-identical output. Other documented wins include a datacenter memory optimization freeing 300-plus TiB. Google says deployments followed automated tests, simulation, and human review.

On the long-horizon software engineering benchmark DeepSWE v1.1, Argon posts 77.9%, ahead of GPT-6 Astra at 74.1, Claude Opus 5.5 at 74.2, and Claude Fable 5.1 at 67.4 - a same-version comparison only. Its output ceiling is 1M tokens per response, about 16x the previous generation. On safety, Google claims the strongest resilience to indirect prompt injection of its generation, with chain-of-thought and action monitoring that can abort execution mid-run, plus sandbox isolation - relevant when a model runs unattended for days.

Honest weaknesses, from Google's own table: FrontierSWE v2 at 55.0 trails Astra's 65.5, and Terminal-bench 4.0 at 57.4 trails Opus 5.5's 66.4. The bigger problem is access. Argon is in controlled release through the Fairwind Program - 650-plus partners spanning government, critical infrastructure, and core platforms, with trusted defenders receiving a version without cyber guardrails. The general public still locked out, the paid API not open, no date given. The intro pricing of $2 per million input and $10 per million output tokens ($4/$20 afterward) stays on paper until that changes.

The reference row: what MIT-licensed open weights buy you

DeepSeek-V4-Flash-Vision-Exp anchors the table as the reference row: 59.3% on DeepSWE v1.1, the same benchmark version as Argon's 77.9%. It is not competitive on the score; what it offers instead is total verifiability: MIT-licensed open weights you can download, run locally, and benchmark yourself - the only row where the score is limited by your hardware rather than a vendor's willingness to let you in. For how this family scores on coding benchmarks specifically, see our DeepSeek V4 Flash benchmark review.

Ling vs Qwen: two opposite answers to one question

The most common real-world question - Ling or Qwen? - deserves an architecture-level answer, because the two teams made opposite bets. Both are 512-expert MoE models. The difference is where they spend their attention budget.

Ling-3.1-flash leans hard into linear attention: 7 layers of KDA (a linear kernel-delta attention) against a single layer of Gated MLA. Qwen3.8-Max keeps a larger share of full attention: 3 layers of Gated DeltaNet against 1 layer of Gated Attention. The trade-off is textbook: linear-heavy stacks get dramatically cheaper as context grows - what a 17-hour run with a huge working set needs - while more full-attention layers preserve retrieval precision over long dependencies, what a 16-day run needs to catch its own mistakes. Neither bet is wrong; they optimize for different shapes of long work. For how these two families compare on pure coding and reasoning scores, see our flagship coding comparison.

In practice, though, architecture sets the cost curve, not the reliability curve. Qwen's 16-day run worked because of scaffolding - the issue state machine, the watchdog, and the CI self-healing loop layered on top of the model. Any overnight deployment lives or dies on that outer loop as much as on the weights.

The half-open problem

Qwen3.8-Max's open weights come with asterisks that matter. The released model, Qwen3.8-2.4T-A95B, is text-only, exposes a native 262,144-token context, and forces thinking mode - setting enable_thinking to false throws an exception rather than silently complying. The API version is a different product: it adds vision, defaults to 1M context, and ships built-in tools. What you download is not what you call.

Then there is the license, and this is the red line for any commercial plan: it is a custom qwen3.8-max license with revenue sharing for large commercial deployments - not Apache 2.0. The revenue-share threshold and ratio have not been finalized; the official position is that they are still being worked out. Only the 27B sibling model carries the Apache-2.0 license.

Critics call this half-open: a text-only, forced-thinking release under an unfinished commercial license, with the flagship capabilities reserved for the paid API. Defenders make a fair counterpoint: the full 2.4T weights are downloadable at all, which several competitors - Argon most obviously - do not offer; the custom license is more permissive than a closed API for most small and mid-size commercial uses; and publishing an unfinished threshold beats silently changing terms later. Both readings fit the same facts; budget for the uncertainty if your plan depends on where the threshold lands.

How to pick your overnight worker

Match the claim to your verification appetite. If you want evidence you can audit yourself, Qwen's public oh-my-cli trace is the strongest record in this comparison, and self-hosting the MIT-licensed DeepSeek-V4-Flash-Vision-Exp removes vendor trust entirely. If you want to test long-run autonomy this week without spending anything, Ling's free window - until around October 13 - is the cheapest on-ramp, with the 256K trial cap as the trade. If you are inside the Fairwind Program, Argon's numbers justify trying it; if not, no benchmark SOTA changes the fact that you cannot buy it yet. For where the previous flagship generation stood, our Claude Sonnet 5.5 release coverage provides the baseline these three are measured against.

A closing rule of thumb: prefer cases with re-runnable tests or public traces over blog numbers, however impressive. Everything in this piece is self-reported. The real difference between the contenders is how much of the claim you can check yourself.

FAQ

Q1: How long can AI agents work autonomously as of late 2026? A1: Documented cases range from about 17 hours (Ling-3.1-flash writing a compiler from scratch) to 16 days (Qwen3.8-Max building oh-my-cli). All figures are officially self-reported; the strongest one is auditable commit by commit on GitHub. The real ceiling depends on scaffolding as much as on the model itself.

Q2: Is the Ling-3.1-flash 17-hour compiler result independently verified? A2: No. It is officially self-reported with no third-party reproduction yet. However, the vendor published a re-runnable test suite (178 of 182 tests passing), so outsiders can reproduce part of the claim.

Q3: Can the Qwen3.8-Max open weights be used commercially? A3: Yes, but under a custom qwen3.8-max license with revenue sharing for large commercial use - it is not an Apache 2.0 release. The threshold and ratio are not finalized. Only the 27B sibling model carries the Apache-2.0 license, and the open weights are text-only with forced thinking, unlike the API version.

Q4: Is Gemini 4 Argon available to the general public? A4: Not yet. It is in a controlled release through the Fairwind Program with 650-plus partners in government, critical infrastructure, and core platforms. The paid API is not open and Google has given no public date.

Q5: Ling-3.1-flash or Qwen3.8-Max for long coding tasks? A5: Ling is free to trial until around October 13 and is cheap on long context thanks to its linear-attention-heavy architecture. Qwen offers auditable long-run traces and a more capable API with vision and tools, at higher cost. A practical path: test on Ling's free window first, then compare survivors against Qwen's API.

Search Keywords

  • ling 3.1 flash vs qwen3.8 max
  • how long can ai agents work autonomously
  • best ai for long running coding tasks

This article was drafted with AI assistance and reviewed by a human editor.

This article is AI-assisted and human-edited. Last updated: 2026-10-07

FAQ

How long can AI agents work autonomously as of late 2026?
Documented cases range from about 17 hours (Ling-3.1-flash writing a compiler from scratch) to 16 days (Qwen3.8-Max building oh-my-cli). All figures are officially self-reported; the strongest one is auditable commit by commit on GitHub. The real ceiling depends on scaffolding as much as on the model itself.
Is the Ling-3.1-flash 17-hour compiler result independently verified?
No. It is officially self-reported with no third-party reproduction yet. However, the vendor published a re-runnable test suite (178 of 182 tests passing), so outsiders can reproduce part of the claim.
Can the Qwen3.8-Max open weights be used commercially?
Yes, but under a custom qwen3.8-max license with revenue sharing for large commercial use - it is not an Apache 2.0 release. The threshold and ratio are not finalized. Only the 27B sibling model carries the Apache-2.0 license, and the open weights are text-only with forced thinking, unlike the API version.
Is Gemini 4 Argon available to the general public?
Not yet. It is in a controlled release through the Fairwind Program with 650-plus partners in government, critical infrastructure, and core platforms. The paid API is not open and Google has given no public date.
Ling-3.1-flash or Qwen3.8-Max for long coding tasks?
Ling is free to trial until around October 13 and is cheap on long context thanks to its linear-attention-heavy architecture. Qwen offers auditable long-run traces and a more capable API with vision and tools, at higher cost. A practical path: test on Ling's free window first, then compare survivors against Qwen's API.

Related

Hardcore Reviews

One Million Output Tokens: Which Model Finishes the Job

A four-way comparison of coding/general models (October 5, 2026 basis; complements rather than repeats the site's late-September pricing review). The axis this time is single-response output length: Gemini 4 Argon stretched output from 64K to 1 million tokens (officially the industry's largest), changing the battlefield for long tasks. Four positions: the controlled flagship Argon (released September 30; officially first place on 13 of 18 benchmarks and a 77.9% DeepSWE v1.1 SOTA, yet in a controlled rollout the general public cannot buy; Artificial Analysis independently scores it slightly lower; introductory $2/$10 rising to $4/$20), the reigning efficiency benchmark Claude Sonnet 5.5 (September 28; Terminal-Bench 4.0 at 70.6%; $2/$10 on three clouds), the cache-economics pick GPT-6.1 Sol ($2/$0.10/$10 - cached input 95% cheaper), and the open-source reference GLM-5.3 (about 750GB of weights on ModelScope for self-hosting, domestic API $1.40/$4.40, open-source coding SOTA with CyberGym at 84.5%). Red lines honored: GLM's Terminal Bench 3.0 and Sonnet's Terminal-Bench 4.0 are different benchmark versions and are never compared directly; both official and independent assessments of Argon are presented. No benchmarks were re-run and no winner is crowned - conclusions follow the need: wait for Argon's wider access, pick Sonnet 5.5 for balanced value, Sol for cache-heavy agent pipelines, Sonnet for multi-cloud, GLM-5.3 for open self-hosting.

Oct 5, 202610 min read
Hardcore Reviews

30-Second Club: Kling 4.0, Seedance 2.5, Wan 3.0, LTX-2.5

A four-way video-model comparison (October 3, 2026 basis; complements rather than repeats the site's August five-way closed-cloud review and the Seedance 2.5 vs MiniMax H3 head-to-head). Framed by the observation that 30-second native single-pass output has become the entry ticket among flagship models, it picks one model per quadrant - cloud flagship, reigning benchmark, cloud value, open self-hosted: Kling 4.0 (October launch; 10-bit HDR, 2-minute extension, 15 references; specs announced but unverified), Seedance 2.5 (live since July 31; 30-second native output, 50 references, sub-second local editing; production deployments at XCMG, XPeng and Differential Intelligence Flight), Wan 3.0 (live since August 24; API from 0.3 CNY per second, five office document formats direct-to-video, no agent capabilities so orchestration is DIY), and LTX-2.5 (the only open weights; 66 GiB, 8-step distillation, ltx-trainer; custom community license with a USD 10 million revenue threshold). With no benchmark runs of its own, the piece refuses to crown a winner and concludes by need: Seedance for a working flagship today, Kling 4.0 to watch for pro specs, Wan 3.0 for value and batch work, LTX-2.5 for data privacy. Sora is excluded over conflicting sources and Veo unverified.

Oct 3, 202610 min read
Hardcore Reviews

Half Price vs One-Fifth: The Flagship Alternative Shake-Up

A value-focused comparison of flagship-adjacent coding models as of October 2026: Claude Sonnet 5.5 (official basis via machine-intelligence press: 30%+ faster, Terminal-Bench 4.0 up from 10.3% to 70.6%, priced at half of Opus 5.5) vs GPT-6.1 Sol (official basis: near-Astra intelligence at one-fifth the API price, with a paid Ultrafast speed tier) vs Claude Opus 5.5 (the reference point, $4/$20 per million tokens) vs DeepSeek V4.1 (the open-weights reference). Discipline: harnesses differ so cross-model scores cannot be compared - this piece only uses relative-to-own-flagship ratios, does not run its own evals, and does not compute Sonnet 5.5's unpublished dollar pricing. Four scenario verdicts: budget-conscious daily coding, flagship-ceiling complex work, and compliance-driven private deployment each have a winner. Complements the site's 2026-08 comprehensive and flagship-reasoning comparisons.

Oct 1, 202610 min read