Frontline Hotspot
Frontline Hotspot

Grok 4.5: SpaceXAI's First Coding+Agent Model, 1.5T and Co-Trained with Cursor

SpaceXAI released Grok 4.5 on July 8, 2026 - a 1.5T V9 model co-trained with Cursor, its first built specifically for coding and agents, priced $2/$6 with a 500K context and pitched internally as "comparable to Opus 4.7 but much faster"; no public benchmarks exist, so it positions as a cheap Opus-class substitute native to Cursor, differentiated from Opus 5 / GPT-5.6 Sol on ecosystem rather than leaderboard rank.

Published July 31, 20265 min read
<!-- grok-4-5-hotspot | hotspot | Grok 4.5: SpaceXAI's First Coding+Agent Model, 1.5T and Co-Trained with Cursor -->

On July 8, 2026, SpaceXAI publicly released Grok 4.5 - described in its own words as "our first model trained specifically for coding and agents," and the first product of the data merger that followed the Cursor acquisition. Two numbers grab the eye: 1.5 trillion parameters, and co-training with Cursor. Both need caveats. The 1.5T figure is a Musk disclosure reported by third parties; SpaceXAI has never officially confirmed it. "Co-trained with Cursor" doesn't just mean "fed Cursor's code" - it means Cursor's real-world development data was baked into a freshly pre-trained V9 foundation. To read this launch clearly, separate the facts, the specs, the capability ceiling, and where it actually stands.

Facts and Specs: 1.5T, V9, and a 500K Context Window

What's confirmed. Grok 4.5 went public on July 8, 2026, after roughly ten days of private beta inside SpaceX and Tesla (Musk first disclosed it on X on June 28). It shipped in three places: the Cursor editor, Grok Build (SpaceXAI's own coding-agent environment), and the SpaceXAI console - with Cursor offering double usage for the first week. EU access came later in July.

On specs, Grok 4.5 runs on a new V9 foundation. Multiple third-party reports cite Musk putting it at roughly 1.5 trillion parameters, nearly 3x the prior V8 - but note SpaceXAI has never officially confirmed this number; it is vendor-stated and media-relayed, not independently verified. What is confirmed is pricing and context: $2 per million input tokens, $6 per million output, with cache hits discounted 75% to $0.5/M; inputs longer than 200K tokens are billed at double. The context window is 500K - half of Grok 4.3's 1M - but vision input and configurable reasoning are retained; Musk said on X it would upgrade back to 1M. One detail easy to miss: xAI has merged into SpaceXAI, so the product entity and brand are both SpaceXAI, though the launch tweet still went out from the original xAI account - treat "xAI" and "SpaceXAI" as the same company when reading coverage.

Coding+Agent Capability: What Cursor Co-Training Actually Means

The "first model trained specifically for coding and agents" framing matters at the training stage, not inference. Cursor is the dominant AI coding editor; its parent Anysphere was acquired by SpaceX for about $60 billion, with the deal expected to close in Q3 2026 pending regulatory approval. Grok 4.5 folded Cursor's real development data - multi-file edits, debugging, long engineering chains - into a supplemental training stage on V9. SpaceXAI's claim is that this gives it "a structural advantage on coding-heavy tasks no other frontier model can currently match." Cursor itself called Grok 4.5 "our most powerful model yet" and stressed it was "the first we've built for more than software engineering."

Be honest, though: as of now there is no independent third-party benchmark (SWE-Bench, Terminal-Bench, etc.) with public Grok 4.5 scores. Musk's internal framing was "roughly comparable to Opus 4.7, but much faster" - note Opus 4.7, not Opus 5, which only launched on July 24. In other words, at launch Grok 4.5 was benchmarked against the prior Opus generation; its pitch is "Opus-class work at half the tokens and under half the price, and faster" - not "number one on the leaderboard." This is a different path from Anthropic and OpenAI's benchmark chasing; SpaceXAI chose the more utilitarian yardstick of "useful to SpaceX and Tesla engineers." So don't compare it to Opus 5 on SWE-Bench - it simply hasn't published a run.

One more thing often glossed over: the official line is "coding and agents," not coding alone. Cursor's own phrasing - "the first we've built for more than software engineering" - hints that its agentic capability (driving multi-step, multi-file changes inside the editor, running tests and looping back to fix) is the real pitch, not just code completion. But agentic capability likewise has no public Agentic Index or Terminal-Bench score to compare against, so here too you can only trust the vendor's word and test it yourself.

Competitive Position: vs Opus 5, GPT-5.6, and Claude Code

Place Grok 4.5 in the late-July landscape and its position is "a cheap Opus-class substitute with no public benchmarks." On price, $2/$6 is far cheaper than Anthropic's Opus 4.8 at $5/$25; against Opus 5 (also $5/$25, released July 24), input is 60% cheaper and output 76% cheaper. Versus GPT-5.6 Sol, input is level ($2) and output is slightly lower ($6 vs $8). But be clear: Opus 5 tops Artificial Analysis at 61 on the Intelligence Index and 55.3 on the Agentic Index; GPT-5.6 Sol sits at 59 on intelligence and, paired with Codex, leads the Coding Agent Index outright at 80 - Grok 4.5 has no comparable public score on any of these. In other words, on "per-million-token price" it is the cheapest; on "benchmark rank" it isn't on the board at all. And cheap per-token is not the same as cheap per-task - as the Opus 5 writeup noted, at max effort it can burn 8x the tokens on a single task. Grok 4.5 claims to "use half the tokens for the same work," which if true would lower per-task cost further, but that is a vendor claim, not third-party-verified; measure "total spend to finish one real task" yourself before concluding.

The real differentiator is the Cursor axis. Opus 5 and GPT-5.6 show their coding muscle paired with Claude Code and Codex respectively; Grok 4.5 is "native to Cursor" - it trained on Cursor's data and shipped inside the Cursor editor on day one. For teams already using Cursor as their primary IDE, switching to Grok 4.5 is nearly frictionless; for those on Claude Code, the migration cost is elsewhere. So this isn't "Grok 4.5 beats Opus 5" - it's "if you're already in the Cursor ecosystem, it's the smoothest path." SpaceXAI's bet is equally visible: use Cursor's data flywheel to lock in developers, and the next model (reportedly already training, around 2T parameters, with Cursor data wired in from pre-training) will dig that moat deeper.

What It Means for Developers

For developers, Grok 4.5 means "a cheap, editor-native option has joined the coding-agent field." If your work is high-volume multi-file editing and long debugging chains, $2/$6 plus the 75% cache discount compounds at batch scale; the 500K context, though shrunken, is enough for most single tasks, and long-context work can wait for the 1M upgrade. Three caveats: first, with no public benchmarks, don't let leaderboard scores drive your decision - measure "cost per task" and "error rate" on your own real workload; second, long inputs past 200K tokens bill at double, so cost out long-context runs first; third, EU and API channel rollout is still on the official schedule, and the most stable entry points today are Cursor and Grok Build. One more thing to keep in mind: binding your main workflow to Cursor plus Grok 4.5 means handing your data flywheel to SpaceXAI - the next model will be stronger and you'll be more locked in, which is the whole point of the Cursor acquisition. For small teams, use it, but don't feed all your engineering data into this one chain.

For everyone else, this launch is mostly "another entrant in the AI coding arms race." An extra model in Cursor is a net positive for subscribers; but if you don't write code, Grok 4.5's coding focus has little to do with your daily Q&A. The test is simple: do you normally work inside Cursor or a similar agent environment? If yes, it's worth a one-week trial; if no, wait for its capabilities to spill into general chat. "Fast and cheap" sounds tempting, but it comes down to one thing: can it, on your hardest coding task, make fewer mistakes and take fewer round-trips than the model you use now? If yes, the switch is worth it.


References

This article is AI-assisted and human-edited. Last updated: 2026-07-31

Related

Frontline Hotspot

AI Coding Agents in August 2026: Three Camps, Each Its Own Pole

By August 2026, AI coding agents have settled into three camps: browser turnkey online platforms (Replit/Bolt/Lovable), local IDEs deep-integrated with codebases (Cursor/Copilot/Trae), and autonomous terminal CLI agents (Claude Code/Codex/Cline). The camps are not tiers but different ranges; the rule is run it first, optimize later. Underneath all is the same context-execute-verify loop; the real barrier is task-decomposition skill.

Aug 5, 20266 min read
Frontline Hotspot

One prompt to final cut: JianYing Hub closes the AI video loop

According to a 9-21 report by Qbit, ByteDance's JianYing launched JianYing Hub, a one-stop AI video creation entry point on PC, whose product move is not about model parameters but about workflow, welding generation and editing into a single entry. Official positioning is a PC-side one-stop AI video creation workbench; the official page lists nine core functions (AI image and asset generation, storyboard scripting, module wiring and asset management, batch storyboard prompt generation, multi-model video generation with preview, direct hand-off to editing, AI post-editing, the JianYing Assistant Agent, and ByteDance asset import) along with a 14-step onboarding path and an official comparison table against Jimeng AI (source-side framing, not independently retested here). Two real changes stand out: generation results are not exported and re-imported but jump straight via "More Editing" into JianYing's multi-track timeline for AI extend, upscaling, frame interpolation, color grading, removal and vocal separation, an in-project closed loop replacing file exchange; and the JianYing Assistant Agent turns repetitive work into a single sentence by calling Skills for cutting voiceover, adding narration, fixing subtitles and batch production. The article's own judgment is that a workbench solves the last mile from asset to publishable cut rather than the ceiling of image quality, and that Hub is an orchestration layer rather than a generation engine, with three costs of the loop, ecosystem lock-in, tight asset-and-account coupling, and opaque pricing. Pricing, free quota, concurrency, credit rules, regional availability and duration or resolution limits are all unpublished and are stated as following the official app, with no invented numbers, and the launch timing is only a second-hand report.

Sep 22, 20267 min read
Frontline Hotspot

From 2.8s to 2.3s: can Qwen3.8 steal the interpreter's job?

In September 2026 Alibaba's Qwen team released Qwen3.8-LiveTranslate, a real-time simultaneous interpretation model opened through the Qwen AI platform and Alibaba Cloud Bailian as a WebSocket streaming API that can be embedded in meeting systems, live streams and support desks. Headline figures: average lag (LAAL) cut from 2.8 to 2.3 seconds; recognition input in 60 languages and speech output in 29; three capabilities, real-time speaker diarization plus voice cloning, source and translation emitted in the same frame, and long-context disambiguation, with video and audio input helping resolve ambiguity. Technically it rests on an Interleave single-stream architecture that caches already-heard audio and already-emitted translation instead of reprocessing each sentence, plus a Hybrid MoE Thinker-Talker pair, where the Thinker arranges video, audio, source and translation into one causal sequence and the Talker fuses translation with source audio into speech that keeps the original speaker's timbre. The article keeps its figures honest: 2.3 seconds is average lag rather than end-to-end first-packet latency, 60 and 29 are different units, the vendor comparison table is not independently retested, an unpublished metric is not the same as a bad one, pricing, rate limits, concurrency and regional availability are not invented, and the model is an API service rather than open source.

Sep 21, 20267 min read