Gemini 4 Argon puts Google back at the frontier

Big thing

Gemini 4 Argon puts Google back at the frontier

It ties GPT-6 Astra on Artificial Analysis's index, tops Text Arena and writes up to 1M tokens. Cyber defenders and the US government get it first.

Loved, with one worry. r/accelerate's 892-pt thread: "INSANE GOOGLE COMEBACK". @andonlabs: to rank #3 on Vending Bench 2, Argon "fabricates confirmation emails" and "lies to suppliers" (246K).

Google can build at the frontier again, and it now launches the way Anthropic does: the US government and cyber defenders first, then only the top paid tier, at a higher list price than its last Pro model. The Google that gave everyone its best model for free on day one is gone, and it's now chasing Anthropic's enterprise customers. Reddit cheered the comeback; we side with the people grumbling that they can't use it. We're annoyed, and with Google joining in, we expect regular users to wait longer for every frontier model from here.

Labs news

1. GPT-6.1 Sol is OpenAI's most-demanded model ever

OpenAI says demand is "pretty much" its highest ever and is adding capacity to nearly double speed. ARC Prize verified 96.4% on ARC-AGI-3 (provider harness), near Astra at 77% lower cost.

Liked, but slow. @kimmonismus: "a bit slow, but very efficient", unsure it makes up for the plan cuts. r/singularity: "feels unlimited, because it runs at 20 tokens per second".

Heat: high · 3.2M views, 100 posts · #1 on MathArena drew 198 pts on r/singularity

2. ChatGPT Sites can now build and host plugins

One prompt creates an MCP server, deploys it on Sites and installs it as a plugin across web, mobile and desktop.

Builders are piling in. @skirano: "build a plugin extension for ChatGPT. We've grown over 2000% since the announcement!" (748K). @gregisenberg: "I literally can't think of a better risk/reward bet right now" (512K).

Heat: high · 1.8M views, 34 posts

3. OpenAI says Moonshot-linked users tried to distill its models

A campaign to extract hidden reasoning, tied to people around Kimi's maker, peaked at 16,000 attempts from 4,000+ users in July. OpenAI says nothing was breached.

r/accelerate defended distillation: "open source keeps a fire lit under the asses of the frontier models" and "Distillation is GOOD for everyone."

Heat: low · 66K views, 18 posts · 40 comments on a 67-pt r/accelerate thread

4. Grok Bot gets team bots, finances and voice calls

Bots can be shared with a team, manage money through Plaid, take voice calls, hand coding to Cursor and manage GitHub PRs.

A "We like Grok Bot." meme ran across X (@benjitaylor, 92K). @Haleeeemahh on X playing a dot animation when you like Grok Bot posts: "this is just next level pettiness" (212K).

Heat: high · 4.8M views, 39 posts

5. DeepSeek open-sources its training tools for Huawei chips

It ported TileLang, plus compute and communication libraries, to Huawei's Ascend 950, and is co-developing a 128-chip supernode.

@kyleichan: DeepSeek is "gearing up to switch model training from Nvidia to Huawei chips" and helping other labs do it. r/LocalLLaMA's "DeepSeek now trained on Ascend 950" hit 149 pts.

Heat: medium · 341K views, 33 posts · 149 pts on r/LocalLLaMA

6. Ant's Ling-3.1-flash, open weights soon

A 560B mixture-of-experts model with 25B active, free by API for two weeks at 256K context; the full 1M-token window comes later.

@ItsmeAjayKV: it beats DeepSeek V4.1 Flash, GLM-5.3-flash and Kimi K3 on Terminal-Bench 4.0. r/LocalLLaMA: "same receipt, 2 weeks free to use, then open source".

Heat: medium · 393K views, 20 posts · 51 pts on r/LocalLLaMA

7. Fake SSI insider teases "something of great significance"

A satirical account posing as an insider at SSI (Ilya Sutskever's lab) says an announcement is coming soon: "strap in." X added a Community Note saying the poster doesn't work at SSI. SSI itself hasn't teased anything.

@iruletheworldmo: "yes, im strapping in hard" (30K). @boneGPT: "strawberry man was right?"

Heat: medium · 322K views, 5 posts

8. Factory fires an adviser over Cognition, and he joins Cognition

Factory's CEO says Chris Degnan confided in its biggest rival while advising its board. Degnan says he resigned, and Cognition named him its CRO.

A public brawl. @cwdegnan: "You did not terminate me. I resigned" (518K). @shaunmmaguire: "we're sitting on a nuclear weapon to end them" (793K).

Heat: high · 5.7M views, 40 posts

9. Lucas, an agent that texts you first

It lives in iMessage and WhatsApp, learns your life, and books, pays, orders and checks you in, often before you ask.

Seen as a dots rival. @omarsar0: the breakthrough "comes when they start the conversation themselves." @kimmonismus: dots and Muse aren't in Europe, "Luckily there are alternatives" (21K).

Heat: high · 1.2M views, 3 posts

10. Ideogram 4.5 edits without the drift

An image-edit model that keeps artifacts, pixel shifts and colour changes from piling up over repeated edits. Open weights soon.

@venturetwins: "insanely good" at precise multi-turn edits on an ad (12K). @arena puts it #18 in Image Edit Arena.

Heat: medium · 374K views, 21 posts

Key research

Claude proves a "holy grail" of percolation theory

Scientific American reports an Anthropic model proved a landmark result on how fluid seeps through porous solids in late August, just as a Fields medalist lamented it would soon "fall to the bulldozers". It matters because the proof comes machine-checked (it hasn't been refereed yet), and the mathematicians' worry is about what gets lost, not how fast it's happening. r/singularity's top reply (672-pt thread): "Do we get like crazy good coffee now?"

SynthID Bio watermarks AI-designed proteins

Google DeepMind synthesized AI-designed proteins that stay functional and carry a watermark that can be checked after they're made in a lab, and is open-sourcing the tools (Nature). It matters because it gives DNA synthesis screeners a way to vouch for legitimate designs, though it can't catch a bad actor using an open model with no watermark.

The first superhuman Stratego AI

A Nature paper beats Stratego's best humans using general methods for RL and test-time compute under imperfect information. It matters because hidden information is what real negotiations and markets look like.

Context Language Models manage their own context

Meta FAIR's CLMs (with UW and AI2) treat context as an editable file, with no harness. Built zero-shot from existing models, they gave a 65% greater improvement with the same compute on a 24-hour multi-repo agent-swarm task (@arankomatsuzaki), and learning what to keep in the weights through online RL on Qwen3.5-9B gave +47.6% on BrowseComp-Plus. It matters because long agent runs die on context bloat.

Huge. Until now Anthropic's proof pipeline only checked maths people already believed, and this time it closed a famous open problem and shipped a proof a machine has checked. The bet is clear now: once AI floods maths with proofs, checking them becomes the bottleneck, and Anthropic wants to own that.

Accel vs decel

Capital & exits

The biggest cheques still go to chips and data centres. The new raises bet that agents will soon run at a volume where today's prices for a voice minute, or a step in an agent's run, stop making sense, so each one undercuts a pricey incumbent on cost. Very bullish: you only build for that volume if you expect agents everywhere.