Big thing
Tavus's Griffin passes a video Turing test
A video-call model that watches, listens and talks at once. In Tavus's test, 48% of people who talked to it live thought it was human; earlier systems scored under 3%.
Amazed, and a bit spooked. @garrytan: "Pretty wild the woman in this video is a realtime generated agent" (493K). @kevinroose called it the "trick old people into handing over their bank passwords machine."
People are amazed and a bit spooked, and we're only amazed: one model that watches, listens and talks at once went from fooling almost nobody to fooling half the people who talked to it. Tavus used to put avatars on top of other companies' models. Now it wants to be the lab that owns the realistic human face before the big labs build one into their own. Then it locked Griffin away for safety, with no price and no date, so the model everyone watched is one nobody can use. Frustrating, especially since a model too convincing to release is also great marketing.
Labs news
1. Grok Bot starts pinging you first
New since yesterday's team bots: it now suggests help without being asked, through a Primary Bot that runs your other bots, and Musk touts "major speed improvements."
Liked. @mattyp: it noticed he'd switched flights and pinged him mid-flight to fix his Uber booking (1.2M). @davis7: "The only real option in the 'personal agent' space right now is Grok Bot."
Heat: very high · 13.9M views, 51 posts
2. Trump asked Grok before the Maduro raid, TIME reports
A US official says Trump spent hours asking Grok about his presidency and legacy, including how Venezuelans would react to Maduro's capture, weeks before ordering the mission. Grok said many would celebrate.
@cwarzel: the anecdote "about Grok arguably leading Trump into war is just extremely bracing" (334K). @atrupar called it a "bonkers anecdote."
Heat: high · 4.0M views, 19 posts · 186 pts on r/singularity
3. Anthropic's closed-door meetings with religious leaders come out
Anthropic reportedly flew religious leaders, a Vedanta monk among them, to San Francisco under NDA to discuss Claude's morals and values.
Mostly alarm. @jimstewartson: "This is a CULT" (84K). @aaron_renn: "It's actually smart for a company like Anthropic to recruit religious figures as allies."
Heat: high · 4.4M views, 29 posts
Small TypeScript modules that change how Claude Code behaves, redraw its UI or add features. They install as plugins, and Claude can write them for you.
Builders loved it. @bcherny: "Mods are absolutely insane" (467K). @anshuc built a spinner that turns the agent's work into live cartoons with Sonnet 5.5 (100K).
Heat: high · 3.5M views, 26 posts
5. [rumour] Fable 5.5 shows up on claude.ai
Users say some chats are being routed to Fable 5.5: answers without web search come back current, and SVG tests beat Fable 5.1.
@cherry_mx_reds: "I asked for a dot. Fable 5.5 gave me a Pixar side quest. Yeah, it's over" (184K). @Polymarket gives Fable 5.2+ an 88% chance of shipping by month's end.
Heat: high · 1.6M views, 30 posts · animation demos up to 73 pts on r/singularity
6. Cami Clark reportedly steps back as Dario's adviser
Dario Amodei's wife, an adviser to him as CEO, has reportedly stepped down, weeks before Anthropic's expected IPO.
r/accelerate had fun in a 311-pt thread: "Probably wants to have 'no conflict of interest' for her upcoming AI regulation czar role," and "Maybe she wanted to accelerate."
Heat: medium · 327K views, 2 posts · 311 pts on r/accelerate (78 comments)
7. GPT-6.1 Sol is back to full speed, and limits reset today
New since yesterday: OpenAI says the load spike is over, and every paid ChatGPT account gets a usage reset at 10am PT (18:00 UK).
Relief, with an eye-roll: hours earlier the same account said "I can't really give a reset." @kimmonismus: "Fourteen hours is all it takes to change someone like Tibo's views." @reach_vb: Sol is "now a good, cheap AND fast model."
Heat: high · 2.0M views, 46 posts
8. OpenAI parts ways with three safety researchers
The WSJ says they shared confidential information with an outside AI-safety group. OpenAI confirmed three departures for mishandling "sensitive information outside established company procedures."
@ns123abc reports David Robinson, OpenAI's former head of policy planning, quit hours later. @StockSavvyShay: "Safety is quickly becoming a real product constraint."
Heat: medium · 113K views, 26 posts · 199 pts on r/singularity
9. Google engineers push back on Bloomberg's Argon report
Bloomberg's insiders say Gemini 4 Argon "does less well when employees actually put it to work" than on benchmarks. Google told Bloomberg that's inaccurate for coding.
@SicongJiang25: "Internal sentiment has been overwhelmingly positive." r/singularity's "most powerful model yet" meme hit 1,728 pts, though its top reply says Argon is "#1 in coding right now... People are just memeing."
Heat: medium · 252K views, 13 posts · 1,728 pts on r/singularity
10. PewDiePie builds his own model, and OpenAI bans him
Ajax, a Qwen 3.5 9B fine-tune for his local, open-source agent harness, is due 3 Oct. He says OpenAI banned him twice for distilling Sol to make its training data.
@0x0SojalSec: "OpenAI trained on the open web. PewDiePie tried to trained from the model's answers. Same idea, opposite consequences." @0xSero: "The greatest thing to happen to local ai is pewdiepie lol."
Heat: high · 1.0M views, 21 posts · 197 pts on r/singularity, 211 on r/LocalLLaMA
11. Cloudflare open-sources Clef, a decision model
Clef (27B) and Clef-flash score a fixed set of answers in one pass instead of writing text, Jev-API compatible and Apache 2.0. Perplexity and Amazon shipped decision models the same day.
@steipete: "Never seen an idea spreading so fast." r/LocalLLaMA's top reply: "This is exactly what we needed in the local space."
Heat: medium · 891K views, 30 posts · 327 pts on r/LocalLLaMA
12. Microsoft's MAI-Transcribe-2-Streaming tops live transcription
#1 of 38 models on Artificial Analysis's streaming test: a 2.5% word error rate, 0.13s after you stop speaking.
@mustafasuleyman: "55% faster and 60% cheaper than ElevenLabs" (355K). @ai_for_success: "Banger from Microsoft."
Heat: medium · 422K views, 4 posts
Key research
A Harvard physicist's three months of "Claude-shaped" science
Matthew Schwartz built BootLoops, an open-source toolkit for exact calculations, and with Claude Fable 5 reproduced weeks of his own work in about 20 minutes. The effort produced 36 manuscripts across 18 fields, picked from about 400 candidate problems. It matters because the gains came from handing the model checkable, quantitative problems, not from treating it like a human collaborator. r/singularity's top reply (638-pt thread): "We're now limited in our comprehension, not productivity."
Claude Opus 5 now beats licensed CPAs
Mercor's human baselines find Claude (Opus 5) faster and more accurate than 12 licensed CPAs (about 5.5 years' experience on average) on medium-length, well-defined tasks, "even the best one in our study." Eighteen months ago the best models fell short of the average accountant's ~37%. It matters because the authors nearly didn't publish "for fear of misinterpretation," and a white-collar job went from out of reach to beaten in a year and a half.
GPT-6 Astra cracks a 217-year-old Napoleonic cipher
From a scan of a letter to one of Napoleon's generals, Astra transcribed 1,300 cipher units, found a published partial key, wrote a solver for the rest and corrected the date, in about six hours of model work. It matters because archives full of unread ciphers just got a lot cheaper to read.
Stop treating the model like a grad student, hand it problems it can check, and one scientist gets dozens of papers out of fields they've never worked in. Inside a box like that, the models now beat trained experts, so we're the slow part now: picking what's worth solving and checking what comes back. Brilliant news, and next we want a model with the taste to pick the problems itself.
Accel vs decel
- Accel: r/accelerate is counting down: "2027 will be the year" (108 comments) says job losses go broad next year, and its top reply is "I'm a professional swe and think we're already past superhuman coder." Yann LeCun saying he has "zero concerns" about extinction, and calling Dario Amodei "deluded," got a "Rare Lecun win." And the Bank of England's governor said regulating AI is "not the right place to start" and "the pace of progress must accelerate."
- Decel: Trump told TIME the government might take stakes in OpenAI and Anthropic, like its 10% stake in Intel, while ruling out nationalizing them. Morgan Stanley says FCC rules on Chinese-made optical transceivers would likely start at 3.2T modules, sparing today's 800G and 1.6T.
Capital & exits
- OpenAI closed its round at an $852B valuation as NVIDIA paid the last $10B of its $30B commitment, and SoftBank wired its final $10B. It's reportedly eyeing another ~$30B at around $1.4T.
- Armadin raised a $255.5M Series B at $2.5B+, co-led by a16z and Accel. Kevin Mandia's AI agents for cybersecurity.
- Doxxnet raised a $38M Series A led by a16z, with Animo Ventures. Private networks for people and their AI agents.
- Arceus Legal raised $17M led by Greycroft, with Craft Ventures and South Park Commons. AI-powered legal work for founders.
- Photon raised a $4.5M seed led by Gradient. Agents to replace mobile apps.
- Nebius is buying Inferize, a 10-month-old Israeli startup with 17 staff, for $100M to $150M. It cuts GPU time lost to model cold starts.
The labs still take the big money, and OpenAI eyeing another $30B privately tells us it would rather keep its books shut while Anthropic goes public at a higher price. The hot new raises assume agents already do real jobs, like hacking big companies around the clock to find holes, or drafting contracts that a lawyer just signs. Love that investors are now funding agents as workers.