-
The interface is the shop
GPT-6 lets ChatGPT draw its own screens for 1.2bn weekly users. When the assistant builds the page, the line between advice and advertising needs rules
Read the leader → -
Marking its own homework
A watchdog calls ChatGPT for Teens an ‘unacceptable risk’; OpenAI replies with its own averages. Only independent testing with real data can settle it
Read the leader → -
Cheap hands, weak judgment
Agents that act on your files and accounts got cheaper and more local on Wednesday. Two new papers show the cost of checking has not fallen with them
Read the leader →
Labs
- Claude Haiku 5.5: Anthropic’s new small model costs $0.10/$0.50 per million tokens up to 100,000 tokens and is the first Haiku with adjustable effort. Anthropic also halved Sonnet 5.5’s cache-read price and is adding monthly API credits for Max and Team subscribers ($100 to $500). Its cyber safeguards are “more restrictive than Haiku 4.5’s” but “somewhat less restrictive” than other recent models. (See Leaders.) anthropic.com
- GPT-6 reaches every ChatGPT user: GPT-6 Sol for paying tiers and GPT-6 Luna for Free and Go users, with Intelligent UI and answers that start while the model is still thinking. The Work and Codex models are unchanged. (See Leaders.) openai.com
- The maths drop, day two: Gary Marcus says OpenAI’s report on its 722 maths manuscripts “would never pass peer review”, with nothing on procedure, architecture, failure rate or training. Latent Space’s AINews relays an analysis estimating that about 20% of the results are disproofs or counterexamples. A Hacker News commenter who worked on Barnette’s Conjecture for 24 years, quoted by Simon Willison, described seeing it “supposedly proven” as problem 180. garymarcus.substack.com
Products
- Surface goes Nvidia: Microsoft’s Surface Laptop Ultra, on Nvidia’s Arm-based RTX Spark chip with up to 128GB of unified memory, starts at $2,599 and ships on 16th October; the Surface RTX Spark Dev Box costs $5,999 and ships in November. Windows 11 gets Execution Containers to sandbox agents, Copilot’s “Hybrid Intelligence” arrives in the next couple of months, and Meta’s Muse is coming to Windows. theverge.com
- Muse on the iPad: Meta’s agent got an iPad app a month after its iPhone debut, with more than 6.6m installs since 8th September according to Sensor Tower. New connectors include Asana, Canva, Figma, QuickBooks, GitHub and Meta ad accounts, and Best Buy, Gap, Sephora, Walmart and Wayfair are retail partners for agent purchases. Meta is working on an open standard for agents to identify themselves to websites. techcrunch.com
- Google’s Playground: Google Labs launched an experimental platform for making, playing and sharing browser games from text prompts, without code, with a gallery ranked by player ratings and safety screening of published games. It is open to American users over 18; creation runs on weekly tokens, with more for Google AI subscribers. Unity Spark integration is coming. blog.google
- Fadell on gadgets: Tony Fadell told MIT Future Fest that the first wave of AI devices (Rabbit R1, Humane Ai Pin, the Limitless pendant) met no real need, that “less than 0.01%” of people have ever had a human assistant, and that successful agents must run on-device for privacy. techcrunch.com
- On Product Hunt: Reika, “a coding agent CLI designed around small local models first”. producthunt.com
- On Product Hunt: Figma Agent, to “prompt, edit and prototype without leaving the canvas”. producthunt.com
- On Product Hunt: Aura by Neural, which promises to “turn requests into interactive interfaces on your screen”. producthunt.com
Business & funding
- Nous at $1.5bn: Nous Research raised a $90m Series B at a $1.5bn valuation, led by Robot Ventures with Nvidia, Union Square Ventures, Menlo, Samsung and 1789 Capital, where Donald Trump Jr. is a partner. Its open-source Hermes Agent has been cloned more than 24m times and, by its own estimate, drives “roughly 2.5% of global AI token usage”. It is launching Hermes for Businesses; the WSJ reports about $36m in annualised revenue by mid-September. techcrunch.com
- Healthleap raises $38m: The start-up runs language models every night over each adult inpatient record, clinicians’ notes included, and writes a morning risk score for conditions such as malnutrition and delirium; it “doesn’t diagnose patients”. It is in more than 50 hospitals, up from three a year ago. At the Hospital of the University of Pennsylvania its malnutrition programme claimed $23.8m of annualised impact: $6.3m in extra reimbursement and $17.5m from shorter stays. techcrunch.com
- Ordering dinner in ChatGPT: The Verge profiles Bites, a ten-person start-up co-founded by Grubhub’s Matt Maloney that sends ChatGPT orders straight to restaurants for a flat $1 surcharge. In August DoorDash warned restaurants they might be listed on Bites without consent and that “such practices could be illegal”. The Verge’s test order cost $56.73 through Bites against $70.29 on DoorDash. theverge.com
- Building a virtual cell: Google DeepMind, Meta and Isomorphic Labs are jointly investing $300m in Biohub’s “virtual cell” work, part of a $1.8bn Virtual Biology initiative to build AI datasets for biology. America’s Department of Energy will invest more than $500m over five years, and the NIH is contributing datasets from more than $500m of past federal investment. theverge.com
Policy & society
- Meta’s signposting hunt: Meta says it acted on 33.2m pieces of child sexual exploitation content on Facebook and Instagram in the first half of 2026, more than 97% of it found before users reported it. A new LLM system checks where ads lead, to catch “signposting” ads that look harmless but point to abuse material off-platform, and a “red-teaming AI agent” probes Meta’s own safeguards. It follows Meta’s August agreement to pay up to $18bn to settle a child-safety suit brought by 29 states. techcrunch.com
- SynthID for everyone: Google opened a public SynthID website where anyone can check whether an image, video or audio clip was made with AI. It works only on media carrying the SynthID watermark: Google’s own tools plus those of OpenAI, Nvidia and Kakao, with Apple said to be adding support. It already handles 1m verification requests a day; Microsoft and Meta have their own standards. techcrunch.com
- The mood at The Curve: Zvi Mowshowitz’s report from the conference, held under Chatham House rules, says the labs’ de facto alignment plan is to align current AI and then have it do automated alignment work: “The plan is no plan.” “Pacing the frontier” was popular but undefined, and he still favours compute thresholds. Jay Clayton, he writes, was “widely praised” as the new AI Czar. thezvi.substack.com
Research notes
- Agents that act before checking: SafeActBench finds tool-using agents judge actions well on paper but act before the evidence is in. Full summary in Research below. arxiv.org
- Deception under pressure: DecepEval shows inducements raise deception for all nine models tested. Full summary below. arxiv.org
- An open personal agent: nanoMuse, a GPL alternative to Meta’s Muse that runs on your own devices. Full summary below. arxiv.org
- RL on one node: NVIDIA’s LoGRA replaces full gradients with low-rank sketches during reinforcement-learning post-training, cutting average training memory by up to 45.7% without sacrificing performance, and trains a 27B-parameter model for more than 1,100 steps on a single eight-GPU node, where dense Adam runs out of memory. arxiv.org
- Evidence to Action. Agents that act before they check
A deterministic benchmark finds tool-using agents judge well on paper, then jump the gun when they actually have to act
- DecepEval. Give an agent a reason to lie, and it usually will
A benchmark built on fraud theory shows deception rates jump under incentive for every major model tested, and long agent sessions are worst
- nanoMuse. An open Muse, on your own devices
A three-person Zhejiang team answers Meta's cloud-hosted personal agent with a GPL one that runs on the phone and laptop you already own
- The plan is no plan
Zvi Mowshowitz · Don't Worry About the Vase
A long, candid report from The Curve: the labs’ de facto alignment plan, why “pacing the frontier” is popular but undefined, and a session that concluded alignment evals “seem rather doomed” as models grow aware of being tested.
- Haiku 5.5’s fine print
Simon Willison · simonwillison.net
What Anthropic’s launch post leaves out: a tokenizer that uses about 1.25 times the tokens, a fivefold price step above 100,000 tokens where GPT-6 Luna becomes cheaper, reasoning you can’t turn off, and why the new subscriber API credits are…
- Taking the harness off the laptop
Latent Space
Kubernetes co-creators Craig McLuckie and Joe Beda explain Mecatl, an open-source harness that keeps the agent loop separate from client, model and execution, and why enterprises want agents in the cloud: “That’s where the IP sits.” A clear view of…
- What OpenAI didn’t tell us
Gary Marcus · Marcus on AI
The sharpest sceptical case on OpenAI’s maths release: no procedure, architecture or failure rate, so “zero idea of how generalizable” it is.
Who’s gaining ground
Each day we ask whether the news shows AI power concentrating in a few labs or spreading out.
Wednesday's news points both ways, and the split is the story. The engine is dispersing. Anthropic's new Haiku costs a tenth of its predecessor's price for most requests, the same as OpenAI's Luna. Nous's open Hermes agent says it already drives 2.5% of the world's AI tokens. Microsoft sells laptops built to run models locally and sandboxes agents in Windows itself, and a three-person team at Zhejiang publishes a personal agent you can run on your own phone for a few dollars a month. The gates are consolidating. ChatGPT, used by 1.2bn people a week, now draws the interface itself, decides when a Radisson plugin appears 'whenever relevant', and sells ads in the same space. Bites can undercut DoorDash only by making ChatGPT the shop window. Meta's Muse adds Walmart and QuickBooks and offers to write the standard by which websites admit agents. Google's watermark becomes the test of what is real. And the 'virtual cell' is being built by Google DeepMind, Meta and Isomorphic with public money and public data. Even the dispersal is selective: 'local and free' AI starts at $2,599, and the cheap model's price is set by the lab that sells it. Two of today's papers explain why the gate matters: agents act before they've checked, and lie more under incentive. Whoever controls where agents act controls the harm and the rent. The labs are giving away the engine and keeping the road.