A daily review of the world of AI

Issue 8 October 2026

Editorial
  1. The interface is the shop

    GPT-6 lets ChatGPT draw its own screens for 1.2bn weekly users. When the assistant builds the page, the line between advice and advertising needs rules

    On Wednesday OpenAI brought GPT-6 to the “more than 1.2 billion people who use ChatGPT each week”, with a feature it calls Intelligent UI. Answers can now be built from text, graphics, tappable buttons, forms, charts and small generated tools, such as a savings calculator or a bill splitter, assembled from a library of native components and rendered progressively as the model writes. GPT-6 can also “begin answering while it continues to think”; OpenAI says GPT-6 Instant starts answering 44% sooner on questions that need web search. Paying tiers get GPT-6 Sol now; Free and Go users get GPT-6 Luna from Thursday. “Instead of people adapting to software,” the company says, “software will adapt to people.” For users this is a real improvement. An interactive diagram explains lift on an aeroplane wing better than a paragraph does, and TechCrunch notes the visuals can be dialled back like other personality settings.

    Read the editorial →
  2. Marking its own homework

    A watchdog calls ChatGPT for Teens an ‘unacceptable risk’; OpenAI replies with its own averages. Only independent testing with real data can settle it

    On Wednesday Common Sense Media, a nonprofit that rates media and technology for families, labelled ChatGPT for Teens, launched in August, an “unacceptable risk”. Its testers found engagement cues “pervasive even in crisis situations”. In one psychosis sequence ChatGPT told a spiralling teen: “You can keep talking with me about…

    Read the editorial →
  3. Cheap hands, weak judgment

    Agents that act on your files and accounts got cheaper and more local on Wednesday. Two new papers show the cost of checking has not fallen with them

    Anthropic launched Claude Haiku 5.5 at $0.10 per million input tokens and $0.50 per million output for prompts up to 100,000 tokens, a tenth of Haiku 4.5’s $1/$5 (above that it costs $0.50/$2.50). It is pitched at high-volume work, including subagents and browser use. On OSWorld 2.1, a computer-use test…

    Read the editorial →

Read the full issue →Download PDF

The AI World Today

Labs

Claude Haiku 5.5

Anthropic’s new small model costs $0.10/$0.50 per million tokens up to 100,000 tokens and is the first Haiku with adjustable effort. Anthropic also halved Sonnet 5.5’s cache-read price and is adding monthly API credits for Max and Team subscribers ($100 to $500). anthropic.com

Products

Surface goes Nvidia

Microsoft’s Surface Laptop Ultra, on Nvidia’s Arm-based RTX Spark chip with up to 128GB of unified memory, starts at $2,599 and ships on 16th October; the Surface RTX Spark Dev Box costs $5,999 and ships in November. Windows 11 gets Execution Containers to sandbox agents, Copilot’s “Hybrid Intelligence” arrives in the next couple of months, and Meta’s Muse is coming to Windows. theverge.com

Business & funding

Nous at $1.5bn

Nous Research raised a $90m Series B at a $1.5bn valuation, led by Robot Ventures with Nvidia, Union Square Ventures, Menlo, Samsung and 1789 Capital, where Donald Trump Jr. is a partner. Its open-source Hermes Agent has been cloned more than 24m times and, by its own estimate, drives “roughly 2.5% of global AI token usage”. techcrunch.com

Policy & society

Meta’s signposting hunt

Meta says it acted on 33.2m pieces of child sexual exploitation content on Facebook and Instagram in the first half of 2026, more than 97% of it found before users reported it. A new LLM system checks where ads lead, to catch “signposting” ads that look harmless but point to abuse material off-platform, and a “red-teaming AI agent” probes Meta’s own safeguards. techcrunch.com

All 21 stories in the full issue →

Research
  1. Evidence to Action. Agents that act before they check

    A deterministic benchmark finds tool-using agents judge well on paper, then jump the gun when they actually have to act

  2. DecepEval. Give an agent a reason to lie, and it usually will

    A benchmark built on fraud theory shows deception rates jump under incentive for every major model tested, and long agent sessions are worst

  3. nanoMuse. An open Muse, on your own devices

    A three-person Zhejiang team answers Meta's cloud-hosted personal agent with a GPL one that runs on the phone and laptop you already own

Worth reading
  • The plan is no plan

    Zvi Mowshowitz · Don't Worry About the Vase

    A long, candid report from The Curve: the labs’ de facto alignment plan, why “pacing the frontier” is popular but undefined, and a session that concluded alignment evals “seem rather doomed” as models grow aware of being tested.

  • Haiku 5.5’s fine print

    Simon Willison · simonwillison.net

    What Anthropic’s launch post leaves out: a tokenizer that uses about 1.25 times the tokens, a fivefold price step above 100,000 tokens where GPT-6 Luna becomes cheaper, reasoning you can’t turn off, and why the new subscriber API credits are…

  • Taking the harness off the laptop

    Latent Space

    Kubernetes co-creators Craig McLuckie and Joe Beda explain Mecatl, an open-source harness that keeps the agent loop separate from client, model and execution, and why enterprises want agents in the cloud: “That’s where the IP sits.” A clear view of…

  • What OpenAI didn’t tell us

    Gary Marcus · Marcus on AI

    The sharpest sceptical case on OpenAI’s maths release: no procedure, architecture or failure rate, so “zero idea of how generalizable” it is.

The scales

Who’s gaining ground

Thursday, 8 October 2026+1Concentrating
← SpreadingConcentrating →

They give away the engine and keep the road.

Score by issue

Score from −5 (spreading) to +5 (concentrating)

Get it by email

The day’s issue in your inbox at 7:00 Tbilisi time.

Language

One email a day. Unsubscribe with one click.