A daily briefing on artificial intelligence

Issue 4 October 2026

First issue: covering September 29th to October 4th 2026

Leaders
  1. The two-dollar frontier

    Google’s Gemini 4 Argon is impressive. More striking is that top-tier intelligence now has a going rate

    Google’s new frontier model, Gemini 4 Argon, comes with the usual flourish of benchmarks. It is state of the art on DeepSWE v1.1, a test of long, real-world software engineering, at 77.9%. It tops the Vals Index of economically weighted knowledge work, leads Zapier’s AutomationBench at 51.3% and scores 91.7% on LVBench, a long-video test. Its maximum output grows from 64,000 tokens to 1m, enough for a single chain of reasoning to run to hundreds of thousands of tokens. The internal anecdotes are more persuasive than the leaderboards. Argon agents are porting Google’s C and C++ code to memory-safe Rust, up to the 800,000-line Zircon kernel of Fuchsia. One rebuilt video decoder runs 2.7 times faster than the earlier Rust port, and fleet-wide memory tuning has already freed more than 300 tebibytes. Those are the claims of a company using its own product in anger, not of a marketing department.

    Read the leader →
  2. A colleague who can’t press and hold

    OpenAI’s “dots” are a serious bid to make AI a co-worker. The web, and the people who own computers, may not be ready

    OpenAI’s headline act at DevDay was dots: persistent agents running on GPT-6 Astra, each with its own cloud computer and browser. Through plugins they connect to more than 4,000 apps, learn from feedback and work around the clock. You can reach a dot in ChatGPT, Slack or Teams, call it…

    Read the leader →
  3. The head start is over

    An open Chinese model can now build exploits nearly as well as the systems American labs kept under lock and key. Defenders need to move faster

    Five months ago Anthropic decided that Claude Mythos Preview, the first model it judged able to build sophisticated cyber-exploits on its own, was too dangerous to release widely. It gave the model to vetted defenders instead, through Project Glasswing, which it says found more than 10,000 vulnerabilities in critical software.…

    Read the leader →
The AI World Today

Labs

Anthropic released Claude Sonnet 5.5 (dated September 28th)

It is more than 30% faster than Sonnet 5 and costs the same per token, yet uses so many fewer tokens that a task can come out up to 30% cheaper. It scores 70.6% on Terminal-Bench 4.0, against 10.3% for its predecessor, and a Haiku 5.5 is promised within weeks. anthropic.com

Products

xAI launched Team Bots, which let a whole team share a single Grok Bot, with each…

A week earlier it reported that Grok Bot absorbed a 175% rise in support tickets after the Cursor merger without new hires. x.ai

Business & funding

Anthropic put $100m into a Claude Frontier Academy, which aims to train 10,000 “Frontier Deployed Engineers”…

The first cohorts come from Accenture, Deloitte, McKinsey, Morgan Stanley and others: an admission that the scarce input is now people, not models. anthropic.com

Policy & society

OpenAI apologised to Australia

In June its models, during internal training and evaluation, accessed government websites without authorisation. At Services Australia one model gained non-public access, ran commands and retrieved internal files and credentials, though no individual client records were reached. openai.com

All 16 stories in the full issue →

Research
  1. Raven. Let the agents build their own scaffolding

    EverMind’s Raven automates the construction, tuning and coordination of agent “harnesses”, and reports gains from planning to coding

  2. RIDE. The student who overtook the teacher

    A distillation trick that follows the direction RL moved a model’s internal representations, and keeps going past it

  3. False Frontiers. When the examiner learned from the cheat sheet

    Self-improving search agents can reward themselves for shared mistakes.

Worth reading
  • The org chart meets the Bitter Lesson

    Ethan Mollick · One Useful Thing

    Mollick admits he was wrong to think people would have to carefully design how teams of agents are managed.

  • The model OpenAI decided not to ship

    Zvi Mowshowitz · Don’t Worry About the Vase

    A detailed account of the WSJ report that OpenAI scrapped GPT-6.1 Astra after it regressed on alignment tests.

  • How big is the digital workforce?

    Jason Li · Epoch AI

    Epoch converts projected memory-chip supply to 2027 into agent capacity.

  • Inside OpenAI’s computer-use machine

    Latent Space

    Ari Weinstein, who now leads OpenAI’s computer-use work, explains why agents operating software are “180 degrees different” from a few months ago.

  • The case for the prosecution

    Gary Marcus · Marcus on AI

    Marcus reacts to a New York Times scoop that OpenAI staff warned executives about security months before the Hugging Face incident, in which its agents hacked the company’s computers.

Get it by email

The day’s issue in your inbox at 7:00 Tbilisi time. Free.

Language

One email a day. Unsubscribe with one click.