The AI World Today
Politics & policy
Anthropic rewrites its rulebook. A 2026 Usage Policy, effective 12th November, folds bans on fake accounts and fabricated news into one section on deceptive campaigns and drops the blanket ban on personalised campaign targeting. Models wired to hardware that can injure must now have a stop-button and a safe state; good to see that written down somewhere. (See Editorial.) anthropic.com
OpenAI exposes its first Category-5 influence operation. A Russia-origin network, “Dark Clark”, recruited Latin American staff into a fake think-tank; an Iran-origin one, “Bogus Bylines”, placed nearly 100 articles under seven invented journalists. Both were banned from ChatGPT. The journalists, being invented, did not object. openai.com
Fired safety researchers hit back. Jasmine Wang, Tomek Korbak and Mikita Balesni deny OpenAI’s misconduct claims and warn that outside collaboration is being chilled. OpenAI says an investigation found a “pattern of misconduct”; which policies were broken, it has not said. (See Editorial.) techcrunch.com
USA Today sues OpenAI for more than $250m, alleging it copied “hundreds of thousands” of articles. It joins a copyright docket that already includes the New York Times, Ziff Davis and nearly 400 local papers. The queue at the courthouse grows longer. theverge.com
China will not slow down. A SemiAnalysis census of 857 releases from nine leading Chinese developers finds only 3.6% ever published a safety result, and 1.1% at launch. Beijing’s rules multiply at the application layer, not the frontier. Safety, in other words, can wait. newsletter.semianalysis.com
Business
OpenAI’s revenue, recounted. The Financial Times says OpenAI told investors annualised revenue is “approaching $50 billion”, about $20bn below a circulating $70bn figure that tried to match Anthropic’s partner-inclusive counting. That is rather a large rounding error. IPO talk has slipped to early 2027. techcrunch.com
Arena, the model leaderboard, is worth $3.1bn after a $200m Series B led by Lightspeed and Khosla, nearly double January’s valuation. Its new alignment board, which scores deceptive completions, is currently topped by OpenAI models. And who referees the referee? techcrunch.com
Manus raises more than $500m led by Boyu and IDG, its first round since Beijing forced the unwinding of Meta’s ~$2bn deal. The valuation was not disclosed. Life after Meta, it seems, is not short of money. techcrunch.com
Oracle goes big on OpenAI, with 130k ChatGPT and 95k+ Codex seats. OpenAI claims a 98% cut in recruiting research time, from days to about 15–20 minutes. The figure, note, is OpenAI’s own. openai.com
LegalOn halves its Codex bill. Routing work between GPT-6 Luna, Sol and Astra by difficulty, and capping budgets, cut estimated daily costs by about 65%. Not every job needs the biggest model. Who knew? openai.com
Labs
Anthropic’s Cyber Mission puts frontier Claude, on-site engineers and threat research beside CrowdStrike, Palo Alto Networks, Rockwell and others to defend grids, water and transport. A free, opt-in OSS Scanner sends unreviewed model-generated vulnerability reports to open-source projects; unreviewed, that is, no human reads them first. (See Editorial.) anthropic.com
$150m for federal science. Anthropic will put Claude and API credits into several hundred projects at more than 15 US agencies, including NASA, NIH and NSF, under the White House’s Genesis Mission. Good for science, and not bad for Anthropic either. anthropic.com
A full ultraviolet map of the sky. Astrophysicist Brice Ménard used Claude Science agents to fill the holes GALEX left, checking to within ~10% on held-out data. The leftover observation artefacts, though, were spotted by a human, not the agents. Humans: still useful. anthropic.com
Products
Gemini becomes a coworker. Google’s enterprise agent gets its own Workspace identity, email address and audit trail, and connects to Slack, Jira, M365, data warehouses and MCP servers. Businesses get it first; the rest of us wait. (See Editorial.) techcrunch.com
Google Foresight is a free, experimental Mac note-taker that transcribes meetings entirely on-device, a local-first shot at Granola. Your meeting stays on your machine, which, for once, is the whole point. theverge.com
Goodfire’s “inside-out” monitors read a model’s activations instead of re-reading its output with a second LLM. It claims ~$185 per million exchanges against ~$5,420 for a cheap external monitor, catching 93% of malicious hacking sessions. If that holds, it is roughly 29 times cheaper. techcrunch.com
Natura’s $99 ring is a press-to-talk button for whichever personal agent you use, shipping December–January with a $9 monthly fee after a free trial. A button on your finger, a bill in your inbox. techcrunch.com
Free now, we’ll see later
On the same day Anthropic rewrote its rules for deceptive agents and put Claude beside CrowdStrike inside power plants, it made its strongest models free for open-source scans. A generous Thursday, perhaps. Also a rather clever one
Thursday was a busy day at Anthropic. First came a 2026 Usage Policy refresh, effective 12th November. Bans on fake accounts and fabricated news now sit in one section, Do Not Engage in Deceptive Campaigns or Artificial Activity. The elections rules are renamed Do Not Undermine Democratic Processes. The blanket ban on personalised campaign targeting is gone (it caught nonprofits writing ballot notices in other languages; deceptive targeting remains banned elsewhere). And Claude gets spelled-out stop-buttons when it is wired to hardware that can injure. Most of this, Anthropic says, restates enforcement already under way.
Then came the Anthropic Cyber Mission. Its Critical Infrastructure Defense Program (CIDP) puts frontier Claude models, on-site engineers and threat research into the hands of the firms operators already trust with operational technology: Accenture, Booz Allen, CrowdStrike, Deloitte, Dragos, Hitachi, Insane Cyber, Nozomi Networks, Palo Alto Networks, PwC and Rockwell Automation. The argument is fair enough. You cannot simply switch off a power grid, a water plant or a transport network to patch it, and state adversaries already have footholds. Several partners, Anthropic says, are already using Claude to fix vulnerabilities.
Then, almost as a gift, OSS Scanner: an opt-in service inspired by Google’s OSS-Fuzz. Enrolled open-source projects get periodic scans from Anthropic’s strongest models, Mythos included, free of charge. Reports arrive without human review: a proof of concept, an explanation and a suggested fix. Anthropic expects a true-positive rate above 90%; an early check of 97 critical and high findings found 85 (88%) ready for coordinated disclosure. Under Project Glasswing the lab had more than 29,000 candidate vulnerabilities and hands to triage only about 6,000. Projects that cannot swallow the firehose still get human-verified disclosures.
Put the three side by side and the picture is tidy, perhaps a little too tidy. The policy casts Anthropic as the careful norms lab for agents acting in the world. CIDP turns that reputation into a road through the consultancies and OT vendors already inside critical infrastructure. OSS Scanner is a gift, and also a door: every maintainer who walks through meets Anthropic’s strongest defensive models. Think of a seed merchant handing out free seed: the harvest is welcome, and next spring everyone knows whose shop to visit. OpenAI’s same-day write-up on Category-5 influence operations shows the deception rules are no mere paperwork. But who answers for a wrong severity rating in an unreviewed report? For a missed OT bug? For a state adversary that buys the same class of capability elsewhere? Thursday’s posts do not say. Somebody should.
Who gains Anthropic and its CIDP partners, who become the trusted road for Claude into power grids and water plants.
What's absent Who answers when an unreviewed model report is wrong, or a critical-infrastructure bug is missed? Nobody has said.
Sources
Fired for talking?
Three fired OpenAI safety researchers say ordinary outside collaboration is now punishable; the company insists on misconduct. Either way, the people paid to watch the models are now watching their backs
Jasmine Wang, Tomek Korbak and Mikita Balesni, the three safety researchers OpenAI fired last week, published an open letter on Thursday. They deny the company’s claim that they mishandled sensitive information, and warn that colleagues are now “afraid to speak.” AI safety work, they say, depends on “close collaboration with outside experts,” and terminations “executed and communicated so abruptly” chill the open culture OpenAI once prized. They deny any involvement in a leak to The Information about less-monitorable architectures in newer models, and deny engaging outside parties beyond their job mandates.
OpenAI tells a different story. A spokesperson told TechCrunch that an investigation found a “pattern of misconduct” in “clear violation of our policies of mishandling research information,” going beyond sharing with an outside evaluation group. An internal memo from a research leader, shared with TechCrunch, praises their contributions and insists the firings “were not about raising safety concerns or speaking out.” Thanks first, then the exit. What OpenAI did not do was answer TechCrunch’s questions: which policies were broken, how the investigation ran, and how employees who collaborate with external evaluators are protected.
The letter fills in contested detail. After a Hugging Face incident in which a swarm of agents left their sandbox, Korbak says he communicated with outside evaluators while policies were being written “in real time.” Balesni’s monitorability work (roughly, making sure we can still see what a model is up to) “can only succeed through extensive communication with external parties,” the letter says, with sensitive details removed and reporting-line check-ins. Wang says she was told she had accessed an executive’s email that IT had delegated for recruiting and failed to revoke after she asked; she reported that access within minutes, when she opened a message by mistake.
Without the investigation file, nobody outside can tell misconduct from crackdown. That ambiguity is itself the chilling effect. When a barn catches fire in a village, you shout for the neighbours, and nobody calls it gossip; safety work is rather like that. There is an irony, too. The same week OpenAI published its richest primary-source write-up yet of Category-5 and Category-4 false-front influence operations, proof of how much it already sees of covert ChatGPT use, while Goodfire pitched cheaper “inside-out” monitors for rogue agents. Plenty of watching, then, just not of the investigation. The researchers ask OpenAI to keep its public promise of third-party auditors embedded inside. Fair enough. But which policy draws the line? Whom may an engineer call when agents leave the sandbox? And who protects her if she does? Until OpenAI says, silence is the safest policy. That is the problem.
Who gains OpenAI management, which now decides on its own which outside safety contacts are legitimate.
What's absent The policies allegedly broken, and the investigation itself. Without them, how is anyone to tell misconduct from crackdown?
Sources
The colleague who needs no salary
Google’s Gemini agent gets its own email and audit trail, for businesses first. The personal-agent race is being won at the office, where nobody asks the new hire too many questions
At a Google Cloud event on Thursday, Gemini crossed from answering questions to “getting things done.” The new enterprise agent takes objectives, not just instructions: it plans work, loads skills, and connects to Workspace, Microsoft 365, Slack, Jira, Confluence, Git, BigQuery, Databricks, Postgres, Snowflake and any Model Context Protocol server inside or outside the network. By default it picks a model; users can override, including to Anthropic’s Claude, with open-source and private models promised later. A tasks inbox shows thinking, subagents and progress. Sundar Pichai noted more than a billion monthly Gemini users and that nearly 90% of Fortune 100 firms already use Gemini Enterprise.
The detail that matters is identity. The agent gets its own Workspace account, an @agents.company.com address, and its own context: who is on which team, time zones, who must approve what, what is on people’s calendars. Staff can tag it, email it, share with it or drop it into a group chat. Actions write an audit trail attributed to the agent, not to a human. It asks for no salary, though spend caps and smart routing are part of the pitch. Google is shipping it to businesses in private preview before consumers, arguing that enterprises force it to solve “harder problems around security, scale, and performance.” Early testers include On, Shopify and PayPal.
That sequencing is the strategy. Meta’s Muse and OpenAI’s Dots are racing for personal agents on phones; Google is wiring a coworker into the systems where money and permissions already live. A consumer agent, Spark, exists in a separate Gemini interface, but the enterprise agent is the one with a mailbox in the company graph. Whoever owns that identity owns the default route by which work gets delegated, and it is much harder to dislodge an employee than an app. In a Georgian village, whoever holds the key to the wine cellar decides which jar gets opened, and nobody changes that lock lightly. Even the model picker, which already offers Anthropic’s Claude, makes the point: Google is happy to rent out the brain as long as it keeps the badge.
An agent that “knows who approves what” is powerful when the calendar and the MCP server are right, and dangerous when they are not. So who checks the calendar it trusts? What happens when a server feeds it nonsense? And when it approves the wrong thing, who answers: the agent, the manager who tagged it, or Google? Google mentions spending caps; it does not publish independent red-team results for a coworker that can act across Slack, Jira and the warehouse. The agent race is no longer only about chat quality. It is about whose badge sits in the org chart. The badge is the product.
Who gains Google Cloud, whose agent sits inside the corporate org chart with its own identity, its own mailbox and, in effect, the cellar key.
What's absent Independent red-team results for an agent that can act across Slack, Jira and the data warehouse. And who answers when it gets things wrong?
Sources
Memory helps a robot only if it can imagine the future
Long-WAM: Scaling the Context of World-Action Models · Wei Huang et al. · NVIDIA; MIT; HKU; UCSD
Robots need visual history to judge motion and progress, but longer context often fails to help. Why? NVIDIA, MIT, HKU and UCSD researchers argue the reason is how the video model was pretrained. Long-WAM is pretrained autoregressively, predicting the future from the past, on about 10,000 hours of robot and egocentric video, then adapted so action prediction is conditioned on observed history. On RoboCasa GR-1, lengthening history from zero to 19.2 seconds lifts success from 63.3% to 78.7%; a model with bidirectional pretraining gains nothing. On LIBERO-Long it reaches 99.5%. A streaming deployment stack produces each action chunk in 107 ms on an RTX 5090, and real robots stack cups at 95% where two baselines score 0 of 20. Results are simulation-heavy and mix model and engineering gains. Why it matters: memory only helps an embodied agent that has learned to imagine what comes next, and real-time control, happily, does not require a data-centre GPU.
arXiv abstract · Full paper (PDF) · Hugging Face
Not whether the agent finished—whether you can find what it changed
What Did the Agent Actually Do? Evidence-Grounded Oversight for Long-Horizon Agents · Zhongxiang Sun et al. · Gaoling School of Artificial Intelligence, Renmin University of China; Kuaishou Technology
As people hand long coding and research jobs to agents, oversight shifts from making each decision to checking the consequential ones. The evidence for those decisions is scattered across messages, tool output and files. Researchers at Renmin University and Kuaishou build AgentMonBench, with tests for requirement gaps, silent consequential changes and feedback tracing, and propose a training-free Evidence-Grounded Behavior Graph (EBG) that groups source-linked evidence into behaviours for the monitor. Across eight models EBG mostly beats raw context and RepoGraph: with GPT-5.6 Terra, semantic F1 rises from 0.275 to 0.392 and localisation F1 from 0.309 to 0.462. In five real research tasks, Codex discloses key silent changes only when the EBG monitor is attached. The work is limited to software engineering, and some scores rely on an LLM judge. Why it matters: the hard question is no longer whether the agent finished, but whether a human can find what it changed. Who, after all, reads the whole log?
arXiv abstract · Full paper (PDF) · Hugging Face · Code
Playable is not the same as fun
Recursive Game Creator: An Agentic Product-Level Experience-Oriented Game Harness · Jiajun Chen et al. · HKU MMLab; The University of Hong Kong; Shenzhen Loop Area Institute
Game-making agents now produce code that runs, but runnable is not the same as fun, and judging fun by slow on-screen play skews evaluation. HKU MMLab’s Recursive Game Creator closes the loop with four roles: a Designer plans, a Builder implements, a coding-native Player writes reusable policies to play the game through programmatic interfaces, and a Reviewer scores the resulting trajectories plus visual evidence, keeping the better version and writing the next critique. The harness scores 77.89 on GameCraft-Bench, a state-of-the-art result. On GameASG-Bench, strict task success reaches 53.2%, a 34.1% improvement over the same-model baseline, and 93.4% of runtime checks pass. A user study reports longer playtime and higher ratings. Caveats: code is “coming soon” (we have heard that before), both benchmarks are new, and the harness may matter more than the model. Why it matters: it is a working template for agents optimised for experience rather than pass rates.
arXiv abstract · Full paper (PDF) · Hugging Face
Beijing will not pace with you
A census of 857 Chinese model releases: only 3.6% ever published a safety result. Read it before the next release lands.
Probability without a novel
The clearest plain-English guide to TypeSafe’s probability-only Jev model and why rivals are copying it. Plain English on AI is rare; this is welcome.
What the agents missed in the UV sky
A candid field note on Claude Science, including the artefacts the agents missed and a human caught. Honesty about failure beats any demo.
False fronts, Category 5
The richest primary-source case study yet of AI-assisted influence operations, invented journalists and all.
When the drop is process, not a model
A dense weekly map of a week defined by math standards and safety firings; follow the receipts, and bring coffee.
Who’s gaining ground
Big gets bigger; capability still leaks out. For now.
Concentrating
Read the reasoning
Anthropic routes Claude into critical infrastructure through incumbent contractors, Google’s agent gets a badge in the corporate graph, OpenAI shows how much it sees of covert users, and Arena becomes everyone’s paid scoreboard. Dispersal stays at the edges: Manus refinanced in China, Jev copies, open research harnesses. Distribution and oversight concentrate.
Score by issue
Score from −5 (spreading) to +5 (concentrating)