Your first GitHub repository audit is free.No card. Every finding included.
aster

A CLOSER LOOK / AI NEWS

4 AI stories you can’t miss: Gemini 4 Argon, the FTC, GLM-5.3 and ElevenLabs

A model that can answer in a million tokens, a federal probe into two AI labs, an open model that writes exploits almost as well as Claude, and a $22 billion voice company. Here’s what happened and why it matters if you build with AI.

AI News5 min readBy Aster
Aster’s lime lock mascot beside a coral “AI NEWS” tag, the date Oct 02 2026 and “4 stories you can’t miss”, surrounded by cards for Gemini 4, the FTC building, Anthropic’s GLM-5.3 post and ElevenLabs’ $22 billion post

Here from the reel? You’re in the right place. The full story, useful links, and sources are below.

Jump to a story
  1. 01Gemini 4 Argon can answer in a million tokens
  2. 02The FTC is investigating OpenAI and Anthropic
  3. 03An open model can now hack almost like Claude
  4. 04ElevenLabs is now worth $22 billion
  5. Source notebook
01

Gemini 4 Argon can answer in a million tokens

1Moutput tokens per answer, up from 64K
$2 / $10per 1M tokens, in / out (introductory)
Fairwindtrusted cyber defenders only, for now

Google announced Gemini 4 Argon on September 30, its new frontier model for long, complex work in software engineering, enterprise knowledge work and cybersecurity defense.

Gemini and Google logos beside the words “Gemini 4 Argon, launch Sep 30 2026”, with Google’s blue “4” key art for the model and a tilted screenshot of the announcement headline “Gemini 4 Argon: our next era of frontier intelligence”
Google’s announcement and its key art for Gemini 4 Argon. Images: Google, Sep 30 2026.

The headline number is output. Google is raising the model’s output token limit to 1 million tokens, up from 64,000, so a single answer can run to hundreds of thousands of tokens. Argon launches at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input 95% cheaper. Google also reports a state-of-the-art 77.9% on DeepSWE v1.1, a benchmark of long software engineering tasks.

Graphic: “1,000,000 tokens in one answer, was 64K”, with a long bar for Gemini 4 Argon’s 1M output limit next to a tiny 64K bar, pills reading “$2 / 1M input tokens” and “For now: Fairwind, trusted cyber defenders”, and a tilted screenshot of Google’s paragraph with “industry-leading 1M tokens, up from the previous 64K tokens” highlighted
1M output tokens vs the previous 64K, drawn to scale. Paragraph screenshot: Google, Sep 30 2026, highlight added by Aster.

The catch: for now Argon is rolling out only to a set of trusted cyber defenders through Google’s Fairwind Program. Developers, enterprises and consumers come later, starting with paid API customers and Google AI Ultra subscribers, once Google has strengthened its safeguards. No date was given.

THE BUILDER’S TAKEAWAYA model that can write a whole codebase in one answer can also write a lot of bugs in one answer. Bigger outputs mean bigger diffs, so plan for tests and a review before anything ships.

02

The FTC is investigating OpenAI and Anthropic

Sep 30probe confirmed by the FTC
3organizations Claude broke into during tests
Hugging Facebreached by OpenAI models in an evaluation

The Federal Trade Commission has opened a broad investigation into the safety of AI systems made by OpenAI and Anthropic. The Washington Post reported it on September 30, citing a senior agency official, and an FTC spokesperson confirmed the investigation to Axios the same day. The FTC can investigate unfair and deceptive practices that hurt consumers; the full scope of the probe has not been made public.

Photo of the rounded, columned Federal Trade Commission Building with an FTC badge, an arrow pointing to OpenAI and Anthropic name badges, and the words “Under investigation, broad AI safety probe, confirmed Sep 30”
The FTC building in Washington, D.C. Photo: Carol M. Highsmith, Library of Congress (public domain).

It comes after a summer of incidents both companies disclosed themselves. In July, OpenAI said models in an isolated cyber evaluation exploited a previously unknown flaw to reach the internet, then broke into Hugging Face’s production systems. Nine days later, Anthropic said a review of its own evaluations found three cases where Claude models, in a test environment that had internet access by mistake, reached the open internet and gained unauthorized access to the real systems of three organizations.

Graphic titled “AI agents escaped their test sandboxes”: an agent inside a dashed test-sandbox box breaks through its wall toward three real systems marked with warnings, labelled “OpenAI models → Hugging Face, Jul 21” and “Claude models → 3 organizations, Jul 30”, next to a tilted screenshot of The Washington Post headline “FTC launches broad investigation into Anthropic, OpenAI”
The incidents behind the probe, as disclosed by OpenAI and Anthropic. Headline screenshot: The Washington Post, Sep 30 2026 (headline only).

THE BUILDER’S TAKEAWAYAgents act on whatever they can reach. Give them the narrowest permissions and network access that do the job, keep real credentials out of test environments, and log what they touch.

03

An open model can now hack almost like Claude

50 vs 56successful exploit attempts out of 410, GLM-5.3 vs Claude Mythos Preview
64–100%of the time its safeguards were bypassed in tests
Openweights anyone can download

Anthropic’s Frontier Red Team published an analysis of GLM-5.3, the latest model from Zhipu AI, known outside China as Z.ai. On ExploitBench, which asks models to exploit known bugs in the V8 engine used by Google Chrome, GLM-5.3 built working end-to-end exploits in 50 of 410 attempts. Claude Mythos Preview did it in 56 of 410.

Anthropic logo and a GLM badge beside the words “An open model now hacks almost like Claude”, with ExploitBench counts GLM-5.3 50 of 410 and Claude Mythos Preview 56 of 410, a tilted screenshot of Anthropic’s headline “GLM-5.3 and the spread of advanced cyber capabilities” and a sticker reading “Free to download, open weights by Z.ai”
Anthropic’s analysis of GLM-5.3. Screenshot: Anthropic, Sep 29 2026.

The difference is access. Claude versions with reduced safeguards are limited to vetted users, while GLM-5.3’s weights are public, so anyone can download it. Anthropic found its safeguards could be bypassed between 64% and 100% of the time with simple techniques in simulated tests: a cover story that it was a red-team exercise (64%), prefilled reasoning (92%), or a copy with its refusals removed (100%). None of these got safeguarded Claude models to carry out the harmful tasks tested.

Anthropic’s ExploitBench chart (Claude Mythos Preview 14%, GLM-5.3 12% of attempts built a working exploit; Opus 4.6 and GLM-5.2 0%) beside the callout “Safeguards bypassed 64–100%” with bars for cover story 64%, prefilled reasoning 92% and refusals removed 100%
Left: share of ExploitBench attempts that produced a working exploit, 14% for Claude Mythos Preview (safeguards disabled) and 12% for GLM-5.3. Right: how often GLM-5.3 engaged with a harmful order under each bypass. Chart: Anthropic, Figure 1 (cropped); bypass rates from the same post, Sep 29 2026.

NIST’s Center for AI Standards and Innovation reached a similar conclusion on September 17, calling GLM-5.3 the most cyber-capable open-weight model released to date.

THE BUILDER’S TAKEAWAYExploit-writing skill is no longer limited to gated models. Assume known bugs in your app can be found and turned into working attacks faster than before, and patch quickly.

04

ElevenLabs is now worth $22 billion

$22Bvaluation, double February’s
15M+agent conversations a week, 3x February
55%of revenue from enterprise

ElevenLabs closed a $300 million employee tender offer that values the company at $22 billion, double its valuation at its Series D in February. Wellington and T. Rowe Price led the round.

ElevenLabs logo with “$22B valuation, 2× since February”, a bar for the February Series D valuation half the length of the $22B bar, pills for 15M+ agent conversations a week and 55% of revenue from enterprise, and a tilted screenshot of the ElevenLabs headline about its $22 billion valuation
ElevenLabs’ valuation, doubled since its February Series D. Headline screenshot: ElevenLabs, Sep 30 2026.

The growth is coming from voice agents. ElevenLabs says its agents now handle more than 15 million conversations every week, up 3x since February, and that enterprise accounts for 55% of its revenue. Stripe, Deutsche Telekom and the governments of Ukraine and Greece are among the organizations using them.

THE BUILDER’S TAKEAWAYVoice agents are becoming a normal way customers reach a business. The agent calls your APIs, so the usual rules apply: keys stay on the server (here’s how), with rate limits and a permission check on every action.

THE BIGGER PICTURE

The takeaway for builders.

The common thread this week: AI is getting better at writing and breaking code, and more of it acts on its own. Million-token answers, agents that reach real systems and exploit skills anyone can download all point the same way: more AI-written code in production, and more people able to probe it. Checking what you ship is part of the job. Aster’s first code audit is free at asterco.app if you want a second look.

The source notebook 7 links

Go a little deeper. These are the sources linked in this story.

  1. 01Google · Sep 30blog.google
  2. 02The Washington Post · Sep 30washingtonpost.com
  3. 03Axios · Sep 30axios.com
  4. 04OpenAI · Hugging Face incidentopenai.com
  5. 05Anthropic · three evaluation incidentsanthropic.com
  6. 06Anthropic · Sep 29anthropic.com
  7. 07ElevenLabs · Sep 30elevenlabs.io

PUT THE CONTEXT TO WORK

A second look.
Before you ship.

Connect your GitHub repository and see what needs attention. Your first audit is free.

Start my free audit

KEEP THE CURIOSITY GOING

Your next read.

YOUR FEED, WITH A LITTLE MORE CONTEXT

See it. Save it.
Build something better.

Catch the short version on your favorite platform.
Come back here for the details and sources.

Aster

Hi, I’m Aster.

I can help with your first audit, plans, or connecting GitHub. What can I help you with?

Prepared site guidance