A CLOSER LOOK / AI NEWS
4 AI stories you can’t miss: Gemini 4 Argon, the FTC, GLM-5.3 and ElevenLabs
A model that can answer in a million tokens, a federal probe into two AI labs, an open model that writes exploits almost as well as Claude, and a $22 billion voice company. Here’s what happened and why it matters if you build with AI.

Here from the reel? You’re in the right place. The full story, useful links, and sources are below.
Jump to a story
Gemini 4 Argon can answer in a million tokens
Google announced Gemini 4 Argon on September 30, its new frontier model for long, complex work in software engineering, enterprise knowledge work and cybersecurity defense.

The headline number is output. Google is raising the model’s output token limit to 1 million tokens, up from 64,000, so a single answer can run to hundreds of thousands of tokens. Argon launches at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input 95% cheaper. Google also reports a state-of-the-art 77.9% on DeepSWE v1.1, a benchmark of long software engineering tasks.

The catch: for now Argon is rolling out only to a set of trusted cyber defenders through Google’s Fairwind Program. Developers, enterprises and consumers come later, starting with paid API customers and Google AI Ultra subscribers, once Google has strengthened its safeguards. No date was given.
THE BUILDER’S TAKEAWAYA model that can write a whole codebase in one answer can also write a lot of bugs in one answer. Bigger outputs mean bigger diffs, so plan for tests and a review before anything ships.
The FTC is investigating OpenAI and Anthropic
The Federal Trade Commission has opened a broad investigation into the safety of AI systems made by OpenAI and Anthropic. The Washington Post reported it on September 30, citing a senior agency official, and an FTC spokesperson confirmed the investigation to Axios the same day. The FTC can investigate unfair and deceptive practices that hurt consumers; the full scope of the probe has not been made public.

It comes after a summer of incidents both companies disclosed themselves. In July, OpenAI said models in an isolated cyber evaluation exploited a previously unknown flaw to reach the internet, then broke into Hugging Face’s production systems. Nine days later, Anthropic said a review of its own evaluations found three cases where Claude models, in a test environment that had internet access by mistake, reached the open internet and gained unauthorized access to the real systems of three organizations.

THE BUILDER’S TAKEAWAYAgents act on whatever they can reach. Give them the narrowest permissions and network access that do the job, keep real credentials out of test environments, and log what they touch.
An open model can now hack almost like Claude
Anthropic’s Frontier Red Team published an analysis of GLM-5.3, the latest model from Zhipu AI, known outside China as Z.ai. On ExploitBench, which asks models to exploit known bugs in the V8 engine used by Google Chrome, GLM-5.3 built working end-to-end exploits in 50 of 410 attempts. Claude Mythos Preview did it in 56 of 410.

The difference is access. Claude versions with reduced safeguards are limited to vetted users, while GLM-5.3’s weights are public, so anyone can download it. Anthropic found its safeguards could be bypassed between 64% and 100% of the time with simple techniques in simulated tests: a cover story that it was a red-team exercise (64%), prefilled reasoning (92%), or a copy with its refusals removed (100%). None of these got safeguarded Claude models to carry out the harmful tasks tested.

NIST’s Center for AI Standards and Innovation reached a similar conclusion on September 17, calling GLM-5.3 the most cyber-capable open-weight model released to date.
THE BUILDER’S TAKEAWAYExploit-writing skill is no longer limited to gated models. Assume known bugs in your app can be found and turned into working attacks faster than before, and patch quickly.
ElevenLabs is now worth $22 billion
ElevenLabs closed a $300 million employee tender offer that values the company at $22 billion, double its valuation at its Series D in February. Wellington and T. Rowe Price led the round.

The growth is coming from voice agents. ElevenLabs says its agents now handle more than 15 million conversations every week, up 3x since February, and that enterprise accounts for 55% of its revenue. Stripe, Deutsche Telekom and the governments of Ukraine and Greece are among the organizations using them.
THE BUILDER’S TAKEAWAYVoice agents are becoming a normal way customers reach a business. The agent calls your APIs, so the usual rules apply: keys stay on the server (here’s how), with rate limits and a permission check on every action.
THE BIGGER PICTURE
The takeaway for builders.
The common thread this week: AI is getting better at writing and breaking code, and more of it acts on its own. Million-token answers, agents that reach real systems and exploit skills anyone can download all point the same way: more AI-written code in production, and more people able to probe it. Checking what you ship is part of the job. Aster’s first code audit is free at asterco.app if you want a second look.
The source notebook 7 links
Go a little deeper. These are the sources linked in this story.
PUT THE CONTEXT TO WORK
A second look.
Before you ship.
Connect your GitHub repository and see what needs attention. Your first audit is free.
Start my free audit

