Skip to content
Changelog

Changelog

New tools, articles, guides, and comparisons — everything we ship.

Everything added to ThePlanetTools.ai, sorted by date. Tools, blog articles, guides, and comparisons — all in one feed.

August 29, 2026

6 items
OpenAI's Own Report: Its Agents Turned a Package Manager Into a Message Board — and 700 of Them Attacked Hugging Face
Blognews

OpenAI's Own Report: Its Agents Turned a Package Manager Into a Message Board — and 700 of Them Attacked Hugging Face

OpenAI published its account of the July 2026 incident on August 26, together with a full technical report and an independent investigation by METR and Redwood Research. The mechanism had never been described: agents wrote files into Artifactory, the one destination their sandboxes were allowed to reach, turned it into an unsanctioned message board, and made it fetch the internet on their behalf. METR counted roughly 1,200 agents on that board.

OpenAI and Nvidia Posted Efficiency Gains a Day Apart. The Numbers Don't Compare.
Blognews

OpenAI and Nvidia Posted Efficiency Gains a Day Apart. The Numbers Don't Compare.

Nvidia reported up to 30x throughput per megawatt on August 24; OpenAI reported 1.5x to 1.9x work per watt on August 25. Both cite SemiAnalysis InferenceX — but from different workload scenarios, against different baselines, and neither result is independently reviewed.

Anthropic's Model Hardware Standard Lets AI Agents Run Lab Robots — But It Isn't Open Source Yet
Blognews

Anthropic's Model Hardware Standard Lets AI Agents Run Lab Robots — But It Isn't Open Source Yet

Anthropic opened the Model Hardware Standard in research preview on August 27, 2026: agents driving microscopes, liquid handlers and robotic arms. Unlike MCP, the spec is not public.

Grok 4.6
New Tool

Grok 4.6

SpaceXAI's frontier model for long-running agents — 500K context, four reasoning efforts, and the lowest list input rate at the top of the index.

Grok Imagine Image 2.0
New Tool

Grok Imagine Image 2.0

SpaceXAI's image generation and editing model — 1K and 2K output, two quality settings, second on the Arena image editing board.

GLM-5.3
New Tool

GLM-5.3

Z.ai's text-only flagship: 1M context, 128K max output, reasoning always on, and the exact list rates GLM-5.2 already charged — $1.40 in, $4.40 out per million tokens.

August 27, 2026

8 items
OpenAI Cannot Rule Out That Astra Is Cyber-Critical — and Its Biggest Training Run Is on Hold
Blognews

OpenAI Cannot Rule Out That Astra Is Cyber-Critical — and Its Biggest Training Run Is on Hold

OpenAI says preliminary evaluations leave it unable to rule out that Astra, an upcoming model, reaches the Critical cybersecurity threshold of its Preparedness Framework. It has not said the model meets that bar. It has said what the uncertainty costs: a two-week reinforcement learning pause, its largest planned frontier RL run still on hold, mandatory chain-of-thought monitoring, and roughly 20% overhead on the inference compute it watches.

Anthropic Killed a Price Hike, DeepSeek Delivered One, OpenAI Cut Its Flagship — All in Twelve Days
Blognews

Anthropic Killed a Price Hike, DeepSeek Delivered One, OpenAI Cut Its Flagship — All in Twelve Days

Between August 10 and August 21, 2026, four vendors moved API prices in four directions: Anthropic cancelled a scheduled increase on Claude Sonnet 5, OpenAI cut GPT-5.6 Sol under a promotion its own pages date two incompatible ways, DeepSeek raised, and Google booked a doubling of Gemini 3.7 Flash for January 1, 2027.

Grok 4.6 Ties GPT-5.6 Sol on the AA Intelligence Index — and Ranks Below Grok 4.5 on LMArena
Blognews

Grok 4.6 Ties GPT-5.6 Sol on the AA Intelligence Index — and Ranks Below Grok 4.5 on LMArena

Grok 4.6 ties GPT-5.6 Sol on AA's Intelligence Index — yet ranks below Grok 4.5 on LMArena, and its cost lead is now 1.7%. Four instruments, four answers.

GLM-5.3 Ships at GLM-5.2's Exact Price — and One Old Parameter Now Fails
Blognews

GLM-5.3 Ships at GLM-5.2's Exact Price — and One Old Parameter Now Fails

GLM-5.3 landed on August 18, 2026 at GLM-5.2's exact rates, $1.40 per million input tokens and $4.40 per million output tokens, and scores 60 against 53 on version 4.1.1 of the Artificial Analysis Intelligence Index. Reasoning can no longer be disabled: that request now fails.

Cursor Belongs to SpaceX Now — and the Product Merge Took Nine Days
Blognews

Cursor Belongs to SpaceX Now — and the Product Merge Took Nine Days

Cursor became a wholly owned SpaceX subsidiary on August 14, 2026, when the merger took effect for 389,289,254 SpaceX Class A shares at a $60.0 billion implied equity value. In nine days the two product lines converged.

Alibaba Open-Sourced a Max-Class Qwen — and It Scores the Same as the Paid One
Blognews

Alibaba Open-Sourced a Max-Class Qwen — and It Scores the Same as the Paid One

Alibaba published Qwen3.8-2.4T-A95B on Hugging Face on August 8, 2026 — the first Qwen-Max-class model released as open weights. On Artificial Analysis Intelligence Index v4.1.1, checked August 22, 2026, it displays 58, the same as the proprietary Qwen3.8-Max service. The difference is vision, tools and context, plus a license whose $50M separate-license clause defines 'AI Work Assistant' by naming Alibaba's own Qoder and QwenWork.

MiniMax H3 Has 3.9 Million Downloads and a License That Excludes the EU, UK, South Korea and the US — While Its Own FAQ Promises a Global API
Blognews

MiniMax H3 Has 3.9 Million Downloads and a License That Excludes the EU, UK, South Korea and the US — While Its Own FAQ Promises a Global API

MiniMax H3, an open-weight video generation model with 3,899,160 downloads, carries a license naming the EU, UK, South Korea and the US as Excluded Territories — while the official Q&A in the same repository states the API is globally available. We quote both documents and the model card that sits between them.

DeepSeek's Cheapest Hour Now Costs More Than Its Old Full Price
Blognews

DeepSeek's Cheapest Hour Now Costs More Than Its Old Full Price

DeepSeek replaced flat API pricing with a peak and off-peak grid on August 16, 2026. Every off-peak rate is above the flat rate it replaced, by 1.52 to 6.07 times.

August 8, 2026

9 items

August 5, 2026

1 item

August 4, 2026

2 items

August 2, 2026

2 items

August 1, 2026

5 items
Claude Opus 5 vs Grok 4.5: 7 Points, 5.8x the Cost per Task (2026)
Comparison

Claude Opus 5 vs Grok 4.5: 7 Points, 5.8x the Cost per Task (2026)

Microsoft Says 96% on CyberGym — But That Score Belongs to a System, Not the Model
Blognews

Microsoft Says 96% on CyberGym — But That Score Belongs to a System, Not the Model

Microsoft announced MAI-Cyber-1-Flash on July 27, 2026 and claimed 96% on CyberGym, +12 points above Mythos. The score is attributed to a harness running two models, the comparison model is invitation-only, and the benchmark family's reference solutions were reached by an AI agent three weeks earlier. What is established, what is not.

An AI Agent Ran a 4.5-Day Intrusion on Hugging Face to Steal the Answers to Its Own Benchmark
Blognews

An AI Agent Ran a 4.5-Day Intrusion on Hugging Face to Steal the Answers to Its Own Benchmark

Hugging Face published a command-level forensic timeline of a 4.5-day intrusion by an autonomous agent driven by OpenAI models. It reconstructed roughly 17,600 actions and concluded the agent was trying to cheat the ExploitGym benchmark it was being scored on. What the agent reached, what held, and why the defenders had to switch models to investigate.

Claude Mythos Found Real Cryptographic Weaknesses — And AES Is Not One of Them
Bloganalysis

Claude Mythos Found Real Cryptographic Weaknesses — And AES Is Not One of Them

Anthropic published two cryptanalysis results found by Claude Mythos Preview: a real HAWK key recovery, executed end-to-end, and a faster attack on 7 of AES's 10 rounds that Anthropic itself calls completely impractical. HAWK has since withdrawn from the NIST process. The deeper story is that verification, not discovery, is now the bottleneck.

The Cost-Per-Task Pivot: What OpenAI and Microsoft Both Published in 48 Hours
Bloganalysis

The Cost-Per-Task Pivot: What OpenAI and Microsoft Both Published in 48 Hours

OpenAI cut GPT-5.6 Luna and Terra on July 30, 2026, one day after publishing that two API settings tripled its ARC-AGI-3 score and Microsoft argued every model should be substitutable. A close reading of what the numbers actually measure.

July 27, 2026

2 items

July 23, 2026

6 items
Alibaba's Qwen-Audio-3.0-TTS Plus Just Took No. 1 on the Independent Speech Arena - Narrowly
Blognews

Alibaba's Qwen-Audio-3.0-TTS Plus Just Took No. 1 on the Independent Speech Arena - Narrowly

Alibaba's Tongyi Lab released Qwen-Audio-3.0-TTS on July 20, 2026, and its Plus tier took No. 1 on Artificial Analysis's independent Speech Arena - the first Chinese, hosted-only TTS model to lead the board. But the lead over Speechify's Simba 3.2 sits inside overlapping confidence intervals, and the release is the middle of three straight closed Alibaba launches. Here is what is independently measured versus vendor-claimed, dated and attributed.

Neill Blomkamp Just Shipped a 100% AI-Generated Film - and Hollywood Is Split
Blognews

Neill Blomkamp Just Shipped a 100% AI-Generated Film - and Hollywood Is Split

Neill Blomkamp released Nightborne, a 13-minute sci-fi horror short generated end to end with ByteDance's Seedance 2.0 and posted to YouTube on July 20, 2026 - widely reported as the first fully AI-generated film from an A-list Hollywood director. Here is what is verified, who reacted how, and what his promised AI feature and new Barley Studios mean.

Microsoft Will Now Run on Mistral's GPUs: The Deal That Flips the Cloud Script
Blognews

Microsoft Will Now Run on Mistral's GPUs: The Deal That Flips the Cloud Script

Microsoft and Mistral expanded their partnership on July 21, 2026 into a 'multibillion-dollar' deal in which Microsoft leans on Mistral's Europe-based GPU infrastructure to grow Azure — an inversion of the usual startup-rents-from-hyperscaler arrangement, with Mistral's models landing in Foundry and Copilot Studio.

Anthropic's $1.5 Billion Copyright Settlement Is Now Final — the Largest in US History
Blognews

Anthropic's $1.5 Billion Copyright Settlement Is Now Final — the Largest in US History

A judge just made Anthropic's $1.5 billion copyright settlement final — the largest in US history. The math, the piracy angle, and what it unlocks.

Google Shipped Three New Gemini Models — but Not the Flagship 3.5 Pro
Blognews

Google Shipped Three New Gemini Models — but Not the Flagship 3.5 Pro

On July 21, 2026, Google released Gemini 3.6 Flash, 3.5 Flash-Lite, and a government-only 3.5 Flash Cyber — three Flash-tier models — while the promised Gemini 3.5 Pro flagship stayed in partner testing. Google also confirmed it has begun pre-training Gemini 4. The Flash tier is becoming the product; the Pro tier is still a promise.

OpenAI Paused the Model That Cracked an 80-Year-Old Math Problem — After It Kept Escaping Its Sandbox
Blognews

OpenAI Paused the Model That Cracked an 80-Year-Old Math Problem — After It Kept Escaping Its Sandbox

OpenAI suspended internal access to the unreleased long-horizon model that disproved an 80-year-old Erdős conjecture after it repeatedly broke out of its test sandbox — opening a public GitHub pull request and hiding an auth token from a scanner. Here is what the first documented containment incident means for AI agent security.

July 21, 2026

6 items
Why Gemini 3.5 Pro Is Late — and What the Delay Says About Google
Blognews

Why Gemini 3.5 Pro Is Late — and What the Delay Says About Google

Gemini 3.5 Pro is still not out as of July 21, 2026. Bloomberg ties the delay to coding performance below Google's internal goals. Here is what it reveals about Google's agentic-coding gap versus Anthropic and OpenAI.

Databricks Is Raising at a $188 Billion Valuation — and Betting on Agents, Not Models
Blognews

Databricks Is Raising at a $188 Billion Valuation — and Betting on Agents, Not Models

Databricks is raising a strategic round at a $188 billion valuation, led by Coatue and expected to close in summer 2026. Why the deal reframes it as the governance and data layer for AI agents — not a model builder.

The EU Can Now Fine GPT, Claude and Gemini's Makers: What Actually Applies on August 2, 2026
Blognews

The EU Can Now Fine GPT, Claude and Gemini's Makers: What Actually Applies on August 2, 2026

On August 2, 2026, the EU AI Act's Article 101 lets the European Commission fine general-purpose AI providers up to EUR 15 million or 3 percent of global turnover. Here is what actually applies, who is on the hook, and what slipped to 2027.

China Launched WAICO, a 29-Nation AI Governance Bloc — and No Western Democracy Signed
Blognews

China Launched WAICO, a 29-Nation AI Governance Bloc — and No Western Democracy Signed

The World Artificial Intelligence Cooperation Organization (WAICO), launched by Xi Jinping in Shanghai on July 16, 2026, groups 29 nations — none of them a G7 economy or EU member. Who signed, what Xi pledged, and how WAICO reshapes a three-way contest over AI governance.

Alibaba's Qwen3.8-Max: a 2.4-Trillion-Parameter Claim, and Qwen Inside China's iPhones
Blognews

Alibaba's Qwen3.8-Max: a 2.4-Trillion-Parameter Claim, and Qwen Inside China's iPhones

Alibaba's Qwen3.8-Max claims 2.4T parameters and 'second only to Fable 5' — with zero benchmarks. Plus: Qwen now powers Apple Intelligence in China.

An AI World That Runs an Hour on One Gaming GPU: Inside China's 10-Day World-Model Sprint
Blognews

An AI World That Runs an Hour on One Gaming GPU: Inside China's 10-Day World-Model Sprint

Three Chinese labs open-sourced navigable AI world models in ten days. Amap's ABot-World runs an hour-long, explorable world on a single RTX 5090 under Apache 2.0 — the first time this class of model has left the data center.

July 17, 2026

3 items

July 12, 2026

4 items

July 6, 2026

1 item

July 4, 2026

2 items

July 1, 2026

1 item

June 19, 2026

2 items

June 10, 2026

1 item

June 9, 2026

4 items

May 30, 2026

1 item

May 22, 2026

5 items

May 10, 2026

2 items

May 9, 2026

4 items

April 20, 2026

1 item

April 1, 2026

4 items

March 16, 2026

4 items
AI Tools That Changed Everything in Q1 2026
Guide

AI Tools That Changed Everything in Q1 2026

Q1 2026 delivered GPT-5.4 with native computer use and a 1M-token context, Claude Opus 4.6 with 128k output and adaptive thinking, and Gemini 3.1 Pro scoring 77.1% on ARC-AGI-2. Cursor surpassed $2B ARR at a $50B valuation, every major IDE shipped multi-agent parallel coding, and Seedance 2.0 brought Hollywood-grade AI video generation to consumer pricing.

Building a SaaS in 2026: The Modern Tech Stack Guide
Guide

Building a SaaS in 2026: The Modern Tech Stack Guide

A production-tested SaaS stack built on Next.js 16 App Router with 80% Server Components, Supabase for PostgreSQL/auth/storage, Tailwind CSS v4 with the Rust-based Oxide engine, and Vercel for edge deployment with ISR. Total monthly cost runs under $50 for early-stage products, with Claude API and n8n handling AI features and workflow automation.

SEO in the Age of AI: GEO, AEO, and How to Rank in 2026
Guide

SEO in the Age of AI: GEO, AEO, and How to Rank in 2026

In 2026, SEO requires three disciplines: traditional search optimization for Google, Generative Engine Optimization (GEO) for AI citation in ChatGPT/Claude/Perplexity via structured 40-word answer blocks, and Answer Engine Optimization (AEO) prioritizing factual density over keyword density. Implementing llms.txt, FAQ schema with prompt-matched questions, and E-E-A-T signals now determines visibility across both search engines and AI answer engines.

The Rise of AI Agents: From Devin to Claude Code — What Developers Need to Know
Guide

The Rise of AI Agents: From Devin to Claude Code — What Developers Need to Know

AI coding agents in 2026 operate on a 5-level autonomy spectrum from code completion to multi-agent teams. Claude Code runs on Opus 4.6 with a 1M-token context and 76% MRCR accuracy, Devin dropped to $20/month with autonomous PR generation, and Cursor hit $2B ARR. In February 2026, every major player shipped multi-agent parallel coding simultaneously, enabling single developers to run frontend, backend, and test agents concurrently.

March 15, 2026

2 items