Back to Archive
Wednesday, September 2, 2026
8 stories3 min read

Today's Highlights

1

Anthropic Releases Fable 5.1 with 75% Cheaper Cache Reads

Model ReleaseDevelopment ToolsAI Infrastructure

Anthropic has launched Claude Fable 5.1 and Mythos 5.1, targeting coding and knowledge work respectively. Fable 5.1 scores 52.6% on Terminal-Bench-Science 0.1, up from 24.7% in the previous version; performance on Terminal-Bench 4.0 improves from 42% to 55.8%, and it leads CursorBench 3.2 with 73.4%. Base pricing remains unchanged, while cache read costs drop from $1 per million tokens to $0.25. The update is now live on Claude Code, Claude Platform, and Cursor.

Read full article
2

OpenAI Teases Astra, Cybersecurity Capability Reaches Critical Level

Model ReleaseCybersecurityAI Safety

OpenAI has previewed its upcoming cybersecurity model Astra, stating its capabilities have reached the highest level 「Critical」 in the company's Preparedness Framework, indicating potential for significant harm if misused. OpenAI notes that safeguards evolve alongside capabilities, and evaluation metrics and methods will be publicly shared upon release. Although Astra is ready for deployment, the company plans to slow down development of future models to allow time for safety and alignment research, real-world feedback, and societal adaptation. Specific launch date and usage details remain undisclosed.

Read full article
3

Gemini Introduces Agentic Video, Reducing Token Usage by Up to 88%

MultimodalVideo UnderstandingAPI

Google DeepMind has introduced Agentic Video mode for its latest Gemini models and Gemini API, designed for long-video understanding. Instead of scanning entire videos at fixed frame rates, the system dynamically selects key frames and performs cross-modal reasoning using visual, audio, and text information. Official benchmarks claim improved analysis quality with up to an 88% reduction in token consumption. This mode can be enabled per video via the API, and documentation references models such as Gemini 3.7 Flash, making it suitable for cost-sensitive, long-duration video analysis tasks.

Read full article
4

AllenAI Releases BenchMIRT, Maintaining Rankings with Just 10% of Questions

Model EvaluationAI SafetyResearch

AllenAI has released BenchMIRT, a method based on multidimensional item response theory that audits large model benchmarks question-by-question, disentangling mixed signals like safety and general reasoning. Without predefined labels, the approach identifies two primary dimensions—safety and reasoning—from 16 diverse benchmarks, revealing that low scores on bias tests such as BBQ may stem more from reasoning difficulty. Experiments show that retaining only 10% to 50% of high-information questions typically preserves original model rankings. Prediction accuracy for unseen questions reaches 79%, surpassing the 70% achieved using overall average scores.

Read full article
5

Anthropic Launches Enterprise Frontier Safeguards with Zero Data Retention

Enterprise AIData SecurityAI Governance

Anthropic has introduced Enterprise Frontier Safeguards for enterprise clients using frontier models in sensitive operations, with phased deployment starting this fall. The solution combines zero data retention, adversarial attack protection, and customer control: enterprises can keep cloud storage, access logs, and security signals within their own environments. Machine audits return only predefined risk findings without transmitting customer content back. Anthropic demonstrated the architecture with Uber and Visa, emphasizing that customers retain full control over log transmission and security decisions.

Read full article
6

NVIDIA and CrowdStrike Build Closed-Loop Offensive-Defensive Agent System

CybersecurityAI AgentNVIDIA

NVIDIA and CrowdStrike have demonstrated an adaptive agent-based cybersecurity system that continuously discovers and patches defense gaps through automated red-blue teaming. The workflow covers attack execution, telemetry analysis, detection rule generation, backtesting, and live retesting against novel attacks, iterating until no viable attack paths remain in the modeled environment. The system uses a dual-model architecture: Nemotron 3 Ultra orchestrates the defense workflow, while a fine-tuned Nemotron 3 Super generates and refines detection queries to improve reliability of detection rules in real-world environments.

Read full article
7

OpenClaw Releases 2.0, Integrating Over 16,000 Pull Requests

Open SourceAI AgentDevelopment Tools

The open-source agent project OpenClaw has released version 2.0, incorporating more than 16,000 pull requests—a broad engineering update. The release focuses on overhauling components including memory management, model integration, security mechanisms, installation流程, browser experience, and team collaboration, aiming to consolidate previously fragmented agent capabilities into a unified version. Multiple summaries highlight that the update emphasizes systemic restructuring of the agent runtime stack and developer experience rather than introducing a single new model or feature. No official release page, compatibility scope, or migration requirements have been published yet.

8

Codex Desktop Includes 1.7GB Runtime for Document Processing

OpenAIDevelopment ToolsLocal Execution

Analysis of OpenAI Codex's desktop application reveals a ~1.7GB standalone runtime bundled in its cache directory, fully packing Python, Node.js, and LibreOffice, along with native binaries such as Poppler and git. The app's document plugin locates and invokes these components via skill descriptions, enabling file reading, conversion, and processing of PDFs, Word documents, and spreadsheets without relying on user-installed system environments. This design enhances execution consistency and offline capability but significantly increases local storage usage.

Read full article

Don't Miss Tomorrow's Insights

Join thousands of professionals who start their day with AI Daily Brief