September 03, 2026 10:05 AM
Meta released Muse Spark 1.3 with improved coding and agentic performance, alongside changes intended to make the model easier to use in production. It has begun rolling out through Muse Code and the Meta Model API, with its highest reasoning mode awaiting additional safety testing.
Read MoreSeptember 03, 2026 10:05 AM
Meta is moving closer to launching its agent super app under the launch name Muse. A waitlist is now available for the iOS app. Meta has added a setting for computer use on its desktop app. The company appears to be testing a model variant that supports computer control.
Read MoreSeptember 03, 2026 10:05 AM
ArtificialAnalysis' intelligence vs. cost plot, which shows the cheapest model that can achieve each intelligence score, is misleading. It uses a logarithmic scale on the cost axis, which means viewers can't appreciate the immensity of the price difference between the cheap models and the heavy ones, nor can they realize how inconsequential the price differences are between the cheap models. It also lists open models at their datacenter pricing, which is always very expensive compared to local hardware. Most people don't need frontier-level intelligence and would be satisfied with Chinese open source models.
Read MoreSeptember 03, 2026 10:05 AM
One of the most tantalizing phrases in model development is 'new scaling axis'. Every time the industry has found a new scaling axis, it has unlocked a large boost in model effectiveness. The idea of test-time training is an appealing one as it unlocks a new scaling axis. New research has discovered some exciting techniques, but the continual learning problem has still yet to be solved.
Read MoreSeptember 03, 2026 10:05 AM
Anthropic is planning to bring METR inside for an independent review of its recent security incidents involving AI agents. While the company has paused its highest-risk RL efforts, it is also sharing research in which it intentionally created a reward-seeking version of Claude. It's hard to slow down even when it's in your own commercial interests. It seems like Anthropic will take at least short-to-medium term and prosaic alignment tasks a lot more seriously, and devote substantial resources to these efforts.
Read MoreSeptember 03, 2026 10:05 AM
This post lays out a detailed architecture for agent harnesses, covering state management, runtimes, control planes, inference, tools, interfaces, and language choices. The central argument is that unavoidable complexity should be absorbed by core abstractions rather than repeatedly pushed onto extensions and users.
Read MoreSeptember 03, 2026 10:05 AM
Gemini 3.8 Flash has improved coding, agentic, and multi-step reasoning performance at the same introductory pricing as 3.7 Flash. A specialized Flash Cyber variant has been released for vulnerability detection and automated patching through a restricted defender program.
Read MoreSeptember 03, 2026 10:05 AM
Meta has developed an AI agent that serves as a "second brain," codifying and preserving expert knowledge, enabling it to be easily accessed within organizations. This AI integrates a two-layer system: a structured, auditable knowledge architecture separates knowledge from reasoning processes, and a self-improvement loop that incorporates expert feedback without retraining models. This approach enhances productivity, allowing experts to focus on complex tasks while also maintaining consistent, high-quality outputs across large-scale assessments.
Read MoreSeptember 03, 2026 10:05 AM
Agents' new capabilities make it practical for teams to provide and manage their own infrastructure at scale. Cursor's cloud agents can now execute on dynamically scheduled pools of machines inside private networks. Agents are still started and managed from Cursor, but teams have more control over where agents execute and what infrastructure they use. This enables agents to work next to internal services and source control, run on custom hardware, or use operating systems and build pipelines that are difficult to package as a Cloud Agent build.
Read MoreSeptember 03, 2026 10:05 AM
Click here to learn more
Read MoreSeptember 03, 2026 10:05 AM
OpenAI's model is a looped transformer, which means it reuses layers in the transformer block to increase capacity without adding parameters. This can significantly increase the size of the model without increasing the amount of storage and RAM needed to host it. However, it also increases the costs of running the model as embedded text has to run through more layers. The looped transformer aspect is just a small architectural tweak and is not the reason why Astra is a really good model.
Read MoreSeptember 03, 2026 10:05 AM
World models could become the next major AI paradigm by helping systems represent environments, predict outcomes, simulate possibilities, plan, and act. The convergence of Yann LeCun, Demis Hassabis, and Fei-Fei Li suggests growing momentum around models built for decisions, not just generation.
Read MoreSeptember 03, 2026 10:05 AM
Nvidia and CrowdStrike have introduced a new family of agentic AI models dubbed SafeMind that can both find and close attack paths for customers.
Read MoreSeptember 03, 2026 10:05 AM
TxBench-AB is a novel AI benchmark that assesses LLMs' effectiveness in biomedical research.
Read MoreSeptember 03, 2026 10:05 AM
Shamez Hemani, a senior OpenAI data center employee who left the company in April to join Meta's dedicated compute team, is now a member of Anthropic's technical staff.
Read MoreSeptember 03, 2026 10:05 AM
This tool identifies the text watermark Claude uses to produce files.
Read MoreSeptember 03, 2026 10:05 AM
Unit 42 investigated a ransomware attack using frontier AI, where a human attacker breached an enterprise network with unprecedented speed.
Read More