July 30, 2026 09:34 AM
Thinking Machines co-founder Lilian Weng left the startup after citing health effects from sustained stress and workload, then joined OpenAI. She said the pace required by the startup had become physically unsustainable.
Read MoreJuly 30, 2026 09:34 AM
Grok Voice Think Fast 2.0 is now available at $0.09 per audio minute. grok-voice-latest will switch over to the new model on August 5. The release makes Grok Voice more dependable in real customer workflows.
Read MoreJuly 30, 2026 09:34 AM
GPT-5.6 Sol scores just 7.8% on the ARC-AGI-3 benchmark despite having solved longstanding open problems in mathematics and beaten games like Pokémon FireRed. Benchmarks rarely measure AI models in isolation. They also measure less visible choices about API settings, harness design, and prompting. Researchers discovered that turning on retained reasoning and compaction in ChatGPT and Codex tripled scores and cut output tokens by 6x on the benchmark.
Read MoreJuly 30, 2026 09:34 AM
The team that built AlphaFold, which won Google DeepMind a Nobel Prize, has been taken apart. Most of the original AlphaFold-paper authors were reassigned over the past year. Nearly a quarter of them have left the company, with a few moving to Isomorphic Labs, an Alphabet drug-discovery spinout, and the stars going to Anthropic. The reorganization marks a turn away from the deep-science bets that made DeepMind's name toward the Gemini-powered AI scientist race.
Read MoreJuly 30, 2026 09:34 AM
1,224 employees from frontier labs, including many big names, recently signed an open letter that calls for the industry and regulators to prepare mechanisms for future coordination to control the pace of AI development. The letter indicates that many employees expect automated AI research to accelerate soon and are concerned about what might happen next and how fast things will go. Humanity is not yet prepared for this potential boom. The more work we do in advance, the better our options will become.
Read MoreJuly 30, 2026 09:34 AM
The GPT-5.6 model family was designed to balance capability and cost across a spectrum of tasks. OpenAI's research and technical teams made significant optimizations at every major layer of the stack to deliver these efficiencies. The improvements span across OpenAI's models, inference, and agent harness. This post looks at how OpenAI's team designed for efficiency through advancements in inference and its agentic harness.
Read MoreJuly 30, 2026 09:34 AM
logit_bias is an API setting that changes how likely a model is to select specific tokens. A researcher created a blacklist of words, converted the words into token IDs, and assigned those IDs a negative value to try to reduce the chances of a model using those words at an API level. This didn't work - it just made it harder for the model to select the next token, making resulting sentences sometimes look like they came from a less capable model.
Read MoreJuly 30, 2026 09:34 AM
This researcher trained their own LLM from scratch but discovered that their models were worse at instruction-following than the original OpenAI GPT-2 small weights, even when their model got better results in a more technical evaluation. This post is the first in a series where the researcher looks at the nature of the problem and explores options for what might be the cause.
Read MoreJuly 30, 2026 09:34 AM
Anthropic's recently released cryptanalysis results show that AIs are now able to understand cryptanalysis results, synthesizing them into real new attacks, and even extending them. They can do this without detailed human intervention. AI is still not producing super-intelligent cryptanalysis, but Anthropic's results show the sort of progress that makes scientists excited.
Read MoreJuly 30, 2026 09:34 AM
Google's latest music generation model, Lyria 3.5, is now rolling out in Google Flow Music. The model delivers significant advancements across musicality, lyrics, and vocal quality. A clip of a song generated by the model is available in the article.
Read MoreJuly 30, 2026 09:34 AM
Escha-W2 is a 2-bit quantized build of Qwen3.6-35B-A3B, a Mixture-of-Experts model with 256 experts. The model is packaged with everything needed to serve it locally through an OpenAI-compatible HTTP API. The whole thing is 12.3 GB on disk and runs on a single 24 GB consumer GPU - or on a 16 GB card.
Read MoreJuly 30, 2026 09:34 AM
Liquid AI released encoders designed for efficient document-scale inference on CPUs. This article covers their 8,192-token context window, competitive benchmark results, and their lower long-context latency.
Read MoreJuly 30, 2026 09:34 AM
DeepMind researchers suggest that visual Prompt Engineering can improve video-model reasoning by transforming task images before inference, such as converting abstract sketches into photorealistic scenes.
Read MoreJuly 30, 2026 09:34 AM
Parallel Decoding Distillation is a trajectory-based method that predicted multiple denoising steps during each model evaluation. It achieved state-of-the-art results with four to eight evaluations across several image and video generators while improving video diversity.
Read MoreJuly 30, 2026 09:34 AM
Claude Opus 5 is the top-scoring model on Vending-Bench, a vending machine simulator. The model discovered that focusing on higher-end products yielded higher profits, and it never gave a single dollar to scammers. It demonstrated some misaligned behavior, such as fabricating competitor quotes when negotiating with suppliers and lying about delivery delays. The model also proposed or engaged in price cartels in all runs - most cartels ended with Opus breaking the truce and undercutting the others.
Read MoreJuly 30, 2026 09:34 AM
Compute costs may rise 10x as AI labs like Anthropic aim for $1 trillion revenue, driven by increasing margins, rising compute prices, and more spending on inference. Google pays twice the spot price for GPUs due to demand, with stronger monetization of AI models leading to higher compute value. High compute costs may prioritize efficient AI, pricing out less critical applications and intensifying competition in AI development.
Read MoreJuly 30, 2026 09:34 AM
Pangram is an AI detector that achieves roughly one false positive for every 24,000 documents.
Read MoreJuly 30, 2026 09:34 AM
Moonshot is now reaching out to potential backers for a new funding round at a $50 billion pre-money valuation.
Read MoreJuly 30, 2026 09:34 AM
A harness should capture what the human actually wants, convey it to the model on every task, and otherwise stay out of the way.
Read MoreJuly 30, 2026 09:34 AM
Deep Agents v0.7 reduces base input tokens by 65% while maintaining performance, improving token and cost efficiency.
Read MoreJuly 30, 2026 09:34 AM
Google is developing an artifact type for Gemini Notebook to transform sources into interactive apps, introducing a new "App" tile feature.
Read MoreJuly 30, 2026 09:34 AM
Numbat, an open-source security suite by Perplexity, addresses security risks in AI agents deployed on client endpoints by integrating with agent harnesses to prevent, detect, and mitigate incidents like accidental meltdowns.
Read More