Top Stories

OpenAI's Jalapeño inference accelerator moves toward deployment
IMAP

August 26, 2026 09:26 AM

OpenAI reported first results from Jalapeño, an inference accelerator designed around low-latency agent workloads, and plans to deploy it in its own infrastructure by year-end. A large connected system keeps prompt processing and token generation close together, while AI helped design circuits and program kernels.

Read More
Perplexity partners with Nvidia to launch Portable Computer, a fully local AI agent with zero token costs
IMAP

August 26, 2026 09:26 AM

Perplexity's Portable Computer is a version of its agentic Computer platform that runs entirely on hardware users already own. The model, user data, and work can all stay on local machines with no billing credits. Every task starts on device by default, and the system asks for permission before sending any individual step to a more powerful model in the cloud. Portable Computer is now available for Pro, Max, Enterprise Pro, and Enterprise Max subscribers on Linux, with Windows support coming in September. Users will need an RTX GPU with at least 24GB of VRAM to run the agent.

Read More
Vocab Break
IMAP

August 26, 2026 09:26 AM

Claude's current tokenizer appears to only have about 15,000 entries. This is surprising, as the trend had seemed to be that more is better in this space. One theory is that Anthropic has been working around a bottleneck caused by the final softmax layer. This article takes a closer look at how Anthropic might be achieving this and the effects it might have on model training.

Read More
OpenAI and Anthropic Could Dominate Global AI Compute
IMAP

August 26, 2026 09:26 AM

Dylan Patel discusses how OpenAI and Anthropic could control most usable AI compute by 2028 as their ability to monetize FLOPs lets them outbid competitors. The conversation also covered rising AI capex, potential sovereign debt risks, and the economic forces pushing the industry toward greater centralization.

Read More
OpenAI's Jalapeño Optimizes Inference Throughput and Token Latency
IMAP

August 26, 2026 09:26 AM

OpenAI's Jalapeño inference chip was designed around its own model workloads, with early benchmarks showing higher peak throughput per kilowatt and lower token latency than the commercial systems tested on GPT-OSS 120B.

Read More
Moats in the age of floods
IMAP

August 26, 2026 09:26 AM

Abundant frontier intelligence will not eliminate application-layer moats - it shifts value toward companies that translate models into real outcomes. Durable winners will own coordination, workflow data, customer transformation, narrative, higher-level abstractions, outcome-based economics, and structural necessity.

Read More
Open Omnimodal World Models
IMAP

August 26, 2026 09:26 AM

EchoWM is an omnimodal world model that follows continuous 6-DoF camera trajectories while jointly generating 720p video, environmental sound, music, and speech. It supports first- and third-person interaction and uses progressive plus autoregressive training for synchronized long-horizon generation.

Read More
Short-Lived Credentials for AI Agents
IMAP

August 26, 2026 09:26 AM

Vercel Connect replaces long-lived API tokens with runtime-issued credentials that are scoped to individual tasks and expire automatically. Its generally available release added more than 100 connectors along with a unified integration model and production governance controls.

Read More
Granite 4.2 LLMs: How They're Built
IMAP

August 26, 2026 09:26 AM

Granite 4.2 models by IBM are dense, decoder-only reasoning LLMs available in 3B, 8B, and 30B sizes. Trained on 15T tokens, they use a five-phase strategy that includes a multi-stage RL pipeline and support native tool calling with a THINKING/NON-THINKING switch. The 8B and 30B models learn agentic behavior through RL stages in real environments, enhancing capabilities such as code editing and web searching.

Read More
‘The world seems to be ready': An interview with OpenAI head of product Thibault Sottiaux
IMAP

August 26, 2026 09:26 AM

Thibault Sottiaux is widely known as the guy who resets people's token limits whenever OpenAI's Codex hits a growth milestone. OpenAI plans to bring the same treatment to ChatGPT Work, a platform for white-collar workers leveraging AI agents. This article features an interview with Sottiaux on Codex, winning over skeptics, discovery as a product design philosophy, and the cost of intelligence.

Read More
Anthropic merges Claude chat and Cowork memory, on by default
IMAP

August 26, 2026 09:26 AM

Anthropic has merged Claude and Claude Cowork's memory systems, so the platforms now remember the same conversations. The feature is on by default. Claude now adds topics to memory while users are still conversing. Everything that is remembered is stored as a list of files under Topics in the memory settings. Users can read, edit, or delete each one individually.

Read More
Apple introduces M6 and M5 Ultra for local AI compute
IMAP

August 26, 2026 09:26 AM

Apple announced M6 and M5 Ultra chips for Mac mini and Mac Studio, expanding the model work those machines can run locally.

Read More
OpenAI's Head of Data Centers Has Left the Company
IMAP

August 26, 2026 09:26 AM

Chris Malone, the executive who was overseeing OpenAI's data-center build-out, left the company last week.

Read More
Applied Compute Agent Cloud
IMAP

August 26, 2026 09:26 AM

Applied Compute has launched AC2, a platform enabling AI teams to train, serve, and improve custom models.

Read More
Keenable builds a web index and query layer for AI agents
IMAP

August 26, 2026 09:26 AM

Keenable emerged from stealth with a claimed index of more than 100 billion documents, an API already used by unnamed AI labs, and a planned query language for combining evidence across sources.

Read More
Apply here
IMAP

August 26, 2026 09:26 AM

Jacob Turner

Read More
create your own role
IMAP

August 26, 2026 09:26 AM

Jacob Turner

Read More
Inc.'s Best Bootstrapped businesses
IMAP

August 26, 2026 09:26 AM

Jacob Turner

Read More