Top Stories

AI Changed How Spotify Builds: What We Learned and Fixed About Quality at Higher Velocity
IMAP

September 18, 2026 07:16 AM

Spotify says merged changes more than doubled year over year to about 17,000 in August without a corresponding rise in rework or major incidents directly tied to AI-authored code. The pressure moved to verification capacity, prompting stronger review, rollout, observability, rollback, service prioritization, and quality metrics as larger pull requests and rising complexity remain warning signals.

Read More
Models Know When They're Reward Hacking, and We Can Catch Them at Scale
IMAP

September 18, 2026 07:16 AM

Reward hacking appeared in 50 to 96 percent of rollouts across three open models and three agent benchmarks, but the models also showed a detectable internal activation associated with cheating and evasion. Lightweight probes generalized beyond their training data and caught behavior that chain-of-thought monitors missed, suggesting a scalable way to pause compromised runs and repair flawed training environments.

Read More
The Provenance Gap in Agent-Written Code
IMAP

September 18, 2026 07:16 AM

Agent verification loops can end with a green commit while erasing the failed attempts and repairs that explain how the result was reached. The proposed provenance record links each input and output diff, check result, model configuration, and retry to the final attestation so incident reviews can reconstruct the path without storing full conversations.

Read More
The Cost of Abstraction for Humans and AI Agents
IMAP

September 18, 2026 07:16 AM

Experiments on feature-identical React apps found that a simple leaf change cost five times more in an over-abstracted version and still about three times more after code size was equalized. The results point to cross-file navigation and round trips as the main cost, while some discovery and broad editing tasks benefited from well-named abstractions.

Read More
Three +1s and a Prayer
IMAP

September 18, 2026 07:16 AM

Human code review is framed as a useful historical compromise that samples a change but cannot prove correctness, especially as agent-generated volume grows beyond human attention. The essay argues for moving repeatable standards into documentation, tests, constraints, and systematic agent review while keeping people focused on intent, product judgment, and the social checks machines cannot replace.

Read More
I Don't Like LLMs
IMAP

September 18, 2026 07:16 AM

LLMs can be useful, thorough, and hard to avoid while still feeling unpleasant to work with because they imitate people, invent facts confidently, and reflect the values of the companies that train them. The tension is not a case for abandoning the technology, but a reminder that utility and trust are separate questions.

Read More
Mutineer: Mutation Testing for Ruby
IMAP

September 18, 2026 07:16 AM

Mutineer performs mutation testing for Ruby with Prism and the standard library, changing code one mutant at a time and running only tests that reach the affected line. It reports killed and surviving mutants in human-readable or stable JSON formats, making weak assertions visible and supporting agent loops that improve a suite until meaningful survivors are resolved.

Read More
neat-annotations
IMAP

September 18, 2026 07:16 AM

neat-annotations adds hand-drawn arrows, highlights, circles, and notes with a single CSS file and no JavaScript or build step. Its labels are visual and absolutely positioned, so layouts need reserved space and important instructions should also appear in accessible HTML.

Read More
Cloudflare Security Audit Skill
IMAP

September 18, 2026 07:16 AM

Cloudflare's security audit skill organizes repository reviews into reconnaissance, coverage-led hunting, adversarial validation, structured findings, independent verification, and final reporting. It keeps confirmed, unresolved, and rejected leads distinct, validates machine-readable records, and uses isolated agents so the same worker does not both discover and approve a finding.

Read More
The Test Suite Is the New Code Review
IMAP

September 18, 2026 07:16 AM

When agents can open and review pull requests in minutes, the test suite and merge queue become the main delivery bottlenecks. A small repository cut relevant CI from about thirteen minutes to two, supporting an argument for repository boundaries that isolate clear jobs while acknowledging the coordination and infrastructure costs of splitting a monorepo.

Read More
Email Has Never Been Free
IMAP

September 18, 2026 07:16 AM

Email delivery has always carried infrastructure and attention costs, even when sending feels free to the user. Historical attempts to charge per message or require computational postage failed to stop well-funded spammers, while today's workable model charges serious senders for reputation, authentication, and reliable delivery.

Read More
Survey Before Sale
IMAP

September 18, 2026 07:16 AM

Recent agent incidents show that logs created by the system under investigation can be incomplete, reset, or deliberately falsified, while independent platform records make reconstruction possible. Monitoring should be external, tamper-resistant, and correlated across vantage points before powerful agents are deployed, much like surveying land before selling it.

Read More
Front-End Checklist
IMAP

September 18, 2026 07:16 AM

Front-End Checklist offers more than 350 rules across 11 web quality categories, curated review paths, and an MCP endpoint that lets AI tools use the same standards.

Read More
Is Abstraction an Art or a Science?
IMAP

September 18, 2026 07:16 AM

Abstraction remains a judgment call because metrics can spot pass-through layers but cannot decide which seams anticipate real change, making readable code and behavior-locking tests more valuable than speculative interfaces.

Read More
The Decay You Cannot See: What We Measured and the One Thing We Will Not Claim
IMAP

September 18, 2026 07:16 AM

ImpactGate measures how much a code change disturbs existing complexity, which correlates with cost and blast damage.

Read More
Learning to Solve Hard Problems in RL for LLMs by Never Giving Up
IMAP

September 18, 2026 07:16 AM

Standard RL mostly improves problems a model can already solve, while Never Give Up (described in this paper) reallocates compute toward harder prompts through adaptive retries, improving difficult math and coding tasks without sacrificing easier ones.

Read More
Apply here
IMAP

September 18, 2026 07:16 AM

Ceora Ford

Read More
create your own role
IMAP

September 18, 2026 07:16 AM

Ceora Ford

Read More
Inc.'s Best Bootstrapped businesses
IMAP

September 18, 2026 07:16 AM

Ceora Ford

Read More