AnthropicFigmaLinearVercelStripeShopifyPostHogDatadogRampSeatGeekSierraHexTurbopufferAnthropicFigmaLinearVercelStripeShopifyPostHogDatadogRampSeatGeekSierraHexTurbopuffer
NVDA$1,234.56+1.23%
NVDA$1,234.56+1.23%
NVDA$1,234.56+1.23%
NVDA$1,234.56+1.23%
NVDA$1,234.56+1.23%
NVDA$1,234.56+1.23%
NVDA$1,234.56+1.23%
NVDA$1,234.56+1.23%
NVDA$1,234.56+1.23%
NVDA$1,234.56+1.23%
The Claude Code Daily
Plugin eval terminal output

Claude Code Launches `claude plugin eval` Command

A new `claude plugin eval` command lets you measure the actual value your plugin is adding. You create test cases, run your plugin against them, score the results, then run the same cases without the plugin to see the delta.

ClaudeDevs·Fri, Sep 11 3:56pm ET

Plugin Eval Reports Scores in Terminal and HTML

Running `claude plugin eval` scores each test case with and without the plugin in the terminal, and generates a full HTML report. Accounts that support it also get the report published as a private artifact.

ClaudeDevs·Fri, Sep 11 3:56pm ET

Bcherny's Guide to Claude Code Quality in Production

Production code written by Claude should clear a higher bar than human-written code, not a lower one. The recommended stack at Anthropic includes lint rules, tests, Claude-driven end-to-end tests, daily fuzzers, and automated code and security reviews. If quality slips, the levers are: use Opus 5 or Fable 5.1, raise effort to high or xhigh, invest in CLAUDE.md, steer more actively, or wait for the next model.

bcherny·Thu, Sep 10 9:10pm ET

Anthropic Threat Intelligence Report Flags Dual-Use AI Risks

Anthropic's latest Threat Intelligence report examines how the same capabilities that make models useful, like coding ability and biology assistance, also make them dangerous. The report argues safeguards and monitoring are essential as models grow more capable.

bcherny·Fri, Sep 11 1:25am ET
More Stories

Code Review Bar Should Match Blast Radius

Bcherny endorses the principle that review rigor should scale with risk: throwaway scripts can ship without scrutiny, but anything touching auth or money warrants line-by-line review.

bcherny·Thu, Sep 10 9:19pm ET

Bcherny Reviews Diffs on Critical Code Only

Bcherny reads diffs when code touches something critical or important to get right. He doesn't write the code himself; his job is reviewing what Claude produces in high-stakes areas.

bcherny·Thu, Sep 10 8:42pm ET

Bcherny on AI Risk: Blast Radius Matters

The relevant question isn't whether a tool can be misused (anything can) but how much damage it can do when it is. The "blast radius" framing is bcherny's lens for thinking about AI risk.

bcherny·Fri, Sep 11 1:39am ET

Bcherny Pushes Back on Fear-Mongering Accusation

Responding to a user who accused Anthropic of running a fear campaign to drive AI regulation, bcherny points to the Threat Intelligence report itself as the source of information and encourages readers to form their own conclusions.

bcherny·Fri, Sep 11 1:14am ET

Claude Code Desktop Gets Keep-Awake Session Toggle

A new toggle in the Claude Code desktop app prevents your computer from sleeping during a session, including idle time between turns, so long-running tasks don't get interrupted.

lydiahallie·Fri, Sep 11 12:56pm ET

Get the daily digest in your inbox

Every weekday morning. Unsubscribe anytime.