Production code written by Claude should clear a higher bar than human-written code, not a lower one. The recommended stack at Anthropic includes lint rules, tests, Claude-driven end-to-end tests, daily fuzzers, and automated code and security reviews. If quality slips, the levers are: use Opus 5 or Fable 5.1, raise effort to high or xhigh, invest in CLAUDE.md, steer more actively, or wait for the next model.
bcherny·Thu, Sep 10 9:10pm ET
Anthropic's latest Threat Intelligence report examines how the same capabilities that make models useful, like coding ability and biology assistance, also make them dangerous. The report argues safeguards and monitoring are essential as models grow more capable.
bcherny·Fri, Sep 11 1:25am ET