Field note
Anthropic published two wet-lab-adjacent results on 18 August. In the first, Claude models designed protein binders from scratch against 15 targets and succeeded against 14, with between 22% and 35% of individual designs binding successfully depending on setup, against the 10-15% Anthropic describes as typical for protein design campaigns today; some of the strongest designs bound several times more tightly than the best previously published result. In the second, Claude Opus 5 was given a contract lab's raw NMR and LC-MS files and a two-sentence prompt, and returned finished analytical results.
The interesting variable is not the hit rate but what the hit rate is measured against. Designing binders is an early drug-design task that has historically taken a specialist weeks or months per target, so a per-design success rate that roughly doubles changes how many candidates a lab can afford to test, rather than removing the lab from the loop.
That is the shape of the shift tracked in Arcee teaching an open model to do science and in Firecrawl adding 41 million life-science papers to its research index: the model is becoming a throughput multiplier on existing scientific pipelines, which is how the AI capability hub reads results like these.
These are vendor-reported results, the comparison baseline is Anthropic's characterisation of typical campaigns, and the analytical-chemistry example is a single illustrative run. The condition to watch is whether the binder results appear in peer-reviewed form with the failed target included, because a 14-of-15 headline is only interpretable next to the target that did not work.