Patrick Rowe

May 2026

Getting a model into a screening pipeline

The gap between a good validation number and a decision someone will act on.

The validation number is the easy part

A model that ranks well on a held-out split has cleared the lowest bar in the process. The question that decides whether it gets used is different: when this model and an experienced scientist disagree, who is right, and how would anyone know?

Calibration beats accuracy

Outline: why a well-calibrated mediocre model is more useful in a screening cascade than a sharp uncalibrated one, because the downstream decision is a threshold and a threshold needs a probability that means something.

The split is the experiment

Outline: random splits flatter models. Domain-appropriate splits, whether scaffold, temporal or target-level, are what tell you whether the thing will work on next month’s chemistry.

Trust is built at the boundary

Outline: what it takes for a screening team to act on a prediction. Abstention, uncertainty that is legible, and being visibly right about the cases the team already knows the answer to.

What to write next

  • A concrete cascade with real numbers, if any can be cleared for external use.
  • The failure case: a model that was accurate, well-engineered, and never used.