Patrick Rowe

What a machine-learned potential actually learns

Descriptors, smoothness, and why transferability is mostly a data problem.

The descriptor is the whole game

A potential can only distinguish two environments that its descriptor distinguishes. Everything downstream is constrained by that one choice: the regression, the training set, the validation protocol. It is also the part most papers spend the least time on.

Smoothness is not a nice-to-have

Outline: why discontinuities in the descriptor become discontinuities in the force, and what that does to a molecular dynamics trajectory that has to conserve energy for a few million steps.

Transferability is mostly a data problem

Outline: the argument that models fail outside their training distribution in ways that look like model failures and are actually sampling failures. Needs a worked example. The amorphous carbon case is the obvious one, and the hero figure on the landing page is already the right illustration.

What to write next

  • A concrete descriptor comparison, with a figure generated from a committed script.
  • The honest version of “how do you know when you are extrapolating”.
  • Set math: true is already on; add the actual equations.