Patrick Rowe

Jun 2026

SMILES is a strange language to model

Canonicalisation, invalid strings, and what tokenisation costs you.

It looks like text and it is not

SMILES is a serialisation of a graph. Two strings that differ everywhere can denote the same molecule, and two strings that differ by one character can denote molecules that behave nothing alike. A language model trained on it inherits both problems.

Canonicalisation is a modelling decision

Outline: training on canonical SMILES only versus augmenting with randomised traversals; what each does to the effective size of the training set and to what the model learns about the underlying graph.

Invalid strings are information

Outline: validity rate as a metric is less useful than it looks. A model at 95% validity and a model at 99% can be differently wrong, and the interesting question is what the invalid 1% has in common.

What tokenisation costs

Outline: character-level versus atom-level versus learned subwords, and the ring-closure digits that break all three.

What to write next

  • Measured numbers from the model behind the molecular generation project, once those exist and can be quoted honestly.
  • A figure showing the failure modes, generated from a committed script.