Antibody selection and engineering workflows
Antibody discovery has the opposite problem to most of computational chemistry. It is not short of data and it is not short of candidates. A single campaign produces far more binders than anyone can characterise properly, and the real constraint is assay time. The useful question is not which of these is best but which of these is worth measuring, and can you decide early enough for the answer to change what happens next.
That framing is what the work was about. Three years of it at AbCellera in Vancouver, across more than thirty clinical programmes.
What the work was
The work had three layers, and the least glamorous one mattered most.
Data. High-throughput antibody discovery runs on lab automation, and lab automation produces data that is voluminous, heterogeneous and rarely assembled with modelling in mind. A large fraction of the effort went into the pipelines and the data architecture underneath everything else, because a scoring function is worth nothing if the assay results it is trained on cannot be joined reliably to the sequences that produced them.
Simulation. Structure-based work on the molecules themselves: the interface a binder makes with its target, and the properties that decide whether a good binder can be manufactured, formulated and dosed at all.
Scoring. Models over both, aimed at ranking and triage rather than at prediction for its own sake. The output that matters is a shortlist a team will actually run.
What is visible from outside
Most of this is unpublished, as industrial work generally is. Two pieces of it are public, and both are on this site:
Selecting CD3-binding antibodies for T-cell engagers is the clearest published example of the underlying approach: profile a large panel functionally, find the axes that actually organise it, then select across those axes rather than picking favourites. The peptide-MHC T-cell engager programme is one of the campaigns that consumed the output.
Between them those two records carry three conference posters, and all three are hosted here in full:
- Selecting CD3 binders for a T-cell engager platform (PDF, 2.9 MB). SITC 2023, poster 1367. The selection method itself, and the only one of the three where I am named as an equal contributor.
- MAGE-A4 peptide-MHC T-cell engagers (PDF, 2.3 MB). SITC 2023, poster 1395.
- T-cell engagers against MAGE-A4 (PDF, 3.4 MB). AACR 2024, poster 2373. The same programme a year further on.
I am not going to describe the internal platform beyond that, or name targets or partners. The posters are the part that has been cleared for the outside world, and they are a fair sample of the shape of the work.
Where this sits
Read against the earlier pages here, this looks like a change of field. It is less of one than it appears. Ranking a panel of antibodies by a cheap model, against the possibility that the cheap model is confidently wrong somewhere you have not looked, is the carbon potential question with proteins substituted in. The units change, but the underlying failure modes do not.