Every experienced GC–MS analyst runs a small physics simulation in their head. A library search proposes a compound for a peak, and before accepting it they check the clock: does something with that structure plausibly elute there, on this column, under this program?
That sanity check can catch otherwise plausible identifications, yet in many workflows it remains largely manual. It is also a natural supervised-learning problem, which makes it one of the highest-leverage places to put a model in an analytical lab.
Why indices and not times
Absolute retention times are a moving target. They depend on column dimensions, carrier-gas conditions, temperature program, column aging, and other method and instrument changes, which makes raw retention times difficult to compare directly across systems.
The field solved much of this normalization problem with retention indices. Kováts introduced the original retention-index system for isothermal GC in 1958; van den Dool and Kratz later generalized it to linear temperature-programmed separations. Both express retention relative to reference compounds, typically an n-alkane series.
Retention indices are therefore substantially more transferable than raw retention times, but they are not universal molecular constants. Stationary-phase chemistry still matters, and temperature and other chromatographic conditions can introduce additional variation. A reference RI measured on a chemically similar phase is much more informative than one measured on an unrelated column.
The retention-data landscape is better than it first appears. NIST maintains the largest curated GC retention-index collection: NIST26 contains about 528,000 RI values covering more than 216,000 compounds. At that scale, retention data become useful not only as a reference for analysts, but also as a foundation for predictive models.
There is an important catch: coverage is uneven.
A compound may have several measurements on common stationary phases and none on the exact phase used by a particular laboratory. Detailed chromatographic conditions are available only where they were reported in the underlying literature. And the full NIST library is a commercial reference product rather than a single standardized, open machine-learning dataset.
This is why finding “the retention index for this compound” can be easy while finding the relevant retention index for this compound on this particular column can still be surprisingly difficult.
For comparison, liquid chromatography has the METLIN Small Molecule Retention Time (SMRT) dataset, which released experimentally measured reverse-phase retention times for 80,038 molecules specifically to support model development. GC has no single directly equivalent open release at the same level of standardization, but its accumulated retention-index collections still provide a substantial foundation for structure-to-retention modeling.
Where the models are
Structure-based retention prediction has a long history in quantitative structure–retention relationships, or QSRR.
Traditional approaches represent a molecule using calculated physicochemical or topological descriptors and fit a regression or machine-learning model to retention data. Gradient-boosted trees, support-vector methods, and related descriptor-based approaches remain useful, particularly when datasets are modest or interpretability and training cost matter.
At larger scale, deep-learning models have shown that molecular structure alone contains enough information to predict retention surprisingly well.
Convolutional neural networks operating on molecular representations have predicted retention indices on common non-polar stationary phases with errors on the order of tens of RI units. Graph neural networks trained on more than 100,000 NIST compounds have likewise achieved mean errors of roughly a few tens of RI units. Work by Matyushin and colleagues extended deep-learning approaches to polar and mid-polar phases, reporting mean absolute errors of approximately 16–50 RI units depending on stationary phase and test set.
More recently, NIST developed AIRI, a deep-learning model for standard semipolar columns. AIRI reported a mean absolute error of 15.1 RI units and, importantly for practical use, estimates uncertainty for individual predictions.
NIST has also used AIRI predictions to extend retention-index coverage for compounds in its EI mass-spectral resources, illustrating the transition from retention prediction as a benchmark problem to retention prediction as part of an identification workflow.
The reported error numbers should not be compared too literally across publications. Datasets, stationary phases, train/test splits, and validation strategies differ.
The broader result is what matters: molecular structure contains enough information for predicted retention to become a useful additional identification signal.
It does not replace experimental retention data. It does not need to.
The layer the papers rarely discuss: the lab’s own archive
Public and commercial reference collections provide breadth. A laboratory’s own analytical history provides specificity.
Every confirmed identification can associate a molecular structure with an observed retention index, stationary phase, and chromatographic method under conditions directly relevant to that laboratory.
Those measurements are more than historical records. They can become calibration data.
A general structure-to-retention model can be adapted to a particular column family or method using the laboratory’s own confirmed measurements. This is especially valuable for stationary phases with sparse public reference data. Published work has shown that transfer and second-stage modeling approaches can leverage predictions from data-rich stationary phases to build useful models for phases with much smaller experimental datasets.
The amount of local data required, and the accuracy improvement it provides, is method-dependent and should be measured rather than assumed. But the principle is straightforward: structured analytical history can turn a generic model into one better matched to the laboratory where it is actually being used.
This is also an argument for structuring analytical history in the first place.
A confirmed identification is not only the result of one run. Properly captured, it becomes evidence that can improve every future run on the same method.
And this is where the operational problem becomes more interesting. Broad structure-to-retention models are increasingly capable. Matching those models to the actual stationary phase, method, and analytical history of a laboratory is where much of the practical value begins.
Evidence, never oracle
How prediction enters the workflow determines whether analysts trust it.
For each chromatographic peak, candidate structures from a spectral-library search can be assigned a retention-plausibility score based on the difference between observed and predicted retention, ideally taking the model’s uncertainty into account. That information can then be combined with spectral similarity and other available evidence to re-rank the candidate list.
A candidate with a convincing mass spectrum but highly implausible retention can move down the list.
When several structures produce similar spectra, a candidate whose retention is exactly where the model expects it can move up.
The important point is not that the model makes the identification. It helps the analyst prioritize the evidence.
Uncertainty matters here. A prediction ten RI units away from an observation means something very different when the expected model error is five units than when it is fifty. NIST’s AIRI work explicitly estimates uncertainty for individual predictions, pointing toward the kind of behavior practical systems should expose rather than hide.
The same machinery can assist with other chromatographic tasks. Predicted retention overlap, combined with spectral or deconvolution evidence, can help flag possible unresolved co-elutions. Models trained or calibrated for a target stationary phase can also provide approximate expected elution regions when transferring compound panels to a new method.
None of this touches the confirmation standards of a validated or regulated laboratory, and it shouldn’t.
Predicted retention does not turn a tentative identification into a confirmed one. It does not replace authentic reference standards where those are required. It is another independent piece of evidence that can make the path to the right answer shorter.
What it changes is where the analyst’s morning goes: from paging through a raw library dump to reviewing a shorter, physically plausible, uncertainty-annotated list.
Multiply that across every run, every analyst, and every site, and retention prediction stops being a paper topic and starts becoming infrastructure.
That is the standard we think structure-to-property models in analytical chemistry should be held to now: not benchmark error, but ambiguity removed and hours returned to the lab.
References
Kováts, E. (1958). Gas-chromatographische Charakterisierung organischer Verbindungen. Teil 1: Retentionsindices aliphatischer Halogenide, Alkohole, Aldehyde und Ketone. Helvetica Chimica Acta, 41, 1915–1932. DOI: 10.1002/hlca.19580410703.
van den Dool, H.; Kratz, P. D. (1963). A generalization of the retention index system including linear temperature programmed gas–liquid partition chromatography. Journal of Chromatography, 11, 463–471. DOI: 10.1016/S0021-9673(01)80947-X.
Babushok, V. I.; Linstrom, P. J.; Reed, J. J.; Zenkevich, I. G.; Brown, R. L.; Mallard, W. G.; Stein, S. E. (2007). Development of a database of gas chromatographic retention properties of organic compounds. Journal of Chromatography A, 1157(1–2), 414–421. DOI: 10.1016/j.chroma.2007.05.044.
Matyushin, D. D.; Sholokhova, A. Yu.; Buryak, A. K. (2019). A deep convolutional neural network for the estimation of gas chromatographic retention indices. Journal of Chromatography A, 1607, 460395. DOI: 10.1016/j.chroma.2019.460395.
Qu, C. et al. (2021). Predicting Kováts Retention Indices Using Graph Neural Networks. Journal of Chromatography A, 1646, 462100. DOI: 10.1016/j.chroma.2021.462100.
Matyushin, D. D.; Sholokhova, A. Yu.; Buryak, A. K. (2021). Deep Learning Based Prediction of Gas Chromatographic Retention Indices for a Wide Variety of Polar and Mid-Polar Liquid Stationary Phases. International Journal of Molecular Sciences, 22, 9194. DOI: 10.3390/ijms22179194.
Geer, L. Y.; Stein, S. E.; Mallard, W. G.; Slotta, D. J. (2024). AIRI: Predicting Retention Indices and Their Uncertainties Using Artificial Intelligence. Journal of Chemical Information and Modeling, 64. DOI: 10.1021/acs.jcim.3c01758.
Domingo-Almenara, X. et al. (2019). The METLIN small molecule dataset for machine learning-based retention time prediction. Nature Communications, 10, 5811. DOI: 10.1038/s41467-019-13680-7.