Walk into the QC or R&D lab of almost any flavor, fragrance or specialty chemicals company and you will find the same thing: a GC–MS bench producing detailed compositional data every day, and a folder structure on the instrument PC where much of that data effectively goes to die.
The run gets interpreted once, usually by one person. The conclusion goes into a report, spreadsheet, LIMS or ERP field. The chromatogram, spectra, candidate identities, rejected alternatives and reasoning that produced that conclusion remain behind in the instrument software.
Five years later, someone asks whether a supplier’s lots have been drifting, whether an impurity appeared before a complaint started, or whether the lab has seen this unknown before. The data technically exists. Getting an answer from it is another matter.
We have spent a lot of hours sitting next to analysts trying to understand why this happens, because on paper it shouldn’t. The instruments are excellent. The data exists. The problem is everything in between.
Where the data gets stuck
Acquisitions are born inside vendor ecosystems: Agilent .D directories, Thermo RAW files, and other proprietary formats containing spectra, chromatograms, acquisition metadata and method information.
Getting data out of those ecosystems is possible. Open interchange formats such as mzML are well established, and tools such as ProteoWizard’s msConvert can read many common vendor formats and convert them into open representations. Proprietary-format access can still depend on vendor-supplied libraries and platform-specific software, so the low-level interoperability problem is better described as manageable than solved.
But file conversion is not the interesting part. What conversion gives you is analytical measurements. What an R&D team needs is several abstraction levels higher: peaks and their integrations; component spectra after deconvolution; compound hypotheses and the evidence behind them; compositions that can be compared across runs, instruments and years. Between those two layers sits real engineering, and it is where much of laboratory digitization gets difficult.
The chain that has to work
Peak detection and deconvolution come first. Complex matrices routinely contain partially or fully co-eluting compounds, low-abundance signals under larger peaks, and integrations that require judgment. AMDIS remains a foundational reference for spectral deconvolution, but the important operational requirement is not that every instrument produce an identical peak table. It is that processing be reproducible, inspectable and versioned.
Retention provides another axis of evidence. Absolute retention times shift with column dimensions and condition, flow, trims and temperature programs. Retention indices, calculated against reference compounds such as n-alkanes, make historical comparisons more useful, although they are not universal constants and still depend on stationary phase and chromatographic conditions.
Identification then combines spectral evidence with retention behavior and chemical context. Library matching against NIST and other reference collections is mature technology, but a match factor is a similarity score, not an identity. An analyst may reject a high-scoring candidate because of a diagnostic ion, choose a lower-ranked candidate, manually separate co-eluting compounds, or leave a component unknown. Those decisions are exactly the information an automated system should preserve.
A useful analytical layer therefore stores more than the final compound name. It preserves the acquisition, peak and spectrum, candidate list and scores, retention evidence, analyst confirmations and rejections, manual interventions, and the processing and library versions that produced them. The analyst’s last click is one of the most valuable data points in the whole pipeline, and almost every lab throws it away.
What this unlocks
The immediate benefits are mundane and worth a lot of analyst time: find every analysis of a raw material over five years, compare an incoming supplier lot with its historical compositional envelope, investigate when an impurity first appeared, or find previous analyses of an unknown instead of starting again from zero. Questions that today require opening folders, workstation methods, PDFs, spreadsheets and old reports become queries. The larger prize is that the layer becomes trainable.
Confirmed and rejected identifications become supervision for models that rank candidates on future runs. Analyst corrections become examples of where automated deconvolution fails. Historical component profiles support anomaly detection on incoming lots. Retention measurements can help adapt general prediction models to the stationary phases and methods a laboratory actually uses.
The important distinction is that the model is no longer being trained to imitate a library match factor. It can learn from what experienced analysts actually decided, together with the evidence they used to make those decisions. And because every result is linked back to its acquisition and processing history, the system can answer not only what was identified, but why.
The compounding asset
Our bet is not that AI replaces the analyst. It is almost the opposite. The valuable dataset is created every time an analyst makes a difficult decision that existing software cannot make reliably. If that decision disappears into a PDF or remains trapped inside workstation software, the next analyst has to solve the same problem again. If the evidence, decision and context are captured together, the laboratory gets to learn from it. Once. Then again on the next run. And eventually automatically.
For flavor and fragrance companies, that analytical history can eventually be connected to another dataset they already generate: sensory observations. Linking composition with what perfumers and panels actually perceive does not make olfaction simple—mixture perception is highly nonlinear—but it creates the possibility of learning which compositional changes actually matter perceptually.
The companies that structure their analytical history now will be in a position to build compound-identification, quality, anomaly-detection and eventually sensory models on years of proprietary measurements and expert decisions. Their competitors may own the same instruments and subscribe to the same spectral libraries. They will not own the same dataset. The instruments were never the bottleneck; the layer between the instrument and the question always was.
References
Martens, L. et al. (2011). mzML—a community standard for mass spectrometry data. Molecular & Cellular Proteomics, 10(1), R110.000133. DOI: 10.1074/mcp.R110.000133.
Kessner, D.; Chambers, M.; Burke, R.; Agus, D.; Mallick, P. (2008). ProteoWizard: open source software for rapid proteomics tools development. Bioinformatics, 24(21), 2534–2536. DOI: 10.1093/bioinformatics/btn323.
Stein, S. E. (1999). An integrated method for spectrum extraction and compound identification from gas chromatography/mass spectrometry data. Journal of the American Society for Mass Spectrometry, 10(8), 770–781. DOI: 10.1016/S1044-0305(99)00047-1.
Stein, S. E.; Scott, D. R. (1994). Optimization and testing of mass spectral library search algorithms for compound identification. Journal of the American Society for Mass Spectrometry, 5(9), 859–866. DOI: 10.1016/1044-0305(94)87009-8.
Kováts, E. (1958). Gas-chromatographische Charakterisierung organischer Verbindungen. Teil 1: Retentionsindices aliphatischer Halogenide, Alkohole, Aldehyde und Ketone. Helvetica Chimica Acta, 41, 1915–1932. DOI: 10.1002/hlca.19580410703.
van den Dool, H.; Kratz, P. D. (1963). A generalization of the retention index system including linear temperature programmed gas–liquid partition chromatography. Journal of Chromatography, 11, 463–471. DOI: 10.1016/S0021-9673(01)80947-X.
Matyushin, D. D.; Sholokhova, A. Yu.; Buryak, A. K. (2021). Deep Learning Based Prediction of Gas Chromatographic Retention Indices for a Wide Variety of Polar and Mid-Polar Liquid Stationary Phases. International Journal of Molecular Sciences, 22(17), 9194. DOI: 10.3390/ijms22179194.
Geer, L. Y.; Stein, S. E.; Mallard, W. G.; Slotta, D. J. (2024). AIRI: Predicting Retention Indices and Their Uncertainties Using Artificial Intelligence. Journal of Chemical Information and Modeling, 64(3), 690–696. DOI: 10.1021/acs.jcim.3c01758.
Lee, B. K. et al. (2023). A Principal Odor Map Unifies Diverse Tasks in Olfactory Perception. Science, 381(6661), 999–1006. DOI: 10.1126/science.ade4401.