EconPapers    
Economics at your fingertips  
 

State-of-the-art augmented NLP transformer models for direct and single-step retrosynthesis

Igor V. Tetko (), Pavel Karpov, Ruud Deursen and Guillaume Godin ()
Additional contact information
Igor V. Tetko: Institute of Structural Biology, Helmholtz Zentrum München—Research Center for Environmental Health (GmbH)
Pavel Karpov: Institute of Structural Biology, Helmholtz Zentrum München—Research Center for Environmental Health (GmbH)
Ruud Deursen: Firmenich International SA, D-Lab by Firmenich
Guillaume Godin: Firmenich International SA, D-Lab by Firmenich

Nature Communications, 2020, vol. 11, issue 1, 1-11

Abstract: Abstract We investigated the effect of different training scenarios on predicting the (retro)synthesis of chemical compounds using text-like representation of chemical reactions (SMILES) and Natural Language Processing (NLP) neural network Transformer architecture. We showed that data augmentation, which is a powerful method used in image processing, eliminated the effect of data memorization by neural networks and improved their performance for prediction of new sequences. This effect was observed when augmentation was used simultaneously for input and the target data simultaneously. The top-5 accuracy was 84.8% for the prediction of the largest fragment (thus identifying principal transformation for classical retro-synthesis) for the USPTO-50k test dataset, and was achieved by a combination of SMILES augmentation and a beam search algorithm. The same approach provided significantly better results for the prediction of direct reactions from the single-step USPTO-MIT test set. Our model achieved 90.6% top-1 and 96.1% top-5 accuracy for its challenging mixed set and 97% top-5 accuracy for the USPTO-MIT separated set. It also significantly improved results for USPTO-full set single-step retrosynthesis for both top-1 and top-10 accuracies. The appearance frequency of the most abundantly generated SMILES was well correlated with the prediction outcome and can be used as a measure of the quality of reaction prediction.

Date: 2020
References: Add references at CitEc
Citations: View citations in EconPapers (8)

Downloads: (external link)
https://www.nature.com/articles/s41467-020-19266-y Abstract (text/html)

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: https://EconPapers.repec.org/RePEc:nat:natcom:v:11:y:2020:i:1:d:10.1038_s41467-020-19266-y

Ordering information: This journal article can be ordered from
https://www.nature.com/ncomms/

DOI: 10.1038/s41467-020-19266-y

Access Statistics for this article

Nature Communications is currently edited by Nathalie Le Bot, Enda Bergin and Fiona Gillespie

More articles in Nature Communications from Nature
Bibliographic data for series maintained by Sonal Shukla () and Springer Nature Abstracting and Indexing ().

 
Page updated 2025-03-19
Handle: RePEc:nat:natcom:v:11:y:2020:i:1:d:10.1038_s41467-020-19266-y