<?xml version='1.0' encoding='utf-8'?>
<?xml-stylesheet type="text/xsl" href="/v2/static/oai2.xsl"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-10-10T11:51:54Z</responseDate>
  <request identifier="oai:figshare.com:article/34021752" metadataPrefix="oai_dc" verb="GetRecord">https://api.figshare.com/v2/oai</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:figshare.com:article/34021752</identifier>
        <datestamp>2026-09-29T07:05:32Z</datestamp>
        <setSpec>category_4</setSpec>
        <setSpec>category_8</setSpec>
        <setSpec>category_13</setSpec>
        <setSpec>category_21</setSpec>
        <setSpec>category_734</setSpec>
        <setSpec>category_931</setSpec>
        <setSpec>category_132</setSpec>
        <setSpec>category_133</setSpec>
        <setSpec>category_135</setSpec>
        <setSpec>portal_63</setSpec>
        <setSpec>item_type_3</setSpec>
        <setSpec>month_year_09_2026</setSpec>
      </header>
      <metadata>
        <oai_dc:dc xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"  xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
          <dc:title>Enzymatic Reaction
Feasibility Classification Using
Machine Learning Methods</dc:title>
          <dc:creator>Xin Wang (91924)</dc:creator>
          <dc:creator>Hongyan Yin (5066285)</dc:creator>
          <dc:creator>Yekai Shen (22351183)</dc:creator>
          <dc:creator>Yushan Zhu (809017)</dc:creator>
          <dc:creator>Igor V. Tetko (13525)</dc:creator>
          <dc:creator>Aixia Yan (1957624)</dc:creator>
          <dc:subject>Biochemistry</dc:subject>
          <dc:subject>Microbiology</dc:subject>
          <dc:subject>Genetics</dc:subject>
          <dc:subject>Biotechnology</dc:subject>
          <dc:subject>Biological Sciences not elsewhere classified</dc:subject>
          <dc:subject>Information Systems not elsewhere classified</dc:subject>
          <dc:subject>Infectious Diseases</dc:subject>
          <dc:subject>Plant Biology</dc:subject>
          <dc:subject>Computational  Biology</dc:subject>
          <dc:subject>stereo ” dataset</dc:subject>
          <dc:subject>step biosynthetic pathway</dc:subject>
          <dc:subject>previously reported deeprfc</dc:subject>
          <dc:subject>numerous biosynthetic pathways</dc:subject>
          <dc:subject>improved predictive performance</dc:subject>
          <dc:subject>identifying feasible reactions</dc:subject>
          <dc:subject>identify feasible reactions</dc:subject>
          <dc:subject>highest predictive performance</dc:subject>
          <dc:subject>extreme gradient boosting</dc:subject>
          <dc:subject>deep neural network</dc:subject>
          <dc:subject>864 feasible reactions</dc:subject>
          <dc:subject>optimal individual model</dc:subject>
          <dc:subject>“ stereo ”</dc:subject>
          <dc:subject>systematic reaction preprocessing</dc:subject>
          <dc:subject>infeasible reactions based</dc:subject>
          <dc:subject>individual model 1a</dc:subject>
          <dc:subject>c_ecfp4 representation achieved</dc:subject>
          <dc:subject>consensus model erfc</dc:subject>
          <dc:subject>achieving mcc values</dc:subject>
          <dc:subject>model 1a</dc:subject>
          <dc:subject>reaction representation</dc:subject>
          <dc:subject>“ non</dc:subject>
          <dc:subject>reaction rules</dc:subject>
          <dc:subject>reaction centers</dc:subject>
          <dc:subject>tuned chemberta</dc:subject>
          <dc:subject>three datasets</dc:subject>
          <dc:subject>test sets</dc:subject>
          <dc:subject>successfully validated</dc:subject>
          <dc:subject>stereochemical information</dc:subject>
          <dc:subject>results indicate</dc:subject>
          <dc:subject>openly available</dc:subject>
          <dc:subject>flexible framework</dc:subject>
          <dc:subject>experimental validation</dc:subject>
          <dc:subject>equal number</dc:subject>
          <dc:subject>effective classifiers</dc:subject>
          <dc:subject>drfp ),</dc:subject>
          <dc:subject>direct input</dc:subject>
          <dc:subject>collected 75</dc:subject>
          <dc:subject>biocatalysis databases</dc:subject>
          <dc:subject>aided retrobiosynthesis</dc:subject>
          <dc:subject>agnostic version</dc:subject>
          <dc:description>With the advancement of computer-aided retrobiosynthesis,
numerous
biosynthetic pathways have been predicted, exceeding the capacity
of experimental validation. Effective classifiers are needed to identify
feasible reactions. In this study, we collected 75,864 feasible reactions
from biocatalysis databases and generated an equal number of infeasible
reactions based on reaction rules. Following atom mapping focused
on reaction centers and systematic reaction preprocessing, three datasets
for training were constructed: the “stereo” dataset,
which retained reaction stereochemical information; the “non-stereo”
dataset, which was a stereochemistry-agnostic version of the “stereo”
dataset; and the “mixed” dataset, which comprised both.
We established a total of 22 individual enzymatic reaction feasibility
classification models, which include: eXtreme Gradient Boosting (XGBoost)
and Deep Neural Network (DNN) models utilizing Reaction Fingerprints
(RXNFP), Differential Reaction Fingerprint (DRFP), and our constructed
Combined ECFP4 Reaction Fingerprints (c_ECFP4) for reaction representation,
and Transformer models and fine-tuned ChemBERTa-77M-MLM (ChemMLM)
models using reaction SMILES strings as the direct input. The results
indicate that models utilizing the c_ECFP4 representation achieved
the highest predictive performance, which effectively captured underlying
enzymatic reaction mechanisms. Among them, Model 1A-M (based on XGBoost
and “mixed” dataset) was identified as the optimal individual
model, achieving Matthews Correlation Coefficient (MCC) values of
0.865 and 0.853 and Area Under Curve (AUC) values of 0.981 and 0.980
on the “stereo” and “non-stereo” test
sets, respectively. Furthermore, a consensus model enzymatic reaction
feasibility classification (ERFC) integrating four reaction representations
further improved predictive performance, achieving MCC values of 0.894
and 0.886 and an AUC of 0.986 on both test sets. Moreover, both models
(Model 1A-M and ERFC) successfully validated a five-step biosynthetic
pathway, demonstrating higher prediction accuracy than the previously
reported DeepRFC and DORA-XGB models in identifying feasible reactions.
All data, the individual Model 1A-M, and the consensus model ERFC
are openly available, offering a reliable and flexible framework for
predicting enzymatic reaction feasibility in the presence or absence
of stereochemical information.</dc:description>
          <dc:date>2026-09-29T00:00:00Z</dc:date>
          <dc:type>Dataset</dc:type>
          <dc:type>Dataset</dc:type>
          <dc:identifier>10.1021/acs.jcim.6c02294.s005</dc:identifier>
          <dc:relation>https://figshare.com/articles/dataset/Enzymatic_Reaction_Feasibility_Classification_Using_Machine_Learning_Methods/34021752</dc:relation>
          <dc:rights>CC BY-NC 4.0</dc:rights>
        </oai_dc:dc>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
