<?xml version='1.0' encoding='utf-8'?>
<?xml-stylesheet type="text/xsl" href="/v2/static/oai2.xsl"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-10-10T11:31:29Z</responseDate>
  <request identifier="oai:figshare.com:article/34028306" metadataPrefix="oai_dc" verb="GetRecord">https://api.figshare.com/v2/oai</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:figshare.com:article/34028306</identifier>
        <datestamp>2026-09-30T00:06:38Z</datestamp>
        <setSpec>category_13</setSpec>
        <setSpec>category_873</setSpec>
        <setSpec>category_734</setSpec>
        <setSpec>category_931</setSpec>
        <setSpec>category_69</setSpec>
        <setSpec>category_106</setSpec>
        <setSpec>portal_63</setSpec>
        <setSpec>item_type_3</setSpec>
        <setSpec>month_year_09_2026</setSpec>
      </header>
      <metadata>
        <oai_dc:dc xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"  xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
          <dc:title>Language of Toxicity:
An eXplainable Artificial Intelligence
Approach</dc:title>
          <dc:creator>Nicola Amoroso (4591525)</dc:creator>
          <dc:creator>Ester Pantaleo (12400655)</dc:creator>
          <dc:creator>Fulvio Ciriaco (5569355)</dc:creator>
          <dc:creator>Francesca Cutropia (24130888)</dc:creator>
          <dc:creator>Nicola Gambacorta (9514363)</dc:creator>
          <dc:creator>Fabrizio Mastrolorito (14264646)</dc:creator>
          <dc:creator>Alfonso Monaco (4739133)</dc:creator>
          <dc:creator>Angelica Orfino (17699096)</dc:creator>
          <dc:creator>Roberto Bellotti (4591528)</dc:creator>
          <dc:creator>Orazio Nicolotti (168458)</dc:creator>
          <dc:subject>Genetics</dc:subject>
          <dc:subject>Chemical Sciences not elsewhere classified</dc:subject>
          <dc:subject>Biological Sciences not elsewhere classified</dc:subject>
          <dc:subject>Information Systems not elsewhere classified</dc:subject>
          <dc:subject>Inorganic Chemistry</dc:subject>
          <dc:subject>Science Policy</dc:subject>
          <dc:subject>small molecules represents</dc:subject>
          <dc:subject>relevant molecular substructures</dc:subject>
          <dc:subject>predefined molecular descriptors</dc:subject>
          <dc:subject>often strongly imbalanced</dc:subject>
          <dc:subject>gated recurrent unit</dc:subject>
          <dc:subject>chemical safety assessment</dc:subject>
          <dc:subject>two important challenges</dc:subject>
          <dc:subject>based framework offers</dc:subject>
          <dc:subject>architecture also exploits</dc:subject>
          <dc:subject>drug repurposing applications</dc:subject>
          <dc:subject>related structural patterns</dc:subject>
          <dc:subject>model proposed combines</dc:subject>
          <dc:subject>capture chemical patterns</dc:subject>
          <dc:subject>unified architecture</dc:subject>
          <dc:subject>two languages</dc:subject>
          <dc:subject>proposed attention</dc:subject>
          <dc:subject>inspired framework</dc:subject>
          <dc:subject>drug development</dc:subject>
          <dc:subject>drug design</dc:subject>
          <dc:subject>sequential nature</dc:subject>
          <dc:subject>roc curve</dc:subject>
          <dc:subject>repeated cross</dc:subject>
          <dc:subject>potentially limiting</dc:subject>
          <dc:subject>nontoxic chemicals</dc:subject>
          <dc:subject>model provided</dc:subject>
          <dc:subject>future investigations</dc:subject>
          <dc:subject>future improvement</dc:subject>
          <dc:subject>fundamental challenge</dc:subject>
          <dc:subject>different scales</dc:subject>
          <dc:subject>capture complex</dc:subject>
          <dc:subject>canonical smiles</dc:subject>
          <dc:subject>average area</dc:subject>
          <dc:subject>attention mechanism</dc:subject>
          <dc:subject>accurate description</dc:subject>
          <dc:description>Toxicity prediction
in small molecules represents a fundamental
challenge in drug development and chemical safety assessment. Traditional
approaches heavily rely on predefined molecular descriptors or fingerprints,
potentially limiting the ability to capture complex and nonlinear
structure–activity relationships. Here, we present a descriptor-free,
language-inspired framework that can be applied to different toxicity
prediction tasks within a unified architecture. The model proposed
combines a multiscale Convolutional Neural Network (CNN) layer to
capture chemical patterns at different scales and a Gated Recurrent
Unit (GRU) layer to capture the sequential nature of these patterns.
This architecture also exploits an attention mechanism that computes
attention weights across the sequence, enabling the model to focus
on the most relevant molecular substructures for toxicity prediction.
Toxic and nontoxic chemicals, represented by canonical SMILES, are
investigated as the words of two languages which have to be discriminated;
using eight different end points, the model provided an accurate description
of toxicity patterns, with an average Area under the ROC curve (AUC)
of 0.83 (min: 0.70, max: 0.94) under repeated cross-validation. The
models were trained on relatively small data sets (∼1000 samples)
and often strongly imbalanced, two important challenges that highlight
the opportunities for future improvement; moreover, the proposed attention-based
framework offers a representation of the molecular regions influencing
model predictions, providing a basis for future investigations into
toxicity-related structural patterns and potentially supporting hypothesis
generation in drug design or drug repurposing applications.</dc:description>
          <dc:date>2026-09-29T00:00:00Z</dc:date>
          <dc:type>Dataset</dc:type>
          <dc:type>Dataset</dc:type>
          <dc:identifier>10.1021/acs.jcim.5c03214.s003</dc:identifier>
          <dc:relation>https://figshare.com/articles/dataset/Language_of_Toxicity_An_eXplainable_Artificial_Intelligence_Approach/34028306</dc:relation>
          <dc:rights>CC BY-NC 4.0</dc:rights>
        </oai_dc:dc>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
