<?xml version='1.0' encoding='utf-8'?>
<?xml-stylesheet type="text/xsl" href="/v2/static/oai2.xsl"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-10-06T13:04:49Z</responseDate>
  <request identifier="oai:figshare.com:article/34025802" metadataPrefix="oai_dc" verb="GetRecord">https://api.figshare.com/v2/oai</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:figshare.com:article/34025802</identifier>
        <datestamp>2026-09-29T17:25:38Z</datestamp>
        <setSpec>category_7</setSpec>
        <setSpec>category_734</setSpec>
        <setSpec>category_931</setSpec>
        <setSpec>portal_5</setSpec>
        <setSpec>item_type_3</setSpec>
        <setSpec>month_year_09_2026</setSpec>
      </header>
      <metadata>
        <oai_dc:dc xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"  xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
          <dc:title>&lt;p&gt;S2 Table; Figs S1 and S2; Tables S2-S11.&lt;/p&gt;</dc:title>
          <dc:creator>Kurban Kotan (25138206)</dc:creator>
          <dc:creator>Serdar Kırışoğlu (25138209)</dc:creator>
          <dc:subject>Medicine</dc:subject>
          <dc:subject>Biological Sciences not elsewhere classified</dc:subject>
          <dc:subject>Information Systems not elsewhere classified</dc:subject>
          <dc:subject>wilson score method</dc:subject>
          <dc:subject>present work addresses</dc:subject>
          <dc:subject>health informatics pipelines</dc:subject>
          <dc:subject>gaussian naïve bayes</dc:subject>
          <dc:subject>full mcnemar results</dc:subject>
          <dc:subject>driven missingness patterns</dc:subject>
          <dc:subject>competing methods degrade</dc:subject>
          <dc:subject>chronic kidney disease</dc:subject>
          <dc:subject>38 percentage points</dc:subject>
          <dc:subject>configuration mean accuracy</dc:subject>
          <dc:subject>classifier accuracy differences</dc:subject>
          <dc:subject>accuracy remains stable</dc:subject>
          <dc:subject>regression training set</dc:subject>
          <dc:subject>appropriate imputation strategy</dc:subject>
          <dc:subject>results provide mechanism</dc:subject>
          <dc:subject>perfect downstream classification</dc:subject>
          <dc:subject>heart disease dataset</dc:subject>
          <dc:subject>div &gt;&lt; p</dc:subject>
          <dc:subject>reported accuracy values</dc:subject>
          <dc:subject>chit achieves near</dc:subject>
          <dc:subject>missing data mechanism</dc:subject>
          <dc:subject>mcar ), missing</dc:subject>
          <dc:subject>logistic regression</dc:subject>
          <dc:subject>multiple imputation</dc:subject>
          <dc:subject>imputation selection</dc:subject>
          <dc:subject>hdd ),</dc:subject>
          <dc:subject>first mechanism</dc:subject>
          <dc:subject>taken together</dc:subject>
          <dc:subject>stratified comparison</dc:subject>
          <dc:subject>statistical significance</dc:subject>
          <dc:subject>standard deviation</dc:subject>
          <dc:subject>specific guidance</dc:subject>
          <dc:subject>scarce features</dc:subject>
          <dc:subject>record prioritisation</dc:subject>
          <dc:subject>overall rates</dc:subject>
          <dc:subject>mped ).</dc:subject>
          <dc:subject>iterative enrichment</dc:subject>
          <dc:subject>fig 2</dc:subject>
          <dc:subject>experimental platforms</dc:subject>
          <dc:subject>decision tree</dc:subject>
          <dc:subject>corrected chi</dc:subject>
          <dc:subject>confidence intervals</dc:subject>
          <dc:subject>computed using</dc:subject>
          <dc:subject>based filling</dc:subject>
          <dc:subject>96 ].</dc:subject>
          <dc:subject>80 ].</dc:subject>
          <dc:description>&lt;div&gt;&lt;p&gt;Choosing an appropriate imputation strategy for clinical datasets requires understanding which missing data mechanism is operative, yet most imputation benchmarks conflate performance across mechanisms. The present work addresses this gap by conducting the first mechanism-stratified comparison of three imputation approaches—CHIT, Multiple Imputation by Chained Equations via BayesianRidge (MICE), and Random-Forest-based Iterative Imputation—across Missing Completely at Random (MCAR), Missing at Random (MAR), and Missing Not at Random (MNAR) conditions. Three health datasets serve as experimental platforms: the Chronic Kidney Disease (CKD) dataset, the Heart Disease Dataset (HDD), and the Mice Protein Expression Dataset (MPED). Domain-knowledge-driven missingness patterns are constructed for each mechanism (Fig 2) at approximately 20–25% overall rates. Eight classifiers—KNN, Logistic Regression, SVC, Decision Tree, Random Forest, Gaussian Naïve Bayes, MLP, and a deep neural network—are trained on imputed data following GridSearchCV optimisation. Across all conditions, CHIT achieves near-perfect or perfect downstream classification, with SVC reaching up to 100% accuracy on the CKD dataset under MNAR, and 98.75% under MCAR and MAR. Competing methods degrade by up to 19.38 percentage points under MNAR, while CHIT’s accuracy remains stable. Two structural properties account for this resilience: an iterative enrichment of the regression training set as records are completed, and a within-record prioritisation of the most data-scarce features for model-based filling. Taken together, these results provide mechanism-specific guidance for imputation selection in health informatics pipelines. Statistical significance of classifier accuracy differences between CHIT and MICE was assessed using McNemar’s test (two-tailed, continuity-corrected chi-square) applied to the exact binary prediction vectors from each experiment [n_test = 80]. Ninety-five percent confidence intervals for all reported accuracy values were computed using the Wilson score method [z = 1.96]. The 42-configuration mean accuracy and standard deviation for CHIT are reported in Supplementary &lt;a href="http://www.plosone.org/article/info:doi/10.1371/journal.pone.0359577#pone.0359577.s001" target="_blank"&gt;Table S1&lt;/a&gt;. Full McNemar results with confidence intervals are in Supplementary &lt;a href="http://www.plosone.org/article/info:doi/10.1371/journal.pone.0359577#pone.0359577.s002" target="_blank"&gt;Table S2&lt;/a&gt;.&lt;/p&gt;&lt;/div&gt;</dc:description>
          <dc:date>2026-09-29T17:25:35Z</dc:date>
          <dc:type>Dataset</dc:type>
          <dc:type>Dataset</dc:type>
          <dc:identifier>10.1371/journal.pone.0359577.s002</dc:identifier>
          <dc:relation>https://figshare.com/articles/dataset/_p_S2_Table_Figs_S1_and_S2_Tables_S2-S11_p_/34025802</dc:relation>
          <dc:rights>CC BY 4.0</dc:rights>
        </oai_dc:dc>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
