<?xml version='1.0' encoding='utf-8'?>
<?xml-stylesheet type="text/xsl" href="/v2/static/oai2.xsl"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-10-11T17:31:15Z</responseDate>
  <request identifier="oai:figshare.com:article/33962116" metadataPrefix="oai_dc" verb="GetRecord">https://api.figshare.com/v2/oai</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:figshare.com:article/33962116</identifier>
        <datestamp>2026-09-22T05:38:31Z</datestamp>
        <setSpec>category_450</setSpec>
        <setSpec>portal_316</setSpec>
        <setSpec>item_type_3</setSpec>
        <setSpec>month_year_09_2026</setSpec>
      </header>
      <metadata>
        <oai_dc:dc xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"  xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
          <dc:title>Supplementary file 1_Feedback-driven rule induction and retrieval-augmented bias mitigation for large language models.pdf</dc:title>
          <dc:creator>Abhishek Kumar (156479)</dc:creator>
          <dc:creator>Aishwaryaa Shree Muralitharan (25087876)</dc:creator>
          <dc:creator>Vinesh Kannaa Balaji (25087879)</dc:creator>
          <dc:creator>Bhargavi Renta Chintala (25087882)</dc:creator>
          <dc:subject>Knowledge Representation and Machine Learning</dc:subject>
          <dc:subject>external alignment memory</dc:subject>
          <dc:subject>feedback-guided rule induction</dc:subject>
          <dc:subject>gender bias mitigation</dc:subject>
          <dc:subject>large language models</dc:subject>
          <dc:subject>post-hoc alignment</dc:subject>
          <dc:subject>responsible AI</dc:subject>
          <dc:description>Introduction&lt;p&gt;Large Language Models (LLMs) are becoming widely adopted for reasoning and decision-support tasks, yet they can inherit and reproduce gender-related biases present in their training data. Most existing bias-mitigation strategies depend on fine-tuning, reinforcement learning, prompt engineering, or interventions during pre-training, requiring either parameter updates or extensive manual prompt design. These requirements limit their applicability when the underlying model is available only as a closed-source or black-box system.&lt;/p&gt;Methods&lt;p&gt;This work explores whether bias-mitigation knowledge can instead be separated from the model and reused without altering its parameters. We propose a feedback-driven external alignment memory framework for post-hoc gender bias mitigation that transforms identified bias into concise corrective rules through an iterative feedback process. These rules are stored independently of the model parameters and are retrieved during inference to influence future responses. Retrieval is performed using cosine-similarity matching between an incoming query and previously stored examples, allowing corrective knowledge to be reused while maintaining complete model independence.&lt;/p&gt;Results&lt;p&gt;Evaluation on selected subsets of BBQ, BiasNLI, CoBias, CrowS-Pairs, and WinoBias shows measurable reductions in gender-related bias. In the 120-record evaluation, Gender Assumption (GA) decreases from 15.83 to 7.08%, while Stereotypical Gender Assumption (SGA) decreases from 24.16 to 7.08%. Gender Neutral responses increase from 75.00 to 90.415%, and response quality improves from 4.15 to 4.211. A larger 500-record evaluation further indicates that corrective rules learned earlier continue to provide benefits when applied beyond the original feedback corpus, reducing GA from 13.8 to 12.0% and SGA from 14.4 to 11.0%. During the same evaluation, Gender Neutral responses increase from 78.6 to 83.2%, while response quality improves from 4.00 to 4.05.&lt;/p&gt;Discussion&lt;p&gt;These results indicate that corrective alignment knowledge can be maintained as a reusable external memory and incorporated during inference to improve fairness in black-box large language models without retraining or modifying model parameters.&lt;/p&gt;</dc:description>
          <dc:date>2026-09-22T05:38:31Z</dc:date>
          <dc:type>Dataset</dc:type>
          <dc:type>Dataset</dc:type>
          <dc:identifier>10.3389/frai.2026.1925701.s001</dc:identifier>
          <dc:relation>https://figshare.com/articles/dataset/Supplementary_file_1_Feedback-driven_rule_induction_and_retrieval-augmented_bias_mitigation_for_large_language_models_pdf/33962116</dc:relation>
          <dc:rights>CC BY 4.0</dc:rights>
        </oai_dc:dc>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
