<?xml version='1.0' encoding='utf-8'?>
<?xml-stylesheet type="text/xsl" href="/v2/static/oai2.xsl"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-10-09T11:39:30Z</responseDate>
  <request identifier="oai:figshare.com:article/34039200" metadataPrefix="oai_dc" verb="GetRecord">https://api.figshare.com/v2/oai</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:figshare.com:article/34039200</identifier>
        <datestamp>2026-10-01T05:30:48Z</datestamp>
        <setSpec>category_15</setSpec>
        <setSpec>portal_316</setSpec>
        <setSpec>item_type_3</setSpec>
        <setSpec>month_year_10_2026</setSpec>
      </header>
      <metadata>
        <oai_dc:dc xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"  xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
          <dc:title>Data Sheet 1_MEMOIR-VLM—a multimodal vision-language model for Alzheimer's disease classification and question answering.pdf</dc:title>
          <dc:creator>Sohail Haresh Gidwani (25152636)</dc:creator>
          <dc:creator>Tamoghna Chattopadhyay (18948082)</dc:creator>
          <dc:creator>Sophia I. Thomopoulos (6376262)</dc:creator>
          <dc:creator>Paul M. Thompson (135650)</dc:creator>
          <dc:subject>Neuroscience</dc:subject>
          <dc:subject>Alzheimer's disease</dc:subject>
          <dc:subject>deep learning</dc:subject>
          <dc:subject>missing data</dc:subject>
          <dc:subject>vision-language model</dc:subject>
          <dc:subject>retrieval-augmented generation</dc:subject>
          <dc:subject>visual question answering</dc:subject>
          <dc:subject>diffusion tensor imaging</dc:subject>
          <dc:subject>large language model</dc:subject>
          <dc:description>Introduction&lt;p&gt;Alzheimer's disease (AD) affects an estimated 55 million people worldwide and is projected to nearly double by 2050, creating a need for scalable, non-invasive AI systems that can classify disease stage, support multimodal clinical reasoning, and interact in natural language. Most deep learning work on AD focuses on single-modality classification without natural-language interaction and typically assumes complete input data.&lt;/p&gt;Methods&lt;p&gt;We present MEMOIR-VLM (Multimodal Encoder with Missing-modality and Open-ended Inference for Retrieval), a modular two-stage vision-language framework trained and evaluated on 2,363 ADNI subjects. The first stage is a missing-modality-aware encoder that fuses T1-weighted MRI and DTI fractional anisotropy maps with structured clinical scores via cross-attention fusion with stochastic modality dropout, allowing a single set of weights to operate on any of the seven non-empty modality subsets without imputation. After CLIP/InfoNCE contrastive pretraining, the encoder is fine-tuned for multi-task prediction of 3-way diagnosis, binary diagnosis, CDR-SB, age, and sex. The second stage extends the frozen encoder into a retrieval-augmented generation pipeline for case-based reasoning and open-ended visual question answering.&lt;/p&gt;Results&lt;p&gt;On held-out ADNI test data, with CDR-SB excluded from the inputs to remove a cognitive-score leak, the encoder achieved 91.3% balanced accuracy for CN vs. dementia and 68.2% balanced accuracy for 3-class diagnosis (macro-F1 = 0.681); age was predicted with a 5.96-year MAE. When the diagnosis field was withheld from retrieved captions, no evaluated LLM exceeded a retrieval-only k-NN majority vote (0.673), which itself closely matched the encoder's 3-class accuracy (0.682). The RAG-VQA layer showed strong text fidelity (BERTScore = 0.894; SBERT cosine similarity = 0.811). Zero-shot external validation on OASIS-3 (1,048 subjects, no diffusion data) retained 78.7% balanced accuracy for CN vs. impaired (AUC = 0.889) using identical ADNI-trained weights with the DTI branch masked.&lt;/p&gt;Discussion&lt;p&gt;MEMOIR-VLM combines missing-modality-aware multimodal encoding, multi-task clinical prediction, and retrieval-grounded natural-language reasoning within a single modular framework. The results support using the encoder, optionally with k-NN voting, as the diagnostic component and the RAG-VQA layer as an interpretable natural-language interface that surfaces comparable cases and readable clinical summaries. The external validation results further demonstrate the practical value of the missing-modality mechanism when an imaging modality is unavailable.&lt;/p&gt;</dc:description>
          <dc:date>2026-10-01T05:30:48Z</dc:date>
          <dc:type>Dataset</dc:type>
          <dc:type>Dataset</dc:type>
          <dc:identifier>10.3389/fncom.2026.1902258.s001</dc:identifier>
          <dc:relation>https://figshare.com/articles/dataset/Data_Sheet_1_MEMOIR-VLM_a_multimodal_vision-language_model_for_Alzheimer_s_disease_classification_and_question_answering_pdf/34039200</dc:relation>
          <dc:rights>CC BY 4.0</dc:rights>
        </oai_dc:dc>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
