<?xml version='1.0' encoding='utf-8'?>
<?xml-stylesheet type="text/xsl" href="/v2/static/oai2.xsl"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-10-10T06:15:12Z</responseDate>
  <request identifier="oai:figshare.com:article/33966917" metadataPrefix="oai_dc" verb="GetRecord">https://api.figshare.com/v2/oai</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:figshare.com:article/33966917</identifier>
        <datestamp>2026-09-22T17:25:25Z</datestamp>
        <setSpec>category_21</setSpec>
        <setSpec>category_46</setSpec>
        <setSpec>category_931</setSpec>
        <setSpec>category_64</setSpec>
        <setSpec>category_106</setSpec>
        <setSpec>category_122</setSpec>
        <setSpec>portal_5</setSpec>
        <setSpec>item_type_3</setSpec>
        <setSpec>month_year_09_2026</setSpec>
      </header>
      <metadata>
        <oai_dc:dc xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"  xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
          <dc:title>&lt;p&gt;Detailed study characteristics and outcomes.&lt;/p&gt;</dc:title>
          <dc:creator>Rocco Sheldon (24549748)</dc:creator>
          <dc:creator>Morya Wadodkar (24549751)</dc:creator>
          <dc:creator>Andrew Gan (24549754)</dc:creator>
          <dc:creator>Rachael Yip (25091168)</dc:creator>
          <dc:creator>Yimeng Zhang (283704)</dc:creator>
          <dc:creator>Jyoti Baharani (3584696)</dc:creator>
          <dc:subject>Biotechnology</dc:subject>
          <dc:subject>Immunology</dc:subject>
          <dc:subject>Information Systems not elsewhere classified</dc:subject>
          <dc:subject>Cancer</dc:subject>
          <dc:subject>Science Policy</dc:subject>
          <dc:subject>Mental Health</dc:subject>
          <dc:subject>small sample sizes</dc:subject>
          <dc:subject>primary outcomes due</dc:subject>
          <dc:subject>one study reported</dc:subject>
          <dc:subject>large language models</dc:subject>
          <dc:subject>information fidelity ranging</dc:subject>
          <dc:subject>grade appraisal tool</dc:subject>
          <dc:subject>effectiveness remains fragmented</dc:subject>
          <dc:subject>drafting time compared</dc:subject>
          <dc:subject>1 november 2025</dc:subject>
          <dc:subject>findings regarding readability</dc:subject>
          <dc:subject>tactless patient descriptions</dc:subject>
          <dc:subject>authored control group</dc:subject>
          <dc:subject>outpatient clinic letters</dc:subject>
          <dc:subject>five using synthetic</dc:subject>
          <dc:subject>concerns regarding accuracy</dc:subject>
          <dc:subject>world clinic data</dc:subject>
          <dc:subject>providing insufficient evidence</dc:subject>
          <dc:subject>assessed quality using</dc:subject>
          <dc:subject>limiting patient comprehension</dc:subject>
          <dc:subject>studies evaluating ai</dc:subject>
          <dc:subject>world patient</dc:subject>
          <dc:subject>authored letters</dc:subject>
          <dc:subject>evidence regarding</dc:subject>
          <dc:subject>certainty using</dc:subject>
          <dc:subject>measured patient</dc:subject>
          <dc:subject>five databases</dc:subject>
          <dc:subject>extracted data</dc:subject>
          <dc:subject>clinical accuracy</dc:subject>
          <dc:subject>simplified letters</dc:subject>
          <dc:subject>two studies</dc:subject>
          <dc:subject>studies relied</dc:subject>
          <dc:subject>seven studies</dc:subject>
          <dc:subject>widespread adoption</dc:subject>
          <dc:subject>systematic review</dc:subject>
          <dc:subject>subjective measures</dc:subject>
          <dc:subject>scalable tools</dc:subject>
          <dc:subject>prospero registration</dc:subject>
          <dc:subject>overall certainty</dc:subject>
          <dc:subject>mmat ),</dc:subject>
          <dc:subject>methodological limitations</dc:subject>
          <dc:subject>low across</dc:subject>
          <dc:subject>improvement settings</dc:subject>
          <dc:subject>hypothetical scenarios</dc:subject>
          <dc:subject>generalisability remain</dc:subject>
          <dc:subject>fold reduction</dc:subject>
          <dc:subject>clinical communication</dc:subject>
          <dc:subject>centred research</dc:subject>
          <dc:subject>artificial intelligence</dc:subject>
          <dc:description>&lt;div&gt;&lt;p&gt;Outpatient clinic letters are a cornerstone of clinical communication but frequently exceed recommended reading levels, limiting patient comprehension and engagement. Large Language Models (LLMs) and Artificial Intelligence (AI) have been proposed as scalable tools for generating and simplifying these documents, but the evidence regarding their clinical accuracy, safety, and effectiveness remains fragmented. We conducted a systematic review of studies evaluating AI or LLMs for generating, simplifying, or enhancing outpatient clinic letters. Five databases (PubMed, EMBASE, Web of Science, CENTRAL, and CINAHL) were searched from inception to 1 November 2025. Two independent reviewers screened studies, extracted data, assessed quality using the Mixed Methods Appraisal Tool (MMAT), and certainty using the GRADE appraisal tool. Seven studies were included, comprising two studies using real-world clinic data and five using synthetic or hypothetical scenarios. Findings regarding readability were mixed; while some AI models improved readability scores compared to human-authored letters, most AI-generated content still failed to meet the recommended US Grade 6 reading level. The two studies that measured patient-reported outcomes reported high satisfaction and comprehension with AI-simplified letters, but these studies relied on subjective measures of understanding or lacked a human-authored control group. Clinical accuracy varied substantially, with information fidelity ranging from 10% to 100%. Risks of ‘hallucinations’, inappropriate medical advice, and tactless patient descriptions were documented. One study reported a ten-fold reduction in drafting time compared to human dictation. Overall certainty of evidence was rated very low across all primary outcomes due to methodological limitations, small sample sizes, and reliance on hypothetical scenarios. The current evidence supporting AI and LLMs use for outpatient clinic letters is preliminary and limited, providing insufficient evidence to support routine clinical implementation beyond supervised experimental and quality-improvement settings. Concerns regarding accuracy, safety, and generalisability remain. Further real-world patient-centred research is required before widespread adoption. PROSPERO registration: CRD420251181303.&lt;/p&gt;&lt;/div&gt;</dc:description>
          <dc:date>2026-09-22T17:25:16Z</dc:date>
          <dc:type>Dataset</dc:type>
          <dc:type>Dataset</dc:type>
          <dc:identifier>10.1371/journal.pdig.0001745.t001</dc:identifier>
          <dc:relation>https://figshare.com/articles/dataset/_p_Detailed_study_characteristics_and_outcomes_p_/33966917</dc:relation>
          <dc:rights>CC BY 4.0</dc:rights>
        </oai_dc:dc>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
