<?xml version='1.0' encoding='utf-8'?>
<?xml-stylesheet type="text/xsl" href="/v2/static/oai2.xsl"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-10-07T21:43:55Z</responseDate>
  <request identifier="oai:figshare.com:article/34038465" metadataPrefix="oai_dc" verb="GetRecord">https://api.figshare.com/v2/oai</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:figshare.com:article/34038465</identifier>
        <datestamp>2026-10-01T04:33:49Z</datestamp>
        <setSpec>category_323</setSpec>
        <setSpec>portal_316</setSpec>
        <setSpec>item_type_3</setSpec>
        <setSpec>month_year_10_2026</setSpec>
      </header>
      <metadata>
        <oai_dc:dc xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"  xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
          <dc:title>Table 1_How well do large language models answer postoperative lumbar fusion questions? A blinded comparative analysis of accuracy, quality, readability, and safety-related content.docx</dc:title>
          <dc:creator>Ahmet Kürşat Kara (25152201)</dc:creator>
          <dc:creator>Ozan Işık (25152204)</dc:creator>
          <dc:subject>Surgery</dc:subject>
          <dc:subject>large language models</dc:subject>
          <dc:subject>lumbar fusion</dc:subject>
          <dc:subject>patient safety</dc:subject>
          <dc:subject>postoperative patient education</dc:subject>
          <dc:subject>readability</dc:subject>
          <dc:description>Background&lt;p&gt;Patients recovering from lumbar fusion increasingly seek guidance from large language model (LLM) chatbots when their surgeon is unavailable, but the accuracy, completeness, and safety-related content of such responses have not been compared across models for the postoperative period.&lt;/p&gt;Objective&lt;p&gt;To compare the accuracy, quality, safety-related content, and readability of responses from ChatGPT, Gemini, and Claude to frequently asked postoperative lumbar fusion questions, and to identify the question types for which responses are weakest.&lt;/p&gt;Methods&lt;p&gt;Thirty questions, compiled by three neurosurgeons and grouped into six clinical categories, were submitted to each model, each with a standardized, patient-oriented prompt requesting warning signs. Ninety anonymized, randomized responses were rated independently by three blinded neurosurgeons for accuracy and four quality subdomains (clarity, completeness, relevance, and patient safety) on 4-point scales. Readability (Flesch–Kincaid, Flesch Reading Ease, and Gunning Fog), word count, and responses meeting an operational screening criterion for potential harm were also assessed. Models were compared with Friedman and Holm-corrected pairwise tests.&lt;/p&gt;Results&lt;p&gt;Accuracy was high for all models without significant differences. Quality, clarity, completeness, relevance, and patient safety differed globally, but effect sizes were small (Kendall's W: 0.10–0.21) and, after correction across the six outcome families, only quality and completeness remained significant, with Gemini scoring lower than ChatGPT and Claude. Gemini had the most favorable readability. ChatGPT responses were the shortest (mean 294 vs. 462 and 445 words); length was unrelated to completeness or readability. Two Gemini responses (6.7%) and none from ChatGPT or Claude met a post hoc operational screening criterion for potential harm (a score of ≤2 on accuracy or patient safety from at least one of three raters, without concurrence from the other two); no response was adjudicated as actually harmful. In exploratory question- and category-level analyses, responses tended to be rated highest for standardized questions and lowest for those requiring individualized judgment or interpretation of wound-related symptoms.&lt;/p&gt;Conclusion&lt;p&gt;Responses from contemporary LLMs to common postoperative lumbar fusion questions were generally rated as accurate, with small between-model differences and a model-level divergence between readability and completeness. Because safety-related content was explicitly prompted and no patient outcomes were assessed, these findings describe the quality of safety-related information rather than clinical safety. LLMs may support postoperative patient education but should not replace physician guidance, particularly for warning signs and patient-specific recovery decisions.&lt;/p&gt;</dc:description>
          <dc:date>2026-10-01T04:33:49Z</dc:date>
          <dc:type>Dataset</dc:type>
          <dc:type>Dataset</dc:type>
          <dc:identifier>10.3389/fsurg.2026.1932133.s002</dc:identifier>
          <dc:relation>https://figshare.com/articles/dataset/Table_1_How_well_do_large_language_models_answer_postoperative_lumbar_fusion_questions_A_blinded_comparative_analysis_of_accuracy_quality_readability_and_safety-related_content_docx/34038465</dc:relation>
          <dc:rights>CC BY 4.0</dc:rights>
        </oai_dc:dc>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
