<?xml version='1.0' encoding='utf-8'?>
<?xml-stylesheet type="text/xsl" href="/v2/static/oai2.xsl"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-10-11T09:55:10Z</responseDate>
  <request identifier="oai:figshare.com:article/33970918" metadataPrefix="oai_dc" verb="GetRecord">https://api.figshare.com/v2/oai</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:figshare.com:article/33970918</identifier>
        <datestamp>2026-09-23T05:44:27Z</datestamp>
        <setSpec>category_388</setSpec>
        <setSpec>portal_316</setSpec>
        <setSpec>item_type_3</setSpec>
        <setSpec>month_year_09_2026</setSpec>
      </header>
      <metadata>
        <oai_dc:dc xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"  xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
          <dc:title>Table 1_Benchmarking large language models on a Chinese radiation oncology technology examination-preparation question set: accuracy, consensus, and efficiency for AI assisted education.docx</dc:title>
          <dc:creator>Heling Zhu (25094323)</dc:creator>
          <dc:creator>Yongguang Liang (16762215)</dc:creator>
          <dc:creator>Xinyu Long (6028754)</dc:creator>
          <dc:creator>Yanchen Liu (3691888)</dc:creator>
          <dc:creator>Tingtian Pang (25094326)</dc:creator>
          <dc:creator>Wenbo Li (669721)</dc:creator>
          <dc:creator>Bo Yang (104813)</dc:creator>
          <dc:creator>Lin Yao (332529)</dc:creator>
          <dc:subject>Oncology and Carcinogenesis not elsewhere classified</dc:subject>
          <dc:subject>Chinese national examination</dc:subject>
          <dc:subject>examination preparation</dc:subject>
          <dc:subject>large language models (LLM)</dc:subject>
          <dc:subject>medical education</dc:subject>
          <dc:subject>radiation oncology technology</dc:subject>
          <dc:description>Purpose&lt;p&gt;This study evaluated GPT-4o, GPT-5.4, and two DeepSeek platform configurations using 1, 053 examination-preparation questions from a publicly and commercially available 2025 exercise collection for the Chinese National Radiation Oncology Technology Qualification Examination (Intermediate Level). We assessed accuracy rates, identical-response coverage, accuracy consensus, and model-reported latency estimates.&lt;/p&gt;Methods&lt;p&gt;Four configurations—GPT-4o, GPT-5.4, DS-Fast, and DS-Expert—were tested using an identical zero-shot Chinese prompt. Accuracy rates were summarized with Wilson 95% confidence intervals. Pairwise differences in the primary Chinese-language analysis were assessed using exact McNemar’s tests with Holm adjustment. A descriptive language-controlled analysis evaluated English translations of the same questions using an equivalent English prompt.&lt;/p&gt;Results&lt;p&gt;Overall accuracy rates were 56.7% (95% CI, 53.7–59.7) for GPT-4o, 66.1% (63.2–68.9) for GPT-5.4, 98.4% (97.4–99.0) for DS-Fast, and 99.8% (99.3–99.9) for DS-Expert. All six overall pairwise comparisons remained significant after Holm adjustment. DS-Fast and DS-Expert produced identical responses for 1, 034 of 1, 053 questions (98.2%), all of which were correct. GPT-5.4 and DS-Expert showed 100% conditional accuracy among shared responses, but their identical-response coverage was only 65.9% (694/1, 053). Model-reported latency estimates were lowest for DS-Fast (0.68 ± 1.00 seconds) and highest for DS-Expert (3.69 ± 2.24 seconds). Under the English-translated condition, overall accuracy rates were 62.6%, 61.5%, 66.5%, and 76.9%, respectively. DS-Expert remained the most accurate configuration, although its advantage over the GPT models was reduced.&lt;/p&gt;Conclusion&lt;p&gt;DeepSeek configurations outperformed the GPT models on this Chinese-language examination-preparation question set, with DS-Expert achieving the highest accuracy. However, performance varied substantially by question language. These findings support further evaluation for answer verification and practice-question review, but they do not establish educational effectiveness.&lt;/p&gt;</dc:description>
          <dc:date>2026-09-23T05:44:27Z</dc:date>
          <dc:type>Dataset</dc:type>
          <dc:type>Dataset</dc:type>
          <dc:identifier>10.3389/fonc.2026.1946351.s001</dc:identifier>
          <dc:relation>https://figshare.com/articles/dataset/Table_1_Benchmarking_large_language_models_on_a_Chinese_radiation_oncology_technology_examination-preparation_question_set_accuracy_consensus_and_efficiency_for_AI_assisted_education_docx/33970918</dc:relation>
          <dc:rights>CC BY 4.0</dc:rights>
        </oai_dc:dc>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
