<?xml version='1.0' encoding='utf-8'?>
<?xml-stylesheet type="text/xsl" href="/v2/static/oai2.xsl"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-10-10T21:42:38Z</responseDate>
  <request identifier="oai:figshare.com:article/33872113" metadataPrefix="oai_dc" verb="GetRecord">https://api.figshare.com/v2/oai</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:figshare.com:article/33872113</identifier>
        <datestamp>2026-09-17T05:54:19Z</datestamp>
        <setSpec>category_396</setSpec>
        <setSpec>portal_316</setSpec>
        <setSpec>item_type_3</setSpec>
        <setSpec>month_year_09_2026</setSpec>
      </header>
      <metadata>
        <oai_dc:dc xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"  xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
          <dc:title>Data Sheet 8_A comparative cross-sectional evaluation of generative AI chatbots for patient-oriented bacterial vaginosis health advice: safety, accuracy, guideline concordance, empathy, and readability.pdf</dc:title>
          <dc:creator>Ling Miao (1533910)</dc:creator>
          <dc:creator>Lina Gu (237116)</dc:creator>
          <dc:creator>Zhaole Gong (24548625)</dc:creator>
          <dc:creator>Zhengfeng Gu (24704371)</dc:creator>
          <dc:creator>Xiaoli Qian (10205046)</dc:creator>
          <dc:subject>Reproduction</dc:subject>
          <dc:subject>bacterial vaginosis</dc:subject>
          <dc:subject>chatbot</dc:subject>
          <dc:subject>generative artificial intelligence</dc:subject>
          <dc:subject>large language model</dc:subject>
          <dc:subject>patient education</dc:subject>
          <dc:subject>patient safety</dc:subject>
          <dc:description>Background&lt;p&gt;Bacterial vaginosis (BV) concerns involve intimate symptoms, stigma, diagnostic uncertainty, medication use, pregnancy, sexual health, and self-care. Publicly accessible generative artificial intelligence chatbots offer immediate and potentially non-judgmental information, but the quality and safety of specific responses remain uncertain.&lt;/p&gt;Objective&lt;p&gt;To compare five publicly accessible generative AI chatbot interfaces using standardized, researcher-developed patient-oriented questions about BV. In this study, patient-oriented denotes consumer-facing wording and does not imply direct patient derivation or validation.&lt;/p&gt;Methods&lt;p&gt;Sixty-one standardized English-language questions were submitted once to ChatGPT, Gemini, Microsoft Copilot, DeepSeek, and Doubao in separate single-turn sessions, yielding 305 responses. Five senior obstetrician-gynecologists independently evaluated safety, accuracy, study-specific guideline-anchored concordance, and empathy using a question-specific reference framework. Independent ratings were locked before adjudication and were used to estimate inter-rater agreement. Consensus scores were used for primary comparisons, and a sensitivity analysis used the median of the five locked ratings. Readability was assessed separately using six formula-based indices. Paired comparisons used Cochran's Q, Friedman, McNemar, and Wilcoxon signed-rank tests with Benjamini-Hochberg adjustment.&lt;/p&gt;Results&lt;p&gt;Thirty-one responses (10.2%) were classified as unsafe or potentially unsafe. Observed interface-specific rates ranged from 6.6% (95% CI, 2.6%–15.7%) to 14.8% (95% CI, 8.0%–25.7%), while the matched binary comparison did not detect an overall difference (Cochran's Q = 2.596, df = 4, P = 0.627). Under the documented query conditions, differences were detected in accuracy (Kendall's W = 0.686), study-specific guideline-anchored concordance (W = 0.243), empathy (W = 0.322), and all six readability indices (W range, 0.482–0.679; all P &lt; 0.001). The sensitivity analysis produced the same inferential conclusions. Unsafe content occurred in every interface and was concentrated in recurrent-BV, sexual-health, medication, and self-care scenarios.&lt;/p&gt;Conclusions&lt;p&gt;This exploratory study characterizes a recorded sample of specific chatbot responses, not stable or repeatable performance characteristics. Clinically relevant risks were observed in every interface, and the non-significant safety comparison should not be interpreted as equivalence. Because each researcher-developed prompt was submitted once and the interfaces were queried on different dates in a fixed order, the findings do not establish a stable ranking of underlying model capability, clinical effectiveness, or suitability for unsupervised care.&lt;/p&gt;</dc:description>
          <dc:date>2026-09-17T05:54:19Z</dc:date>
          <dc:type>Dataset</dc:type>
          <dc:type>Dataset</dc:type>
          <dc:identifier>10.3389/frph.2026.1967818.s004</dc:identifier>
          <dc:relation>https://figshare.com/articles/dataset/Data_Sheet_8_A_comparative_cross-sectional_evaluation_of_generative_AI_chatbots_for_patient-oriented_bacterial_vaginosis_health_advice_safety_accuracy_guideline_concordance_empathy_and_readability_pdf/33872113</dc:relation>
          <dc:rights>CC BY 4.0</dc:rights>
        </oai_dc:dc>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
