<?xml version='1.0' encoding='utf-8'?>
<?xml-stylesheet type="text/xsl" href="/v2/static/oai2.xsl"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-10-11T12:16:21Z</responseDate>
  <request identifier="oai:figshare.com:article/33980773" metadataPrefix="oai_dc" verb="GetRecord">https://api.figshare.com/v2/oai</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:figshare.com:article/33980773</identifier>
        <datestamp>2026-09-24T08:53:51Z</datestamp>
        <setSpec>category_24310</setSpec>
        <setSpec>item_type_3</setSpec>
        <setSpec>month_year_09_2026</setSpec>
      </header>
      <metadata>
        <oai_dc:dc xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"  xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
          <dc:title>Pervasive influence of repeat-induced point mutations on single-copy functional genes confirm its importance in the evolution of &lt;i&gt;Neurospora crassa&lt;/i&gt;</dc:title>
          <dc:creator>Huawei Tan (21600428)</dc:creator>
          <dc:subject>Genomics</dc:subject>
          <dc:subject>Neurospora crassa</dc:subject>
          <dc:subject>FGSC2225</dc:subject>
          <dc:subject>Repeat-induced point mutation (RIP)</dc:subject>
          <dc:subject>Highly RIP-Affected Gene</dc:subject>
          <dc:subject>Functional gene</dc:subject>
          <dc:subject>Anti-preservative</dc:subject>
          <dc:description>&lt;p dir="ltr"&gt;1 FGSC2225 Genome&lt;/p&gt;&lt;p&gt;&lt;br&gt;&lt;/p&gt;&lt;p dir="ltr"&gt;FGSC2225-Genome.zip contains the complete genome assembly for Neurospora crassa strain FGSC2225, including the assembled genome sequence, annotated coding sequences (CDS), and predicted peptide sequences.&lt;/p&gt;&lt;p&gt;&lt;br&gt;&lt;/p&gt;&lt;p dir="ltr"&gt;Assembly-Scripts.zip contains the scripts used for the assembly and annotation of the FGSC2225 genome, covering genome survey profiling, repeat annotation, and gene model prediction.&lt;/p&gt;&lt;p&gt;&lt;br&gt;&lt;/p&gt;&lt;p&gt;&lt;br&gt;&lt;/p&gt;&lt;p&gt;&lt;br&gt;&lt;/p&gt;&lt;p dir="ltr"&gt;2 Simulation on Ti/Tv Method&lt;/p&gt;&lt;p&gt;&lt;br&gt;&lt;/p&gt;&lt;p dir="ltr"&gt;To evaluate the efficacy of the Ti/Tv method, we conducted benchmark simulations by artificially introducing RIP-like mutations into the genome of Sordaria macrospora—a close relative of N. crassa that lacks an active RIP system. We simulated RIP-like mutations into the genome of Sordaria macrospora, a close relative of N. crassa that lacks active RIP. We simulated a range of RIP-induced sequence divergences (0.1%, 0.5%, 1%, 2%, 5%, 10%, 15%, and 20%), as well as mixed models containing both RIP-like and non-RIP-like mutations (see Methods for details). Across all conditions, the Ti/Tv method consistently demonstrated higher sensitivity than the canonical RIP index in detecting RIP-affected genes (see our paper Fig. S4b &amp; S4c). Importantly, the modest reduction in specificity relative to the RIP index was effectively mitigated by applying a Ti/Tv threshold of ≥ 10 (see our paper Fig. S4d).&lt;/p&gt;&lt;p&gt;&lt;br&gt;&lt;/p&gt;&lt;p dir="ltr"&gt;Simulation-TiTv-Method.tar.gz contains the example datasets, simulation pipelines, and analysis scripts used in this benchmark evaluation.&lt;/p&gt;&lt;p&gt;&lt;br&gt;&lt;/p&gt;&lt;p&gt;&lt;br&gt;&lt;/p&gt;&lt;p&gt;&lt;br&gt;&lt;/p&gt;&lt;p dir="ltr"&gt;3 Sanger verification data&lt;/p&gt;&lt;p dir="ltr"&gt;3.1 Methods&lt;/p&gt;&lt;p&gt;&lt;br&gt;&lt;/p&gt;&lt;p dir="ltr"&gt;To address potential mapping artifacts and short-read misalignment associated with self-repetitive sequences, we conducted a pairwise BLAST analysis of all coding sequences within the FGSC2489 reference genome. A lenient threshold (E-value &lt; 1e-4 and alignment length &gt; 80 bp) was utilized to identify regions of internal homology, resulting in the identification of 132 HRAGs. To experimentally validate the accuracy of the bioinformatically predicted variants, 8 HRAGs harboring candidate SNPs were selected for further PCR amplification. Primers were designed using Primer 5 (detailed in Sanger_verification_summary_tables.xlsx), and the resulting amplicons were subjected to Sanger sequencing to verify the identified variants. Additionally, to rigorously evaluate the accuracy of novel SNP identification within internal homology domains, the sequence of the Chitinase-1 gene was amplified and verified by Sanger sequencing across both parental strains and three representative offspring.&lt;/p&gt;&lt;p&gt;&lt;br&gt;&lt;/p&gt;&lt;p&gt;&lt;br&gt;&lt;/p&gt;&lt;p&gt;&lt;br&gt;&lt;/p&gt;&lt;p dir="ltr"&gt;3.2 Results&lt;/p&gt;&lt;p&gt;&lt;br&gt;&lt;/p&gt;&lt;p dir="ltr"&gt;To exclude potential mapping artifacts, we selected a subset of HRAG loci that are more susceptible to such issues for experimental validation, including Chitinase-1, which contains short internal repetitive elements and exhibits a high local density of mutations. Such features can increase the likelihood of ambiguous short-read alignment or clustering of apparent SNPs. These loci were therefore validated using Sanger sequencing. All tested variants were fully consistent with short-read–based variant calls (Sanger_verification_summary_tables.xlsx), confirming the robustness of our variant detection.&lt;/p&gt;&lt;p&gt;&lt;br&gt;&lt;/p&gt;&lt;p dir="ltr"&gt;Sanger_verification_data.zip contains supplementary PCR validation results, demonstrating that the substantial observed mutations at coding genes in N. crassa are not artifacts of sequence repetition or read misalignment, but represent genuine, reproducible genetic variation. The complete data matrix summarizing these PCR validation experiments and Sanger sequencing results is provided in Sanger_verification_summary_tables.xlsx.&lt;/p&gt;&lt;p&gt;&lt;br&gt;&lt;/p&gt;&lt;p dir="ltr"&gt;Folder structure&lt;/p&gt;&lt;p dir="ltr"&gt;Sanger_verification_data/&lt;/p&gt;&lt;p dir="ltr"&gt;├── ab1/&lt;/p&gt;&lt;p dir="ltr"&gt;├── ref_fasta/&lt;/p&gt;&lt;p dir="ltr"&gt;├── screenshot/&lt;/p&gt;&lt;p dir="ltr"&gt;├── Readme.txt&lt;/p&gt;&lt;p dir="ltr"&gt;└── Sanger_verification_summary_tables.xlsx&lt;/p&gt;&lt;p&gt;&lt;br&gt;&lt;/p&gt;&lt;p dir="ltr"&gt;Contents&lt;/p&gt;&lt;p&gt;--------&lt;/p&gt;&lt;p dir="ltr"&gt;1. ab1/&lt;/p&gt;&lt;p dir="ltr"&gt;   This folder contains the original Sanger sequencing chromatogram files in AB1 format.&lt;/p&gt;&lt;p dir="ltr"&gt;   These files can be opened with software such as SnapGene Viewer, Chromas, FinchTV, or other chromatogram viewers.&lt;/p&gt;&lt;p dir="ltr"&gt;   They are used to inspect peak quality and manually confirm Sanger validation results.&lt;/p&gt;&lt;p&gt;&lt;br&gt;&lt;/p&gt;&lt;p dir="ltr"&gt;2. ref_fasta/&lt;/p&gt;&lt;p dir="ltr"&gt;   This folder contains FASTA files used for sequence alignment.&lt;/p&gt;&lt;p dir="ltr"&gt;   They can be used to verify whether the detected variant is supported by the sequencing result.&lt;/p&gt;&lt;p&gt;&lt;br&gt;&lt;/p&gt;&lt;p dir="ltr"&gt;3. screenshot/&lt;/p&gt;&lt;p dir="ltr"&gt;   This folder contains screenshots of chromatograms, sequence alignments, or validation results.&lt;/p&gt;&lt;p dir="ltr"&gt;   These images are provided for quick visual inspection and documentation.&lt;/p&gt;&lt;p dir="ltr"&gt;   Visualization of these files was implemented using SnapGene.&lt;/p&gt;&lt;p&gt;&lt;br&gt;&lt;/p&gt;&lt;p dir="ltr"&gt;4. readme.txt&lt;/p&gt;&lt;p dir="ltr"&gt;   This file describes the organization and purpose of the supplementary materials.&lt;/p&gt;&lt;p&gt;&lt;br&gt;&lt;/p&gt;&lt;p dir="ltr"&gt;5. Sanger_verification_summary_tables.xlsx&lt;/p&gt;&lt;p dir="ltr"&gt;   This file data matrix summarizing these PCR validation experiments and Sanger sequencing results&lt;/p&gt;&lt;p&gt;&lt;br&gt;&lt;/p&gt;&lt;p dir="ltr"&gt;Tips&lt;/p&gt;&lt;p&gt;-------------&lt;/p&gt;&lt;p dir="ltr"&gt;Detailed information on the corresponding PCR regions and SNP sites is provided in our paper Supplemental Datasheet 4.&lt;/p&gt;&lt;p&gt;&lt;br&gt;&lt;/p&gt;&lt;p dir="ltr"&gt;The recommended workflow is:&lt;/p&gt;&lt;p dir="ltr"&gt;1. Open the AB1 chromatogram files in the ab1 folder to check raw Sanger sequencing quality.&lt;/p&gt;&lt;p dir="ltr"&gt;2. Compare the corresponding sequences with the files in alignment_fasta.&lt;/p&gt;&lt;p dir="ltr"&gt;3. Refer to the images in screenshot for visual confirmation of the Sanger validation result.&lt;/p&gt;&lt;p&gt;&lt;br&gt;&lt;/p&gt;&lt;p&gt;&lt;br&gt;&lt;/p&gt;&lt;p dir="ltr"&gt;Codes used in this study are also available at GitHub (https://github.com/TanHuawei/RIPed-gene).&lt;/p&gt;</dc:description>
          <dc:date>2026-09-24T08:53:51Z</dc:date>
          <dc:type>Dataset</dc:type>
          <dc:type>Dataset</dc:type>
          <dc:identifier>10.6084/m9.figshare.33980773.v1</dc:identifier>
          <dc:relation>https://figshare.com/articles/dataset/Pervasive_influence_of_repeat-induced_point_mutations_on_single-copy_functional_genes_confirm_its_importance_in_the_evolution_of_i_Neurospora_crassa_i_/33980773</dc:relation>
          <dc:rights>CC BY 4.0</dc:rights>
        </oai_dc:dc>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
