<?xml version='1.0' encoding='utf-8'?>
<?xml-stylesheet type="text/xsl" href="/v2/static/oai2.xsl"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-10-09T18:51:47Z</responseDate>
  <request identifier="oai:figshare.com:article/33991537" metadataPrefix="oai_dc" verb="GetRecord">https://api.figshare.com/v2/oai</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:figshare.com:article/33991537</identifier>
        <datestamp>2026-10-01T04:21:14Z</datestamp>
        <setSpec>category_24184</setSpec>
        <setSpec>item_type_3</setSpec>
        <setSpec>month_year_10_2026</setSpec>
      </header>
      <metadata>
        <oai_dc:dc xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"  xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
          <dc:title>Reference data for nanoID testing</dc:title>
          <dc:creator>dieter tourlousse (4239604)</dc:creator>
          <dc:subject>Bioinformatic methods development</dc:subject>
          <dc:subject>nanoid</dc:subject>
          <dc:subject>reference sequences</dc:subject>
          <dc:subject>simulated ONT reads</dc:subject>
          <dc:description>&lt;p&gt;&lt;br&gt;&lt;/p&gt;&lt;p dir="ltr"&gt;&lt;b&gt;[A] Genome-derived 16S rRNA gene sequences of 50-species gut mock community&lt;/b&gt;&lt;/p&gt;&lt;p dir="ltr"&gt;File: refseqs_50species_gut_mock_V1V9.fasta&lt;/p&gt;&lt;p dir="ltr"&gt;&lt;br&gt;&lt;/p&gt;&lt;p dir="ltr"&gt;&lt;b&gt;[B] Simulated ONT reads&lt;/b&gt;&lt;/p&gt;&lt;p dir="ltr"&gt;Files: simulated_mock_compositions_n_reads.tsv, refseqs_simulations_V1V9.tar.gz, simulated_fastq.tar.gz&lt;/p&gt;&lt;p dir="ltr"&gt;&lt;i&gt;Long amplicon reads were simulated using Badread v0.4.1 (Wick, 2019) with the nanopore2023 quality score model. Simulations were performed without random reads (--random_reads 0), chimeras (--chimeras 0), glitches (--glitches 0,0,0), or adapter sequences (--start_adapter 0,0 and --end_adapter 0,0). Reference 16S rRNA gene sequences were obtained from the Genome Taxonomy Database (GTDB release 220; file ssu_all_r220.fna). To extract high-quality near-full-length sequences, in silico PCR was performed with Cutadapt using the parameters --front 'AGRGTTYGATYHTGGCTCAG...AAGTCGTAACAAGGTARCCG' --discard-untrimmed --overlap 20 --error-rate 2 --action retain --minimum-length 1200 --maximum-length 1800 --max-n 0. Trimmed sequences belonging to GTDB representative genomes corresponding to approximately 500 bacterial species commonly detected in more than 1,000 healthy Japanese individuals were retained. Reads were simulated independently for each sequence, including multiple copies originating from the same genome. &lt;/i&gt;&lt;i&gt;Four Badread identity settings were used to generate reads with different accuracy profiles: &lt;/i&gt;&lt;b&gt;&lt;i&gt;97,98.5,1&lt;/i&gt;&lt;/b&gt;&lt;i&gt; (fastq identifier: badread1), &lt;/i&gt;&lt;b&gt;&lt;i&gt;98,99.5,1&lt;/i&gt;&lt;/b&gt;&lt;i&gt; (badread2), &lt;/i&gt;&lt;b&gt;&lt;i&gt;98.5,99.9,1&lt;/i&gt;&lt;/b&gt;&lt;i&gt; (badread3), and &lt;/i&gt;&lt;b&gt;&lt;i&gt;99,100,1&lt;/i&gt;&lt;/b&gt;&lt;i&gt;(badread4), where the parameters specify the mean, maximum, and standard deviation of the simulated read identity distribution, respectively. Simulated reads subsequently underwent primer trimming with Cutadapt, and reads lacking either primer sequence were discarded to retain only complete amplicons. &lt;/i&gt;&lt;i&gt;The resulting reads were randomly assigned to four simulated mock communities, each comprising 100 species. Each 16S rRNA gene sequence was represented by 300 reads, such that variation in 16S rRNA gene copy number both within and between genomes was reflected in the final read abundances.&lt;/i&gt;&lt;/p&gt;</dc:description>
          <dc:date>2026-10-01T04:21:14Z</dc:date>
          <dc:type>Dataset</dc:type>
          <dc:type>Dataset</dc:type>
          <dc:identifier>10.6084/m9.figshare.33991537.v1</dc:identifier>
          <dc:relation>https://figshare.com/articles/dataset/Reference_data_for_nanoID_testing/33991537</dc:relation>
          <dc:rights>CC BY 4.0</dc:rights>
        </oai_dc:dc>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
