<?xml version='1.0' encoding='utf-8'?>
<?xml-stylesheet type="text/xsl" href="/v2/static/oai2.xsl"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-10-07T05:50:44Z</responseDate>
  <request identifier="oai:figshare.com:article/34001649" metadataPrefix="oai_dc" verb="GetRecord">https://api.figshare.com/v2/oai</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:figshare.com:article/34001649</identifier>
        <datestamp>2026-09-26T01:18:37Z</datestamp>
        <setSpec>category_29398</setSpec>
        <setSpec>category_29401</setSpec>
        <setSpec>category_29410</setSpec>
        <setSpec>item_type_3</setSpec>
        <setSpec>month_year_09_2026</setSpec>
      </header>
      <metadata>
        <oai_dc:dc xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"  xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
          <dc:title>Shanghainese Word Frequency (v1.0): a spoken-weighted frequency list with rank uncertainty</dc:title>
          <dc:creator>Verbavia (25112554)</dc:creator>
          <dc:subject>Computational linguistics</dc:subject>
          <dc:subject>Corpus linguistics</dc:subject>
          <dc:subject>Language documentation and description</dc:subject>
          <dc:subject>Shanghainese</dc:subject>
          <dc:subject>Shanghai Wu</dc:subject>
          <dc:subject>Wu Chinese</dc:subject>
          <dc:subject>wuu</dc:subject>
          <dc:subject>word frequency</dc:subject>
          <dc:subject>frequency list</dc:subject>
          <dc:subject>corpus linguistics</dc:subject>
          <dc:subject>spoken corpus</dc:subject>
          <dc:subject>dialect</dc:subject>
          <dc:subject>lexicon</dc:subject>
          <dc:subject>bootstrap</dc:subject>
          <dc:subject>dispersion</dc:subject>
          <dc:description>&lt;p dir="ltr"&gt;A word frequency list for Shanghainese (Shanghai Wu, ISO 639-3 wuu): 2,524 ranked words, each with counts in conversational speech and in written Wu, a combined frequency per million (0.75 × spoken + 0.25 × written), document range, Gries' DP, and a bootstrap 95% rank interval.&lt;/p&gt;&lt;p dir="ltr"&gt;Limits. The spoken corpus is small: 20 speakers, about 4 hours, 52,412 tokens. Ranks are reliable only to rank 236, and there only to within about ±50%; at no depth are they within ±25%. Below rank 236 the order is indicative only. Split-half Spearman between random halves of the speakers is 0.80 (top 100), 0.77 (top 300) and 0.71 (top 500). Spoken and written counts agree weakly (ρ = 0.36 over 833 shared words), which is why speech is weighted 0.75. Transcription conventions affect some counts: 仔 and 哉 never occur in the conversations while 了 occurs 972 times. No native speaker has reviewed the list.&lt;/p&gt;&lt;p dir="ltr"&gt;Sources (counts only; no source text is redistributed). Spoken: MagicData, ASR-CShhiDiaCSC Chinese Shanghai Dialect Conversational Speech Corpus, via the Hugging Face copy TingChen-ppmc/Shanghai_Dialect_Conversational_Speech_Corpus, revision f10f366. Written: the Wu Chinese Wikipedia dump of 2026-09-01, filtered to colloquial Shanghai-type Wu: 2,405 of 48,393 articles, 459,388 tokens. Historical check, not part of the ranking: Pott (1907) and Edkins (1868) from Project Gutenberg, 34,534 tokens; 86% / 81% / 77% of the top 100 / 300 / 500 words appear there.&lt;/p&gt;&lt;p dir="ltr"&gt;Method. OpenCC conversion and a hand-written spelling map (the transcripts write 吾 for 我, and 伐 for both the question particle and the negator 勿); two context rules for 伐 and 呃 showed about 2–5% error in hand-checked samples of 80. Segmentation uses a 3,113-entry hand-written lexicon with a unigram model re-estimated by hard EM; particles such as 仔, 个 and 伐 are words and are never joined to a neighbour. Each word has a Wu Association romanisation, part of speech, flags and an English gloss written for this dataset. The release rebuilds byte for byte from pinned sources (seed 20260926).&lt;/p&gt;&lt;p dir="ltr"&gt;Files. The zip holds the list (output/shanghainese_word_frequency.tsv), per-source counts, JSON files with every quoted number, the lexicon, the code and the documentation.&lt;/p&gt;&lt;p dir="ltr"&gt;Licence. CC0 1.0: the release contains only counts, statistics and original lexicon, glosses and code. The source corpora keep their own licences (MagicData: CC BY-NC-ND 4.0; Wikipedia: CC BY-SA 4.0) and are not redistributed.&lt;/p&gt;&lt;p dir="ltr"&gt;Compiled by the &lt;a href="https://verbavia.com/words/shanghainese/" target="_blank" rel="noreferrer"&gt;Verbavia project&lt;/a&gt;. With thanks to MagicData, the contributors to the Wu Chinese Wikipedia, F. L. Hawks Pott and J. Edkins.&lt;/p&gt;</dc:description>
          <dc:date>2026-09-26T01:18:37Z</dc:date>
          <dc:type>Dataset</dc:type>
          <dc:type>Dataset</dc:type>
          <dc:identifier>10.6084/m9.figshare.34001649.v1</dc:identifier>
          <dc:relation>https://figshare.com/articles/dataset/Shanghainese_Word_Frequency_v1_0_a_spoken-weighted_frequency_list_with_rank_uncertainty/34001649</dc:relation>
          <dc:rights>CC0</dc:rights>
        </oai_dc:dc>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
