<?xml version='1.0' encoding='utf-8'?>
<?xml-stylesheet type="text/xsl" href="/v2/static/oai2.xsl"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-10-07T19:11:43Z</responseDate>
  <request identifier="oai:figshare.com:article/33104105" metadataPrefix="oai_dc" verb="GetRecord">https://api.figshare.com/v2/oai</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:figshare.com:article/33104105</identifier>
        <datestamp>2026-10-02T14:23:47Z</datestamp>
        <setSpec>category_26971</setSpec>
        <setSpec>category_26557</setSpec>
        <setSpec>category_29173</setSpec>
        <setSpec>item_type_3</setSpec>
        <setSpec>month_year_10_2026</setSpec>
      </header>
      <metadata>
        <oai_dc:dc xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"  xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
          <dc:title>On-Road Vehicular Emission Measurement Data &amp; Analysis leakage codes</dc:title>
          <dc:creator>Mohd Kamil Vakil (24474350)</dc:creator>
          <dc:subject>Pollution and contamination not elsewhere classified</dc:subject>
          <dc:subject>Air pollution modelling and control</dc:subject>
          <dc:subject>Machine learning not elsewhere classified</dc:subject>
          <dc:subject>cross-validation leakage</dc:subject>
          <dc:subject>random forest</dc:subject>
          <dc:subject>vehicle emissions</dc:subject>
          <dc:subject>leave-one-group-out</dc:subject>
          <dc:subject>small-fleet machine learning</dc:subject>
          <dc:description>&lt;p dir="ltr"&gt;Random-forest and other flexible regressors fitted to small-fleet vehicle emission datasets often fail to generalize when validated on held-out vehicles, yet appear accurate under standard random cross-validation. We demonstrate this on a nine-vehicle Indian in-use fleet, showing that adding static vehicle attributes (kerb weight, engine displacement, odometer) raises random-fold R² to 0.60–0.86, but leave-one-vehicle-out R² drops to −0.47 to −1.64. The mechanism: the three attributes jointly identify every vehicle uniquely, so the model functions as a vehicle-identity lookup rather than an emission model; random splitting leaks this identity across train/test boundaries. A decision-tree classifier recovers vehicle identity from these attributes with 100% accuracy. Pooling more training vehicles (2–8) does not close the gap. We recommend leave-one-vehicle-out or leave-one-group-out validation as a minimum reporting standard for any fleet-emission model including vehicle-constant features, and quantify on real data the error that random-split reporting alone would produce.&lt;/p&gt;</dc:description>
          <dc:date>2026-10-02T14:23:47Z</dc:date>
          <dc:type>Dataset</dc:type>
          <dc:type>Dataset</dc:type>
          <dc:identifier>10.6084/m9.figshare.33104105.v2</dc:identifier>
          <dc:relation>https://figshare.com/articles/dataset/On-Road_Vehicular_Emission_Measurement_Data/33104105</dc:relation>
          <dc:rights>CC BY 4.0</dc:rights>
        </oai_dc:dc>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
