<?xml version='1.0' encoding='utf-8'?>
<?xml-stylesheet type="text/xsl" href="/v2/static/oai2.xsl"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-10-07T18:23:55Z</responseDate>
  <request identifier="oai:figshare.com:article/33959791" metadataPrefix="oai_dc" verb="GetRecord">https://api.figshare.com/v2/oai</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:figshare.com:article/33959791</identifier>
        <datestamp>2026-09-21T19:34:18Z</datestamp>
        <setSpec>category_29194</setSpec>
        <setSpec>category_26539</setSpec>
        <setSpec>category_26548</setSpec>
        <setSpec>item_type_3</setSpec>
        <setSpec>month_year_09_2026</setSpec>
      </header>
      <metadata>
        <oai_dc:dc xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"  xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
          <dc:title>replication package</dc:title>
          <dc:creator>KHALED ISMAIL (25076764)</dc:creator>
          <dc:subject>Software architecture</dc:subject>
          <dc:subject>Engineering practice</dc:subject>
          <dc:subject>Systems engineering</dc:subject>
          <dc:subject>archival demonstration</dc:subject>
          <dc:subject>story point</dc:subject>
          <dc:subject>extraction</dc:subject>
          <dc:subject>analysis</dc:subject>
          <dc:subject>conformal scripts</dc:subject>
          <dc:description>&lt;pre&gt;# Replication package&lt;br&gt;&lt;br&gt;## Intelligent Agile Adoption Framework — archival demonstration (Section 5)&lt;br&gt;&lt;br&gt;Khaled Ismail · ORCID 0000-0002-0499-5805&lt;br&gt;&lt;br&gt;### What this tests&lt;br&gt;&lt;br&gt;Proposition 2 of the manuscript is separated into two claims and both are tested here.&lt;br&gt;&lt;br&gt;* **P2a** — the dominant component of squared forecast error in human Agile estimation is&lt;br&gt;dispersion, not bias. **Supported.**&lt;br&gt;* **P2b** — predictive inference reduces that dispersion component. **Not supported** on this&lt;br&gt;corpus, in a result consistent with independent replications (Tawosi et al., 2023, 2024).&lt;br&gt;&lt;br&gt;Pooled human forecast error: SD 2.358 log-units, mean +0.071; the central 80% of issues are&lt;br&gt;realized at 0.05x to 13.5x their planned duration.&lt;br&gt;&lt;br&gt;A third analysis (`p2\_conformal.py`) compares interval calibration methods on identical&lt;br&gt;predictions. Quantile regression attains **64.3%** coverage on a nominal 80% interval;&lt;br&gt;locally adaptive split conformal attains **76.6%**, within 5 points of nominal in 21 of 24&lt;br&gt;projects (Wilcoxon p = 5.2e-6, rank-biserial 0.91), at a 28% cost in interval width. The&lt;br&gt;severe under-coverage is therefore largely an estimator artifact rather than an irreducible&lt;br&gt;property of the data; the residual 3.4-point shortfall (p = 0.039) indicates mild&lt;br&gt;non-exchangeability, i.e. process drift.&lt;br&gt;&lt;br&gt;### Data&lt;br&gt;&lt;br&gt;TAWOS — a versatile dataset of agile open source software projects (Tawosi, Al-Subaihin,&lt;br&gt;Moussa \&amp; Sarro, MSR 2022), DOI 10.1145/3524842.3528029. Dataset DOI 10.5522/04/21308124,&lt;br&gt;Apache License 2.0. 458,232 issues, 39 projects, 12 public Jira repositories.&lt;br&gt;&lt;br&gt;The dataset ships as a MySQL dump. `extract.py` stream-parses it directly and requires no&lt;br&gt;database server.&lt;br&gt;&lt;br&gt;### Pipeline&lt;br&gt;&lt;br&gt;```&lt;br&gt;python3 extract.py                        # TAWOS.sql -&gt; issues.csv (458,232 rows)&lt;br&gt;python3 p2\_analysis.py primary            # total-effort outcome, revised estimates excluded&lt;br&gt;python3 p2\_analysis.py sens\_resolution    # resolution-time sensitivity&lt;br&gt;python3 p2\_conformal.py                   # conformal vs quantile interval calibration&lt;br&gt;python3 delegation\_rule.py                # regenerates Table 3, Table 4 and Table S6&lt;br&gt;python3 figs.py                           # regenerates Figures 1, 2, 3 and S1&lt;br&gt;```&lt;br&gt;&lt;br&gt;### Analysed sample&lt;br&gt;&lt;br&gt;|Specification|Projects|Predictions|&lt;br&gt;|-|-|-|&lt;br&gt;|primary (Total\_Effort\_Minutes)|25|31,356|&lt;br&gt;|sensitivity (Resolution\_Time\_Minutes)|32|43,536|&lt;br&gt;&lt;br&gt;Inclusion: story point &gt; 0; outcome &gt; 0; story point not revised after estimation;&lt;br&gt;at least 150 prior issues before the first prediction; at least 60 evaluable predictions&lt;br&gt;per project.&lt;br&gt;&lt;br&gt;### Method&lt;br&gt;&lt;br&gt;Error is taken in log space, e = ln(actual) − ln(predicted), giving the exact decomposition&lt;br&gt;MSLE = mean(e)² + var(e) into a bias and a dispersion term. Three forecasters are evaluated&lt;br&gt;strictly out of sample under a time-ordered walk-forward within each project:&lt;br&gt;&lt;br&gt;* **H** human — story point × historical minutes-per-point rate (velocity planning)&lt;br&gt;* **R** reference class — historical median duration of prior same-type issues&lt;br&gt;* **M** model — gradient-boosted quantile regression at the 0.10 / 0.50 / 0.90 quantiles&lt;br&gt;&lt;br&gt;Inference across projects uses Wilcoxon signed-rank tests, 10,000-sample bootstrap&lt;br&gt;confidence intervals, and matched-pairs rank-biserial effect sizes. The unit of analysis is&lt;br&gt;the project (n = 25), not the issue.&lt;br&gt;&lt;br&gt;### Known limitations&lt;br&gt;&lt;br&gt;* The outcome is a **duration proxy** derived from issue state transitions, not logged&lt;br&gt;effort. Only 900 issues in the corpus carry logged work, too few to analyse.&lt;br&gt;* The corpus contains **none of the Layer 1 telemetry** the framework specifies — no CI&lt;br&gt;outcomes, work-in-progress levels, dependency depth, review latency or team composition&lt;br&gt;history. The P2b test is therefore a joint test of P1 and P2b, and a null is what P1&lt;br&gt;predicts under thin telemetry.&lt;br&gt;* The model arm is one model class with one feature set.&lt;br&gt;&lt;br&gt;### Environment&lt;br&gt;&lt;br&gt;Python 3.13 · numpy 2.4 · pandas 3.0 · scipy 1.17 · scikit-learn 1.8 · matplotlib.&lt;br&gt;Random seed 20260921 for all bootstrap resampling.&lt;br&gt;&lt;br&gt;### Figure palette&lt;br&gt;&lt;br&gt;The two-colour categorical palette (#2E6FA8, #C2681A) was validated for colour-vision&lt;br&gt;deficiency separation, lightness band, chroma floor and surface contrast before use.&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;br&gt;&lt;/pre&gt;&lt;p&gt;&lt;/p&gt;</dc:description>
          <dc:date>2026-09-21T19:34:18Z</dc:date>
          <dc:type>Dataset</dc:type>
          <dc:type>Dataset</dc:type>
          <dc:identifier>10.6084/m9.figshare.33959791.v1</dc:identifier>
          <dc:relation>https://figshare.com/articles/dataset/replication_package/33959791</dc:relation>
          <dc:rights>CC BY 4.0</dc:rights>
        </oai_dc:dc>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
