<?xml version='1.0' encoding='utf-8'?>
<?xml-stylesheet type="text/xsl" href="/v2/static/oai2.xsl"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-10-06T22:57:43Z</responseDate>
  <request identifier="oai:figshare.com:article/34001037" metadataPrefix="oai_dc" verb="GetRecord">https://api.figshare.com/v2/oai</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:figshare.com:article/34001037</identifier>
        <datestamp>2026-09-30T13:29:25Z</datestamp>
        <setSpec>category_28843</setSpec>
        <setSpec>category_28846</setSpec>
        <setSpec>category_28831</setSpec>
        <setSpec>category_28879</setSpec>
        <setSpec>category_28891</setSpec>
        <setSpec>portal_421</setSpec>
        <setSpec>item_type_8</setSpec>
        <setSpec>month_year_09_2026</setSpec>
      </header>
      <metadata>
        <oai_dc:dc xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"  xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
          <dc:title>TOWARD SPATIAL INTELLIGENCE AND PHYSICALLY REALISTIC WORLD MODELS</dc:title>
          <dc:creator>Lu Ling (25111914)</dc:creator>
          <dc:subject>Knowledge representation and reasoning</dc:subject>
          <dc:subject>Modelling and simulation</dc:subject>
          <dc:subject>Autonomous agents and multiagent systems</dc:subject>
          <dc:subject>Computer vision</dc:subject>
          <dc:subject>Pattern recognition</dc:subject>
          <dc:subject>World Model</dc:subject>
          <dc:subject>Spatial Intelligence</dc:subject>
          <dc:subject>3D Generation</dc:subject>
          <dc:description>&lt;p dir="ltr"&gt;Spatial intelligence requires models to infer and maintain coherent three-dimensional structure, preserve objects and relations across viewpoints, compose interactable environments, and respect geometric and physical constraints. Progress toward such world models is limited by two forms of data scarcity: a shortage of diverse, high-quality real-world spatial observations and the narrow distributions represented by curated interactive 3D scene datasets. This dissertation studies how structured data and reusable priors can address both bottlenecks.&lt;/p&gt;&lt;p dir="ltr"&gt;The first contribution, DL3DV-10K, establishes a high-quality corpus of 10,510 diverse real-world scenes and 51.2 million frames captured as multiview videos, together with a reproducible data-production protocol. Its benchmark reveals novel view synthesis failure modes missed by smaller collections; its scaling experiments show that broader real-world data improves generalizable 3D representation learning; and its adoption across 3D vision, generative video, and hybrid 3D--video systems demonstrates its value as shared infrastructure.&lt;/p&gt;&lt;p dir="ltr"&gt;The second and third contributions study complementary routes to interactive 3D scene generation beyond curated distributions. Scenethesis grounds language planning and visual guidance with explicit 3D assets and geometric and physical constraints, producing editable indoor and outdoor scenes with more coherent relations, support, and collision behavior. I-Scene instead repurposes a pretrained 3D instance generator as a feed-forward scene-level spatial learner. Learning from non-semantic random compositions reduces its dependence on annotated semantic layouts and supports generalization to unseen object arrangements.&lt;/p&gt;&lt;p dir="ltr"&gt;Together, these works show that progress toward spatially intelligent and physically realistic world models depends not only on model design but also on how spatial experience is collected, structured, grounded, and reused. High-quality multiview data supplies transferable observations, while inference-time and learning-based approaches expand interactive 3D environment generation beyond bounded datasets. These findings provide foundations for persistent, interactive, and physically grounded world models.&lt;/p&gt;</dc:description>
          <dc:date>2026-09-30T13:29:25Z</dc:date>
          <dc:type>Text</dc:type>
          <dc:type>Thesis</dc:type>
          <dc:identifier>10.25394/PGS.34001037.v1</dc:identifier>
          <dc:relation>https://figshare.com/articles/thesis/TOWARD_SPATIAL_INTELLIGENCE_AND_PHYSICALLY_REALISTIC_WORLD_MODELS/34001037</dc:relation>
          <dc:rights>CC BY 4.0</dc:rights>
        </oai_dc:dc>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
