<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="3.0" xml:lang="en">
<front>
<journal-meta>
<journal-id journal-id-type="publisher">ISPRS-Archives</journal-id>
<journal-title-group>
<journal-title>The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences</journal-title>
<abbrev-journal-title abbrev-type="publisher">ISPRS-Archives</abbrev-journal-title>
<abbrev-journal-title abbrev-type="nlm-ta">Int. Arch. Photogramm. Remote Sens. Spatial Inf. Sci.</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2194-9034</issn>
<publisher><publisher-name>Copernicus Publications</publisher-name>
<publisher-loc>Göttingen, Germany</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.5194/isprs-archives-XLIX-B2-2026-1319-2026</article-id>
<title-group>
<article-title>SceneReasoner: Decoupled Spatial Tokenization for Indoor Large-Scene Understanding with LLMs</article-title>
</title-group>
<contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Shi</surname>
<given-names>Bohang</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Tang</surname>
<given-names>Shengjun</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Li</surname>
<given-names>Xiaoming</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Wang</surname>
<given-names>Weixi</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Xie</surname>
<given-names>Linfu</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Zhou</surname>
<given-names>Baoding</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
</contrib-group><aff id="aff1">
<label>1</label>
<addr-line>School of Architecture and Urban Planning, Research Institute for Smart Cities, Shenzhen University, Shenzhen, China</addr-line>
</aff>
<aff id="aff2">
<label>2</label>
<addr-line>College of Civil and Transportation Engineering, Shenzhen University, Shenzhen, China</addr-line>
</aff>
<pub-date pub-type="epub">
<day>23</day>
<month>07</month>
<year>2026</year>
</pub-date>
<volume>XLIX-B2-2026</volume>
<fpage>1319</fpage>
<lpage>1325</lpage>
<permissions>
<copyright-statement>Copyright: &#x000a9; 2026 Bohang Shi et al.</copyright-statement>
<copyright-year>2026</copyright-year>
<license license-type="open-access">
<license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri"  xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p>
</license>
</permissions>
<self-uri xlink:href="https://isprs-archives.copernicus.org/articles/XLIX-B2-2026/1319/2026/isprs-archives-XLIX-B2-2026-1319-2026.html">This article is available from https://isprs-archives.copernicus.org/articles/XLIX-B2-2026/1319/2026/isprs-archives-XLIX-B2-2026-1319-2026.html</self-uri>
<self-uri xlink:href="https://isprs-archives.copernicus.org/articles/XLIX-B2-2026/1319/2026/isprs-archives-XLIX-B2-2026-1319-2026.pdf">The full text article is available as a PDF file from https://isprs-archives.copernicus.org/articles/XLIX-B2-2026/1319/2026/isprs-archives-XLIX-B2-2026-1319-2026.pdf</self-uri>
<abstract>
<p>Most existing 3D vision-language models focus on object-level or single-room understanding and perform poorly in large-scale, multi-room indoor environments where task-relevant objects constitute only a small fraction of the total point cloud. When multi-room point clouds are fed directly into an LLM, critical semantic signals are diluted by the vast amount of redundant background, making it difficult for the model to focus on truly relevant regions. We propose SceneReasoner, a decoupled spatial tokenisation framework that addresses this challenge through three core designs: (1) pre-tokenisation text-guided feature weighting that leverages the shared CLIP embedding space between OpenScene point features and text queries to amplify question-relevant point features before any spatial compression occurs; (2) 2D&amp;ndash;3D feature fusion that integrates top-down 2D CLIP features with 3D sparse tokens, supplying the model with appearance semantics&amp;mdash;such as texture, material, and room layout&amp;mdash;absent from raw point clouds; and (3) layer-wise dense feature injection that inserts local dense features into the LLM attention mechanism layer by layer for fine-grained perception of key regions. We evaluate on the XR-Scene benchmark, which covers cross-room question answering and scene captioning over HM3D indoor environments with an average area of 132 m&lt;sup&gt;2&lt;/sup&gt;. SceneReasoner achieves the best CIDEr on XR-SceneCaption (+0.33 over LSceneLLM), the highest METEOR on XR-QA, and competitive ROUGE-L across all three tasks, demonstrating the effectiveness of task-guided spatial tokenisation for large-scene understanding.</p>
</abstract>
<counts><page-count count="7"/></counts>
</article-meta>
</front>
<body/>
<back>
</back>
</article>