<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="3.0" xml:lang="en">
<front>
<journal-meta>
<journal-id journal-id-type="publisher">ISPRS-Archives</journal-id>
<journal-title-group>
<journal-title>The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences</journal-title>
<abbrev-journal-title abbrev-type="publisher">ISPRS-Archives</abbrev-journal-title>
<abbrev-journal-title abbrev-type="nlm-ta">Int. Arch. Photogramm. Remote Sens. Spatial Inf. Sci.</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2194-9034</issn>
<publisher><publisher-name>Copernicus Publications</publisher-name>
<publisher-loc>Göttingen, Germany</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.5194/isprs-archives-XLIX-B2-2026-745-2026</article-id>
<title-group>
<article-title>The Emerging Role of Vision-Language Models in the Automation of Railway Asset Management: A Review and Future Perspective</article-title>
</title-group>
<contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Varghese</surname>
<given-names>Ashley</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Ghorbanalivaki</surname>
<given-names>Mohammadjavad</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Sohn</surname>
<given-names>Gunho</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
</contrib-group><aff id="aff1">
<label>1</label>
<addr-line>Canadian National Railway, Canada</addr-line>
</aff>
<aff id="aff2">
<label>2</label>
<addr-line>Dept. of Earth and Space Science and Engineering, York University, Toronto, Ontario, Canada</addr-line>
</aff>
<pub-date pub-type="epub">
<day>23</day>
<month>07</month>
<year>2026</year>
</pub-date>
<volume>XLIX-B2-2026</volume>
<fpage>745</fpage>
<lpage>753</lpage>
<permissions>
<copyright-statement>Copyright: &#x000a9; 2026 Ashley Varghese et al.</copyright-statement>
<copyright-year>2026</copyright-year>
<license license-type="open-access">
<license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri"  xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p>
</license>
</permissions>
<self-uri xlink:href="https://isprs-archives.copernicus.org/articles/XLIX-B2-2026/745/2026/isprs-archives-XLIX-B2-2026-745-2026.html">This article is available from https://isprs-archives.copernicus.org/articles/XLIX-B2-2026/745/2026/isprs-archives-XLIX-B2-2026-745-2026.html</self-uri>
<self-uri xlink:href="https://isprs-archives.copernicus.org/articles/XLIX-B2-2026/745/2026/isprs-archives-XLIX-B2-2026-745-2026.pdf">The full text article is available as a PDF file from https://isprs-archives.copernicus.org/articles/XLIX-B2-2026/745/2026/isprs-archives-XLIX-B2-2026-745-2026.pdf</self-uri>
<abstract>
<p>The safety, efficiency, and longevity of global railway networks are directly linked to the rigorous inspection and management of their vast inventory of physical assets. Over the past decade, the field has progressed from manual surveys to automated systems leveraging imagery from track-based or aerial platforms. These systems predominantly built on traditional Computer Vision (CV) models have proven effective at detecting a pre-defined set of common assets. However, this progress has exposed a fundamental architectural and operational ceiling: the closed-world assumption. Current models are constrained to a fixed catalogue of classes defined during their training. It makes the model incapable of identifying novel objects or adapting to environmental changes without costly and continuous cycles of data re-annotation, retraining, and redeployment. This review paper argues that Vision-Language Models (VLMs), a paradigm whose rapid maturation is evidenced by recent comprehensive surveys offer a transformative solution. We provide a focused overview of the limitations of current CV systems and map the mechanics of a VLM-powered approach specifically Open-Vocabulary Detection and Reasoning Segmentation directly to the outstanding challenges in rail asset management. Ultimately, the literature suggests that the adoption of VLMs could catalyze a fundamental shift in railway infrastructure management that serves as a key enabler for next-generation Predictive Maintenance and autonomous Digital Twins.</p>
</abstract>
<counts><page-count count="9"/></counts>
</article-meta>
</front>
<body/>
<back>
</back>
</article>