Full-Resolution Residual Networks for Semantic Segmentation in Street Scenes

Pohlen, Tobias; Hermans, Alexander; Mathias, Markus; Leibe, Bastian

Computer Science > Computer Vision and Pattern Recognition

arXiv:1611.08323 (cs)

[Submitted on 24 Nov 2016 (v1), last revised 6 Dec 2016 (this version, v2)]

Title:Full-Resolution Residual Networks for Semantic Segmentation in Street Scenes

Authors:Tobias Pohlen, Alexander Hermans, Markus Mathias, Bastian Leibe

View PDF

Abstract:Semantic image segmentation is an essential component of modern autonomous driving systems, as an accurate understanding of the surrounding scene is crucial to navigation and action planning. Current state-of-the-art approaches in semantic image segmentation rely on pre-trained networks that were initially developed for classifying images as a whole. While these networks exhibit outstanding recognition performance (i.e., what is visible?), they lack localization accuracy (i.e., where precisely is something located?). Therefore, additional processing steps have to be performed in order to obtain pixel-accurate segmentation masks at the full image resolution. To alleviate this problem we propose a novel ResNet-like architecture that exhibits strong localization and recognition performance. We combine multi-scale context with pixel-level accuracy by using two processing streams within our network: One stream carries information at the full image resolution, enabling precise adherence to segment boundaries. The other stream undergoes a sequence of pooling operations to obtain robust features for recognition. The two streams are coupled at the full image resolution using residuals. Without additional processing steps and without pre-training, our approach achieves an intersection-over-union score of 71.8% on the Cityscapes dataset.

Comments:	Changes in v2: Fixed equation (10), fixed legend of Figure 6, fixed legend of Figure 9, added page numbers, fixed minor spelling mistakes
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:1611.08323 [cs.CV]
	(or arXiv:1611.08323v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1611.08323

Submission history

From: Tobias Pohlen [view email]
[v1] Thu, 24 Nov 2016 23:55:28 UTC (4,632 KB)
[v2] Tue, 6 Dec 2016 19:36:19 UTC (4,878 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Full-Resolution Residual Networks for Semantic Segmentation in Street Scenes

Submission history

Access Paper:

References & Citations

1 blog link

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Full-Resolution Residual Networks for Semantic Segmentation in Street Scenes

Submission history

Access Paper:

References & Citations

1 blog link

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators