Language-Grounded Indoor 3D Semantic Segmentation in the Wild

doi:10.1007/978-3-031-19827-4_8

Open AccessJournal ArticleDOI

Language-Grounded Indoor 3D Semantic Segmentation in the Wild

David Rozenberszki

- 01 Jan 2022 -

Lecture Notes in Computer Science

- pp 125-141

Chats0

TLDR

This paper proposed a language-driven pre-training method to encourage learned 3D features that might have limited training examples to lie close to their pre-trained text embeddings, which outperformed state-of-the-art 3D pre-learning for 3D semantic segmentation.

Abstract:

Recent advances in 3D semantic segmentation with deep neural networks have shown remarkable success, with rapid performance increase on available datasets. However, current 3D semantic segmentation benchmarks contain only a small number of categories – less than 30 for ScanNet and SemanticKITTI, for instance, which are not enough to reflect the diversity of real environments (e.g., semantic image understanding covers hundreds to thousands of classes). Thus, we propose to study a larger vocabulary for 3D semantic segmentation with a new extended benchmark on ScanNet data with 200 class categories, an order of magnitude more than previously studied. This large number of class categories also induces a large natural class imbalance, both of which are challenging for existing 3D semantic segmentation methods. To learn more robust 3D features in this context, we propose a language-driven pre-training method to encourage learned 3D features that might have limited training examples to lie close to their pre-trained text embeddings. Extensive experiments show that our approach consistently outperforms state-of-the-art 3D pre-training for 3D semantic segmentation on our proposed benchmark (+9% relative mIoU), including limited-data scenarios with +25% relative mIoU using only 5% annotations.

Language-Grounded Indoor 3D Semantic Segmentation in the Wild

Citations

4DContrast: Contrastive Learning with Dynamic Correspondences for 3D Scene Understanding

References

Microsoft COCO: Common Objects in Context

The Pascal Visual Object Classes (VOC) Challenge

Focal Loss for Dense Object Detection

Indoor segmentation and support inference from RGBD images

Momentum Contrast for Unsupervised Visual Representation Learning

Related Papers (5)

A syntax and semantics linking algorithm for the chinese language

Shape-included label-consistent discriminative dictionary learning: An approach to detect and segment multi-class objects in images

Multiple Kernel Boosting Based Two-level RGBD Image Co-Segmentation

Variational multichannel multiclass segmentation using unsupervised lifting with CNNs

Leaf Segmentation and Classification with a Complicated Background Using Deep Learning