室内空间物体检测与场景布局估计融合方法综述:机遇与挑战

孙绮曼, 孙宇轩, 周百顺, 易波, 王兴伟

小型微型计算机系统 ›› 2026, Vol. 47 ›› Issue (9) : 2199 -2210.

小型微型计算机系统 ›› 2026, Vol. 47 ›› Issue (9) : 2199 -2210. DOI: 10.20009/j.cnki.21-1106/TP.2026-0133
计算机图形与图像

室内空间物体检测与场景布局估计融合方法综述:机遇与挑战

    孙绮曼1, 孙宇轩1, 周百顺2, 易波1, 王兴伟1
作者信息 +

Review of Fusion Methods for Indoor Spatial Object Detection and Scene Layout Estimation:Opportunities and Challenges

    SUN Qiman1, SUN Yuxuan1, ZHOU Baishun2, YI Bo1, WANG Xingwei1
Author information +
文章历史 +

摘要

随着具身智能与人机交互等应用对环境理解能力提出更高要求,室内场景的空间物体检测与布局估计日益成为构建空间智能的关键基础.二者在几何结构与语义上下文层面高度互补,其协同建模已成为室内场景理解的重要方向.本文聚焦室内场景中的物体—布局联合感知问题,系统梳理了该领域从任务分离到协同建模的发展脉络,并将相关方法归纳为早期启发式建模、深度学习联合估计、结构化拓扑推理和生成式序列建模4类.在此基础上,进一步总结常用数据集、代表性研究进展及当前在实时性、泛化性、开放词汇感知和数据依赖等方面面临的主要问题,以期为室内空间智能研究提供参考.

Abstract

As applications such as embodied intelligence and human-computer interaction put forward higher requirements for environmental understanding capabilities,spatial object detection and layout estimation in indoor scenes are increasingly becoming the key basis for building spatial intelligence.The two are highly complementary at the level of geometric structure and semantic context,and their collaborative modeling has become an important direction for indoor scene understanding.This paper focuses on the joint perception of object-layout in indoor scenes,systematically reviews the development of this field from task separation to collaborative modeling,and summarizes the related methods into four categories:early heuristic modeling,deep learning joint estimation,structured topology reasoning and generative sequence modeling.On this basis,the common datasets,representative research progress and the main problems in real-time,generalization,open vocabulary perception and data dependence are further summarized in order to provide reference for the research of indoor space intelligence.

关键词

空间物体检测 / 场景布局估计 / 协同建模 / 空间智能 / 具身智能

Key words

space object detection / scene layout estimation / collaborative modeling / space intelligence / embodied intelligence

引用本文

引用格式 ▾
孙绮曼, 孙宇轩, 周百顺, 易波, 王兴伟. 室内空间物体检测与场景布局估计融合方法综述:机遇与挑战[J]. 小型微型计算机系统, 2026, 47(9): 2199-2210 DOI:10.20009/j.cnki.21-1106/TP.2026-0133

登录浏览全文

4963

注册一个新账户 忘记密码

参考文献

[1] Mao Y S,Zhong J H,Fang C,et al.SpatialLM:training large language models for structured indoor modeling[EB/OL].https://arxiv.org/abs/2506.07491,2025.
[2] Chen X X,Zhao H,Zhou G Y,et al.Pq-transformer:jointly parsing 3d objects and layouts from point clouds[J].IEEE Robotics and Automation Letters,2022,7(2):2519-2526.
[3] Nie Y Y,Han X G,Guo S H,et al.Total3dunderstanding:joint layout,object pose and mesh reconstruction for indoor scenes from a single image[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,2020:55-64.
[4] Avetisyan A,Khanova T,Choy C,et al.Scenecad:predicting object alignments and layouts in rgb-d scans[C]//European Conference on Computer Vision,2020:596-612.
[5] Avetisyan A,Xie C,Howard Jenkins H,et al.Scenescript:reconstructing scenes with an autoregressive structured language model[C]//European Conference on Computer Vision,2024:247-263.
[6] Hoiem D,Efros A,Hebert M.Recovering surface layout from an image[J].International Journal of Computer Vision,2007,75(1):151-172.
[7] Mohan N,Kumar M.Room layout estimation in indoor environment:a review[J].Multimedia Tools and Applications,2022,81(2):1921-1951.
[8] Gao H,Tian B W,Li P F,et al.From semi-supervised to omni-supervised room layout estimation using point clouds[J].arXiv preprint arXiv:2301.13865,2023.
[9] Wang Z Y,Li Y L,Chen X,et al.Uni3detr:unified 3d detection transformer[C]//Advances in Neural Information Processing Systems,2023:39876-39896.
[10] Lee C Y,Badrinarayanan V,Malisiewicz T,et al.Roomnet:end-to-end room layout estimation[C]//Proceedings of the IEEE International Conference on Computer Vision,2017:4865-4874.
[11] Qi C R,Litany O,He K,et al.Deep hough voting for 3D object detection in point clouds[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision,2019:9277-9286.
[12] Zhang X S,He Y,Song C L,et al.Enhanced frustrum multi-scale VoteNet for 3D object detection in cluttered indoor scene[J].Applied Intelligence,2025,55(7):588,doi:10.1007/S10489-025-06492-4.
[13] Yan C J,Shao B Y,Zhao H,et al.3D room layout estimation from a single RGB image[J].IEEE Transactions on Multimedia,2020,22(11):3014-3024.
[14] Rukhovich D,Vorontsova A,Konushin A.Imvoxelnet:image to voxels projection for monocular and multi-view general-purpose 3D object detection[C]//Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision,2022:2397-2406.
[15] Yue Y W,Kontogianni T,Schindler K,et al.Connecting the dots:Floorplan reconstruction using two-level queries[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,2023:845-854.
[16] Mia M S,Adnan M A.Layout anything:one transformer for universal room layout estimation[C]//Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision,2026:1565-1574.
[17] Konushin A,Drozdov N,Gabdullin B,et al.TUN3D:towards real-world scene understanding from inposed images[J].arXiv preprint arXiv:2509.21388,2025.
[18] Xie Y M,Gadelha M,Yang F T,et al.Planarrecon:real-time 3D plane detection and reconstruction from posed monocular videos[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,2022:6219-6228.
[19] Cheng Y F,Yan Y K,Yi X,et al.Semanticadapt:optimization-based adaptation of mixed reality layouts leveraging virtual-physical semantic connections[C]//34th Annual ACM Symposium on User Interface Software and Technology,2021:282-297.
[20] Mehranfar M,Vega Torres M A,Braun A,et al.Automated data-driven method for creating digital building models from dense point clouds and images through semantic segmentation and parametric model fitting[J].Advanced Engineering Informatics,2024,62:102643,doi:10.1016/j.aei.2024.102643.
[21] Xia H C,Su E,Memmel M,et al.Drawer:digital reconstruction and articulation with environment realism[C]//Proceedings of the Computer Vision and Pattern Recognition Conference,2025:21771-21782.
[22] Cheng B,Sheng L,Shi S S,et al.Back-tracing representative points for voting-based 3D object detection in point clouds[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,2021:8963-8972.
[23] Zhang Z W,Sun B,Yang H T,et al.H3dnet:3D object detection using hybrid geometric primitives[C]//European Conference on Computer Vision,2020:311-329.
[24] Wang H Y,Shi S S,Yang Z,et al.Rbgnet:ray-based grouping for 3D object detection[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,2022:1110-1119.
[25] Xie Q,Lai Y K,Wu J,et al.Venet:voting enhancement network for 3D object detection[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision,2021:3712-3721.
[26] Zhu Y,Hui L,Shen Y Q,et al.Spgroup3D:superpoint grouping network for indoor 3d object detection[C]//Proceedings of the AAAI Conference on Artificial Intelligence.2024,38(7):7811-7819.
[27] Wang H Y,Ding L H,Dong S C,et al.Cagroup3D:class-aware grouping for 3D object detection on point clouds[C]//Advances in Neural Information Processing Systems,2022:29975-29988.
[28] Liu Z,Zhang Z,Cao Y,et al.Group-free 3D object detection via transformers[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision,2021:2949-2958.
[29] Misra I,Girdhar R,Joulin A.An end-to-end transformer model for 3D object detection[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision,2021:2906-2917.
[30] Kolodiazhnyi M,Vorontsova A,Skripkin M,et al.Unidet3D:Multi-dataset indoor 3D object detection[C]//Proceedings of the AAAI Conference on Artificial Intelligence,2025,39(4):4365-4373.
[31] Gwak J Y,Choy C,Savarese S.Generative sparse detection networks for 3D single-shot object detection[C]//European Conference on Computer Vision,2020:297-313.
[32] Rukhovich D,Vorontsova A,Konushin A.Fcaf3D:fully convolutional anchor-free 3D object detection[C]//European Conference on Computer Vision,2022:477-493.
[33] Rukhovich D,Vorontsova A,Konushin A.Tr3D towards real-time indoor 3D object detection[C]//IEEE International Conference on Image Processing(ICIP),2023:281-285.
[34] Mousavian A,Anguelov D,Flynn J,et al.3D bounding box estimation using deep learning and geometry[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition,2017:7074-7082.
[35] Qin Z Y,Wang J L,Lu Y.Monogrnet:a geometric reasoning network for monocular 3D object localization[C]//Proceedings of the AAAI Conference on Artificial Intelligence,2019:8851-8858.
[36] Huang S Y,Qi S Y,Xiao Y X,et al.Cooperative holistic scene understanding:unifying 3D object,layout,and camera pose estimation[C]//Advances in Neural Information Processing Systems,2018:207-218.
[37] Huang S Y,Qi S Y,Zhu Y X,et al.Holistic 3D scene parsing and reconstruction from a single rgb image[C]//Proceedings of the European Conference on Computer Vision(ECCV),2018:187-203.
[38] Wang Y,Guizilini V C,Zhang T Y,et al.Detr3D:3D object detection from multi-view images via 3D-to-2D queries[C]//Conference on Robot Learning,PMLR,2022:180-191.
[39] Liu Y F,Wang T C,Zhang X G,et al.Petr:position embedding transformation for multi-view 3D object detection[C]//European Conference on Computer Vision,2022:531-548.
[40] Huang C,Li X Y,Qu Y X,et al.NeRF-DetS:enhanced adaptive spatial-wise sampling and view-wise fusion strategies for NeRF-based indoor multi-view 3D object detection[C]//International Joint Conference on Neural Networks(IJCNN),2025:1-8.
[41] Huang C X,Hou Y N,Ye W C,et al.Nerf-det++:incorporating semantic cues and perspective-aware depth supervision for indoor multi-view 3D detection[J].IEEE Transactions on Image Processing,2025,34:2575-2587,doi:10.1109/TIP.2025.3560240.
[42] Coughlan J M,Yuille A L.Manhattan world:compass direction from a single image by bayesian inference[C]//Proceedings of the 11th IEEE International Conference on Computer Vision,1999:941-947.
[43] Hedau V,Hoiem D,Forsyth D.Recovering the spatial layout of cluttered rooms[C]//IEEE 12th International Conference on Computer Vision,2009:1849-1856.
[44] Ochmann S,Vock R,Klein R.Automatic reconstruction of fully volumetric 3D building models from oriented point clouds[J].ISPRS Journal of Photogrammetry and Remote Sensing,2019,151:251-262,doi:10.1016/j.isprs.jprs.2019.03.017.
[45] Murali S,Speciale P,Oswald M R,et al.Indoor Scan2BIM:building information models of house interiors[C]//IEEE/RSJ International Conference on Intelligent Robots and Systems(IROS),2017:6126-6133.
[46] Liu C,Wu J Y,Furukawa Y.Floornet:a unified framework for floorplan reconstruction from 3D scans[C]//Proceedings of the European Conference on Computer Vision(ECCV),2018:201-217.
[47] Zhang Y D,Song S,Tan P,et al.Panocontext:a whole-room 3D context model for panoramic scene understanding[C]//European Conference on Computer Vision,2014:668-686.
[48] Zou C H,Colburn A,Shan Q,et al.Layoutnet:reconstructing the 3D room layout from a single rgb image[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition,2018:2051-2059.
[49] Sun C,Hsiao C W,Sun M,et al.Horizonnet:learning room layout with 1d representation and pano stretch data augmentation[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,2019:1047-1056.
[50] Yang S T,Wang F E,Peng C H,et al.Dula-net:a dual-projection network for estimating room layouts from a single rgb panorama[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,2019:3363-3372.
[51] Jiang Z G,Xiang Z Z,Xu J H,et al.Lgt-net:indoor panoramic room layout estimation with geometry-aware transformer network[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,2022:1654-1663.
[52] Shen Z J,Zheng Z S,Lin C Y,et al.Disentangling orthogonal planes for indoor panoramic room layout estimation with cross-scale distortion awareness[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,2023:17337-17345.
[53] Solarte B,Wu C H,Liu Y C,et al.360-mlc:multi-view layout consistency for self-training and hyper-parameter tuning[C]//Advances in Neural Information Processing Systems,2022:6133-6146.
[54] Pu G,Zhao Y M,Lian Z H.Pano2room:novel view synthesis from a single indoor panorama[C]//SIGGRAPH Asia Conference Papers,2024:1-11.
[55] Del P L,Bowdish J,Kermgard B,et al.Understanding bayesian rooms using composite 3D object models[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition,2013:153-160.
[56] Chen Y X,Huang S Y,Yuan T,et al.Holistic++ scene understanding:single-view 3D holistic scene parsing and human pose estimation with human-object interaction and physical commonsense[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision,2019:8648-8657.
[57] Kulkarni N,Misra I,Tulsiani S,et al.3D-relnet:joint object and relational network for 3D prediction[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision,2019:2212-2221.
[58] Dai A,Chang A X,Savva M,et al.Scannet:richly-annotated 3D reconstructions of indoor scenes[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition,2017:5828-5839.
[59] Song S,Lichtenberg S P,Xiao J S.Sun rgb-d:a rgb-d scene understanding benchmark suite[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition,2015:567-576.
[60] Baruch G,Chen Z Y,Dehghan A,et al.Arkitscenes:a diverse real-world dataset for 3D indoor scene understanding using mobile rgb-d data[J].arXiv preprint arXiv:2111.08897,2021.
[61] Armeni I,Sener O,Zamir A R,et al.3D semantic parsing of large-scale indoor spaces[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition,2016:1534-1543.
[62] Sun X Y,Wu J J,Zhang X M,et al.Pix3D:dataset and methods for single-image 3D shape modeling[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition,2018:2974-2983.
[63] Meta Reality Labs,Project Aria Team,et al.Hands-on ecocentric research with project aria from meta[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR),2023.
[64] Zhang Y,Wang Y,Cui Y W,et al.3dgeodet:general-purpose geometry-aware image-based 3D object detection[J].IEEE Transactions on Multimedia,2025,27:6235-6247,doi:10.1109/TMM.2025.3581780.
[65] Li Z,Yu H,Ding Y,et al.Go-n3rdet:geometry optimized nerf-enhanced 3D object detector[C]//Proceedings of the Computer Vision and Pattern Recognition Conference,2025:27211-27221.
[66] Jiang X H,Han L J,Liu H,et al.Deep learning based 3D object detection in indoor environments:a review[C]//6th CAA International Conference on Vehicular Control and Intelligence(CVCI),2022:1-6.
[67] Wang Z Y,Li Y L,Liu T C,et al.Ov-uni3detr:towards unified open-vocabulary 3D object detection via cycle-modality propagation[C]//European Conference on Computer Vision,2024:73-89.
[68] Moliner O,Larsson V,Åström K.Sparse multiview open-vocabulary 3D detection[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision,2025:2591-2600.
[69] Kang Z Z,Yang J T,Yang Z,et al.A review of techniques for 3D reconstruction of indoor environments[J].ISPRS International Journal of Geo-Information,2020,9(5):330,doi:10.3390/ijgi9050330.
[70] Fan Z W,Zhang J,Li R J,et al.Vlm-3r:vision-language models augmented with instruction-aligned 3D reconstruction[J].arXiv preprint arXiv:2505.20279,2025.

基金资助

国家自然科学基金项目(62472078,U22A2004)资助;中央高校基本科研业务费项目(N25LPY013)资助;广东实验室人工智能与数字经济开放研究基金项目(GML-KF-24-31)资助;国家电网公司科技项目(2024YF-95)资助.

AI Summary AI Mindmap

0

访问

0

被引

详细

导航
相关文章

AI思维导图

/