Research Paper:
Moving Object Detection from Moving Camera Using Focus of Expansion Likelihood and Segmentation
Masahiro Ogawa*,
, Qi An**
, and Atsushi Yamashita**

*Department of Precision Engineering, Graduate School of Engineering, The University of Tokyo
7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, Japan
Corresponding author
**Department of Human and Engineered Environmental Studies, Graduate School of Frontier Sciences, The University of Tokyo
Kashiwa, Japan
Separating moving and static objects from a moving camera viewpoint is essential for 3D reconstruction, autonomous navigation, and scene understanding in robotics. Existing approaches often rely primarily on optical flow, which struggles to detect moving objects in complex, structured scenes involving camera motion. To address this limitation, we propose Focus of Expansion Likelihood and Segmentation (FoELS), a method based on the core idea of integrating both optical flow and texture information. FoELS computes the focus of expansion (FoE) from optical flow and derives an initial motion likelihood from the outliers of the FoE computation. This likelihood is then fused with a segmentation-based prior to estimate the final moving probability. The method effectively handles challenges including complex structured scenes, rotational camera motion, and parallel motion. Comprehensive evaluations on the DAVIS 2016 and FBMS-59 datasets, along with real-world traffic videos including parallel, cross-direction, opposite-direction, and crowded scenes, demonstrate its effectiveness and state-of-the-art performance.
Sample result of FoELS (our method)
- [1] W. Zhang, X. Sun, and Q. Yu, “Moving object detection under a moving camera via background orientation reconstruction,” Sensors, Vol.20, No.11, Article No.3103, 2020. https://doi.org/10.3390/s20113103
- [2] Z. Hu, K. Uchimura, and S. Kawaji, “Determining motion parameters for vehicle-mounted camera using focus of expansion,” IEEJ Trans. on Industry Applications, Vol.119, No.1, pp. 50-57, 1999 (in Japanese). https://doi.org/10.1541/ieejias.119.50
- [3] Y. Yang, A. Loquercio, D. Scaramuzza, and S. Soatto, “Unsupervised moving object detection via contextual information separation,” Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 879-888, 2019. https://doi.org/10.1109/CVPR.2019.00097
- [4] Z. Hu and K. Uchimura, “Multiple moving objects detection and simultaneous tracking from the time-varied background,” IEEJ Trans. on Industry Applications, Vol.120, No.10, pp. 1134-1142, 2000 (in Japanese). https://doi.org/10.1541/ieejias.120.1134
- [5] A. Dosovitskiy, P. Fischer, E. Ilg, P. Häusser, C. Hazırbas, V. Golkov, P. van der Smagt, D. Cremers, and T. Brox, “FlowNet: Learning optical flow with convolutional networks,” Proc. of the IEEE Int. Conf. on Computer Vision (ICCV), pp. 2758-2766, 2015. https://doi.org/10.1109/ICCV.2015.316
- [6] D. Sun, X. Yang, M.-Y. Liu, and J. Kautz, “PWC-Net: CNNs for optical flow using pyramid, warping, and cost volume,” Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 8934-8943, 2018. https://doi.org/10.1109/CVPR.2018.00931
- [7] Z. Teed and J. Deng, “RAFT: Recurrent all-pairs field transforms for optical flow,” Proc. of the European Conf. on Computer Vision (ECCV), pp. 402-419, 2020. https://doi.org/10.1007/978-3-030-58536-5_24
- [8] J. Chung, C. Gulcehre, K. Cho, and Y. Bengio, “Empirical evaluation of gated recurrent neural networks on sequence modeling,” NIPS 2014 Deep Learning and Representation Learning Workshop, 2014.
- [9] S. Jiang, D. Campbell, Y. Lu, H. Li, and R. Hartley, “Learning to estimate hidden motions with global motion aggregation,” Proc. of the IEEE/CVF Int. Conf. on Computer Vision (ICCV), pp. 9752-9761, 2021. https://doi.org/10.1109/ICCV48922.2021.00963
- [10] Z. Huang, X. Shi, C. Zhang, Q. Wang, K. C. Cheung, H. Qin, J. Dai, and H. Li, “FlowFormer: A transformer architecture for optical flow,” Proc. of the European Conf. on Computer Vision (ECCV), pp. 668-685, 2022. https://doi.org/10.1007/978-3-031-19790-1_40
- [11] H. Shi, Y. Zhou, K. Yang, X. Yin, and K. Wang, “CSFlow: Learning optical flow via cross strip correlation for autonomous driving,” 2022 IEEE Intelligent Vehicles Symp. (IV), pp. 1851-1858, 2022. https://doi.org/10.1109/IV51971.2022.9827341
- [12] X. Shi, Z. Huang, W. Bian, D. Li, M. Zhang, K. C. Cheung, S. See, H. Qin, J. Dai, and H. Li, “VideoFlow: Exploiting temporal cues for multi-frame optical flow estimation,” Proc. of the IEEE/CVF Int. Conf. on Computer Vision (ICCV), pp. 12435-12446, 2023. https://doi.org/10.1109/ICCV51070.2023.01146
- [13] H. Xu, J. Zhang, J. Cai, H. Rezatofighi, F. Yu, D. Tao, and A. Geiger, “Unifying flow, stereo and depth estimation,” IEEE Trans. on Pattern Analysis and Machine Intelligence, Vol.45, No.11, pp. 13941-13958, 2023. https://doi.org/10.1109/TPAMI.2023.3298645
- [14] Q. Dong and Y. Fu, “MemFlow: Optical flow estimation and prediction with memory,” Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2024. https://doi.org/10.1109/CVPR52733.2024.01804
- [15] A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Dollár, and R. Girshick, “Segment anything,” Proc. of the IEEE/CVF Int. Conf. on Computer Vision (ICCV), pp. 3992-4003, 2023. https://doi.org/10.1109/ICCV51070.2023.00371
- [16] L. Ke, M. Ye, M. Danelljan, Y. Liu, Y.-W. Tai, C.-K. Tang, and F. Yu, “Segment anything in high quality,” Advances in Neural Information Processing Systems (NeurIPS), 2023.
- [17] P. Wang, S. Wang, J. Lin, S. Bai, X. Zhou, J. Zhou, X. Wang, and C. Zhou, “ONE-PEACE: Exploring one general representation model toward unlimited modalities,” arXiv preprint, arXiv:2305.11172, 2023. https://doi.org/10.48550/arXiv.2305.11172
- [18] Y. Yuan, X. Chen, and J. Wang, “Object-contextual representations for semantic segmentation,” Proc. of the European Conf. on Computer Vision (ECCV), pp. 173-190, 2020. https://doi.org/10.1007/978-3-030-58539-6_11
- [19] Z. Zhang, H. Cai, and S. Han, “EfficientViT-SAM: Accelerated segment anything model without performance loss,” Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 7859-7863, 2024. https://doi.org/10.1109/CVPRW63382.2024.00782
- [20] H. Cai, J. Li, M. Hu, C. Gan, and S. Han, “Efficientvit: Lightweight multi-scale attention for high-resolution dense prediction,” Proc. of the IEEE/CVF Int. Conf. on Computer Vision (ICCV), pp. 17302-17313, 2023. https://doi.org/10.1109/ICCV51070.2023.01587
- [21] W. Wang, J. Dai, Z. Chen, Z. Huang, Z. Li, X. Zhu, X. Hu, T. Lu, L. Lu, H. Li, X. Wang, and Y. Qiao, “InternImage: Exploring large-scale vision foundation models with deformable convolutions,” Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 14408-14419, 2023. https://doi.org/10.1109/CVPR52729.2023.01385
- [22] N. Ravi, V. Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafson, E. Mintun, J. Pan, K. V. Alwala, N. Carion, C.-Y. Wu, R. Girshick, P. Dollár, and C. Feichtenhofer, “SAM 2: Segment anything in images and videos,” Int. Conf. on Learning Representations (ICLR), 2025.
- [23] N. Carion, L. Gustafson, Y.-T. Hu, S. Debnath, R. Hu, D. Suris et al., “SAM 3: Segment anything with concepts,” Int. Conf. on Learning Representations (ICLR), 2026.
- [24] B. Cheng, I. Misra, A. G. Schwing, A. Kirillov, and R. Girdhar, “Masked-attention mask transformer for universal image segmentation,” Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 1280-1289, 2022. https://doi.org/10.1109/CVPR52688.2022.00135
- [25] J. Jain, J. Li, M. Chiu, A. Hassani, N. Orlov, and H. Shi, “OneFormer: One transformer to rule universal image segmentation,” Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2023. https://doi.org/10.1109/CVPR52729.2023.00292
- [26] D. Rozumnyi, J. Matas, F. Sroubek, M. Pollefeys, and M. R. Oswald, “FMODetect: Robust detection of fast moving objects,” Proc. of the IEEE/CVF Int. Conf. on Computer Vision (ICCV), pp. 3521-3529, 2021. https://doi.org/10.1109/iccv48922.2021.00352
- [27] M.-N. Chapel and T. Bouwmans, “Moving objects detection with a moving camera: A comprehensive review,” Computer Science Review, Vol.38, Article No.100310, 2020. https://doi.org/10.1016/j.cosrev.2020.100310
- [28] B. Hou, Y. Liu, N. Ling, Y. Ren, and L. Liu, “A survey of efficient deep learning models for moving object segmentation,” APSIPA Trans. on Signal and Information Processing, Vol.12, No.1, pp. 1-84, 2023. https://doi.org/10.1561/116.00000140
- [29] X. Zhao, G. Wang, Z. He, and H. Jiang, “A survey of moving object detection methods: A practical perspective,” Neurocomputing, Vol.503, pp. 28-48, 2022. https://doi.org/10.1016/j.neucom.2022.06.104
- [30] J. J. Gibson, “The Perception of the Visual World,” Houghton Mifflin, 1950.
- [31] G. Rahmon, F. Bunyak, G. Seetharaman, and K. Palaniappan, “Motion U-Net: Multi-cue encoder-decoder network for motion segmentation,” Proc. of the 2020 25th Int. Conf. on Pattern Recognition (ICPR), pp. 8125-8132, 2021. https://doi.org/10.1109/ICPR48806.2021.9413211
- [32] H. Seong, S. W. Oh, J.-Y. Lee, S. Lee, S. Lee, and E. Kim, “Hierarchical memory matching network for video object segmentation,” Proc. of the IEEE/CVF Int. Conf. on Computer Vision (ICCV), 2021.
- [33] Z. Yang, Y. Wei, and Y. Yang, “Associating objects with transformers for video object segmentation,” Advances in Neural Information Processing Systems (NeurIPS), Vol.34, pp. 2491-2502, 2021.
- [34] S. W. Oh, J.-Y. Lee, N. Xu, and S. J. Kim, “Video object segmentation using space-time memory networks,” Proc. of the IEEE/CVF Int. Conf. on Computer Vision (ICCV), pp. 9225-9234, 2019. https://doi.org/10.1109/ICCV.2019.00932
- [35] C. Homeyer and C. Schnörr, “On moving object segmentation from monocular video with transformers,” Proc. of the IEEE/CVF Int. Conf. on Computer Vision Workshops (ICCVW), pp. 880-891, 2023.
- [36] Z. Teed and J. Deng, “RAFT-3D: Scene flow using rigid-motion embeddings,” Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 8375-8384, 2021. https://doi.org/10.1109/CVPR46437.2021.00827
- [37] R. Ranftl, A. Bochkovskiy, and V. Koltun, “Vision transformers for dense prediction,” Proc. of the IEEE/CVF Int. Conf. on Computer Vision (ICCV), pp. 12179-12188, 2021. https://doi.org/10.1109/ICCV48922.2021.01196
- [38] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 770-778, 2016. https://doi.org/10.1109/CVPR.2016.90
- [39] K. Izumida, K. Shiiya, H. Takahashi, and S. Derrouich, “Moving objects detection from travelling monocular camera image,” IEEJ Trans. on Electronics, Information and Systems, Vol.122, No.3, pp. 498-505, 2002 (in Japanese). https://doi.org/10.1541/ieejeiss1987.122.3_498
- [40] S. Negahdaripour and B. K. P. Horn, “A direct method for locating the focus of expansion,” Computer Vision, Graphics, and Image Processing, Vol.46, No.3, pp. 303-326, 1989. https://doi.org/10.1016/0734-189X(89)90035-2
- [41] F. Perazzi, J. Pont-Tuset, B. McWilliams, L. Van Gool, M. Gross, and A. Sorkine-Hornung, “A benchmark dataset and evaluation methodology for video object segmentation,” Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2016. https://doi.org/10.1109/CVPR.2016.85
- [42] P. Ochs, J. Malik, and T. Brox, “Segmentation of moving objects by long term video analysis,” IEEE Trans. on Pattern Analysis and Machine Intelligence, Vol.36, No.6, pp. 1187-1200, 2014. https://doi.org/10.1109/TPAMI.2013.242
This article is published under a Creative Commons Attribution-NoDerivatives 4.0 Internationa License.