Multi-Source Fusion and Automatic Predictor Selection for Zero-Shot Video Object Segmentation

Zhao, Xiaoqi; Pang, Youwei; Yang, Jiaxing; Zhang, Lihe; Lu, Huchuan

Computer Science > Computer Vision and Pattern Recognition

arXiv:2108.05076 (cs)

[Submitted on 11 Aug 2021]

Title:Multi-Source Fusion and Automatic Predictor Selection for Zero-Shot Video Object Segmentation

Authors:Xiaoqi Zhao, Youwei Pang, Jiaxing Yang, Lihe Zhang, Huchuan Lu

View PDF

Abstract:Location and appearance are the key cues for video object segmentation. Many sources such as RGB, depth, optical flow and static saliency can provide useful information about the objects. However, existing approaches only utilize the RGB or RGB and optical flow. In this paper, we propose a novel multi-source fusion network for zero-shot video object segmentation. With the help of interoceptive spatial attention module (ISAM), spatial importance of each source is highlighted. Furthermore, we design a feature purification module (FPM) to filter the inter-source incompatible features. By the ISAM and FPM, the multi-source features are effectively fused. In addition, we put forward an automatic predictor selection network (APS) to select the better prediction of either the static saliency predictor or the moving object predictor in order to prevent over-reliance on the failed results caused by low-quality optical flow maps. Extensive experiments on three challenging public benchmarks (i.e. DAVIS$_{16}$, Youtube-Objects and FBMS) show that the proposed model achieves compelling performance against the state-of-the-arts. The source code will be publicly available at \textcolor{red}{\url{this https URL}}.

Comments:	This work was accepted as ACM MM 2021 oral
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2108.05076 [cs.CV]
	(or arXiv:2108.05076v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2108.05076

Submission history

From: Xiaoqi Zhao [view email]
[v1] Wed, 11 Aug 2021 07:37:44 UTC (3,625 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Multi-Source Fusion and Automatic Predictor Selection for Zero-Shot Video Object Segmentation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Multi-Source Fusion and Automatic Predictor Selection for Zero-Shot Video Object Segmentation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators