KP-Mask2Former: A fish occlusion contour recovery method based on KANsformer and PCBAM

Authors

  • Yunfeng Wang College of Computer Science and Artificial Intelligence, Ludong University, Yantai 264025, Shandong, China
  • Beibei Li College of Information and Electrical Engineering, China Agricultural University, Beijing 100083, China
  • Jun Yue College of Computer Science and Artificial Intelligence, Ludong University, Yantai 264025, Shandong, China
  • Yixun Zhao College of Computer Science and Artificial Intelligence, Ludong University, Yantai 264025, Shandong, China
  • Lei Yuan College of Computer Science and Artificial Intelligence, Ludong University, Yantai 264025, Shandong, China
  • Peiheng Li College of Computer Science and Artificial Intelligence, Ludong University, Yantai 264025, Shandong, China
  • Guangjie Kou College of Computer Science and Artificial Intelligence, Ludong University, Yantai 264025, Shandong, China

Keywords:

Fish, instance segmentation, occlusion recovery, attention mechanism, deep learning

Abstract

Recovering complete contours of occluded fish is crucial for accurate fish size estimation and automated counting. Nevertheless, frequent and severe occlusions in high-density aquaculture environments make reliable contour recovery particularly challenging. To address this problem, this study proposes KP-Mask2Former, a fish contour recovery framework specifically designed for occlusion-intensive scenarios. First, a Parallel Convolutional Block Attention Module (PCBAM) is introduced into the patch-merging stages of the backbone to strengthen the representation of fish contour features. Second, Swin KANsformer is incorporated into the backbone to enhance the reconstruction of incomplete boundaries caused by occlusion. Third, a SEAM-Mask Head is developed to suppress interference from occluding instances and improve the recovery of complete fish contours. The proposed framework is evaluated on a self-constructed Occluded Oplegnathus punctatus Dataset (OOD). Experimental results demonstrate that KP-Mask2Former achieves an AP of 87.1%, representing a 4.3-percentage-point improvement over the baseline. It also outperforms representative state-of-the-art methods, including PolySnake, E2EC, PatchDCT, BCNet, Transfiner, MP-Former, and FastInst. Further analyses confirm the superiority of the proposed method in boundary delineation, intra-class occlusion handling, and multi-scale instance segmentation.      

Key words: fish; instance segmentation; occlusion recovery; attention mechanism; deep learning

DOI: 10.25165/j.ijabe.20261904.10285

Citation: Wang Y F, Li B B, Yue J, Zhao Y X, Yuan L, Li P H, et al. KP-Mask2Former: A fish occlusion contour recovery
method based on KANsformer and PCBAM. Int J Agric & Biol Eng, 2026; 19(4): 235–242.

References

[1] Costa C, Loy A, Cataudella S, Davis D, Scardi M. Extracting fish size using dual underwater cameras. Aquacultural Engineering, 2006; 35(3): 218–227.

[2] Guo Y, Aggrey S E, Oladeinde A, Johnson J, Chai L. A machine vision-based method optimized for restoring broiler chicken images occluded by feeding and drinking equipment. Animals, 2021; 11(1): 123.

[3] Huang E, Mao A, Hou J, Wu Y, Xu W, Ceballos M-C, et al. Occlusion-resistant instance segmentation of piglets in farrowing pens using center clustering network. Computers and Electronics in Agriculture, 2023; 210: 107950.

[4] Meng H, Jin S, Liu W T, Qian C, Lin M X, Ouyang W L, et al. 3D interacting hand pose estimation by hand de-occlusion and removal. Proceedings of the European Conference on Computer Vision, 2022; pp.380–397. DOI: 10.1007/978-3-031-20068-7_22

[5] Ju Y-J, Lee G-H, Hong J-H, Lee S-W. Complete face recovery GAN: Unsupervised joint face rotation and de-occlusion from a single-view image. Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 2022; pp.3711–3721. doi: 10.1109/WACV51458.2022.00124

[6] Yin X, Huang D, Fu Z, Wang Y, Chen L. Segmentation-reconstruction-guided facial image de-occlusion. Proceedings of the 2023 IEEE 17th International Conference on Automatic Face and Gesture Recognition (FG), 2023; pp.1–8. doi: 10.1109/FG57933.2023.10042570

[7] Zhou Q, Wang S Y, Wang Y T, Huang Z L, Wang X G. Human de-occlusion: Invisible perception and recovery for humans. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021; pp.3691–3701. doi: 10.1109/CVPR46437.2021.00369

[8] Li L, Zhang T F, Kang Z F, Jiang X K. Mask-FPAN: semi-supervised face parsing in the wild with de-occlusion and UV GAN. Computers & Graphics, 2023; 116: 185–193.

[9] Zhan X, Pan X, Dai B, Liu Z, Lin D, Loy C-C. Self-supervised scene de-occlusion. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020; pp.3784–3792. doi: 10.1109/CVPR42600.2020.00384

[10] Huang J Z, Yu X, An D, Ning X, Liu J C, Tiwari P. Uniformity and deformation: A benchmark for multi-fish real-time tracking in the farming. Expert Systems with Applications, 2025; 264: 125653.

[11] Ghiasi G, Cui Y, Srinivas A, Qian R, Lin T-Y, Cubuk E-D, et al. Simple copy-paste is a strong data augmentation method for instance segmentation. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021; pp.2918–2928. doi: 10.1109/CVPR46437.2021.00294

[12] Cheng B, Misra I, Schwing A G, Kirillov A, Girdhar R. Masked-attention mask transformer for universal image segmentation. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022: 1290–1299. doi: 10.1109/cvpr52688.2022.00135

[13] Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez A N, et al. Attention is all you need. Advances in Neural Information Processing Systems, 2017; 30.

[14] He K M, Zhang X Y, Ren S Q, Sun J. Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2016; pp.770–778. doi: 10.1109/CVPR.2016.90

[15] Liu Z, Lin Y T, Cao Y, Hu H, Wei Y X, Zhang Z, et al. Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021; pp.10012–10022. doi: 10.1109/ICCV48922.2021.00986.

[16] Zhu X Z, Su W J, Lu L W, Li B, Wang X G, Dai J F. Deformable detr: Deformable transformers for end-to-end object detection. arXiv 2020. DOI: 10.48550/arXiv.2010.04159.

[17] Woo S, Park J, Lee J-Y, Kweon I S. Cbam: Convolutional block attention module. Proceedings of the European Conference on Computer Vision (ECCV), 2018; pp.3–19. doi: 10.1007/978-3-030-01234-2_1

[18] Hu J, Shen L, Sun G. Squeeze-and-excitation networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018; pp.7132–7141. doi: 10.1109/TPAMI.2019.2913372

[19] Liu Z, Wang Y, Vaidya S, Ruehle F, Halverson J, Soljačić M, et al. KAN: Kolmogorov-arnold networks. International Conference on Learning Representations, 2025; pp.70367–70413.

[20] Chollet F. Xception: Deep learning with depthwise separable convolutions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017; pp.1251–1258. DOI: 10.1109/CVPR.2017.195

[21] Yu Z, Huang H, Chen W, Su Y, Liu Y, Wang X. YOLO-Facev2: A scale and occlusion aware face detector. Pattern Recognition, 2024; 155: 10.

[22] Feng H, Zhou K Y, Zhou W G, Yin Y F, Deng J J, Sun Q, et al. Recurrent generic contour-based instance segmentation with progressive learning. IEEE Transactions on Circuits and Systems for Video Technology, 2024. https://arxiv.org/html/2301.08898v2

[23] Zhang T, Wei S Q, Ji S P. E2EC: An end-to-end contour-based method for high-quality high-speed instance segmentation. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022; pp.4443–4452. DOI: 10.48550/arXiv.2203.04074

[24] Wen Q R, Yang J R, Yang X, Liang K W. Patchdct: Patch refinement for high quality instance segmentation. Proceedings of the Eleventh International Conference on Learning Representations, 2022. DOI: 10.48550/arXiv.2302.02693

[25] Ke L, Tai Y-W, Tang C-K. Deep occlusion-aware instance segmentation with overlapping bilayers. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021; pp.4019–4028. doi: 10.1109/CVPR46437.2021.00401

[26] Ke L, Danelljan M, Li X, Tai Y-W, Tang C-K, Yu F. Mask transfiner for high-quality instance segmentation. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022; pp.4412–4421. doi: 10.1109/CVPR52688.2022.00437

[27] Zhang H, Li F, Xu H Z, Huang S J, Liu S L, Ni L M, et al. Mp-Former: Mask-piloted transformer for image segmentation. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023; pp.18074–18083. doi: 10.1109/CVPR52729.2023.01733

[28] He J J, Li P Y, Geng Y F, Xie X S. Fastinst: A simple query-based model for real-time instance segmentation. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023; pp.23663–23672. doi: 10.1109/CVPR52729.2023.02266

[29] Wang X Y, Yu H H, Zhang C, Liu Z N, Luo Z, Chen Y Y, et al. An underwater image dataset for occlusion-aware fish instance segmentation. Scientific Data, 2026; 13(1). doi: 10.1038/s41597-026-06898-w

[30] Li Y, Yang H H, Xing H R, Fu L Y, Gu N R, Liu S E. A Transformer-Mamba framework for cross-regional long-term forecasting of key water quality in aquaculture. Water Research, 2026; 302: 126135.

Downloads

Published

2026-09-03

How to Cite

(1)
Wang, Y.; Li, B.; Yue, J.; Zhao, Y.; Yuan, L.; Li, P.; Kou, G. KP-Mask2Former: A Fish Occlusion Contour Recovery Method Based on KANsformer and PCBAM. Int J Agric & Biol Eng 2026, 19, 235-242.

Issue

Section

Information Technology, Sensors and Control Systems