KP-Mask2Former: A fish occlusion contour recovery method based on KANsformer and PCBAM
Keywords:
Fish, instance segmentation, occlusion recovery, attention mechanism, deep learningAbstract
Recovering complete contours of occluded fish is crucial for accurate fish size estimation and automated counting. Nevertheless, frequent and severe occlusions in high-density aquaculture environments make reliable contour recovery particularly challenging. To address this problem, this study proposes KP-Mask2Former, a fish contour recovery framework specifically designed for occlusion-intensive scenarios. First, a Parallel Convolutional Block Attention Module (PCBAM) is introduced into the patch-merging stages of the backbone to strengthen the representation of fish contour features. Second, Swin KANsformer is incorporated into the backbone to enhance the reconstruction of incomplete boundaries caused by occlusion. Third, a SEAM-Mask Head is developed to suppress interference from occluding instances and improve the recovery of complete fish contours. The proposed framework is evaluated on a self-constructed Occluded Oplegnathus punctatus Dataset (OOD). Experimental results demonstrate that KP-Mask2Former achieves an AP of 87.1%, representing a 4.3-percentage-point improvement over the baseline. It also outperforms representative state-of-the-art methods, including PolySnake, E2EC, PatchDCT, BCNet, Transfiner, MP-Former, and FastInst. Further analyses confirm the superiority of the proposed method in boundary delineation, intra-class occlusion handling, and multi-scale instance segmentation.
Key words: fish; instance segmentation; occlusion recovery; attention mechanism; deep learning
DOI: 10.25165/j.ijabe.20261904.10285
Citation: Wang Y F, Li B B, Yue J, Zhao Y X, Yuan L, Li P H, et al. KP-Mask2Former: A fish occlusion contour recovery
method based on KANsformer and PCBAM. Int J Agric & Biol Eng, 2026; 19(4): 235–242.
References
[1] Costa C, Loy A, Cataudella S, Davis D, Scardi M. Extracting fish size using dual underwater cameras. Aquacultural Engineering, 2006; 35(3): 218–227.
[2] Guo Y, Aggrey S E, Oladeinde A, Johnson J, Chai L. A machine vision-based method optimized for restoring broiler chicken images occluded by feeding and drinking equipment. Animals, 2021; 11(1): 123.
[3] Huang E, Mao A, Hou J, Wu Y, Xu W, Ceballos M-C, et al. Occlusion-resistant instance segmentation of piglets in farrowing pens using center clustering network. Computers and Electronics in Agriculture, 2023; 210: 107950.
[4] Meng H, Jin S, Liu W T, Qian C, Lin M X, Ouyang W L, et al. 3D interacting hand pose estimation by hand de-occlusion and removal. Proceedings of the European Conference on Computer Vision, 2022; pp.380–397. DOI: 10.1007/978-3-031-20068-7_22
[5] Ju Y-J, Lee G-H, Hong J-H, Lee S-W. Complete face recovery GAN: Unsupervised joint face rotation and de-occlusion from a single-view image. Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 2022; pp.3711–3721. doi: 10.1109/WACV51458.2022.00124
[6] Yin X, Huang D, Fu Z, Wang Y, Chen L. Segmentation-reconstruction-guided facial image de-occlusion. Proceedings of the 2023 IEEE 17th International Conference on Automatic Face and Gesture Recognition (FG), 2023; pp.1–8. doi: 10.1109/FG57933.2023.10042570
[7] Zhou Q, Wang S Y, Wang Y T, Huang Z L, Wang X G. Human de-occlusion: Invisible perception and recovery for humans. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021; pp.3691–3701. doi: 10.1109/CVPR46437.2021.00369
[8] Li L, Zhang T F, Kang Z F, Jiang X K. Mask-FPAN: semi-supervised face parsing in the wild with de-occlusion and UV GAN. Computers & Graphics, 2023; 116: 185–193.
[9] Zhan X, Pan X, Dai B, Liu Z, Lin D, Loy C-C. Self-supervised scene de-occlusion. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020; pp.3784–3792. doi: 10.1109/CVPR42600.2020.00384
[10] Huang J Z, Yu X, An D, Ning X, Liu J C, Tiwari P. Uniformity and deformation: A benchmark for multi-fish real-time tracking in the farming. Expert Systems with Applications, 2025; 264: 125653.
[11] Ghiasi G, Cui Y, Srinivas A, Qian R, Lin T-Y, Cubuk E-D, et al. Simple copy-paste is a strong data augmentation method for instance segmentation. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021; pp.2918–2928. doi: 10.1109/CVPR46437.2021.00294
[12] Cheng B, Misra I, Schwing A G, Kirillov A, Girdhar R. Masked-attention mask transformer for universal image segmentation. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022: 1290–1299. doi: 10.1109/cvpr52688.2022.00135
[13] Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez A N, et al. Attention is all you need. Advances in Neural Information Processing Systems, 2017; 30.
[14] He K M, Zhang X Y, Ren S Q, Sun J. Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2016; pp.770–778. doi: 10.1109/CVPR.2016.90
[15] Liu Z, Lin Y T, Cao Y, Hu H, Wei Y X, Zhang Z, et al. Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021; pp.10012–10022. doi: 10.1109/ICCV48922.2021.00986.
[16] Zhu X Z, Su W J, Lu L W, Li B, Wang X G, Dai J F. Deformable detr: Deformable transformers for end-to-end object detection. arXiv 2020. DOI: 10.48550/arXiv.2010.04159.
[17] Woo S, Park J, Lee J-Y, Kweon I S. Cbam: Convolutional block attention module. Proceedings of the European Conference on Computer Vision (ECCV), 2018; pp.3–19. doi: 10.1007/978-3-030-01234-2_1
[18] Hu J, Shen L, Sun G. Squeeze-and-excitation networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018; pp.7132–7141. doi: 10.1109/TPAMI.2019.2913372
[19] Liu Z, Wang Y, Vaidya S, Ruehle F, Halverson J, Soljačić M, et al. KAN: Kolmogorov-arnold networks. International Conference on Learning Representations, 2025; pp.70367–70413.
[20] Chollet F. Xception: Deep learning with depthwise separable convolutions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017; pp.1251–1258. DOI: 10.1109/CVPR.2017.195
[21] Yu Z, Huang H, Chen W, Su Y, Liu Y, Wang X. YOLO-Facev2: A scale and occlusion aware face detector. Pattern Recognition, 2024; 155: 10.
[22] Feng H, Zhou K Y, Zhou W G, Yin Y F, Deng J J, Sun Q, et al. Recurrent generic contour-based instance segmentation with progressive learning. IEEE Transactions on Circuits and Systems for Video Technology, 2024. https://arxiv.org/html/2301.08898v2
[23] Zhang T, Wei S Q, Ji S P. E2EC: An end-to-end contour-based method for high-quality high-speed instance segmentation. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022; pp.4443–4452. DOI: 10.48550/arXiv.2203.04074
[24] Wen Q R, Yang J R, Yang X, Liang K W. Patchdct: Patch refinement for high quality instance segmentation. Proceedings of the Eleventh International Conference on Learning Representations, 2022. DOI: 10.48550/arXiv.2302.02693
[25] Ke L, Tai Y-W, Tang C-K. Deep occlusion-aware instance segmentation with overlapping bilayers. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021; pp.4019–4028. doi: 10.1109/CVPR46437.2021.00401
[26] Ke L, Danelljan M, Li X, Tai Y-W, Tang C-K, Yu F. Mask transfiner for high-quality instance segmentation. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022; pp.4412–4421. doi: 10.1109/CVPR52688.2022.00437
[27] Zhang H, Li F, Xu H Z, Huang S J, Liu S L, Ni L M, et al. Mp-Former: Mask-piloted transformer for image segmentation. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023; pp.18074–18083. doi: 10.1109/CVPR52729.2023.01733
[28] He J J, Li P Y, Geng Y F, Xie X S. Fastinst: A simple query-based model for real-time instance segmentation. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023; pp.23663–23672. doi: 10.1109/CVPR52729.2023.02266
[29] Wang X Y, Yu H H, Zhang C, Liu Z N, Luo Z, Chen Y Y, et al. An underwater image dataset for occlusion-aware fish instance segmentation. Scientific Data, 2026; 13(1). doi: 10.1038/s41597-026-06898-w
[30] Li Y, Yang H H, Xing H R, Fu L Y, Gu N R, Liu S E. A Transformer-Mamba framework for cross-regional long-term forecasting of key water quality in aquaculture. Water Research, 2026; 302: 126135.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 International Journal of Agricultural and Biological Engineering

This work is licensed under a Creative Commons Attribution 4.0 International License.
IJABE is an international peer reviewed, open access journal, adopting Creative Commons Copyright Notices as follows.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).