Fruit detection and picking keypoint localization with state space model and geometric compensation for strawberry harvesting robots
Keywords:
strawberry detection, picking keypoint localization, STRAW-MAMBA, state space model, lightweight modelAbstract
Strawberry harvesting robots are vital for efficient harvesting and labor cost reduction. However, accurate strawberry detection and picking keypoint localization remain challenging due to dense fruit clusters, variable illumination, leaf and stem occlusions, and limited onboard computational resources. This study proposed STRAW-MAMBA, a lightweight state space model (SSM)-based network for accurate strawberry detection and keypoint localization in unconstrained environments. Specifically, the C2f-MAMBA block was developed to enhance global and multi-scale feature extraction for densely clustered or partly obscured strawberries. It redesigned the bottleneck layer of C2f as the Straw-Vim module, which integrates Hidden State Mixer-based State Space Duality (HSM-SSD) for global feature extraction and multi-scale convolutional attention (MSCA) for multi-scale feature extraction. Meanwhile, to preserve more fruit edge feature details, the large-stride convolution was optimized as a Haar wavelet downsampling module, thereby addressing illumination-induced edge blurring of strawberries while improving computational efficiency. In addition, the FasterNet block was superseded to reduce the model size further. Finally, a novel five-point geometric compensation method was proposed to address occlusions, reduce unpickable fruits, and minimize crop waste. Experiments show STRAW-MAMBA achieves 85.5% precision, 84.8% recall, and 89.7% mAP@0.5 for detection, and 88.7% precision, 77.5% recall, and 84.9% mAP@0.5 for keypoint localization, with 93.5 fps and only 1.9 M parameters. Ablation studies confirm each module’s effectiveness, demonstrating suitability for greenhouse strawberry harvesting robots. This study can provide an efficient and robust visual perception framework, significantly advancing the practical deployment of automated harvesting systems in precision agriculture.
Key words: strawberry detection; picking keypoint localization; STRAW-MAMBA; state space model; lightweight model
DOI: 10.25165/j.ijabe.20261904.10528
Citation: Kong D W, Deng Y J, Zhu X H, Hao J T, Yang X W, Zhou X C, et al. Fruit detection and picking keypointlocalization with state space model and geometric compensation for strawberry harvesting robots. Int J Agric & Biol Eng, 2026;19(4): 203–213.
References
[1] Giampieri F, Tulipani S, Alvarez-Suarez J M, Quiles J L, Mezzetti B, Battino M. The strawberry: Composition, nutritional quality, and impact on human health. Nutrition, 2012; 28(1): 9–19.
[2] Bac C W, van Henten E J, Hemming J, Edan Y. Harvesting robots for high‐value crops: State‐of‐the‐art review and challenges ahead. Journal of Field Robotics, 2014; 31(6): 888–911.
[3] Hayashi S, Shigematsu K, Yamamoto S, Kobayashi K, Kohno Y, Kamata J, et al. Evaluation of a strawberry-harvesting robot in a field test. Biosystems Engineering, 2010; 105(2): 160–171.
[4] van Henten E J, Hemming J, van Tuijl B A J, Kornet J G, Meuleman J, Bontsema J, et al. An autonomous robot for harvesting cucumbers in greenhouses. Autonomous Robots, 2002; 13(3): 241–258.
[5] Guo J, Yang Z, Karkee M, Jiang Q J, Feng X P, He Y. Technology progress in mechanical harvest of fresh market strawberries. Computers and Electronics in Agriculture, 2024; 226: 109468.
[6] Huang Z L, Wane S, Parsons S. Towards automated strawberry harvesting: Identifying the picking point. In: Towards Autonomous Robotic Systems, Springer, 2017; pp.222–236. doi:10.1007/978-3-319-64107-2_18.
[7] Hu H M, Kaizu Y, Zhang H D, Xu Y W, Imou K, Li M, et al. Recognition and localization of strawberries from 3D binocular cameras for a strawberry picking robot using coupled YOLO/Mask R-CNN. Int J Agric & Biol Eng, 2022; 15(6): 175–179.
[8] Wang J N, Tan D Z, Sui L M, Guo J, Wang R W. Wolfberry recognition and picking-point localization technology in natural environments based on improved Yolov8n-Pose-LBD. Computers and Electronics in Agriculture, 2024; 227: 109551.
[9] Zhang S Y, Hu K, Sha W, Chen Q, Hou Z M, Weng S Z. Efficient one-stage location method for grape picking points in natural scene by combining detection network and point regression. Computers and Electronics in Agriculture, 2025; 230: 109725.
[10] Peng H X, Liang Q J, Zou X J, Wang H J, Xiong J T, Luo Y L, et al. Synchronous detection method for litchi fruits and picking points of a litchi-picking robot based on improved YOLOv8-pose. Int J Agric & Biol Eng, 2025; 18(4): 266–274.
[11] Fu Y X, Zheng H C, Wang Z B, Huang J Y, Fu W. Detection of multi-class coconut clusters for robotic picking under occlusion conditions. Int J Agric & Biol Eng, 2025; 18(1): 267–278.
[12] Yu Y, Zhang K L, Yang L, Zhang D X. Fruit detection for strawberry harvesting robot in non-structural environment based on Mask-RCNN. Computers and Electronics in Agriculture, 2019; 163: 104846.
[13] Yu Y, Zhang K L, Liu H, Yang L, Zhang D X. Real-time visual localization of the picking points for a ridge-planting strawberry harvesting robot. IEEE Access, 2020; 8: 116556–116568.
[14] Xie H H, Zhang Z J, Zhang K L, Yang L, Zhang D X, Yu Y. Research on the visual location method for strawberry picking points under complex conditions based on composite models. Journal of Science of Food and Agriculture, 2024; 104(14): 8566–8579.
[15] Ma Z H, Dong N S, Gu J Y, Cheng H C, Meng Z C, Du X Q. STRAW-YOLO: A detection method for strawberry fruits targets and key points. Computers and Electronics in Agriculture, 2025; 230: 109853.
[16] Bochkovskiy A, Wang C-Y, Liao H-Y M. YOLOv4: Optimal speed and accuracy of object detection. arXiv, 2020; arXiv: 2004.10934. doi:10.48550/arXiv.2004.10934.
[17] Zhong Z, Zheng L, Kang G L, Li S Z, Yang Y. Random erasing data augmentation. In: Proceedings of the AAAI Conference on Artificial Intelligence, 2020; pp.13001–13008. doi:10.1609/aaai.v34i07.7000.
[18] Cubuk E D, Zoph B, Shlens J, Le Q V. Randaugment: Practical automated data augmentation with a reduced search space. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2020; pp.702-703. doi:10.5555/3495724.3497141.
[19] Gu A, Dao T. Mamba: Linear-time sequence modeling with selective state spaces. arXiv, 2023; arXiv: 2312.00752. doi:10.48550/arXiv.2312.00752.
[20] Dao T, Gu A. Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality. In: Ruslan S, Zico K, Katherine H, Adrian W, Nuria O, Jonathan S, et al., editors. In: Proceedings of the 41st International Conference on Machine Learning; Proceedings of Machine Learning Research: PMLR, 2024; pp.10041–10071. doi:10.48550/arXiv.2405.21060.
[21] Liu S H, Yu W H, Tan Z X, Wang X C. Linfusion: 1 GPU, 1 minute, 16 k image. arXiv, 2024; arXiv: 2409.02097. doi:10.48550/arXiv.2409.02097.
[22] Zhu L H, Liao B C, Zhang Q, Wang X L, Liu W Y, Wang X G. Vision Mamba: Efficient visual representation learning with bidirectional state space model. In: International Conference on Machine Learning: PMLR, 2024; pp.62429–42. doi:10.48550/arXiv.2401.09417.
[23] Lee S, Choi J, Kim H J. EfficientViM: Efficient vision Mamba with hidden state mixer based state space duality. In: Proceedings of the Computer Vision and Pattern Recognition Conference, 2024; pp.14923–14933. doi:10.48550/arXiv.2411.15241.
[24] Guo M-H, Lu C-Z, Hou Q, Liu Z, Cheng M-M, Hu S-M. Segnext: Rethinking convolutional attention design for semantic segmentation. In: The Thirty-Sixth Annual Conference on Neural Information Processing Systems, 2022; pp.1140–1156. doi:10.48550/arXiv.2209.08575.
[25] Xu G P, Liao W T, Zhang X, Li C, He X W, Wu X L. Haar wavelet downsampling: A simple but effective downsampling module for semantic segmentation. Pattern Recognition, 2023; 143: 109819.
[26] Chen J, Kao S-h, He H, Zhuo W, Wen S, Lee C-H, et al. Run, don’t walk: chasing higher FLOPS for faster neural networks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023; pp.12021–12031. doi:10.1109/CVPR52729.2023.01157.
[27] Maji D, Nagori S, Mathew M, Poddar D. YOLO-POSE: Enhancing YOLO for multi person pose estimation using object keypoint similarity loss. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022; pp.2637–2646. doi: 10.1109/CVPRW56347.2022.00297.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 International Journal of Agricultural and Biological Engineering

This work is licensed under a Creative Commons Attribution 4.0 International License.
IJABE is an international peer reviewed, open access journal, adopting Creative Commons Copyright Notices as follows.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).