IMPROVING THE ROBUSTNESS OF SIAMESE NETWORK DETECTION WITH LIMITED ANNOTATION THROUGH HARD EXAMPLE MINING
Abstract
The aim of this study is to develop and experimentally evaluate a method for improving the quality of marine vessel detection based on the Attention RPN architecture under conditions of limited training data, with the integration of inter-class hard negative mining. To achieve this goal, the following were implemented: model adaptation in a fine-tuning mode based on five examples, creation of a specialized dataset consisting of 9 vessel categories (637 training and 175 validation images), and an algorithm for selecting hard negative pairs using cosine similarity. The methodological foundation includes deep learning approaches for object detection based on Siamese neural networks, analysis of feature distributions in embedding space, and algorithms for generating hard negative samples by computing inter-class distances using the cosine metric. Experimental evaluation was carried out on standard few-shot object detection benchmarks. The integration of cross-class hard negative mining reduced the proportion of false positive detections by approximately 18% and increased the mean Average Precision (mAP) by about 6% on average compared to the baseline model without hard negative mining. The greatest improvement (up to 9.1%) was observed for classes with large target objects. The results demonstrate that targeted selection of inter-class hard negative examples significantly improves the robustness of metric-based detectors under conditions of strong class imbalance and high intra-class similarity. The practical significance of the work lies in the applicability of the proposed approach for deploying video surveillance systems in tasks where large training datasets and extensive annotation are unavailable
References
1. Zhao J., Masood R., Seneviratne S. A Review of Computer Vision Methods in Network Security, IEEE Communications Surveys & Tutorials, 2021, Vol. PP, pp. 1-1. doi: 10.1109/COMST.2021.3086475.
2. Casas E., Ramos L.T., Romero C., Rivas-Echeverría F. A Review of Computer Vision Applications for Asset Inspection in the Oil and Gas Industry, Journal of Pipeline Science and Engineering, 2025, Vol. 5(3), pp. 100246. doi: 10.1016/j.jpse.2024.100246.
3. Zhou L., Zhang L., Konz N. Computer Vision Techniques in Manufacturing, IEEE Transactions on Systems, Man, and Cybernetics: Systems, 2022, Vol. 53 (1), pp. 105-117. doi: 10.1109/TSMC.2022.3166397.
4. Suguna A. Computer Vision for Autonomous Driving, International Journal of Innovative Research in Information Security, 2025, Vol. 11, pp. 846-850. doi: 10.26562/ijiris.2025.v1112.03.
5. Dilek E., Dener M. Computer Vision Applications in Intelligent Transportation Systems: A Survey, Sensors, 2023, Vol. 23 (6), pp. 2938. doi: 10.3390/s23062938.
6. Kabir M.M., Rahman A., Hasan M.N., Mridha M.F. Computer Vision Algorithms in Healthcare: Recent Advancements and Future Challenges, Computers in Biology and Medicine, 2025, Vol. 185,
pp. 109531. doi: 10.1016/j.compbiomed.2024.109531.
7. Gao J., Yang Y., Lin P., Park D. Computer Vision in Healthcare Applications, Journal of Healthcare Engineering, 2018, pp. 1-4. doi: 10.1155/2018/5157020.
8. Ren S., He K., Girshick R., Sun J. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks, Advances in Neural Information Processing Systems (NeurIPS), 2015, pp. 91-99.
9. Cong X., Li S., Chen F., Liu C., Meng Y. A Review of YOLO Object Detection Algorithms Based on Deep Learning, Frontiers in Computing and Intelligent Systems, 2023, Vol. 4(2), pp. 17-20. doi: 10.54097/fcis.v4i2.9730.
10. Lei M., Li S., Wu Y., Hu H., Zhou Y., Zheng X., Ding G., Du S., Wu Z., Gao Y. YOLOv13: Real-Time Object Detection with Hypergraph-Enhanced Adaptive Visual Perception, arXiv, 2025. doi: 10.48550/arXiv.2506.17733.
11. Carion N., Massa F., Synnaeve G., Usunier N., Kirillov A., Zagoruyko S. End-to-End Object Detection with Transformers, European Conference on Computer Vision (ECCV). Springer, 2020, pp. 213-229. doi: 10.1007/978-3-030-58452-8_13.
12. Shehzadi T., Hashmi K.A., Liwicki M., Stricker D., Afzal M.Z. Object Detection with Transformers:
A Review, Sensors, 2025, Vol. 25, pp. 6025. doi: 10.3390/s25196025.
13. Lin T.Y., Maire M., Belongie S., Hays J., Perona P., Ramanan D., Dollár P., Zitnick C.L. Microsoft COCO: Common Objects in Context, European Conference on Computer Vision (ECCV). Springer, 2014, pp. 740-755.
14. Everingham M., Van Gool L., Williams C.K.I., Winn J., Zisserman A. The Pascal Visual Object Classes (VOC) Challenge, International Journal of Computer Vision, 2010, Vol. 88 (2), pp. 303-338.
15. Kohler M., Eisenbach M., Gross H.M. Few-Shot Object Detection: A Comprehensive Survey, IEEE Transactions on Neural Networks and Learning Systems, 2023, Vol. PP. doi: 10.1109/TNNLS.2023.3265051.
16. Fan Q., Zhuo W., Tang C.K., Tai C.L. Few-Shot Object Detection with Attention-RPN and Multi-Relation Detector, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni-tion (CVPR), 2020, pp. 4013-4022.
17. OpenMMLab. MMFewShot: A Toolbox for Few Shot Learning Detection and Classification. Available at: https://mmfewshot.readthedocs.io/en/latest/ (accessed 14 January 2026).
18. Girshick R. Fast R-CNN, Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2015, pp. 1440–1448.
19. He K., Zhang X., Ren S., Sun J. Deep Residual Learning for Image Recognition, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770-778. doi: 10.1109/CVPR.2016.90.
20. Robbins H.E. A Stochastic Approximation Method, Annals of Mathematical Statistics, 1951, Vol. 22, pp. 400-407.








