Edge-oriented lightweight attention-enhanced Faster R-CNN: A systematic study of innovative methods for garbage object detection
Municipal solid waste is an increasing challenge to urban management systems all over the world. The pressure has spurred intelligent garbage detection and automatic sorting, but three challenges remain: data heterogeneity across domains, low accuracy of small-object detection in complex scenes, and limited computing resources on edge devices. To overcome these issues, this paper introduces a lightweight attention-enhanced faster region-based convolutional neural network framework for deployment on the edge. The TACO and RealWaste datasets are then combined to create a cross-domain benchmark of 1,000 images across ten unified categories, and the image annotation alignment rate is increased to 96.8% using category remapping and multi-strategy augmentation at the data level. A convolutional block attention module is added to the feature pyramid network neck to further improve feature selectivity, while replacing ResNet-50 with MobileNetV3 as the backbone. For focal loss optimization, the number of parameters is only 9.2 M, and the model achieves an mAP@50 of 71.3%, which is 9.9 percentage points higher than the baseline. The quantized model is deployed to RK3588 via RKNN Toolkit2, and it achieves 35.2 FPS with an INT8 configuration at 3.1 W, which is an 89.2% increase in inference speed compared to the Jetson Nano. The proposed method offers a competitive balance between detection accuracy, model size, and intelligent edge performance, thus providing a feasible technical path for the intelligent waste sorting system.
Ahmad, Z. S. M., Aziz, A. A. N., Lim, S. H., Hasanuddin, Z. S., Hadi, N. A., Hussin, M. S., & Mohd Zain, N. (2025). Impact of image enhancement using contrast-limited adaptive histogram equalization (CLAHE), anisotropic diffusion, and histogram equalization on spine X-ray segmentation with U-Net, Mask R-CNN, and transfer learning. Algorithms, 18(12), Article 796. https://doi.org/10.3390/a18120796
Akintola, G. A., Obiwusi, Y. K., & Olatunde, O. Y. (2025). Integrated deep learning paradigm for comprehensive lung cancer segmentation and classification using Mask R-CNN and CNN models. Franklin Open, 11, Article 100278. https://doi.org/10.1016/j.fraope.2025.100278
Alfred, R., Leo, J., & Kaijage, F. S. (2025). Detectron2-enhanced Mask R-CNN for precise instance segmentation of rice blast disease in Tanzania: Supporting timely intervention and data-driven severity assessment. Smart Agricultural Technology, 12, Article 101301. https://doi.org/10.1016/j.atech.2025.101301
Amin, J., Gul, N., & Sharif, M. (2025). Dual-method for semantic and instance brain tumor segmentation based on UNet and Mask R-CNN using MRI. Neural Computing and Applications, 37(14), 1–19. https://doi.org/10.1007/s00521-025-11013-y
Arias, R. F. J., Castaño, M. J. M., & Espriella, L. D. M. F. M. (2025). An application of deep learning models for the detection of cocoa pods at different ripening stages: An approach with Faster R-CNN and Mask R-CNN. Computation, 13(7), Article 159. https://doi.org/10.3390/computation13070159
Banik, D. (2025). Enhanced detection and segmentation of ground-glass opacities in SARS-CoV-2 patients using Mask R-CNN on chest CT images. Multimedia Tools and Applications, 84(31), 1–16. https://doi.org/10.1007/s11042-025-20720-6
Bashkirova, D., Abdelfattah, M., Zhu, Z., Sherrill, Z., Calli, B., Shi, B., & Saenko, K. (2022). ZeroWaste dataset: Towards deformable object segmentation in cluttered scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 21298–21307). IEEE. https://doi.org/10.1109/cvpr52688.2022.02047
Bejo, K. S., Ibrahim, F. M., & Hanafi, M. (2024). Automatic paddy planthopper detection and counting using Faster R-CNN. Agriculture, 14(9), Article 1567. https://doi.org/10.3390/agriculture14091567
Bhalla, S., Kushwaha, R., & Kumar, A. (2025). HydR-CNN: Advancing underwater object detection using a multi-stage framework with hybrid R-CNN and pyramid vision transformer with augmented convolution. The European Physical Journal Plus, 140(6), Article 556. https://doi.org/10.1140/epjp/s13360-025-06498-4
Brandao, I., Fidalgo, B., & González, S. R. (2025). Storm damage and planting success assessment in Pinus pinaster Aiton stands using Mask R-CNN. Forests, 16(11), Article 1730. https://doi.org/10.3390/f16111730
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., & Zagoruyko, S. (2020). End-to-end object detection with transformers. In A. Vedaldi, H. Bischof, T. Brox, & J.-M. Frahm (Eds.), Computer Vision – ECCV 2020 (pp. 213–229). Springer. https://doi.org/10.1007/978-3-030-58452-8_13
Chai, S., Gao, P., & Li, M. (2025). Mask R-CNN for predicting rib fractures on CT images with interpretability and ChatGPT-based structured outcomes. Expert Systems with Applications, 274, Article 127047. https://doi.org/10.1016/j.eswa.2025.127047
Chakraborty, D., Saha, S., & Halder, B. (2025). Learning to localize image forgery using boundary-preserving Mask R-CNN. Journal of Forensic Sciences, 71(1), 388–404. https://doi.org/10.1111/1556-4029.70203
Chansuparp, M., Pansawat, N., & Wangvoralak, S. (2025). Automated grading of boiled shrimp by color level using image processing techniques and Mask R-CNN with feature pyramid networks. Applied Sciences, 15(19), Article 10632. https://doi.org/10.3390/app151910632
Chau, V. T., Jung, S., & Kim, M. (2025). Indirect estimation of seagrass frontal area for coastal protection: A Mask R-CNN and dual-reference approach. Journal of Marine Science and Engineering, 13(7), Article 1262. https://doi.org/10.3390/jmse13071262
Cheng, G., Yuan, X., Yao, X., Yan, K., Zeng, Q., Xie, X., & Han, J. (2023). Towards large-scale small object detection: Survey and benchmarks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(11), 13467–13488. https://doi.org/10.1109/tpami.2023.3290594
Crespo, A., Moncada, C., & Crespo, F. (2025). An efficient strawberry segmentation model based on Mask R-CNN and TensorRT. Artificial Intelligence in Agriculture, 15(2), 327–337. https://doi.org/10.1016/j.aiia.2025.01.008
Dandıl, E., Baştuğ, T. B., & Yıldırım, S. M. (2024). MaskAppendix: Backbone-enriched Mask R-CNN based on Grad-CAM for automatic appendix segmentation. Diagnostics, 14(21), Article 2346. https://doi.org/10.3390/diagnostics14212346
David, G., & Faure, E. (2025). End-to-end 3D instance segmentation of synthetic data and embryo microscopy images with a 3D Mask R-CNN. Frontiers in Bioinformatics, 4, Article 1497539. https://doi.org/10.3389/fbinf.2024.1497539
Dutta, M., Das, K. U., & Datta, N. (2026). Multilingual scene text recognition: A Faster R-CNN approach for Bengali and English scripts. Expert Systems with Applications, 302, Article 130482. https://doi.org/10.1016/j.eswa.2025.130482
Elhosseini, A. M., Sayed, A. H., & Agamy, E. F. R. (2025). A hybrid YOLOv10–Faster R-CNN framework for mobility-aid detection and traffic optimization in disability-inclusive smart cities. Alexandria Engineering Journal, 129, 1279–1298. https://doi.org/10.1016/j.aej.2025.08.044
Gholinavaz, S., Saeedi, N., & Gharehveran, S. S. (2025). Robustness analysis of YOLO and Faster R-CNN for object detection in realistic weather scenarios with noise augmentation. Scientific Reports, 15(1), Article 44888. https://doi.org/10.1038/s41598-025-28737-5
He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 770–778). IEEE. https://doi.org/10.1109/cvpr.2016.90
Heda, L., & Sahare, P. (2025). Design of an iterative method for crowd behavior analysis integrating Faster R-CNN, YOLOv8, and graph convolutional networks. Signal, Image and Video Processing, 19(7), Article 553. https://doi.org/10.1007/s11760-025-04127-2
Howard, A., Sandler, M., Chu, G., Chen, L.-C., Chen, B., Tan, M., Wang, W., Zhu, Y., Pang, R., Vasudevan, V., Le, Q. V., & Adam, H. (2019). Searching for MobileNetV3. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 1314–1324). IEEE. https://doi.org/10.1109/iccv.2019.00140
Hu, J., Shen, L., & Sun, G. (2018). Squeeze-and-excitation networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 7132–7141). IEEE. https://doi.org/10.1109/cvpr.2018.00745
Huang, R., Ding, J., & Ren, Z. (2025). Research on the segmentation of individual trees and the extraction of structural parameters in eucalyptus plantations based on a TEMA Mask R-CNN model. Journal of Forest Research, 30(4), 316–327. https://doi.org/10.1080/13416979.2025.2479873
Islam, T., Sarker, T. T., & Ahmed, R. K. (2024). Detection and classification of cannabis seeds using RetinaNet and Faster R-CNN. Seeds, 3(3), 456–478. https://doi.org/10.3390/seeds3030031
Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., & Kalenichenko, D. (2018). Quantization and training of neural networks for efficient integer-arithmetic-only inference. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 2704–2713). IEEE. https://doi.org/10.1109/cvpr.2018.00286
Kaviani, M., Leblon, B., & Akilan, T. (2025). Tree health assessment using Mask R-CNN on UAV multi-spectral imagery over apple orchards. Remote Sensing, 17(19), Article 3369. https://doi.org/10.3390/rs17193369
Kriouile, Y., Ancourt, C., & Wolska, W. K. (2024). Nested object detection using Mask R-CNN: Application to bee and varroa detection. Neural Computing and Applications, 36(35), 1–23. https://doi.org/10.1007/s00521-024-10393-x
Kulambayev, B., & Olzhayev, O. (2025). A Mask R-CNN algorithm for automated segmentation of asphalt road cracks. Procedia Computer Science, 269, 39–48. https://doi.org/10.1016/j.procs.2025.08.257
Kwon, K., Im, K. S., & Kim, Y. S. (2024). Estimation of tree diameter at breast height from aerial photographs using a Mask R-CNN and Bayesian regression. Forests, 15(11), Article 1881. https://doi.org/10.3390/f15111881
Lei, M., Zhang, M., & Wang, K. (2026). Research on bidding optimization strategy for virtual power plants with wind–solar–storage systems based on IGDT–DRO. Electrical Engineering, 108(3), Article 208. https://doi.org/10.1007/s00202-026-03544-x
Lin, T.-Y., Dollár, P., Girshick, R., He, K., Hariharan, B., & Belongie, S. (2017a). Feature pyramid networks for object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 936–944). IEEE. https://doi.org/10.1109/cvpr.2017.106
Lin, T.-Y., Goyal, P., Girshick, R., He, K., & Dollár, P. (2017b). Focal loss for dense object detection. In Proceedings of the IEEE International Conference on Computer Vision (ICCV) (pp. 2980–2988). IEEE. https://doi.org/10.1109/iccv.2017.324
Lv, X., Cui, S., Wang, Y., Lu, J., Yu, P., & Wang, K. (2026). Patch time series transformer–based short-term photovoltaic power prediction enhanced by artificial fish. Energies, 19(1), Article 284. https://doi.org/10.3390/en19010284
Marzouk, A. M., & Elkholy, M. (2025). Tracking and estimating the speed of vehicles using an enhanced R-CNN and Deep SORT architecture. Journal of Advances in Information Technology, 16(2), 204–214. https://doi.org/10.12720/jait.16.2.204-214
Nazir, A., & Wani, A. M. (2025). Multi-scale feature enhancement using EfficientNet-B7 and PANet in Faster R-CNN for small object detection. International Journal of Information Technology, 18(1), 1–8. https://doi.org/10.1007/s41870-025-02790-9
Paściak, A., Piwowarczyk, K. P., & Iskander, R. D. (2025). Enhancing meibography based assessment of gland morphology by utilizing an image-rotating Mask R-CNN approach. Biomedical Signal Processing and Control, 109, Article 108045. https://doi.org/10.1016/j.bspc.2025.108045
Prasad, R., Saxena, K., & Maurya, A. (2025). DermoAI-CNN: Leveraging GANs, Mask R-CNN, and attention mechanisms for enhanced skin disease analysis. Foundations of Computing and Decision Sciences, 50(4), 537–569. https://doi.org/10.2478/fcds-2025-0021
Proença, P. F., & Simões, P. (2020). TACO: Trash Annotations in Context for Litter Detection (Version 2). arXiv. https://doi.org/10.48550/ARXIV.2003.06975
Rajendran, V., Subramaniyan, M. U., & Suriyan, K. (2025). Enhanced R-CNN model for traffic sign recognition under diverse environmental conditions. International Journal of Intelligent Transportation Systems Research, 23(2), 1–18. https://doi.org/10.1007/s13177-025-00476-x
Reddy, N. T., Kumar, N., & Ponnappa, P. N. (2025). Intelligent GD&T symbol detection in mechanical drawings: A comparative study of YOLOv11, Faster R-CNN, and RetinaNet for quality assurance. Journal of Intelligent Manufacturing. Advance online publication. https://doi.org/10.1007/s10845-025-02669-3
Reddy, A. S., & G., M. P. (2025). Skin cancer detection using optimized Mask R-CNN and two-fold deep learning classifier framework. Multimedia Tools and Applications, 84(30), 1–28. https://doi.org/10.1007/s11042-024-20377-7
Ren, S., He, K., Girshick, R., & Sun, J. (2017). Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(6), 1137–1149. https://doi.org/10.1109/tpami.2016.2577031
Rifat, S. M. K., Nayeem, J. M., Neon, M. N. M., & Nur-A-Alam, M. (2026). A deep learning-based performance-oriented approach for diabetic foot ulcer identification and localization using Faster R-CNN. In Silico Research in Biomedicine, Article 100202. https://doi.org/10.1016/j.insi.2026.100202
Ritha, N., Hayaty, N., & Apdillah, D. (2025). Deep object detection approaches for identification of seagrass species using Faster R-CNN and SSD. IOP Conference Series: Earth and Environmental Science, 1559(1), Article 012012. https://doi.org/10.1088/1755-1315/1559/1/012012
Salvador, F. A. C., Bilyk, T., & Dartois, A. (2025). High-throughput analysis of dislocation loops in irradiated metals using Mask R-CNN. Micron, 201, Article 103927. https://doi.org/10.1016/j.micron.2025.103927
Shobaki, A. W., & Milanova, M. (2025). A comparative study of YOLO, SSD, Faster R-CNN, and more for optimized eye-gaze writing. Sci, 7(2), Article 47. https://doi.org/10.3390/sci7020047
Single, S., Iranmanesh, S., & Raad, R. (2023). RealWaste: A novel real-life data set for landfill waste classification using deep learning. Information, 14(12), Article 633. https://doi.org/10.3390/info14120633
Sinha, M., Paul, S., & Lala, N. G. M. (2025). Comparative analysis of Mask R-CNN and U-Net architectures using ResNet as backbone for lunar crater detection. Planetary and Space Science, 264, Article 106140. https://doi.org/10.1016/j.pss.2025.106140
Sun, S., Zheng, S., Xu, X., & He, Z. (2025). GD-YOLO: A lightweight model for household waste image detection. Expert Systems with Applications, 279, Article 127525. https://doi.org/10.1016/j.eswa.2025.127525
Wang, B., Liu, D., & Wu, J. (2025). SSNFNet: An enhanced few-shot learning model for efficient poultry farming detection. Animals, 15(15), Article 2252. https://doi.org/10.3390/ani15152252
Wang, D., Shi, L., & Li, Y. (2025). An enhanced Faster R-CNN for high-throughput winter wheat spike monitoring to improve yield prediction and water use efficiency. Agronomy, 15(10), Article 2388. https://doi.org/10.3390/agronomy15102388
Wang, Z. W., Sun, R., & Wang, Z. (2025). Jiyu Faster R-CNN de maolian zitai shibie [Anchor chain attitude recognition based on Faster R-CNN]. Zhongguo Jixie [China Machinery], (28), 33–37. [In Chinese]
Woo, S., Park, J., Lee, J.-Y., & Kweon, I. S. (2018). CBAM: Convolutional block attention module. In V. Ferrari, M. Hebert, C. Sminchisescu, & Y. Weiss (Eds.), Computer Vision – ECCV 2018 (pp. 3–19). Springer. https://doi.org/10.1007/978-3-030-01234-2_1
Xie, R. J., & Ren, R. X. (2025). Jiyu Faster R-CNN de jianzhuwu yaogan tuxiang mubiao jiance [Remote sensing image object detection of buildings based on Faster R-CNN]. Xinxi Jilu Cailiao [Information Recording Materials], 26(9), 113–115. [In Chinese] https://doi.org/10.16009/j.cnki.cn13-1295/tq.2025.09.067
Yang, W. (2025). Jiyu gaijin Mask R-CNN de dianli shebei rege guzhang zhenduan yanjiu [Research on thermal fault diagnosis of power equipment based on improved Mask R-CNN]. Jiamusi Daxue Xuebao (Ziran Kexue Ban) [Journal of Jiamusi University (Natural Science Edition)], 43(7), 61–64. [In Chinese] https://doi.org/10.20232/j.cnki.jmsdxxb.2025.07.034
Zhu, X., Su, W., Lu, L., Li, B., Wang, X., & Dai, J. (2021). Deformable DETR: Deformable transformers for end-to-end object detection. In 9th International Conference on Learning Representations (ICLR 2021). OpenReview. https://openreview.net/forum?id=gZ9hCDWe6ke
Zou, Z., Chen, K., Shi, Z., Guo, Y., & Ye, J. (2023). Object detection in 20 years: A survey. Proceedings of the IEEE, 111(3), 257–276. https://doi.org/10.1109/jproc.2023.3238524
