YOLOv11 optimization for tiny object in crowded scenes
Abstract
Small object detection in crowded urban and aerial scenes remains a critical challenge due to limited pixel information and information loss in deep neural networks. This study introduces a novel optimization framework for YOLOv11, specifically engineered for tiny-scale targets by integrating convolutional block attention modules (CBAM), k-means anchor clustering, and an enhanced feature pyramid network (FPN). Evaluated on the TinyPerson and COCO-mini datasets, the YOLOv11-optimized model achieves significant performance breakthroughs, delivering a +7.3% gain in mean average precision (mAP) and a +10.5% increase in recall over the baseline. Notably, the model achieved a recall of 0.072 on the TinyPerson dataset, with double sensitivity of standard YOLOv11. With a high-speed inference rate of 27.3 FPS, this research demonstrates that strategic architectural refinements can drastically improve small object detection reliability without compromising real-time viability on edge devices.
Keywords
Anchor box optimization; Convolutional block attention module; Data augmentation; Feature pyramid network; Tiny object detection; TinyPerson dataset; YOLOv11
Full Text:
PDFDOI: http://doi.org/10.11591/ijai.v15.i4.pp3452-3463
Refbacks
- There are currently no refbacks.
Copyright (c) 2026 Husna Sarirah Husin, Howard Chong Yun Hao, Pan Yuan Fei, Mohsen Marjani, Suriana Ismail

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
IAES International Journal of Artificial Intelligence (IJ-AI)
ISSN/e-ISSN 2089-4872/2252-8938
This journal is published by the Institute of Advanced Engineering and Science (IAES).