computer_science3 papersavg year 2025weak evidence

We presented a systematic study on small object detection

Research gap analysis derived from 3 computer_science papers in our local library.

The gap

We presented a systematic study on small object detection. Concretely, we exhaustively reviewed hundreds of literature for SOD from the perspective of algorithms and datasets. Moreover, to catalyze the progress of SOD, we constructed two la

Evidence profile

Sourced from the future work of the source papers, classified as general, drawn from work published between 2023 and 2026, spanning 3 journals. Those papers have been cited 480 times in total.

Research trend

Established — well-defined area with open sub-problems.

Supporting evidence — 3 representative gaps

  • Convolutional low-rank adaptation for efficient semantic segmentation in vision transformers (2026) · Scientific Reports · doi

    This work has demonstrated the effectiveness of Low-Rank Adaptation for Convolutions (LoCON) in fine- tuning complex vision models for downstream tasks. By integrating only 150k trainable parameters into DAV2’s decoder, binary human segmentation has been achieved with performance competitive with specialized segmentation models, while maintaining the model’s original depth estimation capabilities. The experimental results on COCO and ImageNet datasets validate that the proposed approach strikes an optimal balance between performance and efficiency. The model achieves an mAP of 0.8969 and mIoU of 0.7917, approaching state-of-the-art performance while using only a fraction of trainable parameters compared to full fine-tuning. This efficiency makes the approach particularly valuable for resource-constrained environments where computational overhead is a critical concern. The success of LoCON adaptation for convolutional networks suggests that similar parameter-efficient techniques could be effective across a broader range of vision architectures beyond transformer-based models. This opens promising avenues for democratizing access to state-of-the-art computer vision capabilities in scenarios where computational resources are limited. Scientific Reports | (2026) 16:27030 | https://doi.org/10.1038/s41598-026-57622-y 9 Future work could explore the application of this methodology to other vision tasks such as object detection and instance segmentation, as well as investigating the scalability of LoCON adaptations to even larger foundation models. Additionally, examining the combination of LoCON with complementary techniques like knowledge distillation could further enhance efficiency while preserving performance. The ability to retain the original model’s depth estimation capabilities while seamlessly adding segmentation functionality demonstrates the flexibility and modularity of the LoCON approach. Rather than requiring the deployment of separate models for different tasks, this unified solution significantly reduces deployment complexity and resource consumption. As multimodal models become increasingly prevalent, this work presents a compelling case for lightweight task adaptation without sacrificing general-purpose utility – enabling broader adoption of foundation models in constrained settings such as mobile devices, robotics, and edge computing platforms. Furthermore, the use of small, isolated subsets of trainable parameters for different downstream tasks makes it feasible to dynamically switch between capabilities–such as segmentation, depth estimation, or even pose detection–by simply loading the corresponding LoCON weights. This modularity facilitates real-time adaptation in multi-task systems without the need to reinitialize or fully retrain the core model, making it ideal for embedded AI applications, on-device personalization, and scalable multi-user deployments where each task or user may require a specialized behavior without incurring additional memory or compute overhead.

    generalfuture work
    Keywords: models locon segmentation adaptation vision tasks performance model capabilities trainable parameters depth estimation approach efficiency
  • Lightweight vision transformer for adversarially robust object detection in autonomous vehicles (2026) · Gümüşhane Üniversitesi Fen Bilimleri Enstitüsü Dergisi · doi

    This paper presented a robust and efficient object detection framework for autonomous driving that integrates a lightweight vision transformer backbone, a Deformable DETR detection head, detection-aware adversarial training, and semi-supervised learning into a unified architecture. The proposed L-ViT framework was motivated by three persistent challenges in autonomous vehicle perception: the trade-off between detection accuracy and computational efficiency, the vulnerability of deep learning models to adversarial perturbations, and the high cost of large-scale annotated driving data. By jointly addressing these challenges, the proposed approach offers a practical and deployable solution for safety-critical intelligent transportation systems. Extensive experiments conducted on the KITTI, BDD100K, and Cityscapes benchmarks demonstrate that the proposed method achieves a favorable balance between accuracy, robustness, and efficiency. Compared to classical baselines such as Faster R-CNN, YOLOv4, and standard DETR, the proposed framework delivers competitive or superior detection accuracy while maintaining real-time inference capability on edge devices. 799 Okul, 2026 • Volume 16 • Issue 3 • Page 789-801 In particular, the lightweight backbone combined with deformable attention enables effective detection of small and distant objects in dense urban scenes, which is essential for safe autonomous navigation. A key contribution of this work is the explicit integration of adversarial robustness into the object detection pipeline. The proposed detection-aware adversarial consistency regularization significantly reduces performance degradation under both digital and physically inspired attacks. Experimental results show that the proposed model experiences an accuracy drop of approximately 13% as shown in Table 3 under strong adversarial perturbations, compared to drops exceeding 25% for conventional detectors. These findings highlight the importance of robustness-aware design in safety-critical autonomous driving systems. Furthermore, the incorporation of semi-supervised learning improves generalization across challenging environmental conditions, yielding a notable mAP improvement over variants trained solely on labeled data. Despite these promising results, several limitations remain. First, although simulated digital and physical perturbations provide a broad robustness assessment, large-scale real-world physical attack evaluations would offer stronger validation. Second, while the proposed framework achieves real-time performance on current edge devices, scaling to higher-resolution inputs or multi-modal perception may require additional optimization. Finally, the effectiveness of semi-supervised learning depends on the quality of pseudo- labels, which may introduce noise under highly ambiguous conditions. Future work will focus on extending the proposed framework in several directions. First, integrating multi- modal sensor data, such as LiDAR and radar, is expected to further enhance robustness under poor visibility and severe occlusion. Second, exploring self-supervised or foundation-model-based pretraining could reduce reliance on extensive labeled datasets. Third, additional model compression techniques, including pruning, quantization, and neural architecture search, will be investigated to support ultra-low-power embedded platforms. Finally, real-world deployment and evaluation in autonomous driving systems will be pursued to bridge the gap between simulation and operational environments. In conclusion, this work advances the state of the art in autonomous driving perception by demonstrating that accuracy, efficiency, and adversarial robustness can be jointly achieved within a single lightweight detection framework. The proposed L-ViT architecture provides a solid foundation for developing reliable and secure perception systems in next-generation intelligent transportation applications.

    generalfuture work
    Keywords: detection proposed framework autonomous adversarial robustness driving accuracy supervised learning perception systems real lightweight aware
  • Towards Large-Scale Small Object Detection: Survey and Benchmarks (2023) · IEEE Transactions on Pattern Analysis and Machine Intelligence · cited 480× · doi

    We presented a systematic study on small object detection. Concretely, we exhaustively reviewed hundreds of literature for SOD from the perspective of algorithms and datasets. Moreover, to catalyze the progress of SOD, we constructed two large-scale benchmarks under driving scenario and aerial scene, dubbed SODA-D and SODA-A. SODA-D com- prises 278433 instances annotated with horizontal boxes, while SODA-A includes 872069 objects with oriented boxes. The well-annotated datasets, to the best of our knowledge, are the first attempt to large-scale benchmarks tailored for small object detection, and could serve as an impartial platform for benchmarking various SOD methods. On top of SODA, we performed a thorough evaluation and com- parison of several representative algorithms. Based on the results, we discuss several potential solutions and directions for future development of SOD task. Effective feature extractor for small objects. As alluded to in the results, deeper backbone networks might not be conducive to extract high-quality feature representations for small objects. Designing an effective backbone, which enjoys powerful feature extraction capability while avoiding high computational cost and information loss, is of paramount importance. High-quality hierarchical representation. FPN is an indispensable part in small object detection. Nevertheless, current feature pyramid architecture is suboptimal for SOD, owing to the heuristic pyramid level assignment strategy, few samples were assigned to higher levels (actually only P2 feature is responsible to the detection during our bench- mark experiments). Consequently, the high-level layers are optimized in an implicit and indirect manner which may hamper the fusion quality. Moreover, detecting on low-level feature maps brings heavy computational burden. Thus, an efficient hierarchical feature architecture tailored for SOD task is in high demand. Optimized label assignment strategy. As we discussed in Sec. 2.3.1 and Sec. B.2 of Appendix, albeit the current label assignment schemes perform well on generic object detec- tion and large objects, they still struggle on the instances of extremely small sizes, neither the overlap-based strategies nor the distribution-based ones. Therefore, designing an optimized strategy to assign sufficient positive samples for size-limited instances can substantially stabilize the training procedure and boost the performance further. Proper evaluation metric for SOD. The multiple IoU thresholds-based evaluation process has been the de facto standard for validating the effectiveness of methods in generic object detection. However, such ubiquitous metric is too stringent for those instances with extremely sizes. In other words, the top priority of small object detection under some specific scenarios is to recognize the objects and obtain their rough locations instead of obsessing how accurate they are. Hence, it is impractical to pursue pre- cise detections of small objec

    generalfuture work
    Keywords: small feature object detection soda objects high instances based large evaluation quality level assignment strategy

Questions about this gap

We presented a systematic study on small object detection. Concretely, we exhaustively reviewed hundreds of literature for SOD from the perspective of algorithms and datasets. More… This is supported by 3 representative gap statements extracted from 3 papers, rated weak evidence.

Explore this gap further

Run this gap as a query across open scholarly engines for the latest related literature.

Working on this gap? Review it with us.

AI Review reads your manuscript in one pass with 8 specialist agents, calibrated on 69K+ real peer reviews.

Related gaps in Computer Science

Command palette

Jump anywhere, run any action.