Open research questions in Advanced Neural Network Applications
58 unresolved questions extracted from the limitations and future-work sections of 554 Advanced Neural Network Applications papers in our library. Each links back to the study that raised it.
What the literature leaves open
Future research will focus on three primary directions: (1) domain-specific fine-tuning of the SqueezeNet classifier using a curated dataset of cropped union regions to elevate prediction confidence; (2) integration of temporal tracking algorithms to enable continuous motion verification across video frames; and (3) deployment and benchmarking on physical edge hardware (e.
A Hybrid YOLOv5-SqueezeNet Framework with Crop-Union Context Verification for Real-Time Motorcycle Rider Identification · 2026 · DOIThis paper presented a robust and efficient object detection framework for autonomous driving that integrates a lightweight vision transformer backbone, a Deformable DETR detection head, detection-aware adversarial training, and semi-supervised learning into a unified architecture. The proposed L-ViT framework was motivated by three persistent challenges in autonomous vehicle perception: the trade-off between detection accuracy and computational efficiency, the vulnerability of deep learning models to adversarial perturbations, and the high cost of large-scale annotated driving data. By jointly addressing these challenges, the proposed approach offers a practical and deployable solution for safety-critical intelligent transportation systems. Extensive experiments conducted on the KITTI, BDD100K, and Cityscapes benchmarks demonstrate that the proposed method achieves a favorable balance between accuracy, robustness, and efficiency. Compared to classical baselines such as Faster R-CNN, YOLOv4, and standard DETR, the proposed framework delivers competitive or superior detection accuracy while maintaining real-time inference capability on edge devices. 799 Okul, 2026 • Volume 16 • Issue 3 • Page 789-801 In particular, the lightweight backbone combined with deformable attention enables effective detection of small and distant objects in dense urban scenes, which is essential for safe autonomous navigation. A key contribution of this work is the explicit integration of adversarial robustness into the object detection pipeline. The proposed detection-aware adversarial consistency regularization significantly reduces performance degradation under both digital and physically inspired attacks. Experimental results show that the proposed model experiences an accuracy drop of approximately 13% as shown in Table 3 under strong adversarial perturbations, compared to drops exceeding 25% for conventional detectors. These findings highlight the importance of robustness-aware design in safety-critical autonomous driving systems. Furthermore, the incorporation of semi-supervised learning improves generalization across challenging environmental conditions, yielding a notable mAP improvement over variants trained solely on labeled data. Despite these promising results, several limitations remain. First, although simulated digital and physical perturbations provide a broad robustness assessment, large-scale real-world physical attack evaluations would offer stronger validation. Second, while the proposed framework achieves real-time performance on current edge devices, scaling to higher-resolution inputs or multi-modal perception may require additional optimization. Finally, the effectiveness of semi-supervised learning depends on the quality of pseudo- labels, which may introduce noise under highly ambiguous conditions. Future work will focus on extending the proposed framework in several directions. First, integrating multi- modal sensor data, such as LiDAR and radar, is expected to further enhance robustness under poor visibility and severe occlusion. Second, exploring self-supervised or foundation-model-based pretraining could reduce reliance on extensive labeled datasets. Third, additional model compression techniques, including pruning, quantization, and neural architecture search, will be investigated to support ultra-low-power embedded platforms. Finally, real-world deployment and evaluation in autonomous driving systems will be pursued to bridge the gap between simulation and operational environments. In conclusion, this work advances the state of the art in autonomous driving perception by demonstrating that accuracy, efficiency, and adversarial robustness can be jointly achieved within a single lightweight detection framework. The proposed L-ViT architecture provides a solid foundation for developing reliable and secure perception systems in next-generation intelligent transportation applications.
Lightweight vision transformer for adversarially robust object detection in autonomous vehicles · 2026 · DOISecond, the study was limited to a single benchmark dataset, and the relative performance of the models may differ on more complex datasets such as CIFAR-10 or ImageNet [8],[11], [12].
A Comparative Study of Deep Learning Models for Image Classification: Simple MLP, Deep MLP, Basic CNN, LeNet CNN · 2026 · DOIFuture research may consider leveraging more efficient train- ing algorithms or optimizing the network architecture to enhance computational efficiency. Future studies may focus on incorpo- rating more contextual information or exploring multimodal learning techniques to improve the model’s adaptability to various clothing types in complex environments.
Garment Image Segmentation Method Based on Wavelet Multi-Scale Features and Improved U-Net · 2026 · DOIFuture work may explore lightweight transformer architectures, hybrid CNN-transformer models, explainable artificial intelligence techniques, and larger-scale datasets to further improve classification efficiency, robustness, and real- time deployment capability.
Vision Transformer-Based Dog Breed Classification with a Hybrid Detection-Classification Framework · 2026 · DOIIndonesian airspace to surveillance context. The remainder of this paper is organized as follows: Section 2 presents a literature review on YOLO architectures and aerial object detection; Section 3 describes the research methodology including the dataset, experimental setup, and evaluation metrics; Section 4 presents the experimental results and discussion; and Section 5 provides conclusions and directions for future research. relevant the II. LITERATURE REVIEW A. Airspace Surveillance and ISR Operations Airspace surveillance represents an integral component of national defense systems aimed at detecting, identifying, and assessing the movement of aerial objects within a country's sovereign territory. Within the operational framework of modern defense, such activities are conducted through the Intelligence, Surveillance, and Reconnaissance (ISR) concept, which integrates three key functions: intelligence as the collection and analysis of strategic information, surveillance as continuous monitoring of specific targets, and reconnaissance as active exploration of operational areas. Contemporary ISR systems integrate diverse sensor platforms, ranging from ground-based radar, reconnaissance aircraft, Unmanned Aerial Vehicles (UAVs), to high-resolution satellite imagery, in order to produce comprehensive situational awareness For Indonesia, implementing ISR presents specific challenges due to its extensive archipelagic geography, the existence of radar coverage gaps in several strategic regions, and the limitations of conventional defense assets, all of which drive the need for new artificial intelligence-based technological approaches to complement existing detection systems. The utilization of machine learning technology on aerial imagery data is regarded as a promising solution because it can produce automated detection capabilities that are fast, consistent, and operable on a continuous basis at a scale that is difficult for human operators to achieve. B. Deep Learning-Based Object Detection Methodologically, deep learning-based object detection approaches can be categorized into two main groups: twostage detectors such as Faster R-CNN and Mask R-CNN, the which separate classification stage, and one-stage detectors such as YOLO, region proposal stage from the TEKNIKA, Volume 15(2), July 2026, pp. 225-233 ISSN 2549-8037, EISSN 2549-8045 DOI: 10.34148/teknika.v15i2.1483 Suganda, F. S. et al.: Comparative Analysis of YOLOv5, YOLOv8, and YOLOv11 for Military Aircraft Detection on Aerial Imagery Supporting Airspace 227 SSD, and RetinaNet, which perform bounding box and class predictions simultaneously in a single forward pass.
Comparative Analysis of YOLOv5, YOLOv8, and YOLOv11 for Military Aircraft Detection on Aerial Imagery Supporting Airspace Surveillance · 2026 · DOIto produce empirical research aims state-of-the-art YOLO architectures in literature the national the most optimal YOLO architecture regarding for implementation in military aircraft detection systems based on aerial imagery, thereby supporting the development of automated airspace surveillance systems within the framework of strengthening Indonesia's national defense.
Comparative Analysis of YOLOv5, YOLOv8, and YOLOv11 for Military Aircraft Detection on Aerial Imagery Supporting Airspace Surveillance · 2026 · DOIThis study elucidates a novel cross-domain zero-shot framework tailored for semantic segmentation in non-geometric terrains, essentially augmenting the cross-modal alignment proficiency of the EVA-CLIP architecture. The methodology Table 6. Comparison with state-of-the-art methods on Rellis-3D dataset.
Cross-domain zero-shot semantic segmentation for unstructured environments via EVA-CLIP model, ensemble prompt engineering, and optimized text-image matching · 2026 · DOIIn particular, developing fully sparse training methods, where dense tensors are never materialized, is a major open problem [27] that STen allows researchers to make progress on.
Building upon the limitations dis- cussed, future work will focus on exploring geograph- ically separated data splits to rigorously evaluate zero- shot generalization, extending the framework to multi- modal remote sensing data, and developing lightweight architectures for efficient real-time deployment.
AN IMPROVED EVIT NETWORK FOR SEMANTIC SEGMENTATION OF HIGH-RESOLUTION REMOTE SENSING IMAGERY · 2026 · DOIWhile the proposed EViT-based architecture demonstrates strong performance, a few limitations re- main for future exploration. First, regarding the experimental design, our study strictly utilizes the official data splits (training, valida- tion, and testing) provided by standard benchmarks ( WHU Aerial, Massachusetts, GF-7). While adhering to these splits is essential to ensure fair comparisons with existing state-of-the-art methods, we acknowledge that remote sensing imagery inherently contains strong spa- tial autocorrelation. Because nearby patches may share similar textures and roof types, such evaluations can sometimes produce overly optimistic metrics. This rep- resents a broader limitation in current benchmarks for fully assessing zero-shot generalization across entirely unseen urban morphologies. Second, practical applica- tion boundaries exist. Performance is inherently con- strained by mixed-pixel effects in lower-resolution im- ages. Furthermore, the model is currently optimized for optical imagery, and its specific parameter complexity may pose challenges for deployment in strictly resource- constrained or real-time edge devices. Generalization to multimodal data and the development of lightweight variants will be the focus of future work. CONCLUSION This paper presents an improved EViT-based se- mantic segmentation network for high-resolution remote sensing imagery. By integrating GCC-MSA, LGC, CoAt, SGSPP, GEP, and SPGM, the hybrid architecture simultaneously models local details and global seman- tics. Experiments on three building datasets (WHU Aer- ial, Massachusetts, GF-7) show that the proposed method outperforms most existing models in metrics such as IoU. Ablation studies confirm the importance of the dual-branch structure and the contribution of each key module. The work validates the effectiveness of CNN-Transformer hybrid architectures for building ex- traction and provides an extensible solution for precise surface object extraction from high-resolution remote sensing imagery. Building upon the limitations dis- cussed, future work will focus on exploring geograph- ically separated data splits to rigorously evaluate zero- shot generalization, extending the framework to multi- modal remote sensing data, and developing lightweight architectures for efficient real-time deployment.
AN IMPROVED EVIT NETWORK FOR SEMANTIC SEGMENTATION OF HIGH-RESOLUTION REMOTE SENSING IMAGERY · 2026 · DOIIn addition, addressing the issue of insufficient information interaction among traditional detection heads, which can lead to the loss of partial information, we propose a Lightweight Detail-Enhanced Convolutional Detection Head (LDECD).
Insulator Defect Detection Based on Multi-scale Feature Fusion and Local Topology Reconstruction · 2026 · DOIthe dataset the model’s generalization ability remain. First, for strong proposed method performance in safety violation detection at power operation sites, size is violation samples, which relatively may in unseen affect the framework is mainly designed for scenarios. Second, in operation environments, power fine- other domains may industrial additional current tuning or domain adaptation. the still method relies on supervised annotations and may experience performance degradation under severe occlusion or low-light conditions.
Violation detection in power operation sites based on multi-scale detection and few-shot learning · 2026 · DOIReferences Despite the promising results achieved by MDAB-UNet, sev- eral limitations should be acknowledged. First, while our method demonstrates good performance on the LoveDA and WHDLD datasets, both datasets are primarily composed of optical remote sensing images. The adaptability of MDAB-UNet to other sensor types (e.g., SAR, hyperspectral) or images captured under challenging envi- ronmental conditions remains to be validated. Second, the current study does not systematically inves- tigate the model’s robustness to input noise or image degra- dation, which are common in real-world remote sensing applications. Third, regarding real-time inference capability, we acknowl- edge that the current MDAB-UNet, while achieving a 1. Li, J., Cai, Y., Li, Q., Kou, M., Zhang, T.: A review of remote sensing image segmentation by deep learning methods. Int. J. Digit. Earth (2024). https://doi.org/10.1080/17538947.2024.2328827 2. Wang, L., Fang, S., Meng, X., Li, R.: Building extraction with vision transformer. IEEE Trans. Geosci. Remote Sens. 60, 1–11 (2022). https://doi.org/10.1109/TGRS.2022.3186634 3. Zhang, R., Chen, J., Feng, L., Li, S., Yang, W., Guo, D.: A refined pyramid scene parsing network for polarimetric SAR image seman- tic segmentation in agricultural areas. IEEE Geosci. Remote Sens. Lett. 19, 1–5 (2022). https://doi.org/10.1109/LGRS.2021.3086117 4. Han, W., Li, J., Wang, S., Zhang, X., Dong, Y., Fan, R., Zhang, X., Wang, L.: Geological remote sensing interpretation using deep learning feature and an adaptive multisource data fusion network. IEEE Trans. Geosci. Remote Sens. 60, 1–14 (2022). https://doi. org/10.1109/TGRS.2022.3183080 5. Ma, Z., Mei, G.: Deep learning for geological hazards analysis: data, models, applications, and opportunities. Earth-Sci. Rev. 223, 103858 (2021). https://doi.org/10.1016/j.earscirev.2021.103858 123 103 Page 16 of 17 Journal of Real-Time Image Processing (2026) 23:103 6. Ferraioli, G.: Multichannel InSAR building edge detection. IEEE Trans. Geosci. Remote Sens. 48(3), 1224–1231 (2010). https://doi. org/10.1109/TGRS.2009.2029338 7. Yuan, J., Wang, D., Li, R.: Remote sensing image segmentation by combining spectral and texture features. IEEE Trans. Geosci. Remote Sens. 52(1), 16–24 (2014). https://doi.org/10.1109/TGRS. 2012.2234755 8. Simonyan, K., Zisserman, A.: Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv preprint, abs/1409.1556, https://arxiv.org/abs/1409.1556 (2014) 9. Long, J., Shelhamer, E., Darrell, T.: Fully convolutional networks for semantic segmentation. In: Proceedings of the IEEE Confer- ence on Computer Vision and Pattern Recognition, pp. 3431–3440. https://doi.org/10.1109/CVPR.2015.7298965 (2015) 10. Ronneberger, O., Fischer, P., Brox, T.: U-Net: Convolutional Net- works for Biomedical Image Segmentation. In: Medical Image Computing and Computer-Assisted Intervention, pp. 234–241. https://doi.org/10.1007/978-3-319-24574-4_28 (2015) 11.
MDAB-UNet: a lightweight and efficient improved UNet architecture for remote sensing image segmentation · 2026 · DOI• Adversarial training and evolutionary optimization increase training time and require more computational resources com- pared to traditional segmentation models. • The GA search adds extra upfront computation, but it helps find optimal settings that improve the model’s accuracy and stability. • Like other deep learning methods, this model relies on large, high-quality labeled datasets to ensure good generaliza- tion and avoid performance drops on different scanners.
A novel hybrid framework integrating GA-driven 3D ResUNetGAN for MRI brain tumor segmentation · 2026 · DOIFuture work will focus on extending the proposed approach to other imaging modalities, such as CT and ultrasound, exploring alternative backbone architectures, incorporating volumetric evaluation metrics, and performing standardized comparisons on the DLDS dataset using physical unit-based metrics, as voxel spacing metadata was not consistently available in the current dataset.
An efficient pyramid scene parsing network with multi-scale feature fusion for liver segmentation in magnetic resonance imaging · 2026 · DOIMost CNN-based FPGA implementations for remote sensing do not benchmark against prior FPGA-based CNN architectures, making it impossible to identify commonly compared network designs and establish reproducible performance baselines. Standardized benchmarking protocols comparing different CNN models and FPGA devices for Earth observation tasks are needed to enable meaningful cross-study comparisons.
Knowledge Distillation (KD) as a compression strategy appears in only one surveyed study despite its strong potential for FPGA-constrained Earth observation deployment. The synergy between KD-guided weight pruning and hardware-aware neural network design for maximizing resource utilization on edge FPGA platforms remains largely unexplored in remote sensing applications.
Coarse-Grained Reconfigurable Arrays (CGRAs) with hardened arithmetic units and systolic array PE configurations remain underexplored for FPGA-enabled ML in Earth observation. A detailed comparative analysis between CGRA-based implementations (e.g., AMD Versal AI Engines) and traditional FPGA-only solutions for highly quantized semantic segmentation models on remote sensing datasets is needed to establish performance benchmarks.
Vision Transformers (ViTs) and recurrent neural networks (RNNs) lack mature FPGA implementations despite their superior performance in remote sensing tasks like temporal change monitoring and flight-path analysis. Current FPGA toolchains (FINN, Vitis AI) do not support transformer and recurrent architectures, including their quantized and pruned variants, creating a barrier to deploying these modern architectures on edge platforms.
Feature estimation and retrieval problems are severely underrepresented in FPGA-enabled ML for remote sensing, with only two studies addressing these tasks despite their practical importance. Systematic investigation of lightweight retrieval architectures on FPGA platforms for Earth observation applications remains largely unexplored.
Data compression and disaster response remain unexplored onboard use cases for FPGA-enabled ML in Earth observation, despite their relevance to the NewSpace era. Lightweight FPGA implementations of state-of-the-art compression methods for SmallSat imaging payloads generating ~640 Mbps data rates present a critical research opportunity to enable real-time alerts for wildfires and post-disaster damage assessments.
YOLOv5s exhibits significantly lower validation performance ([email protected]:0.95 of 0.759) compared to YOLOv8s (0.859), but the paper does not investigate the specific architectural or training factors responsible for this performance degradation on visually similar grocery items.
Enhancing Retail Checkout Efficiency Through a Hybrid YOLOv8-Based Grocery Detection and Billing System · 2026 · DOIOptical character recognition (OCR) integration for price tag detection is mentioned as future work but no framework, accuracy targets, or method for handling variable price tag formats across different grocery items and retailers is specified.
Enhancing Retail Checkout Efficiency Through a Hybrid YOLOv8-Based Grocery Detection and Billing System · 2026 · DOIThe authors mention applying model optimization techniques (pruning, quantization, lightweight architectures) for edge deployment but provide no specific implementation details, performance metrics, or latency measurements for these optimizations on the YOLOv8s model selected for web deployment.
Enhancing Retail Checkout Efficiency Through a Hybrid YOLOv8-Based Grocery Detection and Billing System · 2026 · DOI
Most-cited papers in Advanced Neural Network Applications
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data · 2024 · 1,112 citations
- TransUNet: Rethinking the U-Net architecture design for medical image segmentation through the lens of transformers · Medical Image Analysis · 2024 · 1,065 citations
- YOLOv10: Real-Time End-to-End Object Detection · 2024 · 607 citations
- Slim-neck by GSConv: a lightweight-design for real-time detector architectures · Journal of Real-Time Image Processing · 2024 · 572 citations
- Rep ViT: Revisiting Mobile CNN From ViT Perspective · 2024 · 506 citations
- UNETR++: Delving Into Efficient and Accurate 3D Medical Image Segmentation · IEEE Transactions on Medical Imaging · 2024 · 374 citations
- FFCA-YOLO for Small Object Detection in Remote Sensing Images · IEEE Transactions on Geoscience and Remote Sensing · 2024 · 316 citations
- Fusion of 3D LIDAR and Camera Data for Object Detection in Autonomous Vehicle Applications · IEEE Sensors Journal · 2020 · 296 citations
- Channel prior convolutional attention for medical image segmentation · Computers in Biology and Medicine · 2024 · 293 citations
- Deep learning for medical image segmentation: State-of-the-art advancements and challenges · Informatics in Medicine Unlocked · 2024 · 275 citations
Most recent work
- FPGA-Enabled Machine Learning Applications in Earth Observation: A Systematic Review · ACM Computing Surveys · 2026
- A sparse-to-dense guided fusion framework for three-dimensional object detection in railway environments · Engineering Applications of Artificial Intelligence · 2026
- A method for detecting foreign objects in railway perimeters under low-light conditions based on image enhancement · Engineering Research Express · 2026
- DPF-DETR: Enhancing Drone Image Detection with Density Perception and Multi-Scale Feature Fusion · Remote Sensing · 2026
- CIHM: context-insight hybrid Mamba for efficient medical image segmentation · Visual Intelligence · 2026
- LASFNet: A Lightweight Attention-Guided Self-Modulation Feature Fusion Network for Multimodal Object Detection · IEEE Transactions on Cybernetics · 2026
- Capacity-Aware Lightweight Object Detection for UAV Remote Sensing: Dynamic Coupling Regularity and the SP-YOLO Model Family · Applied Sciences · 2026
- Explainable brain tumor segmentation via attention-guided hybrid CNN–Transformer–Mamba network · Knowledge-Based Systems · 2026
- A Performance Analysis of YOLOv8 and YOLOv10 on PCB Surface Defects · Journal of the Institute of Science and Technology · 2026
- An Efficient Weighted Majority Voting Ensemble Machine Learning Classifier Framework for Image Segmentation · Engineering, Technology & Applied Science Research · 2026
Find a gap in your own Advanced Neural Network Applications sub-topic
This page shows what the Advanced Neural Network Applications literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.
Open the Research Gap Finder →