Open research questions in Advanced Vision and Imaging
122 unresolved questions extracted from the limitations and future-work sections of 533 Advanced Vision and Imaging papers in our library. Each links back to the study that raised it.
What the literature leaves open
It can be observed that although reliable 2D instance segmentation masks have been provided by off-the-shelf 2D models, the network may fail in difficult situations, which indicates that our method is mainly limited by the sparse and noisy nature of radar data itself, rather than insufficient support from 2D semantics.
Weakly Supervised Cross-Modal Learning for 4D Radar Scene Flow Estimation · 2026However, the principles governing how these neurons filter their inputs to generate appropriate responses remain unclear.
Dendritic architecture enables de novo computation of salient motion in the superior colliculus · 2025 · DOIThe gap in current research is the requirement of a large number of input views, which can be resource-intensive. The gap is the inability to render photorealistic street-view panoramas with a small number of input views.
Exploring the use of Lite3R in various applications, such as robotics and computer vision. Investigating the potential of Lite3R to improve the efficiency of 3D reconstruction models in real-world scenarios.
Lite3R: A Model-Agnostic Framework for Efficient Feed-Forward 3D Reconstruction · 2026Modern 3D transformer pipelines face significant challenges, including dense multi-view attention and low-precision execution. Improving the efficiency of 3D reconstruction models is crucial for practical deployment.
Lite3R: A Model-Agnostic Framework for Efficient Feed-Forward 3D Reconstruction · 2026The paper identifies the need for more diverse modalities and more accurate reconstruction. The authors discuss the challenge of free-viewpoint synthesis and long-context generation. The survey highlights the potential misuse of generative capabilities of feed-forward 3D reconstruction models.
Improving model generalizability to reduce computational costs - Developing advanced detection models to distinguish between generated and real content - Establishing new regulations as 3D reconstruction technologies become increasingly accessible
Further improvement of the method for scenes with very large camera movements or tissue deformations. Exploration of real-time applications of the method. Investigation of potential applications in medical imaging and computer vision.
Existing methods struggle with challenging scenes, such as those with deformable tissues and surgical tool occlusion. Inferior initialization of 3D Gaussians in dynamic endoscopic scenes. Limited computational efficiency during training and inference.
Conventional RGB-based stereo matching algorithms suffer from ambiguities caused by metamerism and weak textures. The lack of a co-aperture system and a Measure-and-Complete strategy for generating dense pseudo-ground truth.
4D-Aware Stereo Matching via Implicit Spectral Reconstruction with Multi-modal Training and RGB-only Deployment · 2026 · DOIFirst, the spectral reconstruction is restricted to the 450–700 nm visible band; its performance in the near-infrared and short-wave infrared regimes—where many distinctive material fingerprints reside—remains to be validated. 8 m), and generalization to medium- and long-range scenarios remains an open question.
4D-Aware Stereo Matching via Implicit Spectral Reconstruction with Multi-modal Training and RGB-only Deployment · 2026 · DOIExisting generation paradigms struggle to satisfy the requirements of whole-house tours simultaneously. There is a gap in generating consistent whole-house panoramas with high-fidelity and material consistency.
PanoWorld: A Generative Spatial World Model for Consistent Whole-House Panorama Synthesis · 2026Exploring metrics that quantify the similarity gap between correct and uncertain matches. Evaluating the performance of the proposed approach in different environments and conditions.
Patch Ensembles for Robust Salmon Re-Identification with Weak Trajectory Labels · 2026The lack of labeled data for large-scale salmon re-identification. The high accuracy requirements imposed by commercial net-pens. The need for a realistic evaluation setup for salmon re-identification.
Patch Ensembles for Robust Salmon Re-Identification with Weak Trajectory Labels · 2026Existing methods typically trade off generalization, geometric fidelity, and efficiency. There is a need for an automated correspondence framework that is accurate, robust, and computationally efficient.
SGSoft: Learning Fused Semantic-Geometric Features for 3D Shape Correspondence via Template-Guided Soft Signals · 2026To explore the application of the method to other domains. To investigate the robustness of the method to different types of attacks. To develop more advanced methods for capturing dispersed attribution signals.
Who Generated This 3D Asset? Learning Source Attribution for Generative 3D Models · 2026Prior work has not systematically studied passive source attribution for generated 3D assets. The problem faces challenges such as dispersed attribution signals and realistic deployment constraints.
Who Generated This 3D Asset? Learning Source Attribution for Generative 3D Models · 2026The experiments were run in a simulated environment using the Habitat simulator. The evaluation was limited to six apartment scenes from the ReplicaCAD dataset.
Fixed External Cameras as Common Prior Maps for Active 3D Scene Graph Generation · 2026Extending the proposed method to more complex environments and scenarios. Investigating the use of multiple external cameras as Common Prior Maps. Improving the efficiency of active exploration in various applications.
Fixed External Cameras as Common Prior Maps for Active 3D Scene Graph Generation · 2026The datasets used in this work do not provide ground-truth relational annotations. The system relies on a deterministic geometry-based module for edge generation. The experiments are limited to the Replica and ReplicaCAD datasets.
RGB-only Active 3D Scene Graph Generation for Indoor Mobile Robots · 2026Current approaches to 3D scene graph generation rely on dedicated depth sensors, limiting deployment to specialized robotic platforms. Traditional perception pipelines lack semantic context or relational structure.
RGB-only Active 3D Scene Graph Generation for Indoor Mobile Robots · 2026The method encounters difficulties at very high magnifications (e.g., ×1024), where current vision-language models struggle to infer coherent structures. The system may not perform well with low-quality input images or limited training data.
GaussianZoom: Progressive Zoom-in Generative 3D Gaussian Splatting with Geometric and Semantic Guidance · 2026Investigating more capable content creative zoom-in approaches to enable seamless transitions from cosmic-scale environments down to microscopic and molecular scenes. Improving the performance of the system at very high magnifications (e.g., ×1024). Exploring the application of GaussianZoom to various fields such as architecture, product design, and video games.
GaussianZoom: Progressive Zoom-in Generative 3D Gaussian Splatting with Geometric and Semantic Guidance · 2026One challenge is the lack of a robust background matting framework for Virtual Production. Another challenge is the need to preserve pretrained semantics and improve robustness to background shifts.
CineMatte: Background Matting for Virtual Production and Beyond · 2026The research gap is the lack of a robust background matting framework for Virtual Production. Prior work has limitations, such as boundary artifacts and overfitting.
CineMatte: Background Matting for Virtual Production and Beyond · 2026
Most-cited papers in Advanced Vision and Imaging
- NeRF · Communications of the ACM · 2021 · 6,387 citations
- Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-Shot Cross-Dataset Transfer · IEEE Transactions on Pattern Analysis and Machine Intelligence · 2022 · 1,639 citations
- Depth-first iterative-deepening · Artificial Intelligence · 1985 · 960 citations
- Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation · 2024 · 432 citations
- Dynamic 3D Gaussians: Tracking by Persistent Dynamic View Synthesis · 2024 · 423 citations
- Deformable 3D Gaussians for High-Fidelity Monocular Dynamic Scene Reconstruction · 2024 · 395 citations
- Disentangling Light Fields for Super-Resolution and Disparity Estimation · IEEE Transactions on Pattern Analysis and Machine Intelligence · 2022 · 320 citations
- Interpreting perspective images · Artificial Intelligence · 1983 · 278 citations
- NeRFPlayer: A Streamable Dynamic Scene Representation with Decomposed Neural Radiance Fields · IEEE Transactions on Visualization and Computer Graphics · 2023 · 245 citations
- UniDepth: Universal Monocular Metric Depth Estimation · 2024 · 210 citations
Most recent work
- Leveraging image transformation and optical flow for heterogeneous change detection under co-registration errors · Pattern Recognition · 2026
- Parameter-Efficient Fine-Tuning for Pre-Trained Vision Models: A Survey and Benchmark · International Journal of Computer Vision · 2026
- MMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion · Information Fusion · 2026
- Advances in Feed‐Forward 3D Reconstruction and View Synthesis: A Survey · Computer Graphics Forum · 2026
- 4D-Aware Stereo Matching via Implicit Spectral Reconstruction with Multi-modal Training and RGB-only Deployment · Optics Express · 2026
- SCFP-Depth: Achieving robust self-supervised monocular depth estimation via static compensation and frequency-domain priors · Knowledge-Based Systems · 2026
- EOGS++: Earth Observation Gaussian Splatting with Internal Camera Refinement and Direct Panchromatic Rendering · ISPRS annals of the photogrammetry, remote sensing and spatial information sciences · 2026
- CED: CLIP-guided entropy dynamics for robust test-time adaptation in harsh visual conditions · Pattern Recognition · 2026
- Learning domain-agnostic spatial-angular feature for light field image super-resolution · Pattern Recognition · 2026
- S ee 4D: Pose‐Free 4D Generation via Auto‐Regressive Video Inpainting · Computer Graphics Forum · 2026
Find a gap in your own Advanced Vision and Imaging sub-topic
This page shows what the Advanced Vision and Imaging literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.
Open the Research Gap Finder →