Open research questions in Advanced Vision and Imaging
27 unresolved questions extracted from the limitations and future-work sections of 391 Advanced Vision and Imaging papers in our library. Each links back to the study that raised it.
What the literature leaves open
, training data strictly limited to a single walking trajectory). Ad- ditionally, noting that the PSNR of our current synthesized vir- tual views is limited to approximately 21.
319 backbone in this work, future research will investigate altern- ative and more conventional diffusion formulations to better understand how different denoising architectures affect scene- level point cloud generation under varying LiDAR acquisition conditions. As shown in Table 2, simple feature concatenation yields the weakest over- all performance, indicating that naive fusion is insufficient for robust scene-level conditioning.
First, the spectral reconstruction is restricted to the 450–700 nm visible band; its performance in the near-infrared and short-wave infrared regimes—where many distinctive material fingerprints reside—remains to be validated. 8 m), and generalization to medium- and long-range scenarios remains an open question.
4D-Aware Stereo Matching via Implicit Spectral Reconstruction with Multi-modal Training and RGB-only Deployment · 2026 · DOINevertheless, existing methods remain inaccurate due to two obstacles: (i) training data is scarce and lacks intrinsics diversity; and (ii) benchmarks, including InFlux, have limited scene and camera motion diversity, making it difficult to properly evaluate methods.
InFlux++: Real and Synthetic Data for Estimating Dynamic Camera Intrinsics · 2026We then construct a mixed dataset comprising high-fidelity HDR images to provide realistic HDR priors, and in-the-wild HDR videos to provide dynamic spatio-temporal context.
Video Generation Models Are Inherent Lighting Estimators · 2026Video world models are a foundational generative technology for embodied AI and the Metaverse, yet existing approaches are inherently limited to a single agent observing from a single perspective.
MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data · 2026Since, to the best of our knowledge, frame-wise indoor lidar semantic segmentation has not been previously studied and no labeled datasets or established benchmarks exist for this task, direct quantitative comparison is not possible.
Feasibility of Indoor Frame-Wise Lidar Semantic Segmentation via Distillation from Visual Foundation Model · 2026 · DOIWe present SparseOcc++, a geometry-aware sparse framework that explicitly decouples scene completion from semantic segmentation.
SparseOcc++: Geometry-Aware Sparse Latent Representation for Semantic Occupancy Prediction · 2026While intuitive, this 2D-to-3D prediction is inherently entangled with camera parameters, leading to limited data scalability across heterogeneous driving datasets.
PixelPilot: Scalable Vision-Language-Action Models for End-to-End Autonomous Driving · 2026Dense depth estimation for autonomous driving faces a geometry-scale conflict: depth foundation models deliver pixel-aligned dense visual geometry without reliable metric scale, while projected LiDAR provides metric anchors that are sparse, noisy, and misaligned with image structures.
DrivingDepth: Sparse-Prompted Pixel-wise Scale Correction for Driving Depth Estimation · 2026Event-based optical-flow estimation remains challenging across different fixed-frequency sampling, particularly under low-frequency conditions where observations are sparse and motion displacements are large.
This assumption has not been validated. They have not been tested on actual hardware, with actual datasets, or in actual deployment. It should not replace traditional methods but can provide additional evidence in cases where traditional methods are inconclusive.
The Twin Crises of Modern Vision AI: Progressive Tensor Core Decay and Geometric Blindness -A Zig-Native Architectural Countermeasure by mouad tarif · 2026 · DOIQuery-based transformer decoders are effective for object reconstruction from sparse scientific sensor measurements, but their scalability to high-multiplicity data is limited by fixed, input-independent query sets and costly decoder cross-attention.
Better Queries, Cheaper Attention: Adapting Transformers for Efficient Sparse Reconstruction · 2026Although recent vision foundation models have shown promise, their learned representations often remain insufficiently geometry-consistent, hindering stable feature correspondence and limiting their reliability for downstream navigation tasks.
Geometry-Consistent Endoscopic Representations for Image-Guided Navigation via Structured Foundation Model Adaptation · 2026Existing restoration and video generation methods are insufficient for this task, as they often fail to jointly repair 3DGS-specific artifacts, improve visual realism, and ensure temporal consistency.
RealityBridge: Bridging Editable 3D Gaussian Splatting Driving Simulations and Real-World Videos · 2026Existing reference-based metrics require ground truth, while ground-truth-free metrics such as MEt3R depend on learned reconstruction backbones whose failure modes are poorly characterized.
Can These Views Be One Scene? Evaluating Multiview 3D Consistency when 3D Foundation Models Hallucinate · 2026Structure-from-Motion -- the process of simultaneously estimating camera poses and 3D scene structure from a collection of images -- remains a central challenge in computer vision, with many open problems yet to be solved.
Global Structure-from-Motion Meets Feedforward Reconstruction · 2026While considerable progress has been made in reconstructing 3D models of terrestrial quadrupeds, aquatic animals remain unexplored due to the difficulty of observing them in their natural underwater environment.
Model-based Metric 3D Shape and Motion Reconstruction of Wild Bottlenose Dolphins in Drone-Shot Videos · 2026 · DOIWe validate our approach on three image sets representative of in-orbit operations, demonstrating its effectiveness for offline reconstruction and highlighting its suitability for online reconstruction, an open problem in the field.
NeRF-based Spacecraft Reconstruction from Monocular Imagery Under Illumination Variability and Pose Uncertainty · 2026Generating geometrically consistent videos remains an open challenge: text-to-video diffusion models trained on web-scale data treat geometry only implicitly, leading to object deformation, texture drift, and non-rigid backgrounds under camera motion.
GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation · 2026The natural stereo pair shows that even when intensity changes are sparse the reconstruction preserves the shape although, the interpolant exhibits a tendency to consider spurious stereo matches also as a potential data point.
Visual surface perception: Surface reconstruction from shading and stereo. · 1990
Most-cited papers in Advanced Vision and Imaging
- NeRF · Communications of the ACM · 2021 · 6,387 citations
- Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation · 2024 · 432 citations
- Dynamic 3D Gaussians: Tracking by Persistent Dynamic View Synthesis · 2024 · 423 citations
- Deformable 3D Gaussians for High-Fidelity Monocular Dynamic Scene Reconstruction · 2024 · 395 citations
- UniDepth: Universal Monocular Metric Depth Estimation · 2024 · 210 citations
- DNGaussian: Optimizing Sparse-View 3D Gaussian Radiance Fields with Global-Local Depth Normalization · 2024 · 188 citations
- Compressed 3D Gaussian Splatting for Accelerated Novel View Synthesis · 2024 · 182 citations
- Spacetime Gaussian Feature Splatting for Real-Time Dynamic View Synthesis · 2024 · 181 citations
- VastGaussian: Vast 3D Gaussians for Large Scene Reconstruction · 2024 · 177 citations
- MotionCtrl: A Unified and Flexible Motion Controller for Video Generation · 2024 · 158 citations
Most recent work
- Parameter-Efficient Fine-Tuning for Pre-Trained Vision Models: A Survey and Benchmark · International Journal of Computer Vision · 2026
- Advances in Feed‐Forward 3D Reconstruction and View Synthesis: A Survey · Computer Graphics Forum · 2026
- 4D-Aware Stereo Matching via Implicit Spectral Reconstruction with Multi-modal Training and RGB-only Deployment · Optics Express · 2026
- SCFP-Depth: Achieving robust self-supervised monocular depth estimation via static compensation and frequency-domain priors · Knowledge-Based Systems · 2026
- S <scp>ee</scp> 4D: Pose‐Free 4D Generation via Auto‐Regressive Video Inpainting · Computer Graphics Forum · 2026
- Seeing Through Satellite Images at Street Views · IEEE Transactions on Pattern Analysis and Machine Intelligence · 2026
- Camera augmentation: enabling uncalibrated stereo matching of minimally invasive surgery images by training from the wealth of public synthetic image datasets · International Journal of Computer Assisted Radiology and Surgery · 2026
- Efficient Self-supervised Monocular Depth Estimation via Knowledge Distillation and Light-weighted Attention · Journal of Institute of Control Robotics and Systems · 2026
- Endo-PairGS: pair priors for dynamic endoscopic scene reconstruction · International Journal of Computer Assisted Radiology and Surgery · 2026
- Model-based Metric 3D Shape and Motion Reconstruction of Wild Bottlenose Dolphins in Drone-Shot Videos · International Journal of Computer Vision · 2026
Find a gap in your own Advanced Vision and Imaging sub-topic
This page shows what the Advanced Vision and Imaging literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.
Open the Research Gap Finder →