Computer Science · Research topic

Open research questions in Advanced Vision and Imaging

27 unresolved questions extracted from the limitations and future-work sections of 391 Advanced Vision and Imaging papers in our library. Each links back to the study that raised it.

What the literature leaves open

  • , training data strictly limited to a single walking trajectory). Ad- ditionally, noting that the PSNR of our current synthesized vir- tual views is limited to approximately 21.

    VISTA-GS: MVS-Guided Virtual View Augmentation for Sparse-View 3D Gaussian Splatting · 2026 · DOI
  • 319 backbone in this work, future research will investigate altern- ative and more conventional diffusion formulations to better understand how different denoising architectures affect scene- level point cloud generation under varying LiDAR acquisition conditions. As shown in Table 2, simple feature concatenation yields the weakest over- all performance, indicating that naive fusion is insufficient for robust scene-level conditioning.

    Appearance-aware Scaling Diffusion Model for 3D Point Cloud Upsampling · 2026 · DOI
  • First, the spectral reconstruction is restricted to the 450–700 nm visible band; its performance in the near-infrared and short-wave infrared regimes—where many distinctive material fingerprints reside—remains to be validated. 8 m), and generalization to medium- and long-range scenarios remains an open question.

    4D-Aware Stereo Matching via Implicit Spectral Reconstruction with Multi-modal Training and RGB-only Deployment · 2026 · DOI
  • Nevertheless, existing methods remain inaccurate due to two obstacles: (i) training data is scarce and lacks intrinsics diversity; and (ii) benchmarks, including InFlux, have limited scene and camera motion diversity, making it difficult to properly evaluate methods.

    InFlux++: Real and Synthetic Data for Estimating Dynamic Camera Intrinsics · 2026
  • We then construct a mixed dataset comprising high-fidelity HDR images to provide realistic HDR priors, and in-the-wild HDR videos to provide dynamic spatio-temporal context.

    Video Generation Models Are Inherent Lighting Estimators · 2026
  • Video world models are a foundational generative technology for embodied AI and the Metaverse, yet existing approaches are inherently limited to a single agent observing from a single perspective.

    MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data · 2026
  • Since, to the best of our knowledge, frame-wise indoor lidar semantic segmentation has not been previously studied and no labeled datasets or established benchmarks exist for this task, direct quantitative comparison is not possible.

    Feasibility of Indoor Frame-Wise Lidar Semantic Segmentation via Distillation from Visual Foundation Model · 2026 · DOI
  • We present SparseOcc++, a geometry-aware sparse framework that explicitly decouples scene completion from semantic segmentation.

    SparseOcc++: Geometry-Aware Sparse Latent Representation for Semantic Occupancy Prediction · 2026
  • While intuitive, this 2D-to-3D prediction is inherently entangled with camera parameters, leading to limited data scalability across heterogeneous driving datasets.

    PixelPilot: Scalable Vision-Language-Action Models for End-to-End Autonomous Driving · 2026
  • Dense depth estimation for autonomous driving faces a geometry-scale conflict: depth foundation models deliver pixel-aligned dense visual geometry without reliable metric scale, while projected LiDAR provides metric anchors that are sparse, noisy, and misaligned with image structures.

    DrivingDepth: Sparse-Prompted Pixel-wise Scale Correction for Driving Depth Estimation · 2026
  • Event-based optical-flow estimation remains challenging across different fixed-frequency sampling, particularly under low-frequency conditions where observations are sparse and motion displacements are large.

    TSR-SSM: temporal scale-robust state-space models for efficient event optical flow · 2026 · DOI
  • This assumption has not been validated. They have not been tested on actual hardware, with actual datasets, or in actual deployment. It should not replace traditional methods but can provide additional evidence in cases where traditional methods are inconclusive.

    The Twin Crises of Modern Vision AI: Progressive Tensor Core Decay and Geometric Blindness -A Zig-Native Architectural Countermeasure by mouad tarif · 2026 · DOI
  • Query-based transformer decoders are effective for object reconstruction from sparse scientific sensor measurements, but their scalability to high-multiplicity data is limited by fixed, input-independent query sets and costly decoder cross-attention.

    Better Queries, Cheaper Attention: Adapting Transformers for Efficient Sparse Reconstruction · 2026
  • Although recent vision foundation models have shown promise, their learned representations often remain insufficiently geometry-consistent, hindering stable feature correspondence and limiting their reliability for downstream navigation tasks.

    Geometry-Consistent Endoscopic Representations for Image-Guided Navigation via Structured Foundation Model Adaptation · 2026
  • Existing restoration and video generation methods are insufficient for this task, as they often fail to jointly repair 3DGS-specific artifacts, improve visual realism, and ensure temporal consistency.

    RealityBridge: Bridging Editable 3D Gaussian Splatting Driving Simulations and Real-World Videos · 2026
  • Existing reference-based metrics require ground truth, while ground-truth-free metrics such as MEt3R depend on learned reconstruction backbones whose failure modes are poorly characterized.

    Can These Views Be One Scene? Evaluating Multiview 3D Consistency when 3D Foundation Models Hallucinate · 2026
  • Structure-from-Motion -- the process of simultaneously estimating camera poses and 3D scene structure from a collection of images -- remains a central challenge in computer vision, with many open problems yet to be solved.

    Global Structure-from-Motion Meets Feedforward Reconstruction · 2026
  • While considerable progress has been made in reconstructing 3D models of terrestrial quadrupeds, aquatic animals remain unexplored due to the difficulty of observing them in their natural underwater environment.

    Model-based Metric 3D Shape and Motion Reconstruction of Wild Bottlenose Dolphins in Drone-Shot Videos · 2026 · DOI
  • We validate our approach on three image sets representative of in-orbit operations, demonstrating its effectiveness for offline reconstruction and highlighting its suitability for online reconstruction, an open problem in the field.

    NeRF-based Spacecraft Reconstruction from Monocular Imagery Under Illumination Variability and Pose Uncertainty · 2026
  • Generating geometrically consistent videos remains an open challenge: text-to-video diffusion models trained on web-scale data treat geometry only implicitly, leading to object deformation, texture drift, and non-rigid backgrounds under camera motion.

    GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation · 2026
  • The natural stereo pair shows that even when intensity changes are sparse the reconstruction preserves the shape although, the interpolant exhibits a tendency to consider spurious stereo matches also as a potential data point.

    Visual surface perception: Surface reconstruction from shading and stereo. · 1990

Most-cited papers in Advanced Vision and Imaging

Most recent work

Find a gap in your own Advanced Vision and Imaging sub-topic

This page shows what the Advanced Vision and Imaging literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.

Open the Research Gap Finder →

Related topics in Computer Science

27 open questions have been extracted from the limitations and future-work passages of 391 Advanced Vision and Imaging papers in our library. Each one below links back to the study that raised it, so you can read the original claim in context.

Tools for your next paper

Compare the categoryHonest roundups of the AI research tools, ours listed alongside the alternatives.

Command palette

Jump anywhere, run any action.