Open research questions in Web Data Mining and Analysis
70 unresolved questions extracted from the limitations and future-work sections of 648 Web Data Mining and Analysis papers in our library. Each links back to the study that raised it.
What the literature leaves open
Future research can focus on improving the quality of the retrieval component. Future research can explore the use of the proposed approach in other applications such as text generation and question answering.
The paper identifies a gap in the existing literature on retrieval augmented language models. The gap is that previous methods do not effectively address hallucination in generative language models.
The paper identifies a gap in the integration of Boolean queries with probabilistic retrieval models. The paper notes that complex Boolean queries can be difficult to construct.
To implement the system in a real-world setting. To extend the system to handle other types of document retrievals. To improve the performance of the set of heuristics.
There is a need for a distributed system for document retrieval in office environments. The existing systems do not learn document distribution patterns and user interests and preferences. The existing systems do not customize document retrievals for each user.
Future research can focus on developing more effective weighted retrieval systems. Future research can explore the application of the threshold-value concept to other areas.
Further evaluation of the extended models is needed. The use of other term significance weights should be explored. The application of the extended models to other domains should be considered.
The limitation of probabilistic models of retrieval assuming binary index terms is identified. The need to extend these models to include term significance weights is recognized.
Future research can explore the application of passage retrieval to other domains. The technique can be improved by developing more sophisticated search algorithms. The paper provides a framework for future research in passage retrieval.
The paper identifies a gap in the application of passage retrieval to scientists. The technique has been successfully applied to lawyers, but not to scientists.
Alternative term-weighting and ranking algorithm combinations should be explored. Further experimentation is needed to find ranking procedures that work better.
The challenge facing an information retrieval system is to present a user with references that fulfill their information need. There is a need to develop a better understanding of how the individual components of retrieval systems function.
The paper identifies a gap in the knowledge of how employees of institutions undergoing change are linked to the change process. The study indicates that state mental health employees at the institutional level feel genuinely unable to affect the change process.
The need for efficient retrieval from automated bibliographic data bases. The lack of knowledge on the most suitable combination of data elements for document retrieval.
Relative effectiveness of titles, abstracts, and subject headings for machine retrieval from the COMPENDEX services · 1975 · DOIThe lack of a formal criterion for distinguishing between index and non-index words. The need for a probabilistic model that generalizes the pure random model to account for the observed distribution of content-bearing words.
The potential 'washout' of a few large correlations by a host of smaller ones - The need to increase the size of large correlations at the expense of smaller ones
The need for a more objective and efficient method for selecting the best word subset for analysis - The limitations of frequency as a sole criterion for word selection
There is a need to compare conventional retrieval methods (MEDLARS) with automatic text analysis methods (SMART); There is a lack of understanding of the relative merits of controlled versus free language indexing and manual versus automatic analysis methodology
A new comparison between conventional indexing (MEDLARS) and automatic text processing (SMART) · 1972 · DOIThe need to analyze the conditions under which various methods of sentence selection are successful - The need to develop criteria for selecting sentences to form an abstract
Automatic abstracting and indexing. II. Production of indicative abstracts by application of contextual inference and syntactic coherence criteria · 1971 · DOINot enough is known about the behaviour of automatic keyword classifications. Few systematic experiments have been carried out on the properties of effective keyword classifications.
The paper identifies a gap in the use of controlled vocabulary subject indexing in archives. The author notes that systematic approaches have been made in the library field but not in archives.
The difficulty of analyzing documents in different languages. The need for a complete and accurate multilingual thesaurus. The challenge of evaluating the effectiveness of mixed language processing.
Improving the completeness of the German thesaurus. Evaluating the effectiveness of mixed language processing in other languages. Developing more advanced linguistic tools for document analysis.
Organizing a network of urban observatories in major U.S. cities and urban regions. Conducting policy-oriented research on selected major issues of direct concern to mayors and others on the firing line.
The need for machine translation between Japanese and English. The lack of commercial machine translation systems in Japan.
Special Issue: “Collection of Best Annual Papers” Organized for the 20th Anniversary of the Association for Natural Language Processing · 2014 · DOI
Most-cited papers in Web Data Mining and Analysis
- An analysis of the relative hardness of Reuters‐21578 subsets · Journal of the American Society for Information Science and Technology · 2005 · 110 citations
- Large Language Models can Accurately Predict Searcher Preferences · 2024 · 106 citations
- Seven Failure Points When Engineering a Retrieval Augmented Generation System · 2024 · 103 citations
- Automatic abstracting and indexing. II. Production of indicative abstracts by application of contextual inference and syntactic coherence criteria · Journal of the American Society for Information Science · 1971 · 55 citations
- Web searcher interaction with the Dogpile.com metasearch engine · Journal of the American Society for Information Science and Technology · 2007 · 53 citations
- A survey in indexing and searching XML documents · Journal of the American Society for Information Science and Technology · 2002 · 45 citations
- From ChatGPT to CatGPT · Information Technology and Libraries · 2023 · 45 citations
- Methods for using Bing's AI‐powered search engine for data extraction for a systematic review · Research Synthesis Methods · 2023 · 36 citations
- Document representation in probabilistic models of information retrieval · Journal of the American Society for Information Science · 1981 · 28 citations
- Applying machine classifiers to update searches: Analysis from two case studies · Research Synthesis Methods · 2021 · 25 citations
Most recent work
- What not to index: a flow chart for passing mentions in book indexing · The Indexer · 2026
- ArabicEduCrawler: AI-Assisted Focused Crawling and Corpus Construction for Arabic Educational Web Content · Zenodo (CERN European Organization for Nuclear Research) · 2026
- Ancient GeoCities: A Dataset of Temporally Annotated Web Pages · 2026
- Media Cloud 2.0: An Updated Open Web News Archive · Proceedings of the International AAAI Conference on Web and Social Media · 2026
- TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework · ACM Transactions on Information Systems · 2026
- BC Hansard Index: a roadmap to parliamentary debates · The Indexer · 2026
- Summarization of Web Page using Web Scraping System · International Scientific Journal of Engineering and Management · 2026
- Plans for Evaluating Structured Generative Search Summaries · arXiv · 2026
- Self-Conditioned Positional HNSW for Overlap-Aware Retrieval in Chunked-Document RAG Systems: Method and Industrial Evidence-Quality Audit · arXiv · 2026
- EviProp: Seeded Relevance Diffusion on Chunk-Page Graphs for Long Multimodal Document Retrieval · arXiv · 2026
Find a gap in your own Web Data Mining and Analysis sub-topic
This page shows what the Web Data Mining and Analysis literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.
Open the Research Gap Finder →