Computer Science · Research topic

Open research questions in Web Data Mining and Analysis

70 unresolved questions extracted from the limitations and future-work sections of 648 Web Data Mining and Analysis papers in our library. Each links back to the study that raised it.

What the literature leaves open

  • Future research can focus on improving the quality of the retrieval component. Future research can explore the use of the proposed approach in other applications such as text generation and question answering.

    Active Retrieval Augmented Generation · 2023 · DOI
  • The paper identifies a gap in the existing literature on retrieval augmented language models. The gap is that previous methods do not effectively address hallucination in generative language models.

    Active Retrieval Augmented Generation · 2023 · DOI
  • The paper identifies a gap in the integration of Boolean queries with probabilistic retrieval models. The paper notes that complex Boolean queries can be difficult to construct.

    Boolean queries and term dependencies in probabilistic retrieval models · 1986 · DOI
  • To implement the system in a real-world setting. To extend the system to handle other types of document retrievals. To improve the performance of the set of heuristics.

    An intelligent system for document retrieval in distributed office environments · 1986 · DOI
  • There is a need for a distributed system for document retrieval in office environments. The existing systems do not learn document distribution patterns and user interests and preferences. The existing systems do not customize document retrievals for each user.

    An intelligent system for document retrieval in distributed office environments · 1986 · DOI
  • Future research can focus on developing more effective weighted retrieval systems. Future research can explore the application of the threshold-value concept to other areas.

    A model for a weighted retrieval system · 1981 · DOI
  • Further evaluation of the extended models is needed. The use of other term significance weights should be explored. The application of the extended models to other domains should be considered.

    Document representation in probabilistic models of information retrieval · 1981 · DOI
  • The limitation of probabilistic models of retrieval assuming binary index terms is identified. The need to extend these models to include term significance weights is recognized.

    Document representation in probabilistic models of information retrieval · 1981 · DOI
  • Future research can explore the application of passage retrieval to other domains. The technique can be improved by developing more sophisticated search algorithms. The paper provides a framework for future research in passage retrieval.

    Answer‐passage retrieval by text searching · 1980 · DOI
  • The paper identifies a gap in the application of passage retrieval to scientists. The technique has been successfully applied to lawyers, but not to scientists.

    Answer‐passage retrieval by text searching · 1980 · DOI
  • Alternative term-weighting and ranking algorithm combinations should be explored. Further experimentation is needed to find ranking procedures that work better.

    Automatic ranked output from boolean searches in SIRE · 1977 · DOI
  • The challenge facing an information retrieval system is to present a user with references that fulfill their information need. There is a need to develop a better understanding of how the individual components of retrieval systems function.

    Automatic ranked output from boolean searches in SIRE · 1977 · DOI
  • The paper identifies a gap in the knowledge of how employees of institutions undergoing change are linked to the change process. The study indicates that state mental health employees at the institutional level feel genuinely unable to affect the change process.

    Footnotes (items of interest) · 1977 · DOI
  • The need for efficient retrieval from automated bibliographic data bases. The lack of knowledge on the most suitable combination of data elements for document retrieval.

    Relative effectiveness of titles, abstracts, and subject headings for machine retrieval from the COMPENDEX services · 1975 · DOI
  • The lack of a formal criterion for distinguishing between index and non-index words. The need for a probabilistic model that generalizes the pure random model to account for the observed distribution of content-bearing words.

    Probabilistic models for automatic indexing · 1974 · DOI
  • The potential 'washout' of a few large correlations by a host of smaller ones - The need to increase the size of large correlations at the expense of smaller ones

    SELECT: A computer program to identify associationally rich words for content analysis · 1974 · DOI
  • The need for a more objective and efficient method for selecting the best word subset for analysis - The limitations of frequency as a sole criterion for word selection

    SELECT: A computer program to identify associationally rich words for content analysis · 1974 · DOI
  • There is a need to compare conventional retrieval methods (MEDLARS) with automatic text analysis methods (SMART); There is a lack of understanding of the relative merits of controlled versus free language indexing and manual versus automatic analysis methodology

    A new comparison between conventional indexing (MEDLARS) and automatic text processing (SMART) · 1972 · DOI
  • The need to analyze the conditions under which various methods of sentence selection are successful - The need to develop criteria for selecting sentences to form an abstract

    Automatic abstracting and indexing. II. Production of indicative abstracts by application of contextual inference and syntactic coherence criteria · 1971 · DOI
  • Not enough is known about the behaviour of automatic keyword classifications. Few systematic experiments have been carried out on the properties of effective keyword classifications.

    What makes an automatic keyword classification effective? · 1971 · DOI
  • The paper identifies a gap in the use of controlled vocabulary subject indexing in archives. The author notes that systematic approaches have been made in the library field but not in archives.

    Controlled vocabulary subject indexing of archives∗ · 1971 · DOI
  • The difficulty of analyzing documents in different languages. The need for a complete and accurate multilingual thesaurus. The challenge of evaluating the effectiveness of mixed language processing.

    Automatic processing of foreign language documents · 1970 · DOI
  • Improving the completeness of the German thesaurus. Evaluating the effectiveness of mixed language processing in other languages. Developing more advanced linguistic tools for document analysis.

    Automatic processing of foreign language documents · 1970 · DOI
  • Organizing a network of urban observatories in major U.S. cities and urban regions. Conducting policy-oriented research on selected major issues of direct concern to mayors and others on the firing line.

    Urbandoc: Document Retrieval in Urban Affairs · 1965 · DOI
  • The need for machine translation between Japanese and English. The lack of commercial machine translation systems in Japan.

    Special Issue: “Collection of Best Annual Papers” Organized for the 20th Anniversary of the Association for Natural Language Processing · 2014 · DOI

Most-cited papers in Web Data Mining and Analysis

Most recent work

Find a gap in your own Web Data Mining and Analysis sub-topic

This page shows what the Web Data Mining and Analysis literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.

Open the Research Gap Finder →

Related topics in Computer Science

70 open questions have been extracted from the limitations and future-work passages of 648 Web Data Mining and Analysis papers in our library. Each one below links back to the study that raised it, so you can read the original claim in context.

Tools for your next paper

Compare the categoryHonest roundups of the AI research tools, ours listed alongside the alternatives.

Command palette

Jump anywhere, run any action.