Open research questions in Semantic Web and Ontologies
36 unresolved questions extracted from the limitations and future-work sections of 813 Semantic Web and Ontologies papers in our library. Each links back to the study that raised it.
What the literature leaves open
Improved accuracy and cold-start handling VI. REPRESENTATIVE SYSTEMS FROM THE LITERATURE Several concrete systems illustrate the taxonomy. Semantic recommendation frameworks for e-learning have combined a domain ontology with OWL rules so that learners are matched to materials that fit their field of interest and current requirements, using rule filtering as the recommendation mechanism. Such rule and ontology based designs demonstrate how explicit pedagogical policy can be encoded and applied, and they are robust to cold start because they do not require prior ratings for a new learner. Hybrid knowledge-based systems have coupled ontologies with sequential pattern mining and collaborative filtering. One influential design models the learner and learning resources in an ontology, uses collaborative filtering to predict interest, and mines sequential access patterns to order recommendations, reporting improved relevance over nonsemantic baselines. Related hybrids incorporate learner context and the sequence of previously accessed material, arguing that both the learner situation and the order of study should shape what is recommended next. Learner-preference and sequential-pattern models that operate over multidimensional descriptions of material similarly show how ontological structure and behavioural sequence can be fused. Earlier personalization systems, though not always framed as ontological, established the pattern of identifying learning styles and mining behaviour to adapt recommendations, and they informed subsequent semantic designs. More recent hybrid ontology-based approaches for online learning resource recommendation have refined these ideas, combining collaborative filtering with sequential pattern mining under an ontological representation and reporting measurable improvements on resource recommendation tasks. Across these systems a common architecture appears: an ontology layer that represents domain, learner, and pedagogy, a reasoning or similarity layer that interprets it, and a recommendation layer that ranks or sequences candidates. The reviews consolidate these examples and note recurring benefits in cold-start behaviour and semantic relevance, alongside recurring costs in ontology construction and upkeep,. It is worth reading these results with care. Reported improvements are usually relative to non-semantic baselines on the authors own data, and gains vary with the quality and coverage of the ontology, the size of the interaction log, and the metric used. Systems that fuse an ontology with collaborative filtering and sequential pattern mining tend to report the strongest results, which suggests that the ontology contributes most when it complements rather than replaces behavioural signal,.
Ontologies enter educational recommenders through several complementary mechanisms. The first is semantic similarity. Instead of comparing items by shared keywords, a system can measure how close two concepts are within the ontology graph, using path length, shared ancestors, or information content. This lets a recommender relate resources that use different terminology but describe the same idea, and it enriches both learner and item profiles with concepts inferred from the taxonomy. Hybrid learning material recommenders have combined such semantic profiles with collaborative signals and with the sequential patterns of previously accessed material to produce ordered suggestions,. Volume: 1 | Issue: 7 | July – 2026 | www.peerreviewjournal.in/index.php/prjcs | 4 The second mechanism is reasoning. Because OWL ontologies carry logical axioms, a reasoner can infer implicit facts, check consistency, and apply rules. In a recommender, rules can express pedagogical policy, for example that an advanced module should be withheld until its prerequisites are satisfied, or that a learner who failed an assessment should be offered remedial content. This rule-based filtering narrows candidate items to those that are pedagogically admissible before ranking them. The third mechanism is learner modelling. An ontology gives the learner profile structure and meaning, so that observed behaviour updates a semantically rich model of knowledge state, goals, and preferences rather than a flat vector, which in turn supports more targeted recommendation and clearer explanations. The fourth mechanism concerns sequencing and learning paths. Educational recommendation is often not about a single item but about the next best step in a trajectory toward a goal. Prerequisite relations in a domain ontology define a partial order over content, and reasoning over this order lets a system assemble a coherent path, avoid gaps, and adapt when a learner struggles. This connects to the pedagogical notion of the zone of proximal development, the band of tasks a learner can accomplish with appropriate support, since prerequisite and difficulty annotations let a recommender aim at material that is challenging but reachable. Ontology-based systems have used exactly this structure to generate personalized sequences and to recommend the next resource in a course,. Reviews note that prerequisite and competency reasoning is among the clearest advantages ontologies bring over purely data-driven recommenders. V. A TAXONOMY OF ONTOLOGY-BASED EDUCATIONAL RECOMMENDERS Surveying the literature suggests a taxonomy along the dominant role the ontology plays. In semantic enrichment approaches, the ontology mainly supplies similarity and profile expansion that feed an otherwise conventional contentbased or collaborative engine. In reasoning and rule-based approaches, an inference layer filters or ranks candidates according to logical constraints and pedagogical policy. In learner-modelling approaches, the ontology structures the student model and drives adaptation. In learning-path approaches, prerequisite and competency relations generate ordered sequences. Finally, hybrid ontology plus machine learning approaches combine ontological knowledge with statistical methods such as collaborative filtering, clustering, or sequential pattern mining, and improved hybrid designs of this kind have reported gains on online learning resource recommendation tasks. These categories are not mutually exclusive, and many systems occupy more than one, but the taxonomy clarifies where the ontology contributes. Figure 1 sketches the representative growth of publications on ontology-based educational recommenders over the last decade and a half, reflecting the sustained interest documented by the surveys rather than a precise bibliometric count. Table 2 details the taxonomy, associating each category with the ontology role, a typical technique, and representative lines of work identified in the reviewed literature. Fig. 1. Representative (schematic) growth in publications on ontology-based educational recommenders. Values are illustrative of the upward trend reported across the surveyed reviews, not an exact bibliometric measurement. Table 2.
Conclusion This paper introduced T-TExTS, a recommendation system that assists high school English Literature teachers in selecting texts that are diverse in genre, theme, subtheme, and author, yet similar in pedagogical merits, using a domain-specific knowledge graph. The system makes three contributions, each of which we treat as essential to the overall claim. First, a pedagogy-grounded ontology was constructed using the KNARM methodology, capturing not only genres and themes but also instructional and qualitative pedagogical elements (levels of meaning, text structure, language conventionality and clarity, knowledge demands) that prior text-recommendation systems have not modeled together. Second, this ontology was instantiated as a knowledge graph and used to drive a recommendation pipeline based on random-walk graph embeddings. Third, we conducted a comparative evaluation of four embedding strategies (DeepWalk, biased random walk, hybrid, and Node2Vec) across three dataset sizes (98, 196, and 351 texts) and two weight configurations, yielding the main empirical findings of this paper. 1 3T-TExTS (Teaching Text Expansion for Teacher Scaffolding): Enhancing… 70 Page 26 of 30 The central empirical finding is that traversal-level expert weighting alone does not outperform algorithmic structural tuning on either AUC or any ranking metric, but combining the two preserves both pedagogical interpretability and competitive ranking quality. Node2Vec, which operates on the same expert-curated graph as the biased random walk but applies no additional traversal-level weighting, achieves the highest AUC at every dataset size (0.9642–0.9750) and the strongest ranking metrics on the larger graphs. The hybrid embedding, which concatenates DeepWalk and biased random walk representations, retains a high AUC across all scales (0.9122– 0.9350) while staying within a few percentage points of Node2Vec on every ranking metric. This makes the hybrid model the most practical choice for deployment in a teacher-facing tool: it inherits the structural robustness of uniform exploration while exposing which expert-assigned weights influenced the result, providing the transparency that classroom-facing applications typically require. These findings extend prior work on KG-based recommendation in a specific direction. They demonstrate that high-fidelity, expert-curated knowledge graphs offer a viable foundation for specialized educational scaffolding tools at the modest data scales typical of curriculum-domain ontologies, and that the relative ordering of embedding methods that we observed is stable across three dataset sizes ranging from 98 to 351 texts.
T-TExTS (Teaching Text Expansion for Teacher Scaffolding): Enhancing Text Selection in High School Literature through Knowledge Graph-Based Recommendation · 2026 · DOITechnology combination analysis revealed equally mature research streams for OPC UA–KG and AAS–KG integration (14 publications each), yet architectures combining all three technologies remain scarce (4 publications).
Semantic Interoperability in Industry 4.0: A Systematic Mapping Study on Integrating Knowledge Graphs with AAS and OPC UA · 2026 · DOIFuture work should investigate alternative embedding strategies or heuristics that maintain semantic grounding while reducing computational footprint. Future work should explore: (1) augmenting the knowledge graph with timezone metadata at level; (2) developing specialized temporal reasoning modules that infer UTC offsets from conversational context (e.
TST is proposed as a constructive ontological framework and predicate grammar substrate, not as a completed formal logic, finalized implementation standard, or exhaustive metaphysical system. Its purpose in this paper is to establish the proposition and value of a minimally categorical, maximally constructive approach to ontology. Several areas require future development. 14.1 Admissibility Algebra TST depends on the idea that not all State transitions preserve identity continuity. A Thing persists only when its changing States remain admissibly connected. This is expressed informally through: 𝐶(𝑇, 𝑆1, 𝑆2, 𝑡1, 𝑡2) However, the framework still requires a more formal admissibility algebra: a system of rules for determining when a transition is identity-preserving, identity-breaking, or ambiguous. Such an algebra would need to account for: • bounded distinguishability, • structural continuity, • functional continuity, • causal continuity, • semantic continuity, • spatiotemporal continuity, • institutional continuity, • domain-specific admissibility rules. This is one of the most important future tasks because the Continuity Identity Principle depends on it. 14.2 Continuity Metrics Relatedly, TST requires formal continuity metrics. A bridge may be repaired and remain the same bridge. A road may be resurfaced and remain the same road. A legal organization may change officers and remain the same organization. But each case relies on different continuity criteria. Future work should develop ways to measure or evaluate continuity across domains, including: • physical continuity, • legal continuity, • semantic continuity, • operational continuity, • • representational continuity, identity continuity. These metrics may vary by domain, but TST needs a general framework for representing them. 14.3 Formal Semantics TST presently defines a constructive predicate grammar but does not yet provide a complete formal semantics. Future work should formalize TST using one or more suitable systems, such as: • Common Logic, • • temporal first-order logic, typed predicate logic, • modal logic for potentiality, • causal logic, • description logic extensions, • or hybrid logic systems. The goal would be to specify the truth conditions, inference rules, typing constraints, and admissibility requirements for TST predicates and operators. 14.4 OWL / RDF Representation TST should also be represented in semantic web formats. An OWL/RDF profile would need to model: • Thing, • State, • StateType, • TimeIndex, • RelationalState, • Transition, • Interpretation, • PotentialState, • EpistemicState, • CausalTrace, • Representation, • EmbeddingState. However, TST should not be prematurely reduced to OWL. OWL is a representation layer, not the ontology itself.
Things in States Through Time: A Constructive Ontological Framework for Predicate Logic and Derived Semantic Structure · 2026 · DOILLMs, such as GPT-3.5 with likely over 200 billion parameters, show a marked advantage in data extraction quality and ease due to their pretraining on extensive text corpora, even without fine-tuning on domainspecific datasets. The efficacy of pre-training is highlighted by GPT-3.5’s proficiency in recognizing material and chemical entities. The LlaMa-2 model, with 70 billion parameters, demonstrates comparatively limited capability in recognizing chemical entities and establishing correct entity relationships, hinting at consideration for potential improvement through fine-tuning on labeled datasets. The challenges specific to polymer literature for NER-based models are marked by the absence of a standardized naming convention for polymers and the requirement for manual efforts to identify entity relationships. Despite the promising performance of GPT-3.5, various limitations still exist for the extraction of data from polymer literature and their applications in polymer informatics. We discuss some specific issues and our goals for improvement below. (cid:129) Manual conversion is necessary to transform the extracted material names into machine-readable formats such as Simplified Molecular Input Line Entry System (SMILES) strings to make the datasets informatics-ready. Despite the incorporation of LLMs into the data extraction pipeline, the parsing of chemical structures from figures remains a significant challenge, particularly in polymer-related studies. This is due largely to the fact that polymer structures are often exclusively presented in figures, which obstructs the direct extraction and conversion of polymer chemistry into machine-readable SMILES strings from the text. In the future, integration of large-scale computer vision models with LLMs to efficiently identify and extract polymer molecules depicted in figures will enable immediate use of the extracted data for training ML models without the need for additional manual processing. (cid:129) The intricate nature of scientific texts, particularly in introducing material names across different sections and using abbreviations, makes establishing correct relationships between entities mentioned in different paragraphs or even sentences a difficult task. Our current pipelines extract data that is described completely in a specific paragraph by looking for all the required named entities (i.e., ‘material’, ‘property’, ‘value’ and ‘unit’) to establish correct relationships. However, the properties of polymers often rely on further information, such as molecular weights, temperature, synthesis and processing conditions, and morphology. This additional data also needs to be extracted from multiple paragraphs, while ensuring the preservation of valid relationships. Using a specific example of Fig.
Originality/value To the best of the author’s knowledge, the idea of subject classification of LOs through the reuse of search query terms combined with SKOS-based matching and expansion has not been investigated before in a federated scholarly setting.
With the dominance of service-oriented architecture, many enterprises have started providing their distributed web services in IoT as their core business system interface and business manner. Multimedia systems interact with various resources in the form of data, information and knowledge. Provision of resources as various services, such as data as a service, information as a service and knowledge as a serinformation and vice, specifically according to data, knowledge, seems to be an effective way of leveraging the current influence of the IoT trend.
Transformation-based processing of typed resources for multimedia sources in the IoT environment · 2019 · DOIis, that For every entity in the real world, we propose to measure the individual components of an entity element from a multidimensional perspective, to redescribe the entity element with different components. Types to which each component of all entities belong, constitute the final fully typed dimensional profile of each entity. In a network of connected entities in summarized scenarios, provided that the network is becoming complete, more obvious is that the entities are defining each other as comprising types of other types or being composed types of other entities. Figure 3 describes the architecture of our proposed fully typed and multidimensional system. Leaf nodes in the same layer belong to one dimension of typed data. Each node represents a kind of data type. An edge represents the ancestral relationship between different types. Each leaf node in the last layer represents the most granular type, which no longer has subtypes. Figure 4 shows an example of modeling a person according to a fully typed and multidimensional system. We treat the typed data of a person in two dimensions: a dynamic dimension and a static dimension. A person’s gender, height, weight and profession are currently static-typed data, and one’s hobby and pulse are dynamic-typed data. A dress code can have subtypes, such as color, style, and so on. After obtaining the fully typed and multidimensional expression of entities, we can use the newly defined entity element to build our GraphDIK. The formal expression is as follows: 123 Fig.
Transformation-based processing of typed resources for multimedia sources in the IoT environment · 2019 · DOIPersistent semantic identity in WordNetAlthough rarely studied, the persistence of semantic identity in the WordNet lexical database is crucial for the interoperability of all the resources that use WordNet data.
Abstract While Web Question Answering System (WQAS) has made great progress in Internet currently, a major limitation of the information sources that current WQASs are using is limited to static page texts.
Through our reviews of previous researches, we find several inadequacies such as low utilization ratio of WordNet and the lack of standardized evaluation and give some suggestions for future works.
PrOnto prototype is as far limited to assessment of the similarity between concepts by exact matching their label names and by the comparison of keywords associated with them.
Personal Ontologies for Knowledge Acquisition and Sharing in Collaborative PrOnto Framework · 2010 · DOIIn this paper the framework for ontology construction for a research institute has been proposed. The framework is organized in a distributed hierarchical structure, with lo- cal ontologies associated with individual employees and an integrated higher level group ontologies with concepts and relations promoted from individual profiles. Three main steps of ontology construction have been outlined, namely topic generation from documents, individual reflection on ontological profile and cross-level agreement between in- terested parties. Special, superior role of group leader has been emphasized. Some preliminary results for simple, but robust topic generation method have been presented. We believe the framework may be a better choice for a re- search institute trying to develop its own ontology of re- search topics for integration and management of knowledge resources than adaptation of any well-known domain on- tology or creation of global ontology by domain experts. Reflection and agreement stages have themselves an addi- tional value as they are driving processes of exploring sci- entific neighbourhood (individual reflection) and exchang- ing knowledge through debate (cross-level agreement). As such they may be seen as supporting creativity in scientific environment. What must be stressed here is that development is at the early stage and far from complete. There is still much of work to be done. More sophisticated methods for topic ex- traction from documents are to be tested, detailed specifica- tion of reflection and agreement phases and implementation of software component with appropriate human-computer interface is still to be worked out.
Ontology Creation Process in Knowledge Management Support System for a Research Institute · 2008 · DOICurrently, OWL Versioning in SKOS does not account for this refinement, lumping or other transformation of concepts (and their relationships) between different versions of concept schemes.
Identify versions of concept schemes Identify one-to-one changes of concepts between schemes The second function of OWL Versioning as included in OWL Overview does not account for a change in the concept, except where one concept (for example, “bananas”) wholly replaces another concept (for example, “plantains”).
Although the procedure used in semantic modelling to decompose scientific statements into welfare-relevant attributes has been applied to multiple contexts, the procedure of information extraction has not been described in sufficient detail to allow new modellers to easily understand the procedure.
Semantic modelling of animal welfare explained – Part 2: The basis of welfare weighting and usage of scientific information · 2026 · DOIMicrodata embedded in Web pages for purposes of facilitating indexing and search engine optimization are a potential source to augment KGs under some assumptions of complementarity and quality that have not been thoroughly explored to date.
Enhancing knowledge graphs with microdata and LLMs: the case of Schema.org and Wikidata in touristic information · 2024 · DOIAs little is known about where computational linguistics integration is gaining momentum beyond academic research and software engineering, the purpose of this systematic mapping study is to map in what areas of distributed knowledge management it is perceived to gain traction.
Systematic Mapping of Computational Linguistics in Distributed Knowledge Based Systems and Management · 2024 · DOIHowever, building good-quality and reliable datasets of PRs from a textbook is still an open issue, not just for automated annotation methods but even for manual annotation.
The analysis of the workflows allowed us to identify that there is no consensus regarding the stages, their nomenclatures, and technologies, besides presenting superficial discussions.
Workflow models for aggregating cultural heritage data on the web: A systematic literature review · 2021 · DOIStandards are not yet established for these domains, and hence they are difficult to describe and present, and methods are needed that will reflect the changes that will occur as the domains develop and mature.
This paper is a think piece about the possible future of bibliographic control; it provides a brief introduction to the Semantic Web and defines related terms, and it discusses granularity and structure issues and the lack of standards for the efficient display and indexing of bibliographic data.
Most-cited papers in Semantic Web and Ontologies
- Cobots in knowledge work · Journal of Business Research · 2020 · 275 citations
- Is Semantic Information Meaningful Data? · Philosophy and Phenomenological Research · 2005 · 253 citations
- A Survey of Knowledge Graph Reasoning on Graph Types: Static, Dynamic, and Multi-Modal · IEEE Transactions on Pattern Analysis and Machine Intelligence · 2024 · 212 citations
- A review of the semantic web field · Communications of the ACM · 2021 · 210 citations
- Less Data, More Knowledge: Building Next-Generation Semantic Communication Networks · IEEE Communications Surveys & Tutorials · 2024 · 181 citations
- Transformation-based processing of typed resources for multimedia sources in the IoT environment · Wireless Networks · 2019 · 93 citations
- Ontology-based knowledge management · Computer · 2002 · 80 citations
- Data extraction from polymer literature using large language models · Communications Materials · 2024 · 57 citations
- ESGReveal: An LLM-based approach for extracting structured data from ESG reports · Journal of Cleaner Production · 2024 · 50 citations
- Tracing the development of mapping knowledge domains · Scientometrics · 2021 · 43 citations
Most recent work
- VECTAETOS™ Canonical Ontology License v2.0 · Zenodo (CERN European Organization for Nuclear Research) · 2026
- Knowledge graphs generation from cultural heritage texts: combining LLMs and ontological engineering for scholarly debates · Journal of Documentation · 2026
- View Interactive Network of Ontologies Matching for Brazilian Gravimetric Data Integration · Zenodo (CERN European Organization for Nuclear Research) · 2026
- Ontology for Heritage Semantic System · Zenodo (CERN European Organization for Nuclear Research) · 2026
- Bridging “Nature” and “Spirit”: The CRMhs Ontology for the Integration of Heritage Science and Cultural Heritage Data · Heritage · 2026
- Things in States Through Time: A Constructive Ontological Framework for Predicate Logic and Derived Semantic Structure · Zenodo (CERN European Organization for Nuclear Research) · 2026
- Metaheuristics for Ontology-Based Information Extraction Rule Learning · Journal of Web Engineering · 2026
- LLMs for industrial databases: an agro-food production plant use case · Frontiers in Artificial Intelligence · 2026
- Ontology-driven software engineering using LLMs for knowledge graphs in engineering biology · bioRxiv · 2026
- Design of an improved validation model for semantic interoperability in IoT using MOVFRGD ECDTSA and CCEAICM process · Applied Network Science · 2026
Find a gap in your own Semantic Web and Ontologies sub-topic
This page shows what the Semantic Web and Ontologies literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.
Open the Research Gap Finder →