Open research questions in Machine Learning in Materials Science
105 unresolved questions extracted from the limitations and future-work sections of 563 Machine Learning in Materials Science papers in our library. Each links back to the study that raised it.
What the literature leaves open
However, since the downstream demonstration in the present work focuses on bilayer water on Au(111), despite previous success of discovery approaches generally for a wide class of systems18,20,49, the broader applicability of the style translation augmentation to other molecular systems remains to be systematically validated in future ARTICLE IN PRESS The same consideration applies to experimental conditions.
The construction of machine-learned interatomic potentials (MLIPs) is often limited by the cost of generating large density-functional-theory (DFT) training datasets.
Data-Efficient Training of Linear ACE Potentials through Leverage-Guided Subset Selection of ASSYST Structure Pools · 2026Future work will explore its extension to variable-length sequences [68], integration of additional sensing modalities (e. However, the present validation is limited to a single graded build with a stratified window-level split for training, validation, and testing.
Learning composition-sensitive signatures in multi-material PBF-LB: a lightweight, modality-aware, explainable graph-attention sensor fusion framework for in-situ monitoring of graded 316L–CuCrZr alloys · 2026 · DOIHowever, as a comprehensive understanding of the property space defined by known functional molecules is lacking, assessing whether new AI/human-designed molecules truly surpass existing ones is challenging.
MolAtlas: a visualization framework for molecular property distributions to guide functional molecule development · 2026 · DOIAlthough recent advances in generative artificial intelligence have enabled the sampling of chemically plausible compositions and structures, a fundamental challenge remains: the objective misalignment between the likelihood-based sampling in generative modelling and the targeted focus on underexplored regions where novel compounds reside.
Guiding generative models to uncover diverse and novel crystals via reinforcement learning · 2026 · DOIFacing the challenge of scarce data for metal oxides, we employed Unsupervised Domain Adapta- tion (UDA) to successfully transfer knowledge from the well-established chemical space of metal hydrides.
Prediction and investigation of transition metal oxides for hydrogen storage application via machine learning-assisted density functional theory · 2026 · DOIStructural optimization for on-target potency purely based on ligand information from SAR knowledge and without external scoring has not been established.
While the data are limited, the observed trends appear robust, and the resulting fits provide reliable estimates even for values of N N beyond current computational feasibility.
Multidimensional modeling of the number of local minimum-energy structures and the energy of putative global minima · 2026 · DOIThe hybrid NODE is evaluated against a discrete-time feedforward neural network and a purely data-driven NODE under sparse data conditions, with models trained on as few as ten measurements under both regular and irregular sampling.
Hybrid Neural Ordinary Differential Equations for Data-Efficient Polymerization Modeling with Incomplete Kinetics · 2026In spite of its success, standard MFML schemes rely on pre-defined scaling factors to determine sparse data ratio across fidelities, often generating redundant multifidelity data resulting in a loss of efficiency.
Improvise, Adapt, Overcome: An On-The-Fly Multifidelity Algorithm for Efficient Machine Learning · 2026While much development has been done, there are still open challenges in how methods bridge across different time and/or length scales, how we inform, validate, and understand complex material behaviour, and more recently how emerging machine learning and artificial intelligence tools can enhance our multiscale modelling approaches.
Roadmap on novel computational approaches for bridging length and time scales: addressing challenges in modeling processes, characterization, and performance of metals and alloys · 2026 · DOIThe design of high-strength and high-thermal-conductivity Mg/Al alloys has transitioned from empirical discovery to computational materials engineering. This review bridges physical metallurgy and modern informatics strategies. The design concept is to precipitate solute atoms from the matrix, thereby restoring lattice integrity to facilitate electron transport, while simultaneously enhancing mechanical strength through the introduction of a second phase. This approach provides a viable pathway for mitigating the inherent trade-off between strength and electrical conductivity. We systematically analyze the computational toolkit driving this paradigm shift: CALPHAD delineates thermodynamic boundaries, DFT elucidates intrinsic transport physics, and ML accelerates the exploration of high-dimensional design spaces. While single-objective methods have laid the foundation, our review highlights that Multi-Objective Optimization frameworks represent a more robust strategy. These approaches utilize physics-based parameters as high-fidelity descriptors, enabling efficient mapping of the Pareto frontier to balance conflicting properties. Despite significant progress, several critical frontiers remain for next-generation alloy design. Foremost among these is the challenge of data scarcity and high-dimensional feature representation. High-quality datasets for TC are rare and heterogeneous. In addition, the high dimensionality of composition and processing spaces renders purely data-driven approaches inefficient. Future work would likely benefit from Transfer Learning (transferring knowledge from data-rich Al systems to Mg systems) and active learning loops, which are expected to maximize information gain from limited experiments. Concurrently, expert involvement in data curation, feature validation, and physical consistency checking remains indispensable, serving as a critical safeguard for data reliability and model robustness. A second major frontier involves enhancing physical consistency through explainable AI. Black-box models face limitations in accurately reflecting rigorous physical metallurgy principles. To enhance generalizability and build confidence in predictive models, the adoption of explainable AI methods, such as SHAP, is Page 22 of 27 Chen et al. J. Mater. Inf. 2026, 6, 31 essential. By integrating algorithmic approaches with expert domain knowledge, the patterns identified by AI can be interpreted from a domain-informed perspective and translated into meaningful scientific insights and technological innovations, thereby ensuring that model outputs remain consistent with established metallurgical principles. Finally, a pivotal objective for future research is the paradigm shift from discriminative screening and forward prediction to generative inverse design. Emerging generative AI techniques, such as GANs and VAEs, hold significant potential for effectively navigating the continuous latent space of material design. However, the future development of these technologies depends not only on the deployment of complex algorithms, but more importantly on the establishment of large-scale, high-fidelity datasets grounded in physical principles and curated with expert guidance. Such efforts are essential to ensure that AI-generated insights align with scientific reality. DECLARATIONS Authors’ contributions Made substantial contributions to the conception and design of this review, as well as writing and editing: Chen, Y.; Chen, H.; Luo, Q. Made substantial contributions to literature collation, figure preparation, and writing: Chen, Y.; Hu, L.; Tang, K.; Liu, B.; Zhang, Y.; Hu, B. Performed data acquisition and provided administrative, technical, and material support: Luo, Q.; Li, Q.
Computational strategies for the design of high-strength and high-thermal-conductivity casting Mg/Al alloys · 2026 · DOIIn summary, we have presented examples for how high-accu- racy simulations of molecules and chemical reactions can be per- formed using a combination of machine learning with semiclas- sical approximations to quantum dynamics. In future work, we plan to go beyond the approximations of instanton theory using path-integral molecular dynamics. [71,72] In principle, this method allows for the exact calculation of tunnelling splittings (within the errors of the underlying PES). Although it is significantly more expensive than the instanton approach, the computational cost can be met due to the development of efficient ML-PESs. In this way, we can have a direct test of the accuracy of the electronic-structure methods and the machine-learning procedures against experimen- tal measurements from high-resolution spectroscopy. In addition, the combined approach can be extended to both perturbatively corrected instanton rate theory [73] and nonadiabatic chemical reactions using golden-rule instanton theory. [74] In prin- ciple, machine learning can be used not just for potential energy surfaces, but also for nonadiabatic couplings. There is, however, an extra complication that nonadiabatic couplings are double val- ued (just as x2 = 4 has two solutions, x = ± 2) and change sign after winding around a conical intersection. This complication can be avoided by evaluating the outer product of the nonadiabatic cou- pling vector with itself, resulting in a single-valued entity that can be learned using standard methods. [75,76] The rapid evolution of the interface between ML and com- putational chemistry represents a powerful symbiosis rather than a replacement of traditional theory. By learning and inter- polating potential energy surfaces, ML models allow research- ers to bypass computationally expensive calculations, granting access to complex molecular systems that were previously out of reach. However, such data-driven acceleration works best when coupled with rigorous theoretical frameworks that include the underlying physics. In particular, we need high-accuracy electronic-structure theories to generate the training data and rigorous semiclassical theories to determine the dynamics on the resulting ML-PESs. Ultimately, while ML dramatically re- duces the computational cost, the expertise of the theoretician remains essential to correctly define physical problems, develop improved theories, and interpret the new phenomena revealed by such hybrid approaches.
High-Accuracy Molecular Simulations with Machine-Learning Potentials and Semiclassical Approximations to Quantum Dynamics · 2026 · DOIGenerative molecular artificial intelligence has matured con- siderably over the past decade and several companies and research groups are now actively deploying it in drug discovery and mate- rials design pipelines. In this column, we introduced how genera- tive models learn distributions over chemical space, how masked diffusion generates molecules by filling in masked fragments and how both ideas come together in models such as GenMol. One important caveat is that a high predicted property score is not a synthesis plan, since models can propose structures that are dif- ficult or impossible to make, [14] and many structures presented as novel turn out to be close variants of known compounds. [15] These tools are best used with ‘chemical judgment’, scepticism even. The best way to build it is to use these tools, test their limits and pay attention to where they fail. Have fun creating!
Thanks to the rise of machine learning in recent years, much progress has been made in the field of MLIPs. Increasing com- putational resources have made the generation of large training sets possible, such as the OMol25 dataset, and thus the devel- opment of foundational MLIPs has become feasible. In addi- tion, many attempts to explicitly include physics and long-range interactions have further advanced the field. However, there is a trade-off between the complexity of the model and the com- putational cost such that most models focus on one of the two aspects. As MLIPs are computationally more expensive than classi- cal force fields, multiscale ML/MM approaches provide a power- ful strategy for the simulation of larger systems in the condensed phase – ideally with an electrostatic embedding scheme to include the polarization of the ML zone by the MM environment. Due to the reduced computational cost of ML/MM compared to QM/ MM simulations, a much larger QM (or ML) zone can be treated at quantum accuracy (e.g. the entire enzyme in the study of en- zymatic reactions), which removes the need to handle covalent 04_Riniker_05_26_251985.indd 301 04_Riniker_05_26_251985.indd 301 04.05.26 11:55 04.05.26 11:55 302 CHIMIA 2026, 80, No. 5 The InTerface of aI and compuTaTIonal chemIsTry In swITzerland bonds between QM (ML) and MM particles and thus avoids the associated artifacts. In the future, we expect to see improvements in the accura- cy and general applicability of MLIPs in ML/MM simulations through advances in the training sets, architectural developments (e.g. physics-inspired or physics-augmented models, reduced complexity for increased speed), and/or going beyond electrostat- ic embedding. Polarizable-embedding schemes using polarizable force fields for the MM zone are known for QM/MM simulations, but to the best of our knowledge, no published ML/MM model has implemented for this scheme to-date. Overall, we anticipate that these developments will make ML/MM simulations of large systems and reactions in solution more accessible, providing a promising direction for the future.
Leveraging the Potential of Machine-Learning Interatomic Potentials for QM/MM Simulations · 2026 · DOIFive priorities emerge from the analysis above which were summarized in Figure 6. Experimentalists can strengthen the field by reporting provenance, failed runs, and validation context in machine-readable form. Model developers can increase usefulness by optimizing for extrapolation and calibrated uncertainty. Database builders can preserve operating conditions, negative data, and versioned metadata so that catalytic knowledge accumulates across papers. Device and reactor engineers can define the decisive endpoint early enough to shape the search space. Sustainability analysts can match proxy choice to durability, cost, criticality, and the intended decision horizon. Figure 6. Priorities for the next phase of artificial intelligence in the field of energy catalysis.
Artificial Intelligence in Energy Catalysis: From Catalyst Screening to Decision-Making · 2026 · DOICurrent challenges and inherent limitations Our analysis through the six-level framework reveals that the current research gaps in agentic MSE fall into two distinct categories, defined by the nature of the tasks involved. Cognition-centric challenges These challenges primarily emerge in tasks such as information retrieval, property prediction, and simulation, which operate in the digital domain and rely on the reasoning capabilities of LLMs. Despite rapid progress, current systems are constrained by the intrinsic limitations of LLMs when applied to scientific domains. MSE data are often sparse, heterogeneous, and highly structured, yet LLMs typically process such data as ungrounded text. Consequently, retrieval agents may miss critical context, property predictors may extrapolate beyond physical validity, and simulation planners may generate workflows that are linguistically coherent but numerically unstable. Fundamentally, these systems often lack robust mechanisms for enforcing physical laws, estimating uncertainty, and recovering from failures, hindering their progression to higher levels of autonomous reasoning. Execution-centric challenges This second category primarily appears in experimental synthesis and materials characterization, where the dominant difficulty shifts from language-based reasoning to physical interaction and real-world control. In materials science research, the execution bottleneck is driven by the heterogeneous and non-standardized Zhu et al. J. Mater. Inf. 2026, 6, 32 Page 25 of 37 nature of laboratory environments, including diverse software–hardware interfaces, inconsistent data formats, and the intrinsic variability of material samples. These factors introduce substantial noise and uncertainty into experimental processes, making reliable execution significantly more challenging than in purely digital settings. Such execution-centric settings also expose reliability problems in instruction adherence. Recent AFM automation studies have shown that LLM agents can take extra actions beyond the given protocol, sometimes acting as if they rely on prior context or memory rather than the current instruction - a behavior referred to as “sleepwalking”. This behavior can appear as risky physical actions beyond authorized limits or as functional code that exceeds the specified requirements, reflecting instruction drift during execution. Such behavior raises clear concerns regarding operational safety and the validity of closed-loop experiments. Recent advances in collaborative robotics and automated laboratories have led to the development of middleware frameworks, hardware standardization efforts, and communication protocols (e.g., SiLA, ChemOS[25,222], Robot Operating System), which provide an important technical pathway for device coordination and standardization in materials science labs.
A survey of agentic materials science and engineering: where are we and where are we going? · 2026 · DOIAmong GNN architectures spanning 2018–2025, single-scale GNNs dominate the literature but are provably limited by the 1-WL expressivity ceiling at their chosen cutoff radius.
A Survey on Graph Neural Networks for Crystal Property Prediction: Architectures, Expressivity, Uncertainty Quantification, Out-of-Distribution Generalization · 2026 · DOIIn this study, we proposed a digital modeling approach that integrates artificial intelligence and high throughput computational techniques to optimize steel materials production for sports equipment manufacturing. By employing the Manifold Guided Optimization Network (MGON), which combines deep learning, manifold learning, and probabilistic modeling, we successfully navigated the high dimensional design space of steel materials. The MGON framework, with its Counterfactual Constraint Encoder, Agent Driven Event Planner, and Probabilistic Uncertainty Filter, enabled efficient exploration of material properties while adhering to manufacturing constraints and performance objectives.
Digital modeling approach for optimizing steel materials production in sports equipment manufacturing using artificial intelligence and high-throughput computational techniques · 2026 · DOI(1) Alignment and interpretation: MatterChat’s behavioural success on property tasks may reflect learned correlations rather than deep semantic internalization of graph-based structural semantics. This limits interpretability and compositional reasoning involving struc- tural concepts. Addressing this requires explicit representation-level alignment objectives—such as contrastive losses, modality matching or shared embedding projections—to ensure the LLM fully grounds language in atomic representations59–62. (2) Data and reasoning: current training relies on single-turn ques- tion–answer pairs, lacking the multistep reasoning and cross-modal inference chaining essential for expert enquiry63–65. Future develop- ments should transition towards multiturn, multimodal dialogue trajectories. Techniques like phased instruction tuning66–68 and least-to-most prompting69 offer promising pathways for stepwise scientific problem-solving grounded in material structures. (3) Hallucination and reliability: frozen LLM backbones are susceptible to hallucinations in which language priors dominate structural information70–72. Although RAG provides initial contex- tual grounding73, future modular enhancements are necessary. These include multimodal fusion techniques (for example, mixture of features)74, domain-adaptive fine tuning on expert corpora75,76 and hallucination-aware training objectives25,77–79. Finally, post hoc correction frameworks—including fact-checking and self-revision loops—can further enhance the reliability of open-ended scientific responses80,81. Finally, although the current work prioritizes structure-informed reasoning, MatterChat’s modular architecture is designed for future extensibility to text-only materials benchmarks33,35–37. Its interchange- able components provide a flexible framework for potential systematic evaluation on tasks like synthesis question–answer classification from abstracts in future studies, offering a pathway to further bridge the gap between linguistic and structure-aware understanding.
Summary of Key Advances The catalyst design has entered a new phase of sophistication. This has been advanced by the capability of producing active sites with atomic accuracy by using single-atom catalysts, intermetallics, and defect engineering. It is also backed by profound knowledge of mechanisms made possible through operando characterization that shows dynamics of catalysts under operative condition. Experiment design and catalyst optimization are now directed by predictive computational frameworks that are premised on density functional theory, microkinetic modeling and machine learning. The increasing focus on the concept of sustainability has facilitated the use of electro- and photocatalytic methods which offer routes towards decarbon based chemical production. Additionally, the interdisciplinary collaboration of basic science and applied engineering issues allow to both design rationally, conduct tests, and scale up catalysts in order to develop a synergistic strategy that involves theoretical knowledge, experimental approaches, and application procedures. 7.2 Remaining Challenges Catalyst design still has a number of challenges despite the tremendous progress. It is important to bridge the complexity gap since most computational and experimental studies are done on well-defined model systems, whereas industrial catalysts are subject to complex, multicomponent, impure, and pressure gradient conditions as well as deactivation effects. The stability of catalysts and their deactivation remain to be significant concerns. The behavior of performance under conditions of interest in the industry, such as sintering, coking, poisoning, leaching, and phase transformations, is to be explored further under long-term conditions. Selectivity in complex reactions has not been easily attained, especially in reactions like the reduction of CO 2 or biomass whereby there are various competing pathways. Scaling advanced catalysts such as single-atom © 2026 The Author(s). Published by IJCOPE Journal. Website: https://ijcope.org/ 17 International Journal of Creative and Open Research in Engineering and Management ISSN: 3108-1754 (Online) Volume 02 Issue 04 April-2026 | Impact Factor: 3.5 catalysts, shape-controlled nanocrystals and metal-organic frameworks also has issues associated to reproducibility, cost, and throughput. Lastly, testing on catalysts needs to be standardized and reproducible, particularly in developing directions such as electrocatalytic nitrogen reduction, where the protocol and benchmarking are still absent. 7.3 Emerging Opportunities The area of catalysis is experiencing a number of exciting opportunities.
Smart Catalyst Design: Integrating Structure–Activity Relationships with Computational and Data-Driven Approaches · 2026 · DOIThe comparison between adaptive sampling methods and dropout-based residual selection strategies mentions that dropout incurs repeated residual evaluation overhead, but provides no direct computational comparison or accuracy trade-off analysis. Quantitative benchmarking of energy adaptive PINNs against dropout-based residual selection on the same phase transition test cases would clarify the efficiency-accuracy frontier.
Example 4.3 introduces a 2D phase transition problem where behavior changes significantly between times 0.9 and 1.0 as interfaces recede. The paper does not evaluate whether the energy adaptive or residual adaptive methods effectively detect and respond to these late-time regime shifts, nor does it test adaptive sampling performance on spatiotemporal problems with non-uniform temporal evolution rates.
The residual adaptive method fails to capture dual-interface structures in Example 4.2, losing the exact solution's interfacial topology in favor of a residual-wise simpler interpolant. The paper does not investigate why residual-based metrics fail at capturing multiple interfaces simultaneously or propose modifications to the residual adaptive sampling distribution to better represent complex interface geometries in phase field equations.
Example 4.2 identifies a specific failure mode where PDE loss at periodic boundaries conflicts with boundary loss terms when interfaces pass through the boundary, causing elevated PDE loss. The paper does not propose a modified weighting scheme or loss function decomposition to separately handle conflicting boundary and PDE constraints in energy adaptive PINNs for periodic domain problems.
Most-cited papers in Machine Learning in Materials Science
- Structured information extraction from scientific text with large language models · Nature Communications · 2024 · 594 citations
- Augmenting large language models with chemistry tools · Nature Machine Intelligence · 2024 · 565 citations
- Leveraging large language models for predictive chemistry · Nature Machine Intelligence · 2024 · 321 citations
- ChatMOF: an artificial intelligence system for predicting and generating metal-organic frameworks using large language models · Nature Communications · 2024 · 208 citations
- Harmonizing physical and deep learning modeling: A computationally efficient and interpretable approach for property prediction · Scripta Materialia · 2024 · 201 citations
- General-purpose machine-learned potential for 16 elemental metals and their alloys · Nature Communications · 2024 · 175 citations
- Exploring the use of large language models (LLMs) in chemical engineering education: Building core course problem models with Chat-GPT · Education for Chemical Engineers · 2023 · 158 citations
- Recent Advances in Machine Learning‐Assisted Multiscale Design of Energy Materials · Advanced Energy Materials · 2024 · 148 citations
- From Characterization to Discovery: Artificial Intelligence, Machine Learning and High-Throughput Experiments for Heterogeneous Catalyst Design · ACS Catalysis · 2024 · 147 citations
- Crystal structure generation with autoregressive large language modeling · Nature Communications · 2024 · 130 citations
Most recent work
- MOPAC: An open-source semiempirical molecular orbital program · The Journal of Open Source Software · 2026
- 3D convolutional neural network–driven ring size distribution prediction in magnesium-doped 45S5 bioactive glasses: An integrated molecular dynamics–machine learning framework · Materials Research Bulletin · 2026
- Artificial Intelligence-Driven Development and Characterization of Nanomedicine · BioNanoScience · 2026
- Enhanced climbing image nudged elastic band method with Hessian eigenmode alignment · Frontiers in Chemistry · 2026
- Scalable Boltzmann generators for equilibrium sampling of large-scale materials · Nature Communications · 2026
- Learning data-efficient coarse-grained molecular dynamics from forces and noise · Nature Communications · 2026
- NMR-Solver: automated structure elucidation via large-scale spectral matching and physics-guided fragment optimization · Nature Communications · 2026
- Angular relational knowledge distillation of machine learning interatomic potentials for scalable catalyst exploration · npj Computational Materials · 2026
- Accelerating the Design of High-Entropy Alloys for Hydrogen Storage via Generative Adversarial Networks · Inorganic Chemistry · 2026
- Accurate Chemistry Collection: Coupled cluster atomization energies for broad chemical space · Scientific Data · 2026
Find a gap in your own Machine Learning in Materials Science sub-topic
This page shows what the Machine Learning in Materials Science literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.
Open the Research Gap Finder →