Social Sciences · Research topic

Open research questions in Artificial Intelligence in Law

68 unresolved questions extracted from the limitations and future-work sections of 525 Artificial Intelligence in Law papers in our library. Each links back to the study that raised it.

What the literature leaves open

  • AI systems are increasingly governed by natural language rules, yet a key challenge arising from reliance on language remains underexplored: interpretive ambiguity.

    Statutory construction and interpretation for AI · 2026 · DOI
  • Future studies should focus on the empirical validation of the proposed invariant and variable modular frameworks across various regional law colleges, alongside the development of standardized, automated criteria for detecting and evaluating the precise margins of acceptable AI usage in students' creative and research endeavors.

    Application of generative artificial intelligence in forming professionally oriented digital competence of future lawyers in college · 2026 · DOI
  • The paper proposes an integrated legal AI system that incorporates case outcome prediction and legal question answering. This shows that transparent models are capable of competitive performance (91.3% accuracy) without sacrificing interpretability. On 26,688 Indian Supreme Court cases, logistic regression with confidence calibration has a higher performance (r = 0.73) by 9.3 percentage points over existing work. RAG-based legal question answering has 86% correctness with 7% hallucinations, which is 70% less than existing baseline LLMs. The work is shown to be accessible to non-expert citizens with 4.0/5 user satisfaction across three languages: English, Hindi, and Tamil, along with interactive explainability visualizations. Key Takeaways: (1) Interpretability in ML can be a viable approach to high-stakes legal prediction with careful feature engineering and confidence calibration. (2) RAG architecture can significantly reduce LLM hallucinations in domain-specific settings. (3) Transparent AI systems with confidence gauges, feature importance, and similar case retrieval can achieve 92% professional comprehension for trustworthy deployment. (4) Achieving 79% to 82% translation approval for multilingual accessibility also proves the feasibility of equitable legal AI in India’s linguistic diversity. User evaluation with 35 participants, consisting of 15 law students, 12 legal professionals, and 8 general users, recorded an overall satisfaction score of 4.0/5. The most valuable features, as rated by users, are multilingual support at 4.5/5, explainability visualizations at 4.3/5, prediction accuracy at 4.1/5, and answer quality at 3.9/5. Comparison of expert ratings with existing legal AI tools shows that the proposed system has the highest Indian law specificity with a favorable balance between clarity and completeness.

    Bridging Legal Language Barriers Using Explainable AI: Outcome Prediction and Multilingual Knowledge based answer retrieval for Indian Law · 2026 · DOI
  • Translation-based entity alignment (Sun et al., 2017; Efficiently models one-to-one relations in KGs by Primarily focuses on local structures. Limited ability to Zhu et al., 2017, 2019) translating a relationships from the source to the capture complex relationships. target entity. GNN-based entity alignment (Cao et al., 2019; Guo et al., Designed to handle the complexity of Limited ability to capture distant relationship 2018, 2021; Sang et al., 2023) heterogeneous KGs by capturing the semantics and relies on high-quality seeds. Requires neighborhood of an entity. hyperparameter tuning for optimal performance.

    Integrating knowledge graphs and contextual patterns for legal concept alignment · 2026 · DOI
  • VII. Support for Multilingual and Voice First • • Fine-Tuning Legal LLM • Case tracking and personalized legal • Mobile app and off-line capabilities • Integration with government/ NGO ecosystem VIII. REFERENCES S. Gupta, R. Mehta, and P. Agarwal, "AI in Legal Domain: Semantic Understanding of Law Documents," in Proc. IEEE Conference on Computational Linguistics, 2021. P. Ramesh and S. Kulkarni, "Building Conversational Legal Chatbots Using Machine Learning," IEEE Transactions on Computational Social Systems, 2022. A. Kumar and V. Narayan, "Semantic Search Techniques for Legal Information Retrieval," Springer International Conference on Data Engineering and Applications, 2020. N. Desai and R. Sinha, "Bridging Legal Awareness through AI: Opportunities and Challenges," International Journal of AI Applications and Innovation, vol. 14, no. 2, pp. 45–62, 2023. C. Trivedi, S. Kumar, I. Mohd, R. Bhalla, N. A. Lone, and D. Dogra, "Leveraging AI-Driven Chatbots for Legal Literacy," IEEE Access, vol. 12, pp. 78213–78229, 2024. Nikita, E. Srivastav, A. Patel, A. Singh, R. Sharma, D. P. Rana, and R. G. Mehta, "LAWBOT: A Smart User Indian Legal Chatbot using Machine Learning Framework," in Proc. International Conference on Emerging Technologies in Computing, 2024. K. D. Ashley, Artificial Intelligence and Legal Analytics: New Tools for Law Practice in the Digital Age. Cambridge University Press, 2017. M. Medvedeva, M. Vols, and M. Wieling, "Using Machine Learning to Predict Decisions of the European Court of Human Rights," Artificial Intelligence and Law, vol. 28, no. 2, pp. 237–266, 2020. E. S. Kamarudin and M. Ismail, "Promoting Civic Engagement through Digital Platforms: The Case for Legal Awareness," International Journal of Law and Information Technology, vol. 28, no. 3, pp. 201– 224, 2020. S. P. Smith, "Combating Legal Misinformation: The Role of Technology in Public Education," Journal of Legal Communication and Rhetoric, vol. 17, no. 1, pp. 89–112, 2020. N. Reimers and I. Gurevych, "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks," in Proc. EMNLP 2019, 2019. Meta AI, "LLaMA: Open and Efficient Foundation Language Models," arXiv preprint arXiv:2302.13971, 2023.

    JustiFind: An Intelligent Legal Aid and Awareness System · 2026 · DOI
  • While the proposed pedagogical framework draws on multiple empirical studies and offers a theoretically coherent approach to teaching polysemous legal adjectives, several limitations constrain the generalizability and practical implementation of the findings. 1 3K. Chiknaverova Scope and representativeness of empirical evidence. The studies informing this model originate from diverse instructional contexts, learner populations, and profi- ciency levels, but few focus specifically on legal English or the teaching of polyse- mous legal adjectives. Much of the supporting evidence derives from general EFL adjective instruction or from corpus-based analyses rather than controlled experi- mental research on legal discourse. Consequently, while the principles of inductive discovery, systemic sequencing, and corpus-informed input are transferable, their exact efficacy in specialized legal settings remains to be empirically verified. Contextual and curricular constraints. Implementation of a hybrid model that combines inductive tasks, corpus consultation, and authentic case analysis presup- poses access to legal corpora, instructor expertise in both applied linguistics and law, and sufficient instructional time to accommodate multi-stage activities. Such resources may not be available in all institutional settings, particularly in programs with large class sizes, rigid curricula, or limited technological infrastructure. Variability in learner profiles. Learners’ linguistic and disciplinary backgrounds, professional goals, and familiarity with legal systems can significantly influence how they process polysemous adjectives. For example, law students with prior exposure to legal terminology may benefit more from prototype-based sense mapping than general EFL learners preparing for non-legal careers. The present model does not fully account for such heterogeneity, and further research is needed to identify learner variables that mediate instructional outcomes. Methodological constraints of existing studies. Many of the cited studies rely on small sample sizes, quasi-experimental designs, or short-term interventions, limit- ing claims about long-term retention, transfer to real-world legal tasks, or scalability across different educational contexts. In addition, most research measures gains in accuracy or collocational knowledge but provides limited data on pragmatic appro- priacy or interpretive flexibility—skills that are critical for legal communication. Acknowledging these limitations underscores the need for longitudinal, context- specific studies that evaluate the effectiveness of blended inductive–systemic instruc- tion in legal English classrooms and explore learner-specific variables.

    Adjectival Polysemy in Legal English: Semantic Analysis and Pedagogical Implications · 2026 · DOI
  • The DRILL Shared Task highlighted how NLP methods can effectively support text retrieval across a wide range of legal domains. In two months, from the call for participation to the end of the private test phase, we received over 1,100 submissions from more than 50 teams, underscoring both the community’s enthusiasm xx T. H. Y. Vuong et al. / VNU Journal of Science: Com. Science & Com. Eng. Vol. 42, No. xx (2026) xx–xx and the importance of this research challenge. The strongest systems combined LLMs with fine-tuned pre-trained models and consistently outperformed the baselines. However, our error analysis reveals that commonsense knowledge, extremely long context, and temporal relations are the main bottlenecks of the current systems, suggesting avenues for future work. We are the DRILL Shared pleased to conclude that Task was organized successfully, offering a solid foundation for further developments in legal NLP for Vietnamese law.

    DRILL Shared Task 2025: The Challenge of Deep Retrieval in the Expansive Legal Landscape · 2026 · DOI
  • The ongoing incorporation of artificial intelligence into legal frameworks necessitates proactive governance, institutional adjustment, and interdisciplinary cooperation. Future 18 advancements must emphasize legitimacy, human oversight, and systemic resilience over mere technological progress. 7.1 Hybrid Decision-Making A hybrid decision-making framework—often described as a “human-in-the-loop” model—represents the most viable path forward. In this structure: • AI systems conduct large-scale data analysis, pattern recognition, and precedent mapping. • Human judges and lawyers retain interpretative authority and ethical discretion. • Algorithmic outputs function as advisory tools rather than determinative mandates. This method retains the efficacy and analytical capabilities of AI while ensuring normative accountability within human organizations. Hybrid systems can alleviate deficiencies on both fronts: AI diminishes irregularity and errors due to weariness, while humans address algorithmic inflexibility and contextual limitations. Empirical study indicates that the precision and equity of decisions are enhanced when algorithmic recommendations are subjected to critical assessment rather than being uncritically accepted. Consequently, it is imperative to educate legal professionals about the capabilities and limitations of AI. Legal education may progressively include technology proficiency, algorithmic auditing concepts, and multidisciplinary ethics. 7.2 Regulatory Frameworks Strong regulatory frameworks are essential to ensuring that AI implementation adheres to constitutional principles and human rights standards. International and regional governance projects offer valuable frameworks. The European Commission has delineated ethical principles for reliable AI, highlighting the importance of human oversight, transparency, and accountability. The Institute of Electrical and Electronics Engineers has established ethical guidelines around algorithmic accountability and transparency.

    Artificial Intelligence and Human Intelligence in Legal Systems · 2026 · DOI
  • In this work, we proposed a legal argumentation model that emulates adversarial legal reasoning through the application of fine-tuned generative models with retrieval. The approach consisted of gathering real-world data, triplets extraction, and constructing a counter-argument generator with legal logic and rhetorical devices. The system demonstrated its ability to engage in formal legal arguments and presented initial results regarding the feasibility of applying such simulations to qualitative outcome simulation. Through an advancement that transcends superficial text generation, the model demonstrated strategic legal reasoning rooted in established precedents and procedural standards. It presents a valuable instrument for legal professionals to evaluate the viability of arguments, foresee counter-arguments, and replicate intri- cate legal interactions before the initiation of courtroom proceedings. Despite such contributions, the system is not without its intrinsic limitations. Its coverage domain is limited to a certain subset of Indian civil and criminal cases, and this can limit its applicability in other jurisdictions. Furthermore, the model is yet to prove the ability to reason explicitly regarding the hierarchical nature of statutes, and its per- formance is still subject to the initial clarity and organization of inputs by humans. Additionally, the RAG system may require a larger set of documents to improve retrieval quality. The computational proxies used in the persuasiveness metric, such as entailment, sentiment analysis, and word count ratio, are recognized as proxies for complex legal reasoning concepts. Although each of them is conceptually associated with its respective dimension, there is no exact match to how persuasion works in a legal setting. The validity of these proxies for legal persuasion can only be seen within the framework of a diagnostic support system rather than a legal evaluator. Further research should concentrate on building legal persuasion datasets to replace or improve these general-purpose proxies While the current implementation employs GPT-3.5-turbo as the underlying language model, the proposed framework is model- agnostic. Future work may explore the integration of more recent or specialized large language models, which could further enhance legal reasoning depth, factual ground- ing, and argument quality. Anticipating future inquiries, numerous pathways present promising opportunities for subsequent investigation: AI-driven legal argument strength analyzer & counter-argument…1 3 ● Argument Weighting Metric: The first measure suggested in Section 5.6 can be formulated as a scoring function that measures the strength of counter-arguments and makes probabilistic claims about case outcomes. ● Case-Law Graphs: Building citation networks and legal knowledge graphs would allow the system to find arguments in terms of a nested hierarchy of precedents, with greater depth and interpretability. ● Interactive Courtroom Simulations: An interactive and user-centered system would enable legal professionals to rehearse courtroom dialogue, experiment with various argumentative approaches, and receive tactical guidance. ● Multilingual Extension: Expanding the system to accommodate regional Indian languages and other jurisdictions would extend its applicability and enhance us- ability for a multilingual user population. Essentially, this research is the basis of artificial intelligence-based legal decision- making support. It offers scope to create high-tech tools to reshape the process of crafting litigation tactics, legal studies, and support that judicial systems can harness in the modern digital age. A Prompt Templates Counter-Argument Generation Prompt T. Harde et al.1 3 Data Availability In line with Open Science principles, the code and evaluation artifacts used in this study have been fully implemented and are currently being prepared for public release. The implementation includes the counter-argument generation pipeline, multi-turn debate simulation framework, argument and citation scoring components. The public repository link will be released upon acceptance of the manu- script. This approach balances transparency with responsible data sharing considerations.

    AI-driven legal argument strength analyzer & counter-argument generator · 2026 · DOI
  • The research introduced ContractIQ, which is an artificial intelligence contract intelligence platform that automatically performs contract analysis and clause extraction and legal risk assessment. The system uses Natural Language Processing together with GPT-4omini and RAG to deliver precise legal contract analysis that considers contextual information. The model successfully extracts essential contractual sections while conducting semantic searches and producing fact-based replies that include citations. The testing results reveal that the system achieves better contract analysis performance because it demonstrates high success in retrieving information and operates at quick processing speeds. The automated verification and risk detection modules help identify missing clauses and ambiguous language, which helps legal professionals make better decisions for their cases. RAG integration establishes better system reliability because it allows system responses to be verified through existing source documents.

    ContractIQ: An AI-Based Contract Intelligence Platform for Automated Verification and Legal Risk Analysis · 2026 · DOI
  • In: Proceedings of the 28th international conference on computational linguistics, pp 6229–6239 Sansone C, Sperlí G (2022) Legal information retrieval systems: State-of-the-art and open issues.

    Comma: A multi-task and multi-lingual dataset of constitutional verdicts · 2026 · DOI
  • No evaluation was conducted on whether the findings regarding instruction tuning effectiveness, retrieval augmentation necessity, and hybrid method advantages for citation prediction generalize to other knowledge-intensive legal tasks beyond citation prediction (e.g., case outcome prediction, legal argument relevance ranking).

    Legal citation prediction with LLMs: a comparative evaluation of instruction tuning, retrieval, and jurisdiction-specific pre-training on the AusLaw citation benchmark · 2026 · DOI
  • The study evaluated hybrid ensemble methods (voting-based combinations of LLM and retrieval) for reducing hallucination in legal citation prediction, but did not systematically explore alternative ensemble architectures, weighted voting schemes, or confidence-based selection mechanisms that might further improve accuracy-grounding trade-offs.

    Legal citation prediction with LLMs: a comparative evaluation of instruction tuning, retrieval, and jurisdiction-specific pre-training on the AusLaw citation benchmark · 2026 · DOI
  • The comparison between sparse lexical retrieval and dense vector retrieval methods for legal citation prediction is incomplete. The paper notes that sparse lexical retrieval remains competitive with dense methods but does not provide detailed ablation studies or cross-jurisdiction evaluation to determine when each approach should be preferred.

    Legal citation prediction with LLMs: a comparative evaluation of instruction tuning, retrieval, and jurisdiction-specific pre-training on the AusLaw citation benchmark · 2026 · DOI
  • While the paper demonstrates that Reason-of-Citation (RoC) aggregations outperform full case texts and catchwords for retrieval quality in citation prediction, no systematic investigation was conducted comparing different RoC granularity levels or exploring optimal semantic chunking strategies for legal text representation.

    Legal citation prediction with LLMs: a comparative evaluation of instruction tuning, retrieval, and jurisdiction-specific pre-training on the AusLaw citation benchmark · 2026 · DOI
  • The AusLaw Citation Benchmark exhibits severe class imbalance with a long tail distribution of rarely cited cases. Current models achieve below 60% accuracy for cases with fewer than 20 citations despite near-100% accuracy on frequently cited cases (>100 citations). Targeted approaches for improving generalization to under-represented cases in legal citation prediction need investigation.

    Legal citation prediction with LLMs: a comparative evaluation of instruction tuning, retrieval, and jurisdiction-specific pre-training on the AusLaw citation benchmark · 2026 · DOI
  • The study did not implement the restoration phase required to preserve general-purpose LLM capabilities (chat functionality, reasoning across domains, alignment safeguards) when domain specialization is applied through pre-training. Research is needed to develop training curricula that integrate domain-specific legal material with diverse reasoning and generic capabilities to create fully capable domain-augmented legal LLMs.

    Legal citation prediction with LLMs: a comparative evaluation of instruction tuning, retrieval, and jurisdiction-specific pre-training on the AusLaw citation benchmark · 2026 · DOI
  • The jurisdiction-specific pre-training experiment was limited to only 5 epochs due to infrastructure constraints on the 0.5B Australian legal corpus. A complete pre-training regime with substantially more epochs is needed to fully evaluate whether extended Australian-specific pre-training can further improve legal citation prediction performance beyond the preliminary 52.0% accuracy achieved by Cite-AusLawLM-7B.

    Legal citation prediction with LLMs: a comparative evaluation of instruction tuning, retrieval, and jurisdiction-specific pre-training on the AusLaw citation benchmark · 2026 · DOI
  • The system architecture aligns proof generation with discrete audit events rather than per-query inference, but the paper does not empirically validate the latency and throughput characteristics of this batch-oriented approach when multiple litigants request simultaneous demographic threshold verifications in a high-volume prelitigation marketplace.

    An agentic AI marketplace for prelitigation analyses with ZKP-integrated ethical verifications · 2026 · DOI
  • The inequality constraint enforcement using range constraints or bit-decomposition is mentioned abstractly without specifying the exact bit-width decomposition strategy, the number of additional constraints introduced, or how different constraint encoding techniques affect total circuit size and prover performance in practical demographic threshold evaluations.

    An agentic AI marketplace for prelitigation analyses with ZKP-integrated ethical verifications · 2026 · DOI
  • The paper estimates Ethereum on-chain verification cost at 2.3×10⁵ gas units but provides no empirical benchmarking on actual Ethereum testnet or mainnet, nor comparative cost analysis for permissioned blockchain deployments (e.g., Hyperledger Fabric) where verification overhead may differ significantly.

    An agentic AI marketplace for prelitigation analyses with ZKP-integrated ethical verifications · 2026 · DOI
  • The prototype simulation assumes static, pre-compiled circuits with fixed demographic predicates. The paper does not address how the circuit design must evolve to support dynamic predicate modification across audit intervals or how proof generation time scales when predicates change between successive marketplace verification events.

    An agentic AI marketplace for prelitigation analyses with ZKP-integrated ethical verifications · 2026 · DOI
  • The on-chain commitment strategy for anchoring dataset integrity via Merkle root is described as optional but lacks formal analysis of when on-chain commitment is necessary versus when off-chain auditor signatures alone suffice, and no evaluation of Merkle tree depth impact on circuit constraints for realistic dataset sizes in prelitigation AI contexts.

    An agentic AI marketplace for prelitigation analyses with ZKP-integrated ethical verifications · 2026 · DOI
  • The paper theoretically describes selective disclosure signature schemes (BBS+ over BLS12-381) for binding metadata provenance but does not specify the exact signature encoding strategy within the arithmetic circuit or empirically measure the constraint overhead of signature verification against predicate evaluation in the integrated protocol.

    An agentic AI marketplace for prelitigation analyses with ZKP-integrated ethical verifications · 2026 · DOI
  • The batched aggregation and recursive proof composition techniques mentioned for bounding circuit size in large-scale demographic threshold verification are not formally specified or empirically validated. The paper lacks concrete algorithms for decomposing aggregation across tens of thousands of metadata entries while maintaining consistent verification complexity.

    An agentic AI marketplace for prelitigation analyses with ZKP-integrated ethical verifications · 2026 · DOI

Most-cited papers in Artificial Intelligence in Law

Most recent work

Find a gap in your own Artificial Intelligence in Law sub-topic

This page shows what the Artificial Intelligence in Law literature already flags as unresolved. To narrow it to your specific question, run the guided finder — it searches the gap library on demand and checks candidates against 250M+ OpenAlex works.

Open the Research Gap Finder →

Related topics in Social Sciences

68 open questions have been extracted from the limitations and future-work passages of 525 Artificial Intelligence in Law papers in our library. Each one below links back to the study that raised it, so you can read the original claim in context.

Tools for your next paper

Compare the categoryHonest roundups of the AI research tools, ours listed alongside the alternatives.

Command palette

Jump anywhere, run any action.