LegalOSS82.1ktracked
Legal AI & NLP · Benchmarks & Datasets

LLM-and-Law

Jeryi-Sun/LLM-and-Law

Continually updated LLM and law papers

LLM-and-Law

Awesome License: MIT

This repository tracks papers and resources about large language models (LLMs) in the legal domain.

Daily update 2026-07-23 (cs.CL): scanned 70 papers (date=2026-07-23), 0 new papers met the strict "legal task + LLM semantics" inclusion rule.

Category increments today — Applications +0, Legal Reasoning Models +0, Legal Agent +0, Legal Problems +0, Data Resources +0, Law LLMs +0, Evaluation +0.

Contents

Applications of Large Language Models in Legal Tasks

  • Legal Prompt Engineering for Multilingual Legal Judgement Prediction

  • Can GPT-3 Perform Statutory Reasoning?

  • Legal Prompting: Teaching a Language Model to Think Like a Lawyer

  • Large Language Models as Fiduciaries: A Case Study Toward Robustly Communicating With Artificial Intelligence Through Legal Standards

  • ChatGPT Goes to Law School

  • ChatGPT, Professor of Law

  • ChatGPT & Generative AI Systems as Quasi-Expert Legal Advice Lawyers - Case Study Considering Potential Appeal Against Conviction of Tom Hayes

  • ‘Words Are Flowing Out Like Endless Rain Into a Paper Cup’: ChatGPT & Law School Assessments

  • ChatGPT by OpenAI: The End of Litigation Lawyers?

  • Law Informs Code: A Legal Informatics Approach to Aligning Artificial Intelligence with Humans

  • ChatGPT may Pass the Bar Exam soon, but has a Long Way to Go for the LexGLUE benchmark paper

  • How Ready are Pre-trained Abstractive Models and LLMs for Legal Case Judgement Summarization? paper

  • Explaining Legal Concepts with Augmented Large Language Models (GPT-4) paper

  • Garbage in, garbage out: Zero-shot detection of crime using Large Language Models paper

  • Legal Summarisation through LLMs: The PRODIGIT Project paper

  • Black-Box Analysis: GPTs Across Time in Legal Textual Entailment Task paper

  • PolicyGPT: Automated Analysis of Privacy Policies with Large Language Models

  • Reformulating Domain Adaptation of Large Language Models as Adapt-Retrieve-Revise paper

  • Precedent-Enhanced Legal Judgment Prediction with LLM and Domain-Model Collaboration paper

  • From Text to Structure: Using Large Language Models to Support the Development of Legal Expert Systems paper

  • Boosting legal case retrieval by query content selection with large language models paper

  • LLMediator: GPT-4 Assisted Online Dispute Resolution paper

  • Employing Label Models on ChatGPT Answers Improves Legal Text Entailment Performance paper

  • LLaMandement: Large Language Models for Summarization of French Legislative Proposals paper

  • Logic Rules as Explanations for Legal Case Retrieval paper Our new paper, welcome to pay attention !!!

  • Enhancing Legal Document Retrieval: A Multi-Phase Approach with Large Language Models paper

  • BLADE: Enhancing Black-box Large Language Models with Small Domain-Specific Models paper

  • A Survey on Large Language Models for Critical Societal Domains: Finance, Healthcare, and Law paper

  • Archimedes-AUEB at SemEval-2024 Task 5: LLM explains Civil Procedure paper

  • More Than Catastrophic Forgetting: Integrating General Capabilities For Domain-Specific LLMs paper

  • Legal Documents Drafting with Fine-Tuned Pre-Trained Large Language Model paper

  • Knowledge-Infused Legal Wisdom: Navigating LLM Consultation through the Lens of Diagnostics and Positive-Unlabeled Reinforcement Learning paper

  • GOLDCOIN: Grounding Large Language Models in Privacy Laws via Contextual Integrity Theory paper

  • Enabling Discriminative Reasoning in Large Language Models for Legal Judgment Prediction paper

  • Large Language Models for Judicial Entity Extraction: A Comparative Study paper

  • Applicability of Large Language Models and Generative Models for Legal Case Judgement Summarization paper

  • LawLLM: Law Large Language Model for the US Legal System paper

  • Legal syllogism prompting: Teaching large language models for legal judgment prediction paper

  • Optimizing Numerical Estimation and Operational Efficiency in the Legal Domain through Large Language Models paper

  • Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools paper

  • KRAG Framework for Enhancing LLMs in the Legal Domain paper

  • Legal Evaluations and Challenges of Large Language Models paper

  • Analyzing Images of Legal Documents: Toward Multi-Modal LLMs for Access to Justice paper

  • Automating Legal Concept Interpretation with LLMs: Retrieval, Generation, and Evaluation paper

  • Domaino1s: Guiding LLM Reasoning for Explainable Answers in High-Stakes Domains paper

  • RELexED: Retrieval-Enhanced Legal Summarization with Exemplar Diversity paper

  • Courtroom-LLM: A Legal-Inspired Multi-LLM Framework for Resolving Ambiguous Text Classifications paper

  • LegalViz: Legal Text Visualization by Text To Diagram Generation paper

  • Investigating the Shortcomings of LLMs in Step-by-Step Legal Reasoning paper

  • Elevating Legal LLM Responses: Harnessing Trainable Logical Structures and Semantic Knowledge with Legal Reasoning paper

  • JuDGE: Benchmarking Judgment Document Generation for Chinese Legal System paper

  • LARGE: Legal Retrieval Augmented Generation Evaluation Tool paper

  • Tasks and Roles in Legal AI: Data Curation, Annotation, and Verification paper

  • The Hidden Structure -- Improving Legal Document Understanding Through Explicit Text Formatting paper

  • Conditioning Large Language Models on Legal Systems? Detecting Punishable Hate Speech paper

  • CORE-KG: An LLM-Driven Knowledge Graph Construction Framework for Human Smuggling Networks paper

  • LLM-based Embedders for Prior Case Retrieval paper

  • Towards Automated Regulatory Compliance Verification in Financial Auditing with Large Language Models paper

  • Using Large Language Models for Legal Decision-Making in Austrian Value-Added Tax Law: An Experimental Study paper

  • ASVRI-Legal: Fine-Tuning LLMs with Retrieval Augmented Generation for Enhanced Legal Regulation paper

  • LLMs in Interpreting Legal Documents paper

  • HalluGraph: Auditable Hallucination Detection for Legal RAG Systems via Knowledge Graph Alignment paper

  • Evaluating Legal Reasoning Traces with Legal Issue Tree Rubrics paper

  • Jurisdiction as Structural Barrier: How Privacy Policy Organization May Reduce Visibility of Substantive Disclosures paper

  • Enforcing Monotonic Progress in Legal Cross-Examination: Preventing Long-Horizon Stagnation in LLM-Based Inquiry paper

  • Anchored Decoding: Provably Reducing Copyright Risk for Any Language Model paper

  • Synthesizing the Virtual Advocate: A Multi-Persona Speech Generation Framework for Diverse Linguistic Jurisdictions in Indic Languages paper

  • AI-Assisted Moot Courts: Simulating Justice-Specific Questioning in Oral Arguments paper

  • LLM-Assisted Causal Structure Disambiguation and Factor Extraction for Legal Judgment Prediction paper

  • De Jure: Iterative LLM Self-Refinement for Structured Extraction of Regulatory Rules paper

  • ContextLens: Modeling Imperfect Privacy and Safety Context for Legal Compliance paper

  • Eskwai for Students: Generative AI Assistant for Legal Education in Ghana paper

  • Learning Faster with Better Tokens: Parameter-Efficient Vocabulary Adaptation for Specialized Text Summarization paper

  • Exploring Lightweight Large Language Models for Court View Generation paper

  • Retrieval-Based Multi-Label Legal Annotation: Extensible, Data-Efficient and Hallucination-Free paper

  • Chunking German Legal Code paper

  • Generating Legal Commentaries from Case Databases via Retrieval, Clustering, and Generation paper

  • Decompose-and-Refine: Structured Legal Question Answering with Parametric Retrieval paper

  • Retrieval-Augmented Detection of Potentially Abusive Clauses in Chilean Terms of Service paper

  • Traceable by Design: An LLM Pipeline and Dashboard for EU Regulatory Consultation Analysis paper

  • ImmigrationQA: A Source-Grounded Dataset and Small-Model Adaptation for U.S. Immigration Law paper

  • Citation Grounding: Detecting and Reducing LLM Citation Hallucinations via Legal Citation Graphs paper

  • Enhancing BiGRU with a KAN Block for Legal Document Classification and Summarization paper

  • On Wednesdays, We Ask Questions: Optimizing "Active Listening" in Automated Legal Triage and Referral paper

  • Re-Ranking Through an Attribution Lens for Citation Quality in Legal QA paper

  • EURO-5K: When Does Domain Pretraining Matter? Benchmarking Transformers for EU Reporting Obligation Extraction paper

  • DAR: Deontic Reasoning with Agentic Harnesses paper

  • GARL: Game-Theoretic Reinforcement Learning for Multi-Agent Strategic Prioritisation paper

  • Civil Court Simulation with Large Language Models paper

  • LexRubric: A Rubric-Guided Diagnostic Benchmark for Open-Ended Legal Tasks paper

  • From Statute to Control Flow: Span-Grounded Deontic Trees for Defeasible Scope Parsing paper

  • Retrieval Augmented Generation Framework for the Nepali Legal Domain Question Answering paper

  • FineREX: Fine-Tuned NER-RE for Human Smuggling Knowledge Graphs paper

  • Reinforcement learning to improve large language model-based automated code compliance systems paper

  • Measuring & Mitigating Over-Alignment for LLMs in Multilingual Criminal Law Courts paper

  • Peeking Inside LLMs: Leveraging Internal Artifacts of LLMs for Enhancing Reliability in Legal Classification paper

  • A Tree-of-Thoughts Inspired Hybrid Approach for Legal Case Judgement Summarization using LLMs paper

  • PRecG: Legal Precedent Retrieval with Graph Neural Networks and Rhetorical Role Segmentation paper

  • Cross-Architecture LLM Ensembles, Feature-Based Reranking and Retrieval-Augmented Prompting for Legal Information Processing paper

  • When Reasoning Hurts Legal Drafting: The Verbalization Bottleneck in Patent Claim Generation paper

  • The Price of Reasoning: Cost-Quality Tradeoffs in Reinforcement Learning for Neural Machine Translation paper

  • Reasoning Before Translation: Enhancing Legal Machine Translation with Structured Reasoning paper

  • AILQA: Evaluating AI-Driven Legal Question Answering Systems for the Indian Legal System paper

  • Enabling Multilingual Privacy Policy Audits: Large-Scale Analysis of Spanish Mobile Apps paper

Legal Reasoning Models

  • LexPro-1.0 Technical Report paper
  • LexPam: Legal Procedure Awareness-Guided Mathematical Reasoning paper
  • Mining Hidden Thoughts from Texts: Evaluating Continual Pretraining with Synthetic Data for LLM Reasoning paper
  • SyLeR: A Framework for Explicit Syllogistic Legal Reasoning in Large Language Models paper
  • Contractual Deepfakes: Can Large Language Models Generate Contracts? paper
  • LePREC: Reasoning as Classification over Structured Factors for Assessing Relevance of Legal Issues paper
  • Falkor-IRAC: Graph-Constrained Generation for Verified Legal Reasoning in Indian Judicial AI paper
  • Bridging Legal Interpretation and Formal Logic: Faithfulness, Assumption, and the Future of AI Legal Reasoning paper
  • TypedCSIP: Typed Counterfactual Pretraining for Chinese Legislative Conflict Classification paper
  • By Their Fruits You Will Know Them: Comparing Formalizations of Law by the Decisions They Encode paper
  • Scaling Multi-Hop Training Data via Graph-Constrained Path Selection paper
  • DAR: Deontic Reasoning with Agentic Harnesses paper
  • Civil Court Simulation with Large Language Models paper
  • From Statute to Control Flow: Span-Grounded Deontic Trees for Defeasible Scope Parsing paper
  • Generative Chinese Statute Retrieval paper

Legal Agent

  • When Correct Isn't Usable: Improving Structured Output Reliability in Small Language Models paper
  • Accurate Legal Reasoning at Scale: Neuro-Symbolic Offloading and Structural Auditability for Robust Legal Adjudication paper
  • SimuCourt: Building Judicial Decision-Making Agents with Real-world Judgement Documents paper
  • Can Large Language Models Grasp Legal Theories? Enhance Legal Reasoning with Insights from Multi-Agent Collaboration paper
  • Multi-Agent Simulator Drives Language Models for Legal Intensive Interaction paper
  • CLAIM: An Intent-Driven Multi-Agent Framework for Analyzing Manipulation in Courtroom Dialogues paper
  • PAKTON: A Multi-Agent Framework for Question Answering in Long Legal Agreements paper
  • LLM Agents in Law: Taxonomy, Applications, and Challenges paper
  • Dissecting Judicial Reasoning in U.S. Copyright Damage Awards paper
  • Multimodal Multi-Agent Empowered Legal Judgment Prediction paper
  • LegalMALR:Multi-Agent Query Understanding and LLM-Based Reranking for Chinese Statute Retrieval paper
  • LegalOne: A Family of Foundation Models for Reliable Legal Reasoning paper
  • Domain-Partitioned Hybrid RAG for Legal Reasoning: Toward Modular and Explainable Legal AI for India paper
  • LexChronos: An Agentic Framework for Structured Event Timeline Extraction in Indian Jurisprudence paper
  • Fanar-Sadiq: A Multi-Agent Architecture for Grounded Islamic QA paper
  • Dual Hierarchical Dialogue Policy Learning for Legal Inquisitive Conversational Agents paper
  • Can LLMs Time Travel? Enhancing Temporal Consistency in Legal Agentic Search through Reinforcement Learning paper
  • SIA: Self Improving AI with Harness & Weight Updates paper
  • LegalGraphRAG: Multi-Agent Graph Retrieval-Augmented Generation for Reliable Legal Reasoning paper
  • DAR: Deontic Reasoning with Agentic Harnesses paper
  • GARL: Game-Theoretic Reinforcement Learning for Multi-Agent Strategic Prioritisation paper
  • Civil Court Simulation with Large Language Models paper
  • From Statute to Control Flow: Span-Grounded Deontic Trees for Defeasible Scope Parsing paper
  • DeXposure-Claw: An Agentic System for DeFi Risk Supervision paper

Legal Problems of Large Language Models

  • Towards WinoQueer: Developing a Benchmark for Anti-Queer Bias in Large Language Models

  • Persistent Anti-Muslim Bias in Large Language Models

  • Understanding the Capabilities, Limitations, and Societal Impact of Large Language Models

  • The Dark Side of ChatGPT: Legal and Ethical Challenges from Stochastic Parrots and Hallucination

  • The GPTJudge: Justice in a Generative AI World paper

  • Is the U.S. Legal System Ready for AI's Challenges to Human Values? paper

  • Questioning Biases in Case Judgment Summaries: Legal Datasets or Large Language Models? paper

  • Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models

  • A Legal Framework for Natural Language Processing Model Training in Portugal paper

  • LegalBench-RAG: A Benchmark for Retrieval-Augmented Generation in the Legal Domain paper

  • Bias in Large Language Models: Origin, Evaluation, and Mitigation paper

  • An Information Theoretic Approach to Operationalize Right to Data Protection paper

  • Protecting Privacy in Multimodal Large Language Models with MLLMU-Bench paper

  • Standardizing Intelligence: Aligning Generative AI for Regulatory and Operational Compliance paper

  • TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law paper

  • Large Language Models as Search Engines: Societal Challenges paper

  • From Unfamiliar to Familiar: Detecting Pre-training Data via Gradient Deviations in Large Language Models paper

  • TimeMark: A Trustworthy Time Watermarking Framework for Exact Generation-Time Recovery from AIGC paper

  • To See is Not to Learn: Protecting Multimodal Data from Unauthorized Fine-Tuning of Large Vision-Language Model paper

  • Divergence Decoding: Inference-Time Unlearning via Auxiliary Models paper

  • Citation Grounding: Detecting and Reducing LLM Citation Hallucinations via Legal Citation Graphs paper

  • Which Institutional Frameworks Do Chatbots Assume? Auditing Jurisdictional Defaults in Multilingual LLMs paper

  • Measuring & Mitigating Over-Alignment for LLMs in Multilingual Criminal Law Courts paper

  • Who Checks the Citations? Benchmarking Legal Hallucination Detection paper

  • Peeking Inside LLMs: Leveraging Internal Artifacts of LLMs for Enhancing Reliability in Legal Classification paper

  • Position: The Term "Machine Unlearning" Is Overused in LLMs paper

  • Open but Incompatible: A License Compatibility Analysis of Corpora for Low-Resource African Languages paper

  • Probing Stylistic Appropriation using Large Language Models: An Evaluation Framework for Copyright Infringement under EU Law paper

  • The Perplexity Trap: When Patent Law Makes Human Writing Look Like AI paper

Data Resources for Large Language Models in Law

  • CAIL2018: A Large-Scale Legal Dataset for Judgment Prediction
  • When does pretraining help? assessing self-supervised learning for law and the casehold dataset of 53,000+ legal holdings
  • LeCaRD: a legal case retrieval dataset for Chinese law system
  • LeXFiles and LegalLAMA: Facilitating English Multinational Legal Language Model Development
  • Legal Extractive Summarization of U.S. Court Opinions
  • Awesome Chinese Legal Resources github
  • MultiLegalPile: A 689GB Multilingual Legal Corpus paper
  • The Cambridge Law Corpus: A Corpus for Legal AI Research
  • TransformLLM: Adapting Large Language Models via LLM-Transformed Reading Comprehension Text paper
  • Natural Language Processing for the Legal Domain: A Survey of Tasks, Datasets, Models, and Challenges paper
  • ChineseSafe: A Chinese Benchmark for Evaluating Safety in Large Language Models paper
  • Augmenting Legal Decision Support Systems with LLM-based NLI for Analyzing Social Media Evidence paper
  • CaseSumm: A Large-Scale Dataset for Long-Context Summarization from U.S. Supreme Court Opinions paper
  • From Answers to Guidance: A Proactive Dialogue System for Legal Documents paper
  • A Greek Government Decisions Dataset for Public-Sector Analysis and Insight [paper](https://arxiv.org/abs/2512.05647
  • Data Augmented Pipeline for Legal Information Extraction and Reasoning paper
  • Domain-Adaptation through Synthetic Data: Fine-Tuning Large Language Models for German Law paper
  • Knowledge Augmented Entity and Relation Extraction for Legal Documents with Hypergraph Neural Network paper
  • Quecto-V1: Empirical Analysis of 8-bit Quantized Small Language Models for On-Device Legal Retrieval paper
  • SinhaLegal: A Benchmark Corpus for Information Extraction and Analysis in Sinhala Legislative Texts paper
  • LAMUS: A Large-Scale Corpus for Legal Argument Mining from U.S. Caselaw using LLMs paper
  • RoLegalGEC: Legal Domain Grammatical Error Detection and Correction Dataset for Romanian paper
  • Automatic Construction of a Legal Citation Graph from 100 Million Ukrainian Court Decisions: Large-Scale Extraction, Topological Analysis, and Ontology-Driven Clustering paper
  • Temporal Decay of Co-Citation Predictability: A 20-Year Statute Retrieval Benchmark from 396M Ukrainian Court Citations paper
  • IMLJD: A Computational Dataset for Indian Matrimonial Litigation Analysis paper
  • Fine-grained Claim-level RAG Benchmark for Law paper
  • A Reproducible Universal Dependencies-Style Pipeline for Katharevousa Greek Parliamentary Text paper
  • Temporal Concept Drift in Legal Judgment Prediction: Neural Baselines Across Three Epochs of Ukrainian Court Decisions paper
  • Bundesrecht: An Open Library and Corpus for German Statutory Reference Processing paper
  • CanLegalRAGBench: Evaluating Retrieval-Augmented Generation on Canadian Case Law paper
  • Scaling Multi-Hop Training Data via Graph-Constrained Path Selection paper
  • ImmigrationQA: A Source-Grounded Dataset and Small-Model Adaptation for U.S. Immigration Law paper
  • Citation Grounding: Detecting and Reducing LLM Citation Hallucinations via Legal Citation Graphs paper
  • Enhancing BiGRU with a KAN Block for Legal Document Classification and Summarization paper
  • Digging Up Citations: FOSSIL, a Dataset and Workflow for Reference Extraction in Law and the Humanities paper
  • EURO-5K: When Does Domain Pretraining Matter? Benchmarking Transformers for EU Reporting Obligation Extraction paper
  • HKJudge: A Legal Discourse-Annotated Corpus for Interpreting What Courts Find, How They Reason, and What They Rule paper
  • LexRubric: A Rubric-Guided Diagnostic Benchmark for Open-Ended Legal Tasks paper
  • From Statute to Control Flow: Span-Grounded Deontic Trees for Defeasible Scope Parsing paper
  • Retrieval Augmented Generation Framework for the Nepali Legal Domain Question Answering paper
  • REDACT: A Systematically Controlled Multilingual Benchmark for Personal Information Detection paper
  • FineREX: Fine-Tuned NER-RE for Human Smuggling Knowledge Graphs paper
  • Who Checks the Citations? Benchmarking Legal Hallucination Detection paper
  • Open but Incompatible: A License Compatibility Analysis of Corpora for Low-Resource African Languages paper
  • Grounded Event Extraction from SEC 8-K Filings with a Fine-Grained Taxonomy paper
  • Enabling Multilingual Privacy Policy Audits: Large-Scale Analysis of Spanish Mobile Apps paper

Law LLMs

  • LawGPT_zh github
  • LaWGPT github
  • Lawyer LLaMA github
  • LexiLaw github
  • LexGPT 0.1: pre-trained GPT-J models with Pile of Law paper
  • TOWARDS THE EXPLOITATION OF LLM-BASED CHATBOT FOR PROVIDING LEGAL SUPPORT TO PALESTINIAN COOPERATIVES paper
  • ChatLaw: Open-Source Legal Large Language Model with Integrated External Knowledge Bases paper
  • DISC-LawLLM github
  • InternLM-Law: An Open Source Chinese Legal Large Language Model paper
  • SaulLM-54B & SaulLM-141B: Scaling Up Domain Adaptation for the Legal Domain paper
  • LAWGPT: KNOWLEDGE-GUIDED DATA GENERATION AND ITS APPLICATION TO LEGAL LLM paper
  • HKGAI-V1: Towards Regional Sovereign Large Language Model for Hong Kong paper
  • HUKUKBERT: Domain-Specific Language Model for Turkish Law paper
  • PoliLegalLM: A Technical Report on a Large Language Model for Political and Legal Affairs paper
  • ImmigrationQA: A Source-Grounded Dataset and Small-Model Adaptation for U.S. Immigration Law paper
  • Legal Domain Adaptation of Modern BERT Models paper

Evaluation

  • Measuring Massive Multitask Chinese Understanding paper

  • LawBench: Benchmarking Legal Knowledge of Large Language Models github

  • Large Language Models are legal but they are not: Making the case for a powerful LegalLLM paper

  • Better Call GPT, Comparing Large Language Models Against Lawyers paper

  • Evaluating GPT-3.5's Awareness and Summarization Abilities for European Constitutional Texts with Shared Topics paper

  • Evaluation Ethics of LLMs in Legal Domain paper

  • GPTs and Language Barrier: A Cross-Lingual Legal QA Examination paper

  • LawBench: Benchmarking Legal Knowledge of Large Language Models paper

  • LexEval: A Comprehensive Chinese Legal Benchmark for Evaluating Large Language Models paper

  • LegalAgentBench: Evaluating LLM Agents in Legal Domain paper

  • Measuring Faithfulness and Abstention: An Automated Pipeline for Evaluating LLM-Generated 3-ply Case-Based Legal Arguments paper

  • LegalEval-Q: A New Benchmark for The Quality Evaluation of LLM-Generated Legal Text paper

  • When Fairness Isn't Statistical: The Limits of Machine Learning in Evaluating Legal Reasoning paper

  • LexTime: A Benchmark for Temporal Ordering of Legal Events paper

  • LLM-based HSE Compliance Assessment: Benchmark, Performance, and Advancements paper

  • VLQA: The First Comprehensive, Large, and High-Quality Vietnamese Dataset for Legal Question Answering paper

  • The Massive Legal Embedding Benchmark (MLEB) paper

  • LexGenius: An Expert-Level Benchmark for Large Language Models in Legal General Intelligence paper

  • Legal-DC: Benchmarking Retrieval-Augmented Generation for Legal Documents paper

  • Adaptive Cost-Efficient Evaluation for Reliable Patent Claim Validation paper

  • From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text paper

  • Exploiting LLM-as-a-Judge Disposition on Free Text Legal QA via Prompt Optimization paper

  • Tokenizer Fertility and Zero-Shot Performance of Foundation Models on Ukrainian Legal Text: A Comparative Study paper

  • Reasoners or Translators? Contamination-aware Evaluation and Neuro-Symbolic Robustness in Tax Law paper

  • Validate Your Authority: Benchmarking LLMs on Multi-Label Precedent Treatment Classification paper

  • LP-Eval: Rubric and Dataset for Measuring the Quality of Legal Proposition Generation paper

  • GradeLegal: Automated Grading for German Legal Cases paper

  • Asking For An Old Friend: Diagnosing and Mitigating Temporal Failure Modes in LLM-based Statutory Question Answering paper

  • A Two-Phase Stability Study of LLM Judges and Bar Council Examiners on Thai Bar-Exam Free-Form Essays paper

  • JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment paper

  • The Cases LJP Never Sees: Prosecution Decision Prediction for More Complete Criminal Liability Assessment paper

  • BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law paper

  • Multi-Legal-Bench: Evaluating LLMs on Legal Reasoning Across Jurisdictions, Languages, and Legal Traditions paper

  • UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning paper

  • LegalBench-BR: A Benchmark for Evaluating Large Language Models on Brazilian Legal Decision Classification paper

  • IndiaFinBench: An Evaluation Benchmark for Large Language Model Performance on Indian Financial Regulatory Text paper

  • RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity Corpora paper

  • ImmigrationQA: A Source-Grounded Dataset and Small-Model Adaptation for U.S. Immigration Law paper

  • CanLegalRAGBench: Evaluating Retrieval-Augmented Generation on Canadian Case Law paper

  • Citation Grounding: Detecting and Reducing LLM Citation Hallucinations via Legal Citation Graphs paper

  • Which Institutional Frameworks Do Chatbots Assume? Auditing Jurisdictional Defaults in Multilingual LLMs paper

  • On Wednesdays, We Ask Questions: Optimizing "Active Listening" in Automated Legal Triage and Referral paper

  • Re-Ranking Through an Attribution Lens for Citation Quality in Legal QA paper

  • EURO-5K: When Does Domain Pretraining Matter? Benchmarking Transformers for EU Reporting Obligation Extraction paper

  • HKJudge: A Legal Discourse-Annotated Corpus for Interpreting What Courts Find, How They Reason, and What They Rule paper

  • LexRubric: A Rubric-Guided Diagnostic Benchmark for Open-Ended Legal Tasks paper

  • From Statute to Control Flow: Span-Grounded Deontic Trees for Defeasible Scope Parsing paper

  • Retrieval Augmented Generation Framework for the Nepali Legal Domain Question Answering paper

  • REDACT: A Systematically Controlled Multilingual Benchmark for Personal Information Detection paper

  • DeXposure-Claw: An Agentic System for DeFi Risk Supervision paper

  • Measuring & Mitigating Over-Alignment for LLMs in Multilingual Criminal Law Courts paper

  • Who Checks the Citations? Benchmarking Legal Hallucination Detection paper

  • Peeking Inside LLMs: Leveraging Internal Artifacts of LLMs for Enhancing Reliability in Legal Classification paper

  • Probing Stylistic Appropriation using Large Language Models: An Evaluation Framework for Copyright Infringement under EU Law paper

  • Do LLMs Fabricate Legal Citations? A Bilingual Benchmark on Saudi Data Protection Law and the GDPR paper

  • The Perplexity Trap: When Patent Law Makes Human Writing Look Like AI paper

The survey paper is shown in paper

Acknowledgement

Please cite the following papers as references if this repository helps you (*ゝω・)ノThanks!

@article{sun2023short,
  title={A short survey of viewing large language models in legal aspect},
  author={Sun, Zhongxiang},
  journal={arXiv preprint arXiv:2303.09136},
  year={2023}
}
  • AILQA: Evaluating AI-Driven Legal Question Answering Systems for the Indian Legal System paper
  • Enabling Multilingual Privacy Policy Audits: Large-Scale Analysis of Spanish Mobile Apps paper