Legal AI & NLP
Models, pipelines, and applications for legal language.
Legal-domain BERT trained on the Pile of Law
AI-powered legal research assistant providing legal advisory, case outcome prediction, and report generation from PDF documents using FastAPI, Streamlit, and Groq LLM.
Open-source legal AI agent skills for PKULaw, citations, and DOCX workflows
Corporate law RAG with novel retrieval
Generate tracked-changes Word docs and redline PDFs from a contract review, the same deliverables lawyers actually send.
The Harvard USPTO Patent Dataset (HUPD) is a large-scale, well-structured, and multi-purpose corpus of English-language patent applications filed to the United States Patent and Trademark Office (USPTO) between 2004 and 2018. With more than 4.5 million patent documents, HUPD is two to three times larger than comparable corpora. Unlike other NLP patent datasets, HUPD contains the inventor-submitted versions of patent applications, not the final versions of granted patents, allowing us to study patentability at the time of filing using NLP methods for the first time.
Brazilian open legal method for agents
An MCP server that provides comprehensive US legislation.
Open-source prototype for deterministic, evidence-grounded screening under the EU AI Act, GDPR, and EU Data Act, with bilingual workflows, versioned rules, and traceable legal evidence.
Swiss sovereign legal AI. Fork of Mike, extended with MCP client and local model support (Apertus). Named after Emilie Kempin-Spyri.
Zero dependency sentence boundary detection tuned for legal text, from the LexNLP authors, so chunking for legal retrieval does not split on citations and abbreviations.
Open framework for Saudi legal AI
AI-powered legal signature page extractor and packet generator. Automates the creation of wet-ink signature packs for M&A and financing transactions.
Rust framework for parsing legal text
Dataset Card for "IL-TUR" Dataset Description Summary "IL-TUR": Benchmark for Indian Legal Text Understanding and Reasoning is a collaborative effort to establish a modern benchmark for training and evaluating AI/NLP models on Indian Law. IL-TUR consists of 8 foundational tasks, requiring different types of understanding and skills. Apart from English, some tasks involve Indic languages. This dataset repository has been created to unify the data… See the full description on the dataset page: https://huggingface.co/datasets/Exploration-Lab/IL-TUR.
Build one of these? Verify ownership through GitHub and take over your project's page: tagline, categories, maintainer's note.