Retrieval & RAG
Embeddings, retrieval pipelines, and RAG over legal corpora.
A sophisticated RAG intelligence engine for Singaporean laws, policies, and history. Comes with a triple-AI failover backend (Gemini/Llama/Groq), semantic embeddings using FAISS, and an Apple-inspired interactive UI. Designed with precision and high availability in mind.
AI-powered legal research assistant providing legal advisory, case outcome prediction, and report generation from PDF documents using FastAPI, Streamlit, and Groq LLM.
Corporate law RAG with novel retrieval
Open-source Legal AI workspace for evidence-grounded legal drafting, matter analysis and verifiable answers.
The code used to evaluate embedding models on the Massive Legal Embedding Benchmark (MLEB).
Local-first research and drafting tool
Self-hosted document chat for attorneys
Open-source prototype for deterministic, evidence-grounded screening under the EU AI Act, GDPR, and EU Data Act, with bilingual workflows, versioned rules, and traceable legal evidence.
Zero dependency sentence boundary detection tuned for legal text, from the LexNLP authors, so chunking for legal retrieval does not split on citations and abbreviations.
UK Government i.AI MCP server that embeds Hansard debates and written parliamentary questions into Qdrant so agents can search UK legislative records semantically.
The French Cour de cassation's own search API behind JUDILIBRE: Elasticsearch full text and faceted retrieval over pseudonymised supreme court and appellate decisions.
Rust search engine for French and EU law: BM25 plus vector retrieval on Postgres with ParadeDB and VectorChord, citation graph resolution, a REST API, and a hosted MCP endpoint.
German Dataset for Legal Information Retrieval
Sovereign French legal AI platform
Self-hosted contract lifecycle management with seven LangGraph agents for review, drafting, redlining and negotiation, over pgvector and Elasticsearch hybrid retrieval.
Source-backed administrative consequence graph
Open legal intelligence for Mexico. 30,000+ laws across federal, state, and municipal sources · 3.5M+ searchable articles · trilingual UI (Spanish · English · Classical Nahuatl) · first-party RAG chat · OpenAPI + SDKs · built so citizens, accountants, lawyers, and AI agents can navigate Mexican law as code.
An MCP (Model Context Protocol) server providing AI agents access to CourtListener's comprehensive legal database, featuring semantic search, hybrid search, citation verification, and research tools.
LegalSummarization An MTEB dataset Massive Text Embedding Benchmark The dataset consistes of 439 pairs of contracts and their summarizations from https://tldrlegal.com and https://tosdr.org/. Task category t2t Domains Legal, Written Reference https://github.com/lauramanor/legal_summarization How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/legal_summarization.
Build one of these? Verify ownership through GitHub and take over your project's page: tagline, categories, maintainer's note.