MFV Legal Tranche — Rights-Cleared Legal Reasoning Dataset

The MFV Legal Tranche is a rights-cleared, provenance-verified legal reasoning dataset: 759 retrieval cards, 118 canonical packages, and 291 documented corrections derived from two years of California civil litigation that settled favorably against a top-20 defense firm. Licensed for AI training, fine-tuning, and evaluation. A free sample tranche is available on AWS Data Exchange; the full corpus is licensed at $299 for 12 months.

Dataset statistics

MetricValue
Retrieval cards759
Canonical packages118
Documented corrections291
Optimized prompts3
Validated quality lift20%
Validation methodTwo independent AI judges

Provenance

Single author, 100% owned, zero scraped content, no third-party rights. Every document originates from two years of pro se California civil litigation conducted by the author against a top-20 defense firm. The origin of the material is documented in a public court record: Superior Court of California, County of Los Angeles, Case No. 25STCV14949. Post-litigation-era buyers price licensing risk into every training-data purchase — a single-author corpus with a verifiable public-record origin removes that risk entirely: there is no scraped content to trace, no third-party rights to clear, and no ambiguity about who owns the data.

The corpus itself is sanitized for public sale: party names and personally identifying details are redacted throughout, so the dataset ships clean for training use while the public court record independently verifies its origin.

What it's for

Fine-tuning legal AI models. The tranche captures how legal reasoning actually unfolds in live litigation — claims constructed, opposing arguments answered, errors caught and corrected — rather than the static case-law text most legal corpora contain. Fine-tuning on reasoning traces with documented corrections teaches a model the process of legal argument, not just its vocabulary.

RAG evaluation. The 759 retrieval cards are self-contained, human-authored units of legal knowledge with known relationships to the 118 canonical packages they support. That structure makes them natural ground truth for evaluating whether a retrieval pipeline surfaces the right authority for a given legal question — a labeled retrieval corpus that did not come from synthetic generation.

Legal-reasoning benchmarks. The 291 documented corrections are before-and-after pairs: a flawed legal work product and its repaired version. Correction pairs are benchmark material — they let you test whether a model can spot the same defects a human litigator had to find and fix under adversarial pressure.

Frequently asked questions

Is this dataset rights-cleared for commercial AI training?
Yes. The entire corpus was written by a single author who owns 100% of it. It contains no scraped content and no third-party-licensed material, so there are no upstream rights to clear. The license explicitly covers commercial AI training, fine-tuning, and evaluation.
How was the 20% quality lift measured?
Model outputs produced with and without the tranche's material were scored by two independent AI judges, and the tranche-assisted outputs measured 20% higher in quality. The evaluation methodology and judge configuration are shared with licensees.
What formats are the files in?
The corpus is delivered as a ZIP archive through AWS Data Exchange, containing structured plain-text records: 759 retrieval cards, 118 canonical packages, 291 correction pairs, and 3 optimized prompts. A file manifest is included with the free sample.
Where does the data come from?
Two years of pro se California civil litigation conducted by the author, which settled favorably against a top-20 defense firm. It is human-generated working legal product from a real case — not scraped case law and not synthetic output from another model.
Are party names or personal details included in the data?
No. The corpus is sanitized: party names and personally identifying details are redacted throughout, so it can be used for training without PII handling obligations. Provenance is verified independently through the public court record (LA Superior Court Case No. 25STCV14949) rather than through identifying content in the files.
Is there a free sample before buying?
Yes. A free sample tranche is offered on AWS Data Exchange, and a gated sample is offered on Hugging Face. Both include representative retrieval cards and a correction pair so you can evaluate structure and quality before licensing the full corpus.
How is the full corpus licensed and priced?
The full corpus is licensed at $299 for 12 months through AWS Data Exchange. AWS handles billing, delivery, and entitlement, so procurement works like any other AWS marketplace purchase.

License & access