MFV Legal Tranche — Rights-Cleared Legal Reasoning Dataset

The MFV Legal Tranche is a rights-cleared, provenance-verified legal reasoning dataset: 767 contribution cards, 119 packages, and 291 correction events across 2,389 files, derived from California civil litigation that settled favorably against a top-20 defense firm. Licensed for AI training, fine-tuning, and evaluation. A free sample is available on Hugging Face and AWS Data Exchange; the full corpus is licensed from $299/month.

Dataset statistics

MetricValue
Packages119
Contribution cards767
Correction events291
Files2,389
Legal failure types10
Craft dimensions8
Lifecycle stages7
Severity levels4
Validated quality lift20%
Validation methodTwo independent AI judges

Measured performance

Frontier-model output augmented with the dataset's retrieval library and calibrated system prompts, scored against baseline output on the 8-dimension rubric by two independent AI judges from different model families (Claude Sonnet 4.6 and GPT-4o). Overall quality lift: 20% under both judges.

DimensionLift (Sonnet 4.6 judge)Lift (GPT-4o judge)Score reached
Jurisdictional specificity+36%+31%4.97/5
Citation accuracy+22%+29%4.70/5
Evidence integration+20%+27%4.20/5
Argument persuasiveness4.67/5 under both judges

Floor and ceiling: under Sonnet 4.6 (the stricter judge), 4.18/5 overall with 7 of 8 dimensions at strong professional quality (4.0+); under GPT-4o, 4.72/5 overall with all 8 dimensions above 4.0 and 7 of 8 at near-exceptional quality (4.5+).

Provenance

Single author, 100% owned, zero scraped content, no third-party rights. Every document originates from two years of pro se California civil litigation conducted by the author against a top-20 defense firm. The origin of the material is documented in a public court record: Superior Court of California, County of Los Angeles, Case No. 25STCV14949. Post-litigation-era buyers price licensing risk into every training-data purchase — a single-author corpus with a verifiable public-record origin removes that risk entirely: there is no scraped content to trace, no third-party rights to clear, and no ambiguity about who owns the data.

The corpus itself is sanitized for public sale: party names and personally identifying details are redacted throughout, so the dataset ships clean for training use while the public court record independently verifies its origin.

What it's for

Fine-tuning legal AI models. The tranche captures how legal reasoning actually unfolds in live litigation — claims constructed, opposing arguments answered, errors caught and corrected — rather than the static case-law text most legal corpora contain. Fine-tuning on reasoning traces with documented corrections teaches a model the process of legal argument, not just its vocabulary.

RAG evaluation. The 767 contribution cards are self-contained, human-authored units of legal knowledge with known relationships to the 119 packages they support. That structure makes them natural ground truth for evaluating whether a retrieval pipeline surfaces the right authority for a given legal question — a labeled retrieval corpus that did not come from synthetic generation.

Legal-reasoning benchmarks. The 291 correction events are before-and-after pairs: a flawed legal work product and its repaired version. Correction pairs are benchmark material — they let you test whether a model can spot the same defects a human litigator had to find and fix under adversarial pressure.

Frequently asked questions

Is this dataset rights-cleared for commercial AI training?
Yes. The entire corpus was written by a single author who owns 100% of it. It contains no scraped content and no third-party-licensed material, so there are no upstream rights to clear. The license explicitly covers commercial AI training, fine-tuning, and evaluation.
How was the 20% quality lift measured?
Model outputs produced with and without the tranche's material were scored by two independent AI judges, and the tranche-assisted outputs measured 20% higher in quality. The evaluation methodology and judge configuration are shared with licensees.
What formats are the files in?
The corpus is delivered as a ZIP archive through AWS Data Exchange containing 2,389 files across 119 packages: 767 contribution cards, 291 correction events, retrieval-ready JSON, evaluation harnesses and rubrics, and buyer due-diligence documentation. A field dictionary is included with the free sample.
Where does the data come from?
Two years of pro se California civil litigation conducted by the author, which settled favorably against a top-20 defense firm. It is human-generated working legal product from a real case — not scraped case law and not synthetic output from another model.
Are party names or personal details included in the data?
No. The corpus is sanitized: party names and personally identifying details are redacted throughout, so it can be used for training without PII handling obligations. Provenance is verified independently through the public court record (LA Superior Court Case No. 25STCV14949) rather than through identifying content in the files.
Is there a free sample before buying?
Yes. A free sample tranche is offered on AWS Data Exchange, and a gated sample is offered on Hugging Face. Both include representative retrieval cards and a correction pair so you can evaluate structure and quality before licensing the full corpus.
How is the full corpus licensed and priced?
Licensed from $299/month through AWS Data Exchange. AWS handles billing, delivery, and entitlement, so procurement works like any other AWS marketplace purchase; specific terms and durations are set out in the AWS Data Exchange offer.

License & access

Also from Meta-Flywheel Ventures: the Professional Screenwriting AI Evaluation & Correction Dataset — a second rights-cleared judgment dataset from the same documented pipeline.