LegalForecastBench

This benchmark gives models motion to dismiss briefs from real federal cases and asks the model to predict how the judge will resolve each issue in the motion.

Why this benchmark matters

Testing models’ ability to predict motion outcomes is valuable because it is an objectively verifiable, high-level judgment task. In real-world litigation, predicting how a judge is likely to react to a contemplated litigation strategy or a potential motion is usually performed by the most senior and highly paid lawyers on a litigation team. Unlike many legal work-product tasks, moreover, the outcome is an externally observable judicial disposition.

This benchmark is also contamination resistant by design. It evaluates models’ ability to predict judicial decisions made after their documented training-data cutoffs, reducing the risk that the models have simply memorized the decision in pre-training. I’ve built a pipeline that ingests new cases, so that as models’ training cutoffs advance, I can create a new version that uses the most recent decisions at that point in time.

Cycle 1

The Cycle 1 report will be published soon. See the methods for the prediction units, scoring, and eligibility controls.

Current research: LegalForecastGym

I am also building LegalForecastGym, a reinforcement learning environment for federal motion-to-dismiss forecasting. I am post-training open-weight models to assess whether training on expert-labeled litigation prediction tasks improves model performance and prediction calibration in a way that generalizes. I am particularly interested in exploring whether the expert-labeled data provides meaningful training signal that could generalize beyond prediction tasks.