Your read on the merits. Our read on the clock.

Push a case's resolution out and the same recovery is a lower annualized return — and the reserve behind it stays locked the whole way. Duration moves IRR, lockup, and reserves directly, yet it's the one input most books still carry without a measured distribution behind it. The Tertius Duration Engine supplies that distribution — trained on 9.1M federal cases, graded in public below against a million more it never saw — so your judgment on the merits runs on a calibrated clock.

Held-out test cases

1,370,792

never trained on

Median error

113 d

Inside the 80% band

78%

target 80% · adjudicable on 408,120 of the 1,370,792 · raw (0.1, 0.9) band: 75% on 630,526

Model version

tertius-acta-6

2026-07-25

We publish our error. These numbers are censoring-honest — still-open cases count, slow cases are never dropped from the test set, and any shortfall from the 80% target is shown, not smoothed. Training window: 1960–2026 filings; calibration held out 2020+ filings from a pre-cut fit. Full breakdown by case type and district on the calibration page.

The Tertius Duration Engine

Language models read the complaint; the Engine counts the record. Extraction — pulling case type, court, amount, and posture off page one — is language work, and we use the best tools for it. The number is not: every duration Tertius reports comes from a proprietary survival model, fit by maximum likelihood on millions of real dockets, with its error record published. Here's what it does that a spreadsheet, a rule of thumb, or a generic model can't.

1

Built for litigation's real shape

Cases settle in waves, then grind on for years. A bell curve can't draw that. Our engine fits the one distribution family built for exactly this shape — the same math used to model machine failure and drug survival, tuned for a federal docket.

2

A different tail for every case type

Prisoner petitions resolve fast and tight. Antitrust drags for years. One global shape would blur both. The Engine fits a separate tail per case type — so your patent case's p90 isn't quietly borrowed from cases that behave nothing like it.

3

Corrected for the era it's in

Dockets have sped up and slowed down for decades — e-filing, COVID backlogs, rule changes. The Engine fits a continuous filing-year trend, so a case filed today is priced off the current docket regime instead of a thirty-year average. It tracks long-run drift; no model pins a single year's shock.

4

Never forgets what's already survived

A case alive for two years cannot resolve on day 200 — that outcome is gone. The Engine conditions on time already survived, so an open case gets the forecast for where it actually stands, not the naive one for a case freshly filed — and that conditional path is scored on held-out data at 1, 2, and 3 years elapsed.

5

Reads the filing, not just the label

Pro se cases resolve ~20% faster; docket-wide, a jury demand runs about a third slower (the size varies sharply by case type — near zero for patent, roughly double for antitrust; per-type effects are on the model roadmap), a class action about a sixth. The Engine conditions on what the complaint already tells you — pro se status, class allegations, jury demand, jurisdiction basis — so two 'contract' cases stop getting the same number.

6

Knows how cases end, not just when

For every case type: the share that settle, get dismissed, reach judgment, or get swept into an MDL — each path with its own clock, estimated only on cohorts old enough to be fully observed — 2.4 million labeled resolutions. Your merits view stays yours; now it has a base rate.

Checked against the government's own numbers

The federal judiciary publishes one timing statistic — median months to disposition, per district (AO Table C-5). We replicated their exact metric from our cleaned data across three separate years and matched their statisticians district-by-district to a fraction of a month, with trial counts agreeing within 2%. Every large gap traced to a measurable cause: mass-MDL waves we exclude by design, so an ordinary case's forecast never inherits 200,000 earplug claims settling at once. The full method, sources, and reproduction script ship in the repo — and everything the government doesn't publish (forecasts, tails, path probabilities, portfolios) is graded on the calibration page instead.

From clock to capital

A duration curve is academic until it's money. The Engine runs a Monte Carlo across every open case's fitted distribution — thousands of simulated paths — and turns it into an IRR fan chart, a capital-lockup curve, and a recycling schedule, with every correlation assumption stated explicitly rather than buried.

Win rate and recovery multiple are your inputs, labeled as yours. We don't sell a merits guess — we sell the clock, priced honestly, so the judgment you already trust is running on a measured number.

Scored on your book

The production model is trained exclusively on public federal court records — no customer data is in it. The outcomes you record do something more useful to you: they grade Tertius against your own realized results, on your own book, out of sample — the accuracy number no vendor benchmark can fake. They also count toward clearly-labeled future recalibration milestones; if a future model version ever trains on customer-contributed outcomes, its calibration page will say so before it ships. Captions, parties, and amounts never cross customer boundaries.

One rule never bends: a resolution date has to appear in the actual docket text, or nothing gets written. Zero invented dates. Scoring only works if the data is clean.

Graded everywhere, not just on average

An average hides a model that nails your biggest category and guesses at the rest. Here it's broken out — every case type, sample size included, nothing smoothed over. Coverage varies by case type and some groups sit below the 80% target; they're published anyway, because a number you can't audit is a number you shouldn't price with. One rule enforced below: band coverage can only be judged on cases whose follow-up outlasts the predicted band, and where fewer than 500 cases clear that bar, we print the shortfall instead of the percentage — a coverage number scored on a sliver of old filings isn't calibration, it's selection. The served band's quantile levels are themselves recalibrated per case type to hit the 80% target — they are not literal 10th/90th percentiles (the exact served levels ship in the artifact's interval_recal field). The raw-band column scores the model's own fixed (0.1, 0.9) interval, whose denominator the band recalibration cannot shrink — the two columns together show what we ship and what the raw distribution earns.

Case typen (holdout)Median abs. errorInside 80% band (served)Raw (0.1, 0.9) band
Prisoner260,814136 d74%on 18,90874%on 71,962
Civil rights167,197112 d79%on 80,68478%on 86,293
Contract163,637147 d80%on 52,17375%on 79,903
Employment160,290143 d79%on 83,56276%on 85,724
Tort personal injury116,947155 d83%on 6,73978%on 33,591
Social security95,556131 d69%on 67,56264%on 66,263
Tort other95,22483 d84%on 9,64877%on 62,316
Immigration84,52763 d80%on 40,80481%on 43,125
Other statutory71,816101 d81%on 4,50271%on 39,705
Ip other54,88381 d81%on 31,95081%on 36,819
Product liability33,0511.3 yrthin gated samplen=260 — too thin to scorefollow-up outlasts the served band for 260 of 33,051 cases35%on only 31 — provisional
Real property23,228101 d76%on 9,97176%on 5,294
Ip patent21,73560 dthin gated samplen=160 — too thin to scorefollow-up outlasts the served band for 160 of 21,735 cases72%on 13,151
Securities8,755141 dthin gated samplen=395 — too thin to scorefollow-up outlasts the served band for 395 of 8,755 cases66%on 4,923
Ip trade secret3,651188 d73%on 80069%on 1,392
Antitrust is missing from this table, not from the docket. Under this artifact's follow-up gate, the served antitrust band reaches so far past available follow-up that no cohort can be scored — the last scoreable snapshot (tertius-acta-3) measured 60.5% on just 38 cases, well below target. Treat antitrust bands as unvalidated: the next artifact scores antitrust against the fixed raw band, whose denominator the gate cannot shrink.

Which model do these numbers grade?

The record above is a temporal holdout: a model fit only on pre-cutoff filings, scored on later filings it never saw. The model serving forecasts is refit on all data through the snapshot — same pipeline, more recent information — so the holdout record certifies the method, not the exact deployed coefficients. To grade the method on the newest evaluable regime, the identical pipeline is refit on filings before 2024 and scored on 2024+ filings none of it saw:

2024+ filings scored

570,530

unseen by that fit

Median error

67 d

Inside the 80% band

77%

served band, adjudicable on 33,248 · raw band 76% on 57,037

C-index

0.656

ranking skill, censoring-aware

Known limitations

A model you'd underwrite with is a model whose edges are documented. These are ours, stated here rather than discovered in diligence.

Very large demands are a weak regime

The public federal data top-codes amounts demanded at $9.999M, so the training data cannot distinguish a $12M case from a $500M one. The Engine pools everything above $1M into a single bucket — a case with more than $10M at stake gets a '$1M+' forecast, not a distinct >$10M one. If your book lives above that line, backtest it there before relying on the numbers.

Calibration varies by case type

The 80% band is a target, not a guarantee: coverage differs across case types, and some groups run below target. The per-group table above and the calibration page publish every group's coverage with its sample size — check your categories, not the average.

Long-pending cases are the hardest

Forecasts for cases already years into their life are scored separately, at 1, 2, and 3 years elapsed — and coverage declines as elapsed time grows. The Engine's conditional forecasts are recalibrated against what surviving cohorts actually did (the raw parametric conditioning overstates remaining time on this data). Median forecasts track those cohorts closely, but band coverage past two years is scored on thin samples and runs below target — the dynamic scores are published so you can see the gap.

Federal district courts only — and docket time, not cash time

Coverage is U.S. federal district court civil cases; state-court matters are flagged on import and not forecast. And the Engine predicts when the district-court docket terminates — settlements can pay on schedules and judgments can be appealed and collected later, so cash receipt can lag the forecast resolution.

Your book is selected — the baseline is not

The Engine's baseline is estimated on the whole federal docket; a funded portfolio is cases chosen for merits, and selection can shift timing. That's a feature to use, not a flaw to hide: the blind backtest scores the model on YOUR book — the exact population your future book is drawn from — and cases you evaluated but passed on are welcome in the backtest set too.

MDL and mass-tort timelines are flagged, not point-forecast

An MDL's life after transfer is driven by bellwethers, settlement matrices, and judicial management — not by anything on page one of a complaint. The Engine forecasts the probability a case gets swept into an MDL and excludes mass-MDL waves from ordinary cohorts by design, but it does not pretend to point-forecast a post-transfer timeline. Don't project mass-tort portfolio liquidity from filing covariates — with us or anyone.

A baseline, not an oracle

The Engine doesn't know your case's merits — by design. What it gives you is the objective baseline rate: how long cases with these procedural characteristics actually run, measured on the full record. Your merits judgment is your edge, and this is what it should be measured against — if your thesis requires a case to resolve eighteen months faster than the historical baseline, that's now a claim you have to justify, price, and underwrite explicitly instead of one that hides inside a point estimate. Baselines standardize the argument; they don't replace the judgment.

Don't trust the scatter plot. Test it on your own book.

Import your last 20 closed cases. The Engine forecasts them blind and shows you, case by case, how close it would have been.

Run your last 20 closed cases through it

Tertius produces statistical timing estimates — not legal advice, never merits or damages. Federal civil only.