200 cases · candidate raft-sintel against raft-things · kitti2015-train

Your new model looks better overall.See exactly where it got worse.

Average case error fell from 5.40 to 1.52 px, a 71.8% improvement. One case out of 200 still got worse, by more than the 0.3 px limit this review was run with. A mean would never have told you which one.

Two consecutive frames from a KITTI driving sequence, alternatingNo overlay
Two consecutive frames, looped. Both checkpoints estimate the motion between them, pixel by pixel.

About the 71.8%: these are two public RAFT checkpoints trained on different data, and the candidate’s training set includes KITTI. Much of that improvement is a domain gap measured in-sample, so it is not a release decision. The example is here for the rest of it: a run that looks like a clean win, one case that got worse anyway, and a receipt that names it. See the run, its receipt and all 200 per-case numbers

Average case error
5.401.52 px-71.8%
Worse on error
1 of 200 cases000145_10 +0.41 px
Unstable on trajectory
0 of 200 casesevery candidate was still shrinking its updates at the last iteration
See the comparison

Every panel is output from the run itself.

Release review

Two model versions, case by case. Error regressions and unsettled answers flagged; the checks stay with you.

raft-kittiWorked example
Local workspace
Currentraft-things
Candidateraft-sintel
kitti2015-train200 paired cases · mean endpoint error (px) · 200 with trajectories

Not ready: 1 error regression.

Average case error
5.401.52px-71.8%
CaseError change
1 of 200 casesNeeds a decision first
RegressionSettled000145_10

000145_10

kitti
Current model5.48px
Candidate5.89px
+0.41px
Changed. Where the candidate disagrees with the current checkpoint. Disagreement needs no ground truth, so this panel works on unlabelled data: it marks what changed, not what is wrong. The full sheet adds the two input frames.
Worse. Error against ground truth, candidate minus current. Red is where the candidate got worse, and it is almost all one object: the pole nearest the camera. This is where the +0.41 px comes from. The full sheet adds each model's own error map.
Both flow fields, and the labels. Both predictions and the ground truth, at one scale. KITTI labels 21% of the pixels here, which is why the error maps are sparse and why the Changed panel above covers ground the error panel cannot.
Every iteration. All twelve refinement iterations, for both models: the candidate on top, the current one below. The last quarter is framed, because that is the window the stability limits are measured over.
How much each model kept moving. Both models settled here: 4% late revision against a 25% limit, no reversals. So this case is flagged on error alone, and the trajectory says the candidate converged confidently to a worse answer rather than failing to converge.

Rendered by rb, the command line tool, during the run, unmodified. Every panel carries its own scale, so you can check what a colour means without leaving the picture.

No per-frame series for this case.

Add paired frame-error arrays to see where the models diverge.

Error increased by 0.41 px, exceeding your 0.3 px limit.

Refinement trajectory update per iteration · px12 iterations
last third0.019.639.3Iteration 1: candidate update 35.051 px, current 30.429 pxIteration 2: candidate update 10.330 px, current 13.263 pxIteration 3: candidate update 6.715 px, current 9.071 pxIteration 4: candidate update 5.084 px, current 5.502 pxIteration 5: candidate update 3.816 px, current 3.386 pxIteration 6: candidate update 2.241 px, current 3.003 pxIteration 7: candidate update 1.592 px, current 2.684 pxIteration 8: candidate update 1.357 px, current 2.069 pxIteration 9: candidate update 1.004 px, current 1.696 pxIteration 10: candidate update 0.788 px, current 1.409 pxIteration 11: candidate update 0.593 px, current 1.172 pxIteration 12: candidate update 0.513 px, current 0.966 px1Iteration12
Current · 7% lateCandidate · 4% late

Late revision: share of all refinement in the last third of iterations (limit 25%). Reversal: an update larger than the one before it by more than 5% (limit 2).

The candidate settled: 4% late revision, 0 reversals. Current model: 7%. It converged to a worse answer: a data or training gap rather than instability.

Keep this case in your next release check.
A real run, loaded as the worked exampleLower is better · equal-weight case averages · stability from label-free trajectories

Make the next comparison yours.

Export once from your evaluator. Import, investigate, and keep checking.

Join the waitlist for the team workspace.

The command line tool is free and complete. The team workspace is not built yet. When it is, it will add:

  • The approved baseline that CI measures the next candidate against
  • The release policy, under version control, with its history
  • Every gate verdict, kept across builds
  • A second person who can open the review at all

Join the waitlist and we will email you when it is ready. We also use the list to decide what to build first.

No card and no account required. Two emails at most.