What the model finds, and what it misses
Scores are computed with the organisers' evaluate.py on our own labels of the four sample videos
(dev_labels.json), using the frozen submission output
predictions_samples.json. The organisers' hidden test set will score differently.
Annotated sample videos
Rendered by our own tooling from the tracks: boxes with track ids, scene geometry (stop line red, solid lines yellow, crossings white), the signal state read from the lamp, active events and the Part B risk bar. The timeline shows predictions against our labels.
Scores per class
Mean F1 over tIoU 0.3 / 0.5 / 0.7
Score A is the mean of these bars across every class in labels ∪ predictions. A labelled class we never predict (congestion) counts as 0.
Examples of each detected class
For each class, the prediction with the highest temporal IoU against our labels. Click to play it.
Failure cases
For each class, the longest false positive and the longest missed label at tIoU 0.5, with what went wrong.
Operator view
What a traffic centre would see across the four clips: how often each violation happens and for how long.