This repository contains the implementation of the paper:
STRIDE: Evaluating Spatiotemporal Reasoning in Driving Edge Cases
Lieqi Liu1*, Rui Gao1*, Jia-Chen Gu1, Wenbo Hu1, Zhaobin Mo2, Ahmadreza Moradipari2, Nejib Ammar2, Wei Wang1, [Nanyun Peng]1
1University of California, Los Angeles 2Toyota InfoTech Labs
*Equal contribution
Figure 1: Overview of STRIDE (data · tasks · comparison). Vector version: STRIDE_v4.pdf.
- 2026.09: Initial release of STRIDE annotations (nuScenes + Waymo), evaluation toolkit, and baseline leaderboard.
- 2026.09: Mini demonstration subsets released under
STRIDE/Mini.
STRIDE ships question annotations (and nuScenes group sidecars) in this repository and on Hugging Face (uclanlp/STRIDE). Raw sensor data must be downloaded from upstream datasets; this repo provides the scripts to turn that data into the media layout STRIDE expects.
- nuScenes trainval — follow the official download / ToU.
- Waymo Open Dataset validation images — follow the official terms.
nuScenes — stitch 6-camera grids, install sidecars, render selected-vehicle query overlays:
pip install nuscenes-devkit matplotlib # needed for selected-vehicle overlays
python scripts/prepare_nuscenes_media.py \
--nuscenes-root /path/to/nuscenes/trainval \
--output-dir /path/to/formatted_scenes
export STRIDE_NUSCENES_FORMATTED_SCENES=/path/to/formatted_scenes
python scripts/verify_stride_media.py --split nuscenes --require-overlaysThis installs STRIDE/nuScenes/formatted_metadata/ into the output directory, stitches the 2×3 multi-camera grids (*_grid.jpg), and writes {group}_selected_vehicle_render.jpg query overlays from the shipped annotation tokens.
Waymo — no stitching and no SAM: verify relative paths only.
python scripts/prepare_waymo_media.py \
--waymo-image-root /path/to/waymo/val/images
export STRIDE_WAYMO_IMAGE_ROOT=/path/to/waymo/val/images
python scripts/verify_stride_media.py --split waymoTarget-object boxes were computed offline (SAM during construction) and are stored as bbox_xyxy in STRIDE/Waymo/questions.json. TRJ ranking options ship both ego-frame candidate_trajectories and preprojected candidate_trajectories_image. At eval time stride.visual.load_waymo_images draws the red box / labeled curves — users do not run SAM or need Waymo calibration files.
| Split | Size | Image Source | Annotation | Notes |
|---|---|---|---|---|
| nuScenes | 1150 | nuScenes trainval | STRIDE/nuScenes |
23 templates × 50; GT-revised open OEQs |
| Waymo | 1200 | Waymo Open Dataset val | STRIDE/Waymo |
24 templates × 50; bbox_xyxy precomputed |
| Mini (nuScenes) | 3 | same as nuScenes | STRIDE/Mini |
Schema demo only |
| Mini (Waymo) | 4 | same as Waymo | STRIDE/Mini |
Schema demo only |
Note that:
- Questions + nuScenes group metadata ship in this repo and on Hugging Face (including Waymo
bbox_xyxy/ trajectory overlays). - Raw images are not vendored (upstream licenses / size). Grids + nuScenes query overlays are built locally with
prepare_nuscenes_media.py. - Always run
verify_stride_media.pybefore inference so missing frames fail fast. - STRIDE Mini is for schema demonstration only, not for reporting scores.
- Scoring (
python -m stride.cli score) needs only prediction JSON + shipped questions; inference needs the prepared media roots above.
After setup, the layout looks like:
├── STRIDE
│ ├── nuScenes
│ │ ├── questions.json
│ │ └── formatted_metadata/ -- shipped group sidecars
│ ├── Waymo/questions.json -- relative CAM_FRONT paths + overlays
│ ├── Mini/
│ └── statistics/
├── $STRIDE_NUSCENES_FORMATTED_SCENES -- built by prepare_nuscenes_media.py
│ ├── scene_012/
│ │ ├── 010_*_grid.jpg
│ │ ├── group_002_vehicle_annotations.json
│ │ └── group_002_selected_vehicle_render.jpg
│ └── scene_name_to_formatted_dir.json
└── $STRIDE_WAYMO_IMAGE_ROOT -- Waymo val/images
└── <segment_id>/CAM_FRONT/*.jpg
STRIDE is organized around the complementary dimensions of space and time, plus driving context for ego motion and the broader environment (paper §3.1). It comprises 39 question templates in six task families, pairing MCQs that test intermediate judgments with OEQs that assess integration into driving explanations and decisions.
- Spatial perception (
SP-*). Basic object-level properties that are directly observable from the scene, including relative position, distance, direction of motion, and lane occupancy. - Spatial understanding (
SU-*). Interprets the driving significance of these spatial relationships—whether an object restricts the drivable region, affects lane-change feasibility, constrains the ego vehicle’s motion, or poses a potential risk.
- Temporal memory (
TM-*). Reconstructs how the current spatial state developed from previous observations: previously visible or disappeared objects, earlier locations, occlusion and reappearance, and recent motion trends. - Temporal extrapolation (
TE-*). Anticipates how the current scene may evolve, including future object motion, lane occupancy, collision timing, and likely subsequent events.
- Trajectory prediction (
TRJ-*). Assesses the ego vehicle’s past and future motion (ranking, language justifications, and waypoint regression). - Scene-context awareness (
SC-*). Assesses environmental conditions, road structure, and surrounding activity that provide context for interpreting individual objects and their interactions.
| Family | Split coverage | Formats |
|---|---|---|
| Spatial perception | nuScenes + Waymo | MCQ + OEQ |
| Spatial understanding | nuScenes | MCQ + OEQ |
| Temporal memory | nuScenes + Waymo | MCQ |
| Temporal extrapolation | nuScenes + Waymo | MCQ + OEQ |
| Trajectory prediction | nuScenes + Waymo | MCQ + OEQ + Geo (TRJ-5/TRJ-6) |
| Scene-context awareness | nuScenes + Waymo | OEQ (and Waymo MCQs for SC-*) |
Inputs. Each question uses a five-frame temporal context. nuScenes uses surrounding multi-camera grids; Waymo uses CAM_FRONT. When a target object exists, it is highlighted by a red box on the query frame.
The annotation files contain question-answering pairs as follows (fields may vary slightly by split):
{
"meta": {
"benchmark": "STRIDE",
"split": "nuscenes", -- or waymo
"n_questions": 1150,
"...": "..."
},
"tasks": [
{
"id": "SP-1", -- template id
"task": "space-perception", -- task family
"question_format": "MCQ", -- MCQ or OEQ
"scoring_type": "mcq", -- mcq | open_oeq_text | trajectory_waypoints
"question": <str>,
"choices": { -- MCQ only
"A": <str>,
"B": <str>,
"..."
},
"ground_truth": <str>, -- letter, free text, or waypoint JSON
"scene_id": <str>,
"group_id": <str>,
"object_id": <str>, -- when a target object exists
"source_group_file": <str>, -- nuScenes: relative group annotation path
"image_paths": [<str>, ...], -- Waymo: 5 paths relative to WAYMO_IMAGE_ROOT
"bbox_xyxy": [<float>, ...], -- Waymo: precomputed target box (no SAM at eval)
"candidate_trajectories": {...}, -- Waymo TRJ ranking: ego-frame polylines (m)
"candidate_trajectories_image": {...} -- Waymo TRJ ranking: preprojected pixels
},
...
]
}
- Install dependencies:
pip install -e .
# optional BLEURT scoring
pip install -e ".[bleurt]"- Load a split and inspect tasks:
from stride.aggregate import load_tasks
payload, tasks = load_tasks("nuscenes") # or "waymo", "nuscenes_mini", "waymo_mini"
print(payload["meta"])
print(tasks[0]["id"], tasks[0]["question"])To help users run models against STRIDE, we provide helpers under evaluation/:
- Make sure the directory organization follows Data Preparation.
- Build a blank prediction file:
python evaluation/make_prediction_template.py \
--split nuscenes \
--output runs/my_model_responses.json- Fill each
model_response:
| Format | Expected model_response |
|---|---|
| MCQ | Single letter A / B / … (also accepts ANSWER: A, {"answer":"A"}) |
| OEQ | Free-form text |
| TRJ-5 / TRJ-6 | JSON list of [x, y] waypoints (5 or 4 points) |
- Optionally run the OpenAI-compatible multimodal runner:
export OPENAI_API_KEY=...
export STRIDE_NUSCENES_FORMATTED_SCENES=/path/to/formatted_scenes
python evaluation/run_openai.py \
--split nuscenes \
--model gpt-4.1 \
--limit 10 \
--output runs/gpt41_responses.jsonPrimary ranking key: MCQ accuracy. Open-ended and trajectory metrics are reported separately (no single fused score). Random-guess MCQ baselines are ≈19.8% (nuScenes) and ≈20.3% (Waymo).
As in the overview figure, STRIDE is substantially harder than prior driving VQA suites: even GPT-6-Astra drops from ~85% on previous benchmarks to 45.5% MCQ on STRIDE (nuScenes).
| Rank | Model | Type | MCQ ↑ | BLEURT ↑ | Vision judge ↑ | ADE ↓ | mini-FDE ↓ |
|---|---|---|---|---|---|---|---|
| 1 | gpt-6-astra | proprietary | 45.5% | −0.303 | 6.41 | 0.60 | 0.98 |
| 2 | GPT-5.5 | proprietary | 39.3% | −0.289 | 6.68 | 0.65 | 0.97 |
| 3 | gemini-3.6-flash | proprietary | 38.9% | −0.238 | 6.50 | 0.69 | 1.01 |
| 4 | gpt-5.6-sol | proprietary | 38.5% | −0.323 | 5.65 | 0.69 | 0.99 |
| 5 | gemini-3.1-pro-preview | proprietary | 37.8% | −0.301 | 4.53 | 1.42 | 1.84 |
| 6 | Qwen3.6 35B-A3B | VLM | 31.6% | −0.359 | 5.32 | 7.42 | 5.45 |
| 7 | claude-sonnet-5 | proprietary | 29.4% | −0.249 | 6.42 | 0.87 | 1.30 |
| 8 | Qwen2.5-VL 3B | VLM | 27.4% | — | 3.62 | — | — |
| 9 | Qwen3-VL 30B-A3B-Thinking | VLM | 26.8% | −0.321 | 5.45 | 7.12 | 3.42 |
| 10 | Qwen3-VL 8B-Thinking | VLM | 23.4% | −0.287 | 5.50 | 8.00 | 3.80 |
| 11 | Dolphins | expert | 18.5% | −0.435 | 3.24 | 8.16 | 3.81 |
| 12 | Qwen3.6 27B | VLM | 17.5% | −0.402 | 5.99 | 5.81 | 4.06 |
| — | UniAD 2.0 | expert (TRJ-5 only) | — | — | — | 0.56 | 0.79 |
| Rank | Model | Type | MCQ ↑ | BLEURT ↑ | Vision judge ↑ | ADE ↓ | mini-FDE ↓ |
|---|---|---|---|---|---|---|---|
| 1 | gpt-6-astra | proprietary | 41.2% | −0.517 | 6.77 | 13.64 | 12.03 |
| 2 | gemini-3.1-pro-preview | proprietary | 39.7% | −0.558 | 4.03 | 14.22 | 14.21 |
| 3 | claude-sonnet-5 | proprietary | 39.2% | −0.504 | 7.01 | 13.16 | 10.27 |
| 4 | gemini-3.6-flash | proprietary | 38.8% | −0.288 | 6.20 | 13.53 | 11.37 |
| 5 | GPT-5.5 | proprietary | 38.7% | −0.494 | 7.00 | 13.84 | 12.57 |
| 6 | gpt-5.6-sol | proprietary | 37.7% | −0.594 | 6.48 | 13.81 | 12.39 |
| 7 | Qwen2.5-VL 3B | VLM | 37.4% | — | 2.57 | — | — |
| 8 | Qwen3.6 35B-A3B | VLM | 34.3% | −0.539 | 5.83 | 18.55 | 8.86 |
| 9 | Qwen3-VL 8B-Thinking | VLM | 32.0% | −0.663 | 4.72 | 22.36 | 19.37 |
| 10 | Qwen3-VL 30B-A3B-Thinking | VLM | 31.8% | — | 4.98 | 11.83 | 6.18 |
| 11 | Qwen3-VL 8B-Instruct | VLM | 30.1% | −0.685 | 4.96 | 19.64 | 12.65 |
| 12 | Qwen3.6 27B | VLM | 28.4% | −0.496 | 5.87 | 20.37 | 19.25 |
| 13 | Dolphins | expert | 22.2% | −0.629 | 1.86 | 18.43 | 8.60 |
Machine-readable tables: evaluation/leaderboard.
Check STRIDE Evaluation for more details.
@misc{liu2026stride,
title = {{STRIDE}: Evaluating Spatiotemporal Reasoning in Driving Edge Cases},
author = {Liu, Lieqi and Gao, Rui and Gu, Jia-Chen and Hu, Wenbo and Mo, Zhaobin and Moradipari, Ahmadreza and Ammar, Nejib and Wang, Wei and Peng, Nanyun},
year = {2026},
note = {Preprint},
howpublished = {\url{https://github.com/PlusLabNLP/STRIDE}}
}