Controllable and Photorealistic Pedestrian Risky Motion
Generation for End-to-End Driving Safety Evaluation

ControlPed turns real-world vehicle–pedestrian conflicts into controllable 3D human motions and renders them as photorealistic multi-camera video, so end-to-end driving models can be tested on the situations that matter most.

Anonymous Authors
Controllable, photorealistic risky pedestrian motions rendered for driving evaluation.

Seven leading end-to-end driving models were evaluated on 88 rendered safety-critical scenarios. Their mean HDScore fell from 88.8 on the reconstructed original scenes to 47.4 under dangerous pedestrian behaviors.

Original scene 88.8
Safety-critical 47.4

Abstract

Evaluating end-to-end autonomous driving under rare, safety-critical vehicle–pedestrian interactions requires photorealistic, sensor-level scenarios. However, trajectory-based scenario generators cannot synthesize raw visual observations, whereas video-based approaches lack controllability.

To bridge this gap, we present ControlPed, a novel framework that combines trajectory-level conflict synthesis with 3D Gaussian Splatting (3DGS) to generate photorealistic, motion-controllable safety-critical scenarios. Built upon HazardPed, a dataset derived from 10,352 traffic videos comprising 422 conflict trajectories, HD maps, and 857 annotated 3D human motions, ControlPed first generates conflict trajectories, lifts them into 3D human motion sequences via text-conditioned motion diffusion, and finally renders multi-view sensor observations using animatable 3DGS avatars.

Safety evaluation in 88 rendered photorealistic scenarios reveals that seven leading end-to-end driving models suffer a severe performance drop, with their mean HDScore plunging from 88.8 to 47.4, exposing major failure modes under dangerous pedestrian behaviors. The dataset and testing benchmarks will be released to facilitate safety assessment of vehicle–pedestrian interactions.

Methodology overview

ControlPed is a three-stage framework. It couples conflict data collection and a conflict trajectory generator with a text-driven motion diffusion model (a–c). Animatable avatars are then composited into 3DGS scenes to render photorealistic multi-camera videos (d), and the synthesized scenarios are used to evaluate end-to-end driving (e).

1

Conflict trajectory generation

Generate a risky future path from observed history, surrounding vehicles and pedestrians, HD map and a goal heading.

2

Text-driven motion synthesis

Lift the 2D path into full-body 3D motion with a diffusion model, guided by an action description.

3

Photorealistic rendering

Composite animatable gaussian avatars into reconstructed scenes and rasterize multi-view video.

ControlPed overview

HazardPed dataset

HazardPed is a dataset of real-world vehicle–pedestrian conflicts. It provides conflict trajectories with corresponding HD maps, together with human motions paired with text descriptions.

10,352 source traffic videos 422 conflict trajectories with HD maps 857 annotated 3D human motions
Conflict videosRaw videos are not provided
Conflict trajectoriesExtracted paths aligned to HD maps
Text–motion pairs3D human motions with text descriptions

Conflict trajectory generation

Given a pedestrian's observed history, the trajectory generator predicts its future path conditioned on the surrounding vehicles and pedestrians, local map semantics, and a goal heading that specifies the intended direction of motion. The model is first pre-trained on normal pedestrian behavior from nuScenes and then fine-tuned on the conflict trajectories in HazardPed.

Trajectory generation architecture

Qualitative comparison

OursAdvBMTCCDiffCounterScene

Text-driven motion generation

Given a generated 2D conflict trajectory, the motion model lifts it into a full 3D human motion sequence. Generation is guided by an action category and its text description, such as "lie down on the ground" or "run across the road". The trajectory is written directly into the root-position channels of the motion and kept fixed throughout the diffusion process, so motion synthesis becomes an inpainting problem: the denoiser fills in the remaining body motion while the generated pedestrian stays strictly on the target path.

Motion generation architecture

Qualitative comparison

OursMDMCondMDIPriorMDMOmniControlStableMoFusion

Ablation study

Oursw/o adaLNw/o reyaww/o geow/o cat

End-to-end driving safety evaluation

We reconstruct 88 nuScenes driving scenes with OmniRe and represent pedestrians as drivable Gaussian human avatars. The Gaussians of the background, vehicles, and inserted pedestrians are composited and rasterized into multi-view camera videos, which are fed to seven end-to-end driving models for safety evaluation.

Quantitative results

Model RC ↑ NC ↑ DAC ↑ TTC ↑ COM ↑ HDScore ↑
DiffDrive 99.062.8 99.086.2 99.999.8 94.884.9 99.099.4 94.651.9
SSR 100.062.7 98.684.6 99.699.6 94.484.0 99.599.4 95.151.1
VAD 99.663.7 97.684.3 99.499.1 92.183.4 99.499.9 92.750.8
LAW 100.062.8 97.784.2 98.798.4 92.682.5 99.9100.0 92.849.7
LTF 98.860.7 96.784.4 99.098.9 93.983.7 90.389.9 90.047.7
UniAD 80.747.2 90.179.2 90.890.7 79.972.6 87.283.7 62.430.2
SparseDrive 100.062.2 98.585.2 99.399.2 93.883.6 99.599.9 94.350.6
Mean drop −36.6−12.9−0.2−9.5−0.4−41.4
Reconstructed original scene Safety-critical scenario Bold marks the best original-scene score per metric

Rendered safety-critical scenarios

Pedestrian suddenly appearing
Pedestrian falling down
Pedestrian running forward
Pedestrian crossing the road