Multi-Person Human Motion Forecasting in Complex Scenes

Serdar Ozsoy 1,2
Lars Doorenbos 1,2
Juergen Gall 1,2
1University of Bonn, 2Lamarr Institute for Machine Learning and Artificial Intelligence
GCPR 2026

Abstract

Accurately forecasting the movement of people in complex scenes requires reasoning over the past and present state of the entire environment. In this context, effectively incorporating object information and social interactions into a unified framework remains particularly challenging. To address this, we propose Object-Conditioned Social Diffusion (OCSD), a conditional diffusion model that integrates motion history, multi-person interactions, and object cues into a single framework. OCSD uses an object-conditioning mechanism that modulates denoising at every timestep, enabling fine-grained human-object reasoning, and a social encoder that models the interactions between all humans in the scene. As a result, our model naturally handles varying group sizes, complex social interactions, and supports sampling multiple plausible futures. Extensive experiments show that OCSD achieves state-of-the-art results on the Humans in Kitchens (HiK) and HOI-M3 benchmarks. It reduces the two-second path error by 121.5 mm (31.3%) on HiK and 130.5 mm (33.2%) on HOI-M3 compared to prior work, and produces more realistic long-term forecasts.

Video Samples

HOI-M3: 2 seconds

HIK: 2 seconds

HIK: 10 seconds

Limitations

Citation

@misc{ozsoy2026ocsd,
author = "{Serdar Ozsoy, Lars Doorenbos, Juergen Gall}",
title = "Multi-Person Human Motion Forecasting in Complex Scenes",
year = "2026",
}