PEVA: Predicting Egocentric Video from Whole-Body Actions
Researchers from BAIR (Berkeley AI) introduce PEVA, a model that predicts egocentric video frames from human whole-body actions. It uses an autoregressive conditional diffusion transformer trained on the Nymeria dataset, enabling atomic action synthesis, counterfactual simulation, and long video generation. PEVA can also be used for visual planning by optimizing action sequences.
Berkeley Artificial Intelligence Research (BAIR)
