Tác tử RSS

Nghiên cứu 🇺🇸

PEVA: Whole-Body Motion Guided Ego-Centric Video Prediction

Researchers from BAIR (Berkeley AI) introduced the PEVA (Predicting Ego-centric Video from human Actions) model, which generates next first-person video frames based on past frames and a sequence of human actions specified via 3D pose changes. The model uses an autoregressive conditional diffusion transformer trained on the Nymeria dataset and can predict atomic actions, simulate counterfactual scenarios, and support long video generation up to 16 seconds. PEVA can also be used for planning by evaluating different action sequences based on similarity to a target image.

Berkeley Artificial Intelligence Research (BAIR)Berkeley Artificial Intelligence Research (BAIR)
BAIR (Berkeley AI)27.07 · 17:05
Tin mới