paper-conference

Video Generation Models are General-Purpose Vision Learners
A general-purpose perception model built from a pretrained video diffusion backbone and steered across vision tasks with text instructions.
Moving Off-the-Grid: Scene-Grounded Video Representations
A self-supervised video model whose latent tokens move off the image grid to represent and track scene elements consistently through time.
An Investigation of Model-Free Planning
An empirical study showing that a standard model-free reinforcement-learning agent can learn behaviors commonly associated with planning.
Multi-Object Representation Learning with Iterative Variational Inference
An unsupervised iterative-inference method that jointly segments scenes into objects and learns disentangled object representations.