Publications

(2026). COBALT: Censored Optimization and Bayesian Active Learning Techniques. UAI 2026.

PDF

(2026). Video Generation Models are General-Purpose Vision Learners. ECCV 2026.

PDF Code

(2026). A Mixed Diet Makes DINO An Omnivorous Vision Encoder. arXiv.

PDF Code

(2026). How to Spin an Object: First, Get the Shape Right. CVPRW 2026.

PDF Code

(2025). OpenWorldSAM: Extending SAM2 for Universal Image Segmentation with Language Prompts. arXiv.

PDF Code

(2025). Generalist forecasting with frozen video models via latent diffusion. arXiv.

PDF

(2025). From Image to Video: An Empirical Study of Diffusion Representations. ICCV 2025.

PDF

(2024). Scaling 4D Representations. arXiv.

PDF Code

(2024). Neural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion Models. NeurIPS 2024.

PDF Project

(2024). Moving Off-the-Grid: Scene-Grounded Video Representations. NeurIPS 2024.

PDF

(2024). PaliGemma: A versatile 3B VLM for transfer. arXiv.

PDF Code

(2024). Leveraging VLM-Based Pipelines to Annotate 3D Objects. In ICML.

Cite Project

(2021). SIMONe: View-Invariant, Temporally-Abstracted Object Representations via Unsupervised Video Decomposition. In NeurIPS.

Cite Project Slides Video

(2020). AlignNet: Unsupervised Entity Alignment. arXiv.

PDF

(2019). Multi-Object Representation Learning with Iterative Variational Inference. ICML 2019.

PDF

(2019). An Investigation of Model-Free Planning. ICML 2019.

PDF

(2019). MONet: Unsupervised Scene Decomposition and Representation. arXiv.

PDF