article-journal

Generalist forecasting with frozen video models via latent diffusion
A generalist forecasting framework that predicts future frozen visual features and decodes them for multiple downstream tasks.
Scaling 4D Representations
A study showing that large-scale self-supervised video models improve consistently on spatial and temporal 4D vision tasks as model size grows.
PaliGemma: A versatile 3B VLM for transfer
An open vision-language model designed as a versatile base model for transfer across a wide variety of open-world tasks.
AlignNet: Unsupervised Entity Alignment
An unsupervised module for aligning object representations across time so entities can be tracked between segmented frames.