<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Conferences | Academic</title><link>http://rkabra.com/category/conferences/</link><atom:link href="http://rkabra.com/category/conferences/index.xml" rel="self" type="application/rss+xml"/><description>Conferences</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Mon, 18 Dec 2023 00:00:00 +0000</lastBuildDate><image><url>http://rkabra.com/media/icon_hu0b7a4cb9992c9ac0e91bd28ffd38dd00_9727_512x512_fill_lanczos_center_3.png</url><title>Conferences</title><link>http://rkabra.com/category/conferences/</link></image><item><title>NeurIPS 2023 Recap</title><link>http://rkabra.com/post/neurips-2023-recap/</link><pubDate>Mon, 18 Dec 2023 00:00:00 +0000</pubDate><guid>http://rkabra.com/post/neurips-2023-recap/</guid><description>&lt;p>This is a slice of topics I’ve been interested in of late, including VLMs, scene understanding, object-centric representations, and generative models. The posters below reflect about 2% of NeurIPS. While I have attempted to capture what is trending, there are entire fields that aren’t reflected here. On the whole the conference was less hype-y than last year. There was a lot of focus on evaluation, new ways of using existing models, and generating new data.&lt;/p>
&lt;h3 id="text-image">Text-image alignment, attribution, and synthetic data&lt;/h3>
&lt;p>&lt;a href="https://dreamsim-nights.github.io/" target="_blank" rel="noopener">DreamSim&lt;/a>: introduces a human-judgment dataset called NIGHTS to capture “mid-level perceptual similarities.” The dataset is collected using a two-alternative forced choice given a reference image. An ensemble of CLIP, OpenCLIP, and DINO is LoRA-tuned on NIGHTS, showing an increase in the DreamSim perceptual metric.
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/fu_s_hu867588328072ff9938b5d5efe5a42e29_4228619_9f4411db9f09479e9cd6686c2132f6a7.webp 400w,
/post/neurips-2023-recap/posters/fu_s_hu867588328072ff9938b5d5efe5a42e29_4228619_e603a407961d8590d814cd502269551a.webp 760w,
/post/neurips-2023-recap/posters/fu_s_hu867588328072ff9938b5d5efe5a42e29_4228619_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/fu_s_hu867588328072ff9938b5d5efe5a42e29_4228619_9f4411db9f09479e9cd6686c2132f6a7.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Cycle consistency for diffusion models–they derive a set of losses which can be used on unpaired data. Method: they pass in x_0 to enable “reconstruction,” which allows driving away from a given class.
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/xu_s_huaa997b12c17ef6982eeaa0fe1a6ff7f1_3785830_d847051c98d663be94b59e7bcf1dc1c0.webp 400w,
/post/neurips-2023-recap/posters/xu_s_huaa997b12c17ef6982eeaa0fe1a6ff7f1_3785830_06e9da76913136e2d496762c2fac7216.webp 760w,
/post/neurips-2023-recap/posters/xu_s_huaa997b12c17ef6982eeaa0fe1a6ff7f1_3785830_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/xu_s_huaa997b12c17ef6982eeaa0fe1a6ff7f1_3785830_d847051c98d663be94b59e7bcf1dc1c0.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Attack to alter images imperceptibly so VLMs (MiniGPT-4, BLIP-2) are totally confused when captioning them. Uses a transfer-based attack to maximize similarity of adversarial image, followed by a query-based attack to further maximize similarity between generated caption and target string.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/zhao_y_hu151568f5c5ed138b56c5b3a10cdb6018_4169042_50d61d8ae080c44e47f205e4f6998c58.webp 400w,
/post/neurips-2023-recap/posters/zhao_y_hu151568f5c5ed138b56c5b3a10cdb6018_4169042_8d3dcb332b560cec23f609c0258a42bf.webp 760w,
/post/neurips-2023-recap/posters/zhao_y_hu151568f5c5ed138b56c5b3a10cdb6018_4169042_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/zhao_y_hu151568f5c5ed138b56c5b3a10cdb6018_4169042_50d61d8ae080c44e47f205e4f6998c58.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Text-image alignment and improved compositional prompting. The DA-Score evaluates alignment using VQA feedback. The proposed Eval-and-Refine method improves alignment and is model-agnostic; it works with Stable Diffusion XL.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/singh_hua0a49ff8394c50d68b78bf16ca4054e8_2476201_f06f0695c080f6ea889d52cae65843ca.webp 400w,
/post/neurips-2023-recap/posters/singh_hua0a49ff8394c50d68b78bf16ca4054e8_2476201_83471e202c4dba2e6cb7c91603262839.webp 760w,
/post/neurips-2023-recap/posters/singh_hua0a49ff8394c50d68b78bf16ca4054e8_2476201_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/singh_hua0a49ff8394c50d68b78bf16ca4054e8_2476201_f06f0695c080f6ea889d52cae65843ca.webp"
width="760"
height="397"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>&lt;a href="https://virajprabhu.github.io/lance-web/" target="_blank" rel="noopener">LANCE&lt;/a>: Synthetic image generation by structured prompt variation.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/prabhu_hua35a34e6fc89fea14e57846ae9cfc704_2066543_74739a45fb38c11dbe38d6112bfe0ad9.webp 400w,
/post/neurips-2023-recap/posters/prabhu_hua35a34e6fc89fea14e57846ae9cfc704_2066543_f87014f8aeffb47c35c0a73de78c2635.webp 760w,
/post/neurips-2023-recap/posters/prabhu_hua35a34e6fc89fea14e57846ae9cfc704_2066543_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/prabhu_hua35a34e6fc89fea14e57846ae9cfc704_2066543_74739a45fb38c11dbe38d6112bfe0ad9.webp"
width="760"
height="388"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>“Guided imagination” to expand small datasets. They perturb latent features, optimizing the perturbation to maximize “class informativeness” and a KL-based sample diversity score.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/zhang_y_huf9b5a164c004cc0cd402053dfe0db1cb_3135297_5b2108e1db555adc9d77ac15f6010cc3.webp 400w,
/post/neurips-2023-recap/posters/zhang_y_huf9b5a164c004cc0cd402053dfe0db1cb_3135297_ccecfa516f604fdd469f2a7aba5bc57e.webp 760w,
/post/neurips-2023-recap/posters/zhang_y_huf9b5a164c004cc0cd402053dfe0db1cb_3135297_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/zhang_y_huf9b5a164c004cc0cd402053dfe0db1cb_3135297_5b2108e1db555adc9d77ac15f6010cc3.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Ask GPT to generate confusing descriptions for better negative examples. They use fine-grained (LLM-generated?) descriptions as true positives. Those aren&amp;rsquo;t verified but are supposedly innocuous. Perhaps there&amp;rsquo;s a risk true positives are wrong and the generated negatives are true? Not convinced how well this can work, but the authors won a CVPR challenge.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/li_l_h_hu3a345345c55c1a3b0df073fe9d6fae10_1968080_1b263bcf8240c19c3cd3dddf6e55cb88.webp 400w,
/post/neurips-2023-recap/posters/li_l_h_hu3a345345c55c1a3b0df073fe9d6fae10_1968080_f5a1f29408a6442e73f237b73d14d59f.webp 760w,
/post/neurips-2023-recap/posters/li_l_h_hu3a345345c55c1a3b0df073fe9d6fae10_1968080_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/li_l_h_hu3a345345c55c1a3b0df073fe9d6fae10_1968080_1b263bcf8240c19c3cd3dddf6e55cb88.webp"
width="760"
height="543"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Text-image alignment benchmark. Uses an NLI entailment model at Google (Q-squared).&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/yarom_hu29513a7bab22fe7f72139941ef2fe3f7_3791513_ca5c6660b6e1e3a5180d7521ef26ec94.webp 400w,
/post/neurips-2023-recap/posters/yarom_hu29513a7bab22fe7f72139941ef2fe3f7_3791513_ead58846dd640bf217a540c67a9f783d.webp 760w,
/post/neurips-2023-recap/posters/yarom_hu29513a7bab22fe7f72139941ef2fe3f7_3791513_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/yarom_hu29513a7bab22fe7f72139941ef2fe3f7_3791513_ca5c6660b6e1e3a5180d7521ef26ec94.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Diffusion self-guidance showing impressive ability to make compositional corrections in generated images.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/epstein_hu8cb067339d7c27b91b9679c3aac4d8dc_3857457_73ac20d38a4bc5eed6272d6ba8d9d9f7.webp 400w,
/post/neurips-2023-recap/posters/epstein_hu8cb067339d7c27b91b9679c3aac4d8dc_3857457_faab700ae8b20904c97b22328664fedf.webp 760w,
/post/neurips-2023-recap/posters/epstein_hu8cb067339d7c27b91b9679c3aac4d8dc_3857457_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/epstein_hu8cb067339d7c27b91b9679c3aac4d8dc_3857457_73ac20d38a4bc5eed6272d6ba8d9d9f7.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Compositional image generation allowing style transfer, negative prompting, etc. New benchmark MCC-250. Image fidelity and text-image alignment measured using FID and CLIP.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/bellagente_hu36d8113e56b2225fe83901e0ef75bc00_3880769_c1a7170c79478863255092dc9551f0ef.webp 400w,
/post/neurips-2023-recap/posters/bellagente_hu36d8113e56b2225fe83901e0ef75bc00_3880769_85e341752b5a32c9a67baa87f4b98f9c.webp 760w,
/post/neurips-2023-recap/posters/bellagente_hu36d8113e56b2225fe83901e0ef75bc00_3880769_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/bellagente_hu36d8113e56b2225fe83901e0ef75bc00_3880769_c1a7170c79478863255092dc9551f0ef.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>First instruction-following VLM called Llava. Helps improve factuality.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/liu_h_hu460e9f8be350761ac625877a3d74d7a9_4202505_b29ee33df32cc961f929e9f784601d04.webp 400w,
/post/neurips-2023-recap/posters/liu_h_hu460e9f8be350761ac625877a3d74d7a9_4202505_05463ad1480d300f5af5e2db391dad30.webp 760w,
/post/neurips-2023-recap/posters/liu_h_hu460e9f8be350761ac625877a3d74d7a9_4202505_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/liu_h_hu460e9f8be350761ac625877a3d74d7a9_4202505_b29ee33df32cc961f929e9f784601d04.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>An object-centric benchmark to assess generated images for text-image alignment. They study multiple compositional tasks including attribute binding but also counting, position, colors. They show high agreement with human judgment, better than a SoTA CLIPScore. And they show diffusion models are still bad, particularly with position. DeepFloyd IF-XL was the best model (even relative to Stable Diffusion XL).&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/ghosh_hu86bec41e0939df1923e015e6980f04e3_2010022_3c3e04e9450d01a014714e2f75e6c59c.webp 400w,
/post/neurips-2023-recap/posters/ghosh_hu86bec41e0939df1923e015e6980f04e3_2010022_758906b7b71b7e9761be5582834a7d58.webp 760w,
/post/neurips-2023-recap/posters/ghosh_hu86bec41e0939df1923e015e6980f04e3_2010022_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/ghosh_hu86bec41e0939df1923e015e6980f04e3_2010022_3c3e04e9450d01a014714e2f75e6c59c.webp"
width="760"
height="547"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>&lt;a href="https://universal-simulator.github.io/unisim/" target="_blank" rel="noopener">UniSim&lt;/a>, a video diffusion model which can “simulate realistic experience” of humans/agents interacting with their environment. Can simulate both high-level (e.g., open the drawer) and low-level instructions (e.g., move to x,y). Can be used to train both vision-language planners and low-level RL policies.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/yang_s_hu46b2026c833571f301faf0b4442efba0_385439_27b3e547945d39d86bfc9c710158a910.webp 400w,
/post/neurips-2023-recap/posters/yang_s_hu46b2026c833571f301faf0b4442efba0_385439_372d448f7fdf13cc9852c5afc4f6dd71.webp 760w,
/post/neurips-2023-recap/posters/yang_s_hu46b2026c833571f301faf0b4442efba0_385439_1200x1200_fit_q75_h2_lanczos_3.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/yang_s_hu46b2026c833571f301faf0b4442efba0_385439_27b3e547945d39d86bfc9c710158a910.webp"
width="760"
height="377"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Open-vocabulary part segmentation. They clean and relabel the Pascal-Part and ADE20K-Part datasets.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/wei_m_hu383d5483e2097aee531031d2fcdf36ad_3491048_e29cfe51471f9cf7e6c96b421feb0de4.webp 400w,
/post/neurips-2023-recap/posters/wei_m_hu383d5483e2097aee531031d2fcdf36ad_3491048_0038b078d352dd509a400bbecc769cc8.webp 760w,
/post/neurips-2023-recap/posters/wei_m_hu383d5483e2097aee531031d2fcdf36ad_3491048_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/wei_m_hu383d5483e2097aee531031d2fcdf36ad_3491048_e29cfe51471f9cf7e6c96b421feb0de4.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Quantifying image difficulty based on human viewing time to identify category. They show images to participants for 17ms, 50ms, etc and collect 7 responses. When the majority of participants get it right, that is the MVT. Since this is correlated with VLM performance, do we need to run this exercise again in the future? The authors show CLIP ViT models have the best performance on the hardest images. The authors also mentioned ImageNetX, which comes with labels for lighting and other conditions.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/mayo_hua0a49ff8394c50d68b78bf16ca4054e8_2370341_15acff4cdaf6c4c098ead2087848b60a.webp 400w,
/post/neurips-2023-recap/posters/mayo_hua0a49ff8394c50d68b78bf16ca4054e8_2370341_275a7f6ec9529dd836df667f1cdd0f4f.webp 760w,
/post/neurips-2023-recap/posters/mayo_hua0a49ff8394c50d68b78bf16ca4054e8_2370341_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/mayo_hua0a49ff8394c50d68b78bf16ca4054e8_2370341_15acff4cdaf6c4c098ead2087848b60a.webp"
width="760"
height="401"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Two-object attribute binding: What proportion of feature combinations/”rules” do models need to be trained on to be able to generate the full feature space. They test CNN and ViT based encoders, either using Slot Attention or just a VAE. The hardest version is on ClevrTex.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/kim_y_hu8aa55ac9e3704011f3db91ffed4fd3dc_3867568_e4c2545fe9c4e6f59a884bf9e3fa8ca1.webp 400w,
/post/neurips-2023-recap/posters/kim_y_hu8aa55ac9e3704011f3db91ffed4fd3dc_3867568_c7ee5ef54d5854f882f4cc785777fd0e.webp 760w,
/post/neurips-2023-recap/posters/kim_y_hu8aa55ac9e3704011f3db91ffed4fd3dc_3867568_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/kim_y_hu8aa55ac9e3704011f3db91ffed4fd3dc_3867568_e4c2545fe9c4e6f59a884bf9e3fa8ca1.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>LLMs are great at scoring object-level text-to-image generation!&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/lu_y_hu5745edd01569002f09e96e125268da7e_3664199_fb70b05b6b06d81a117d3482aec1d604.webp 400w,
/post/neurips-2023-recap/posters/lu_y_hu5745edd01569002f09e96e125268da7e_3664199_b859430ed40aaad62126c9b3e79982f7.webp 760w,
/post/neurips-2023-recap/posters/lu_y_hu5745edd01569002f09e96e125268da7e_3664199_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/lu_y_hu5745edd01569002f09e96e125268da7e_3664199_fb70b05b6b06d81a117d3482aec1d604.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Augmentation to increase dataset diversity while maintaining visual consistency. Utilizes captioning models/LLMs to “extract task-agnostic concepts from training data.” Then augments the training data via language-guided image editing. “To maintain data integrity, a model trained on the original dataset filters out minimal image edits and those which corrupt class-relevant information.”&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/dunlap_hua0a49ff8394c50d68b78bf16ca4054e8_1347492_4cb124ae5e031f35a336beb6ada075f9.webp 400w,
/post/neurips-2023-recap/posters/dunlap_hua0a49ff8394c50d68b78bf16ca4054e8_1347492_b157f8c3b1ba650adaeb0385caa93166.webp 760w,
/post/neurips-2023-recap/posters/dunlap_hua0a49ff8394c50d68b78bf16ca4054e8_1347492_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/dunlap_hua0a49ff8394c50d68b78bf16ca4054e8_1347492_4cb124ae5e031f35a336beb6ada075f9.webp"
width="760"
height="695"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>A captioning benchmark/dataset. The authors collect human data using a rating game to achieve community consensus. The aim is to ensure objects, global context, and actions are correctly described in the caption. This produces a caption dataset called VICR. The authors further train a baseline ViLBERT model on VICR and an existing Flickr8k-Expert dataset. VICR training outperforms on all metrics including BLEU, METEOR, ROUGE, and CLIPScore.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/narins_hud2561b53b0c32c24592cb9162dd8c8ee_1587337_464cb397adfde162a02f9aafccfaf2c9.webp 400w,
/post/neurips-2023-recap/posters/narins_hud2561b53b0c32c24592cb9162dd8c8ee_1587337_15dabb1864d3d5061cb6c449f4ca1099.webp 760w,
/post/neurips-2023-recap/posters/narins_hud2561b53b0c32c24592cb9162dd8c8ee_1587337_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/narins_hud2561b53b0c32c24592cb9162dd8c8ee_1587337_464cb397adfde162a02f9aafccfaf2c9.webp"
width="760"
height="599"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Cola: A benchmark to test object attribute binding with hard distractors. Derived from GQA/Visual Genome. The authors advocate multimodal adaptation (i.e., fine-tuning the cross-attention layers) over tuning other parts of the network.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/ray_hu2c94550eb5365fd1b879639f25f11951_4067906_b18b9300a8cb436a476292f98c483a26.webp 400w,
/post/neurips-2023-recap/posters/ray_hu2c94550eb5365fd1b879639f25f11951_4067906_ae335d12dae676d7cd6d40c29e97704e.webp 760w,
/post/neurips-2023-recap/posters/ray_hu2c94550eb5365fd1b879639f25f11951_4067906_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/ray_hu2c94550eb5365fd1b879639f25f11951_4067906_b18b9300a8cb436a476292f98c483a26.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Measuring visual perception alignment between humans and models. They take into account cases when humans will abstain from prediction. In uncertain cases, they expect models to imitate human judgements.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/lee_j_hu8187872added461913daf3aaf3727614_3963846_e82c02e69bd0b102e6a625ee8eba83b3.webp 400w,
/post/neurips-2023-recap/posters/lee_j_hu8187872added461913daf3aaf3727614_3963846_1bd049ea0730abe297b25e7e9cff39c7.webp 760w,
/post/neurips-2023-recap/posters/lee_j_hu8187872added461913daf3aaf3727614_3963846_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/lee_j_hu8187872added461913daf3aaf3727614_3963846_e82c02e69bd0b102e6a625ee8eba83b3.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Multimodal Information Bottleneck for attribution. Outperforms other methods like GradCAM, Saliency, KernelSHAP, RISE, Chefer et al to attribute CLIP on CC and MS-CXR images.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/wang_y_hu940658a6f928a5eb4a43279a55f4b76e_2170339_7d09a25ee927776c6a02b3d99503e8b2.webp 400w,
/post/neurips-2023-recap/posters/wang_y_hu940658a6f928a5eb4a43279a55f4b76e_2170339_6c5322a1024b09cfa7d3361e46637771.webp 760w,
/post/neurips-2023-recap/posters/wang_y_hu940658a6f928a5eb4a43279a55f4b76e_2170339_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/wang_y_hu940658a6f928a5eb4a43279a55f4b76e_2170339_7d09a25ee927776c6a02b3d99503e8b2.webp"
width="760"
height="578"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;h3 id="objects-slots-segments">Objects/slots/segments&lt;/h3>
&lt;p>SOLV: Object discovery on real life videos (YouTube) without any additional modalities. They compare with the SoTA, DINOSAUR.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/aydemir_hua0a49ff8394c50d68b78bf16ca4054e8_1724194_dd786f81bfa5aa1178298662aad4261a.webp 400w,
/post/neurips-2023-recap/posters/aydemir_hua0a49ff8394c50d68b78bf16ca4054e8_1724194_b6043f2ecb98c7841f5567cbf60f4bbd.webp 760w,
/post/neurips-2023-recap/posters/aydemir_hua0a49ff8394c50d68b78bf16ca4054e8_1724194_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/aydemir_hua0a49ff8394c50d68b78bf16ca4054e8_1724194_dd786f81bfa5aa1178298662aad4261a.webp"
width="760"
height="557"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>VideoSAUR: combines Recurrent Slot Attention with DINOSAUR and adds a temporal similarity loss. Outperforms STEVE on MoVI-E but perhaps slightly worse than SOLV (above)?&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/zadaianchuk_hu31cddbefb4032adb8c069c53e8c18441_2080338_89ca0fad9d4c4bf5ca66e5c36b2a5cb5.webp 400w,
/post/neurips-2023-recap/posters/zadaianchuk_hu31cddbefb4032adb8c069c53e8c18441_2080338_60858f97107b0183886195ba22ab4a8d.webp 760w,
/post/neurips-2023-recap/posters/zadaianchuk_hu31cddbefb4032adb8c069c53e8c18441_2080338_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/zadaianchuk_hu31cddbefb4032adb8c069c53e8c18441_2080338_89ca0fad9d4c4bf5ca66e5c36b2a5cb5.webp"
width="760"
height="486"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>In-context compositional generation using cross-attention between slots and analogy-based instructions.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/dedhia_huc0c4ba88f8ed47fc22a6d800617eef90_3583747_eb6a080d24b580654a375379ddac6e5a.webp 400w,
/post/neurips-2023-recap/posters/dedhia_huc0c4ba88f8ed47fc22a6d800617eef90_3583747_4d9bc09026f47c7a7b9be5725a8522b7.webp 760w,
/post/neurips-2023-recap/posters/dedhia_huc0c4ba88f8ed47fc22a6d800617eef90_3583747_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/dedhia_huc0c4ba88f8ed47fc22a6d800617eef90_3583747_eb6a080d24b580654a375379ddac6e5a.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Cross-attention over slot attention slots can work if diffusing over a pretrained CNN latent space. Plus clustering all slots can produce an interesting concept library (except number of clusters needs to be set).&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/jiang_hu949eb471688cb488b5c77d4325bfaa36_4340098_f2969fabf184be9ebbff2db4d2f39c8b.webp 400w,
/post/neurips-2023-recap/posters/jiang_hu949eb471688cb488b5c77d4325bfaa36_4340098_04960b81a750aeddaa6164aba9f328e4.webp 760w,
/post/neurips-2023-recap/posters/jiang_hu949eb471688cb488b5c77d4325bfaa36_4340098_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/jiang_hu949eb471688cb488b5c77d4325bfaa36_4340098_f2969fabf184be9ebbff2db4d2f39c8b.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>SAMCLR––use SAM segments to sample views for contrastive learning. This helps ensure crops contain sufficient (semantic) overlap.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/missaoui_hu947b06f7f5a561d7b397ccb62f8e864a_2874822_634249af7e3dabb9c89b2014e20fdaef.webp 400w,
/post/neurips-2023-recap/posters/missaoui_hu947b06f7f5a561d7b397ccb62f8e864a_2874822_b4c0491f40669a982f32ad1b499fd3e6.webp 760w,
/post/neurips-2023-recap/posters/missaoui_hu947b06f7f5a561d7b397ccb62f8e864a_2874822_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/missaoui_hu947b06f7f5a561d7b397ccb62f8e864a_2874822_634249af7e3dabb9c89b2014e20fdaef.webp"
width="760"
height="529"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Rotating features––inspiration from neuroscience where magnitude indicates presence of feature while orientation captures semantics. Does not solve BinarizedMNIST. Also did not try spherical features only.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/lowe_hu978b7028d7b63ba9059a3810281a6cb1_1867220_507c152f50b3d4960898cbf8566ddffd.webp 400w,
/post/neurips-2023-recap/posters/lowe_hu978b7028d7b63ba9059a3810281a6cb1_1867220_4842ec9c704387fab7f1355299f997be.webp 760w,
/post/neurips-2023-recap/posters/lowe_hu978b7028d7b63ba9059a3810281a6cb1_1867220_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/lowe_hu978b7028d7b63ba9059a3810281a6cb1_1867220_507c152f50b3d4960898cbf8566ddffd.webp"
width="760"
height="505"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Reusable Slotwise Mechanisms evaluated on CLEVRER and Physion. Compared with SwitchFormer, SlotFormer, and NPS (but also similar to Recurrent Independent Mechanisms?). They emphasize the role of a communication bottleneck between slots.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/nguyen_hufc55757a4e6b860ece907791a6c8018a_3764263_b7b8dfd199f9ae7ffabdd91ccbd2a4d9.webp 400w,
/post/neurips-2023-recap/posters/nguyen_hufc55757a4e6b860ece907791a6c8018a_3764263_d9a2cef9ee490f953e093a826caa5286.webp 760w,
/post/neurips-2023-recap/posters/nguyen_hufc55757a4e6b860ece907791a6c8018a_3764263_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/nguyen_hufc55757a4e6b860ece907791a6c8018a_3764263_b7b8dfd199f9ae7ffabdd91ccbd2a4d9.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>A novel self-supervised framework and model mimicking the mutual exclusivity bias from developmental psychology. Supports multi-object multi-view representation learning, and novel classes (as opposed to object discovery).&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/thai_hu065cc8b7b2741dc41ef871314acac316_2376327_323e251abd1592b959da843ea29a7115.webp 400w,
/post/neurips-2023-recap/posters/thai_hu065cc8b7b2741dc41ef871314acac316_2376327_3b5340f2b860f59e84eb31e67e572cc5.webp 760w,
/post/neurips-2023-recap/posters/thai_hu065cc8b7b2741dc41ef871314acac316_2376327_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/thai_hu065cc8b7b2741dc41ef871314acac316_2376327_323e251abd1592b959da843ea29a7115.webp"
width="760"
height="601"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Explicit symbolic labeling helps generalization on relational composition. The authors predict the next image on a Sticky Shapeworld dataset.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/sehgal_hub92782e96a1da13aeebec012a58c9cfd_3415652_13dbe8752dbf9c8b72c19babc7452862.webp 400w,
/post/neurips-2023-recap/posters/sehgal_hub92782e96a1da13aeebec012a58c9cfd_3415652_8d5a40c0cd20c0e6b86037b3490e5310.webp 760w,
/post/neurips-2023-recap/posters/sehgal_hub92782e96a1da13aeebec012a58c9cfd_3415652_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/sehgal_hub92782e96a1da13aeebec012a58c9cfd_3415652_13dbe8752dbf9c8b72c19babc7452862.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>A study of occlusions in video action detection. They have 9 severity levels based on background and foreground complexity. They show emergent segregation in capsules.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/modi_hub22eb7be181ada204f90128ff8210b5f_1538549_d49fba9a2bfe44677f18390aef016435.webp 400w,
/post/neurips-2023-recap/posters/modi_hub22eb7be181ada204f90128ff8210b5f_1538549_33e95e35ec32e5462bcd34c6c28d0eb3.webp 760w,
/post/neurips-2023-recap/posters/modi_hub22eb7be181ada204f90128ff8210b5f_1538549_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/modi_hub22eb7be181ada204f90128ff8210b5f_1538549_d49fba9a2bfe44677f18390aef016435.webp"
width="760"
height="496"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;h3 id="contrastive">Contrastive learning&lt;/h3>
&lt;p>Contextual pretraining using a buffer of data enables semantic segmentation, depth prediction, and in-context decoding via retrieval.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/balazevic_huf975ae7ca048e757d356ca77debf20ff_2991206_d8507d4affb0348862ba5bd38f77dbe3.webp 400w,
/post/neurips-2023-recap/posters/balazevic_huf975ae7ca048e757d356ca77debf20ff_2991206_622e048fd91d3c65b5e1fd7e4cc84d8a.webp 760w,
/post/neurips-2023-recap/posters/balazevic_huf975ae7ca048e757d356ca77debf20ff_2991206_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/balazevic_huf975ae7ca048e757d356ca77debf20ff_2991206_d8507d4affb0348862ba5bd38f77dbe3.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>FactorCL: Contrastive learning taking into account shared and unique information between X1 and X2. They propose using a conditional InfoNCE lower bound to include information and conditional CLUB upper bound to exclude information.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/liang_p_hu2f441730169f0dc1e039b7a6ebb8881d_3789961_598dec155a2f0ae4dbce82e30cd4679c.webp 400w,
/post/neurips-2023-recap/posters/liang_p_hu2f441730169f0dc1e039b7a6ebb8881d_3789961_d06e122831f0125f6540fea67980f44f.webp 760w,
/post/neurips-2023-recap/posters/liang_p_hu2f441730169f0dc1e039b7a6ebb8881d_3789961_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/liang_p_hu2f441730169f0dc1e039b7a6ebb8881d_3789961_598dec155a2f0ae4dbce82e30cd4679c.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>A comparison of video contrastive (MoCo and SimCLR), non-contrastive (BYOL, SimSiam, DINO), generative, and supervised methods under distribution shifts affecting content, viewpoint, actor, etc. Contrastive or Siamese methods learn better viewpoint invariance.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/sarkar_hu9189a779b7e5539e9a90d0e9dcd914e6_2928307_b3b9d536d1bd6fe6d92ceea62c9485a7.webp 400w,
/post/neurips-2023-recap/posters/sarkar_hu9189a779b7e5539e9a90d0e9dcd914e6_2928307_fb9cdf24945e4145e94c178849003783.webp 760w,
/post/neurips-2023-recap/posters/sarkar_hu9189a779b7e5539e9a90d0e9dcd914e6_2928307_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/sarkar_hu9189a779b7e5539e9a90d0e9dcd914e6_2928307_b3b9d536d1bd6fe6d92ceea62c9485a7.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Implicit contrastive learning using an asymmetric loss between a source encoder and target encoder (stop-gradient). The asymmetry helps push negative examples apart “in the service of pull”.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/lee_b_hu065cc8b7b2741dc41ef871314acac316_2753538_8458b23b5c2773a92e121e2de297aded.webp 400w,
/post/neurips-2023-recap/posters/lee_b_hu065cc8b7b2741dc41ef871314acac316_2753538_8e99b764d8f3070182e7ea8051a9c06a.webp 760w,
/post/neurips-2023-recap/posters/lee_b_hu065cc8b7b2741dc41ef871314acac316_2753538_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/lee_b_hu065cc8b7b2741dc41ef871314acac316_2753538_8458b23b5c2773a92e121e2de297aded.webp"
width="760"
height="453"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;h3 id="other-theory">Other Theory&lt;/h3>
&lt;p>[Best paper] Mirage. Emergence by scaling is a phenomenon of tracking nonlinear metrics such as accuracy. With smoother metrics, performance scales smoothly as a function of model size.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/schaeffer_hu028332e40727407c11cf6cf428db6e9e_572796_933a9ae934fb3dfb7f399e9af12e648d.webp 400w,
/post/neurips-2023-recap/posters/schaeffer_hu028332e40727407c11cf6cf428db6e9e_572796_fc4225eda8c35ca68ee1c31e894477d2.webp 760w,
/post/neurips-2023-recap/posters/schaeffer_hu028332e40727407c11cf6cf428db6e9e_572796_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/schaeffer_hu028332e40727407c11cf6cf428db6e9e_572796_933a9ae934fb3dfb7f399e9af12e648d.webp"
width="760"
height="725"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Fine-tuning by preserving the pairwise angle of weights rather than additive LoRA. This seems like a more constrained form of fine-tuning, but the authors showed it&amp;rsquo;s enough to fine-tune this way from random weights! The hypersphere energy has been shown to charecterize generalization. They also have new work on language models.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/qiu_hu32375ea25b4d6a5b6415bd2538b3a542_2048995_c237d2be5890216ad56a4dd9afb791f4.webp 400w,
/post/neurips-2023-recap/posters/qiu_hu32375ea25b4d6a5b6415bd2538b3a542_2048995_c25a2bf6493bb2b638f4eea9cb590ba3.webp 760w,
/post/neurips-2023-recap/posters/qiu_hu32375ea25b4d6a5b6415bd2538b3a542_2048995_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/qiu_hu32375ea25b4d6a5b6415bd2538b3a542_2048995_c237d2be5890216ad56a4dd9afb791f4.webp"
width="760"
height="561"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Tree of Thoughts implements deliberate problem solving via search in LLMs.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/yao_s_hu1ee8d7730e57fad5fef70fb9952686e4_3057410_aa3ce4abc19714f804d74bca42a35c2a.webp 400w,
/post/neurips-2023-recap/posters/yao_s_hu1ee8d7730e57fad5fef70fb9952686e4_3057410_bb5ad9033a1c363095994092cfbb2ac5.webp 760w,
/post/neurips-2023-recap/posters/yao_s_hu1ee8d7730e57fad5fef70fb9952686e4_3057410_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/yao_s_hu1ee8d7730e57fad5fef70fb9952686e4_3057410_aa3ce4abc19714f804d74bca42a35c2a.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Tree VAEs to enable hierarchical clustering. The tree structure is discovered using an iterative growing schedule.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/manduchi_hu8705c562f94dbcab9622ce8d148f9c76_2795503_5bc5bb3f594eebb4e63396d7a07bf05b.webp 400w,
/post/neurips-2023-recap/posters/manduchi_hu8705c562f94dbcab9622ce8d148f9c76_2795503_a13aea9f7ae25fa8ef3aa5901e87b724.webp 760w,
/post/neurips-2023-recap/posters/manduchi_hu8705c562f94dbcab9622ce8d148f9c76_2795503_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/manduchi_hu8705c562f94dbcab9622ce8d148f9c76_2795503_5bc5bb3f594eebb4e63396d7a07bf05b.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Identifiability theory for nonlinear ICA: what structural sparsity assumptions lead to identifiability of the generative process given independent sources.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/zheng_y_hu3a345345c55c1a3b0df073fe9d6fae10_2405594_e2500750a79dd7120c2f531ba91b21ff.webp 400w,
/post/neurips-2023-recap/posters/zheng_y_hu3a345345c55c1a3b0df073fe9d6fae10_2405594_16cc88bfa4dc19cc604196d3bb805467.webp 760w,
/post/neurips-2023-recap/posters/zheng_y_hu3a345345c55c1a3b0df073fe9d6fae10_2405594_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/zheng_y_hu3a345345c55c1a3b0df073fe9d6fae10_2405594_e2500750a79dd7120c2f531ba91b21ff.webp"
width="760"
height="591"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Causal generalization of ICA.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/wendong_huf16b0a65963cbf378fbacd6eb33e4751_1876653_d299ef383005da81b55e4d468bd53bd1.webp 400w,
/post/neurips-2023-recap/posters/wendong_huf16b0a65963cbf378fbacd6eb33e4751_1876653_fec1439a1358f4f4e2011f3e955fdaf7.webp 760w,
/post/neurips-2023-recap/posters/wendong_huf16b0a65963cbf378fbacd6eb33e4751_1876653_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/wendong_huf16b0a65963cbf378fbacd6eb33e4751_1876653_d299ef383005da81b55e4d468bd53bd1.webp"
width="760"
height="616"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>An assessment of the permutation conjecture when pretrained ReLU convolutional filters are flipped (horizontally mirrored). They compare data augmentation and the use of an invariance loss versus the perfect baseline (a graph conv network, GCN). Flipped features are not perfectly equal, suggesting it would make sense to average outputs when permutation equivariance is desired.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/bokman_hu947b06f7f5a561d7b397ccb62f8e864a_2491428_4da82ec34cdb932cafcb7bdf0ba5ad62.webp 400w,
/post/neurips-2023-recap/posters/bokman_hu947b06f7f5a561d7b397ccb62f8e864a_2491428_b9359bbcb55e278d7f1de2a749079a08.webp 760w,
/post/neurips-2023-recap/posters/bokman_hu947b06f7f5a561d7b397ccb62f8e864a_2491428_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/bokman_hu947b06f7f5a561d7b397ccb62f8e864a_2491428_4da82ec34cdb932cafcb7bdf0ba5ad62.webp"
width="760"
height="590"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Quantifying the extent to which Graph NNs model interaction between vertices. This leads to an edge sparsification algorithm with performance benefits.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/razin_hu31cddbefb4032adb8c069c53e8c18441_1331654_6ebd4ab2cf3f516436ddda003ae5a913.webp 400w,
/post/neurips-2023-recap/posters/razin_hu31cddbefb4032adb8c069c53e8c18441_1331654_8d7a9ceaf086695db161e3b904d7e6e4.webp 760w,
/post/neurips-2023-recap/posters/razin_hu31cddbefb4032adb8c069c53e8c18441_1331654_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/razin_hu31cddbefb4032adb8c069c53e8c18441_1331654_6ebd4ab2cf3f516436ddda003ae5a913.webp"
width="760"
height="479"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Convolutional State Space Models: help improve long horizon moving MNIST generation over 600 frames. Also evaluated on Minecraft and Habitat long-range benchmarks.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/smith_hua0a49ff8394c50d68b78bf16ca4054e8_2052568_21cc8afae159455d333cc52996e3c97d.webp 400w,
/post/neurips-2023-recap/posters/smith_hua0a49ff8394c50d68b78bf16ca4054e8_2052568_be21a9be1e444c6fd5cab85047fbdf70.webp 760w,
/post/neurips-2023-recap/posters/smith_hua0a49ff8394c50d68b78bf16ca4054e8_2052568_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/smith_hua0a49ff8394c50d68b78bf16ca4054e8_2052568_21cc8afae159455d333cc52996e3c97d.webp"
width="760"
height="407"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;h3 id="3d">3D&lt;/h3>
&lt;p>A 3D feature extractor based on three different methods to inject 3D information into LLMs: features from direct reconstruction (of point cloud and images), grandSLAM, and NeRFs. On the language side, they generate language data using a pipeline based on ChatGPT and BLIP. Evaluation on ScanQA.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/hong_y_hu0df79afed70143aa713a98f14be0919e_2625659_731f4ee4ecae181cfa466ae6408d1021.webp 400w,
/post/neurips-2023-recap/posters/hong_y_hu0df79afed70143aa713a98f14be0919e_2625659_87218948faf441dbc0c1986ec2ced281.webp 760w,
/post/neurips-2023-recap/posters/hong_y_hu0df79afed70143aa713a98f14be0919e_2625659_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/hong_y_hu0df79afed70143aa713a98f14be0919e_2625659_731f4ee4ecae181cfa466ae6408d1021.webp"
width="760"
height="532"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>&lt;a href="https://sd-complements-dino.github.io/" target="_blank" rel="noopener">Fusing Stable Diffusion and DINO&lt;/a> features followed by zero-shot evaluation can outperform SoTA methods on dense and semantic correspondence tasks.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/zhang_j_hua35a34e6fc89fea14e57846ae9cfc704_2222229_5ae4ffa137dee026c03c523494a7c65f.webp 400w,
/post/neurips-2023-recap/posters/zhang_j_hua35a34e6fc89fea14e57846ae9cfc704_2222229_fd915f9b7763cbf625835586b285a116.webp 760w,
/post/neurips-2023-recap/posters/zhang_j_hua35a34e6fc89fea14e57846ae9cfc704_2222229_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/zhang_j_hua35a34e6fc89fea14e57846ae9cfc704_2222229_5ae4ffa137dee026c03c523494a7c65f.webp"
width="760"
height="419"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Autodecoding to learn latents to diffuse on for 3D voxel diffusion. Evaluation based on FID only because they didn&amp;rsquo;t want to reconstruct objects. Some normalization trick involving the median and IQ distance to make stats of 3D diffusion latents look better. MVImgNet contains real objects but might still be partial views because the objects are placed against something during imaging.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/ntavelis_hu86bec41e0939df1923e015e6980f04e3_2818456_53ca6e4318b3b6abe38b857be01f45ac.webp 400w,
/post/neurips-2023-recap/posters/ntavelis_hu86bec41e0939df1923e015e6980f04e3_2818456_bd07d4988ad640d8072a9a6d42b31b55.webp 760w,
/post/neurips-2023-recap/posters/ntavelis_hu86bec41e0939df1923e015e6980f04e3_2818456_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/ntavelis_hu86bec41e0939df1923e015e6980f04e3_2818456_53ca6e4318b3b6abe38b857be01f45ac.webp"
width="760"
height="432"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>sVORF: Slot-guided volumetric radiance fields. (Follow up on uORF?). Impressive segmentation of CLEVR shadows.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/qi_d_hub5587884cf0cee06657c7a59d408e6d1_3676066_ba3979f4f01cc34d69cd37a21a3cb67e.webp 400w,
/post/neurips-2023-recap/posters/qi_d_hub5587884cf0cee06657c7a59d408e6d1_3676066_ef7861d961f30496da3ce507d4bbd22d.webp 760w,
/post/neurips-2023-recap/posters/qi_d_hub5587884cf0cee06657c7a59d408e6d1_3676066_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/qi_d_hub5587884cf0cee06657c7a59d408e6d1_3676066_ba3979f4f01cc34d69cd37a21a3cb67e.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>&lt;a href="https://www.tmonnier.com/DBW/" target="_blank" rel="noopener">Differentiable Blocks World&lt;/a>: represent shapes from multi-view images using primitive blocks.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/monnier_huae63baab1e6299f5082e1d44c56d49cd_3995548_f32960a611769e41e8faf65620c6159e.webp 400w,
/post/neurips-2023-recap/posters/monnier_huae63baab1e6299f5082e1d44c56d49cd_3995548_2aa1029b27d8105a8cc0568980c923a1.webp 760w,
/post/neurips-2023-recap/posters/monnier_huae63baab1e6299f5082e1d44c56d49cd_3995548_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/monnier_huae63baab1e6299f5082e1d44c56d49cd_3995548_f32960a611769e41e8faf65620c6159e.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>&lt;a href="https://frozenburning.github.io/projects/primdiffusion/" target="_blank" rel="noopener">PrimDiffusion&lt;/a>: Diffusion on volumetric primitives to generate 3D humans. Also permits texture transfer (e.g., clothing) and 3D inpainting.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/chen_z_hu2e18505ed1e27cf1f6c748b8e8fe41f5_4232316_e5d3b352c4605dcb86cbbad2d2b50ef1.webp 400w,
/post/neurips-2023-recap/posters/chen_z_hu2e18505ed1e27cf1f6c748b8e8fe41f5_4232316_7dc4257bdcf7c0c167737bfb23997e4f.webp 760w,
/post/neurips-2023-recap/posters/chen_z_hu2e18505ed1e27cf1f6c748b8e8fe41f5_4232316_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/chen_z_hu2e18505ed1e27cf1f6c748b8e8fe41f5_4232316_e5d3b352c4605dcb86cbbad2d2b50ef1.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Supervising on raw Lidar (photon histograms) is better than simple depth-supervised Nerf. It helps distribute radiance over the full distribution of depth rather than targeting a specific depth. One caveat is all photon data was collected in a dark room.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/malik_hu441e214a60d9f5ad54035457be60c892_4187741_9d09b76e7eb50160770bb855d5792d5b.webp 400w,
/post/neurips-2023-recap/posters/malik_hu441e214a60d9f5ad54035457be60c892_4187741_fa563cb2d2b35426bae8d75e063e85e5.webp 760w,
/post/neurips-2023-recap/posters/malik_hu441e214a60d9f5ad54035457be60c892_4187741_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/malik_hu441e214a60d9f5ad54035457be60c892_4187741_9d09b76e7eb50160770bb855d5792d5b.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>A modular look at 3D-aware image synthesis. They look at point embedding, feature decoding, volume rendering, upsampling, and pose sampling techniques separately. They reproduce models from 2020-2022, showing EG3D was a huge improvement over previous models. They show “volume” point embedding alone (over MLP or Tri-plane representations) gives the best FID on Cats and Cars, while being competitive on FFHQ.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/wang_q_hu7365387379434051c4daf923a45b3771_4192987_5b208e88b2da35f77de77d1f97d21061.webp 400w,
/post/neurips-2023-recap/posters/wang_q_hu7365387379434051c4daf923a45b3771_4192987_fa939504f1a865b7c93edb3d0e712edd.webp 760w,
/post/neurips-2023-recap/posters/wang_q_hu7365387379434051c4daf923a45b3771_4192987_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/wang_q_hu7365387379434051c4daf923a45b3771_4192987_5b208e88b2da35f77de77d1f97d21061.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>OpenMask3D: takes posed RGB-D frames and a reconstructed 3D geometry to produce open-vocabulary 3D instance segmentations. They first get class-agnostic masks, then compute mask features by taking top-k views, and finally use CLIP to query all 3D instances.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/takmaz_hud9ff233e8e02e30a8ad10a12484c3611_2646822_22fc74bfa0522978ca79dcf64c4611a3.webp 400w,
/post/neurips-2023-recap/posters/takmaz_hud9ff233e8e02e30a8ad10a12484c3611_2646822_b0836b604963e4b55298249b7b023f38.webp 760w,
/post/neurips-2023-recap/posters/takmaz_hud9ff233e8e02e30a8ad10a12484c3611_2646822_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/takmaz_hud9ff233e8e02e30a8ad10a12484c3611_2646822_22fc74bfa0522978ca79dcf64c4611a3.webp"
width="588"
height="760"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>3D room layout generation from RGB video and “manual 2D annotations of structural elements and their visible parts” (just segmentations?). The method combines point tracking, edge matching, and perpendicularity constraints. The authors release a video and CAD dataset called CAD-Estate.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/rozumnyi_hu647c4ae3830339327e9fdf9773dac300_3332022_e05d38631fb6c9976909d2d322641305.webp 400w,
/post/neurips-2023-recap/posters/rozumnyi_hu647c4ae3830339327e9fdf9773dac300_3332022_896bd731cd7dc6bbe1a0452d85d001c6.webp 760w,
/post/neurips-2023-recap/posters/rozumnyi_hu647c4ae3830339327e9fdf9773dac300_3332022_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/rozumnyi_hu647c4ae3830339327e9fdf9773dac300_3332022_e05d38631fb6c9976909d2d322641305.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>NAVI: a dataset of multi-view and in-the-wild image collections of 36 objects and 10k images. The authors show an improvement over using COLMAP poses. By aligning to 3D, they also provide dense geometric correspondences.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/jampani_hu115c38faba30f8b726a8f31d98be9030_3007379_b42024751bdac1425f2cea3abf390edb.webp 400w,
/post/neurips-2023-recap/posters/jampani_hu115c38faba30f8b726a8f31d98be9030_3007379_df74a07a71f5bb4c1d9c6a87b94451d3.webp 760w,
/post/neurips-2023-recap/posters/jampani_hu115c38faba30f8b726a8f31d98be9030_3007379_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/jampani_hu115c38faba30f8b726a8f31d98be9030_3007379_b42024751bdac1425f2cea3abf390edb.webp"
width="760"
height="396"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>An extension of EPIC Kitchens with 3D geometry (pointclouds and cameras) for 19M frames. Includes VISOR annotations for segmenting objects and hands.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/tschernezki_hu31cddbefb4032adb8c069c53e8c18441_1974145_66249faa3ed6fcf0a542b4a6d4c254d3.webp 400w,
/post/neurips-2023-recap/posters/tschernezki_hu31cddbefb4032adb8c069c53e8c18441_1974145_7560078e5b9b85c3431a240e642271fd.webp 760w,
/post/neurips-2023-recap/posters/tschernezki_hu31cddbefb4032adb8c069c53e8c18441_1974145_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/tschernezki_hu31cddbefb4032adb8c069c53e8c18441_1974145_66249faa3ed6fcf0a542b4a6d4c254d3.webp"
width="760"
height="483"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Virtual objects pasted into real scenes, with 3D bounding boxes and plausible physical locations. This synthetic data helps achieve SoTA on monocular 3D detection.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/ge_y_hu71215ad1fcd09b7825443f3f7d49b8fe_3406259_338421515e80d4b2f89a0399a1089612.webp 400w,
/post/neurips-2023-recap/posters/ge_y_hu71215ad1fcd09b7825443f3f7d49b8fe_3406259_370672b06e9cc77d092944aa9925435a.webp 760w,
/post/neurips-2023-recap/posters/ge_y_hu71215ad1fcd09b7825443f3f7d49b8fe_3406259_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/ge_y_hu71215ad1fcd09b7825443f3f7d49b8fe_3406259_338421515e80d4b2f89a0399a1089612.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Planning over partially observed 3D scene graphs using language models. The method PROPHE-C is based on MC trajectory sampling to hallucinate possible 3DSGs.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/remy_hu31a8a95040bb19b3ddac5ce899e4f993_3000573_9b19aa6a1626d5c5fbb932e3a0de4628.webp 400w,
/post/neurips-2023-recap/posters/remy_hu31a8a95040bb19b3ddac5ce899e4f993_3000573_2efdc3d752078604ebaec440726d1050.webp 760w,
/post/neurips-2023-recap/posters/remy_hu31a8a95040bb19b3ddac5ce899e4f993_3000573_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/remy_hu31a8a95040bb19b3ddac5ce899e4f993_3000573_9b19aa6a1626d5c5fbb932e3a0de4628.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Diffusion over articulation trees (graphs) to generate articulated objects.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/lei_j_hu5790d4509a1fe90141692ac19185cfac_4177079_6e25b12b99267f7febe3667e8229aa20.webp 400w,
/post/neurips-2023-recap/posters/lei_j_hu5790d4509a1fe90141692ac19185cfac_4177079_11644754c4defc2cac48b545188ddabd.webp 760w,
/post/neurips-2023-recap/posters/lei_j_hu5790d4509a1fe90141692ac19185cfac_4177079_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/lei_j_hu5790d4509a1fe90141692ac19185cfac_4177079_6e25b12b99267f7febe3667e8229aa20.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;h3 id="diffusion-and-gans">Diffusion and GANs (for images)&lt;/h3>
&lt;p>Text-to-panorama generation using synchronized joint diffusions. They compute the coherence for multiple diffusion processes in &lt;em>advance&lt;/em>.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/lee_y_hu9f01ab06a85535f84fd61e3e45d761dd_2295906_5893bcbf4f2e53b280e6826bec248236.webp 400w,
/post/neurips-2023-recap/posters/lee_y_hu9f01ab06a85535f84fd61e3e45d761dd_2295906_33160f2f4128bb5abf8a2286ad8ba480.webp 760w,
/post/neurips-2023-recap/posters/lee_y_hu9f01ab06a85535f84fd61e3e45d761dd_2295906_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/lee_y_hu9f01ab06a85535f84fd61e3e45d761dd_2295906_5893bcbf4f2e53b280e6826bec248236.webp"
width="760"
height="338"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Diffusion on bounded ranges using beta distributions in both directions. Optimized using KL-divergence upper bounds rather than reweighted ELBOs. They also compare with categorical or count-based diffusion processes.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/zhou_m_hu216f6e010bd369f8907ae9ca6afe3cfd_1820647_ac0a98f026c07ede7246e1452e72fec9.webp 400w,
/post/neurips-2023-recap/posters/zhou_m_hu216f6e010bd369f8907ae9ca6afe3cfd_1820647_1ba65ff720f968b942128c9bbc8656af.webp 760w,
/post/neurips-2023-recap/posters/zhou_m_hu216f6e010bd369f8907ae9ca6afe3cfd_1820647_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/zhou_m_hu216f6e010bd369f8907ae9ca6afe3cfd_1820647_ac0a98f026c07ede7246e1452e72fec9.webp"
width="760"
height="628"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Precision-recall curves for generative models. Each point on Fig 1(b) is a f-divergence that can be optimized. We might choose a specific one based on the quality/diversity trade-off we’re looking for.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/verine_hu6ef5883331fb681a38a84d8e860f75d0_2348175_01ff0c0a9b48b5b7ca57be115897c07b.webp 400w,
/post/neurips-2023-recap/posters/verine_hu6ef5883331fb681a38a84d8e860f75d0_2348175_8c2f25dd0caded9326a128a1e5fc4a9f.webp 760w,
/post/neurips-2023-recap/posters/verine_hu6ef5883331fb681a38a84d8e860f75d0_2348175_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/verine_hu6ef5883331fb681a38a84d8e860f75d0_2348175_01ff0c0a9b48b5b7ca57be115897c07b.webp"
width="760"
height="490"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Particle-based framework to unify GANs and score-based diffusion. They study particle models with and without generators (which enable interaction of particles) versus various gradient vector fields that the generated particles may follow: Wasserstein, Log Ratio, and Discriminator-based.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/franceschi_hu2037d40e356ecf8aad44890a47538ad6_4172718_d2a4e59ca5c5f8973f118b58c92bdbae.webp 400w,
/post/neurips-2023-recap/posters/franceschi_hu2037d40e356ecf8aad44890a47538ad6_4172718_3168d71519f598398e0b3c1ca086c1c0.webp 760w,
/post/neurips-2023-recap/posters/franceschi_hu2037d40e356ecf8aad44890a47538ad6_4172718_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/franceschi_hu2037d40e356ecf8aad44890a47538ad6_4172718_d2a4e59ca5c5f8973f118b58c92bdbae.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Classification using diffusion models using a labels x timesteps score matrix and timesteps weighting function.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/clark_hu3d08974d51ff5e88e98b21ed3696c805_1425452_a8a9dd2fa8d7887a27c150bc3e5e2899.webp 400w,
/post/neurips-2023-recap/posters/clark_hu3d08974d51ff5e88e98b21ed3696c805_1425452_7d0bba51ecf1c97d1a3be3906a24a7f8.webp 760w,
/post/neurips-2023-recap/posters/clark_hu3d08974d51ff5e88e98b21ed3696c805_1425452_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/clark_hu3d08974d51ff5e88e98b21ed3696c805_1425452_a8a9dd2fa8d7887a27c150bc3e5e2899.webp"
width="760"
height="621"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Reweighted diffusion losses (which look different from the ELBO) can be seen as ELBO + additive Gaussian data augmentation.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/kingma_hu53e2b58bc195e8d236d17ba5382559d9_2032506_436531bf3a3910b61981569758913cc5.webp 400w,
/post/neurips-2023-recap/posters/kingma_hu53e2b58bc195e8d236d17ba5382559d9_2032506_a872328e273741962eb9e1fc323d93e7.webp 760w,
/post/neurips-2023-recap/posters/kingma_hu53e2b58bc195e8d236d17ba5382559d9_2032506_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/kingma_hu53e2b58bc195e8d236d17ba5382559d9_2032506_436531bf3a3910b61981569758913cc5.webp"
width="760"
height="474"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Diffusion models for dense vision tasks. Using large-scale synthetic data boosts performance. Evaluation on KITTI flow and NYU depth.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/saxena_hucbe0000f3d8642ac7b9195ce4bdea8f0_1788060_28ce1882c75fa5a7ba699d96c4f23045.webp 400w,
/post/neurips-2023-recap/posters/saxena_hucbe0000f3d8642ac7b9195ce4bdea8f0_1788060_10d1a49edbd8ce5273390644cf429839.webp 760w,
/post/neurips-2023-recap/posters/saxena_hucbe0000f3d8642ac7b9195ce4bdea8f0_1788060_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/saxena_hucbe0000f3d8642ac7b9195ce4bdea8f0_1788060_28ce1882c75fa5a7ba699d96c4f23045.webp"
width="760"
height="396"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;h3 id="concepts-disentanglement">Concepts and disentanglement&lt;/h3>
&lt;p>Multiple classification taxonomies on CLEVR. MAEs perform the worst on unsupervised clustering (of frozen features) on shape. All models are bad at counting––it’s better to train a ResNet from scratch.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/vaze_hu32375ea25b4d6a5b6415bd2538b3a542_2708036_4d87e6bec54da3ffc60304c20a16b7b8.webp 400w,
/post/neurips-2023-recap/posters/vaze_hu32375ea25b4d6a5b6415bd2538b3a542_2708036_30e11768b918b86b0347578993856c1e.webp 760w,
/post/neurips-2023-recap/posters/vaze_hu32375ea25b4d6a5b6415bd2538b3a542_2708036_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/vaze_hu32375ea25b4d6a5b6415bd2538b3a542_2708036_4d87e6bec54da3ffc60304c20a16b7b8.webp"
width="760"
height="434"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Concept algebra as an alternative to prompt tuning to generate images. Every prompt induces a concept distribution which can be mapped to a vector space. Concepts are then modified on their own subspaces.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/wang_z_hu947b06f7f5a561d7b397ccb62f8e864a_2796223_9d778a24a0c9dd6b85ba82aeae8d40d9.webp 400w,
/post/neurips-2023-recap/posters/wang_z_hu947b06f7f5a561d7b397ccb62f8e864a_2796223_6430fffd0c8cb83569edfde901f02110.webp 760w,
/post/neurips-2023-recap/posters/wang_z_hu947b06f7f5a561d7b397ccb62f8e864a_2796223_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/wang_z_hu947b06f7f5a561d7b397ccb62f8e864a_2796223_9d778a24a0c9dd6b85ba82aeae8d40d9.webp"
width="760"
height="535"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Concepts as dictionary learning and concept importance as attribution to image pixels. Measures like TCAV and Sobol. Visualization of concepts (e.g., what makes an espresso) available &lt;a href="https://serre-lab.github.io/Lens/" target="_blank" rel="noopener">in this demo&lt;/a>.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/fel_hu065cc8b7b2741dc41ef871314acac316_3079722_75859473f943034306dcac805da945d1.webp 400w,
/post/neurips-2023-recap/posters/fel_hu065cc8b7b2741dc41ef871314acac316_3079722_53ce7c4d1b5b1c96b3fc6010f8f8ac98.webp 760w,
/post/neurips-2023-recap/posters/fel_hu065cc8b7b2741dc41ef871314acac316_3079722_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/fel_hu065cc8b7b2741dc41ef871314acac316_3079722_75859473f943034306dcac805da945d1.webp"
width="760"
height="448"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Quantization works better than regularization to learn disentangled latent representations?&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/hsu_hu3a345345c55c1a3b0df073fe9d6fae10_2451871_e72459ecb1b0928af995c40bd8baffae.webp 400w,
/post/neurips-2023-recap/posters/hsu_hu3a345345c55c1a3b0df073fe9d6fae10_2451871_8e9139ee726e888a3d26e1e78377f8ed.webp 760w,
/post/neurips-2023-recap/posters/hsu_hu3a345345c55c1a3b0df073fe9d6fae10_2451871_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/hsu_hu3a345345c55c1a3b0df073fe9d6fae10_2451871_e72459ecb1b0928af995c40bd8baffae.webp"
width="760"
height="588"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>A theory of compositional generalization. The claims are: it requires a compositional support (i.e., training set supporting all configurations) that is also sufficient (i.e., enables reconstruction of all components).&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/wiedemer_hu3a345345c55c1a3b0df073fe9d6fae10_2996187_29e08ea57c7e277c98f61b6802fe7b58.webp 400w,
/post/neurips-2023-recap/posters/wiedemer_hu3a345345c55c1a3b0df073fe9d6fae10_2996187_6a797cbfdb15eacb8ca152957ba5a540.webp 760w,
/post/neurips-2023-recap/posters/wiedemer_hu3a345345c55c1a3b0df073fe9d6fae10_2996187_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/wiedemer_hu3a345345c55c1a3b0df073fe9d6fae10_2996187_29e08ea57c7e277c98f61b6802fe7b58.webp"
width="760"
height="447"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;h3 id="novelties">Fun/novelties/new datasets&lt;/h3>
&lt;p>Tracking human positions based on the acoustic profile of the room.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/wang_m_hu115c38faba30f8b726a8f31d98be9030_2670853_bb951acb45e5a251ffa1f9167eb662b0.webp 400w,
/post/neurips-2023-recap/posters/wang_m_hu115c38faba30f8b726a8f31d98be9030_2670853_1f1b74b66d825dbe06415b25c1fe0994.webp 760w,
/post/neurips-2023-recap/posters/wang_m_hu115c38faba30f8b726a8f31d98be9030_2670853_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/wang_m_hu115c38faba30f8b726a8f31d98be9030_2670853_bb951acb45e5a251ffa1f9167eb662b0.webp"
width="760"
height="402"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>A fun dataset of chimpanzee actions/behaviors. Permits detection, pose estimation, and spatiotemporal action detection.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/ma_x_hu29b7fe8438b0b95b1373bcb3def8678e_3257593_3bdf0c3b0e7a2d9bc6fe4f735ef96e65.webp 400w,
/post/neurips-2023-recap/posters/ma_x_hu29b7fe8438b0b95b1373bcb3def8678e_3257593_360f5c1ad3144d7afda6cffc3cbe1ff0.webp 760w,
/post/neurips-2023-recap/posters/ma_x_hu29b7fe8438b0b95b1373bcb3def8678e_3257593_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/ma_x_hu29b7fe8438b0b95b1373bcb3def8678e_3257593_3bdf0c3b0e7a2d9bc6fe4f735ef96e65.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Efficiency gains in voting unlocked using an explainable randomization step added to deterministic voting rules.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/ebadian_hu8d01b7813c2e9c3ad2eea83ee823b238_3031828_f885707b1b746b65ea9593087af739e3.webp 400w,
/post/neurips-2023-recap/posters/ebadian_hu8d01b7813c2e9c3ad2eea83ee823b238_3031828_e2c8f8ecddd8a587a7edf10a080a5af5.webp 760w,
/post/neurips-2023-recap/posters/ebadian_hu8d01b7813c2e9c3ad2eea83ee823b238_3031828_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/ebadian_hu8d01b7813c2e9c3ad2eea83ee823b238_3031828_f885707b1b746b65ea9593087af739e3.webp"
width="760"
height="642"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Geographically diverse images per object class compared to ImageNet. They identify gaps in model performance (ResNet50, CLIP) from region to region. They also train models on their dataset, and show improvements when testing on a different geo-diverse dataset.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/ramaswamy_hu6449ba66d208d87c3bd7215aabad5e1f_4238457_da59651c5f6b26f223448153334e0c1e.webp 400w,
/post/neurips-2023-recap/posters/ramaswamy_hu6449ba66d208d87c3bd7215aabad5e1f_4238457_83e6ea893aa754be22293318b462bf4b.webp 760w,
/post/neurips-2023-recap/posters/ramaswamy_hu6449ba66d208d87c3bd7215aabad5e1f_4238457_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/ramaswamy_hu6449ba66d208d87c3bd7215aabad5e1f_4238457_da59651c5f6b26f223448153334e0c1e.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>DataComp: a large-scale multimodal dataset that outperforms LAION.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/gadre_hu224a6c9c799c1c621674561257b65479_4342152_d81e816af1f6adcd1b7e0deed180a66e.webp 400w,
/post/neurips-2023-recap/posters/gadre_hu224a6c9c799c1c621674561257b65479_4342152_ecb1d21239894b1a3d069b8b0ac6e395.webp 760w,
/post/neurips-2023-recap/posters/gadre_hu224a6c9c799c1c621674561257b65479_4342152_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/gadre_hu224a6c9c799c1c621674561257b65479_4342152_d81e816af1f6adcd1b7e0deed180a66e.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;h1 id="workshops">Workshops&lt;/h1>
&lt;p>Object placement in scenes by first hallucinating scenes around objects. A model is trained to calculate placement plausibility based on positive (generated) examples and negative (bad crop) examples. The model can be run on grid points on the target scene to find the best object placement.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/yuan_l_hua568ae46cfc75e6781dbb76fefba2ddb_3221936_8779ce4e59a2c33072388912d6fb18c3.webp 400w,
/post/neurips-2023-recap/posters/yuan_l_hua568ae46cfc75e6781dbb76fefba2ddb_3221936_7ac7913c40b5490b4b67f11248822264.webp 760w,
/post/neurips-2023-recap/posters/yuan_l_hua568ae46cfc75e6781dbb76fefba2ddb_3221936_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/yuan_l_hua568ae46cfc75e6781dbb76fefba2ddb_3221936_8779ce4e59a2c33072388912d6fb18c3.webp"
width="573"
height="760"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>DA-Fusion: dreambooth style tokens which allow generating variations of an image for data augmentation. Keeping the token fixed, it is possible to alter the image latent at some stage t of the diffusion generative process.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/trabucco_hu2dd811786fe2af575ca96c0c840e6148_3999027_e4070f083cdfbe7137d3dfcb6e1340ab.webp 400w,
/post/neurips-2023-recap/posters/trabucco_hu2dd811786fe2af575ca96c0c840e6148_3999027_a13518292656a57fbaa93df035b5cb36.webp 760w,
/post/neurips-2023-recap/posters/trabucco_hu2dd811786fe2af575ca96c0c840e6148_3999027_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/trabucco_hu2dd811786fe2af575ca96c0c840e6148_3999027_e4070f083cdfbe7137d3dfcb6e1340ab.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Diffusion models do disentangle when shown only diagonal elements of a feature grid, if they are stopped early during training.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/scimeca_hu31cddbefb4032adb8c069c53e8c18441_2139671_931242f8300f207270dc079317e05f8b.webp 400w,
/post/neurips-2023-recap/posters/scimeca_hu31cddbefb4032adb8c069c53e8c18441_2139671_015d01f51cb58ad0142c1faa53099ef1.webp 760w,
/post/neurips-2023-recap/posters/scimeca_hu31cddbefb4032adb8c069c53e8c18441_2139671_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/scimeca_hu31cddbefb4032adb8c069c53e8c18441_2139671_931242f8300f207270dc079317e05f8b.webp"
width="760"
height="507"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>SAD: segmentation over RGBD by running SAM on colormap of depth, and then fusing with an OVSeg segmentation of the RGB image.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/cen_j_hu40c960659d6daa53b848eed8f25a17c3_2828611_3c033c82f29672d0aaae2145908a70c0.webp 400w,
/post/neurips-2023-recap/posters/cen_j_hu40c960659d6daa53b848eed8f25a17c3_2828611_a9f60ac4fafdfdb51a3a9d2caa0ae129.webp 760w,
/post/neurips-2023-recap/posters/cen_j_hu40c960659d6daa53b848eed8f25a17c3_2828611_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/cen_j_hu40c960659d6daa53b848eed8f25a17c3_2828611_3c033c82f29672d0aaae2145908a70c0.webp"
width="573"
height="760"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Distance Learner: a new paradigm (same vein as maximum margin?) for classification. Yields out-of-domain understanding because the classifier becomes more uncertain away from the data manifold.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/chetan_huae9785cf4f69c487acbedd5aca3becc8_3893824_7f4cdccff1809aa58d3ca674474ba05d.webp 400w,
/post/neurips-2023-recap/posters/chetan_huae9785cf4f69c487acbedd5aca3becc8_3893824_b3bbfc9bd216095e92af91e245e2b0bf.webp 760w,
/post/neurips-2023-recap/posters/chetan_huae9785cf4f69c487acbedd5aca3becc8_3893824_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/chetan_huae9785cf4f69c487acbedd5aca3becc8_3893824_7f4cdccff1809aa58d3ca674474ba05d.webp"
width="573"
height="760"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Investigating shape bias in ResNets and ViTs.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/benarous_hubd6cc2a4a182e46b9766c8021978a7ef_2635279_1534aeafc62554614426f6a3278a8114.webp 400w,
/post/neurips-2023-recap/posters/benarous_hubd6cc2a4a182e46b9766c8021978a7ef_2635279_c4485fc15da3b5dc2530246200008ac8.webp 760w,
/post/neurips-2023-recap/posters/benarous_hubd6cc2a4a182e46b9766c8021978a7ef_2635279_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/benarous_hubd6cc2a4a182e46b9766c8021978a7ef_2635279_1534aeafc62554614426f6a3278a8114.webp"
width="760"
height="529"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Synthetic data generation optimized using one-shot feedback from an existing classifier to balance classes. Improves classification performance (in a newly trained model) on underrepresented classes in ImageNet-LT and NICO++.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/askari_hu222d06e1030f55b036be330bef29021e_2642811_0c1b55e027594b8a4368d292ac00de7f.webp 400w,
/post/neurips-2023-recap/posters/askari_hu222d06e1030f55b036be330bef29021e_2642811_1f1b87dd2fe97a4f8f6a2a5b8e0622ca.webp 760w,
/post/neurips-2023-recap/posters/askari_hu222d06e1030f55b036be330bef29021e_2642811_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/askari_hu222d06e1030f55b036be330bef29021e_2642811_0c1b55e027594b8a4368d292ac00de7f.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>f-GANs settle scores: assuming optimal discriminators, they show the optimal generator must satisfy a score-based equality. The Reverse KL leaves only one term (based on the score) in the product. Better results on mode collapse using ScoreGAN.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/neurips-2023-recap/posters/asokan_hu4984a2f5b218d6fe87f1039f4820c87a_3259289_abda278e23226df3c03c9e87569bb986.webp 400w,
/post/neurips-2023-recap/posters/asokan_hu4984a2f5b218d6fe87f1039f4820c87a_3259289_829eb769e564a7c7c36aed359d3d1917.webp 760w,
/post/neurips-2023-recap/posters/asokan_hu4984a2f5b218d6fe87f1039f4820c87a_3259289_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/neurips-2023-recap/posters/asokan_hu4984a2f5b218d6fe87f1039f4820c87a_3259289_abda278e23226df3c03c9e87569bb986.webp"
width="538"
height="760"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;h1 id="heading">&lt;/h1>
&lt;h1 id="favorite-talks">Favorite talks&lt;/h1>
&lt;p>Chelsea Finn (Stanford) at SSL workshop on (1) improving factuality using “semantic entropy,” a model intrinsic quantity, and (2) dealing with cut-off dates using meta-learning (CAMELS) over new information.&lt;/p>
&lt;p>Adji Bousso Dieng (Princeton) at SyntheticData4ML on Vendi scores to assess dataset diversity (better than averaging a similarity matrix).&lt;/p>
&lt;p>Björn Ommer (LMU, StableDiffusion) keynote on capturing long-range dependencies (e.g., by diffusing on a latent space).&lt;/p>
&lt;p>Linda Smith (IndianaU) keynote on how babies might learn long-tailed knowledge from one-shot episodic cues.&lt;/p>
&lt;h1 id="appendix">Appendix&lt;/h1>
&lt;p>If you’re interested, I did a similar recap for ICML 2023 &lt;a href="../icml-2023-recap/">here&lt;/a>.&lt;/p></description></item><item><title>ICML 2023 Recap</title><link>http://rkabra.com/post/icml-2023-recap/</link><pubDate>Mon, 31 Jul 2023 00:00:00 +0000</pubDate><guid>http://rkabra.com/post/icml-2023-recap/</guid><description>&lt;h2 id="vision">Vision&lt;/h2>
&lt;h3 id="vision-transformers">Vision Transformers&lt;/h3>
&lt;ul>
&lt;li>
&lt;p>VIT-22B models have much better alignment with human visual perception: 87% shape bias versus 20-30% in prior models. Prior models were much more texture-biased.
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/dehghani_hudc7af41ffc108184cae762807ed083de_2214572_330e5936243166137d144fd9fd6c2d29.webp 400w,
/post/icml-2023-recap/posters/dehghani_hudc7af41ffc108184cae762807ed083de_2214572_65ca39b47940763c57accc10046606c5.webp 760w,
/post/icml-2023-recap/posters/dehghani_hudc7af41ffc108184cae762807ed083de_2214572_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/dehghani_hudc7af41ffc108184cae762807ed083de_2214572_330e5936243166137d144fd9fd6c2d29.webp"
width="573"
height="760"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>A hierarchical VIT i.e. non-uniform feature size through the depth of the network. Also removes unnecessary bells and whistles from prior work by learning those biases instead.
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/ryali_hu36213a861e41b9751f40e65098f79f10_2826680_7e1866ed98df499e4cf0bcb54082cb6e.webp 400w,
/post/icml-2023-recap/posters/ryali_hu36213a861e41b9751f40e65098f79f10_2826680_1a62f06dad110bda05a1de7c02d68896.webp 760w,
/post/icml-2023-recap/posters/ryali_hu36213a861e41b9751f40e65098f79f10_2826680_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/ryali_hu36213a861e41b9751f40e65098f79f10_2826680_7e1866ed98df499e4cf0bcb54082cb6e.webp"
width="760"
height="467"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>VIT with global attention interspersed with regular attention
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/hatamizadeh_hu0580d66a5a78a7794ec37d9c4b3b565d_2374484_261c5bc5f3ceec546be1b2b8a3614e83.webp 400w,
/post/icml-2023-recap/posters/hatamizadeh_hu0580d66a5a78a7794ec37d9c4b3b565d_2374484_a283845de6cd18d24c2a0d92ead0aa92.webp 760w,
/post/icml-2023-recap/posters/hatamizadeh_hu0580d66a5a78a7794ec37d9c4b3b565d_2374484_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/hatamizadeh_hu0580d66a5a78a7794ec37d9c4b3b565d_2374484_261c5bc5f3ceec546be1b2b8a3614e83.webp"
width="760"
height="571"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;/ul>
&lt;h3 id="2d">2D&lt;/h3>
&lt;ul>
&lt;li>
&lt;p>Use both text and vision to improve classification of novel classes.
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/kaul_hua163f89b40cccd29bb7e98cd329e7d26_2614143_e67514b0b5e1faf7c4c6f6876f5bbb86.webp 400w,
/post/icml-2023-recap/posters/kaul_hua163f89b40cccd29bb7e98cd329e7d26_2614143_8cd620d4230cb31f71c8f03a54e30dcd.webp 760w,
/post/icml-2023-recap/posters/kaul_hua163f89b40cccd29bb7e98cd329e7d26_2614143_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/kaul_hua163f89b40cccd29bb7e98cd329e7d26_2614143_e67514b0b5e1faf7c4c6f6876f5bbb86.webp"
width="760"
height="488"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Learning a displacement field to learn the correspondence between photos and sketches
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/lu_x_huf5b01844fcbb97fac8af942f386e62b9_2210912_6acefce23c8ebdd6199ab2fba29be202.webp 400w,
/post/icml-2023-recap/posters/lu_x_huf5b01844fcbb97fac8af942f386e62b9_2210912_49d8b24997a5354f8f7966610c7dc086.webp 760w,
/post/icml-2023-recap/posters/lu_x_huf5b01844fcbb97fac8af942f386e62b9_2210912_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/lu_x_huf5b01844fcbb97fac8af942f386e62b9_2210912_6acefce23c8ebdd6199ab2fba29be202.webp"
width="760"
height="400"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Interpretable subspaces in image representations extracted using CLIP
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/kalibhat_hu516c825a9b4dec4b063137726e6df4ce_3533602_071c603a80aea2001f107eff42cc5ad4.webp 400w,
/post/icml-2023-recap/posters/kalibhat_hu516c825a9b4dec4b063137726e6df4ce_3533602_ffc4f7f828103be5e169286e2927fa13.webp 760w,
/post/icml-2023-recap/posters/kalibhat_hu516c825a9b4dec4b063137726e6df4ce_3533602_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/kalibhat_hu516c825a9b4dec4b063137726e6df4ce_3533602_071c603a80aea2001f107eff42cc5ad4.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Measuring &lt;em>compositionality&lt;/em> and &lt;em>invertibility&lt;/em> for object-centric representations
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/brady_hu3068ca9d7339a7b83bb7434cf896c6ba_3090171_4e101c2377b53fcf4f91787cfe3d393f.webp 400w,
/post/icml-2023-recap/posters/brady_hu3068ca9d7339a7b83bb7434cf896c6ba_3090171_c336ff5a8ff9fc6f49c011c91522d431.webp 760w,
/post/icml-2023-recap/posters/brady_hu3068ca9d7339a7b83bb7434cf896c6ba_3090171_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/brady_hu3068ca9d7339a7b83bb7434cf896c6ba_3090171_4e101c2377b53fcf4f91787cfe3d393f.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Multi-view self-supervised learning analyzed using Mutual Information.
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/galvez_hu1a2af76e6e573ec28d5ac585b8e7f687_3703118_8151a904688d463104c354ba201422ed.webp 400w,
/post/icml-2023-recap/posters/galvez_hu1a2af76e6e573ec28d5ac585b8e7f687_3703118_de4523cca446cff64e354d4cc5baf63a.webp 760w,
/post/icml-2023-recap/posters/galvez_hu1a2af76e6e573ec28d5ac585b8e7f687_3703118_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/galvez_hu1a2af76e6e573ec28d5ac585b8e7f687_3703118_8151a904688d463104c354ba201422ed.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Class collapse and feature suppression during contrastive learning
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/xue_y_hub64bb1b562f5121fbcfaa10b7658610c_1269941_3d15ea655c237f8916dee1dbbc9bfd6f.webp 400w,
/post/icml-2023-recap/posters/xue_y_hub64bb1b562f5121fbcfaa10b7658610c_1269941_1a758a5f5351dcdbfb619030ac56a38f.webp 760w,
/post/icml-2023-recap/posters/xue_y_hub64bb1b562f5121fbcfaa10b7658610c_1269941_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/xue_y_hub64bb1b562f5121fbcfaa10b7658610c_1269941_3d15ea655c237f8916dee1dbbc9bfd6f.webp"
width="760"
height="409"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>The latest on hyperbolic representations.
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/desai_huaaeada71cdf32c27baad3f4d7f9127b6_3509140_9c726cd6c40e6c0941153eefdefad971.webp 400w,
/post/icml-2023-recap/posters/desai_huaaeada71cdf32c27baad3f4d7f9127b6_3509140_1c2b062213cab1d27ada0c307af51530.webp 760w,
/post/icml-2023-recap/posters/desai_huaaeada71cdf32c27baad3f4d7f9127b6_3509140_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/desai_huaaeada71cdf32c27baad3f4d7f9127b6_3509140_9c726cd6c40e6c0941153eefdefad971.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;/ul>
&lt;h3 id="3d">3D&lt;/h3>
&lt;ul>
&lt;li>
&lt;p>Spherical CNNs (rotation equivariant) scaled to 5e6 convolutions and 1e7-1e9 feature maps
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/esteves_hu787f49937e2d18381b46e8599038ec42_2745503_a7808e52b66b0f03887783b95e5528f5.webp 400w,
/post/icml-2023-recap/posters/esteves_hu787f49937e2d18381b46e8599038ec42_2745503_123b435bb80cdbbcb022bf595a8d4582.webp 760w,
/post/icml-2023-recap/posters/esteves_hu787f49937e2d18381b46e8599038ec42_2745503_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/esteves_hu787f49937e2d18381b46e8599038ec42_2745503_a7808e52b66b0f03887783b95e5528f5.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Object pose canonicalization measured for &lt;em>stability&lt;/em> and &lt;em>consistency&lt;/em>. They also train on multiple object classes.
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/kim_s_hu36213a861e41b9751f40e65098f79f10_2616288_d833867bd8abd05192d0b6d5bc608bdb.webp 400w,
/post/icml-2023-recap/posters/kim_s_hu36213a861e41b9751f40e65098f79f10_2616288_848e30bc50c5cbb3b65b7bac716063b6.webp 760w,
/post/icml-2023-recap/posters/kim_s_hu36213a861e41b9751f40e65098f79f10_2616288_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/kim_s_hu36213a861e41b9751f40e65098f79f10_2616288_d833867bd8abd05192d0b6d5bc608bdb.webp"
width="760"
height="434"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Signed distance functions learnt “provably.”
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/bethune_huf5b01844fcbb97fac8af942f386e62b9_1974637_243120c316d765d842b38c67739c4eac.webp 400w,
/post/icml-2023-recap/posters/bethune_huf5b01844fcbb97fac8af942f386e62b9_1974637_19d62035ad1f612210809ff081e5c686.webp 760w,
/post/icml-2023-recap/posters/bethune_huf5b01844fcbb97fac8af942f386e62b9_1974637_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/bethune_huf5b01844fcbb97fac8af942f386e62b9_1974637_243120c316d765d842b38c67739c4eac.webp"
width="760"
height="383"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;/ul>
&lt;h3 id="video">Video&lt;/h3>
&lt;ul>
&lt;li>
&lt;p>Keypoint learning in videos.
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/younes_hu12ab96f88193939af5457d4a61c51cb2_2977724_6de2cd824f9f3035ff94efeb42415f84.webp 400w,
/post/icml-2023-recap/posters/younes_hu12ab96f88193939af5457d4a61c51cb2_2977724_2dc29f81aee79b052617160edd0ed510.webp 760w,
/post/icml-2023-recap/posters/younes_hu12ab96f88193939af5457d4a61c51cb2_2977724_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/younes_hu12ab96f88193939af5457d4a61c51cb2_2977724_6de2cd824f9f3035ff94efeb42415f84.webp"
width="760"
height="546"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Efficient episodic recall (aka “video search”).
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/ramakrishnan_hu85e83a101b12c409d568d931f67985a7_3085809_a48de891f935a3a1a7226f4724276aa5.webp 400w,
/post/icml-2023-recap/posters/ramakrishnan_hu85e83a101b12c409d568d931f67985a7_3085809_4717aa95267070c0f9db9ff36a820f7e.webp 760w,
/post/icml-2023-recap/posters/ramakrishnan_hu85e83a101b12c409d568d931f67985a7_3085809_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/ramakrishnan_hu85e83a101b12c409d568d931f67985a7_3085809_a48de891f935a3a1a7226f4724276aa5.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;/ul>
&lt;h2 id="generative-models">Generative Models&lt;/h2>
&lt;ul>
&lt;li>
&lt;p>Electrostatics-based generative model with better FID numbers than diffusion
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/xu_y_hu06bc956626cface11985b13c9db38fb6_2216444_2f2163faf0a21e05e3ac56ece48e3c82.webp 400w,
/post/icml-2023-recap/posters/xu_y_hu06bc956626cface11985b13c9db38fb6_2216444_dc378deb04406b53ce6141d8e744f5c7.webp 760w,
/post/icml-2023-recap/posters/xu_y_hu06bc956626cface11985b13c9db38fb6_2216444_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/xu_y_hu06bc956626cface11985b13c9db38fb6_2216444_2f2163faf0a21e05e3ac56ece48e3c82.webp"
width="760"
height="418"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Animated 3D models without any additional dataset.
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/singer_hu879bc88e65bfcdca17fd8a86ed988c2b_2985429_1f3e34145e428d56eeaf32bd2b196ff0.webp 400w,
/post/icml-2023-recap/posters/singer_hu879bc88e65bfcdca17fd8a86ed988c2b_2985429_19791ca26956d3b681afec0280c9fdb6.webp 760w,
/post/icml-2023-recap/posters/singer_hu879bc88e65bfcdca17fd8a86ed988c2b_2985429_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/singer_hu879bc88e65bfcdca17fd8a86ed988c2b_2985429_1f3e34145e428d56eeaf32bd2b196ff0.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Diffusion without upsamplers. Harder to train and inefficient.
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/hoogeboom_hu24f87e6d104f60e3670a3e0413076241_2321432_5defe0c2417344bfd2283292f38a0237.webp 400w,
/post/icml-2023-recap/posters/hoogeboom_hu24f87e6d104f60e3670a3e0413076241_2321432_5ea2974cb569c55ad3ce9fc29a4e5bec.webp 760w,
/post/icml-2023-recap/posters/hoogeboom_hu24f87e6d104f60e3670a3e0413076241_2321432_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/hoogeboom_hu24f87e6d104f60e3670a3e0413076241_2321432_5defe0c2417344bfd2283292f38a0237.webp"
width="760"
height="661"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Consistency models: diffusion without multi-step denoising.
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/song_y_hu06bc956626cface11985b13c9db38fb6_1433703_cc6dd2a8a8d9758624714547038c8688.webp 400w,
/post/icml-2023-recap/posters/song_y_hu06bc956626cface11985b13c9db38fb6_1433703_9a6b038f316f277b76724075b7d719b9.webp 760w,
/post/icml-2023-recap/posters/song_y_hu06bc956626cface11985b13c9db38fb6_1433703_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/song_y_hu06bc956626cface11985b13c9db38fb6_1433703_cc6dd2a8a8d9758624714547038c8688.webp"
width="688"
height="760"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Diffusion models evaluated on one-shot drawing task.
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/boutin_hu39249f2c6b3c7513d0b89f535064e539_2444441_3dec451cf6ab93ac7fa715abdc6a5fba.webp 400w,
/post/icml-2023-recap/posters/boutin_hu39249f2c6b3c7513d0b89f535064e539_2444441_cd9dfc142c5fd91ebe2544057ac5044b.webp 760w,
/post/icml-2023-recap/posters/boutin_hu39249f2c6b3c7513d0b89f535064e539_2444441_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/boutin_hu39249f2c6b3c7513d0b89f535064e539_2444441_3dec451cf6ab93ac7fa715abdc6a5fba.webp"
width="760"
height="404"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>NeRF from fewer samples using geometric invariances.
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/kwak_hu6d7048cf4f07c816b2542c0c70a2c6a6_2575135_9e351f7952fb158f90336db6ed0b1954.webp 400w,
/post/icml-2023-recap/posters/kwak_hu6d7048cf4f07c816b2542c0c70a2c6a6_2575135_32e23935c33b92dec5d67c4ffb35b923.webp 760w,
/post/icml-2023-recap/posters/kwak_hu6d7048cf4f07c816b2542c0c70a2c6a6_2575135_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/kwak_hu6d7048cf4f07c816b2542c0c70a2c6a6_2575135_9e351f7952fb158f90336db6ed0b1954.webp"
width="688"
height="760"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;/ul>
&lt;h2 id="world-modelsrl">World Models/RL&lt;/h2>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/wu_p_hu9bdd3fcc3a121ce18340deb727406f87_2188354_19dc5d3b488dbdc03a4f7131fbff9992.webp 400w,
/post/icml-2023-recap/posters/wu_p_hu9bdd3fcc3a121ce18340deb727406f87_2188354_f959b400e99cea8edc5d25b0f04d4e5b.webp 760w,
/post/icml-2023-recap/posters/wu_p_hu9bdd3fcc3a121ce18340deb727406f87_2188354_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/wu_p_hu9bdd3fcc3a121ce18340deb727406f87_2188354_19dc5d3b488dbdc03a4f7131fbff9992.webp"
width="760"
height="584"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/freed_hu1408dadf568e761a6cd23bdfcbd11895_3208000_4ca72e213e6c31bd98040d8b6775d837.webp 400w,
/post/icml-2023-recap/posters/freed_hu1408dadf568e761a6cd23bdfcbd11895_3208000_34a7db919866d60da526c9e5bc70fa4c.webp 760w,
/post/icml-2023-recap/posters/freed_hu1408dadf568e761a6cd23bdfcbd11895_3208000_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/freed_hu1408dadf568e761a6cd23bdfcbd11895_3208000_4ca72e213e6c31bd98040d8b6775d837.webp"
width="573"
height="760"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/seo_y_huf5b01844fcbb97fac8af942f386e62b9_2310125_fd8bb9027d4a59ec1946f5a7348a1576.webp 400w,
/post/icml-2023-recap/posters/seo_y_huf5b01844fcbb97fac8af942f386e62b9_2310125_80c3da0107304c035d59334fa5441b83.webp 760w,
/post/icml-2023-recap/posters/seo_y_huf5b01844fcbb97fac8af942f386e62b9_2310125_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/seo_y_huf5b01844fcbb97fac8af942f386e62b9_2310125_fd8bb9027d4a59ec1946f5a7348a1576.webp"
width="760"
height="380"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/nottingham_hub9d239cb4417fa59651302f4bf86a934_3697280_085420b7355b5263fc0f0c5ce95e3218.webp 400w,
/post/icml-2023-recap/posters/nottingham_hub9d239cb4417fa59651302f4bf86a934_3697280_5f0e62bb5c313ce97220ad7a4655baf0.webp 760w,
/post/icml-2023-recap/posters/nottingham_hub9d239cb4417fa59651302f4bf86a934_3697280_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/nottingham_hub9d239cb4417fa59651302f4bf86a934_3697280_085420b7355b5263fc0f0c5ce95e3218.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/gmelin_hu5723f632d9330704783f8a404f0f204e_2173729_2acacf0a1b96f49f7559712e09e9c894.webp 400w,
/post/icml-2023-recap/posters/gmelin_hu5723f632d9330704783f8a404f0f204e_2173729_3f6d45047940415d212aa48e83113d99.webp 760w,
/post/icml-2023-recap/posters/gmelin_hu5723f632d9330704783f8a404f0f204e_2173729_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/gmelin_hu5723f632d9330704783f8a404f0f204e_2173729_2acacf0a1b96f49f7559712e09e9c894.webp"
width="738"
height="760"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/ghosh_hu9fb790e0a18b972bbd43b93125f8ff2d_1757020_f01280c88c1146687408d094246bf7fc.webp 400w,
/post/icml-2023-recap/posters/ghosh_hu9fb790e0a18b972bbd43b93125f8ff2d_1757020_d0631fe0da72ea2bd7efd829760ea929.webp 760w,
/post/icml-2023-recap/posters/ghosh_hu9fb790e0a18b972bbd43b93125f8ff2d_1757020_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/ghosh_hu9fb790e0a18b972bbd43b93125f8ff2d_1757020_f01280c88c1146687408d094246bf7fc.webp"
width="708"
height="760"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;h2 id="transformers">Transformers&lt;/h2>
&lt;ul>
&lt;li>
&lt;p>Beautiful work showing transformers have a “lower-degree” bias toward polynomial terms of lower degree, which is somewhat counterintuitive given their pairwise attention mechanism.
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/abbe_hu833ba20c9e983529dcb24b105bafe0b3_2796995_8165533d793a366f25e1d5a320386eee.webp 400w,
/post/icml-2023-recap/posters/abbe_hu833ba20c9e983529dcb24b105bafe0b3_2796995_aaad139e79ca73df961c035050d8c6d6.webp 760w,
/post/icml-2023-recap/posters/abbe_hu833ba20c9e983529dcb24b105bafe0b3_2796995_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/abbe_hu833ba20c9e983529dcb24b105bafe0b3_2796995_8165533d793a366f25e1d5a320386eee.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Improving the focal loss by taking into account the second highest predicted logit, rather than naively maximizing entropy.
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/tao_l_hu39249f2c6b3c7513d0b89f535064e539_1973676_28c37c454ca205d305c7e9c4e8ff1362.webp 400w,
/post/icml-2023-recap/posters/tao_l_hu39249f2c6b3c7513d0b89f535064e539_1973676_fb950ebb64d220fe9606e6e3a499fd91.webp 760w,
/post/icml-2023-recap/posters/tao_l_hu39249f2c6b3c7513d0b89f535064e539_1973676_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/tao_l_hu39249f2c6b3c7513d0b89f535064e539_1973676_28c37c454ca205d305c7e9c4e8ff1362.webp"
width="760"
height="450"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Do early layers generalize while later layers memorize? Apparently not–memorization can be localized to a small number of neurons dispersed across layers.
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/maini_hub64bb1b562f5121fbcfaa10b7658610c_1844564_d0acb00cad558edaa1e3149f781ceb11.webp 400w,
/post/icml-2023-recap/posters/maini_hub64bb1b562f5121fbcfaa10b7658610c_1844564_e02743cb8bfaa698a02db3bb750d1c1f.webp 760w,
/post/icml-2023-recap/posters/maini_hub64bb1b562f5121fbcfaa10b7658610c_1844564_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/maini_hub64bb1b562f5121fbcfaa10b7658610c_1844564_d0acb00cad558edaa1e3149f781ceb11.webp"
width="760"
height="371"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Characterizing training trajectories of different representation learning tasks
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/ramesh_hu0fd1af34dd49a514d4c76ed0903da650_2504861_590c8314a509b273ac9ec52285baaef8.webp 400w,
/post/icml-2023-recap/posters/ramesh_hu0fd1af34dd49a514d4c76ed0903da650_2504861_5cc4944ef183a883887238ba041acb87.webp 760w,
/post/icml-2023-recap/posters/ramesh_hu0fd1af34dd49a514d4c76ed0903da650_2504861_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/ramesh_hu0fd1af34dd49a514d4c76ed0903da650_2504861_590c8314a509b273ac9ec52285baaef8.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Is local flatness desirable for generalization? Not necessarily. There are more promising indicators such as SGD-based disagreement on unlabelled data.
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/andriushchenko_hu5cb92f6f09d1e1b192bdf666a3458368_2596402_9c6a3f7a7c4f2342616c87a12eeb547e.webp 400w,
/post/icml-2023-recap/posters/andriushchenko_hu5cb92f6f09d1e1b192bdf666a3458368_2596402_46ce5c1586218c5a77192e857a405411.webp 760w,
/post/icml-2023-recap/posters/andriushchenko_hu5cb92f6f09d1e1b192bdf666a3458368_2596402_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/andriushchenko_hu5cb92f6f09d1e1b192bdf666a3458368_2596402_9c6a3f7a7c4f2342616c87a12eeb547e.webp"
width="760"
height="653"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Category-theory view of disentanglement
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/zhang_y_hue35f580e0c09b2b5b48aff41532292af_2522046_b0d2abb4e613343a99a4a620b0a107c5.webp 400w,
/post/icml-2023-recap/posters/zhang_y_hue35f580e0c09b2b5b48aff41532292af_2522046_4370e5f6f3d675cde90e3fad52f2a0ce.webp 760w,
/post/icml-2023-recap/posters/zhang_y_hue35f580e0c09b2b5b48aff41532292af_2522046_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/zhang_y_hue35f580e0c09b2b5b48aff41532292af_2522046_b0d2abb4e613343a99a4a620b0a107c5.webp"
width="760"
height="594"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Using category theory to show that foundation models cannot be used for everything, but CLIP-like algorithms do have “creativity”
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/yuan_y_hu24e1143b5ff7d99db7e37faaca27ad44_3912948_95eabb2cac5ae2437c6031716f99e7ad.webp 400w,
/post/icml-2023-recap/posters/yuan_y_hu24e1143b5ff7d99db7e37faaca27ad44_3912948_f45a9c4b3c4cb42721b9c2b9993eb2ca.webp 760w,
/post/icml-2023-recap/posters/yuan_y_hu24e1143b5ff7d99db7e37faaca27ad44_3912948_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/yuan_y_hu24e1143b5ff7d99db7e37faaca27ad44_3912948_95eabb2cac5ae2437c6031716f99e7ad.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;/ul>
&lt;h2 id="novel-architectures">Novel architectures&lt;/h2>
&lt;ul>
&lt;li>
&lt;p>Super simple long convolutions
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/fu_d_hua7c2110322b8a8941ac069d58cc85390_3590383_1343ddc07cd817ab8fac058e62f25ea0.webp 400w,
/post/icml-2023-recap/posters/fu_d_hua7c2110322b8a8941ac069d58cc85390_3590383_db36af6d094945bbcb525cbde1951ab2.webp 760w,
/post/icml-2023-recap/posters/fu_d_hua7c2110322b8a8941ac069d58cc85390_3590383_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/fu_d_hua7c2110322b8a8941ac069d58cc85390_3590383_1343ddc07cd817ab8fac058e62f25ea0.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Differentiable “if blocks”
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/faber_hu4ce546542d1668c41a9b667ed6cb6303_1867144_b7b395b8c8913b19098aafca8c2359af.webp 400w,
/post/icml-2023-recap/posters/faber_hu4ce546542d1668c41a9b667ed6cb6303_1867144_ab62624defb84e3cfe138f1d31c0c55a.webp 760w,
/post/icml-2023-recap/posters/faber_hu4ce546542d1668c41a9b667ed6cb6303_1867144_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/faber_hu4ce546542d1668c41a9b667ed6cb6303_1867144_b7b395b8c8913b19098aafca8c2359af.webp"
width="760"
height="731"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Differentiable tree operations
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/soulos_hu5bc6605ae5e323894cb44573bfc0c07f_3170930_42a39fd494c4af732c9f1ff4f6852202.webp 400w,
/post/icml-2023-recap/posters/soulos_hu5bc6605ae5e323894cb44573bfc0c07f_3170930_237d014528559276e490f55b6700996b.webp 760w,
/post/icml-2023-recap/posters/soulos_hu5bc6605ae5e323894cb44573bfc0c07f_3170930_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/soulos_hu5bc6605ae5e323894cb44573bfc0c07f_3170930_42a39fd494c4af732c9f1ff4f6852202.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Continuous spatiotemporal transformers
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/fonseca_hu734b17d0b7b67fbec45f3cfbed41927b_3880462_832d151d950358e0d3a146c617699ca1.webp 400w,
/post/icml-2023-recap/posters/fonseca_hu734b17d0b7b67fbec45f3cfbed41927b_3880462_9b92bc0aa2c7b84184636902341f95cd.webp 760w,
/post/icml-2023-recap/posters/fonseca_hu734b17d0b7b67fbec45f3cfbed41927b_3880462_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/fonseca_hu734b17d0b7b67fbec45f3cfbed41927b_3880462_832d151d950358e0d3a146c617699ca1.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;/ul>
&lt;h2 id="graphs">Graphs&lt;/h2>
&lt;ul>
&lt;li>
&lt;p>Compositionality via learnt pooling from a multi-view graph to a latent graph
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/liu_t_hu2d56bc0107e205eeacd443b416f8db4f_3043824_09cbd70756b064c0139bfe65f8a9f54e.webp 400w,
/post/icml-2023-recap/posters/liu_t_hu2d56bc0107e205eeacd443b416f8db4f_3043824_3d7aee4f8ee60b48aef5eb8c64a294a8.webp 760w,
/post/icml-2023-recap/posters/liu_t_hu2d56bc0107e205eeacd443b416f8db4f_3043824_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/liu_t_hu2d56bc0107e205eeacd443b416f8db4f_3043824_09cbd70756b064c0139bfe65f8a9f54e.webp"
width="760"
height="547"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Positional encodings to take advantage of edge directions
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/geisler_hu39249f2c6b3c7513d0b89f535064e539_1880315_4d7d3ceb805ea7b7579ba12ab24e5c51.webp 400w,
/post/icml-2023-recap/posters/geisler_hu39249f2c6b3c7513d0b89f535064e539_1880315_57c92f375e02e460ca384cb9dc79758c.webp 760w,
/post/icml-2023-recap/posters/geisler_hu39249f2c6b3c7513d0b89f535064e539_1880315_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/geisler_hu39249f2c6b3c7513d0b89f535064e539_1880315_4d7d3ceb805ea7b7579ba12ab24e5c51.webp"
width="760"
height="649"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;/ul>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/ma_l_hu6fc5e2a70522bc7ab83b183ade6beed9_2277322_96755b073e4c84a528535875a8c2b2fa.webp 400w,
/post/icml-2023-recap/posters/ma_l_hu6fc5e2a70522bc7ab83b183ade6beed9_2277322_6ab549c131b74d12465968a752f7a290.webp 760w,
/post/icml-2023-recap/posters/ma_l_hu6fc5e2a70522bc7ab83b183ade6beed9_2277322_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/ma_l_hu6fc5e2a70522bc7ab83b183ade6beed9_2277322_96755b073e4c84a528535875a8c2b2fa.webp"
width="760"
height="393"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;h2 id="adversarial-attacks">Adversarial attacks&lt;/h2>
&lt;ul>
&lt;li>Independent component analysis to design an attack on federated learning
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/kariyappa_hu36213a861e41b9751f40e65098f79f10_2039771_7b9fb736054004b9a99561d335d7baea.webp 400w,
/post/icml-2023-recap/posters/kariyappa_hu36213a861e41b9751f40e65098f79f10_2039771_c1faa1994c873a07d4adba832a3d667f.webp 760w,
/post/icml-2023-recap/posters/kariyappa_hu36213a861e41b9751f40e65098f79f10_2039771_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/kariyappa_hu36213a861e41b9751f40e65098f79f10_2039771_7b9fb736054004b9a99561d335d7baea.webp"
width="760"
height="583"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/li>
&lt;/ul>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/khaddaj_hu3b607321febd34ec7bcae4851ca7e03c_1778496_9e4bd8d8337284f8a2829d46935af2ba.webp 400w,
/post/icml-2023-recap/posters/khaddaj_hu3b607321febd34ec7bcae4851ca7e03c_1778496_dcfbf1f7f7a7a73d89b8b33d2fa012ab.webp 760w,
/post/icml-2023-recap/posters/khaddaj_hu3b607321febd34ec7bcae4851ca7e03c_1778496_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/khaddaj_hu3b607321febd34ec7bcae4851ca7e03c_1778496_9e4bd8d8337284f8a2829d46935af2ba.webp"
width="760"
height="534"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;h2 id="curiosities">Curiosities&lt;/h2>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/becker_hu36213a861e41b9751f40e65098f79f10_1611576_3b38aab8968b85c3c9da8b92db96d88f.webp 400w,
/post/icml-2023-recap/posters/becker_hu36213a861e41b9751f40e65098f79f10_1611576_caa05d8207b8ea3625dbc26b594515b9.webp 760w,
/post/icml-2023-recap/posters/becker_hu36213a861e41b9751f40e65098f79f10_1611576_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/becker_hu36213a861e41b9751f40e65098f79f10_1611576_3b38aab8968b85c3c9da8b92db96d88f.webp"
width="760"
height="554"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;ul>
&lt;li>Implicit neural representations (using spatial coordinates C or environmental features E or both) to predict presence of wildlife species.
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/cole_hu6eb93efe4fc0ce9ebf77e9a7da59aa5d_3362794_9ea39230e18d0c454d0e44ab5a26e987.webp 400w,
/post/icml-2023-recap/posters/cole_hu6eb93efe4fc0ce9ebf77e9a7da59aa5d_3362794_e75b40e7c29149f470bba942ab86b9cf.webp 760w,
/post/icml-2023-recap/posters/cole_hu6eb93efe4fc0ce9ebf77e9a7da59aa5d_3362794_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/cole_hu6eb93efe4fc0ce9ebf77e9a7da59aa5d_3362794_9ea39230e18d0c454d0e44ab5a26e987.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/li>
&lt;/ul>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/huang_l_hu98d32e1306c1fae9c61ab31387eb590f_3836569_1602da20b54fc72a417b7162e467cd89.webp 400w,
/post/icml-2023-recap/posters/huang_l_hu98d32e1306c1fae9c61ab31387eb590f_3836569_2c0bcc695641f94062bbe7b6af834065.webp 760w,
/post/icml-2023-recap/posters/huang_l_hu98d32e1306c1fae9c61ab31387eb590f_3836569_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/huang_l_hu98d32e1306c1fae9c61ab31387eb590f_3836569_1602da20b54fc72a417b7162e467cd89.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;ul>
&lt;li>
&lt;p>ML on Mars for source separation to detect marsquakes!
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/siahkoohi_hu06bc956626cface11985b13c9db38fb6_2486262_f531c58cb5ba975cc1bc22379d1cecf6.webp 400w,
/post/icml-2023-recap/posters/siahkoohi_hu06bc956626cface11985b13c9db38fb6_2486262_ef5970b6368f45bb4b2371034949121a.webp 760w,
/post/icml-2023-recap/posters/siahkoohi_hu06bc956626cface11985b13c9db38fb6_2486262_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/siahkoohi_hu06bc956626cface11985b13c9db38fb6_2486262_f531c58cb5ba975cc1bc22379d1cecf6.webp"
width="760"
height="424"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Template + score/filter prompts for a dataset without access to labels.
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/allingham_hu9bdd3fcc3a121ce18340deb727406f87_1984151_2f2cf1872c49214481aa8f589d258c57.webp 400w,
/post/icml-2023-recap/posters/allingham_hu9bdd3fcc3a121ce18340deb727406f87_1984151_b4d116ccb19607d4100004ff889bb50b.webp 760w,
/post/icml-2023-recap/posters/allingham_hu9bdd3fcc3a121ce18340deb727406f87_1984151_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/allingham_hu9bdd3fcc3a121ce18340deb727406f87_1984151_2f2cf1872c49214481aa8f589d258c57.webp"
width="760"
height="542"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>A simple initialization trick for VIT-Tiny
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/trockman_hu63a7579c67908f95fb2c2b7b8dfc79b4_2370189_39b3813128678f96a2c890776053cdd3.webp 400w,
/post/icml-2023-recap/posters/trockman_hu63a7579c67908f95fb2c2b7b8dfc79b4_2370189_ce7c9e202bc6535f50482b337200f849.webp 760w,
/post/icml-2023-recap/posters/trockman_hu63a7579c67908f95fb2c2b7b8dfc79b4_2370189_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/trockman_hu63a7579c67908f95fb2c2b7b8dfc79b4_2370189_39b3813128678f96a2c890776053cdd3.webp"
width="545"
height="760"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;li>
&lt;p>How to fine-tune ML models in an “open-source” fashion: fine-tune in parallel and then merge
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="alt_text" srcset="
/post/icml-2023-recap/posters/rame_hu91333ccc833615fa7167921a58bb6b44_2273397_565e4f6c3653be05bcbd0625c7a6af27.webp 400w,
/post/icml-2023-recap/posters/rame_hu91333ccc833615fa7167921a58bb6b44_2273397_8544531d4d2feb32868bf8044fe4fe12.webp 760w,
/post/icml-2023-recap/posters/rame_hu91333ccc833615fa7167921a58bb6b44_2273397_1200x1200_fit_q75_h2_lanczos.webp 1200w"
src="http://rkabra.com/post/icml-2023-recap/posters/rame_hu91333ccc833615fa7167921a58bb6b44_2273397_565e4f6c3653be05bcbd0625c7a6af27.webp"
width="760"
height="573"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;/ul></description></item></channel></rss>