Stony Brook University | Department of Computer Science
Fall 2026
| Instructor |
|
| Office Hours | TBA |
| Teaching Assistant | TBA |
| TA Office Hours | TBA |
| Class Meetings | Tuesday/Thursday - 12:30pm-1:50pm - LGT ENGR LAB 152, West Campus |
| Meeting Dates | August 24, 2026 - December 17, 2026 |
In this graduate-level special topics course, we will explore recent advancements in vision-language models (VLMs), focusing on their architectures, applications, and ongoing research. We will study VLMs applied to both images and videos, addressing tasks such as visual reasoning, classification by description, image and video captioning, text-to-image and text-to-video generation, and visual question answering. We will analyze current challenges, including representation learning, domain shifts, and cultural biases, with special attention to compositionality and embodied AI. We will emphasize on the critical analysis of these techniques, with a focus on reading and discussion of recent work. This course will primarily consist of student presentations on assigned conference and journal publications, and a semester-long project, aiming to equip students with a comprehensive understanding of the current state and future directions of VLMs.
We will use SOLAR for all course materials, including readings, assignments, and announcements.
This is a graduate-level course. Students are expected to have a solid mathematics background and strong programming skills. Students are also expected to have a strong foundation in machine learning and deep learning (e.g., CSE 512, CSE 527, CSE 538, or equivalent). Proficiency in Python and experience with a deep learning framework (e.g., PyTorch, TensorFlow) are required.
Upon successful completion of this course, students will be able to:
This is a graduate-level course centered around reading, presenting, and discussing research papers. A significant portion of this course is dedicated to a semester-long research project.
| Paper Presentations | 30% |
| Class Participation & Discussion | 10% |
| Reviewing an LLM-Generated Paper Review (4 total) | 10% |
| Semester Project | 50% |
Logistics: Students sign up in Week 2 to claim one paper from the approved pool. Upload a PDF of your slide deck to SOLAR 48 hours before your in-class talk, prepare a presentation followed by Q&A, and conclude with 2–3 discussion topics that tie the work to course themes (TBD - subject to change).
Students will use an LLM to generate a review of a paper, identify any inaccuracies or biases of the generated review, and provide personal/constructive feedback.
This schedule is tentative and may be adjusted. All papers will be linked from the course website. Students are expected to have read the assigned papers *before* the class session.
| Week | Date | Topic | Readings | Due Activities |
|---|---|---|---|---|
| 1 | 08/25 |
Early work: pre-Transformers: Visual representations Sequence-to-sequence models Neural image captioning VQA, multimodal fusion, and language shortcuts |
+ Krizhevsky et al. "ImageNet Classification with Deep Convolutional Neural Networks". NeurIPS, 2012. + He et al. "Deep Residual Learning for Image Recognition". CVPR, 2016. + Ren et al. "Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks". NeurIPS, 2015. |
|
| 08/27 |
+ Sutskever et al. "Sequence to Sequence Learning with Neural Networks". NeurIPS, 2014. + Bahdanau et al. "Neural Machine Translation by Jointly Learning to Align and Translate". ICLR, 2015. + Luong et al. "Effective Approaches to Attention-based Neural Machine Translation". EMNLP, 2015. |
|||
| 2 | 09/01 |
+ Vinyals et al. "Show and Tell: A Neural Image Caption Generator". CVPR, 2015. + Karpathy and Fei-Fei. "Deep Visual-Semantic Alignments for Generating Image Descriptions". CVPR, 2015. + Xu et al. "Show, Attend and Tell: Neural Image Caption Generation with Visual Attention". ICML, 2015. |
Find Your Project Team Members | |
| 09/03 |
+ Antol et al. "VQA: Visual Question Answering". ICCV, 2015. + Fukui et al. "Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding". EMNLP, 2016. + Goyal et al. "Making the V in VQA Matter". CVPR, 2017. + Anderson et al. "Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering". CVPR, 2018. |
|||
| 3 | 09/08 | Transformers, ViTs, LLMs and VLMs |
+ Vaswani et al. "Attention Is All You Need". NeurIPS, 2017. + Devlin et al. "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding". NAACL, 2019. + Brown et al. "Language Models are Few-Shot Learners". NeurIPS, 2020. |
Select Paper to Present |
| 09/10 |
+ Dosovitskiy et al. "An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale". ICLR, 2021. + Touvron et al. "Training Data-Efficient Image Transformers and Distillation Through Attention". ICML, 2021. + Caron et al. "Emerging Properties in Self-Supervised Vision Transformers". ICCV, 2021. |
|||
| 4 | 09/15 |
+ Radford et al. "Learning Transferable Visual Models From Natural Language Supervision". ICML, 2021. + Jia et al. "Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision". ICML, 2021. + Li et al. "Align Before Fuse: Vision and Language Representation Learning with Momentum Distillation". NeurIPS, 2021. |
||
| 09/17 |
+ Tsimpoukelli et al. "Multimodal Few-Shot Learning with Frozen Language Models". NeurIPS, 2021. + Alayrac et al. "Flamingo: A Visual Language Model for Few-Shot Learning". NeurIPS, 2022. + Li et al. "BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models". ICML, 2023. + Liu et al. "Visual Instruction Tuning". NeurIPS, 2023. |
|||
| 5 | 09/22 | Compositionality & Visual Reasoning | + Thrush, Tristan, et al. "Winoground: Probing vision and language models for visio-linguistic compositionality." CVPR, 2022. | Project Proposal Due |
| 09/24 |
+ Wu, Wenshan, et al. "Mind’s Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models." NeurIPS, 2024. + Li, Chengzu, et al. "Imagine while Reasoning in Space: Multimodal Visualization-of-Thought." ICML, 2025. |
|||
| 6 | 09/29 | Domain Shift & Generalization | + Wortsman, Mitchell, et al. "Robust fine-tuning of zero-shot models." CVPR, 2022. | |
| 10/01 | + Lafon, Marc, et al. "GalLoP: Learning global and local prompts for vision-language models." ECCV, 2024. | |||
| 7 | 10/06 | + Zhang, Yabin, et al. "Dual memory networks: A versatile adaptation approach for vision-language models." CVPR, 2024. | ||
| 10/08 | Cultural Bias & Fairness | + Nayak, Shravan, et al. "Benchmarking vision language models for cultural understanding." EMNLP, 2024. | ||
| 8 | 10/13 | Fall Break - No class | ||
| 10/15 | Cultural Bias & Fairness | + Lee, Tony, et al. "VHELM: A holistic evaluation of vision language models." NeurIPS, 2024. | ||
| 9 | 10/20 | Embodied AI & Agents | + Driess, Danny, et al. "Palm-e: An embodied multimodal language model." ICML, 2023. | |
| 10/22 | + Zitkovich, Brianna, et al. "Rt-2: Vision-language-action models transfer web knowledge to robotic control." Conference on Robot Learning. PMLR, 2023. | |||
| 10 | 10/27 | + Kim, Moo Jin, et al. "Openvla: An open-source vision-language-action model." CoRL, 2024. | Midterm Report Due | |
| 10/29 | + Wang, Weizhen, et al. "Embodied scene understanding for vision language models via metavqa." Proceedings of the Computer Vision and Pattern Recognition Conference. 2025. | |||
| 11 | 11/03 | VLM Reliability, Hallucination & Interpretability | + Guan, Tianrui, et al. "HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models." CVPR, 2024. | |
| 11/05 | + Zhang, Zhi, et al. "Cross-modal Information Flow in Multimodal Large Language Models." CVPR, 2025. | |||
| 12 | 11/10 | + Hou, Xuehe, et al. "VES-RFT: Rewarding Visual Evidence Sensitivity to Mitigate Hallucinations in Large Vision-Language Models." CVPR, 2026. | ||
| 11/12 | Long-Form Video Understanding | + Fu, Chaoyou, et al. "Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis." CVPR, 2025. | ||
| 13 | 11/17 | + Liu, Shuming, et al. "BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding." CVPR, 2025. | ||
| 11/19 | World Models | + Yang, Mengjiao, et al. "UniSim: Learning interactive real-world simulators." ICLR, 2024. | ||
| 14 | 11/24 | + Bruce, Jake, et al. "Genie: Generative interactive environments." Forty-first International Conference on Machine Learning. 2024. | ||
| 11/26 | Thanksgiving Break - No class | |||
| 15 | 12/01 | Efficiency & Scaling Strategies | + Rajbhandari, Samyam, et al. "Zero: Memory optimizations toward training trillion parameter models." SC20: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 2020. | |
| 12/03 | + Dao, Tri, et al. "FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness." NeurIPS, 2022. + Dao, Tri. "FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning." ICLR, 2024. |
|||
| Finals | 12/15 11:15am-1:45pm |
Final Exam Period | Final Project Presentations | Final Project Report & Code Due |
The semester project is a core component of this course. It provides an opportunity to explore a research topic in depth. Projects can be done individually or in groups of three. The goal is to produce a conference-quality paper (6-8 pages in a format like CVPR or NeurIPS).
Each student must pursue his or her academic goals honestly and be personally accountable for all submitted work. Representing another person's work as your own is always a serious offense. All submitted work must be your own. For the semester project, collaboration within your group is expected, but all work must be original to the group. Any use of external code or ideas must be properly cited. Any violation of academic integrity will be reported to the Academic Judiciary and can result in failure of the course. Faculty is required to report any suspected instances of academic dishonesty to the Academic Judiciary. For more comprehensive information on academic integrity, including categories of academic dishonesty please refer to the academic judiciary website.
If you have a physical, psychological, medical, or learning disability that may impact your course work, please contact the Student Accessibility Support Center, Stony Brook Union Suite 107, (631) 632-6748, or at sasc@stonybrook.edu. They will determine with you what accommodations are necessary and appropriate. All information and documentation is confidential.
Stony Brook University expects students to respect the rights, privileges, and property of other people. Faculty are required to report to the Office of Student Conduct and Community Standards any disruptive behavior that interrupts their ability to teach, compromises the safety of the learning environment, or inhibits students' ability to learn. Faculty in the HSC Schools and the School of Medicine are required to follow their school-specific procedures. Further information about most academic matters can be found in the Undergraduate Bulletin, the Undergraduate Class Schedule, and the Faculty-Employee Handbook.