Publications

(2024). Shaking Up VLMs: Comparing Transformers and Structured State Space Models for Vision & Language Modeling. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024.
(2024). Repairs in a Block World: A New Benchmark for Handling User Corrections with Multi-Modal Language Models. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024.
(2024). PIXAR: Auto-Regressive Language Modeling in Pixel Space. Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11-16, 2024.
(2024). Lost in Space: Probing Fine-grained Spatial Understanding in Vision and Language Resamplers. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Short Papers, NAACL 2024, Mexico City, Mexico, June 16-21, 2024.
(2024). LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks. CoRR.
(2024). Learning To See But Forgetting To Follow: Visual Instruction Tuning Makes LLMs More Prone To Jailbreak Attacks. CoRR.
(2024). Investigating the Role of Instruction Variety and Task Difficulty in Robotic Manipulation Tasks. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024.
(2024). Human - Large Language Model Interaction: The dawn of a new era or the end of it all?. Companion of the 2024 ACM/IEEE International Conference on Human-Robot Interaction, HRI 2024, Boulder, CO, USA, March 11-15, 2024.
(2024). Enhancing Continual Learning in Visual Question Answering with Modality-Aware Feature Distillation. CoRR.
(2024). AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding. Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, November 12-16, 2024.