Publications

(2026). Decoupled Early Exits for Task-Dependent Compute Allocation in Flow-Matching VLAs. CoRR.
(2026). GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes. TMLR 2026.
(2026). Do Composed Image Retrieval Benchmarks Require Multimodal Composition?. NeurIPS 2026.
(2026). MIXAR: Scaling Autoregressive Pixel-based Language Models to Multiple Languages and Scripts. CoRR.
(2026). VoyagerVision: Investigating the Role of Multi-modal Information for Open-ended Learning Systems. Advances in Intelligent Systems and Computing ((AISC,volume 1468)).
(2026). VLM-RobustBench: A Comprehensive Benchmark for Robustness of Vision-Language Models. ICML 2026.
(2026). Same Answer, Different Representations: Hidden instability in VLMs. CoRR.
(2026). Retrievit: In-context Retrieval Capabilities of Transformers, State Space Models, and Hybrid Architectures. TMLR 2026.
(2026). AgriPath: A Systematic Exploration of Architectural Trade-offs for Crop Disease Classification. TMLR 2026.
(2025). CROPE: Evaluating In-Context Adaptation of Vision and Language Models to Culture-Specific Concepts. NAACL 2025.