Article-Journal

VoyagerVision: Investigating the Role of Multi-modal Information for Open-ended Learning Systems
VoyagerVision: Investigating the Role of Multi-modal Information for Open-ended Learning Systems

An open-ended multimodal embodied agents for the Minecraft game.

Jan 2, 2026

VLM-RobustBench: A Comprehensive Benchmark for Robustness of Vision-Language Models
VLM-RobustBench: A Comprehensive Benchmark for Robustness of Vision-Language Models

A comprehensive study of Vision-Language Models’ behavioural robustness.

Jan 1, 2026

Same Answer, Different Representations: Hidden instability in VLMs

Add the full text or supplementary notes for the publication here using Markdown formatting.

Jan 1, 2026

Retrievit: In-context Retrieval Capabilities of Transformers, State Space Models, and Hybrid Architectures

Add the full text or supplementary notes for the publication here using Markdown formatting.

Jan 1, 2026

AgriPath: A Systematic Exploration of Architectural Trade-offs for Crop Disease Classification

Add the full text or supplementary notes for the publication here using Markdown formatting.

Jan 1, 2026

Movie Facts and Fibs (MF $^ 2$): A Benchmark for Long Movie Understanding

Add the full text or supplementary notes for the publication here using Markdown formatting.

Jan 1, 2025

Visually Grounded Language Learning: A Review of Language Games, Datasets, Tasks, and Models
Visually Grounded Language Learning: A Review of Language Games, Datasets, Tasks, and Models

Jan 1, 2024

LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Jan 1, 2024

Learning To See But Forgetting To Follow: Visual Instruction Tuning Makes LLMs More Prone To Jailbreak Attacks

Jan 1, 2024

Enhancing Continual Learning in Visual Question Answering with Modality-Aware Feature Distillation

Jan 1, 2024