Article-Journal

Decoupled Early Exits for Task-Dependent Compute Allocation in Flow-Matching VLAs
Decoupled Early Exits for Task-Dependent Compute Allocation in Flow-Matching VLAs

Decoupled early exits across backbone depth, action-expert depth and denoising steps for task-dependent compute allocation in flow-matching VLAs.

Sep 29, 2026

GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes
GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes

A benchmark for real-time collaboration between multimodal agents on Keep Talking and Nobody Explodes.

Sep 4, 2026

MIXAR: Scaling Autoregressive Pixel-based Language Models to Multiple Languages and Scripts
MIXAR: Scaling Autoregressive Pixel-based Language Models to Multiple Languages and Scripts

A generative pixel-based language model trained on eight languages across multiple scripts.

Apr 13, 2026

VoyagerVision: Investigating the Role of Multi-modal Information for Open-ended Learning Systems
VoyagerVision: Investigating the Role of Multi-modal Information for Open-ended Learning Systems

An open-ended multimodal embodied agents for the Minecraft game.

Jan 2, 2026

VLM-RobustBench: A Comprehensive Benchmark for Robustness of Vision-Language Models
VLM-RobustBench: A Comprehensive Benchmark for Robustness of Vision-Language Models

A comprehensive study of Vision-Language Models’ behavioural robustness.

Jan 1, 2026

Same Answer, Different Representations: Hidden instability in VLMs

Add the full text or supplementary notes for the publication here using Markdown formatting.

Jan 1, 2026

Retrievit: In-context Retrieval Capabilities of Transformers, State Space Models, and Hybrid Architectures

Add the full text or supplementary notes for the publication here using Markdown formatting.

Jan 1, 2026

AgriPath: A Systematic Exploration of Architectural Trade-offs for Crop Disease Classification

Add the full text or supplementary notes for the publication here using Markdown formatting.

Jan 1, 2026

Movie Facts and Fibs (MF $^ 2$): A Benchmark for Long Movie Understanding

Add the full text or supplementary notes for the publication here using Markdown formatting.

Jan 1, 2025

Visually Grounded Language Learning: A Review of Language Games, Datasets, Tasks, and Models
Visually Grounded Language Learning: A Review of Language Games, Datasets, Tasks, and Models

Jan 1, 2024