Filippo Momentè, Alessandro Suglia, Mario Giulianelli, Ambra Ferrari, Alexander Koller, Oliver Lemon, David Schlangen, Raquel Fernández, Raffaella Bernardi(2025).
Triangulating LLM Progress through Benchmarks, Games, and Cognitive Tests.
Findings of the Association for Computational Linguistics: EMNLP 2025, Suzhou, China, November 4-9, 2025.
Nicola Horst, Davide Mazzaccara, Antonia Schmidt, Michael Sullivan, Filippo Momentè, Luca Franceschetti, Philipp Sadler, Sherzod Hakimov, Alberto Testoni, Raffaella Bernardi, Raquel Fernández, Alexander Koller, Oliver Lemon, David Schlangen, Mario Giulianelli, Alessandro Suglia(2025).
Playpen: An Environment for Exploring Learning From Dialogue Game Feedback.
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025.
Emmanouil Zaranis, António Farinhas, Saul Santos, Beatriz Canaverde, Miguel Moura Ramos, Aditya K Surikuchi, André Viveiros, Baohao Liao, Elena Bueno-Benito, Nithin Sivakumaran, Others(2025).
Movie Facts and Fibs (MF $^ 2$): A Benchmark for Long Movie Understanding.
arXiv preprint arXiv:2506.06275.
Yintao Tai, Xiyang Liao, Alessandro Suglia, Antonio Vergari(2024).
PIXAR: Auto-Regressive Language Modeling in Pixel Space.
Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11-16, 2024.
Georgios Pantazopoulos, Alessandro Suglia, Oliver Lemon, Arash Eshghi(2024).
Lost in Space: Probing Fine-grained Spatial Understanding in Vision and Language Resamplers.
Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Short Papers, NAACL 2024, Mexico City, Mexico, June 16-21, 2024.
Anna Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott, Raquel Fernández, Albert Gatt, Esam Ghaleb, Mario Giulianelli, Michael Hanna, Alexander Koller, André F. T. Martins, Philipp Mondorf, Vera Neplenbroek, Sandro Pezzelle, Barbara Plank, David Schlangen, Alessandro Suglia, Aditya K. Surikuchi, Ece Takmaz, Alberto Testoni(2024).
LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks.
CoRR.