MLLMs Construction Company
Investigating multimodal LLMs' communicative skills in a collaborative building task.
How effective are the communication choices of Multimodal Large Language Models when pursuing a common goal? Can they make use of common human dialogical patterns?
We address these questions by engaging two agents based on the Mistral model in a collaborative building task, where one has to instruct the other how to build a specific target structure. The work investigates whether different prompting techniques with varying degrees of multimodality influence the performance of MLLM-based agents.
Published at CLiC-it 2025.