MLLMs Construction Company

Investigating multimodal LLMs' communicative skills in a collaborative building task.

How effective are the communication choices of Multimodal Large Language Models when pursuing a common goal? Can they make use of common human dialogical patterns?

We address these questions by engaging two agents based on the Mistral model in a collaborative building task, where one has to instruct the other how to build a specific target structure. The work investigates whether different prompting techniques with varying degrees of multimodality influence the performance of MLLM-based agents.

Published at CLiC-it 2025.

r3lativo/MLLMs-construction-company