Conferences, 17th International Conference on Computational Methods

Font Size: 
Autonomous Prompt Optimization for Motion Pictogram Generation Based on Multimodal LLM Evaluation
Natsumi Okatani, Ryuji Shioya, Yasushi Nakabayashi

Last modified: 2026-05-29

Abstract


Pictograms in public spaces are widely used as a means of conveying information without relying on language; however, they have inherent limitations in communicating dynamic processes and complex action instructions. While automatic generation of motion pictograms using video generation AI is a promising approach, quality variation due to the stochastic nature of generative AI output and the human cost of manual verification remain barriers to practical deployment.

This study proposes a framework in which an MLLM (Multimodal Large Language Model) is integrated as an evaluation agent to recursively perform automatic quality assessment and autonomous optimization of motion generation prompt. Built upon a TI2V (Text-guided Image-to-Video) generation method [1][2], the framework assigns the MLLM the role of an expert in semiotics and human-centered design, enabling scoring of generated motion pictograms along two axes: Semantic Fidelity and Expressive Quality. When scores fall below the threshold, the MLLM verbalizes the cause of failure and feeds it back into the next prompt generation. This loop achieves quality improvement without human intervention.

Validation experiments were conducted using Luma Dream Machine [3] (Ray2) for video generation and Google Gemini 3 [4] for evaluation and optimization, targeting five pictograms with a maximum of four iterations (Iteration 0–3). Outputs judged C (Unusable) by the AI evaluation were frequently rated relatively low in human evaluation as well, suggesting the validity of the system as a filtering mechanism. At the same time, erroneous judgments due to MLLM hallucination were also observed, indicating that improving evaluation reliability remains a future challenge. The proposed framework demonstrates that an AI-driven feedback loop can control the uncertainty inherent in motion pictogram generation, and holds potential for application to the automated generation of public pictograms where safety and accuracy are paramount.


Keywords


Motion Pictogram, Multimodal LLM, Autonomous Prompt Optimization, Video Generation AI, Agentic AI

Conference registration is required in order to view papers.