Publications

論文誌 (国際) Evaluating the cultural alignment of multilingual LLMs in typical Japanese workplace scenarios

Zhiwei Gao (NAIST), Nobuyuki Shimizu, Sumio Fujita, Shaowen Peng (NAIST), Shoko Wakamiya (NAIST), Eiji Aramaki (NAIST)

PLOS ONE

2026.7.27

Large Language Models (LLMs) are increasingly used in cross-cultural communication, but their ability to align with local cultural norms remains unclear. This study investigates how multilingual LLMs adapt to Japanese workplace culture, using Hofstede’s six cultural dimensions as the theoretical framework. We evaluate five LLMs—LLM-jp, Phi, Llama, Qwen, and GLM—through two methods: (1) qualitative analysis based on the Valence-Arousal-Dominance (VAD) emotion model to examine whether model outputs reflect culturally expected emotions, and (2) quantitative evaluation using a structured questionnaire rated by Japanese crowd workers. Based on human ratings, we calculate the Japanese Workplace Cultural Alignment Score (JWCAS) for each model. The results show that GLM and LLM-jp achieve relatively strong alignment with Japanese cultural expectations, while models like Llama show lower consistency. We also analyze the correlation between human and LLM-judge evaluations, finding moderate agreement overall, but with significant variation across models and cultural dimensions. These findings show the need for evaluation methods that consider culture and point to future work for making LLMs more culturally adaptable.

Paper : Evaluating the cultural alignment of multilingual LLMs in typical Japanese workplace scenarios新しいタブまたはウィンドウで開く (外部サイト)