Categories
October Surprise 2024

MAPS: A Multilingual Benchmark for Global Agent Performance and Security

Despite growing interest in evaluating agentic AI, existing benchmarks focus exclusively on English, leaving multilingual settings unexplored. To address this gap, we propose MAPS, a multilingual benchmark suite designed to evaluate agentic AI systems across diverse languages and tasks.