As image editing models become more powerful and widely used, there is a growing need to evaluate their performance in real-world scenarios. However, existing benchmarks are often limited to simple single-image tasks, which do not adequately capture the complexity of practical use cases or differentiate between models effectively.
To address this gap, researchers from Hugging Face have introduced CPI-Bench, a new benchmark designed for real-world image editing. It comprises three core subsets: CPI-General-Bench, which covers diverse editing tasks and pioneers multi-image editing evaluation; CPI-Practical-Bench, focusing on high-frequency real-user application scenarios; and CPI-Intelligent-Bench, dedicated to reasoning-based editing tasks that demand high-level understanding.
Evaluation of mainstream image editing models using CPI-Bench shows that it enhances performance differentiation among models. The benchmark provides a comprehensive and reliable quantification of gaps in general editing capabilities, practical deployment efficacy, and advanced reasoning-based editing, offering guidance for future model optimization.
Importantly, ranking analysis reveals that CPI-Bench achieves the highest alignment with the Arena Image Edit Leaderboard, indicating that it faithfully captures human evaluators' preferences and perceptual judgments, serving as a robust proxy for real-world user experience.