TerraVis: Towards Evaluation of World-Grounded Visual Consistency in Text-to-Image Generation via MLLM Workflows
Abstract
Recent text-to-image models have made substantial progress in photorealism, aesthetics, and text-image alignment. Yet visually appealing images can still violate real-world plausibility, exhibiting malformed object structures, impossible anatomy, physically implausible interactions, or inconsistent spatial relationships. Such failures are not well captured by existing fidelity, aesthetics, preference, or alignment metrics. To address this gap, we introduce TerraVis, a framework for evaluating world-grounded visual consistency in generated images. TerraVis defines a structured taxonomy of world-consistency violations spanning object-, interaction-, and scene-level failures, and employs a multi-stage evaluation framework to identify and quantify them. Given an image, TerraVis first uses an MLLM to assess its eligibility for evaluation, then detects violations across 18 taxonomy-defined types and classifies them as minor or major to derive an overall world-consistency score. Across diverse open-source and proprietary text-to-image models on two widely used benchmarks, TerraVis achieves the strongest correlation with human judgments of world consistency among existing metrics. Our benchmark results further show that models that achieve strong performance on conventional metrics can still exhibit substantial world-consistency failures. These findings highlight world consistency as a complementary evaluation dimension and demonstrate that TerraVis enables systematic quantification, diagnosis, and comparison of such failures. Our code is publicly available at https://github.com/ShyFoo/TerraVis.
Community
Excited to share TerraVis, accepted to NeurIPS 2026 (Evaluations & Datasets Track)! ๐๐๐
While existing T2I metrics primarily focus on prompt alignment and visual quality, TerraVis addresses a long-standing gap in text-to-image evaluation: world consistency, whether generated objects, interactions, and scenes obey real-world constraints.
๐ 18 interpretable violation types with evidence and severity
๐ 13.5K human annotations for evaluation
๐ Try our online evaluator: https://shyfoo.github.io/TerraVis-web/
Code, dataset, and demo are publicly available. Feedback and discussions are very welcome!
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Blending Concepts: Benchmarking Visual Metaphor Generation in Text-to-Image Models (2026)
- GenScale: A Benchmark for Relative Object Scale in Image Generation and Editing (2026)
- GenPuzzle: Benchmarking Visual Reasoning in Image Generation Models (2026)
- Evaluating the Evaluators: Diagnosing Large Multimodal Models for AI-Generated Image Assessment (2026)
- T2LSC-Bench: Benchmarking Localized Semantic Control in Text-to-Image Generation (2026)
- FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling (2026)
- TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.02959 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper