学术研究★★★arXiv · 2026-07-20
The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric
Human judgments of visual similarity are context-dependent, and existing metrics fail to distinguish between different aspects. This paper introduces a text-prompted perceptual metric for images using a large-scale dataset annotated with multiple aspects of similarity.
📌 Key points
- Human visual similarity judgments are context-dependent
- Existing metrics fail to distinguish between different aspects like shape and co
- A large-scale dataset of image triplets annotated with multiple similarity aspec
- Frontier vision-language models show significant performance gaps in this task
本页为 gitzw.com 基于公开来源的 AI 中文解读,非原文转载。