EN Submit a tool
Paper

Why Multimodal Large Models Struggle to Understand Humor: A Survey of Methods, Datasets, Evaluation, and Challenges

Published: Source: HuggingFace Daily Papers (Community Hot Papers)

ShareXFacebookTelegramWhatsApp

A survey systematically reviews the structural blind spots of multimodal large language models (MLLMs) in understanding visual humor such as memes, comics, and satirical images: the core difficulty lies not in multimodal alignment, but in reasoning about non-literal mechanisms, shared cultural knowledge, and communicative intent. The survey organizes literature by three progressive capability levels—recognition, interpretation and reasoning, and generation—and points out that current progress is limited by shortcut evaluations, insufficient cultural coverage, weak evidence foundations, and safety and ownership issues.

Read the original (opens in a new tab)

News stream data aggregated by AI HOT

Related newsLatest in this category