Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

CaRGo-T: Graph-of-Thought Framework Boosts Multimodal Humor Comprehension in VLMs

AI By Crimson AI Hugging Face Papers 28 August 2026 · 00:00 1 views
Share: X Telegram

Researchers propose CaRGo-T, a causal reasoning graph-of-thought framework that improves humor understanding and detection in vision-language models by representing causal relationships as graph-based structures, achieving gains of 1-20% on humor understanding and 1-3% on humor detection.

CaRGo-T: Graph-of-Thought Framework Boosts Multimodal Humor Comprehension in VLMs

Key points

Large-scale vision-language models (VLMs) have shown impressive versatility across multimodal tasks, yet humor comprehension remains a challenge. Humor often relies on subtle interactions among entities, events, and context across image and text, requiring complex reasoning chains that conventional prompting or linear chain-of-thought methods struggle to capture.

To address this, researchers introduce CaRGo-T (Causal Reasoning Graph-of-Thought), a framework that models the causal and contextual relationships underlying multimodal humor as a lightweight graph-based reasoning structure. The graph is serialized into a code-based representation generated by a VLM, which can then be interpreted by the same or a different VLM to produce final predictions in zero-shot or in-context learning settings.

Evaluated on four datasets covering satire, sarcasm, and memes, CaRGo-T consistently outperforms existing reasoning-based baselines with state-of-the-art commercial and open-source VLMs. The framework achieves performance gains of approximately 1-20% on humor understanding and 1-3% on humor detection. Further analysis using mutual information shows that CaRGo-T's reasoning representations contain more information relevant to the target output than baseline approaches.

The code is available on GitHub, and the paper is published on arXiv.

TaskPerformance Gain
Humor Understanding~1-20%
Humor Detection~1-3%
Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 29 Aug 2026

Hugging Face Audit: 110 of 124 AI Evaluations Fail to Support Their Claims

A new commit-bound census of 124 Inspect Evals units reveals that 110 stop before deterministic inference due to missing historica...

4
Research paper
Hugging Face 29 Aug 2026

Aphanta: New Framework Diagnoses When Image Editing Boosts Multimodal Reasoning

Hugging Face researchers introduce Aphanta, a diagnostic framework that evaluates when image-editing intermediates improve multimo...

5
Research paper
Hugging Face 29 Aug 2026

EditaLive! Enables Real-Time Character Video Editing for Live Streaming

Hugging Face researchers introduce EditaLive, a framework for real-time human-centric video editing in live streams, achieving sta...

4