Sarcasm is not only about what someone says. It is also about how they say it.
A literal compliment can become ironic through timing, pitch, intensity, or voice quality. Human listeners combine these cues with the words and conversational context to infer the speaker's intent. Multimodal large language models (MLLMs) can now process speech and text together, but it remains unclear whether they reason about prosody in the same flexible way—or whether they rely on simpler acoustic shortcuts.
Our new paper investigates this question through spoken sarcasm detection in English and Mandarin Chinese. Rather than evaluating only aggregate accuracy, we decompose the contribution of text, vocal content, and prosodic structure; identify which acoustic features characterize model errors; and then manipulate those features directly to test whether they cause the errors.
The result is a consistent diagnostic pattern: current MLLMs often respond to a cross-linguistic stereotype of expressive speech—elevated pitch and irregular pauses—even when that stereotype does not match the authentic sarcasm cues in either language.
Why sarcasm is a useful stress test
Speech models can perform well on transcription while still struggling with pragmatic meaning. Sarcasm makes that distinction especially visible because the intended meaning often emerges from tension between lexical content and vocal delivery.
Prosodic cues are also highly context- and language-dependent. Longer duration is a relatively stable marker of sarcasm across languages, but the direction of pitch is not. Studies of English have reported both higher and lower fundamental frequency, and tonal languages introduce additional interactions between pitch and lexical meaning.
A model that has genuinely learned pragmatic prosody should therefore respond to combinations of cues in their linguistic context. A model using a surface heuristic may instead equate any conspicuously expressive voice with irony.
Our study asks not simply, “Can the model classify sarcasm?” but a more revealing question: What does the model actually hear when it decides that an utterance is sarcastic?
Separating words, voice, and prosody
We evaluate two openly available speech-capable models zero-shot:
- Qwen2.5-Omni-7B, referred to as Omni 2.5; and
- Qwen3-Omni-30B-A3B-Thinking, referred to as Omni 3.0.
The evaluation uses two balanced multimodal sarcasm corpora: MUStARD++, containing 1,202 English audio–transcript pairs from sitcoms, and MSCD, containing 2,705 Mandarin Chinese samples from televised stand-up comedy.
To isolate the information available to the model, we construct five modality conditions:
- Text-only: the transcript without audio;
- Speech-only: the complete voice signal without a transcript;
- Prosody-only: low-pass-filtered audio that preserves pitch, rhythm, and intensity patterns while removing most lexical information;
- Bimodal: the full audio paired with the transcript; and
- Biprosody: the transcript paired with prosody-only audio.
Because Bimodal and Biprosody share the same transcript, their difference isolates the contribution of vocal semantics. Comparing either condition with Text-only reveals what the audio channel adds. Comparing Speech-only with the transcript-paired conditions reveals what happens when literal words and vocal delivery are interpreted together.
We also extract 66 acoustic features covering fundamental frequency, intensity and energy, rhythm and timing, and voice quality. These measurements let us compare the prosody associated with genuine sarcasm against the prosody associated with model mistakes.
What audio changes
Adding audio produces modest improvements in overall F1—between 1.0 and 6.8 percentage points over Text-only in the Bimodal condition—but the aggregate score hides an important shift in errors.
For Omni 3.0, adding full audio to the transcript increases false positives by 9.8% of English samples and 8.6% of Chinese samples on average across splits. Biprosody produces a nearly identical effect: +12.0% in English and +7.6% in Chinese. In other words, the filtered signal—without intelligible words—retains most of the audio effect.
The model trades missed sarcasm for false alarms. More non-sarcastic utterances are labeled sarcastic whenever a transcript is paired with a salient prosodic contour.
Prosody alone is not enough. Omni 3.0's Prosody-only F1 falls to 22.2% in English and 14.3% in Chinese, below the 50% chance baseline. Omni 2.5 also shows no reliable prosody-only discrimination. The models do not appear to decode sarcasm directly from prosodic structure; instead, the transcript and audio together activate a perceived mismatch between literal words and expressive delivery.
What sarcasm actually sounds like in the two corpora
The acoustic ground truth differs substantially between English and Mandarin Chinese.
In the Chinese corpus, the strongest sarcasm markers are temporal and pitch-structural: longer total pause duration, longer utterance duration, more complex pitch contours, and less regular rhythm. Mean pitch itself is not a leading discriminator.
In the English corpus, the strongest signals are intensity-related, followed by lower pitch and longer duration. The pitch direction is particularly important: genuine English sarcasm in this dataset tends to have a lower fundamental frequency.
The two languages therefore do not share a simple “sarcastic tone.” Mandarin sarcasm is characterized more by pitch movement and timing, while English sarcasm is associated with energy patterns and suppressed pitch. Duration is one of the few cues that generalizes consistently.
The false-positive signature
Model errors do not follow either of these ground-truth profiles.
Non-sarcastic utterances that become false positives after audio is added are acoustically much closer to correctly classified non-sarcastic controls than to genuine sarcasm. Yet the models' audio encoders place these same utterances in an ambiguous region between sarcastic and non-sarcastic representations.
Across both languages, the false positives share two surface properties:
- elevated pitch, and
- irregular or conspicuous pausing.
In Mandarin, the model focuses on the wrong aspect of the relevant acoustic domains: pitch level rather than pitch-contour complexity, and pause irregularity rather than total duration. In English, the mismatch is even sharper—the model responds to elevated pitch even though genuine sarcasm in the corpus is marked by lower pitch.
This is the core heuristic uncovered by the analysis. The models appear to have learned a generic association between expressive delivery and irony rather than a language-grounded representation of sarcastic intent.
Causal verification: changing only pitch and pauses
Correlation alone cannot establish that these cues cause the mistakes. To test the mechanism directly, we take a separate set of non-sarcastic control utterances that the models classify correctly and modify only the two dimensions associated with the false-positive profile.
Using PSOLA-based speech manipulation, we raise pitch by 8.8% and alter existing pause durations within the range observed in naturally occurring false positives. The transcript, speaker content, and recording context remain unchanged. No new pauses are inserted, and intelligibility and perceived audio quality remain largely stable.
The intervention produces large increases in false positives across both languages and models. In Mandarin, false-positive rates rise to between 19.35% and 40.65%, depending on the model and modality condition. In English, they reach between 33.55% and 60.53%.
Changing only pitch level and pause structure is therefore sufficient to make many correctly classified non-sarcastic utterances look sarcastic to the models.
The reverse test supports the same conclusion. Moving naturally occurring false positives toward control-like pitch and pause targets restores the correct classification in 29.5–56.3% of cases.
The effect transfers beyond one model family
To test whether the heuristic is specific to Qwen's Omni architecture, we apply the same diagnostic procedure to Gemini 3 Flash Preview.
The overall pattern repeats. Adding audio increases false positives relative to Text-only in both languages, with increases of 13.2 and 13.6 percentage points for English Bimodal and Biprosody conditions. The manipulation template derived from the Omni models transfers to Gemini without modification, flipping 4.7–17.1% of previously correct predictions across conditions and languages.
The lower transfer rate reflects differences in model architecture and audio encoding, but the direction is consistent. Elevated pitch and irregular pausing form a broader vulnerability rather than an artifact of a single model family.
Why this matters
Aggregate benchmark scores can conceal the mechanism a model uses to arrive at its predictions. In this study, audio sometimes improves F1, but it does so while systematically increasing false alarms. Without directional error analysis, acoustic profiling, and causal intervention, that trade-off would be easy to miss.
The findings also have implications beyond sarcasm. Speech-based assistants increasingly make inferences about emotion, intent, urgency, confidence, and other paralinguistic properties. A model that treats generic expressiveness as evidence for a specific mental or pragmatic state may behave unpredictably across speakers, cultures, and languages.
Future training should therefore align prosodic representations with language-specific pragmatic cues instead of rewarding broad surface correlations. Contrastive examples with identical transcripts and different vocal realizations could teach models that pitch and pausing do not carry fixed meanings. Prosody-aware reinforcement learning offers another path toward more grounded acoustic reasoning.
Looking ahead
This work moves multimodal sarcasm evaluation from performance measurement to mechanism diagnosis. The central lesson is that hearing audio is not the same as understanding prosody.
Models may appear to benefit from an additional modality while still relying on a shallow expectation of what sarcasm “should” sound like. Reliable multimodal reasoning will require evaluations—and training methods—that distinguish authentic, context-dependent pragmatic cues from convenient acoustic stereotypes.