
Research
Mapping Affective Polarization in YouTube Shorts: A Data-Driven Analysis of Political Communication During the 2023–2024 Israel–Hamas War
By Daniel Miehling
The Problem
Miehling addresses a methodological and empirical gap in the study of online political communication, particularly in the context of highly polarized conflicts such as the Israel–Hamas war. The author argues that digital communication consisting of user-generated content is often shaped by emotive cues that signal ideological alignment and provide insight into polarization and sentiment dynamics. However, much of the existing research on communication focuses on small-scale qualitative studies, which cannot capture such patterns on a large scale.
A related problem is that computational methods capable of analyzing large volumes of text — including a technique called Aspect-Based Sentiment Analysis (ABSA), which assesses sentiment toward specific entities mentioned in text (e.g., "Israel," "Hamas," "Palestinians") rather than just the overall mood of a passage — have not been sufficiently adapted to politically charged domains. ABSA is widely used in commercial settings (for product reviews, for example); its application to political communication remains comparatively limited. Most existing computational studies focus on micro-blogging platforms such as Twitter/X, leaving algorithmically driven, visually oriented environments like YouTube Shorts understudied despite their growing importance.
The author argues that YouTube Shorts play an increasingly important role in understanding accelerated communication domains, in which user-generated and state-funded media content shape the digital climate mediated by recommendation algorithms. Under these conditions, affective polarization — the emotional and moral alignment of users toward collective actors like Israel, Zionists, or Palestinians — becomes central to engagement. The paper argues that scalable tools for systematically mapping these evaluative patterns — in which individuals dislike and distrust those with opposing political views – remain underdeveloped in such platform-specific contexts.
Approach and Findings
The study draws on a year-long corpus of 2,370 YouTube Shorts and over 3.4 million comments and replies from 787,157 distinct users, collected from four state-funded outlets representing different geopolitical perspectives: Al Jazeera English (Qatar), TRT World (Turkey), BBC News (UK), and Deutsche Welle (Germany). TRT and Al Jazeera contribute the overwhelming majority of comments (roughly 2.2 million and 1.2 million, respectively), and the author notes that the findings are disproportionately shaped by these two outlets. A follow-up rescrape conducted two months after initial collection revealed that a significant portion of content — particularly on BBC videos — had been deleted or made inaccessible.
The method combines two computational steps. First, comments are processed using dependency parsing, a linguistic technique that identifies syntactic relationships between words by assigning each word a grammatical head and dependency relation. This allows the system to identify entity-specific linguistic contexts within a comment, so that multiple entities mentioned in the same comment can be analyzed independently. Second, ABSA assigns sentiment (positive, neutral, negative) to each of these entity-specific units.
A manually annotated dataset of 5,147 such units was used to fine-tune and evaluate the model. Domain experts labeled each segment for sentiment polarity, applying conservative judgment in cases of sarcasm or implicit meaning. Two annotators agreed at a high rate (Cohen's κ = 0.86; Krippendorff's α = 0.85). The annotated data were used to fine-tune a transformer-based language model (DeBERTa-v3-large-absa-v1.1), which achieved an overall model-performance score (macro F1) of 0.78 — strong performance given the noisy, informal nature of user-generated content.
A complementary approach uses a curated list of 720 morally and ideologically charged tokens (e.g., "terrorism," "genocide," "apartheid," "Nazism") applied to the dependency-parsed text, so that this language can be linked to specific actors in the comments.
The author reports that across all four outlets, user comments frame Hamas overwhelmingly negatively (63–73%), with very little positive sentiment. Comments about Israel are also predominantly negative, particularly on Al Jazeera and TRT, with BBC and DW showing somewhat more balance but still substantial negativity. References to Zionists are overwhelmingly negative, with over 88% having a negative tone across outlets. This negativity is often comparable to or exceeds the negativity directed towards Hamas. Palestine and Palestinians receive the most positive sentiment, particularly on Al Jazeera and TRT, while references to Jews are more neutral on average. The author interprets these distributions as evidence of a polarized moral landscape in the comments, in which user discourse positions Israel as aggressor and Palestinians as victims.
The seed-list analysis identifies asymmetries in how morally charged language is distributed across entities in the comments. Israel is linked to 26,599 terror-related tokens and approximately 18,000 mentions of "genocide" — figures the author reports as rivaling or exceeding Hamas-linked associations (26,153 terror-related; 7,625 genocide-related). The author notes that "genocide" functions differently depending on the target: predominantly as an accusation when linked to Israel ("Israel is committing genocide") and as a marker of victimhood when linked to Palestinians ("genocide against Palestinians"). Hamas is also frequently connected to "resistance," "jihad," and "martyr" in the comments — a pattern the author reads as discursive framing that can normalize militant action.
Qualitative analysis describes how these patterns are intensified through specific discursive tactics. Across the corpus, 18.1% of classified sentences contain at least one emoji. Users employ visual cues (emojis), lexical amplification (all-caps, stacked evaluative nouns such as "APARTHEID GENOCIDAL-TERROR ZIONIST"), and conceptual ambiguity to encode ideological positions — the author illustrates the last with the contrast between "Free Palestine from Hamas" and "Free pagers for Palestine," which share lexis but carry opposing pragmatic meaning. The author also identifies what the paper calls "selective approval," in which Jewish identity is conditionally validated in user comments based on alignment with anti-Israel or anti-Zionist positions ("good Jew / bad Zionist"). Recurring antisemitic tropes documented in the comments include Holocaust inversion ("ZioNazi"), perpetrator–victim reversal, and use of "Zionist" as a proxy for "Jew." The author also documents anti-Palestinian sentiment in the comments, including dehumanizing language, denialist rhetoric, the conflation of Hamas with the civilian population, and portrayals of Palestinians as culturally inferior or inherently violent. The author argues that polarization is bidirectional rather than one-sided. However, the distribution of sentiment suggests that such expressions occur on a much smaller scale than those directed toward Israel and Zionists.
The author identifies three distinct discursive tactics in the corpus tied to potential terrorism endorsement: denial and rejection of hostile evidence ("Can you prove…"), moral reframing of militant violence as legitimate resistance ("Hamas is resistance by the people of Palestine"), and explicit advocacy of violence. The paper notes that rhetorical strategies, including sarcasm, irony, and coded language, complicate automated detection.
Implications
The author argues that large-scale computational approaches such as ABSA, combined with longitudinal aggregation and lexicon-based methods, can capture both emotional intensity and ideological alignment in highly polarized digital environments. The paper suggests that computational communication research is well-positioned to move beyond surface-level sentiment detection toward identifying more subtle and context-specific forms of discourse, including coded language, dog whistles, and indirect expressions of hostility. At the same time, the author notes that the prevalence of such forms indicates that polarization is often embedded in moralized and affective language that may appear benign or empathetic — complicating both detection and moderation.
The longitudinal patterns lead the author to argue that affective polarization in this corpus is not merely reactive to specific events but exhibits a degree of stability over time. Sentiment in the comments fluctuates with major developments, but consistent negative framing of certain groups (Israel and Zionists) and positive framing of others (Palestinians) persists. The author also notes that a portion of hostile comments appears unrelated to the specific video content — antisemitic and racist tropes surface even under videos with neutral topics, which the paper interprets as evidence that antagonistic sentiment toward Jews can persist independently of immediate context.
The study situates online discourse within broader offline dynamics. The author notes that user-generated content engages with Western protest movements and campus encampments organized in solidarity with Gaza, and reports that some commenters frame Jews who identify as Zionists in terms drawn from anti-colonial discourse ("racist white oppressors," "colonizers"). The author argues that narratives chanted in street protests circulate into digital spaces and often intensify there, with users converging around anti-Israel rhetoric, BDS-related framings, and social-justice-framed hate speech and Holocaust distortion. Comparative evidence from other platforms is cited: a CyberWell report indicates that 98% of such content violates community standards on Meta, while 45% remains visible on TikTok.
The author notes a methodological limitation worth highlighting: the analysis is confined to textual content, while affective polarization on YouTube Shorts is fundamentally bimodal — future work would need to integrate systematic analysis of the video content itself to fully account for how audiovisual elements shape user reactions.