
Research
Routine Distortion: Why We Urgently Need to Expand Research on AI-Facilitated Holocaust Misrepresentation Beyond the AI Slop
By Mykola Makhortykh, Elizaveta Kuznetsova
The Problem
Generative AI is increasingly implicated in the production and circulation of distorted representations of the Holocaust, ranging from fabricated historical images — such as the wave of AI-generated images of Nazi concentration camps that circulated on social media in summer 2025, apparently produced to game platforms’ monetization programs — to chatbot systems like Grok (which briefly generated outputs referring to itself as "MechaHitler") or applications such as Historical Figures, which let users converse with chatbots imitating Nazi officials who then express remorse and claimed they tried to prevent the Holocaust. The authors note that these outputs can misrepresent events, blur lines of responsibility, or reframe perpetrators and victims in ways that deviate from established historical understanding. Even when such content does not formally meet the IHRA criteria for Holocaust distortion — minimization, blurring responsibility, or presenting the Holocaust as a positive event — it can still mischaracterize the genocide, distort historical facts, and weaken public trust in authentic evidence. The authors also note that AI is being adopted by institutions such as Yad Vashem and projects like Dimensions in Testimony for Holocaust memory preservation, education, and archival enhancement, though these constructive applications tend to receive less visibility than cases of misuse.
The authors argue that public and scholarly attention has concentrated on highly visible, viral instances of AI slop, which risks obscuring more routine, embedded, and less conspicuous forms of distortion arising through everyday uses of generative AI in historical representation. As a result, understanding of AI's impact on Holocaust memory becomes skewed toward exceptional incidents rather than structural transformations in how historical narratives are produced and circulated.
Makhortykh & Kuznetsova therefore propose to move beyond the focus on viral cases by examining less visible, routine practices of AI use, in order to better understand how generative AI is reshaping Holocaust memory in more gradual and embedded ways.
Approach and Findings
The paper examines how generative AI is reshaping the informational ecosystem through which people access and interpret Holocaust history, focusing on routine, everyday information-seeking rather than viral cases of misuse. The authors emphasize a broader shift from retrieving information sources to receiving synthesized responses, which they argue fundamentally alters how historical knowledge is encountered online. This shift has been described as a "Google Zero" or zero-click reality, in which former gatekeepers of information effectively become machines for ready answers — a transformation the authors note is already reflected in documented declines in traffic to reference sites such as Wikipedia.
The authors present empirical evidence from survey data across Switzerland, the United States, and Germany, indicating a rapid and sustained increase in the use of generative AI for historical inquiry, including Holocaust-related topics. A series of Swiss surveys reports the share of respondents using chatbots for historical information consumption at approximately 5% in January 2024, 7% in December 2024, and 14% in February 2026. For Holocaust-specific queries, 18.6% of US respondents and 12.7% of German respondents reported having used generative AI chatbots. The authors note dramatically higher rates among student populations reported by external studies: 88% of UK undergraduates reported using generative AI in 2025, up from 53% the previous year. This puts generative AI among the key platforms used for accessing historical information, alongside search engines, online encyclopedias, and traditional media.
Despite this growing reliance, the authors argue that multiple studies indicate generative AI systems frequently produce unreliable outputs in response to Holocaust-related queries. For instance, a study on the accuracy of generative AI chatbots for prompts about the Holocaust in Ukraine suggests that in best-case scenarios, only around 70% of prompts receive factually correct responses, with accuracy dropping substantially for prompts in regional languages such as Russian and Ukrainian. Beyond factual errors, chatbots exhibit hallucinations — confidently stated but invented information — including references to fabricated war crime trials, invented memory laws, and made-up witness testimonies. The authors offer the Nachtigall battalion as an illustrative case: ChatGPT in Russian claimed the battalion did not exist during WWII but was founded in 2014 following Russia's hybrid warfare against Ukraine; Google's Bard indiscriminately attributed Second World War atrocities to it, in one instance inventing an "Alfred Eichmann" as a Nachtigall member. The authors argue that these hallucinations often emerge from ordinary, non-adversarial prompts and that they reflect a structural feature of current generative AI rather than a transient bug. Because text-generative models predict the next word based on context — rather than understanding what they are saying — the authors argue that hallucination, and therefore distortion, is likely inevitable. Other studies showcase that outputs are further destabilized by random variation in how the AI models respond to identical prompts, and by inconsistent framing — for example, some AI responses treat the 1936 Olympics as a precursor to Nazi atrocities, while others omit the Holocaust and frame the event as a triumph over racism centered on Jesse Owens.
The authors argue that these risks are amplified by the level of trust users place in generative AI systems, even when outputs are uncertain or incorrect. The combination of widespread adoption, perceived reliability, and structural limitations, they argue, creates conditions in which routine interactions with AI can subtly shape understandings of Holocaust history.
Implications
Makhortykh & Kuznetsova argue that routine, low-visibility forms of AI-driven Holocaust distortion are unlikely to be eliminated through any single technological, educational, or regulatory intervention. They propose a shift away from attempts at total prevention toward pragmatic strategies that reduce risk and improve the reliability of high-frequency routine interactions. A central implication they draw is that most distortion emerges from a relatively limited set of recurring topics and informational needs — definitions of the Holocaust, victim and survivor numbers, perpetrator motivations — and that efforts should therefore prioritize accurate responses to the hundred or thousand most common Holocaust-related queries, while strengthening safeguards around well-known forms of distortion. The authors frame this pragmatically: better to address 85% of the problem with a realistic fix than to aim for 95% with a solution that may not work.
The authors argue for the need for high-quality, expert-curated knowledge infrastructures to support generative AI systems, potentially through approaches such as retrieval-augmented generation (RAG) — a method that gives AI systems direct access to vetted external databases when answering questions. This, they contend, requires sustained collaboration between AI developers and domain experts in Holocaust history, antisemitism studies, digital humanities, and human-computer interaction, as well as institutions responsible for Holocaust remembrance. The authors note that the dynamic nature of both Holocaust memory and AI systems calls for continuous empirical monitoring to ensure that evolving models and user behaviors do not introduce new forms of distortion over time.
Beyond technical solutions, the authors emphasize complementary societal interventions — particularly AI literacy initiatives that help users critically engage with generative AI systems in domains such as history and cultural memory — and coordination across fragmented efforts by NGOs, research centers, and heritage institutions, exemplified by the Landecker Digital Memory Lab's initiatives. The authors stress the urgency of acting now, arguing that growing reliance on generative AI risks reinforcing a feedback loop in which synthetic content increasingly becomes training data for future models — a phenomenon known as model collapse — potentially degrading information quality and undermining trust in historical authenticity more broadly.