Rewriting the Past: Can AI Model Literary History?

Author: Denis Avetisyan


A new study explores how artificial intelligence can be used to simulate literary outputs and explore ‘what if?’ scenarios in the history of literature.

The study contrasts outputs from artificial intelligence with human writing across diverse genres, demonstrating how prompting complexity-ranging from basic to nuanced-significantly shapes the characteristics of generated text and reveals the inherent differences between algorithmic and organic composition.
The study contrasts outputs from artificial intelligence with human writing across diverse genres, demonstrating how prompting complexity-ranging from basic to nuanced-significantly shapes the characteristics of generated text and reveals the inherent differences between algorithmic and organic composition.

This research demonstrates the potential of large language models for simulation-based experiments in literary analysis, leveraging text generation and embedding similarity to address questions of causal inference and counterfactual history.

While reconstructing literary history relies on incomplete evidence and subjective interpretation, this paper, ‘AI as a Tool for Simulation-Based Experiments in Literary Studies’, explores the potential of large language models to simulate cultural production and enable controlled, large-scale experimentation. We demonstrate initial success in generating literary text and validating AI as a proxy for human authors, offering a novel approach to counterfactual analysis and literary research. However, significant challenges remain in achieving historical accuracy and stylistic diversity-what new insights might emerge as these models become increasingly sophisticated tools for understanding the complex dynamics of literary culture?


The Echo Chamber of Interpretation

For generations, understanding literature has depended on close reading and nuanced interpretation – a process inherently subjective and time-consuming. While invaluable for appreciating individual works, this qualitative approach presents significant limitations when attempting to discern broader patterns across numerous texts. Identifying evolving stylistic trends, tracing the prevalence of specific themes, or comparing the works of hundreds of authors demands a scope beyond what traditional methods can realistically achieve. The very strength of literary analysis – its focus on individual meaning and contextual understanding – becomes a bottleneck when researchers seek to map the literary landscape at scale, revealing the need for complementary, quantitatively-driven techniques to unlock insights hidden within vast textual datasets.

The traditional methods of literary analysis, while valuable, are fundamentally constrained by their reliance on subjective interpretation. This inherent subjectivity doesn’t invalidate nuanced readings, but it does create significant obstacles when attempting to discern broader patterns or conduct rigorous comparative studies. Because interpretations can vary widely even amongst experts, establishing definitive trends in literary style, thematic development, or the influence of one author upon another becomes exceedingly difficult. This limitation is particularly pronounced when dealing with large corpora of text, where the sheer volume makes exhaustive, consensus-based qualitative assessment impractical. Consequently, researchers often find themselves navigating a landscape of potentially conflicting interpretations, hindering the development of truly data-driven insights into the evolution of literature and the complex interplay of ideas within it.

The limitations of traditional literary scholarship – reliant on nuanced, yet subjective, readings – are prompting a shift towards computational methods. Researchers are now developing techniques to analyze extensive collections of texts, moving beyond individual interpretation to identify broad patterns in literary style, thematic development, and historical change. This involves applying tools from natural language processing and data mining to quantify aspects of writing previously assessed qualitatively – such as vocabulary richness, sentence structure, and the prevalence of specific motifs. By treating literary works as data points, scholars aim to reveal underlying trends and connections that might otherwise remain obscured, offering a more objective and comprehensive understanding of the literary landscape and its evolution over time. This quantitative approach doesn’t seek to replace traditional analysis, but rather to complement it, providing new avenues for inquiry and a wider lens through which to examine the complexities of literature.

Narrative texts are discernibly clustered within a dimension-reduced embedding space, demonstrating the effectiveness of the representation.
Narrative texts are discernibly clustered within a dimension-reduced embedding space, demonstrating the effectiveness of the representation.

The Literary Engine: Fabricating Voices at Scale

The generation of synthetic text for research is accomplished utilizing Large Language Models (LLM), with the current implementation leveraging the GPT-5 architecture. This model, a deep learning neural network, is employed to produce text based on probabilistic predictions of word sequences. Input to GPT-5 consists of initial prompts and parameters defining desired text characteristics, such as length, style, and subject matter. The resulting output is raw text data, which is then subject to further analysis and processing as part of our experimental workflow. GPT-5 was selected for its demonstrated capacity to generate coherent and contextually relevant text at scale, facilitating the creation of large datasets for controlled experimentation.

Prompt engineering for Large Language Models (LLMs) involves carefully constructing input text to elicit desired outputs, functioning as a primary control mechanism for both stylistic and thematic elements. The precision of these prompts-including detailed instructions regarding tone, vocabulary, sentence structure, and subject matter-directly impacts the generated text’s characteristics. Iterative refinement of prompts, often involving A/B testing with varying parameters, is essential to achieve consistent and predictable results. Furthermore, techniques such as few-shot learning-providing the LLM with examples of the desired output style-and the use of constraint-based prompts-explicitly defining boundaries for the generated content-are commonly employed to fine-tune the LLM’s behavior and ensure alignment with specific research objectives.

Synthetic author biographies are constructed datasets detailing fictional author backgrounds, including dates of birth, education, key life events, and prevailing philosophical or political viewpoints. These biographies are then incorporated into the prompting sequence for the Large Language Model (LLM), providing a defined contextual framework. The LLM utilizes this biographical information to inform stylistic choices, thematic preferences, and narrative perspectives within the generated text. By manipulating parameters within these synthetic biographies – such as altering an author’s documented exposure to specific literary movements or historical events – we can observe quantifiable shifts in the characteristics of the generated output, thereby increasing the fidelity of the simulation and providing richer contextual data for analysis.

The generation of synthetic text using LLMs enables the creation of a controlled experimental environment for literary research. By manipulating parameters such as authorial biography and prompt structure, researchers can isolate variables and systematically test hypotheses regarding the development of literary trends. This approach facilitates the investigation of authorial influence – specifically, how attributed characteristics impact textual output – without the confounding factors present in historical literary production. The resulting data allows for quantitative analysis of stylistic and thematic elements, providing a means to validate or refute theories concerning literary evolution and the relationship between author and text.

Analysis of word usage in human- and AI-authored pulp western fiction reveals distinct linguistic patterns, highlighting differences in authorial style.
Analysis of word usage in human- and AI-authored pulp western fiction reveals distinct linguistic patterns, highlighting differences in authorial style.

Mapping the Literary Genome: Quantification as Revelation

Document embedding techniques transform text into numerical vectors – specifically, high-dimensional arrays of floating-point numbers – allowing for computational analysis of textual similarity and difference. These techniques, such as those utilizing transformer models, map words and entire documents into a vector space where the geometric distance between vectors reflects semantic relatedness. This representation enables quantitative comparison of texts regardless of length or complexity, and facilitates the application of statistical methods – including cosine similarity and clustering – to identify patterns and relationships within and between corpora of human-authored and AI-generated content. The resulting vectors capture underlying semantic features, moving beyond simple keyword matching to assess conceptual similarity.

Fightin’ Words Analysis is a quantitative lexical analysis technique used to pinpoint statistically significant differences in word usage between two or more text corpora. The method operates by calculating the frequency of individual words or n-grams within each corpus and then applying statistical tests – typically chi-squared or log-likelihood ratio – to determine if observed differences in frequency are unlikely to have occurred by chance. Identified “fightin’ words” – those exhibiting statistically significant variance – serve as stylistic fingerprints, characterizing the distinct lexical profiles of each corpus and enabling researchers to quantify stylistic divergence or similarity. This allows for objective comparison of authorial styles, genre conventions, or, as demonstrated in our research, the characteristics of human versus AI-generated text.

Application of document embedding and fightin’ words analysis to the CONLIT Corpus, a large collection of literary texts, and our generated texts allows for a quantitative evaluation of AI-generated literature’s authenticity and diversity. This methodology facilitates the comparison of stylistic and thematic features between human and AI-authored works. Specifically, the cosine similarity between vector representations of texts within the corpus serves as a metric for assessing textual homogeneity; lower similarity scores indicate greater diversity. By benchmarking AI-generated texts against the CONLIT Corpus, we can determine the extent to which AI is replicating existing literary styles or generating novel content, and thus evaluate its contribution to literary diversity.

Quantitative analysis of the CONLIT corpus and generated texts demonstrates the feasibility of measuring narrative homogeneity and identifying stylistic characteristics within textual datasets. Initial results, based on document embedding and cosine similarity calculations, reveal a similarity score of 0.635 for human-authored texts and 0.580 for AI-generated texts nominated for literary prizes, both within the same genre. This indicates a marginally greater degree of diversity in the AI-generated outputs when compared to the human-authored corpus, suggesting that current AI models, at least within this dataset, do not entirely replicate the stylistic consistency observed in human writing.

Rewriting the Past: The Contingency of Literary History

Simulation-based literary experimentation offers a novel approach to understanding literary history by moving beyond traditional analysis and into the realm of “what if.” This technique utilizes artificial intelligence to recreate and then subtly alter the conditions surrounding a literary work or author, effectively rewriting history to observe the potential consequences. Researchers can, for example, simulate a literary landscape where a key author never existed, or where a significant historical event unfolded differently, and then analyze the resulting changes in generated text. By systematically manipulating these variables, the method allows for the testing of hypotheses regarding causal relationships – determining whether a particular event genuinely influenced a literary trend, or if it was merely a coincidence. This isn’t about predicting the past, but about creating controlled experiments within a digital environment to gain a deeper understanding of the complex forces that shape creative expression and literary evolution.

The capacity to reshape literary outputs based on altered historical landscapes hinges on sophisticated techniques like temporal conditioning and model editing. Temporal conditioning involves subtly influencing the generative model with contextual cues representing specific eras, prompting it to adopt the linguistic styles, prevalent themes, and even biases of that time. Complementing this, model editing allows for direct manipulation of the model’s internal parameters, effectively ‘teaching’ it alternative literary lineages or suppressing knowledge of certain authors or movements. This isn’t simply about stylistic mimicry; the aim is to simulate how a different intellectual climate might have shaped the very content of literary works, exploring how the absence of a key influence, or the prominence of a forgotten one, could have steered creative expression in entirely new directions. By meticulously adjusting these parameters, researchers can generate texts that aren’t just written in a different era, but feel genuinely of that era, offering a powerful lens for understanding the contingent nature of literary history.

The capacity to selectively erase information from large language models, a process known as model unlearning, offers a unique lens through which to examine the subtle forces shaping literary history. Researchers are leveraging this technique to simulate the effects of censorship, lost texts, or the suppression of artistic movements, effectively reconstructing literary landscapes devoid of specific influences. By systematically removing knowledge of certain authors, historical events, or even stylistic conventions from the model, it becomes possible to observe how the generated text adapts – revealing which elements were contingent upon those lost factors and which remain resilient. This approach doesn’t simply identify gaps in knowledge; it illuminates the often-invisible scaffolding that supports creative expression, demonstrating how easily literary evolution can be diverted by the absence of key inspirations or the deliberate silencing of voices.

The capacity to computationally model literary history allows for a rigorous investigation of cause and effect in the development of creative traditions. Recent experimentation demonstrates that increasingly intricate prompts – those detailing nuanced historical contexts and authorial influences – yield AI-generated texts exhibiting greater stylistic diversity, as measured by a cosine similarity of 0.580, compared to simpler prompts which score higher at 0.682. This suggests the model responds to complexity by generating less predictable outputs, mirroring the organic evolution of literary style. Notably, a cross-similarity score of 0.541 between AI-generated texts and those recognized with literary prizes indicates a surprising degree of alignment in stylistic qualities, hinting at underlying, quantifiable principles governing aesthetic appeal and potentially offering new avenues for understanding literary judgment.

The pursuit of simulating literary outputs, as detailed in the study, echoes a fundamental truth about complex systems: one cannot truly build them, only cultivate their potential. The LLM, in this context, isn’t a constructor of narratives, but a garden where counterfactual histories might bloom, however imperfectly. As Blaise Pascal observed, “The eloquence of the body is in its movements, and the eloquence of the mind is in its order.” The ‘order’ the LLM attempts to impose on textual possibility reveals the inherent chaos of creation, a system constantly tending toward both coherence and collapse. The limitations in historical accuracy and output diversity aren’t failures of the tool, but prophecies of the system’s inevitable divergence from perfect simulation, reminding one that even the most sophisticated model remains a shadow of the infinitely complex reality it attempts to represent.

What Lies Ahead?

The simulation of literary output, as demonstrated, isn’t a path to definitive answers, but a carefully constructed echo chamber. Each successful counterfactual isn’t a revelation of ‘what if,’ but a testament to the model’s ability to convincingly appear to explore it. The illusion of historical accuracy will prove a persistent ghost in these machines; every prompt engineered towards realism is, implicitly, a negotiation with the inherent biases encoded within the language itself. The system isn’t a tool for uncovering hidden truths, but a magnifying glass for the assumptions already present.

The true challenge won’t be generating more text, but discerning meaningful variation. Diversity of output isn’t merely a matter of tweaking parameters; it demands a reckoning with the very notion of ‘style’ – a fluid, context-dependent phenomenon ill-suited to static embedding spaces. The promise of experimental literary research hinges on the ability to move beyond novelty and toward genuine insight, a task requiring far more than clever algorithms.

One suspects that this line of inquiry will reveal, as all architectures eventually do, that the freedom to simulate is inextricably linked to the constraints of maintenance. The cost of ‘what if’ will be paid in endless prompt refinement, data curation, and the slow, creeping realization that order is just a temporary cache between failures. The system will grow, certainly, but not towards control-towards a more complex, more beautiful, and ultimately, more unpredictable form of chaos.


Original article: https://arxiv.org/pdf/2606.02293.pdf

Contact the author: https://www.linkedin.com/in/avetisyan/

2026-06-02 16:03