Understanding how authors weave references to other texts into their work is a key part of literary scholarship, but it can be extremely challenging—especially when those references are indirect, paraphrased, or translated. A newly published research paper tackles this problem by focusing on biblical allusions in Karen Blixen’s Seven Gothic Tales, using advanced computational methods to detect these subtle intertextual connections. This study is important because it explores how artificial intelligence can assist scholars in uncovering layers of meaning in literature that might otherwise go unnoticed.
Key Takeaways
- The researchers created a benchmark dataset of 189 annotated biblical references identified in Seven Gothic Tales, drawing from expert commentary.
- They tested various text retrieval methods, including traditional keyword-based approaches (TF-IDF and BM25) and modern AI sentence encoders, to find these references within over 31,000 verses of Danish Bible translations.
- Linguistically normalized BM25—a method that accounts for variations in language—performed strongly without prior training, especially in retrieving direct quotations.
- Fine-tuning a large Danish language model significantly improved detection of subtle allusions, more than doubling its accuracy on these challenging references.
To address the challenge of spotting biblical references in Blixen’s stories, the team first compiled a comprehensive set of known references from scholarly annotations. These included direct quotations, paraphrases, and more vague allusions, all written in Danish translations of the Bible from Blixen’s era. The goal was to see how well different computational techniques could retrieve these references from the vast body of biblical text.
The researchers compared traditional information retrieval methods like TF-IDF and BM25, which rely on matching words and phrases, with newer “dense” models—AI systems that transform sentences into numerical representations (called embeddings) to capture deeper semantic meaning. They also experimented with linguistic normalization, a process that standardizes words to reduce the impact of spelling or grammatical variations, which is especially helpful when working with historical language and translations.
One key innovation was fine-tuning a large Danish sentence encoder using “hard negatives” — examples that are difficult to distinguish from true references. This training helped the model better recognize subtle, indirect biblical allusions that simple keyword matching might miss. The team evaluated performance by measuring how often the correct biblical verse appeared among the top 10 results returned by each method.
Interestingly, while traditional BM25 with linguistic normalization provided a strong baseline—successfully retrieving all direct quotations in the top 10 results—the fine-tuned AI model excelled at finding more nuanced allusions. However, the researchers caution that neither approach is perfect. When literary scholars reviewed some model “false positives” (cases the model flagged but weren’t in the original annotations), they found several meaningful references that experts had not previously documented. This suggests the models can serve as helpful “co-readers,” proposing candidate references for scholars to explore further, rather than replacing expert analysis.
These findings highlight both the promise and the limits of computational tools in literary studies. By automating parts of the intertextual search process, AI can save time and uncover hidden connections, but human expertise remains essential for interpretation and validation. Going forward, combining AI retrieval models with expert close reading could open new avenues for understanding complex literary works and their rich networks of influence.
Based on research published on arXiv by András Kovács, Alexander Conroy, Daniel Hershcovich et al..
