The style that was learned: large language models, imitation, and the ethics of literary influence

Authors

  • Nomvuselelo N. Nxumalo University of Eswatini
  • Seroja Ainun Nadhifah

DOI:

https://doi.org/10.64595/80qxyz22

Keywords:

Authorship, Large language models, Literary influence, Literary style imitation, Textual reproduction

Abstract

Background: Large language models have transformed literary imitation into a scalable practice, complicating distinctions among stylistic influence, textual reproduction, authorship, and ethical appropriation. Objective: This study examines how author-conditioned generation produces stylistic convergence, preserves or transforms source relations, and acquires ethical legitimacy under documented research conditions. Method: This study applies paired stylometry, intertextual reproduction auditing, and an ethical literary influence matrix to ten public-domain literary excerpts, ten controlled model outputs, five methodological documents, and seven contextual sources. Results: Stylometric comparison reveals graded convergence, with recognizable authorial cues coexisting with substantial divergence in sentence architecture and other distributed features. Reproduction auditing identifies low lexical overlap, no shared four-grams, no source-character or source-event transfer, and consistently high transformation across all pairs. Ethical coding classifies risk as low because sources are public domain, authors are deceased, prompts and exemplars are disclosed, and human control is documented. Implication: These findings indicate that literary resemblance should not be treated automatically as authorship, memorization, plagiarism, or infringement, but evaluated through layered textual and contextual evidence. Novelty: This study advances a multidimensional framework that theorizes learned style as mediated literary influence shaped by convergence, transformation, provenance, and accountable use across computational, literary, legal, and contemporary ethical domains.

Downloads

Download data is not yet available.

References

Abdillah, Y. A. (2026). Poetics of algorithmic excess: Digital aesthetics in Indonesia’s Twitter poetry bot. Lingua Technica: Journal of Digital Literary Studies, 2(1), 86–101. https://doi.org/10.64595/lingtech.v2i1.138

Alvero, A., Lee, J., Regla-Vargas, A., Kizilcec, R. F., Joachims, T., & Antonio, A. (2024). Large language models, social demography, and hegemony: Comparing authorship in human and synthetic text. Journal of Big Data, 11. https://doi.org/10.1186/s40537-024-00986-7

Chen, T., Asai, A., Mireshghallah, N., Min, S., Grimmelmann, J., Choi, Y., Hajishirzi, H., Zettlemoyer, L. S., & Koh, P. W. (2024). CopyBench: Measuring literal and non-literal reproduction of copyright-protected text in language model generation. arXiv. https://doi.org/10.48550/arXiv.2407.07087

Chen, Z., & Moscholios, S. (2024). Using prompts to guide large language models in imitating a real person’s language style. arXiv. https://doi.org/10.48550/arXiv.2410.03848

Choksi, M. Z., & Goedicke, D. (2023). Whose text is it anyway? Exploring BigCode, intellectual property, and ethics. arXiv. https://doi.org/10.48550/arXiv.2304.02839

Dathathri, S., See, A., Ghaisas, S., Huang, P.-S., McAdam, R., Welbl, J., Bachani, V., Kaskasoli, A., Stanforth, R., Matejovicova, T., Hayes, J., Vyas, N., Merey, M. A., Brown-Cohen, J., Bunel, R., Balle, B., Cemgil, T., Ahmed, Z., Stacpoole, K., . . . Kohli, P. (2024). Scalable watermarking for identifying large language model outputs. Nature, 634, 818–823. https://doi.org/10.1038/s41586-024-08025-4

DeHaan, S., Liu, Y., Bollen, J., & Blanco, S. A. (2025). GPT editors, not authors: The stylistic footprint of LLMs in academic preprints. arXiv. https://doi.org/10.48550/arXiv.2505.17327

Fawaid, A. (2025). Mapping the field of digital literary studies: Concepts, methods, and emerging directions. Lingua Technica: Journal of Digital Literary Studies, 1(1), 1–13. https://doi.org/10.64595/99c6t274

Karamolegkou, A., Li, J., Zhou, L., & Søgaard, A. (2023). Copyright violations and large language models. arXiv. https://doi.org/10.48550/arXiv.2310.13771

Khan, A., Wang, A., Hager, S., & Andrews, N. (2023). Learning to generate text in arbitrary writing styles. arXiv. https://doi.org/10.48550/arXiv.2312.17242

Konen, K., Jentzsch, S., Diallo, D., Schutt, P., Bensch, O., Baff, R. E., Opitz, D., & Hecking, T. (2024). Style vectors for steering generative large language models. arXiv. https://doi.org/10.48550/arXiv.2402.01618

Liu, X., Sun, T., Xu, T., Wu, F., Wang, C., Wang, X., & Gao, J. (2024). SHIELD: Evaluation and defense strategies for copyright compliance in LLM text generation. arXiv. https://doi.org/10.48550/arXiv.2406.12975

Mann, S. P., Earp, B., Møller, N., Vynn, S., & Savulescu, J. (2023). AUTOGEN: A personalized large language model for academic enhancement—Ethics and proof of principle. The American Journal of Bioethics, 23, 28–41. https://doi.org/10.1080/15265161.2023.2233356

McCoy, T. R., Smolensky, P., Linzen, T., Gao, J., & Celikyilmaz, A. (2023). How much do language models copy from their training data? Evaluating linguistic novelty in text generation using RAVEN. Transactions of the Association for Computational Linguistics, 11, 652–670. https://doi.org/10.1162/tacl_a_00567

Mijatović, A., Žuljević, M., Ursić, L., & Marušić, A. (2026). Responsible use of large language models in manuscript preparation. Current Protocols, 6. https://doi.org/10.1002/cpz1.70300

Mikros, G. (2025). Beyond the surface: Stylometric analysis of GPT-4o’s capacity for literary style imitation. Digital Scholarship in the Humanities, 40, 587–600. https://doi.org/10.1093/llc/fqaf035

Müller, F., Görge, R., Bernzen, A. K., Pirk, J. C., & Poretschkin, M. (2024). LLMs and memorization: On quality and specificity of copyright compliance. arXiv. https://doi.org/10.48550/arXiv.2405.18492

Nguyen, T., Hu, Y., & Le, T. (2025). Unraveling interwoven roles of large language models in authorship privacy: Obfuscation, mimicking, and verification. arXiv, 14903–14919. https://doi.org/10.48550/arXiv.2505.14195

Project Gutenberg. (n.d.). The Project Gutenberg license. https://www.gutenberg.org/policy/license.html

Sag, M. (2023). Copyright safety for generative AI. SSRN Electronic Journal. https://doi.org/10.2139/ssrn.4438593

Smith, R. (2025). Licensing of text for generative AI: Learnings from non-AI licensing practices. The Columbia Journal of Law & the Arts, 48(4). https://doi.org/10.52214/jla.v48i4.13926

Sourati, Z., Ziabari, A. S., & Dehghani, M. (2025). The homogenizing effect of large language models on human expression and thought. Trends in Cognitive Sciences. https://doi.org/10.48550/arXiv.2508.01491

Toshevska, M., & Gievska, S. (2025). LLM-based text style transfer: Have we taken a step forward? IEEE Access, 13, 44707–44721. https://doi.org/10.1109/ACCESS.2025.3548967

Tripto, N., Venkatraman, S., Macko, D., Móro, R., Srba, I., Uchendu, A., Le, T., & Lee, D. (2023). A ship of Theseus: Curious cases of paraphrasing in LLM-generated texts. arXiv. https://doi.org/10.48550/arXiv.2311.08374

Wahle, J. P., Ruas, T., Kirstein, F., & Gipp, B. (2022). How large language models are transforming machine-paraphrase plagiarism. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (pp. 952–963). Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.emnlp-main.62

Wang, Z., Tripto, N., Park, S., Li, Z., & Zhou, J. (2025). Catch me if you can? Not yet: LLMs still struggle to imitate the implicit writing styles of everyday authors. arXiv, 10040–10055. https://doi.org/10.48550/arXiv.2509.14543

Weerasinghe, J., Seepersaud, O., Smothers, G., Jose, J., & Greenstadt, R. (2025). Be sure to use the same writing style: Applying authorship verification on large-language-model-generated texts. Applied Sciences, 15(5), Article 2467. https://doi.org/10.3390/app15052467

Xu, J., Li, S., Xu, Z., & Zhang, D. (2024). Do LLMs know to respect copyright notice? In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 20604–20619). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.emnlp-main.1147

Downloads

Published

31-07-2026

How to Cite

Nomvuselelo N. Nxumalo, & Seroja Ainun Nadhifah. (2026). The style that was learned: large language models, imitation, and the ethics of literary influence. Lingua Technica: Journal of Digital Literary Studies, 2(2), 123-140. https://doi.org/10.64595/80qxyz22

Similar Articles

11-20 of 20

You may also start an advanced similarity search for this article.