Corpus Realism
"It's easier to imagine the end of human readers than the end of writing."

In May 2023, the Organization for Transformative Works — the nonprofit behind the open-source fanfiction repository Archive of Our Own — published a post explaining what it could and couldn’t do in response to its writers’ work being scraped for AI training.
They begin with a clear list of actions taken to address their community’s concerns — namely, rate limiting and traffic monitoring — followed by an equally upfront concession: “Putting systems in place that attempt to block all scraping would be difficult or impossible without also blocking legitimate uses of the site.” The post then shares another measure already taken as consolation: after discovering that data from AO3 had been incorporated into Common Crawl, a popular dataset used to train AI, the site’s organizers added code back in December 2022 “requesting Common Crawl not scrape the Archive again.” Then, after all these demonstrations of competence and transparency, this:
“We cannot go back in time to stop data collection that already occurred, or remove AO3’s content from existing datasets, as much as we may dislike that it happened. All we can do is attempt to reduce such collection in the future.”
Few online institutions can rival OTW’s nearly two-decade track record of meaningful advocacy on behalf of amateur writers, and it would be a mistake to read this as capitulation. Everything else in this document is clearly the product of an organization that has no plans of giving up on those it is built to support. Still, the shape of this particular statement warrants careful inspection: resistance and concession arriving in the same line, with the latter doing significant load-bearing work. The past is gone, while the corpus is frustratingly closed to withdrawals and remains even-more-frustratingly open to deposits. All that’s left is the future, and our first glimpse of it arrives with limits already imposed.
What is corpus realism?
Mark Fisher called capitalist realism the condition in which it’s easier to imagine the end of the world than the end of capitalism — a phrase he credited to Fredric Jameson and Slavoj Žižek1 — and the operative word here is “condition.” Almost no one defends capitalist realism as a position, because it’s not the kind of position you can choose to take. Rather, it’s more like a horizon of possibility that limits what we can picture. The condition I want to talk about belongs to the same family.
Corpus realism is the condition in which it is impossible to imagine any writing that will not eventually be assimilated by machines.
The term corpus serves double duty here: a corpus is the body of work a writer produces themselves, but it’s also the term for the vast, ecumenical pile of data that LLMs are trained from.
The empirical half of this concept is not hard to check. Recent crawl data from Graphite puts the share of new online articles that are primarily AI-generated at roughly half—a trend which can easily be read as the opening act of the “textpocalypse” that Matthew Kirschenbaum warned about back in March 2023. At the same time, every reasonable measure of human reading activity — that is, long-form reading for leisure — seems to be pointing the other way. An iScience study logging twenty years of the American Time Use Survey found that the share of Americans who read for pleasure on a given day fell from 28% in 2004 to 16% in 2023. In December 2025, YouGov found that four in ten Americans read no books at all—while 82% of all books still being read were read by the top 19% of readers. Together, these figures paint a picture straight out of the Library of Babel: there is more writing than ever before, addressed to fewer and fewer people who actually bother to read it.
To demonstrate the other half — corpus realism as something experienced, rather than believed — here’s a thought experiment: try to imagine publishing a piece of writing that stays outside the corpus, not hidden, paywalled or locked to registered users, but open and freely available. Whatever you reach for will be a countermeasure, which runs into a familiar asymmetry. This kind of defense has to hold everywhere, against everyone, forever, but a crawl only has to work once—and if it does, there’s no take-backs.
Nearly three years after OTW asked Common Crawl to stop crawling AO3, Alex Reisner reported in The Atlantic that the foundation behind the dataset had told publishers their content was being removed on request while its own logs showed no content files had been modified since 2016. Sites like AO3 may well have had their requests to limit scraping honored, but anything that’s already been scraped is in there forever—or so Common Crawl’s response went, because its archive is stored in a format that is immutable by design, from which nothing can be deleted. For better or worse, we are now in a world where, as Holly Herndon and Mat Dryhurst titled the monograph tie-in to their 2024 Serpentine exhibition, all media is training data.
This is what makes this concept a state of realism, rather than mere description or metaphor. Belief in corpus realism is fully compatible with hating it, and the clearest statement of that on record comes from one of the organizations most vocally devoted to fighting it. Like capitalist realism, it is a condition that none of us have chosen to experience.
Why corpus “realism”?
If Fisher’s term is worth borrowing here, it deserves to be borrowed whole. The point of diagnosing realism, in the way that Fisher uses that word, is to point out that the necessity it names is at least partly false. Of the three layers of “exit” on the table in this case, only one feels like a viable candidate.
First, there’s individual exit — staying publicly readable while staying out of the corpus — which, as discussed above, is impossible. There’s also the possibility of withdrawing into print runs and private channels — which, to be clear, is eminently possible — but this isn’t an exit as much as a temporary abstention or reduction; your defense is only as good as the people in your channel, and all it takes is one screenshot or photocopy for you to be back at square one. What corpus realism forecloses is a much less discussed third thing: the idea that we might collectively renegotiate the terms of AI assimilation itself.
Nothing at this third layer is impossible; the infrastructure to make it happen exists, from collective licensing to data trusts. It’s just hard to imagine, which is corpus realism doing its thing. The foundation behind Common Crawl has gone on record in defense of its arguably underhanded data-scraping and retention tactics by claiming that its archive is immutable by design—but “by design” also means “by decision,” and decisions can be reversed.
Assimilation-without-consent is a policy default in the guise of physics, and the people who see the condition most clearly are already treating it accordingly. In the same OTW post that concedes the past, the foundation shares that they’ve argued to no less than the U.S. Copyright Office that users “should be allowed to opt out from having their works incorporated into AI training sets.” Meanwhile, Herndon and Dryhurst are also co-founders of Spawning, a startup that develops genuine consent infrastructure for training data.
Why “corpus realism”?
The house that Fisher built already has more than one AI-themed room. Dan McQuillan has used the term “AI realism“ to describe the feeling that AI’s deployment can never be seriously questioned or reversed, but only reformed. This, too, is a real condition, but it’s distinct from the one I’m talking about. If AI realism is the inability to imagine a future without AI, corpus realism is the inability to imagine not being assimilated by it. Unlike AI realism, corpus realism survives even in those who have still proudly never used a chatbot, because assimilation doesn’t require participation. All you have to do is publish or publicly post something.
On the reception side of the AI writing equation, Hannes Bajohr has added a second room: the “post-artificial” text. Bajohr uses this phrase to describe the slow but steady erosion of every human reader’s millennia-old default assumption that any unknown piece of writing was produced by another human, eventually reaching a tipping point where the provenance of any text becomes irrelevant. Corpus realism adopts a similar line of thinking on behalf of writers. Bajohr says you can no longer assume a human wrote a given piece of text. I’m saying you can no longer assume a machine won’t — or hasn’t already — absorbed it.
That sense of inevitable assimilation cements this as a subspecies of Fisher’s thinking. Like capitalist realism, corpus realism is also a matter of enclosure. Critics like Nick Couldry and Ulises Mejias have tracked this same movement as “data colonialism”: a commons of published text converted into private training data, at a scale where the enclosed party doesn’t realize what’s going on until it’s too late to do anything—if there was ever anything for them to meaningfully do. When AO3’s fanfic writers found out they’d been scraped, they weren’t being shut out of a market. They were simply discovering that they no longer had the ability to write in public any other way.
Diagnosis, tragedy and calling
At the heart of corpus realism is a problem of attribution, and that goes much deeper than citing sources by name. Unlike search engines, which are pretty good at reading everything and coming back with receipts, LLMs metabolize data a little differently. When a piece of text is assimilated into a model, there’s no way of tying that specific text back to any subsequent piece of text it’s used to shape—let alone when a model develops a cadence based on a million pieces of text written by a million different human writers. The problem with AI attribution is not that it’s withheld — LLMs frequently invoke certain authors by name, including, curiously enough, Mark Fisher himself — but that it’s often structurally impossible. From this problem of attribution, we get another Borges-style conundrum: if everything is assimilated and nothing is attributable, writing acquires infinite reach and zero credit. How you react to that depends on what you value more.
If you value credit, it may read as a tragedy—your work has been stripped, along with your name, and your labor is now feeding something built (or, depending on your tolerance for AI hype, hollowly marketed) to be your replacement. However, if you value reach, it could read as a calling: at a time when “nobody reads anymore,” every word you write is now guaranteed at least one reader, forever. The difference between these two conclusions is not due to a misinterpretation of facts: both stem from the same condition and underlying variables. In the same way a leftist organizer can invoke capitalist realism as a sign of capitalism’s entropic potential while a hustlepreneur can invoke it as a reason to “lock in” more, corpus realism can be used by two people in complete agreement about what is happening and complete disagreement about what to do about it.
You’re right—this is water
This is not my first brush with corpus realism: much of my recent work has dealt with it at length, before I had a name for it.
In a recent essay for INC Longforms, I argued that 4chan greentext might represent a viably “AI-proof” form of writing because nothing is ever stored on the page where a counterfeiter could reach it. What I didn’t realize then, but realize now, is that the underlying question of greentext is not what is authentic, but what survives assimilation: given that text will be consumed, what property should that text have?
In another recent piece, I explored the new trend of AI forecasting whitepapers seemingly written for what I called a “second reader” — one recently went as far as publishing a version of itself explicitly optimized “for AI agents & crawlers” — and concluded it was a mild case of hyperstition. I still think this is true, but now I know that this channel exists only because everyone has started to assume assimilation in advance. No one builds a door for a guest they’re not expecting.
Most recently, in a third piece, I examined the curious new range of post-AI proof-of-person rituals that some writers have begun to perform to appease another reader: AI-powered detector programs like Pangram. In passing, I noted that because these maneuvers are public, they are all “already bound for the next batch of training data,” stuck in yet another AI-powered arms race. What I couldn’t say then, but can now, is why the race can’t be won: in the corpus, it’s all just another deposit.
Form that survives assimilation, address to the assimilator, and evasion assimilated are all symptoms of the same condition in action, and I don’t think I would have been able to name the water if I hadn’t spent a little time swimming in it.
By now, some of you may have noticed the presence of a fairly large elephant in the room. Corpus realism rests on a truth as immutable as Common Crawl’s archive — assimilation does not guarantee attribution — and here I am, trying to coin a term, in a totally transparent bid to have that term metabolized by The Discourse with my name still attached. I don’t think this undermines what I’m trying to get at, though; if anything, it feels like another example of the concept in practice.
If everything is assimilated, the only remaining question is how much of yourself to knowingly offer up. Those who see a calling answer with volume, while those who see a tragedy answer with nothing, or as little as possible. Corpus realism doesn’t tell you which answer is right—only that every decision that led to this one was made without our collective consent.
Continuing the chain of attribution, Jameson introduced the line as something “someone once said,” which means that the founding quotation of capitalist realism is, itself, unattributed.



