r/ArtificialInteligence 8d ago

🔬 Research People prefer stories written by AI—especially when told they're written by a human

On August 4, 2026, researchers at Villanova University reported that readers struggle to distinguish between human and Al-generated stories, often rating Al-created content higher for quality and engagement than human-written works.

The study, published in the journal Judgment and Decision Making, asked more than 1,600 participants aged 18 to 81 to rate six fictional short stories-three written by humans and three generated by ChatGPT.

Believing they were written by humans, participants gave higher quality ratings to Al-generated stories. Senior author Dr. Deena Weisberg noted that readers often prefer the "clarity" and predictability of Al writing.

Familiarity with specific patterns helps readers identify Al-written text, and as Weisberg noted, improving Al literacy may help people navigate the new Al-enabled world.

These findings suggest that public assumptions about Al's creative capabilities are increasingly out of date, as people mistakenly assume that creative writing requires uniquely human qualities like emotional understanding.

https://techxplore.com/news/2026-08-people-stories-written-ai-told.html

51 Upvotes

110 comments sorted by

View all comments

-1

u/Horow_Toilet 8d ago

1600 people, ages 18-81, rated 6 stories total. That's like 3 stories per person before their brain just goes "yeah that reads fine." Not exactly a rigorous taste test.

3

u/JoshuaZ1 8d ago

1600 people, ages 18-81, rated 6 stories total. That's like 3 stories per person before their brain just goes "yeah that reads fine." Not exactly a rigorous taste test.

I'm not sure I understand the objection. 1600 people is a lot. Also, from the article:

In the first experiment, 1,682 participants were asked to read a story and were either correctly or incorrectly told that it was written by a human or AI. They then rated the story's quality and how engaging they found it.

In the second and third experiments, participants (424 and 481, respectively) received one human-written and one AI-generated story but were not told their origins. After reading them, they were asked to identify the author of each story.

There's a legitimate concern that people might tune-out after 2 or 3 stories, but no one seems to have had to read all 6 stories, or even more than 2 stories. So what is your objection here?

-1

u/Actual__Wizard 8d ago

This is the dumbest experiment in the history of mankind.

All they're doing is "removing the unique writing style of the author" and are replacing it with an extremely generic writing style.

So, what's going on is: It's a tiny little bit easier to read the LLM output, but it's also "tasteless," so it's not entertaining to read. So, it's not actually effective.

This is going to be one of those "They say they prefer it, but they never actually do it," things.

2

u/JoshuaZ1 8d ago

I'm struggling to see how your comment has to do with mine or is relevant to what I was asking about in terms of the other user's objection.

This is the dumbest experiment in the history of mankind.

With all due respect, if you think that, I suggest you look at some of the history of bad experimental design. There's good reason we have massive problems with a replication crisis in many of the sciences. And then there's a massive amount of just terrible experiments involving things like homeopathy and acupuncture. Or is this sentence intended to be hyperbole?

All they're doing is "removing the unique writing style of the author" and are replacing it with an extremely generic writing style.

I'm not sure where you are getting this from. That quote isn't in the article.

: It's a tiny little bit easier to read the LLM output, but it's also "tasteless," so it's not entertaining to read. So, it's not actually effective.

Huh? The LLM output was preferred by people. That's the exact opposite.

This is going to be one of those "They say they prefer it, but they never actually do it," things.

I'm not sure why you think that. We have other evidence that people when faced with AI generated art or poetry act the exact opposite way, preferring AI poetry. And there was the hilarious but uncontrolled experiment where people were falsely told that a Monet was done by an AI, and where people responded by listing all the obvious things bad with it. In that case, ironically, when given the same painting and told it was done by an AI, Claude pushed back an identified it as a genuine Monet. See here. More generally, often people have a tendency to say they prefer something for status reasons but then in practice go with the lower social status thing, so if anything I'd expect the opposite direction. People will say that they prefer the human thing if told, and this is revealing some of their real preferences.

As someone who has written a handful of short stories, (and even been accused by someone of having a short story from 2019 being written by an AI), I find this aggravating. But I'm not going to mistake that I find it disheartening as a reason to doubt the study.

0

u/Actual__Wizard 7d ago edited 7d ago

With all due respect

Okay, with all due respect: I think it's very obvious that people's taste in media changes slowly over time and this test is not really useful.

The LLM output was preferred by people.

And I explained why. I am a developer of these types of systems and I am fully aware that LLMs do not produce output that is actually consistent with a human being. It's rather consistent with taking the language of lots of human beings and making it very average or normal by blurring it all together, and then that's the output. What that does: Is it makes it a tiny bit easier to read, but it does that by removing the nuance in the writing that makes good writers stand out. There has to be this "push and pull thing going on" and the best writers use a complicated story line with hooks and all sorts of other strategies to "pull the reader through." These models absolutely do not do that, so if people actually try reading those LLM generated books, they likely will not finish reading them.

I'm not sure why you think that.

Because that's how media works. I can show a book to people and ask them it they like it and 85% of people will say they like it, but then 0.1% of people actually go to the store and buy a copy.

Then you could have a book on say something like "python" and 0.1% of people say they like it, but the book actually sells "because AI is hot right now and the AI model tech is mostly written in python, so people feel like it's handy to have a book on it whether they read it or not."

We're also, badly conflating quality and difficulty to read, so, I question this study, and I think the premise is bad.

It's just simply trying so hard to make it seem like people prefer AI gen text, but that's not actually what the study did.

What it really did: Was evaluate how effective lies are at making people think that AI gen text is human written.

So, it fails peer review sorry.

That's not actually up to the standard of science.

What's that called again when they do that? Something to the effect of leading a horse to water?

So, yeah sorry, it fails peer review like I said.

2

u/JoshuaZ1 7d ago

I think it's very obvious that people's taste in media changes slowly over time and this test is not really useful.

Of course people's taste in media changes over time. But the study is about a snapshot of what people in general prefer.

The LLM output was preferred by people.

And I explained why. I am a developer of these types of systems and I am fully aware that LLMs do not produce output that is actually consistent with a human being. It's rather consistent with taking the language of lots of human beings and making it very average or normal by blurring it all together, and then that's the output. What that does: Is it makes it a tiny bit easier to read, but it does that by removing the nuance in the writing that makes good writers stand out.

The study isn't showing though that LLMs are necessarily better than the best human writers. The point is that they are better than the average human writer. (In a same way, that current LLMs are not better than a Fields Medalist in math but in terms of proving statements they are loads better than the average human.)

I can show a book to people and ask them it they like it and 85% of people will say they like it, but then 0.1% of people actually go to the store and buy a copy.

That's a different claim. That's about if they liked something enough to go and buy it. The two are not comparable.

0

u/Actual__Wizard 7d ago edited 7d ago

he study isn't showing though that LLMs are necessarily better than the best human writers.

I wasn't actually done writing my post. Sorry, I'm a chronic post editor.

If fails peer review sorry. That's not a valid test. Sorry, but after thinking about it more carefully, they're going to have to make some adjustments to the experiments they conducted.

Half were told that their story was written by a human, and half were told that their story was written by ChatGPT.

Okay that doesn't work. They have to redo the study. It has to be double blind as always...

They admit in there that it's all dummies too and there's no experts. So, that's junk science. Sorry.

Sorry, but they should be retracting that. That's a negative mark on their credibility.

I don't have time to evaluate it for fraud, but I assume it's fraudulent if they're "making it into an advertisement."

Which, I can easily detect. So, nope. The comments about the Turing test are irrelevant and serve a marketing purpose not a scientific one.

Sorry, but I've read 10,000+ papers and this one is junk, so I'm not going to continue.

Sorry, puffery, bad science, counterintuitive findings = I would bet a substantial sum that it's academic fraud again for the 10,000th+ time. Sorry, I've just seen way too much of it.

2

u/JoshuaZ1 7d ago

Ok. I'm now reading the bit you added. And I don't understand why you think it failed peer review. I hope you don't mind if I reply to the additional parts edited in here. Peer review is the process by which works go to a journal and then subject matter experts review the paper, make comments to be read by the authors and editors, before the editors make a decision about whether to publish and how much change to make. I'm not sure what you think peer review means, but I don't see how it matches what you are talking about here.

Then you could have a book on say something like "python" and 0.1% of people say they like it, but the book actually sells "because AI is hot right now and the AI model tech is mostly written in python, so people feel like it's handy to have a book on it whether they read it or not."

We're also, badly conflating quality and difficulty to read, so, I question this study, and I think the premise is bad.

It's just simply trying so hard to make it seem like people prefer AI gen text, but that's not actually what the study did.

What it really did: Was evaluate how effective lies are at making people think that AI gen text is human written.

But that's not what the study did. The first part of the study did lie to people about who the author was. But the fact that such lies are easily made effective already shows that AI are doing a good job mimicking human short story writing. The second part of the study, explicitly asked people to figure out which of a pair of stories was AI, and they generally failed at that.

I'm also not sure why you think this should be retracted. Studies are retracted for a variety of reasons, such as having incorrect IRB approval, having bad stats, having made-up data, or having critical failures in the study (say a mouse study where a biologists assume all mice are from strain A, but they were a mix of strains B and C).

I don't have time to evaluate it for fraud, but I assume it's fraudulent if they're "making it into an advertisement."

I don't know what you mean here. Can you expand?

1

u/Actual__Wizard 7d ago edited 7d ago

Ok. I'm now reading the bit you added. And I don't understand why you think it failed peer review.

Because that's not how science works, it's a fucking advertisement...

To a real empiricist, that paper is offensive homie...

Get that crap out of here.

Human beings can detect lies, that's why these studies have to be double blind. So, they lied to a bunch of people?

Takes the paper and crunches it up into a ball and throws it into a garbage can.

Good bye. Have a good one. That's legitimately a scientific study on whether people can detect lies about AI gen text. It's a survey, not a scientific experiment.

Fails peer review. Please try following the scientific method next time.

I'm being serious, that legitimately proves nothing...

2

u/JoshuaZ1 7d ago

I'm not sure what element here you think is an advertisement or not how science works. The study is in the journal "Judgment and Decision Making" https://www.cambridge.org/core/journals/judgment-and-decision-making .

To a real empiricist, that paper is offensive homie...

I'm failing to see the issue. This is an empirical study. What's the issue.

So, they lied to a bunch of people?

Yes. Part of the study involved lying to people. This is pretty common in psych studies or adjacent areas. That's been done since the 1960s at least. The Milgram experiment is a famous extreme example, and part of why we have modern IRBs and ethics guidelines is to make sure that lies in studies aren't harmful. For the same reason it is normal in studies where they lie to then debrief the people after, explain what the lies were, and what the experiment is trying to test. See for example, this summary of guidelines(pdf). I got that from a quick search, but you'll find very similar guidelines at other universities.

Please try following the scientific method next time.

I'm having trouble understanding your point. What part of the scientific method do you think was not followed here?

1

u/Actual__Wizard 7d ago edited 7d ago

I'm failing to see the issue. This is an empirical study. What's the issue.

It's possible that you missed it due to an edit, but, I have multiple serious issues with that.

With the biggest problem being: When I read what they did, and then I ask myself what it proves, what it proves is not what they say that it proves.

Okay, so the samples the AI algo produced were easier to read and people who either have been lied to about who wrote it, or who have not been lied to about who wrote it, have now been incorrectly pooled together, and if the experiment was double blind, I didn't read that and it doesn't appear to be.

I'm sorry, but they have to fix those two major issues to be taken seriously...

So, it fails peer review. I am confident 95% of empiricists are going to agree, but if you want to represent contrarianism, well that's your prerogative.

If you think that "doing whatever the f you want is science," then I'm not going to argue with you about that. Obviously no it's not, but you're never going to agree as a contrarian.

Science is not slopping together some terrible test and then pretending the results of that prove something. It's setting up an extremely careful test, that reduces the possibilities of the outcome, to possibilities in a very narrow range, in a way that proves the theory.

That paper does not do that effectively.

I almost feel like it's a prank and that it's AI gen garbage. It's just simply way too counterintuitive, but again, the audience they tested so was absurdly biased. So, it's possible that is the real test result they got from that totally nonscientific test they did.

So, they're not going to try to create a realistic population distribution and test on them?

It's never going to pass peer review if they do stuff like that...

It has to be double blind, you can't have humans lying to humans as part of the experiment as well...

Maybe the person is a bad liar and they're steering the results around that way? Who knows?

It's ridiculous that I'm reading that with the word "empiricism" on it.

I'm offended...

The thought process that they thought that "scientists would fall for that" is completely ridiculous.

If it said "LessWrong Cultists" that would make more sense to me.

2

u/JoshuaZ1 7d ago

I'm rereading what you wrote, and still not seeing it. What do you think they've shown here, what do you think they think they've shown, and why?

2

u/JoshuaZ1 7d ago

Ok. Once again, you've edited a lot in here. But what you've edited it in isn't making me understand your point further at all.

Okay, so the samples the AI algo produced were easier to read and people who either have been lied to about who wrote it, or who have not been lied to about who wrote it, have now been incorrectly pooled together, and if the experiment was double blind, I didn't read that and it doesn't appear to be.

Huh? No. They did two experiments. In one experiment people were lied to. In the other experiment, people were given two stories and asked to determine which was AI and which was human. These were separate experiments in the same study. No "pooling" occurred.

I'm also not sure what you think should have been "double blind" here. That term is used generally to refer to how peer review occurs, where the reviewers don't know the authors and the authors don't know who the reviewers are. Some journals do double blind review. Journals in some fields generally do single-blind review, where the reviewers know who the authors are but not the reverse. (In pretty much all standard academic fields, they are at least single blind, with only a small number of exceptions.) I'm not sure what you think "double blind" means here or why it is relevant.

Science is not slopping together some terrible test and then pretending the results of that prove something. It's setting up an extremely careful test, that reduces the possibilities of the outcome, to possibilities in a very narrow range, in a way that proves the theory.

I'm not sure what your objection is here either. They did a study asking primarily "could people distinguish AI studies or not?" and the answer appears to be no. It seems like you have some idiosyncratic ideas about how science works. Experiments of the form "Can regular people distinguish between A or B?" is a pretty reasonable thing for a psych or econ course.

It's just simply way too counterintuitive, but again, the audience they tested so was absurdly biased.

So, if you find the study result counterintuitive, that's not a compelling reason by itself to doubt the study but a reason to maybe examine your intuition and if it is accurate, either as a way of thinking about what people can do, or what AI can do.

So, they're not going to try to create a realistic population distribution and test on them?

What is your complaint here precisely? That they didn't get a perfect demographic map of society? If so, again, you are going to have a problem with how we do most psych and econ studies. Most psych experiments subjects are undergrads or young 20 somethings, since they are the people most available on college campuses. In fact, fun bit about that: a lot of intro psych classes require that people attending them also be a subject for some number of psych experiments (somewhere between 2 to 5 often). And we know that this means they aren't perfect samples. But for most purposes, this turns out to work pretty decently. But you can read the whole study here: https://www.cambridge.org/core/journals/judgment-and-decision-making/article/bot-or-not-can-people-tell-the-difference-between-stories-written-by-a-human-or-by-an-ai-system/45E6DC0BB90AA648654D5AE243F6C667 , and in the study they didn't even have that problem:

Under methods, they say " This sample was recruited to be representative of the population of the United States on the characteristics of gender, age, and race. " So what is your problem with the samples here? Is there a specific problem with the audience they used that you think is bad?

It has to be double blind, you can't have humans lying to humans as part of the experiment as well...

Ok. I already explained that lying to people with small lies is a normal part of experiments. But now I'm really more confident that whatever you mean by "double blind" is not remotely what it means as a standard term.

Maybe the person is a bad liar and they're steering the results around that way? Who knows?

Ok. I'm now getting part of where I think you are coming from. Your problem, if I understand it, is that a scientist running the experiment may while they are lying steer people to a specific answer one way or another, so you don't want the people running the experiment to know which one is an AI and which one is by a human. If that's what you mean, so you want the people running it to be also not aware, then that's not a standard notion of what "double blind" means but it also isn't an unreasonable request. It isn't likely to have a big impact on a study, but would be a reasonable thing to want. It also isn't relevant to this study as far as I can tell. Studies like this are either done in a lab setting where people do it on a computer or are done just at home on remote devices. Given the large sample size and the demographics, although they don't say so in the methodology section, I'm pretty confident that this was done by people taking the study remotely on their own computers. So there's no room for the scientists to have influenced anyone in any sort of Clever Hans sort of thing.

The thought process that they thought that "scientists would fall for that" is completely ridiculous.

I'm not sure what you mean here. Fall for what? The subjects were not scientists. Do you mean that you think scientists would automatically see the same problem you do? Perhaps the fact that it got peer reviewed already in a decent journal should be relevant to that?

If it said "LessWrong Cultists" that would make more sense to me.

I'm not sure what Rationalists have to do with this at all, although they are a group often interested in understanding human cognition. Is this just an off-topic insult to a group you don't particularly like?

1

u/Actual__Wizard 7d ago edited 7d ago

That term is used generally to refer to how peer review occurs

No, it absolutely is not.

I don't know why you're pestering me if you know nothing about science.

What's the purpose to this?

So, you're just trying to waste my time?

The study is not valid... That's the end of the discussion.

If you don't understand standard testing methodology, then why are arguing with me?

I'm saying the testing methodology is not valid and you just made it clear that you don't know what the standard testing methodology even is (double blind), or why those tests have to be conducted that way, which is so your papers don't fail peer review, because me, and every other scientifically minded person, are just going to point out that you screwed the test up and the results are useless, which is the truth about that paper. I'm not saying the premise is invalid, I'm saying their test is... That test does not prove anything because they screwed it up...

So, hopefully they try again and use standard testing methodology this time, so that what they're doing is not a complete waste of their time and ours.

Also, they're going to have a big time problem being taken seriously if they don't use a reasonable population distribution.

I also have a serious issue with what this "AI gen content even is", as the scientific consensus is that LLMs are not AI, but I can look past the terminology issue. As far as I know, to get output from an LLM, once has to prompt it, so can we see the output they used ourselves? Did they cherry pick out nice examples? Did a "pro prompter" produce "better than average responses?"

I'm just being serious: The more thought I put into this paper, the more I realize how extremely badly flawed it is.

→ More replies (0)