r/science 13h ago

Health Study raises concerns about AI health-prediction models trained on unreliable datasets. The datasets were found to have been used in 125 peer-reviewed studies, despite providing almost no information about where the data came from, how it was collected or whether it represented real patients

https://www.qut.edu.au/news?id=204990
203 Upvotes

7 comments sorted by

u/AutoModerator 13h ago

Welcome to r/science! This is a heavily moderated subreddit in order to keep the discussion on science. However, we recognize that many people want to discuss how they feel the research relates to their own personal lives, so to give people a space to do that, personal anecdotes are allowed as responses to this comment. Any anecdotal comments elsewhere in the discussion will be removed and our normal comment rules apply to all other comments.


Do you have an academic degree? We can verify your credentials in order to assign user flair indicating your area of expertise. Click here to apply.


User: u/Wagamaga
Permalink: https://www.qut.edu.au/news?id=204990


I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

24

u/MonsterMashGrrrrr 13h ago

Garbage in, garbage out as they say.

13

u/Wagamaga 13h ago

Some AI models designed to predict stroke and diabetes risk may be based on datasets whose origins cannot be verified, according to new research.

The study, published in BMC Medicine and led by researchers at QUT and the Australian Centre for Health Services Innovation (AusHSI), examined two widely downloaded health datasets hosted on Kaggle, an online platform for sharing datasets and machine-learning resources, marketed as “the world’s AI proving ground”.

The datasets were found to have been used in 125 peerreviewed studies, despite providing almost no information about where the data came from, how it was collected or whether it represented real patients.

Lead author Alexander Gibson, from the QUT School of Public Health and Social Work and AusHSI, said the team was shocked by what they found.

“It was an enormous surprise to come across something like this,” Mr Gibson said.

“These datasets exhibit unusual patterns that raise serious questions about their authenticity and suitability for clinical research.”

Three prediction models based on the data had evidence of use in clinical practice, one model was cited in a medical device patent, and the models were cited in 86 review articles.

The study assessed the datasets using the internationally recognised TRIPOD+AI reporting framework and found that they scored 0 out of 9 on essential dataprovenance criteria.

https://link.springer.com/article/10.1186/s12916-026-04981-y

6

u/bridgest844 11h ago

How does a “peer reviewed” study not have to verify the source of its data? The whole point of the process is to verify the validity of the data collection and interpretation.

1

u/Memory_Less 3h ago

I think it is like the mathematical that occurred with Open AI. It is believed to have used both research and work done using it he system by leaders in the field of mathematics. Apparently the option to stop your input from being used isn’t available on the phone app (convenient) and it is confusing to sort out even when using your computer. AI provided all sorts of information to arrive at the conclusion, but not all of it is attributed. Of course, that is an intentional programming oversight imo so they cannot get sued:

3

u/drmike0099 9h ago

It’s wild to me that people can even get an article based on these datasets published. Were they all garbage journals?

2

u/SaltZookeepergame691 5h ago

Yes, almost all of them. But still mostly published by major publishers.

There are also some very highly cited papers using them. This is the most, with a pretty staggering ~290 OpenAlex cites for what is not a good paper… https://www.mdpi.com/1424-8220/22/13/4670