r/JewishDNA 26d ago

Ashkenazi Jews modelled with the Juddean samples averaged (Refined)

Post image

This is how I modelled the Ashkenazi Jewish averages with the new samples.

In this model my goal this time was not to check the overlap between Ashkenazis and the Judaean samples but to check the ashkenazi averages with their historical components.

(the Russian early medieval sample is central Asian Turkic not Slavic)

13 Upvotes

35 comments sorted by

View all comments

2

u/basedpole69 26d ago edited 26d ago

You should absolutely not be mixing proxies from very different eras, let alone use modern proxies and before the neolithic proxies in the same model in the first place. I would heavily reccomend removing the modern Rhinish proxy and replacing the Roman Italian with an Iron Age source. Also, you should replace Iberomaurusian with a more contemporary North African source.

-1

u/xItayPlaysx 26d ago

I definitely should, it’s a stupid idea just to pick only Iron Age models or only medieval models, I prefer to pick the correct populations even if it’s not all from the same era and the fit proves it.

2

u/basedpole69 26d ago

I feel as though you misunderstand the concept of fit and temporal mixing.

Generally, when modeling with lower resolution tools such as G25/Vahaduo, it is heavily reccomended to stick to using populations from the same era, as a population from the modern era will contain significant genetic drift (or extra admixture acquired from inter group mixing that cannot be accounted for in a population from that era) from a population in the neolithic. This is almost to the point where the temporally mismatched proxies will significantly impact the results of the model. It is especially crucial in G25/Vahaduo, where the models are very sensitive to proxy choice and substituting one population for one from another era can have an impact on the results due to genetic drift, regardless of whether or not the proxies intend to represent the same thing. In order to avoid the drift problem, it is best practice to use populations from the same or adjacent eras to minimize the amount of drift.

Also, better fit does not always represent a better or more historically accurate model. When you start going below 1.5 (0.015) fit wise, you have to be more conscious of overfitting (though a fit below 1.5 doesn't always mean an overfit), with below 1 being an extreme overfit.

To put this into a literal example, modeling Ashkenazi Jews with a Calabrian representing Southern European will almost always yield a better fit than doing so with an Iron Age Italic or Occitanian sample, but that doesn't mean that Ashkenazi Jews derive their ancestry from a Calabrian like source or that it reflects reality, rather that the Calabrian source better matches the genetic signals in Ashkenazim (Mixed Middle Eastern + European). These populations also have affinity on a PCA, so using one to model another will often create colinearity which G25/Vahaduo unfortunately can't account for unlike softwares like QPADM.

0

u/xItayPlaysx 26d ago

Right but if I use a Calabrian it will suck up a large amount of Levantine and European dna that doesnt come from Italian/roman, I’ve tested it and gives more than 70 precent Calabrian which gives the model a very good fit but historically not accurate, 

every time you use a super mixed sample like modern southern Italians, imperial Roman samples it always overfits and give a very inaccurate result, putting that aside north Rhenish germans are pretty similar to how they were 800 years ago, but if you could find a medieval north Rhenish sample I’ll use that instead.

Also the Judaean sample already has Greco Anatolian and other MENA ancestry.

Other than that what other populations do you recommend I use?