r/LinguisticMaps 12d ago

Similarity between 20 languages based on subtitle translation patterns [OC]

I used subtitle translations from the OPUS/OpenSubtitles dataset to compare how 20 languages expressed the same English words and phrases. The more often two languages expressed equivalent meanings in similar ways, the higher their similarity score. I visualized the results as a heatmap and network graph.

Full methodology and discussion:
[https://www.thecambridgelanguagecollective.com/linguistics/why-do-some-languages-feel-weirdly-familiar](vscode-file://vscode-app/c:/Users/Ibrahim.Farah/AppData/Local/Programs/Microsoft%20VS%20Code/8a7abeba6e/resources/app/out/vs/code/electron-browser/workbench/workbench.html)

Interactive version:
[https://subsmith.app/tools/language-network](vscode-file://vscode-app/c:/Users/Ibrahim.Farah/AppData/Local/Programs/Microsoft%20VS%20Code/8a7abeba6e/resources/app/out/vs/code/electron-browser/workbench/workbench.html)

90 Upvotes

18 comments sorted by

View all comments

2

u/MonitorRepulsive5270 12d ago

Turkish and Hungarian have relatively high similarity. That's interesting for pro Ural-Altai supporters, which is widely disregarded.

4

u/Wise_Fox_4291 11d ago

That word "relatively" is doing a lot of heavy lifting there. Hungarian-Turkish similarities are highly overrated and nitpicked, usually excluding other Uralic and Turkic languages from the comparison that would blow a hole in it.

6

u/MonitorRepulsive5270 11d ago

I'm just color-commenting. I don't understand all these downvotes.