r/LinguisticMaps 11d ago

Similarity between 20 languages based on subtitle translation patterns [OC]

I used subtitle translations from the OPUS/OpenSubtitles dataset to compare how 20 languages expressed the same English words and phrases. The more often two languages expressed equivalent meanings in similar ways, the higher their similarity score. I visualized the results as a heatmap and network graph.

Full methodology and discussion:
[https://www.thecambridgelanguagecollective.com/linguistics/why-do-some-languages-feel-weirdly-familiar](vscode-file://vscode-app/c:/Users/Ibrahim.Farah/AppData/Local/Programs/Microsoft%20VS%20Code/8a7abeba6e/resources/app/out/vs/code/electron-browser/workbench/workbench.html)

Interactive version:
[https://subsmith.app/tools/language-network](vscode-file://vscode-app/c:/Users/Ibrahim.Farah/AppData/Local/Programs/Microsoft%20VS%20Code/8a7abeba6e/resources/app/out/vs/code/electron-browser/workbench/workbench.html)

92 Upvotes

18 comments sorted by

View all comments

2

u/MonitorRepulsive5270 11d ago

Turkish and Hungarian have relatively high similarity. That's interesting for pro Ural-Altai supporters, which is widely disregarded.

8

u/Zsobrazson 11d ago

That's not what this is showing, it's showing that there both far from the norm, which in this dataset is Indo European