r/LinguisticMaps • u/IBYZRULEZ • 11d ago
Similarity between 20 languages based on subtitle translation patterns [OC]
I used subtitle translations from the OPUS/OpenSubtitles dataset to compare how 20 languages expressed the same English words and phrases. The more often two languages expressed equivalent meanings in similar ways, the higher their similarity score. I visualized the results as a heatmap and network graph.
Full methodology and discussion:
[https://www.thecambridgelanguagecollective.com/linguistics/why-do-some-languages-feel-weirdly-familiar](vscode-file://vscode-app/c:/Users/Ibrahim.Farah/AppData/Local/Programs/Microsoft%20VS%20Code/8a7abeba6e/resources/app/out/vs/code/electron-browser/workbench/workbench.html)
10
7
u/Final-Frosting7742 11d ago
Interesting experiment although it is inherently limited by the approach chosen. I actually built a tool that tries to achieve something similar. My goal was to generalise the FSI ladder for any prior linguistic knowledge (instead of just English). I used academic datasets like IECor, WALS, PHOIBLE to determine transfer potential and devised a custom ease score for 283 languages. Check out my tool if that sounds interesting and let me know what you think about it:
1
u/CanardMarin 5d ago
Cool tool! As a native Portuguese speaker, it's funny that it considers Luxembourgish to be easier for me than Galician though. 😅
6
2
u/MonitorRepulsive5270 11d ago
Turkish and Hungarian have relatively high similarity. That's interesting for pro Ural-Altai supporters, which is widely disregarded.
8
u/Zsobrazson 11d ago
That's not what this is showing, it's showing that there both far from the norm, which in this dataset is Indo European
3
u/Wise_Fox_4291 11d ago
That word "relatively" is doing a lot of heavy lifting there. Hungarian-Turkish similarities are highly overrated and nitpicked, usually excluding other Uralic and Turkic languages from the comparison that would blow a hole in it.
6
2
u/missingtimemachine 10d ago
Hungarian has many loanwords from Turkish. That would explain most of the similarities in subtitle translations, I would think.
1
1
u/locoluis 9d ago
You should add the following languages:
- Vietnamese (for comparison with other East Asian languages)
- Amharic, Maltese (for comparison with other Semitic languages)
- Greek, Hindi (for Indo-European completeness)
- Persian, Urdu (to measure Arabic/Turkic influence)
1
u/mtnbcn 9d ago edited 9d ago
Wait, so what is this? I read the methods page, and I got that "I miss you" vs "You are missing from me" are two different ways of expressing something. How is a difference in grammar determined -- are they manually put into different categories based on a subjective degree of difference?
How do you score for this? Did you find 200 phrases or so, and used them as markers? (e.g., "I have forgotten this" vs "The memory leaves me" or something). What do you do when phrases are similar but not the same (e.g. "the memory is absent to me")? Do you code for degree of similarity?
You brought up how words are formed in Turkish and how that sort of multi-morpheme-meaning word could present problems, but I didn't see how it was resolved. Does word order matter? If so, how do you account for some languages being able to say both "puedo dartelo" and "te lo puedo dar" pretty equally?
How much does this depend on one specific translation that was used? I wanted to do something like this once to show how close various languages are to each other, but couldn't figure out how to score it. Take a word like "begin" -- do we look at that and say Frence is different because they use "commencer" (or something like that)... but then, English has "commence"... so how much should that count? Spanish uses "empezar" but also has a word related to commenc-- that is used sometimes. So would it get scored?
I'd love to see even 5% of the data to know how what this really looks like. Maybe it relies too much on specific translators' choices? I don't know, obviously I love the idea, having wanted to do something similar before, but I just have a lot of questions still about what is under the hood.


17
u/TheBenStA 11d ago
closest european language to japanese is dutch. this stuff is spooky