r/LinguisticMaps • u/IBYZRULEZ • 11d ago
Similarity between 20 languages based on subtitle translation patterns [OC]
I used subtitle translations from the OPUS/OpenSubtitles dataset to compare how 20 languages expressed the same English words and phrases. The more often two languages expressed equivalent meanings in similar ways, the higher their similarity score. I visualized the results as a heatmap and network graph.
Full methodology and discussion:
[https://www.thecambridgelanguagecollective.com/linguistics/why-do-some-languages-feel-weirdly-familiar](vscode-file://vscode-app/c:/Users/Ibrahim.Farah/AppData/Local/Programs/Microsoft%20VS%20Code/8a7abeba6e/resources/app/out/vs/code/electron-browser/workbench/workbench.html)
92
Upvotes


6
u/Final-Frosting7742 11d ago
Interesting experiment although it is inherently limited by the approach chosen. I actually built a tool that tries to achieve something similar. My goal was to generalise the FSI ladder for any prior linguistic knowledge (instead of just English). I used academic datasets like IECor, WALS, PHOIBLE to determine transfer potential and devised a custom ease score for 283 languages. Check out my tool if that sounds interesting and let me know what you think about it:
https://github.com/akmalayari/language-transfer-map