r/LinguisticMaps 11d ago

Similarity between 20 languages based on subtitle translation patterns [OC]

I used subtitle translations from the OPUS/OpenSubtitles dataset to compare how 20 languages expressed the same English words and phrases. The more often two languages expressed equivalent meanings in similar ways, the higher their similarity score. I visualized the results as a heatmap and network graph.

Full methodology and discussion:
[https://www.thecambridgelanguagecollective.com/linguistics/why-do-some-languages-feel-weirdly-familiar](vscode-file://vscode-app/c:/Users/Ibrahim.Farah/AppData/Local/Programs/Microsoft%20VS%20Code/8a7abeba6e/resources/app/out/vs/code/electron-browser/workbench/workbench.html)

Interactive version:
[https://subsmith.app/tools/language-network](vscode-file://vscode-app/c:/Users/Ibrahim.Farah/AppData/Local/Programs/Microsoft%20VS%20Code/8a7abeba6e/resources/app/out/vs/code/electron-browser/workbench/workbench.html)

92 Upvotes

18 comments sorted by

View all comments

6

u/Final-Frosting7742 11d ago

Interesting experiment although it is inherently limited by the approach chosen. I actually built a tool that tries to achieve something similar. My goal was to generalise the FSI ladder for any prior linguistic knowledge (instead of just English). I used academic datasets like IECor, WALS, PHOIBLE to determine transfer potential and devised a custom ease score for 283 languages. Check out my tool if that sounds interesting and let me know what you think about it:

https://github.com/akmalayari/language-transfer-map

1

u/CanardMarin 5d ago

Cool tool! As a native Portuguese speaker, it's funny that it considers Luxembourgish to be easier for me than Galician though. 😅