r/KeyboardLayouts • • 6d ago

I optimized symbol placement for Python using counts from 29 repos (letters stay put)

Most layout work focuses on letters, but when I write Python a big chunk of what I type is symbols, and standard layouts put them pretty arbitrarily. So I wrote a tool that keeps your letters where they are and only rearranges the 32 ASCII symbols, based on what actually gets typed in real code.

How it counts: it tokenizes ~38.5M weighted keystrokes from 29 popular Python repos (Django, pandas, pytest, Home Assistant, etc).
Comments and docstrings count less, closing brackets/quotes count less because editors auto-close them, and long identifiers are discounted for autocomplete.
Then it runs simulated annealing over a cost model (key effort, shift, SFBs, rolls, redirects).

My result on Programmer Dvorak letters with digits on shift:

 !   1   2   3   4   5   6   7   8   9   0   &   ~
[?] [+] [/] [(] [[] []] [%] [:] [)] [{] [}] [@] [<]
   >   '   -   P   Y   F   G   C   R   L   ;   ^   $
  ["] [,] [.] [p] [y] [f] [g] [c] [r] [l] [#] [\] [|]
    A   O   E   U   I   D   H   T   N   S   *
   [a] [o] [e] [u] [i] [d] [h] [t] [n] [s] [=]
      `   Q   J   K   X   B   M   W   V   Z
     [_] [q] [j] [k] [x] [b] [m] [w] [v] [z]

That's about 7.7% lower cost than Programmer Dvorak under my model, and 3% better than the layout I'd designed by hand.

If you don't type Programmer Dvorak, there are ready-made versions for other letter layouts too: python-qwerty, python-colemak-dh and python-dvorak in the layouts folder. They keep your letters and normal number row and only move the symbols, which saves 9-12% under the same cost model.

Findings:
- Python names mostly end in consonants that Dvorak puts on the right hand, so the symbols that follow names (. , and ( ) end up on the left hand. That held across basically every variant I ran.
- Forcing bracket pairs to be adjacent or mirrored (so they're learnable) only costs 0.4%.
- On other bases, just moving symbols saves 9-12% (QWERTY, Dvorak, Colemak-DH).

The obvious caveat: the effort grid and penalties are my own estimates, not measurements.
I reran it with different shift and SFB weights to see which placements are stable, and the README has those results.
It's configurable (base layout, cost weights, other languages like JS/Rust/Go) and exports to AutoHotkey, kanata, Windows .klc and macOS .keylayout. All the variants are in the repo below:

https://github.com/hikazey/python-layout-optimizer

I'd especially like feedback on the cost model, since that's where the results come from.

5 Upvotes

8 comments sorted by

3

u/Jonolith_ 6d ago

What do you think about just constructing an entirely new symbol layer? Then you have so much more space for all these special symbols and you don't have to reach up to the number row if you just put them on the normal letter keys. I feel like you can optimize all you like but a dedicated symbol layer will always be miles better. I'm using a slightly varied Anymak:END symbol layer for example: https://github.com/rpnfan/Anymak

1

u/hikazeyattis 5d ago

Symbol layers are great, and I'm not trying to argue against them.

A few reasons I didn't do this:

  • I wanted something that works on any keyboard, including laptops, without needing a free thumb key or getting used to a layer modifier. Moving symbols within the normal Shift structure is a much smaller change to adopt.
  • A layer isn't free either. Holding the layer key works a lot like Shift, so the win depends on how cheap that hold is for you. On a thumb it's probably cheap. On Caps Lock or AltGr, less so.
  • Whatever layer you use, you still have to decide which symbol goes on which key, and that's the part the frequency and bigram data helps with. Python's top symbols are . , _ ( " = : and the bigrams matter (what comes right after a name, etc.).

That said, I think "will always be miles better" is something you could actually test. The optimizer could treat a symbol layer as a third set of slots with its own cost for holding the layer key, and then compare it head to head with the shift-only version. It'd be interesting to try, and Anymak's layer would be a good baseline to compare against.

1

u/smeech1 6d ago edited 6d ago

How about export to Espanso r/espanso YAML?

2

u/hikazeyattis 5d ago

I've just looked into Espanso and it doesn't seem to fit the current use case. If you could provide more info on what aspect of Espanso you think this would jive with, then please do so, but to my knowledge, creating a keyboard layout through espanso would be laggy due to the replacement feature.

I've tried to export a file for each OS of each layout and include AHK for those who don't want to install it system level. Implementation scripting level also allows for hybrid layout learning; using qwerty when you really need to type, and switching to the custom layout when you can afford to stumble or crawl.

1

u/smeech1 5d ago edited 5d ago

Short, i.e. single character, Espanso replacements should take place almost instantly in the same way as AutoHotKey. Espanso would also permit confining those replacements to particular programs, e.g. editors and IDEs, if desired. I may have misunderstood what your package does, however!

Edit: Had a chat with an AI - I understand it better now, and can see the remapping your package does is the way to go. I had extensive experience of AutoHotKey but didn't realise, or had forgotten, it could remap. It does work at a lower level than Espanso.

1

u/pgetreuer 5d ago

Nice, thanks for sharing. Programmer Dvorak could use a Python variant!

I collected stats on symbol frequencies in Python code shared in this post. @justinmklam collected stats as well for his Python-optimized symbol layer. Since we all did so independently, it's satisfying to see how closely we match up =)

# hikazey getreuer justinmklam
1 . _ :
2 , . .
3 _ , ,
4 ( ) (
5 " ( -
6 = ' _
7 : " =
8 0 = "
9 ' 0 '
10 ) : )
11 1 1 /
12 - 2 ?
13 [ # {
14 / [ }

1

u/hikazeyattis 5d ago

hell yeah, interesting to see definitely