Sometimes, Less is More

Since releasing Prophet 5.0, I’ve started experimenting with the neural network architecture. The results are surprising.

As explained in my release announcement, the network architecture released with 5.0 is 2x(768 -> 1536) -> 2. That means there are 768 x 1536 = 1,179,648 weights between the input and hidden layer, since the weights for both halves of the input are shared.

My first experiment was to reduce the size of the hidden layer, from 1536 neurons (per half) to 1024 (subsequently reducing the number of shared weights to 768 x 1024 = 786,432), and then retrain on the same dataset. The combined results of two gauntlets:

Rank Name               Elo    +    - games score oppo. draws 
   1 toad-3.0            72    9    9  4000   50%     0   28% 
   2 luna-2.0.0          54    9    9  4000   47%     0   29% 
   3 fatalii-0.9.0       49    9    9  4000   46%     0   29% 
   4 jikchess-0.02       47    9    9  4000   46%     0   23% 
   5 admete-1.5.0        22    9    9  4000   43%     0   25% 
   6 prophet-nn-25       19    4    4 24000   63%     0   25% 
   7 maverick-1.5         8    9   10  4000   41%     0   24% 
   8 smol-1.61          -11    9    9  4000   38%     0   26% 
   9 prophet-5.0        -19    4    4 24000   58%     0   26% 
  10 tantabus-2.0.0     -33   10   10  4000   35%     0   23% 
  11 sofcheck-0.9.1     -40    9   10  4000   34%     0   26% 
  12 loki-3.5.0         -51   10   10  4000   33%     0   24% 
  13 fornax-4.0         -59   10   10  4000   32%     0   24% 
  14 barbarossa-0.6.0   -60   10   10  4000   32%     0   20% 

Wow! But, what happens if I were to reduce it even further, to say, just 512 neurons per half (768 x 512 = 393,216 shared weights) ? The combined results of all three gauntlets are below:

Rank Name               Elo    +    - games score oppo. draws 
   1 toad-3.0            70    7    7  6000   47%     0   29% 
   2 luna-2.0.0          52    7    7  6000   44%     0   29% 
   3 fatalii-0.9.0       51    7    7  6000   44%     0   30% 
   4 jikchess-0.02       45    8    8  6000   43%     0   22% 
   5 prophet-nn-26       40    4    4 24000   68%     0   25% 
   6 admete-1.5.0        19    8    8  6000   39%     0   25% 
   7 maverick-1.5         5    8    8  6000   37%     0   23% 
   8 prophet-nn-25       -1    4    4 24000   63%     0   25% 
   9 smol-1.61           -8    8    8  6000   36%     0   26% 
  10 tantabus-2.0.0     -29    8    8  6000   33%     0   23% 
  11 sofcheck-0.9.1     -39    8    8  6000   32%     0   26% 
  12 prophet-5.0        -39    4    4 24000   58%     0   26% 
  13 loki-3.5.0         -51    8    8  6000   30%     0   23% 
  14 barbarossa-0.6.0   -57    8    9  6000   29%     0   20% 
  15 fornax-4.0         -58    8    8  6000   29%     0   24% 

Another 40 elo jump! It’s incredible that reducing the number of weights by 2/3 actually increases the strength so dramatically, but I guess it makes sense. Before I added in the NNUE stuff, I was focused on just getting the trainer to produce a network that outperformed the hand crafted evaluation in fixed depth matches. A wider hidden layer allows for more information to be encoded, and since I was testing at fixed depth the computational overhead was not an issue. When running in a timed match however, where speed matters, the overhead of all those additional computations begins to outweigh the increased accuracy.

Will further reductions bring in even more elo? Or did I perhaps overshoot and reduce too much already? Would an additional hidden layer help at all? How does the amount of training data impact the optimal network architecture? I happen to have another batch of lichess games labeled (using Prophet’s HCE at depth 5). Does the mirrored input really help? So many questions! One thing that does seem certain though is that 5.1 will be another jump forward in terms of elo!