Since releasing Prophet 5.0, I’ve started experimenting with the neural network architecture. The results are surprising.
As explained in my release announcement, the network architecture released with 5.0 is 2x(768 -> 1536) -> 2. That means there are 768 x 1536 = 1,179,648 weights between the input and hidden layer, since the weights for both halves of the input are shared.
My first experiment was to reduce the size of the hidden layer, from 1536 neurons (per half) to 1024 (subsequently reducing the number of shared weights to 768 x 1024 = 786,432), and then retrain on the same dataset. The combined results of two gauntlets:
Rank Name Elo + - games score oppo. draws
1 toad-3.0 72 9 9 4000 50% 0 28%
2 luna-2.0.0 54 9 9 4000 47% 0 29%
3 fatalii-0.9.0 49 9 9 4000 46% 0 29%
4 jikchess-0.02 47 9 9 4000 46% 0 23%
5 admete-1.5.0 22 9 9 4000 43% 0 25%
6 prophet-nn-25 19 4 4 24000 63% 0 25%
7 maverick-1.5 8 9 10 4000 41% 0 24%
8 smol-1.61 -11 9 9 4000 38% 0 26%
9 prophet-5.0 -19 4 4 24000 58% 0 26%
10 tantabus-2.0.0 -33 10 10 4000 35% 0 23%
11 sofcheck-0.9.1 -40 9 10 4000 34% 0 26%
12 loki-3.5.0 -51 10 10 4000 33% 0 24%
13 fornax-4.0 -59 10 10 4000 32% 0 24%
14 barbarossa-0.6.0 -60 10 10 4000 32% 0 20%
Wow! But, what happens if I were to reduce it even further, to say, just 512 neurons per half (768 x 512 = 393,216 shared weights) ? The combined results of all three gauntlets are below:
Rank Name Elo + - games score oppo. draws
1 toad-3.0 70 7 7 6000 47% 0 29%
2 luna-2.0.0 52 7 7 6000 44% 0 29%
3 fatalii-0.9.0 51 7 7 6000 44% 0 30%
4 jikchess-0.02 45 8 8 6000 43% 0 22%
5 prophet-nn-26 40 4 4 24000 68% 0 25%
6 admete-1.5.0 19 8 8 6000 39% 0 25%
7 maverick-1.5 5 8 8 6000 37% 0 23%
8 prophet-nn-25 -1 4 4 24000 63% 0 25%
9 smol-1.61 -8 8 8 6000 36% 0 26%
10 tantabus-2.0.0 -29 8 8 6000 33% 0 23%
11 sofcheck-0.9.1 -39 8 8 6000 32% 0 26%
12 prophet-5.0 -39 4 4 24000 58% 0 26%
13 loki-3.5.0 -51 8 8 6000 30% 0 23%
14 barbarossa-0.6.0 -57 8 9 6000 29% 0 20%
15 fornax-4.0 -58 8 8 6000 29% 0 24%
Another 40 elo jump! It’s incredible that reducing the number of weights by 2/3 actually increases the strength so dramatically, but I guess it makes sense. Before I added in the NNUE stuff, I was focused on just getting the trainer to produce a network that outperformed the hand crafted evaluation in fixed depth matches. A wider hidden layer allows for more information to be encoded, and since I was testing at fixed depth the computational overhead was not an issue. When running in a timed match however, where speed matters, the overhead of all those additional computations begins to outweigh the increased accuracy.
Will further reductions bring in even more elo? Or did I perhaps overshoot and reduce too much already? Would an additional hidden layer help at all? How does the amount of training data impact the optimal network architecture? I happen to have another batch of lichess games labeled (using Prophet’s HCE at depth 5). Does the mirrored input really help? So many questions! One thing that does seem certain though is that 5.1 will be another jump forward in terms of elo!