Goodbye JNI, Hello FFM

I’ve finally done it. I’ve ripped out all of the Java Native Interface (JNI) code and replaced it with Java’s Foreign Function and Memory API (FFM). This has dramatically simplified the codebase and the build process, as the entire submodule dedicated to JNI code was removed. None of that is necessary with FFM. From a software architecture perspective, I’m really very happy with it.

That’s the good news. Now the bad news. It doesn’t solve the problem I wrote about a few weeks ago in The Overhead of JNI , where I showed that chess4j + Prophet as a native library was nearly 2x the speed of chess4j by itself, but sadly chess4j + Prophet is only 0.4x the speed of Prophet by itself. I attributed the difference to the overhead of JNI which is a layer of code that serves as a bridge between Java and native code (C in this case). I was wrong.

The majority of that penalty wasn’t the JNI layer at all. It was the fact that, when Prophet is used as a native library, it has to be compiled as a shared library. This is true whether you use JNI or FFM. When I benchmarked Prophet by itself, the C code was compiled as a static library. Here are some new benchmarks showing the same positions.

PositionProphet knps (static)Prophet knps (shared)chess4j knpschess4j + Prophet knps (using FFM)
initial position5405252110682580
r2qr1k1/p4ppp/5n2/2PpN3/8/2N1P1P1/1PP3BP/R4RK1 b – –5244268410212644
2r1r1k1/p4ppp/8/2Pp4/N7/4qNP1/1PP2RBP/4R2K b – –5355283610522845
rnbqkbnr/pppppppp/8/8/8/8/PPPPPPPP/RNBQKBNR w KQkq –5497262110462549
rn1qk2r/ppp2ppp/3bpn2/3p4/6b1/1P2P1Q1/PBPP1PPP/RN2KBNR w KQkq –498323019632287
rnbqkbnr/pp2pppp/8/3p4/8/2N2N2/PPPP1PPP/R1BQKB1R b KQkq –524525229902470
r2qkb1r/1p1b1ppp/p3pn2/3p2B1/3P4/2N2N2/PPP2PPP/R2Q1RK1 b kq –5406260410232554
2kr1b1r/1pqb1p2/p3p2p/3pNp2/3P4/2NQ4/PPP2PPP/R3R1K1 b – –456423679762508
2k3rr/1pqb4/p2bpp1p/3p1p2/NP1P4/P2Q1N1P/2P2PP1/R3R1K1 b – –479624899312432
2k4r/1pqN4/p2b3p/3Q1p2/1P6/P4p1P/2P2Pr1/R3RK2 b – –4565326712383295

If you compare the third column (standalone Prophet as a shared library) against the fifth column (chess4j + Prophet using FFM), there is very little difference. There is almost no overhead to speak of. However, if you compare the second column (standalone Prophet as a static library) vs the third column (standalone Prophet as a shared library), the difference is 1.95x. That was the culprit all along.

I’m not entirely sure why the shared library should be so much slower than a static library. I doubt that the overhead of dynamic linking is to blame. I suspect it has more to do with reduced opportunity for the compiler to perform certain optimizations.

Regardless of the reason, this means that if I want Prophet to be more competitive in programmer tournaments, I’ll have to add some additional features rather than relying on the chess4j + Prophet integration. Mainly, Prophet will need pondering and an opening book of its own. So, I will be adding those to the task board, even though I won’t prioritize them right away. For now I’m going to continue to experiment with NNUE networks.