Author Archives: jswafford

The Cost of a “Free” Hash Probe

I recently ran an experiment to test probing the hash table from the quiescence phase of the search. The argument for it is: qsearch is effectively a depth-0 search, so anything already sitting in the hash table was put there by a full-width search at depth 1 or deeper, which by definition did more work than qsearch would do on its own. If that entry’s bound is good enough to cause a cutoff, it’s a cutoff you get essentially for free.

So I added it to chess4j. The change was small: at the top of quiescence search, probe the main hash table.

Elo difference: -35.9 +/- 9.8, LOS: 0.0 %
SPRT: llr -2.95 (-100.1%), lbound -2.94, ubound 2.94 - H0 was accepted

Not a wash — a clear loss, and a fairly big one at that. But why?

My best guess … consider that the quiescence nodes make up a large majority of the search tree – something like 2/3.

Also consider that when you hit the quiescence phase, depth is already at 0. Those lines are usually fairly shallow.

Add a hash lookup to every one of those enormous numbers of qsearch nodes, and the nodes-per-second hit apparently outweighs whatever cutoffs it was finding.

Perhaps the outcome would be different with a different hash replacement scheme. chess4j is “always replace,” which isn’t the most cache friendly scheme to begin with.

For the sake of completeness I tried a couple of variations of gating the lookup based on depth. The first attempt was to only probe once qsearch is at least 2 plies deep.

Elo difference: 1.6 +/- 3.5, LOS: 81.4 %, DrawRatio: 42.6 %

It’s statistically neutral, but neutral is a massive improvement over -36 Elo. So naturally I pushed it further: if 2 plies helps, does 4 help more?

Elo difference: -15.2 +/- 6.5, LOS: 0.0 %
SPRT: llr -2.95 (-100.2%), lbound -2.94, ubound 2.94 - H0 was accepted

Nope. Worse than 2 plies, though still better than probing everywhere.

End result: I don’t have a version of “probe the hash table from qsearch” that’s actually worth merging into chess4j. Sometimes the answer an experiment gives you is just “no.”

Profile-Guided Optimization for Prophet: A Regression That Wasn’t

Co-authored by Claude based on context and transcripts collected during a real-time debugging workflow.

Profile-Guided Optimization is supposed to be one of the safer performance wins available to a C/C++ project. You build once with instrumentation, run a representative workload, feed the resulting profile back into a second build, and the compiler makes better inlining, branch-layout, and register-allocation decisions based on how the code actually runs rather than static heuristics. In the worst case it does nothing. It should never make a chess engine play worse chess — a PGO build searches the same tree with the same evaluation, just (hopefully) a bit faster.

So when I wired up scripts/build-pgo.sh for Prophet and set a fresh PGO build loose against a plain Release build in a real match, I expected a small, boring win. Instead I got this:

Score of prophet-dev vs prophet-latest: 100 - 236 - 201  [0.373] 537
     prophet-dev playing White: 93 - 35 - 141  [0.608] 269
     prophet-dev playing Black: 7 - 201 - 60  [0.138] 268
     White vs Black: 294 - 42 - 201  [0.735] 537
Elo difference: -89.9 +/- 23.5, LOS: 0.0 %, DrawRatio: 37.4 %
SPRT: llr -1.33 (-45.2%), lbound -2.94, ubound 2.94

prophet-dev was the PGO build. prophet-latest was a plain Release build of the same commit. Same source, same neural network weights, same everything except an extra compiler pass — and PGO was losing by the better part of 90 Elo. Granted, the error bars are pretty high with just 537 games, but it’s clear the PGO build is just worse. It took the better part of two weeks, several false leads, and a full concurrency sweep to get to the bottom of it.

Ruling out the code

A result like that raises a correctness question before anything else: is something actually wrong, not just slower? The elimination process was mostly writing small scripts to compare the two binaries directly rather than through hundreds of games of noise:

  • Compile flags — confirmed identical between the PGO and baseline builds (same -Wall -Werror -Wextra, same -flto, same -mavx2 -mbmi2; the only difference is -fprofile-use vs. nothing).
  • Fixed-depth search determinism  — 11 positions, fixed-depth (d=10), comparing best move and score. Matched on every one.
  • Sustained throughput — fixed-time (10s) searches on four representative positions, run alone, no contention. PGO was faster on every single one.
  • Code size — I had a theory that PGO’s more aggressive inlining might bloat the binary enough to hurt icache behavior under load. The .text segment was actually smaller for the PGO build (79,162 bytes vs. 101,983 bytes), which killed that theory outright.

Every controlled, single-process test said the two binaries were functionally identical and the PGO build was, if anything, a little faster. And yet the real match said PGO was losing by 90 Elo. At that point I began to suspect I had a test environment problem.

The actual variable

The turning point was the realization that every test I’d run had compared the two binaries running alone. The real match hadn’t — it was configured with -concurrency 8, and cutechess-cli spawns two engine processes per game. On an 8-core/16-thread Ryzen 5800X, -concurrency 8 means up to sixteen engine processes competing for sixteen logical threads: full SMT saturation, zero headroom, all the time. (Note: Prophet does not use “pondering mode,” where the engine can search on the opponent’s time, so CPU contention between active threads was never the issue.)

I hadn’t reproduced that in any of my isolated tests. So I wrote one that did — launching eight simultaneous prophet-dev searches and eight simultaneous prophet-latest searches at once, matching the real concurrency — and found… nothing. Both binaries degraded under contention by roughly the same amount (1.05x in PGO’s favor, if anything). Same-vs-same contention didn’t reproduce the asymmetry either.

The real test was just running the actual match at lower concurrency. A control run at effectively -concurrency 1 immediately flipped the sign: 58.0% at 44 games, settling to 53.0% (Elo +20.9 ± 47.4) at 100 games — noisy, but nothing like the -90 Elo collapse. A longer confirmation run at -concurrency 4 over several days put it beyond doubt:

Score of prophet-dev vs prophet-latest: 2012 - 1706 - 2899  [0.523] 6617
Elo difference: 16.1 +/- 6.3, LOS: 100.0 %, DrawRatio: 43.8 %
SPRT: llr 2.95 (100.0%), lbound -2.94, ubound 2.94 - H1 was accepted

PGO wasn’t a 90-Elo regression. It was a genuine, statistically solid ~16-Elo improvement — once the test stopped saturating the machine.

The concurrency sweep

-concurrency N spawns 2N processes; the machine has 8 physical cores. N=4 is exactly one process per physical core — no SMT sharing required — so the obvious next question was whether the degradation was a hard cliff at that boundary or a gradual slope above it. The first pass:

ConcurrencyProcessesGamesScoreElo
486,6170.523+16.1 ± 6.3
5105,1810.471-20.5 ± 7.5
6129570.336-118.0 ± 18.4
8165370.373-89.9 ± 23.5

Fine right at the physical core count, a steep penalty just above it, roughly plateauing once the machine is thoroughly oversubscribed. Then I ran N=3 — six processes, comfortably under the 8-core boundary, which the theory said should be at least as clean as N=4 — and got this:

RunGamesScoreElo
N=37,4600.523+16.3 ± 6.3

That landed almost exactly on top of N=4‘s +16.1, which made sense — the machine was not oversubscribed.

Rerunning everything

With N=3 confirmed, I reran N=5, N=6, and N=8 at a fixed 10,000 games apiece. (Larger sample sizes give tighter confidence intervals.) Every one of them came back with a smaller, more moderate magnitude than the original small-sample read:

ConcurrencyProcessesGamesScoreElo (original)Elo (rerun, fixed 10k)
51010,0000.418-20.5 ± 7.5-57.4
61210,0000.407-118.0 ± 18.4-65.1 ± 5.5
81610,0000.241-89.9 ± 23.5-199.3 ± 6.8

N=6 in particular moved a lot — from a small-sample -118.0 down to a confirmed -65.1 — which fit the pattern of small samples reading more extreme than the truth. N=8, on the other hand, moved the other way, from -89.9 up to a much more severe, tightly-bounded -199.3.

That set up an obvious next test: N=7, at fourteen processes — two threads still idle, one core’s worth short of full saturation. If the curve was smooth, N=7 should sit between N=6’s -65.1 and N=8’s -199.3. Instead:

Score of prophet-dev vs prophet-latest: 1105 - 7004 - 1891  [0.205] 10000
Elo difference: -235.4 +/- 7.1, LOS: 0.0 %, DrawRatio: 18.9 %

Worse than N=8! Not close, either — the confidence intervals didn’t overlap. That’s a genuinely strange result. I had a mechanism ready to explain it: with idle threads available, the Linux scheduler has somewhere to migrate processes to when it rebalances load, and each migration costs a cold cache — while at full saturation there’s nowhere to migrate to, so assignments settle and stay put, giving a slower but more stable steady state. It’s a plausible story. It also should be replicated before I believed it. So I reran N=7:

Score of prophet-dev vs prophet-latest: 1538 - 5977 - 2485  [0.278] 10000
Elo difference: -165.8 +/- 6.3, LOS: 0.0 %, DrawRatio: 24.9 %

Two full 10,000-game runs, identical settings, landing 70 Elo apart with non-overlapping confidence intervals. Pooling both runs (20,000 games: 2643-12981-4376) gives a combined estimate of roughly -198.7 Elo — which lands almost exactly on top of N=8’s -199.3. The tidy, counterintuitive “N=7 is worse than full saturation” finding dissolved the moment I checked it twice. N=7 and N=8 are most likely the same number, both sitting in the same severe-collapse plateau, and my scheduler-migration theory was a good explanation for a result that turned out not to exist.

The complete picture

ConcurrencyProcessesGamesScoreElo (dev vs. baseline)
367,4600.523+16.3 ± 6.3
486,6170.523+16.1 ± 6.3
51010,0000.418-57.4
61210,0000.407-65.1 ± 5.5
71420,000 (2 runs, pooled)0.242≈ -198.7
81610,0000.241-199.3 ± 6.8

The result that emerges is clean, and slightly favorable to PGO, at or below the physical core count (3-4); a real but moderate penalty just above it (5-6); and a severe, decisive collapse once you’re deep into SMT-shared territory (7-8).

Where this leaves the feature

The original question is answered: PGO is a real, if modest, improvement for Prophet — about +16 Elo, confirmed twice independently, at a concurrency setting that respects the physical core count. Left purely on the merits, it would ship.

It isn’t shipping, at least not as a default. The reason is everything above: getting a trustworthy read out of this required a full concurrency sweep and multiple reruns. I don’t control what concurrency other people testing or building Prophet will run their matches at, and I can’t hand someone a build flag with a footnote that says “only trust this if you keep concurrency at or below your physical core count and ran it more than once.” That’s not a maintainable claim to stand behind. PROFILE_GUIDED stays off by default in CMakeLists.txt – but scripts/build-pgo.sh and the comparison scripts I wrote along the way stay in the tree as opt-in tooling, in case the calculus changes later or someone wants to reproduce any of this.

Lesson (re-re-re…) learned

It’s interesting that a Profile Guided Optimization build degrades in performance relative to the “normal” build as resource contention increases. The bigger lesson– your test environment should match real world conditions as closely as possible.

Side note: as stated in Prophet 5.2 released, I’ve decided not to use AI to make any algorithmic changes in Prophet, but I did leverage Claude for this experiment. (And I will heavily leverage AI going forward in chess4j.)

As always, the latest can be downloaded from the Prophet Github site.


chess4j 6.3 is out. Prophet climbs a little higher on CCRL.

Short update – chess4j 6.3 is out. Nothing earth shattering, just a minor update, mainly to update the net and get the code merged up.

As always, the latest can be downloaded from the Github site.

Also, I see that Prophet 5.2 is on the CCRL Blitz list at 2753 (ranked 292). 2800 is on the horizon!

Next up: I want to get back to making some algorithmic improvements to the search, starting with another look at root move ordering. (While continuing to improve the net.)

Prophet 5.2 released

At long last, Prophet 5.2 is released! Development has been quiet for a while as I was settling into a new job and dealing with other commitments, but I’ve managed to finally get this one over the line. Nothing earth shattering, but a couple of nice improvements:

  • The neural network has been improved. I haven’t changed the architecture from the last release, but the quality of the training data has improved. I’m training off approximately 200m positions, each searched to a depth of 5 (+quiescence). The position at the end of the principal variation is then scored using a 35k fixed node search.
  • I finally have a Native WIndows build! This has been a challenge for me, as I don’t typically do development work on Windows.

My testing shows Prophet 5.2 to be about +33 elo over 5.1:


ResultSet-EloRating>ratings
Rank Name              Elo    +    - games score oppo. draws 
   1 zevra-2.6         103    5    5 16066   62%    16   16% 
   2 Tcheran-5.1        54    4    4 16064   55%    19   31% 
   3 prophet-5.2        33    3    3 38400   51%    24   30% 
   4 casacnchess-0.9    27    4    4 16065   51%    20   36% 
   5 aramis-1.4.0       27    4    4 16064   51%    20   30% 
   6 Lux-4.2            20    5    5 16064   50%    21   27% 
   7 lishex-1.1.1        7    5    5 16066   48%    22   27% 
   8 prophet-5.1         0    3    3 34117   47%    24   26% 
   9 Supernova-2.4     -10    5    5 16064   45%    23   21% 
  10 drosophila-1.6    -32    5    5 16064   42%    24   25% 

Other thoughts: in my day job, we are leaning more and more into AI driven development. The tool of choice at work is Claude Code. I’ve also dabbled with Codex. I used Codex to help with the compatibility issues in getting an MSVC build. I’ve also had both Claude and Codex do an analysis of the codebase to suggest spreed imrpovements, mainly out of curiosity; I wanted to compare the quality of the suggestions. (I haven’t actioned any of those suggestions yet.) The question now is- do I really want to use AI tools?

I’ve put a lot of thought into this actually, and I’ve decided not to leverage AI tools for Prophet, or at least not for anything that changes the engine’s behavior (e.g. algorithmic). As stated above, Codex did help produce an MSVC build. However, this is a creative endeavor for me, and part of the joy of it is “staying close to the code.” I also think it’s a little unfair perhaps to use AI tools and then enter the program into competition.

However, chess4j, my Java program, isn’t meant to be competitive. It’s just a learning testbed. I WILL use AI tools to improve chess4j. That gives an opportunity to experiment with the latest AI coding agents without feeling as though I’m cheating when entering into a competition.

In short- no AI coding agents to improve Prophet’s playing strength. Yes to AI coding for chess4j.

MSVC Build for Prophet

Prophet has been a Linux / gcc engine forever. To run on Windows, I resorted to using Cygwin. It’s been on my to do list to build Prophet natively under Windows using MSVC. The problem is, I’m not well versed with MSVC. Over the holidays I decided to see if OpenAI’s coding agent Codex could help.

My first attempt was literally just telling Codex to make the codebase build cleanly under MSVC. With some minor coaxing it finally did, but the resulting binary just … crashed. I’m actually not sure what went wrong. So, I reset and tried to make a series of smaller changes.

The first step was to create some abstractions, particularly around threading. With help from Codex, I unified thread and mutex usage behind explicit wrappers so thread creation, locking, and joining behave identically on Windows and POSIX systems.

From there, I updated CMake for MSVC and tried to compile. Not quite there yet.

MSVC’s lack of support for C99 variable-length arrays required some refactoring. Where sizes had known upper bounds, I replaced VLAs with fixed-size arrays. Where they were truly dynamic, I moved allocations to the heap.

I also had to address a series of MSVC CRT warnings promoted to errors. Functions like setbuf, freopen, and getenv were replaced with their Windows-safe equivalents (setvbuf, freopen_s, _dupenv_s), wrapped in #ifdef _WIN32 where appropriate so non-Windows builds remain untouched.

Testing needed similar attention. Capturing stdout and locating test resources relied on POSIX assumptions. I introduced small helpers to abstract device paths and resource locations, allowing tests to run cleanly on both Linux and Windows without cluttering test code with platform checks.

There were some C/C++ linkage issues in the test suite. The engine is written in C, but tests compile as C++, which led to name-mangling problems for global symbols. Centralizing these declarations in a shared header wrapped with extern "C" eliminated a whole class of MSVC linker errors.

All in all, it still wasn’t a trivial affair, but Codex did help me get it across the finish line.

Chess programming is, in my view, a creative endeavor. That being the case, I wouldn’t use a coding agent for anything that actually modified the engine’s behavior, but this feels like an appropriate use.

As I write this, all the work is still on a branch for further testing, but I anticipate it will make its way into a release in the near future.

Prophet 5.1 Breaks 2700 on the CCRL Blitz Ratings List

Prophet 5.1 has hit another milestone, crossing the 2700 Elo mark on the CCRL Blitz Ratings List. The latest list shows Prophet at 2712 Elo, another big step forward and a nice validation of the neural-network work that started with version 5.0.

The focus of development is still squarely on the neural network — refining the architecture, expanding the training data, and experimenting with new approaches to label the data.

Breaking 2700 is a great milestone, but there’s plenty of headroom left.

Java Vector API results

I finally got around to investigating the Java Vector API. This feature has been in the incubator for ages. It’s purpose is to optimize vector operations using SIMD instructions, presumably much faster than the performance from scalar operations. Lots of people have reported some pretty fantastic speedups, but unfortunately I’m not seeing them.

The task was simple. I wanted to replace this block of code in the inference phase:

        for (int i=0;i<NN_SIZE_L2;i++) {
            int sum = B1[i];
            for (int j=0;j<(NN_SIZE_L1*2);j++) {
                sum += W1[i * (NN_SIZE_L1*2) + j] * L1[j];
            }
            L2[i] = sum;
        }

With something that looks like this:

VectorSpecies<Integer> INT_SPEC = IntVector.SPECIES_256;
for (int i=0;i<NN_SIZE_L2;i++) {
IntVector sum32 = IntVector.zero(INT_SPEC);
for (int j=0;j<(NN_SIZE_L1*2);j+=INT_SPEC.length()) {
IntVector inp = IntVector.fromArray(INT_SPEC, L1, j);
IntVector wei = IntVector.fromArray(INT_SPEC, W1, i * (NN_SIZE_L1 * 2) + j);
IntVector dot = inp.mul(wei);
sum32 = sum32.add(dot);
}
L2[i] = sum32.reduceLanes(VectorOperators.ADD) + B1[i];
}

Notice how the inner loop gets incremented by INT_SPEC.length() , which happens to be 8 (lanes) in this case, since ints in Java are 32 bits, and the entire vector is 256 bits wide. The idea is that we will load up the vector and perform those multiplication operations in parallel, then do an ‘add’ reduction at the end. It works fine, but unfortunately is actually slightly *slower* than the non-vectorized implementation. I tried variations using ByteVectors and ShortVectors, but they were even worse.

I’ll take another look when it comes out of the incubation phase, and/or when I get a machine that supports AVX-512 instructions, but for now I’ll be sticking with the non-vectorized version.

Prophet 5.1 and chess4j 6.2 released

I am very pleased to announce the latest releases of both of my chess programs. This is another big improvement for Prophet, around 96 elo! I had the nice “problem” of having to find an entire new set of sparring partners!

Rank Name               Elo    +    - games score oppo. draws 
   1 Lux-4.2            142    4    4 18600   60%    70   30% 
   2 zevra-2.6          137    5    5 18600   59%    71   19% 
   3 Tcheran-5.1        136    4    4 18600   59%    71   28% 
   4 casacnchess-0.9    133    4    4 18600   59%    71   39% 
   5 aramis-1.4.0       124    4    4 18600   57%    71   30% 
   6 prophet-5.1         96    2    2 57361   52%    83   30% 
   7 lishex-1.1.1        87    4    4 18600   52%    73   28% 
   8 drosophila-1.6      83    4    4 18600   51%    73   28% 
   9 Supernova-2.4       72    4    4 18600   50%    74   26% 
  10 blunder-8.5.5       49    4    4 18600   46%    75   31% 
  11 mess-0.3.0          27    4    4 18600   43%    76   27% 
  12 toad-3.0            15    5    5 13000   59%   -51   27% 
  13 luna-2.0.0           2    5    5 13000   58%   -50   28% 
  14 fatalii-0.9.0        1    5    5 13000   57%   -50   28% 
  15 prophet-5.0          0    2    2 81361   42%    53   27% 
  16 jikchess-0.02      -13    5    5 13000   55%   -49   21% 
  17 admete-1.5.0       -42    5    5 13000   51%   -47   26% 
  18 maverick-1.5       -52    5    5 13000   49%   -46   24% 
  19 smol-1.61          -65    5    5 13000   47%   -45   27% 
  20 tantabus-2.0.0     -75    5    5 13000   46%   -44   26% 
  21 sofcheck-0.9.1     -94    5    5 13000   43%   -43   26% 
  22 loki-3.5.0        -106    5    5 13000   41%   -42   25% 
  23 barbarossa-0.6.0  -107    5    5 13000   41%   -42   22% 
  24 fornax-4.0        -118    5    5 13000   39%   -41   25% 

The gains in Prophet can be attributed 100% to the changes in the neural network architecture described in Sometimes, Less is More, and in improvements to the training process itself by incorporating quantization error. There were also some non-functional changes: a massive cleanup of the public headers of the core library. A lot of data structures and methods were pulled into internal headers. I also cleaned up the Doxygen documentation, did some reformatting to make the codebase more consistent, and cleaned up / added some options to the CMake file. So all in all, this is a nice update for Prophet!

chess4j uses the same neural network, so it also benefits. In fact, for chess4j, this is the first version where NNUE seems to actually be an improvement over the handcrafted evaluation! I haven’t extensively tested it, but it appears to be a 10-20 elo improvement over HCE. Previously, with the larger network, the overhead was just too high.

As I wrote in Goodbye JNI, Hello FFM, chess4j no longer uses the Java Native Interface stuff. It’s all been replaced with Java’s new Foreign Function and Memory API. FFM is so much simpler and safer. It’s easier to write, easier to test, easier to maintain. With FFM the integration code itself is all Java, where it’s native code in JNI. This eliminates the need for an entire submodule of the project! The project structure and build process are just – simpler.

All that said, I’ve decided to no longer publicly support the chess4j + Prophet integration. I started this integration about six years ago, and wrote about it in chess4j + Prophet4 POC . It was a challenging project, and I learned a lot doing it, but my goals have changed. I was focused on chess4j, and making a Java program faster by dropping down into native code. Now, I’d like to try to make Prophet more competitive as a standalone engine, and that means adding some features that are only available in chess4j directly into Prophet. But, the more of those features are added, the less that integration makes sense. So, that integration will live on, but for private use as a debugging tool.

The short term roadmap:

  • Better training data. My current dataset of ~500 million positions is labeled by using a depth 5 (+ quiescence search). I think there is a lot of room for improvement here.
  • Windows builds for Prophet. My development environment is on Linux, but I’d like to get a native Windows compile using MSVC, rather than relying on Cygwin. I’m guessing it would be faster.
  • On the Java side, investigate the Vector API to determine if there is any application in chess4j’s NNUE or eval tuning code.
  • Improve root move ordering

As always, lots to do!

Goodbye JNI, Hello FFM

I’ve finally done it. I’ve ripped out all of the Java Native Interface (JNI) code and replaced it with Java’s Foreign Function and Memory API (FFM). This has dramatically simplified the codebase and the build process, as the entire submodule dedicated to JNI code was removed. None of that is necessary with FFM. From a software architecture perspective, I’m really very happy with it.

That’s the good news. Now the bad news. It doesn’t solve the problem I wrote about a few weeks ago in The Overhead of JNI , where I showed that chess4j + Prophet as a native library was nearly 2x the speed of chess4j by itself, but sadly chess4j + Prophet is only 0.4x the speed of Prophet by itself. I attributed the difference to the overhead of JNI which is a layer of code that serves as a bridge between Java and native code (C in this case). I was wrong.

The majority of that penalty wasn’t the JNI layer at all. It was the fact that, when Prophet is used as a native library, it has to be compiled as a shared library. This is true whether you use JNI or FFM. When I benchmarked Prophet by itself, the C code was compiled as a static library. Here are some new benchmarks showing the same positions.

PositionProphet knps (static)Prophet knps (shared)chess4j knpschess4j + Prophet knps (using FFM)
initial position5405252110682580
r2qr1k1/p4ppp/5n2/2PpN3/8/2N1P1P1/1PP3BP/R4RK1 b – –5244268410212644
2r1r1k1/p4ppp/8/2Pp4/N7/4qNP1/1PP2RBP/4R2K b – –5355283610522845
rnbqkbnr/pppppppp/8/8/8/8/PPPPPPPP/RNBQKBNR w KQkq –5497262110462549
rn1qk2r/ppp2ppp/3bpn2/3p4/6b1/1P2P1Q1/PBPP1PPP/RN2KBNR w KQkq –498323019632287
rnbqkbnr/pp2pppp/8/3p4/8/2N2N2/PPPP1PPP/R1BQKB1R b KQkq –524525229902470
r2qkb1r/1p1b1ppp/p3pn2/3p2B1/3P4/2N2N2/PPP2PPP/R2Q1RK1 b kq –5406260410232554
2kr1b1r/1pqb1p2/p3p2p/3pNp2/3P4/2NQ4/PPP2PPP/R3R1K1 b – –456423679762508
2k3rr/1pqb4/p2bpp1p/3p1p2/NP1P4/P2Q1N1P/2P2PP1/R3R1K1 b – –479624899312432
2k4r/1pqN4/p2b3p/3Q1p2/1P6/P4p1P/2P2Pr1/R3RK2 b – –4565326712383295

If you compare the third column (standalone Prophet as a shared library) against the fifth column (chess4j + Prophet using FFM), there is very little difference. There is almost no overhead to speak of. However, if you compare the second column (standalone Prophet as a static library) vs the third column (standalone Prophet as a shared library), the difference is 1.95x. That was the culprit all along.

I’m not entirely sure why the shared library should be so much slower than a static library. I doubt that the overhead of dynamic linking is to blame. I suspect it has more to do with reduced opportunity for the compiler to perform certain optimizations.

Regardless of the reason, this means that if I want Prophet to be more competitive in programmer tournaments, I’ll have to add some additional features rather than relying on the chess4j + Prophet integration. Mainly, Prophet will need pondering and an opening book of its own. So, I will be adding those to the task board, even though I won’t prioritize them right away. For now I’m going to continue to experiment with NNUE networks.

The Overhead of JNI

Almost six years ago I wrote chess4j + Prophet4 POC , where I laid out an idea for a proof-of-concept to marry up my Java engine chess4j and my C engine Prophet. The idea was to use native C code, which typically executes much faster than Java code, for the core engine, but to use the higher level language for all the ancillary features. Of course, if using a higher level language was the only goal, C++ would have been the obvious choice. This is a hobby though, and one of the goals for chess4j was to explore Java based technologies. So, I wrote a Java Native Interface (JNI) layer as the bridge between the two codebases.

The learning curve for JNI is fairly steep, and it’s cumbersome to work with, but in the end it came together and works well. As predicted it gave chess4j a nice speed boost, and I haven’t had to write some of the features in chess4j into Prophet. Prophet has no opening book, no pondering support, no test suite support, and no HCE auto-tuner since all I have to do is run chess4j with the Prophet engine to leverage those features.

As I mentioned, running Prophet as a native library within chess4j does give chess4j a nice speed boost. But now I am asking the inverse question: how much slower is chess4j + Prophet than just plain Prophet? So, I put together a small test and was unpleasantly surprised by the results. Below are the results of doing a 30 second search.

PositionProphet knpschess4j knpschess4j + Prophet knps
initial position542810062058
r2qr1k1/p4ppp/5n2/2PpN3/8/2N1P1P1/1PP3BP/R4RK1 b – –523910262086
2r1r1k1/p4ppp/8/2Pp4/N7/4qNP1/1PP2RBP/4R2K b – –501310662127
rnbqkbnr/pppppppp/8/8/8/8/PPPPPPPP/RNBQKBNR w KQkq –555711672115
rn1qk2r/ppp2ppp/3bpn2/3p4/6b1/1P2P1Q1/PBPP1PPP/RN2KBNR w KQkq –47749981800
rnbqkbnr/pp2pppp/8/3p4/8/2N2N2/PPPP1PPP/R1BQKB1R b KQkq –517610491984
r2qkb1r/1p1b1ppp/p3pn2/3p2B1/3P4/2N2N2/PPP2PPP/R2Q1RK1 b kq –531610852052
2kr1b1r/1pqb1p2/p3p2p/3pNp2/3P4/2NQ4/PPP2PPP/R3R1K1 b – –478310131941
2k3rr/1pqb4/p2bpp1p/3p1p2/NP1P4/P2Q1N1P/2P2PP1/R3R1K1 b – –44409441924
2k4r/1pqN4/p2b3p/3Q1p2/1P6/P4p1P/2P2Pr1/R3RK2 b – –573612402555

On average, loading Prophet as a native library into chess4j is 1.95x the speed of chess4j by itself. That’s great! However, chess4j + Prophet is about 0.4x the speed of Prophet alone. That’s no so great.

So, why does this matter? It goes back to the goals for each project. chess4j was started as a test bed to learn about Java based technologies. I’ve never meant for it to be a competitive program, so from that perspective it really doesn’t matter. That said, I would like Prophet to be a competitive engine.

Prophet has been on the Computer Chess Ratings List for some time. Prophet 5.0 is rated 2648 in Blitz as I write this. They test Prophet as a standalone engine, so this issue has absolutely no bearing on that. The CCRL team uses a generic opening book, not the engine’s opening book, and they disable pondering, so it’s not an issue for Prophet to be missing those features. Sometimes I enter programmer tournaments though, where programmers run their programs on their own machines in whatever configuration they’d like. Opening books do matter here. Pondering matters. So, in these tournaments I run chess4j + Prophet to take advantage of those features. But, apparently I’m taking a massive penalty in terms of speed to do so. What to do?

If I want to compete in programmer tournaments with Prophet, I’m going to have to find a way to significantly reduce that overhead, or add opening book and pondering support directly into Prophet. I’m currently migrating away from JNI to Java’s Foreign Function and Memory API (FFM). FFM is Java’s new approach to interacting with native code. At the very least, it will significantly simplify the codebase and the build process. I’m hopeful it will also be more performant. I’m guesstimating this migration will be complete towards the end of Sept. or early Oct. Once it is, I’ll repeat the test above and make some decisions.

UPDATE 9/16/25: SEE FOLLOW UP IN Goodbye JNI, Hello FFM . (TL;DR : JNI was not the issue. shared vs static compilation of the Prophet library was.)