Profile-Guided Optimization for Prophet: A Regression That Wasn’t

Co-authored by Claude based on context and transcripts collected during a real-time debugging workflow.

Profile-Guided Optimization is supposed to be one of the safer performance wins available to a C/C++ project. You build once with instrumentation, run a representative workload, feed the resulting profile back into a second build, and the compiler makes better inlining, branch-layout, and register-allocation decisions based on how the code actually runs rather than static heuristics. In the worst case it does nothing. It should never make a chess engine play worse chess — a PGO build searches the same tree with the same evaluation, just (hopefully) a bit faster.

So when I wired up scripts/build-pgo.sh for Prophet and set a fresh PGO build loose against a plain Release build in a real match, I expected a small, boring win. Instead I got this:

Score of prophet-dev vs prophet-latest: 100 - 236 - 201  [0.373] 537
     prophet-dev playing White: 93 - 35 - 141  [0.608] 269
     prophet-dev playing Black: 7 - 201 - 60  [0.138] 268
     White vs Black: 294 - 42 - 201  [0.735] 537
Elo difference: -89.9 +/- 23.5, LOS: 0.0 %, DrawRatio: 37.4 %
SPRT: llr -1.33 (-45.2%), lbound -2.94, ubound 2.94

prophet-dev was the PGO build. prophet-latest was a plain Release build of the same commit. Same source, same neural network weights, same everything except an extra compiler pass — and PGO was losing by the better part of 90 Elo. Granted, the error bars are pretty high with just 537 games, but it’s clear the PGO build is just worse. It took the better part of two weeks, several false leads, and a full concurrency sweep to get to the bottom of it.

Ruling out the code

A result like that raises a correctness question before anything else: is something actually wrong, not just slower? The elimination process was mostly writing small scripts to compare the two binaries directly rather than through hundreds of games of noise:

  • Compile flags — confirmed identical between the PGO and baseline builds (same -Wall -Werror -Wextra, same -flto, same -mavx2 -mbmi2; the only difference is -fprofile-use vs. nothing).
  • Fixed-depth search determinism  — 11 positions, fixed-depth (d=10), comparing best move and score. Matched on every one.
  • Sustained throughput — fixed-time (10s) searches on four representative positions, run alone, no contention. PGO was faster on every single one.
  • Code size — I had a theory that PGO’s more aggressive inlining might bloat the binary enough to hurt icache behavior under load. The .text segment was actually smaller for the PGO build (79,162 bytes vs. 101,983 bytes), which killed that theory outright.

Every controlled, single-process test said the two binaries were functionally identical and the PGO build was, if anything, a little faster. And yet the real match said PGO was losing by 90 Elo. At that point I began to suspect I had a test environment problem.

The actual variable

The turning point was the realization that every test I’d run had compared the two binaries running alone. The real match hadn’t — it was configured with -concurrency 8, and cutechess-cli spawns two engine processes per game. On an 8-core/16-thread Ryzen 5800X, -concurrency 8 means up to sixteen engine processes competing for sixteen logical threads: full SMT saturation, zero headroom, all the time. (Note: Prophet does not use “pondering mode,” where the engine can search on the opponent’s time, so CPU contention between active threads was never the issue.)

I hadn’t reproduced that in any of my isolated tests. So I wrote one that did — launching eight simultaneous prophet-dev searches and eight simultaneous prophet-latest searches at once, matching the real concurrency — and found… nothing. Both binaries degraded under contention by roughly the same amount (1.05x in PGO’s favor, if anything). Same-vs-same contention didn’t reproduce the asymmetry either.

The real test was just running the actual match at lower concurrency. A control run at effectively -concurrency 1 immediately flipped the sign: 58.0% at 44 games, settling to 53.0% (Elo +20.9 ± 47.4) at 100 games — noisy, but nothing like the -90 Elo collapse. A longer confirmation run at -concurrency 4 over several days put it beyond doubt:

Score of prophet-dev vs prophet-latest: 2012 - 1706 - 2899  [0.523] 6617
Elo difference: 16.1 +/- 6.3, LOS: 100.0 %, DrawRatio: 43.8 %
SPRT: llr 2.95 (100.0%), lbound -2.94, ubound 2.94 - H1 was accepted

PGO wasn’t a 90-Elo regression. It was a genuine, statistically solid ~16-Elo improvement — once the test stopped saturating the machine.

The concurrency sweep

-concurrency N spawns 2N processes; the machine has 8 physical cores. N=4 is exactly one process per physical core — no SMT sharing required — so the obvious next question was whether the degradation was a hard cliff at that boundary or a gradual slope above it. The first pass:

ConcurrencyProcessesGamesScoreElo
486,6170.523+16.1 ± 6.3
5105,1810.471-20.5 ± 7.5
6129570.336-118.0 ± 18.4
8165370.373-89.9 ± 23.5

Fine right at the physical core count, a steep penalty just above it, roughly plateauing once the machine is thoroughly oversubscribed. Then I ran N=3 — six processes, comfortably under the 8-core boundary, which the theory said should be at least as clean as N=4 — and got this:

RunGamesScoreElo
N=37,4600.523+16.3 ± 6.3

That landed almost exactly on top of N=4‘s +16.1, which made sense — the machine was not oversubscribed.

Rerunning everything

With N=3 confirmed, I reran N=5, N=6, and N=8 at a fixed 10,000 games apiece. (Larger sample sizes give tighter confidence intervals.) Every one of them came back with a smaller, more moderate magnitude than the original small-sample read:

ConcurrencyProcessesGamesScoreElo (original)Elo (rerun, fixed 10k)
51010,0000.418-20.5 ± 7.5-57.4
61210,0000.407-118.0 ± 18.4-65.1 ± 5.5
81610,0000.241-89.9 ± 23.5-199.3 ± 6.8

N=6 in particular moved a lot — from a small-sample -118.0 down to a confirmed -65.1 — which fit the pattern of small samples reading more extreme than the truth. N=8, on the other hand, moved the other way, from -89.9 up to a much more severe, tightly-bounded -199.3.

That set up an obvious next test: N=7, at fourteen processes — two threads still idle, one core’s worth short of full saturation. If the curve was smooth, N=7 should sit between N=6’s -65.1 and N=8’s -199.3. Instead:

Score of prophet-dev vs prophet-latest: 1105 - 7004 - 1891  [0.205] 10000
Elo difference: -235.4 +/- 7.1, LOS: 0.0 %, DrawRatio: 18.9 %

Worse than N=8! Not close, either — the confidence intervals didn’t overlap. That’s a genuinely strange result. I had a mechanism ready to explain it: with idle threads available, the Linux scheduler has somewhere to migrate processes to when it rebalances load, and each migration costs a cold cache — while at full saturation there’s nowhere to migrate to, so assignments settle and stay put, giving a slower but more stable steady state. It’s a plausible story. It also should be replicated before I believed it. So I reran N=7:

Score of prophet-dev vs prophet-latest: 1538 - 5977 - 2485  [0.278] 10000
Elo difference: -165.8 +/- 6.3, LOS: 0.0 %, DrawRatio: 24.9 %

Two full 10,000-game runs, identical settings, landing 70 Elo apart with non-overlapping confidence intervals. Pooling both runs (20,000 games: 2643-12981-4376) gives a combined estimate of roughly -198.7 Elo — which lands almost exactly on top of N=8’s -199.3. The tidy, counterintuitive “N=7 is worse than full saturation” finding dissolved the moment I checked it twice. N=7 and N=8 are most likely the same number, both sitting in the same severe-collapse plateau, and my scheduler-migration theory was a good explanation for a result that turned out not to exist.

The complete picture

ConcurrencyProcessesGamesScoreElo (dev vs. baseline)
367,4600.523+16.3 ± 6.3
486,6170.523+16.1 ± 6.3
51010,0000.418-57.4
61210,0000.407-65.1 ± 5.5
71420,000 (2 runs, pooled)0.242≈ -198.7
81610,0000.241-199.3 ± 6.8

The result that emerges is clean, and slightly favorable to PGO, at or below the physical core count (3-4); a real but moderate penalty just above it (5-6); and a severe, decisive collapse once you’re deep into SMT-shared territory (7-8).

Where this leaves the feature

The original question is answered: PGO is a real, if modest, improvement for Prophet — about +16 Elo, confirmed twice independently, at a concurrency setting that respects the physical core count. Left purely on the merits, it would ship.

It isn’t shipping, at least not as a default. The reason is everything above: getting a trustworthy read out of this required a full concurrency sweep and multiple reruns. I don’t control what concurrency other people testing or building Prophet will run their matches at, and I can’t hand someone a build flag with a footnote that says “only trust this if you keep concurrency at or below your physical core count and ran it more than once.” That’s not a maintainable claim to stand behind. PROFILE_GUIDED stays off by default in CMakeLists.txt – but scripts/build-pgo.sh and the comparison scripts I wrote along the way stay in the tree as opt-in tooling, in case the calculus changes later or someone wants to reproduce any of this.

Lesson (re-re-re…) learned

It’s interesting that a Profile Guided Optimization build degrades in performance relative to the “normal” build as resource contention increases. The bigger lesson– your test environment should match real world conditions as closely as possible.

Side note: as stated in Prophet 5.2 released, I’ve decided not to use AI to make any algorithmic changes in Prophet, but I did leverage Claude for this experiment. (And I will heavily leverage AI going forward in chess4j.)

As always, the latest can be downloaded from the Prophet Github site.