The per-CPU TB jump cache has held 4096 entries since it was introduced.
That is too small for guests running large programs: an emulated compiler
misses often enough that the fallback qht lookup shows up prominently in
a profile.
Measured with qemu-alpha running an emulated alpha gcc 16.2.0 compiling
the SQLite 3.45.1 amalgamation (255k lines, -O2) on an x86-64 host. The
compile performs 34.2 billion TB executions, of which 8.4 billion take
the indirect dispatch path.
Sizing curve, on top of the preceding patch, instructions retired and
wall clock:
12 bits ( 64 KiB): 1,563,829,403,943 133.13s
14 bits ( 256 KiB): 1,493,865,985,972 -4.47% 124.89s -6.19%
16 bits ( 1 MiB): 1,469,729,281,442 -6.02% 120.97s -9.13%
18 bits ( 4 MiB): 1,462,262,363,257 -6.49% 120.16s -9.74%
16 bits is the knee. 18 buys another 0.47% of instructions for four times
the memory, and since instructions retired does not account for the data
cache pressure of a 4 MiB table, that 0.47% is probably not real: the
wall clock difference between 16 and 18 bits is 0.67%, against a
run-to-run spread of the same order.
In a perf profile the mechanism is visible directly: tb_htable_lookup(),
which is where qht_lookup_custom() lands once it is inlined in an LTO
build, falls from 5.73% of samples to 1.52%.
The cost is memory: the cache grows from 64 KiB to 1 MiB per vCPU. That
is easy to justify for a single-vCPU linux-user process and less obvious
for system emulation with many vCPUs, so this may want to be sized by
target or made tunable rather than raised unconditionally.
Signed-off-by: Matt Turner <[email protected]>
---
accel/tcg/tb-jmp-cache.h | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git ./accel/tcg/tb-jmp-cache.h ./accel/tcg/tb-jmp-cache.h
index c3a505e394..268dacd7ba 100644
--- ./accel/tcg/tb-jmp-cache.h
+++ ./accel/tcg/tb-jmp-cache.h
@@ -12,7 +12,7 @@
#include "qemu/rcu.h"
#include "exec/cpu-common.h"
-#define TB_JMP_CACHE_BITS 12
+#define TB_JMP_CACHE_BITS 16
#define TB_JMP_CACHE_SIZE (1 << TB_JMP_CACHE_BITS)
/*
--
2.54.0