September 27, 2026
What if the prefetcher is wrong for my code?
Machine defaults are someone else's guess. I put Intel's hardware prefetchers on a per-core switch in /proc so I can benchmark goblin-store, goblin-core, and even the Python interpreter against a CPU that stops guessing.
Every machine you buy arrives already tuned. Someone picked the memory timings, the governor, the prefetcher settings, and a hundred other defaults, and they picked them for a generic workload that is probably not yours. We treat those choices as the shape of the hardware, but they are choices, and a performance decision is never free. Even the good ones cost something; they just spend it somewhere you were not looking.
The hardware prefetcher is my favorite example, because it looks like pure upside. A modern Intel core watches your memory accesses and, when it thinks it sees a pattern, pulls the next lines into cache before you ask for them. When the guess is right, the data is waiting and the load is free. There are four of these predictors on the Xeon this runs on: an L2 stream prefetcher, an L2 adjacent-line prefetcher, and two at L1, the DCU next-line and DCU IP-history prefetchers. All of them are on by default, and most of the time nobody thinks about them.
Here is the part that is easy to forget. The cache is finite. Every line the prefetcher speculatively pulls in has to take the place of a line that was already there. If the prediction is right, you win. If it is wrong, you pay twice: once in memory bandwidth for a fetch you never needed, and again — the expensive part — because you evicted a line that was still live and turned a future hit into a miss. The cost of a bad prefetch is not the fetch. It is the thing you threw away to make room for it. Prefetching buys latency hiding at the opportunity cost of the cache contents it invalidates.
For a workload with long, linear, predictable scans, that trade is a bargain. For a workload that chases pointers, or streams data past exactly once, or walks a hash table in a deliberately unpredictable order, the predictor can be systematically wrong, and then it is not a subsidy but a tax you cannot see on any profiler that only counts your instructions.
These knobs are not secret. On some machines you can toggle the prefetchers in
the BIOS. But the BIOS setting is machine-wide and it means a reboot to change,
which makes it almost useless for the thing I actually want to do, which is
measure. You cannot A/B a setting you have to reboot to flip, and you certainly
cannot give two processes on the same machine different settings. So the
question I wanted to answer was narrower and more practical: what happens if
this feature is exposed to the user, per core, in /proc, so it is trivial to
benchmark against?
So I built it. It is a small patch to a mainline Linux kernel. Each task carries
a mask of which prefetchers it wants disabled; the kernel writes Intel's
MSR_MISC_FEATURE_CONTROL (0x1A4) at context switch, so the setting follows
the process onto whatever CPU it lands on. Userspace controls it through a
directory that appears under every process:
ls /proc/self/prefetch_disable/
# l1_ip l1_stream l2_adjacent l2_stream
echo 1 > /proc/self/prefetch_disable/l1_stream # turn the L1 next-line prefetcher off
echo 1 > /proc/self/prefetch_disable/l1_ip # and the L1 IP-history one
cat /proc/self/prefetch_disable/l2_stream # 0 == still enabled
Each file is a 1 (that prefetcher disabled for this task) or a 0 (enabled).
A process can set its own, or, with the right permissions, another process's,
which matters when the thing you want to slow down is not the thing holding the
shell.
| file | prefetcher |
|---|---|
l2_stream |
L2 hardware stream prefetcher |
l2_adjacent |
L2 adjacent cache-line (128-byte pair) |
l1_stream |
L1 data-cache next-line prefetcher |
l1_ip |
L1 data-cache IP-history prefetcher |
There is one genuinely interesting wrinkle. The register is per-core, and a core has two hyperthreads that share it. So "this process wants the L2 prefetcher off" is not a question one thread can answer alone. I resolved it the conservative way: a prefetcher is disabled on the core if either sibling asks for it to be off. That is the right default for isolating an effect, and it is exactly the configuration you want when you deliberately place two processes on the two threads of one core and watch them interfere. To keep the whole thing cheap enough to leave on, each core caches the value it has already programmed, so the actual register only gets written when the effective setting changes, not on every context switch.
Alongside it I keep an unmodified kernel, identical in every other way, as the baseline. That is the only honest way to separate three different things: the cost of the feature existing, the cost of using it, and the effect of the prefetchers themselves.
Which brings me to why I wanted this at all. I have two pieces of storage software I care about, goblin-store.dev and goblin-core.dev, and their access patterns are not generic. Storage engines spend their lives deciding, very deliberately, which bytes to touch and in what order. That is precisely the kind of code where the hardware's guess about what I will read next might be helping me, or might be quietly evicting the pages I was about to use. I want to run them against a machine that has stopped guessing and see which way it moves. And once the knob exists, it is hard not to point it at other things — the Python interpreter, for one, which is nothing but pointer chasing dressed up as a language.
There is a humbling version of this and a funny version, and they might be the same version. I have spent years tuning this software with the prefetchers on, which means every layout decision I have made was made in their presence. Maybe they have been rescuing a data structure I would otherwise have to fix. Maybe they have been fighting a careful access pattern the whole time and I have been optimizing around the damage. Wouldn't it be something if we had been making the wrong choices all along, not because we reasoned badly, but because we were reasoning on top of a default nobody chose on purpose for us?
I do not have the numbers yet. The tool is built and the machine is coming up on the patched kernel now. But the setup is the argument I wanted to make first: the defaults are a guess, the guess has a cost, and the least I can do before trusting it is put it on a switch I can flip.