Re: [RFC PATCH] mm/mglru: dynamically protect readahead fault folios under refault pressure

From: Kairui Song

Date: Thu Oct 01 2026 - 14:11:40 EST


On Fri, Oct 2, 2026 at 1:21 AM Ababneh, Ehab <ehab.ababneh@xxxxxxxxx> wrote:
>
> Hi Kairui,

...

> > > > > Results with MGLRU-FG:
> > > > > Op rate: 55.0k - 56.5k op/s

...

> > > > > Latency 99th percentile: 6.8 - 6.9 ms
> > > > >
> > > > > For reference, here's where the other variants landed on the same setup:
> > > > >
> > > > > Regression (6cbdd9726fb5, "mm/mglru: use folio_mark_accessed to
> > > > > replace folio_set_active"):
> > > > > p99 ~9.2-9.5 ms, throughput ~41.8k-43.6k op/s
> > > > >
> > > > > Revert of 6cbdd9726fb5:
> > > > > p99 ~5.5-5.6 ms, throughput ~51.9k-53.1k op/s
> > > > >
> > > > > My dynamic readahead-credit fix:
> > > > > p99 ~5.8 ms, throughput ~51.9k-52.7k op/s
> > > > >
> > > > > So MGLRU-FG recovers most of the latency regression, but the
> > > > > revert and my fix still recover more of it: p99 with MGLRU-FG is
> > > > > noticeably higher than with the revert and my fix (6.8-6.9 ms vs
> > > > > ~5.5-5.8 ms), though it's still a big improvement over the
> > > > > regression's 9.2-9.5 ms
> >
> > Did you test the V2 or V1 of MGLRU-FG? V2 includes Barry's aging
> > optimization; I'm not sure if this latency is caused by aging or maybe some
> > other change.
> >
> > And is there any swap device you used, or is the test using the same base
> > commit? MM also changes significantly between different baselines.
>
> My setup runs Cassandra 4.1.5 on Java 11. The host has two Intel Xeon
> Platinum 8468 sockets, 96 physical CPU cores (192 logical CPUs), two
> NUMA nodes, and 440 GiB of system-visible RAM. Cassandra instances
> use CPU affinity, with memory interleaved across NUMA nodes.
>
> I run four Cassandra instances, each backed by a dedicated 3.5 TB
> NVMe drive. Each drive has a separate mount, and holds its instance’s
> data files, commit log, and saved caches.
>
> The stress workload uses a pre-populated table. Its schema, partitioning,
> replication, compaction, and compression settings, along with the stress
> profile and concurrency, are recorded because they can affect performance.
>
> I used more than one kernel version across my test runs. The most recent
> was Linux 7.3.0-rc5. Within each set of setups being compared, I kept the
> kernel version consistent.
>
> For the MGLRU-FG tests, I used the ryncsn/b4/mglru-fg-v1.8 branch. The
> MGLRU-FG implementation is commit 6bcf4ae70b5e; the branch tip
> is a619ab0105d9.

Thanks for the info! That v1.8 is missing recent upstream changes
indeed, so aging could be a bit slower, causing the jitter. I also
found a few other minor improvements doable, I'll also try to run some
Cassandra benchmarks too. Higher throughput and worse P99 is an
interesting effect, I think it means the LRU is doing better at
protecting the working set and maybe spending a bit more effort on
reclaim / aging, which could be optimized away. I will try to
reproduce and optimize for that as well while keep the series updated.