Re: [PATCH v4 00/11] s390: More this_cpu_*() changes
From: Mete Durlu
Date: Tue Sep 22 2026 - 11:44:10 EST
On 21/09/2026 17:57, Heiko Carstens wrote:
[..snip..]
v1:
Most of this is only about cleaning up the percpu code after preemptible
this_cpu_*() operations have been implemented. The first ten patches are
all more or less trivial cleanup patches trying to make the code shorter
and more readable.
The only non-trivial patch is the last one, which converts s390's
this_cpu_*() operations to use a similar scheme like Mark Rutland
provided it for arm64 [1]. This allows to simplify the irq entry and exit
path, however at the cost of slightly worse code for this_cpu_*()
operations.
The simplified irq entry and exit code seems to be worth it. Usable
performance numbers are not available yet, however I don't expect big
difference to before.
[1] https://lore.kernel.org/all/20260904161758.376504-1-mark.rutland@xxxxxxx/
Thanks,
Heiko
Heiko Carstens (11):
s390/percpu: Fix comment typo
s390/percpu: Add sanity check to GEN_MVIY macro
s390/lowcore: Remove _AC() from LOWCORE_ALT_ADDRESS
s390/percpu: Let MVIY_PERCPU() calculate alternative displacement
s390/percpu/lowcore: Add and use LC_PERCPU lowcore offset defines
s390/percpu: Rename inline assembly symbolic names
s390/percpu: Use __PCPU_BEGIN() and __PCPU_END() for inline assemblies
s390/percpu: Use percpu code section for this_cpu_cmpxchg128()
s390/percpu: Use percpu code section for this_cpu_xchg()
s390/percpu: Use percpu code section for this_cpu_cmpxchg()
s390/percpu: Rework to simplify percpu_entry() and percpu_exit()
arch/s390/include/asm/entry-percpu.h | 71 ++----
arch/s390/include/asm/lowcore.h | 5 +-
arch/s390/include/asm/percpu.h | 354 +++++++++++++++++----------
arch/s390/kernel/irq.c | 10 +-
arch/s390/kernel/nmi.c | 4 +-
arch/s390/kernel/traps.c | 4 +-
6 files changed, 248 insertions(+), 200 deletions(-)
Hi all,
I did some benchmarking for this patch series.
Used a shared LPAR with 24 CPUs (12 cores w SMT2).
Base commit: f0100363d8c3 ("Merge tag 'xfs-fixes-7.3-rc5' of
gitolite.kernel.org:/pub/scm/fs/xfs/xfs-linux")
Some of the benchmark runs show +-5% standard deviation which can
be attributed to s390 being a virtualized system. Each run has been
repeated five times to circumvent the effects of standard deviation.
(In the odd case where baseline has -5% and patched run has +5%
deviation, the comparison can show ~+10% improvements which is highly
misleading).
Overall, the benchmarks do not show any real difference between baseline
and patched kernels;
1-) Hackbench
========================================================================
Repeated hackbench runs with different group and fd count combinations;
[1, 2, 4, 8]
$ hackbench -T -p -l $loops -g $g -f $f
no real difference.
2-) Stress-ng
========================================================================
Repeated stress-ng runs with range of stressors: [6, $(nproc)]
# 3d matrix operations
$ stress-ng --matrix-3d $cpu --matrix-3d-method mult --timeout 10
# repeated mmap() and munmap() operations and also writing to the
# allocated memory.
$ stress-ng --vm $cpu --vm-bytes 128M --timeout 10
# starts workers that each fork off 32 child processes. Each child
# tries to allocate some memory. Child processes use madvise and memset
# to produce VM activity.
$ stress-ng --mmapfork $cpu --mmapfork-bytes 128M --timeout 10
no real difference.
3-) Openblas Benchmark
========================================================================
Repeated matrix operations with different size of 3d matrices.
$ make -j$(nproc) \
USE_OPENMP=1 \
NUM_THREADS=$(nproc)\
NOFORTRAN=1 > /dev/null 2>&1
$ make -C benchmark sgemm.goto \
USE_OPENMP=1 \
NUM_THREADS=$(nproc) \
NOFORTRAN=1 > /dev/null 2>&1
$ OPENBLAS_LOOPS=50 sgemm.goto
no real difference between runs.
4-) Compiling kernel
========================================================================
Linux kernel compilation
$ make clean
$ make mrproper
$ make defconfig
$ time make -j$(nproc)
no real difference between runs.