Re: [RFC PATCH v3 1/1] arch: arm64: Implement unaligned atomic emulation
From: André Almeida
Date: Thu Oct 01 2026 - 13:55:23 EST
On 2026-09-30 15:29, Will Deacon wrote:
> On Wed, Sep 30, 2026 at 07:29:22AM -0300, André Almeida wrote:
>> Em 30/09/2026 04:48, Will Deacon escreveu:
>> > On Tue, Sep 29, 2026 at 11:01:38PM -0300, André Almeida wrote:
>> > > Implement support for emulating unaligned atomic operations on arm64.
>> > > User applications that wish to enable support for this should use the
>> > > pctrl() flag `PR_ARM64_UNALIGN_ATOMIC_EMULATE`.
>> > >
>> > > Signed-off-by: André Almeida <andrealmeid@xxxxxxxxxx>
>> > > ---
>> > > arch/arm64/Kconfig | 6 +
>> > > arch/arm64/include/asm/exception.h | 1 +
>> > > arch/arm64/include/asm/processor.h | 5 +
>> > > arch/arm64/include/asm/rwonce.h | 14 +-
>> > > arch/arm64/include/asm/thread_info.h | 1 +
>> > > arch/arm64/kernel/Makefile | 3 +-
>> > > arch/arm64/kernel/process.c | 15 +
>> > > arch/arm64/kernel/unaligned_atomic.c | 521 +++++++++++++++++++++++++++
>> > > arch/arm64/mm/fault.c | 10 +
>> > > include/uapi/linux/prctl.h | 5 +
>> > > kernel/sys.c | 7 +-
>> > > 11 files changed, 579 insertions(+), 9 deletions(-)
>> > > create mode 100644 arch/arm64/kernel/unaligned_atomic.c
>> >
>> > No.
>> >
>> > I already explained to you why this doesn't work:
>> >
>> > https://lore.kernel.org/r/aV1YnOetDHhKe4hz@willie-the-truck
>>
>> Indeed, last time you raised some points, and then Ryan replied them. Is
>> there any specific point that doesn't work? I couldn't find a reply for
>> Ryan's answers:
>>
>> https://lore.kernel.org/all/CABnRqDf5EQUoXu=pJ6mj4-JfwAzEfcAE2cYrNzJANFycx7cMUA@xxxxxxxxxxxxxx/
>
> So rather than get involved in the discussion, you did nothing for almost
> a year and then resent the exact same patch? Why?
>
> I don't think the implementation is correct and I don't think we should
> be emulating this either. I hope I made that clear last year. Ryan
> thinks it's "fine" due to the locking, but I don't see how that helps
> with the example I gave.
Good, now we're talking. Now is clear that the most problematic part is
the split lock "emulation".
I agree that your example would cause tearing, and it would create a
memory state that would never happen with such instructions, or the
equivalent
of them in x86. Now I wonder which options do we have to have some sort
of bus
locking here. Giving that this instruction is already super expensive in
x86 anyways,
we don't need to be super quick as well, but we don't want a system
freeze as well:
- Somehow pause the other CPUs while the two instructions are happening.
I wonder if
we need to pause all of them or just the ones that share the specific
cache, or is the
LLC always shared amongst all cores?
- The kernel side lock doesn't serialize all users, so maybe there could
be a Giant Lock
in userspace for all threads sharing this resource. My cover letter is
outdated in that
regard, giving that the new set_robust_list2() will enable multiple
robust lists per process.
- Could flush instructions be useful somehow how?
While I try to test some of this ideas, any other ideas that you might
have
about locking two cache lines would be super useful.
Thanks again for your time!
André