Re: [PATCH v4 3/4] pps: Always use ktime_get_snapshot_id() for pps_get_ts()
From: David Woodhouse
Date: Fri Oct 02 2026 - 05:09:27 EST
On Fri, 2026-10-02 at 09:04 +0200, Rodolfo Giometti wrote:
> On 01/10/2026 15:14, Miroslav Lichvar wrote:
> > On Mon, Sep 28, 2026 at 09:58:02AM +0200, Rodolfo Giometti wrote:
> > > On Sat, 2026-09-26 at 21:38 +0100, David Woodhouse wrote:
> > > > So yes, the specific code path you're looking at *does* get slightly
> > > > longer (50ns to the counter read instead of 30ns).
> > > [...]
> > > > However, they are *entirely* in the noise, as there's about 600 ns of
> > > > hardware and 2-4 *microseconds* of software latency before we even get
> > > > there.
> > >
> > > Thanks for measuring it, and on real hardware with a real edge. That
> > > answers my concern: ~20 ns of constant cost against microseconds of
> > > latency upstream of the handler is not something PPS can see, and
> > > trading it for the removal of a non-constant error is the right
> > > trade.
> >
> > FWIW, there are some polling versions of the pps-gpio driver (out of
> > tree), which provide much more stable measurements by trading
> > interrupts for higher CPU use and where this additional delay might be
> > visible.
> >
>
> If polling gives that much better stability, I'd be glad to have it in
> drivers/pps/clients, as a mode of pps-gpio or as a client of its own.
> Its timestamping needs would then be part of the discussion about the
> timestamping interface started in the 2/4 thread.
Found it: https://github.com/mlichvar/pps-gpio-poll/blob/master/pps-gpio-poll.c
I kicked off a test run last night, results at
• https://david.woodhou.se/ntptest-r64/spin0-tickful-1hz/
For comparison, 1Hz normal tickful and my entry.S (backdate) hack are:
• https://david.woodhou.se/ntptest-r64/tickful-1hz/
• https://david.woodhou.se/ntptest-r64/backdate-tickful-1hz/
I'm looking mostly at the 'Pulse arrival phase' charts, and the full
data sets are linked from the bottom if you want to play more.
> Do you see the extra delay there as a shift of the offset or as more
> jitter?
The polling mode addresses both. For the *shift* it'll eliminate the
actual hardware interrupt path, in this case through the EINT, any
debouncing, and the GIC (and disabled interrupts). I previously did a
loopback GPIO test which measured that at around 800ns, which
*includes* the MMIO write to GPIO too. If needed, I *could* try to set
up something where one CPU polls while the other takes the interrupt
(and captures the counter in entry.S), for a better measurement. It'd
be very hardware-specific though.
Polling will also eliminate the multi-microsecond path we observed from
the exception to the pps_gpio_irq_hardirq() handler, which for my test
board was a median shift of ~2.2µs (p99 ~5.6µs, p100 13.6µs). That's
the periodic run; the tickless kernel was 2-3µs higher than that.
For the jitter... it's still quite wide, which I think is largely down
to the time it takes to do the GPIO read itself (qv). My entry.S
capture is much tighter, not that I can see a way to merge that for
real. (And my testing is currently done on an idle system from an
initramfs, so the system doing actual *work* and occasionally disabling
interrupts might change the picture even for the efficacy of the
entry.S hack, let alone its practicality.)
I've modified the spin code to take the shape I outlined in email last
night, just capturing the *counter* bracketing the GPIO read and then
converting it to 'time' later. I'll have results in a few hours.
This is the table showing the per-pulse phase delta (ns) for the three
modes linked above, for now:
┌───────────────────────────┬─────┬──────┬──────┬──────┬─────┐
│ capture method │ p50 │ p95 │ p99 │ max │ σ │
├───────────────────────────┼─────┼──────┼──────┼──────┼─────┤
│ pps-gpio (IRQ) │ 84 │ 1206 │ 4013 │ 4889 │ 688 │
│ polling, stamp after edge │ 206 │ 536 │ 665 │ 1106 │ 286 │
│ entry.S counter capture │ 46 │ 109 │ 135 │ 206 │ 58 │
└───────────────────────────┴─────┴──────┴──────┴──────┴─────┘
Having typed the above to look at the case for the gpio-poll driver, I
now realise "the extra delay" in your question was perhaps referring to
the 10-20ns extra that ktime_get_snapshot_id() would take. That should
be fairly much constant, in the current gpio-poll driver.
It's fairly much noise even compared with the jitter in the current
gpio-poll mechanism, purely because the latency of the gpio_read()
itself limits the precision. (My friend estimates that as 250-350ns per
read based on other measurements, which would mean σ ~95ns but I'll
have a proper measurement of that when I harvest the currently running
test.)
There's *also* an occasional extra delay (0.5-1.5µs est.) when *either*
timestamp function is used, if it hits the seqcount retry. In fact,
even *checking* the seqcount, when it *doesn't* change, has a cost in
the same nanoseconds range as the ktime_get_snapshot_id() switch.
None of which has to live in the critical path at all, hence my comment
about 'not even trying' and the variant I'm testing right now which
*only* grabs the counter, and does the actual work of converting to a
timestamp later.
Next step, if we wanted, might be to add a 'wait_for_edge' method to
the GPIO which literally does let the driver spin on the MMIO. (And
even if implemented in GPIO core with a fallback, could reduce the
overall sampling period).
Attachment:
smime.p7s
Description: S/MIME cryptographic signature