RE: [PATCH v5 26/27] Documentation: vfio-pci: Document CXL Type-2 device passthrough
From: Manish Honap
Date: Mon Sep 21 2026 - 05:51:45 EST
Hello Gregory,
Thank you for the careful read of the documentation. I will rework the
sections you flagged as below:
- Address model
I will update the wording to mention that VMM does know the GPA: it builds
the guest's CFMWS window and so chooses the guest-physical range the device
can inhabit. What the host owns is the HPA and the placement within it, not
knowledge of the GPA.
I will add the two models you described, fixed placement for an accelerator
that needs exact physical placement versus no fixed placement for pooled or
compressed memory, with a small HPA/GPA example, and note that the current
series under discussion implements the fixed-offset case.
- Virtual decoders
I will rename the The "Guest decoder and commit" section to "Virtual
decoders" and add a short flow showing a guest vdecoder write being absorbed
and a live read returning COMMITTED from the host-locked physical decoder.
- "The kernel virtualizes ..."
I will update it to say vfio-pci-core (its config-space permission hooks).
- guest IOAS, stage-2 mapping, and struct-page-less coherent memory
I will define these terms where first used and include that the range is
handed to the device whole and never onlined as system RAM. I will also add
details why a stale stage-2 mapping must not outlive the HDM window, and how
the VMM rebuilds it.
- Reset
I will reword this section. FLRs are virtualized so a guest reset can't
inflict physical CXL.mem/decoder effects on the host rather than loose
wording "must not take an FLR"
I can share the revised text ahead of the v6 posting if that is easier to
review.
Thanks,
Manish
> -----Original Message-----
> From: Gregory Price <gourry@xxxxxxxxxx>
> Sent: Thursday, September 17, 2026 1:03 AM
> To: Manish Honap <mhonap@xxxxxxxxxx>
> Cc: alex@xxxxxxxxxxx; jgg@xxxxxxxx; Ankit Agrawal <ankita@xxxxxxxxxx>;
> jic23@xxxxxxxxxx; dave.jiang@xxxxxxxxx; alejandro.lucero-palau@xxxxxxx;
> Srirangan Madhavan <smadhavan@xxxxxxxxxx>; corbet@xxxxxxx;
> skhan@xxxxxxxxxxxxxxxxxxx; dave@xxxxxxxxxxxx; alison.schofield@xxxxxxxxx;
> vishal.l.verma@xxxxxxxxx; iweiny@xxxxxxxxxx; ming.li@xxxxxxxxxxxx; Yishai
> Hadas <yishaih@xxxxxxxxxx>; Shameer Kolothum Thodi
> <skolothumtho@xxxxxxxxxx>; kevin.tian@xxxxxxxxx; bhelgaas@xxxxxxxxxx;
> dmatlack@xxxxxxxxxx; kees@xxxxxxxxxx; gustavoars@xxxxxxxxxx; Neo Jia
> <cjia@xxxxxxxxxx>; Krishnakant Jaju <kjaju@xxxxxxxxxx>; Vikram Sethi
> <vsethi@xxxxxxxxxx>; Zhi Wang <zhiw@xxxxxxxxxx>; linux-
> doc@xxxxxxxxxxxxxxx; linux-kernel@xxxxxxxxxxxxxxx; kvm@xxxxxxxxxxxxxxx;
> linux-cxl@xxxxxxxxxxxxxxx; linux-pci@xxxxxxxxxxxxxxx; linux-
> kselftest@xxxxxxxxxxxxxxx; linux-hardening@xxxxxxxxxxxxxxx
> Subject: Re: [PATCH v5 26/27] Documentation: vfio-pci: Document CXL Type-
> 2 device passthrough
>
> External email: Use caution opening links or attachments
>
>
> On Thu, Sep 17, 2026 at 12:05:39AM +0530, mhonap@xxxxxxxxxx wrote:
> > From: Manish Honap <mhonap@xxxxxxxxxx>
> >
>
> 1) Thank you so much for writing documentation, i truly appreciate this.
>
> 2) I apologize in advance for my terseness, I know writing is hard,
> please do not interpret this as disliking your writing or series.
>
> > +Address model
> > +=============
> > +
> > +The HDM memory is a coherent host physical range (HPA). The host
> > +kernel resolves that range before the guest sees the device, and owns
> > +it for the bind lifetime. The guest only chooses where the memory
> > +appears in its own physical address space (GPA), by programming a
> > +virtual endpoint HDM decoder. The guest never reprograms the physical
> decoder.
> > +
> > +The kernel holds the HPA and does not see the GPA. The guest programs
> > +a GPA and does not see the HPA. The VMM holds the device fd, reads
> > +the committed base from the decoder-register region described below,
> > +and maps the HPA-backed HDM region at the GPA the guest committed.
> > +The base the guest reads back is the GPA, not the HPA.
> > +
>
> I think this must be slightly inaccurate / imprecise wording.
>
> The host *must* provide some form of physical memory window to the guest
> at initialization time, otherwise the guest has no way to know - at boot time -
> that there's even a window of memory it can use.
>
> That's what the CFMWS is. This is initialized by the hypervisor - which is
> controlled by the host.
>
> So the host (at least the VMM) must know, for the region the entire device
> *could* inhabit, what that GPA is - because it's the one that makes the CFMWS
> for the guest.
>
> If this is not the case, then something is missing from this documentation to
> explain why.
>
>
> If you're actually trying to say is that the GPA's programmed into the virtual
> decoders are largely symbolic - this at best feels a bit inaccurate and simply an
> implementation detail.
>
> The host's virtio device could enforce ....:
>
> Host Range:
> CFMWS HPA - [0x10000, 0x20000]
> | |
> Guest Range: | |
> CFMWS GPA - [0x50000, 0x60000]
>
> In that case, you'd get the following translation...
> vdecoder0.0 - [0x58000, 0x60000]
> CFMWS GPA - [0x58000, 0x60000]
> CFMWS HPA - [0x18000, 0x20000]
>
> Or the virtio device could not enforce that and let the host page-fault just hand
> it a random page from the actual CXL device.
>
> vdecoder0.0 - [0x58000, 0x60000]
> CFMWS GPA - [0x58000, 0x60000]
> |
> No discrete host mapping
>
>
> The former makes sense if the device (accelerator) requires exact physical
> placement to do its accelerator nonsense.
>
> The latter makes sense if the device (accelerator) doesn't care about placement
> (compressed memory).
>
> This is not saying we need support both out of the box, but we shouldn't lock
> ourselves into the former unless there's some reason why the latter is not
> reasonable.
>
> Can you please help document what the actual expected behavior is with
> examples in the Address model section so it's easier to understand the intent?
> That will help quite a bit.
>
> > +Guest decoder and commit
> > +========================
> > +
> > +The guest programs its virtual endpoint decoder through the trapped
> > +region: it writes a base (a GPA), a size, and then the COMMIT bit.
> > +The host already resolved and committed the physical placement before
> > +the guest ran,
>
> So the host does know GPA, just not exact placement.
>
> > so a live read of the decoder always shows COMMITTED and the
> > +guest's commit poll completes. The physical decoder is never
> > +rewritten; the guest's writes are absorbed.
> > +
>
> Rather clunky, round-about way to say "The guest decoders are
> virtualized". If possible, it would be nice to formalize this concept
> into "Virtual Decoders" - since that's what this is.
>
> With that concept i think you can probably generate some nice diagrams that
> show how the guest vdecoder's interact with the host drivers.
>
> > +The VMM observes the commit, reads the committed base, and maps the
> > +HDM region at that GPA.
> > +
>
> So the host does know the GPA.
>
> > +CXL Device DVSEC
> > +================
> > +
> > +The kernel virtualizes the CXL Device DVSEC body through the
> > +config-space permission hooks. Reads and writes inside the DVSEC body
> > +use a per-open shadow; a guest write stays in the shadow and does not
> reach hardware.
> > +Accesses outside the DVSEC body go to the device as usual.
> > +
>
> "The kernel" - what part? vfio-pci ? the vmm ?
>
>
> > +DMA and iommufd
> > +===============
> > +
> > +A Type-2 accelerator issues ATS-translated DMA to addresses inside
> > +its own HDM window, so that range must be present in the guest IOAS
> > +that backs the nested stage-2 translation.
>
> Type-2, ATS, DMA, HDM window, guest IOAS, stage-2 translation
>
> I think the only thing i don't know in this sentence is "guest IOAS" and it's still
> hurting my brain to read.
>
> Are all accelerators expect to have this particular interaction, or just yours?
>
> > The HDM range is struct-page-less coherent
> > +memory, which a userspace-VA ``IOMMU_IOAS_MAP`` cannot pin.
> > +
>
> The hardest part about writing about virtualization is keeping a consistent
> mental model from section to section.
>
> which userspace? guest? host? (i presume guest here)
>
> `struct-page-less coherent memory`
> e.g. the host never hotplugs this, it hands the entire region
> directly to the VFIO device, right?
>
> I think this would be nice to spell out somewhere.
>
> > +The HDM memory region is therefore exportable as a dma-buf:
> > +``VFIO_DEVICE_FEATURE_DMA_BUF`` on that region returns an fd that
> > +iommufd maps with ``IOMMU_IOAS_MAP_FILE``, mapping the physical
> range
> > +without a VA or a page pin. The dma-buf is revoked whenever the
> > +mapping is torn down (reset, power transition, teardown), so a stale
> > +stage-2 mapping cannot outlive the HDM window.
> > +
>
> For the sake of readers, I think either a little bit more information on this
> "stage-2 mapping" concept is needed to make sense of what's going on here
> and why it mustn't outlive the HDM window.
>
> > +Reset
> > +=====
> > +
> > +A CXL Type-2 function must not take a Function Level Reset: an FLR
> > +resets the coherent CXL.mem state and the HDM decoder. The PCI core
> > +reflects this by preferring the CXL reset over FLR, so a function
> > +reset of a CXL device runs the CXL DVSEC reset sequence, which resets
> > +the function and then restores the HDM decoder and the PCI config state.
> > +
>
> I think what you're trying to say is that FLRs are never passed to the device
> because it can cause physical device effects that defeat the purpose of the
> virtualization, yes?
>
> So we virtualize FLRs...
>
> > +A guest requests a reset by writing Initiate CXL Reset in the DVSEC.
> > +That write only stamps completion in the shadow. The real reset runs
> > +at the vfio reset points (the reset ioctl and a virtualized FLR through config
> space):
>
> As you describe here.
>
> So it's not that an accelerator "must not take an FLR" - it's that FLRs are
> virtualized to prevent deleterious effects on the host / hardware.
>
> Am I misunderstanding this?
>
> > +the kernel zaps the HDM mapping and revokes the dma-buf, then runs
> > +the CXL reset, which always clears the device memory, and restores
> > +and re-samples the decoder afterwards. A CXL port masks Secondary Bus
> > +Reset by default, so a ``VFIO_DEVICE_PCI_HOT_RESET`` does not reach
> > +the endpoint and the HDM state is untouched. If the port has SBR
> > +unmasked the reset can decommit the decoder without restoring it, so
> > +the reset_done handler gates HDM access; a ``VFIO_DEVICE_RESET`` then
> > +runs the CXL reset sequence and restores it.
> > +
> > +The decoder register region is served by live reads of the hardware
> > +decoder with guest writes absorbed: the decoder is committed and
> > +locked by the host, so a guest can neither decommit nor reprogram it,
> > +and the kernel keeps no shadow of the decoder state.
>
> This is basically what I said at the beginning - it must either be that the host
> provides locked auto-decoders at boot, or it must provide proper virtualization
> so that the decoders settings are fully virtualized.
>
> Seems it's the former, and that makes sense. Please correct me if i'm
> misunderstanding.
>
> > After a reset the kernel restores and
> > +re-samples the firmware-committed decoder, so the geometry the guest
> > +reads back is unchanged. A VMM that dropped its HDM mapping, for
> > +example across a reset or a D3hot->D0 transition, must rescan the
> > +decoder and rebuild its
> > +stage-2 mapping before it resumes HDM access.
> > +
>
> Yeah i think we need a bit more information about this stage-2 mapping
> rebuild to make sense of this. Maybe I'm just not read-up enough on this
> particular setup - is there another part of the docs you can link to that talk
> about this, or are you able to share some details as to what this rebuild
> process looks like?
>
> ~Gregory