Re: [PATCH v3 06/15] Introduce structured tag value definition

From: Herve Codina

Date: Fri Sep 18 2026 - 04:27:18 EST


Hi David,

On Fri, 18 Sep 2026 14:41:11 +1000
David Gibson <david@xxxxxxxxxxxxxxxxxxxxx> wrote:

> On Thu, Sep 17, 2026 at 09:04:50AM +0200, Herve Codina wrote:
> > Hi David,
> >
> > On Wed, 16 Sep 2026 15:21:15 +1000
> > David Gibson <david@xxxxxxxxxxxxxxxxxxxxx> wrote:
> >
> > > On Mon, Sep 14, 2026 at 12:19:37PM +0200, Herve Codina wrote:
> > > > Hi David,
> > > >
> > > > On Sat, 12 Sep 2026 12:34:24 +1000
> > > > David Gibson <david@xxxxxxxxxxxxxxxxxxxxx> wrote:
> > > >
> > > > ...
> > > >
> > > > > > Do you mean that we should avoid the DATA_LEN_ENCODING and always have the
> > > > > > 32-bit value right after the tag to give the size for all "skippable" tags?
> > > > >
> > > > > Yes.
> > > > >
> > > >
> > > > I did a test using a dts file available in kernel sources. I used (arbitrary
> > > > choice) juno.dts [0].
> > > >
> > > > Without any new tags, the size of the compiled dtb is 27067 bytes.
> > > >
> > > > With new metadata tags identifying phandles in properties (FDT_PROPDATA_PHANDLE),
> > > > the size of the dtb becomes 29027 bytes and so 29027 - 27067 = 1960 bytes for
> > > > those FDT_PROPDATA_PHANDLE tags (+7.2%).
> > > >
> > > > The tags used are composed of:
> > > > 32-bit: FDT_PROPDATA_PHANDLE value encoding 1 x 32-bit for data
> > > > 32-bit: offset in the property where a phandle is present.
> > > >
> > > > Removing the '1 x 32-bit' information from the tag value and adding a 32-bit
> > > > 'length' in all cases will lead 3 x 32-bit values for a FDT_PROPDATA_PHANDLE
> > > > tag (tag + length + offset) instead of the 2 x 32-bit (tag + offset).
> > > >
> > > > Back to juno.dts instead of 1960 bytes, the FDT_PROPDATA_PHANDLE will need
> > > > 1960 * 3 / 2 = 2640 bytes (+9.7%). This leads to around +2.5% of the whole
> > > > dtb just to have the 32-bit for length. This +2.5% can be easily avoided.
> > > >
> > > > Also, I will not be surprised to see more tags in the future adding some more
> > > > metadata information and so increasing dtb sizes.
> > > >
> > > > Quite often you have mentioned memory constraints system where libfdt should
> > > > be as small as possible. On those system, the dtb itself is embedded in the
> > > > binary close to libfdt. The size of dtb should be taken into account.
> > >
> > > Yeah, those proportions are high enough that I think it's worth it.
> > >
> > > > If the SAFE_SKIP bit is removed, I even plan to use this now free bit in the
> > > > length encoding part:
> > > > 0b000: No data
> > > > 0b001: 1 fdt32
> > > > 0b010: 2 fdt32
> > > > ...
> > > > 0b110: 6 fdt32
> > > > 0b111: On additional fdt32 to encode the length of data.
> > > >
> > > > IHMO, length encoding bits in tag value definition should be kept and used
> > > > for all tags where the length is fixed and can be encoded using
> > > > these bits.
> > >
> > > Well, I'm convinced we want some sort of compact encoding of the
> > > length, but I think we can do better than the current proposal. It
> > > seems implausible to me that we'll need 2^29 different metadata tags,
> > > so I think we can spend some more of the tag bits on the length. How about:
> > >
> > > 0x80000000 structured tag bit
> > > 0x7fff0000 tag type
> > > 0x0000ffff tag length
> > >
> > > So we have up to 2^15 (32k) different structured tags each with a
> > > length of [0..65534] bytes (length==65535 reserved for those that need
> > > a full 32-bit length word).
> > >
> > > I believe that will avoid the extra length word for everything you
> > > have currently drafted.
> > >
> >
> > Yes, this will avoid the extra length field. The drawback is the that the
> > tag value is no more a well fixed value. Each time we have to check the tag
> > value we have to filter out the tag length.
> >
> > For instance:
> > - FDT_PROPDATA_PHANDLE
> > fixed data size 4 bytes for offset
> > tag value: 0x80010004
> >
> > - FDT_PROPDATA_PHANDLE_REF
> > data: 4 bytes for offset + N bytes for a string
> > tag value 0x8002ssss with ssss for the size
> >
> > This will lead to code like this:
> > tag = fdt_next_tag();
> > if (tag == FDT_PROPDATA_PHANDLE)
> > /* Do something */
> >
> > if (TAG_GET_ID(tag) == FDT_PROPDATA_PHANDLE_REF)
> > /* Do something */
>
> True. But.. a similar problem kind of exists with the original
> proposed encoding too: we *expect* a tag with fixed 4-byte contents to
> use the "1 cell" flags, but we need to consider the case of encoding
> it as VARLEN with a length field of 4. We could choose to make that
> forbidden, but we'd still need to consider who's responsible for
> enforcing that.
>
> Similarly, if a variable length metadata tag happens to have length 4
> or 8 in a particular place, is it valid to encode it with the 1-cell
> or 2-cell flag?

My position was: if we expect a tag with "1-cell" flag, using varlen
field encoding is considered as an other tag and so either an error
or a skippable unknown tag.

The same apply for varlen defined tag. Even if the data is, let's
say 8 bytes, the tag cannot be moved to a "2-cells" tag.

The kind of data length encoding (1-cell, 2-cells, varlength) is done
when the tag is defined and cannot be changed.

Who is responsible for enforcing that ?
I would say the documentation of the tag should clearly set the data
encoding used for the tag and the documentation of the "skippable"
format should say that data encoding is fixed for a given tag. It is
set when the tag is defined and any changes at runtime should be
considered as a different tag.

Of course we can introduce dynamic length encoding to set the data length
encoding in the tag according to the exact data length found at runtime.

>
> As a variant on my proposal, I'd also be fine with dividing the
> structed tags into several classes with bits indicating which is
> which. Either:
>
> * "short" vs "long": "short" always has the length within the tag
> word (and so cannot exceed 64k, or however many bits we set aside)
> whereas long always has a length word
> * "fixed" vs "variable", fixed length tag types always have the same
> length, so the length can be considered part of the tag. Variable
> would have a length word.

Well, only strings, or more generally arrays, need a varlen. For those
item, I would use the varlen word and so "long" in your definition.

For all others where the sizeof(data) is well known when the tag is
defined, I would use "fixed" and "long" only if sizeof(data) cannot
be encoded by "fixed" (lengh > limit of dedicated bits).

Without any additional bits for any category, all of these fit with the
following length encoding:
000...00: No data
000...01: 1 x 32-bit
111...10: N x 32-bit
111...11: varlen word

>
> Not sure if that makes things any easier, but they might, and I'd be
> fine with either option (or some combination). In any of these cases
> we do need to spell out what the requirements are: for dtb writers,
> for dtb readers and for whoever defines a new tag.
>
> > Further more, in the code you will both a mix of both construction:
> > while (tag == FDT_NOP || tag == FDT_BEGIN_NODE ||
> > TAG_GET_ID(tag) == FDT_BEGIN_NODE_REF);
> >
> > with #define TAG_GET_ID(tag) ((tag) & 0xffff0000)
> >
> > Also when we write dtbs either in libfdt or dtc, tags value
> > have to be built with the length when needed.
> > #define TAG_VALUE(tag_id, length) (((tag_id) & 0xffff0000) || \
> > ((length) & 0x0000ffff))
> >
> > I am totally fine with that but we need to have it in mind.
>
> Right, that's not a deal breaker for me - especially since I think
> we'll need some similar stuff even with the original proposal.
>
> > Of course, I can encode FDT_PROPDATA_PHANDLE_REF with 0x8002ffff + length
> > field but we lose all the benefits
> >
> > Ready to see TAG_GET_ID(tag) and TAG_VALUE(tag_id, length) when needed in
> > the code?
>
> I think so, though that could change depending on what it ends up
> looking like in practice.
>

Best regards,
Hervé