Re: [PATCH 3/5] perf symbol: Grow the Rust demangle buffer geometrically

From: Ian Rogers

Date: Fri Oct 09 2026 - 16:40:00 EST


On Thu, Oct 8, 2026 at 7:22 PM Michal Pluta <michalpl2003@xxxxxxxxx> wrote:
>
> On Mon, Sep 28, 2026 at 2:43 PM Ian Rogers <irogers@xxxxxxxxxx> wrote:
> >
> > It seems changing the scope of this variable is unnecessary.
> >
> > Thanks for digging into this problem and exploring a fix! Previously,
> > we guessed the demangled length was twice the mangled length, then
> > added 32 bytes for each retry. These were numbers I pulled out of thin
> > air, so I'm glad you've found them to be wrong :-). Could we estimate
> > the initial demangled size better? Could you get data from Bevy and
> > Typst? I'm a little concerned that a demangled symbol of say just over
> > 2KB might require 4KB with this change, instead of 2KB + 32bytes.
>
> Thank you for the reviews. I'll move the buf_len declaration back to
> where it was before.
>
> For context, the other demanglers in perf don't size their buffer by
> retrying. The OCaml one allocates len + 1, as the output is never longer
> than the input. The Java one allocates strlen * 3 + 1 and truncates. The
> C++ demangler (libiberty) starts with an empty buffer and doubles it,
> but it writes through a callback so growing doesn't restart the
> formatting. The Rust one writes into a caller provided buffer and has to
> start over each time it is too small.
>
> I collected some data from release builds of Bevy and Typst and a few
> other large(r) Rust projects. For each binary I took the defined symbols
> and counted the number of attempts that the buffer loop in
> dso__demangle_sym() makes for them. I ran both the current +32 loop and
> the doubling one from this series in a small test program. No symbol
> went over the 1 MiB limit.
>
> Builds (mostly v0 symbols, plus up to 148 C++ ones per binary from
> linked libraries):
>
> bevy 0f38358f573a cargo +1.98.1 build --example 3d_scene
> --release --config profile.release.strip=false
> datafusion acf5c89e6454 cargo +1.98.1 build --locked -p datafusion-cli
> --profile release-nonlto
> materialize 54a6aca7b5f1 bin/environmentd +1.98.1 --build-only --optimized
> (clusterd and environmentd)
> polars 750bbaa9054c cargo +nightly-2026-09-01 build --locked
> -p polars-dylib --profile fast-release
> --config profile.fast-release.debug=0
> --config profile.fast-release.strip=false
> risingwave be77a6149e97 cargo +nightly-2026-06-21 build --locked
> -p risingwave_cmd_all --bin risingwave --release
> --config profile.release.debug=0
> --config profile.release.strip=false
> typst 9dfd3a08500b cargo +1.98.1 build --locked -p typst-cli --release
> (v0.15.1) --config profile.release.package.typst-cli.strip=false
>
> Firstly some data on how often initial guesses are too small and how
> many retries they cause. "retried symbols" is the share of symbols whose
> output did not fit in the initial buffer. The other columns count the
> extra attempts after the first one ("total"), and the most any single
> symbol needed ("max"):
> symbols retried +32 loop doubling
> symbols total max total max
> materialize clusterd 295,365 5.66% 3,500,913 17,844 27,960 8
> risingwave 1,051,376 0.87% 297,151 1,561 10,180 5
> polars libpolars_dylib 276,797 5.61% 247,424 236 16,149 3
> bevy 3d_scene 218,042 1.09% 205,030 1,135 3,493 5
> materialize environmentd 377,085 0.51% 117,485 1,539 2,302 5
> datafusion-cli 175,020 1.10% 24,304 125 1,957 3
> typst 36,504 0.12% 907 66 50 2
>
> Only 0.1% to 5.7% don't fit in the initial buffer, but with +32 this is
> very costly. In clusterd it adds up to 3.5 million extra attempts with
> one symbol needing 17,844 retries. Doubling does the same job in 27,960
> total retries and no symbol needs more than 8 retries.
>
> You asked whether a better initial guess would help, so here I keep the
> +32 loop and change only the initial buffer. Its size is the mangled
> length multiplied by a factor N and rounded up to a power of two. The
> code uses 2x today, so 1x is a smaller guess than today and 4x and 8x
> are larger ones. Each value is the total extra attempts for that binary:
>
> N 1x 2x 4x 8x
> materialize clusterd 4,507,480 3,500,913 2,460,193 1,528,046
> risingwave 1,306,091 297,151 71,074 25,133
> polars libpolars_dylib 1,254,174 247,424 14,214 305
> bevy 3d_scene 357,175 205,030 94,936 29,978
> materialize environmentd 271,354 117,485 59,697 22,265
> datafusion-cli 179,241 24,304 1,003 29
> typst 7,066 907 15 0
>
> A bigger guess reduces the retries for every binary, but the symbols
> that still miss are the ones with many backrefs. In every binary, all
> the symbols that missed the initial buffer have backrefs, with a median
> of 53-138 expansions compared to 1-4 for the symbols that fit.
>
> Running the same experiment with the doubling loop from this series:
>
> 1x 2x 4x 8x
> materialize clusterd 70,399 27,960 11,231 4,073
> risingwave 106,878 10,180 1,016 128
> polars libpolars_dylib 84,502 16,149 634 9
> bevy 3d_scene 15,742 3,493 1,115 192
> materialize environmentd 22,668 2,302 365 73
> datafusion-cli 15,366 1,957 38 1
> typst 2,310 50 5 0
>
> On the 2KB example you mentioned, a symbol that barely misses a power of
> two could get a buffer up to twice what it needs, but the next patch
> shrinks it to the exact string length before returning, so the extra
> memory is only held during the call. I'll move the realloc patch ahead
> of this one.
>
> One other way I tried to avoid retries is a minimum size for the initial
> buffer. Here the initial buffer is the larger of this minimum and the
> current 2x guess, and each cell is the total extra attempts.
>
> First with the +32 loop:
> current 1KiB 4KiB 16KiB 64KiB
> materialize clusterd 3,500,913 3,500,867 3,181,650 1,425,379 216,505
> risingwave 297,151 296,929 151,909 14,567 0
> polars libpolars_dylib 247,424 246,715 15,413 0 0
> bevy 3d_scene 205,030 204,818 132,911 17,038 0
> materialize environmentd 117,485 117,455 79,073 37,531 0
> datafusion-cli 24,304 24,277 866 0 0
> typst 907 898 2 0 0
>
> And with doubling:
> current 1KiB 4KiB 16KiB 64KiB
> materialize clusterd 27,960 27,945 18,824 3,260 183
> risingwave 10,180 10,117 1,887 63 0
> polars libpolars_dylib 16,149 16,043 686 0 0
> bevy 3d_scene 3,493 3,472 1,574 64 0
> materialize environmentd 2,302 2,282 419 88 0
> datafusion-cli 1,957 1,939 27 0 0
> typst 50 47 1 0 0
>
> As we can see, a minimum does help but it depends a lot on the size
> chosen and would mean we incur a large allocation per-symbol albeit only
> during the call.
>
> Sorry for the long reply, just wanted to get all the data and rationale
> across.
>
> Fun fact: the largest symbol I encountered was in materialize's
> clusterd. It is 1,091 bytes mangled and demangles to 575,074 bytes.

Thanks, Michal and thanks for the long reply and data!
I wasn't sure if you saw any patterns when comparing the unmangled and
mangled sizes? What is the typical rust symbol size? I'm concerned
that reallocating to return memory to the heap might just be treated
as a no-op, so guessing the size correctly would be great. It would
also be good if rather than just failing to parse the parser could
compute the output size, so we retry a maximum of once (snprintf
behaves like this).

Thanks,
Ian

> Thanks,
> Michal