Re: [PATCH v2] pmdomain: imx8m-blk-ctrl: Serialize power on/off across sibling domains
From: Ming Qian(OSS)
Date: Thu Sep 17 2026 - 22:09:26 EST
Hi Frank,
On Thu, Sep 17, 2026 at 11:42:22AM -0500, Frank Li wrote:
> On Thu, Sep 17, 2026 at 04:07:54PM +0900, Ming Qian wrote:
> > On i.MX8MP the VPU blk-ctrl exposes G1, G2 and VC8000E as three separate
> > genpds, each serialized only by its own genpd lock, so their power_on and
> > power_off callbacks can run concurrently on SMP.
> >
> > The sequences are not independent: they share the VPUMIX bus domain, the
> > VPU_NOC and the ADB400 handshake. On power up the GPC cannot ack-verify
> > the ADB400 handshake - the ack only completes once blk-ctrl sets the bus
> > clk-en bit - so it just waits a fixed delay instead of polling hskack. A
> > sibling transition landing inside another domain's reset window disturbs
> > that shared clock and handshake timing, the victim's reset does not take
> > effect, and its block registers read back all zeros: the decoder times
> > out or the encoder fails its format check.
> >
> > Serialize the blk-ctrl reset sequence with a per-blk-ctrl mutex; the
> > driver deliberately avoids a genpd hierarchy, so the genpd core gives no
> > cross-sibling serialization.
> >
> > Fixes: a1a5f15f7f6c ("soc: imx: imx8m-blk-ctrl: add i.MX8MP VPU blk ctrl")
> > Signed-off-by: Ming Qian <ming.qian@xxxxxxxxxxx>
> > ---
> > Problem:
> > On i.MX8MP, running the VC8000E encoder and the G1/G2 decoders
> > concurrently rarely and non-deterministically leaves a VPU stuck in
> > reset: its block registers read back all zeros. A decoder then times out
> > or the encoder fails its format check.
> >
> > Root cause:
> > The VPU blk-ctrl exposes G1, G2 and VC8000E as three separate genpds,
> > serialized only by the per-genpd lock, so on SMP their power_on/power_off
> > callbacks can run concurrently. The sequences share the VPUMIX bus
> > domain, the VPU_NOC and the ADB400 handshake. On power up the GPC does
> > not ack-verify the ADB400 handshake - the ack only completes once
> > blk-ctrl sets the bus clk-en bit - so it just waits a fixed delay. A
> > sibling transition landing inside another domain's reset window disturbs
> > that shared clock and handshake timing, the victim's reset fails to take
> > effect, and its block is left in reset with registers reading zero.
> >
> > Fix:
> > Serialize the blk-ctrl reset sequence with a per-blk-ctrl mutex, so a
> > sibling domain cannot run its sequence while another is inside its reset
> > window. The driver deliberately avoids a genpd hierarchy, so the genpd
> > core provides no cross-sibling serialization.
> >
> > Test:
> > i.MX8MP, Android 6.18 kernel, concurrent H.264 decode and encode. Without
> > this patch the failure reproduces within one to two hours. With it the
> > same test ran overnight, over 14 hours, without a single occurrence.
> > ---
> > Changes in v2:
> > - Replace guard(mutex) with explicit mutex_lock()/mutex_unlock():
> > power_on() already unwinds errors with goto, and cleanup.h asks not to
> > mix goto and scope-based cleanup in one function (sashiko-bot).
> > - Shorten the commit message to the essentials and move the detailed
> > hardware analysis into this cover letter (Frank Li).
> > - Link to v1: https://patch.msgid.link/20260915-imx8mp-blk-ctrl-v1-1-b3b4e6e7e676@xxxxxxxxxxx
> >
> > To: Ulf Hansson <ulfh@xxxxxxxxxx>
> > To: Frank Li <Frank.Li@xxxxxxx>
> > To: Sascha Hauer <s.hauer@xxxxxxxxxxxxxx>
> > To: Pengutronix Kernel Team <kernel@xxxxxxxxxxxxxx>
> > To: Fabio Estevam <festevam@xxxxxxxxx>
> > To: Shawn Guo <shawnguo@xxxxxxxxxx>
> > To: Peng Fan <peng.fan@xxxxxxx>
> > Cc: linux-pm@xxxxxxxxxxxxxxx
> > Cc: imx@xxxxxxxxxxxxxxx
> > Cc: linux-arm-kernel@xxxxxxxxxxxxxxxxxxx
> > Cc: linux-kernel@xxxxxxxxxxxxxxx
> > ---
> > drivers/pmdomain/imx/imx8m-blk-ctrl.c | 23 ++++++++++++++++++++++-
> > 1 file changed, 22 insertions(+), 1 deletion(-)
> >
> > diff --git a/drivers/pmdomain/imx/imx8m-blk-ctrl.c b/drivers/pmdomain/imx/imx8m-blk-ctrl.c
> > index 479789009c7f..270f43229fe7 100644
> > --- a/drivers/pmdomain/imx/imx8m-blk-ctrl.c
> > +++ b/drivers/pmdomain/imx/imx8m-blk-ctrl.c
> > @@ -15,6 +15,7 @@
> > #include <linux/pm_runtime.h>
> > #include <linux/regmap.h>
> > #include <linux/clk.h>
> > +#include <linux/mutex.h>
> >
> > #include <dt-bindings/power/imx8mm-power.h>
> > #include <dt-bindings/power/imx8mn-power.h>
> > @@ -34,6 +35,12 @@ struct imx8m_blk_ctrl {
> > struct regmap *regmap;
> > struct imx8m_blk_ctrl_domain *domains;
> > struct genpd_onecell_data onecell_data;
> > + /*
> > + * Serializes the blk-ctrl reset/clock sequence across sibling domains;
> > + * their transitions interact through the shared VPUMIX bus domain,
> > + * VPU_NOC and the not-ack-verified ADB400 handshake (ERR050531).
> > + */
> > + struct mutex power_lock;
> > };
> >
> > struct imx8m_blk_ctrl_domain_data {
> > @@ -98,12 +105,14 @@ static int imx8m_blk_ctrl_power_on(struct generic_pm_domain *genpd)
> > struct imx8m_blk_ctrl *bc = domain->bc;
> > int ret;
> >
> > + mutex_lock(&bc->power_lock);
> > +
> > /* make sure bus domain is awake */
> > ret = pm_runtime_get_sync(bc->bus_power_dev);
> > if (ret < 0) {
> > pm_runtime_put_noidle(bc->bus_power_dev);
> > dev_err(bc->dev, "failed to power up bus domain\n");
> > - return ret;
> > + goto unlock;
>
> can you use auto cleanup guard()
>
> Frank
>
v1 did exactly that, and sashiko-bot flagged it, because
imx8m_blk_ctrl_power_on() unwinds its errors with goto:
https://patch.msgid.link/20260915102028.2CFFD1F000FF@xxxxxxxxxxxxxxx
According to the cleanup subsystem guidelines (include/linux/cleanup.h),
using goto and scope-based cleanup helpers shouldn't be mixed in the same
function:
* Lastly, given that the benefit of cleanup helpers is removal of
* "goto", and that the "goto" statement can jump between scopes, the
* expectation is that usage of "goto" and cleanup helpers is never
* mixed in the same function. I.e. for a given routine, convert all
* resources that need a "goto" cleanup to scope-based cleanup, or
* convert none of them.
So v2 changed imx8m_blk_ctrl_power_on() back from guard() to
mutex_lock(), and imx8m_blk_ctrl_power_off() follows the same style for
consistency.
Regards,
Ming
> > }
> >
> > /* put devices into reset */
> > @@ -148,12 +157,16 @@ static int imx8m_blk_ctrl_power_on(struct generic_pm_domain *genpd)
> > /* disable upstream clocks */
> > clk_bulk_disable_unprepare(data->num_clks, domain->clks);
> >
> > + mutex_unlock(&bc->power_lock);
> > +
> > return 0;
> >
> > clk_disable:
> > clk_bulk_disable_unprepare(data->num_clks, domain->clks);
> > bus_put:
> > pm_runtime_put(bc->bus_power_dev);
> > +unlock:
> > + mutex_unlock(&bc->power_lock);
> >
> > return ret;
> > }
> > @@ -164,6 +177,8 @@ static int imx8m_blk_ctrl_power_off(struct generic_pm_domain *genpd)
> > const struct imx8m_blk_ctrl_domain_data *data = domain->data;
> > struct imx8m_blk_ctrl *bc = domain->bc;
> >
> > + mutex_lock(&bc->power_lock);
> > +
> > /* put devices into reset and disable clocks */
> > if (data->mipi_phy_rst_mask)
> > regmap_clear_bits(bc->regmap, BLK_MIPI_RESET_DIV, data->mipi_phy_rst_mask);
> > @@ -177,6 +192,8 @@ static int imx8m_blk_ctrl_power_off(struct generic_pm_domain *genpd)
> > /* allow bus domain to suspend */
> > pm_runtime_put(bc->bus_power_dev);
> >
> > + mutex_unlock(&bc->power_lock);
> > +
> > return 0;
> > }
> >
> > @@ -202,6 +219,10 @@ static int imx8m_blk_ctrl_probe(struct platform_device *pdev)
> >
> > bc->dev = dev;
> >
> > + ret = devm_mutex_init(dev, &bc->power_lock);
> > + if (ret)
> > + return ret;
> > +
> > bc_data = of_device_get_match_data(dev);
> >
> > base = devm_platform_ioremap_resource(pdev, 0);
> >
> > ---
> > base-commit: 27953c044974baf7e24dee3e9342fe0103dea80c
> > change-id: 20260911-imx8mp-blk-ctrl-c46f26783073
> >
> >