[PATCH v2 1/4] blk-iocost: charge flushes as pageless random writes

From: Tao Cui

Date: Wed Sep 16 2026 - 05:05:21 EST


From: Tao Cui <cuitao@xxxxxxxxxx>

Standalone flushes issued by blkdev_issue_flush() are represented as
dataless REQ_OP_WRITE | REQ_PREFLUSH bios, which
calc_vtime_cost_builtin() prices at zero. The flush component of
flush-heavy workloads such as database commits, journal flushes, and
metadata sync is thus neither charged nor throttled: a cgroup at 1% weight
can issue ~510k flushes per 12s, monopolizing the device while iocost
reports zero usage.

The same is true for the flush component of data-bearing
REQ_OP_WRITE | REQ_PREFLUSH bios, e.g. journal commit writes: they are
charged for their data only, and the cache flush the flush machine runs
ahead of it is free.

Charge the flush component of any REQ_PREFLUSH bio on top of its data
cost, priced as a pageless random write (LCOEF_WRANDIO), which provides
an approximation of the device time consumed by a flush. For profiles
where WRANDIO clamps to zero (ssd_dfl / ssd_fast), use a one-page floor
(LCOEF_WPAGE). A dataless flush bio falls out of the switch with zero
data cost and picks up the same surcharge, so standalone and pre-flush
forms are priced the same way. After this patch, the same 1%-weight
cgroup is limited to 24 flushes per 12s; on ext4, write+fsync workloads
are correctly accounted through the journal layer (~2.2us per flush on
the ssd_fast profile).

A standalone flush must also not update iocg->cursor: its bi_sector
(usually 0) is not a data position, so setting the cursor from it
would misclassify the following READ/WRITE bios, and a zero cursor
defeats the !iocg->cursor sentinel in calc_vtime_cost_builtin().
Skip the cursor update for dataless bios.

Fixes: 7caa47151ab2 ("blkcg: implement blk-iocost")
Signed-off-by: Tao Cui <cuitao@xxxxxxxxxx>
---
Changes in v2:
- Skip the iocg->cursor update for dataless flush bios, which would
otherwise corrupt the seq/rand classification of the following IOs
(reported in review of v1).
- Charge the flush component of data-bearing REQ_PREFLUSH bios too;
v1 only priced standalone flushes (reported in review of v1).
---
block/blk-iocost.c | 20 +++++++++++++++++---
1 file changed, 17 insertions(+), 3 deletions(-)

diff --git a/block/blk-iocost.c b/block/blk-iocost.c
index 2745bffcd5eef..082f26d6e27b6 100644
--- a/block/blk-iocost.c
+++ b/block/blk-iocost.c
@@ -2532,8 +2532,20 @@ static void calc_vtime_cost_builtin(struct bio *bio, struct ioc_gq *iocg,
u64 pages = max_t(u64, bio_sectors(bio) >> IOC_SECT_TO_PAGE_SHIFT, 1);
u64 seek_pages = 0;
u64 cost = 0;
+ u64 flush_cost = 0;

- /* Can't calculate cost for empty bio */
+ /*
+ * A WRITE|REQ_PREFLUSH bio carries a flush component: the flush
+ * machine runs a cache flush for it, either standalone (dataless)
+ * or ahead of the data. Charge the flush on top of the data cost,
+ * priced as a pageless random write with a one-page floor so fast
+ * profiles still charge something. Flush bios are never merged.
+ */
+ if (!is_merge && (bio->bi_opf & REQ_PREFLUSH))
+ flush_cost = max(ioc->params.lcoefs[LCOEF_WRANDIO],
+ ioc->params.lcoefs[LCOEF_WPAGE]);
+
+ /* Can't calculate data cost for empty bio */
if (!bio->bi_iter.bi_size)
goto out;

@@ -2566,7 +2578,7 @@ static void calc_vtime_cost_builtin(struct bio *bio, struct ioc_gq *iocg,
}
cost += pages * coef_page;
out:
- *costp = cost;
+ *costp = cost + flush_cost;
}

static u64 calc_vtime_cost(struct bio *bio, struct ioc_gq *iocg, bool is_merge)
@@ -2708,7 +2720,9 @@ static void ioc_rqos_throttle(struct rq_qos *rqos, struct bio *bio)
if (!iocg_activate(iocg, &now))
return;

- iocg->cursor = bio_end_sector(bio);
+ /* dataless bios have no meaningful position for seq/rand detection */
+ if (bio->bi_iter.bi_size)
+ iocg->cursor = bio_end_sector(bio);
vtime = atomic64_read(&iocg->vtime);
cost = adjust_inuse_and_calc_cost(iocg, vtime, abs_cost, &now);

--
2.43.0