[PATCH v7 04/31] ext4: recheck extent status tree before block allocation

From: Zhang Yi

Date: Fri Oct 09 2026 - 06:38:53 EST


From: Zhang Yi <yi.zhang@xxxxxxxxxx>

After acquiring i_data_sem in write mode, recheck that the mapping
found via the extent status tree or disk query has not changed. A
racing truncate may have trimmed the extent between the earlier lookup
and the write lock acquisition, since writeback does not hold i_rwsem
or the folio locks covering the full extent. This could cause
ext4_map_create_blocks() to allocate blocks beyond the truncated range,
potentially leading to quota leaks in the upcomming iomap buffered
writeback path, because the iomap writeback infrastructure caches
extents beyond the folio range.

Therefore, if we find a valid extent and the sequence number has
changed, retry the entire lookup to obtain the correct trimmed mapping.

Note that a retry only happens when i_es_seq has actually changed, which
requires another thread to hold i_data_sem in write mode, so each retry
implies real lock contention, and the writer yields on the rwsem
slowpath rather than spinning. Moreover, the read-write lock also tries
its best to guarantee fairness, so even though no retry limit is added
here, livelock should not occur in theory. Also, the retry here mirrors
the existing write_ops->iomap_valid() + IOMAP_F_STALE mechanism in
iomap, which also rechecks the sequence number and retries without any
upper bound, and has not shown livelock problems in practice.

In addition, if ext4_map_query_blocks() fails, return the error
immediately. Otherwise the following sequence number comparison would
use a map->m_seq that may not have been updated, and silently
continuing on an error is not advisable.

Suggested-by: Jan Kara <jack@xxxxxxx>
Link: https://lore.kernel.org/linux-ext4/b0781809-4759-4e12-be17-71555b764f48@xxxxxxxxx/
Signed-off-by: Zhang Yi <yi.zhang@xxxxxxxxxx>
Reviewed-by: Ojaswin Mujoo <ojaswin@xxxxxxxxxxxxx>
---
fs/ext4/inode.c | 18 ++++++++++++++++++
1 file changed, 18 insertions(+)

diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c
index c0f8ff495dd82..739f8d3914653 100644
--- a/fs/ext4/inode.c
+++ b/fs/ext4/inode.c
@@ -734,6 +734,7 @@ int ext4_map_blocks(handle_t *handle, struct inode *inode,
else
ext4_check_map_extents_env(inode);

+create_retry:
/* Lookup extent status tree firstly */
if (ext4_es_lookup_extent(inode, map->m_lblk, NULL, &es, &map->m_seq)) {
if (ext4_es_is_written(&es) || ext4_es_is_unwritten(&es)) {
@@ -784,6 +785,8 @@ int ext4_map_blocks(handle_t *handle, struct inode *inode,
down_read(&EXT4_I(inode)->i_data_sem);
retval = ext4_map_query_blocks(handle, inode, map, flags);
up_read((&EXT4_I(inode)->i_data_sem));
+ if (retval < 0)
+ return retval;

found:
if (retval > 0 && map->m_flags & EXT4_MAP_MAPPED) {
@@ -820,6 +823,21 @@ int ext4_map_blocks(handle_t *handle, struct inode *inode,
* with create == 1 flag.
*/
down_write(&EXT4_I(inode)->i_data_sem);
+
+ /*
+ * Check the validity of the mapping found via the extent status
+ * tree or the disk query. A racing truncate may have changed the
+ * extent, since writeback may not hold i_rwsem or the folio locks
+ * covering the full extent(e.g., the iomap writeback path may
+ * allocate an extent that extends beyond the length of the locked
+ * folio in advance).
+ */
+ if (map->m_seq != READ_ONCE(EXT4_I(inode)->i_es_seq)) {
+ up_write(&EXT4_I(inode)->i_data_sem);
+ map->m_flags = 0;
+ map->m_len = orig_mlen;
+ goto create_retry;
+ }
retval = ext4_map_create_blocks(handle, inode, map, flags);
up_write((&EXT4_I(inode)->i_data_sem));

--
2.52.0