[5.10,CANDIDATE,8/9] xfs: logging the on disk inode LSN can make it go backwards

From: Dave Chinner <dchinner@redhat.com>

From: Dave Chinner <dchinner@redhat.com>

commit 32baa63d82ee3f5ab3bd51bae6bf7d1c15aed8c7 upstream.

When we log an inode, we format the "log inode" core and set an LSN
in that inode core. We do that via xfs_inode_item_format_core(),
which calls:

	xfs_inode_to_log_dinode(ip, dic, ip->i_itemp->ili_item.li_lsn);

to format the log inode. It writes the LSN from the inode item into
the log inode, and if recovery decides the inode item needs to be
replayed, it recovers the log inode LSN field and writes it into the
on disk inode LSN field.

Now this might seem like a reasonable thing to do, but it is wrong
on multiple levels. Firstly, if the item is not yet in the AIL,
item->li_lsn is zero. i.e. the first time the inode it is logged and
formatted, the LSN we write into the log inode will be zero. If we
only log it once, recovery will run and can write this zero LSN into
the inode.

This means that the next time the inode is logged and log recovery
runs, it will *always* replay changes to the inode regardless of
whether the inode is newer on disk than the version in the log and
that violates the entire purpose of recording the LSN in the inode
at writeback time (i.e. to stop it going backwards in time on disk
during recovery).

Secondly, if we commit the CIL to the journal so the inode item
moves to the AIL, and then relog the inode, the LSN that gets
stamped into the log inode will be the LSN of the inode's current
location in the AIL, not it's age on disk. And it's not the LSN that
will be associated with the current change. That means when log
recovery replays this inode item, the LSN that ends up on disk is
the LSN for the previous changes in the log, not the current
changes being replayed. IOWs, after recovery the LSN on disk is not
in sync with the LSN of the modifications that were replayed into
the inode. This, again, violates the recovery ordering semantics
that on-disk writeback LSNs provide.

Hence the inode LSN in the log dinode is -always- invalid.

Thirdly, recovery actually has the LSN of the log transaction it is
replaying right at hand - it uses it to determine if it should
replay the inode by comparing it to the on-disk inode's LSN. But it
doesn't use that LSN to stamp the LSN into the inode which will be
written back when the transaction is fully replayed. It uses the one
in the log dinode, which we know is always going to be incorrect.

Looking back at the change history, the inode logging was broken by
commit 93f958f9c41f ("xfs: cull unnecessary icdinode fields") way
back in 2016 by a stupid idiot who thought he knew how this code
worked. i.e. me. That commit replaced an in memory di_lsn field that
was updated only at inode writeback time from the inode item.li_lsn
value - and hence always contained the same LSN that appeared in the
on-disk inode - with a read of the inode item LSN at inode format
time. CLearly these are not the same thing.

Before 93f958f9c41f, the log recovery behaviour was irrelevant,
because the LSN in the log inode always matched the on-disk LSN at
the time the inode was logged, hence recovery of the transaction
would never make the on-disk LSN in the inode go backwards or get
out of sync.

A symptom of the problem is this, caught from a failure of
generic/482. Before log recovery, the inode has been allocated but
never used:

xfs_db> inode 393388
xfs_db> p
core.magic = 0x494e
core.mode = 0
....
v3.crc = 0x99126961 (correct)
v3.change_count = 0
v3.lsn = 0
v3.flags2 = 0
v3.cowextsize = 0
v3.crtime.sec = Thu Jan  1 10:00:00 1970
v3.crtime.nsec = 0

After log recovery:

xfs_db> p
core.magic = 0x494e
core.mode = 020444
....
v3.crc = 0x23e68f23 (correct)
v3.change_count = 2
v3.lsn = 0
v3.flags2 = 0
v3.cowextsize = 0
v3.crtime.sec = Thu Jul 22 17:03:03 2021
v3.crtime.nsec = 751000000
...

You can see that the LSN of the on-disk inode is 0, even though it
clearly has been written to disk. I point out this inode, because
the generic/482 failure occurred because several adjacent inodes in
this specific inode cluster were not replayed correctly and still
appeared to be zero on disk when all the other metadata (inobt,
finobt, directories, etc) indicated they should be allocated and
written back.

The fix for this is two-fold. The first is that we need to either
revert the LSN changes in 93f958f9c41f or stop logging the inode LSN
altogether. If we do the former, log recovery does not need to
change but we add 8 bytes of memory per inode to store what is
largely a write-only inode field. If we do the latter, log recovery
needs to stamp the on-disk inode in the same manner that inode
writeback does.

I prefer the latter, because we shouldn't really be trying to log
and replay changes to the on disk LSN as the on-disk value is the
canonical source of the on-disk version of the inode. It also
matches the way we recover buffer items - we create a buf_log_item
that carries the current recovery transaction LSN that gets stamped
into the buffer by the write verifier when it gets written back
when the transaction is fully recovered.

However, this might break log recovery on older kernels even more,
so I'm going to simply ignore the logged value in recovery and stamp
the on-disk inode with the LSN of the transaction being recovered
that will trigger writeback on transaction recovery completion. This
will ensure that the on-disk inode LSN always reflects the LSN of
the last change that was written to disk, regardless of whether it
comes from log recovery or runtime writeback.

Fixes: 93f958f9c41f ("xfs: cull unnecessary icdinode fields")
Signed-off-by: Dave Chinner <dchinner@redhat.com>
Reviewed-by: Darrick J. Wong <djwong@kernel.org>
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Signed-off-by: Amir Goldstein <amir73il@gmail.com>
---
 fs/xfs/libxfs/xfs_log_format.h  | 11 +++++++++-
 fs/xfs/xfs_inode_item_recover.c | 39 ++++++++++++++++++++++++---------
 2 files changed, 39 insertions(+), 11 deletions(-)

Message ID	20220726092125.3899077-9-amir73il@gmail.com (mailing list archive)
State	New, archived
Headers	show Return-Path: <fstests-owner@kernel.org> X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by smtp.lore.kernel.org (Postfix) with ESMTP id A23FEC433EF for <linux-fstests@archiver.kernel.org>; Tue, 26 Jul 2022 09:21:43 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S238008AbiGZJVm (ORCPT <rfc822;linux-fstests@archiver.kernel.org>); Tue, 26 Jul 2022 05:21:42 -0400 Received: from lindbergh.monkeyblade.net ([23.128.96.19]:58008 "EHLO lindbergh.monkeyblade.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S231708AbiGZJVm (ORCPT <rfc822;fstests@vger.kernel.org>); Tue, 26 Jul 2022 05:21:42 -0400 Received: from mail-ej1-x62f.google.com (mail-ej1-x62f.google.com [IPv6:2a00:1450:4864:20::62f]) by lindbergh.monkeyblade.net (Postfix) with ESMTPS id BC8F331391; Tue, 26 Jul 2022 02:21:40 -0700 (PDT) Received: by mail-ej1-x62f.google.com with SMTP id fy29so24950128ejc.12; Tue, 26 Jul 2022 02:21:40 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20210112; h=from:to:cc:subject:date:message-id:in-reply-to:references :mime-version:content-transfer-encoding; bh=MqQlMmk2tv+YSwSAAWhsJzSfK9jP24nEPqYd+Fzrkjo=; b=i1FEsv+k8+LD8t19b21uGfJOqlEv98PrafOKtZ+95wFOA/84MIzV+fh67i5NN3J5Bi kYCRKn6kh7bXCyEbrCL9yOEkHT1XwODOnwdEMP9/U+CQKh0H1TSQcfMpmFUVU330L7Tu JgsMLa0rPlMXIlJdJ4gfGj/Zz1YXmqfCsXmWb5PBST4If2cPTJB3zzRBQYod2E7rLP3N IThk+uEG8tVwxpHo33g1pRUZn3Arc6MNJ0J8a+S15WbIquGJrr/b0dcr2pWnYvkqlxFK PWFcJraQHlEEYzyynXTPBOWeALSlcB+fHW+hrRuDo6wVPqonolAzHHo80XBJnbNs8qAH EhQQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20210112; h=x-gm-message-state:from:to:cc:subject:date:message-id:in-reply-to :references:mime-version:content-transfer-encoding; bh=MqQlMmk2tv+YSwSAAWhsJzSfK9jP24nEPqYd+Fzrkjo=; b=uakARgjRS8pPqqT6/vtlR67BqTilRdyUspsilRz2scwhdv771aNY20vodLV6gMj7+K HRWCJpwRWwFcRZFu4y4+H67hk2HG8zm/u1VM0C2HcbkrVCooakR9Z1w9ZPDlhAjRxaet IoSQxAtgyvCiGjBkVgw5qTxT8OTWLZsouHq4k3tYRBgs9sfFysgImXY134LUnI0R5H6a 2DLRmIqtuoPKGrRoezuN4gQSk+lD9iGDFxfZzluj/R+1dRM1seO7QdMnpJrGm7ojFfy9 XlOe4PUWKkcw5UH4MlxpKGiM8er4zFTNHru+YAQBZKfQ2Z2vW/Gsq1KRsO6UwlzCBkwG iASg== X-Gm-Message-State: AJIora/7RvlZRpo5wEI8HBNR2JHrM0XbPpD1TPX4K/4SQCIysqvFkoPI Cw+No912QxW3C813MDpud9D+Uzzyn2OLrg== X-Google-Smtp-Source: AGRyM1t7QReYKXl1qCtSKYPO4xZWz3jvJ2E1+qO036UFW3z5TE8ZIt/gfqxCkRRqRfMlSqS6Tx0SLw== X-Received: by 2002:a17:907:e94:b0:72b:700e:21d9 with SMTP id ho20-20020a1709070e9400b0072b700e21d9mr13657619ejc.665.1658827298896; Tue, 26 Jul 2022 02:21:38 -0700 (PDT) Received: from amir-ThinkPad-T480.kpn (2a02-a45a-4ae9-1-7aa-6650-a0dd-61a2.fixed6.kpn.net. [2a02:a45a:4ae9:1:7aa:6650:a0dd:61a2]) by smtp.gmail.com with ESMTPSA id w17-20020a056402071100b0043aa17dc199sm8161528edx.90.2022.07.26.02.21.37 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 26 Jul 2022 02:21:38 -0700 (PDT) From: Amir Goldstein <amir73il@gmail.com> To: "Darrick J . Wong" <djwong@kernel.org> Cc: Leah Rumancik <leah.rumancik@gmail.com>, Chandan Babu R <chandan.babu@oracle.com>, linux-xfs@vger.kernel.org, fstests@vger.kernel.org, Dave Chinner <dchinner@redhat.com> Subject: [PATCH 5.10 CANDIDATE 8/9] xfs: logging the on disk inode LSN can make it go backwards Date: Tue, 26 Jul 2022 11:21:24 +0200 Message-Id: <20220726092125.3899077-9-amir73il@gmail.com> X-Mailer: git-send-email 2.25.1 In-Reply-To: <20220726092125.3899077-1-amir73il@gmail.com> References: <20220726092125.3899077-1-amir73il@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: 8bit Precedence: bulk List-ID: <fstests.vger.kernel.org> X-Mailing-List: fstests@vger.kernel.org
Series	xfs stable candidate patches for 5.10.y (from v5.13+) \| expand [5.10,CANDIDATE,0/9] xfs stable candidate patches for 5.10.y (from v5.13+) [5.10,CANDIDATE,1/9] xfs: refactor xfs_file_fsync [5.10,CANDIDATE,2/9] xfs: xfs_log_force_lsn isn't passed a LSN [5.10,CANDIDATE,3/9] xfs: prevent UAF in xfs_log_item_in_current_chkpt [5.10,CANDIDATE,4/9] xfs: fix log intent recovery ENOSPC shutdowns when inactivating inodes [5.10,CANDIDATE,5/9] xfs: force the log offline when log intent item recovery fails [5.10,CANDIDATE,6/9] xfs: hold buffer across unpin and potential shutdown processing [5.10,CANDIDATE,7/9] xfs: remove dead stale buf unpin handling code [5.10,CANDIDATE,8/9] xfs: logging the on disk inode LSN can make it go backwards [5.10,CANDIDATE,9/9] xfs: Enforce attr3 buffer recovery order

[5.10,CANDIDATE,8/9] xfs: logging the on disk inode LSN can make it go backwards

Commit Message

Patch