[09/45] xfs: Fix CIL throttle hang when CIL space used going backwards

From: Dave Chinner <dchinner@redhat.com>

From: Dave Chinner <dchinner@redhat.com>

A hang with tasks stuck on the CIL hard throttle was reported and
largely diagnosed by Donald Buczek, who discovered that it was a
result of the CIL context space usage decrementing in committed
transactions once the hard throttle limit had been hit and processes
were already blocked.  This resulted in the CIL push not waking up
those waiters because the CIL context was no longer over the hard
throttle limit.

The surprising aspect of this was the CIL space usage going
backwards regularly enough to trigger this situation. Assumptions
had been made in design that the relogging process would only
increase the size of the objects in the CIL, and so that space would
only increase.

This change and commit message fixes the issue and documents the
result of an audit of the triggers that can cause the CIL space to
go backwards, how large the backwards steps tend to be, the
frequency in which they occur, and what the impact on the CIL
accounting code is.

Even though the CIL ctx->space_used can go backwards, it will only
do so if the log item is already logged to the CIL and contains a
space reservation for it's entire logged state. This is tracked by
the shadow buffer state on the log item. If the item is not
previously logged in the CIL it has no shadow buffer nor log vector,
and hence the entire size of the logged item copied to the log
vector is accounted to the CIL space usage. i.e.  it will always go
up in this case.

If the item has a log vector (i.e. already in the CIL) and the size
decreases, then the existing log vector will be overwritten and the
space usage will go down. This is the only condition where the space
usage reduces, and it can only occur when an item is already tracked
in the CIL. Hence we are safe from CIL space usage underruns as a
result of log items decreasing in size when they are relogged.

Typically this reduction in CIL usage occurs from metadata blocks
being free, such as when a btree block merge occurs or a directory
enter/xattr entry is removed and the da-tree is reduced in size.
This generally results in a reduction in size of around a single
block in the CIL, but also tends to increase the number of log
vectors because the parent and sibling nodes in the tree needs to be
updated when a btree block is removed. If a multi-level merge
occurs, then we see reduction in size of 2+ blocks, but again the
log vector count goes up.

The other vector is inode fork size changes, which only log the
current size of the fork and ignore the previously logged size when
the fork is relogged. Hence if we are removing items from the inode
fork (dir/xattr removal in shortform, extent record removal in
extent form, etc) the relogged size of the inode for can decrease.

No other log items can decrease in size either because they are a
fixed size (e.g. dquots) or they cannot be relogged (e.g. relogging
an intent actually creates a new intent log item and doesn't relog
the old item at all.) Hence the only two vectors for CIL context
size reduction are relogging inode forks and marking buffers active
in the CIL as stale.

Long story short: the majority of the code does the right thing and
handles the reduction in log item size correctly, and only the CIL
hard throttle implementation is problematic and needs fixing. This
patch makes that fix, as well as adds comments in the log item code
that result in items shrinking in size when they are relogged as a
clear reminder that this can and does happen frequently.

The throttle fix is based upon the change Donald proposed, though it
goes further to ensure that once the throttle is activated, it
captures all tasks until the CIL push issues a wakeup, regardless of
whether the CIL space used has gone back under the throttle
threshold.

This ensures that we prevent tasks reducing the CIL slightly under
the throttle threshold and then making more changes that push it
well over the throttle limit. This is acheived by checking if the
throttle wait queue is already active as a condition of throttling.
Hence once we start throttling, we continue to apply the throttle
until the CIL context push wakes everything on the wait queue.

We can use waitqueue_active() for the waitqueue manipulations and
checks as they are all done under the ctx->xc_push_lock. Hence the
waitqueue has external serialisation and we can safely peek inside
the wait queue without holding the internal waitqueue locks.

Many thanks to Donald for his diagnostic and analysis work to
isolate the cause of this hang.

Reported-and-tested-by: Donald Buczek <buczek@molgen.mpg.de>
Signed-off-by: Dave Chinner <dchinner@redhat.com>
Reviewed-by: Brian Foster <bfoster@redhat.com>
Reviewed-by: Chandan Babu R <chandanrlinux@gmail.com>
Reviewed-by: Darrick J. Wong <djwong@kernel.org>
---
 fs/xfs/xfs_buf_item.c   | 37 ++++++++++++++++++-------------------
 fs/xfs/xfs_inode_item.c | 14 ++++++++++++++
 fs/xfs/xfs_log_cil.c    | 22 +++++++++++++++++-----
 3 files changed, 49 insertions(+), 24 deletions(-)

Message ID	20210305051143.182133-10-david@fromorbit.com (mailing list archive)
State	Superseded
Headers	show Return-Path: <linux-xfs-owner@kernel.org> X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org X-Spam-Level: X-Spam-Status: No, score=-16.8 required=3.0 tests=BAYES_00, HEADER_FROM_DIFFERENT_DOMAINS,INCLUDES_CR_TRAILER,INCLUDES_PATCH, MAILING_LIST_MULTI,SPF_HELO_NONE,SPF_PASS,URIBL_BLOCKED,USER_AGENT_GIT autolearn=ham autolearn_force=no version=3.4.0 Received: from mail.kernel.org (mail.kernel.org [198.145.29.99]) by smtp.lore.kernel.org (Postfix) with ESMTP id 7B792C433E0 for <linux-xfs@archiver.kernel.org>; Fri, 5 Mar 2021 05:29:44 +0000 (UTC) Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by mail.kernel.org (Postfix) with ESMTP id 4D9EC65005 for <linux-xfs@archiver.kernel.org>; Fri, 5 Mar 2021 05:29:44 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S229446AbhCEF3n (ORCPT <rfc822;linux-xfs@archiver.kernel.org>); Fri, 5 Mar 2021 00:29:43 -0500 Received: from mail107.syd.optusnet.com.au ([211.29.132.53]:38680 "EHLO mail107.syd.optusnet.com.au" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S229448AbhCEF3m (ORCPT <rfc822;linux-xfs@vger.kernel.org>); Fri, 5 Mar 2021 00:29:42 -0500 Received: from dread.disaster.area (pa49-179-130-210.pa.nsw.optusnet.com.au [49.179.130.210]) by mail107.syd.optusnet.com.au (Postfix) with ESMTPS id 5FD84ECB0AE for <linux-xfs@vger.kernel.org>; Fri, 5 Mar 2021 16:29:41 +1100 (AEDT) Received: from discord.disaster.area ([192.168.253.110]) by dread.disaster.area with esmtp (Exim 4.92.3) (envelope-from <david@fromorbit.com>) id 1lI2kg-00Fbo9-Ad for linux-xfs@vger.kernel.org; Fri, 05 Mar 2021 16:11:50 +1100 Received: from dave by discord.disaster.area with local (Exim 4.94) (envelope-from <david@fromorbit.com>) id 1lI2kg-000lZB-2r for linux-xfs@vger.kernel.org; Fri, 05 Mar 2021 16:11:50 +1100 From: Dave Chinner <david@fromorbit.com> To: linux-xfs@vger.kernel.org Subject: [PATCH 09/45] xfs: Fix CIL throttle hang when CIL space used going backwards Date: Fri, 5 Mar 2021 16:11:07 +1100 Message-Id: <20210305051143.182133-10-david@fromorbit.com> X-Mailer: git-send-email 2.28.0 In-Reply-To: <20210305051143.182133-1-david@fromorbit.com> References: <20210305051143.182133-1-david@fromorbit.com> MIME-Version: 1.0 Content-Transfer-Encoding: 8bit X-Optus-CM-Score: 0 X-Optus-CM-Analysis: v=2.3 cv=F8MpiZpN c=1 sm=1 tr=0 cx=a_idp_d a=JD06eNgDs9tuHP7JIKoLzw==:117 a=JD06eNgDs9tuHP7JIKoLzw==:17 a=dESyimp9J3IA:10 a=20KFwNOVAAAA:8 a=pGLkceISAAAA:8 a=VwQbUJbxAAAA:8 a=jMqSS0pQBfBzRtNEyY0A:9 a=AjGcO6oz07-iQ99wixmX:22 Precedence: bulk List-ID: <linux-xfs.vger.kernel.org> X-Mailing-List: linux-xfs@vger.kernel.org
Series	xfs: consolidated log and optimisation changes \| expand [00/45,v3] xfs: consolidated log and optimisation changes [01/45] xfs: initialise attr fork on inode create [02/45] xfs: log stripe roundoff is a property of the log [03/45] xfs: separate CIL commit record IO [04/45] xfs: remove xfs_blkdev_issue_flush [05/45] xfs: async blkdev cache flush [06/45] xfs: CIL checkpoint flushes caches unconditionally [07/45] xfs: remove need_start_rec parameter from xlog_write() [08/45] xfs: journal IO cache flush reductions [09/45] xfs: Fix CIL throttle hang when CIL space used going backwards [10/45] xfs: reduce buffer log item shadow allocations [11/45] xfs: xfs_buf_item_size_segment() needs to pass segment offset [12/45] xfs: optimise xfs_buf_item_size/format for contiguous regions [13/45] xfs: xfs_log_force_lsn isn't passed a LSN [14/45] xfs: AIL needs asynchronous CIL forcing [15/45] xfs: CIL work is serialised, not pipelined [16/45] xfs: type verification is expensive [17/45] xfs: No need for inode number error injection in __xfs_dir3_data_check [18/45] xfs: reduce debug overhead of dir leaf/node checks [19/45] xfs: factor out the CIL transaction header building [20/45] xfs: only CIL pushes require a start record [21/45] xfs: embed the xlog_op_header in the unmount record [22/45] xfs: embed the xlog_op_header in the commit record [23/45] xfs: log tickets don't need log client id [24/45] xfs: move log iovec alignment to preparation function [25/45] xfs: reserve space and initialise xlog_op_header in item formatting [26/45] xfs: log ticket region debug is largely useless [27/45] xfs: pass lv chain length into xlog_write() [28/45] xfs: introduce xlog_write_single() [29/45] xfs:_introduce xlog_write_partial() [30/45] xfs: xlog_write() no longer needs contwr state [31/45] xfs: CIL context doesn't need to count iovecs [32/45] xfs: use the CIL space used counter for emptiness checks [33/45] xfs: lift init CIL reservation out of xc_cil_lock [34/45] xfs: rework per-iclog header CIL reservation [35/45] xfs: introduce per-cpu CIL tracking sructure [36/45] xfs: implement percpu cil space used calculation [37/45] xfs: track CIL ticket reservation in percpu structure [38/45] xfs: convert CIL busy extents to per-cpu [39/45] xfs: Add order IDs to log items in CIL [40/45] xfs: convert CIL to unordered per cpu lists [41/45] xfs: move CIL ordering to the logvec chain [42/45] xfs: __percpu_counter_compare() inode count debug too expensive [43/45] xfs: avoid cil push lock if possible [44/45] xfs: xlog_sync() manually adjusts grant head space [45/45] xfs: expanding delayed logging design with background material

[09/45] xfs: Fix CIL throttle hang when CIL space used going backwards

Commit Message

Patch