[24/28] xfs: reclaim inodes from the LRU

From: Dave Chinner <dchinner@redhat.com>

From: Dave Chinner <dchinner@redhat.com>

Replace the AG radix tree walking reclaim code with a list_lru
walker, giving us both node-aware and memcg-aware inode reclaim
at the XFS level. This requires adding an inode isolation function to
determine if the inode can be reclaim, and a list walker to
dispose of the inodes that were isolated.

We want the isolation function to be non-blocking. If we can't
grab an inode then we either skip it or rotate it. If it's clean
then we skip it, if it's dirty then we rotate to give it time to be
cleaned before it is scanned again.

This congregates the dirty inodes at the tail of the LRU, which
means that if we start hitting a majority of dirty inodes either
there are lots of unlinked inodes in the reclaim list or we've
reclaimed all the clean inodes and we're looped back on the dirty
inodes. Either way, this is an indication we should tell kswapd to
back off.

The non-blocking isolation function introduces a complexity for the
filesystem shutdown case. When the filesystem is shut down, we want
to free the inode even if it is dirty, and this may require
blocking. We already hold the locks needed to do this blocking, so
what we do is that we leave inodes locked - both the ILOCK and the
flush lock - while they are sitting on the dispose list to be freed
after the LRU walk completes.  This allows us to process the
shutdown state outside the LRU walk where we can block safely.

Because we now are reclaiming inodes from the context that it needs
memory in (memcg and/or node), direct reclaim throttling within the
high level reclaim code in now much more effective. Hence we don't
wait on IO for either kswapd or direct reclaim. However, we have to
tell kswapd to back off if we start hitting too many dirty inodes.
This implies we've wrapped around the LRU and don't have many clean
inodes left to reclaim, so it needs to wait a while for the AIL
pushing to clean some of the remaining reclaimable inodes.

Keep in mind we don't have to care about inode lock order or
blocking with inode locks held here because a) we are using
trylocks, and b) once marked with XFS_IRECLAIM they can't be found
via the LRU and inode cache lookups will abort and retry. Hence
nobody will try to lock them in any other context that might also be
holding other inode locks.

Also convert xfs_reclaim_all_inodes() to use a LRU walk to free all
the reclaimable inodes in the filesystem.

Signed-off-by: Dave Chinner <dchinner@redhat.com>
---
 fs/xfs/xfs_icache.c | 404 +++++++++++++-------------------------------
 fs/xfs/xfs_icache.h |  18 +-
 fs/xfs/xfs_inode.h  |  18 ++
 fs/xfs/xfs_super.c  |  46 ++++-
 4 files changed, 190 insertions(+), 296 deletions(-)

Message ID	20191031234618.15403-25-david@fromorbit.com (mailing list archive)
State	New, archived
Headers	show Return-Path: <SRS0=3TFX=YY=vger.kernel.org=linux-fsdevel-owner@kernel.org> Received: from mail.kernel.org (pdx-korg-mail-1.web.codeaurora.org [172.30.200.123]) by pdx-korg-patchwork-2.web.codeaurora.org (Postfix) with ESMTP id 626D4912 for <patchwork-linux-fsdevel@patchwork.kernel.org>; Thu, 31 Oct 2019 23:47:06 +0000 (UTC) Received: from vger.kernel.org (vger.kernel.org [209.132.180.67]) by mail.kernel.org (Postfix) with ESMTP id 3617C217D9 for <patchwork-linux-fsdevel@patchwork.kernel.org>; Thu, 31 Oct 2019 23:47:06 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1729091AbfJaXqt (ORCPT <rfc822;patchwork-linux-fsdevel@patchwork.kernel.org>); Thu, 31 Oct 2019 19:46:49 -0400 Received: from mail105.syd.optusnet.com.au ([211.29.132.249]:56547 "EHLO mail105.syd.optusnet.com.au" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1728729AbfJaXqf (ORCPT <rfc822;linux-fsdevel@vger.kernel.org>); Thu, 31 Oct 2019 19:46:35 -0400 Received: from dread.disaster.area (pa49-180-67-183.pa.nsw.optusnet.com.au [49.180.67.183]) by mail105.syd.optusnet.com.au (Postfix) with ESMTPS id D03C03A23B1; Fri, 1 Nov 2019 10:46:25 +1100 (AEDT) Received: from discord.disaster.area ([192.168.253.110]) by dread.disaster.area with esmtp (Exim 4.92.3) (envelope-from <david@fromorbit.com>) id 1iQK8x-0007Cg-Ps; Fri, 01 Nov 2019 10:46:19 +1100 Received: from dave by discord.disaster.area with local (Exim 4.92.3) (envelope-from <david@fromorbit.com>) id 1iQK8x-00042K-Nv; Fri, 01 Nov 2019 10:46:19 +1100 From: Dave Chinner <david@fromorbit.com> To: linux-xfs@vger.kernel.org Cc: linux-fsdevel@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org Subject: [PATCH 24/28] xfs: reclaim inodes from the LRU Date: Fri, 1 Nov 2019 10:46:14 +1100 Message-Id: <20191031234618.15403-25-david@fromorbit.com> X-Mailer: git-send-email 2.24.0.rc0 In-Reply-To: <20191031234618.15403-1-david@fromorbit.com> References: <20191031234618.15403-1-david@fromorbit.com> MIME-Version: 1.0 Content-Transfer-Encoding: 8bit X-Optus-CM-Score: 0 X-Optus-CM-Analysis: v=2.2 cv=D+Q3ErZj c=1 sm=1 tr=0 a=3wLbm4YUAFX2xaPZIabsgw==:117 a=3wLbm4YUAFX2xaPZIabsgw==:17 a=jpOVt7BSZ2e4Z31A5e1TngXxSK0=:19 a=MeAgGD-zjQ4A:10 a=20KFwNOVAAAA:8 a=CWdGKF79LG_jA4S27w8A:9 a=3olIEYc3oiszChvJ:21 a=oXdPoYE73i3MdTZ7:21 Sender: linux-fsdevel-owner@vger.kernel.org Precedence: bulk List-ID: <linux-fsdevel.vger.kernel.org> X-Mailing-List: linux-fsdevel@vger.kernel.org
Series	mm, xfs: non-blocking inode reclaim \| expand [00/28] mm, xfs: non-blocking inode reclaim [01/28] xfs: Lower CIL flush limit for large logs [02/28] xfs: Throttle commits on delayed background CIL push [03/28] xfs: don't allow log IO to be throttled [04/28] xfs: Improve metadata buffer reclaim accountability [05/28] xfs: correctly acount for reclaimable slabs [06/28] xfs: factor common AIL item deletion code [07/28] xfs: tail updates only need to occur when LSN changes [08/28] xfs: factor inode lookup from xfs_ifree_cluster [09/28] mm: directed shrinker work deferral [10/28] shrinkers: use defer_work for GFP_NOFS sensitive shrinkers [11/28] mm: factor shrinker work calculations [12/28] shrinker: defer work only to kswapd [13/28] shrinker: clean up variable types and tracepoints [14/28] mm: reclaim_state records pages reclaimed, not slabs [15/28] mm: back off direct reclaim on excessive shrinker deferral [16/28] mm: kswapd backoff for shrinkers [17/28] xfs: synchronous AIL pushing [18/28] xfs: don't block kswapd in inode reclaim [19/28] xfs: reduce kswapd blocking on inode locking. [20/28] xfs: kill background reclaim work [21/28] xfs: use AIL pushing for inode reclaim IO [22/28] xfs: remove mode from xfs_reclaim_inodes() [23/28] xfs: track reclaimable inodes using a LRU list [24/28] xfs: reclaim inodes from the LRU [25/28] xfs: remove unusued old inode reclaim code [26/28] xfs: use xfs_ail_push_all in xfs_reclaim_inodes [27/28] rwsem: introduce down/up_write_non_owner [28/28] xfs: rework unreferenced inode lookups

[24/28] xfs: reclaim inodes from the LRU

Commit Message

Comments

Patch