From patchwork Tue Oct 24 23:35:01 2023 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Nhat Pham X-Patchwork-Id: 13435340 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by smtp.lore.kernel.org (Postfix) with ESMTP id A3FF0C07545 for ; Tue, 24 Oct 2023 23:35:06 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id 1B33F6B02F8; Tue, 24 Oct 2023 19:35:06 -0400 (EDT) Received: by kanga.kvack.org (Postfix, from userid 40) id 163716B02F9; Tue, 24 Oct 2023 19:35:06 -0400 (EDT) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id 02B236B02FA; Tue, 24 Oct 2023 19:35:05 -0400 (EDT) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0017.hostedemail.com [216.40.44.17]) by kanga.kvack.org (Postfix) with ESMTP id E91C46B02F8 for ; Tue, 24 Oct 2023 19:35:05 -0400 (EDT) Received: from smtpin15.hostedemail.com (a10.router.float.18 [10.200.18.1]) by unirelay08.hostedemail.com (Postfix) with ESMTP id B45371406CB for ; Tue, 24 Oct 2023 23:35:05 +0000 (UTC) X-FDA: 81381962970.15.65E23FA Received: from mail-pj1-f46.google.com (mail-pj1-f46.google.com [209.85.216.46]) by imf08.hostedemail.com (Postfix) with ESMTP id E6EBF16000F for ; Tue, 24 Oct 2023 23:35:03 +0000 (UTC) Authentication-Results: imf08.hostedemail.com; dkim=pass header.d=gmail.com header.s=20230601 header.b=R21GFA3E; dmarc=pass (policy=none) header.from=gmail.com; spf=pass (imf08.hostedemail.com: domain of nphamcs@gmail.com designates 209.85.216.46 as permitted sender) smtp.mailfrom=nphamcs@gmail.com ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1698190504; a=rsa-sha256; cv=none; b=C8OlX73GQ40trM8leXDA909AvXUwvwOwYCK0il88WG7X1o8yCcUe1nGX08hu4WL8Rswnbg GsRfO3OYIZUoloGIlxNHNdPgqBCM/UrFygiJiWyAB0/rSH4D/O2j97C2nTHxOW2St65E/G IVu/95PcuBiqCDrveDvRjvw1ZDb6eJ4= ARC-Authentication-Results: i=1; imf08.hostedemail.com; dkim=pass header.d=gmail.com header.s=20230601 header.b=R21GFA3E; dmarc=pass (policy=none) header.from=gmail.com; spf=pass (imf08.hostedemail.com: domain of nphamcs@gmail.com designates 209.85.216.46 as permitted sender) smtp.mailfrom=nphamcs@gmail.com ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1698190504; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-transfer-encoding:content-transfer-encoding: in-reply-to:references:dkim-signature; bh=b4Ij5pQG5szV54GYUJBKO1wv/LZcmkWQRS+LtnyDhXg=; b=HnY+xKCbz8WXL2Yu2Ccjspzak9Fe+DIOJdAUoPszL3LhLc5ewTqXPswgxEcMypYgHsOoGR 8A/3AC6myq2RsXsLKWspHDCMMWDmW5EeKk9rGljdCpRGQ73PoXIQs+tsCEKk53EUShK26+ K+XzeXlNQHEKYvAlmrI1eugO+zStwW4= Received: by mail-pj1-f46.google.com with SMTP id 98e67ed59e1d1-27d0a173e61so3487873a91.0 for ; Tue, 24 Oct 2023 16:35:03 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1698190503; x=1698795303; darn=kvack.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to; bh=b4Ij5pQG5szV54GYUJBKO1wv/LZcmkWQRS+LtnyDhXg=; b=R21GFA3EsPyOTBW5no4rbir02FgZHpi50KJbusMF+2+laqmRq2+MpWsmQUHK6ylZi2 ITZ0YRNj0HLpV6AoqfXaL60kx5+kc+GlvvQ6kxIbcsbMjUlhqh4QEojwVu7zg8rHcU/x 0VzTcAbugRD12nxgSYJ7hy9VeTMPu1VpGKjO/ifUQzDseoszIGPIaJ7XE47ciUapvn3w 1cyswTmWTr4Wj5XcwcVxV1HAVzAL9J1peNFtgsQjD7cDRcmRsM+IVykIumkl5AvbYjAu jfVYQ2Huh/5YIJtVksx4Rr5Qd+ulUMcPrdhib/u1kpTcBZ/qVpUgv6yE5b1YNY/KN/LD xxrg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1698190503; x=1698795303; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to; bh=b4Ij5pQG5szV54GYUJBKO1wv/LZcmkWQRS+LtnyDhXg=; b=Yr0l34Ovx0dRVlKs7yRqe5+dZXDRcgGhyySxc2q6neHJHZEbj5pAUWYjzNJMSRhIR7 rPyLZ3FIEMp/BmIDTh4l0Xl99D2u0io/S9kPtfABs9EN8f8AlnIV0nnnzoXpP7zHiWVV 00inajj2itIAMYaqvetVzmhqW3TE58PUqastP6tIJZaNQxG1KE+F5x/rEuWbYQelmWu/ U1xH/QDjwamPBwavKHNsCYYIGj2V6gCO6h/ozbDAjEDAvKJAywMw6VYEAGglhL5d7tgX 8c9Yemwc0WxZhRQ2WyUxJRT+oHnzhg4+fuiNpipAHVzgvZ7SS+TlgJ1Np1bchPWUcFlO PGjA== X-Gm-Message-State: AOJu0YxeGeviHuXLU+m7LIDyVbshOPj1pdXMI/86Mt1YVw+sGzVE5rbv SSIHK2i6/Bu58vjtrU2VC2U= X-Google-Smtp-Source: AGHT+IHnCRDw8I8GGeR+MlaaT+kwS4k/fxb61omhQzxfL7F7XleUyS/Kj/7XfTbFNLKOmrMdoESqhg== X-Received: by 2002:a17:90b:70b:b0:263:1f1c:ef4d with SMTP id s11-20020a17090b070b00b002631f1cef4dmr11664717pjz.10.1698190502525; Tue, 24 Oct 2023 16:35:02 -0700 (PDT) Received: from localhost (fwdproxy-prn-004.fbsv.net. [2a03:2880:ff:4::face:b00c]) by smtp.gmail.com with ESMTPSA id x89-20020a17090a6c6200b0027d06ddc06bsm10378267pjj.33.2023.10.24.16.35.02 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 24 Oct 2023 16:35:02 -0700 (PDT) From: Nhat Pham To: akpm@linux-foundation.org Cc: tj@kernel.org, lizefan.x@bytedance.com, hannes@cmpxchg.org, cerasuolodomenico@gmail.com, yosryahmed@google.com, sjenning@redhat.com, ddstreet@ieee.org, vitaly.wool@konsulko.com, mhocko@kernel.org, roman.gushchin@linux.dev, shakeelb@google.com, muchun.song@linux.dev, hughd@google.com, corbet@lwn.net, konrad.wilk@oracle.com, senozhatsky@chromium.org, rppt@kernel.org, linux-mm@kvack.org, kernel-team@meta.com, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, david@ixit.cz Subject: [RFC PATCH] memcontrol: implement swap bypassing Date: Tue, 24 Oct 2023 16:35:01 -0700 Message-Id: <20231024233501.2639043-1-nphamcs@gmail.com> X-Mailer: git-send-email 2.34.1 MIME-Version: 1.0 X-Rspam-User: X-Rspamd-Server: rspam06 X-Rspamd-Queue-Id: E6EBF16000F X-Stat-Signature: bs7jritgoic3pboo7i68azq9ay6tpk56 X-HE-Tag: 1698190503-64116 X-HE-Meta: U2FsdGVkX1+GgYL1fFlm4Mr3lIYw6HMgeciKp7ruNkuHhQ5TJUGf4FJrnQwfpL2wXQ2y7MW/1kfBgA7PD5t0NOqxg5idLI0GpMOELC6chphk2odMw1JBhds9gD/kolZOlKXhrfhIFZ3WjuAFGWDhe3T8ek8wAt9XdD3vy8j4mWHeB93ay9mwvc3+QJXcYhmKsR7YZAC8HuRKEeWZ/GrYtmpa4d37QqUyd97htjJCm0l6lUksEg/D/EBROHAxL5xAU+I7O49sFx2x49UOVotYZdffxXZ9NAu+A+1uxqfwJ6PUkAtbkvNj9LbO9nNQhDWLCKi62FPiqFOnuCrgC6c3oVyp1OoH0/b0uWOsbWE3wXc81oW2ZqaG/ITM2V0QG9KUUTjmpW9hSue37cb7Lqbg3/tQD2MYQdyejIOL6YrJi/xn46cE7A2qBKATClK4zl0iPTWmzFkT82CSLo+54paw4Y/nEZ+Yi/aL4tlKj33FYp8NHL7/yv+/Q8oxCHehNcpai4ec52kTDevc04soedS8NZi1ugcxY3YLqPwsnlGfSv+ONQQc2la4iZulzdPFRCyAO3IyKsAzPmkMndyVd74X7TmYaR6m314+YqwIZk5mVwX26Z+J/xf1kdgOIbOSLEOfqgWIyYz8Gieb7KxwRrzc/aK3fogQQcJJ2tuaDreW0qz0mRnkBpj2HQou0u6nnq1MwyZAGfKavRN4JiJXSyG3HGf5n1T8+uKqE98Q9i7xGr4D56OtNT9+PDIRsCo6fC4BA97Wcd8AMyKSo4bw5QhD59pnLCV662WcXAnl4SPPaGiz/ex1+MKjNgy3cqGZLoQ+kQRJVOLNAZqH3M4ohRhdRb2PwlSPb/LSK4taz8uc6pmHECEG64jF3TbJtw8uBxnge9er+x//XHurolit5plTeimPX5jK1VkreqI53CYuSXGpioKeD/UwIRsOeYJhixUqr/21szdMmlrkUSska4D iZzfbJft Uo8bRMtCu9ystje1iJDsdGp0mZYOKczT2vHWNwoSNO1KTZDezNIKi2qst0FAchqswKYn+eUoYOhZyD6wOV6idKTrRHBDB81K83wLR6WwDkYAS6CxjVRXHzTDWOTrQ0qUUOVQV56fAlKkPQ+ku4tugZJOPSl+QJWN7FjLkboOw06LBkirpSHiiutxSKBWE6ZdxbwdSgdoTI5qpkzoE57gzRfGcvqEHc5EokDofDZd4nOuDw8RyVYctQ4wpfXG6/0omwQ9iXL1SeFufaL6L48dbqRcjQRglJWrIcNkDgE5grB+x7KExcLCczeNrz1zEYl9qw+iCkwZtHjVGZpgtUdNf9U/EmOzUJWZcRXhnLdkie7hxzIUF/O+G7Rld4bo24pDKCM/QKpAnhEcRH3WSaS5sEKA+Xk5sRRr3EW2eRLRYa0bukKzlU7vyDUrXPX0C4iEzu7Wk4O04J/7HMlwvGDalS53X2mKIo6ktR/KS4eWPXl4COZWi3dXEGgnsaaBf1WY4+n0T X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: List-Subscribe: List-Unsubscribe: During our experiment with zswap, we sometimes observe swap IOs due to occasional zswap store failures and writebacks. These swapping IOs prevent many users who cannot tolerate swapping from adopting zswap to save memory and improve performance where possible. This patch adds the option to bypass swap entirely: do not swap when an zswap store attempt fail, and do not write pages in the zswap pool back to swap. The feature is disabled by default (to preserve the existing behavior), and can be enabled on a cgroup-basis via a new cgroup file. Note that this is subtly different from setting memory.swap.max to 0, as it still allows for pages to be stored in the zswap pool (which itself consumes swap space in its current form). This is the second attempt (spiritual successor) of the following patch: https://lore.kernel.org/linux-mm/20231017003519.1426574-2-nphamcs@gmail.com/ and should be applied on top of the zswap shrinker series: https://lore.kernel.org/linux-mm/20231024203302.1920362-1-nphamcs@gmail.com/ Suggested-by: Johannes Weiner Signed-off-by: Nhat Pham --- Documentation/admin-guide/cgroup-v2.rst | 11 +++++ Documentation/admin-guide/mm/zswap.rst | 6 +++ include/linux/memcontrol.h | 20 ++++++++++ mm/memcontrol.c | 53 +++++++++++++++++++++++++ mm/page_io.c | 6 +++ mm/shmem.c | 8 +++- mm/zswap.c | 9 +++++ 7 files changed, 111 insertions(+), 2 deletions(-) diff --git a/Documentation/admin-guide/cgroup-v2.rst b/Documentation/admin-guide/cgroup-v2.rst index 606b2e0eac4b..34306d70b3f7 100644 --- a/Documentation/admin-guide/cgroup-v2.rst +++ b/Documentation/admin-guide/cgroup-v2.rst @@ -1657,6 +1657,17 @@ PAGE_SIZE multiple when read back. higher than the limit for an extended period of time. This reduces the impact on the workload and memory management. + memory.swap.bypass.enabled + A read-write single value file which exists on non-root + cgroups. The default value is "0". + + When this is set to 1, all swapping attempts are disabled. + Note that this is subtly different from setting memory.swap.max to + 0, as it still allows for pages to be written to the zswap pool + (which also consumes swap space in its current form). However, + zswap store failure will not lead to swapping, and zswap writebacks + will be disabled altogether. + memory.zswap.current A read-only single value file which exists on non-root cgroups. diff --git a/Documentation/admin-guide/mm/zswap.rst b/Documentation/admin-guide/mm/zswap.rst index 522ae22ccb84..b7bf481a3e25 100644 --- a/Documentation/admin-guide/mm/zswap.rst +++ b/Documentation/admin-guide/mm/zswap.rst @@ -153,6 +153,12 @@ attribute, e. g.:: Setting this parameter to 100 will disable the hysteresis. +Some users cannot tolerate the swapping that comes with zswap store failures +and zswap writebacks. Swapping can be disabled entirely (without disabling +zswap itself) on a cgroup-basis as follows: + + echo 1 > /sys/fs/cgroup//memory.swap.bypass.enabled + When there is a sizable amount of cold memory residing in the zswap pool, it can be advantageous to proactively write these cold pages to swap and reclaim the memory for other use cases. By default, the zswap shrinker is disabled. diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h index c1846e57011b..e481c5c609f2 100644 --- a/include/linux/memcontrol.h +++ b/include/linux/memcontrol.h @@ -221,6 +221,9 @@ struct mem_cgroup { unsigned long zswap_max; #endif + /* bypass swap (on zswap failure and writebacks) */ + bool swap_bypass_enabled; + unsigned long soft_limit; /* vmpressure notifications */ @@ -1157,6 +1160,13 @@ unsigned long mem_cgroup_soft_limit_reclaim(pg_data_t *pgdat, int order, gfp_t gfp_mask, unsigned long *total_scanned); +static inline bool mem_cgroup_swap_bypass_enabled(struct mem_cgroup *memcg) +{ + return memcg && READ_ONCE(memcg->swap_bypass_enabled); +} + +bool mem_cgroup_swap_bypass_folio(struct folio *folio); + #else /* CONFIG_MEMCG */ #define MEM_CGROUP_ID_SHIFT 0 @@ -1615,6 +1625,16 @@ unsigned long mem_cgroup_soft_limit_reclaim(pg_data_t *pgdat, int order, { return 0; } + +static inline bool mem_cgroup_swap_bypass_enabled(struct mem_cgroup *memcg) +{ + return false; +} + +static inline bool mem_cgroup_swap_bypass_folio(struct folio *folio) +{ + return false; +} #endif /* CONFIG_MEMCG */ static inline void __inc_lruvec_kmem_state(void *p, enum node_stat_item idx) diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 568d9d037a59..f231cf2f745b 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -7928,6 +7928,28 @@ bool mem_cgroup_swap_full(struct folio *folio) return false; } +bool mem_cgroup_swap_bypass_folio(struct folio *folio) +{ + struct obj_cgroup *objcg = get_obj_cgroup_from_folio(folio); + struct mem_cgroup *memcg; + bool ret; + + if (!objcg) + return false; + + if (mem_cgroup_disabled()) { + obj_cgroup_put(objcg); + return false; + } + + memcg = get_mem_cgroup_from_objcg(objcg); + ret = mem_cgroup_swap_bypass_enabled(memcg); + + mem_cgroup_put(memcg); + obj_cgroup_put(objcg); + return ret; +} + static int __init setup_swap_account(char *s) { pr_warn_once("The swapaccount= commandline option is deprecated. " @@ -8013,6 +8035,31 @@ static int swap_events_show(struct seq_file *m, void *v) return 0; } +static int swap_bypass_enabled_show(struct seq_file *m, void *v) +{ + struct mem_cgroup *memcg = mem_cgroup_from_seq(m); + + seq_printf(m, "%d\n", READ_ONCE(memcg->swap_bypass_enabled)); + return 0; +} + +static ssize_t swap_bypass_enabled_write(struct kernfs_open_file *of, + char *buf, size_t nbytes, loff_t off) +{ + struct mem_cgroup *memcg = mem_cgroup_from_css(of_css(of)); + int swap_bypass_enabled; + ssize_t parse_ret = kstrtoint(strstrip(buf), 0, &swap_bypass_enabled); + + if (parse_ret) + return parse_ret; + + if (swap_bypass_enabled != 0 && swap_bypass_enabled != 1) + return -ERANGE; + + WRITE_ONCE(memcg->swap_bypass_enabled, swap_bypass_enabled); + return nbytes; +} + static struct cftype swap_files[] = { { .name = "swap.current", @@ -8042,6 +8089,12 @@ static struct cftype swap_files[] = { .file_offset = offsetof(struct mem_cgroup, swap_events_file), .seq_show = swap_events_show, }, + { + .name = "swap.bypass.enabled", + .flags = CFTYPE_NOT_ON_ROOT, + .seq_show = swap_bypass_enabled_show, + .write = swap_bypass_enabled_write, + }, { } /* terminate */ }; diff --git a/mm/page_io.c b/mm/page_io.c index cb559ae324c6..0c84e1592c39 100644 --- a/mm/page_io.c +++ b/mm/page_io.c @@ -201,6 +201,12 @@ int swap_writepage(struct page *page, struct writeback_control *wbc) folio_end_writeback(folio); return 0; } + + if (mem_cgroup_swap_bypass_folio(folio)) { + folio_mark_dirty(folio); + return AOP_WRITEPAGE_ACTIVATE; + } + __swap_writepage(&folio->page, wbc); return 0; } diff --git a/mm/shmem.c b/mm/shmem.c index cab053831fea..6ce1d4a7a48b 100644 --- a/mm/shmem.c +++ b/mm/shmem.c @@ -1514,8 +1514,12 @@ static int shmem_writepage(struct page *page, struct writeback_control *wbc) mutex_unlock(&shmem_swaplist_mutex); BUG_ON(folio_mapped(folio)); - swap_writepage(&folio->page, wbc); - return 0; + /* + * Seeing AOP_WRITEPAGE_ACTIVATE here indicates swapping is disabled on + * zswap store failure. Note that in that case the folio is already + * re-marked dirty by swap_writepage() + */ + return swap_writepage(&folio->page, wbc); } mutex_unlock(&shmem_swaplist_mutex); diff --git a/mm/zswap.c b/mm/zswap.c index c40697f07ba3..f19e26d647a3 100644 --- a/mm/zswap.c +++ b/mm/zswap.c @@ -535,6 +535,9 @@ static unsigned long zswap_shrinker_scan(struct shrinker *shrinker, struct zswap_pool *pool = shrinker->private_data; bool encountered_page_in_swapcache = false; + if (mem_cgroup_swap_bypass_enabled(sc->memcg)) + return SHRINK_STOP; + nr_protected = atomic_long_read(&lruvec->zswap_lruvec_state.nr_zswap_protected); lru_size = list_lru_shrink_count(&pool->list_lru, sc); @@ -565,6 +568,9 @@ static unsigned long zswap_shrinker_count(struct shrinker *shrinker, struct lruvec *lruvec = mem_cgroup_lruvec(memcg, NODE_DATA(sc->nid)); unsigned long nr_backing, nr_stored, nr_freeable, nr_protected; + if (mem_cgroup_swap_bypass_enabled(memcg)) + return 0; + #ifdef CONFIG_MEMCG_KMEM cgroup_rstat_flush(memcg->css.cgroup); nr_backing = memcg_page_state(memcg, MEMCG_ZSWAP_B) >> PAGE_SHIFT; @@ -890,6 +896,9 @@ static int shrink_memcg(struct mem_cgroup *memcg) struct zswap_pool *pool; int nid, shrunk = 0; + if (mem_cgroup_swap_bypass_enabled(memcg)) + return -EINVAL; + /* * Skip zombies because their LRUs are reparented and we would be * reclaiming from the parent instead of the dead memcg.