[v6,05/17] mm: Assign memcg-aware shrinkers bitmap to memcg

Message ID	152663295709.5308.12103481076537943325.stgit@localhost.localdomain (mailing list archive)
State	New, archived
Headers	show Return-Path: <owner-linux-mm@kvack.org> Received-SPF: pass (google.com: domain of ktkhai@virtuozzo.com designates 104.47.1.128 as permitted sender) client-ip=104.47.1.128; Subject: [PATCH v6 05/17] mm: Assign memcg-aware shrinkers bitmap to memcg From: Kirill Tkhai <ktkhai@virtuozzo.com> To: akpm@linux-foundation.org, vdavydov.dev@gmail.com, shakeelb@google.com, viro@zeniv.linux.org.uk, hannes@cmpxchg.org, mhocko@kernel.org, ktkhai@virtuozzo.com, tglx@linutronix.de, pombredanne@nexb.com, stummala@codeaurora.org, gregkh@linuxfoundation.org, sfr@canb.auug.org.au, guro@fb.com, mka@chromium.org, penguin-kernel@I-love.SAKURA.ne.jp, chris@chris-wilson.co.uk, longman@redhat.com, minchan@kernel.org, ying.huang@intel.com, mgorman@techsingularity.net, jbacik@fb.com, linux@roeck-us.net, linux-kernel@vger.kernel.org, linux-mm@kvack.org, willy@infradead.org, lirongqing@baidu.com, aryabinin@virtuozzo.com Date: Fri, 18 May 2018 11:42:37 +0300 Message-ID: <152663295709.5308.12103481076537943325.stgit@localhost.localdomain> In-Reply-To: <152663268383.5308.8660992135988724014.stgit@localhost.localdomain> References: <152663268383.5308.8660992135988724014.stgit@localhost.localdomain> User-Agent: StGit/0.18 MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: 7bit Received-SPF: None (protection.outlook.com: virtuozzo.com does not designate permitted sender hosts) X-Microsoft-Exchange-Diagnostics: =?utf-8?B?MTtEQjZQUjA4MDFNQjEzMzY7MjM6Uld0Wm5zaDAwRDI3cjhlTmRrWFRDR0Zu?= =?utf-8?B?SnBtR3J2V3VzemdTdHVJOThXYlFQL2ZQRGk1aFpLeGQzR2lrSjh0eCtwaUgx?= =?utf-8?B?TXduR1orc1E2Um4yODlPekpyM09rREVzQ1Nza2czZEhHbFErZGV0OWdUbHdU?= =?utf-8?B?SzZvcFNWVnF5NUV4R0dCWWJyMUZramQ3bFg5WHFDcXByL3JnTFBORXZkTGpx?= =?utf-8?B?T01VWGZQanU2ZjZJQTd6ZzdrQW1lQWhGcUlXcmNYUmN6eVM2SXk5VlQ1OGx2?= =?utf-8?B?ZXNPN0JjeUc4TlR6NTgvMnN6ZGxRcFJMZUR3dFVLUUR5MkhrQ0twK2pMbW13?= =?utf-8?B?N0VzU214TUhwL2tYTU9XVWRXZ2VSWERGZVdOa0R6S3lRdDRSR1FrU1l5WFN2?= =?utf-8?B?UDZGQXFwRWhZVlJtcnV0Sno0akVBWlBFZlhFQW5qempNN0EybzQ2aVpkTDhi?= =?utf-8?B?UFRNbzIxVkozbnBOSEMwcTAyYWRhSHFjLzNXeFhqYTZZNlVYNmpwQm5UL0lX?= =?utf-8?B?OEl0RUdUWmNoQnZNOUUyWnRrYmdSdjYzcC9WK2R5N1RUSElzZjc1YU1nT2hV?= =?utf-8?B?NkhRWExic2VFOEY1MG1qdlV4aytkL0hhK2pZam1xcG50MUlSRGs4c2Ewc2pr?= =?utf-8?B?WjVMeVVMdjd0TW5ERFAwQnpzSDVNeFEweWpzRlFxS3owZ3Y1L0kxS3ZHL0k5?= =?utf-8?B?QWE2Y21jYWtyNEovcjNXdDAxaEYvdFNYMFZtZGxkbTlJbEdudDBKQ2xOYWw3?= =?utf-8?B?cG9NWEVTNWNTejZrSno0WmdMeTFXV1VpRzBRR2NqVFBvWnFnUkJSQURKcjVj?= =?utf-8?B?SHlVMmdEMittK0txa0NTdE1mdWNJS05LZEdPK2FTUnJVbUJGYmRXMlZxVHAz?= =?utf-8?B?d2p0bG90S3BDaTBLUmEvNzhqQm9Kb2FpWVZMYTNKckF1dm5MWTMvSTFKY0Jw?= =?utf-8?B?OGJnVlF6UU81cW5TZGM4SjFJZlRwWTJaZHdnRmhtTWx4dU5pS0QrbTNHQTVY?= =?utf-8?B?VitIaml2WjR4K1lzUzNpRXhRbG1mTlFEU09oekZsVjBLSjZGTzRabXVpaWM4?= =?utf-8?B?YzRQWnAxcHZVZ3ltVWxHdUVZMlA4RHZ3YjR3UHFjRzZPaUt4U1M3VEFCTE81?= =?utf-8?B?SFJDaXpzSzEyM0JhUTI4Z1YyTC84U2hxSHVpSWlUYkxKcnc0alNVWjczQ1I4?= =?utf-8?B?aEpLd3RkUXZPM0xodVlaTEFuTmJTWGsva2kwMzAwdDlQYmFaY3V5WW1DdGps?= =?utf-8?B?dlFkZjQ1Y3lubkVtMmpZQ1o2U0FlS0xpNEllc1VaeHlHNkh4UVBiM1E5RFZ4?= =?utf-8?B?dHIyVUpDYlVxTjU5N1IwRFFQbGxRbHVmVHJ4NjMxM2xhR0RZQ0kyQkVUaE9k?= =?utf-8?B?Q3cvSTFSUWtFOEwwUDNONkxkb21hNUNOcTZrbnp4MXRRM09XRzZob0djeDhB?= =?utf-8?B?NW14bHlIRS9WMTVpNTZvUXpRNmZIQlVOU1dwUXhXOEZZc0xoT3pkVmlsMWRx?= =?utf-8?B?enVOT3hETDZkUk5zUUsrckJMeFN1SHZRdTRSWFg1akVCMWJNclRtaFJaUC9U?= =?utf-8?B?MUVCMHdLb01YNlNWSFAwM2hGWjh4cTlQd3hHck93c1UxSDFCbGswS3U1K2tj?= =?utf-8?B?cGluQUM1eWxoQ2ppZnE3TEYwQVlZNmNtUVFtTmRFSGtxVjI5bWZkNDNFREdj?= =?utf-8?B?Yno5K2ZDcXM4K1RzK0o3WHpjbityN1Q0T1NOS1krOUhVVGpjbEIxMnoxeTgx?= =?utf-8?B?SW1xRkRlMDd5Skt5akdWWndFWVFlUVVSM2doeXZGQmxhTVJxUWg3TURoL0hw?= =?utf-8?Q?a3jZGzHIOX4vXwJ?= X-Microsoft-Antispam-Message-Info: eunsHif6556Z15WBHLcg3c20tPP0xU1KNwYpze7HuQ8luClxY4YA+Gfw8+sW4G/ILbgBK2DneEMR1M2tnPwaLCVfU+vUNH48CbldfbHj44CyDuJGyuOgi/n5giwF+KbGVzk7g/rrOXuflsG5SzgP+wsG/m5u+XSLq8mmpyDzSXwcmCS/5hshhXCjDTUdnxSm X-Microsoft-Exchange-Diagnostics: 1; DB6PR0801MB1336; 6:C3W//2yD55mnwYWEG/bkzjrRN69fLlYxJgSVV9X5cTKhd2mjp1wkiPnMBC57mFHHU14J6oZ/OBcgpf62tSJjqX4810V0RvxdtNJbAmSdDziJ//2VRU66BS9qKPmrbASI0Ilsj7zMr+312+4mq7gyrUDuLAKhsL3btf/JmUD58WEjg2qhBTr1l8i50PvVEYKxL9ayoQj90CBCvS16OtBUneq6fgGnJDeJ2f+VqALf+3LdKRCnSzgeEKRI1begZt9hfkprlEFDP1qXW8u4Cq5tAVtLo1FPWxWJTr5zVLH2ndU7kxMYUvrmhw8fkkAehGKVPKH45BdaniLIRQWcosQz+2bATQghJtS/dAO6RezQMQcb7eJjV4OulafYb8msWX0Qq3Ygf+LG8eVQ4pYfLjIMhLn/m43rGHIBTbxJEjHu2VTcXgVETBQG1RaULXXOQ0cNheASTG+82jxBPHkJiT9wPw==; 5:tGsKZ39sjG/A+tOvFjbXPIGMkyDYvX+3h4NcJdFTk3LtdWvW13IHs8cqQyv1O/t6q5VOn6HMHijcm1J5pduAVZr33Gahv1dLUIgqT5Oz7S72MQ30Oi1xjKvrMbNwaEGSyZFdwced7GnY1l3Sf0ES8+zIsjsguPUoQGre0HNrfv4=; 24:L8cw4PrlhVK2tRfpbgIJ+QZP2gcb+JIQpC3SaCJ4fTdB77FRFUqX4lVnjnz/+BqK+TpVmSrXtIEqa6VYxCexkfEvij+huyd4I9T/cl6vG3E= SpamDiagnosticOutput: 1:99 SpamDiagnosticMetadata: NSPM X-Microsoft-Exchange-Diagnostics: 1; DB6PR0801MB1336; 7:gPSjoNcWTJ5JyVodX+61UNURlrTCLyKlPYw0ixzmSr0abPK0IuXa0mjHjf0PmwyxXZwVZnpNjmJXtnJTBQfDiileWKJgOqKAIq9uatWEkVLqzih2FyYLSFdMCHjYJf2hclg/eLJxO/r9gLlX3U2ElXlas8ifMCBlydzCkuFK5qL0Yb0TeLn/1JyBb3GP3zf3o/4Fs1a2Thgwu92pDcvFJfBzzaejFsNaaEcqeHioW54Ekd3kg+P54vaK+qNvhdEj; 20:xIyqOMsn4GHnsRq48i3alwcGEenmhSim0nGRAaNXEq1jPB7axYCQsbHHvkcVYM/bhUmJ06C8YrRToRFyL+Jh4rXB+CGdd94lg3QUcUHzPT5kK1gsewBkDJahdbv5V4oj2L7fIYBtk80srTN4xw939WkM/XW6uUBCR1L+r07En3s= X-MS-Office365-Filtering-Correlation-Id: 4bc14661-dadf-4f49-f111-08d5bc9b51c0 Sender: owner-linux-mm@kvack.org Precedence: bulk

Message ID

152663295709.5308.12103481076537943325.stgit@localhost.localdomain (mailing list archive)

State

New, archived

Headers

Received-SPF: pass (google.com: domain of ktkhai@virtuozzo.com designates
	104.47.1.128 as permitted sender) client-ip=104.47.1.128; 
Subject: [PATCH v6 05/17] mm: Assign memcg-aware shrinkers bitmap to memcg
From: Kirill Tkhai <ktkhai@virtuozzo.com>
To: akpm@linux-foundation.org, vdavydov.dev@gmail.com, shakeelb@google.com, 
	viro@zeniv.linux.org.uk, hannes@cmpxchg.org, mhocko@kernel.org,
	ktkhai@virtuozzo.com, tglx@linutronix.de, pombredanne@nexb.com,
	stummala@codeaurora.org, gregkh@linuxfoundation.org,
	sfr@canb.auug.org.au, 
	guro@fb.com, mka@chromium.org, penguin-kernel@I-love.SAKURA.ne.jp,
	chris@chris-wilson.co.uk, longman@redhat.com, minchan@kernel.org,
	ying.huang@intel.com, mgorman@techsingularity.net, jbacik@fb.com,
	linux@roeck-us.net, linux-kernel@vger.kernel.org, linux-mm@kvack.org, 
	willy@infradead.org, lirongqing@baidu.com, aryabinin@virtuozzo.com
Date: Fri, 18 May 2018 11:42:37 +0300
Message-ID: <152663295709.5308.12103481076537943325.stgit@localhost.localdomain>
In-Reply-To: <152663268383.5308.8660992135988724014.stgit@localhost.localdomain>
References: <152663268383.5308.8660992135988724014.stgit@localhost.localdomain>
User-Agent: StGit/0.18
MIME-Version: 1.0
Content-Type: text/plain; charset="utf-8"
Content-Transfer-Encoding: 7bit
Received-SPF: None (protection.outlook.com: virtuozzo.com does not designate
	permitted sender hosts)
SpamDiagnosticOutput: 1:99
SpamDiagnosticMetadata: NSPM
X-MS-Exchange-CrossTenant-OriginalArrivalTime: 18 May 2018 08:42:39.9545
	(UTC)
X-MS-Exchange-CrossTenant-Network-Message-Id: 4bc14661-dadf-4f49-f111-08d5bc9b51c0
X-MS-Exchange-CrossTenant-FromEntityHeader: Hosted
X-MS-Exchange-CrossTenant-Id: 0bc7f26d-0264-416e-a6fc-8352af79c58f
X-MS-Exchange-Transport-CrossTenantHeadersStamped: DB6PR0801MB1336
X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4
Sender: owner-linux-mm@kvack.org
Precedence: bulk
X-Loop: owner-majordomo@kvack.org
List-ID: <linux-mm.kvack.org>
X-Virus-Scanned: ClamAV using ClamSMTP

Commit Message

Kirill Tkhai May 18, 2018, 8:42 a.m. UTC

Imagine a big node with many cpus, memory cgroups and containers.
Let we have 200 containers, every container has 10 mounts,
and 10 cgroups. All container tasks don't touch foreign
containers mounts. If there is intensive pages write,
and global reclaim happens, a writing task has to iterate
over all memcgs to shrink slab, before it's able to go
to shrink_page_list().

Iteration over all the memcg slabs is very expensive:
the task has to visit 200 * 10 = 2000 shrinkers
for every memcg, and since there are 2000 memcgs,
the total calls are 2000 * 2000 = 4000000.

So, the shrinker makes 4 million do_shrink_slab() calls
just to try to isolate SWAP_CLUSTER_MAX pages in one
of the actively writing memcg via shrink_page_list().
I've observed a node spending almost 100% in kernel,
making useless iteration over already shrinked slab.

This patch adds bitmap of memcg-aware shrinkers to memcg.
The size of the bitmap depends on bitmap_nr_ids, and during
memcg life it's maintained to be enough to fit bitmap_nr_ids
shrinkers. Every bit in the map is related to corresponding
shrinker id.

Next patches will maintain set bit only for really charged
memcg. This will allow shrink_slab() to increase its
performance in significant way. See the last patch for
the numbers.

Signed-off-by: Kirill Tkhai <ktkhai@virtuozzo.com>
---
 include/linux/memcontrol.h |   14 +++++
 mm/memcontrol.c            |  120 ++++++++++++++++++++++++++++++++++++++++++++
 mm/vmscan.c                |   10 ++++
 3 files changed, 144 insertions(+)

Comments

Vladimir Davydov May 20, 2018, 7:27 a.m. UTC | #1

On Fri, May 18, 2018 at 11:42:37AM +0300, Kirill Tkhai wrote:
> Imagine a big node with many cpus, memory cgroups and containers.
> Let we have 200 containers, every container has 10 mounts,
> and 10 cgroups. All container tasks don't touch foreign
> containers mounts. If there is intensive pages write,
> and global reclaim happens, a writing task has to iterate
> over all memcgs to shrink slab, before it's able to go
> to shrink_page_list().
> 
> Iteration over all the memcg slabs is very expensive:
> the task has to visit 200 * 10 = 2000 shrinkers
> for every memcg, and since there are 2000 memcgs,
> the total calls are 2000 * 2000 = 4000000.
> 
> So, the shrinker makes 4 million do_shrink_slab() calls
> just to try to isolate SWAP_CLUSTER_MAX pages in one
> of the actively writing memcg via shrink_page_list().
> I've observed a node spending almost 100% in kernel,
> making useless iteration over already shrinked slab.
> 
> This patch adds bitmap of memcg-aware shrinkers to memcg.
> The size of the bitmap depends on bitmap_nr_ids, and during
> memcg life it's maintained to be enough to fit bitmap_nr_ids
> shrinkers. Every bit in the map is related to corresponding
> shrinker id.
> 
> Next patches will maintain set bit only for really charged
> memcg. This will allow shrink_slab() to increase its
> performance in significant way. See the last patch for
> the numbers.
> 
> Signed-off-by: Kirill Tkhai <ktkhai@virtuozzo.com>
> ---
>  include/linux/memcontrol.h |   14 +++++
>  mm/memcontrol.c            |  120 ++++++++++++++++++++++++++++++++++++++++++++
>  mm/vmscan.c                |   10 ++++
>  3 files changed, 144 insertions(+)
> 
> diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h
> index 996469bc2b82..e51c6e953d7a 100644
> --- a/include/linux/memcontrol.h
> +++ b/include/linux/memcontrol.h
> @@ -112,6 +112,15 @@ struct lruvec_stat {
>  	long count[NR_VM_NODE_STAT_ITEMS];
>  };
>  
> +/*
> + * Bitmap of shrinker::id corresponding to memcg-aware shrinkers,
> + * which have elements charged to this memcg.
> + */
> +struct memcg_shrinker_map {
> +	struct rcu_head rcu;
> +	unsigned long map[0];
> +};
> +
>  /*
>   * per-zone information in memory controller.
>   */
> @@ -125,6 +134,9 @@ struct mem_cgroup_per_node {
>  
>  	struct mem_cgroup_reclaim_iter	iter[DEF_PRIORITY + 1];
>  
> +#ifdef CONFIG_MEMCG_KMEM
> +	struct memcg_shrinker_map __rcu	*shrinker_map;
> +#endif
>  	struct rb_node		tree_node;	/* RB tree node */
>  	unsigned long		usage_in_excess;/* Set to the value by which */
>  						/* the soft limit is exceeded*/
> @@ -1261,6 +1273,8 @@ static inline int memcg_cache_id(struct mem_cgroup *memcg)
>  	return memcg ? memcg->kmemcg_id : -1;
>  }
>  
> +extern int memcg_expand_shrinker_maps(int new_id);
> +
>  #else
>  #define for_each_memcg_cache_index(_idx)	\
>  	for (; NULL; )
> diff --git a/mm/memcontrol.c b/mm/memcontrol.c
> index 023a1e9c900e..317a72137b95 100644
> --- a/mm/memcontrol.c
> +++ b/mm/memcontrol.c
> @@ -320,6 +320,120 @@ EXPORT_SYMBOL(memcg_kmem_enabled_key);
>  
>  struct workqueue_struct *memcg_kmem_cache_wq;
>  
> +static int memcg_shrinker_map_size;
> +static DEFINE_MUTEX(memcg_shrinker_map_mutex);
> +
> +static void memcg_free_shrinker_map_rcu(struct rcu_head *head)
> +{
> +	kvfree(container_of(head, struct memcg_shrinker_map, rcu));
> +}
> +
> +static int memcg_expand_one_shrinker_map(struct mem_cgroup *memcg,
> +					 int size, int old_size)

Nit: No point in passing old_size here. You can instead use
memcg_shrinker_map_size directly.

> +{
> +	struct memcg_shrinker_map *new, *old;
> +	int nid;
> +
> +	lockdep_assert_held(&memcg_shrinker_map_mutex);
> +
> +	for_each_node(nid) {
> +		old = rcu_dereference_protected(
> +				memcg->nodeinfo[nid]->shrinker_map, true);

Nit: Sometimes you use mem_cgroup_nodeinfo() helper, sometimes you
access mem_cgorup->nodeinfo directly. Please, be consistent.

> +		/* Not yet online memcg */
> +		if (!old)
> +			return 0;
> +
> +		new = kvmalloc(sizeof(*new) + size, GFP_KERNEL);
> +		if (!new)
> +			return -ENOMEM;
> +
> +		/* Set all old bits, clear all new bits */
> +		memset(new->map, (int)0xff, old_size);
> +		memset((void *)new->map + old_size, 0, size - old_size);
> +
> +		rcu_assign_pointer(memcg->nodeinfo[nid]->shrinker_map, new);
> +		if (old)
> +			call_rcu(&old->rcu, memcg_free_shrinker_map_rcu);
> +	}
> +
> +	return 0;
> +}
> +
> +static void memcg_free_shrinker_maps(struct mem_cgroup *memcg)
> +{
> +	struct mem_cgroup_per_node *pn;
> +	struct memcg_shrinker_map *map;
> +	int nid;
> +
> +	if (mem_cgroup_is_root(memcg))
> +		return;
> +
> +	for_each_node(nid) {
> +		pn = mem_cgroup_nodeinfo(memcg, nid);
> +		map = rcu_dereference_protected(pn->shrinker_map, true);
> +		if (map)
> +			kvfree(map);
> +		rcu_assign_pointer(pn->shrinker_map, NULL);
> +	}
> +}
> +
> +static int memcg_alloc_shrinker_maps(struct mem_cgroup *memcg)
> +{
> +	struct memcg_shrinker_map *map;
> +	int nid, size, ret = 0;
> +
> +	if (mem_cgroup_is_root(memcg))
> +		return 0;
> +
> +	mutex_lock(&memcg_shrinker_map_mutex);
> +	size = memcg_shrinker_map_size;
> +	for_each_node(nid) {
> +		map = kvzalloc(sizeof(*map) + size, GFP_KERNEL);
> +		if (!map) {

> +			memcg_free_shrinker_maps(memcg);

Nit: Please don't call this function under the mutex as it isn't
necessary. Set 'ret', break the loop, then check 'ret' after releasing
the mutex, and call memcg_free_shrinker_maps() if it's not 0.

> +			ret = -ENOMEM;
> +			break;
> +		}
> +		rcu_assign_pointer(memcg->nodeinfo[nid]->shrinker_map, map);
> +	}
> +	mutex_unlock(&memcg_shrinker_map_mutex);
> +
> +	return ret;
> +}
> +
> +int memcg_expand_shrinker_maps(int nr)

Nit: Please pass the new shrinker id to this function, not the max
number of shrinkers out there - this will look more intuitive. And
please add a comment to this function. Something like:

  Make sure memcg shrinker maps can store the given shrinker id.
  Expand the maps if necessary.

> +{
> +	int size, old_size, ret = 0;
> +	struct mem_cgroup *memcg;
> +
> +	size = DIV_ROUND_UP(nr, BITS_PER_BYTE);

Note, this will turn into DIV_ROUND_UP(id + 1, BITS_PER_BYTE) then.

> +	old_size = memcg_shrinker_map_size;

Nit: old_size won't be needed if you make memcg_expand_one_shrinker_map
use memcg_shrinker_map_size directly.

> +	if (size <= old_size)
> +		return 0;
> +
> +	mutex_lock(&memcg_shrinker_map_mutex);
> +	if (!root_mem_cgroup)
> +		goto unlock;
> +
> +	for_each_mem_cgroup(memcg) {
> +		if (mem_cgroup_is_root(memcg))
> +			continue;
> +		ret = memcg_expand_one_shrinker_map(memcg, size, old_size);
> +		if (ret)
> +			goto unlock;
> +	}
> +unlock:
> +	if (!ret)
> +		memcg_shrinker_map_size = size;
> +	mutex_unlock(&memcg_shrinker_map_mutex);
> +	return ret;
> +}
> +#else /* CONFIG_MEMCG_KMEM */
> +static int memcg_alloc_shrinker_maps(struct mem_cgroup *memcg)
> +{
> +	return 0;
> +}
> +static void memcg_free_shrinker_maps(struct mem_cgroup *memcg) { }
>  #endif /* CONFIG_MEMCG_KMEM */
>  
>  /**
> @@ -4482,6 +4596,11 @@ static int mem_cgroup_css_online(struct cgroup_subsys_state *css)
>  {
>  	struct mem_cgroup *memcg = mem_cgroup_from_css(css);
>  
> +	if (memcg_alloc_shrinker_maps(memcg)) {
> +		mem_cgroup_id_remove(memcg);
> +		return -ENOMEM;
> +	}
> +
>  	/* Online state pins memcg ID, memcg ID pins CSS */
>  	atomic_set(&memcg->id.ref, 1);
>  	css_get(css);
> @@ -4534,6 +4653,7 @@ static void mem_cgroup_css_free(struct cgroup_subsys_state *css)
>  	vmpressure_cleanup(&memcg->vmpressure);
>  	cancel_work_sync(&memcg->high_work);
>  	mem_cgroup_remove_from_trees(memcg);
> +	memcg_free_shrinker_maps(memcg);
>  	memcg_free_kmem(memcg);
>  	mem_cgroup_free(memcg);
>  }
> diff --git a/mm/vmscan.c b/mm/vmscan.c
> index 3de12a9bdf85..f09ea20d7270 100644
> --- a/mm/vmscan.c
> +++ b/mm/vmscan.c
> @@ -171,6 +171,7 @@ static DECLARE_RWSEM(shrinker_rwsem);
>  
>  #ifdef CONFIG_MEMCG_KMEM
>  static DEFINE_IDR(shrinker_idr);

> +static int memcg_shrinker_nr_max;

Nit: Please rename it to shrinker_id_max and make it store max shrinker
id, not the max number shrinkers that have ever been allocated. This
will make it easier to understand IMO.

Also, this variable doesn't belong to this patch as you don't really
need it to expaned mem cgroup maps. Let's please move it to patch 3
(the one that introduces shrinker_idr).

>  
>  static int prealloc_memcg_shrinker(struct shrinker *shrinker)
>  {
> @@ -181,6 +182,15 @@ static int prealloc_memcg_shrinker(struct shrinker *shrinker)
>  	ret = id = idr_alloc(&shrinker_idr, shrinker, 0, 0, GFP_KERNEL);
>  	if (ret < 0)
>  		goto unlock;

> +
> +	if (id >= memcg_shrinker_nr_max) {
> +		if (memcg_expand_shrinker_maps(id + 1)) {
> +			idr_remove(&shrinker_idr, id);
> +			goto unlock;
> +		}
> +		memcg_shrinker_nr_max = id + 1;
> +	}
> +

Then we'll have here:

	if (memcg_expaned_shrinker_maps(id)) {
		idr_remove(shrinker_idr, id);
		goto unlock;
	}

and from patch 3:

	shrinker_id_max = MAX(shrinker_id_max, id);

>  	shrinker->id = id;
>  	ret = 0;
>  unlock:
>

Kirill Tkhai May 21, 2018, 10:16 a.m. UTC | #2

On 20.05.2018 10:27, Vladimir Davydov wrote:
> On Fri, May 18, 2018 at 11:42:37AM +0300, Kirill Tkhai wrote:
>> Imagine a big node with many cpus, memory cgroups and containers.
>> Let we have 200 containers, every container has 10 mounts,
>> and 10 cgroups. All container tasks don't touch foreign
>> containers mounts. If there is intensive pages write,
>> and global reclaim happens, a writing task has to iterate
>> over all memcgs to shrink slab, before it's able to go
>> to shrink_page_list().
>>
>> Iteration over all the memcg slabs is very expensive:
>> the task has to visit 200 * 10 = 2000 shrinkers
>> for every memcg, and since there are 2000 memcgs,
>> the total calls are 2000 * 2000 = 4000000.
>>
>> So, the shrinker makes 4 million do_shrink_slab() calls
>> just to try to isolate SWAP_CLUSTER_MAX pages in one
>> of the actively writing memcg via shrink_page_list().
>> I've observed a node spending almost 100% in kernel,
>> making useless iteration over already shrinked slab.
>>
>> This patch adds bitmap of memcg-aware shrinkers to memcg.
>> The size of the bitmap depends on bitmap_nr_ids, and during
>> memcg life it's maintained to be enough to fit bitmap_nr_ids
>> shrinkers. Every bit in the map is related to corresponding
>> shrinker id.
>>
>> Next patches will maintain set bit only for really charged
>> memcg. This will allow shrink_slab() to increase its
>> performance in significant way. See the last patch for
>> the numbers.
>>
>> Signed-off-by: Kirill Tkhai <ktkhai@virtuozzo.com>
>> ---
>>  include/linux/memcontrol.h |   14 +++++
>>  mm/memcontrol.c            |  120 ++++++++++++++++++++++++++++++++++++++++++++
>>  mm/vmscan.c                |   10 ++++
>>  3 files changed, 144 insertions(+)
>>
>> diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h
>> index 996469bc2b82..e51c6e953d7a 100644
>> --- a/include/linux/memcontrol.h
>> +++ b/include/linux/memcontrol.h
>> @@ -112,6 +112,15 @@ struct lruvec_stat {
>>  	long count[NR_VM_NODE_STAT_ITEMS];
>>  };
>>  
>> +/*
>> + * Bitmap of shrinker::id corresponding to memcg-aware shrinkers,
>> + * which have elements charged to this memcg.
>> + */
>> +struct memcg_shrinker_map {
>> +	struct rcu_head rcu;
>> +	unsigned long map[0];
>> +};
>> +
>>  /*
>>   * per-zone information in memory controller.
>>   */
>> @@ -125,6 +134,9 @@ struct mem_cgroup_per_node {
>>  
>>  	struct mem_cgroup_reclaim_iter	iter[DEF_PRIORITY + 1];
>>  
>> +#ifdef CONFIG_MEMCG_KMEM
>> +	struct memcg_shrinker_map __rcu	*shrinker_map;
>> +#endif
>>  	struct rb_node		tree_node;	/* RB tree node */
>>  	unsigned long		usage_in_excess;/* Set to the value by which */
>>  						/* the soft limit is exceeded*/
>> @@ -1261,6 +1273,8 @@ static inline int memcg_cache_id(struct mem_cgroup *memcg)
>>  	return memcg ? memcg->kmemcg_id : -1;
>>  }
>>  
>> +extern int memcg_expand_shrinker_maps(int new_id);
>> +
>>  #else
>>  #define for_each_memcg_cache_index(_idx)	\
>>  	for (; NULL; )
>> diff --git a/mm/memcontrol.c b/mm/memcontrol.c
>> index 023a1e9c900e..317a72137b95 100644
>> --- a/mm/memcontrol.c
>> +++ b/mm/memcontrol.c
>> @@ -320,6 +320,120 @@ EXPORT_SYMBOL(memcg_kmem_enabled_key);
>>  
>>  struct workqueue_struct *memcg_kmem_cache_wq;
>>  
>> +static int memcg_shrinker_map_size;
>> +static DEFINE_MUTEX(memcg_shrinker_map_mutex);
>> +
>> +static void memcg_free_shrinker_map_rcu(struct rcu_head *head)
>> +{
>> +	kvfree(container_of(head, struct memcg_shrinker_map, rcu));
>> +}
>> +
>> +static int memcg_expand_one_shrinker_map(struct mem_cgroup *memcg,
>> +					 int size, int old_size)
> 
> Nit: No point in passing old_size here. You can instead use
> memcg_shrinker_map_size directly.

This is made for the readability. All the actions with global variable
is made in the same function -- memcg_expand_shrinker_maps(), all
the actions with local variables are also in the same -- memcg_expand_one_shrinker_map().
Accessing memcg_shrinker_map_size in memcg_expand_one_shrinker_map()
looks not intuitive and breaks modularity. 

>> +{
>> +	struct memcg_shrinker_map *new, *old;
>> +	int nid;
>> +
>> +	lockdep_assert_held(&memcg_shrinker_map_mutex);
>> +
>> +	for_each_node(nid) {
>> +		old = rcu_dereference_protected(
>> +				memcg->nodeinfo[nid]->shrinker_map, true);
> 
> Nit: Sometimes you use mem_cgroup_nodeinfo() helper, sometimes you
> access mem_cgorup->nodeinfo directly. Please, be consistent.

Ok, will change.
 
>> +		/* Not yet online memcg */
>> +		if (!old)
>> +			return 0;
>> +
>> +		new = kvmalloc(sizeof(*new) + size, GFP_KERNEL);
>> +		if (!new)
>> +			return -ENOMEM;
>> +
>> +		/* Set all old bits, clear all new bits */
>> +		memset(new->map, (int)0xff, old_size);
>> +		memset((void *)new->map + old_size, 0, size - old_size);
>> +
>> +		rcu_assign_pointer(memcg->nodeinfo[nid]->shrinker_map, new);
>> +		if (old)
>> +			call_rcu(&old->rcu, memcg_free_shrinker_map_rcu);
>> +	}
>> +
>> +	return 0;
>> +}
>> +
>> +static void memcg_free_shrinker_maps(struct mem_cgroup *memcg)
>> +{
>> +	struct mem_cgroup_per_node *pn;
>> +	struct memcg_shrinker_map *map;
>> +	int nid;
>> +
>> +	if (mem_cgroup_is_root(memcg))
>> +		return;
>> +
>> +	for_each_node(nid) {
>> +		pn = mem_cgroup_nodeinfo(memcg, nid);
>> +		map = rcu_dereference_protected(pn->shrinker_map, true);
>> +		if (map)
>> +			kvfree(map);
>> +		rcu_assign_pointer(pn->shrinker_map, NULL);
>> +	}
>> +}
>> +
>> +static int memcg_alloc_shrinker_maps(struct mem_cgroup *memcg)
>> +{
>> +	struct memcg_shrinker_map *map;
>> +	int nid, size, ret = 0;
>> +
>> +	if (mem_cgroup_is_root(memcg))
>> +		return 0;
>> +
>> +	mutex_lock(&memcg_shrinker_map_mutex);
>> +	size = memcg_shrinker_map_size;
>> +	for_each_node(nid) {
>> +		map = kvzalloc(sizeof(*map) + size, GFP_KERNEL);
>> +		if (!map) {
> 
>> +			memcg_free_shrinker_maps(memcg);
> 
> Nit: Please don't call this function under the mutex as it isn't
> necessary. Set 'ret', break the loop, then check 'ret' after releasing
> the mutex, and call memcg_free_shrinker_maps() if it's not 0.

No, it must be called under the mutex. See the race with memcg_expand_one_shrinker_map().
NULL maps are not expanded, and this is the indicator we use to differ memcg, which is
not completely online. If the allocations in memcg_alloc_shrinker_maps() fail at nid == 1,
then freeing of nid == 0 can race with expanding.
 
>> +			ret = -ENOMEM;
>> +			break;
>> +		}
>> +		rcu_assign_pointer(memcg->nodeinfo[nid]->shrinker_map, map);
>> +	}
>> +	mutex_unlock(&memcg_shrinker_map_mutex);
>> +
>> +	return ret;
>> +}
>> +
>> +int memcg_expand_shrinker_maps(int nr)
> 
> Nit: Please pass the new shrinker id to this function, not the max
> number of shrinkers out there - this will look more intuitive. And
> please add a comment to this function. Something like:
> 
>   Make sure memcg shrinker maps can store the given shrinker id.
>   Expand the maps if necessary.
> 
>> +{
>> +	int size, old_size, ret = 0;
>> +	struct mem_cgroup *memcg;
>> +
>> +	size = DIV_ROUND_UP(nr, BITS_PER_BYTE);
> 
> Note, this will turn into DIV_ROUND_UP(id + 1, BITS_PER_BYTE) then.
> 
>> +	old_size = memcg_shrinker_map_size;
> 
> Nit: old_size won't be needed if you make memcg_expand_one_shrinker_map
> use memcg_shrinker_map_size directly.

This is made deliberately. Please, see above.

>> +	if (size <= old_size)
>> +		return 0;
>> +
>> +	mutex_lock(&memcg_shrinker_map_mutex);
>> +	if (!root_mem_cgroup)
>> +		goto unlock;
>> +
>> +	for_each_mem_cgroup(memcg) {
>> +		if (mem_cgroup_is_root(memcg))
>> +			continue;
>> +		ret = memcg_expand_one_shrinker_map(memcg, size, old_size);
>> +		if (ret)
>> +			goto unlock;
>> +	}
>> +unlock:
>> +	if (!ret)
>> +		memcg_shrinker_map_size = size;
>> +	mutex_unlock(&memcg_shrinker_map_mutex);
>> +	return ret;
>> +}
>> +#else /* CONFIG_MEMCG_KMEM */
>> +static int memcg_alloc_shrinker_maps(struct mem_cgroup *memcg)
>> +{
>> +	return 0;
>> +}
>> +static void memcg_free_shrinker_maps(struct mem_cgroup *memcg) { }
>>  #endif /* CONFIG_MEMCG_KMEM */
>>  
>>  /**
>> @@ -4482,6 +4596,11 @@ static int mem_cgroup_css_online(struct cgroup_subsys_state *css)
>>  {
>>  	struct mem_cgroup *memcg = mem_cgroup_from_css(css);
>>  
>> +	if (memcg_alloc_shrinker_maps(memcg)) {
>> +		mem_cgroup_id_remove(memcg);
>> +		return -ENOMEM;
>> +	}
>> +
>>  	/* Online state pins memcg ID, memcg ID pins CSS */
>>  	atomic_set(&memcg->id.ref, 1);
>>  	css_get(css);
>> @@ -4534,6 +4653,7 @@ static void mem_cgroup_css_free(struct cgroup_subsys_state *css)
>>  	vmpressure_cleanup(&memcg->vmpressure);
>>  	cancel_work_sync(&memcg->high_work);
>>  	mem_cgroup_remove_from_trees(memcg);
>> +	memcg_free_shrinker_maps(memcg);
>>  	memcg_free_kmem(memcg);
>>  	mem_cgroup_free(memcg);
>>  }
>> diff --git a/mm/vmscan.c b/mm/vmscan.c
>> index 3de12a9bdf85..f09ea20d7270 100644
>> --- a/mm/vmscan.c
>> +++ b/mm/vmscan.c
>> @@ -171,6 +171,7 @@ static DECLARE_RWSEM(shrinker_rwsem);
>>  
>>  #ifdef CONFIG_MEMCG_KMEM
>>  static DEFINE_IDR(shrinker_idr);
> 
>> +static int memcg_shrinker_nr_max;
> 
> Nit: Please rename it to shrinker_id_max and make it store max shrinker
> id, not the max number shrinkers that have ever been allocated. This
> will make it easier to understand IMO.
>
> Also, this variable doesn't belong to this patch as you don't really
> need it to expaned mem cgroup maps. Let's please move it to patch 3
> (the one that introduces shrinker_idr).
> 
>>  
>>  static int prealloc_memcg_shrinker(struct shrinker *shrinker)
>>  {
>> @@ -181,6 +182,15 @@ static int prealloc_memcg_shrinker(struct shrinker *shrinker)
>>  	ret = id = idr_alloc(&shrinker_idr, shrinker, 0, 0, GFP_KERNEL);
>>  	if (ret < 0)
>>  		goto unlock;
> 
>> +
>> +	if (id >= memcg_shrinker_nr_max) {
>> +		if (memcg_expand_shrinker_maps(id + 1)) {
>> +			idr_remove(&shrinker_idr, id);
>> +			goto unlock;
>> +		}
>> +		memcg_shrinker_nr_max = id + 1;
>> +	}
>> +
> 
> Then we'll have here:
> 
> 	if (memcg_expaned_shrinker_maps(id)) {
> 		idr_remove(shrinker_idr, id);
> 		goto unlock;
> 	}
> 
> and from patch 3:
> 
> 	shrinker_id_max = MAX(shrinker_id_max, id);

So, shrinker_id_max contains "the max number shrinkers that have ever been allocated" minus 1.
The only difference to existing logic is "minus 1", which will be needed to reflect in
shrink_slab_memcg()->for_each_set_bit()...

To have "minus 1" instead of "not to have minus 1" looks a little subjective.

> 
>>  	shrinker->id = id;
>>  	ret = 0;
>>  unlock:
>>

Thanks,
Kirill

Vladimir Davydov May 21, 2018, 6:40 p.m. UTC | #3

On Mon, May 21, 2018 at 01:16:40PM +0300, Kirill Tkhai wrote:
> >> +static int memcg_expand_one_shrinker_map(struct mem_cgroup *memcg,
> >> +					 int size, int old_size)
> > 
> > Nit: No point in passing old_size here. You can instead use
> > memcg_shrinker_map_size directly.
> 
> This is made for the readability. All the actions with global variable
> is made in the same function -- memcg_expand_shrinker_maps(), all
> the actions with local variables are also in the same -- memcg_expand_one_shrinker_map().
> Accessing memcg_shrinker_map_size in memcg_expand_one_shrinker_map()
> looks not intuitive and breaks modularity. 

I guess it depends on how you look at it. Anyway, it's nitpicking so I
won't mind if you leave it as is.

> >> +static int memcg_alloc_shrinker_maps(struct mem_cgroup *memcg)
> >> +{
> >> +	struct memcg_shrinker_map *map;
> >> +	int nid, size, ret = 0;
> >> +
> >> +	if (mem_cgroup_is_root(memcg))
> >> +		return 0;
> >> +
> >> +	mutex_lock(&memcg_shrinker_map_mutex);
> >> +	size = memcg_shrinker_map_size;
> >> +	for_each_node(nid) {
> >> +		map = kvzalloc(sizeof(*map) + size, GFP_KERNEL);
> >> +		if (!map) {
> > 
> >> +			memcg_free_shrinker_maps(memcg);
> > 
> > Nit: Please don't call this function under the mutex as it isn't
> > necessary. Set 'ret', break the loop, then check 'ret' after releasing
> > the mutex, and call memcg_free_shrinker_maps() if it's not 0.
> 
> No, it must be called under the mutex. See the race with memcg_expand_one_shrinker_map().
> NULL maps are not expanded, and this is the indicator we use to differ memcg, which is
> not completely online. If the allocations in memcg_alloc_shrinker_maps() fail at nid == 1,
> then freeing of nid == 0 can race with expanding.

Ah, I see, you're right.

> >> diff --git a/mm/vmscan.c b/mm/vmscan.c
> >> index 3de12a9bdf85..f09ea20d7270 100644
> >> --- a/mm/vmscan.c
> >> +++ b/mm/vmscan.c
> >> @@ -171,6 +171,7 @@ static DECLARE_RWSEM(shrinker_rwsem);
> >>  
> >>  #ifdef CONFIG_MEMCG_KMEM
> >>  static DEFINE_IDR(shrinker_idr);
> > 
> >> +static int memcg_shrinker_nr_max;
> > 
> > Nit: Please rename it to shrinker_id_max and make it store max shrinker
> > id, not the max number shrinkers that have ever been allocated. This
> > will make it easier to understand IMO.
> >
> > Also, this variable doesn't belong to this patch as you don't really
> > need it to expaned mem cgroup maps. Let's please move it to patch 3
> > (the one that introduces shrinker_idr).
> > 
> >>  
> >>  static int prealloc_memcg_shrinker(struct shrinker *shrinker)
> >>  {
> >> @@ -181,6 +182,15 @@ static int prealloc_memcg_shrinker(struct shrinker *shrinker)
> >>  	ret = id = idr_alloc(&shrinker_idr, shrinker, 0, 0, GFP_KERNEL);
> >>  	if (ret < 0)
> >>  		goto unlock;
> > 
> >> +
> >> +	if (id >= memcg_shrinker_nr_max) {
> >> +		if (memcg_expand_shrinker_maps(id + 1)) {
> >> +			idr_remove(&shrinker_idr, id);
> >> +			goto unlock;
> >> +		}
> >> +		memcg_shrinker_nr_max = id + 1;
> >> +	}
> >> +
> > 
> > Then we'll have here:
> > 
> > 	if (memcg_expaned_shrinker_maps(id)) {
> > 		idr_remove(shrinker_idr, id);
> > 		goto unlock;
> > 	}
> > 
> > and from patch 3:
> > 
> > 	shrinker_id_max = MAX(shrinker_id_max, id);
> 
> So, shrinker_id_max contains "the max number shrinkers that have ever been allocated" minus 1.
> The only difference to existing logic is "minus 1", which will be needed to reflect in
> shrink_slab_memcg()->for_each_set_bit()...
> 
> To have "minus 1" instead of "not to have minus 1" looks a little subjective.

OK, leave 'nr' then, but please consider my other comments:

 - rename memcg_shrinker_nr_max to shrinker_nr_max so that the variable
   name is consistent with shrinker_idr

 - move shrinker_nr_max to patch 3 as you don't need it for expanding
   memcg shrinker maps

 - don't use shrinker_nr_max to check whether we need to expand memcg
   maps - simply call memcg_expand_shrinker_maps() and let it decide -
   this will neatly isolate all the logic related to memcg shrinker map
   allocation in memcontrol.c

diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h
index 996469bc2b82..e51c6e953d7a 100644
--- a/include/linux/memcontrol.h
+++ b/include/linux/memcontrol.h
@@ -112,6 +112,15 @@  struct lruvec_stat {
 	long count[NR_VM_NODE_STAT_ITEMS];
 };
 
+/*
+ * Bitmap of shrinker::id corresponding to memcg-aware shrinkers,
+ * which have elements charged to this memcg.
+ */
+struct memcg_shrinker_map {
+	struct rcu_head rcu;
+	unsigned long map[0];
+};
+
 /*
  * per-zone information in memory controller.
  */
@@ -125,6 +134,9 @@  struct mem_cgroup_per_node {
 
 	struct mem_cgroup_reclaim_iter	iter[DEF_PRIORITY + 1];
 
+#ifdef CONFIG_MEMCG_KMEM
+	struct memcg_shrinker_map __rcu	*shrinker_map;
+#endif
 	struct rb_node		tree_node;	/* RB tree node */
 	unsigned long		usage_in_excess;/* Set to the value by which */
 						/* the soft limit is exceeded*/
@@ -1261,6 +1273,8 @@  static inline int memcg_cache_id(struct mem_cgroup *memcg)
 	return memcg ? memcg->kmemcg_id : -1;
 }
 
+extern int memcg_expand_shrinker_maps(int new_id);
+
 #else
 #define for_each_memcg_cache_index(_idx)	\
 	for (; NULL; )
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index 023a1e9c900e..317a72137b95 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -320,6 +320,120 @@  EXPORT_SYMBOL(memcg_kmem_enabled_key);
 
 struct workqueue_struct *memcg_kmem_cache_wq;
 
+static int memcg_shrinker_map_size;
+static DEFINE_MUTEX(memcg_shrinker_map_mutex);
+
+static void memcg_free_shrinker_map_rcu(struct rcu_head *head)
+{
+	kvfree(container_of(head, struct memcg_shrinker_map, rcu));
+}
+
+static int memcg_expand_one_shrinker_map(struct mem_cgroup *memcg,
+					 int size, int old_size)
+{
+	struct memcg_shrinker_map *new, *old;
+	int nid;
+
+	lockdep_assert_held(&memcg_shrinker_map_mutex);
+
+	for_each_node(nid) {
+		old = rcu_dereference_protected(
+				memcg->nodeinfo[nid]->shrinker_map, true);
+		/* Not yet online memcg */
+		if (!old)
+			return 0;
+
+		new = kvmalloc(sizeof(*new) + size, GFP_KERNEL);
+		if (!new)
+			return -ENOMEM;
+
+		/* Set all old bits, clear all new bits */
+		memset(new->map, (int)0xff, old_size);
+		memset((void *)new->map + old_size, 0, size - old_size);
+
+		rcu_assign_pointer(memcg->nodeinfo[nid]->shrinker_map, new);
+		if (old)
+			call_rcu(&old->rcu, memcg_free_shrinker_map_rcu);
+	}
+
+	return 0;
+}
+
+static void memcg_free_shrinker_maps(struct mem_cgroup *memcg)
+{
+	struct mem_cgroup_per_node *pn;
+	struct memcg_shrinker_map *map;
+	int nid;
+
+	if (mem_cgroup_is_root(memcg))
+		return;
+
+	for_each_node(nid) {
+		pn = mem_cgroup_nodeinfo(memcg, nid);
+		map = rcu_dereference_protected(pn->shrinker_map, true);
+		if (map)
+			kvfree(map);
+		rcu_assign_pointer(pn->shrinker_map, NULL);
+	}
+}
+
+static int memcg_alloc_shrinker_maps(struct mem_cgroup *memcg)
+{
+	struct memcg_shrinker_map *map;
+	int nid, size, ret = 0;
+
+	if (mem_cgroup_is_root(memcg))
+		return 0;
+
+	mutex_lock(&memcg_shrinker_map_mutex);
+	size = memcg_shrinker_map_size;
+	for_each_node(nid) {
+		map = kvzalloc(sizeof(*map) + size, GFP_KERNEL);
+		if (!map) {
+			memcg_free_shrinker_maps(memcg);
+			ret = -ENOMEM;
+			break;
+		}
+		rcu_assign_pointer(memcg->nodeinfo[nid]->shrinker_map, map);
+	}
+	mutex_unlock(&memcg_shrinker_map_mutex);
+
+	return ret;
+}
+
+int memcg_expand_shrinker_maps(int nr)
+{
+	int size, old_size, ret = 0;
+	struct mem_cgroup *memcg;
+
+	size = DIV_ROUND_UP(nr, BITS_PER_BYTE);
+	old_size = memcg_shrinker_map_size;
+	if (size <= old_size)
+		return 0;
+
+	mutex_lock(&memcg_shrinker_map_mutex);
+	if (!root_mem_cgroup)
+		goto unlock;
+
+	for_each_mem_cgroup(memcg) {
+		if (mem_cgroup_is_root(memcg))
+			continue;
+		ret = memcg_expand_one_shrinker_map(memcg, size, old_size);
+		if (ret)
+			goto unlock;
+	}
+unlock:
+	if (!ret)
+		memcg_shrinker_map_size = size;
+	mutex_unlock(&memcg_shrinker_map_mutex);
+	return ret;
+}
+#else /* CONFIG_MEMCG_KMEM */
+static int memcg_alloc_shrinker_maps(struct mem_cgroup *memcg)
+{
+	return 0;
+}
+static void memcg_free_shrinker_maps(struct mem_cgroup *memcg) { }
 #endif /* CONFIG_MEMCG_KMEM */
 
 /**
@@ -4482,6 +4596,11 @@  static int mem_cgroup_css_online(struct cgroup_subsys_state *css)
 {
 	struct mem_cgroup *memcg = mem_cgroup_from_css(css);
 
+	if (memcg_alloc_shrinker_maps(memcg)) {
+		mem_cgroup_id_remove(memcg);
+		return -ENOMEM;
+	}
+
 	/* Online state pins memcg ID, memcg ID pins CSS */
 	atomic_set(&memcg->id.ref, 1);
 	css_get(css);
@@ -4534,6 +4653,7 @@  static void mem_cgroup_css_free(struct cgroup_subsys_state *css)
 	vmpressure_cleanup(&memcg->vmpressure);
 	cancel_work_sync(&memcg->high_work);
 	mem_cgroup_remove_from_trees(memcg);
+	memcg_free_shrinker_maps(memcg);
 	memcg_free_kmem(memcg);
 	mem_cgroup_free(memcg);
 }
diff --git a/mm/vmscan.c b/mm/vmscan.c
index 3de12a9bdf85..f09ea20d7270 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -171,6 +171,7 @@  static DECLARE_RWSEM(shrinker_rwsem);
 
 #ifdef CONFIG_MEMCG_KMEM
 static DEFINE_IDR(shrinker_idr);
+static int memcg_shrinker_nr_max;
 
 static int prealloc_memcg_shrinker(struct shrinker *shrinker)
 {
@@ -181,6 +182,15 @@  static int prealloc_memcg_shrinker(struct shrinker *shrinker)
 	ret = id = idr_alloc(&shrinker_idr, shrinker, 0, 0, GFP_KERNEL);
 	if (ret < 0)
 		goto unlock;
+
+	if (id >= memcg_shrinker_nr_max) {
+		if (memcg_expand_shrinker_maps(id + 1)) {
+			idr_remove(&shrinker_idr, id);
+			goto unlock;
+		}
+		memcg_shrinker_nr_max = id + 1;
+	}
+
 	shrinker->id = id;
 	ret = 0;
 unlock:

[v6,05/17] mm: Assign memcg-aware shrinkers bitmap to memcg

Commit Message

Comments

Patch