From patchwork Thu Nov 16 02:24:08 2023
Content-Type: text/plain; charset="utf-8"
MIME-Version: 1.0
Content-Transfer-Encoding: 7bit
X-Patchwork-Submitter: Yosry Ahmed <yosryahmed@google.com>
X-Patchwork-Id: 13457530
Return-Path: <owner-linux-mm@kvack.org>
X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on
	aws-us-west-2-korg-lkml-1.web.codeaurora.org
Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17])
	by smtp.lore.kernel.org (Postfix) with ESMTP id 4D9BBC2BB3F
	for <linux-mm@archiver.kernel.org>; Thu, 16 Nov 2023 02:24:23 +0000 (UTC)
Received: by kanga.kvack.org (Postfix)
	id 8D3CF6B03DE; Wed, 15 Nov 2023 21:24:22 -0500 (EST)
Received: by kanga.kvack.org (Postfix, from userid 40)
	id 85BC56B03E2; Wed, 15 Nov 2023 21:24:22 -0500 (EST)
X-Delivered-To: int-list-linux-mm@kvack.org
Received: by kanga.kvack.org (Postfix, from userid 63042)
	id 6AF576B03E4; Wed, 15 Nov 2023 21:24:22 -0500 (EST)
X-Delivered-To: linux-mm@kvack.org
Received: from relay.hostedemail.com (smtprelay0013.hostedemail.com
 [216.40.44.13])
	by kanga.kvack.org (Postfix) with ESMTP id 4E0E76B03DE
	for <linux-mm@kvack.org>; Wed, 15 Nov 2023 21:24:22 -0500 (EST)
Received: from smtpin09.hostedemail.com (a10.router.float.18 [10.200.18.1])
	by unirelay06.hostedemail.com (Postfix) with ESMTP id 2D978B571C
	for <linux-mm@kvack.org>; Thu, 16 Nov 2023 02:24:22 +0000 (UTC)
X-FDA: 81462223164.09.E4A9051
Received: from mail-yw1-f202.google.com (mail-yw1-f202.google.com
 [209.85.128.202])
	by imf03.hostedemail.com (Postfix) with ESMTP id 525ED20007
	for <linux-mm@kvack.org>; Thu, 16 Nov 2023 02:24:20 +0000 (UTC)
Authentication-Results: imf03.hostedemail.com;
	dkim=pass header.d=google.com header.s=20230601 header.b=LOt6xWGY;
	spf=pass (imf03.hostedemail.com: domain of
 3U31VZQoKCNQOEIHO07C436EE6B4.2ECB8DKN-CCAL02A.EH6@flex--yosryahmed.bounces.google.com
 designates 209.85.128.202 as permitted sender)
 smtp.mailfrom=3U31VZQoKCNQOEIHO07C436EE6B4.2ECB8DKN-CCAL02A.EH6@flex--yosryahmed.bounces.google.com;
	dmarc=pass (policy=reject) header.from=google.com
ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed;
 d=hostedemail.com;
	s=arc-20220608; t=1700101460;
	h=from:from:sender:reply-to:subject:subject:date:date:
	 message-id:message-id:to:to:cc:cc:mime-version:mime-version:
	 content-type:content-type:content-transfer-encoding:
	 in-reply-to:in-reply-to:references:references:dkim-signature;
	bh=LSZ+R0SqSzDzKoZBVcJ1mA+eTEuDRTwjKPgQcB0rBEk=;
	b=WMdnnO1vkZFnz+vJD2ac3f0uo2W/Iq+pP9+k2EqKXLqNLKfdbzqo9hYhmKjTwM4aHmKMCv
	sBiOKzZyuxpBtJvbUtkuJ04Xft0ERT+/57C05f3TTd+vCmdEbgacDOVhfYd8XljZWzBL1+
	i22TfupEqa8f724O7EJ5IWXn1Q3NmVQ=
ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1700101460; a=rsa-sha256;
	cv=none;
	b=yjNzrDyCdK+Prk1u+6HukLiUKXzv5d2DCPr6MFGaixzYFlLmeRiF8mqwveuSL3TJym28Uq
	PuQ5VZt1KJiU7ZdhMl/Sz76BFnDjCjl0BgrxbOjlXiffaSNt6BQ8XZzvv7+tKgb9ubcX+8
	u8AtDwVplgmWTBQQxPFUqpDb0ze67GU=
ARC-Authentication-Results: i=1;
	imf03.hostedemail.com;
	dkim=pass header.d=google.com header.s=20230601 header.b=LOt6xWGY;
	spf=pass (imf03.hostedemail.com: domain of
 3U31VZQoKCNQOEIHO07C436EE6B4.2ECB8DKN-CCAL02A.EH6@flex--yosryahmed.bounces.google.com
 designates 209.85.128.202 as permitted sender)
 smtp.mailfrom=3U31VZQoKCNQOEIHO07C436EE6B4.2ECB8DKN-CCAL02A.EH6@flex--yosryahmed.bounces.google.com;
	dmarc=pass (policy=reject) header.from=google.com
Received: by mail-yw1-f202.google.com with SMTP id
 00721157ae682-5b1ff96d5b9so4089357b3.1
        for <linux-mm@kvack.org>; Wed, 15 Nov 2023 18:24:19 -0800 (PST)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
        d=google.com; s=20230601; t=1700101459; x=1700706259; darn=kvack.org;
        h=cc:to:from:subject:message-id:references:mime-version:in-reply-to
         :date:from:to:cc:subject:date:message-id:reply-to;
        bh=LSZ+R0SqSzDzKoZBVcJ1mA+eTEuDRTwjKPgQcB0rBEk=;
        b=LOt6xWGYNDMUJa75L98M6OJlnPs7jSVyxTc/eHuCpX/b6eLFXfRd53q6HL8AZJpsSD
         gQbBrKMuqU+xijW00n/tNvrs26XXHyI3tML8Yd+Dc+F4u/uy1H9/xgUg8gdaTJs+T4iF
         SZaOQydq5QNhJsWgs2H16gsC/IzaE/yTOn3MlFCP9Y20d8K2BmmpJEt26AGn2aW4eo9/
         ahqgSMNzFvr+RCVdk4mZ/9Rka5zkaUfq2aL9SvrgbxlhdOX9MtKG03vQJOhIbr7kvKGv
         KXV20Ux8RHOcSYfIURHE3iGB6B4SlSW8IROy1IYm2iLwFWqJo27QgM7DNA/+WlHZ/p0Z
         Qgog==
X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
        d=1e100.net; s=20230601; t=1700101459; x=1700706259;
        h=cc:to:from:subject:message-id:references:mime-version:in-reply-to
         :date:x-gm-message-state:from:to:cc:subject:date:message-id:reply-to;
        bh=LSZ+R0SqSzDzKoZBVcJ1mA+eTEuDRTwjKPgQcB0rBEk=;
        b=NP3VlNfGGlpbW1K+rgJwrHClfQBNyaY1H/yvJtzfZrkemP53pnxfT/YD5D8LO6/pLg
         sPqBcx12NlRsMJPWF9BQwjHmMAzyJLZnLY7G2HUEUzI3nJ1szLQ+QlP6RiTKtHuUbfve
         VDv6t7xtiL5uthkRyQMMMpx6nXYmL9anO7oIC3MkiYzTIL3P37w8Qq+zFqFcDkQalKZW
         DGxGGB9cX8kpIHj0PWc5o3D2vP3BquJGlrDuJkGHvFGdgid+38aBU2A2LQsyfcwtAJnv
         7vA7vE3Tv02xrPtG7D5NdOojTxOtMB+3jQVQ3HoAinf4ydyQlRKmOJdTqNSsYnqwNtOs
         +Aww==
X-Gm-Message-State: AOJu0YwcYrrhUK3ZMosCYmoOIbUpfgc8WPkaR2YwM2wN4bJzbnFFH70f
	1x8JJNLzFLtXLDX3/eoZUwWvgSq361jFSpeE
X-Google-Smtp-Source: 
 AGHT+IEFkRjPRA+t8bnXs4r4alzEC6UzKUxHSRyn9fc3Cf7ttKBPYuT2Hd+c68LlEN9cCFYG3gXViVMxsm4UDLTD
X-Received: from yosry.c.googlers.com ([fda3:e722:ac3:cc00:20:ed76:c0a8:29b4])
 (user=yosryahmed job=sendgmr) by 2002:a81:4f90:0:b0:5a7:b543:7f0c with SMTP
 id d138-20020a814f90000000b005a7b5437f0cmr404902ywb.10.1700101459465; Wed, 15
 Nov 2023 18:24:19 -0800 (PST)
Date: Thu, 16 Nov 2023 02:24:08 +0000
In-Reply-To: <20231116022411.2250072-1-yosryahmed@google.com>
Mime-Version: 1.0
References: <20231116022411.2250072-1-yosryahmed@google.com>
X-Mailer: git-send-email 2.43.0.rc0.421.g78406f8d94-goog
Message-ID: <20231116022411.2250072-4-yosryahmed@google.com>
Subject: [PATCH v3 3/5] mm: memcg: make stats flushing threshold per-memcg
From: Yosry Ahmed <yosryahmed@google.com>
To: Andrew Morton <akpm@linux-foundation.org>
Cc: Johannes Weiner <hannes@cmpxchg.org>, Michal Hocko <mhocko@kernel.org>,
  Roman Gushchin <roman.gushchin@linux.dev>,
 Shakeel Butt <shakeelb@google.com>,  Muchun Song <muchun.song@linux.dev>,
 Ivan Babrou <ivan@cloudflare.com>, Tejun Heo <tj@kernel.org>,  "
	=?utf-8?q?Michal_Koutn=C3=BD?= " <mkoutny@suse.com>,
 Waiman Long <longman@redhat.com>, kernel-team@cloudflare.com,
  Wei Xu <weixugc@google.com>, Greg Thelen <gthelen@google.com>,
  Domenico Cerasuolo <cerasuolodomenico@gmail.com>, linux-mm@kvack.org,
 cgroups@vger.kernel.org,  linux-kernel@vger.kernel.org,
 Yosry Ahmed <yosryahmed@google.com>
X-Rspamd-Queue-Id: 525ED20007
X-Rspam-User: 
X-Stat-Signature: fkq4mj618kzrd396o8oxnhgjhey3zso5
X-Rspamd-Server: rspam03
X-HE-Tag: 1700101460-336297
X-HE-Meta: 
 U2FsdGVkX1+n/OV/V9cN3IEUeOOD/d/ugb9eI3sRTuFatIU33ClD/AZb2J3qKUiYo3Q70vE9cHZQ3PfZDC+B2FRkz1wGbMwVpg637WWFklt9wZHUIYp39+9mSqAWDwQa+44DcThp+Prm/e4kaMkD4ql5HCvvzbFXclT1m7VhF70W8TxyqSVKgRyrjlLiT9P3g4SSU4KmZu15NYwRKpFHsjzKKv09GXzxih9xpmK7Fy8mxDXwTZZhoTuRkpzGI8YWKAK/m8kBKWfD+/kaULo22kYlRwt8ixR6tK3nimaBnjXGNnmWy97VFejVvCdbcSe/0qFURccRH3pEovuDf5hZ91EVdxaxbVcm4VI9KRzyCMHwk3Sq4PxprkroKRrF1oCK0XdW8R/6tgPrJs0/68pswW8xzPnkDAfNef6U3b6teXmXIMWYJcwmP0oaXbwrDWONzAAlLcmoGIrABU9i+W7CvpyP3L6uKyR7A+el5R0pE5Om/ladXF8jmfNa9ZRs2w+dnDyAhPt4p1c8WgSJdTrDwAtuJ9SCMpJfyyjpc7ap5haCX6TwcYEkfJI/T2sVEtJ8acQVaGiCyDNjMBfCArzto1ymW5yHCiqiPJNx+U7SMQlaubZa1H1QtEthJ7c2OUzFn5GcmXr/2mPOoPBMzHW5DkuqNUbdR7y5oNNsQoKcxyqbWQmVNNjbONmSDWHA+Ge/Vm+1ZpIMRX7aAiDN0xVGl08c8685xMfx90CcRJhtQeDuQaaaYsr+U38WB90id0BTV1TCIBO1wQyNNF5r3gSbE6Ph0lttACpdajJCWFvkRlbAAp1q6grzAIN+PBl9tmF8lS+28iR6OSZRhnOg1twhD1rJdBf1NI4N+FNwM5F2oU7mKnuGwj5bEg9oL9+rkX+F0+15z8nkD6H+Ls9IaPmYcC79/kIg4Z5nRwcL4OKG0DZvmDia9Gtq4w/RtV1YGFL+6Irv/oXsRPBzYL3aS7i
 Smv89ht2
 rhx0Z/w2LyLYsTABmygwCZNVJnPjxvR+UE9WwgcDwHOtGO34QtKIlTcot8LaZPpaw/BZvQtQ4hGHDXry66wO2Q93inaK2fQ1dVcQzyHcgPBQoz0+fSKJnZHTODBSS2X3uLLYXJ3WuWYl/Serje2rPBon+fHN+91afkqf6nCclhCeJcOX9GDqqFYlM3T8jfwZ7XgcwMkoVvh8EIWgD+jYtH/r3NYiaMaOtdpN4nLnOtb/RZJIQST0AUpj/117l9wT/Ny6kZkMbm4RSiiy7YULUQfb2Na52e+d6rlD/BwZTYBRaml/TjpY0MoSRsCwy5jPiv3PeRuccqeqma41Kz88yYkynoi4d9VjrqLJcEhuSbYJNUh6iNLiFFLs8RNm1O9IDuQ1L+EG+Wu2olxcK+8f3nPAInndV/y0f5k7beYut4RXMODOINOHSbMMOMICfI88QkDispzchQtrLZ48K+KsslPt4mG0BpE3ARQtCZ5/xjH3Wv1AsmUcgoBZC8hg9wqXFRO5OKKmeFZ2CS6Nr9qMdzX0JpylIUeJyrkzrc+/jFrFNi+FQ/hmDaSzTDlY/HTBH9dLqlNjkh9iYNFciq3FquWQswK/IFNv6+bIG9il4C9/Wdc96mC1xvWx573m50IYAhz1xPitn4kXzUZ7SZDYhf3AylIap8/LVKw/7Vy/nHR+pFxGCM8eQLGty5HUPLJ1SGOC/ew0MRp05dUXiuAVoNPBLbfddbRukGvG0yAgD0/KH+tfNc3JTDcu0ric7KfpoRO/cA/1GEIJPoorM9xWjPQVdUQ==
X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4
Sender: owner-linux-mm@kvack.org
Precedence: bulk
X-Loop: owner-majordomo@kvack.org
List-ID: <linux-mm.kvack.org>
List-Subscribe: <mailto:majordomo@kvack.org>
List-Unsubscribe: <mailto:majordomo@kvack.org>

A global counter for the magnitude of memcg stats update is maintained
on the memcg side to avoid invoking rstat flushes when the pending
updates are not significant. This avoids unnecessary flushes, which are
not very cheap even if there isn't a lot of stats to flush. It also
avoids unnecessary lock contention on the underlying global rstat lock.

Make this threshold per-memcg. The scheme is followed where percpu (now
also per-memcg) counters are incremented in the update path, and only
propagated to per-memcg atomics when they exceed a certain threshold.

This provides two benefits:
(a) On large machines with a lot of memcgs, the global threshold can be
reached relatively fast, so guarding the underlying lock becomes less
effective. Making the threshold per-memcg avoids this.

(b) Having a global threshold makes it hard to do subtree flushes, as we
cannot reset the global counter except for a full flush. Per-memcg
counters removes this as a blocker from doing subtree flushes, which
helps avoid unnecessary work when the stats of a small subtree are
needed.

Nothing is free, of course. This comes at a cost:
(a) A new per-cpu counter per memcg, consuming NR_CPUS * NR_MEMCGS * 4
bytes. The extra memory usage is insigificant.

(b) More work on the update side, although in the common case it will
only be percpu counter updates. The amount of work scales with the
number of ancestors (i.e. tree depth). This is not a new concept, adding
a cgroup to the rstat tree involves a parent loop, so is charging.
Testing results below show no significant regressions.

(c) The error margin in the stats for the system as a whole increases
from NR_CPUS * MEMCG_CHARGE_BATCH to NR_CPUS * MEMCG_CHARGE_BATCH *
NR_MEMCGS. This is probably fine because we have a similar per-memcg
error in charges coming from percpu stocks, and we have a periodic
flusher that makes sure we always flush all the stats every 2s anyway.

This patch was tested to make sure no significant regressions are
introduced on the update path as follows. The following benchmarks were
ran in a cgroup that is 2 levels deep (/sys/fs/cgroup/a/b/):

(1) Running 22 instances of netperf on a 44 cpu machine with
hyperthreading disabled. All instances are run in a level 2 cgroup, as
well as netserver:
  # netserver -6
  # netperf -6 -H ::1 -l 60 -t TCP_SENDFILE -- -m 10K

Averaging 20 runs, the numbers are as follows:
Base: 40198.0 mbps
Patched: 38629.7 mbps (-3.9%)

The regression is minimal, especially for 22 instances in the same
cgroup sharing all ancestors (so updating the same atomics).

(2) will-it-scale page_fault tests. These tests (specifically
per_process_ops in page_fault3 test) detected a 25.9% regression before
for a change in the stats update path [1]. These are the
numbers from 10 runs (+ is good) on a machine with 256 cpus:

             LABEL            |     MEAN    |   MEDIAN    |   STDDEV   |
------------------------------+-------------+-------------+-------------
  page_fault1_per_process_ops |             |             |            |
  (A) base                    | 270249.164  | 265437.000  | 13451.836  |
  (B) patched                 | 261368.709  | 255725.000  | 13394.767  |
                              | -3.29%      | -3.66%      |            |
  page_fault1_per_thread_ops  |             |             |            |
  (A) base                    | 242111.345  | 239737.000  | 10026.031  |
  (B) patched                 | 237057.109  | 235305.000  | 9769.687   |
                              | -2.09%      | -1.85%      |            |
  page_fault1_scalability     |             |             |
  (A) base                    | 0.034387    | 0.035168    | 0.0018283  |
  (B) patched                 | 0.033988    | 0.034573    | 0.0018056  |
                              | -1.16%      | -1.69%      |            |
  page_fault2_per_process_ops |             |             |
  (A) base                    | 203561.836  | 203301.000  | 2550.764   |
  (B) patched                 | 197195.945  | 197746.000  | 2264.263   |
                              | -3.13%      | -2.73%      |            |
  page_fault2_per_thread_ops  |             |             |
  (A) base                    | 171046.473  | 170776.000  | 1509.679   |
  (B) patched                 | 166626.327  | 166406.000  | 768.753    |
                              | -2.58%      | -2.56%      |            |
  page_fault2_scalability     |             |             |
  (A) base                    | 0.054026    | 0.053821    | 0.00062121 |
  (B) patched                 | 0.053329    | 0.05306     | 0.00048394 |
                              | -1.29%      | -1.41%      |            |
  page_fault3_per_process_ops |             |             |
  (A) base                    | 1295807.782 | 1297550.000 | 5907.585   |
  (B) patched                 | 1275579.873 | 1273359.000 | 8759.160   |
                              | -1.56%      | -1.86%      |            |
  page_fault3_per_thread_ops  |             |             |
  (A) base                    | 391234.164  | 390860.000  | 1760.720   |
  (B) patched                 | 377231.273  | 376369.000  | 1874.971   |
                              | -3.58%      | -3.71%      |            |
  page_fault3_scalability     |             |             |
  (A) base                    | 0.60369     | 0.60072     | 0.0083029  |
  (B) patched                 | 0.61733     | 0.61544     | 0.009855   |
                              | +2.26%      | +2.45%      |            |

All regressions seem to be minimal, and within the normal variance for
the benchmark. The fix for [1] assumes that 3% is noise -- and there
were no further practical complaints), so hopefully this means that such
variations in these microbenchmarks do not reflect on practical
workloads.

(3) I also ran stress-ng in a nested cgroup and did not observe any
obvious regressions.

[1]https://lore.kernel.org/all/20190520063534.GB19312@shao2-debian/

Suggested-by: Johannes Weiner <hannes@cmpxchg.org>
Signed-off-by: Yosry Ahmed <yosryahmed@google.com>
Tested-by: Domenico Cerasuolo <cerasuolodomenico@gmail.com>
---
 mm/memcontrol.c | 50 +++++++++++++++++++++++++++++++++----------------
 1 file changed, 34 insertions(+), 16 deletions(-)

diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index 5ae2a8f04be45..74db05237775d 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -630,6 +630,9 @@ struct memcg_vmstats_percpu {
 	/* Cgroup1: threshold notifications & softlimit tree updates */
 	unsigned long		nr_page_events;
 	unsigned long		targets[MEM_CGROUP_NTARGETS];
+
+	/* Stats updates since the last flush */
+	unsigned int		stats_updates;
 };
 
 struct memcg_vmstats {
@@ -644,6 +647,9 @@ struct memcg_vmstats {
 	/* Pending child counts during tree propagation */
 	long			state_pending[MEMCG_NR_STAT];
 	unsigned long		events_pending[NR_MEMCG_EVENTS];
+
+	/* Stats updates since the last flush */
+	atomic64_t		stats_updates;
 };
 
 /*
@@ -663,9 +669,7 @@ struct memcg_vmstats {
  */
 static void flush_memcg_stats_dwork(struct work_struct *w);
 static DECLARE_DEFERRABLE_WORK(stats_flush_dwork, flush_memcg_stats_dwork);
-static DEFINE_PER_CPU(unsigned int, stats_updates);
 static atomic_t stats_flush_ongoing = ATOMIC_INIT(0);
-static atomic_t stats_flush_threshold = ATOMIC_INIT(0);
 static u64 flush_last_time;
 
 #define FLUSH_TIME (2UL*HZ)
@@ -692,26 +696,37 @@ static void memcg_stats_unlock(void)
 	preempt_enable_nested();
 }
 
+
+static bool memcg_should_flush_stats(struct mem_cgroup *memcg)
+{
+	return atomic64_read(&memcg->vmstats->stats_updates) >
+		MEMCG_CHARGE_BATCH * num_online_cpus();
+}
+
 static inline void memcg_rstat_updated(struct mem_cgroup *memcg, int val)
 {
+	int cpu = smp_processor_id();
 	unsigned int x;
 
 	if (!val)
 		return;
 
-	cgroup_rstat_updated(memcg->css.cgroup, smp_processor_id());
+	cgroup_rstat_updated(memcg->css.cgroup, cpu);
+
+	for (; memcg; memcg = parent_mem_cgroup(memcg)) {
+		x = __this_cpu_add_return(memcg->vmstats_percpu->stats_updates,
+					  abs(val));
+
+		if (x < MEMCG_CHARGE_BATCH)
+			continue;
 
-	x = __this_cpu_add_return(stats_updates, abs(val));
-	if (x > MEMCG_CHARGE_BATCH) {
 		/*
-		 * If stats_flush_threshold exceeds the threshold
-		 * (>num_online_cpus()), cgroup stats update will be triggered
-		 * in __mem_cgroup_flush_stats(). Increasing this var further
-		 * is redundant and simply adds overhead in atomic update.
+		 * If @memcg is already flush-able, increasing stats_updates is
+		 * redundant. Avoid the overhead of the atomic update.
 		 */
-		if (atomic_read(&stats_flush_threshold) <= num_online_cpus())
-			atomic_add(x / MEMCG_CHARGE_BATCH, &stats_flush_threshold);
-		__this_cpu_write(stats_updates, 0);
+		if (!memcg_should_flush_stats(memcg))
+			atomic64_add(x, &memcg->vmstats->stats_updates);
+		__this_cpu_write(memcg->vmstats_percpu->stats_updates, 0);
 	}
 }
 
@@ -730,13 +745,12 @@ static void do_flush_stats(void)
 
 	cgroup_rstat_flush(root_mem_cgroup->css.cgroup);
 
-	atomic_set(&stats_flush_threshold, 0);
 	atomic_set(&stats_flush_ongoing, 0);
 }
 
 void mem_cgroup_flush_stats(void)
 {
-	if (atomic_read(&stats_flush_threshold) > num_online_cpus())
+	if (memcg_should_flush_stats(root_mem_cgroup))
 		do_flush_stats();
 }
 
@@ -750,8 +764,8 @@ void mem_cgroup_flush_stats_ratelimited(void)
 static void flush_memcg_stats_dwork(struct work_struct *w)
 {
 	/*
-	 * Always flush here so that flushing in latency-sensitive paths is
-	 * as cheap as possible.
+	 * Deliberately ignore memcg_should_flush_stats() here so that flushing
+	 * in latency-sensitive paths is as cheap as possible.
 	 */
 	do_flush_stats();
 	queue_delayed_work(system_unbound_wq, &stats_flush_dwork, FLUSH_TIME);
@@ -5784,6 +5798,10 @@ static void mem_cgroup_css_rstat_flush(struct cgroup_subsys_state *css, int cpu)
 			}
 		}
 	}
+	statc->stats_updates = 0;
+	/* We are in a per-cpu loop here, only do the atomic write once */
+	if (atomic64_read(&memcg->vmstats->stats_updates))
+		atomic64_set(&memcg->vmstats->stats_updates, 0);
 }
 
 #ifdef CONFIG_MMU