[v6,4/6] mm: Shuffle initial free memory to improve memory-side-cache utilization

Randomization of the page allocator improves the average utilization of
a direct-mapped memory-side-cache. Memory side caching is a platform
capability that Linux has been previously exposed to in HPC
(high-performance computing) environments on specialty platforms. In
that instance it was a smaller pool of high-bandwidth-memory relative to
higher-capacity / lower-bandwidth DRAM. Now, this capability is going to
be found on general purpose server platforms where DRAM is a cache in
front of higher latency persistent memory [1].

Robert offered an explanation of the state of the art of Linux
interactions with memory-side-caches [2], and I copy it here:

    It's been a problem in the HPC space:
    http://www.nersc.gov/research-and-development/knl-cache-mode-performance-coe/

    A kernel module called zonesort is available to try to help:
    https://software.intel.com/en-us/articles/xeon-phi-software

    and this abandoned patch series proposed that for the kernel:
    https://lkml.org/lkml/2017/8/23/195

    Dan's patch series doesn't attempt to ensure buffers won't conflict, but
    also reduces the chance that the buffers will. This will make performance
    more consistent, albeit slower than "optimal" (which is near impossible
    to attain in a general-purpose kernel).  That's better than forcing
    users to deploy remedies like:
        "To eliminate this gradual degradation, we have added a Stream
         measurement to the Node Health Check that follows each job;
         nodes are rebooted whenever their measured memory bandwidth
         falls below 300 GB/s."

A replacement for zonesort was merged upstream in commit cc9aec03e58f
"x86/numa_emulation: Introduce uniform split capability". With this
numa_emulation capability, memory can be split into cache sized
("near-memory" sized) numa nodes. A bind operation to such a node, and
disabling workloads on other nodes, enables full cache performance.
However, once the workload exceeds the cache size then cache conflicts
are unavoidable. While HPC environments might be able to tolerate
time-scheduling of cache sized workloads, for general purpose server
platforms, the oversubscribed cache case will be the common case.

The worst case scenario is that a server system owner benchmarks a
workload at boot with an un-contended cache only to see that performance
degrade over time, even below the average cache performance due to
excessive conflicts. Randomization clips the peaks and fills in the
valleys of cache utilization to yield steady average performance.

Here are some performance impact details of the patches:

1/ An Intel internal synthetic memory bandwidth measurement tool, saw a
3X speedup in a contrived case that tries to force cache conflicts. The
contrived cased used the numa_emulation capability to force an instance
of the benchmark to be run in two of the near-memory sized numa nodes.
If both instances were placed on the same emulated they would fit and
cause zero conflicts.  While on separate emulated nodes without
randomization they underutilized the cache and conflicted unnecessarily
due to the in-order allocation per node.

2/ A well known Java server application benchmark was run with a heap
size that exceeded cache size by 3X. The cache conflict rate was 8% for
the first run and degraded to 21% after page allocator aging. With
randomization enabled the rate levelled out at 11%.

3/ A MongoDB workload did not observe measurable difference in
cache-conflict rates, but the overall throughput dropped by 7% with
randomization in one case.

4/ Mel Gorman ran his suite of performance workloads with randomization
enabled on platforms without a memory-side-cache and saw a mix of some
improvements and some losses [3].

While there is potentially significant improvement for applications that
depend on low latency access across a wide working-set, the performance
may be negligible to negative for other workloads. For this reason the
shuffle capability defaults to off unless a direct-mapped
memory-side-cache is detected. Even then, the page_alloc.shuffle=0
parameter can be specified to disable the randomization on those
systems.

Outside of memory-side-cache utilization concerns there is potentially
security benefit from randomization. Some data exfiltration and
return-oriented-programming attacks rely on the ability to infer the
location of sensitive data objects. The kernel page allocator,
especially early in system boot, has predictable first-in-first out
behavior for physical pages. Pages are freed in physical address order
when first onlined.

Quoting Kees:
    "While we already have a base-address randomization
     (CONFIG_RANDOMIZE_MEMORY), attacks against the same hardware and
     memory layouts would certainly be using the predictability of
     allocation ordering (i.e. for attacks where the base address isn't
     important: only the relative positions between allocated memory).
     This is common in lots of heap-style attacks. They try to gain
     control over ordering by spraying allocations, etc.

     I'd really like to see this because it gives us something similar
     to CONFIG_SLAB_FREELIST_RANDOM but for the page allocator."

While SLAB_FREELIST_RANDOM reduces the predictability of some local slab
caches it leaves vast bulk of memory to be predictably in order
allocated.  However, it should be noted, the concrete security benefits
are hard to quantify, and no known CVE is mitigated by this
randomization.

Introduce shuffle_free_memory(), and its helper shuffle_zone(), to
perform a Fisher-Yates shuffle of the page allocator 'free_area' lists
when they are initially populated with free memory at boot and at
hotplug time.

The shuffling is done in terms of CONFIG_SHUFFLE_PAGE_ORDER sized free
pages where the default CONFIG_SHUFFLE_PAGE_ORDER is MAX_ORDER-1 i.e.
10, 4MB this trades off randomization granularity for time spent
shuffling.  MAX_ORDER-1 was chosen to be minimally invasive to the page
allocator while still showing memory-side cache behavior improvements,
and the expectation that the security implications of finer granularity
randomization is mitigated by CONFIG_SLAB_FREELIST_RANDOM.

The performance impact of the shuffling appears to be in the noise
compared to other memory initialization work. Also the bulk of the work
is done in the background as a part of deferred_init_memmap().

This initial randomization can be undone over time so a follow-on patch
is introduced to inject entropy on page free decisions. It is reasonable
to ask if the page free entropy is sufficient, but it is not enough due
to the in-order initial freeing of pages. At the start of that process
putting page1 in front or behind page0 still keeps them close together,
page2 is still near page1 and has a high chance of being adjacent. As
more pages are added ordering diversity improves, but there is still
high page locality for the low address pages and this leads to no
significant impact to the cache conflict rate.

[1]: https://itpeernetwork.intel.com/intel-optane-dc-persistent-memory-operating-modes/
[2]: https://lkml.org/lkml/2018/9/22/54
[3]: https://lkml.org/lkml/2018/10/12/309

Cc: Michal Hocko <mhocko@suse.com>
Cc: Kees Cook <keescook@chromium.org>
Cc: Dave Hansen <dave.hansen@linux.intel.com>
Signed-off-by: Dan Williams <dan.j.williams@intel.com>
---
 include/linux/list.h    |   17 ++++
 include/linux/mmzone.h  |    4 +
 include/linux/shuffle.h |   47 ++++++++++
 init/Kconfig            |   36 ++++++++
 mm/Makefile             |    7 +-
 mm/memblock.c           |   16 +++
 mm/memory_hotplug.c     |    3 +
 mm/page_alloc.c         |    3 +
 mm/shuffle.c            |  215 +++++++++++++++++++++++++++++++++++++++++++++++
 9 files changed, 346 insertions(+), 2 deletions(-)
 create mode 100644 include/linux/shuffle.h
 create mode 100644 mm/shuffle.c

Message ID	154510702402.1941238.1616430879354317384.stgit@dwillia2-desk3.amr.corp.intel.com (mailing list archive)
State	New, archived
Headers	show Return-Path: <owner-linux-mm@kvack.org> Received: from mail.wl.linuxfoundation.org (pdx-wl-mail.web.codeaurora.org [172.30.200.125]) by pdx-korg-patchwork-2.web.codeaurora.org (Postfix) with ESMTP id 1DB176C5 for <patchwork-linux-mm@patchwork.kernel.org>; Tue, 18 Dec 2018 04:36:26 +0000 (UTC) Received: from mail.wl.linuxfoundation.org (localhost [127.0.0.1]) by mail.wl.linuxfoundation.org (Postfix) with ESMTP id 0B0172A77F for <patchwork-linux-mm@patchwork.kernel.org>; Tue, 18 Dec 2018 04:36:26 +0000 (UTC) Received: by mail.wl.linuxfoundation.org (Postfix, from userid 486) id F28042A785; Tue, 18 Dec 2018 04:36:25 +0000 (UTC) X-Spam-Checker-Version: SpamAssassin 3.3.1 (2010-03-16) on pdx-wl-mail.web.codeaurora.org X-Spam-Level: X-Spam-Status: No, score=-2.9 required=2.0 tests=BAYES_00,MAILING_LIST_MULTI, RCVD_IN_DNSWL_NONE autolearn=ham version=3.3.1 Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by mail.wl.linuxfoundation.org (Postfix) with ESMTP id 5958A2A77F for <patchwork-linux-mm@patchwork.kernel.org>; Tue, 18 Dec 2018 04:36:24 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id 56EB38E0006; Mon, 17 Dec 2018 23:36:23 -0500 (EST) Delivered-To: linux-mm-outgoing@kvack.org Received: by kanga.kvack.org (Postfix, from userid 40) id 51D3A8E0001; Mon, 17 Dec 2018 23:36:23 -0500 (EST) X-Original-To: int-list-linux-mm@kvack.org X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id 433638E0006; Mon, 17 Dec 2018 23:36:23 -0500 (EST) X-Original-To: linux-mm@kvack.org X-Delivered-To: linux-mm@kvack.org Received: from mail-pg1-f197.google.com (mail-pg1-f197.google.com [209.85.215.197]) by kanga.kvack.org (Postfix) with ESMTP id E6CE88E0001 for <linux-mm@kvack.org>; Mon, 17 Dec 2018 23:36:22 -0500 (EST) Received: by mail-pg1-f197.google.com with SMTP id r13so12602933pgb.7 for <linux-mm@kvack.org>; Mon, 17 Dec 2018 20:36:22 -0800 (PST) X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20161025; h=x-original-authentication-results:x-gm-message-state:subject:from :to:cc:date:message-id:in-reply-to:references:user-agent :mime-version:content-transfer-encoding; bh=FBwD6C9LHwrJwuepoGcOzerrwadbddjaA3KChCu6Cno=; b=PXK9y/N2fVjWh4tYZZ4mYYJPoTZ/lDo80OqBVvV6ZLn9JzSv4GECYAc/876rKX54mp ZWOYXVG1Xk5lQTIre7HXt4R/GI/SlqMaLcX0S5vs2Cqcw+dwdMdmpbTImrB7aciqKdTr p5LyHK9iFJFrJt9gPsTvV4MLm/Kh+3R9Cl70w+J+5dTsLdpSB4U9HPrS8GY4uKYVSH9y fbw0ZHQBKEEEyTGlX6AIIHZUFn/gjjGrYTDB9SN17ALB8j/Lasku/xHk3L7BqKlJeZ1G Z9/9/T+KsLhrZl/bhHx/V8ZQv4Td0gyFWoNRlI64tJutlfB+6FKHoC7RzQx+a+lN/dfJ 4c3w== X-Original-Authentication-Results: mx.google.com; spf=pass (google.com: domain of dan.j.williams@intel.com designates 134.134.136.100 as permitted sender) smtp.mailfrom=dan.j.williams@intel.com; dmarc=pass (p=NONE sp=NONE dis=NONE) header.from=intel.com X-Gm-Message-State: AA+aEWY8j+PUPki1D7FIdCJS2giW550y8sBZZGTHQMFv49swkHv/7Pvn amuuZ0xDS6obWJOH+ag1HtffppNGC6ALBuGWKgpoPllni/eKPLIGwnd2kVqDp2KGILma+UU+2wa 5yjCSc7xc1FHQpe4kpzW/RxV1W5PEV+tqHPVEwvxvj/5r9xIVgzzrKGSI5qPwi84xKg== X-Received: by 2002:a17:902:2c03:: with SMTP id m3mr14344655plb.6.1545107782442; Mon, 17 Dec 2018 20:36:22 -0800 (PST) X-Google-Smtp-Source: AFSGD/Xoiizt1MqvCd/Ju7Zo6ZLsw8jTJw+yrJzm0dr41/WZW88mht+4OlVpsJEE6gEcFedBbxSI X-Received: by 2002:a17:902:2c03:: with SMTP id m3mr14344606plb.6.1545107781064; Mon, 17 Dec 2018 20:36:21 -0800 (PST) ARC-Seal: i=1; a=rsa-sha256; t=1545107781; cv=none; d=google.com; s=arc-20160816; b=TxYn5htnavv7YdzLgXmAg5FZ2LkVQADLtoPg9Azy4no78AP/n/dTRXaRY5y10Ktyjl 4g3VqC2A9frdGx0vRpaiZc86PFtwfwu9wHdfQP04r5sSPqaXW2DJY8m0LWE2j1hRskSS 5OkJfqY9BBfJUmJVwgbsB8092KpmEBGdVXyDIxmn1Y6Px2kwFbXjWvsVXb2MXbJi8DBp kVoTVTFtwc2uMaLul2isH86YqgwzwG0+3TSnwM4qCkaiYwn7Ua3+y2mbSsxvoAvZBbbk /xPNCsU+X7qegFyKCDeWc3jqf3rM/AP2uF8YFf9JZk8Sl9wD+vmbjdHp9rGUf1V3sV4J Mqsw== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=content-transfer-encoding:mime-version:user-agent:references :in-reply-to:message-id:date:cc:to:from:subject; bh=FBwD6C9LHwrJwuepoGcOzerrwadbddjaA3KChCu6Cno=; b=dwB29C4XBDGNDNf7ntZ7XKEl4jhDNYAhd0cH3U3oSRWkqWtTJlqnjaa9ynoGMPXA/x oUK54lj9aRm8v1J3/utvgpCV+4xZPeUiNlEWZFRcoOYbl56cc2QAy1xxyexPULO7gEHH gvXI4HsANM1mfJEOmsnqtCx+/yP+tapReeyqEWAnO1EW4xx/orTFwh5sISKoKtZ3C6P8 XRWQPVSGYx5iHpptvTahcoGDkz0pQ8dZnjBDg4RU1Ny6mRohfQxZpdhUCrUgQTgYFFoI ycXO5hLEux4m9jV2vASt/OInBUD9hlBwM0OMHX6CKMNTk761T2D/9prRKxjbWluJ2ml/ /vkg== ARC-Authentication-Results: i=1; mx.google.com; spf=pass (google.com: domain of dan.j.williams@intel.com designates 134.134.136.100 as permitted sender) smtp.mailfrom=dan.j.williams@intel.com; dmarc=pass (p=NONE sp=NONE dis=NONE) header.from=intel.com Received: from mga07.intel.com (mga07.intel.com. [134.134.136.100]) by mx.google.com with ESMTPS id r14si12554650pfh.229.2018.12.17.20.36.20 for <linux-mm@kvack.org> (version=TLS1_2 cipher=ECDHE-RSA-AES128-GCM-SHA256 bits=128/128); Mon, 17 Dec 2018 20:36:21 -0800 (PST) Received-SPF: pass (google.com: domain of dan.j.williams@intel.com designates 134.134.136.100 as permitted sender) client-ip=134.134.136.100; Authentication-Results: mx.google.com; spf=pass (google.com: domain of dan.j.williams@intel.com designates 134.134.136.100 as permitted sender) smtp.mailfrom=dan.j.williams@intel.com; dmarc=pass (p=NONE sp=NONE dis=NONE) header.from=intel.com X-Amp-Result: SKIPPED(no attachment in message) X-Amp-File-Uploaded: False Received: from fmsmga001.fm.intel.com ([10.253.24.23]) by orsmga105.jf.intel.com with ESMTP/TLS/DHE-RSA-AES256-GCM-SHA384; 17 Dec 2018 20:36:20 -0800 X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="5.56,367,1539673200"; d="scan'208";a="130830957" Received: from dwillia2-desk3.jf.intel.com (HELO dwillia2-desk3.amr.corp.intel.com) ([10.54.39.16]) by fmsmga001.fm.intel.com with ESMTP; 17 Dec 2018 20:36:19 -0800 Subject: [PATCH v6 4/6] mm: Shuffle initial free memory to improve memory-side-cache utilization From: Dan Williams <dan.j.williams@intel.com> To: akpm@linux-foundation.org Cc: Michal Hocko <mhocko@suse.com>, Kees Cook <keescook@chromium.org>, Dave Hansen <dave.hansen@linux.intel.com>, peterz@infradead.org, linux-mm@kvack.org, x86@kernel.org, linux-kernel@vger.kernel.org, mgorman@suse.de Date: Mon, 17 Dec 2018 20:23:44 -0800 Message-ID: <154510702402.1941238.1616430879354317384.stgit@dwillia2-desk3.amr.corp.intel.com> In-Reply-To: <154510700291.1941238.817190985966612531.stgit@dwillia2-desk3.amr.corp.intel.com> References: <154510700291.1941238.817190985966612531.stgit@dwillia2-desk3.amr.corp.intel.com> User-Agent: StGit/0.18-2-gc94f MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: 7bit X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: <linux-mm.kvack.org> X-Virus-Scanned: ClamAV using ClamSMTP
Series	mm: Randomize free memory \| expand [v6,0/6] mm: Randomize free memory [v6,1/6] acpi: Create subtable parsing infrastructure [v6,2/6] acpi: Add HMAT to generic parsing tables [v6,3/6] acpi/numa: Set the memory-side-cache size in memblocks [v6,4/6] mm: Shuffle initial free memory to improve memory-side-cache utilization [v6,5/6] mm: Move buddy list manipulations into helpers [v6,6/6] mm: Maintain randomization of page free lists

[v6,4/6] mm: Shuffle initial free memory to improve memory-side-cache utilization

Commit Message

Comments

Patch