From patchwork Wed Nov 15 13:27:24 2023 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Ryan Roberts X-Patchwork-Id: 13456674 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by smtp.lore.kernel.org (Postfix) with ESMTP id B20C0C48BF9 for ; Wed, 15 Nov 2023 13:28:01 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id E4ED46B0339; Wed, 15 Nov 2023 08:28:00 -0500 (EST) Received: by kanga.kvack.org (Postfix, from userid 40) id DFECC6B033A; Wed, 15 Nov 2023 08:28:00 -0500 (EST) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id CF8096B033B; Wed, 15 Nov 2023 08:28:00 -0500 (EST) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0012.hostedemail.com [216.40.44.12]) by kanga.kvack.org (Postfix) with ESMTP id BDE2C6B0339 for ; Wed, 15 Nov 2023 08:28:00 -0500 (EST) Received: from smtpin08.hostedemail.com (a10.router.float.18 [10.200.18.1]) by unirelay02.hostedemail.com (Postfix) with ESMTP id 8BA0E12025A for ; Wed, 15 Nov 2023 13:28:00 +0000 (UTC) X-FDA: 81460266720.08.DCFE320 Received: from foss.arm.com (foss.arm.com [217.140.110.172]) by imf20.hostedemail.com (Postfix) with ESMTP id 9892D1C0021 for ; Wed, 15 Nov 2023 13:27:58 +0000 (UTC) Authentication-Results: imf20.hostedemail.com; dkim=none; spf=pass (imf20.hostedemail.com: domain of ryan.roberts@arm.com designates 217.140.110.172 as permitted sender) smtp.mailfrom=ryan.roberts@arm.com; dmarc=pass (policy=none) header.from=arm.com ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1700054879; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-transfer-encoding:content-transfer-encoding: in-reply-to:references; bh=5N01B5Vi5ytOHE5iHt+9JwgPUylfOl/J0fUt0HC/TfU=; b=hb2Obv4WO+c2SsZcYvCOBcvkOZ1JRbZPecGE76YrBjaypW3YC4Iqi8B5xOCKBP2idi10pB ejzIxjDidYTRXHWep8/u/tbzOopqiRZJQcOhJ9z/5dSq+LLt55ws2fMYPhybY3fJCXkGH/ BA695TOnalT5gPx9Kfu3HgzzcYmjBDU= ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1700054879; a=rsa-sha256; cv=none; b=DL35L18QQ0eCBT/uX0vunmIOBA6WT99F+b2S8/n2Dok29HMKfdy9kxhBwd7baT8GjKVZd5 r1BxtRl+RoqvztzsMyE05iJwt0m1nNlBDyxSbLLYDGNaDXjyhQrnerzS/M5Gu6ZS0/YCaE tkRxlQ8qVs6REJljG39lyniD3UCY8yc= ARC-Authentication-Results: i=1; imf20.hostedemail.com; dkim=none; spf=pass (imf20.hostedemail.com: domain of ryan.roberts@arm.com designates 217.140.110.172 as permitted sender) smtp.mailfrom=ryan.roberts@arm.com; dmarc=pass (policy=none) header.from=arm.com Received: from usa-sjc-imap-foss1.foss.arm.com (unknown [10.121.207.14]) by usa-sjc-mx-foss1.foss.arm.com (Postfix) with ESMTP id 0C07BDA7; Wed, 15 Nov 2023 05:28:43 -0800 (PST) Received: from e125769.cambridge.arm.com (e125769.cambridge.arm.com [10.1.196.26]) by usa-sjc-imap-foss1.foss.arm.com (Postfix) with ESMTPSA id 769C23F7B4; Wed, 15 Nov 2023 05:27:54 -0800 (PST) From: Ryan Roberts To: Andrew Morton , Matthew Wilcox , Yin Fengwei , David Hildenbrand , Yu Zhao , Catalin Marinas , Anshuman Khandual , Yang Shi , "Huang, Ying" , Zi Yan , Luis Chamberlain , Itaru Kitayama , "Kirill A. Shutemov" , John Hubbard , David Rientjes , Vlastimil Babka , Hugh Dickins , Kefeng Wang Cc: Ryan Roberts , linux-mm@kvack.org, linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org Subject: [PATCH v7 00/10] Small-sized THP for anonymous memory Date: Wed, 15 Nov 2023 13:27:24 +0000 Message-Id: <20231115132734.931023-1-ryan.roberts@arm.com> X-Mailer: git-send-email 2.25.1 MIME-Version: 1.0 X-Rspamd-Queue-Id: 9892D1C0021 X-Rspam-User: X-Stat-Signature: efikrmmkmtz3j3p5yyw5yh9pb65qano3 X-Rspamd-Server: rspam03 X-HE-Tag: 1700054878-521014 X-HE-Meta: U2FsdGVkX1/0W1k65Es1F8PaEj0yqp7Qft0ngwrYt69AbKChkKWdJ2QR1lj6HoeTG/M0wN25nxpkEqVgbA8VqeNp6rKQ/PlQTcnfoAGPghOcebjeGfpS6lAaCp+mnOrcd0GYJmauUGyhE/ZplPxcgnGh4n+TZXSrOEj+s8rmFgQO+SIJkIwx6EANjPinJPRK/BRyf8hqcTTuDXf/gE6PCzsNICKmQZ8VwqmeEE7ZhGgo5NNeIs0iZS9M7ZV3Jeq09A1oIEsbD7w2mExSgDFyX6HBUJfYbVZuE9hgrGwQcgR7QfYFo1YtRtCGQgBnKc9nxhNq0D/TSLLC+X3KRgfxPgSgKfYwRQLqTTHeRiqjj96KsksvZEwABGncUpPhG8sKfbmdxFOWF8+h/J+04D+24SfX4ZyoXFJ3NVelKQIgIUXyUwTNrjvogOnyBnUucxoN4TCUIuVL+cEkTzHgAEsnBC9mrSy1F+TywLOPpkaqembNc5DpN1Ia9j3ee0kk/VvlKLgpaG8OtG3nd1x8jPqxdgJS2s+7Cm5D/P0R2LEFGo+Fg5BoytxB6ExF9upHdUz5/J3BZbH4o1yrs7XocCgQ48DsCc6MURv4UyuTnWhTTLeKU1XLQctC598kqBhqJajCJd94qWxlMAi1U08okcLhf/Su1UJa9GTfXlJJWCE5VmTlk6nfOAJ4UZg3usW0GNmiAG0+eIQojgnv4jFrp0Ysoiq4vqU1pXuv4Gp5U39TzQmWMn1PcZ5oWZlQrppXocH9Z5Zo6XylGjyhlIWrzG5INjYPl57Fb4Etrh9VK+CFdbL6VuO0VPPBwmic1uOJG7Ku9SdLhAtFtrvaqmQQUW1Gd4jait5GJKOJCcDxQTXm7PUOuPvp26yyl3OthBAMNP7/HkUEgVCGAAlFwAlb4bqCDT4Rce/WLvhloxf6Ue+PXE88E5Ewfq2NLtzmrflG8bmW2/vPy24OOwUY3wvlPwo dmhOuRp6 zCVuA0etNvk13pgtaIySSgtISBkuLWxib3NYgthNqOnHPWC72+5Ugq7NF4YGrB7Gv8Wk/yQLsGj6j4Xx6YtwaN4pjuu42ho+qecAEMcPTZmhEJ/ShE9FTUNhIovuM7eW7RBOH0hGcq+sjdVpe2LIzTixH29zJGhIiRHoTbD8hEdAlLdr96t1p/pYgjrc1w3ZZgEA7MJbnfF39T0fPAtQNiYz8Xz9kM6TxWUIA1aMYR61mCCnowa85Ix9BIRccm414NDgJRps9eSj5RzC6ONSVsvxjha21Bm5vF7XDXCvuUhFVi5zXulaxKT9y3A+v1DI4+och87TGgbKmEDkOW1kww7GcmzMHRY/nN6ytcxzPR+IHtRBlHd/vcJsEFaptqLaEEe+I12KoRGQHUj1UdhQf5SMIkRT8kuD4X10NmE8EGtyW24LAL7roKJgcoJS2jXo4zayg X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: List-Subscribe: List-Unsubscribe: Hi All, This is v7 of a series to implement small-sized THP for anonymous memory (previously called "large anonymous folios"). The objective of this is to improve performance by allocating larger chunks of memory during anonymous page faults: 1) Since SW (the kernel) is dealing with larger chunks of memory than base pages, there are efficiency savings to be had; fewer page faults, batched PTE and RMAP manipulation, reduced lru list, etc. In short, we reduce kernel overhead. This should benefit all architectures. 2) Since we are now mapping physically contiguous chunks of memory, we can take advantage of HW TLB compression techniques. A reduction in TLB pressure speeds up kernel and user space. arm64 systems have 2 mechanisms to coalesce TLB entries; "the contiguous bit" (architectural) and HPA (uarch). The major change in this revision is the migration to a new sysfs interface as recommended by David Hildenbrand - thanks to David for the suggestion! This interface is inspired by the existing per-hugepage-size sysfs interface used by hugetlb, provides full backwards compatibility with the existing PMD-size THP interface, and provides a base for future extensibility. See [7] for detailed discussion of the interface. By default, the existing behaviour (and performance) is maintained. The user must explicitly enable small-sized THP to see the performance benefit. The series has also become heavy with mm selftest changes: These all relate to enlightenment of cow and khugepaged tests to explicitly test with small-sized THP. This series is based on mm-unstable (60df8b4235f5). Prerequisites ============= Some work items identified as being prerequisites are listed on page 3 at [8]. The summary is: | item | status | |:------------------------------|:------------------------| | mlock | In mainline (v6.7) | | madvise | In mainline (v6.6) | | compaction | v1 posted [9] | | numa balancing | Investigated: see below | | user-triggered page migration | In mainline (v6.7) | | khugepaged collapse | In mainline (NOP) | On NUMA balancing, which currently ignores any PTE-mapped THPs it encounters, John Hubbard has investigated this and concluded that it is A) not clear at the moment what a better policy might be for PTE-mapped THP and B) questions whether this should really be considered a prerequisite given no regression is caused for the default "small-sized THP disabled" case, and there is no correctness issue when it is enabled - its just a potential for non-optimal performance. (John please do elaborate if I haven't captured this correctly!) If there are no disagreements about removing numa balancing from the list, then that just leaves compaction which is in review on list at the moment. I really would like to get this series (and its remaining comapction prerequisite) in for v6.8. I accept that it may be a bit optimistic at this point, but lets see where we get to with review? Testing ======= The series includes patches for mm selftests to enlighten the cow and khugepaged tests to explicitly test with small-order THP, in the same way that PMD-order THP is tested. The new tests all pass, and no regressions are observed in the mm selftest suite. I've also run my usual kernel compilation and java script benchmarks without any issues. Refer to my performance numbers posted with v6 [6]. (These are for small-sized THP only - they do not include the arm64 contpte follow-on series). John Hubbard at Nvidia has indicated dramatic 10x performance improvements for some workloads at [10]. (Observed using v6 of this series as well as the arm64 contpte series). Kefeng Wang at Huawei has also indicated he sees improvements at [11] although there are some latency regressions also. Changes since v6 [6] ==================== - Refactored vmf_pte_range_changed() to remove uffd special-case (suggested by JohnH) - Dropped accounting patch (#3 in v6) (suggested by DavidH) - Continue to account *PMD-sized* THP only for now - Can add more counters in future if needed - Page cache large folios haven't needed any new counters yet - Pivot to sysfs ABI proposed by DavidH - per-size directories in a similar shape to that used by hugetlb - Dropped "recommend" keyword patch (#6 in v6) (suggested by DavidH, Yu Zhou) - For now, users need to understand implicitly which sizes are beneficial to their HW/SW - Dropped arch_wants_pte_order() patch (#7 in v6) - No longer needed due to dropping patch "recommend" keyword patch - Enlightened khugepaged mm selftest to explicitly test with small-size THP - Scrubbed commit logs to use "small-sized THP" consistently (suggested by DavidH) Changes since v5 [5] ==================== - Added accounting for PTE-mapped THPs (patch 3) - Added runtime control mechanism via sysfs as extension to THP (patch 4) - Minor refactoring of alloc_anon_folio() to integrate with runtime controls - Stripped out hardcoded policy for allocation order; its now all user space controlled (although user space can request "recommend" which will configure the HW-preferred order) Changes since v4 [4] ==================== - Removed "arm64: mm: Override arch_wants_pte_order()" patch; arm64 now uses the default order-3 size. I have moved this patch over to the contpte series. - Added "mm: Allow deferred splitting of arbitrary large anon folios" back into series. I originally removed this at v2 to add to a separate series, but that series has transformed significantly and it no longer fits, so bringing it back here. - Reintroduced dependency on set_ptes(); Originally dropped this at v2, but set_ptes() is in mm-unstable now. - Updated policy for when to allocate LAF; only fallback to order-0 if MADV_NOHUGEPAGE is present or if THP disabled via prctl; no longer rely on sysfs's never/madvise/always knob. - Fallback to order-0 whenever uffd is armed for the vma, not just when uffd-wp is set on the pte. - alloc_anon_folio() now returns `struct folio *`, where errors are encoded with ERR_PTR(). The last 3 changes were proposed by Yu Zhao - thanks! Changes since v3 [3] ==================== - Renamed feature from FLEXIBLE_THP to LARGE_ANON_FOLIO. - Removed `flexthp_unhinted_max` boot parameter. Discussion concluded that a sysctl is preferable but we will wait until real workload needs it. - Fixed uninitialized `addr` on read fault path in do_anonymous_page(). - Added mm selftests for large anon folios in cow test suite. Changes since v2 [2] ==================== - Dropped commit "Allow deferred splitting of arbitrary large anon folios" - Huang, Ying suggested the "batch zap" work (which I dropped from this series after v1) is a prerequisite for merging FLXEIBLE_THP, so I've moved the deferred split patch to a separate series along with the batch zap changes. I plan to submit this series early next week. - Changed folio order fallback policy - We no longer iterate from preferred to 0 looking for acceptable policy - Instead we iterate through preferred, PAGE_ALLOC_COSTLY_ORDER and 0 only - Removed vma parameter from arch_wants_pte_order() - Added command line parameter `flexthp_unhinted_max` - clamps preferred order when vma hasn't explicitly opted-in to THP - Never allocate large folio for MADV_NOHUGEPAGE vma (or when THP is disabled for process or system). - Simplified implementation and integration with do_anonymous_page() - Removed dependency on set_ptes() Changes since v1 [1] ==================== - removed changes to arch-dependent vma_alloc_zeroed_movable_folio() - replaced with arch-independent alloc_anon_folio() - follows THP allocation approach - no longer retry with intermediate orders if allocation fails - fallback directly to order-0 - remove folio_add_new_anon_rmap_range() patch - instead add its new functionality to folio_add_new_anon_rmap() - remove batch-zap pte mappings optimization patch - remove enabler folio_remove_rmap_range() patch too - These offer real perf improvement so will submit separately - simplify Kconfig - single FLEXIBLE_THP option, which is independent of arch - depends on TRANSPARENT_HUGEPAGE - when enabled default to max anon folio size of 64K unless arch explicitly overrides - simplify changes to do_anonymous_page(): - no more retry loop [1] https://lore.kernel.org/linux-mm/20230626171430.3167004-1-ryan.roberts@arm.com/ [2] https://lore.kernel.org/linux-mm/20230703135330.1865927-1-ryan.roberts@arm.com/ [3] https://lore.kernel.org/linux-mm/20230714160407.4142030-1-ryan.roberts@arm.com/ [4] https://lore.kernel.org/linux-mm/20230726095146.2826796-1-ryan.roberts@arm.com/ [5] https://lore.kernel.org/linux-mm/20230810142942.3169679-1-ryan.roberts@arm.com/ [6] https://lore.kernel.org/linux-mm/20230929114421.3761121-1-ryan.roberts@arm.com/ [7] https://lore.kernel.org/linux-mm/6d89fdc9-ef55-d44e-bf12-fafff318aef8@redhat.com/ [8] https://drive.google.com/file/d/1GnfYFpr7_c1kA41liRUW5YtCb8Cj18Ud/view?usp=sharing&resourcekey=0-U1Mj3-RhLD1JV6EThpyPyA [9] https://lore.kernel.org/linux-mm/20231113170157.280181-1-zi.yan@sent.com/ [10] https://lore.kernel.org/linux-mm/c507308d-bdd4-5f9e-d4ff-e96e4520be85@nvidia.com/ [11] https://lore.kernel.org/linux-mm/479b3e2b-456d-46c1-9677-38f6c95a0be8@huawei.com/ Thanks, Ryan Ryan Roberts (10): mm: Allow deferred splitting of arbitrary anon large folios mm: Non-pmd-mappable, large folios for folio_add_new_anon_rmap() mm: thp: Introduce per-size thp sysfs interface mm: thp: Support allocation of anonymous small-sized THP selftests/mm/kugepaged: Restore thp settings at exit selftests/mm: Factor out thp settings management selftests/mm: Support small-sized THP interface in thp_settings selftests/mm/khugepaged: Enlighten for small-sized THP selftests/mm/cow: Generalize do_run_with_thp() helper selftests/mm/cow: Add tests for anonymous small-sized THP Documentation/admin-guide/mm/transhuge.rst | 74 +++- Documentation/filesystems/proc.rst | 6 +- fs/proc/task_mmu.c | 3 +- include/linux/huge_mm.h | 102 +++-- mm/huge_memory.c | 263 +++++++++++-- mm/khugepaged.c | 16 +- mm/memory.c | 112 +++++- mm/page_vma_mapped.c | 3 +- mm/rmap.c | 32 +- tools/testing/selftests/mm/Makefile | 4 +- tools/testing/selftests/mm/cow.c | 215 +++++++---- tools/testing/selftests/mm/khugepaged.c | 410 ++++----------------- tools/testing/selftests/mm/run_vmtests.sh | 2 + tools/testing/selftests/mm/thp_settings.c | 349 ++++++++++++++++++ tools/testing/selftests/mm/thp_settings.h | 80 ++++ 15 files changed, 1160 insertions(+), 511 deletions(-) create mode 100644 tools/testing/selftests/mm/thp_settings.c create mode 100644 tools/testing/selftests/mm/thp_settings.h --- 2.25.1