From patchwork Mon Dec 11 15:56:13 2023 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: David Hildenbrand X-Patchwork-Id: 13487434 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by smtp.lore.kernel.org (Postfix) with ESMTP id 8A479C4167B for ; Mon, 11 Dec 2023 15:57:04 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id 233816B00F9; Mon, 11 Dec 2023 10:57:04 -0500 (EST) Received: by kanga.kvack.org (Postfix, from userid 40) id 1E3506B00FA; Mon, 11 Dec 2023 10:57:04 -0500 (EST) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id 0D2E86B00FB; Mon, 11 Dec 2023 10:57:04 -0500 (EST) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0014.hostedemail.com [216.40.44.14]) by kanga.kvack.org (Postfix) with ESMTP id F134E6B00F9 for ; Mon, 11 Dec 2023 10:57:03 -0500 (EST) Received: from smtpin08.hostedemail.com (a10.router.float.18 [10.200.18.1]) by unirelay02.hostedemail.com (Postfix) with ESMTP id BB95A1206F6 for ; Mon, 11 Dec 2023 15:57:03 +0000 (UTC) X-FDA: 81554991126.08.A82E1EF Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) by imf15.hostedemail.com (Postfix) with ESMTP id 11013A0006 for ; Mon, 11 Dec 2023 15:57:01 +0000 (UTC) Authentication-Results: imf15.hostedemail.com; dkim=pass header.d=redhat.com header.s=mimecast20190719 header.b=DQr6R278; spf=pass (imf15.hostedemail.com: domain of david@redhat.com designates 170.10.133.124 as permitted sender) smtp.mailfrom=david@redhat.com; dmarc=pass (policy=none) header.from=redhat.com ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1702310222; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-transfer-encoding:content-transfer-encoding: in-reply-to:references:dkim-signature; bh=gIYNTkt4lRg5dQRj6ZTd0KNVXXqltY+UcILlQ0Orr/Y=; b=KE3UalE0Z1naRpnkIrh9Pb17x6vi7UuW8gLPaL0zdIoNqx2y05L9aYLi/CuRGfS1Esvefs 0UClCktVUdjrFtiGu0cdfb56aG1Ow4Tys7zGq2OmRQixi5+aiE4vB3adXf0DxCw6Bmyd7h NXGPH1Ajjsb5RWSAtleQZBelt2LmPoY= ARC-Authentication-Results: i=1; imf15.hostedemail.com; dkim=pass header.d=redhat.com header.s=mimecast20190719 header.b=DQr6R278; spf=pass (imf15.hostedemail.com: domain of david@redhat.com designates 170.10.133.124 as permitted sender) smtp.mailfrom=david@redhat.com; dmarc=pass (policy=none) header.from=redhat.com ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1702310222; a=rsa-sha256; cv=none; b=ocbqYH8S5TLNpn9x2RRxsqquWACYdGFw4XkW10mH11B2hZaWaYK+rZfe+6FkdkB3w2Uufr 3yS7ys+f+KNA1ANkt4AsAzGsXzTadDDUXZrF6Phn4tpSxlV2JMEuvFtCZnQmagU7hYskz3 m+IZ0D0lXPMK1kSu8s1mIqpUzLjoTIY= DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1702310221; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-transfer-encoding:content-transfer-encoding; bh=gIYNTkt4lRg5dQRj6ZTd0KNVXXqltY+UcILlQ0Orr/Y=; b=DQr6R278h5b1+ulf157zwoBwRkowhyCrZka8t1gEMo4T8rgz6y0ghpxQPlUG477oOLYBYV +iv2p30GFrCOmLvAU4MCXYGWS+bW9V644EvcmMFheDI8qHmzDb047V273O5IprIX8x28d+ D8I9NqxDFpwRlGHre7G0xOfnInoO1OY= Received: from mimecast-mx02.redhat.com (mimecast-mx02.redhat.com [66.187.233.88]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-138-p_br7NpMPgawWyGWGkfahg-1; Mon, 11 Dec 2023 10:56:56 -0500 X-MC-Unique: p_br7NpMPgawWyGWGkfahg-1 Received: from smtp.corp.redhat.com (int-mx03.intmail.prod.int.rdu2.redhat.com [10.11.54.3]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mimecast-mx02.redhat.com (Postfix) with ESMTPS id EB317185A789; Mon, 11 Dec 2023 15:56:55 +0000 (UTC) Received: from t14s.redhat.com (unknown [10.39.192.166]) by smtp.corp.redhat.com (Postfix) with ESMTP id 7596E1121312; Mon, 11 Dec 2023 15:56:53 +0000 (UTC) From: David Hildenbrand To: linux-kernel@vger.kernel.org Cc: linux-mm@kvack.org, David Hildenbrand , Andrew Morton , "Matthew Wilcox (Oracle)" , Hugh Dickins , Ryan Roberts , Yin Fengwei , Mike Kravetz , Muchun Song , Peter Xu Subject: [PATCH v1 00/39] mm/rmap: interface overhaul Date: Mon, 11 Dec 2023 16:56:13 +0100 Message-ID: <20231211155652.131054-1-david@redhat.com> MIME-Version: 1.0 X-Scanned-By: MIMEDefang 3.4.1 on 10.11.54.3 X-Rspamd-Queue-Id: 11013A0006 X-Rspam-User: X-Stat-Signature: kbgra75r18dgmqz5j38gcmpcreyysxkm X-Rspamd-Server: rspam01 X-HE-Tag: 1702310221-583714 X-HE-Meta: U2FsdGVkX1+gh2BlTrP3F5RegJTAwVrTPlHJqPT7QnitlBIY2OKucIkjVMsBoV3lIJ/ByLB1GsUHA7dHGGPke84zsPRm1vAiNijuVzyxoW4US1iEO9h3cPzm14Uwy13nCnj9ZzTZmm9JZIQNny/5gZ1JR9IQ49NEJBWvYzNeqNXEJJwdGJTotSE8gi+yzlqO1y1RLRrtPkQott1tI8Nvt2owqTUmW2/MUKMS2loIehnTNE/mGt+umvy+XjkkDpo2u24FXbq9k1RRRwKFPllghaFkZ4fn+WKVUKGu/8mcGEmG4FJCGRy0x0wYufzerF6P3SZG3HfQG3jQaH7SmftukaHV9Z+TbgXOZ3a4iD2n73sY/iSLideacUASKYDpe4bukYkhGxfHnfCtApB6GPPmlJ7zkqDscfRSbDBln8kmAkb0qZ4lzZ2hegRnFQO3LEcNfxUTOjA5HqnLiyK4EJ0bwFndIKokjZyUUzPL0sK0wVXlg59P27/9ppKMY/4K1frVEWjY2uWI4slw8rQ9FQDJa6qwcauTxkcQucfxkUm/ZWM3leq+IdxrdwFi7dWOPW3W47a/TocTBK2gMvdkH/CI7zweOsIbILOppASen48NkNnoTMsNlVvIPW+bEtccFu9pBuWMvz6mK3vFjUsxhGzWyofwXqIyi+Z/TPNaAO9VwpJVtiAY3UDpetsssy3TSx2y76/4mW6RPMDWLdEzmLvnVILht4h4WCmdVTRaH8QKLFM6pZIRtpzVd2y0H6cGwW7arW7DsSQ1iZWWUuCvyDP6Fj3euxkMosXqfR8ywfN44LjYL8DgYUqEuQJQ7ozwJmaDNPB7ihj/58kfb3a0EJCneIn12ykcYnyHe8ULF5WlAdNhaonih97leWUOUMan5M3yt6zgUGQCvN2Xt4a++JEkwEr3sFCPhvFWAX6n6/z5HK8wy480aprGsCmehph6hLkwYAt9QlSC1ssSxlQ8Loh 1wJ1ojEZ z288vFghu8eGn6zn2sjRff3xk+Uwe705CdIFj3Ggz0QM2i6YWdrrbyWeFNq+LrY0w7qXJHsDNiJ5tkK+kVvm8gnmCXRgCkbYsgUMQMhAQh4kTlHM/8tnrTWtTg8Piwo7kDtqGxVRd28d4CrWa0g2ribI8SmAzRzPbxbVFB+VXy86ZWO8BnuVHRlHw5L+IVxB2Q6XKSNsuUdW1YAzR6C3wsznlvHl5JcMu2+Ywy95uPbw5tE57mqX8iDawo5InSlP0B+3Xw4Cf5rQAHsFm5YaTt8LlknPyhq38JkDpRajipLyBn01jZEUYt0zkb6Qo7240pktzPhpHLqDU1tg2OMQ/YRmBjonKsF32/6nOqV5XTVUnCgo5RDpWURVd43YRsyu1+H5qtRvVoF6G3H+o88XdkEnB1+ost+XGEhFBfLWjv2mS6/e2Kfed9O0IOGeeVA8Rs627Tt8+PKSFXoH2Ej9Ox/lvWsKlhAtxY4xh X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: List-Subscribe: List-Unsubscribe: This series overhauls the rmap interface, to get rid of the "bool compound" / RMAP_COMPOUND parameter with the goal of making the interface less error prone, more future proof, and more natural to extend to "batching". Also, this converts the interface to always consume folio+subpage, which speeds up operations on large folios. Further, this series adds PTE-batching variants for 4 rmap functions, whereby only folio_add_anon_rmap_ptes() is used for batching in this series when PTE-remapping a PMD-mapped THP. folio_remove_rmap_ptes(), folio_try_dup_anon_rmap_ptes() and folio_dup_file_rmap_ptes() will soon come in handy[1,2]. This series performs a lot of folio conversion along the way. Most of the added LOC in the diff are only due to documentation. As we're moving to a pte/pmd interface where we clearly express the mapping granularity we are dealing with, we first get the remainder of hugetlb out of the way, as it is special and expected to remain special: it treats everything as a "single logical PTE" and only currently allows entire mappings. Even if we'd ever support partial mappings, I strongly assume the interface and implementation will still differ heavily: hopefull we can avoid working on subpages/subpage mapcounts completely and only add a "count" parameter for them to enable batching. New (extended) hugetlb interface that operates on entire folio: * hugetlb_add_new_anon_rmap() -> Already existed * hugetlb_add_anon_rmap() -> Already existed * hugetlb_try_dup_anon_rmap() * hugetlb_try_share_anon_rmap() * hugetlb_add_file_rmap() * hugetlb_remove_rmap() New "ordinary" interface for small folios / THP:: * folio_add_new_anon_rmap() -> Already existed * folio_add_anon_rmap_[pte|ptes|pmd]() * folio_try_dup_anon_rmap_[pte|ptes|pmd]() * folio_try_share_anon_rmap_[pte|pmd]() * folio_add_file_rmap_[pte|ptes|pmd]() * folio_dup_file_rmap_[pte|ptes|pmd]() * folio_remove_rmap_[pte|ptes|pmd]() folio_add_new_anon_rmap() will always map at the largest granularity possible (currently, a single PMD to cover a PMD-sized THP). Could be extended if ever required. In the future, we might want "_pud" variants and eventually "_pmds" variants for batching. I ran some simple microbenchmarks on an Intel(R) Xeon(R) Silver 4210R: measuring munmap(), fork(), cow, MADV_DONTNEED on each PTE ... and PTE remapping PMD-mapped THPs on 1 GiB of memory. For small folios, there is barely a change (< 1%). For PTE-mapped THP: * PTE-remapping a PMD-mapped THP is more than 10% faster. * fork() is more than 4% faster. * MADV_DONTNEED is 2% faster * COW when writing only a single byte on a COW-shared PTE is 1% faster * munmap() barely changes (< 1%). [1] https://lkml.kernel.org/r/20230810103332.3062143-1-ryan.roberts@arm.com [2] https://lkml.kernel.org/r/20231204105440.61448-1-ryan.roberts@arm.com --- Based on current mm/mm-unstable. Compile-tested with/wihout THP on x86-64 and with defconig on a bunch more. Tested on x86-64. RFC -> v1: * Rebased on top of mm-unstable (containing mTHP) * Use switch()-case and _always_inline for helper functions * Fixed some (intermittend) compile issues and some smaller stuff * folio_try_dup_anon_rmap_[pte|ptes|pmd]() rewrite * Pass nr_pages consistently as "int" * Simplify sanity checks * Added RBs Cc: Andrew Morton Cc: "Matthew Wilcox (Oracle)" Cc: Hugh Dickins Cc: Ryan Roberts Cc: Yin Fengwei Cc: Mike Kravetz Cc: Muchun Song Cc: Peter Xu David Hildenbrand (39): mm/rmap: rename hugepage_add* to hugetlb_add* mm/rmap: introduce and use hugetlb_remove_rmap() mm/rmap: introduce and use hugetlb_add_file_rmap() mm/rmap: introduce and use hugetlb_try_dup_anon_rmap() mm/rmap: introduce and use hugetlb_try_share_anon_rmap() mm/rmap: add hugetlb sanity checks mm/rmap: convert folio_add_file_rmap_range() into folio_add_file_rmap_[pte|ptes|pmd]() mm/memory: page_add_file_rmap() -> folio_add_file_rmap_[pte|pmd]() mm/huge_memory: page_add_file_rmap() -> folio_add_file_rmap_pmd() mm/migrate: page_add_file_rmap() -> folio_add_file_rmap_pte() mm/userfaultfd: page_add_file_rmap() -> folio_add_file_rmap_pte() mm/rmap: remove page_add_file_rmap() mm/rmap: factor out adding folio mappings into __folio_add_rmap() mm/rmap: introduce folio_add_anon_rmap_[pte|ptes|pmd]() mm/huge_memory: batch rmap operations in __split_huge_pmd_locked() mm/huge_memory: page_add_anon_rmap() -> folio_add_anon_rmap_pmd() mm/migrate: page_add_anon_rmap() -> folio_add_anon_rmap_pte() mm/ksm: page_add_anon_rmap() -> folio_add_anon_rmap_pte() mm/swapfile: page_add_anon_rmap() -> folio_add_anon_rmap_pte() mm/memory: page_add_anon_rmap() -> folio_add_anon_rmap_pte() mm/rmap: remove page_add_anon_rmap() mm/rmap: remove RMAP_COMPOUND mm/rmap: introduce folio_remove_rmap_[pte|ptes|pmd]() kernel/events/uprobes: page_remove_rmap() -> folio_remove_rmap_pte() mm/huge_memory: page_remove_rmap() -> folio_remove_rmap_pmd() mm/khugepaged: page_remove_rmap() -> folio_remove_rmap_pte() mm/ksm: page_remove_rmap() -> folio_remove_rmap_pte() mm/memory: page_remove_rmap() -> folio_remove_rmap_pte() mm/migrate_device: page_remove_rmap() -> folio_remove_rmap_pte() mm/rmap: page_remove_rmap() -> folio_remove_rmap_pte() Documentation: stop referring to page_remove_rmap() mm/rmap: remove page_remove_rmap() mm/rmap: convert page_dup_file_rmap() to folio_dup_file_rmap_[pte|ptes|pmd]() mm/rmap: introduce folio_try_dup_anon_rmap_[pte|ptes|pmd]() mm/huge_memory: page_try_dup_anon_rmap() -> folio_try_dup_anon_rmap_pmd() mm/memory: page_try_dup_anon_rmap() -> folio_try_dup_anon_rmap_pte() mm/rmap: remove page_try_dup_anon_rmap() mm: convert page_try_share_anon_rmap() to folio_try_share_anon_rmap_[pte|pmd]() mm/rmap: rename COMPOUND_MAPPED to ENTIRELY_MAPPED Documentation/mm/transhuge.rst | 4 +- Documentation/mm/unevictable-lru.rst | 4 +- include/linux/mm.h | 6 +- include/linux/rmap.h | 398 +++++++++++++++++++----- kernel/events/uprobes.c | 2 +- mm/filemap.c | 10 +- mm/gup.c | 2 +- mm/huge_memory.c | 85 +++--- mm/hugetlb.c | 21 +- mm/internal.h | 12 +- mm/khugepaged.c | 17 +- mm/ksm.c | 15 +- mm/memory-failure.c | 4 +- mm/memory.c | 60 ++-- mm/migrate.c | 12 +- mm/migrate_device.c | 41 +-- mm/mmu_gather.c | 2 +- mm/rmap.c | 433 ++++++++++++++++----------- mm/swapfile.c | 2 +- mm/userfaultfd.c | 2 +- 20 files changed, 740 insertions(+), 392 deletions(-)