From patchwork Fri Mar 29 05:33:51 2024 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: "Ho-Ren (Jack) Chuang" X-Patchwork-Id: 13610116 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by smtp.lore.kernel.org (Postfix) with ESMTP id 0EDADCD1283 for ; Fri, 29 Mar 2024 05:34:05 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id 7A3846B0087; Fri, 29 Mar 2024 01:34:04 -0400 (EDT) Received: by kanga.kvack.org (Postfix, from userid 40) id 753E76B0095; Fri, 29 Mar 2024 01:34:04 -0400 (EDT) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id 5F5936B0096; Fri, 29 Mar 2024 01:34:04 -0400 (EDT) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0011.hostedemail.com [216.40.44.11]) by kanga.kvack.org (Postfix) with ESMTP id 402336B0087 for ; Fri, 29 Mar 2024 01:34:04 -0400 (EDT) Received: from smtpin02.hostedemail.com (a10.router.float.18 [10.200.18.1]) by unirelay09.hostedemail.com (Postfix) with ESMTP id F298281178 for ; Fri, 29 Mar 2024 05:34:03 +0000 (UTC) X-FDA: 81948960366.02.0C8F4A9 Received: from mail-qt1-f180.google.com (mail-qt1-f180.google.com [209.85.160.180]) by imf14.hostedemail.com (Postfix) with ESMTP id A870D100004 for ; Fri, 29 Mar 2024 05:34:01 +0000 (UTC) Authentication-Results: imf14.hostedemail.com; dkim=pass header.d=bytedance.com header.s=google header.b="b/Er/z3Y"; dmarc=pass (policy=quarantine) header.from=bytedance.com; spf=pass (imf14.hostedemail.com: domain of horenchuang@bytedance.com designates 209.85.160.180 as permitted sender) smtp.mailfrom=horenchuang@bytedance.com ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1711690442; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-transfer-encoding:content-transfer-encoding: in-reply-to:references:dkim-signature; bh=kzD9bp4lfPKjq881EZTr8X35H33UpZ0U2oJTGnZVwgs=; b=mjqtWfGFm6TM7tI6Ur+G0MRil4WlMUIvRh8x00UR1dGxxFhRfMK3NaNAIUmHtVXc2DWFYu e+eoLl0KpwPHuAT1TP/8nkvFJU9VQdB8e0ZpyvLNWqmvjwIWxgOsOuikkgAIcPfcgemYTQ v+cs81BVHMiu9uyzG3A9fZ6nT4hX7fA= ARC-Authentication-Results: i=1; imf14.hostedemail.com; dkim=pass header.d=bytedance.com header.s=google header.b="b/Er/z3Y"; dmarc=pass (policy=quarantine) header.from=bytedance.com; spf=pass (imf14.hostedemail.com: domain of horenchuang@bytedance.com designates 209.85.160.180 as permitted sender) smtp.mailfrom=horenchuang@bytedance.com ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1711690442; a=rsa-sha256; cv=none; b=33UIvIrjzfh/xGXLGhgrD0y370ZXKvkChs/XajLlZIIxFlUgKySTQ+AaLur09xH6eZOBXD 4VWrmV6HTQ8sehWaOROY374sKLGzBj/Iv1N2Na5oSpzYiog3V2GJudriB2bleNosjyZMI2 YojceHw4/jwmET+9pySbp5Wt+jm0uhI= Received: by mail-qt1-f180.google.com with SMTP id d75a77b69052e-430a0d6c876so10415381cf.3 for ; Thu, 28 Mar 2024 22:34:01 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=bytedance.com; s=google; t=1711690440; x=1712295240; darn=kvack.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to; bh=kzD9bp4lfPKjq881EZTr8X35H33UpZ0U2oJTGnZVwgs=; b=b/Er/z3YaRl9/CJSMxLYT2HPderrVA6XJjpzdQmXN1Mm734Tp1auKt+0i1kkGkdGrI fsUSlydgLuZ64HIisOoA4jNNeQ2WfbknhL5Kb755gBbj74jAC4NZ67GpNSpr0s7uqZg9 e3aZDeKDeQdnZnKqofL/pvkivEzTe65nUJNLAW/WDpqgkx76gtGeCnljqfwhpnvA3TKI hiIiRqB88s31xdUVE/w/pDZQhZSLg8rxByLLp/kQqiWzGwptNrK0327wjiUvYebBE1OA X6NA/JttWTdoxnhqDlsMZTtaNLtbKotxQLMuj/O2ZOf/QrpaGFa6NLgNrQwxSuHu9t73 VL7A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1711690440; x=1712295240; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to; bh=kzD9bp4lfPKjq881EZTr8X35H33UpZ0U2oJTGnZVwgs=; b=cc87thztC5huewEDJ/Aedt9VGpcWuclquR5u6q+aGKrj4/WBk4Dh3TvN14yYvpqh0N anMp3i1yNlByM4CZjTOT7P71AowBJNAuqG1w/IwbqtQptGt8VEa2aVnmXeI7P4GQ12he G/6xHLjLjbpZcqsJhvqgxJlh8nBijvO0TVcUYQ8WQqMneNbnoGEJkF9b+lEWDKr3QPVR TLrPPCHoya79IdX4lPeTD+heLNB4hj0yI86b+g0cxtlmZDa17BR+QjwtpGprl+N5EPyd Lys8O4qQKvNRcZ2J3yp5VJrqyaJC/m6i3FJMzOtRQAdY8qFeCncTy2W13xif1kFFnD7r MaFQ== X-Forwarded-Encrypted: i=1; AJvYcCUkN64XJDsmdRBZUwX8UchtOcjA43rXJvh+9xuDDDAkwAMLK1tvuwR1/iAzGx8px6ERJ6xkfDXaXnzesz6Xv0hl+1Y= X-Gm-Message-State: AOJu0YwafQxMtmeVZyj71R/TUKqkieQMg2F4jd7aEpqHJn+xXu2d3Jmv 42XJt74uPl0o815HrZRj1SgGW0U299s7CvVxHk8HgDMg43aSapH6Aq4ktivr3Ik= X-Google-Smtp-Source: AGHT+IEPCB5hxd4bBd3zlijyd5nzuuMHJJeNzrT0TX98AmIg0/OG0kaxC+8voIkqbdWtPzxYHjpC/A== X-Received: by 2002:a05:622a:40e:b0:432:c50a:3d65 with SMTP id n14-20020a05622a040e00b00432c50a3d65mr152679qtx.36.1711690440563; Thu, 28 Mar 2024 22:34:00 -0700 (PDT) Received: from n231-228-171.byted.org ([147.160.184.85]) by smtp.gmail.com with ESMTPSA id jd25-20020a05622a719900b00430bf59ebccsm1293700qtb.11.2024.03.28.22.33.59 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 28 Mar 2024 22:34:00 -0700 (PDT) From: "Ho-Ren (Jack) Chuang" To: "Huang, Ying" , "Gregory Price" , aneesh.kumar@linux.ibm.com, mhocko@suse.com, tj@kernel.org, john@jagalactic.com, "Eishan Mirakhur" , "Vinicius Tavares Petrucci" , "Ravis OpenSrc" , "Alistair Popple" , "Srinivasulu Thanneeru" , Dan Williams , Vishal Verma , Dave Jiang , Andrew Morton , nvdimm@lists.linux.dev, linux-cxl@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org Cc: "Ho-Ren (Jack) Chuang" , "Ho-Ren (Jack) Chuang" , "Ho-Ren (Jack) Chuang" , qemu-devel@nongnu.org Subject: [PATCH v9 0/2] Improved Memory Tier Creation for CPUless NUMA Nodes Date: Fri, 29 Mar 2024 05:33:51 +0000 Message-Id: <20240329053353.309557-1-horenchuang@bytedance.com> X-Mailer: git-send-email 2.20.1 MIME-Version: 1.0 X-Rspam-User: X-Rspamd-Server: rspam12 X-Rspamd-Queue-Id: A870D100004 X-Stat-Signature: u6ekstamusnuyombbucnjkh55r1x389o X-HE-Tag: 1711690441-603221 X-HE-Meta: U2FsdGVkX18vrO2Rrb3sPsyiaIomSmmNovW9C8NI1DK7ZfNYe8ifEeflNLN/b2uUW/30z0Vcrnv4j+HCO/2pGY8exQFUIXly/UPFVwvMucW2IbTfPUgycI3fdYbbLTgfSibsNI56hW4CUukRyPVxVNz7dc1t7Qrz8d7LBBHCB8+HoIcXPxbyxp9MtCRHuxzuX2pDjKVw9bDapMsgA7EtdbSOLqTq2FWx1VTNNhcnHEGQ2oK6f+izGvEyfngXwJH/WjUsxgoAbWTrDq1888vEDvQgGyUiS+DlmGP//fwCekXxeaOMgqx5SV8aj4TxzexOLVtxBqE5yvRyaxQmbslxmJJAQtCT4OE2CMyDGPEQQ/sKtBSxMwAtqUJPd6HvEDaCVfLvrSxhl0zAkk7txcvezNDahZsdFgOAzcXRuPXVHUvfgFnNIJEp/tpV8qvxOl4K+J6zVq6btyWrRVBWbuZ3fEZcQdlK1FL2u213ThO5flGdMhO9vX5oE27XJYsjz1VC3I5ycPlcuDMlRaTtNJQcOAWq/Hhzv7YApUR5JmX8q7VmwLgwOD3+SXQZmTw2yoDqequrngdjFqFvw3cG2ES/dog9J+aTwPjSk4daA2OiL2ICEMBQOiyW0ZQR2JgtuAb0vLxfbe6aood9diCk3ww9OuAbJdEFVoZB9RrOcoTyWU2OF4kbdcqu6ekCbuG/ogOCoCm+snhkeiEhoLV010pWbRqAN+IXjN2L9zDaC6J7QNvX2MaurdrcOPYrRhfjwZsX5jpa8JxqMfn4osjGzT49E18PDWJeIwArWjWAWcBlVJUcErQcEapJEeOdQCl6aLTyclt3/2RNrWu30OVysWZh67yUCvIt1QWVJFxZyOTtQR+MzJttguFJmu3OONGTQ/2EL6FCJM3yN48tEsfWJd1+Yu1wPXVDkFc8+PJ6sBFBY86crMwDFHGtokDT4PLh+T2w4atOlgGpFXDGUArcdxM xsmB8LPE SZE+cWNvhzqnI0+gTWpMn3bdGS/XZkB1YLW8o9l0yFxOrGdVk8oNNvXBRhUQF8yj51oHXQ9zYATVnJpCr/GddUi8jpH+SoU4UUXLXDEoh+Lce+qd5rdpg+ATfknHODb/Mnw6t7NI8rq4ToJDr//zWTmjhGkNhl3Jv1oxB07k5aHaj0vsDGEYU47T4Z0aWOnfrbUwluQc7Fpp6III= X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: List-Subscribe: List-Unsubscribe: When a memory device, such as CXL1.1 type3 memory, is emulated as normal memory (E820_TYPE_RAM), the memory device is indistinguishable from normal DRAM in terms of memory tiering with the current implementation. The current memory tiering assigns all detected normal memory nodes to the same DRAM tier. This results in normal memory devices with different attributions being unable to be assigned to the correct memory tier, leading to the inability to migrate pages between different types of memory. https://lore.kernel.org/linux-mm/PH0PR08MB7955E9F08CCB64F23963B5C3A860A@PH0PR08MB7955.namprd08.prod.outlook.com/T/ This patchset automatically resolves the issues. It delays the initialization of memory tiers for CPUless NUMA nodes until they obtain HMAT information and after all devices are initialized at boot time, eliminating the need for user intervention. If no HMAT is specified, it falls back to using `default_dram_type`. Example usecase: We have CXL memory on the host, and we create VMs with a new system memory device backed by host CXL memory. We inject CXL memory performance attributes through QEMU, and the guest now sees memory nodes with performance attributes in HMAT. With this change, we enable the guest kernel to construct the correct memory tiering for the memory nodes. -v9: * Address corner cases in `memory_tier_late_init`. Thank Ying's comments. -v8: * Fix email format * https://lore.kernel.org/lkml/20240329004815.195476-1-horenchuang@bytedance.com/T/#u -v7: * Add Reviewed-by: "Huang, Ying" -v6: Thanks to Ying's comments, * Move `default_dram_perf_lock` to the function's beginning for clarity * Fix double unlocking at v5 * https://lore.kernel.org/lkml/20240327072729.3381685-1-horenchuang@bytedance.com/T/#u -v5: Thanks to Ying's comments, * Add comments about what is protected by `default_dram_perf_lock` * Fix an uninitialized pointer mtype * Slightly shorten the time holding `default_dram_perf_lock` * Fix a deadlock bug in `mt_perf_to_adistance` * https://lore.kernel.org/lkml/20240327041646.3258110-1-horenchuang@bytedance.com/T/#u -v4: Thanks to Ying's comments, * Remove redundant code * Reorganize patches accordingly * https://lore.kernel.org/lkml/20240322070356.315922-1-horenchuang@bytedance.com/T/#u -v3: Thanks to Ying's comments, * Make the newly added code independent of HMAT * Upgrade set_node_memory_tier to support more cases * Put all non-driver-initialized memory types into default_memory_types instead of using hmat_memory_types * find_alloc_memory_type -> mt_find_alloc_memory_type * https://lore.kernel.org/lkml/20240320061041.3246828-1-horenchuang@bytedance.com/T/#u -v2: Thanks to Ying's comments, * Rewrite cover letter & patch description * Rename functions, don't use _hmat * Abstract common functions into find_alloc_memory_type() * Use the expected way to use set_node_memory_tier instead of modifying it * https://lore.kernel.org/lkml/20240312061729.1997111-1-horenchuang@bytedance.com/T/#u -v1: * https://lore.kernel.org/lkml/20240301082248.3456086-1-horenchuang@bytedance.com/T/#u Ho-Ren (Jack) Chuang (2): memory tier: dax/kmem: introduce an abstract layer for finding, allocating, and putting memory types memory tier: create CPUless memory tiers after obtaining HMAT info drivers/dax/kmem.c | 20 +----- include/linux/memory-tiers.h | 13 ++++ mm/memory-tiers.c | 125 ++++++++++++++++++++++++++++++----- 3 files changed, 124 insertions(+), 34 deletions(-)