From patchwork Fri Sep 14 20:34:57 2018 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Yang Shi X-Patchwork-Id: 10601217 Return-Path: Received: from mail.wl.linuxfoundation.org (pdx-wl-mail.web.codeaurora.org [172.30.200.125]) by pdx-korg-patchwork-2.web.codeaurora.org (Postfix) with ESMTP id E253913AD for ; Fri, 14 Sep 2018 20:35:48 +0000 (UTC) Received: from mail.wl.linuxfoundation.org (localhost [127.0.0.1]) by mail.wl.linuxfoundation.org (Postfix) with ESMTP id C782B2A6C1 for ; Fri, 14 Sep 2018 20:35:48 +0000 (UTC) Received: by mail.wl.linuxfoundation.org (Postfix, from userid 486) id B96AF2BA53; Fri, 14 Sep 2018 20:35:48 +0000 (UTC) X-Spam-Checker-Version: SpamAssassin 3.3.1 (2010-03-16) on pdx-wl-mail.web.codeaurora.org X-Spam-Level: X-Spam-Status: No, score=-2.9 required=2.0 tests=BAYES_00,MAILING_LIST_MULTI, RCVD_IN_DNSWL_NONE,UNPARSEABLE_RELAY autolearn=ham version=3.3.1 Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by mail.wl.linuxfoundation.org (Postfix) with ESMTP id 7417A2A6C1 for ; Fri, 14 Sep 2018 20:35:47 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id 439B78E0005; Fri, 14 Sep 2018 16:35:46 -0400 (EDT) Delivered-To: linux-mm-outgoing@kvack.org Received: by kanga.kvack.org (Postfix, from userid 40) id 3EA258E0001; Fri, 14 Sep 2018 16:35:46 -0400 (EDT) X-Original-To: int-list-linux-mm@kvack.org X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id 2D8CA8E0005; Fri, 14 Sep 2018 16:35:46 -0400 (EDT) X-Original-To: linux-mm@kvack.org X-Delivered-To: linux-mm@kvack.org Received: from mail-pl1-f199.google.com (mail-pl1-f199.google.com [209.85.214.199]) by kanga.kvack.org (Postfix) with ESMTP id E0A238E0001 for ; Fri, 14 Sep 2018 16:35:45 -0400 (EDT) Received: by mail-pl1-f199.google.com with SMTP id k18-v6so4809203pls.12 for ; Fri, 14 Sep 2018 13:35:45 -0700 (PDT) X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20161025; h=x-original-authentication-results:x-gm-message-state:from:to:cc :subject:date:message-id:in-reply-to:references; bh=nSY/IrqhwC+U7DlP23Iq5/2IjFexFdx3dmUCEmqeqXU=; b=GFZCSLtZC5fT/x/V5TI3rvjV6D/3rMQ/Acp3/T4+T6gcGXz/pqoQ2lMOwQBEG1Ktj5 I3wLat4nyJBhuSV5w62lhXAI0VZyqn3RUonE+GDTxevbIxTtpGyRRPZhYbioDH1VHzw0 R4eCVxBP0CZjAIHPdV98vqBjLzskwgm4Le90/NMHdl9RSfZ0j5AXlgx/X6oweSC9UQLE 1TVCyr8wNfducz4KHZGMPgd/R5D+3T1x5cb2BGutvBjIJs1iI/0kiUyUikzzVTc+RR40 8a+C7GZ/KnhQ/s7zrDutyLlNzVbcQMlzqrDv04d0yn3yuL5YPYuhYW0XKzCsAwWZqMfF X80A== X-Original-Authentication-Results: mx.google.com; spf=pass (google.com: domain of yang.shi@linux.alibaba.com designates 47.88.44.37 as permitted sender) smtp.mailfrom=yang.shi@linux.alibaba.com; dmarc=pass (p=NONE sp=NONE dis=NONE) header.from=alibaba.com X-Gm-Message-State: APzg51B5livz3GUDTxJP3eQ8+1YCWzbw6haV/DwUiyNfgc4Mp5SOditG kyQM78XqljDsy399fSjsekPz6Qjaa3Vunzj0cRgUFPUVuaOmi2RlNPvRpbrnya6rJ/XV5WdF1Jt gEliZCsYsEqTCBjQymgaatZ47mEybrlbVViqli0W40+SG5uhDhOFelJY061UrBcOLdg== X-Received: by 2002:a17:902:d706:: with SMTP id w6-v6mr13917853ply.158.1536957345562; Fri, 14 Sep 2018 13:35:45 -0700 (PDT) X-Google-Smtp-Source: ANB0VdYI+jcKlfxNo4lMC8VxrNzMhStPfe2SrFPCSiMcfdGlIjRJCtsQ4onm5QHwkwSp7Ku1r/aE X-Received: by 2002:a17:902:d706:: with SMTP id w6-v6mr13917813ply.158.1536957344441; Fri, 14 Sep 2018 13:35:44 -0700 (PDT) ARC-Seal: i=1; a=rsa-sha256; t=1536957344; cv=none; d=google.com; s=arc-20160816; b=LqWNU5OQ1aCxi4PtrUZc9p+PMmETcagqo8Fl3sci44lDkE7NahJadtgS7Frd+XpEWK GCvmhYsJPCk03lLv3RVt11P5nWfR0fw8XqFuWaWF3PPo13QQ1BSOByMp5lcOMT+NN+gg eLWNrY7wgZB01a1sU7DTnzUCnRijsfmrNp9ITqGI9rtvMW+FQeNfV7TkM19qozg1ZlH6 WSVrD68UAa4hN7vxClOyyEFH/up3XKxnJoWcTBJ+B4k7D6grfI8A5PEZCuVHT0rTxItn W3e+gxp7QmUdIvvhe42BBWoyeW7kpNIJC00QW3eu7SsRXeJJUmozEPAkhMgKcCm7wzW9 suHw== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=references:in-reply-to:message-id:date:subject:cc:to:from; bh=nSY/IrqhwC+U7DlP23Iq5/2IjFexFdx3dmUCEmqeqXU=; b=xCYIeW7fkGwyKTXBTVjmMDG8h83fCyxWMGG0o1wliSY/72dVcvOuzlz5RFand8VKvH WOS6Xab9C4HvtY2ZQqEPmhYISpMpKpsqv9NZSAz/jOmrCu7yXQyasc3JveiLdONDM8+l 2Dqw63SS7X8u/N5hjLsbHCb95FTsL+oUhcWu1Rx/0uZLRlbfhq3bUX6fQ7gFsH136/Rs 6sqWISQLp99BgHZ3oENuCWOc94Dd3W31awXzKZGXEkO0ZF2rVvdHns0yccj9utzGxBgB TMPEm5+EMk8ZMmFCN4tVZjKPyhqvMH76z9e9jERrkn2hlH7I6D69SiiNIz0KBdYX/Mmz 1N9w== ARC-Authentication-Results: i=1; mx.google.com; spf=pass (google.com: domain of yang.shi@linux.alibaba.com designates 47.88.44.37 as permitted sender) smtp.mailfrom=yang.shi@linux.alibaba.com; dmarc=pass (p=NONE sp=NONE dis=NONE) header.from=alibaba.com Received: from out4437.biz.mail.alibaba.com (out4437.biz.mail.alibaba.com. [47.88.44.37]) by mx.google.com with ESMTPS id a9-v6si7635273pgf.380.2018.09.14.13.35.43 for (version=TLS1_2 cipher=ECDHE-RSA-AES128-GCM-SHA256 bits=128/128); Fri, 14 Sep 2018 13:35:44 -0700 (PDT) Received-SPF: pass (google.com: domain of yang.shi@linux.alibaba.com designates 47.88.44.37 as permitted sender) client-ip=47.88.44.37; Authentication-Results: mx.google.com; spf=pass (google.com: domain of yang.shi@linux.alibaba.com designates 47.88.44.37 as permitted sender) smtp.mailfrom=yang.shi@linux.alibaba.com; dmarc=pass (p=NONE sp=NONE dis=NONE) header.from=alibaba.com X-Alimail-AntiSpam: AC=PASS;BC=-1|-1;BR=01201311R941e4;CH=green;FP=0|-1|-1|-1|0|-1|-1|-1;HT=e01f04455;MF=yang.shi@linux.alibaba.com;NM=1;PH=DS;RN=12;SR=0;TI=SMTPD_---0T8i4EXv_1536957308; Received: from e19h19392.et15sqa.tbsite.net(mailfrom:yang.shi@linux.alibaba.com fp:SMTPD_---0T8i4EXv_1536957308) by smtp.aliyun-inc.com(127.0.0.1); Sat, 15 Sep 2018 04:35:15 +0800 From: Yang Shi To: mhocko@kernel.org, willy@infradead.org, ldufour@linux.vnet.ibm.com, vbabka@suse.cz, kirill@shutemov.name, akpm@linux-foundation.org Cc: dave.hansen@intel.com, oleg@redhat.com, srikar@linux.vnet.ibm.com, yang.shi@linux.alibaba.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org Subject: [RFC v10 PATCH 1/3] mm: mmap: zap pages with read mmap_sem in munmap Date: Sat, 15 Sep 2018 04:34:57 +0800 Message-Id: <1536957299-43536-2-git-send-email-yang.shi@linux.alibaba.com> X-Mailer: git-send-email 1.8.3.1 In-Reply-To: <1536957299-43536-1-git-send-email-yang.shi@linux.alibaba.com> References: <1536957299-43536-1-git-send-email-yang.shi@linux.alibaba.com> X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: X-Virus-Scanned: ClamAV using ClamSMTP When running some mmap/munmap scalability tests with large memory (i.e. > 300GB), the below hung task issue may happen occasionally. INFO: task ps:14018 blocked for more than 120 seconds. Tainted: G E 4.9.79-009.ali3000.alios7.x86_64 #1 "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message. ps D 0 14018 1 0x00000004 ffff885582f84000 ffff885e8682f000 ffff880972943000 ffff885ebf499bc0 ffff8828ee120000 ffffc900349bfca8 ffffffff817154d0 0000000000000040 00ffffff812f872a ffff885ebf499bc0 024000d000948300 ffff880972943000 Call Trace: [] ? __schedule+0x250/0x730 [] schedule+0x36/0x80 [] rwsem_down_read_failed+0xf0/0x150 [] call_rwsem_down_read_failed+0x18/0x30 [] down_read+0x20/0x40 [] proc_pid_cmdline_read+0xd9/0x4e0 [] ? do_filp_open+0xa5/0x100 [] __vfs_read+0x37/0x150 [] ? security_file_permission+0x9b/0xc0 [] vfs_read+0x96/0x130 [] SyS_read+0x55/0xc0 [] entry_SYSCALL_64_fastpath+0x1a/0xc5 It is because munmap holds mmap_sem exclusively from very beginning to all the way down to the end, and doesn't release it in the middle. When unmapping large mapping, it may take long time (take ~18 seconds to unmap 320GB mapping with every single page mapped on an idle machine). Zapping pages is the most time consuming part, according to the suggestion from Michal Hocko [1], zapping pages can be done with holding read mmap_sem, like what MADV_DONTNEED does. Then re-acquire write mmap_sem to cleanup vmas. But, some part may need write mmap_sem, for example, vma splitting. So, the design is as follows: acquire write mmap_sem lookup vmas (find and split vmas) deal with special mappings detach vmas downgrade_write zap pages free page tables release mmap_sem The vm events with read mmap_sem may come in during page zapping, but since vmas have been detached before, they, i.e. page fault, gup, etc, will not be able to find valid vma, then just return SIGSEGV or -EFAULT as expected. If the vma has VM_HUGETLB | VM_PFNMAP, they are considered as special mappings. They will be handled by without downgrading mmap_sem in this patch since they may update vm flags. But, with the "detach vmas first" approach, the vmas have been detached when vm flags are updated, so it sounds safe to update vm flags with read mmap_sem for this specific case. So, VM_HUGETLB and VM_PFNMAP will be handled by using the optimized path in the following separate patches for bisectable sake. Unmapping uprobe areas may need update mm flags (MMF_RECALC_UPROBES). However it is fine to have false-positive MMF_RECALC_UPROBES according to uprobes developer. With the "detach vmas first" approach we don't have to re-acquire mmap_sem again to clean up vmas to avoid race window which might get the address space changed since downgrade_write() doesn't release the lock to lead regression, which simply downgrades to read lock. And, since the lock acquire/release cost is managed to the minimum and almost as same as before, the optimization could be extended to any size of mapping without incurring significant penalty to small mappings. For the time being, just do this in munmap syscall path. Other vm_munmap() or do_munmap() call sites (i.e mmap, mremap, etc) remain intact due to some implementation difficulties since they acquire write mmap_sem from very beginning and hold it until the end, do_munmap() might be called in the middle. But, the optimized do_munmap would like to be called without mmap_sem held so that we can do the optimization. So, if we want to do the similar optimization for mmap/mremap path, I'm afraid we would have to redesign them. mremap might be called on very large area depending on the usecases, the optimization to it will be considered in the future. With the patches, exclusive mmap_sem hold time when munmap a 80GB address space on a machine with 32 cores of E5-2680 @ 2.70GHz dropped to us level from second. munmap_test-15002 [008] 594.380138: funcgraph_entry: | __vm_munmap() { munmap_test-15002 [008] 594.380146: funcgraph_entry: !2485684 us | unmap_region(); munmap_test-15002 [008] 596.865836: funcgraph_exit: !2485692 us | } Here the excution time of unmap_region() is used to evaluate the time of holding read mmap_sem, then the remaining time is used with holding exclusive lock. [1] https://lwn.net/Articles/753269/ Suggested-by: Michal Hocko Suggested-by: Kirill A. Shutemov Suggested-by: Matthew Wilcox Cc: Laurent Dufour Cc: Vlastimil Babka Cc: Andrew Morton Signed-off-by: Yang Shi Reviewed-by: Matthew Wilcox --- mm/mmap.c | 59 ++++++++++++++++++++++++++++++++++++++++++++++++----------- 1 file changed, 48 insertions(+), 11 deletions(-) diff --git a/mm/mmap.c b/mm/mmap.c index 5f2b2b1..2879b19 100644 --- a/mm/mmap.c +++ b/mm/mmap.c @@ -2687,8 +2687,8 @@ int split_vma(struct mm_struct *mm, struct vm_area_struct *vma, * work. This now handles partial unmappings. * Jeremy Fitzhardinge */ -int do_munmap(struct mm_struct *mm, unsigned long start, size_t len, - struct list_head *uf) +static int __do_munmap(struct mm_struct *mm, unsigned long start, size_t len, + struct list_head *uf, bool downgrade) { unsigned long end; struct vm_area_struct *vma, *prev, *last; @@ -2770,25 +2770,47 @@ int do_munmap(struct mm_struct *mm, unsigned long start, size_t len, mm->locked_vm -= vma_pages(tmp); munlock_vma_pages_all(tmp); } + + /* + * Unmapping vmas, which have VM_HUGETLB or VM_PFNMAP, + * need get done with write mmap_sem held since they may + * update vm_flags. + */ + if (downgrade && + (tmp->vm_flags & (VM_HUGETLB | VM_PFNMAP))) + downgrade = false; + tmp = tmp->vm_next; } } - /* - * Remove the vma's, and unmap the actual pages - */ + /* Detatch vmas from rbtree */ detach_vmas_to_be_unmapped(mm, vma, prev, end); - unmap_region(mm, vma, prev, start, end); + /* + * mpx unmap need to be handled with write mmap_sem. It is safe to + * deal with it before unmap_region(). + */ arch_unmap(mm, vma, start, end); + if (downgrade) + downgrade_write(&mm->mmap_sem); + + unmap_region(mm, vma, prev, start, end); + /* Fix up all other VM information */ remove_vma_list(mm, vma); - return 0; + return downgrade ? 1 : 0; } -int vm_munmap(unsigned long start, size_t len) +int do_munmap(struct mm_struct *mm, unsigned long start, size_t len, + struct list_head *uf) +{ + return __do_munmap(mm, start, len, uf, false); +} + +static int __vm_munmap(unsigned long start, size_t len, bool downgrade) { int ret; struct mm_struct *mm = current->mm; @@ -2797,17 +2819,32 @@ int vm_munmap(unsigned long start, size_t len) if (down_write_killable(&mm->mmap_sem)) return -EINTR; - ret = do_munmap(mm, start, len, &uf); - up_write(&mm->mmap_sem); + ret = __do_munmap(mm, start, len, &uf, downgrade); + /* + * Returning 1 indicates mmap_sem is down graded. + * But 1 is not legal return value of vm_munmap() and munmap(), reset + * it to 0 before return. + */ + if (ret == 1) { + up_read(&mm->mmap_sem); + ret = 0; + } else + up_write(&mm->mmap_sem); + userfaultfd_unmap_complete(mm, &uf); return ret; } + +int vm_munmap(unsigned long start, size_t len) +{ + return __vm_munmap(start, len, false); +} EXPORT_SYMBOL(vm_munmap); SYSCALL_DEFINE2(munmap, unsigned long, addr, size_t, len) { profile_munmap(addr); - return vm_munmap(addr, len); + return __vm_munmap(addr, len, true); } From patchwork Fri Sep 14 20:34:58 2018 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Yang Shi X-Patchwork-Id: 10601213 Return-Path: Received: from mail.wl.linuxfoundation.org (pdx-wl-mail.web.codeaurora.org [172.30.200.125]) by pdx-korg-patchwork-2.web.codeaurora.org (Postfix) with ESMTP id C6C45933 for ; Fri, 14 Sep 2018 20:35:34 +0000 (UTC) Received: from mail.wl.linuxfoundation.org (localhost [127.0.0.1]) by mail.wl.linuxfoundation.org (Postfix) with ESMTP id AE9582A6C1 for ; Fri, 14 Sep 2018 20:35:34 +0000 (UTC) Received: by mail.wl.linuxfoundation.org (Postfix, from userid 486) id A2A702BA53; Fri, 14 Sep 2018 20:35:34 +0000 (UTC) X-Spam-Checker-Version: SpamAssassin 3.3.1 (2010-03-16) on pdx-wl-mail.web.codeaurora.org X-Spam-Level: X-Spam-Status: No, score=-2.9 required=2.0 tests=BAYES_00,MAILING_LIST_MULTI, RCVD_IN_DNSWL_NONE,UNPARSEABLE_RELAY autolearn=ham version=3.3.1 Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by mail.wl.linuxfoundation.org (Postfix) with ESMTP id E4D882A6C1 for ; Fri, 14 Sep 2018 20:35:33 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id E4CEE8E0004; Fri, 14 Sep 2018 16:35:32 -0400 (EDT) Delivered-To: linux-mm-outgoing@kvack.org Received: by kanga.kvack.org (Postfix, from userid 40) id DFC318E0001; Fri, 14 Sep 2018 16:35:32 -0400 (EDT) X-Original-To: int-list-linux-mm@kvack.org X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id CEC2A8E0004; Fri, 14 Sep 2018 16:35:32 -0400 (EDT) X-Original-To: linux-mm@kvack.org X-Delivered-To: linux-mm@kvack.org Received: from mail-pl1-f198.google.com (mail-pl1-f198.google.com [209.85.214.198]) by kanga.kvack.org (Postfix) with ESMTP id 9E8FC8E0001 for ; Fri, 14 Sep 2018 16:35:32 -0400 (EDT) Received: by mail-pl1-f198.google.com with SMTP id b6-v6so4831765pls.16 for ; Fri, 14 Sep 2018 13:35:32 -0700 (PDT) X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20161025; h=x-original-authentication-results:x-gm-message-state:from:to:cc :subject:date:message-id:in-reply-to:references; bh=bUqtduUu3Gd8GjNLNmNtoVSTPH2DemyZc+Wdmdk61AM=; b=QOB/jfgBZExmFV9xITKFYHsy4K2T7fhSA/Blh+hw/G4mm84WiwPC4Ja/YeANB4ZKep w/xbWWL1lBHC46nOnWN29inuXvPP09D9QS7UdPi8x0grSMe6xHL9Euso9FKrautuzSST CXlzg4Z11e8wPdh8xQzrglARjlaSZkdItwII2aJ8mhAWFQa1My3V0vY4xozgPrr8kTc+ 2HoA8V9HzE1Uz0GFrbIrefjEIE+7S5xwwn1AjaX7v0ZxNgQyKo8xTKhj1Z+x21O5temL pgfSehgDFCnxTkSKQsGMlB5M98xcrGURMvb5xi1bSNBKyfThlfaBtK+5JOYYs6vHSX/R 2FQQ== X-Original-Authentication-Results: mx.google.com; spf=pass (google.com: domain of yang.shi@linux.alibaba.com designates 115.124.30.133 as permitted sender) smtp.mailfrom=yang.shi@linux.alibaba.com; dmarc=pass (p=NONE sp=NONE dis=NONE) header.from=alibaba.com X-Gm-Message-State: APzg51CACh/JYrGD4iyHWOrVly/crDnR6knqrwG7IiY0P7BuPnSzIB59 iTDi966POR+qmqwVHWkz8wRHVQ3eIFxZSc90pg8zUVOqgD2eSDTEO31GVTS5taMxTBh3Qsk2c1Z /MWFGxGBG3QEYBs7AbxpDbizOvp9DTUH60w2Es4E6qq92DU2BdoXaaJuKdNFjttVpRg== X-Received: by 2002:a63:f44d:: with SMTP id p13-v6mr13743132pgk.257.1536957332300; Fri, 14 Sep 2018 13:35:32 -0700 (PDT) X-Google-Smtp-Source: ANB0VdaS+moQ4uxqFu3CCFA0s8l3UMlCJWGb/4cgFf30zkmsxpruf4NUDbWh2RXdlw5oWA2xxEon X-Received: by 2002:a63:f44d:: with SMTP id p13-v6mr13743090pgk.257.1536957331042; Fri, 14 Sep 2018 13:35:31 -0700 (PDT) ARC-Seal: i=1; a=rsa-sha256; t=1536957331; cv=none; d=google.com; s=arc-20160816; b=e7lQzC1+wdC1yHD60U1NapG+DZqRVqGAkuMl9MD0eDAqC9GLTUdWnQqGvV1NmWvZ2F O2uFeluLaBwUwKL7/4BFKLjj2LpcP5WeMWuapNVM5McNX3eOQW+zEb/HkvVdwTxU3uqN 0u9R4InVcD3harmYkZYEg3Eyju8merckAqfXyW7y4aZl63X2N2OBozeH0vAQFd3+/rcH FDp2UCcO8+kV8wHPbxK/CL5Gmg63p72z3zZ+8LVoTr6uoUHgwAqtZUN5yjUiMAV4s8Ud GPaW5BZOs1KAE5tPjR11yDr4VKh/NxVMzA2vJFk23QEDu48qQdj74DRThdaaHFveZi3o lU2Q== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=references:in-reply-to:message-id:date:subject:cc:to:from; bh=bUqtduUu3Gd8GjNLNmNtoVSTPH2DemyZc+Wdmdk61AM=; b=JBPlHqhoFK5FtXa1fIFALhQpJMVnuEyFsuc7U9uJr4d0dmuJ96KrGlgCwSkXgAMgbS /tgR4SJKrENi2FDPyDiQHQknO7ETtqs+RuY0H3YUNMQpvdBnvf+ZVk7h1/DzvMDmDs/e PoSoOclwQwntIcJM7zgMe5RCKwLO8tI6v84yy8/WtLXpaGokx4zXSvYm45iKuHLthPg7 EmRo7jysp9VZAAQub1Cy9/uWnTCbOFJ7yUsmTbASsALXPUsWiysO0Qnv8m3e5QJI26qv Tg4UM3FotpXjje3sQ5Wz31s6VBsbdfxjfQ95o1sfabSJhC6ssFz1awr+ekpRZuE/Bqm8 SPTA== ARC-Authentication-Results: i=1; mx.google.com; spf=pass (google.com: domain of yang.shi@linux.alibaba.com designates 115.124.30.133 as permitted sender) smtp.mailfrom=yang.shi@linux.alibaba.com; dmarc=pass (p=NONE sp=NONE dis=NONE) header.from=alibaba.com Received: from out30-133.freemail.mail.aliyun.com (out30-133.freemail.mail.aliyun.com. [115.124.30.133]) by mx.google.com with ESMTPS id a9-v6si8184601pgj.224.2018.09.14.13.35.30 for (version=TLS1_2 cipher=ECDHE-RSA-AES128-GCM-SHA256 bits=128/128); Fri, 14 Sep 2018 13:35:31 -0700 (PDT) Received-SPF: pass (google.com: domain of yang.shi@linux.alibaba.com designates 115.124.30.133 as permitted sender) client-ip=115.124.30.133; Authentication-Results: mx.google.com; spf=pass (google.com: domain of yang.shi@linux.alibaba.com designates 115.124.30.133 as permitted sender) smtp.mailfrom=yang.shi@linux.alibaba.com; dmarc=pass (p=NONE sp=NONE dis=NONE) header.from=alibaba.com X-Alimail-AntiSpam: AC=PASS;BC=-1|-1;BR=01201311R131e4;CH=green;FP=0|-1|-1|-1|0|-1|-1|-1;HT=e01e07402;MF=yang.shi@linux.alibaba.com;NM=1;PH=DS;RN=12;SR=0;TI=SMTPD_---0T8i4EXv_1536957308; Received: from e19h19392.et15sqa.tbsite.net(mailfrom:yang.shi@linux.alibaba.com fp:SMTPD_---0T8i4EXv_1536957308) by smtp.aliyun-inc.com(127.0.0.1); Sat, 15 Sep 2018 04:35:15 +0800 From: Yang Shi To: mhocko@kernel.org, willy@infradead.org, ldufour@linux.vnet.ibm.com, vbabka@suse.cz, kirill@shutemov.name, akpm@linux-foundation.org Cc: dave.hansen@intel.com, oleg@redhat.com, srikar@linux.vnet.ibm.com, yang.shi@linux.alibaba.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org Subject: [RFC v10 PATCH 2/3] mm: unmap VM_HUGETLB mappings with optimized path Date: Sat, 15 Sep 2018 04:34:58 +0800 Message-Id: <1536957299-43536-3-git-send-email-yang.shi@linux.alibaba.com> X-Mailer: git-send-email 1.8.3.1 In-Reply-To: <1536957299-43536-1-git-send-email-yang.shi@linux.alibaba.com> References: <1536957299-43536-1-git-send-email-yang.shi@linux.alibaba.com> X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: X-Virus-Scanned: ClamAV using ClamSMTP When unmapping VM_HUGETLB mappings, vm flags need to be updated. Since the vmas have been detached, so it sounds safe to update vm flags with read mmap_sem. Cc: Michal Hocko Cc: Vlastimil Babka Signed-off-by: Yang Shi Reviewed-by: Matthew Wilcox --- mm/mmap.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/mm/mmap.c b/mm/mmap.c index 2879b19..991e066 100644 --- a/mm/mmap.c +++ b/mm/mmap.c @@ -2777,7 +2777,7 @@ static int __do_munmap(struct mm_struct *mm, unsigned long start, size_t len, * update vm_flags. */ if (downgrade && - (tmp->vm_flags & (VM_HUGETLB | VM_PFNMAP))) + (tmp->vm_flags & VM_PFNMAP)) downgrade = false; tmp = tmp->vm_next; From patchwork Fri Sep 14 20:34:59 2018 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Yang Shi X-Patchwork-Id: 10601215 Return-Path: Received: from mail.wl.linuxfoundation.org (pdx-wl-mail.web.codeaurora.org [172.30.200.125]) by pdx-korg-patchwork-2.web.codeaurora.org (Postfix) with ESMTP id D1E0A13AD for ; Fri, 14 Sep 2018 20:35:36 +0000 (UTC) Received: from mail.wl.linuxfoundation.org (localhost [127.0.0.1]) by mail.wl.linuxfoundation.org (Postfix) with ESMTP id B3B522A6C1 for ; Fri, 14 Sep 2018 20:35:36 +0000 (UTC) Received: by mail.wl.linuxfoundation.org (Postfix, from userid 486) id A7C7A2BA53; Fri, 14 Sep 2018 20:35:36 +0000 (UTC) X-Spam-Checker-Version: SpamAssassin 3.3.1 (2010-03-16) on pdx-wl-mail.web.codeaurora.org X-Spam-Level: X-Spam-Status: No, score=-2.9 required=2.0 tests=BAYES_00,MAILING_LIST_MULTI, RCVD_IN_DNSWL_NONE,UNPARSEABLE_RELAY autolearn=ham version=3.3.1 Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by mail.wl.linuxfoundation.org (Postfix) with ESMTP id 5190E2A6C1 for ; Fri, 14 Sep 2018 20:35:36 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id 2DA9B8E0003; Fri, 14 Sep 2018 16:35:33 -0400 (EDT) Delivered-To: linux-mm-outgoing@kvack.org Received: by kanga.kvack.org (Postfix, from userid 40) id 28B158E0001; Fri, 14 Sep 2018 16:35:33 -0400 (EDT) X-Original-To: int-list-linux-mm@kvack.org X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id 135F78E0005; Fri, 14 Sep 2018 16:35:33 -0400 (EDT) X-Original-To: linux-mm@kvack.org X-Delivered-To: linux-mm@kvack.org Received: from mail-pl1-f198.google.com (mail-pl1-f198.google.com [209.85.214.198]) by kanga.kvack.org (Postfix) with ESMTP id C695B8E0003 for ; Fri, 14 Sep 2018 16:35:32 -0400 (EDT) Received: by mail-pl1-f198.google.com with SMTP id b6-v6so4831769pls.16 for ; Fri, 14 Sep 2018 13:35:32 -0700 (PDT) X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20161025; h=x-original-authentication-results:x-gm-message-state:from:to:cc :subject:date:message-id:in-reply-to:references; bh=md78xmCAKo3IDwBaEhlInV8jFl8Cij6aMMzrkrh08fU=; b=HCWBkRpI2QZO1BPWyZKLHzizD9r+vY0SagnGO3z+Bn3jE5l6njhJ22L/a0cQmP+1ED fGG8TtKAHJi9ZP7IornUVQ5TkrEDDY7pTQGAqH9pVkC5mp2WHMojZEL6OZi0llv5znZe AvBG5YZMMRJWxLFnd6he7X8idDYd/hdEqJwSMKTmYHVvxOQXJjJfqKm42HS1GNLiOlwt 0/yayxqS1HPgkROxNC0YudPG6G33BBBLmhmJm0VR1hZEXa4CiqObancdZDYbY6NgeT4O upMt78IJfc5jX684wcumLQejrIMlTh/i2sB7hqKSPGL1B4yW9bM1G9i9pCqYClm9MdxR Qhog== X-Original-Authentication-Results: mx.google.com; spf=pass (google.com: domain of yang.shi@linux.alibaba.com designates 115.124.30.133 as permitted sender) smtp.mailfrom=yang.shi@linux.alibaba.com; dmarc=pass (p=NONE sp=NONE dis=NONE) header.from=alibaba.com X-Gm-Message-State: APzg51Cr1r+NKOh86McCwernA3Ms3f9vF/J4264Y3JMdN/qou6PHWEko RGyp4taKuIBDwv3optoqzZiMamj3CkWixNrO1fW3tV9AYGbH/n547jc6Y6Thj94G/V3dDJim6qd T/E4+zDaAD2jPm4fiXpDZfWxCclahpm2mNTdLlS7m40M181jF/zYQHb9/a161N0xfcg== X-Received: by 2002:aa7:831b:: with SMTP id t27-v6mr14236011pfm.81.1536957332494; Fri, 14 Sep 2018 13:35:32 -0700 (PDT) X-Google-Smtp-Source: ANB0VdYjLw6hRQ0qG+oAdDOhirJribY/55jrxGohljuLj8E14AZ4xm7sD4pw1MB6/ZYl64CC3GQP X-Received: by 2002:aa7:831b:: with SMTP id t27-v6mr14235957pfm.81.1536957331283; Fri, 14 Sep 2018 13:35:31 -0700 (PDT) ARC-Seal: i=1; a=rsa-sha256; t=1536957331; cv=none; d=google.com; s=arc-20160816; b=mbgIdqZIW0AGYNQTvy0yI0s7OywShdokIlZP/KFZIMvYtvmZyiL4PjDPifOBO6PSHp 4hYrYf2rg88KGOklrsUnVkSFzrE38XLImGEvswNX5BcBF7zNSj0y9Jp5dblDFMPChwgo Q3PE7nGLByI9E6cUc5IQpIwzXPX91jtNJ7ZV45T/ukI8x4a17eEmLvFEAPFpYk5bE5xN aW7Iyb/c2wkyzVkxBqFtlJlsrJxb2PItJ4FHguOZ3UYREYqSf9m6IMaIsSxT4GYfYGio z5a707Qe1lqdr8VeGYBVFv3P0I94ZI6xcg0M506CzsSmyaNf6jo6IqjjAlfcDtOBzjle T3KA== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=references:in-reply-to:message-id:date:subject:cc:to:from; bh=md78xmCAKo3IDwBaEhlInV8jFl8Cij6aMMzrkrh08fU=; b=Jxkq+LT05IxwAbYWlsT1wViiuSMH+quo3bM5mjZvUbdqfrVDr/xVEBxapeKun4iQnk PgARGOoW5v3diOQoTbj8Pwi9iOUr41rO85BiqNVQG20aIhhXt96iJbs6IxtRgki3kZma vNmMnrqL8tColgWAQMVAINpelVADq6kdDrGwcGBSAqs2byhdXFgzHSa73MU/UISpKxAO Z4xIATlPw+pX+8jP8EjFKiMNxtPvmsr4+w/nayLv9WouZWdyB1/2Bg5xTVxI9Ix3ML0f B1QzhvvbcKT6P0yF6BOqEUzOdoHK/0ZhREHhFZZX//DVcgCMrdj+I9dR6ZjVuHLV1l0H hsnQ== ARC-Authentication-Results: i=1; mx.google.com; spf=pass (google.com: domain of yang.shi@linux.alibaba.com designates 115.124.30.133 as permitted sender) smtp.mailfrom=yang.shi@linux.alibaba.com; dmarc=pass (p=NONE sp=NONE dis=NONE) header.from=alibaba.com Received: from out30-133.freemail.mail.aliyun.com (out30-133.freemail.mail.aliyun.com. [115.124.30.133]) by mx.google.com with ESMTPS id y5-v6si6912247pll.89.2018.09.14.13.35.30 for (version=TLS1_2 cipher=ECDHE-RSA-AES128-GCM-SHA256 bits=128/128); Fri, 14 Sep 2018 13:35:31 -0700 (PDT) Received-SPF: pass (google.com: domain of yang.shi@linux.alibaba.com designates 115.124.30.133 as permitted sender) client-ip=115.124.30.133; Authentication-Results: mx.google.com; spf=pass (google.com: domain of yang.shi@linux.alibaba.com designates 115.124.30.133 as permitted sender) smtp.mailfrom=yang.shi@linux.alibaba.com; dmarc=pass (p=NONE sp=NONE dis=NONE) header.from=alibaba.com X-Alimail-AntiSpam: AC=PASS;BC=-1|-1;BR=01201311R121e4;CH=green;FP=0|-1|-1|-1|0|-1|-1|-1;HT=e01f04427;MF=yang.shi@linux.alibaba.com;NM=1;PH=DS;RN=12;SR=0;TI=SMTPD_---0T8i4EXv_1536957308; Received: from e19h19392.et15sqa.tbsite.net(mailfrom:yang.shi@linux.alibaba.com fp:SMTPD_---0T8i4EXv_1536957308) by smtp.aliyun-inc.com(127.0.0.1); Sat, 15 Sep 2018 04:35:16 +0800 From: Yang Shi To: mhocko@kernel.org, willy@infradead.org, ldufour@linux.vnet.ibm.com, vbabka@suse.cz, kirill@shutemov.name, akpm@linux-foundation.org Cc: dave.hansen@intel.com, oleg@redhat.com, srikar@linux.vnet.ibm.com, yang.shi@linux.alibaba.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org Subject: [RFC v10 PATCH 3/3] mm: unmap VM_PFNMAP mappings with optimized path Date: Sat, 15 Sep 2018 04:34:59 +0800 Message-Id: <1536957299-43536-4-git-send-email-yang.shi@linux.alibaba.com> X-Mailer: git-send-email 1.8.3.1 In-Reply-To: <1536957299-43536-1-git-send-email-yang.shi@linux.alibaba.com> References: <1536957299-43536-1-git-send-email-yang.shi@linux.alibaba.com> X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: X-Virus-Scanned: ClamAV using ClamSMTP When unmapping VM_PFNMAP mappings, vm flags need to be updated. Since the vmas have been detached, so it sounds safe to update vm flags with read mmap_sem. Cc: Michal Hocko Cc: Vlastimil Babka Signed-off-by: Yang Shi Reviewed-by: Matthew Wilcox --- mm/mmap.c | 9 --------- 1 file changed, 9 deletions(-) diff --git a/mm/mmap.c b/mm/mmap.c index 991e066..04c1d3b 100644 --- a/mm/mmap.c +++ b/mm/mmap.c @@ -2771,15 +2771,6 @@ static int __do_munmap(struct mm_struct *mm, unsigned long start, size_t len, munlock_vma_pages_all(tmp); } - /* - * Unmapping vmas, which have VM_HUGETLB or VM_PFNMAP, - * need get done with write mmap_sem held since they may - * update vm_flags. - */ - if (downgrade && - (tmp->vm_flags & VM_PFNMAP)) - downgrade = false; - tmp = tmp->vm_next; } }