From patchwork Mon Jan 10 16:36:22 2022 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Jan Beulich X-Patchwork-Id: 12708964 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from lists.xenproject.org (lists.xenproject.org [192.237.175.120]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.lore.kernel.org (Postfix) with ESMTPS id EAF8CC433F5 for ; Mon, 10 Jan 2022 16:36:47 +0000 (UTC) Received: from list by lists.xenproject.org with outflank-mailman.255485.437824 (Exim 4.92) (envelope-from ) id 1n6xen-00067w-7P; Mon, 10 Jan 2022 16:36:29 +0000 X-Outflank-Mailman: Message body and most headers restored to incoming version Received: by outflank-mailman (output) from mailman id 255485.437824; Mon, 10 Jan 2022 16:36:29 +0000 Received: from localhost ([127.0.0.1] helo=lists.xenproject.org) by lists.xenproject.org with esmtp (Exim 4.92) (envelope-from ) id 1n6xen-00067p-41; Mon, 10 Jan 2022 16:36:29 +0000 Received: by outflank-mailman (input) for mailman id 255485; Mon, 10 Jan 2022 16:36:27 +0000 Received: from se1-gles-flk1-in.inumbo.com ([94.247.172.50] helo=se1-gles-flk1.inumbo.com) by lists.xenproject.org with esmtp (Exim 4.92) (envelope-from ) id 1n6xel-0004yp-BB for xen-devel@lists.xenproject.org; Mon, 10 Jan 2022 16:36:27 +0000 Received: from de-smtp-delivery-102.mimecast.com (de-smtp-delivery-102.mimecast.com [194.104.111.102]) by se1-gles-flk1.inumbo.com (Halon) with ESMTPS id 746e844d-7233-11ec-81c1-a30af7de8005; Mon, 10 Jan 2022 17:36:26 +0100 (CET) Received: from EUR05-VI1-obe.outbound.protection.outlook.com (mail-vi1eur05lp2170.outbound.protection.outlook.com [104.47.17.170]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id de-mta-16-x9AfBvtDNsu14kHz6D9I_w-2; Mon, 10 Jan 2022 17:36:24 +0100 Received: from VI1PR04MB5600.eurprd04.prod.outlook.com (2603:10a6:803:e7::16) by VE1PR04MB6477.eurprd04.prod.outlook.com (2603:10a6:803:11e::14) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.20.4867.9; Mon, 10 Jan 2022 16:36:23 +0000 Received: from VI1PR04MB5600.eurprd04.prod.outlook.com ([fe80::5951:a489:1cf0:19fe]) by VI1PR04MB5600.eurprd04.prod.outlook.com ([fe80::5951:a489:1cf0:19fe%6]) with mapi id 15.20.4867.011; Mon, 10 Jan 2022 16:36:23 +0000 X-BeenThere: xen-devel@lists.xenproject.org List-Id: Xen developer discussion List-Unsubscribe: , List-Post: List-Help: List-Subscribe: , Errors-To: xen-devel-bounces@lists.xenproject.org Precedence: list Sender: "Xen-devel" X-Inumbo-ID: 746e844d-7233-11ec-81c1-a30af7de8005 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=suse.com; s=mimecast20200619; t=1641832586; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=7JJks6iLXnqZbdVJXHnHne9F/os9YAY04CAdOZJgakY=; b=RuV+ANhIwcNhbL2dX9vsllUqL4+2K6WhPQ6uEn4g4Kzns7tAId4R6HCygEvPd9/uRamWeU li2P0QvQpfVSQ9MyP/3w4o2KMGE7WURYX3yD9UvuUsJ5o/WX6zfnRMxzZndch9VSrUMsx3 TheFKYV4XUGMVt1mX7jwTUKqQ988vPY= X-MC-Unique: x9AfBvtDNsu14kHz6D9I_w-2 ARC-Seal: i=1; a=rsa-sha256; s=arcselector9901; d=microsoft.com; cv=none; b=iHTRMtmmEE2hLOLG23RwPscowjFBQuH5dS2hfHm1QYgtBPTdtTPBImxZ1gceNzOv5oQrgZ22b3FwLVtOmcKqnwCTN/kSzwv2z8NEmk/Wh5etLv9G5HxEaeuwpTmrgVaij9dwG7zkedHAV9CuQG9N4Ae4mN+G4pKKAGFhJLIUC4AoAZd6b+wW21kEzJxcw4NfHSqy1jnBJ0MLA43VVZVnW+isTnTZuVhxkE+Q63Lu4aRjP7/8c4GuEvrds46yRjWdXi+2It3ApYXV9kKa4zmPWxEwPFtwYRdCEi7zOdZ99zH+YxIzvcY/hfKpsoyX9ArGSLGu/Y5N1ufjYZns2ONDnQ== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector9901; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=7JJks6iLXnqZbdVJXHnHne9F/os9YAY04CAdOZJgakY=; b=BRvapkZ/kMBNywaVBV17H5cMnXaz9vDnqiwMe9vVfwChD+sPE5d/f5qOnkRcDnM+alAsijXvLf5vVqDxgLx6pj5VSTLgyVfSmrpkh9OKRiJBW6lgwMxDzg1mDQhVGyAdP/p2enFRpf7LbWaTsBkmiS6OnBXOKE1PLXGuOIoBo2UolBxY2umJL2HOovOPyH2WJcDzAUNQLeXihUm+jFhr9vSf5jPoyby0W8bUxZolzrgE4byT72YO5B8lrpcWafOvpg2/qKcJaxIzMeL/uWVz8ECHOzCr8IwKAG4vmQc+3uIEtDkFnSq7gDdY8LzHPboaXKAmh4po5SDgwS+hoG4KZg== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass smtp.mailfrom=suse.com; dmarc=pass action=none header.from=suse.com; dkim=pass header.d=suse.com; arc=none Authentication-Results: dkim=none (message not signed) header.d=none;dmarc=none action=none header.from=suse.com; Message-ID: <807a48fe-3829-d976-75dc-1767d32fb0f4@suse.com> Date: Mon, 10 Jan 2022 17:36:22 +0100 User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:91.0) Gecko/20100101 Thunderbird/91.4.1 Subject: [PATCH v3 20/23] VT-d: free all-empty page tables Content-Language: en-US From: Jan Beulich To: "xen-devel@lists.xenproject.org" Cc: Andrew Cooper , Paul Durrant , =?utf-8?q?Roger_Pau_Monn=C3=A9?= , Kevin Tian References: <76cb9f26-e316-98a2-b1ba-e51e3d20f335@suse.com> In-Reply-To: <76cb9f26-e316-98a2-b1ba-e51e3d20f335@suse.com> X-ClientProxiedBy: FR0P281CA0064.DEUP281.PROD.OUTLOOK.COM (2603:10a6:d10:49::8) To VI1PR04MB5600.eurprd04.prod.outlook.com (2603:10a6:803:e7::16) MIME-Version: 1.0 X-MS-PublicTrafficType: Email X-MS-Office365-Filtering-Correlation-Id: 79597edb-0028-4dd3-1446-08d9d4575705 X-MS-TrafficTypeDiagnostic: VE1PR04MB6477:EE_ X-Microsoft-Antispam-PRVS: X-MS-Oob-TLC-OOBClassifiers: OLM:8882; X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0; X-Microsoft-Antispam-Message-Info: 8Uzz8Z54+phrzr25gIj42tCtUg1AEJcT+uHpkUKEMgh3eW/aUOLr4XSgz6QyS+d62XUWJob3/eXOFZvuu0VEMk4oZZDX0w/zDaPMINX0773C2/1vSaQgTgy26dS4QxKWv+9YsdVBZWhYS20lPzBZu8+REA6O4r9lsvIJNzDTA9TAM+1p50/AFaPmT0OqI0Khn3UM1qf2B5U6zE3fRpvstwvyrsdkcWqHtWX2DVKLgLw18GRTVBeVrGxfjjz0+Mt1l9Xn9ctc25RcZvWekCjHq2yWT4J/6RXA4Ys7EWEfTg05fGuhj9QGWVH2d0L531yBst4P/9tE6f7BaQVL3W7HV/afLEvnT7yY5IIRocIOV1miAiKhN1rytsLWR6EiPrsCHjfugZBGgLzksfrrYmKEJVrkTpI4IwgANcC3tua9yYG3DEhObeeSJ2JBJK64TUQZQLs8DNSYnSvWAdUXfzCm69yfUst0GWThwXx6X64DVJMZMZxngqE/GEz/HYw0NrUjhIp2HYzh3QEUlq8YkqFYZHPGO/DKe5GwOadbeauyJzPjkC7k5+wGvTNBkRyKGW1TzczuW2ZTfsiWFT63Lo/+tTF3LNwS4m7AFh2j1ROiDI5ohdvUH7/vbmuLp4zykZ995IXOBT3nKENbs/6jrDweXdwQJxZZHuofbj/6Qj8Xu6mXZdjwLoqyCTgg+/yqb7uISx/H0Gi1AV674J4WhtLahsM5G/QZQqwj4xYPhad6o4A= X-Forefront-Antispam-Report: CIP:255.255.255.255;CTRY:;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:VI1PR04MB5600.eurprd04.prod.outlook.com;PTR:;CAT:NONE;SFS:(366004)(8936002)(6916009)(86362001)(36756003)(8676002)(508600001)(6486002)(31686004)(186003)(26005)(2616005)(6512007)(66556008)(6506007)(66476007)(2906002)(66946007)(4326008)(31696002)(83380400001)(38100700002)(54906003)(5660300002)(316002)(43740500002)(45980500001);DIR:OUT;SFP:1101; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: =?utf-8?q?uuGWHwtOk/vihIzfwXTEhgf2MZxv?= =?utf-8?q?V1NankDc0IOeJZQo5HxeocM2uBqvvR2qOLWW3Ohp0WFOZ1M4DWqJq+AxttSod2smj?= =?utf-8?q?qFJjT/OT28O+SYwCedGGl0OQVp2iLoT8st4TGe9xD0I9io9aqU2mitddGW7dHPd3K?= =?utf-8?q?qSutRmPQkYds7itqBY/QD/Edg5kuiLwMiU94CX9nkSRNKoAFn3KElm99C+i4yMsAN?= =?utf-8?q?q3G15mEC93VTX58MrshYMRrzGEs77NWxcshQLoLtTcoLBjxAz5Pe7qIQx7DyQbRVt?= =?utf-8?q?/iOPYVTA7vpDxZBO0Kxyy6nKxujt2E99ks1oDSm18EMmQiGRtaGbi3N5yb2SpEasx?= =?utf-8?q?Ukol5aRa2lFZGlq+7mIPR4NbRAKX3+Q7qNx9iDqhZhXoTsyZAE69poXSy86f5s24r?= =?utf-8?q?6o925e2bwzGBcz1uygbgMmemfgoQY9gwkHiWwrqgFXiygjxT5LLzlf1RYtWf9ibwJ?= =?utf-8?q?+fks+izIF/8RUwsx1HKiQ4L/2Bp8yOEHcHyTU2iSHc/vrDYe79RwxqLU9nO3+n1dP?= =?utf-8?q?O0nHcCx8XWfU/5+UBwX0W8aQ+xPjQew+B5lqr3aCc7fbcOh4zPx/xRa9GeGJ05voK?= =?utf-8?q?UpS6lMVTvVuMjDQ+hofe+iDBzJmpt+eZJCA2W8wUrG18yYpi0RMd8LVBfiGiAzI3k?= =?utf-8?q?pzQnuc+utQA50sleMdphBXaEUETfJRow6hGbifJnob9Qbfd7/ZeLR5KU1EvAGLgGe?= =?utf-8?q?UN9KEzwXeMgQs7ilV1EpVwjS9N4Six6BggT0c/rrzvbVOltKq4O2hoFaO96yXOuym?= =?utf-8?q?zpzW9e8dq9q0n8h1YmIT8zRAbeW7Lt+tbpft8HUKEjDD8cfQ34qsAr6FCnrpzarXp?= =?utf-8?q?4P5vwQ94uvWu7YLmza1XKQg7DuNwwc1l2+Sq3bOc+Zq4isal4z/EUWkd4wfCaHV0I?= =?utf-8?q?KAx/FgAlN+izVH8ORHiTnToyPcHCvnHUwRd5DGk0goRr/wpnlyAmGSdLEWpqiN+Yi?= =?utf-8?q?8r+/xEiq0jgyTd4PGTwKMk+qsatcgIBMQYYtfnESCZ8ZR4QpgLfYS5AQDW3b4GHgp?= =?utf-8?q?V3ccksJp2hNYlURkowJZ8pXdymmOHxG/YMuVJUxxiz92Q/qvMwLrUSBhSv/OUdRLb?= =?utf-8?q?TCoiUc9aVNyIaZVpoemL9SeB3Elp6puoRBg1VMIE2CrP+gqxJFDYX9eKVc70NvQVY?= =?utf-8?q?Bor7PwIAOfdip7P+p+dwZXE8m2V3uNTbVe6hox98gp9vI4InZMyuu4Ph4UBXiLnK9?= =?utf-8?q?1z1OV03gD+uSi/2aa1mlv+JtHhFzOSNGOW98CGxbR3bbqfICtK6q8m2hGEcdFvzKp?= =?utf-8?q?GKqSeO6PQ/goa34jrcbNGiWR95+dw1o7yTohOrOayM+d8sEv8/C0Q1/1EoE8untNk?= =?utf-8?q?844fKah+pH/Skbwop9D3b++da8dTtwrb8+dpelSnuHSQKkX3SL9y5HsS/OiQc12Gl?= =?utf-8?q?24ySfexWg7vI8tFS92wOZqkAc663Lg0cTjq/etSFiXjV7ryg4wb+OH1/Ebn9nuZUq?= =?utf-8?q?9pBiYWRbAbT1/1SVy5c4tMDRqO8jVBgdF0w765FFb4PJMrpnoGsjbSaIgYWaOkBFh?= =?utf-8?q?CAeE+DswCJZrm4BessIESbzfB8NGPr2JZ3yvV480LHKQNt/0v6W2nwQ=3D?= X-OriginatorOrg: suse.com X-MS-Exchange-CrossTenant-Network-Message-Id: 79597edb-0028-4dd3-1446-08d9d4575705 X-MS-Exchange-CrossTenant-AuthSource: VI1PR04MB5600.eurprd04.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Internal X-MS-Exchange-CrossTenant-OriginalArrivalTime: 10 Jan 2022 16:36:23.7529 (UTC) X-MS-Exchange-CrossTenant-FromEntityHeader: Hosted X-MS-Exchange-CrossTenant-Id: f7a17af6-1c5c-4a36-aa8b-f5be247aa4ba X-MS-Exchange-CrossTenant-MailboxType: HOSTED X-MS-Exchange-CrossTenant-UserPrincipalName: 9RtBamhs2V1al6wFByJPRBYTUuBH5OWDHW8G5RiM8MV8708RCPu4hnE6zJTCoGB5mKlCo/0EeiBZQMPrmwLUfQ== X-MS-Exchange-Transport-CrossTenantHeadersStamped: VE1PR04MB6477 When a page table ends up with no present entries left, it can be replaced by a non-present entry at the next higher level. The page table itself can then be scheduled for freeing. Note that while its output isn't used there yet, pt_update_contig_markers() right away needs to be called in all places where entries get updated, not just the one where entries get cleared. Note further that while pt_update_contig_markers() updates perhaps several PTEs within the table, since these are changes to "avail" bits only I do not think that cache flushing would be needed afterwards. Such cache flushing (of entire pages, unless adding yet more logic to me more selective) would be quite noticable performance-wise (very prominent during Dom0 boot). Signed-off-by: Jan Beulich --- v3: Properly bound loop. Re-base over changes earlier in the series. v2: New. --- The hang during boot on my Latitude E6410 (see the respective code comment) was pretty close after iommu_enable_translation(). No errors, no watchdog would kick in, just sometimes the first few pixel lines of the next log message's (XEN) prefix would have made it out to the screen (and there's no serial there). It's been a lot of experimenting until I figured the workaround (which I consider ugly, but halfway acceptable). I've been trying hard to make sure the workaround wouldn't be masking a real issue, yet I'm still wary of it possibly doing so ... My best guess at this point is that on these old IOMMUs the ignored bits 52...61 aren't really ignored for present entries, but also aren't "reserved" enough to trigger faults. This guess is from having tried to set other bits in this range (unconditionally, and with the workaround here in place), which yielded the same behavior. --- a/xen/drivers/passthrough/vtd/iommu.c +++ b/xen/drivers/passthrough/vtd/iommu.c @@ -42,6 +42,9 @@ #include "vtd.h" #include "../ats.h" +#define CONTIG_MASK DMA_PTE_CONTIG_MASK +#include + /* dom_io is used as a sentinel for quarantined devices */ #define QUARANTINE_SKIP(d) ((d) == dom_io && !dom_iommu(d)->arch.vtd.pgd_maddr) @@ -452,6 +455,9 @@ static uint64_t addr_to_dma_page_maddr(s write_atomic(&pte->val, new_pte.val); iommu_sync_cache(pte, sizeof(struct dma_pte)); + pt_update_contig_markers(&parent->val, + address_level_offset(addr, level), + level, PTE_kind_table); } if ( --level == target ) @@ -879,9 +885,31 @@ static int dma_pte_clear_one(struct doma old = *pte; dma_clear_pte(*pte); + iommu_sync_cache(pte, sizeof(*pte)); + + while ( pt_update_contig_markers(&page->val, + address_level_offset(addr, level), + level, PTE_kind_null) && + ++level < min_pt_levels ) + { + struct page_info *pg = maddr_to_page(pg_maddr); + + unmap_vtd_domain_page(page); + + pg_maddr = addr_to_dma_page_maddr(domain, addr, level, flush_flags, + false); + BUG_ON(pg_maddr < PAGE_SIZE); + + page = map_vtd_domain_page(pg_maddr); + pte = &page[address_level_offset(addr, level)]; + dma_clear_pte(*pte); + iommu_sync_cache(pte, sizeof(*pte)); + + *flush_flags |= IOMMU_FLUSHF_all; + iommu_queue_free_pgtable(domain, pg); + } spin_unlock(&hd->arch.mapping_lock); - iommu_sync_cache(pte, sizeof(struct dma_pte)); unmap_vtd_domain_page(page); @@ -2037,8 +2065,21 @@ static int __must_check intel_iommu_map_ } *pte = new; - iommu_sync_cache(pte, sizeof(struct dma_pte)); + + /* + * While the (ab)use of PTE_kind_table here allows to save some work in + * the function, the main motivation for it is that it avoids a so far + * unexplained hang during boot (while preparing Dom0) on a Westmere + * based laptop. + */ + pt_update_contig_markers(&page->val, + address_level_offset(dfn_to_daddr(dfn), level), + level, + (hd->platform_ops->page_sizes & + (1UL << level_to_offset_bits(level + 1)) + ? PTE_kind_leaf : PTE_kind_table)); + spin_unlock(&hd->arch.mapping_lock); unmap_vtd_domain_page(page);