From patchwork Tue Oct 10 20:33:20 2023 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Taylor Blau X-Patchwork-Id: 13416036 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by smtp.lore.kernel.org (Postfix) with ESMTP id 578E8CD8CB6 for ; Tue, 10 Oct 2023 20:33:31 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S234279AbjJJUd3 (ORCPT ); Tue, 10 Oct 2023 16:33:29 -0400 Received: from lindbergh.monkeyblade.net ([23.128.96.19]:60576 "EHLO lindbergh.monkeyblade.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S229437AbjJJUd1 (ORCPT ); Tue, 10 Oct 2023 16:33:27 -0400 Received: from mail-qk1-x72d.google.com (mail-qk1-x72d.google.com [IPv6:2607:f8b0:4864:20::72d]) by lindbergh.monkeyblade.net (Postfix) with ESMTPS id 1C8E58E for ; Tue, 10 Oct 2023 13:33:23 -0700 (PDT) Received: by mail-qk1-x72d.google.com with SMTP id af79cd13be357-77574c5979fso388089085a.3 for ; Tue, 10 Oct 2023 13:33:23 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ttaylorr-com.20230601.gappssmtp.com; s=20230601; t=1696970002; x=1697574802; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=itRyMwkKKeI1echHpIQE/jBk1Zn2bEmG+N8VmsMr7h4=; b=HcQz9Z20gjjyuTcotSrBDUDMJrYO+LDHWcwqKEwrngmuDfLl/qdMpXbOul7NIpM+K9 9B6JkOB+GZ7cyx/DiK8rLdfkEF2wnjriFg19F3mv4uYxyrnsdcQRJtD99Sudnm8RsbKP DTh9OuaWq5WT6VsQwOXQr3onHY4QIic2eDK+/3ULQvec70lwxGZsDW81hkArN7vxm2N/ y3O/rgsK0H4ALeqH+z69JH2t6BewPy7X5a3W1TSnxbq/W52YL+oCNhmKQuiHVI++MVd1 7py5I8vLDeBQaKUH5hX4SsWT2+974ndePnlvtoHTRTf7dxXAT/DdGEuIgPzw+rWFp8fj ABUQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1696970002; x=1697574802; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=itRyMwkKKeI1echHpIQE/jBk1Zn2bEmG+N8VmsMr7h4=; b=vc3THQVZUTvhC4id81RoD3463HsK7ERLjs5RYJADr5ABsjJoIUubkYWpspiWrGtVHH zFE2U8YRGZDK8vInxC8ruvd5psnt1VNCJ05r094UjdvANPIb9OSPpeKvourBA9uf/jnR LDHzDa90aiLMOCE75HjrXSicueMMw9MQSmjs8QnXRQ7tu47lMG3DrH4mVHiuuXX6Y0MC FT02dwIfkyxEnv9G+p2WKuowjrlRart7zSJ9FxtUH7batLbchhMRttVQyBkBiVsfo6bL tA4a9ePdKfAHSsZ9pu8f97wtaHzYwsGIy5K+PbcNeLwjPs5v6MwXSVpm//Ypb2ATVRXV MzIw== X-Gm-Message-State: AOJu0YwFUw4qkcauv7aF9ldouC/8xtA+5NXNVROzQsUGenrqVYsibu+G WpvKol8tViGXPqQvqX1vFrXXXSRkHCV2dW8W1ttvqA== X-Google-Smtp-Source: AGHT+IHug1G+d3UNSu4TKft16mY0rNpFybgBdDLLgG/CVBIt3Yeib5n/n+c320XLC/ZVojOXrf89xQ== X-Received: by 2002:a05:620a:40d2:b0:775:93aa:ce27 with SMTP id g18-20020a05620a40d200b0077593aace27mr23201364qko.3.1696970002033; Tue, 10 Oct 2023 13:33:22 -0700 (PDT) Received: from localhost (104-178-186-189.lightspeed.milwwi.sbcglobal.net. [104.178.186.189]) by smtp.gmail.com with ESMTPSA id vr3-20020a05620a55a300b007756fe0bb17sm4615151qkn.19.2023.10.10.13.33.21 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 10 Oct 2023 13:33:21 -0700 (PDT) Date: Tue, 10 Oct 2023 16:33:20 -0400 From: Taylor Blau To: git@vger.kernel.org Cc: Jonathan Tan , Junio C Hamano , Jeff King , SZEDER =?utf-8?b?R8OhYm9y?= Subject: [PATCH v3 01/17] t/t4216-log-bloom.sh: harden `test_bloom_filters_not_used()` Message-ID: References: MIME-Version: 1.0 Content-Disposition: inline In-Reply-To: Precedence: bulk List-ID: X-Mailing-List: git@vger.kernel.org The existing implementation of test_bloom_filters_not_used() asserts that the Bloom filter sub-system has not been initialized at all, by checking for the absence of any data from it from trace2. In the following commit, it will become possible to load Bloom filters without using them (e.g., because `commitGraph.changedPathVersion` is incompatible with the hash version with which the commit-graph's Bloom filters were written). When this is the case, it's possible to initialize the Bloom filter sub-system, while still not using any Bloom filters. When this is the case, check that the data dump from the Bloom sub-system is all zeros, indicating that no filters were used. Signed-off-by: Taylor Blau --- t/t4216-log-bloom.sh | 14 +++++++++++++- 1 file changed, 13 insertions(+), 1 deletion(-) diff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh index fa9d32facf..487fc3d6b9 100755 --- a/t/t4216-log-bloom.sh +++ b/t/t4216-log-bloom.sh @@ -81,7 +81,19 @@ test_bloom_filters_used () { test_bloom_filters_not_used () { log_args=$1 setup "$log_args" && - ! grep -q "statistics:{\"filter_not_present\":" "$TRASH_DIRECTORY/trace.perf" && + + if grep -q "statistics:{\"filter_not_present\":" "$TRASH_DIRECTORY/trace.perf" + then + # if the Bloom filter system is initialized, ensure that no + # filters were used + data="statistics:{" + data="$data\"filter_not_present\":0," + data="$data\"maybe\":0," + data="$data\"definitely_not\":0," + data="$data\"false_positive\":0}" + + grep -q "$data" "$TRASH_DIRECTORY/trace.perf" + fi && test_cmp log_wo_bloom log_w_bloom } From patchwork Tue Oct 10 20:33:23 2023 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 8bit X-Patchwork-Submitter: Taylor Blau X-Patchwork-Id: 13416037 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by smtp.lore.kernel.org (Postfix) with ESMTP id 9790BCD8CB7 for ; Tue, 10 Oct 2023 20:33:33 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S234244AbjJJUdc (ORCPT ); Tue, 10 Oct 2023 16:33:32 -0400 Received: from lindbergh.monkeyblade.net ([23.128.96.19]:60590 "EHLO lindbergh.monkeyblade.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S231787AbjJJUd1 (ORCPT ); Tue, 10 Oct 2023 16:33:27 -0400 Received: from mail-qv1-xf2f.google.com (mail-qv1-xf2f.google.com [IPv6:2607:f8b0:4864:20::f2f]) by lindbergh.monkeyblade.net (Postfix) with ESMTPS id 3017391 for ; Tue, 10 Oct 2023 13:33:26 -0700 (PDT) Received: by mail-qv1-xf2f.google.com with SMTP id 6a1803df08f44-65b0557ec77so35571036d6.0 for ; Tue, 10 Oct 2023 13:33:26 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ttaylorr-com.20230601.gappssmtp.com; s=20230601; t=1696970005; x=1697574805; darn=vger.kernel.org; h=in-reply-to:content-transfer-encoding:content-disposition :mime-version:references:message-id:subject:cc:to:from:date:from:to :cc:subject:date:message-id:reply-to; bh=2hiQDzvzOH04ATF2Y09DUGzs8UCVCslAIdd4WYxjmYo=; b=u4LI34ed9Zcp3HtDYKY1D+1x0NSfOIgCbT9D6ebfHze9RJiVfQrjQ1Kig4VbcwlK9B H/dw5aFaNfPfVCtJf75nO3yeRqJl//apIrVI6xyqV17A6aXJ9PzuQzpVLri3Z2ccxJg6 q/Bi6R1DxAjbJ+oqvPxpMDBEfCCZLSkSHjzZdLNkvrx2jmtS46Yw+jz7w1kiyiylRyPp Fs59hFr9DygjjtiMTWHsYa7w1mJHgiYI1ap2ka8tCL8hspkeHFLJ0coUiKaxVk2qw8VH fhQGR52nirNl/L2H6mcVqFgWgtyv3KFRTIagDALqNIaYWa/cWQ6O/xjOLBltRO1ghRFP qJ4A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1696970005; x=1697574805; h=in-reply-to:content-transfer-encoding:content-disposition :mime-version:references:message-id:subject:cc:to:from:date :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to; bh=2hiQDzvzOH04ATF2Y09DUGzs8UCVCslAIdd4WYxjmYo=; b=GC+kntc8NY7tjTjDQ8Fl8hyS1xFB7dtWv8FMxImiE5GewdOiRgcrY5XKL7Ez9FKobb /kiCErHQfYimFMrIdyqR50j/jO5DJ+WP1oEh3Vytps/qQabvwI+UjqbYNvHvtSJ/RcyE /uUHEc5KMDX7N61VFRWB8thbXofCcqPOzpIHsavi5V33+7lKdRjqH4JVT9rfF8L+aQPA UKbjMcR+plvahUTyPShxC2hDFE19MJEiqVOiCjls8ylTNw+/ABvGdarcy7dbllkGIHq/ 6mSiW0FjzQfLwWRzWG/udH87fugquAOefbsBYcL3WB4To+KcWBOjlsGCFn4EqRiQVqq7 KEMQ== X-Gm-Message-State: AOJu0YySxVpQ80hi5x1UfLB1JwkoR+D/IIbKPgTF154le6tpTfBu0wqq jNqNJwPqCSlfgOWNdyz3muDWnfV1ZIKu4/gNuV4e0Q== X-Google-Smtp-Source: AGHT+IGratRQGuwe/vxjPJzHMuZ/W6xeu/34oQ2gme1bnFdlc4JMD8fIjkQJWKp4qQo18ENlcsRzAw== X-Received: by 2002:a0c:e3d1:0:b0:66d:2a3:2abe with SMTP id e17-20020a0ce3d1000000b0066d02a32abemr831138qvl.32.1696970005123; Tue, 10 Oct 2023 13:33:25 -0700 (PDT) Received: from localhost (104-178-186-189.lightspeed.milwwi.sbcglobal.net. [104.178.186.189]) by smtp.gmail.com with ESMTPSA id x15-20020a0ce0cf000000b0065b2e561c17sm5018774qvk.123.2023.10.10.13.33.24 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 10 Oct 2023 13:33:24 -0700 (PDT) Date: Tue, 10 Oct 2023 16:33:23 -0400 From: Taylor Blau To: git@vger.kernel.org Cc: Jonathan Tan , Junio C Hamano , Jeff King , SZEDER =?utf-8?b?R8OhYm9y?= Subject: [PATCH v3 02/17] revision.c: consult Bloom filters for root commits Message-ID: <7d0fa9354328799606f7e9a85fb5cedd06da466c.1696969994.git.me@ttaylorr.com> References: MIME-Version: 1.0 Content-Disposition: inline In-Reply-To: Precedence: bulk List-ID: X-Mailing-List: git@vger.kernel.org The commit-graph stores changed-path Bloom filters which represent the set of paths included in a tree-level diff between a commit's root tree and that of its parent. When a commit has no parents, the tree-diff is computed against that commit's root tree and the empty tree. In other words, every path in that commit's tree is stored in the Bloom filter (since they all appear in the diff). Consult these filters during pathspec-limited traversals in the function `rev_same_tree_as_empty()`. Doing so yields a performance improvement where we can avoid enumerating the full set of paths in a parentless commit's root tree when we know that the path(s) of interest were not listed in that commit's changed-path Bloom filter. Suggested-by: SZEDER Gábor Original-patch-by: Jonathan Tan Signed-off-by: Taylor Blau --- revision.c | 26 ++++++++++++++++++++++---- t/t4216-log-bloom.sh | 8 ++++++-- 2 files changed, 28 insertions(+), 6 deletions(-) diff --git a/revision.c b/revision.c index e789834dd1..dae569a547 100644 --- a/revision.c +++ b/revision.c @@ -834,17 +834,28 @@ static int rev_compare_tree(struct rev_info *revs, return tree_difference; } -static int rev_same_tree_as_empty(struct rev_info *revs, struct commit *commit) +static int rev_same_tree_as_empty(struct rev_info *revs, struct commit *commit, + int nth_parent) { struct tree *t1 = repo_get_commit_tree(the_repository, commit); + int bloom_ret = 1; if (!t1) return 0; + if (nth_parent == 1 && revs->bloom_keys_nr) { + bloom_ret = check_maybe_different_in_bloom_filter(revs, commit); + if (!bloom_ret) + return 1; + } + tree_difference = REV_TREE_SAME; revs->pruning.flags.has_changes = 0; diff_tree_oid(NULL, &t1->object.oid, "", &revs->pruning); + if (bloom_ret == 1 && tree_difference == REV_TREE_SAME) + count_bloom_filter_false_positive++; + return tree_difference == REV_TREE_SAME; } @@ -882,7 +893,7 @@ static int compact_treesame(struct rev_info *revs, struct commit *commit, unsign if (nth_parent != 0) die("compact_treesame %u", nth_parent); old_same = !!(commit->object.flags & TREESAME); - if (rev_same_tree_as_empty(revs, commit)) + if (rev_same_tree_as_empty(revs, commit, nth_parent)) commit->object.flags |= TREESAME; else commit->object.flags &= ~TREESAME; @@ -978,7 +989,14 @@ static void try_to_simplify_commit(struct rev_info *revs, struct commit *commit) return; if (!commit->parents) { - if (rev_same_tree_as_empty(revs, commit)) + /* + * Pretend as if we are comparing ourselves to the + * (non-existent) first parent of this commit object. Even + * though no such parent exists, its changed-path Bloom filter + * (if one exists) is relative to the empty tree, using Bloom + * filters is allowed here. + */ + if (rev_same_tree_as_empty(revs, commit, 1)) commit->object.flags |= TREESAME; return; } @@ -1059,7 +1077,7 @@ static void try_to_simplify_commit(struct rev_info *revs, struct commit *commit) case REV_TREE_NEW: if (revs->remove_empty_trees && - rev_same_tree_as_empty(revs, p)) { + rev_same_tree_as_empty(revs, p, nth_parent)) { /* We are adding all the specified * paths from this parent, so the * history beyond this parent is not diff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh index 487fc3d6b9..322640feeb 100755 --- a/t/t4216-log-bloom.sh +++ b/t/t4216-log-bloom.sh @@ -87,7 +87,11 @@ test_bloom_filters_not_used () { # if the Bloom filter system is initialized, ensure that no # filters were used data="statistics:{" - data="$data\"filter_not_present\":0," + # unusable filters (e.g., those computed with a + # different value of commitGraph.changedPathsVersion) + # are counted in the filter_not_present bucket, so any + # value is OK there. + data="$data\"filter_not_present\":[0-9][0-9]*," data="$data\"maybe\":0," data="$data\"definitely_not\":0," data="$data\"false_positive\":0}" @@ -174,7 +178,7 @@ test_expect_success 'setup - add commit-graph to the chain with Bloom filters' ' test_bloom_filters_used_when_some_filters_are_missing () { log_args=$1 - bloom_trace_prefix="statistics:{\"filter_not_present\":3,\"maybe\":6,\"definitely_not\":9" + bloom_trace_prefix="statistics:{\"filter_not_present\":3,\"maybe\":6,\"definitely_not\":10" setup "$log_args" && grep -q "$bloom_trace_prefix" "$TRASH_DIRECTORY/trace.perf" && test_cmp log_wo_bloom log_w_bloom From patchwork Tue Oct 10 20:33:26 2023 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 8bit X-Patchwork-Submitter: Taylor Blau X-Patchwork-Id: 13416040 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by smtp.lore.kernel.org (Postfix) with ESMTP id BB45ACD8CB4 for ; Tue, 10 Oct 2023 20:33:46 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S234410AbjJJUdj (ORCPT ); Tue, 10 Oct 2023 16:33:39 -0400 Received: from lindbergh.monkeyblade.net ([23.128.96.19]:60654 "EHLO lindbergh.monkeyblade.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S231787AbjJJUdd (ORCPT ); Tue, 10 Oct 2023 16:33:33 -0400 Received: from mail-qk1-x733.google.com (mail-qk1-x733.google.com [IPv6:2607:f8b0:4864:20::733]) by lindbergh.monkeyblade.net (Postfix) with ESMTPS id 4897D9E for ; Tue, 10 Oct 2023 13:33:29 -0700 (PDT) Received: by mail-qk1-x733.google.com with SMTP id af79cd13be357-7757f2d3956so23865385a.0 for ; Tue, 10 Oct 2023 13:33:29 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ttaylorr-com.20230601.gappssmtp.com; s=20230601; t=1696970008; x=1697574808; darn=vger.kernel.org; h=in-reply-to:content-transfer-encoding:content-disposition :mime-version:references:message-id:subject:cc:to:from:date:from:to :cc:subject:date:message-id:reply-to; bh=mUZYz73TCkmr9Uvn+COH52Jc2nmCrspZUfXeIu/zvWU=; b=wYWc8Zhvq/hsLumHrReEIHnPpzgKpW4UE7hGYiEAtXQQEJASNLS+IM+1TfsRWoRQ/e iuB9pHwXsD/E09AiuPCf26I+y8v4KHeFViQwjRCfOgOrCfTRBz2Yb+NrBZXvyfMikv1k dvrYOifOLCzIFKfQLg7lfK/IQbLLLx240aIq+qY2QkP1qHxXTBz8yICqHMQWeXEUnFEk zA+f8OUehfp9JTCHKjiTfNrmXs8CKCx8dew5dIKWYMeR1y1GQkXQPmnGtEDd+DABtYQI y8EKLKC7RgEufrS/O73vgrIiDTbRzGUxnDzK1eOb8zyKtVqbI72nKTkAo0YEDBbyLdCu NfjA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1696970008; x=1697574808; h=in-reply-to:content-transfer-encoding:content-disposition :mime-version:references:message-id:subject:cc:to:from:date :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to; bh=mUZYz73TCkmr9Uvn+COH52Jc2nmCrspZUfXeIu/zvWU=; b=L6U/oMet6Bly2LBlnV6h8EwQlEh/K8sSOoba/Mg3jNLEaQDh1JFfOj5bVN0ca+znF+ CSxdRSlRd+7snK7B57yXCaEv7ebFjS8H3kod7CWQK+MV0L3w3XOAqKj5CVnR7lZ/gT2U 1pmovSEH93X0566F4GdaLRZoeK4PSTJkexRg/FL42gewpN+PeCB3TyxbKLxzg58x3SBJ VASPxyYF4e/NtWgF5pRmh4zOKoXNFDetisPWcfyubjVJXa3s7PDv0TQ2ICvd46f6o/zt 99VPH2QYQX3KEF48zz+rpzMnDf8GJtpP1cNwujwOHvxXl0tGDDSqLU99nrFIf5DOWxiy qm2A== X-Gm-Message-State: AOJu0YzsCvhGj2y+gfyTGH2HK8AbNq48vGAYn9kKK+GuCFHZDhvHmOhE s9yJZIGiguQt4Vpm4pcJk0TDXZT+aMjBPUoz9x0FWw== X-Google-Smtp-Source: AGHT+IFXnfj/j7f2NUxDivvNJb5Qi2+9g1/U8+M5jkNhUKElrDV9o3o9oA8PNA8fq3JJPUul1N/ovQ== X-Received: by 2002:a05:620a:454f:b0:774:19a4:117a with SMTP id u15-20020a05620a454f00b0077419a4117amr24134408qkp.19.1696970008185; Tue, 10 Oct 2023 13:33:28 -0700 (PDT) Received: from localhost (104-178-186-189.lightspeed.milwwi.sbcglobal.net. [104.178.186.189]) by smtp.gmail.com with ESMTPSA id o24-20020a05620a111800b0075ca4cd03d4sm4618697qkk.64.2023.10.10.13.33.27 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 10 Oct 2023 13:33:27 -0700 (PDT) Date: Tue, 10 Oct 2023 16:33:26 -0400 From: Taylor Blau To: git@vger.kernel.org Cc: Jonathan Tan , Junio C Hamano , Jeff King , SZEDER =?utf-8?b?R8OhYm9y?= Subject: [PATCH v3 03/17] commit-graph: ensure Bloom filters are read with consistent settings Message-ID: <2ecc0a2d58432b149d73a3e2abfa948eb1f0aa0b.1696969994.git.me@ttaylorr.com> References: MIME-Version: 1.0 Content-Disposition: inline In-Reply-To: Precedence: bulk List-ID: X-Mailing-List: git@vger.kernel.org The changed-path Bloom filter mechanism is parameterized by a couple of variables, notably the number of bits per hash (typically "m" in Bloom filter literature) and the number of hashes themselves (typically "k"). It is critically important that filters are read with the Bloom filter settings that they were written with. Failing to do so would mean that each query is liable to compute different fingerprints, meaning that the filter itself could return a false negative. This goes against a basic assumption of using Bloom filters (that they may return false positives, but never false negatives) and can lead to incorrect results. We have some existing logic to carry forward existing Bloom filter settings from one layer to the next. In `write_commit_graph()`, we have something like: if (!(flags & COMMIT_GRAPH_NO_WRITE_BLOOM_FILTERS)) { struct commit_graph *g = ctx->r->objects->commit_graph; /* We have changed-paths already. Keep them in the next graph */ if (g && g->chunk_bloom_data) { ctx->changed_paths = 1; ctx->bloom_settings = g->bloom_filter_settings; } } , which drags forward Bloom filter settings across adjacent layers. This doesn't quite address all cases, however, since it is possible for intermediate layers to contain no Bloom filters at all. For example, suppose we have two layers in a commit-graph chain, say, {G1, G2}. If G1 contains Bloom filters, but G2 doesn't, a new G3 (whose base graph is G2) may be written with arbitrary Bloom filter settings, because we only check the immediately adjacent layer's settings for compatibility. This behavior has existed since the introduction of changed-path Bloom filters. But in practice, this is not such a big deal, since the only way up until this point to modify the Bloom filter settings at write time is with the undocumented environment variables: - GIT_TEST_BLOOM_SETTINGS_BITS_PER_ENTRY - GIT_TEST_BLOOM_SETTINGS_NUM_HASHES - GIT_TEST_BLOOM_SETTINGS_MAX_CHANGED_PATHS (it is still possible to tweak MAX_CHANGED_PATHS between layers, but this does not affect reads, so is allowed to differ across multiple graph layers). But in future commits, we will introduce another parameter to change the hash algorithm used to compute Bloom fingerprints itself. This will be exposed via a configuration setting, making this foot-gun easier to use. To prevent this potential issue, validate that all layers of a split commit-graph have compatible settings with the newest layer which contains Bloom filters. Reported-by: SZEDER Gábor Original-test-by: SZEDER Gábor Signed-off-by: Taylor Blau --- commit-graph.c | 25 +++++++++++++++++ t/t4216-log-bloom.sh | 64 ++++++++++++++++++++++++++++++++++++++++++++ 2 files changed, 89 insertions(+) diff --git a/commit-graph.c b/commit-graph.c index 1a56efcf69..ae0902f7f4 100644 --- a/commit-graph.c +++ b/commit-graph.c @@ -498,6 +498,30 @@ static int validate_mixed_generation_chain(struct commit_graph *g) return 0; } +static void validate_mixed_bloom_settings(struct commit_graph *g) +{ + struct bloom_filter_settings *settings = NULL; + for (; g; g = g->base_graph) { + if (!g->bloom_filter_settings) + continue; + if (!settings) { + settings = g->bloom_filter_settings; + continue; + } + + if (g->bloom_filter_settings->bits_per_entry != settings->bits_per_entry || + g->bloom_filter_settings->num_hashes != settings->num_hashes) { + g->chunk_bloom_indexes = NULL; + g->chunk_bloom_data = NULL; + FREE_AND_NULL(g->bloom_filter_settings); + + warning(_("disabling Bloom filters for commit-graph " + "layer '%s' due to incompatible settings"), + oid_to_hex(&g->oid)); + } + } +} + static int add_graph_to_chain(struct commit_graph *g, struct commit_graph *chain, struct object_id *oids, @@ -614,6 +638,7 @@ struct commit_graph *load_commit_graph_chain_fd_st(struct repository *r, } validate_mixed_generation_chain(graph_chain); + validate_mixed_bloom_settings(graph_chain); free(oids); fclose(fp); diff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh index 322640feeb..f49a8f2fbf 100755 --- a/t/t4216-log-bloom.sh +++ b/t/t4216-log-bloom.sh @@ -420,4 +420,68 @@ test_expect_success 'Bloom generation backfills empty commits' ' ) ' +graph=.git/objects/info/commit-graph +graphdir=.git/objects/info/commit-graphs +chain=$graphdir/commit-graph-chain + +test_expect_success 'setup for mixed Bloom setting tests' ' + repo=mixed-bloom-settings && + + git init $repo && + for i in one two three + do + test_commit -C $repo $i file || return 1 + done +' + +test_expect_success 'split' ' + # Compute Bloom filters with "unusual" settings. + git -C $repo rev-parse one >in && + GIT_TEST_BLOOM_SETTINGS_NUM_HASHES=3 git -C $repo commit-graph write \ + --stdin-commits --changed-paths --split in && + git -C $repo commit-graph write --stdin-commits --no-changed-paths \ + --split=no-merge in && + git -C $repo commit-graph write --stdin-commits --changed-paths \ + --split=no-merge expect 2>err && + git -C $repo log --oneline --no-decorate -- file >actual 2>err && + test_cmp expect actual && + grep "disabling Bloom filters for commit-graph layer .$layer." err +' + +test_expect_success 'merge graph layers with incompatible Bloom settings' ' + # Ensure that incompatible Bloom filters are ignored when + # generating new layers. + git -C $repo commit-graph write --reachable --changed-paths 2>err && + grep "disabling Bloom filters for commit-graph layer .$layer." err && + + test_path_is_file $repo/$graph && + test_dir_is_empty $repo/$graphdir && + + # ...and merging existing ones. + git -C $repo -c core.commitGraph=false log --oneline --no-decorate -- file \ + >expect 2>err && + GIT_TRACE2_PERF="$(pwd)/trace.perf" \ + git -C $repo log --oneline --no-decorate -- file >actual 2>err && + + test_cmp expect actual && cat err && + grep "statistics:{\"filter_not_present\":0" trace.perf && + ! grep "disabling Bloom filters" err +' + test_done From patchwork Tue Oct 10 20:33:30 2023 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Taylor Blau X-Patchwork-Id: 13416038 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by smtp.lore.kernel.org (Postfix) with ESMTP id CE89ECD8CB4 for ; Tue, 10 Oct 2023 20:33:43 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S234284AbjJJUdm (ORCPT ); Tue, 10 Oct 2023 16:33:42 -0400 Received: from lindbergh.monkeyblade.net ([23.128.96.19]:60690 "EHLO lindbergh.monkeyblade.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S234393AbjJJUde (ORCPT ); Tue, 10 Oct 2023 16:33:34 -0400 Received: from mail-qv1-xf36.google.com (mail-qv1-xf36.google.com [IPv6:2607:f8b0:4864:20::f36]) by lindbergh.monkeyblade.net (Postfix) with ESMTPS id 6E0F594 for ; Tue, 10 Oct 2023 13:33:32 -0700 (PDT) Received: by mail-qv1-xf36.google.com with SMTP id 6a1803df08f44-65b07651b97so37442076d6.1 for ; Tue, 10 Oct 2023 13:33:32 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ttaylorr-com.20230601.gappssmtp.com; s=20230601; t=1696970011; x=1697574811; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=M8raC1pwRNPDQkT90yACQAJIG2zJj06sfDjsWA9vqOU=; b=sJT0KNgHZyV1JglWBV/ah8JdjZLFtkui2F5mzG6Vmaa9twFrgZYeyGODly27YHz/15 CG1ofPIyHvaz4DhIp4+wjsIUEyY6+mOrRJKSImtUO4R5NSOkPBSpaDT3AzxxUa6/rmjj kSzwHWu9/zFdjdjvr7lqcIK5+KI5zDlHPF6itxIPv6jEnustaRtdFLMOfKC8MQ7q0c10 MaqrEKvTYvSO0ZD8zCPHwf94rwvN3Zah2KQFjvz3fUCYHjz1q5XH5qUTGHNYtQ+Q4wgE 4QSjZYIXLdKUlOwTptjtO5vKX8wpGVIOPD/JZQ9kIyVTi+RqHcwRp7up8DNEjGf5IQ5d A7Hw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1696970011; x=1697574811; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=M8raC1pwRNPDQkT90yACQAJIG2zJj06sfDjsWA9vqOU=; b=GQRBhOsBVFfPQVhLf6M2sUrCDZ9SYsfor5MIDcLjYFog7toi/xSeHR+NUHF+GoHqRi 7dHzCLZys3YEPDQLACSboXnrZdC8HZzM4EcQxvu9ThCnBNo/Ow5n3puDpPkGjSuwShxK aUTTg1Ijyxug6QwjJmzjNHy1Y8B5im3G9Myg8RyyIvgCyLuZzlyqpy3W3xG8ItJ2VntE gjTG1ueF+XWri9OKu6dkZfUkgJAlQMhgGZM7rLK9G59jdqARepKvrBNf6dIXOa0bZ4qV YWho8Rhm4gWJOUw5TxizqTi8YrXa+YMxg4+RowmTgzO9fExeZs+c7uLnxBzqcKTaLJk7 6A5A== X-Gm-Message-State: AOJu0YxqscSuqEsg0JdhuWuxifdry+pWL9vcz/qiX3o8pMWcyp52Lt8c BpmfkJKAGJ/UIBOBWjKbFZhoFa9EE9YSURQH5rmMNQ== X-Google-Smtp-Source: AGHT+IH7Cifs0+Hq0a6X63OCrwl/LJQORO4UNZdn2RYGFul2FO3j+nPKqM43xJDPdbgLPTG7k0jh5g== X-Received: by 2002:a05:6214:104d:b0:658:3a12:9949 with SMTP id l13-20020a056214104d00b006583a129949mr20366515qvr.53.1696970011429; Tue, 10 Oct 2023 13:33:31 -0700 (PDT) Received: from localhost (104-178-186-189.lightspeed.milwwi.sbcglobal.net. [104.178.186.189]) by smtp.gmail.com with ESMTPSA id ct5-20020a056214178500b0066cf09f5ba9sm942231qvb.131.2023.10.10.13.33.31 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 10 Oct 2023 13:33:31 -0700 (PDT) Date: Tue, 10 Oct 2023 16:33:30 -0400 From: Taylor Blau To: git@vger.kernel.org Cc: Jonathan Tan , Junio C Hamano , Jeff King , SZEDER =?utf-8?b?R8OhYm9y?= Subject: [PATCH v3 04/17] gitformat-commit-graph: describe version 2 of BDAT Message-ID: <17703ed89ae9a60a26161682ce1c5a468d4ab3e0.1696969994.git.me@ttaylorr.com> References: MIME-Version: 1.0 Content-Disposition: inline In-Reply-To: Precedence: bulk List-ID: X-Mailing-List: git@vger.kernel.org From: Jonathan Tan The code change to Git to support version 2 will be done in subsequent commits. Signed-off-by: Jonathan Tan Signed-off-by: Junio C Hamano Signed-off-by: Taylor Blau --- Documentation/gitformat-commit-graph.txt | 9 ++++++--- 1 file changed, 6 insertions(+), 3 deletions(-) diff --git a/Documentation/gitformat-commit-graph.txt b/Documentation/gitformat-commit-graph.txt index 31cad585e2..3e906e8030 100644 --- a/Documentation/gitformat-commit-graph.txt +++ b/Documentation/gitformat-commit-graph.txt @@ -142,13 +142,16 @@ All multi-byte numbers are in network byte order. ==== Bloom Filter Data (ID: {'B', 'D', 'A', 'T'}) [Optional] * It starts with header consisting of three unsigned 32-bit integers: - - Version of the hash algorithm being used. We currently only support - value 1 which corresponds to the 32-bit version of the murmur3 hash + - Version of the hash algorithm being used. We currently support + value 2 which corresponds to the 32-bit version of the murmur3 hash implemented exactly as described in https://en.wikipedia.org/wiki/MurmurHash#Algorithm and the double hashing technique using seed values 0x293ae76f and 0x7e646e2 as described in https://doi.org/10.1007/978-3-540-30494-4_26 "Bloom Filters - in Probabilistic Verification" + in Probabilistic Verification". Version 1 Bloom filters have a bug that appears + when char is signed and the repository has path names that have characters >= + 0x80; Git supports reading and writing them, but this ability will be removed + in a future version of Git. - The number of times a path is hashed and hence the number of bit positions that cumulatively determine whether a file is present in the commit. - The minimum number of bits 'b' per entry in the Bloom filter. If the filter From patchwork Tue Oct 10 20:33:33 2023 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Taylor Blau X-Patchwork-Id: 13416039 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by smtp.lore.kernel.org (Postfix) with ESMTP id 0D120CD8CB6 for ; Tue, 10 Oct 2023 20:33:46 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S234416AbjJJUdo (ORCPT ); Tue, 10 Oct 2023 16:33:44 -0400 Received: from lindbergh.monkeyblade.net ([23.128.96.19]:60666 "EHLO lindbergh.monkeyblade.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S234483AbjJJUdh (ORCPT ); Tue, 10 Oct 2023 16:33:37 -0400 Received: from mail-qk1-x733.google.com (mail-qk1-x733.google.com [IPv6:2607:f8b0:4864:20::733]) by lindbergh.monkeyblade.net (Postfix) with ESMTPS id AFD66CA for ; Tue, 10 Oct 2023 13:33:35 -0700 (PDT) Received: by mail-qk1-x733.google.com with SMTP id af79cd13be357-77574c2cffdso24862385a.0 for ; Tue, 10 Oct 2023 13:33:35 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ttaylorr-com.20230601.gappssmtp.com; s=20230601; t=1696970014; x=1697574814; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=HrhldJ5BiV5CHevI9UjIhPIAgr48MFPlcVr2jUR0coQ=; b=ttdlZTDDL5ovJdsJ1FU5HYSLnI64X4ua97gG63/c69oBEsNwH+FU5FswbY3tMfhmYK rZj4Wr61R4iKqOG9jQONZIBif670QT0YRR+Ws83suvZi87ZWyT84iolHdcyBfC7t/Z/Q B6vW2r1WZ205Lr665SL3OB/8ibiaP8NpVeICYXU9CFlzEALka2MAqzgrXurUVIIqnr0K ULvNtLUVWjyDmRqkO+ZN/IZwZFsvqXdsbx4VGgWC2wYsPNDQOze2IQRPjIyNydNUzuj8 nAbJj2bffiCW1EoRhA0iTfy1+z1LjaAOhDU7iiEHOlnVl6lIzppQW8640L3C2Ft8PXPX YXEg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1696970014; x=1697574814; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=HrhldJ5BiV5CHevI9UjIhPIAgr48MFPlcVr2jUR0coQ=; b=P79Ujf8vksExWObdhwKMWCe4bH2DlJ6boqjyDdx9p2HPX1UyDBom+uWq5VnGCcutzH TMlwjcDwIxVtHZYqBg97i0VaorQftqSUe84VErRYoRU4w4EUmGCVxeC1KzYlG++EOshz cFUuDzlME+E1NFjf+ZlZYFKISoB5TSLocx08vH9tCOXg+KZ+aqJ48TOTxQdAISEuvCki aX2MkPqcPcvVspRcMHnx39SvlfiJGVBnmlzUsAkip6ehCotD9h8qLu1YexscUGDcprDt col5I5o4V44grZMwwwTc4kD1ssh9jDhNTF6/I11RvXoxwGpfkQlHVzpcsou+0zwtA9+D PH4Q== X-Gm-Message-State: AOJu0YxVz1xATkIjRMeg1Cv8FiDOJ/F8706EKfG7cvD7EIgYg6V9yQNJ YTxs/rYpVk7QuyOFFkscjn2zpM34egQPxhGJOBTQxg== X-Google-Smtp-Source: AGHT+IGQClRCTwmQGRSXBHJbh41t9mXyvwWrvGiX76vsZZNTOdRypzHweC2NuqXHIQZjCmW+9xWaqQ== X-Received: by 2002:a05:620a:2956:b0:765:734b:1792 with SMTP id n22-20020a05620a295600b00765734b1792mr22399786qkp.23.1696970014639; Tue, 10 Oct 2023 13:33:34 -0700 (PDT) Received: from localhost (104-178-186-189.lightspeed.milwwi.sbcglobal.net. [104.178.186.189]) by smtp.gmail.com with ESMTPSA id z26-20020ae9c11a000000b0076ca9f79e1fsm4617602qki.46.2023.10.10.13.33.34 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 10 Oct 2023 13:33:34 -0700 (PDT) Date: Tue, 10 Oct 2023 16:33:33 -0400 From: Taylor Blau To: git@vger.kernel.org Cc: Jonathan Tan , Junio C Hamano , Jeff King , SZEDER =?utf-8?b?R8OhYm9y?= Subject: [PATCH v3 05/17] t/helper/test-read-graph.c: extract `dump_graph_info()` Message-ID: <94552abf455c6d341a0811333ae4edb4a8cea259.1696969994.git.me@ttaylorr.com> References: MIME-Version: 1.0 Content-Disposition: inline In-Reply-To: Precedence: bulk List-ID: X-Mailing-List: git@vger.kernel.org Prepare for the 'read-graph' test helper to perform other tasks besides dumping high-level information about the commit-graph by extracting its main routine into a separate function. Signed-off-by: Taylor Blau Signed-off-by: Jonathan Tan Signed-off-by: Junio C Hamano Signed-off-by: Taylor Blau --- t/helper/test-read-graph.c | 31 ++++++++++++++++++------------- 1 file changed, 18 insertions(+), 13 deletions(-) diff --git a/t/helper/test-read-graph.c b/t/helper/test-read-graph.c index 8c7a83f578..3375392f6c 100644 --- a/t/helper/test-read-graph.c +++ b/t/helper/test-read-graph.c @@ -5,20 +5,8 @@ #include "bloom.h" #include "setup.h" -int cmd__read_graph(int argc UNUSED, const char **argv UNUSED) +static void dump_graph_info(struct commit_graph *graph) { - struct commit_graph *graph = NULL; - struct object_directory *odb; - - setup_git_directory(); - odb = the_repository->objects->odb; - - prepare_repo_settings(the_repository); - - graph = read_commit_graph_one(the_repository, odb); - if (!graph) - return 1; - printf("header: %08x %d %d %d %d\n", ntohl(*(uint32_t*)graph->data), *(unsigned char*)(graph->data + 4), @@ -57,6 +45,23 @@ int cmd__read_graph(int argc UNUSED, const char **argv UNUSED) if (graph->topo_levels) printf(" topo_levels"); printf("\n"); +} + +int cmd__read_graph(int argc UNUSED, const char **argv UNUSED) +{ + struct commit_graph *graph = NULL; + struct object_directory *odb; + + setup_git_directory(); + odb = the_repository->objects->odb; + + prepare_repo_settings(the_repository); + + graph = read_commit_graph_one(the_repository, odb); + if (!graph) + return 1; + + dump_graph_info(graph); UNLEAK(graph); From patchwork Tue Oct 10 20:33:36 2023 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Taylor Blau X-Patchwork-Id: 13416041 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by smtp.lore.kernel.org (Postfix) with ESMTP id 247FCCD8CB4 for ; Tue, 10 Oct 2023 20:33:50 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1343647AbjJJUds (ORCPT ); Tue, 10 Oct 2023 16:33:48 -0400 Received: from lindbergh.monkeyblade.net ([23.128.96.19]:60516 "EHLO lindbergh.monkeyblade.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S234452AbjJJUdm (ORCPT ); Tue, 10 Oct 2023 16:33:42 -0400 Received: from mail-qt1-x835.google.com (mail-qt1-x835.google.com [IPv6:2607:f8b0:4864:20::835]) by lindbergh.monkeyblade.net (Postfix) with ESMTPS id 1C369A4 for ; Tue, 10 Oct 2023 13:33:39 -0700 (PDT) Received: by mail-qt1-x835.google.com with SMTP id d75a77b69052e-4180b417309so35404631cf.0 for ; Tue, 10 Oct 2023 13:33:39 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ttaylorr-com.20230601.gappssmtp.com; s=20230601; t=1696970018; x=1697574818; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=9hlMWRvbuibDzGg8s6Dw+jhGsMoYGMxXXFJm4MZRO3c=; b=1QeINkwGeqZmNF+GFi3qmQXrMdO7THFLUQyqzMIoToPjITmo9GFyo9KL/lzJ3+dIlg QOe4idplqlNIFK7828tmMpejOh2RZXKXx2vdgi7YG6i3Qc1n2y0anoMv0loD/y/7U5fU 19LtTW6Hj74+t7NyGUIUvbs3W9RFIOdOSsBfy6VFv2QTjl/lluvLTQLglbAlvOh4wwpe MD2WEvtVldWleJhTVAttj+gEkTTYVwWk8uYL+N6sjf9+GVBWfcWDDiRo7PmidHPWKtsC 5JiLeeA+VFN2GrHmh1p3/iqHtpwSOlUiyza4+AX+N55XJWLzQqaFuHH+XOOyJQSF5Oxw CM+A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1696970018; x=1697574818; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=9hlMWRvbuibDzGg8s6Dw+jhGsMoYGMxXXFJm4MZRO3c=; b=SFlf8QsDFaMYv/c3Lauj8VDHw4TqwKU5hvadPjA+niaLZHW02l4Svg0yJeJo6btUEi E3ur8Kf3M5E/gijg3C3AHEJ8g0ZHvUiEborTbnoSeVLKt8DkvfB3hwUzQhxZI+72tNmv oXwNB3/L6+9RTHCsnDS3XIEi7uhULyp+V1BdR9Qm6WkhAYu5ZeJF5t4Q8wKZdvn5YlBu bK0JM/7/HTXbvc1HsLdh/HhcwtSvWKKdEXkopAC0WU8Zw43pQ3i9cSJ0ubb/LqVD3gal OQAspefEORslPvE6/0jGMoZ4cupSJHjZrQAG4d19muGHf/8ZyhUyOehP2h66Bc8ywzG5 M9TQ== X-Gm-Message-State: AOJu0YyCxa/oBIBhZPp48ZaiuyvTi8mbuk+CLRUPsKYTGVU1t7X+82Fg o7idVkykibvTyCkVIfYJsD8ACaRerBeoQ/rb6XzWsg== X-Google-Smtp-Source: AGHT+IFG8jssftf8I1p4+1Ws0o5CYJqvMmmyfTdMD5DDz6I2v7XJ6fdpdCdm2e4KEvl87D2e4fN0Xw== X-Received: by 2002:a05:622a:15d0:b0:418:1a08:729 with SMTP id d16-20020a05622a15d000b004181a080729mr20286750qty.10.1696970017784; Tue, 10 Oct 2023 13:33:37 -0700 (PDT) Received: from localhost (104-178-186-189.lightspeed.milwwi.sbcglobal.net. [104.178.186.189]) by smtp.gmail.com with ESMTPSA id y12-20020a05622a120c00b0041991642c62sm4774605qtx.73.2023.10.10.13.33.37 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 10 Oct 2023 13:33:37 -0700 (PDT) Date: Tue, 10 Oct 2023 16:33:36 -0400 From: Taylor Blau To: git@vger.kernel.org Cc: Jonathan Tan , Junio C Hamano , Jeff King , SZEDER =?utf-8?b?R8OhYm9y?= Subject: [PATCH v3 06/17] bloom.h: make `load_bloom_filter_from_graph()` public Message-ID: <3d81efa27b9b7c5b0cc2c77d080c7b65a4e46faa.1696969994.git.me@ttaylorr.com> References: MIME-Version: 1.0 Content-Disposition: inline In-Reply-To: Precedence: bulk List-ID: X-Mailing-List: git@vger.kernel.org Prepare for a future commit to use the load_bloom_filter_from_graph() function directly to load specific Bloom filters out of the commit-graph for manual inspection (to be used during tests). Signed-off-by: Taylor Blau Signed-off-by: Jonathan Tan Signed-off-by: Junio C Hamano Signed-off-by: Taylor Blau --- bloom.c | 6 +++--- bloom.h | 5 +++++ 2 files changed, 8 insertions(+), 3 deletions(-) diff --git a/bloom.c b/bloom.c index aef6b5fea2..3e78cfe79d 100644 --- a/bloom.c +++ b/bloom.c @@ -29,9 +29,9 @@ static inline unsigned char get_bitmask(uint32_t pos) return ((unsigned char)1) << (pos & (BITS_PER_WORD - 1)); } -static int load_bloom_filter_from_graph(struct commit_graph *g, - struct bloom_filter *filter, - uint32_t graph_pos) +int load_bloom_filter_from_graph(struct commit_graph *g, + struct bloom_filter *filter, + uint32_t graph_pos) { uint32_t lex_pos, start_index, end_index; diff --git a/bloom.h b/bloom.h index adde6dfe21..1e4f612d2c 100644 --- a/bloom.h +++ b/bloom.h @@ -3,6 +3,7 @@ struct commit; struct repository; +struct commit_graph; struct bloom_filter_settings { /* @@ -68,6 +69,10 @@ struct bloom_key { uint32_t *hashes; }; +int load_bloom_filter_from_graph(struct commit_graph *g, + struct bloom_filter *filter, + uint32_t graph_pos); + /* * Calculate the murmur3 32-bit hash value for the given data * using the given seed. From patchwork Tue Oct 10 20:33:39 2023 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Taylor Blau X-Patchwork-Id: 13416042 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by smtp.lore.kernel.org (Postfix) with ESMTP id C0474CD8CB7 for ; Tue, 10 Oct 2023 20:34:14 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1343758AbjJJUdv (ORCPT ); Tue, 10 Oct 2023 16:33:51 -0400 Received: from lindbergh.monkeyblade.net ([23.128.96.19]:60604 "EHLO lindbergh.monkeyblade.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S231201AbjJJUdn (ORCPT ); Tue, 10 Oct 2023 16:33:43 -0400 Received: from mail-qt1-x830.google.com (mail-qt1-x830.google.com [IPv6:2607:f8b0:4864:20::830]) by lindbergh.monkeyblade.net (Postfix) with ESMTPS id 130E09E for ; Tue, 10 Oct 2023 13:33:42 -0700 (PDT) Received: by mail-qt1-x830.google.com with SMTP id d75a77b69052e-41b2bf4e9edso2204871cf.1 for ; Tue, 10 Oct 2023 13:33:42 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ttaylorr-com.20230601.gappssmtp.com; s=20230601; t=1696970021; x=1697574821; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=MX+pavq5/sxT1Zqj6yeMZBcPmMUxfU9Mt6wuoUTDNHA=; b=FMxfoRPggHlaWFO7g5IYvg0yCvf9HVbeVo+YIZlh91KfyJCqrs1pESxd7LZLUPtt/T Tn5HUcFwzSzzy/1g77wsZoK6eArbqiIZ6gLiRcHN00ACpvMDJvIadLBqxvXtSgLjMvwd Nnbl3hACvwH5blaZhn4fkuPUlNdk8p59ENd9eQFJ+xO2PQC9kIa1SPjAF0QbtOAZKhKR P8h16Ap5C1ukNhFCgR78i9Bfou6rIiPBmHeerdDCRxb7xBDPLN+/NQ2CQu6LOAloD32A 08ihiYOJYwACQTPxpAAMMVkfeDvjPQIi27/4U5J5eDND3CDvrGvJio4lVl4u4vdg6fR/ Ay9Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1696970021; x=1697574821; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=MX+pavq5/sxT1Zqj6yeMZBcPmMUxfU9Mt6wuoUTDNHA=; b=iOX1xe8Rxc27JrIL/uo0xX60LH4W8cY6PE2DJnsLMdNnB1oRaBatRAL0kelLnM3PEU IBJ0s4AF5+eQJfRT770pey4mkzI+HQPuaF0WimjDGC3qhESODLL7w5ZXaNfskOV4ctYN gNRyB8G9yh/TTJt1sGxjqeRKaeFmu0csV4FrsrppCX1fMBPNiXn1NvH3EPE5iwUzJK+7 rgeBf3J3d1jg/s2WrWMzaWB74mydpx/QoNfS6VhU6FeqfD3yVAqfyq66e9UDWhlacHnA dPVam9EggWbpFsfNDoMX9RUz1kW09wIbcgxrzHP94Hgh8mrVIdRdAtT7k02qrsEsC6IP VyVw== X-Gm-Message-State: AOJu0YxAMFdlgUd8aqXcYRC4Zl3eDW8LQnPVSQx75VGJeb5MetJK6SGZ +LdCVR2eaZd7EC+NcfiTvvDwVonQo6cCnwKs3t1BPA== X-Google-Smtp-Source: AGHT+IEMw0S+xa7FB43lx3uoWqi5mq/Waj67oyVyqyM3BxasPcOMskioQGZrS3pM3lQHdkMo5yZYRQ== X-Received: by 2002:ac8:5b02:0:b0:418:11ab:1bf7 with SMTP id m2-20020ac85b02000000b0041811ab1bf7mr20775256qtw.30.1696970021013; Tue, 10 Oct 2023 13:33:41 -0700 (PDT) Received: from localhost (104-178-186-189.lightspeed.milwwi.sbcglobal.net. [104.178.186.189]) by smtp.gmail.com with ESMTPSA id g10-20020ac84b6a000000b00417dd1dd0adsm4798765qts.87.2023.10.10.13.33.40 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 10 Oct 2023 13:33:40 -0700 (PDT) Date: Tue, 10 Oct 2023 16:33:39 -0400 From: Taylor Blau To: git@vger.kernel.org Cc: Jonathan Tan , Junio C Hamano , Jeff King , SZEDER =?utf-8?b?R8OhYm9y?= Subject: [PATCH v3 07/17] t/helper/test-read-graph: implement `bloom-filters` mode Message-ID: References: MIME-Version: 1.0 Content-Disposition: inline In-Reply-To: Precedence: bulk List-ID: X-Mailing-List: git@vger.kernel.org Implement a mode of the "read-graph" test helper to dump out the hexadecimal contents of the Bloom filter(s) contained in a commit-graph. Signed-off-by: Taylor Blau Signed-off-by: Jonathan Tan Signed-off-by: Junio C Hamano Signed-off-by: Taylor Blau --- t/helper/test-read-graph.c | 44 +++++++++++++++++++++++++++++++++----- 1 file changed, 39 insertions(+), 5 deletions(-) diff --git a/t/helper/test-read-graph.c b/t/helper/test-read-graph.c index 3375392f6c..da9ac8584d 100644 --- a/t/helper/test-read-graph.c +++ b/t/helper/test-read-graph.c @@ -47,10 +47,32 @@ static void dump_graph_info(struct commit_graph *graph) printf("\n"); } -int cmd__read_graph(int argc UNUSED, const char **argv UNUSED) +static void dump_graph_bloom_filters(struct commit_graph *graph) +{ + uint32_t i; + + for (i = 0; i < graph->num_commits + graph->num_commits_in_base; i++) { + struct bloom_filter filter = { 0 }; + size_t j; + + if (load_bloom_filter_from_graph(graph, &filter, i) < 0) { + fprintf(stderr, "missing Bloom filter for graph " + "position %"PRIu32"\n", i); + continue; + } + + for (j = 0; j < filter.len; j++) + printf("%02x", filter.data[j]); + if (filter.len) + printf("\n"); + } +} + +int cmd__read_graph(int argc, const char **argv) { struct commit_graph *graph = NULL; struct object_directory *odb; + int ret = 0; setup_git_directory(); odb = the_repository->objects->odb; @@ -58,12 +80,24 @@ int cmd__read_graph(int argc UNUSED, const char **argv UNUSED) prepare_repo_settings(the_repository); graph = read_commit_graph_one(the_repository, odb); - if (!graph) - return 1; + if (!graph) { + ret = 1; + goto done; + } - dump_graph_info(graph); + if (argc <= 1) + dump_graph_info(graph); + else if (!strcmp(argv[1], "bloom-filters")) + dump_graph_bloom_filters(graph); + else { + fprintf(stderr, "unknown sub-command: '%s'\n", argv[1]); + ret = 1; + } +done: UNLEAK(graph); - return 0; + return ret; } + + From patchwork Tue Oct 10 20:33:42 2023 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Taylor Blau X-Patchwork-Id: 13416050 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by smtp.lore.kernel.org (Postfix) with ESMTP id A4B90CD8CB7 for ; Tue, 10 Oct 2023 20:34:28 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1343904AbjJJUe1 (ORCPT ); Tue, 10 Oct 2023 16:34:27 -0400 Received: from lindbergh.monkeyblade.net ([23.128.96.19]:56010 "EHLO lindbergh.monkeyblade.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1343808AbjJJUd4 (ORCPT ); Tue, 10 Oct 2023 16:33:56 -0400 Received: from mail-vk1-xa30.google.com (mail-vk1-xa30.google.com [IPv6:2607:f8b0:4864:20::a30]) by lindbergh.monkeyblade.net (Postfix) with ESMTPS id 869AFB0 for ; Tue, 10 Oct 2023 13:33:45 -0700 (PDT) Received: by mail-vk1-xa30.google.com with SMTP id 71dfb90a1353d-49d8dd34f7bso1997465e0c.3 for ; Tue, 10 Oct 2023 13:33:45 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ttaylorr-com.20230601.gappssmtp.com; s=20230601; t=1696970024; x=1697574824; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=z+Flt4mcDLnGAqCFrEF92nC3MImuUuxlNM0YsbFqOR4=; b=B/sTBdiKLAGB4MKkmKoPRE3gwq3jxLwWwWM+AXHrELLSVnRuQTmZc21Wdk3KBrBrNQ Pm93A0FMEt14KvuqCmNieawNV6lBkOpvRp5UF/ThCybYYVQy1Z7ar4MIETk/W71bqsXq b/rXFoDvtld844IRelromC863RZOIdIO1w4PioBco3UdSg8cJtW1sHG3QxduMucbbJJa afe4iIZDY0X2Yxwg0bi6qvcGLTpSTOhk8848I1cmLF3YO1Hju7lpGVUZIZkqEDrcHu8M gbpgW2CGp2E1rdagcO0zR6w89jFoLan8iL7Yl0QE/f1jvxd4Ou2LhIGvE02xZa3kVcBU v5fA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1696970024; x=1697574824; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=z+Flt4mcDLnGAqCFrEF92nC3MImuUuxlNM0YsbFqOR4=; b=qVdcxDdW5kpu/OPdPmyWLlg1mvJdpIgFtp1KIyfozRAE3gHIzfDrXESWMj3iNxBkO9 6kMmrh3M6vYojXG9JDmIwt9dT0c9scr8VfLc6ZfmaYlq82gkX+pED2eqw/d7ywBrxLVH javNYCeKYNtldH7vld5UFmBMSCnLOu+pSCwGx+/NtiZTFQ0fOMfUMS8572O/4JuwQfzD eKvI3sbXxivXWSdpR7kXArq90aUigQl2WzUwU4EfJySj864qu9ecdeHXKCtU6++kH5e9 bpPitLMeHyl2ZtOz8YwAES0GcjliybqwBdLEsH7+3e5WkMhzTbHlfxbLPSppUDPbTSsS i6xw== X-Gm-Message-State: AOJu0Yw8mwgqPCtVtukqb7rDt86b3+8M73/1Tr6HEmM3ysvBcdKae0TY XgiAhOhQML4Wi5yxWxrtdQP3yKDyBWVgIafN0zYcPg== X-Google-Smtp-Source: AGHT+IFatPozsq80wnHur68RgmD5QE05xs1GGdAgXwLbppqdpNyTMevo8hMZm+79jWsOWc4mLVjLvw== X-Received: by 2002:a1f:ed02:0:b0:4a0:6fd4:4333 with SMTP id l2-20020a1fed02000000b004a06fd44333mr6823106vkh.13.1696970024404; Tue, 10 Oct 2023 13:33:44 -0700 (PDT) Received: from localhost (104-178-186-189.lightspeed.milwwi.sbcglobal.net. [104.178.186.189]) by smtp.gmail.com with ESMTPSA id e6-20020ac845c6000000b004181234dd1dsm4716256qto.96.2023.10.10.13.33.43 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 10 Oct 2023 13:33:44 -0700 (PDT) Date: Tue, 10 Oct 2023 16:33:42 -0400 From: Taylor Blau To: git@vger.kernel.org Cc: Jonathan Tan , Junio C Hamano , Jeff King , SZEDER =?utf-8?b?R8OhYm9y?= Subject: [PATCH v3 08/17] t4216: test changed path filters with high bit paths Message-ID: References: MIME-Version: 1.0 Content-Disposition: inline In-Reply-To: Precedence: bulk List-ID: X-Mailing-List: git@vger.kernel.org From: Jonathan Tan Subsequent commits will teach Git another version of changed path filter that has different behavior with paths that contain at least one character with its high bit set, so test the existing behavior as a baseline. Signed-off-by: Jonathan Tan Signed-off-by: Junio C Hamano Signed-off-by: Taylor Blau Signed-off-by: Junio C Hamano Signed-off-by: Taylor Blau --- t/t4216-log-bloom.sh | 52 ++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 52 insertions(+) diff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh index f49a8f2fbf..da67c40134 100755 --- a/t/t4216-log-bloom.sh +++ b/t/t4216-log-bloom.sh @@ -484,4 +484,56 @@ test_expect_success 'merge graph layers with incompatible Bloom settings' ' ! grep "disabling Bloom filters" err ' +get_first_changed_path_filter () { + test-tool read-graph bloom-filters >filters.dat && + head -n 1 filters.dat +} + +# chosen to be the same under all Unicode normalization forms +CENT=$(printf "\302\242") + +test_expect_success 'set up repo with high bit path, version 1 changed-path' ' + git init highbit1 && + test_commit -C highbit1 c1 "$CENT" && + git -C highbit1 commit-graph write --reachable --changed-paths +' + +test_expect_success 'setup check value of version 1 changed-path' ' + ( + cd highbit1 && + echo "52a9" >expect && + get_first_changed_path_filter >actual && + test_cmp expect actual + ) +' + +# expect will not match actual if char is unsigned by default. Write the test +# in this way, so that a user running this test script can still see if the two +# files match. (It will appear as an ordinary success if they match, and a skip +# if not.) +if test_cmp highbit1/expect highbit1/actual +then + test_set_prereq SIGNED_CHAR_BY_DEFAULT +fi +test_expect_success SIGNED_CHAR_BY_DEFAULT 'check value of version 1 changed-path' ' + # Only the prereq matters for this test. + true +' + +test_expect_success 'setup make another commit' ' + # "git log" does not use Bloom filters for root commits - see how, in + # revision.c, rev_compare_tree() (the only code path that eventually calls + # get_bloom_filter()) is only called by try_to_simplify_commit() when the commit + # has one parent. Therefore, make another commit so that we perform the tests on + # a non-root commit. + test_commit -C highbit1 anotherc1 "another$CENT" +' + +test_expect_success 'version 1 changed-path used when version 1 requested' ' + ( + cd highbit1 && + test_bloom_filters_used "-- another$CENT" + ) +' + test_done From patchwork Tue Oct 10 20:33:46 2023 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Taylor Blau X-Patchwork-Id: 13416043 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by smtp.lore.kernel.org (Postfix) with ESMTP id 8710FCD8CB4 for ; Tue, 10 Oct 2023 20:34:17 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S234481AbjJJUeQ (ORCPT ); Tue, 10 Oct 2023 16:34:16 -0400 Received: from lindbergh.monkeyblade.net ([23.128.96.19]:56196 "EHLO lindbergh.monkeyblade.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1343978AbjJJUeA (ORCPT ); Tue, 10 Oct 2023 16:34:00 -0400 Received: from mail-oi1-x22b.google.com (mail-oi1-x22b.google.com [IPv6:2607:f8b0:4864:20::22b]) by lindbergh.monkeyblade.net (Postfix) with ESMTPS id 88F07D7 for ; Tue, 10 Oct 2023 13:33:48 -0700 (PDT) Received: by mail-oi1-x22b.google.com with SMTP id 5614622812f47-3ae2ec1a222so4283241b6e.2 for ; Tue, 10 Oct 2023 13:33:48 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ttaylorr-com.20230601.gappssmtp.com; s=20230601; t=1696970027; x=1697574827; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=w9ubiTD+WGpWF9zefVQgtwv/mgUKuh8XWcZTKgx5Q8A=; b=OVMSrjXTkzMOR0/PjiQf2JLVcn6n6TNnvsoY5OEbmlS/NsoAbx8c1VNj6+tTqOFaGS 7xWjrPcyfzjxiK5wefkp8jApeXiTmYNa/95X0iLZ2WqrXXUhCrE2XyV0sOF0ctgRz9Ud FUE06GpEcKnD6twmentyrz+wQ8vq+uM4WSwa0mt7OgzjWaHNt75rQLo/CCwYmyLtpcwN QVEGAYE6W0/dekRC3Upf3KvnSEnxev9lTrgXFOsixM0hHf4vthb0Kwj4Ix6urAoAeKIn DXFv7NIXCX3+3LSU9Xbd7T6owao+/2KNWW/i6NF1aE97CV/0J+HAbIdPJK8v/kX3CuzT 5syg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1696970027; x=1697574827; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=w9ubiTD+WGpWF9zefVQgtwv/mgUKuh8XWcZTKgx5Q8A=; b=dsqkDeeTkY7xUhY5qRHwP3yxz+n8sShHnN8AolTOFntEHKriWgNqTW7RMB55aaI9JZ hYwVEHaSkM618mpq8gVo5CWXcw7rtYaGhMbxx2f0rl9PxUvAmWm2xjm9hHE90rPnB0HE LEa/UUyNQs1mAVnsXumEc8wK7zvh02IUNw29vIm72ds6nZGPgONQoYxzWbuHP5N/PdTO MxCFK4lN1opkwlCpawKh0aWWOfFL8KsFSSNFH3V1QBcctQ6X++uV/p+HQ6fHeR2NNCxD qhMIfky2gKLZufOYEcnaTYf0IvpGSsvuLt8zqsASbjbsuyO4WvA1V6HBoY8K+AYN/4hM ENGQ== X-Gm-Message-State: AOJu0Yy1DDWAbbszMAjrs1vF7SxmZ2A7JMyPufrjFngMGwSDody2R5Qx i/ou5gbuOtEjiDe6Iovwdf/L0zo4p8di2mIY8jZ2bw== X-Google-Smtp-Source: AGHT+IGAQf7XNn9EhJEZcVbjF26BBbL4JTJX1VrQJMn655j5wkohHvTMWl2jxvJmq+3mf7hqfAanEw== X-Received: by 2002:a05:6808:199f:b0:3a8:83df:d5a4 with SMTP id bj31-20020a056808199f00b003a883dfd5a4mr25446854oib.59.1696970027645; Tue, 10 Oct 2023 13:33:47 -0700 (PDT) Received: from localhost (104-178-186-189.lightspeed.milwwi.sbcglobal.net. [104.178.186.189]) by smtp.gmail.com with ESMTPSA id j13-20020a0cf50d000000b0065b1f90ff8csm5135160qvm.40.2023.10.10.13.33.47 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 10 Oct 2023 13:33:47 -0700 (PDT) Date: Tue, 10 Oct 2023 16:33:46 -0400 From: Taylor Blau To: git@vger.kernel.org Cc: Jonathan Tan , Junio C Hamano , Jeff King , SZEDER =?utf-8?b?R8OhYm9y?= Subject: [PATCH v3 09/17] repo-settings: introduce commitgraph.changedPathsVersion Message-ID: References: MIME-Version: 1.0 Content-Disposition: inline In-Reply-To: Precedence: bulk List-ID: X-Mailing-List: git@vger.kernel.org From: Jonathan Tan A subsequent commit will introduce another version of the changed-path filter in the commit graph file. In order to control which version to write (and read), a config variable is needed. Therefore, introduce this config variable. For forwards compatibility, teach Git to not read commit graphs when the config variable is set to an unsupported version. Because we teach Git this, commitgraph.readChangedPaths is now redundant, so deprecate it and define its behavior in terms of the config variable we introduce. This commit does not change the behavior of writing (Git writes changed path filters when explicitly instructed regardless of any config variable), but a subsequent commit will restrict Git such that it will only write when commitgraph.changedPathsVersion is a recognized value. Signed-off-by: Jonathan Tan Signed-off-by: Junio C Hamano Signed-off-by: Taylor Blau --- Documentation/config/commitgraph.txt | 23 ++++++++++++++++++++--- commit-graph.c | 2 +- oss-fuzz/fuzz-commit-graph.c | 2 +- repo-settings.c | 6 +++++- repository.h | 2 +- 5 files changed, 28 insertions(+), 7 deletions(-) diff --git a/Documentation/config/commitgraph.txt b/Documentation/config/commitgraph.txt index 30604e4a4c..2dc9170622 100644 --- a/Documentation/config/commitgraph.txt +++ b/Documentation/config/commitgraph.txt @@ -9,6 +9,23 @@ commitGraph.maxNewFilters:: commit-graph write` (c.f., linkgit:git-commit-graph[1]). commitGraph.readChangedPaths:: - If true, then git will use the changed-path Bloom filters in the - commit-graph file (if it exists, and they are present). Defaults to - true. See linkgit:git-commit-graph[1] for more information. + Deprecated. Equivalent to commitGraph.changedPathsVersion=-1 if true, and + commitGraph.changedPathsVersion=0 if false. (If commitGraph.changedPathVersion + is also set, commitGraph.changedPathsVersion takes precedence.) + +commitGraph.changedPathsVersion:: + Specifies the version of the changed-path Bloom filters that Git will read and + write. May be -1, 0 or 1. ++ +Defaults to -1. ++ +If -1, Git will use the version of the changed-path Bloom filters in the +repository, defaulting to 1 if there are none. ++ +If 0, Git will not read any Bloom filters, and will write version 1 Bloom +filters when instructed to write. ++ +If 1, Git will only read version 1 Bloom filters, and will write version 1 +Bloom filters. ++ +See linkgit:git-commit-graph[1] for more information. diff --git a/commit-graph.c b/commit-graph.c index ae0902f7f4..ea677c87fb 100644 --- a/commit-graph.c +++ b/commit-graph.c @@ -411,7 +411,7 @@ struct commit_graph *parse_commit_graph(struct repo_settings *s, graph->read_generation_data = 1; } - if (s->commit_graph_read_changed_paths) { + if (s->commit_graph_changed_paths_version) { pair_chunk(cf, GRAPH_CHUNKID_BLOOMINDEXES, &graph->chunk_bloom_indexes); read_chunk(cf, GRAPH_CHUNKID_BLOOMDATA, diff --git a/oss-fuzz/fuzz-commit-graph.c b/oss-fuzz/fuzz-commit-graph.c index 2992079dd9..325c0b991a 100644 --- a/oss-fuzz/fuzz-commit-graph.c +++ b/oss-fuzz/fuzz-commit-graph.c @@ -19,7 +19,7 @@ int LLVMFuzzerTestOneInput(const uint8_t *data, size_t size) * possible. */ the_repository->settings.commit_graph_generation_version = 2; - the_repository->settings.commit_graph_read_changed_paths = 1; + the_repository->settings.commit_graph_changed_paths_version = 1; g = parse_commit_graph(&the_repository->settings, (void *)data, size); repo_clear(the_repository); free_commit_graph(g); diff --git a/repo-settings.c b/repo-settings.c index 525f69c0c7..db8fe817f3 100644 --- a/repo-settings.c +++ b/repo-settings.c @@ -24,6 +24,7 @@ void prepare_repo_settings(struct repository *r) int value; const char *strval; int manyfiles; + int read_changed_paths; if (!r->gitdir) BUG("Cannot add settings for uninitialized repository"); @@ -54,7 +55,10 @@ void prepare_repo_settings(struct repository *r) /* Commit graph config or default, does not cascade (simple) */ repo_cfg_bool(r, "core.commitgraph", &r->settings.core_commit_graph, 1); repo_cfg_int(r, "commitgraph.generationversion", &r->settings.commit_graph_generation_version, 2); - repo_cfg_bool(r, "commitgraph.readchangedpaths", &r->settings.commit_graph_read_changed_paths, 1); + repo_cfg_bool(r, "commitgraph.readchangedpaths", &read_changed_paths, 1); + repo_cfg_int(r, "commitgraph.changedpathsversion", + &r->settings.commit_graph_changed_paths_version, + read_changed_paths ? -1 : 0); repo_cfg_bool(r, "gc.writecommitgraph", &r->settings.gc_write_commit_graph, 1); repo_cfg_bool(r, "fetch.writecommitgraph", &r->settings.fetch_write_commit_graph, 0); diff --git a/repository.h b/repository.h index 5f18486f64..f71154e12c 100644 --- a/repository.h +++ b/repository.h @@ -29,7 +29,7 @@ struct repo_settings { int core_commit_graph; int commit_graph_generation_version; - int commit_graph_read_changed_paths; + int commit_graph_changed_paths_version; int gc_write_commit_graph; int fetch_write_commit_graph; int command_requires_full_index; From patchwork Tue Oct 10 20:33:49 2023 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Taylor Blau X-Patchwork-Id: 13416049 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by smtp.lore.kernel.org (Postfix) with ESMTP id 11A3BCD8CB6 for ; Tue, 10 Oct 2023 20:34:27 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1343876AbjJJUe0 (ORCPT ); Tue, 10 Oct 2023 16:34:26 -0400 Received: from lindbergh.monkeyblade.net ([23.128.96.19]:56160 "EHLO lindbergh.monkeyblade.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1344071AbjJJUeD (ORCPT ); Tue, 10 Oct 2023 16:34:03 -0400 Received: from mail-qt1-x834.google.com (mail-qt1-x834.google.com [IPv6:2607:f8b0:4864:20::834]) by lindbergh.monkeyblade.net (Postfix) with ESMTPS id EDEE8100 for ; Tue, 10 Oct 2023 13:33:52 -0700 (PDT) Received: by mail-qt1-x834.google.com with SMTP id d75a77b69052e-41b09c75bd5so28511281cf.3 for ; Tue, 10 Oct 2023 13:33:52 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ttaylorr-com.20230601.gappssmtp.com; s=20230601; t=1696970031; x=1697574831; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=xb93fYpq1DKlPHplGFI9AGn+OfCpcaE0KCgH3pAwbdw=; b=k8qUBnuPOr18tGFlZyRXF2Y1k6b1Acxg0Jjus+W/e/Fnw5sQajuXFGtZPJ9vSgZc8+ VHZ1s7J9y4AegrBzMI4ZAtAYlXLFjEdCxGyCCBEhToIYVUiZb1YOUoqlT8pxLjjQpfQ8 m1LEkKcoJ3vjzq2Wg8vjwbdsulPyFpanGs6A4zsjGWFnqyxtdd7+Vlu7dkI+631sltJb jCFS0KEzyTRpGYdbPjXYA6sFmHGXmWUhG66WoOWArX9sY9Gt0tdsvgIVjogb5FYZgtZm Wonf/78F9EaA6kVLVusMp934sNw7R8i3M1uKnIpDSMRa8CzSX0cAnKwGC0RxOMc6LVf2 cBNQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1696970031; x=1697574831; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=xb93fYpq1DKlPHplGFI9AGn+OfCpcaE0KCgH3pAwbdw=; b=rWcWaSW590kEzcyUB3qXkdwRO2ht4UI6s7pBC9WstgOrkwjSV7wpLx+m+s7Lh642Wj ygiHnXr8Drmy6BYjctL2TBu16VbKbdyARQey6QjA6EhCbeMxtDSk2iz9oFMmqMxTiCP+ tFGGQethI/4+sdA2e5FMzQm9G4ODSwWdiCLauHW7NrNltz5q3FOwuKSMdl8G9wrGRHjc B5YqWqt5q7AzGmG3cdapmBsi7Um9uzNFdDcBsRHIF+KM/w8WYxTlDXZbGt4AwSiYdIiN zsOUMpDjm+DBlMmhDKw31wySEdWgbvaj7MhK5HKDj6oX5K96mPg+41e2M4O6jEoFKO52 ZRLg== X-Gm-Message-State: AOJu0YxWzt1HfptBkHoWRaGKcLwjhLz53tXzdQ/YNoPxM827/tfGBRlh vApeAcXyvdTaRn2M3v0441/0z1ju5Dzqj3RIh52xdw== X-Google-Smtp-Source: AGHT+IEI7PnUtbYVzm+jutwNEZIq8Q0nefA182XxZ9a+uhxyRAiXtmCMTtFgUNKCUNfhk2BYF/OFvQ== X-Received: by 2002:a05:622a:2c9:b0:417:9e55:617f with SMTP id a9-20020a05622a02c900b004179e55617fmr23570876qtx.62.1696970031022; Tue, 10 Oct 2023 13:33:51 -0700 (PDT) Received: from localhost (104-178-186-189.lightspeed.milwwi.sbcglobal.net. [104.178.186.189]) by smtp.gmail.com with ESMTPSA id ku15-20020a05622a0a8f00b00419732075b4sm4762952qtb.84.2023.10.10.13.33.50 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 10 Oct 2023 13:33:50 -0700 (PDT) Date: Tue, 10 Oct 2023 16:33:49 -0400 From: Taylor Blau To: git@vger.kernel.org Cc: Jonathan Tan , Junio C Hamano , Jeff King , SZEDER =?utf-8?b?R8OhYm9y?= Subject: [PATCH v3 10/17] commit-graph: new filter ver. that fixes murmur3 Message-ID: <61d44519a5ffaf2c040198cf8d80d05a09de5de5.1696969994.git.me@ttaylorr.com> References: MIME-Version: 1.0 Content-Disposition: inline In-Reply-To: Precedence: bulk List-ID: X-Mailing-List: git@vger.kernel.org From: Jonathan Tan The murmur3 implementation in bloom.c has a bug when converting series of 4 bytes into network-order integers when char is signed (which is controllable by a compiler option, and the default signedness of char is platform-specific). When a string contains characters with the high bit set, this bug causes results that, although internally consistent within Git, does not accord with other implementations of murmur3 (thus, the changed path filters wouldn't be readable by other off-the-shelf implementatios of murmur3) and even with Git binaries that were compiled with different signedness of char. This bug affects both how Git writes changed path filters to disk and how Git interprets changed path filters on disk. Therefore, introduce a new version (2) of changed path filters that corrects this problem. The existing version (1) is still supported and is still the default, but users should migrate away from it as soon as possible. Because this bug only manifests with characters that have the high bit set, it may be possible that some (or all) commits in a given repo would have the same changed path filter both before and after this fix is applied. However, in order to determine whether this is the case, the changed paths would first have to be computed, at which point it is not much more expensive to just compute a new changed path filter. So this patch does not include any mechanism to "salvage" changed path filters from repositories. There is also no "mixed" mode - for each invocation of Git, reading and writing changed path filters are done with the same version number; this version number may be explicitly stated (typically if the user knows which version they need) or automatically determined from the version of the existing changed path filters in the repository. There is a change in write_commit_graph(). graph_read_bloom_data() makes it possible for chunk_bloom_data to be non-NULL but bloom_filter_settings to be NULL, which causes a segfault later on. I produced such a segfault while developing this patch, but couldn't find a way to reproduce it neither after this complete patch (or before), but in any case it seemed like a good thing to include that might help future patch authors. The value in t0095 was obtained from another murmur3 implementation using the following Go source code: package main import "fmt" import "github.com/spaolacci/murmur3" func main() { fmt.Printf("%x\n", murmur3.Sum32([]byte("Hello world!"))) fmt.Printf("%x\n", murmur3.Sum32([]byte{0x99, 0xaa, 0xbb, 0xcc, 0xdd, 0xee, 0xff})) } Signed-off-by: Jonathan Tan Signed-off-by: Junio C Hamano Signed-off-by: Taylor Blau Signed-off-by: Junio C Hamano Signed-off-by: Taylor Blau --- Documentation/config/commitgraph.txt | 5 +- bloom.c | 69 +++++++++++++++++- bloom.h | 8 +- commit-graph.c | 32 ++++++-- t/helper/test-bloom.c | 9 ++- t/t0095-bloom.sh | 8 ++ t/t4216-log-bloom.sh | 105 +++++++++++++++++++++++++++ 7 files changed, 223 insertions(+), 13 deletions(-) diff --git a/Documentation/config/commitgraph.txt b/Documentation/config/commitgraph.txt index 2dc9170622..acc74a2f27 100644 --- a/Documentation/config/commitgraph.txt +++ b/Documentation/config/commitgraph.txt @@ -15,7 +15,7 @@ commitGraph.readChangedPaths:: commitGraph.changedPathsVersion:: Specifies the version of the changed-path Bloom filters that Git will read and - write. May be -1, 0 or 1. + write. May be -1, 0, 1, or 2. + Defaults to -1. + @@ -28,4 +28,7 @@ filters when instructed to write. If 1, Git will only read version 1 Bloom filters, and will write version 1 Bloom filters. + +If 2, Git will only read version 2 Bloom filters, and will write version 2 +Bloom filters. ++ See linkgit:git-commit-graph[1] for more information. diff --git a/bloom.c b/bloom.c index 3e78cfe79d..ebef5cfd2f 100644 --- a/bloom.c +++ b/bloom.c @@ -66,7 +66,64 @@ int load_bloom_filter_from_graph(struct commit_graph *g, * Not considered to be cryptographically secure. * Implemented as described in https://en.wikipedia.org/wiki/MurmurHash#Algorithm */ -uint32_t murmur3_seeded(uint32_t seed, const char *data, size_t len) +uint32_t murmur3_seeded_v2(uint32_t seed, const char *data, size_t len) +{ + const uint32_t c1 = 0xcc9e2d51; + const uint32_t c2 = 0x1b873593; + const uint32_t r1 = 15; + const uint32_t r2 = 13; + const uint32_t m = 5; + const uint32_t n = 0xe6546b64; + int i; + uint32_t k1 = 0; + const char *tail; + + int len4 = len / sizeof(uint32_t); + + uint32_t k; + for (i = 0; i < len4; i++) { + uint32_t byte1 = (uint32_t)(unsigned char)data[4*i]; + uint32_t byte2 = ((uint32_t)(unsigned char)data[4*i + 1]) << 8; + uint32_t byte3 = ((uint32_t)(unsigned char)data[4*i + 2]) << 16; + uint32_t byte4 = ((uint32_t)(unsigned char)data[4*i + 3]) << 24; + k = byte1 | byte2 | byte3 | byte4; + k *= c1; + k = rotate_left(k, r1); + k *= c2; + + seed ^= k; + seed = rotate_left(seed, r2) * m + n; + } + + tail = (data + len4 * sizeof(uint32_t)); + + switch (len & (sizeof(uint32_t) - 1)) { + case 3: + k1 ^= ((uint32_t)(unsigned char)tail[2]) << 16; + /*-fallthrough*/ + case 2: + k1 ^= ((uint32_t)(unsigned char)tail[1]) << 8; + /*-fallthrough*/ + case 1: + k1 ^= ((uint32_t)(unsigned char)tail[0]) << 0; + k1 *= c1; + k1 = rotate_left(k1, r1); + k1 *= c2; + seed ^= k1; + break; + } + + seed ^= (uint32_t)len; + seed ^= (seed >> 16); + seed *= 0x85ebca6b; + seed ^= (seed >> 13); + seed *= 0xc2b2ae35; + seed ^= (seed >> 16); + + return seed; +} + +static uint32_t murmur3_seeded_v1(uint32_t seed, const char *data, size_t len) { const uint32_t c1 = 0xcc9e2d51; const uint32_t c2 = 0x1b873593; @@ -131,8 +188,14 @@ void fill_bloom_key(const char *data, int i; const uint32_t seed0 = 0x293ae76f; const uint32_t seed1 = 0x7e646e2c; - const uint32_t hash0 = murmur3_seeded(seed0, data, len); - const uint32_t hash1 = murmur3_seeded(seed1, data, len); + uint32_t hash0, hash1; + if (settings->hash_version == 2) { + hash0 = murmur3_seeded_v2(seed0, data, len); + hash1 = murmur3_seeded_v2(seed1, data, len); + } else { + hash0 = murmur3_seeded_v1(seed0, data, len); + hash1 = murmur3_seeded_v1(seed1, data, len); + } key->hashes = (uint32_t *)xcalloc(settings->num_hashes, sizeof(uint32_t)); for (i = 0; i < settings->num_hashes; i++) diff --git a/bloom.h b/bloom.h index 1e4f612d2c..138d57a86b 100644 --- a/bloom.h +++ b/bloom.h @@ -8,9 +8,11 @@ struct commit_graph; struct bloom_filter_settings { /* * The version of the hashing technique being used. - * We currently only support version = 1 which is + * The newest version is 2, which is * the seeded murmur3 hashing technique implemented - * in bloom.c. + * in bloom.c. Bloom filters of version 1 were created + * with prior versions of Git, which had a bug in the + * implementation of the hash function. */ uint32_t hash_version; @@ -80,7 +82,7 @@ int load_bloom_filter_from_graph(struct commit_graph *g, * Not considered to be cryptographically secure. * Implemented as described in https://en.wikipedia.org/wiki/MurmurHash#Algorithm */ -uint32_t murmur3_seeded(uint32_t seed, const char *data, size_t len); +uint32_t murmur3_seeded_v2(uint32_t seed, const char *data, size_t len); void fill_bloom_key(const char *data, size_t len, diff --git a/commit-graph.c b/commit-graph.c index ea677c87fb..db623afd09 100644 --- a/commit-graph.c +++ b/commit-graph.c @@ -314,17 +314,26 @@ static int graph_read_oid_lookup(const unsigned char *chunk_start, return 0; } +struct graph_read_bloom_data_context { + struct commit_graph *g; + int *commit_graph_changed_paths_version; +}; + static int graph_read_bloom_data(const unsigned char *chunk_start, size_t chunk_size, void *data) { - struct commit_graph *g = data; + struct graph_read_bloom_data_context *c = data; + struct commit_graph *g = c->g; uint32_t hash_version; - g->chunk_bloom_data = chunk_start; hash_version = get_be32(chunk_start); - if (hash_version != 1) + if (*c->commit_graph_changed_paths_version == -1) { + *c->commit_graph_changed_paths_version = hash_version; + } else if (hash_version != *c->commit_graph_changed_paths_version) { return 0; + } + g->chunk_bloom_data = chunk_start; g->bloom_filter_settings = xmalloc(sizeof(struct bloom_filter_settings)); g->bloom_filter_settings->hash_version = hash_version; g->bloom_filter_settings->num_hashes = get_be32(chunk_start + 4); @@ -412,10 +421,14 @@ struct commit_graph *parse_commit_graph(struct repo_settings *s, } if (s->commit_graph_changed_paths_version) { + struct graph_read_bloom_data_context context = { + .g = graph, + .commit_graph_changed_paths_version = &s->commit_graph_changed_paths_version + }; pair_chunk(cf, GRAPH_CHUNKID_BLOOMINDEXES, &graph->chunk_bloom_indexes); read_chunk(cf, GRAPH_CHUNKID_BLOOMDATA, - graph_read_bloom_data, graph); + graph_read_bloom_data, &context); } if (graph->chunk_bloom_indexes && graph->chunk_bloom_data) { @@ -2441,6 +2454,13 @@ int write_commit_graph(struct object_directory *odb, } if (!commit_graph_compatible(r)) return 0; + if (r->settings.commit_graph_changed_paths_version < -1 + || r->settings.commit_graph_changed_paths_version > 2) { + warning(_("attempting to write a commit-graph, but " + "'commitgraph.changedPathsVersion' (%d) is not supported"), + r->settings.commit_graph_changed_paths_version); + return 0; + } CALLOC_ARRAY(ctx, 1); ctx->r = r; @@ -2453,6 +2473,8 @@ int write_commit_graph(struct object_directory *odb, ctx->write_generation_data = (get_configured_generation_version(r) == 2); ctx->num_generation_data_overflows = 0; + bloom_settings.hash_version = r->settings.commit_graph_changed_paths_version == 2 + ? 2 : 1; bloom_settings.bits_per_entry = git_env_ulong("GIT_TEST_BLOOM_SETTINGS_BITS_PER_ENTRY", bloom_settings.bits_per_entry); bloom_settings.num_hashes = git_env_ulong("GIT_TEST_BLOOM_SETTINGS_NUM_HASHES", @@ -2482,7 +2504,7 @@ int write_commit_graph(struct object_directory *odb, g = ctx->r->objects->commit_graph; /* We have changed-paths already. Keep them in the next graph */ - if (g && g->chunk_bloom_data) { + if (g && g->bloom_filter_settings) { ctx->changed_paths = 1; ctx->bloom_settings = g->bloom_filter_settings; } diff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c index aabe31d724..3cbc0a5b50 100644 --- a/t/helper/test-bloom.c +++ b/t/helper/test-bloom.c @@ -50,6 +50,7 @@ static void get_bloom_filter_for_commit(const struct object_id *commit_oid) static const char *bloom_usage = "\n" " test-tool bloom get_murmur3 \n" +" test-tool bloom get_murmur3_seven_highbit\n" " test-tool bloom generate_filter [...]\n" " test-tool bloom get_filter_for_commit \n"; @@ -64,7 +65,13 @@ int cmd__bloom(int argc, const char **argv) uint32_t hashed; if (argc < 3) usage(bloom_usage); - hashed = murmur3_seeded(0, argv[2], strlen(argv[2])); + hashed = murmur3_seeded_v2(0, argv[2], strlen(argv[2])); + printf("Murmur3 Hash with seed=0:0x%08x\n", hashed); + } + + if (!strcmp(argv[1], "get_murmur3_seven_highbit")) { + uint32_t hashed; + hashed = murmur3_seeded_v2(0, "\x99\xaa\xbb\xcc\xdd\xee\xff", 7); printf("Murmur3 Hash with seed=0:0x%08x\n", hashed); } diff --git a/t/t0095-bloom.sh b/t/t0095-bloom.sh index b567383eb8..c8d84ab606 100755 --- a/t/t0095-bloom.sh +++ b/t/t0095-bloom.sh @@ -29,6 +29,14 @@ test_expect_success 'compute unseeded murmur3 hash for test string 2' ' test_cmp expect actual ' +test_expect_success 'compute unseeded murmur3 hash for test string 3' ' + cat >expect <<-\EOF && + Murmur3 Hash with seed=0:0xa183ccfd + EOF + test-tool bloom get_murmur3_seven_highbit >actual && + test_cmp expect actual +' + test_expect_success 'compute bloom key for empty string' ' cat >expect <<-\EOF && Hashes:0x5615800c|0x5b966560|0x61174ab4|0x66983008|0x6c19155c|0x7199fab0|0x771ae004| diff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh index da67c40134..8f8b5d4966 100755 --- a/t/t4216-log-bloom.sh +++ b/t/t4216-log-bloom.sh @@ -536,4 +536,109 @@ test_expect_success 'version 1 changed-path used when version 1 requested' ' ) ' +test_expect_success 'version 1 changed-path not used when version 2 requested' ' + ( + cd highbit1 && + git config --add commitgraph.changedPathsVersion 2 && + test_bloom_filters_not_used "-- another$CENT" + ) +' + +test_expect_success 'version 1 changed-path used when autodetect requested' ' + ( + cd highbit1 && + git config --add commitgraph.changedPathsVersion -1 && + test_bloom_filters_used "-- another$CENT" + ) +' + +test_expect_success 'when writing another commit graph, preserve existing version 1 of changed-path' ' + test_commit -C highbit1 c1double "$CENT$CENT" && + git -C highbit1 commit-graph write --reachable --changed-paths && + ( + cd highbit1 && + git config --add commitgraph.changedPathsVersion -1 && + echo "options: bloom(1,10,7) read_generation_data" >expect && + test-tool read-graph >full && + grep options full >actual && + test_cmp expect actual + ) +' + +test_expect_success 'set up repo with high bit path, version 2 changed-path' ' + git init highbit2 && + git -C highbit2 config --add commitgraph.changedPathsVersion 2 && + test_commit -C highbit2 c2 "$CENT" && + git -C highbit2 commit-graph write --reachable --changed-paths +' + +test_expect_success 'check value of version 2 changed-path' ' + ( + cd highbit2 && + echo "c01f" >expect && + get_first_changed_path_filter >actual && + test_cmp expect actual + ) +' + +test_expect_success 'setup make another commit' ' + # "git log" does not use Bloom filters for root commits - see how, in + # revision.c, rev_compare_tree() (the only code path that eventually calls + # get_bloom_filter()) is only called by try_to_simplify_commit() when the commit + # has one parent. Therefore, make another commit so that we perform the tests on + # a non-root commit. + test_commit -C highbit2 anotherc2 "another$CENT" +' + +test_expect_success 'version 2 changed-path used when version 2 requested' ' + ( + cd highbit2 && + test_bloom_filters_used "-- another$CENT" + ) +' + +test_expect_success 'version 2 changed-path not used when version 1 requested' ' + ( + cd highbit2 && + git config --add commitgraph.changedPathsVersion 1 && + test_bloom_filters_not_used "-- another$CENT" + ) +' + +test_expect_success 'version 2 changed-path used when autodetect requested' ' + ( + cd highbit2 && + git config --add commitgraph.changedPathsVersion -1 && + test_bloom_filters_used "-- another$CENT" + ) +' + +test_expect_success 'when writing another commit graph, preserve existing version 2 of changed-path' ' + test_commit -C highbit2 c2double "$CENT$CENT" && + git -C highbit2 commit-graph write --reachable --changed-paths && + ( + cd highbit2 && + git config --add commitgraph.changedPathsVersion -1 && + echo "options: bloom(2,10,7) read_generation_data" >expect && + test-tool read-graph >full && + grep options full >actual && + test_cmp expect actual + ) +' + +test_expect_success 'when writing commit graph, do not reuse changed-path of another version' ' + git init doublewrite && + test_commit -C doublewrite c "$CENT" && + git -C doublewrite config --add commitgraph.changedPathsVersion 1 && + git -C doublewrite commit-graph write --reachable --changed-paths && + git -C doublewrite config --add commitgraph.changedPathsVersion 2 && + git -C doublewrite commit-graph write --reachable --changed-paths && + ( + cd doublewrite && + echo "c01f" >expect && + get_first_changed_path_filter >actual && + test_cmp expect actual + ) +' + test_done From patchwork Tue Oct 10 20:33:52 2023 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Taylor Blau X-Patchwork-Id: 13416047 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by smtp.lore.kernel.org (Postfix) with ESMTP id 7E506CD8CB4 for ; Tue, 10 Oct 2023 20:34:24 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1343810AbjJJUeX (ORCPT ); Tue, 10 Oct 2023 16:34:23 -0400 Received: from lindbergh.monkeyblade.net ([23.128.96.19]:56142 "EHLO lindbergh.monkeyblade.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1344066AbjJJUeC (ORCPT ); Tue, 10 Oct 2023 16:34:02 -0400 Received: from mail-qv1-xf2e.google.com (mail-qv1-xf2e.google.com [IPv6:2607:f8b0:4864:20::f2e]) by lindbergh.monkeyblade.net (Postfix) with ESMTPS id 41D00116 for ; Tue, 10 Oct 2023 13:33:55 -0700 (PDT) Received: by mail-qv1-xf2e.google.com with SMTP id 6a1803df08f44-65af726775eso2455776d6.0 for ; Tue, 10 Oct 2023 13:33:55 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ttaylorr-com.20230601.gappssmtp.com; s=20230601; t=1696970034; x=1697574834; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=YWjvb9Lx38HM1iTNYnkLxQZJ8NbC5I79r8imatKSbIk=; b=cDAyUPIYG5lWlcKCyAjsUih8avoqQwylWBkMTwu/E3CMJ6s530qk/RsVRbq/qcj6Sl za3Q+z0Vr8oyxiS8ltF92H25twS7UyajrKkNORWEYsQqyOXVilgY5wX2zvWgxnrHZYmv RwhlyfujTuomk0exnxxcxhmNdLX2c7YAVz2svfZRK1aOQMDj0TzpdUp/o2f2qMvb9u9O ZmpGCbYBsZPgD7KVEVuh8f4ts8XHy52sP2yEiL2LKeoAL2VimetfkzuJP6A+D1E8uvHk vpXjqs7XCGL35t1xdo4nQVYEN2PIAafgqvY1dzOKb/m70doDyypr42sWjqgZVzvRHHh6 WJnw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1696970034; x=1697574834; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=YWjvb9Lx38HM1iTNYnkLxQZJ8NbC5I79r8imatKSbIk=; b=GvIsXjdINNhWXBZLHa1fmChsc82KKW7a6M/DZ2Vtn8vvQPYXyxCdKfn7ipTZ42D5SQ BiSvre6/QCAxtvOKCgyP2fLGiEu5SpIU/ejywV/7zfVqmr7Q8Gb/n376tE1YcPpEDK3I aR9Z/8XhjMFrNKWA0xNUbNBLClMd7ZfK5c6ngKWzB/Kn3uKYOOZc97988GVbVXv4Yil9 hG9DpgbXMfY8wgpaVYJF+k20VWaq+T9iEnDFmVkFIDOnWgld+t/7MOq8BArsRmQAzoTh fh/0WlsbnnuIWzTsr7I8XrZME2I7K/KzOHIk/O6ugcY+6EtRqDbgwv3bnDYts7lq2eFz i/lA== X-Gm-Message-State: AOJu0YyN1Yz1eiVY17LSpURW/QdA8Q/Hf1FnOfX+9MCkIyClen7wO7GM DlkhjmDFtVdXUPOqr/DWunMxt/pbzWVcS9BeAD/Gsw== X-Google-Smtp-Source: AGHT+IEAFBBL9gpZc346P+9dE9ZR8kXsGvsCU0TNDVBEJlg+vj0PlXc+8nT1I1n3oj0znznbSrsgYQ== X-Received: by 2002:a05:6214:2a4e:b0:62d:ddeb:3770 with SMTP id jf14-20020a0562142a4e00b0062dddeb3770mr23512649qvb.0.1696970034179; Tue, 10 Oct 2023 13:33:54 -0700 (PDT) Received: from localhost (104-178-186-189.lightspeed.milwwi.sbcglobal.net. [104.178.186.189]) by smtp.gmail.com with ESMTPSA id t10-20020a0ca68a000000b0065afcf19e23sm5015025qva.62.2023.10.10.13.33.53 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 10 Oct 2023 13:33:53 -0700 (PDT) Date: Tue, 10 Oct 2023 16:33:52 -0400 From: Taylor Blau To: git@vger.kernel.org Cc: Jonathan Tan , Junio C Hamano , Jeff King , SZEDER =?utf-8?b?R8OhYm9y?= Subject: [PATCH v3 11/17] bloom: annotate filters with hash version Message-ID: References: MIME-Version: 1.0 Content-Disposition: inline In-Reply-To: Precedence: bulk List-ID: X-Mailing-List: git@vger.kernel.org In subsequent commits, we will want to load existing Bloom filters out of a commit-graph, even when the hash version they were computed with does not match the value of `commitGraph.changedPathVersion`. In order to differentiate between the two, add a "version" field to each Bloom filter. Signed-off-by: Taylor Blau --- bloom.c | 11 ++++++++--- bloom.h | 1 + 2 files changed, 9 insertions(+), 3 deletions(-) diff --git a/bloom.c b/bloom.c index ebef5cfd2f..9b6a30f6f6 100644 --- a/bloom.c +++ b/bloom.c @@ -55,6 +55,7 @@ int load_bloom_filter_from_graph(struct commit_graph *g, filter->data = (unsigned char *)(g->chunk_bloom_data + sizeof(unsigned char) * start_index + BLOOMDATA_CHUNK_HEADER_SIZE); + filter->version = g->bloom_filter_settings->hash_version; return 1; } @@ -240,11 +241,13 @@ static int pathmap_cmp(const void *hashmap_cmp_fn_data UNUSED, return strcmp(e1->path, e2->path); } -static void init_truncated_large_filter(struct bloom_filter *filter) +static void init_truncated_large_filter(struct bloom_filter *filter, + int version) { filter->data = xmalloc(1); filter->data[0] = 0xFF; filter->len = 1; + filter->version = version; } struct bloom_filter *get_or_compute_bloom_filter(struct repository *r, @@ -329,13 +332,15 @@ struct bloom_filter *get_or_compute_bloom_filter(struct repository *r, } if (hashmap_get_size(&pathmap) > settings->max_changed_paths) { - init_truncated_large_filter(filter); + init_truncated_large_filter(filter, + settings->hash_version); if (computed) *computed |= BLOOM_TRUNC_LARGE; goto cleanup; } filter->len = (hashmap_get_size(&pathmap) * settings->bits_per_entry + BITS_PER_WORD - 1) / BITS_PER_WORD; + filter->version = settings->hash_version; if (!filter->len) { if (computed) *computed |= BLOOM_TRUNC_EMPTY; @@ -355,7 +360,7 @@ struct bloom_filter *get_or_compute_bloom_filter(struct repository *r, } else { for (i = 0; i < diff_queued_diff.nr; i++) diff_free_filepair(diff_queued_diff.queue[i]); - init_truncated_large_filter(filter); + init_truncated_large_filter(filter, settings->hash_version); if (computed) *computed |= BLOOM_TRUNC_LARGE; diff --git a/bloom.h b/bloom.h index 138d57a86b..330a140520 100644 --- a/bloom.h +++ b/bloom.h @@ -55,6 +55,7 @@ struct bloom_filter_settings { struct bloom_filter { unsigned char *data; size_t len; + int version; }; /* From patchwork Tue Oct 10 20:33:55 2023 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Taylor Blau X-Patchwork-Id: 13416045 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by smtp.lore.kernel.org (Postfix) with ESMTP id 55201CD8CB4 for ; Tue, 10 Oct 2023 20:34:21 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1343768AbjJJUeU (ORCPT ); Tue, 10 Oct 2023 16:34:20 -0400 Received: from lindbergh.monkeyblade.net ([23.128.96.19]:56096 "EHLO lindbergh.monkeyblade.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1344150AbjJJUeE (ORCPT ); Tue, 10 Oct 2023 16:34:04 -0400 Received: from mail-qt1-x829.google.com (mail-qt1-x829.google.com [IPv6:2607:f8b0:4864:20::829]) by lindbergh.monkeyblade.net (Postfix) with ESMTPS id 69343B0 for ; Tue, 10 Oct 2023 13:33:58 -0700 (PDT) Received: by mail-qt1-x829.google.com with SMTP id d75a77b69052e-4195fddd6d7so2497121cf.0 for ; Tue, 10 Oct 2023 13:33:58 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ttaylorr-com.20230601.gappssmtp.com; s=20230601; t=1696970037; x=1697574837; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=vky+ZwHmD3rWBOf6rqDZIiCrZcxnIkK2Pg5TwQcxMGM=; b=ShZs/IJWVnk+WiSuVdzK1E3HGq4oub/n2ZBzKf6WIRNRp0GqMqzfggA2TwyCmrRxuC RlE7iSN16yVcT1hXkQcz5R8UL0OVyLR9M6AFATr91Z7TAlTjzGTtk6PwvBlYO5t/nAL6 dDH6L1yxOo8wSjHcw0gnOvk3+9sNqUruPbiwlIUSdbwYug86i1sg+Qwb9Q+fvO/h8sEA ZGb3EYea5p3MMw1xz/53vk59C+uZbg6GTXZzobgqSf2LzlBB6pMNUHVp+nhHKxlLbLPt kdvLK7ljEuBnFiTzodf0fGj+x73WP0XwpvZ7q+1QyszsgIkTQXURu8ipacwRZwftSt6W xFNA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1696970037; x=1697574837; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=vky+ZwHmD3rWBOf6rqDZIiCrZcxnIkK2Pg5TwQcxMGM=; b=TN+GNyhWEwhxhqOQVPqwqcqDTNmseCmHtgJDLwCq1mMc83baYz7VAFQGKg06PK67Ic gDnmGBCNZQeuHxS3OVTiyyM9NXeosu3C2BSsP2gHPgg/8Nqvut76/W4BGdou9bO/G+Dc SBTpm81imxPqJRr0LfM8LN8Ww6huorOQxwySopAX4JR07jrZkaiWH8R6aPEWPzR5h1Bn c3rTR5eVhwckI6IzBSPw1Zd4IvUIsP+WumVL4rrUI+SaTqweuapXi03FKfVPXZO9lqhz xwkYLIKAiWVRlgdA4KctWyywlcrgrTk8IR2vXfnyyZsdAZwLHUYHZuiHDM5w4DTR/TQ6 U4LA== X-Gm-Message-State: AOJu0YzQUjrvLpSyN3psSNYu9sbzNt11D3nWeNYRYatTZ6NtQASLqf2i rxUkV934BEXc3ryWoZnWTiuy2Rigr3YRaI33Vp+Aqw== X-Google-Smtp-Source: AGHT+IF5srpGevBBOjhs2H0ODTZ+os1UaZXggwQBwKGIwNscsXY0yMOnVYZ3jtfKAFckLX9HyWks5Q== X-Received: by 2002:a05:622a:10f:b0:403:a662:a3c1 with SMTP id u15-20020a05622a010f00b00403a662a3c1mr22923360qtw.29.1696970037211; Tue, 10 Oct 2023 13:33:57 -0700 (PDT) Received: from localhost (104-178-186-189.lightspeed.milwwi.sbcglobal.net. [104.178.186.189]) by smtp.gmail.com with ESMTPSA id o20-20020ac86d14000000b0041950c7f6d8sm4719147qtt.60.2023.10.10.13.33.56 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 10 Oct 2023 13:33:56 -0700 (PDT) Date: Tue, 10 Oct 2023 16:33:55 -0400 From: Taylor Blau To: git@vger.kernel.org Cc: Jonathan Tan , Junio C Hamano , Jeff King , SZEDER =?utf-8?b?R8OhYm9y?= Subject: [PATCH v3 12/17] bloom: prepare to discard incompatible Bloom filters Message-ID: <2ba10a4b4b890d3c75f128a972a93889edc4f60e.1696969994.git.me@ttaylorr.com> References: MIME-Version: 1.0 Content-Disposition: inline In-Reply-To: Precedence: bulk List-ID: X-Mailing-List: git@vger.kernel.org Callers use the inline `get_bloom_filter()` implementation as a thin wrapper around `get_or_compute_bloom_filter()`. The former calls the latter with a value of "0" for `compute_if_not_present`, making `get_bloom_filter()` the default read-only path for fetching an existing Bloom filter. Callers expect the value returned from `get_bloom_filter()` is usable, that is that it's compatible with the configured value corresponding to `commitGraph.changedPathsVersion`. This is OK, since the commit-graph machinery only initializes its BDAT chunk (thereby enabling it to service Bloom filter queries) when the Bloom filter hash_version is compatible with our settings. So any value returned by `get_bloom_filter()` is trivially useable. However, subsequent commits will load the BDAT chunk even when the Bloom filters are built with incompatible hash versions. Prepare to handle this by teaching `get_bloom_filter()` to discard filters that are incompatible with the configured hash version. Callers who wish to read incompatible filters (e.g., for upgrading filters from v1 to v2) may use the lower level routine, `get_or_compute_bloom_filter()`. Signed-off-by: Taylor Blau --- bloom.c | 20 +++++++++++++++++++- bloom.h | 20 ++++++++++++++++++-- 2 files changed, 37 insertions(+), 3 deletions(-) diff --git a/bloom.c b/bloom.c index 9b6a30f6f6..739fa093ba 100644 --- a/bloom.c +++ b/bloom.c @@ -250,6 +250,23 @@ static void init_truncated_large_filter(struct bloom_filter *filter, filter->version = version; } +struct bloom_filter *get_bloom_filter(struct repository *r, struct commit *c) +{ + struct bloom_filter *filter; + int hash_version; + + filter = get_or_compute_bloom_filter(r, c, 0, NULL, NULL); + if (!filter) + return NULL; + + prepare_repo_settings(r); + hash_version = r->settings.commit_graph_changed_paths_version; + + if (!(hash_version == -1 || hash_version == filter->version)) + return NULL; /* unusable filter */ + return filter; +} + struct bloom_filter *get_or_compute_bloom_filter(struct repository *r, struct commit *c, int compute_if_not_present, @@ -275,7 +292,8 @@ struct bloom_filter *get_or_compute_bloom_filter(struct repository *r, filter, graph_pos); } - if (filter->data && filter->len) + if ((filter->data && filter->len) && + (!settings || settings->hash_version == filter->version)) return filter; if (!compute_if_not_present) return NULL; diff --git a/bloom.h b/bloom.h index 330a140520..bfe389e29c 100644 --- a/bloom.h +++ b/bloom.h @@ -110,8 +110,24 @@ struct bloom_filter *get_or_compute_bloom_filter(struct repository *r, const struct bloom_filter_settings *settings, enum bloom_filter_computed *computed); -#define get_bloom_filter(r, c) get_or_compute_bloom_filter( \ - (r), (c), 0, NULL, NULL) +/* + * Find the Bloom filter associated with the given commit "c". + * + * If any of the following are true + * + * - the repository does not have a commit-graph, or + * - the repository disables reading from the commit-graph, or + * - the given commit does not have a Bloom filter computed, or + * - there is a Bloom filter for commit "c", but it cannot be read + * because the filter uses an incompatible version of murmur3 + * + * , then `get_bloom_filter()` will return NULL. Otherwise, the corresponding + * Bloom filter will be returned. + * + * For callers who wish to inspect Bloom filters with incompatible hash + * versions, use get_or_compute_bloom_filter(). + */ +struct bloom_filter *get_bloom_filter(struct repository *r, struct commit *c); int bloom_filter_contains(const struct bloom_filter *filter, const struct bloom_key *key, From patchwork Tue Oct 10 20:33:59 2023 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Taylor Blau X-Patchwork-Id: 13416046 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by smtp.lore.kernel.org (Postfix) with ESMTP id 7A720CD8CB7 for ; Tue, 10 Oct 2023 20:34:22 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S234411AbjJJUeV (ORCPT ); Tue, 10 Oct 2023 16:34:21 -0400 Received: from lindbergh.monkeyblade.net ([23.128.96.19]:56438 "EHLO lindbergh.monkeyblade.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1344178AbjJJUeF (ORCPT ); Tue, 10 Oct 2023 16:34:05 -0400 Received: from mail-qt1-x82a.google.com (mail-qt1-x82a.google.com [IPv6:2607:f8b0:4864:20::82a]) by lindbergh.monkeyblade.net (Postfix) with ESMTPS id A78E8C6 for ; Tue, 10 Oct 2023 13:34:01 -0700 (PDT) Received: by mail-qt1-x82a.google.com with SMTP id d75a77b69052e-417f872fb94so43021841cf.0 for ; Tue, 10 Oct 2023 13:34:01 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ttaylorr-com.20230601.gappssmtp.com; s=20230601; t=1696970040; x=1697574840; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=8ypW6Fo5acSQYVAcI+hiGciE9kVHJqjoj87AiDRIe7Q=; b=bcshmg4qSQHGP6EYQMlvEiitlVkbWouOsitT40beNVrm62dDXd7ryPulZo+oKgFjd1 Fnoa/2DNP1atuvSQyHmsrMusHvjFLtMZ6RVXYuVzdruvqwimOXGXMF1/qqF5C9gXhiiB Y5WuXcSmTVwBIFDRf4srHhnAGmCM2Tc/cf7n7PXBRTOylaFzloKEDFAkhQHG2l2dS4Ns odlT3t+1HuTbqyI4aZVD8VEbObsTfm/Xput4hBZ2kfuiIP2wRhKi5zGac9jpDEroKZic HjfrucYwkmRJW4o6YdTErXTwHwtDwsVljOYxYzixvn2ab4PfBY09eYprmTbJ8bRKaJkF V1Bg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1696970040; x=1697574840; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=8ypW6Fo5acSQYVAcI+hiGciE9kVHJqjoj87AiDRIe7Q=; b=Pv/EoRILZuZQ/44mjBUVKLTc+3cKGL4GM7phasUrdKEekrTNTcJcZAsi543ghU/LOb yv26l5LHK/cy9mbAXdn1NcgFupi5EyLpRuLI4hT7F42XQ7ArCLoBoIBsV2uHGmXYlEiF E+Wk9Jg6SrJsBv1FfozzyIqNAcn0Fd8gkVVBgCgyk1r6/fVzv49WhZx6v0g14YDgB6qN 33SytpYAaAdknrRFesBS847dgD9vBkMYB99ZZtQqVXX067+eqOuoUGmVhLBqXMXxCjBS bARFOWPsjGUSh1g6y+/ua599qFemz2H2tb8GuaTDv0FOu/iCoVquE7u6ok6FmmG5z7cr 6LWA== X-Gm-Message-State: AOJu0YzkDnHnBPEiQeSt35YmhYNEmY4FGxDhWi44+OybyPgzq6Cm06lH k51aF7Q9MuHShzpkQf3qy4F+SOng6/1JzQs0YVKMlQ== X-Google-Smtp-Source: AGHT+IETD5tu7JLnakpFFHgyStz+AapPtqoDkh9FgC66W4VEWWvObJEIM1osz2yJ3BaAiyfBHg3iDg== X-Received: by 2002:ac8:5e4c:0:b0:413:5e4d:bf40 with SMTP id i12-20020ac85e4c000000b004135e4dbf40mr20916509qtx.68.1696970040611; Tue, 10 Oct 2023 13:34:00 -0700 (PDT) Received: from localhost (104-178-186-189.lightspeed.milwwi.sbcglobal.net. [104.178.186.189]) by smtp.gmail.com with ESMTPSA id e1-20020ac81301000000b0040331a24f16sm4767114qtj.3.2023.10.10.13.34.00 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 10 Oct 2023 13:34:00 -0700 (PDT) Date: Tue, 10 Oct 2023 16:33:59 -0400 From: Taylor Blau To: git@vger.kernel.org Cc: Jonathan Tan , Junio C Hamano , Jeff King , SZEDER =?utf-8?b?R8OhYm9y?= Subject: [PATCH v3 13/17] commit-graph.c: unconditionally load Bloom filters Message-ID: <09d8669c3a074e7a2ace9d650a345244b2362f7e.1696969994.git.me@ttaylorr.com> References: MIME-Version: 1.0 Content-Disposition: inline In-Reply-To: Precedence: bulk List-ID: X-Mailing-List: git@vger.kernel.org In 9e4df4da07 (commit-graph: new filter ver. that fixes murmur3, 2023-08-01), we began ignoring the Bloom data ("BDAT") chunk for commit-graphs whose Bloom filters were computed using a hash version incompatible with the value of `commitGraph.changedPathVersion`. Now that the Bloom API has been hardened to discard these incompatible filters (with the exception of low-level APIs), we can safely load these Bloom filters unconditionally. We no longer want to return early from `graph_read_bloom_data()`, and similarly do not want to set the bloom_settings' `hash_version` field as a side-effect. The latter is because we want to wait until we know which Bloom settings we're using (either the defaults, from the GIT_TEST variables, or from the previous commit-graph layer) before deciding what hash_version to use. If we detect an existing BDAT chunk, we'll infer the rest of the settings (e.g., number of hashes, bits per entry, and maximum number of changed paths) from the earlier graph layer. The hash_version will be inferred from the previous layer as well, unless one has already been specified via configuration. Once all of that is done, we normalize the value of the hash_version to either "1" or "2". Signed-off-by: Taylor Blau --- commit-graph.c | 19 ++++++++++--------- 1 file changed, 10 insertions(+), 9 deletions(-) diff --git a/commit-graph.c b/commit-graph.c index db623afd09..fa3b58e762 100644 --- a/commit-graph.c +++ b/commit-graph.c @@ -327,12 +327,6 @@ static int graph_read_bloom_data(const unsigned char *chunk_start, uint32_t hash_version; hash_version = get_be32(chunk_start); - if (*c->commit_graph_changed_paths_version == -1) { - *c->commit_graph_changed_paths_version = hash_version; - } else if (hash_version != *c->commit_graph_changed_paths_version) { - return 0; - } - g->chunk_bloom_data = chunk_start; g->bloom_filter_settings = xmalloc(sizeof(struct bloom_filter_settings)); g->bloom_filter_settings->hash_version = hash_version; @@ -2473,8 +2467,7 @@ int write_commit_graph(struct object_directory *odb, ctx->write_generation_data = (get_configured_generation_version(r) == 2); ctx->num_generation_data_overflows = 0; - bloom_settings.hash_version = r->settings.commit_graph_changed_paths_version == 2 - ? 2 : 1; + bloom_settings.hash_version = r->settings.commit_graph_changed_paths_version; bloom_settings.bits_per_entry = git_env_ulong("GIT_TEST_BLOOM_SETTINGS_BITS_PER_ENTRY", bloom_settings.bits_per_entry); bloom_settings.num_hashes = git_env_ulong("GIT_TEST_BLOOM_SETTINGS_NUM_HASHES", @@ -2506,10 +2499,18 @@ int write_commit_graph(struct object_directory *odb, /* We have changed-paths already. Keep them in the next graph */ if (g && g->bloom_filter_settings) { ctx->changed_paths = 1; - ctx->bloom_settings = g->bloom_filter_settings; + + /* don't propagate the hash_version unless unspecified */ + if (bloom_settings.hash_version == -1) + bloom_settings.hash_version = g->bloom_filter_settings->hash_version; + bloom_settings.bits_per_entry = g->bloom_filter_settings->bits_per_entry; + bloom_settings.num_hashes = g->bloom_filter_settings->num_hashes; + bloom_settings.max_changed_paths = g->bloom_filter_settings->max_changed_paths; } } + bloom_settings.hash_version = bloom_settings.hash_version == 2 ? 2 : 1; + if (ctx->split) { struct commit_graph *g = ctx->r->objects->commit_graph; From patchwork Tue Oct 10 20:34:02 2023 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Taylor Blau X-Patchwork-Id: 13416048 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by smtp.lore.kernel.org (Postfix) with ESMTP id D5134CD8CB7 for ; Tue, 10 Oct 2023 20:34:25 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1343829AbjJJUeY (ORCPT ); Tue, 10 Oct 2023 16:34:24 -0400 Received: from lindbergh.monkeyblade.net ([23.128.96.19]:56150 "EHLO lindbergh.monkeyblade.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1344204AbjJJUeI (ORCPT ); Tue, 10 Oct 2023 16:34:08 -0400 Received: from mail-qt1-x82e.google.com (mail-qt1-x82e.google.com [IPv6:2607:f8b0:4864:20::82e]) by lindbergh.monkeyblade.net (Postfix) with ESMTPS id C8944D8 for ; Tue, 10 Oct 2023 13:34:04 -0700 (PDT) Received: by mail-qt1-x82e.google.com with SMTP id d75a77b69052e-4195fddd6d7so2497971cf.0 for ; Tue, 10 Oct 2023 13:34:04 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ttaylorr-com.20230601.gappssmtp.com; s=20230601; t=1696970044; x=1697574844; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=hy3AUi1FbnH1wytiKQkQYCfah7452QZi9oIeAiyMxsA=; b=mde5Vb8okEpY/l1vs1vDjKxwM7mD/pn2WvtJnKKYl33jfkvDMPClNdIfZIyWlr/IPA eGzFxg6rOO+JNRe5PO+N1yjUkhRyS8KtNK7dSC6mWlfunAOIPpbOi5PuchuHNPAeG6+P fPtkDlUqm52MQF5Cu3EaI5+N84OhObCeRO8JevmPsUKudR/MDCcAslRXyG97KqPha2BL jjtqhH5PwjXwm+D4BxRTDFGJ1XpfUKiU4rSxwzuCAvcyfnQKbM4xmc9WydGCDTNnrFBL bJKU3nbUJttr9tpn7rBdDHa3QzK27EOBm8Y3yq6P60EuKCP89Xf97bgznpiG/M8cggZ+ Gl8A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1696970044; x=1697574844; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=hy3AUi1FbnH1wytiKQkQYCfah7452QZi9oIeAiyMxsA=; b=RRnGFIS+pdJ3hxp/9clzyIpfw/3N9U2KYucEhm4tnZvYmX0ecfS5tI87oySXTnXVkR yCUrY4SRzpNEmansskUPNlo9lc8qzyLw+ircZJ9SVu4rnu3dq4daEr5jehJT8qa0AKUO 4LEIgpJpyegseB41kSRGlsNKOkFpDVzgcZIYRpVQB1Y6Zzp9hriKJAIWhzlrfihD0lZf +J36uT1V1bPvXcnQu5Yia2l3ElDupMafDlx+dyB1n4EbkCPlR1stSwkulks9CK3Zsyd8 sYx/uOG2Jkisi52U7PVnxduma1+txyK33Pycg7yYMMcm7Z3kG1Q5MVa8UCLGP1HgPu2c AH+A== X-Gm-Message-State: AOJu0Yzpc1GkRhWbf6753ATMNVOw/IrEWFzuWP+jTP7GuDzV2zQDfh0A X+TCDeNMWQz5iSdFjVsttZ41mAt7MkVg35Q1Xko/pg== X-Google-Smtp-Source: AGHT+IFECsWAXtLKYdn2nQth2zCTrMoLBp5EY+U+hmyS4k2t8e6FyV5XC0kJE5+SfNAapF1RuzYeGg== X-Received: by 2002:a05:622a:34a:b0:418:1437:303b with SMTP id r10-20020a05622a034a00b004181437303bmr23948161qtw.27.1696970043707; Tue, 10 Oct 2023 13:34:03 -0700 (PDT) Received: from localhost (104-178-186-189.lightspeed.milwwi.sbcglobal.net. [104.178.186.189]) by smtp.gmail.com with ESMTPSA id pj30-20020a05620a1d9e00b00775afce4235sm4571876qkn.131.2023.10.10.13.34.03 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 10 Oct 2023 13:34:03 -0700 (PDT) Date: Tue, 10 Oct 2023 16:34:02 -0400 From: Taylor Blau To: git@vger.kernel.org Cc: Jonathan Tan , Junio C Hamano , Jeff King , SZEDER =?utf-8?b?R8OhYm9y?= Subject: [PATCH v3 14/17] commit-graph: drop unnecessary `graph_read_bloom_data_context` Message-ID: <0d4f9dc4ee58feb81928f92f6f8ac465e49083c0.1696969994.git.me@ttaylorr.com> References: MIME-Version: 1.0 Content-Disposition: inline In-Reply-To: Precedence: bulk List-ID: X-Mailing-List: git@vger.kernel.org The `graph_read_bloom_data_context` struct was introduced in an earlier commit in order to pass pointers to the commit-graph and changed-path Bloom filter version when reading the BDAT chunk. The previous commit no longer writes through the changed_paths_version pointer, making the surrounding context structure unnecessary. Drop it and pass a pointer to the commit-graph directly when reading the BDAT chunk. Noticed-by: Jonathan Tan Signed-off-by: Taylor Blau --- commit-graph.c | 14 ++------------ 1 file changed, 2 insertions(+), 12 deletions(-) diff --git a/commit-graph.c b/commit-graph.c index fa3b58e762..e0fc62e110 100644 --- a/commit-graph.c +++ b/commit-graph.c @@ -314,16 +314,10 @@ static int graph_read_oid_lookup(const unsigned char *chunk_start, return 0; } -struct graph_read_bloom_data_context { - struct commit_graph *g; - int *commit_graph_changed_paths_version; -}; - static int graph_read_bloom_data(const unsigned char *chunk_start, size_t chunk_size, void *data) { - struct graph_read_bloom_data_context *c = data; - struct commit_graph *g = c->g; + struct commit_graph *g = data; uint32_t hash_version; hash_version = get_be32(chunk_start); @@ -415,14 +409,10 @@ struct commit_graph *parse_commit_graph(struct repo_settings *s, } if (s->commit_graph_changed_paths_version) { - struct graph_read_bloom_data_context context = { - .g = graph, - .commit_graph_changed_paths_version = &s->commit_graph_changed_paths_version - }; pair_chunk(cf, GRAPH_CHUNKID_BLOOMINDEXES, &graph->chunk_bloom_indexes); read_chunk(cf, GRAPH_CHUNKID_BLOOMDATA, - graph_read_bloom_data, &context); + graph_read_bloom_data, graph); } if (graph->chunk_bloom_indexes && graph->chunk_bloom_data) { From patchwork Tue Oct 10 20:34:05 2023 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Taylor Blau X-Patchwork-Id: 13416044 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by smtp.lore.kernel.org (Postfix) with ESMTP id F3887CD8CB6 for ; Tue, 10 Oct 2023 20:34:19 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S234511AbjJJUeT (ORCPT ); Tue, 10 Oct 2023 16:34:19 -0400 Received: from lindbergh.monkeyblade.net ([23.128.96.19]:56142 "EHLO lindbergh.monkeyblade.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1344214AbjJJUeJ (ORCPT ); Tue, 10 Oct 2023 16:34:09 -0400 Received: from mail-qk1-x732.google.com (mail-qk1-x732.google.com [IPv6:2607:f8b0:4864:20::732]) by lindbergh.monkeyblade.net (Postfix) with ESMTPS id E2C20DE for ; Tue, 10 Oct 2023 13:34:07 -0700 (PDT) Received: by mail-qk1-x732.google.com with SMTP id af79cd13be357-7741c2fae49so399677585a.0 for ; Tue, 10 Oct 2023 13:34:07 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ttaylorr-com.20230601.gappssmtp.com; s=20230601; t=1696970047; x=1697574847; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=nid45Zo12ttb0THUqfggGnjWqUBDgJZZ2rA2o/8mJoc=; b=D5v/VRXah50gykoECRBmqU/TEV1qtqisBzoneGKyD+jcKPgNcedVQTvhGMF94+MPfh 2sawpFhhEjQ/xDsOF5N+7Ryzpivj2kSK8nLMejX5PIZQO+ElOfcltHs5kkGpx79i98Pk gp2WdNSNAEsZP+kQ1QVbh51iQBi8is5Z0pWM7sUJ+PkbSEy8Jhc6AOvpPN9nPwd+5O0b Pfw1cMn6O6ZZxiFx7D4jOCBHdLkekChoPAsBTa6eJQLCcGrV093rtJ2tWTrTjXNGZV5G 2Nwq8CEMFv/rEabAxJYQhnPttqaPBuiLxf3yl4vhSEmSY9tAmRM2IeWLtwSWNzIDcClo F3SQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1696970047; x=1697574847; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=nid45Zo12ttb0THUqfggGnjWqUBDgJZZ2rA2o/8mJoc=; b=RcMH6wUT/8DadvbQM5nTj3CaNMuheInjV8o2g562T6WJjdq9CqOzpxNFj1sIR+OM+E m1VcqB57avyekcnIiEaDzn4RcE65wWWV9i/5pZQDQmPcl0jqkkjgOO+Xnw0yWgyJsqK9 K+rkrLWDZDN2CTvbFbWtBk7DkEbFJezqG15p7Uba8EZfy9yiGLQazGb10/HW38qFESEk Qj8ZDhiWfMyWFk+fM35ePuSmzxv/l74cGxOqL2AzegEwlecr3p7h0SddWdTSN9Tiwgxu iKn5yLGeh5xg5moOgnlLNkYBA2ygWmxlrt1sagjnYXhrFGf6YkxxJdeHCQhdlnjlCYOG Dg7A== X-Gm-Message-State: AOJu0YxL/U6etxNt6QgrItnDDq+SHOnfuG5wUaO6rNIESTSqUnIzaYxq Y2klP1tIvCoZxnz28QDO26ZMgtlPwZ7/VP7O7zy3ag== X-Google-Smtp-Source: AGHT+IFOu4ZtB4miM0mmYLnmvhDguzDB1DIdsVnEfcM8pdWSvK/augCjGlQQEeCs7reZDA8igFSA9Q== X-Received: by 2002:a05:620a:f14:b0:765:a8b2:18dd with SMTP id v20-20020a05620a0f1400b00765a8b218ddmr22796864qkl.20.1696970046909; Tue, 10 Oct 2023 13:34:06 -0700 (PDT) Received: from localhost (104-178-186-189.lightspeed.milwwi.sbcglobal.net. [104.178.186.189]) by smtp.gmail.com with ESMTPSA id d1-20020a05620a136100b00774309d3e89sm4638793qkl.7.2023.10.10.13.34.06 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 10 Oct 2023 13:34:06 -0700 (PDT) Date: Tue, 10 Oct 2023 16:34:05 -0400 From: Taylor Blau To: git@vger.kernel.org Cc: Jonathan Tan , Junio C Hamano , Jeff King , SZEDER =?utf-8?b?R8OhYm9y?= Subject: [PATCH v3 15/17] object.h: fix mis-aligned flag bits table Message-ID: <1f7f27bc47eba053731e75fb7a4cb330065f6caf.1696969994.git.me@ttaylorr.com> References: MIME-Version: 1.0 Content-Disposition: inline In-Reply-To: Precedence: bulk List-ID: X-Mailing-List: git@vger.kernel.org Bit position 23 is one column too far to the left. Signed-off-by: Taylor Blau --- object.h | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/object.h b/object.h index 114d45954d..db25714b4e 100644 --- a/object.h +++ b/object.h @@ -62,7 +62,7 @@ void object_array_init(struct object_array *array); /* * object flag allocation: - * revision.h: 0---------10 15 23------27 + * revision.h: 0---------10 15 23------27 * fetch-pack.c: 01 67 * negotiator/default.c: 2--5 * walker.c: 0-2 From patchwork Tue Oct 10 20:34:08 2023 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 8bit X-Patchwork-Submitter: Taylor Blau X-Patchwork-Id: 13416051 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by smtp.lore.kernel.org (Postfix) with ESMTP id C5D3BCD8CB6 for ; Tue, 10 Oct 2023 20:34:30 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1343928AbjJJUe3 (ORCPT ); Tue, 10 Oct 2023 16:34:29 -0400 Received: from lindbergh.monkeyblade.net ([23.128.96.19]:44294 "EHLO lindbergh.monkeyblade.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1344237AbjJJUeN (ORCPT ); Tue, 10 Oct 2023 16:34:13 -0400 Received: from mail-qk1-x72c.google.com (mail-qk1-x72c.google.com [IPv6:2607:f8b0:4864:20::72c]) by lindbergh.monkeyblade.net (Postfix) with ESMTPS id 10FEF8E for ; Tue, 10 Oct 2023 13:34:11 -0700 (PDT) Received: by mail-qk1-x72c.google.com with SMTP id af79cd13be357-77433e7a876so401702185a.3 for ; Tue, 10 Oct 2023 13:34:11 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ttaylorr-com.20230601.gappssmtp.com; s=20230601; t=1696970050; x=1697574850; darn=vger.kernel.org; h=in-reply-to:content-transfer-encoding:content-disposition :mime-version:references:message-id:subject:cc:to:from:date:from:to :cc:subject:date:message-id:reply-to; bh=tt8FzCHKkAdTF+bRXZq+jmaSPnCV1reVD8wibNA1QZE=; b=ePhlNbu+E880E/9nhQ1dzENluNp1ODkdcgyuXFdrdUOeSFaPF1IoaQrm0bloymVThq EBHw7TiapPffRhJCpv+O3u1IWe1UqXd+bk5wzHzhbWjB0sDcYqs73EVh24YkQKMa5IGC qmaKlvk6jAKZ+Gk4HFuhGxuBQI9VeH2CDtngwYxaEjztnJa3V09MjWQZJFrW8Q0KV8Qm hyF/D2mtSi5mKcgvrqCILJMix/xChk+rI14vrdvV8LNApIYvlgh9r/B3JqN8xyggUSPE RwROxlS5HapTVoqc/EZph0Pm63StfUXKjfvWVJwS45Ao65BmpJLbdbJTf7dvZNJJT5b1 lmgA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1696970050; x=1697574850; h=in-reply-to:content-transfer-encoding:content-disposition :mime-version:references:message-id:subject:cc:to:from:date :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to; bh=tt8FzCHKkAdTF+bRXZq+jmaSPnCV1reVD8wibNA1QZE=; b=XqAQemZ4aZ0vkQ0NHSTQqi9sJNek1LqgckFfS/YOsyQaQ8/HbLa96wQwxgbGqxt6C8 VONwBca4RRviIda56Kszc+NfAuhFsOfZp+kPPO4/YuXiyjqBSvqlqKXs9D2vmhbH1TiT B0dFdQdK7+82MXuOPrPy762+q/ry2DDMoM1yPvlqjpUoazopu956uOesjnpMhdBXTB+Q C6a8dYCTeX/gZhPpHzCPcwbrwa3Asgb49/juOCxg6OC3OHrWro/tUgNfGuTkC98/TjLA mESKTa/Bj9RpN4iOCVY5FEQDMhugi+T63eBD8j4yvz2bYzsA0kpLY8Bv5IieB1Utwh2P VrAA== X-Gm-Message-State: AOJu0YwKTSGPlRrheNp78C4/oVSZdsalsosmbCVByfAVJTLqKzsH5lRH IGxbtqE6rCvjI0SxzpbHp+qw6qv1fgxkdVZIMYd64w== X-Google-Smtp-Source: AGHT+IHKaGSFgDncOvbXHHsk/DxXYZkrGKlb2UoKNyYt5MD7lkzaIovhI7FEzFwl5dO9XZp5Ilnm4g== X-Received: by 2002:ad4:5310:0:b0:653:5960:8959 with SMTP id y16-20020ad45310000000b0065359608959mr18951325qvr.41.1696970049928; Tue, 10 Oct 2023 13:34:09 -0700 (PDT) Received: from localhost (104-178-186-189.lightspeed.milwwi.sbcglobal.net. [104.178.186.189]) by smtp.gmail.com with ESMTPSA id k29-20020a0cb25d000000b006585069e894sm5067350qve.109.2023.10.10.13.34.09 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 10 Oct 2023 13:34:09 -0700 (PDT) Date: Tue, 10 Oct 2023 16:34:08 -0400 From: Taylor Blau To: git@vger.kernel.org Cc: Jonathan Tan , Junio C Hamano , Jeff King , SZEDER =?utf-8?b?R8OhYm9y?= Subject: [PATCH v3 16/17] commit-graph: reuse existing Bloom filters where possible Message-ID: References: MIME-Version: 1.0 Content-Disposition: inline In-Reply-To: Precedence: bulk List-ID: X-Mailing-List: git@vger.kernel.org In 9e4df4da07 (commit-graph: new filter ver. that fixes murmur3, 2023-08-01), a bug was described where it's possible for Git to produce non-murmur3 hashes when the platform's "char" type is signed, and there are paths with characters whose highest bit is set (i.e. all characters >= 0x80). That patch allows the caller to control which version of Bloom filters are read and written. However, even on platforms with a signed "char" type, it is possible to reuse existing Bloom filters if and only if there are no changed paths in any commit's first parent tree-diff whose characters have their highest bit set. When this is the case, we can reuse the existing filter without having to compute a new one. This is done by marking trees which are known to have (or not have) any such paths. When a commit's root tree is verified to not have any such paths, we mark it as such and declare that the commit's Bloom filter is reusable. Note that this heuristic only goes in one direction. If neither a commit nor its first parent have any paths in their trees with non-ASCII characters, then we know for certain that a path with non-ASCII characters will not appear in a tree-diff against that commit's first parent. The reverse isn't necessarily true: just because the tree-diff doesn't contain any such paths does not imply that no such paths exist in either tree. So we end up recomputing some Bloom filters that we don't strictly have to (i.e. their bits are the same no matter which version of murmur3 we use). But culling these out is impossible, since we'd have to perform the full tree-diff, which is the same effort as computing the Bloom filter from scratch. But because we can cache our results in each tree's flag bits, we can often avoid recomputing many filters, thereby reducing the time it takes to run $ git commit-graph write --changed-paths --reachable when upgrading from v1 to v2 Bloom filters. To benchmark this, let's generate a commit-graph in linux.git with v1 changed-paths in generation order[^1]: $ git clone git@github.com:torvalds/linux.git $ cd linux $ git commit-graph write --reachable --changed-paths $ graph=".git/objects/info/commit-graph" $ mv $graph{,.bak} Then let's time how long it takes to go from v1 to v2 filters (with and without the upgrade path enabled), resetting the state of the commit-graph each time: $ git config commitGraph.changedPathsVersion 2 $ hyperfine -p 'cp -f $graph.bak $graph' -L v 0,1 \ 'GIT_TEST_UPGRADE_BLOOM_FILTERS={v} git.compile commit-graph write --reachable --changed-paths' On linux.git (where there aren't any non-ASCII paths), the timings indicate that this patch represents a speed-up over recomputing all Bloom filters from scratch: Benchmark 1: GIT_TEST_UPGRADE_BLOOM_FILTERS=0 git.compile commit-graph write --reachable --changed-paths Time (mean ± σ): 124.873 s ± 0.316 s [User: 124.081 s, System: 0.643 s] Range (min … max): 124.621 s … 125.227 s 3 runs Benchmark 2: GIT_TEST_UPGRADE_BLOOM_FILTERS=1 git.compile commit-graph write --reachable --changed-paths Time (mean ± σ): 79.271 s ± 0.163 s [User: 74.611 s, System: 4.521 s] Range (min … max): 79.112 s … 79.437 s 3 runs Summary 'GIT_TEST_UPGRADE_BLOOM_FILTERS=1 git.compile commit-graph write --reachable --changed-paths' ran 1.58 ± 0.01 times faster than 'GIT_TEST_UPGRADE_BLOOM_FILTERS=0 git.compile commit-graph write --reachable --changed-paths' On git.git, we do have some non-ASCII paths, giving us a more modest improvement from 4.163 seconds to 3.348 seconds, for a 1.24x speed-up. On my machine, the stats for git.git are: - 8,285 Bloom filters computed from scratch - 10 Bloom filters generated as empty - 4 Bloom filters generated as truncated due to too many changed paths - 65,114 Bloom filters were reused when transitioning from v1 to v2. [^1]: Note that this is is important, since `--stdin-packs` or `--stdin-commits` orders commits in the commit-graph by their pack position (with `--stdin-packs`) or in the raw input (with `--stdin-commits`). Since we compute Bloom filters in the same order that commits appear in the graph, we must see a commit's (first) parent before we process the commit itself. This is only guaranteed to happen when sorting commits by their generation number. Signed-off-by: Taylor Blau --- bloom.c | 90 ++++++++++++++++++++++++++++++++++++++++++-- bloom.h | 1 + commit-graph.c | 5 +++ object.h | 1 + t/t4216-log-bloom.sh | 35 ++++++++++++++++- 5 files changed, 127 insertions(+), 5 deletions(-) diff --git a/bloom.c b/bloom.c index 739fa093ba..24dd874e46 100644 --- a/bloom.c +++ b/bloom.c @@ -7,6 +7,9 @@ #include "commit-graph.h" #include "commit.h" #include "commit-slab.h" +#include "tree.h" +#include "tree-walk.h" +#include "config.h" define_commit_slab(bloom_filter_slab, struct bloom_filter); @@ -250,6 +253,73 @@ static void init_truncated_large_filter(struct bloom_filter *filter, filter->version = version; } +#define VISITED (1u<<21) +#define HIGH_BITS (1u<<22) + +static int has_entries_with_high_bit(struct repository *r, struct tree *t) +{ + if (parse_tree(t)) + return 1; + + if (!(t->object.flags & VISITED)) { + struct tree_desc desc; + struct name_entry entry; + + init_tree_desc(&desc, t->buffer, t->size); + while (tree_entry(&desc, &entry)) { + size_t i; + for (i = 0; i < entry.pathlen; i++) { + if (entry.path[i] & 0x80) { + t->object.flags |= HIGH_BITS; + goto done; + } + } + + if (S_ISDIR(entry.mode)) { + struct tree *sub = lookup_tree(r, &entry.oid); + if (sub && has_entries_with_high_bit(r, sub)) { + t->object.flags |= HIGH_BITS; + goto done; + } + } + + } + +done: + t->object.flags |= VISITED; + } + + return !!(t->object.flags & HIGH_BITS); +} + +static int commit_tree_has_high_bit_paths(struct repository *r, + struct commit *c) +{ + struct tree *t; + if (repo_parse_commit(r, c)) + return 1; + t = repo_get_commit_tree(r, c); + if (!t) + return 1; + return has_entries_with_high_bit(r, t); +} + +static struct bloom_filter *upgrade_filter(struct repository *r, struct commit *c, + struct bloom_filter *filter, + int hash_version) +{ + struct commit_list *p = c->parents; + if (commit_tree_has_high_bit_paths(r, c)) + return NULL; + + if (p && commit_tree_has_high_bit_paths(r, p->item)) + return NULL; + + filter->version = hash_version; + + return filter; +} + struct bloom_filter *get_bloom_filter(struct repository *r, struct commit *c) { struct bloom_filter *filter; @@ -292,9 +362,23 @@ struct bloom_filter *get_or_compute_bloom_filter(struct repository *r, filter, graph_pos); } - if ((filter->data && filter->len) && - (!settings || settings->hash_version == filter->version)) - return filter; + if (filter->data && filter->len) { + struct bloom_filter *upgrade; + if (!settings || settings->hash_version == filter->version) + return filter; + + /* version mismatch, see if we can upgrade */ + if (compute_if_not_present && + git_env_bool("GIT_TEST_UPGRADE_BLOOM_FILTERS", 1)) { + upgrade = upgrade_filter(r, c, filter, + settings->hash_version); + if (upgrade) { + if (computed) + *computed |= BLOOM_UPGRADED; + return upgrade; + } + } + } if (!compute_if_not_present) return NULL; diff --git a/bloom.h b/bloom.h index bfe389e29c..e3a9b68905 100644 --- a/bloom.h +++ b/bloom.h @@ -102,6 +102,7 @@ enum bloom_filter_computed { BLOOM_COMPUTED = (1 << 1), BLOOM_TRUNC_LARGE = (1 << 2), BLOOM_TRUNC_EMPTY = (1 << 3), + BLOOM_UPGRADED = (1 << 4), }; struct bloom_filter *get_or_compute_bloom_filter(struct repository *r, diff --git a/commit-graph.c b/commit-graph.c index e0fc62e110..571f38335a 100644 --- a/commit-graph.c +++ b/commit-graph.c @@ -1109,6 +1109,7 @@ struct write_commit_graph_context { int count_bloom_filter_not_computed; int count_bloom_filter_trunc_empty; int count_bloom_filter_trunc_large; + int count_bloom_filter_upgraded; }; static int write_graph_chunk_fanout(struct hashfile *f, @@ -1716,6 +1717,8 @@ static void trace2_bloom_filter_write_statistics(struct write_commit_graph_conte ctx->count_bloom_filter_trunc_empty); trace2_data_intmax("commit-graph", ctx->r, "filter-trunc-large", ctx->count_bloom_filter_trunc_large); + trace2_data_intmax("commit-graph", ctx->r, "filter-upgraded", + ctx->count_bloom_filter_upgraded); } static void compute_bloom_filters(struct write_commit_graph_context *ctx) @@ -1757,6 +1760,8 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx) ctx->count_bloom_filter_trunc_empty++; if (computed & BLOOM_TRUNC_LARGE) ctx->count_bloom_filter_trunc_large++; + } else if (computed & BLOOM_UPGRADED) { + ctx->count_bloom_filter_upgraded++; } else if (computed & BLOOM_NOT_COMPUTED) ctx->count_bloom_filter_not_computed++; ctx->total_bloom_filter_data_size += filter diff --git a/object.h b/object.h index db25714b4e..2e5e08725f 100644 --- a/object.h +++ b/object.h @@ -75,6 +75,7 @@ void object_array_init(struct object_array *array); * commit-reach.c: 16-----19 * sha1-name.c: 20 * list-objects-filter.c: 21 + * bloom.c: 2122 * builtin/fsck.c: 0--3 * builtin/gc.c: 0 * builtin/index-pack.c: 2021 diff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh index 8f8b5d4966..a321d8d713 100755 --- a/t/t4216-log-bloom.sh +++ b/t/t4216-log-bloom.sh @@ -221,6 +221,10 @@ test_filter_trunc_large () { grep "\"key\":\"filter-trunc-large\",\"value\":\"$1\"" $2 } +test_filter_upgraded () { + grep "\"key\":\"filter-upgraded\",\"value\":\"$1\"" $2 +} + test_expect_success 'correctly report changes over limit' ' git init limits && ( @@ -629,10 +633,19 @@ test_expect_success 'when writing another commit graph, preserve existing versio test_expect_success 'when writing commit graph, do not reuse changed-path of another version' ' git init doublewrite && test_commit -C doublewrite c "$CENT" && + git -C doublewrite config --add commitgraph.changedPathsVersion 1 && - git -C doublewrite commit-graph write --reachable --changed-paths && + GIT_TRACE2_EVENT="$(pwd)/trace2.txt" \ + git -C doublewrite commit-graph write --reachable --changed-paths && + test_filter_computed 1 trace2.txt && + test_filter_upgraded 0 trace2.txt && + git -C doublewrite config --add commitgraph.changedPathsVersion 2 && - git -C doublewrite commit-graph write --reachable --changed-paths && + GIT_TRACE2_EVENT="$(pwd)/trace2.txt" \ + git -C doublewrite commit-graph write --reachable --changed-paths && + test_filter_computed 1 trace2.txt && + test_filter_upgraded 0 trace2.txt && + ( cd doublewrite && echo "c01f" >expect && @@ -641,4 +654,22 @@ test_expect_success 'when writing commit graph, do not reuse changed-path of ano ) ' +test_expect_success 'when writing commit graph, reuse changed-path of another version where possible' ' + git init upgrade && + + test_commit -C upgrade base no-high-bits && + + git -C upgrade config --add commitgraph.changedPathsVersion 1 && + GIT_TRACE2_EVENT="$(pwd)/trace2.txt" \ + git -C upgrade commit-graph write --reachable --changed-paths && + test_filter_computed 1 trace2.txt && + test_filter_upgraded 0 trace2.txt && + + git -C upgrade config --add commitgraph.changedPathsVersion 2 && + GIT_TRACE2_EVENT="$(pwd)/trace2.txt" \ + git -C upgrade commit-graph write --reachable --changed-paths && + test_filter_computed 0 trace2.txt && + test_filter_upgraded 1 trace2.txt +' + test_done From patchwork Tue Oct 10 20:34:11 2023 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Taylor Blau X-Patchwork-Id: 13416057 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by smtp.lore.kernel.org (Postfix) with ESMTP id A536BCD8CB4 for ; Tue, 10 Oct 2023 20:34:33 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1343941AbjJJUec (ORCPT ); Tue, 10 Oct 2023 16:34:32 -0400 Received: from lindbergh.monkeyblade.net ([23.128.96.19]:59088 "EHLO lindbergh.monkeyblade.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S229644AbjJJUeQ (ORCPT ); Tue, 10 Oct 2023 16:34:16 -0400 Received: from mail-qk1-x72b.google.com (mail-qk1-x72b.google.com [IPv6:2607:f8b0:4864:20::72b]) by lindbergh.monkeyblade.net (Postfix) with ESMTPS id 5C4398E for ; Tue, 10 Oct 2023 13:34:14 -0700 (PDT) Received: by mail-qk1-x72b.google.com with SMTP id af79cd13be357-7741c5bac51so359153385a.1 for ; Tue, 10 Oct 2023 13:34:14 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ttaylorr-com.20230601.gappssmtp.com; s=20230601; t=1696970053; x=1697574853; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=5VrBPDmOjWxdllfTw2xH8Ary/m77UeoLJiNh2idn7Tg=; b=LZrg42pef0uPyMbDRKR9Xy8M2GsZuFS9OPlVwQefIGqV1aYPcTzkXVz+v3HyS0v3Ix eU4PQwQEnniS+9cb2EmqGDOIbYS6ew2RukcrdZlxvtx0gGTugvo8dHaUgl4eZyxonPn/ pzHt3wXxo5RKOH3UuZj9b/cpJPHId6d4gM9VCTyuCyA8jj9N+trFmB+y2QfHIy4iAsBI RUhCcSc2TrElNX8qygmvoISB0kFsiztsxrbUaq5ZAJ16XJxbbuQWks6be9DKEIsyoXns NDMjIs2AwKGOxiZcZlE/D+fYOqB8qllKei153JiEo/l02cXHNVeUQoR/NJaFci2y0lVG igEw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1696970053; x=1697574853; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=5VrBPDmOjWxdllfTw2xH8Ary/m77UeoLJiNh2idn7Tg=; b=iBEwYxsicRwzWnyjC5V/AC8/JgD2dl4A2qvt2zC2cHxuEEgQpNsoaSgl5NxT7moyIu 42fAFKeZ1XevRX+rH5QsijGwaZ1vw3FZG87VYM7RfxxFWG+GSRHqXt9U5TODxfSw+/jW OvQZPj/Hi5tiDZ4wZNssMslva10W2bl1P6qPqOVDsCnk9jJnc8vuaCvyi1j1QwluECrf BGKk6dTUq7zJ4aa8sEV9sxK9VAbhIOVBw/7+f+DDMpAygqv3Sch/1Ia7tYLO0LO71H4P zAc3HoDwQOQrhkWhAhDNynd22TbeEo332t+jHayy2cNQgwRfgUQDQXlLZg7Zx233MsGN 6Rqg== X-Gm-Message-State: AOJu0YwZv+AOldLWgwiIroBVl5Pm7J+oBkbOrdpFQtjk3LBh3yKxQK7s pjJ9neV03Cg9J5a7CmI0vSo9T2u6I7ubAMTjYtvWAA== X-Google-Smtp-Source: AGHT+IG5Y2KUKyp7QWO6BA3VoqcvZYLBSh+/8URNymrkL/hyRMuJug0z2DT2uxvCYxz7sp+f9MKrMw== X-Received: by 2002:a05:620a:2952:b0:773:bdb3:1318 with SMTP id n18-20020a05620a295200b00773bdb31318mr22105446qkp.15.1696970052883; Tue, 10 Oct 2023 13:34:12 -0700 (PDT) Received: from localhost (104-178-186-189.lightspeed.milwwi.sbcglobal.net. [104.178.186.189]) by smtp.gmail.com with ESMTPSA id pj40-20020a05620a1da800b007743360b3fasm4632735qkn.34.2023.10.10.13.34.12 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 10 Oct 2023 13:34:12 -0700 (PDT) Date: Tue, 10 Oct 2023 16:34:11 -0400 From: Taylor Blau To: git@vger.kernel.org Cc: Jonathan Tan , Junio C Hamano , Jeff King , SZEDER =?utf-8?b?R8OhYm9y?= Subject: [PATCH v3 17/17] bloom: introduce `deinit_bloom_filters()` Message-ID: References: MIME-Version: 1.0 Content-Disposition: inline In-Reply-To: Precedence: bulk List-ID: X-Mailing-List: git@vger.kernel.org After we are done using Bloom filters, we do not currently clean up any memory allocated by the commit slab used to store those filters in the first place. Besides the bloom_filter structures themselves, there is mostly nothing to free() in the first place, since in the read-only path all Bloom filter's `data` members point to a memory mapped region in the commit-graph file itself. But when generating Bloom filters from scratch (or initializing truncated filters) we allocate additional memory to store the filter's data. Keep track of when we need to free() this additional chunk of memory by using an extra pointer `to_free`. Most of the time this will be NULL (indicating that we are representing an existing Bloom filter stored in a memory mapped region). When it is non-NULL, free it before discarding the Bloom filters slab. Suggested-by: Jonathan Tan Signed-off-by: Taylor Blau Signed-off-by: Junio C Hamano Signed-off-by: Taylor Blau --- bloom.c | 16 +++++++++++++++- bloom.h | 3 +++ commit-graph.c | 4 ++++ 3 files changed, 22 insertions(+), 1 deletion(-) diff --git a/bloom.c b/bloom.c index 24dd874e46..ff131893cd 100644 --- a/bloom.c +++ b/bloom.c @@ -59,6 +59,7 @@ int load_bloom_filter_from_graph(struct commit_graph *g, sizeof(unsigned char) * start_index + BLOOMDATA_CHUNK_HEADER_SIZE); filter->version = g->bloom_filter_settings->hash_version; + filter->to_free = NULL; return 1; } @@ -231,6 +232,18 @@ void init_bloom_filters(void) init_bloom_filter_slab(&bloom_filters); } +static void free_one_bloom_filter(struct bloom_filter *filter) +{ + if (!filter) + return; + free(filter->to_free); +} + +void deinit_bloom_filters(void) +{ + deep_clear_bloom_filter_slab(&bloom_filters, free_one_bloom_filter); +} + static int pathmap_cmp(const void *hashmap_cmp_fn_data UNUSED, const struct hashmap_entry *eptr, const struct hashmap_entry *entry_or_key, @@ -247,7 +260,7 @@ static int pathmap_cmp(const void *hashmap_cmp_fn_data UNUSED, static void init_truncated_large_filter(struct bloom_filter *filter, int version) { - filter->data = xmalloc(1); + filter->data = filter->to_free = xmalloc(1); filter->data[0] = 0xFF; filter->len = 1; filter->version = version; @@ -449,6 +462,7 @@ struct bloom_filter *get_or_compute_bloom_filter(struct repository *r, filter->len = 1; } CALLOC_ARRAY(filter->data, filter->len); + filter->to_free = filter->data; hashmap_for_each_entry(&pathmap, &iter, e, entry) { struct bloom_key key; diff --git a/bloom.h b/bloom.h index e3a9b68905..d20e64bfbb 100644 --- a/bloom.h +++ b/bloom.h @@ -56,6 +56,8 @@ struct bloom_filter { unsigned char *data; size_t len; int version; + + void *to_free; }; /* @@ -96,6 +98,7 @@ void add_key_to_filter(const struct bloom_key *key, const struct bloom_filter_settings *settings); void init_bloom_filters(void); +void deinit_bloom_filters(void); enum bloom_filter_computed { BLOOM_NOT_COMPUTED = (1 << 0), diff --git a/commit-graph.c b/commit-graph.c index 571f38335a..7ccec429cb 100644 --- a/commit-graph.c +++ b/commit-graph.c @@ -787,6 +787,7 @@ static void close_commit_graph_one(struct commit_graph *g) void close_commit_graph(struct raw_object_store *o) { close_commit_graph_one(o->commit_graph); + deinit_bloom_filters(); o->commit_graph = NULL; } @@ -2588,6 +2589,9 @@ int write_commit_graph(struct object_directory *odb, res = write_commit_graph_file(ctx); + if (ctx->changed_paths) + deinit_bloom_filters(); + if (ctx->split) mark_commit_graphs(ctx);