[v2,2/2] f2fs: Support case-insensitive file name lookups

Modeled after commit b886ee3e778e ("ext4: Support case-insensitive file
name lookups")

"""
This patch implements the actual support for case-insensitive file name
lookups in f2fs, based on the feature bit and the encoding stored in the
superblock.

A filesystem that has the casefold feature set is able to configure
directories with the +F (F2FS_CASEFOLD_FL) attribute, enabling lookups
to succeed in that directory in a case-insensitive fashion, i.e: match
a directory entry even if the name used by userspace is not a byte per
byte match with the disk name, but is an equivalent case-insensitive
version of the Unicode string.  This operation is called a
case-insensitive file name lookup.

The feature is configured as an inode attribute applied to directories
and inherited by its children.  This attribute can only be enabled on
empty directories for filesystems that support the encoding feature,
thus preventing collision of file names that only differ by case.

* dcache handling:

For a +F directory, F2Fs only stores the first equivalent name dentry
used in the dcache. This is done to prevent unintentional duplication of
dentries in the dcache, while also allowing the VFS code to quickly find
the right entry in the cache despite which equivalent string was used in
a previous lookup, without having to resort to ->lookup().

d_hash() of casefolded directories is implemented as the hash of the
casefolded string, such that we always have a well-known bucket for all
the equivalencies of the same string. d_compare() uses the
utf8_strncasecmp() infrastructure, which handles the comparison of
equivalent, same case, names as well.

For now, negative lookups are not inserted in the dcache, since they
would need to be invalidated anyway, because we can't trust missing file
dentries.  This is bad for performance but requires some leveraging of
the vfs layer to fix.  We can live without that for now, and so does
everyone else.

* on-disk data:

Despite using a specific version of the name as the internal
representation within the dcache, the name stored and fetched from the
disk is a byte-per-byte match with what the user requested, making this
implementation 'name-preserving'. i.e. no actual information is lost
when writing to storage.

DX is supported by modifying the hashes used in +F directories to make
them case/encoding-aware.  The new disk hashes are calculated as the
hash of the full casefolded string, instead of the string directly.
This allows us to efficiently search for file names in the htree without
requiring the user to provide an exact name.

* Dealing with invalid sequences:

By default, when a invalid UTF-8 sequence is identified, ext4 will treat
it as an opaque byte sequence, ignoring the encoding and reverting to
the old behavior for that unique file.  This means that case-insensitive
file name lookup will not work only for that file.  An optional bit can
be set in the superblock telling the filesystem code and userspace tools
to enforce the encoding.  When that optional bit is set, any attempt to
create a file name using an invalid UTF-8 sequence will fail and return
an error to userspace.

* Normalization algorithm:

The UTF-8 algorithms used to compare strings in f2fs is implemented
in fs/unicode, and is based on a previous version developed by
SGI.  It implements the Canonical decomposition (NFD) algorithm
described by the Unicode specification 12.1, or higher, combined with
the elimination of ignorable code points (NFDi) and full
case-folding (CF) as documented in fs/unicode/utf8_norm.c.

NFD seems to be the best normalization method for F2FS because:

  - It has a lower cost than NFC/NFKC (which requires
    decomposing to NFD as an intermediary step)
  - It doesn't eliminate important semantic meaning like
    compatibility decompositions.

Although:

- This implementation is not completely linguistic accurate, because
different languages have conflicting rules, which would require the
specialization of the filesystem to a given locale, which brings all
sorts of problems for removable media and for users who use more than
one language.
"""

Signed-off-by: Daniel Rosenberg <drosen@google.com>
---
 fs/f2fs/dir.c    | 133 ++++++++++++++++++++++++++++++++++++++++++-----
 fs/f2fs/f2fs.h   |  18 +++++--
 fs/f2fs/file.c   |  10 +++-
 fs/f2fs/hash.c   |  34 +++++++++++-
 fs/f2fs/inline.c |   6 +--
 fs/f2fs/inode.c  |   4 +-
 fs/f2fs/namei.c  |  21 ++++++++
 fs/f2fs/super.c  |   5 ++
 8 files changed, 208 insertions(+), 23 deletions(-)

Message ID	20190717031408.114104-3-drosen@google.com (mailing list archive)
State	New, archived
Headers	show Return-Path: <linux-fsdevel-owner@kernel.org> Received: from mail.wl.linuxfoundation.org (pdx-wl-mail.web.codeaurora.org [172.30.200.125]) by pdx-korg-patchwork-2.web.codeaurora.org (Postfix) with ESMTP id ADE93912 for <patchwork-linux-fsdevel@patchwork.kernel.org>; Wed, 17 Jul 2019 03:14:48 +0000 (UTC) Received: from mail.wl.linuxfoundation.org (localhost [127.0.0.1]) by mail.wl.linuxfoundation.org (Postfix) with ESMTP id 99922286FE for <patchwork-linux-fsdevel@patchwork.kernel.org>; Wed, 17 Jul 2019 03:14:48 +0000 (UTC) Received: by mail.wl.linuxfoundation.org (Postfix, from userid 486) id 8B4E728702; Wed, 17 Jul 2019 03:14:48 +0000 (UTC) X-Spam-Checker-Version: SpamAssassin 3.3.1 (2010-03-16) on pdx-wl-mail.web.codeaurora.org X-Spam-Level: X-Spam-Status: No, score=-15.5 required=2.0 tests=BAYES_00,DKIM_SIGNED, DKIM_VALID,DKIM_VALID_AU,MAILING_LIST_MULTI,RCVD_IN_DNSWL_HI, USER_IN_DEF_DKIM_WL autolearn=ham version=3.3.1 Received: from vger.kernel.org (vger.kernel.org [209.132.180.67]) by mail.wl.linuxfoundation.org (Postfix) with ESMTP id 4F387286FE for <patchwork-linux-fsdevel@patchwork.kernel.org>; Wed, 17 Jul 2019 03:14:47 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1728799AbfGQDOq (ORCPT <rfc822;patchwork-linux-fsdevel@patchwork.kernel.org>); Tue, 16 Jul 2019 23:14:46 -0400 Received: from mail-vk1-f202.google.com ([209.85.221.202]:34688 "EHLO mail-vk1-f202.google.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1728766AbfGQDOp (ORCPT <rfc822;linux-fsdevel@vger.kernel.org>); Tue, 16 Jul 2019 23:14:45 -0400 Received: by mail-vk1-f202.google.com with SMTP id g68so7397565vkb.1 for <linux-fsdevel@vger.kernel.org>; Tue, 16 Jul 2019 20:14:44 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20161025; h=date:in-reply-to:message-id:mime-version:references:subject:from:to :cc; bh=QR1Zb7qY1Bzd2Adyvt2Lbuj1XsNeCB4FTQJmJ+6mSVA=; b=MXv8BVjHHvCmvZ+pCZjOJbKY+vjqL+zW3aNP3OIzGKUBwWWdoDTZw1LNW5vPDIW/2P 0plvgj7rlJr8B7yud5BGcNNb/oAXlJgkw+f1gF9FPcP5vOpd8Hy+6RsDaHoGY6At1VSt cNloPof76a0Xgkv3q12ciGl6rn7NuptByiazmcQsPnOzuRTTRv8dN5+XaLJ1MYkmBsCY RtiJXa3mwyw02IiQ6/l1SRyuGgI0fpD8OFHjmlzPk7sVEc8v/Eyd4PQoG/xfraY3y0g/ 8JJKRwIun0G+PkD0A/ZHXslazFVCDN0sk+N+DHLy1fB3PkCcySh8MIqkP0OgSDIlvLa9 LsgQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20161025; h=x-gm-message-state:date:in-reply-to:message-id:mime-version :references:subject:from:to:cc; bh=QR1Zb7qY1Bzd2Adyvt2Lbuj1XsNeCB4FTQJmJ+6mSVA=; b=H+723FTiAqc9VXVzal85XEwElNyZQKLm02CLuEnQr/39ibut3PUa8fPp0QEA7BAmyl xQ+eDLAPHCal6lH51bEcQuyxqpMZ77c6nLofj/xmyaWZV0wEBNEwhscbVhetA2tY275o bVspoleJ5FMCkQ2CW/8TfhqlVr2xnJ8lTkg79BfDJn7d5QICzCXmKhzao8tbnro5a4ak x7793V8ZdkDs6++WETNm/6Xru92ch+jKoFQTUZp+Uq0YytirxdKrxVCk/obagNpS6Srp jPLsrrIloGu7GnvWcb9vUJLU+V4N82g37lCKesXYMfhTnSGaXxtjIRtX8MReqeQR6Gus HVYg== X-Gm-Message-State: APjAAAXP6zoTSoa6yR8weRq5MSV23aeBWGyx7OL5hIO72TP3KyqVuXah 8KTjKX/m2xPvOkmjsT1Su6sqLkLAtok= X-Google-Smtp-Source: APXvYqw8P8fxHqYF9yREUrV6ojI20C+a9GPYcUS4dpehZs+nMDmIBBeQbl0sP48KllNd5saMeaFnbM6H7WY= X-Received: by 2002:ab0:2ead:: with SMTP id y13mr7571416uay.13.1563333283795; Tue, 16 Jul 2019 20:14:43 -0700 (PDT) Date: Tue, 16 Jul 2019 20:14:08 -0700 In-Reply-To: <20190717031408.114104-1-drosen@google.com> Message-Id: <20190717031408.114104-3-drosen@google.com> Mime-Version: 1.0 References: <20190717031408.114104-1-drosen@google.com> X-Mailer: git-send-email 2.22.0.510.g264f2c817a-goog Subject: [PATCH v2 2/2] f2fs: Support case-insensitive file name lookups From: Daniel Rosenberg <drosen@google.com> To: Jaegeuk Kim <jaegeuk@kernel.org>, Chao Yu <yuchao0@huawei.com>, Jonathan Corbet <corbet@lwn.net>, linux-f2fs-devel@lists.sourceforge.net Cc: linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-fsdevel@vger.kernel.org, kernel-team@android.com, Daniel Rosenberg <drosen@google.com> Content-Type: text/plain; charset="UTF-8" Sender: linux-fsdevel-owner@vger.kernel.org Precedence: bulk List-ID: <linux-fsdevel.vger.kernel.org> X-Mailing-List: linux-fsdevel@vger.kernel.org X-Virus-Scanned: ClamAV using ClamSMTP
Series	Casefolding in F2FS \| expand [v2,0/2] Casefolding in F2FS [v2,1/2] f2fs: include charset encoding information in the superblock [v2,2/2] f2fs: Support case-insensitive file name lookups

[v2,2/2] f2fs: Support case-insensitive file name lookups

Commit Message

Comments

Patch