[2/2] drm/i915/guc: Add delay to disable scheduling after pin count goes to zero

From: Matthew Brost <matthew.brost@intel.com>

From: Matthew Brost <matthew.brost@intel.com>

Add a delay, configurable via debugs (default 100ms), to disable
scheduling of a context after the pin count goes to zero. Disable
scheduling is somewhat costly operation so the idea is a delay allows
the resubmit something before doing this operation. This delay is only
done if the context isn't close and less than 3/4 of the guc_ids are in
use.

As temporary WA disable this feature for the selftests. Selftests are
very timing sensitive and any change in timing can cause failure. A
follow up patch will fixup the selftests to understand this delay.

Alan Previn: Matt Brost first introduced this series back in Oct 2021.
However no real world workload with measured performance impact was
available to prove the intended results. Today, this series is being
republished in response to a real world workload that benefited greatly
from it along with measured performance improvement.

Workload description: 36 containers were created on a DG2 device where
each container was performing a combination of 720p 3d game rendering
and 30fps video encoding. The workload density was configured in way
that guaranteed each container to ALWAYS be able to render and
encode no less than 30fps with a predefined maximum render + encode
latency time. That means that the totality of all 36 containers and its
workloads were not saturating the utilized hw engines to its max
(in order to maintain just enough headrooom to meet the minimum fps and
latencies of incoming container submissions).

Problem statement: It was observed that the CPU utilization of the CPU
core that was pinned to i915 soft IRQ work was experiencing severe load.
Using tracelogs and an instrumentation patch to count specific i915 IRQ
events, it was confirmed that the majority of the CPU cycles were caused
by the gen11_other_irq_handler() -> guc_irq_handler() code path. The vast
majority of the cycles was determined to be processing a specific G2H IRQ
which was INTEL_GUC_ACTION_SCHED_CONTEXT_MODE_DONE. This IRQ is send by
the GuC in response to the i915 KMD sending the H2G requests
INTEL_GUC_ACTION_SCHED_CONTEXT_MODE_SET to the GuC. That request is sent
when the context is idle to unpin the context from any GuC access. The
high CPU utilization % symptom was limiting the density scaling.

Root Cause Analysis: Because the incoming execution buffers were spread
across 36 different containers (each with multiple contexts) but the
system in totality was NOT saturated to the max, it was assumed that each
context was constantly idling between submissions. This was causing thrashing
of unpinning a context from GuC at one moment, followed by repinning it
due to incoming workload the very next moment. Both of these event-pairs
were being triggered across multiple contexts per container, across all
containers at the rate of > 30 times per sec per context.

Metrics: When running this workload without this patch, we measured an average
of ~69K INTEL_GUC_ACTION_SCHED_CONTEXT_MODE_DONE events every 10 seconds or
~10 million times over ~25+ mins. With this patch, the count reduced to ~480
every 10 seconds or about ~28K over ~10 mins. The improvement observed is
~99% for the average counts per 10 seconds.

Signed-off-by: Matthew Brost <matthew.brost@intel.com>
Acked-by: Alan Previn <alan.previn.teres.alexis@intel.com>
---
 drivers/gpu/drm/i915/gem/i915_gem_context.c   |   2 +-
 drivers/gpu/drm/i915/gt/intel_context.h       |   9 ++
 drivers/gpu/drm/i915/gt/intel_context_types.h |   8 ++
 drivers/gpu/drm/i915/gt/uc/intel_guc.h        |  10 ++
 .../gpu/drm/i915/gt/uc/intel_guc_debugfs.c    |  28 ++++
 .../gpu/drm/i915/gt/uc/intel_guc_submission.c | 132 ++++++++++++++----
 drivers/gpu/drm/i915/i915_selftest.h          |   2 +
 drivers/gpu/drm/i915/i915_trace.h             |  10 ++
 8 files changed, 175 insertions(+), 26 deletions(-)

Message ID	20220628055130.1117146-3-alan.previn.teres.alexis@intel.com (mailing list archive)
State	New, archived
Headers	show Return-Path: <intel-gfx-bounces@lists.freedesktop.org> X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from gabe.freedesktop.org (gabe.freedesktop.org [131.252.210.177]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.lore.kernel.org (Postfix) with ESMTPS id 18DBCC43334 for <intel-gfx@archiver.kernel.org>; Tue, 28 Jun 2022 05:50:51 +0000 (UTC) Received: from gabe.freedesktop.org (localhost [127.0.0.1]) by gabe.freedesktop.org (Postfix) with ESMTP id CA6051139B5; Tue, 28 Jun 2022 05:50:49 +0000 (UTC) Received: from mga03.intel.com (mga03.intel.com [134.134.136.65]) by gabe.freedesktop.org (Postfix) with ESMTPS id 5D2B9113994 for <intel-gfx@lists.freedesktop.org>; Tue, 28 Jun 2022 05:50:48 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1656395448; x=1687931448; h=from:to:subject:date:message-id:in-reply-to:references: mime-version:content-transfer-encoding; bh=lu4SosXBH8V4BbfOpChxE8v5MQ31fN/XcuplAzsRbMA=; b=e3D0SPB0HB41tn+WPfcf85qitpdNkOo0Iv79ISlxdsIZRxSFM/QJI4e8 Xa1HlEjJs00ePrCP6iD/2NUsUHTEN0Edb6oqVFyFu6+gUzzw1BWl3Eum4 90BOcTqfdoV9HKHtwD25uAx25iM/7GpSzrEqEm7Yi0a5DV3adXVUwwrau seVtqylTuKuwxq9Qh5UO1SHoYkHz1pPqAFC38KxfAEhOcNEA5GVft3Bx2 /3Nn2zGBSyXzlFo/LUzE5X5B86y+4934f22959L+LkouIEMfBzpr1m60o WJODmjoi2psOHCy2tLfcrTXvG8bzp7gfbZZO7pEG5jHWFF2ESVmX6ZQy9 g==; X-IronPort-AV: E=McAfee;i="6400,9594,10391"; a="282733269" X-IronPort-AV: E=Sophos;i="5.92,227,1650956400"; d="scan'208";a="282733269" Received: from fmsmga004.fm.intel.com ([10.253.24.48]) by orsmga103.jf.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 27 Jun 2022 22:50:47 -0700 X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="5.92,227,1650956400"; d="scan'208";a="657993731" Received: from aalteres-desk.fm.intel.com ([10.80.57.53]) by fmsmga004.fm.intel.com with ESMTP; 27 Jun 2022 22:50:47 -0700 From: Alan Previn <alan.previn.teres.alexis@intel.com> To: intel-gfx@lists.freedesktop.org Date: Mon, 27 Jun 2022 22:51:30 -0700 Message-Id: <20220628055130.1117146-3-alan.previn.teres.alexis@intel.com> X-Mailer: git-send-email 2.25.1 In-Reply-To: <20220628055130.1117146-1-alan.previn.teres.alexis@intel.com> References: <20220628055130.1117146-1-alan.previn.teres.alexis@intel.com> MIME-Version: 1.0 Content-Transfer-Encoding: 8bit Subject: [Intel-gfx] [Intel-gfx 2/2] drm/i915/guc: Add delay to disable scheduling after pin count goes to zero X-BeenThere: intel-gfx@lists.freedesktop.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: Intel graphics driver community testing & development <intel-gfx.lists.freedesktop.org> List-Unsubscribe: <https://lists.freedesktop.org/mailman/options/intel-gfx>, <mailto:intel-gfx-request@lists.freedesktop.org?subject=unsubscribe> List-Archive: <https://lists.freedesktop.org/archives/intel-gfx> List-Post: <mailto:intel-gfx@lists.freedesktop.org> List-Help: <mailto:intel-gfx-request@lists.freedesktop.org?subject=help> List-Subscribe: <https://lists.freedesktop.org/mailman/listinfo/intel-gfx>, <mailto:intel-gfx-request@lists.freedesktop.org?subject=subscribe> Errors-To: intel-gfx-bounces@lists.freedesktop.org Sender: "Intel-gfx" <intel-gfx-bounces@lists.freedesktop.org>
Series	Delay disabling scheduling on a context \| expand [0/2] Delay disabling scheduling on a context [1/2] drm/i915/selftests: Use correct selfest calls for live tests [2/2] drm/i915/guc: Add delay to disable scheduling after pin count goes to zero

[2/2] drm/i915/guc: Add delay to disable scheduling after pin count goes to zero

Commit Message

Comments

Patch