From patchwork Fri Apr  8 07:35:25 2016
Content-Type: text/plain; charset="utf-8"
MIME-Version: 1.0
Content-Transfer-Encoding: 8bit
X-Patchwork-Submitter: Dario Faggioli <dario.faggioli@citrix.com>
X-Patchwork-Id: 8781641
Return-Path: <xen-devel-bounces@lists.xen.org>
X-Original-To: patchwork-xen-devel@patchwork.kernel.org
Delivered-To: patchwork-parsemail@patchwork2.web.kernel.org
Received: from mail.kernel.org (mail.kernel.org [198.145.29.136])
	by patchwork2.web.kernel.org (Postfix) with ESMTP id D3C83C0553
	for <patchwork-xen-devel@patchwork.kernel.org>;
	Fri,  8 Apr 2016 07:38:28 +0000 (UTC)
Received: from mail.kernel.org (localhost [127.0.0.1])
	by mail.kernel.org (Postfix) with ESMTP id 377CA2026D
	for <patchwork-xen-devel@patchwork.kernel.org>;
	Fri,  8 Apr 2016 07:38:27 +0000 (UTC)
Received: from lists.xenproject.org (lists.xenproject.org [192.237.175.120])
	(using TLSv1.2 with cipher AES128-GCM-SHA256 (128/128 bits))
	(No client certificate requested)
	by mail.kernel.org (Postfix) with ESMTPS id 72380201B9
	for <patchwork-xen-devel@patchwork.kernel.org>;
	Fri,  8 Apr 2016 07:38:25 +0000 (UTC)
Received: from localhost ([127.0.0.1] helo=lists.xenproject.org)
	by lists.xenproject.org with esmtp (Exim 4.84_2)
	(envelope-from <xen-devel-bounces@lists.xen.org>)
	id 1aoQxG-0008IW-6N; Fri, 08 Apr 2016 07:35:46 +0000
Received: from mail6.bemta5.messagelabs.com ([195.245.231.135])
	by lists.xenproject.org with esmtp (Exim 4.84_2)
	(envelope-from <prvs=899e9c19a=dario.faggioli@citrix.com>)
	id 1aoQxE-0008IQ-6f
	for xen-devel@lists.xenproject.org; Fri, 08 Apr 2016 07:35:44 +0000
Received: from [85.158.139.211] by server-7.bemta-5.messagelabs.com id
	1E/E0-22167-F4F57075; Fri, 08 Apr 2016 07:35:43 +0000
X-Brightmail-Tracker: H4sIAAAAAAAAA+NgFprKKsWRWlGSWpSXmKPExsXitHSDva5fPHu
	4wc1FXBbft0xmcmD0OPzhCksAYxRrZl5SfkUCa8beu1NYC7aeYqy4d2QiUwPj4z2MXYycHBIC
	IRJn//9gA7F5BQwlFk17xwxiCwsUSkxpaWEHsdkEDCTe7NjLCmKLCLhKXL99B8jm4GAWCJO4c
	ZIFJMwioCIxY/I0sFZOAQ2JD7MXM4HYQgKTGSVm/TIDsfkFJCVuffkIVsMsUClx/OM+VogT9C
	WOz1sBdYKgxMmZT1ggetUkZsy9DFXDLXH79FTmCYz8s5C0z0LSAhF3kGj585wJwtaUaN3+mx3
	C1pZYtvA1M4QdL/F31ywoW1FiSvdDqJoMifbp+6BsW4l1695DzbSR2HR1ASOELS+x/e0c5gWM
	3KsY1YtTi8pSi3Qt9ZKKMtMzSnITM3N0DQ1M9XJTi4sT01NzEpOK9ZLzczcxAqOIAQh2MK5td
	T7EKMnBpCTKKxnOHi7El5SfUpmRWJwRX1Sak1p8iFGDg0Ngwtm505mkWPLy81KVJHgd4oDqBI
	tS01Mr0jJzgHEOUyrBwaMkwmsPkuYtLkjMLc5Mh0idYlSUEuetAUkIgCQySvPg2mCp5RKjrJQ
	wLyPQUUI8BalFuZklqPKvGMU5GJWEeT1BpvBk5pXATX8FtJgJaPEFfjaQxSWJCCmpBkY/nU2/
	77hsE809nnT7VY1srWfJnOyX/ZuWKW2se1i6s+rVrOCJKzeui5O1dlq5S61vV3VD1GvBe0du/
	1ZjzV/jNTn7m2xCllJGZcnnjnkdVsfnKu9iL9o/MzFZWufU4jnWYRaxClpVdsp5Sz6sft386M
	FumWATo1+OvJGzpIzyPbiaNKss8pVYijMSDbWYi4oTAcnuQTcoAwAA
X-Env-Sender: prvs=899e9c19a=dario.faggioli@citrix.com
X-Msg-Ref: server-13.tower-206.messagelabs.com!1460100937!33385223!1
X-Originating-IP: [66.165.176.63]
X-SpamReason: No, hits=0.0 required=7.0 tests=sa_preprocessor:
	VHJ1c3RlZCBJUDogNjYuMTY1LjE3Ni42MyA9PiAzMDYwNDg=\n,
	received_headers: No Received headers
X-StarScan-Received: 
X-StarScan-Version: 8.28; banners=-,-,-
X-VirusChecked: Checked
Received: (qmail 31195 invoked from network); 8 Apr 2016 07:35:41 -0000
Received: from smtp02.citrix.com (HELO SMTP02.CITRIX.COM) (66.165.176.63)
	by server-13.tower-206.messagelabs.com with RC4-SHA encrypted SMTP;
	8 Apr 2016 07:35:41 -0000
X-IronPort-AV: E=Sophos;i="5.24,449,1454976000";
	d="asc'?scan'208";a="352416426"
Message-ID: <1460100925.13871.6.camel@citrix.com>
From: Dario Faggioli <dario.faggioli@citrix.com>
To: Juergen Gross <jgross@suse.com>, <xen-devel@lists.xenproject.org>
Date: Fri, 8 Apr 2016 09:35:25 +0200
In-Reply-To: <57073128.3030803@suse.com>
References: <20160408011204.10762.14241.stgit@Solace.fritz.box>
	<20160408012420.10762.61178.stgit@Solace.fritz.box>
	<57073128.3030803@suse.com>
Organization: Citrix Inc.
X-Mailer: Evolution 3.18.5.2 (3.18.5.2-1.fc23) 
MIME-Version: 1.0
X-DLP: MIA2
Cc: George Dunlap <george.dunlap@eu.citrix.com>,
	Uma Sharma <uma.sharma523@gmail.com>
Subject: Re: [Xen-devel] [PATCH v3 08/11] xen: sched: allow for choosing
	credit2 runqueues configuration at boot
X-BeenThere: xen-devel@lists.xen.org
X-Mailman-Version: 2.1.18
Precedence: list
List-Id: Xen developer discussion <xen-devel.lists.xen.org>
List-Unsubscribe: <http://lists.xen.org/cgi-bin/mailman/options/xen-devel>,
	<mailto:xen-devel-request@lists.xen.org?subject=unsubscribe>
List-Post: <mailto:xen-devel@lists.xen.org>
List-Help: <mailto:xen-devel-request@lists.xen.org?subject=help>
List-Subscribe: <http://lists.xen.org/cgi-bin/mailman/listinfo/xen-devel>,
	<mailto:xen-devel-request@lists.xen.org?subject=subscribe>
Errors-To: xen-devel-bounces@lists.xen.org
Sender: "Xen-devel" <xen-devel-bounces@lists.xen.org>
X-Spam-Status: No, score=-4.2 required=5.0 tests=BAYES_00, RCVD_IN_DNSWL_MED,
	UNPARSEABLE_RELAY autolearn=unavailable version=3.3.1
X-Spam-Checker-Version: SpamAssassin 3.3.1 (2010-03-16) on mail.kernel.org
X-Virus-Scanned: ClamAV using ClamSMTP

On Fri, 2016-04-08 at 06:18 +0200, Juergen Gross wrote:
> On 08/04/16 03:24, Dario Faggioli wrote:
> > 
> > In fact, credit2 uses CPU topology to decide how to arrange
> > its internal runqueues. Before this change, only 'one runqueue
> > per socket' was allowed. However, experiments have shown that,
> > for instance, having one runqueue per physical core improves
> > performance, especially in case hyperthreading is available.
> > 
> > In general, it makes sense to allow users to pick one runqueue
> > arrangement at boot time, so that:
> >  - more experiments can be easily performed to even better
> >    assess and improve performance;
> >  - one can select the best configuration for his specific
> >    use case and/or hardware.
> > 
> > This patch enables the above.
> > 
> > Note that, for correctly arranging runqueues to be per-core,
> > just checking cpu_to_core() on the host CPUs is not enough.
> > In fact, cores (and hyperthreads) on different sockets, can
> > have the same core (and thread) IDs! We, therefore, need to
> > check whether the full topology of two CPUs matches, for
> > them to be put in the same runqueue.
> > 
> > Note also that the default (although not functional) for
> > credit2, since now, has been per-socket runqueue. This patch
> > leaves things that way, to avoid mixing policy and technical
> > changes.
> > 
> > Finally, it would be a nice feature to be able to select
> > a particular runqueue arrangement, even when creating a
> > Credit2 cpupool. This is left as future work.
> > 
> > Signed-off-by: Dario Faggioli <dario.faggioli@citrix.com>
> > Signed-off-by: Uma Sharma <uma.sharma523@gmail.com>
>
> Some nits below.
> 
Thanks for the quick review!

A revised version of this patch is provided here (both inlined and
attached), and a branch with the remaining to be committed patches of
this series, and with this patch changed as you suggest, is available
at:

 git://xenbits.xen.org/people/dariof/xen.git rel/sched/credit2/fix-runq-and-haff-v4
 http://xenbits.xen.org/gitweb/?p=people/dariof/xen.git;a=shortlog;h=refs/heads/rel/sched/credit2/fix-runq-and-haff-v4

Regards,
Dario
Reviewed-by: Juergen Gross <jgross@suse.com>
Reviewed-by: George Dunlap <george.dunlap@citrix.com>
---
commit 7f491488bbff1cc3af021cd29fca7e0fba321e02
Author: Dario Faggioli <dario.faggioli@citrix.com>
Date:   Tue Sep 29 14:05:09 2015 +0200

    xen: sched: allow for choosing credit2 runqueues configuration at boot
    
    In fact, credit2 uses CPU topology to decide how to arrange
    its internal runqueues. Before this change, only 'one runqueue
    per socket' was allowed. However, experiments have shown that,
    for instance, having one runqueue per physical core improves
    performance, especially in case hyperthreading is available.
    
    In general, it makes sense to allow users to pick one runqueue
    arrangement at boot time, so that:
     - more experiments can be easily performed to even better
       assess and improve performance;
     - one can select the best configuration for his specific
       use case and/or hardware.
    
    This patch enables the above.
    
    Note that, for correctly arranging runqueues to be per-core,
    just checking cpu_to_core() on the host CPUs is not enough.
    In fact, cores (and hyperthreads) on different sockets, can
    have the same core (and thread) IDs! We, therefore, need to
    check whether the full topology of two CPUs matches, for
    them to be put in the same runqueue.
    
    Note also that the default (although not functional) for
    credit2, since now, has been per-socket runqueue. This patch
    leaves things that way, to avoid mixing policy and technical
    changes.
    
    Finally, it would be a nice feature to be able to select
    a particular runqueue arrangement, even when creating a
    Credit2 cpupool. This is left as future work.
    
    Signed-off-by: Dario Faggioli <dario.faggioli@citrix.com>
    Signed-off-by: Uma Sharma <uma.sharma523@gmail.com>
    ---
    Cc: George Dunlap <george.dunlap@eu.citrix.com>
    Cc: Uma Sharma <uma.sharma523@gmail.com>
    Cc: Juergen Gross <jgross@suse.com>
    ---
    Changes from v3:
     * fix type and other issue in comments;
       use ARRAY_SIZE when iterating the parameter string array.
    
    Changes from v2:
     * valid strings  are now in an array, that we scan during
       parameter parsing, as suggested during review.
    
    Cahnges from v1:
     * fix bug in parameter parsing, and start using strcmp()
       for that, as requested during review.

commit 7f491488bbff1cc3af021cd29fca7e0fba321e02
Author: Dario Faggioli <dario.faggioli@citrix.com>
Date:   Tue Sep 29 14:05:09 2015 +0200

    xen: sched: allow for choosing credit2 runqueues configuration at boot
    
    In fact, credit2 uses CPU topology to decide how to arrange
    its internal runqueues. Before this change, only 'one runqueue
    per socket' was allowed. However, experiments have shown that,
    for instance, having one runqueue per physical core improves
    performance, especially in case hyperthreading is available.
    
    In general, it makes sense to allow users to pick one runqueue
    arrangement at boot time, so that:
     - more experiments can be easily performed to even better
       assess and improve performance;
     - one can select the best configuration for his specific
       use case and/or hardware.
    
    This patch enables the above.
    
    Note that, for correctly arranging runqueues to be per-core,
    just checking cpu_to_core() on the host CPUs is not enough.
    In fact, cores (and hyperthreads) on different sockets, can
    have the same core (and thread) IDs! We, therefore, need to
    check whether the full topology of two CPUs matches, for
    them to be put in the same runqueue.
    
    Note also that the default (although not functional) for
    credit2, since now, has been per-socket runqueue. This patch
    leaves things that way, to avoid mixing policy and technical
    changes.
    
    Finally, it would be a nice feature to be able to select
    a particular runqueue arrangement, even when creating a
    Credit2 cpupool. This is left as future work.
    
    Signed-off-by: Dario Faggioli <dario.faggioli@citrix.com>
    Signed-off-by: Uma Sharma <uma.sharma523@gmail.com>
    ---
    Cc: George Dunlap <george.dunlap@eu.citrix.com>
    Cc: Uma Sharma <uma.sharma523@gmail.com>
    Cc: Juergen Gross <jgross@suse.com>
    ---
    Changes from v3:
     * fix type and other issue in comments;
       use ARRAY_SIZE when iterating the parameter string array.
    
    Changes from v2:
     * valid strings  are now in an array, that we scan during
       parameter parsing, as suggested during review.
    
    Cahnges from v1:
     * fix bug in parameter parsing, and start using strcmp()
       for that, as requested during review.

diff --git a/docs/misc/xen-command-line.markdown b/docs/misc/xen-command-line.markdown
index ca77e3b..0047f94 100644
--- a/docs/misc/xen-command-line.markdown
+++ b/docs/misc/xen-command-line.markdown
@@ -469,6 +469,25 @@ combination with the `low_crashinfo` command line option.
 ### credit2\_load\_window\_shift
 > `= <integer>`
 
+### credit2\_runqueue
+> `= core | socket | node | all`
+
+> Default: `socket`
+
+Specify how host CPUs are arranged in runqueues. Runqueues are kept
+balanced with respect to the load generated by the vCPUs running on
+them. Smaller runqueues (as in with `core`) means more accurate load
+balancing (for instance, it will deal better with hyperthreading),
+but also more overhead.
+
+Available alternatives, with their meaning, are:
+* `core`: one runqueue per each physical core of the host;
+* `socket`: one runqueue per each physical socket (which often,
+            but not always, matches a NUMA node) of the host;
+* `node`: one runqueue per each NUMA node of the host;
+* `all`: just one runqueue shared by all the logical pCPUs of
+         the host
+
 ### dbgp
 > `= ehci[ <integer> | @pci<bus>:<slot>.<func> ]`
 
diff --git a/xen/common/sched_credit2.c b/xen/common/sched_credit2.c
index a61a45a..d43f67a 100644
--- a/xen/common/sched_credit2.c
+++ b/xen/common/sched_credit2.c
@@ -81,10 +81,6 @@
  * Credits are "reset" when the next vcpu in the runqueue is less than
  * or equal to zero.  At that point, everyone's credits are "clipped"
  * to a small value, and a fixed credit is added to everyone.
- *
- * The plan is for all cores that share an L2 will share the same
- * runqueue.  At the moment, there is one global runqueue for all
- * cores.
  */
 
 /*
@@ -193,6 +189,63 @@ static int __read_mostly opt_overload_balance_tolerance = -3;
 integer_param("credit2_balance_over", opt_overload_balance_tolerance);
 
 /*
+ * Runqueue organization.
+ *
+ * The various cpus are to be assigned each one to a runqueue, and we
+ * want that to happen basing on topology. At the moment, it is possible
+ * to choose to arrange runqueues to be:
+ *
+ * - per-core: meaning that there will be one runqueue per each physical
+ *             core of the host. This will happen if the opt_runqueue
+ *             parameter is set to 'core';
+ *
+ * - per-socket: meaning that there will be one runqueue per each physical
+ *               socket (AKA package, which often, but not always, also
+ *               matches a NUMA node) of the host; This will happen if
+ *               the opt_runqueue parameter is set to 'socket';
+ *
+ * - per-node: meaning that there will be one runqueue per each physical
+ *             NUMA node of the host. This will happen if the opt_runqueue
+ *             parameter is set to 'node';
+ *
+ * - global: meaning that there will be only one runqueue to which all the
+ *           (logical) processors of the host belong. This will happen if
+ *           the opt_runqueue parameter is set to 'all'.
+ *
+ * Depending on the value of opt_runqueue, therefore, cpus that are part of
+ * either the same physical core, the same physical socket, the same NUMA
+ * node, or just all of them, will be put together to form runqueues.
+ */
+#define OPT_RUNQUEUE_CORE   0
+#define OPT_RUNQUEUE_SOCKET 1
+#define OPT_RUNQUEUE_NODE   2
+#define OPT_RUNQUEUE_ALL    3
+static const char *const opt_runqueue_str[] = {
+    [OPT_RUNQUEUE_CORE] = "core",
+    [OPT_RUNQUEUE_SOCKET] = "socket",
+    [OPT_RUNQUEUE_NODE] = "node",
+    [OPT_RUNQUEUE_ALL] = "all"
+};
+static int __read_mostly opt_runqueue = OPT_RUNQUEUE_SOCKET;
+
+static void parse_credit2_runqueue(const char *s)
+{
+    unsigned int i;
+
+    for ( i = 0; i < ARRAY_SIZE(opt_runqueue_str); i++ )
+    {
+        if ( !strcmp(s, opt_runqueue_str[i]) )
+        {
+            opt_runqueue = i;
+            return;
+        }
+    }
+
+    printk("WARNING, unrecognized value of credit2_runqueue option!\n");
+}
+custom_param("credit2_runqueue", parse_credit2_runqueue);
+
+/*
  * Per-runqueue data
  */
 struct csched2_runqueue_data {
@@ -1974,6 +2027,22 @@ static void deactivate_runqueue(struct csched2_private *prv, int rqi)
     cpumask_clear_cpu(rqi, &prv->active_queues);
 }
 
+static inline bool_t same_node(unsigned int cpua, unsigned int cpub)
+{
+    return cpu_to_node(cpua) == cpu_to_node(cpub);
+}
+
+static inline bool_t same_socket(unsigned int cpua, unsigned int cpub)
+{
+    return cpu_to_socket(cpua) == cpu_to_socket(cpub);
+}
+
+static inline bool_t same_core(unsigned int cpua, unsigned int cpub)
+{
+    return same_socket(cpua, cpub) &&
+           cpu_to_core(cpua) == cpu_to_core(cpub);
+}
+
 static unsigned int
 cpu_to_runqueue(struct csched2_private *prv, unsigned int cpu)
 {
@@ -2006,7 +2075,10 @@ cpu_to_runqueue(struct csched2_private *prv, unsigned int cpu)
         BUG_ON(cpu_to_socket(cpu) == XEN_INVALID_SOCKET_ID ||
                cpu_to_socket(peer_cpu) == XEN_INVALID_SOCKET_ID);
 
-        if ( cpu_to_socket(cpumask_first(&rqd->active)) == cpu_to_socket(cpu) )
+        if ( opt_runqueue == OPT_RUNQUEUE_ALL ||
+             (opt_runqueue == OPT_RUNQUEUE_CORE && same_core(peer_cpu, cpu)) ||
+             (opt_runqueue == OPT_RUNQUEUE_SOCKET && same_socket(peer_cpu, cpu)) ||
+             (opt_runqueue == OPT_RUNQUEUE_NODE && same_node(peer_cpu, cpu)) )
             break;
     }
 
@@ -2170,6 +2242,7 @@ csched2_init(struct scheduler *ops)
     printk(" load_window_shift: %d\n", opt_load_window_shift);
     printk(" underload_balance_tolerance: %d\n", opt_underload_balance_tolerance);
     printk(" overload_balance_tolerance: %d\n", opt_overload_balance_tolerance);
+    printk(" runqueues arrangement: %s\n", opt_runqueue_str[opt_runqueue]);
 
     if ( opt_load_window_shift < LOADAVG_WINDOW_SHIFT_MIN )
     {