[for-6.2,32/53] target/arm: Implement MVE VCTP

Message ID	20210729111512.16541-33-peter.maydell@linaro.org (mailing list archive)
State	New, archived
Headers	show Return-Path: <SRS0=bq2O=MV=nongnu.org=qemu-devel-bounces+qemu-devel=archiver.kernel.org@kernel.org> DMARC-Filter: OpenDMARC Filter v1.4.1 mail.kernel.org 6AEFA6056C From: Peter Maydell <peter.maydell@linaro.org> To: qemu-arm@nongnu.org, qemu-devel@nongnu.org Subject: [PATCH for-6.2 32/53] target/arm: Implement MVE VCTP Date: Thu, 29 Jul 2021 12:14:51 +0100 Message-Id: <20210729111512.16541-33-peter.maydell@linaro.org> In-Reply-To: <20210729111512.16541-1-peter.maydell@linaro.org> References: <20210729111512.16541-1-peter.maydell@linaro.org> MIME-Version: 1.0 Content-Transfer-Encoding: 8bit Received-SPF: pass client-ip=2a00:1450:4864:20::330; envelope-from=peter.maydell@linaro.org; helo=mail-wm1-x330.google.com X-Spam_score_int: -20 X-Spam_score: -2.1 X-Spam_bar: -- X-Spam_report: (-2.1 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action Precedence: list Errors-To: qemu-devel-bounces+qemu-devel=archiver.kernel.org@nongnu.org Sender: "Qemu-devel" <qemu-devel-bounces+qemu-devel=archiver.kernel.org@nongnu.org>
Series	target/arm: MVE slices 3 and 4 \| expand [for-6.2,00/53] target/arm: MVE slices 3 and 4 [for-6.2,01/53] target/arm: Note that we handle VMOVL as a special case of VSHLL [for-6.2,02/53] target/arm: Print MVE VPR in CPU dumps [for-6.2,03/53] target/arm: Fix MVE VSLI by 0 and VSRI by <dt> [for-6.2,04/53] target/arm: Fix signed VADDV [for-6.2,05/53] target/arm: Fix mask handling for MVE narrowing operations [for-6.2,06/53] target/arm: Fix 48-bit saturating shifts [for-6.2,07/53] target/arm: Fix MVE 48-bit SQRSHRL for small right shifts [for-6.2,08/53] target/arm: Fix calculation of LTP mask when LR is 0 [for-6.2,09/53] target/arm: Factor out mve_eci_mask() [for-6.2,10/53] target/arm: Fix VPT advance when ECI is non-zero [for-6.2,11/53] target/arm: Fix VLDRB/H/W for predicated elements [for-6.2,12/53] target/arm: Implement MVE VMULL (polynomial) [for-6.2,13/53] target/arm: Implement MVE incrementing/decrementing dup insns [for-6.2,14/53] target/arm: Factor out gen_vpst() [for-6.2,15/53] target/arm: Implement MVE integer vector comparisons [for-6.2,16/53] target/arm: Implement MVE integer vector-vs-scalar comparisons [for-6.2,17/53] target/arm: Implement MVE VPSEL [for-6.2,18/53] target/arm: Implement MVE VMLAS [for-6.2,19/53] target/arm: Implement MVE shift-by-scalar [for-6.2,20/53] target/arm: Move 'x' and 'a' bit definitions into vmlaldav formats [for-6.2,21/53] target/arm: Implement MVE integer min/max across vector [for-6.2,22/53] target/arm: Implement MVE VABAV [for-6.2,23/53] target/arm: Implement MVE narrowing moves [for-6.2,24/53] target/arm: Rename MVEGenDualAccOpFn to MVEGenLongDualAccOpFn [for-6.2,25/53] target/arm: Implement MVE VMLADAV and VMLSLDAV [for-6.2,26/53] target/arm: Implement MVE VMLA [for-6.2,27/53] target/arm: Implement MVE saturating doubling multiply accumulates [for-6.2,28/53] target/arm: Implement MVE VQABS, VQNEG [for-6.2,29/53] target/arm: Implement MVE VMAXA, VMINA [for-6.2,30/53] target/arm: Implement MVE VMOV to/from 2 general-purpose registers [for-6.2,31/53] target/arm: Implement MVE VPNOT [for-6.2,32/53] target/arm: Implement MVE VCTP [for-6.2,33/53] target/arm: Implement MVE scatter-gather insns [for-6.2,34/53] target/arm: Implement MVE scatter-gather immediate forms [for-6.2,35/53] target/arm: Implement MVE interleaving loads/stores [for-6.2,36/53] target/arm: Implement MVE VADD (floating-point) [for-6.2,37/53] target/arm: Implement MVE VSUB, VMUL, VABD, VMAXNM, VMINNM [for-6.2,38/53] target/arm: Implement MVE VCADD [for-6.2,39/53] target/arm: Implement MVE VFMA and VFMS [for-6.2,40/53] target/arm: Implement MVE VCMUL and VCMLA [for-6.2,41/53] target/arm: Implement MVE VMAXNMA and VMINNMA [for-6.2,42/53] target/arm: Implement MVE scalar fp insns [for-6.2,43/53] target/arm: Implement MVE fp-with-scalar VFMA, VFMAS [for-6.2,44/53] softfloat: Remove assertion preventing silencing of NaN in default-NaN mode [for-6.2,45/53] target/arm: Implement MVE FP max/min across vector [for-6.2,46/53] target/arm: Implement MVE fp vector comparisons [for-6.2,47/53] target/arm: Implement MVE fp scalar comparisons [for-6.2,48/53] target/arm: Implement MVE VCVT between floating and fixed point [for-6.2,49/53] target/arm: Implement MVE VCVT between fp and integer [for-6.2,50/53] target/arm: Implement MVE VCVT with specified rounding mode [for-6.2,51/53] target/arm: Implement MVE VCVT between single and half precision [for-6.2,52/53] target/arm: Implement MVE VRINT insns [for-6.2,53/53] target/arm: Enable MVE in Cortex-M55

Message ID

20210729111512.16541-33-peter.maydell@linaro.org (mailing list archive)

State

New, archived

Headers

DMARC-Filter: OpenDMARC Filter v1.4.1 mail.kernel.org 6AEFA6056C
From: Peter Maydell <peter.maydell@linaro.org>
To: qemu-arm@nongnu.org,
	qemu-devel@nongnu.org
Subject: [PATCH for-6.2 32/53] target/arm: Implement MVE VCTP
Date: Thu, 29 Jul 2021 12:14:51 +0100
Message-Id: <20210729111512.16541-33-peter.maydell@linaro.org>
In-Reply-To: <20210729111512.16541-1-peter.maydell@linaro.org>
References: <20210729111512.16541-1-peter.maydell@linaro.org>
MIME-Version: 1.0
Content-Transfer-Encoding: 8bit
Received-SPF: pass client-ip=2a00:1450:4864:20::330;
 envelope-from=peter.maydell@linaro.org; helo=mail-wm1-x330.google.com
X-Spam_score_int: -20
X-Spam_score: -2.1
X-Spam_bar: --
X-Spam_report: (-2.1 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1,
 DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1,
 RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001,
 SPF_PASS=-0.001 autolearn=ham autolearn_force=no
X-Spam_action: no action
X-BeenThere: qemu-devel@nongnu.org
X-Mailman-Version: 2.1.23
Precedence: list
List-Id: <qemu-devel.nongnu.org>
List-Unsubscribe: <https://lists.nongnu.org/mailman/options/qemu-devel>,
 <mailto:qemu-devel-request@nongnu.org?subject=unsubscribe>
List-Archive: <https://lists.nongnu.org/archive/html/qemu-devel>
List-Post: <mailto:qemu-devel@nongnu.org>
List-Help: <mailto:qemu-devel-request@nongnu.org?subject=help>
List-Subscribe: <https://lists.nongnu.org/mailman/listinfo/qemu-devel>,
 <mailto:qemu-devel-request@nongnu.org?subject=subscribe>
Errors-To: qemu-devel-bounces+qemu-devel=archiver.kernel.org@nongnu.org
Sender: "Qemu-devel"
 <qemu-devel-bounces+qemu-devel=archiver.kernel.org@nongnu.org>

Series

target/arm: MVE slices 3 and 4 | expand

Commit Message

Peter Maydell July 29, 2021, 11:14 a.m. UTC

Implement the MVE VCTP insn, which sets the VPR.P0 predicate bits so
as to predicate any element at index Rn or greater is predicated.  As
with VPNOT, this insn itself is predicable and subject to beatwise
execution.

The calculation of the mask is the same as is used to determine
ltpmask in mve_element_mask(), but we precalculate masklen in
generated code to avoid having to have 4 helpers specialized by size.

We put the decode line in with the low-overhead-loop insns in
t32.decode because it's logically part of that collection of insn
patterns, even though it is an MVE only insn.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
---
 target/arm/helper-mve.h    |  2 ++
 target/arm/translate-a32.h |  1 +
 target/arm/t32.decode      |  1 +
 target/arm/mve_helper.c    | 20 ++++++++++++++++++++
 target/arm/translate-mve.c |  2 +-
 target/arm/translate.c     | 33 +++++++++++++++++++++++++++++++++
 6 files changed, 58 insertions(+), 1 deletion(-)

diff --git a/target/arm/helper-mve.h b/target/arm/helper-mve.h
index 8cb941912fc..b6cf3f0c94d 100644
--- a/target/arm/helper-mve.h
+++ b/target/arm/helper-mve.h
@@ -121,6 +121,8 @@  DEF_HELPER_FLAGS_4(mve_veor, TCG_CALL_NO_WG, void, env, ptr, ptr, ptr)
 DEF_HELPER_FLAGS_4(mve_vpsel, TCG_CALL_NO_WG, void, env, ptr, ptr, ptr)
 DEF_HELPER_FLAGS_1(mve_vpnot, TCG_CALL_NO_WG, void, env)
 
+DEF_HELPER_FLAGS_2(mve_vctp, TCG_CALL_NO_WG, void, env, i32)
+
 DEF_HELPER_FLAGS_4(mve_vaddb, TCG_CALL_NO_WG, void, env, ptr, ptr, ptr)
 DEF_HELPER_FLAGS_4(mve_vaddh, TCG_CALL_NO_WG, void, env, ptr, ptr, ptr)
 DEF_HELPER_FLAGS_4(mve_vaddw, TCG_CALL_NO_WG, void, env, ptr, ptr, ptr)
diff --git a/target/arm/translate-a32.h b/target/arm/translate-a32.h
index 6f4d65ddb00..88f15df60e8 100644
--- a/target/arm/translate-a32.h
+++ b/target/arm/translate-a32.h
@@ -48,6 +48,7 @@  long neon_element_offset(int reg, int element, MemOp memop);
 void gen_rev16(TCGv_i32 dest, TCGv_i32 var);
 void clear_eci_state(DisasContext *s);
 bool mve_eci_check(DisasContext *s);
+void mve_update_eci(DisasContext *s);
 void mve_update_and_store_eci(DisasContext *s);
 bool mve_skip_vmov(DisasContext *s, int vn, int index, int size);
 
diff --git a/target/arm/t32.decode b/target/arm/t32.decode
index 2d47f31f143..78fadef9d62 100644
--- a/target/arm/t32.decode
+++ b/target/arm/t32.decode
@@ -748,5 +748,6 @@  BL               1111 0. .......... 11.1 ............         @branch24
       # This is DLSTP
       DLS        1111 0 0000 0 size:2 rn:4 1110 0000 0000 0001
     }
+    VCTP         1111 0 0000 0 size:2 rn:4 1110 1000 0000 0001
   ]
 }
diff --git a/target/arm/mve_helper.c b/target/arm/mve_helper.c
index c22a00c5ed6..1752555a218 100644
--- a/target/arm/mve_helper.c
+++ b/target/arm/mve_helper.c
@@ -2218,6 +2218,26 @@  void HELPER(mve_vpnot)(CPUARMState *env)
     mve_advance_vpt(env);
 }
 
+/*
+ * VCTP: P0 unexecuted bits unchanged, predicated bits zeroed,
+ * otherwise set according to value of Rn. The calculation of
+ * newmask here works in the same way as the calculation of the
+ * ltpmask in mve_element_mask(), but we have pre-calculated
+ * the masklen in the generated code.
+ */
+void HELPER(mve_vctp)(CPUARMState *env, uint32_t masklen)
+{
+    uint16_t mask = mve_element_mask(env);
+    uint16_t eci_mask = mve_eci_mask(env);
+    uint16_t newmask;
+
+    assert(masklen <= 16);
+    newmask = masklen ? MAKE_64BIT_MASK(0, masklen) : 0;
+    newmask &= mask;
+    env->v7m.vpr = (env->v7m.vpr & ~(uint32_t)eci_mask) | (newmask & eci_mask);
+    mve_advance_vpt(env);
+}
+
 #define DO_1OP_SAT(OP, ESIZE, TYPE, FN)                                 \
     void HELPER(mve_##OP)(CPUARMState *env, void *vd, void *vm)         \
     {                                                                   \
diff --git a/target/arm/translate-mve.c b/target/arm/translate-mve.c
index cc2e58cfe2f..865d5acbe76 100644
--- a/target/arm/translate-mve.c
+++ b/target/arm/translate-mve.c
@@ -93,7 +93,7 @@  bool mve_eci_check(DisasContext *s)
     }
 }
 
-static void mve_update_eci(DisasContext *s)
+void mve_update_eci(DisasContext *s)
 {
     /*
      * The helper function will always update the CPUState field,
diff --git a/target/arm/translate.c b/target/arm/translate.c
index 80c282669f0..804a53279bd 100644
--- a/target/arm/translate.c
+++ b/target/arm/translate.c
@@ -8669,6 +8669,39 @@  static bool trans_LCTP(DisasContext *s, arg_LCTP *a)
     return true;
 }
 
+static bool trans_VCTP(DisasContext *s, arg_VCTP *a)
+{
+    /*
+     * M-profile Create Vector Tail Predicate. This insn is itself
+     * predicated and is subject to beatwise execution.
+     */
+    TCGv_i32 rn_shifted, masklen;
+
+    if (!dc_isar_feature(aa32_mve, s) || a->rn == 13 || a->rn == 15) {
+        return false;
+    }
+
+    if (!mve_eci_check(s) || !vfp_access_check(s)) {
+        return true;
+    }
+
+    /*
+     * We pre-calculate the mask length here to avoid having
+     * to have multiple helpers specialized for size.
+     * We pass the helper "rn <= (1 << (4 - size)) ? (rn << size) : 16".
+     */
+    rn_shifted = tcg_temp_new_i32();
+    masklen = load_reg(s, a->rn);
+    tcg_gen_shli_i32(rn_shifted, masklen, a->size);
+    tcg_gen_movcond_i32(TCG_COND_LEU, masklen,
+                        masklen, tcg_constant_i32(1 << (4 - a->size)),
+                        rn_shifted, tcg_constant_i32(16));
+    gen_helper_mve_vctp(cpu_env, masklen);
+    tcg_temp_free_i32(masklen);
+    tcg_temp_free_i32(rn_shifted);
+    mve_update_eci(s);
+    return true;
+}
 
 static bool op_tbranch(DisasContext *s, arg_tbranch *a, bool half)
 {

[for-6.2,32/53] target/arm: Implement MVE VCTP

Commit Message

Patch