[PULL,38/45] target/arm: Use gvec for NEON VLD all lanes

Message ID	20181019165735.22511-39-peter.maydell@linaro.org (mailing list archive)
State	New, archived
Headers	show Return-Path: <qemu-devel-bounces+patchwork-qemu-devel=patchwork.kernel.org@nongnu.org> From: Peter Maydell <peter.maydell@linaro.org> To: qemu-devel@nongnu.org Date: Fri, 19 Oct 2018 17:57:28 +0100 Message-Id: <20181019165735.22511-39-peter.maydell@linaro.org> In-Reply-To: <20181019165735.22511-1-peter.maydell@linaro.org> References: <20181019165735.22511-1-peter.maydell@linaro.org> MIME-Version: 1.0 Content-Transfer-Encoding: 8bit Subject: [Qemu-devel] [PULL 38/45] target/arm: Use gvec for NEON VLD all lanes Precedence: list Errors-To: qemu-devel-bounces+patchwork-qemu-devel=patchwork.kernel.org@nongnu.org Sender: "Qemu-devel" <qemu-devel-bounces+patchwork-qemu-devel=patchwork.kernel.org@nongnu.org>
Series	[PULL,01/45] ssi-sd: Make devices picking up backends unavailable with -device \| expand [PULL,01/45] ssi-sd: Make devices picking up backends unavailable with -device [PULL,02/45] target/arm: Add support for VCPU event states [PULL,03/45] target/arm: Move some system registers into a substructure [PULL,04/45] target/arm: V8M should not imply V7VE [PULL,05/45] target/arm: Convert v8 extensions from feature bits to isar tests [PULL,06/45] target/arm: Convert division from feature bits to isar0 tests [PULL,07/45] target/arm: Convert jazelle from feature bit to isar1 test [PULL,08/45] target/arm: Convert t32ee from feature bit to isar3 test [PULL,09/45] target/arm: Convert sve from feature bit to aa64pfr0 test [PULL,10/45] target/arm: Convert v8.2-fp16 from feature bit to aa64pfr0 test [PULL,11/45] target/arm: Improve debug logging of AArch32 exception return [PULL,12/45] target/arm: Make switch_mode() file-local [PULL,13/45] target/arm: Implement HCR.FB [PULL,14/45] target/arm: Implement HCR.DC [PULL,15/45] target/arm: ISR_EL1 bits track virtual interrupts if IMO/FMO set [PULL,16/45] target/arm: Implement HCR.VI and VF [PULL,17/45] target/arm: Implement HCR.PTW [PULL,18/45] target/arm: New utility function to extract EC from syndrome [PULL,19/45] target/arm: Get IL bit correct for v7 syndrome values [PULL,20/45] target/arm: Report correct syndrome for FP/SIMD traps to Hyp mode [PULL,21/45] hw/arm/boot: Increase compliance with kernel arm64 boot protocol [PULL,22/45] target/arm: Hoist address increment for vector memory ops [PULL,23/45] target/arm: Don't call tcg_clear_temp_count [PULL,24/45] target/arm: Use tcg_gen_gvec_dup_i64 for LD[1-4]R [PULL,25/45] target/arm: Promote consecutive memory ops for aa64 [PULL,26/45] target/arm: Mark some arrays const [PULL,27/45] target/arm: Use gvec for NEON VDUP [PULL,28/45] target/arm: Use gvec for NEON VMOV, VMVN, VBIC & VORR (immediate) [PULL,29/45] target/arm: Use gvec for NEON_3R_LOGIC insns [PULL,30/45] target/arm: Use gvec for NEON_3R_VADD_VSUB insns [PULL,31/45] target/arm: Use gvec for NEON_2RM_VMN, NEON_2RM_VNEG [PULL,32/45] target/arm: Use gvec for NEON_3R_VMUL [PULL,33/45] target/arm: Use gvec for VSHR, VSHL [PULL,34/45] target/arm: Use gvec for VSRA [PULL,35/45] target/arm: Use gvec for VSRI, VSLI [PULL,36/45] target/arm: Use gvec for NEON_3R_VML [PULL,37/45] target/arm: Use gvec for NEON_3R_VTST_VCEQ, NEON_3R_VCGT, NEON_3R_VCGE [PULL,38/45] target/arm: Use gvec for NEON VLD all lanes [PULL,39/45] target/arm: Reorg NEON VLD/VST all elements [PULL,40/45] target/arm: Promote consecutive memory ops for aa32 [PULL,41/45] target/arm: Reorg NEON VLD/VST single element to one lane [PULL,42/45] net: cadence_gem: Announce availability of priority queues [PULL,43/45] net: cadence_gem: Announce 64bit addressing support [PULL,44/45] target/arm: Remove writefn from TTBR0_EL3 [PULL,45/45] target/arm: Only flush tlb if ASID changes

Message ID

20181019165735.22511-39-peter.maydell@linaro.org (mailing list archive)

State

New, archived

Headers

From: Peter Maydell <peter.maydell@linaro.org>
To: qemu-devel@nongnu.org
Date: Fri, 19 Oct 2018 17:57:28 +0100
Message-Id: <20181019165735.22511-39-peter.maydell@linaro.org>
In-Reply-To: <20181019165735.22511-1-peter.maydell@linaro.org>
References: <20181019165735.22511-1-peter.maydell@linaro.org>
MIME-Version: 1.0
Content-Transfer-Encoding: 8bit
Subject: [Qemu-devel] [PULL 38/45] target/arm: Use gvec for NEON VLD all
 lanes
Precedence: list
Errors-To: 
 qemu-devel-bounces+patchwork-qemu-devel=patchwork.kernel.org@nongnu.org
Sender: "Qemu-devel"
	<qemu-devel-bounces+patchwork-qemu-devel=patchwork.kernel.org@nongnu.org>

Series

[PULL,01/45] ssi-sd: Make devices picking up backends unavailable with -device | expand

Commit Message

Peter Maydell Oct. 19, 2018, 4:57 p.m. UTC

From: Richard Henderson <richard.henderson@linaro.org>

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20181011205206.3552-18-richard.henderson@linaro.org
[PMM: added parens in ?: expression]
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 target/arm/translate.c | 81 ++++++++++++++----------------------------
 1 file changed, 26 insertions(+), 55 deletions(-)

diff --git a/target/arm/translate.c b/target/arm/translate.c
index e6b06910369..e5d723d03b7 100644
--- a/target/arm/translate.c
+++ b/target/arm/translate.c
@@ -2993,19 +2993,6 @@  static void gen_vfp_msr(TCGv_i32 tmp)
     tcg_temp_free_i32(tmp);
 }
 
-static void gen_neon_dup_u8(TCGv_i32 var, int shift)
-{
-    TCGv_i32 tmp = tcg_temp_new_i32();
-    if (shift)
-        tcg_gen_shri_i32(var, var, shift);
-    tcg_gen_ext8u_i32(var, var);
-    tcg_gen_shli_i32(tmp, var, 8);
-    tcg_gen_or_i32(var, var, tmp);
-    tcg_gen_shli_i32(tmp, var, 16);
-    tcg_gen_or_i32(var, var, tmp);
-    tcg_temp_free_i32(tmp);
-}
-
 static void gen_neon_dup_low16(TCGv_i32 var)
 {
     TCGv_i32 tmp = tcg_temp_new_i32();
@@ -3024,28 +3011,6 @@  static void gen_neon_dup_high16(TCGv_i32 var)
     tcg_temp_free_i32(tmp);
 }
 
-static TCGv_i32 gen_load_and_replicate(DisasContext *s, TCGv_i32 addr, int size)
-{
-    /* Load a single Neon element and replicate into a 32 bit TCG reg */
-    TCGv_i32 tmp = tcg_temp_new_i32();
-    switch (size) {
-    case 0:
-        gen_aa32_ld8u(s, tmp, addr, get_mem_index(s));
-        gen_neon_dup_u8(tmp, 0);
-        break;
-    case 1:
-        gen_aa32_ld16u(s, tmp, addr, get_mem_index(s));
-        gen_neon_dup_low16(tmp);
-        break;
-    case 2:
-        gen_aa32_ld32u(s, tmp, addr, get_mem_index(s));
-        break;
-    default: /* Avoid compiler warnings.  */
-        abort();
-    }
-    return tmp;
-}
-
 static int handle_vsel(uint32_t insn, uint32_t rd, uint32_t rn, uint32_t rm,
                        uint32_t dp)
 {
@@ -4949,6 +4914,7 @@  static int disas_neon_ls_insn(DisasContext *s, uint32_t insn)
     int load;
     int shift;
     int n;
+    int vec_size;
     TCGv_i32 addr;
     TCGv_i32 tmp;
     TCGv_i32 tmp2;
@@ -5118,28 +5084,33 @@  static int disas_neon_ls_insn(DisasContext *s, uint32_t insn)
             }
             addr = tcg_temp_new_i32();
             load_reg_var(s, addr, rn);
-            if (nregs == 1) {
-                /* VLD1 to all lanes: bit 5 indicates how many Dregs to write */
-                tmp = gen_load_and_replicate(s, addr, size);
-                tcg_gen_st_i32(tmp, cpu_env, neon_reg_offset(rd, 0));
-                tcg_gen_st_i32(tmp, cpu_env, neon_reg_offset(rd, 1));
-                if (insn & (1 << 5)) {
-                    tcg_gen_st_i32(tmp, cpu_env, neon_reg_offset(rd + 1, 0));
-                    tcg_gen_st_i32(tmp, cpu_env, neon_reg_offset(rd + 1, 1));
-                }
-                tcg_temp_free_i32(tmp);
-            } else {
-                /* VLD2/3/4 to all lanes: bit 5 indicates register stride */
-                stride = (insn & (1 << 5)) ? 2 : 1;
-                for (reg = 0; reg < nregs; reg++) {
-                    tmp = gen_load_and_replicate(s, addr, size);
-                    tcg_gen_st_i32(tmp, cpu_env, neon_reg_offset(rd, 0));
-                    tcg_gen_st_i32(tmp, cpu_env, neon_reg_offset(rd, 1));
-                    tcg_temp_free_i32(tmp);
-                    tcg_gen_addi_i32(addr, addr, 1 << size);
-                    rd += stride;
+
+            /* VLD1 to all lanes: bit 5 indicates how many Dregs to write.
+             * VLD2/3/4 to all lanes: bit 5 indicates register stride.
+             */
+            stride = (insn & (1 << 5)) ? 2 : 1;
+            vec_size = nregs == 1 ? stride * 8 : 8;
+
+            tmp = tcg_temp_new_i32();
+            for (reg = 0; reg < nregs; reg++) {
+                gen_aa32_ld_i32(s, tmp, addr, get_mem_index(s),
+                                s->be_data | size);
+                if ((rd & 1) && vec_size == 16) {
+                    /* We cannot write 16 bytes at once because the
+                     * destination is unaligned.
+                     */
+                    tcg_gen_gvec_dup_i32(size, neon_reg_offset(rd, 0),
+                                         8, 8, tmp);
+                    tcg_gen_gvec_mov(0, neon_reg_offset(rd + 1, 0),
+                                     neon_reg_offset(rd, 0), 8, 8);
+                } else {
+                    tcg_gen_gvec_dup_i32(size, neon_reg_offset(rd, 0),
+                                         vec_size, vec_size, tmp);
                 }
+                tcg_gen_addi_i32(addr, addr, 1 << size);
+                rd += stride;
             }
+            tcg_temp_free_i32(tmp);
             tcg_temp_free_i32(addr);
             stride = (1 << size) * nregs;
         } else {

[PULL,38/45] target/arm: Use gvec for NEON VLD all lanes

Commit Message

Patch