VPTESTMB, VPTESTMW, VPTESTMD, VPTESTMQ

Logical AND and Set Mask

stableVMJITAOTinstruction

Encodings

OpcodeInstructionOp/En64-bitCompat/LegacyDescription
EVEX.128.66.0F38.W0 26 /rVPTESTMB k2 {k1}, xmm2, xmm3/m128AValidValidBitwise AND of packed byte integers in xmm2 and AVX512BW) OR xmm3/m128 and set mask k2 to reflect the zero/non- AVX10.1 zero status of each element of the result, under writemask k1.
EVEX.256.66.0F38.W0 26 /rVPTESTMB k2 {k1}, ymm2, ymm3/m256AValidValidBitwise AND of packed byte integers in ymm2 and AVX512BW) OR ymm3/m256 and set mask k2 to reflect the zero/non- AVX10.1 zero status of each element of the result, under writemask k1.
EVEX.512.66.0F38.W0 26 /rVPTESTMB k2 {k1}, zmm2, zmm3/m512AValidValidBitwise AND of packed byte integers in zmm2 and OR AVX10.1 zmm3/m512 and set mask k2 to reflect the zero/non- zero status of each element of the result, under writemask k1.
EVEX.128.66.0F38.W1 26 /rVPTESTMW k2 {k1}, xmm2, xmm3/m128AValidValidBitwise AND of packed word integers in xmm2 and AVX512BW) OR xmm3/m128 and set mask k2 to reflect the zero/non- AVX10.1 zero status of each element of the result, under writemask k1.
EVEX.256.66.0F38.W1 26 /rVPTESTMW k2 {k1}, ymm2, ymm3/m256AValidValidBitwise AND of packed word integers in ymm2 and AVX512BW) OR ymm3/m256 and set mask k2 to reflect the zero/non- AVX10.1 zero status of each element of the result, under writemask k1.
EVEX.512.66.0F38.W1 26 /rVPTESTMW k2 {k1}, zmm2, zmm3/m512AValidValidBitwise AND of packed word integers in zmm2 and OR AVX10.1 zmm3/m512 and set mask k2 to reflect the zero/non- zero status of each element of the result, under writemask k1.
EVEX.128.66.0F38.W0 27 /rVPTESTMD k2 {k1}, xmm2, xmm3/m128/m32bcstBValidValidBitwise AND of packed doubleword integers in xmm2 AVX512F) OR and xmm3/m128/m32bcst and set mask k2 to reflect AVX10.1 the zero/non-zero status of each element of the result, under writemask k1.
EVEX.256.66.0F38.W0 27 /rVPTESTMD k2 {k1}, ymm2, ymm3/m256/m32bcstBValidValidBitwise AND of packed doubleword integers in ymm2 AVX512F) OR and ymm3/m256/m32bcst and set mask k2 to reflect AVX10.1 the zero/non-zero status of each element of the result, under writemask k1.
EVEX.512.66.0F38.W0 27 /rVPTESTMD k2 {k1}, zmm2, zmm3/m512/m32bcstBValidValidBitwise AND of packed doubleword integers in zmm2 OR AVX10.1 and zmm3/m512/m32bcst and set mask k2 to reflect the zero/non-zero status of each element of the result, under writemask k1.
EVEX.128.66.0F38.W1 27 /rVPTESTMQ k2 {k1}, xmm2, xmm3/m128/m64bcstBValidValidBitwise AND of packed quadword integers in xmm2 and AVX512F) OR xmm3/m128/m64bcst and set mask k2 to reflect the AVX10.1 zero/non-zero status of each element of the result, under writemask k1.
EVEX.256.66.0F38.W1 27 /rVPTESTMQ k2 {k1}, ymm2, ymm3/m256/m64bcstBValidValidBitwise AND of packed quadword integers in ymm2 and AVX512F) OR ymm3/m256/m64bcst and set mask k2 to reflect the AVX10.1 zero/non-zero status of each element of the result, under writemask k1.
EVEX.512.66.0F38.W1 27 /rVPTESTMQ k2 {k1}, zmm2, zmm3/m512/m64bcstBValidValidBitwise AND of packed quadword integers in zmm2 and OR AVX10.1 zmm3/m512/m64bcst and set mask k2 to reflect the zero/non-zero status of each element of the result, under writemask k1.

Operand encoding

Each mode is a value of the Op/En column above. It says which field of the encoded instruction carries each operand, in the order they are written, and whether the instruction reads it, writes it or both.

A

  1. modrm.reg escrituraModRM byte, reg field (bits 5-3)
  2. evex.vvvv lecturaEVEX prefix, vvvv field (inverted)
  3. modrm.rm lecturaModRM byte, r/m field (bits 2-0); with the SIB byte and the displacement when the mod field asks for them

Tupla: Full Mem

B

  1. modrm.reg escrituraModRM byte, reg field (bits 5-3)
  2. evex.vvvv lecturaEVEX prefix, vvvv field (inverted)
  3. modrm.rm lecturaModRM byte, r/m field (bits 2-0); with the SIB byte and the displacement when the mod field asks for them

Tupla: Full

Measured cost

Loading measurements from arch-data...

Description

Performs a bitwise logical AND operation on the first source operand (the second operand) and second source operand (the third operand) and stores the result in the destination operand (the first operand) under the writemask. Each bit of the result is set to 1 if the bitwise AND of the corresponding elements of the first and second src operands is non-zero; otherwise it is set to 0.

VPTESTMD/VPTESTMQ: The first source operand is a ZMM/YMM/XMM register. The second source operand can be a ZMM/YMM/XMM register, a 512/256/128-bit memory location or a 512/256/128-bit vector broadcasted from a 32/64-bit memory location. The destination operand is a mask register updated under the writemask.

VPTESTMB/VPTESTMW: The first source operand is a ZMM/YMM/XMM register. The second source operand can be a ZMM/YMM/XMM register or a 512/256/128-bit memory location. The destination operand is a mask register updated under the writemask.

Operation

VPTESTMB (EVEX encoded versions)

(KL, VL) = (16, 128), (32, 256), (64, 512)

FOR j := 0 TO KL-1

i := j * 8

IF k1[j] OR *no writemask*

     THEN DEST[j] := (SRC1[i+7:i] BITWISE AND SRC2[i+7:i] != 0)? 1 : 0;

     ELSE DEST[j] = 0                       ; zeroing-masking only

FI;

ENDFOR

DEST[MAX_KL-1:KL] := 0

VPTESTMW (EVEX encoded versions)

(KL, VL) = (8, 128), (16, 256), (32, 512)

FOR j := 0 TO KL-1

i := j * 16

IF k1[j] OR *no writemask*

     THEN DEST[j] := (SRC1[i+15:i] BITWISE AND SRC2[i+15:i] != 0)? 1 : 0;

     ELSE DEST[j] = 0                       ; zeroing-masking only

FI;

ENDFOR

DEST[MAX_KL-1:KL] := 0


VPTESTMD (EVEX encoded versions)

(KL, VL) = (4, 128), (8, 256), (16, 512)

FOR j := 0 TO KL-1

i := j * 32

IF k1[j] OR *no writemask*

     THEN

             IF (EVEX.b = 1) AND (SRC2 *is memory*)

                  THEN DEST[j] := (SRC1[i+31:i] BITWISE AND SRC2[31:0] != 0)? 1 : 0;

                  ELSE DEST[j] := (SRC1[i+31:i] BITWISE AND SRC2[i+31:i] != 0)? 1 : 0;

             FI;

     ELSE DEST[j] := 0                    ; zeroing-masking only

FI;

ENDFOR

DEST[MAX_KL-1:KL] := 0

VPTESTMQ (EVEX encoded versions)

(KL, VL) = (2, 128), (4, 256), (8, 512)

FOR j := 0 TO KL-1

i := j * 64

IF k1[j] OR *no writemask*

     THEN

             IF (EVEX.b = 1) AND (SRC2 *is memory*)

                  THEN DEST[j] := (SRC1[i+63:i] BITWISE AND SRC2[63:0] != 0)? 1 : 0;

                  ELSE DEST[j] := (SRC1[i+63:i] BITWISE AND SRC2[i+63:i] != 0)? 1 : 0;

             FI;

     ELSE DEST[j] := 0                    ; zeroing-masking only

FI;

ENDFOR

DEST[MAX_KL-1:KL] := 0

Intel C/C++ compiler intrinsics

VPTESTMB __mmask64 _mm512_test_epi8_mask( __m512i a, __m512i b);
VPTESTMB __mmask64 _mm512_mask_test_epi8_mask(__mmask64, __m512i a, __m512i b);
VPTESTMW __mmask32 _mm512_test_epi16_mask( __m512i a, __m512i b);
VPTESTMW __mmask32 _mm512_mask_test_epi16_mask(__mmask32, __m512i a, __m512i b);
VPTESTMD __mmask16 _mm512_test_epi32_mask( __m512i a, __m512i b);
VPTESTMD __mmask16 _mm512_mask_test_epi32_mask(__mmask16, __m512i a, __m512i b);
VPTESTMQ __mmask8 _mm512_test_epi64_mask(__m512i a, __m512i b);
VPTESTMQ __mmask8 _mm512_mask_test_epi64_mask(__mmask8, __m512i a, __m512i b);

SIMD Floating-Point Exceptions

None.

Other Exceptions

VPTESTMD/Q: See Table 2-51, "Type E4 Class Exception Conditions." VPTESTMB/W: See Exceptions Type E4.nb in Table 2-51, "Type E4 Class Exception Conditions."

Sources