This is the mail archive of the libc-alpha@sourceware.org mailing list for the glibc project.


Index Nav: [Date Index] [Subject Index] [Author Index] [Thread Index]
Message Nav: [Date Prev] [Date Next] [Thread Prev] [Thread Next]
Other format: [Raw text]

Re: [PATCH v2] aarch64: thunderx2 memcpy branches reordering


Hi Anton,

> Yes, I might if I wasn't in the loop. As I constrained by the finite
> number of register names I need a value in a particular
> register. The fact that it is already in some other register does
> not help me here. We have the window spanning 5 registers
> and we load only 4 registers each iteration. So, how do I update
> the fifth one?

By ensuring the last use of the register happens before you load the
next value into it. After I apply your v3 patch, we get:

        ext     A_v.16b, C_v.16b, D_v.16b, 16-shft;\
        ext     B_v.16b, D_v.16b, E_v.16b, 16-shft;\
1:;\
        stp     A_q, B_q, [dst], #32;\
        prfm    pldl1strm, [src, MEMCPY_PREFETCH_LDR];\
        ldp     C_q, D_q, [src], #32;\
        ext     H_v.16b, E_v.16b, F_v.16b, 16-shft;\
        ext     I_v.16b, F_v.16b, G_v.16b, 16-shft;\
        stp     H_q, I_q, [dst], #32;\
        ext     A_v.16b, G_v.16b, C_v.16b, 16-shft;\
        ldp     F_q, G_q, [src], #32;\
        ext     B_v.16b, C_v.16b, D_v.16b, 16-shft;\
        mov     E_v.16b, D_v.16b;\
        subs    count, count, 64;\
        b.ge    1b;\
2:;\
        stp     A_q, B_q, [dst], #32;\
        ext     H_v.16b, E_v.16b, F_v.16b, 16-shft;\
        ext     I_v.16b, F_v.16b, G_v.16b, 16-shft;\
        b       L(ext_tail);

Now it's easy to see that the mov is redundant, we can execute
ext H_v.16b, D_v.16b, F_v.16b, 16-shft instead! This makes
one ext in the tail redundant, but you need an extra one at entry.

If you also move the final stp into ext_tail, you end up with 16
instructions in total, so now you can align the inner loop too
to maximise prefetch bandwidth.

Note all the software pipelining is overkill for an OoO core. Writing
the simplest possible loop with minimal scheduling should be just as
fast and would allow removal of even more instructions.

Wilco
    

Index Nav: [Date Index] [Subject Index] [Author Index] [Thread Index]
Message Nav: [Date Prev] [Date Next] [Thread Prev] [Thread Next]