This is the mail archive of the
libc-alpha@sourceware.org
mailing list for the glibc project.
Re: [PATCH v2] aarch64: thunderx2 memcpy branches reordering
Hi Anton,
> Yes, I might if I wasn't in the loop. As I constrained by the finite
> number of register names I need a value in a particular
> register. The fact that it is already in some other register does
> not help me here. We have the window spanning 5 registers
> and we load only 4 registers each iteration. So, how do I update
> the fifth one?
By ensuring the last use of the register happens before you load the
next value into it. After I apply your v3 patch, we get:
ext A_v.16b, C_v.16b, D_v.16b, 16-shft;\
ext B_v.16b, D_v.16b, E_v.16b, 16-shft;\
1:;\
stp A_q, B_q, [dst], #32;\
prfm pldl1strm, [src, MEMCPY_PREFETCH_LDR];\
ldp C_q, D_q, [src], #32;\
ext H_v.16b, E_v.16b, F_v.16b, 16-shft;\
ext I_v.16b, F_v.16b, G_v.16b, 16-shft;\
stp H_q, I_q, [dst], #32;\
ext A_v.16b, G_v.16b, C_v.16b, 16-shft;\
ldp F_q, G_q, [src], #32;\
ext B_v.16b, C_v.16b, D_v.16b, 16-shft;\
mov E_v.16b, D_v.16b;\
subs count, count, 64;\
b.ge 1b;\
2:;\
stp A_q, B_q, [dst], #32;\
ext H_v.16b, E_v.16b, F_v.16b, 16-shft;\
ext I_v.16b, F_v.16b, G_v.16b, 16-shft;\
b L(ext_tail);
Now it's easy to see that the mov is redundant, we can execute
ext H_v.16b, D_v.16b, F_v.16b, 16-shft instead! This makes
one ext in the tail redundant, but you need an extra one at entry.
If you also move the final stp into ext_tail, you end up with 16
instructions in total, so now you can align the inner loop too
to maximise prefetch bandwidth.
Note all the software pipelining is overkill for an OoO core. Writing
the simplest possible loop with minimal scheduling should be just as
fast and would allow removal of even more instructions.
Wilco