[PATCH][AArch64] Cleanup memset

Siddhesh Poyarekar siddhesh@gotplt.org
Tue Mar 3 15:23:00 GMT 2020


On 03/03/20 20:12, Wilco Dijkstra wrote:
> Hi Siddhesh,
> 
>> This looks OK in general, although can you please elaborate on the
>> following:
>>
>> - What cores did you test on to conclude that 160 is a better threshold
>> than 256?
> 
> I've mostly done testing on Neoverse N1, Cortex-A72 and Cortex-A53.
> (the latter seems to be always faster with DC ZVA disabled, so the threshold
> doesn't really matter). I wrote a random memset benchmark similar to the
> memcpy one, and performance is unchanged there given there is no change
> in the way small cases are handled.

OK, please mention this in the commit log.

>> - Is the intention to support non-64-byte zva sizes once there is actual
>> hardware that implements it and not bother with it for now?  I agree
>> with the idea if that's the case, just that it would be nice to have
>> that documented in the git commit message.
> 
> Yes, otherwise it's hard to test or prove it helps performance after all. We've
> had issues with the non-64 ZVA sizes before, so it's best to keep it simple.
> 
> I'm also trying to reduce the amount of code and avoid unnecessary proliferation
> of almost identical ifuncs. I think we can remove most of the memset ifuncs,
> it seems we need one version without ZVA and a ZVA version for size 64.

Sounds like a plan.

Thanks,
Siddhesh



More information about the Libc-alpha mailing list