[PATCH][AArch64] Cleanup memset
Andrew Pinski
pinskia@gmail.com
Thu Mar 12 17:06:08 GMT 2020
On Thu, Mar 12, 2020 at 8:46 AM Wilco Dijkstra <Wilco.Dijkstra@arm.com> wrote:
>
> Hi Andrew,
>
> >> Yes, otherwise it's hard to test or prove it helps performance after all. We've
> >> had issues with the non-64 ZVA sizes before, so it's best to keep it simple.
> >>
> >> I'm also trying to reduce the amount of code and avoid unnecessary proliferation
> >> of almost identical ifuncs. I think we can remove most of the memset ifuncs,
> >> it seems we need one version without ZVA and a ZVA version for size 64.
> >
> > I am trying to understand what was the decision here. The main reason
> > is OcteonTX2 does ZVA 128 which was faster than doing one without
> > (OcteonTX1 is similar but has an errata which causes ZVA to be turned
> > off).
>
> So you're saying there may soon be a new microarchitecture which supports
> ZVA size 128? Is there a significant speedup? Is querying the ZVA size fast?
By new you mean just publicly announced within a month that supports
ZVA size of 128; OcteonTX 2 annoucment:
https://www.marvell.com/company/newsroom/marvell-announces-octeon-tx2-family-of-multi-core-infrastructure-processors.html
?
Yes there is significant speedup for using "DC ZVA" as the latency is
one cycle compared to the multiple (8) cycles needed to fill the full
128byte cache line using stp.
Yes querying ZVA size fast, only a 1 cycles latency.
Thanks,
Andrew
>
> Cheers,
> Wilco
>
More information about the Libc-alpha
mailing list