Using linux perf tooling I see that the hot part of the code is the following large loop, which despite heavy use of sse2 instructions only seems to process 32 bytes per iteration.
│ d0:┌─→movr9,QWORD PTR [r14+r15*1] ▒ │ │ movdquxmm3,XMMWORD PTR [r14+r15*1] ▒0,12 │ │ pshufdxmm12,xmm3,0xee ▒2,56 │ │ movqrdx,xmm12 ▒ │ │ movrsi,rdx ▒ │ │ orrsi,r9 ▒0,59 │ │ testrsi,rcx ▒ │ │↓ jne319 ▒2,10 │ │ movrsi,r9 ▒ │ │ movrdi,r9 ▒ │ │ movr8,r9 ▒ │ │ movr10,r9 ▒2,23 │ │ shrr9d,0x8 ▒ │ │ movdxmm12,r9d ▒ │ │ shrr10,0x20 ◆0,12 │ │ pshufdxmm13,xmm3,0x44 ▒1,05 │ │ movdqaxmm14,xmm3 ▒ │ │ psrlqxmm14,0x10 ▒ │ │ psrlqxmm13,0x18 ▒0,12 │ │ movsdxmm13,xmm14 ▒1,75 │ │ movdxmm14,r10d ▒ │ │ shrr8,0x28 ▒ │ │ punpcklqdqxmm3,xmm12 ▒ │ │ movdxmm12,r8d ▒1,52 │ │ andpdxmm13,xmm0 ▒ │ │ pandxmm3,xmm0 ▒ │ │ packuswbxmm3,xmm13 ▒0,35 │ │ pshufdxmm13,xmm14,0x50 ▒1,05 │ │ movdqaxmm14,XMMWORD PTR [rip+0x403c7] ▒ │ │ pandnxmm14,xmm13 ▒ │ │ psllqxmm12,0x28 ▒ │ │ movdqaxmm13,XMMWORD PTR [rip+0x403c3] ▒2,94 │ │ pandnxmm13,xmm12 ▒ │ │ shrrdi,0x30 ▒ │ │ porxmm13,xmm14 ▒0,35 │ │ movdxmm12,edi ▒2,22 │ │ shrrsi,0x38 ▒ │ │ packuswbxmm3,xmm1 ▒ │ │ packuswbxmm3,xmm1 ▒0,47 │ │ porxmm13,xmm3 ▒4,21 │ │ psllqxmm12,0x30 ▒ │ │ movdqaxmm3,xmm4 ▒ │ │ pandnxmm3,xmm12 ▒0,35 │ │ movdxmm12,esi ▒2,47 │ │ movesi,edx ▒ │ │ shresi,0x8 ▒ │ │ pandxmm13,xmm4 ▒0,35 │ │ porxmm3,xmm13 ▒2,10 │ │ pandxmm3,xmm5 ▒ │ │ psllqxmm12,0x38 ▒ │ │ movdqaxmm13,xmm5 ▒ │ │ pandnxmm13,xmm12 ▒2,34 │ │ porxmm13,xmm3 ▒ │ │ movdxmm3,edx ▒ │ │ pshufdxmm3,xmm3,0x44 ▒0,53 │ │ movdqaxmm12,xmm6 ▒2,47 │ │ pandnxmm12,xmm3 ▒ │ │ movdxmm3,esi ▒ │ │ movesi,edx ▒0,23 │ │ shresi,0x10 ▒2,64 │ │ pandxmm13,xmm6 ▒ │ │ porxmm12,xmm13 ▒ │ │ pslldqxmm3,0x9 ▒0,12 │ │ movdqaxmm13,xmm7 ▒2,45 │ │ pandnxmm13,xmm3 ▒ │ │ movdxmm3,esi ▒ │ │ movesi,edx ▒0,51 │ │ shresi,0x18 ▒2,60 │ │ pandxmm12,xmm7 ▒ │ │ porxmm13,xmm12 ▒ │ │ pslldqxmm3,0xa ▒ │ │ movdqaxmm12,xmm8 ▒1,76 │ │ pandnxmm12,xmm3 ▒ │ │ movdxmm3,esi ▒ │ │ movrsi,rdx ▒0,47 │ │ shrrsi,0x20 ▒2,34 │ │ pandxmm13,xmm8 ▒ │ │ porxmm12,xmm13 ▒ │ │ pslldqxmm3,0xb ▒0,23 │ │ movdqaxmm13,xmm9 ▒1,99 │ │ pandnxmm13,xmm3 ▒ │ │ movdxmm3,esi ▒ │ │ movrsi,rdx ▒0,35 │ │ shrrsi,0x28 ▒2,97 │ │ pandxmm12,xmm9 ▒ │ │ porxmm13,xmm12 ▒ │ │ pshufdxmm3,xmm3,0x0 ▒0,12 │ │ movdqaxmm12,xmm10 ▒2,11 │ │ pandnxmm12,xmm3 ▒ │ │ movdxmm3,esi ▒ │ │ shrrdx,0x30 ▒ │ │ pandxmm13,xmm10 ▒1,87 │ │ porxmm12,xmm13 ▒ │ │ pandxmm12,xmm11 ▒ │ │ pslldqxmm3,0xd ▒0,23 │ │ movdqaxmm13,xmm11 ▒2,23 │ │ pandnxmm13,xmm3 ▒ │ │ porxmm13,xmm12 ▒ │ │ pandxmm13,XMMWORD PTR [rip+0x40320] ▒0,12 │ │ movdxmm3,edx ▒2,80 │ │ pslldqxmm3,0xe ▒ │ │ porxmm3,xmm13 ▒ │ │ pandxmm3,XMMWORD PTR [rip+0x4031a] ▒0,23 │ │ movzxedx,BYTE PTR [r14+r15*1+0xf] ▒3,31 │ │ movdxmm12,edx ▒ │ │ pslldqxmm12,0xf ▒ │ │ porxmm12,xmm3 ▒0,12 │ │ movdqaxmm3,xmm12 ▒2,92 │ │ paddbxmm3,XMMWORD PTR [rip+0x31d97] # 1009a0 <anon.cf73386a2f5127d166baeac25be116f0.63.llvm.16014458289627072720+0x459> ▒ │ │ movdqaxmm13,xmm3 ▒ │ │ pminubxmm13,xmm15 ▒0,47 │ │ pcmpeqbxmm13,xmm3 ▒1,53 │ │ pandxmm13,xmm2 ▒0,36 │ │ porxmm13,xmm12 ▒ │ │ movdqu XMMWORD PTR [rax+r15*1],xmm13 ▒0,23 │ │ leardx,[r15+0x10] ▒2,34 │ │ addr15,0x20 ▒0,12 │ │ cmpr15,rbx ▒ │ │ movr15,rdx ▒1,64 │ └──jbe d0
I don't see an easy way to improve the autovectorization of this code, but it should be relatively easy to explicitly vectorize it using portable_simd, and I would like to prepare such a PR if there are no objections. As far as I know, portable_simd is already in use inside core, for example by #103779.
I'm looking into the performance of
to_lowercase/to_uppercaseon mostly ascii strings, using a small microbenchmark added tolibrary/alloc/benches/string.rs.Using linux perf tooling I see that the hot part of the code is the following large loop, which despite heavy use of sse2 instructions only seems to process 32 bytes per iteration.
I don't see an easy way to improve the autovectorization of this code, but it should be relatively easy to explicitly vectorize it using
portable_simd, and I would like to prepare such a PR if there are no objections. As far as I know,portable_simdis already in use insidecore, for example by #103779.