Skip to content

String::to_lowercase does not get vectorized well contrary to code comments #123712

Description

@jhorstmann

I'm looking into the performance of to_lowercase / to_uppercaseon mostly ascii strings, using a small microbenchmark added to library/alloc/benches/string.rs.

#[bench]fnbench_to_lowercase(b:&mutBencher){let s = "Hello there, the quick brown fox jumped over the lazy dog! \ Lorem ipsum dolor sit amet, consectetur. ";
b.iter(|| s.to_lowercase())}

Using linux perf tooling I see that the hot part of the code is the following large loop, which despite heavy use of sse2 instructions only seems to process 32 bytes per iteration.

 │ d0:┌─→movr9,QWORD PTR [r14+r15*1] │ │ movdquxmm3,XMMWORD PTR [r14+r15*1]0,12 │ │ pshufdxmm12,xmm3,0xee2,56 │ │ movqrdx,xmm12 │ │ movrsi,rdx │ │ orrsi,r90,59 │ │ testrsi,rcx │ │↓ jne3192,10 │ │ movrsi,r9 │ │ movrdi,r9 │ │ movr8,r9 │ │ movr10,r92,23 │ │ shrr9d,0x8 │ │ movdxmm12,r9d │ │ shrr10,0x200,12 │ │ pshufdxmm13,xmm3,0x441,05 │ │ movdqaxmm14,xmm3 │ │ psrlqxmm14,0x10 │ │ psrlqxmm13,0x180,12 │ │ movsdxmm13,xmm141,75 │ │ movdxmm14,r10d │ │ shrr8,0x28 │ │ punpcklqdqxmm3,xmm12 │ │ movdxmm12,r8d1,52 │ │ andpdxmm13,xmm0 │ │ pandxmm3,xmm0 │ │ packuswbxmm3,xmm130,35 │ │ pshufdxmm13,xmm14,0x501,05 │ │ movdqaxmm14,XMMWORD PTR [rip+0x403c7] │ │ pandnxmm14,xmm13 │ │ psllqxmm12,0x28 │ │ movdqaxmm13,XMMWORD PTR [rip+0x403c3]2,94 │ │ pandnxmm13,xmm12 │ │ shrrdi,0x30 │ │ porxmm13,xmm140,35 │ │ movdxmm12,edi2,22 │ │ shrrsi,0x38 │ │ packuswbxmm3,xmm1 │ │ packuswbxmm3,xmm10,47 │ │ porxmm13,xmm34,21 │ │ psllqxmm12,0x30 │ │ movdqaxmm3,xmm4 │ │ pandnxmm3,xmm120,35 │ │ movdxmm12,esi2,47 │ │ movesi,edx │ │ shresi,0x8 │ │ pandxmm13,xmm40,35 │ │ porxmm3,xmm132,10 │ │ pandxmm3,xmm5 │ │ psllqxmm12,0x38 │ │ movdqaxmm13,xmm5 │ │ pandnxmm13,xmm122,34 │ │ porxmm13,xmm3 │ │ movdxmm3,edx │ │ pshufdxmm3,xmm3,0x440,53 │ │ movdqaxmm12,xmm62,47 │ │ pandnxmm12,xmm3 │ │ movdxmm3,esi │ │ movesi,edx0,23 │ │ shresi,0x102,64 │ │ pandxmm13,xmm6 │ │ porxmm12,xmm13 │ │ pslldqxmm3,0x90,12 │ │ movdqaxmm13,xmm72,45 │ │ pandnxmm13,xmm3 │ │ movdxmm3,esi │ │ movesi,edx0,51 │ │ shresi,0x182,60 │ │ pandxmm12,xmm7 │ │ porxmm13,xmm12 │ │ pslldqxmm3,0xa │ │ movdqaxmm12,xmm81,76 │ │ pandnxmm12,xmm3 │ │ movdxmm3,esi │ │ movrsi,rdx0,47 │ │ shrrsi,0x202,34 │ │ pandxmm13,xmm8 │ │ porxmm12,xmm13 │ │ pslldqxmm3,0xb0,23 │ │ movdqaxmm13,xmm91,99 │ │ pandnxmm13,xmm3 │ │ movdxmm3,esi │ │ movrsi,rdx0,35 │ │ shrrsi,0x282,97 │ │ pandxmm12,xmm9 │ │ porxmm13,xmm12 │ │ pshufdxmm3,xmm3,0x00,12 │ │ movdqaxmm12,xmm102,11 │ │ pandnxmm12,xmm3 │ │ movdxmm3,esi │ │ shrrdx,0x30 │ │ pandxmm13,xmm101,87 │ │ porxmm12,xmm13 │ │ pandxmm12,xmm11 │ │ pslldqxmm3,0xd0,23 │ │ movdqaxmm13,xmm112,23 │ │ pandnxmm13,xmm3 │ │ porxmm13,xmm12 │ │ pandxmm13,XMMWORD PTR [rip+0x40320]0,12 │ │ movdxmm3,edx2,80 │ │ pslldqxmm3,0xe │ │ porxmm3,xmm13 │ │ pandxmm3,XMMWORD PTR [rip+0x4031a]0,23 │ │ movzxedx,BYTE PTR [r14+r15*1+0xf]3,31 │ │ movdxmm12,edx │ │ pslldqxmm12,0xf │ │ porxmm12,xmm30,12 │ │ movdqaxmm3,xmm122,92 │ │ paddbxmm3,XMMWORD PTR [rip+0x31d97] # 1009a0 <anon.cf73386a2f5127d166baeac25be116f0.63.llvm.16014458289627072720+0x459> ▒ │ │ movdqaxmm13,xmm3 │ │ pminubxmm13,xmm150,47 │ │ pcmpeqbxmm13,xmm31,53 │ │ pandxmm13,xmm20,36 │ │ porxmm13,xmm12 │ │ movdqu XMMWORD PTR [rax+r15*1],xmm130,23 │ │ leardx,[r15+0x10]2,34 │ │ addr15,0x200,12 │ │ cmpr15,rbx │ │ movr15,rdx1,64 │ └──jbe d0 

I don't see an easy way to improve the autovectorization of this code, but it should be relatively easy to explicitly vectorize it using portable_simd, and I would like to prepare such a PR if there are no objections. As far as I know, portable_simd is already in use inside core, for example by #103779.

Metadata

Metadata

Assignees

No one assigned

    Labels

    C-optimizationCategory: An issue highlighting optimization opportunities or PRs implementing suchT-libsRelevant to the library team, which will review and decide on the PR/issue.

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions