Optimize `pcg32_random_r` utilizing more instruction level parallelism #17

xiver77 · 2022-07-14T14:06:47Z

I had also sent you (the author) an email a while ago about the same issue.

The original code has a chain of shift -> xor -> shift, while the modified code has a chain of shift | shift -> xor. Two shifts can run in parallel.

Clang (14.0) does this optimization even with the original code, but GCC (12.1) doesn't, so it produces better code with manual optimization.

You can compare the machine code output (https://godbolt.org/z/5bMjdsYnf).

I had also sent you (the author) an email a while ago about the same issue. The original code has a chain of shift -> xor -> shift, while the modified code has a chain of shift | shift -> xor. Two shifts can run in parallel. Clang (14.0) does this optimization even with the original code, but GCC (12.1) doesn't, so it produces better code with manual optimization. You can compare the machine code output (https://godbolt.org/z/5bMjdsYnf).

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Optimize `pcg32_random_r` utilizing more instruction level parallelism #17

Optimize `pcg32_random_r` utilizing more instruction level parallelism #17

xiver77 commented Jul 14, 2022

Optimize pcg32_random_r utilizing more instruction level parallelism #17

Are you sure you want to change the base?

Optimize pcg32_random_r utilizing more instruction level parallelism #17

Conversation

xiver77 commented Jul 14, 2022

Optimize `pcg32_random_r` utilizing more instruction level parallelism #17

Optimize `pcg32_random_r` utilizing more instruction level parallelism #17