[SystemZ] Add CodeGen support for v4f32 (80b3af7a) · Commits · Roger Ferrer / llvm-epi

Commit 80b3af7a authored May 05, 2015 by Ulrich Weigand

[SystemZ] Add CodeGen support for v4f32

The architecture doesn't really have any native v4f32 operations except
v4f32->v2f64 and v2f64->v4f32 conversions, with only half of the v4f32
elements being used.  Even so, using vector registers for <4 x float>
and scalarising individual operations is much better than generating
completely scalar code, since there's much less register pressure.
It's also more efficient to do v4f32 comparisons by extending to 2
v2f64s, comparing those, then packing the result.

This particularly helps with llvmpipe.

Based on a patch by Richard Sandiford.

llvm-svn: 236523

parent cd808237

Hide whitespace changes

Inline Side-by-side

Please register or to comment