Commit 9c82dc6a authored Jun 30, 2023 by Matt Arsenault

AMDGPU: Always use v_rcp_f16 and v_rsq_f16

These inherited the fast math checks from f32, but the manual suggests
these should be accurate enough for unconditional use. The definition
of correctly rounded is 0.5ulp, but the manual says "0.51ulp". I've
been a bit nervous about changing this as the OpenCL conformance test
does not cover half. Brute force produces identical values compared to
a reference host implementation for all values.

parent 59c311c5

Expand all Show whitespace changes

Inline Side-by-side

Please to comment