Commits · ce2258c1cd5dc9cf20040d1b1e540d80250c1435 · Lorenzo Albano / LLVM bpEVL

Apr 02, 2020
- clang/AMDGPU: Stop setting old denormal subtarget features · ce2258c1
  Matt Arsenault authored Apr 01, 2020
  
  ce2258c1
Mar 28, 2020

[AMDGPU] Add __builtin_amdgcn_workgroup_size_x/y/z · 369e26ca

Yaxun (Sam) Liu authored Mar 25, 2020

The main purpose of introducing these builtins is to add a range
metadata [1, 1025) on the work group size loaded from dispatch
ptr, which cannot be done by source code.

Differential Revision: https://reviews.llvm.org/D76772

369e26ca

Mar 25, 2020

Implement post-commit comments for D75685/rG86e0a6c60627 · fe5c719e

Erich Keane authored Mar 25, 2020

@Anastasia made a pair of comments on D75685 after it was committed
requesting changes to the test.  This patch updates the test based on
her comments.

fe5c719e

Add MS Mangling for OpenCL Pipe types, add mangling test. · 86e0a6c6

Erich Keane authored Mar 05, 2020

SPIRV2.0 Spec only specifies Linux mangling, however our downstream has
use for a Windows mangling for these types.

Unfortunately, the SPIRV
spec specifies a single mangling for all pipe types, despite clang
allowing overloading on these types.  Because of this, this patch
chooses to mangle the read/writability and element type for the windows
mangling.

The windows manglings in the test all demangle according to demangler:
"void __cdecl test1(struct __clang::ocl_pipe<int,1>)
"void __cdecl test2(struct __clang::ocl_pipe<float,0>)
"void __cdecl test2(struct __clang::ocl_pipe<int,1>)
"void __cdecl test3(struct __clang::ocl_pipe<int const,1>)
"void __cdecl test4(struct __clang::ocl_pipe<union
__clang::__vector<unsigned char,3>,1>)
"void __cdecl test5(struct __clang::ocl_pipe<union
__clang::__vector<int,4>,1>)
"void __cdecl test_reserved_read_pipe(struct __clang::_ASCLglobal<struct
Person > * __ptr64,struct __clang::ocl_pipe<struct Person,1>)

Differential Revision: https://reviews.llvm.org/D75685

86e0a6c6

Mar 24, 2020

[CodeGen] Add an alignment attribute to all sret parameters · de98cf92

Erik Pilkington authored Mar 24, 2020

This fixes a miscompile when the parameter is actually underaligned.
rdar://58316406

Differential revision: https://reviews.llvm.org/D74183

de98cf92

Mar 23, 2020
- AMDGPU: Emit llvm.fshr for __builtin_amdgcn_alignbit · 3f533006
  Matt Arsenault authored Mar 19, 2020
```
These are equivalent. The generic rotate builtins do not directly map
to the fshr intrinsic.
```
  3f533006
Mar 09, 2020

Recommit #2 "[Driver] Default to -fno-common for all targets" · 3d9a0445

Sjoerd Meijer authored Mar 09, 2020

After a first attempt to fix the test-suite failures, my first recommit
caused the same failures again. I had updated CMakeList.txt files of
tests that needed -fcommon, but it turns out that there are also
Makefiles which are used by some bots, so I've updated these Makefiles
now too.

See the original commit message for more details on this change:
0a9fc923

3d9a0445

Revert "Recommit "[Driver] Default to -fno-common for all targets"" · f35d112e
Sjoerd Meijer authored Mar 09, 2020
```
This reverts commit 2c36c23f.

Still problems in the test-suite, which I really thought I had fixed...
```
f35d112e

Recommit "[Driver] Default to -fno-common for all targets" · 2c36c23f

Sjoerd Meijer authored Mar 09, 2020

This includes fixes for:
- test-suite: some benchmarks need to be compiled with -fcommon, see D75557.
- compiler-rt: one test needed -fcommon, and another a change, see D75520.

2c36c23f

Mar 06, 2020

Reapply "clang: Treat ieee mode as the default for denormal-fp-math" · 00b2a9df

Matt Arsenault authored Mar 05, 2020

This reverts commit 737394c4.

The fp-model test was failing on platforms that enable denormal flushing
based on -ffast-math. This needs to reset to IEEE, not the default in
these cases.

Change-Id: Ibbad32f66d0d0b89b9c1173a3a96fb1a570ddd89

00b2a9df

Mar 05, 2020

Revert "clang: Treat ieee mode as the default for denormal-fp-math" · 737394c4

Jeremy Morse authored Mar 05, 2020

This reverts commit c64ca930.

This patch tripped a few build bots:

  http://lab.llvm.org:8011/builders/clang-x86_64-debian-fast/builds/24703/
  http://lab.llvm.org:8011/builders/clang-cmake-x86_64-avx2-linux/builds/13465/
  http://lab.llvm.org:8011/builders/clang-with-lto-ubuntu/builds/15994/

Reverting to clear the bots.

737394c4

clang: Treat ieee mode as the default for denormal-fp-math · c64ca930

Matt Arsenault authored Nov 05, 2019

The IR hasn't switched the default yet, so explicitly add the ieee
attributes.

I'm still not really sure how the target default denormal mode should
interact with -fno-unsafe-math-optimizations. The target may have
selected the default mode to be non-IEEE based on the flags or based
on its true behavior, but we don't know which is the case. Since the
only users of a non-IEEE mode without a flag still support IEEE mode,
just reset to IEEE.

c64ca930

Mar 03, 2020

Revert "[Driver] Default to -fno-common for all targets" · 4e363563

Sjoerd Meijer authored Mar 03, 2020

This reverts commit 0a9fc923.

Going to look at the asan failures.

I find the failures in the test suite weird, because they look
like compile time test and I don't understand how that can be
failing, but will have a brief look at that too.

4e363563

[Driver] Default to -fno-common for all targets · 0a9fc923

Sjoerd Meijer authored Mar 03, 2020

This makes -fno-common the default for all targets because this has performance
and code-size benefits and is more language conforming for C code.
Additionally, GCC10 also defaults to -fno-common and so we get consistent
behaviour with GCC.

With this change, C code that uses tentative definitions as definitions of a
variable in multiple translation units will trigger multiple-definition linker
errors. Generally, this occurs when the use of the extern keyword is neglected
in the declaration of a variable in a header file. In some cases, no specific
translation unit provides a definition of the variable. The previous behavior
can be restored by specifying -fcommon.

As GCC has switched already, we benefit from applications already being ported
and existing documentation how to do this. For example:
- https://gcc.gnu.org/gcc-10/porting_to.html
- https://wiki.gentoo.org/wiki/Gcc_10_porting_notes/fno_common

Differential revision: https://reviews.llvm.org/D75056

0a9fc923

Feb 25, 2020
- Make __builtin_amdgcn_dispatch_ptr dereferenceable and align at 4 · a57d9652
  Yaxun (Sam) Liu authored Feb 20, 2020
```
Differential Revision: https://reviews.llvm.org/D75028
```
  a57d9652
Feb 17, 2020

[OpenCL][CUDA][HIP][SYCL] Add norecurse · fb44b9db

Yaxun (Sam) Liu authored Jan 29, 2020

norecurse function attr indicates the function is not called recursively
directly or indirectly.

Add norecurse to OpenCL functions, SYCL functions in device compilation
and CUDA/HIP kernels.

Although there is LLVM pass adding norecurse to functions, it only works
for whole-program compilation. Also FE adding norecurse can make that
pass run faster since functions with norecurse do not need to be checked
again.

Differential Revision: https://reviews.llvm.org/D73651

fb44b9db

Jan 28, 2020

Corrected clang amdgpu-features.cl test for... · 987aa343

Konstantin Pyzhov authored Jan 28, 2020

Corrected clang amdgpu-features.cl test for 6d614a82 (AMDGPU MFMA built-ins)

Differential Revision: https://reviews.llvm.org/D72723

987aa343

Add missing clang tests for 6d614a82 (AMDGPU MFMA built-ins) · ac9b2a62
Konstantin Pyzhov authored Jan 28, 2020
```
Differential Revision: https://reviews.llvm.org/D72723
```
ac9b2a62

Summary: · 6d614a82

Konstantin Pyzhov authored Jan 28, 2020

This CL adds clang declarations of built-in functions for AMDGPU MFMA intrinsics and instructions.
OpenCL tests for new built-ins are included.

Differential Revision: https://reviews.llvm.org/D72723

6d614a82

Jan 18, 2020

Consolidate internal denormal flushing controls · a4451d88

Matt Arsenault authored Nov 01, 2019

Currently there are 4 different mechanisms for controlling denormal
flushing behavior, and about as many equivalent frontend controls.

- AMDGPU uses the fp32-denormals and fp64-f16-denormals subtarget features
- NVPTX uses the nvptx-f32ftz attribute
- ARM directly uses the denormal-fp-math attribute
- Other targets indirectly use denormal-fp-math in one DAGCombine
- cl-denorms-are-zero has a corresponding denorms-are-zero attribute

AMDGPU wants a distinct control for f32 flushing from f16/f64, and as
far as I can tell the same is true for NVPTX (based on the attribute
name).

Work on consolidating these into the denormal-fp-math attribute, and a
new type specific denormal-fp-math-f32 variant. Only ARM seems to
support the two different flush modes, so this is overkill for the
other use cases. Ideally we would error on the unsupported
positive-zero mode on other targets from somewhere.

Move the logic for selecting the flush mode into the compiler driver,
instead of handling it in cc1. denormal-fp-math/denormal-fp-math-f32
are now both cc1 flags, but denormal-fp-math-f32 is not yet exposed as
a user flag.

-cl-denorms-are-zero, -fcuda-flush-denormals-to-zero and
-fno-cuda-flush-denormals-to-zero will be mapped to
-fp-denormal-math-f32=ieee or preserve-sign rather than the old
attributes.

Stop emitting the denorms-are-zero attribute for the OpenCL flag. It
has no in-tree users. The meaning would also be target dependent, such
as the AMDGPU choice to treat this as only meaning allow flushing of
f32 and not f16 or f64. The naming is also potentially confusing,
since DAZ in other contexts refers to instructions implicitly treating
input denormals as zero, not necessarily flushing output denormals to
zero.

This also does not attempt to change the behavior for the current
attribute. The LangRef now states that the default is ieee behavior,
but this is inaccurate for the current implementation. The clang
handling is slightly hacky to avoid touching the existing
denormal-fp-math uses. Fixing this will be left for a future patch.

AMDGPU is still using the subtarget feature to control the denormal
mode, but the new attribute are now emitted. A future change will
switch this and remove the subtarget features.

a4451d88

Jan 17, 2020
- AMDGPU: Update clang test · 9b549f26
  Matt Arsenault authored Jan 16, 2020
  
  9b549f26
Dec 03, 2019

[OpenCL] Fix mangling of single-overload builtins · 6713670b

Sven van Haastregt authored Dec 03, 2019

Commit 9a8d477a ("[OpenCL] Add builtin function attribute
handling", 2019-11-05) stopped Clang from mangling single-overload
builtins, which is incorrect.

6713670b

Nov 05, 2019

[OpenCL] Add builtin function attribute handling · 9a8d477a

Sven van Haastregt authored Nov 05, 2019

Add handling for the "pure", "const" and "convergent" function
attributes for OpenCL builtin functions.

Patch by Pierre Gondois and Sven van Haastregt.

Differential Revision: https://reviews.llvm.org/D64319

9a8d477a

Sep 05, 2019
- AMDGPU: Add builtins for is_shared/is_private · 281f2e2c
  Matt Arsenault authored Sep 05, 2019
```
llvm-svn: 371010
```
  281f2e2c
Aug 27, 2019

AMDGPU: Always emit amdgpu-flat-work-group-size · eac783a9

Matt Arsenault authored Aug 27, 2019

The backend default maximum should be the hardware maximum, so the
frontend should set the implementation defined default maximum.

llvm-svn: 370101

eac783a9

Aug 06, 2019

Builtins: Start adding half versions of math builtins · acd0a53c

Matt Arsenault authored Aug 06, 2019

The implementation of the OpenCL builtin currently library uses 2
different hacks to get to the corresponding IR intrinsics from the
source. This will allow removal of those.

This is the set that is currently used (minus a few vector ones).

llvm-svn: 367973

acd0a53c

Aug 05, 2019
- [OpenCL] Fix vector literal test broken in rL367675. · ab4a5d14
  Anastasia Stulova authored Aug 05, 2019
```
Avoid checking alignment unnecessary that is not portable
among targets.

llvm-svn: 367823
```
  ab4a5d14
Aug 03, 2019

IR: print value numbers for unnamed function arguments · a009a60a

Tim Northover authored Aug 03, 2019

For consistency with normal instructions and clarity when reading IR,
it's best to print the %0, %1, ... names of function arguments in
definitions.

Also modifies the parser to accept IR in that form for obvious reasons.

llvm-svn: 367755

a009a60a

Aug 02, 2019

[OpenCL] Allow OpenCL C style vector initialization in C++ · 8d99a5c0

Anastasia Stulova authored Aug 02, 2019

Allow creating vector literals from other vectors.

 float4 a = (float4)(1.0f, 2.0f, 3.0f, 4.0f);
 float4 v = (float4)(a.s23, a.s01);

Differential revision: https://reviews.llvm.org/D65286

llvm-svn: 367675

8d99a5c0

Jul 31, 2019
- AMDGPU: Add missing builtin declarations · 64d7af09
  Matt Arsenault authored Jul 31, 2019
```
llvm-svn: 367431
```
  64d7af09
Jul 25, 2019

[OpenCL] Rename lang mode flag for C++ mode · 88ed70e2

Anastasia Stulova authored Jul 25, 2019

Rename lang mode flag to -cl-std=clc++/-cl-std=CLC++
or -std=clc++/-std=CLC++.

This aligns with OpenCL C conversion and removes ambiguity
with OpenCL C++. 

Differential Revision: https://reviews.llvm.org/D65102

llvm-svn: 367008

88ed70e2

Jul 22, 2019

Updated the signature for some stack related intrinsics (CLANG) · 8c5e6fa6

Christudasan Devadasan authored Jul 22, 2019

Modified the intrinsics
int_addressofreturnaddress,
int_frameaddress & int_sponentry.
This commit depends on the changes in rL366679

Reviewed By: arsenm

Differential Revision: https://reviews.llvm.org/D64563

llvm-svn: 366683

8c5e6fa6

Jul 17, 2019
- AMDGPU: Add some missing builtins · e56865d4
  Matt Arsenault authored Jul 17, 2019
```
llvm-svn: 366286
```
  e56865d4
Jul 16, 2019

[OpenCL] Fixing sampler initialisations for C++ mode. · 8ece3b67

Neil Hickey authored Jul 16, 2019

Allow conversions between integer and sampler type.

Differential Revision: https://reviews.llvm.org/D64791

llvm-svn: 366212

8ece3b67

Jul 10, 2019

[clang] Preserve names of addrspacecast'ed values. · de811d1f
Vyacheslav Zakharin authored Jul 10, 2019
```
Differential Revision: https://reviews.llvm.org/D63846

llvm-svn: 365666
```
de811d1f

[AMDGPU] Increased the number of implicit argument bytes for both OpenCL and HIP (CLANG). · 18ba9d60

Christudasan Devadasan authored Jul 10, 2019

To enable a new implicit kernel argument,
increased the number of argument bytes from 48 to 56.

Reviewed By: yaxunl

Differential Revision: https://reviews.llvm.org/D63756

llvm-svn: 365643

18ba9d60

Jul 09, 2019

Use the Itanium C++ ABI for the pipe_builtin.cl test · 9b28d9c3

Reid Kleckner authored Jul 09, 2019

Certain OpenCL constructs cannot yet be mangled in the MS C++ ABI.
Add a FIXME for it if anyone cares to implement it.

llvm-svn: 365557

9b28d9c3

[AMDGPU] gfx908 clang target · 0cfd75a0
Stanislav Mekhanoshin authored Jul 09, 2019
```
Differential Revision: https://reviews.llvm.org/D64430

llvm-svn: 365528
```
0cfd75a0

[OpenCL][Sema] Fix builtin rewriting · b00d5f73

Marco Antognini authored Jul 09, 2019

This patch ensures built-in functions are rewritten using the proper
parent declaration.

Existing tests are modified to run in C++ mode to ensure the
functionality works also with C++ for OpenCL while not increasing the
testing runtime.

llvm-svn: 365499

b00d5f73

Jul 08, 2019

Add nofree attribute to CodeGenOpenCL/convergent.cl test · e6ba2254

Brian Homerding authored Jul 08, 2019

The revision at https://reviews.llvm.org/rL365336 added inference of the nofree
attribute.  This revision updates the test to reflect this.

Differential Revision: https://reviews.llvm.org/D49165

llvm-svn: 365341

e6ba2254