Commits · 68e11a6ecac7a37a5772d8c5c9c56c19614fc7f0 · Lorenzo Albano / LLVM bpEVL

Mar 23, 2018

[AMDGPU] Update OpenCL to use 48 bytes of implicit arguments for AMDGPU (CLANG) · 68e11a6e

Tony Tye authored Mar 23, 2018

Add two additional implicit arguments for OpenCL for the AMDGPU target using the AMDHSA runtime to support device enqueue.

Differential Revision: https://reviews.llvm.org/D44696

llvm-svn: 328350

68e11a6e

[AMDGPU] Remove use of OpenCL triple environment and replace with function attribute for AMDGPU · 7a893d4e

Tony Tye authored Mar 23, 2018


- Remove use of the opencl and amdopencl environment member of the target triple for the AMDGPU target.
- Use function attribute to communicate to the AMDGPU backend to add implicit arguments for OpenCL kernels for the AMDHSA OS.

Differential Revision: https://reviews.llvm.org/D43736

llvm-svn: 328349

7a893d4e

[PDB] Make our PDBs look more like MS PDBs. · a6fb536e

Zachary Turner authored Mar 23, 2018

When investigating bugs in PDB generation, the first step is
often to do the same link with link.exe and then compare PDBs.

But comparing PDBs is hard because two completely different byte
sequences can both be correct, so it hampers the investigation when
you also have to spend time figuring out not just which bytes are
different, but also if the difference is meaningful.

This patch fixes a couple of cases related to string table emission,
hash table emission, and the order in which we emit strings that
makes more of our bytes the same as the bytes generated by MS PDBs.

Differential Revision: https://reviews.llvm.org/D44810

llvm-svn: 328348

a6fb536e

[AMDGPU] Remove use of OpenCL triple environment and replace with function... · 1a3f3a2d

Tony Tye authored Mar 23, 2018

[AMDGPU] Remove use of OpenCL triple environment and replace with function attribute for AMDGPU (CLANG)


- Remove use of the opencl and amdopencl environment member of the target triple for the AMDGPU target.
- Use a function attribute to communicate to the AMDGPU backend.

Differential Revision: https://reviews.llvm.org/D43735

llvm-svn: 328347

1a3f3a2d

[Hexagon] Always generate mux out of predicated transfers if possible · 5f7ba9a7

Krzysztof Parzyszek authored Mar 23, 2018

HexagonGenMux would collapse pairs of predicated transfers if it assumed
that the predicated .new forms cannot be created. Turns out that generating
mux is preferable in almost all cases.
Introduce an option -hexagon-gen-mux-threshold that controls the minimum
distance between the instruction defining the predicate and the later of
the two transfers. If the distance is closer than the threshold, mux will
not be generated. Set the threshold to 0 by default.

llvm-svn: 328346

5f7ba9a7

Delete the copy constructor for llvm::yaml::Node · dc86de6b

Jordan Rose authored Mar 23, 2018

The nodes keep a reference back to the original document, but the
document is streamed, not read all into memory at once, and the
position is part of the state. If nodes are ever copied, the document
position can end up being advanced more than once.

This did not reveal any problems in LLVM or Clang but caught a handful
over in Swift!

llvm-svn: 328345

dc86de6b

[Hexagon] Avoid early if-conversion for one sided branches · 80f10e4f
Krzysztof Parzyszek authored Mar 23, 2018
```
Patch by Anand Kodnani.

llvm-svn: 328344
```
80f10e4f
[X86][Btver2] Cleanup TEST instructions to use JFPA (+JFPX on ymms) function unit · 6c63e6c2
Simon Pilgrim authored Mar 23, 2018
```
llvm-svn: 328343
```
6c63e6c2

[HWASan] Port HWASan to Linux x86-64 (LLVM) · 83e78414

Alex Shlyapnikov authored Mar 23, 2018

Summary:
Porting HWASan to Linux x86-64, first of the three patches, LLVM part.

The approach is similar to ARM case, trap signal is used to communicate
memory tag check failure. int3 instruction is used to generate a signal,
access parameters are stored in nop [eax + offset] instruction immediately
following the int3 one.

One notable difference is that x86-64 has to untag the pointer before use
due to the lack of feature comparable to ARM's TBI (Top Byte Ignore).

Reviewers: eugenis

Subscribers: kristof.beyls, llvm-commits

Differential Revision: https://reviews.llvm.org/D44699

llvm-svn: 328342

83e78414

[ARM] Fix "Constant pool entry out of range!" in Thumb1 mode · 41573804

Ana Pazos authored Mar 23, 2018

This patch fixes PR36658, "Constant pool entry out of range!" in Thumb1 mode.

In ARMConstantIslands::optimizeThumb2JumpTables() in Thumb1 mode,
adjustBBOffsetsAfter() is not calculating postOffset correctly by
properly accounting for the padding that is required for the constant pool
that immediately follows the jump table branch  instruction.

Reviewers: t.p.northover, eli.friedman

Reviewed By: t.p.northover

Subscribers: chrib, tstellar, javed.absar, kristof.beyls, llvm-commits

Differential Revision: https://reviews.llvm.org/D44709

llvm-svn: 328341

41573804

[llvm-mca] update the ResourcePressureView after r328335. NFC. · 083960d1
Andrea Di Biagio authored Mar 23, 2018
```
This should have been part of r328335. I forgot to svn add these files.

llvm-svn: 328340
```
083960d1

[Hexagon] Two fixes in early if-conversion · 570c6440

Krzysztof Parzyszek authored Mar 23, 2018

- Fix checking for vector predicate registers.
- Avoid speculating llvm.lifetime.end intrinsic.

Patch by Harsha Jagasia and Brendon Cahoon.

llvm-svn: 328339

570c6440

[X86][Btver2] Cleanup MOVMSK instructions to use JFPA function unit · e5c0a041
Simon Pilgrim authored Mar 23, 2018
```
Add missing non-VEX and (V)PMOVMSKB instructions to the pattern

llvm-svn: 328338
```
e5c0a041

[vfs] Don't bail out after a missing -ivfsoverlay file · 005c2e57

Ben Langmuir authored Mar 23, 2018

This make -ivfsoverlay behave more like other fatal errors (e.g. missing
-include file) by skipping the missing file instead of bailing out of
the whole compilation. This makes it possible for libclang to still
provide some functionallity as well as to correctly produce the fatal
error diagnostic (previously we lost the diagnostic in libclang since
there was no TU to tie it to).

rdar://33385423

llvm-svn: 328337

005c2e57

Fix a block copying problem in LICM · a237866f
Andrew Kaylor authored Mar 23, 2018
```
Differential Revision: https://reviews.llvm.org/D44817

llvm-svn: 328336
```
a237866f

[llvm-mca] Make the resource cost a double. · 51dba7d3

Andrea Di Biagio authored Mar 23, 2018

This is done in preparation for the fix for PR36874.
The number of cycles consumed for each pipe is now a double quantity. This
allows reuse of the resource pressure view to print out instruction tables.

llvm-svn: 328335

51dba7d3

[ADT] Simplify getMemory. NFC · c244a158
Fangrui Song authored Mar 23, 2018
```
llvm-svn: 328334
```
c244a158

[Hexagon] Copy subregisters in HexagonStoreWiden · c98802de

Krzysztof Parzyszek authored Mar 23, 2018

When converting an instruction to the wider version, copy any
subregisters if the original operand has a subregister.

Patch by Brendon Cahoon.

llvm-svn: 328333

c98802de

Add a minimal fix for PR36878. · 4376cffb

Rafael Espindola authored Mar 23, 2018

When looking for the output section and the output offset the
expectation was that the caller had looked at Repl. That works fine
for InputSections, but in the case of MergeInputSections the caller
doesn't have the section that is actually replaced.

The original testcase was failing because getOutputSection was
returning null. The slightly extended testcase also checks that
getOffset also checks Repl.

I will send a refactoring separetelly.

llvm-svn: 328332

4376cffb

[X86][Btver2] Vector permutes use a JFPU01 scheduler pipe and JFPX/JVALU function unit · 256f149b
Simon Pilgrim authored Mar 23, 2018
```
llvm-svn: 328331
```
256f149b
[InstCombine] auto-generate checks; NFC · cd1f3e7a
Sanjay Patel authored Mar 23, 2018
```
llvm-svn: 328329
```
cd1f3e7a
[X86][Btver2] Vector store instructions use a JFPU1 scheduler pipe and JSAGU/JSTC function units · ee282b31
Simon Pilgrim authored Mar 23, 2018
```
llvm-svn: 328328
```
ee282b31
[InstSimplify] regenerate checks, move tests; NFC · 3547dcb3
Sanjay Patel authored Mar 23, 2018
```
llvm-svn: 328327
```
3547dcb3

Re-commit: [MachineLICM] Add functions to MachineLICM to hoist invariant stores · 65359936

Zaara Syeda authored Mar 23, 2018

This patch adds functions to allow MachineLICM to hoist invariant stores.
Currently, MachineLICM does not hoist any store instructions, however
when storing the same value to a constant spot on the stack, the store
instruction should be considered invariant and be hoisted. The function
isInvariantStore iterates each operand of the store instruction and checks
that each register operand satisfies isCallerPreservedPhysReg. The store
may be fed by a copy, which is hoisted by isCopyFeedingInvariantStore.
This patch also adds the PowerPC changes needed to consider the stack
register as caller preserved.

Differential Revision: https://reviews.llvm.org/D40196

llvm-svn: 328326

65359936

[InstCombine] regenerate test checks; NFC · d189b596
Sanjay Patel authored Mar 23, 2018
```
llvm-svn: 328325
```
d189b596
[X86][Btver2] Cleanup DPPS/DPPD instructions to use JFPA/JFPM function units · 1335b9c0
Simon Pilgrim authored Mar 23, 2018
```
llvm-svn: 328324
```
1335b9c0
[InstCombine] reduce code duplication; NFC · 713ca3d3
Sanjay Patel authored Mar 23, 2018
```
llvm-svn: 328323
```
713ca3d3
[InstCombine] improve variable name; NFC · 6de89ce3
Sanjay Patel authored Mar 23, 2018
```
llvm-svn: 328322
```
6de89ce3

[AArch64] Don't reduce the width of loads if it prevents combining a shift · e3b44f9d

John Brawn authored Mar 23, 2018

Loads and stores can only shift the offset register by the size of the value
being loaded, but currently the DAGCombiner will reduce the width of the load
if it's followed by a trunc making it impossible to later combine the shift.

Solve this by implementing shouldReduceLoadWidth for the AArch64 backend and
make it prevent the width reduction if this is what would happen, though do
allow it if reducing the load width will let us eliminate a later sign or zero
extend.

Differential Revision: https://reviews.llvm.org/D44794

llvm-svn: 328321

e3b44f9d

[X86][Btver2] Fix MicroOps counts for DPPS/YMM memory folded instructions · 5792e10f

Simon Pilgrim authored Mar 23, 2018

This was due to a misunderstanding over what llvm calls a micro-op (retirement unit) is actually called a macro-op on the AMD/Jaguar target. Folded loads don't affect num macro ops.

llvm-svn: 328320

5792e10f

[ELF] - Simplify. NFC. · 16f11462
George Rimar authored Mar 23, 2018
```
llvm-svn: 328319
```
16f11462

[X86][Btver2] Cleanup SSE42 PCMPISTR/PCMPESTR string instructions to correctly... · 8619962c

Simon Pilgrim authored Mar 23, 2018

[X86][Btver2] Cleanup SSE42 PCMPISTR/PCMPESTR string instructions to correctly use JFPU1 scheduler pipe followed by JLAGU/JSAGU/JFPA/JVALU function units

Fixes throughput to match Agner/Fam16h-SoG as well.

llvm-svn: 328318

8619962c

Remove the deprecated single-alignment IRBuilder API for memcpy/memmove (NFC) · a0c5f3ef

Daniel Neilson authored Mar 23, 2018

Summary:
This change is part of step six in the series of changes to remove the alignment
argument from memcpy/memmove/memset in favour of alignment attributes. At this
point all users of the IRBuilder APIs for creating a memcpy/memmove call given
a single value for alignment have been updated. We want to discourage usage of
these old APIs in favour of the newer ones that allow for separate source and
destination alignments, so this patch deletes the old API.

Specifically, we remove from IRBuilder:
CallInst *CreateMemCpy(Value *Dst, Value *Src, uint64_t Size, unsigned Align,
bool isVolatile = false, MDNode *TBAATag = nullptr,
MDNode *TBAAStructTag = nullptr,
MDNode *ScopeTag = nullptr,
MDNode *NoAliasTag = nullptr)
CallInst *CreateMemCpy(Value *Dst, Value *Src, Value *Size, unsigned Align,
bool isVolatile = false, MDNode *TBAATag = nullptr,
MDNode *TBAAStructTag = nullptr,
MDNode *ScopeTag = nullptr,
MDNode *NoAliasTag = nullptr)
CallInst *CreateMemMove(Value *Dst, Value *Src, uint64_t Size, unsigned Align,
bool isVolatile = false, MDNode *TBAATag = nullptr,
MDNode *ScopeTag = nullptr,
MDNode *NoAliasTag = nullptr)
CallInst *CreateMemMove(Value *Dst, Value *Src, Value *Size, unsigned Align,
bool isVolatile = false, MDNode *TBAATag = nullptr,
MDNode *ScopeTag = nullptr,
MDNode *NoAliasTag = nullptr)

Steps:
Step 1) Remove alignment parameter and create alignment parameter attributes for
memcpy/memmove/memset. ( rL322965, rC322964, rL322963 )
Step 2) Expand the IRBuilder API to allow creation of memcpy/memmove with differing
source and dest alignments. ( rL323597 )
Step 3) Update Clang to use the new IRBuilder API. ( rC323617 )
Step 4) Update Polly to use the new IRBuilder API. ( rL323618 )
Step 5) Update LLVM passes that create memcpy/memmove calls to use the new IRBuilder API,
and those that use use MemIntrinsicInst::[get|set]Alignment() to use [get|set]DestAlignment()
and [get|set]SourceAlignment() instead. ( rL323886, rL323891, rL324148, rL324273, rL324278,
rL324384, rL324395, rL324402, rL324626, rL324642, rL324653, rL324654, rL324773, rL324774,
rL324781, rL324784, rL324955, rL324960, rL325816, rL327398, rL327421, rL328097 )
Step 6) Remove the single-alignment IRBuilder API for memcpy/memmove, and the
MemIntrinsicInst::[get|set]Alignment() methods.

Reference
http://lists.llvm.org/pipermail/llvm-dev/2015-August/089384.html
http://lists.llvm.org/pipermail/llvm-commits/Week-of-Mon-20151109/312083.html

llvm-svn: 328317

a0c5f3ef

[SLP] Stop counting cost of gather sequences with multiple uses · 6c289a1c

Matthew Simpson authored Mar 23, 2018

When building the SLP tree, we look for reuse among the vectorized tree
entries. However, each gather sequence is represented by a unique tree entry,
even though the sequence may be identical to another one. This means, for
example, that a gather sequence with two uses will be counted twice when
computing the cost of the tree. We should only count the cost of the definition
of a gather sequence rather than its uses. During code generation, the
redundant gather sequences are emitted, but we optimize them away with CSE. So
it looks like this problem just affects the cost model.

Differential Revision: https://reviews.llvm.org/D44742

llvm-svn: 328316

6c289a1c

Remove deprecated MemIntrinsic methods (NFC) · a92bcbb2

Daniel Neilson authored Mar 23, 2018

Summary:
This change is part of step six in the series of changes to remove
the alignment argument from memcpy/memmove/memset in favour of
alignment attributes. At this point all uses of
MemIntrinsicInst::[get|set]Alignment() have been updated, so we now
remove these methods entirely to discourage their use.

Steps:
Step 1) Remove alignment parameter and create alignment parameter attributes for
memcpy/memmove/memset. ( rL322965, rC322964, rL322963 )
Step 2) Expand the IRBuilder API to allow creation of memcpy/memmove with differing
source and dest alignments. ( rL323597 )
Step 3) Update Clang to use the new IRBuilder API. ( rC323617 )
Step 4) Update Polly to use the new IRBuilder API. ( rL323618 )
Step 5) Update LLVM passes that create memcpy/memmove calls to use the new IRBuilder API,
and those that use use MemIntrinsicInst::[get|set]Alignment() to use [get|set]DestAlignment()
and [get|set]SourceAlignment() instead. ( rL323886, rL323891, rL324148, rL324273, rL324278,
rL324384, rL324395, rL324402, rL324626, rL324642, rL324653, rL324654, rL324773, rL324774,
rL324781, rL324784, rL324955, rL324960, rL325816, rL327398, rL327421, rL328097 )
Step 6) Remove the single-alignment IRBuilder API for memcpy/memmove, and the
MemIntrinsicInst::[get|set]Alignment() methods.

Reference
   http://lists.llvm.org/pipermail/llvm-dev/2015-August/089384.html
   http://lists.llvm.org/pipermail/llvm-commits/Week-of-Mon-20151109/312083.html

llvm-svn: 328315

a92bcbb2

[DEBUGINFO] Add flag for DWARF2 to use sections as references. · bff36086

Alexey Bataev authored Mar 23, 2018

Summary:
Some targets does not support labels inside debug sections, but support
references in form `section+offset`. Patch adds initial support
for this.

Reviewers: echristo, probinson, jlebar

Subscribers: llvm-commits, JDevlieghere

Differential Revision: https://reviews.llvm.org/D43943

llvm-svn: 328314

bff36086

[ARM] Support float literals under XO · 4a025cc7

Christof Douma authored Mar 23, 2018

When targeting execute-only and fp-armv8, float constants in a compare
resulted in instruction selection failures. This is now fixed by using
vmov.f32 where possible, otherwise the floating point constant is
lowered into a integer constant that is moved into a floating point
register.

This patch also restores using fpcmp with immediate 0 under fp-armv8.

Change-Id: Ie87229706f4ed879a0c0cf66631b6047ed6c6443
llvm-svn: 328313

4a025cc7

Revert r328307: [IPSCCP] Use constant range information for comparisons of parameters. · f73c3ece
Florian Hahn authored Mar 23, 2018
```
Reverted for now, due to it causing verifier failures.

llvm-svn: 328312
```
f73c3ece

[GlobalISel] Fix legalizer combine to not use illegal input G_EXTRACT. · f5423559

Amara Emerson authored Mar 23, 2018

This was being masked because GISel is enabled by default for -O0 and
the abort was disabled. Modified test to explicitly enable abort.

llvm-svn: 328311

f5423559

[test] Allow for optional No-Op Barrier Pass in O0 pipeline · 4316a262
Matthew Simpson authored Mar 23, 2018
```
llvm-svn: 328310
```
4316a262