Commits · ad34d91343bba205397955cc8e3e82f6ad99b2d8 · Lorenzo Albano / LLVM bpEVL

Jan 19, 2015

[PM] Relax asserts and always try to reconstruct loop simplify form when · ad34d913

Chandler Carruth authored Jan 19, 2015

we can while splitting critical edges.

The only code which called this and didn't require simplified loops to
be preserved is polly, and the code behaves correctly there anyways.
Without this change, it becomes really hard to share this code with the
new pass manager where things like preserving loop simplify form don't
make any sense.

If anyone discovers this code behaving incorrectly, what it *should* be
testing for is whether the loops it needs to be in simplified form are
in fact in that form. It should always be trying to preserve that form
when it exists.

llvm-svn: 226443

ad34d913

SLPVectorizer: limit the number of alias checks to reduce the runtime. · 76cb53a8

Erik Eckstein authored Jan 19, 2015

In case of blocks with many memory-accessing instructions, alias checking can take lot of time
(because calculating the memory dependencies has quadratic complexity).
I chose a limit which resulted in no changes when running the benchmarks.

llvm-svn: 226439

76cb53a8

[PowerPC] Minor correction to r226432 · c3168129

Hal Finkel authored Jan 19, 2015

We don't need to exclude patchpoints from the implicit r2 dependence in
FastISel because it is added as an implicit operand and, thus, should not
confuse that StackMap code.

By inspection / no test case.

llvm-svn: 226434

c3168129

[MIScheduler] Slightly better handling of constrainLocalCopy when both source and dest are local · 54c61ede
Michael Kuperstein authored Jan 19, 2015
```
This fixes PR21792.

Differential Revision: http://reviews.llvm.org/D6823

llvm-svn: 226433
```
54c61ede

[PowerPC] Add r2 as an operand for all calls under both PPC64 ELF V1 and V2 · af51993e

Hal Finkel authored Jan 19, 2015

Our PPC64 ELF V2 call lowering logic added r2 as an operand to all direct call
instructions in order to represent the dependency on the TOC base pointer
value. Restricting this to ELF V2, however, does not seem to make sense: calls
under ELF V1 have the same dependence, and indirect calls have an r2 dependence
just as direct ones. Make sure the dependence is noted for all calls under both
ELF V1 and ELF V2.

llvm-svn: 226432

af51993e

[x86] Change AVX512 intrinsics to take a 8-bit immediate for the comparision... · f4bf9119

Craig Topper authored Jan 19, 2015

[x86] Change AVX512 intrinsics to take a 8-bit immediate for the comparision kind instead of a 32-bit immediate. This better aligns with the emitted instruction. It also matches SSE and AVX1 equivalents. Also add auto upgrade support.

llvm-svn: 226430

f4bf9119

[tinyptrvector] Add in a MutableArrayRef implicit conversion operator to... · b93d3dbc

Michael Gottesman authored Jan 19, 2015

[tinyptrvector] Add in a MutableArrayRef implicit conversion operator to complement the ArrayRef implicit conversion operator.

llvm-svn: 226428

b93d3dbc

[PM] Lift the analyses into the interface for · 0eae1120

Chandler Carruth authored Jan 19, 2015

SplitLandingPadPredecessors and remove the Pass argument from its
interface.

Another step to the utilities being usable with both old and new pass
managers.

llvm-svn: 226426

0eae1120

Change using => typedef to please the MSVC bots. · 26500a56
Michael Gottesman authored Jan 19, 2015
```
llvm-svn: 226425
```
26500a56

Hide the state of TinyPtrVector and remove the single element constructor. · 4125886b

Michael Gottesman authored Jan 19, 2015

There is no reason for this state to be exposed as public. The single element
constructor was superfulous in light of the single element ArrayRef
constructor.

llvm-svn: 226424

4125886b

Reorder. · ded8207c
NAKAMURA Takumi authored Jan 19, 2015
```
llvm-svn: 226419
```
ded8207c
[CMake] examples/Kaleidoscope: Prune redundant libdeps. · 86c8a9cc
NAKAMURA Takumi authored Jan 19, 2015
```
llvm-svn: 226418
```
86c8a9cc
[CMake] Update libdeps in examples/Kaleidoscope/Chapter4. · bd1d5b1a
NAKAMURA Takumi authored Jan 19, 2015
```
llvm-svn: 226417
```
bd1d5b1a

Jan 18, 2015

unique_ptrify the RelInfo parameter to TargetRegistry::createMCSymbolizer · 186db431
David Blaikie authored Jan 18, 2015
```
llvm-svn: 226416
```
186db431

Attempt to fix the MSVC build by working around a layering issue · b619b786

David Blaikie authored Jan 18, 2015

Since MCStreamer isn't part of Support, the dtor can't be called from
here - so just pass by reference instead. This is rather imperfect, but
will hopefully suffice.

llvm-svn: 226415

b619b786

std::unique_ptrify the MCStreamer argument to createAsmPrinter · 9459832e
David Blaikie authored Jan 18, 2015
```
llvm-svn: 226414
```
9459832e
R600: Remove redundant test · 4843f193
Matt Arsenault authored Jan 18, 2015
```
This is already covered in ftrunc.ll

llvm-svn: 226412
```
4843f193
[mips] 'CHECK :' is not a valid check directive. Fixed. · 01dce6c9
Daniel Sanders authored Jan 18, 2015
```
llvm-svn: 226409
```
01dce6c9

[mips] Make whitespace in disassembler tests more consistent. NFC. · 0cb9dc6e

Daniel Sanders authored Jan 18, 2015

The tests for the ISA's should now be approximately diffable. That is, the
output of 'diff valid-mips1.txt valid-mips2.txt' should be emit the lines
for instructions that were added/removed to/from MIPS-I by MIPS-II. This
doesn't work perfectly at the moment due to ordering differences but it
should be close.

llvm-svn: 226408

0cb9dc6e

[mips] Make whitespace of disassembler tests more consistent by removing blank lines. NFC. · 46ad7cbf
Daniel Sanders authored Jan 18, 2015
```
llvm-svn: 226407
```
46ad7cbf
[X86][SSE] Added scalar min/max folding tests. NFC. · 4cf275ea
Simon Pilgrim authored Jan 18, 2015
```
llvm-svn: 226406
```
4cf275ea
[X86][SSE] Added float extract and xmm extract/insert stack folding tests. NFC. · 1d6dcdca
Simon Pilgrim authored Jan 18, 2015
```
llvm-svn: 226405
```
1d6dcdca
[X86][SSE] Added scalar conversion stack folding tests. NFC. · cd26d0b6
Simon Pilgrim authored Jan 18, 2015
```
llvm-svn: 226404
```
cd26d0b6

[PowerPC] Don't hard-code R2 as register when processing TOC relocations · 58884f9f

Hal Finkel authored Jan 18, 2015

Instructions that have high-order TOC relocations always carry R2 as their base
register, so it does not matter whether we take the register from the
instruction or just hard-code it in PPCAsmPrinter. In the future, however, we
might want to apply these relocations to instructions using a different
register, so taking the register from the instruction is a better thing to do.
No change in functionality here, however.

llvm-svn: 226403

58884f9f

[PowerPC] Add some FIXMEs for fastcc and FPR <-> GPR moves · 8ea446b6

Hal Finkel authored Jan 18, 2015

So we don't forget, once we support FPR <-> GPR moves on the P8, we'll likely
want to re-visit this part of the calling convention.

llvm-svn: 226401

8ea446b6

AVX1 stack folding tests. NFC. · fe3bfb80

Simon Pilgrim authored Jan 18, 2015

Begun adding more exhaustive tests - all floating point instructions should now be either tested or have placeholders. We do seem to have a number of missing instructions, I will add a patch for review once the remaining working instructions are added.

I'll then move on to SSE tests and then the integer instructions.

llvm-svn: 226400

fe3bfb80

[PowerPC] Initial PPC64 calling-convention changes for fastcc · f81b6dd7

Hal Finkel authored Jan 18, 2015

The default calling convention specified by the PPC64 ELF (V1 and V2) ABI is
designed to work with both prototyped and non-prototyped/varargs functions. As
a result, GPRs and stack space are allocated for every argument, even those
that are passed in floating-point or vector registers.

GlobalOpt::OptimizeFunctions will transform local non-varargs functions (that
do not have their address taken) to use the 'fast' calling convention.

When functions are using the 'fast' calling convention, don't allocate GPRs for
arguments passed in other types of registers, and don't allocate stack space for
arguments passed in registers. Other changes for the fast calling convention
may be added in the future.

llvm-svn: 226399

f81b6dd7

[PM] Pull the analyses used for another utility routine into its API · b5797b65

Chandler Carruth authored Jan 18, 2015

rather than relying on the pass object.

This one is a bit annoying, but will pay off. First, supporting this one
will make the next one much easier, and for utilities like LoopSimplify,
this is moving them (slowly) closer to not having to pass the pass
object around throughout their APIs.

llvm-svn: 226396

b5797b65

[PM] Sink the specific analyses preserved by SplitBlock into its · 32c52c7e

Chandler Carruth authored Jan 18, 2015

interface, removing Pass from its interface.

This also makes those analyses optional so that passes which don't even
preserve these (or use them) can skip the logic entirely.

llvm-svn: 226394

32c52c7e

[PM] Replace another Pass argument with specific analyses that are · b5c11535

Chandler Carruth authored Jan 18, 2015

optionally updated by MergeBlockIntoPredecessors.

No functionality changed, just refactoring to clear the way for the new
pass manager.

llvm-svn: 226392

b5c11535

[PM] Refactor how the LoopRotation pass access the DominatorTree. · 94209094

Chandler Carruth authored Jan 18, 2015

Instead of querying the pass every where we need to, do that once and
cache a pointer in the pass object. This is both simpler and I'm about
to add yet another place where we need to dig out that pointer.

llvm-svn: 226391

94209094

[PM] Lift the actual analyses used into the inferface rather than · 5eee895c

Chandler Carruth authored Jan 18, 2015

accepting a Pass and querying it for analyses.

This is necessary to allow the utilities to work both with the old and
new pass managers, and I also think this makes the interface much more
clear and helps the reader know what analyses the utility can actually
handle. I plan to repeat this process iteratively to clean up all the
pass utilities.

llvm-svn: 226386

5eee895c

[PM] Now that LoopInfo isn't in the Pass type hierarchy, it is much · 691addc2

Chandler Carruth authored Jan 18, 2015

cleaner to derive from the generic base.

Thise removes a ton of boiler plate code and somewhat strange and
pointless indirections. It also remove a bunch of the previously needed
friend declarations. To fully remove these, I also lifted the verify
logic into the generic LoopInfoBase, which seems good anyways -- it is
generic and useful logic even for the machine side.

llvm-svn: 226385

691addc2

Jan 17, 2015

[PM] Cleanup more warnings my refactoring exposed where now we have · bc045a5a

Chandler Carruth authored Jan 17, 2015

unused variables in a no-asserts build.

I've fixed this by putting the entire loop behind an #ifndef as it
contains nothing other than asserts.

llvm-svn: 226377

bc045a5a

[PM] Remove a dead field. · 24fd029a

Chandler Carruth authored Jan 17, 2015

This was dead even before I refactored how we initialized it, but my
refactoring made it trivially dead and it is now caught by a Clang
warning. This fixes the warning and should clean up the -Werror bot
failures (sorry!).

llvm-svn: 226376

24fd029a

[PM] Split the LoopInfo object apart from the legacy pass, creating · 4f8f307c

Chandler Carruth authored Jan 17, 2015

a LoopInfoWrapperPass to wire the object up to the legacy pass manager.

This switches all the clients of LoopInfo over and paves the way to port
LoopInfo to the new pass manager. No functionality change is intended
with this iteration.

llvm-svn: 226373

4f8f307c

[PowerPC] Don't list R11 as a patchpoint scratch register · c19805a7

Hal Finkel authored Jan 17, 2015

R11's status is the same under both the PPC64 ELF V1 and V2 ABIs: it is
reserved for use as an "environment pointer" for compilation models that
require such a thing. We don't, we also don't need a second scratch register,
and because we support only "local" patchpoint call targets, we might as well
let R11 be used for anyregcc patchpoints.

llvm-svn: 226369

c19805a7

ProgrammersManual.rst: fix a typo · 8888d5b3
Hans Wennborg authored Jan 17, 2015
```
llvm-svn: 226367
```
8888d5b3

Improve DAG combine pass on certain IR vector patterns · 37f316af

Mehdi Amini authored Jan 17, 2015

Loading 2 2x32-bit float vectors into the bottom half of a 256-bit vector
produced suboptimal code in AVX2 mode with certain IR combinations.

In particular, the IR optimizer folded 2f32 + 2f32 -> 4f32, 4f32 + 4f32
(undef) -> 8f32 into a 2f32 + 2f32 -> 8f32, which seems more canonical,
but then mysteriously generated rather bad code; the movq/movhpd combination
didn't match.

The problem lay in the BUILD_VECTOR optimization path. The 2f32 inputs
would get promoted to 4f32 by the type legalizer, eventually resulting
in a BUILD_VECTOR on two 4f32 into an 8f32. The BUILD_VECTOR then, recognizing
these were both half the output size, concatted them and then produced
a shuffle. However, the resulting concat + shuffle was more complex than
it should be; in the case where the upper half of the output is undef, we
probably want to generate shuffle + concat instead.

This enhancement causes the vector_shuffle combine step to recognize this
suboptimal pattern and correct it. I included it there instead of in BUILD_VECTOR
in case the same suboptimal pattern occurs for other reasons.

This results in the optimizer correctly producing the optimal movq + movhpd
sequence for all three variations on this IR, even with AVX2.

I've included a test case.

Radar link: rdar://problem/19287012
Fix for PR 21943.

From: Fiona Glaser <fglaser@apple.com>
llvm-svn: 226360

37f316af

[RuntimeDyld] Tidy up emitCommonSymbols a little. NFC. · 2996895f
Lang Hames authored Jan 17, 2015
```
llvm-svn: 226358
```
2996895f