Commits · 71928e681b1f192ac2e40d7414f29b1bfea6de5d · Roger Ferrer / llvm-epi-0.8

Apr 17, 2012

Add disassembler to MIPS. · 71928e68
Akira Hatanaka authored Apr 17, 2012
```
Patch by Vladimir Medic. 

llvm-svn: 154935
```
71928e68
Goodbye, JSONParser... · 1f8918f6
Manuel Klimek authored Apr 17, 2012
```
llvm-svn: 154930
```
1f8918f6
Remove unused CCIfSubtarget. · 08a0598c
Jay Foad authored Apr 17, 2012
```
llvm-svn: 154921
```
08a0598c
Fix bad EXTRACT_SUBREG in instruction selection for extending-loads on NEON. · a9bcf20d
James Molloy authored Apr 17, 2012
```
llvm-svn: 154915
```
a9bcf20d
Revert "SCEV: When expanding a GEP the final addition to the base pointer has NUW but not NSW." · e364d195
Benjamin Kramer authored Apr 17, 2012
```
This isn't right either, reverting for now.

llvm-svn: 154910
```
e364d195
Don't decode vperm2i128 or vperm2f128 into a shuffle if bit 3 or 7 of the immediate is set. · 354103d8
Craig Topper authored Apr 17, 2012
```
llvm-svn: 154907
```
354103d8

SlotIndexes used to store the index list in a crufty custom linked-list. I can't · aef91783

Lang Hames authored Apr 17, 2012

for the life of me remember why I wrote it this way, but I can't see any good
reason for it now. This patch replaces the custom linked list with an ilist.

This change should preserve the existing numberings exactly, so no generated code
should change (if it does, file a bug!).

llvm-svn: 154904

aef91783

Fix ARM disassembly of VLD2 (single 2-element structure to all lanes) · 29ae5386

Kevin Enderby authored Apr 17, 2012

instructions with writebacks. And add test a case for all opcodes handed by
DecodeVLD2DupInstruction() in ARMDisassembler.cpp .

llvm-svn: 154884

29ae5386

Typo. · 7df0240e
Eric Christopher authored Apr 16, 2012
```
llvm-svn: 154879
```
7df0240e
Make comment here more clear. · a8caa739
Eric Christopher authored Apr 16, 2012
```
llvm-svn: 154878
```
a8caa739
ARM two-operand forms for vhadd and vhsub instructions. · 2bf5f739
Jim Grosbach authored Apr 16, 2012
```
rdar://11252521

llvm-svn: 154875
```
2bf5f739

Temporarily turn off anti-dependency checking · 5333e2e5

Preston Gurd authored Apr 16, 2012

during Post RA scheduling in X86,
until the X86 target is changed to properly set up
post RA liveness.

llvm-svn: 154874

5333e2e5

Add files which were not included by commit 154868. · e185ecb2
Preston Gurd authored Apr 16, 2012
```
llvm-svn: 154872
```
e185ecb2

Implement GDB integration for source level debugging of code JITed using · cc31af93

Preston Gurd authored Apr 16, 2012

the MCJIT execution engine.

The GDB JIT debugging integration support works by registering a loaded
object image with a pre-defined function that GDB will monitor if GDB
is attached. GDB integration support is implemented for ELF only at this
time. This integration requires GDB version 7.0 or newer.

Patch by Andy Kaylor!

 

llvm-svn: 154868

cc31af93

Fix updateTerminator to be resiliant to degenerate terminators where · 1f5580b6

Chandler Carruth authored Apr 16, 2012

both fallthrough and a conditional branch target the same successor.
Gracefully delete the conditional branch and introduce any unconditional
branch needed to reach the actual successor. This fixes memory
corruption in 2009-06-15-RegScavengerAssert.ll and possibly other tests.

Also, while I'm here fix a latent bug I spotted by inspection. I never
applied the same fundamental fix to this fallthrough successor finding
logic that I did to the logic used when there are no conditional
branches. As a consequence it would have selected landing pads had they
be aligned in just the right way here. I don't have a test case as
I spotted this by inspection, and the previous time I found this
required have of TableGen's source code to produce it. =/ I hate backend
bugs. ;]

Thanks to Jim Grosbach for helping me reason through this and reviewing
the fix.

llvm-svn: 154867

1f5580b6

Apr 16, 2012

MC assembly parser handling for trailing comma in macro instantiation. · 1e1d68f1

Jim Grosbach authored Apr 16, 2012

A trailing comma means no argument at all (i.e., as if the comma were not
present), not an empty argument to the invokee.

rdar://11252521

llvm-svn: 154863

1e1d68f1

ARM handle :lower16: and :upper16: after a '#' prefix. · 003607f4
Jim Grosbach authored Apr 16, 2012
```
rdar://11252521

llvm-svn: 154862
```
003607f4
Remove support for the special 'fast' value for fpmath accuracy for the moment. · 9af62982
Duncan Sands authored Apr 16, 2012
```
llvm-svn: 154850
```
9af62982
Fix incorrect atomics codegen introduced in r154705, and extend test to catch it. · 12da79b8
Richard Smith authored Apr 16, 2012
```
llvm-svn: 154845
```
12da79b8
Remove unused variable · e67cdc07
David Blaikie authored Apr 16, 2012
```
llvm-svn: 154841
```
e67cdc07
ARM assembly two-operand forms for VRSHL. · 6068d001
Jim Grosbach authored Apr 16, 2012
```
rdar://11252521

llvm-svn: 154840
```
6068d001
Do not add offset in applyFixup. This has already been accounted for in Value. · 3e9d81f4
Akira Hatanaka authored Apr 16, 2012
```
llvm-svn: 154838
```
3e9d81f4
ARM two-operand aliases for VRHADD instructions. · cd1c000a
Jim Grosbach authored Apr 16, 2012
```
rdar://11252521

llvm-svn: 154832
```
cd1c000a
Hexagon V5 (Floating Point) Support. · 96e8ee17
Sirish Pande authored Apr 16, 2012
```
llvm-svn: 154829
```
96e8ee17

Make it possible to indicate relaxed floating point requirements at the IR level · 05f4df8d

Duncan Sands authored Apr 16, 2012

through the use of 'fpmath' metadata. Currently this only provides a 'fpaccuracy'
value, which may be a number in ULPs or the keyword 'fast', however the intent is
that this will be extended with additional information about NaN's, infinities
etc later. No optimizations have been hooked up to this so far.

llvm-svn: 154822

05f4df8d

Flip the new block-placement pass to be on by default. · 4190b507

Chandler Carruth authored Apr 16, 2012

This is mostly to test the waters. I'd like to get results from FNT
build bots and other bots running on non-x86 platforms.

This feature has been pretty heavily tested over the last few months by
me, and it fixes several of the execution time regressions caused by the
inlining work by preventing inlining decisions from radically impacting
block layout.

I've seen very large improvements in yacr2 and ackermann benchmarks,
along with the expected noise across all of the benchmark suite whenever
code layout changes. I've analyzed all of the regressions and fixed
them, or found them to be impossible to fix. See my email to llvmdev for
more details.

I'd like for this to be in 3.1 as it complements the inliner changes,
but if any failures are showing up or anyone has concerns, it is just
a flag flip and so can be easily turned off.

I'm switching it on tonight to try and get at least one run through
various folks' performance suites in case SPEC or something else has
serious issues with it. I'll watch bots and revert if anything shows up.

llvm-svn: 154816

4190b507

Add a somewhat hacky heuristic to do something different from whole-loop · 8c0b41d6

Chandler Carruth authored Apr 16, 2012

rotation. When there is a loop backedge which is an unconditional
branch, we will end up with a branch somewhere no matter what. Try
placing this backedge in a fallthrough position above the loop header as
that will definitely remove at least one branch from the loop iteration,
where whole loop rotation may not.

I haven't seen any benchmarks where this is important but loop-blocks.ll
tests for it, and so this will be covered when I flip the default.

llvm-svn: 154812

8c0b41d6

Fix style violation in BBVectorize (pointed out by Bill Wendling) · 52ba49f3
Hal Finkel authored Apr 16, 2012
```
llvm-svn: 154810
```
52ba49f3

Tweak the loop rotation logic to check whether the loop is naturally · 8c74c7b1

Chandler Carruth authored Apr 16, 2012

laid out in a form with a fallthrough into the header and a fallthrough
out of the bottom. In that case, leave the loop alone because any
rotation will introduce unnecessary branches. If either side looks like
it will require an explicit branch, then the rotation won't add any, do
it to ensure the branch occurs outside of the loop (if possible) and
maximize the benefit of the fallthrough in the bottom.

llvm-svn: 154806

8c74c7b1

Reapply 'Add reverseColor to raw_ostream'. · 13d16f3b

Benjamin Kramer authored Apr 16, 2012

To be used in printing unprintable source in clang diagnostics.
Patch by Seth Cantrell, with a minor fix for mingw by me.

llvm-svn: 154805

13d16f3b

Revert r154800 which breaks windows builders. · 64104f16
Argyrios Kyrtzidis authored Apr 16, 2012
```
llvm-svn: 154802
```
64104f16
Replace vpermd/vpermps intrinic patterns with custom lowering to target specific nodes. · 4badeb3f
Craig Topper authored Apr 16, 2012
```
llvm-svn: 154801
```
4badeb3f

Add reverseColor to raw_ostream. · d17db2e0

Argyrios Kyrtzidis authored Apr 16, 2012

To be used in printing unprintable source in clang diagnostics.
Patch by Seth Cantrell!

llvm-svn: 154800

d17db2e0

Change type profile for vpermv back to using operand type for the mask... · 26d7a949

Craig Topper authored Apr 16, 2012

Change type profile for vpermv back to using operand type for the mask argument to match intrinsic behavior. Add a bitcast to the lowering code to convert mask from v8i32 to v8f32 for vpermps.

llvm-svn: 154798

26d7a949

Flip the arguments when converting vpermd/vpermps intrinsics into... · c0075aa7

Craig Topper authored Apr 16, 2012

Flip the arguments when converting vpermd/vpermps intrinsics into instructions. The intrinsic has the mask as the last operand, but the instruction has it as the second.

llvm-svn: 154797

c0075aa7

Add a Fixme. · 82b90a38
Bill Wendling authored Apr 16, 2012
```
llvm-svn: 154793
```
82b90a38
Simplify checking for pointer types in BBVectorize (this change was suggested by Duncan). · 8ee309d9
Hal Finkel authored Apr 16, 2012
```
llvm-svn: 154787
```
8ee309d9
Remove dead SD nodes after the combining pass. Fixes PR12201. · e0cf6397
Hal Finkel authored Apr 16, 2012
```
llvm-svn: 154786
```
e0cf6397

Rewrite how machine block placement handles loop rotation. · ccc7e42b

Chandler Carruth authored Apr 16, 2012

This is a complex change that resulted from a great deal of
experimentation with several different benchmarks. The one which proved
the most useful is included as a test case, but I don't know that it
captures all of the relevant changes, as I didn't have specific
regression tests for each, they were more the result of reasoning about
what the old algorithm would possibly do wrong. I'm also failing at the
moment to craft more targeted regression tests for these changes, if
anyone has ideas, it would be welcome.

The first big thing broken with the old algorithm is the idea that we
can take a basic block which has a loop-exiting successor and a looping
successor and use the looping successor as the layout top in order to
get that particular block to be the bottom of the loop after layout.
This happens to work in many cases, but not in all.

The second big thing broken was that we didn't try to select the exit
which fell into the nearest enclosing loop (to which we exit at all). As
a consequence, even if the rotation worked perfectly, it would result in
one of two bad layouts. Either the bottom of the loop would get
fallthrough, skipping across a nearer enclosing loop and thereby making
it discontiguous, or it would be forced to take an explicit jump over
the nearest enclosing loop to earch its successor. The point of the
rotation is to get fallthrough, so we need it to fallthrough to the
nearest loop it can.

The fix to the first issue is to actually layout the loop from the loop
header, and then rotate the loop such that the correct exiting edge can
be a fallthrough edge. This is actually much easier than I anticipated
because we can handle all the hard parts of finding a viable rotation
before we do the layout. We just store that, and then rotate after
layout is finished. No inner loops get split across the post-rotation
backedge because we check for them when selecting the rotation.

That fix exposed a latent problem with our exitting block selection --
we should allow the backedge to point into the middle of some inner-loop
chain as there is no real penalty to it, the whole point is that it
*won't* be a fallthrough edge. This may have blocked the rotation at all
in some cases, I have no idea and no test case as I've never seen it in
practice, it was just noticed by inspection.

Finally, all of these fixes, and studying the loops they produce,
highlighted another problem: in rotating loops like this, we sometimes
fail to align the destination of these backwards jumping edges. Fix this
by actually walking the backwards edges rather than relying on loopinfo.

This fixes regressions on heapsort if block placement is enabled as well
as lots of other cases where the previous logic would introduce an
abundance of unnecessary branches into the execution.

llvm-svn: 154783

ccc7e42b

Merge vpermps/vpermd and vpermpd/vpermq SD nodes. · b86fa404
Craig Topper authored Apr 16, 2012
```
llvm-svn: 154782
```
b86fa404