Commits · ed1fb92cfe3cd3b670b222cdefa9e1b90c0b8e27 · Roger Ferrer / llvm-epi-0.8

Jan 16, 2011

simplify a little · ed1fb92c
Chris Lattner authored Jan 16, 2011
```
llvm-svn: 123573
```
ed1fb92c

if an alloca is only ever accessed as a unit, and is accessed with load/store instructions, · 6fab2e94

Chris Lattner authored Jan 16, 2011

then don't try to decimate it into its individual pieces.  This will just make a mess of the
IR and is pointless if none of the elements are individually accessed.  This was generating
really terrible code for std::bitset (PR8980) because it happens to be lowered by clang
as an {[8 x i8]} structure instead of {i64}.

The testcase now is optimized to:

define i64 @test2(i64 %X) {
  br label %L2

L2:                                               ; preds = %0
  ret i64 %X
}

before we generated:

define i64 @test2(i64 %X) {
  %sroa.store.elt = lshr i64 %X, 56
  %1 = trunc i64 %sroa.store.elt to i8
  %sroa.store.elt8 = lshr i64 %X, 48
  %2 = trunc i64 %sroa.store.elt8 to i8
  %sroa.store.elt9 = lshr i64 %X, 40
  %3 = trunc i64 %sroa.store.elt9 to i8
  %sroa.store.elt10 = lshr i64 %X, 32
  %4 = trunc i64 %sroa.store.elt10 to i8
  %sroa.store.elt11 = lshr i64 %X, 24
  %5 = trunc i64 %sroa.store.elt11 to i8
  %sroa.store.elt12 = lshr i64 %X, 16
  %6 = trunc i64 %sroa.store.elt12 to i8
  %sroa.store.elt13 = lshr i64 %X, 8
  %7 = trunc i64 %sroa.store.elt13 to i8
  %8 = trunc i64 %X to i8
  br label %L2

L2:                                               ; preds = %0
  %9 = zext i8 %1 to i64
  %10 = shl i64 %9, 56
  %11 = zext i8 %2 to i64
  %12 = shl i64 %11, 48
  %13 = or i64 %12, %10
  %14 = zext i8 %3 to i64
  %15 = shl i64 %14, 40
  %16 = or i64 %15, %13
  %17 = zext i8 %4 to i64
  %18 = shl i64 %17, 32
  %19 = or i64 %18, %16
  %20 = zext i8 %5 to i64
  %21 = shl i64 %20, 24
  %22 = or i64 %21, %19
  %23 = zext i8 %6 to i64
  %24 = shl i64 %23, 16
  %25 = or i64 %24, %22
  %26 = zext i8 %7 to i64
  %27 = shl i64 %26, 8
  %28 = or i64 %27, %25
  %29 = zext i8 %8 to i64
  %30 = or i64 %29, %28
  ret i64 %30
}

In this case, instcombine was able to eliminate the nonsense, but in PR8980 enough
PHIs are in play that instcombine backs off.  It's better to not generate this stuff
in the first place.

llvm-svn: 123571

6fab2e94

Use an irbuilder to get some trivial constant folding when doing a store · 7cd8cf7d
Chris Lattner authored Jan 16, 2011
```
of a constant.

llvm-svn: 123570
```
7cd8cf7d

enhance FoldOpIntoPhi in instcombine to try harder when a phi has · d55581de

Chris Lattner authored Jan 16, 2011

multiple uses.  In some cases, all the uses are the same operation,
so instcombine can go ahead and promote the phi.  In the testcase
this pushes an add out of the loop.

llvm-svn: 123568

d55581de

Jan 15, 2011

temporarily revert r123526. While working on a follow-on patch I · af263907
Chris Lattner authored Jan 15, 2011
```
realize that ConstantFoldTerminator doesn't preserve dominfo.

llvm-svn: 123527
```
af263907

fix rdar://8785296 - -fcatch-undefined-behavior generates inefficient code · 8df83c4a

Chris Lattner authored Jan 15, 2011

The basic issue is that isel (very reasonably!) expects conditional branches
to be folded, so CGP leaving around a bunch dead computation feeding
conditional branches isn't such a good idea.  Just fold branches on constants
into unconditional branches.

llvm-svn: 123526

8df83c4a

simplify code, no functionality change. · ee588def
Chris Lattner authored Jan 15, 2011
```
llvm-svn: 123525
```
ee588def

Now that instruction optzns can update the iterator as they go, we can · 1b93be50

Chris Lattner authored Jan 15, 2011

have objectsize folding recursively simplify away their result when it
folds.  It is important to catch this here, because otherwise we won't
eliminate the cross-block values at isel and other times.

llvm-svn: 123524

1b93be50

make the current instruction iterator an ivar, allowing xforms that · 7a277144

Chris Lattner authored Jan 15, 2011

potentially invalidate it (like inline asm lowering) to be sunk into
their proper place, cleaning up a ton of code.

llvm-svn: 123523

7a277144

Generalize LoadAndStorePromoter a bit and switch LICM · b68ec5c3
Chris Lattner authored Jan 15, 2011
```
to use it.

llvm-svn: 123501
```
b68ec5c3

Jan 14, 2011

switch SRoA to use LoadAndStorePromoter instead of its own copy of the code. · b498f9af
Chris Lattner authored Jan 14, 2011
```
llvm-svn: 123457
```
b498f9af
split SROA into two passes: one that uses DomFrontiers (-scalarrepl) · 9987a6f4
Chris Lattner authored Jan 14, 2011
```
and one that uses SSAUpdater (-scalarrepl-ssa)

llvm-svn: 123436
```
9987a6f4

Implement full support for promoting allocas to registers using SSAUpdater · 543384ef

Chris Lattner authored Jan 14, 2011

instead of DomTree/DomFrontier.  This may be interesting for reducing compile 
time.  This is currently disabled, but seems to work just fine.

When this is enabled, we eliminate two runs of dominator frontier, one in the
"early per-function" optimizations and one in the "interlaced with inliner"
function passes.

llvm-svn: 123434

543384ef

Jan 13, 2011

Fix whitespace. · 328e91bb
Bob Wilson authored Jan 13, 2011
```
llvm-svn: 123396
```
328e91bb
Check for empty structs, and for consistency, zero-element arrays. · c8056a95
Bob Wilson authored Jan 13, 2011
```
llvm-svn: 123383
```
c8056a95

Extend SROA to handle arrays accessed as homogeneous structs and vice versa. · 08713d3c

Bob Wilson authored Jan 13, 2011

This is a minor extension of SROA to handle a special case that is
important for some ARM NEON operations. Some of the NEON intrinsics
return multiple values, which are handled as struct types containing
multiple elements of the same vector type. The corresponding return
types declared in the arm_neon.h header have equivalent arrays. We
need SROA to recognize that it can split up those arrays and structs
into separate vectors, even though they are not always accessed with
the same type. SROA already handles loads and stores of an entire
alloca by using insertvalue/extractvalue to access the individual
pieces, and that code works the same regardless of whether the type
is a struct or an array. So, all that needs to be done is to check
for compatible arrays and homogeneous structs.

llvm-svn: 123381

08713d3c

Make SROA more aggressive with allocas containing padding. · 12eec40c

Bob Wilson authored Jan 13, 2011

SROA only split up structs and arrays one level at a time, so padding can
only cause trouble if it is located in between the struct or array elements.

llvm-svn: 123380

12eec40c

Jan 12, 2011
- Use SmallVector instead of SmallPtrSet and avoid non-deterministic behavior. · 30f3ebbc
  Devang Patel authored Jan 12, 2011
```
llvm-svn: 123318
```
  30f3ebbc
- revert 123144, reenabling the rest of memset formation. · dd5f60b7
  Chris Lattner authored Jan 12, 2011
```
llvm-svn: 123302
```
  dd5f60b7
- revert r123146 which disabled code that wasn't the root cause · 654098f4
  Chris Lattner authored Jan 12, 2011
```
of the bootstrap miscompare issue.

llvm-svn: 123299
```
  654098f4
- revert r123149, reenabling an improvement to memcpyopt that wasn't · fa7c29d2
  Chris Lattner authored Jan 12, 2011
```
the source of the bootstrap problem.

llvm-svn: 123298
```
  fa7c29d2
Jan 11, 2011
- Remove the PR8954 workaround. · 12cc296b
  Jakob Stoklund Olesen authored Jan 11, 2011
```
llvm-svn: 123288
```
  12cc296b
- Dial back the speculative fix for PR8954 a bit, so that we only recompute dominators · cb9c4f85
  Cameron Zwarich authored Jan 11, 2011
```
once at the beginning of GVN instead of once per iteration.

llvm-svn: 123278
```
  cb9c4f85
- Attempt to fix the bootstrap buildbot. Rafael says this works for him on x86-64 Linux. · 51eb4039
  Cameron Zwarich authored Jan 11, 2011
```
llvm-svn: 123270
```
  51eb4039
- update memdep when an instruction is deleted. This code isn't · 193ce7c4
  Chris Lattner authored Jan 11, 2011
```
actually reached in the testcase in PR8954, but it's safe and good
practice.

llvm-svn: 123224
```
  193ce7c4
- Fix FoldSingleEntryPHINodes to update memdep and AA when it deletes · f6ae904e
  Chris Lattner authored Jan 11, 2011
```
phi nodes.  It is called from MergeBlockIntoPredecessor which is 
called from GVN, which claims to preserve these.

I'm skeptical that this is the actual problem behind PR8954, but
this is a stab in the right direction.

llvm-svn: 123222
```
  f6ae904e
- random cleanups · dfcfcb49
  Chris Lattner authored Jan 11, 2011
```
llvm-svn: 123221
```
  dfcfcb49
- remove a bogus assertion: the latch block of a loop is not · 63fe78de
  Chris Lattner authored Jan 11, 2011
```
neccesarily an uncond branch to the header.  This fixes 
PR8955 (the assertion tripping).

llvm-svn: 123219
```
  63fe78de
Jan 10, 2011
- another random stab in the dark trying to fix llvm-gcc-i386-linux-selfhost · 88bc848a
  Chris Lattner authored Jan 10, 2011
```
llvm-svn: 123149
```
  88bc848a
- another (more) aggressive attempt to bring llvm-gcc-i386-linux-selfhost · 4662bd4b
  Chris Lattner authored Jan 10, 2011
```
back to life.

llvm-svn: 123146
```
  4662bd4b
- temporarily disable memset formation from memsets in an effort to restore buildbot stability. · 1017fa67
  Chris Lattner authored Jan 09, 2011
```
llvm-svn: 123144
```
  1017fa67
Jan 09, 2011
- fix a few old bugs (found by inspection) where we would zap instructions · caf5c0d0
  Chris Lattner authored Jan 09, 2011
```
without informing memdep.  This could cause nondeterminstic weirdness 
based on where instructions happen to get allocated, and will hopefully
breath some life into some broken testers.

llvm-svn: 123124
```
  caf5c0d0
- LoopInstSimplify preserves LoopSimplify. · a42e5915
  Cameron Zwarich authored Jan 09, 2011
```
llvm-svn: 123117
```
  a42e5915
- reduce indentation. Print <nuw> and <nsw> when dumping SCEV AddRec's · a337f5ec
  Chris Lattner authored Jan 09, 2011
```
that have the bit set.

llvm-svn: 123104
```
  a337f5ec
Jan 08, 2011

fix a latent bug in memcpyoptimizer that my recent patches exposed: it wasn't · 7d6433ae
Chris Lattner authored Jan 08, 2011
```
updating memdep when fusing stores together.  This fixes the crash optimizing
the bullet benchmark.

llvm-svn: 123091
```
7d6433ae
tryMergingIntoMemset can only handle constant length memsets. · ff6ed2ac
Chris Lattner authored Jan 08, 2011
```
llvm-svn: 123090
```
ff6ed2ac

Merge memsets followed by neighboring memsets and other stores into · 9a1d63ba

Chris Lattner authored Jan 08, 2011

larger memsets.  Among other things, this fixes rdar://8760394 and
allows us to handle "Example 2" from http://blog.regehr.org/archives/320,
compiling it into a single 4096-byte memset:

_mad_synth_mute:                        ## @mad_synth_mute
## BB#0:                                ## %entry
	pushq	%rax
	movl	$4096, %esi             ## imm = 0x1000
	callq	___bzero
	popq	%rax
	ret

llvm-svn: 123089

9a1d63ba

fix an issue in IsPointerOffset that prevented us from recognizing that · 5120ebf1
Chris Lattner authored Jan 08, 2011
```
P and P+1 are relative to the same base pointer.

llvm-svn: 123087
```
5120ebf1
enhance memcpyopt to merge a store and a subsequent · 4dc1fd93
Chris Lattner authored Jan 08, 2011
```
memset into a single larger memset.

llvm-svn: 123086
```
4dc1fd93

constify TargetData references. · c638147e

Chris Lattner authored Jan 08, 2011

Split memset formation logic out into its own
"tryMergingIntoMemset" helper function.

llvm-svn: 123081

c638147e