Commits · 6fab2e9418635c435a516e83279b25532e3bb1bf · Roger Ferrer / llvm-epi-0.8

Jan 16, 2011

if an alloca is only ever accessed as a unit, and is accessed with load/store instructions, · 6fab2e94

Chris Lattner authored Jan 16, 2011

then don't try to decimate it into its individual pieces.  This will just make a mess of the
IR and is pointless if none of the elements are individually accessed.  This was generating
really terrible code for std::bitset (PR8980) because it happens to be lowered by clang
as an {[8 x i8]} structure instead of {i64}.

The testcase now is optimized to:

define i64 @test2(i64 %X) {
  br label %L2

L2:                                               ; preds = %0
  ret i64 %X
}

before we generated:

define i64 @test2(i64 %X) {
  %sroa.store.elt = lshr i64 %X, 56
  %1 = trunc i64 %sroa.store.elt to i8
  %sroa.store.elt8 = lshr i64 %X, 48
  %2 = trunc i64 %sroa.store.elt8 to i8
  %sroa.store.elt9 = lshr i64 %X, 40
  %3 = trunc i64 %sroa.store.elt9 to i8
  %sroa.store.elt10 = lshr i64 %X, 32
  %4 = trunc i64 %sroa.store.elt10 to i8
  %sroa.store.elt11 = lshr i64 %X, 24
  %5 = trunc i64 %sroa.store.elt11 to i8
  %sroa.store.elt12 = lshr i64 %X, 16
  %6 = trunc i64 %sroa.store.elt12 to i8
  %sroa.store.elt13 = lshr i64 %X, 8
  %7 = trunc i64 %sroa.store.elt13 to i8
  %8 = trunc i64 %X to i8
  br label %L2

L2:                                               ; preds = %0
  %9 = zext i8 %1 to i64
  %10 = shl i64 %9, 56
  %11 = zext i8 %2 to i64
  %12 = shl i64 %11, 48
  %13 = or i64 %12, %10
  %14 = zext i8 %3 to i64
  %15 = shl i64 %14, 40
  %16 = or i64 %15, %13
  %17 = zext i8 %4 to i64
  %18 = shl i64 %17, 32
  %19 = or i64 %18, %16
  %20 = zext i8 %5 to i64
  %21 = shl i64 %20, 24
  %22 = or i64 %21, %19
  %23 = zext i8 %6 to i64
  %24 = shl i64 %23, 16
  %25 = or i64 %24, %22
  %26 = zext i8 %7 to i64
  %27 = shl i64 %26, 8
  %28 = or i64 %27, %25
  %29 = zext i8 %8 to i64
  %30 = or i64 %29, %28
  ret i64 %30
}

In this case, instcombine was able to eliminate the nonsense, but in PR8980 enough
PHIs are in play that instcombine backs off.  It's better to not generate this stuff
in the first place.

llvm-svn: 123571

6fab2e94

Use an irbuilder to get some trivial constant folding when doing a store · 7cd8cf7d
Chris Lattner authored Jan 16, 2011
```
of a constant.

llvm-svn: 123570
```
7cd8cf7d
remove a dead check, this was needed before we had an explicit veto on uses of phis. · adb1a233
Chris Lattner authored Jan 16, 2011
```
llvm-svn: 123569
```
adb1a233

enhance FoldOpIntoPhi in instcombine to try harder when a phi has · d55581de

Chris Lattner authored Jan 16, 2011

multiple uses.  In some cases, all the uses are the same operation,
so instcombine can go ahead and promote the phi.  In the testcase
this pushes an add out of the loop.

llvm-svn: 123568

d55581de

Spill R4 if it's going to be used to restore SP from FP. · 572756ac
Evan Cheng authored Jan 16, 2011
```
llvm-svn: 123567
```
572756ac
remove the AllowAggressive argument to FoldOpIntoPhi. It is forced to false in the · ea7131a0
Chris Lattner authored Jan 16, 2011
```
first line of the function because it isn't a good idea, even for compares.

llvm-svn: 123566
```
ea7131a0
more cleanups: use the IR builder. · ff2e7377
Chris Lattner authored Jan 16, 2011
```
llvm-svn: 123565
```
ff2e7377
tidy up code. · 25ce2805
Chris Lattner authored Jan 16, 2011
```
llvm-svn: 123564
```
25ce2805
Improve the safety of my globalopt enhancement by ensuring that the bitcast · 4e54efd6
Owen Anderson authored Jan 16, 2011
```
of the stored value to the new store type is always.  Also, add a testcase.

llvm-svn: 123563
```
4e54efd6
fix PR8983, a broken assertion. · 08f43456
Chris Lattner authored Jan 16, 2011
```
llvm-svn: 123562
```
08f43456
Implement AnalyzeBranch in Sparc Backend. · 1b0e2cbf
Venkatraman Govindaraju authored Jan 16, 2011
```
llvm-svn: 123561
```
1b0e2cbf
fix PR8981, a crash trying to form a conditional inc with a floating point compare. · 218092e6
Chris Lattner authored Jan 16, 2011
```
llvm-svn: 123560
```
218092e6
reapply my fix for PR8961 with a tweak to properly handle · 2d186574
Chris Lattner authored Jan 16, 2011
```
multi-instruction sequences like calls.  Many thanks to Jakob for
finding a testcase.

llvm-svn: 123559
```
2d186574
simplify this code, it is still broken but will follow up on llvm-commits. · 8b4952fc
Chris Lattner authored Jan 16, 2011
```
llvm-svn: 123558
```
8b4952fc
Revert "Archive: Replace all internal uses of PathV1 with PathV2. The external... · 2ff30b84
Michael J. Spencer authored Jan 16, 2011
```
Revert "Archive: Replace all internal uses of PathV1 with PathV2. The external API still uses PathV1."

llvm-svn: 123557
```
2ff30b84
Simplify a README.txt entry significantly to expose the core issue. · ef28abef
Chandler Carruth authored Jan 16, 2011
```
llvm-svn: 123556
```
ef28abef
remove the partial specialization pass. It is unmaintained and has bugs. · 1e209b87
Chris Lattner authored Jan 16, 2011
```
llvm-svn: 123554
```
1e209b87

Jan 15, 2011

Archive: Replace all internal uses of PathV1 with PathV2. The external API still uses PathV1. · a0ce7632
Michael J. Spencer authored Jan 15, 2011
```
llvm-svn: 123551
```
a0ce7632
Add an assert so we don't silently miscompile ctpop for bit widths > 128. · bec03ea7
Benjamin Kramer authored Jan 15, 2011
```
llvm-svn: 123549
```
bec03ea7
Support/PathV2: Add identify_magic. · 94b2ab35
Michael J. Spencer authored Jan 15, 2011
```
llvm-svn: 123548
```
94b2ab35

Reimplement CTPOP legalization with the "best" algorithm from · fff2517e

Benjamin Kramer authored Jan 15, 2011

http://graphics.stanford.edu/~seander/bithacks.html#CountBitsSetParallel

In a silly microbenchmark on a 65 nm core2 this is 1.5x faster than the old
code in 32 bit mode and about 2x faster in 64 bit mode. It's also a lot shorter,
especially when counting 64 bit population on a 32 bit target.

I hope this is fast enough to replace Kernighan-style counting loops even when
the input is rather sparse.

llvm-svn: 123547

fff2517e

Support/PathV2: Implement has_magic in terms of get_magic. · 7887466a
Michael J. Spencer authored Jan 15, 2011
```
llvm-svn: 123545
```
7887466a
Support/PathV2: Implement get_magic. · ee1699c3
Michael J. Spencer authored Jan 15, 2011
```
llvm-svn: 123544
```
ee1699c3
Add missing whitespace. · 4a1ff16b
Nick Lewycky authored Jan 15, 2011
```
llvm-svn: 123543
```
4a1ff16b
Make constmerge a two-pass algorithm so that it won't miss merging · 0296a481
Nick Lewycky authored Jan 15, 2011
```
opporuntities. Fixes PR8978.

llvm-svn: 123541
```
0296a481
Try to unbreak selfhost. · ed5f2e50
Benjamin Kramer authored Jan 15, 2011
```
llvm-svn: 123537
```
ed5f2e50
Add a cache that protects mergefunc's internals from more surprises in DenseSet. · 540f9536
Nick Lewycky authored Jan 15, 2011
```
Also, replace tabs with spaces. Yes, it's 2011.

llvm-svn: 123535
```
540f9536

Teach LazyValueInfo that allocas aren't NULL. Over all of llvm-test, this saves · 367f98f0

Nick Lewycky authored Jan 15, 2011

half a million non-local queries, each of which would otherwise have triggered a
linear scan over a basic block.

Also fix a fixme for memory intrinsics which dereference pointers. With this,
we prove that a pointer is non-null because it was dereferenced by an intrinsic
112 times in llvm-test.

llvm-svn: 123533

367f98f0

Allow unnamed_addr on declarations. · 489e505a
Rafael Espindola authored Jan 15, 2011
```
llvm-svn: 123529
```
489e505a
temporarily revert r123526. While working on a follow-on patch I · af263907
Chris Lattner authored Jan 15, 2011
```
realize that ConstantFoldTerminator doesn't preserve dominfo.

llvm-svn: 123527
```
af263907

fix rdar://8785296 - -fcatch-undefined-behavior generates inefficient code · 8df83c4a

Chris Lattner authored Jan 15, 2011

The basic issue is that isel (very reasonably!) expects conditional branches
to be folded, so CGP leaving around a bunch dead computation feeding
conditional branches isn't such a good idea.  Just fold branches on constants
into unconditional branches.

llvm-svn: 123526

8df83c4a

simplify code, no functionality change. · ee588def
Chris Lattner authored Jan 15, 2011
```
llvm-svn: 123525
```
ee588def

Now that instruction optzns can update the iterator as they go, we can · 1b93be50

Chris Lattner authored Jan 15, 2011

have objectsize folding recursively simplify away their result when it
folds.  It is important to catch this here, because otherwise we won't
eliminate the cross-block values at isel and other times.

llvm-svn: 123524

1b93be50

make the current instruction iterator an ivar, allowing xforms that · 7a277144

Chris Lattner authored Jan 15, 2011

potentially invalidate it (like inline asm lowering) to be sunk into
their proper place, cleaning up a ton of code.

llvm-svn: 123523

7a277144

implement an instcombine xform that canonicalizes casts outside of and-with-constant operations. · 9c10d587

Chris Lattner authored Jan 15, 2011

This fixes rdar://8808586 which observed that we used to compile:


union xy {
        struct x { _Bool b[15]; } x;
        __attribute__((packed))
        struct y {
                __attribute__((packed)) unsigned long b0to7;
                __attribute__((packed)) unsigned int b8to11;
                __attribute__((packed)) unsigned short b12to13;
                __attribute__((packed)) unsigned char b14;
        } y;
};

struct x
foo(union xy *xy)
{
        return xy->x;
}

into:

_foo:                                   ## @foo
	movq	(%rdi), %rax
	movabsq	$1095216660480, %rcx    ## imm = 0xFF00000000
	andq	%rax, %rcx
	movabsq	$-72057594037927936, %rdx ## imm = 0xFF00000000000000
	andq	%rax, %rdx
	movzbl	%al, %esi
	orq	%rdx, %rsi
	movq	%rax, %rdx
	andq	$65280, %rdx            ## imm = 0xFF00
	orq	%rsi, %rdx
	movq	%rax, %rsi
	andq	$16711680, %rsi         ## imm = 0xFF0000
	orq	%rdx, %rsi
	movl	%eax, %edx
	andl	$-16777216, %edx        ## imm = 0xFFFFFFFFFF000000
	orq	%rsi, %rdx
	orq	%rcx, %rdx
	movabsq	$280375465082880, %rcx  ## imm = 0xFF0000000000
	movq	%rax, %rsi
	andq	%rcx, %rsi
	orq	%rdx, %rsi
	movabsq	$71776119061217280, %r8 ## imm = 0xFF000000000000
	andq	%r8, %rax
	orq	%rsi, %rax
	movzwl	12(%rdi), %edx
	movzbl	14(%rdi), %esi
	shlq	$16, %rsi
	orl	%edx, %esi
	movq	%rsi, %r9
	shlq	$32, %r9
	movl	8(%rdi), %edx
	orq	%r9, %rdx
	andq	%rdx, %rcx
	movzbl	%sil, %esi
	shlq	$32, %rsi
	orq	%rcx, %rsi
	movl	%edx, %ecx
	andl	$-16777216, %ecx        ## imm = 0xFFFFFFFFFF000000
	orq	%rsi, %rcx
	movq	%rdx, %rsi
	andq	$16711680, %rsi         ## imm = 0xFF0000
	orq	%rcx, %rsi
	movq	%rdx, %rcx
	andq	$65280, %rcx            ## imm = 0xFF00
	orq	%rsi, %rcx
	movzbl	%dl, %esi
	orq	%rcx, %rsi
	andq	%r8, %rdx
	orq	%rsi, %rdx
	ret

We now compile this into:

_foo:                                   ## @foo
## BB#0:                                ## %entry
	movzwl	12(%rdi), %eax
	movzbl	14(%rdi), %ecx
	shlq	$16, %rcx
	orl	%eax, %ecx
	shlq	$32, %rcx
	movl	8(%rdi), %edx
	orq	%rcx, %rdx
	movq	(%rdi), %rax
	ret

A small improvement :-)

llvm-svn: 123520

9c10d587

one more instcombine variant that is needed to work with future changes, · e20dd530
Chris Lattner authored Jan 15, 2011
```
no functionality change currently.

llvm-svn: 123517
```
e20dd530
fix typo · 497459d5
Chris Lattner authored Jan 15, 2011
```
llvm-svn: 123516
```
497459d5
Catch ~x < cst just like ~x < ~y, we currently handle this through · f3c4eeff
Chris Lattner authored Jan 15, 2011
```
means that are about to disappear.

llvm-svn: 123515
```
f3c4eeff
reduce indentation · 311aa63c
Chris Lattner authored Jan 15, 2011
```
llvm-svn: 123514
```
311aa63c
80-col. · cc385c0c
Eric Christopher authored Jan 15, 2011
```
llvm-svn: 123505
```
cc385c0c