Commits · 7dee697faa5361fc953909b8d4547f2d08af5cab · Roger Ferrer / llvm-epi-0.8

Jul 24, 2013

I'm starting to commit KNL backend. I'll push patches one-by-one. This patch... · 8cfb43f7

Elena Demikhovsky authored Jul 24, 2013

I'm starting to commit KNL backend. I'll push patches one-by-one. This patch includes support for the extended register set XMM16-31, YMM16-31, ZMM0-31.
The full ISA you can see here: http://software.intel.com/en-us/intel-isa-extensions

llvm-svn: 187030

8cfb43f7

Jul 12, 2013

Target/X86: Add explicit Win64 and System V/x86-64 calling conventions. · e8f297ca

Charles Davis authored Jul 12, 2013

Summary:
This patch adds explicit calling convention types for the Win64 and
System V/x86-64 ABIs. This allows code to override the default, and use
the Win64 convention on a target that wants to use SysV (and
vice-versa). This is needed to implement the `ms_abi` and `sysv_abi` GNU
attributes.

Reviewers:

CC:

llvm-svn: 186144

e8f297ca

Jun 25, 2013
- Revert "Temporarily enable MI-Sched on X86." · 121124ac
  Andrew Trick authored Jun 25, 2013
```
This reverts commit 98a9b72e8c56dc13a2617de84503a3d78352789c.

llvm-svn: 184823
```
  121124ac
Jun 24, 2013

Temporarily enable MI-Sched on X86. · 5a1e0af8

Andrew Trick authored Jun 24, 2013

Sorry for the unit test churn. I'll try to make the change permanently
next time.

llvm-svn: 184705

5a1e0af8

Apr 25, 2013

This patch adds the X86FixupLEAs pass, which will reduce instruction · 8b7ab4ba

Preston Gurd authored Apr 25, 2013

latency for certain models of the Intel Atom family, by converting
instructions into their equivalent LEA instructions, when it is both
useful and possible to do so.

llvm-svn: 180573

8b7ab4ba

Mar 29, 2013
- Add support of RDSEED defined in AVX2 extension · a486a11d
  Michael Liao authored Mar 28, 2013
```
llvm-svn: 178314
```
  a486a11d
Mar 27, 2013

· 663e6f95

Preston Gurd authored Mar 27, 2013

For the current Atom processor, the fastest way to handle a call
indirect through a memory address is to load the memory address into
a register and then call indirect through the register.

This patch implements this improvement by modifying SelectionDAG to
force a function address which is a memory reference to be loaded
into a virtual register.

Patch by Sriram Murali.

llvm-svn: 178171

663e6f95

Mar 26, 2013
- Add HLE target feature · e344ec91
  Michael Liao authored Mar 26, 2013
```
llvm-svn: 178082
```
  e344ec91
- Add PREFETCHW codegen support · 5173ee03
  Michael Liao authored Mar 26, 2013
```
- Add 'PRFCHW' feature defined in AVX2 ISA extension

llvm-svn: 178040
```
  5173ee03
Feb 16, 2013
- Reinitialize the ivars in the subtarget so that they can be reset with the new features. · 61375d89
  Bill Wendling authored Feb 16, 2013
```
llvm-svn: 175336
```
  61375d89
- Temporary revert of 175320. · e9434778
  Bill Wendling authored Feb 15, 2013
```
llvm-svn: 175322
```
  e9434778
- Reinitialize the ivars in the subtarget. · a060d0ef
  Bill Wendling authored Feb 15, 2013
```
When we're recalculating the feature set of the subtarget, we need to have the
ivars in their initial state.

llvm-svn: 175320
```
  a060d0ef
Feb 15, 2013

Use the 'target-features' and 'target-cpu' attributes to reset the subtarget features. · aef9c37c

Bill Wendling authored Feb 15, 2013

If two functions require different features (e.g., `-mno-sse' vs. `-msse') then
we want to honor that, especially during LTO. We can do that by resetting the
subtarget's features depending upon the 'target-feature' attribute.

llvm-svn: 175314

aef9c37c

Feb 14, 2013
- added basic support for Intel ADX instructions · f809c649
  Kay Tiong Khoo authored Feb 14, 2013
```
-feature flag, instructions definitions, test cases

llvm-svn: 175196
```
  f809c649
Jan 29, 2013

Teach SDISel to combine fsin / fcos into a fsincos node if the following · 0e88c7d8

Evan Cheng authored Jan 29, 2013

conditions are met:
1. They share the same operand and are in the same BB.
2. Both outputs are used.
3. The target has a native instruction that maps to ISD::FSINCOS node or
   the target provides a sincos library call.

Implemented the generic optimization in sdisel and enabled it for
Mac OSX. Also added an additional optimization for x86_64 Mac OSX by
using an alternative entry point __sincos_stret which returns the two
results in xmm0 / xmm1.

rdar://13087969
PR13204

llvm-svn: 173755

0e88c7d8

Jan 25, 2013

In this patch, we teach X86_64TargetMachine that it has a ILP32 · 597fc123

Eli Bendersky authored Jan 25, 2013

(defined by the x32 ABI) mode, in which case its pointers are 32-bits
in size. This knowledge is also added to X86RegisterInfo that now
returns the appropriate registers in getPointerRegClass.

There are many outcomes to this change. In order to keep the patches
separate and manageable, we start by focusing on some simple testable
cases. The patch adds a test with passing a pointer to a function -
focusing on the difference between the two data models for x86-64.
Another test is added for handling of 'sret' arguments (and
functionality is added in X86ISelLowering to make it work).

A note on naming: the "x32 ABI" document refers to the AMD64
architecture (in LLVM it's distinguished by being is64Bits() in the
x86 subtarget) with two variations: the LP64 (default) data model, and
the ILP32 data model. This patch adds predicates to the subtarget
which are consistent with this naming scheme.

llvm-svn: 173503

597fc123

Jan 08, 2013

Pad Short Functions for Intel Atom · a01daace

Preston Gurd authored Jan 08, 2013

The current Intel Atom microarchitecture has a feature whereby
when a function returns early then it is slightly faster to execute
a sequence of NOP instructions to wait until the return address is ready,
as opposed to simply stalling on the ret instruction until
the return address is ready.

When compiling for X86 Atom only, this patch will run a pass,
called "X86PadShortFunction" which will add NOP instructions where less
than four cycles elapse between function entry and return.

It includes tests.

This patch has been updated to address Nadav's review comments
- Optimize only at >= O1 and don't do optimization if -Os is set
- Stores MachineBasicBlock* instead of BBNum
- Uses DenseMap instead of std::map
- Fixes placement of braces

Patch by Andy Zhang.

llvm-svn: 171879

a01daace

Jan 05, 2013

Revert revision 171524. Original message: · 478b6a47

Nadav Rotem authored Jan 05, 2013

URL: http://llvm.org/viewvc/llvm-project?rev=171524&view=rev
Log:
The current Intel Atom microarchitecture has a feature whereby when a function
returns early then it is slightly faster to execute a sequence of NOP
instructions to wait until the return address is ready,
as opposed to simply stalling on the ret instruction
until the return address is ready.

When compiling for X86 Atom only, this patch will run a pass, called
"X86PadShortFunction" which will add NOP instructions where less than four
cycles elapse between function entry and return.

It includes tests.

Patch by Andy Zhang.

llvm-svn: 171603

478b6a47

Jan 04, 2013

The current Intel Atom microarchitecture has a feature whereby when a function · e36b685a

Preston Gurd authored Jan 04, 2013

returns early then it is slightly faster to execute a sequence of NOP
instructions to wait until the return address is ready,
as opposed to simply stalling on the ret instruction
until the return address is ready.

When compiling for X86 Atom only, this patch will run a pass, called
"X86PadShortFunction" which will add NOP instructions where less than four
cycles elapse between function entry and return.

It includes tests.

Patch by Andy Zhang.

llvm-svn: 171524

e36b685a

Jan 02, 2013

Move all of the header files which are involved in modelling the LLVM IR · 9fb823bb

Chandler Carruth authored Jan 02, 2013

into their new header subdirectory: include/llvm/IR. This matches the
directory structure of lib, and begins to correct a long standing point
of file layout clutter in LLVM.

There are still more header files to move here, but I wanted to handle
them in separate commits to make tracking what files make sense at each
layer easier.

The only really questionable files here are the target intrinsic
tablegen files. But that's a battle I'd rather not fight today.

I've updated both CMake and Makefile build systems (I think, and my
tests think, but I may have missed something).

I've also re-sorted the includes throughout the project. I'll be
committing updates to Clang, DragonEgg, and Polly momentarily.

llvm-svn: 171366

9fb823bb

Dec 04, 2012

Make NaCl naming consistent. The triple OSType is called NaCl and is represented · abe54636

Eli Bendersky authored Dec 04, 2012

textually as NativeClient. Also added a link to the native client project for
readers unfamiliar with it.

A Clang patch will follow shortly.

llvm-svn: 169291

abe54636

Sort includes for all of the .h files under the 'lib' tree. These were · 802d7555

Chandler Carruth authored Dec 04, 2012

missed in the first pass because the script didn't yet handle include
guards.

Note that the script is now able to handle all of these headers without
manual edits. =]

llvm-svn: 169224

802d7555

Nov 29, 2012
- I changed hasAVX() to hasFp256() and hasAVX2() to hasInt256() in X86IselLowering.cpp. · eace43bf
  Elena Demikhovsky authored Nov 29, 2012
```
The logic was not changed, only names.

llvm-svn: 168875
```
  eace43bf
Nov 08, 2012

Add support of RTM from TSX extension · 73cffddb

Michael Liao authored Nov 08, 2012

- Add RTM code generation support throught 3 X86 intrinsics:
  xbegin()/xend() to start/end a transaction region, and xabort() to abort a
  tranaction region

llvm-svn: 167573

73cffddb

Oct 08, 2012
- misched: remove the unused getSpecialAddressLatency hook. · 07dced62
  Andrew Trick authored Oct 08, 2012
```
llvm-svn: 165418
```
  07dced62
Oct 02, 2012

Support for generating ELF objects on Windows. · feb805fc

Andrew Kaylor authored Oct 02, 2012

This adds 'elf' as a recognized target triple environment value and overrides the default generated object format on Windows platforms if that value is present.  This patch also enables MCJIT tests on Windows using the new environment value.

llvm-svn: 165030

feb805fc

Sep 26, 2012
- Remove hasNoAVX method. Can just invert hasAVX instead. · 0a928fa3
  Craig Topper authored Sep 26, 2012
```
llvm-svn: 164664
```
  0a928fa3
Sep 04, 2012

Generic Bypass Slow Div · cdf540d5

Preston Gurd authored Sep 04, 2012

- CodeGenPrepare pass for identifying div/rem ops
- Backend specifies the type mapping using addBypassSlowDivType
- Enabled only for Intel Atom with O2 32-bit -> 8-bit
- Replace IDIV with instructions which test its value and use DIVB if the value
is positive and less than 256.
- In the case when the quotient and remainder of a divide are used a DIV
and a REM instruction will be present in the IR. In the non-Atom case
they are both lowered to IDIVs and CSE removes the redundant IDIV instruction,
using the quotient and remainder from the first IDIV. However,
due to this optimization CSE is not able to eliminate redundant
IDIV instructions because they are located in different basic blocks.
This is overcome by calculating both the quotient (DIV) and remainder (REM)
in each basic block that is inserted by the optimization and reusing the result
values when a subsequent DIV or REM instruction uses the same operands.
- Test cases check for the presents of the optimization when calculating
either the quotient, remainder,  or both.

Patch by Tyler Nowicki!

llvm-svn: 163150

cdf540d5

Aug 30, 2012

Introduce 'UseSSEx' to force SSE legacy encoding · bbd10792

Michael Liao authored Aug 30, 2012

- Add 'UseSSEx' to force SSE legacy insn not being selected when AVX is
  enabled.

  As the penalty of inter-mixing SSE and AVX instructions, we need
  prevent SSE legacy insn from being generated except explicitly
  specified through some intrinsics. For patterns supported by both
  SSE and AVX, so far, we force AVX insn will be tried first relying on
  AddedComplexity or position in td file. It's error-prone and
  introduces bugs accidentally.

  'UseSSEx' is disabled when AVX is turned on. For SSE insns inherited
  by AVX, we need this predicate to force VEX encoding or SSE legacy
  encoding only.

  For insns not inherited by AVX, we still use the previous predicates,
  i.e. 'HasSSEx'. So far, these insns fall into the following
  categories:
  * SSE insns with MMX operands
  * SSE insns with GPR/MEM operands only (xFENCE, PREFETCH, CLFLUSH,
    CRC, and etc.)
  * SSE4A insns.
  * MMX insns.
  * x87 insns added by SSE.

2 test cases are modified:

 - test/CodeGen/X86/fast-isel-x86-64.ll
   AVX code generation is different from SSE one. 'vcvtsi2sdq' cannot be
   selected by fast-isel due to complicated pattern and fast-isel
   fallback to materialize it from constant pool.

 - test/CodeGen/X86/widen_load-1.ll
   AVX code generation is different from SSE one after fixing SSE/AVX
   inter-mixing. Exec-domain fixing prefers 'vmovapd' instead of
   'vmovaps'.

llvm-svn: 162919

bbd10792

Aug 24, 2012
- Custom lower FMA intrinsics to target specific nodes and remove the patterns. · 663d160a
  Craig Topper authored Aug 24, 2012
```
llvm-svn: 162534
```
  663d160a
Aug 23, 2012
- Favor FMA3 over FMA4 if both are enabled. · 4a4634d6
  Craig Topper authored Aug 23, 2012
```
llvm-svn: 162454
```
  4a4634d6
Aug 01, 2012
- Whitespace. · 24c19d20
  Chad Rosier authored Aug 01, 2012
```
llvm-svn: 161122
```
  24c19d20
Jun 03, 2012
- Rename FMA3 feature flag to just FMA to match gcc so it can be added to clang. · 79dbb0c6
  Craig Topper authored Jun 03, 2012
```
llvm-svn: 157903
```
  79dbb0c6
May 31, 2012

X86: Rename the CLMUL target feature to PCLMUL. · a0396e45

Benjamin Kramer authored May 31, 2012

It was renamed in gcc/gas a while ago and causes all kinds of
confusion because it was named differently in llvm and clang.

llvm-svn: 157745

a0396e45

Apr 23, 2012

This patch fixes a problem which arose when using the Post-RA scheduler · 9a091475

Preston Gurd authored Apr 23, 2012

on X86 Atom. Some of our tests failed because the tail merging part of
the BranchFolding pass was creating new basic blocks which did not
contain live-in information. When the anti-dependency code in the Post-RA
scheduler ran, it would sometimes rename the register containing
the function return value because the fact that the return value was
live-in to the subsequent block had been lost. To fix this, it is necessary
to run the RegisterScavenging code in the BranchFolding pass.

This patch makes sure that the register scavenging code is invoked
in the X86 subtarget only when post-RA scheduling is being done.
Post RA scheduling in the X86 subtarget is only done for Atom.

This patch adds a new function to the TargetRegisterClass to control
whether or not live-ins should be preserved during branch folding.
This is necessary in order for the anti-dependency optimizations done
during the PostRASchedulerList pass to work properly when doing
Post-RA scheduling for the X86 in general and for the Intel Atom in particular.

The patch adds and invokes the new function trackLivenessAfterRegAlloc()
instead of using the existing requiresRegisterScavenging().
It changes BranchFolding.cpp to call trackLivenessAfterRegAlloc() instead of
requiresRegisterScavenging(). It changes the all the targets that
implemented requiresRegisterScavenging() to also implement
trackLivenessAfterRegAlloc().  

It adds an assertion in the Post RA scheduler to make sure that post RA
liveness information is available when it is needed.

It changes the X86 break-anti-dependencies test to use –mcpu=atom, in order
to avoid running into the added assertion.

Finally, this patch restores the use of anti-dependency checking
(which was turned off temporarily for the 3.1 release) for
Intel Atom in the Post RA scheduler.

Patch by Andy Zhang!

Thanks to Jakob and Anton for their reviews.

llvm-svn: 155395

9a091475

Mar 17, 2012
- Reorder includes in Target backends to following coding standards. Remove some... · b25fda95
  Craig Topper authored Mar 17, 2012
```
Reorder includes in Target backends to following coding standards. Remove some superfluous forward declarations.

llvm-svn: 152997
```
  b25fda95
Feb 19, 2012
- some comment fix for X86 and ARM · e1d61969
  Jia Liu authored Feb 19, 2012
```
llvm-svn: 150902
```
  e1d61969
Feb 18, 2012
- Emacs-tag and some comment fix for all ARM, CellSPU, Hexagon, MBlaze, MSP430,... · b22310fd
  Jia Liu authored Feb 18, 2012
```
Emacs-tag and some comment fix for all ARM, CellSPU, Hexagon, MBlaze, MSP430, PPC, PTX, Sparc, X86, XCore.

llvm-svn: 150878
```
  b22310fd
Feb 07, 2012
- Use LEA to adjust stack ptr for Atom. Patch by Andy Zhang. · 1b81fddd
  Evan Cheng authored Feb 07, 2012
```
llvm-svn: 150008
```
  1b81fddd
Feb 05, 2012

Begin fleshing out more convenience predicates in llvm::Triple and · ebd90c58

Chandler Carruth authored Feb 05, 2012

convert at least one client over to use them. Subsequent patches both to
LLVM and Clang will try to convert more people over to a common set of
predicates.

This round of predicates is focused on OS-categorization predicates.

llvm-svn: 149815

ebd90c58