Showing posts with label INTEL. Show all posts
Showing posts with label INTEL. Show all posts

Friday, August 20, 2010

Bad news for AMD as Intel gains server share



AMD just can't seem to catch a break. After two profitable quarters (amid a multiyear string of losers), a product transition causes it to miss out on the big first-quarter server market rebound that propelled Intel to record profits.
According to a new market share report from IDC, Intel managed to take critical server market share from AMD, with the former company seeing a year-over-year jump from 89.9 percent last year to 93.5 percent this quarter. Meanwhile, AMD's market share dropped from 10.1 percent to 6.5 percent.
AMD's slow transition to the Opteron 6000 series, and the subsequent market share losses, are practically the mirror image of Intel's success in getting its 32nm Westmere and 45nm Nehalem EX server parts into the waiting hands of server makers who finally were ready to open their wallets and start purchasing again.
What makes this situation especially ugly for AMD is the fact that the server market is the company's bread-and-butter. During the worst of the downturn, AMD notoriously jettisoned every part of the company that didn't involve designing x86 processors and GPUs, and it focused in particular on its server business because that was one place where it was still fairly healthy. Server is AMD's absolute core vertical, which means that the company can't really afford too many missteps of this type.
AMD's fortunes could still turn this year, though. The traditional IT upgrade cycle usually happens in the fall, so September will be a big month for the company—at least, it will be if the normal seasonal buying patterns have really returned to the market.
Most PC buying on both the consumer side and corporate side happens in the second half of the year, particularly in the last quarter, as students go back to school and businesses upgrade their machines. This seasonal cyclicality actually halted altogether in 2008, but most of the component suppliers claim to have seen signs that it's returning. Still, we won't know until the fall how much of that normal seasonal surge in buying we'll see this year—the only thing that's certain is that AMD needs that surge to happen, and the company needs to participate in it.

Thursday, August 19, 2010

Intel Buying Security Firm McAfee for $7.68 Billion



Chipmaker Intel is buying security software maker McAfee in a $7.68 billion deal, significantly expanding its online service business.
Chipmaker Intel has announced an agreement to acquire security software maker McAfee in a deal valued at $7.68 billion. The acquisition will not only expand Intel’s security-related business—which, until now, has been mostly related to hardware—but also bolster’s the company’s shift towards offering online and mobile services in addition to hardware as it seeks to diversify its business.
“With the rapid expansion of growth across a vast array of Internet-connected devices, more and more of the elements of our lives have moved online,” said Intel president and CEO Paul Otellini, in a statement. “In the past, energy-efficient performance and connectivity have defined computing requirements. Looking forward, security will join those as a third pillar of what people demand from all computing experiences.”
Intel will pay $48 per share for McAfee, which represents a 60 percent premium over the company’s closing price of $29.93 on Wednesday.
Otellini said Intel and McAfee have been collaborating for the last year and a half, and the partnership will result in products due to reach market in 2011. Neither Intel or McAfee has discussed the nature of those products.
Intel has been keen to establish a bigger footprint in the mobile device market—and building serious security into its chipset solutions might be a good way to get a leg up on its competitors.
“Hardware-enhanced security will lead to breakthroughs in effectively countering the increasingly sophisticated threats of today and tomorrow,” said Intel senior VP of Software and Services RenĂ©e James. “This acquisition is consistent with our software and services strategy to deliver an outstanding computing experience in fast-growing business areas, especially around the move to wireless mobility.”
Although Intel has made some relatively high-profile acquisitions in recent years—including the Havok physics engine used in gaming and embedded systems company Wind River—the company, historically, has not strayed for from its processor and system-on-a-chip businesses. And the company has met with considerable success there, with its market-dominating position drawing litigation from antitrust regulators around the world—Intel justsettled with the FTC and agreed to significantly change the way it markets processors to system manufacturers.

Thursday, August 12, 2010

Intel's Core i7-970 gets reviewed: great for overclocking, still expensive

It may be a cheaper way to join the high-end Core i7 family, but that doesn't mean it's "cheap." Intel's Core i7-970 ($899), which just started shipping to consumers around a month ago, has just undergone a thorough looking-over atHot Hardware, where the six-core chip was tested alongside its more potent (and in turn, more costly) siblings. If you've no interest in dropping over a grand for a Core i7-980X, and you aren't about to lower yourself by purchasing a quad-core Core i7-975, this here chip might just do you proud. In testing, critics found the 970 to be quick, but hardly mind-blowing, when handling more mundane tasks; stir in a few heavily threaded applications, though, and it managed to "sail past" the quad-core contemporaries and "keep pace" with the aforementioned 980X. All told, the silicon managed to perform around 5 percent worse than the 980X, yet it rings up for around 12 percent less. If you've got the workflow to truly take advantage of all six cores, and you can stomach not having the absolute best, it seems as if the 970 strikes a fine balance -- and hey, if you're down with overclocking, you can probably get that 5 percent back with just a mild uptick in your energy bill.

Friday, August 6, 2010

CORE Or Boost? AMD's And Intel's Turbo Features Dissected

Intel arms its Core i5 and Core i7 CPUs with Turbo Boost. AMD's hexa-core Phenom II X6 chips sport Turbo CORE. Both technologies dynamically increase performance based on perceived workloads and available thermal headroom. Which one does the better job?
Automotive turbochargers increase torque and power output, which is why they're used to increase the air-fuel mixture rate per combustion cycle. AMD’s and Intel’s performance-improving technologies don't actually a require an additional piece of hardware bolted on like a turbo would be, but they both invoke the gas compressor namesake anyway.

Instead, both companies' latest six-core models dynamically increase their clock rates to deliver better performance under workload conditions that allow for faster frequencies. We wanted to see whether Intel's Turbo Boost or AMD's Turbo CORE is the better implementation.

Intel was first to offer this performance-enhancing feature. Its Nehalem architecture and the Core i7-900 family first introduced Turbo Boost in late 2008. The technology is capable of accelerating all cores by one clock speed bin (133 MHz) and one or two cores by two speed increments (depending on the particular model). In 2009, the Lynnfield Core i5/i7 quad-core processors for LGA 1156 enabled a more advanced implementation able to accelerate one or two cores by four clock speed increments. The 800-series even bumps clock speed up by five clock speed bins for a single core. One speed bin equals 133 MHz at stock speed, so we’re effectively talking about a 133 to 533 MHz dynamic increase. Turbo Boost is also an available feature on the Clarkdale-based Core i5 dual-core chips.

AMD introduced Turbo CORE with its six-core Phenom II X6 and will keep adding the feature to new models. While Intel's implementation allows the CPU to specifically accelerate one or more cores, AMD’s approach only accelerates three cores in the case of a six-core CPU and only two with quad-core processors.

We grabbed the latest AMD Phenom II X6 and Core i7-980X six-core processors to find out which implementation works best across our benchmark suite in terms of performance and power efficiency. Since the performance level of these two chips is rather different—Intel has more punch—we decided to compare benchmark results with and without the Turbo feature and normalize these to 100% for the non-Turbo results. This way we can compare the relative impact on the respective configurations despite the absolute performance difference. In short, which Turbo implementation gives you more bang for the buck?

Turbo CORE is available on all AMD Phenom II X4 and X6 processors based on the recent 45 nm designs, namely the Thuban six-core and seen-in-the-wild but not-yet-available-at-retail Zosma quad-core models. Should it ever see retail availability, the Phenom II X4 960T at 3.0 GHz nominal speed could accelerate two cores up to 3.4 GHz (+400 MHz) with the thermal headroom available, and if the application load demands the increase. The Phenom II X6 processors increase their clock speeds by 500 MHz, with the exception of the 1090T flagship, which adds 400 MHz to reach from 3.2 to 3.6 GHz.

This implementation can be considered an addition to the Cool’n’Quiet feature, which reduces clock speeds and voltages if there is little work for the processor to do. Once half of the cores are idle, the system reduces their clock speed to the Cool’n’Quiet minimum of 800 MHz. The next step is a voltage increase for the remaining active cores paired with a speed lift of up to 500 MHz, as explained above.

Unfortunately, few workloads would tax exactly three cores by 100%—the conditions needed for AMD’s solution to run at 3.6 GHz. We found that a two-core load scenario is more realistic. This is why the feature works better on a CPU with an even core count, such as the Phenom II X4 960T.

AMD’s Turbo CORE control allows Black Edition processor users to adjust their number of accelerated cores. This makes analysis more complex, but also gives enthusiasts a more powerful tool for fine tuning their systems.

Intel's implementation works best on processors with a lot of scalability inherent to their design, as Turbo Boost covers much broader clock speed ranges. For example, the new six-core "Gulftown," Core i7-980X, is already running close to its thermal ceiling under load. Thus, it's limited to a 266 MHz boost with a single core active, and a modest 133 MHz bump when two or more cores are active. Knowing that Intel’s overclocking headroom is sizable, this is really a pity for enthusiasts. After all, the Phenom II X6 can speed up three cores by up to 400 MHz using a 45 nm process.

Intel’s power gate transistors facilitate cutting power to individual cores. This allows the processor to actually disengage those cores from the overall power envelope, consequently "buying" the overhead needed to increase the remaining cores’ clock speed. The premise here is that fewer cores can run at higher clock speeds before they reach the same thermal output.

While AMD basically reduces clock speed and voltage for inactive cores, Intel can physically shut them down. In theory, this should result in lower power consumption and, paired with the ability to dynamically scale one or more cores up or down, a better overall performance result.

Intel has another advantage that should be mentioned. While AMD's six-core processors access 6 MB of shared L3 cache, Intel's architecture currently offers a massive 12 MB repository. If you switch off individual cores, the remaining active processing units can still access the full 12 MB L3. This should provide advantages for applications that work with limited data and use few threads.

3DMark, a synthetic benchmark, realizes a slight advantage from Intel's architecture and Turbo Boost.

PCMark Vantage clearly shows that Intel’s approach delivers performance gains while AMD’s Turbo Core doesn’t seem to help as much.

iTunes is single-threaded, and is better-accelerated on the Phenom II X6 with Turbo CORE enabled.

The same applies to Lame.

MainConcept is optimized to take advantage of multiple cores, so it benefits more from Turbo Boost, which can kick in even if many cores are taxed.

Once again, we see the multi-threaded advantage in HandBrake, where AMD's processor easily hits its limits on all six cores, preventing Turbo CORE from kicking in.

As expected, switching the Turbo features on or off doesn’t change idle power.

However, peak power increases under Turbo Boost and Turbo CORE. The differences are small, though.

The runtime for our full efficiency suite decreases a bit more on the Intel platform, as there are more applications taking advantage of Intel’s Turbo Boost implementation than AMD’s Turbo CORE.

Average power consumption is much higher on the AMD system with Turbo CORE enabled.

The total power used is exactly the same on the Intel system. This is interesting because the Core i7-980X with Turbo Boost is still faster. AMD’s Turbo CORE-enabled Phenom II X6 delivers more performance, but it requires more power to deliver it.

In the end, the Intel chip's efficiency stays constant. The total power used is exactly the same, but the average power is higher during the workload. As a result, the efficiency is identical. This is like reaching your destination faster in a car without changing your mileage per gallon. AMD’s Turbo implementation sacrifices power efficiency. Runtime decreases, but average power and total power used increase at a higher proportion.

We can only recommend that AMD and Intel continue implementing and developing their Turbo-oriented features. Both do their job in increasing performance. Since the two approaches are different, though, we found that their outcomes in real life are different, as well.

Let’s start with Intel. The six-core, 3.2 GHz Core i7-980X speeds up a single core by 266 MHz if a single-threaded application wants maximum performance, and it can accelerate all six cores by 133 MHz if thermal headroom allows. This is the main difference compared to AMD’s solution, because Intel's Gulftown design can accelerate single-threaded apps, as well as high-end applications. From a multi-core processing standpoint, Turbo Boost makes more sense than Turbo CORE, since all types of workload benefit when compared to nominal clock speed.

AMD’s Turbo CORE only knows one acceleration mode. It increases clock speed for three cores by up to 400 MHz in the case of the Phenom II X6 1090T 3.2 GHz six-core. This means that all applications that utilize no more than three cores experience immediate acceleration. In this case, we found that AMD's performance improvement is higher, as a 400 MHz upgrade is much more noticeable than Intel’s 133/266 MHz speed bump. The downside is nonexistent acceleration if four to six cores are being taxed.

Neither solution is a clear winner. Intel is better for extremely performance-hungry, multi-threaded environments, while AMD's approach provides more benefits for less-threaded environments. The best Turbo technology would be a more granular one, and a perfect Turbo mode would accelerate a single core by even more than AMD’s 400 MHz, two cores by around 400 MHz, three and four cores by less, and all cores by as much as the remaining thermal envelope allows.

Monday, June 28, 2010

Intel: GPUs Only 14x Faster Than CPUs

While Nvidia developers see a 100x speed increase, Intel only sees 14x with some kernels using CUDA.

ZoomA recent paper written by Intel and presented to the International Symposium on Computer Architecture in France claims that Nvidia's GeForce GTX 280 GPU is only 14x faster than its Core i7 960 processor. The paper attempts to debunk claims made by Nvidia developers who saw a 100x performance improvement in some application kernels using CUDA when compared to running them on a CPU.

But is that any surprise? GPUs like the Nvidia GTX 280 have 240 processing core--the average CPU only has six cores. However it's uncertain how Intel came to its "14x" conclusion, as the findings refer to a set of unknown benchmarks--Nvidia even pointed out that they weren't specified in the paper.

"[But] it's actually unclear...what codes were run and how they were compared between the GPU and CPU," said Nvidia spokesperson Andy Keane. "[Still], it wouldn't be the first time the industry has seen Intel using these types of claims with benchmarks."

Playing on the paper's title--Debunking the 100x GPU vs CPU Myth--Keane said that the real myth is that multi-core CPUs are easy for any developer to use and see performance improvements. "In contrast, [our] CUDA parallel computing architecture is a little over 3 years old and already hundreds of consumer, professional and scientific applications are seeing speedups ranging from 10 to 100x using Nvidia GPUs."



Naturally Intel retaliated, saying that Nvidia had taken one small part of the paper out of context and even added that GPU kernel performance is often exaggerated.

"General purpose processors such as the Intel Core i7 or the Intel Xeon are the best choice for the vast majority of applications, be they for the client, general or HPC market segments," said an Intel spokesperson. "This is because of the well-known Intel Architecture programming model, mature tools for software development and more robust performance across a wide range of workloads--not just certain application kernels."

Intel has reportedly acknowledged that application kernels run up to 14 times faster on an Nvidia GeForce GTX 280 compared to a Core i7 960 CPU.

According to Nvidia spokesperson Andy Keane, the kernels were likely tested by Intel on the previous-gen GTX 280 GPU without any optimizations.

Is Nvidia's GeForce 14X faster than Intel's Core i7?"[But] it's actually unclear...what codes were run and how they were compared between the GPU and CPU...[Still], it wouldn't be the first time the industry has seen Intel using these types of claims with benchmarks."

Keane explained that the above-mentioned stats were presented in an Intel paper titled "Debunking the 100x GPU vs CPU Myth" at the International Symposium on Computer Architecture (ISCA) in Saint-Malo, France.

"[Yes], is indeed true that not *all* applications can see this kind of speed up, some just have to make do with an order of magnitude performance increase. But, 100X speed ups and beyond, have been seen by hundreds of developers," Keane told TG Daily in an e-mailed statement. 



"[So], the real myth here is that multi-core CPUs are easy for any developer to use and see performance improvements. In contrast, [our] CUDA parallel computing architecture is a little over 3 years old and already hundreds of consumer, professional and scientific applications are seeing speedups ranging from 10 to 100x using Nvidia GPUs."



Unsurprisingly, an Intel spokesperson TG Daily that Nvidia had taken "one small part of the paper" out of context. 


"While understanding kernel performance can be useful, kernels typically represent only a fraction of the overall work a real application does. As you can see from the data in the paper – claims around the GPU's kernel performance are often exaggerated.

Intel Core i7"[Now], general purpose processors such as the Intel Core i7 or the Intel Xeon are the best choice for the vast majority of applications, be they for the client, general server or HPC market segments. This is because of the well-known Intel Architecture programming model, mature tools for software development and more robust performance across a wide range of workloads - not just certain application kernels.

"[Yes], it is possible to program a graphics processor to compute on non-graphics workloads. But optimal performance is typically achieved only with a high amount of hand optimization, require graphics languages similar to DirectX or OpenGL shader programs or non-industry standard languages. 



"For those HPC application that do benefit from an extremely high level of parallelism, the Intel MIC architecture will be a good choice as it supports standard tools and libraries in standard high level languages like C/C++, FORTRAN, OpenMP, MPI among many other standards."

Tuesday, June 15, 2010

SeaMicro Uses 512 Atom Processors for Internet Server



SeaMicro Uses 512 Atom Processors for Internet Server

SeaMicro has unveiled an Internet-optimized x86 server based on 512 Intel Atom processors usually found in mobile devices. The SeaMicro SM10000 is described as the "ultimate rethink of the volume server" using a quarter of the power and space. The SeaMicro SM10000 has a throughput of 1.28 terabits and uses off-the-shelf OSes and software.
A Silicon Valley startup is using a processor primarily found in mobile Relevant Products/Services devices for a new kind of cloud Relevant Products/Services-computing Relevant Products/Services server Relevant Products/Services. On Monday, SeaMicro unveiled its Internet-optimized x86 server, based on 512 Intel Atom processors.

The model, SM10000, is described by the Santa Clara, Calif.-based company as the "ultimate rethink of the volume server." It said the server is specifically designed for the workloads and traffic patterns on the Internet, and that its approach "dramatically reduces power draw and footprint without requiring any modifications to existing software."

'Fundamental Server Design Mismatch'

The company said the SM10000's key benefits include using a quarter of the power and taking up a quarter of the space as an equivalent, best-in-class volume server. The new unit can run off-the-shelf operating systems and applications without modification, and has an architecture flexible enough to support any CPU.

Other technology innovations include a patented new CPU I/O virtualization Relevant Products/Services, elimination of 90 percent of the components ordinarily required for virtualization, and a supercomputer-style interconnect fabric linking the 512 mini-motherboards into a single system Relevant Products/Services -- which results in a throughput of 1.28 terabits. The architecture supports any protocol, including Ethernet, fibre channel, or data Relevant Products/Services-center Ethernet. The unit's 512 Atom processors run at 1.6 GHz, with one terabyte of DRAM.

SeaMicro is especially promoting the power savings. The company cited reports Relevant Products/Services from Google to the effect that, if current power requirements continue, the cost of energy for a server will, over its lifetime, surpass the initial purchase cost.

SeaMicro said its approach deals with a "fundamental server design mismatch." Servers, it said, were initially designed to solve a "relatively small number of very hard problems," a situation that was changed by the Internet.

Stealth Mode

In a data center focused on the Internet, it said, the challenge is handling many relatively small, independent tasks -- searches, social networking Relevant Products/Services, web page views, e-mail -- and volume servers are not optimized for these kinds of smaller tasks, leading to the kinds of power problems encountered in many data centers.

The company also noted that the CPU only accounts for about a third of the power used in a server, and, in order to achieve major reductions, its SM10000 has reduced non-CPU components. In this design, a high-density, low-power, single-box cluster computer integrates everything normally found in a rack into one unit. This includes computing, storage Relevant Products/Services, networking, server management, and load balancing.

Al Hilwa, program director at IDC, described the new server as an "interesting" use of the Atom processor. He noted that SeaMicro's emphasis that this is an x86, standards-based, "plug and play" server indicates that existing development tools and environments will work without modification, but that has yet to be determined.

If this does work, Hilwa said, "odds are that existing players," like Intel, will begin to develop similar products.

Founded by veterans from such companies as Cisco, Juniper Networks, Sun Microsystems, Intel and Advanced Micro Devices, SeaMicro said the new server is the result of three years of development