November 20, 2013

Cray Greases Rails for Hadoop on a Supercomputer

Isaac Lopez

Supercomputer company Cray has joined the Hadoop parade with a new framework that it says will give customers the ability to more easily implement and run Apache Hadoop using Cray’s XC30 supercomputers.

The new framework, which Cray is targeting towards customers deploying Hadoop in scientific environments, will marry Intel’s Hadoop distribution with Cray’s most advanced HPC systems. The move, says Cray, addresses a demand in the scientific computing space for Hadoop to bend towards the needs of the scientific environment.

“We are seeing increased interest from organizations that are ready to leverage Hadoop for analytics in scientific environments, but what those organizations are finding is that Hadoop does not meet all of their needs for scientific use cases,” said Bill Blake, senior vice president and CTO of Cray in a statement.

“They find Hadoop isn’t optimized for the large hierarchical file formats and I/O libraries needed by scientific applications that run close to the parallel file systems,” he continues, “or leverage the types of fast interconnects and tightly integrated systems deployed on supercomputers for performance and scalability. And they find it difficult to share infrastructure or manage complex workflows than span both scientific compute and analytics workloads, and being able to integrate math models with data models in a single high-performance environment. With the Cray Framework for Hadoop, we’re helping organizations harness the openness and power of Hadoop, while better leveraging the investments and progress they’ve made in scientific computing.”

The new framework is said to address the overall efficiency and performance needs of Hadoop in the scientific environment. Per Cray:

The Cray Framework for Hadoop package includes documented best practices and performance enhancements designed to optimize Hadoop for the Cray XC30 line of supercomputers. Built with an emphasis on Scientific Big Data usage, the framework provides enhanced support for data sets found in scientific and engineering applications, as well as improved support for multi-purpose environments, where organizations are looking to align their scientific, compute and data-intensive workloads. This enables users to gain the utility of the Java-based MapReduce Hadoop programming model on the Cray XC30 system, complementing the proven, HPC-optimized languages and tools of the Cray Programming Environment.

The new Hadoop framework announcement comes in conjunction with the news that Cray has signed a $30 million contract with the University of Stuttgart to expand their XC30 supercomputer, nicknamed “Hornet” at the University’s High Performance Computing Center Stuttgart (HLRS). HLRS started a project this past September, called “Dreamcloud,” which appears aimed at Hadoop in which they say they aim to develop “novel load balancing mechanisms that can be applied during runtime in a wide range of parallel and high performance computing systems.”

Per HLRS:

The well-established HPC schedulers, such as a Portable Batch System (PBS), offer effective in terms of the offered scheduling features algorithms and techniques to manage the execution of computational tasks, i.e., in the HPC terminology – batch jobs, on distributed compute nodes. However, with the emergence of high-level e-Infrastructures, such as Grid and Cloud, the traditional cluster scheduling techniques have proved useful to a limited extent only. The main reason for this is that applications running on those infrastructures require a job scheduler to offer a much more extensive set of features in terms of scalability, fault tolerance, and usability, which the traditional, static (with regard to the application) scheduling techniques are not able to meet. The execution frameworks of new-generation parallel applications, such as Hadoop/MapReduce, require the underlying infrastructure scheduler to be more interactive with regard to the applications, in order to enable more intelligent allocation of resources within and also beyond a batch job, i.e., the property of dynamism.

Cray’s new Hadoop-friendly framework appears to provide support in this regard.

Cray says that the framework, which will contain validated and documented best practices for Apache Hadoop configurations, is available as a free download.

LLNL Introduces Big Data Catalyst

SGI Aims to Carve Space in Commodity Big Data Market

Applications: Research Analytics

Technologies: Frameworks, Systems

Sectors: Biosciences, Science

Vendors: Cray

Only registered users may comment. Register using the form below.

Check off newsletters you would like to receive*
- HPCwire
- EnterpriseTech
- Datanami
- Technology Conferences & Events
- Advanced Computing Job Bank
- Technology Product Showcase
Email*
Name*
First Last
Organization*
Job Function*
Industry*
Country*
City*
State*
Province*
- Please check here to receive valuable email offers from Datanami on behalf of our select partners.

Cray Greases Rails for Hadoop on a Supercomputer

Join the discussion Cancel reply

Only registered users may comment. Register using the form below.

April 18, 2024

April 17, 2024

April 16, 2024

Sponsored Partner Content

Get your Data AI Ready – Celebrate One Year of Deep Dish Data Virtual Series!

Supercharge Your Data Lake with Spark 3.3

Learn How to Build a Custom Chatbot Using a RAG Workflow in Minutes [Hands-on Demo]

Overcome ETL Bottlenecks with Metadata-driven Integration for the AI Era [Free Guide]

Gartner® Hype Cycle™ for Analytics and Business Intelligence 2023

The Art of Mastering Data Quality for AI and Analytics

Leading Solution Providers

Tabor Network

Sponsored Whitepapers

Building an Operational Data Warehouse for Real-time Analytics

Can You Use Kafka as a Database?

Sponsored Multimedia

The Power of DataOps: Bring Automation to Life
No Comments

Tactical Steps for Cloud Migration
No Comments

Immuta Data Access Platform
No Comments

Data Mesh: Fact or Fiction?
No Comments

Contributors

Featured Events

Call & Contact Center Expo

AI & Big Data Expo North America 2024

AI Hardware & Edge AI Summit 2024

CDAO Government 2024

Cray Greases Rails for Hadoop on a Supercomputer

Join the discussion Cancel reply

Only registered users may comment. Register using the form below.

April 18, 2024

April 17, 2024

April 16, 2024

Most Read Features

Most Read News In Brief

Most Read This Just In

Sponsored Partner Content

Leading Solution Providers

Tabor Network

Sponsored Whitepapers

Sponsored Multimedia

Contributors

Featured Events

Share

Copy short link