Please use this identifier to cite or link to this item: https://hdl.handle.net/20.500.14279/13844
Title: PStorM: Profile storage and matching for feedback-based tuning of MapReduce jobs
Authors: Aboulnaga, Ashraf 
Ead, Mostafa 
Babu, Shivnath
Herodotou, Herodotos 
Major Field of Science: Engineering and Technology
Field Category: Electrical Engineering - Electronic Engineering - Information Engineering
Issue Date: 1-Jan-2014
Source: Advances in Database Technology - EDBT 2014: 17th International Conference on Extending Database Technology, Proceedings
Conference: International Conference on Extending Database Technology 
Abstract: The MapReduce programming model has become widely adopted for large scale analytics on big data. MapReduce systems such as Hadoop have many tuning parameters, many of which have a significant impact on performance. The map and reduce functions that make up a MapReduce job are developed using arbitrary programming constructs, which make them black-box in nature and therefore renders it difficult for users and administrators to make good parameter tuning decisions for a submitted MapReduce job. An approach that is gaining popularity is to provide automatic tuning decisions for submitted MapReduce jobs based on feedback from previously executed jobs. This approach is adopted, for example, by the Starfish system. Starfish and similar systems base their tuning decisions on an execution profile of the MapReduce job being tuned. This execution profile contains summary information about the runtime behavior of the job being tuned, and it is assumed to come from a previous execution of the same job. Managing these execution profiles has not been previously studied. This paper presents PStorM, a profile store and matcher that accurately chooses the relevant profiling information for tuning a submitted MapReduce job from the previously collected profiling information. PStorM can identify accurate tuning profiles even for previously unseen MapReduce jobs. PStorM is currently integrated with the Starfish system, although it can be extended to work with any MapReduce tuning system. Experiments on a large number of MapReduce jobs demonstrate the accuracy and efficiency of profile matching. The results of these experiments show that the profiles returned by PStorM result in tuning decisions that are as good as decisions based on exact profiles collected during pervious executions of the tuned jobs. This holds even for previously unseen jobs, which significantly reduces the overhead of feedback-driven profile-based MapReduce tuning.
ISBN: 978-389318065-3
DOI: 10.5441/002/edbt.2014.02
Rights: © Copyright is with the authors
Type: Conference Papers
Affiliation : Amazon 
Qatar Computing Research Institute 
Microsoft Research 
University of Waterloo 
Duke University 
Appears in Collections:Δημοσιεύσεις σε συνέδρια /Conference papers or poster or presentation

CORE Recommender
Show full item record

SCOPUSTM   
Citations 10

8
checked on Mar 14, 2024

Page view(s)

295
Last Week
0
Last month
3
checked on Dec 3, 2024

Google ScholarTM

Check

Altmetric


Items in KTISIS are protected by copyright, with all rights reserved, unless otherwise indicated.