Data Lake for Enterprises
上QQ阅读APP看书,第一时间看更新

Knowing Hadoop distributions

A Big Data ecosystem consists of multiple capabilities, and for every capability in the ecosystem, there are one or more frameworks. Different distributions realize these capabilities in their own specific ways and also have some additional edge over other competitors in the same space.

Figure 01: Hadoop distributions

Shown here are some of the leading distributions of Hadoop framework, wherein Cloudera, Hortonworks, and MapR are the leaders in commercial space while Apache Hadoop is an open source distribution. These commercial offerings, while having their own specific capabilities, are largely based on the specifications of the open source Hadoop framework.

Just to put a few things into the perspective of why a Hadoop distribution should be chosen, unfortunately there is no straight answer for it. However, we can compare these distributions across various dimensions that we may be interested in for evaluation.