Showing posts with label Hadoop. Show all posts
Showing posts with label Hadoop. Show all posts

August 28, 2019

Apache Hadoop Terms/Abbreviations


HDFS - Hadoop Distributed File System
GFS - Google File System
JSON - Java Script Object Notation

NN - NameNode
DN - Data Node
SNN - Secondary NameNode
JT - Job Tracker
TT - Task Tracker
HA NN - Highly Available NameNode (or NN HA - NameNode Highly Available)

REST - Representational State Transfer
HiveQL or HQL - Hive Query Langauge or Hive SQL
CDH - Cloudera’s Distribution Including Apache Hadoop
ZKFC - ZooKeeper Failover Controller
HVE - Hadoop Virtual Extensions
FUSE - Filesystem In Userspace
YARN - Yet Another Resource Negotiator
SerDe - Serialiser and Deserialiser
Hue - Hadoop User Experience

JRE - Java Runtime Engine
JNA - Java Native Access
JVM - Java Virtual Machine
JMX - Java Management Extensions
JAR - Java ARchive

AWS - Amazon Web Services
S3 - Simple Storage Service
EC2 - Elastic Compute Cloud
VPC - Virtual Private Cloud
EMR - Elastic Map Reduce
IAM - Identity Access Management
RDS - Relational Database Service

GCP - Google Cloud Platform
GCE - Google Compute Engine

User Defined Functions (UDFs)
User Defined Aggregates Functions (UDAFs)
User Defined Table Functions (UDTFs)


Related Hadoop Articles: Apache Hadoop   Hadoop Commercial Distributions

April 20, 2019

Apache Hadoop Versions

Hadoop Versions

Hadoop 3
6 Apr 2018: Release 3.1.0 available
25 March 2018: Release 3.0.1 available
13 December 2017: Release 3.0.0 generally available
03 October 2017: Release 3.0.0-beta1 available
07 July 2017: Release 3.0.0-alpha4 available
26 May 2017: Release 3.0.0-alpha3 available
25 January, 2017: Release 3.0.0-alpha2 available
03 September, 2016: Release 3.0.0-alpha1 available

Hadoop 2
14 December, 2017: Release 2.7.5 available
12 December 2017: Release 2.8.3 available
17 November 2017: Release 2.9.0 available
24 October 2017: Release 2.8.2 available
04 August, 2017: Release 2.7.4 available
08 June, 2017: Release 2.8.1 available
22 March 2017: Release 2.8.0 available
08 October, 2016: Release 2.6.5 available
25 August, 2016: Release 2.7.3 available
11 February, 2016: Release 2.6.4 available
25 January, 2016: Release 2.7.2 (stable) available
17 December, 2015: Release 2.6.3 available
28 October, 2015: Release 2.6.2 available
23 September, 2015: Release 2.6.1 available
06 July, 2015: Release 2.7.1 (stable) available
21 April 2015: Release 2.7.0 available
18 November, 2014: Release 2.6.0 available
19 November, 2014: Release 2.5.2 available
12 September, 2014: Release 2.5.1 available
11 August, 2014: Release 2.5.0 available
30 June, 2014: Release 2.4.1 available
07 April, 2014: Release 2.4.0 available
20 February, 2014: Release 2.3.0 available
15 October, 2013: Release 2.2.0 available
23 September, 2013: Release 2.1.1-beta available
25 August, 2013: Release 2.1.0-beta available
23 August, 2013: Release 2.0.6-alpha available
6 June, 2013: Release 2.0.5-alpha available
25 April, 2013: Release 2.0.4-alpha available
Hadoop 2.0.3-alpha (released on 14 February, 2013)
Hadoop 2.0.2-alpha (released on 9 October, 2012)
Hadoop 2.0.1-alpha (released on 26 July, 2012)
Hadoop 2.0.0-alpha (released on 23 May, 2012)

Hadoop 1
1 Aug, 2013: Release 1.2.1 (stable) available
13 May, 2013: Release 1.2.0 available
Hadoop 1.1.2 (released on 15 February, 2013)
Hadoop 1.1.1 (released on 1 December, 2012)
Hadoop 1.1.0 (released on 13 October, 2012)
Hadoop 1.0.4 (released on 12 October, 2012)
Hadoop 1.0.3 (released on 16 May, 2012)
Hadoop 1.0.2 (released on 3 Apr, 2012)
Hadoop 1.0.1 (released on 10 Mar, 2012)
Hadoop 1.0.0 (released on 27 December, 2011)

Hadoop 0
27 June, 2014: Release 0.23.11 available
11 December, 2013: Release 0.23.10 available
8 July, 2013: Release 0.23.9 available
5 June, 2013: Release 0.23.8 available
18 April, 2013: Release 0.23.7 available
Hadoop 0.23.6 (released on 7 February, 2013)
Hadoop 0.23.5 (released on 28 November, 2012)
Hadoop 0.23.4 (released on 15 October, 2012)
Hadoop 0.23.3 (released on 17 September, 2012)
Hadoop 0.23.1 (released on 27 Feb, 2012)
Hadoop 0.22.0 (released on 10 December, 2011)
Hadoop 0.23.0 (released on 11 Nov, 2011)
Hadoop 0.20.205.0 (released on 17 Oct, 2011)
Hadoop 0.20.204.0 (released on 5 Sep, 2011)
Hadoop 0.20.203.0 (released on 11 May, 2011)
Hadoop 0.21.0 (released on 23 August, 2010)
Hadoop 0.20.2 (released on 26 February, 2010)
Hadoop 0.20.1 (released on 14 September, 2009)
Hadoop 0.19.2 (released on 23 July, 2009)
Hadoop 0.20.0 (released on 22 April, 2009)
Hadoop 0.19.1 (released on 24 February, 2009)
Hadoop 0.18.3 (released on 29 January, 2009)
Hadoop 0.19.0 (released on 21 November, 2008)
Hadoop 0.18.2 (released on 3 November, 2008)
Hadoop 0.18.1 (released on 17 September, 2008)
Hadoop 0.18.0 (released on 22 August, 2008)
Hadoop 0.17.2 (released on 19 August, 2008)
Hadoop 0.17.1 (released on 23 June, 2008)
Hadoop 0.17.0 (released on 20 May, 2008)
Hadoop 0.16.4 (released on 5 May, 2008)
Hadoop 0.16.3 (released on 16 April, 2008)
Hadoop 0.16.2 (released on 2 April, 2008)
Hadoop 0.16.1 (released on 13 March, 2008)
Hadoop 0.16.0 (released on 7 February, 2008)
Hadoop 0.15.3 (released on 18 January, 2008)
Hadoop 0.15.2 (released on 2 January, 2008)
Hadoop 0.15.1 (released on 27 November, 2007)
Hadoop 0.14.4 (released on 26 November, 2007)
Hadoop 0.15.0 (released on 29 October 2007)
Hadoop 0.14.3 (released on 19 October, 2007)
Hadoop 0.14.1 (released on 4 September, 2007)


Related Hadoop Articles: Hadoop Commercial Distributions    Hadoop Commands

April 2, 2019

Hadoop Commands

Hadoop CLI Commands

hadoop command [genericOptions] [commandOptions]

hadoop fs
Usage: hadoop fs [generic options]
[-appendToFile <localsrc> ... <dst>]
[-cat [-ignoreCrc] <src> ...]
[-checksum <src> ...]
[-chgrp [-R] GROUP PATH...]
[-chmod [-R] <MODE[,MODE]... | OCTALMODE> PATH...]
[-chown [-R] [OWNER][:[GROUP]] PATH...]
[-copyFromLocal [-f] [-p] [-l] [-d] <localsrc> ... <dst>]
[-copyToLocal [-f] [-p] [-ignoreCrc] [-crc] <src> ... <localdst>]
[-count [-q] [-h] [-v] [-t [<storage type>]] [-u] [-x] <path> ...]
[-cp [-f] [-p | -p[topax]] [-d] <src> ... <dst>]
[-createSnapshot <snapshotDir> [<snapshotName>]]
[-deleteSnapshot <snapshotDir> <snapshotName>]
[-df [-h] [<path> ...]]
[-du [-s] [-h] [-x] <path> ...]
[-expunge]
[-find <path> ... <expression> ...]
[-get [-f] [-p] [-ignoreCrc] [-crc] <src> ... <localdst>]
[-getfacl [-R] <path>]
[-getfattr [-R] {-n name | -d} [-e en] <path>]
[-getmerge [-nl] [-skip-empty-file] <src> <localdst>]
[-help [cmd ...]]
[-ls [-C] [-d] [-h] [-q] [-R] [-t] [-S] [-r] [-u] [<path> ...]]
[-mkdir [-p] <path> ...]
[-moveFromLocal <localsrc> ... <dst>]
[-moveToLocal <src> <localdst>]
[-mv <src> ... <dst>]
[-put [-f] [-p] [-l] [-d] <localsrc> ... <dst>]
[-renameSnapshot <snapshotDir> <oldName> <newName>]
[-rm [-f] [-r|-R] [-skipTrash] [-safely] <src> ...]
[-rmdir [--ignore-fail-on-non-empty] <dir> ...]
[-setfacl [-R] [{-b|-k} {-m|-x <acl_spec>} <path>]|[--set <acl_spec> <path>]]
[-setfattr {-n name [-v value] | -x name} <path>]
[-setrep [-R] [-w] <rep> <path> ...]
[-stat [format] <path> ...]
[-tail [-f] <file>]
[-test -[defsz] <path>]
[-text [-ignoreCrc] <src> ...]
[-touchz <path> ...]
[-truncate [-w] <length> <path> ...]
[-usage [cmd ...]]

hadoop fs -ls / -- Get a directory listing of the HDFS root directory
hadoop fs -ls /user/hive
hadoop fs -ls s3a://satya-hive/trip_data_dec17
hadoop fs -cat /test1/foo.txt -- Display contents of the file residing in HDFS
hadoop fs -cat /test1/foo.txt | wc -l -- Count the number of lines of file in  HDFS
hadoop fs -rm /test1/foo.txt -- Remove the file in HDFS
hadoop fs -rm -r hadoop-test2
hadoop fs -mkdir /test6/ -- Create a directory under the HDFS root directory
hadoop fs -put foo.txt /test6/ -- Copy file foo.txt from local disk to the directory
hadoop fs -put /etc/note.txt /test2/note_fs.txt
hadoop fs -get /user/satya/passwd ./ -- Copy the file back to local disk
hadoop fs -cp hadoop-test1/dwp-payments-april10.csv hadoop-test2
hadoop fs -cp /user/satya/reviews_Home_and_Kitchen_5.json s3a://satya-sparks/reviews_HomeKitchen
hadoop fs -Ddfs.replication=2 -cp hadoop-test2/dwp-payments-april10.csv hadoop-test2/test_with_rep2.csv
hadoop fs -mv hadoop-test1/dwp-payments-april10.csv hadoop-test3
hadoop fs -setrep 5 -R /user/satya/tmp/
hadoop fs -chmod 1777 /tmp
hadoop fs -touch /user/satya/test/foo
hadoop fs -rmr /user/satya/test/foo
hadoop fs -touchz /user/satya/test/bar
hadoop fs -count -q /user/satya
hadoop fs -copyFromLocal  /hirw-starterkit/hdfs/commands/dwp-payments-april10.csv hadoop-test1
hadoop fs -copyToLocal hadoop-test1/dwp-payments-april10.csv .
hadoop -execute start-all.sh

hadoop job -list
hadoop job -kill jobID
hadoop job -list-attempt-ids jobID taskType taskState
hadoop job -kill-task taskAttemptId

hadoop namenode -format

hadoop jar <jar_file> wordcount <output_file>
hadoop jar /opt/hadoop/hadoop-examples-1.0.4.jar wordcount /out/wc_output

hadoop dfsadmin -report
hadoop dfsadmin -setSpaceQuota 10737418240 /user/esammer
hadoop dfsadmin -refreshNodes
hadoop dfsadmin -upgradeProgress status
hadoop dfsadmin -finalizeUpgrade

hadoop fsck
Usage: DFSck <path> [-move | -delete | -openforwrite] [-files [-blocks [-locations | -racks]]]
<path> -- start checking from this path
-move -- move corrupted files to /lost+found
-delete -- delete corrupted files
-files -- print out files being checked
-openforwrite -- print out files opened for write
-blocks -- print out block report
-locations -- print out locations for every block
-racks -- print out network topology for data-node locations
By default fsck ignores files opened for write, use -openforwrite to report such files. They are usually tagged CORRUPT or HEALTHY depending on their block allocation status.
hadoop fsck / -files -blocks -locations
hadoop fsck /user/satya -files -blocks -locations

hadoop distcp -- Distributed Copy (distcp)
distcp [OPTIONS] <srcurl>* <desturl>
OPTIONS:
-p[rbugp] Preserve status
r: replication number
b: block size
u: user
g: group
p: permission
-p alone is equivalent to -prbugp
-i Ignore failures
-log <logdir> Write logs to <logdir>
-m <num_maps> Maximum number of simultaneous copies
-overwrite Overwrite destination
-update Overwrite if src size different from dst size
-skipcrccheck Do not use CRC check to determine if src is different from dest. Relevant only if -update is specified
-f <urilist_uri> Use list at <urilist_uri> as src list
-filelimit <n> Limit the total number of files to be <= n
-sizelimit <n> Limit the total size to be <= n bytes
-delete Delete the files existing in the dst but not in src
-mapredSslConf <f> Filename of SSL configuration for mapper task

NOTE 1: if -overwrite or -update are set, each source URI is interpreted as an isomorphic update to an existing directory.
For example:
hadoop distcp -p -update "hdfs://A:8020/user/foo/bar" "hdfs://B:8020/user/foo/baz"
would update all descendants of 'baz' also in 'bar'; it would *not* update /user/foo/baz/bar

NOTE 2: The parameter <n> in -filelimit and -sizelimit can be specified with symbolic representation. For examples,
1230k = 1230 * 1024 = 1259520
891g = 891 * 1024^3 = 956703965184

hadoop distcp hdfs://A:8020/path/one hdfs://B:8020/path/two
hadoop distcp /path/one /path/two


Related Hadoop Articles: Apache Hadoop    Hadoop Training in India (Bangalore/Hyderabad)


October 7, 2018

Hadoop Training Institutes in India

Apache Hadoop Online/Offline Training Institutes in India (Bangalore/Hyderabad)

Many institutes/companies are providing Hadoop training in India. And it's available like any other course (like Java, .NET) in Ameerpet (Hyderabad), but it's costly.


BigData/Hadoop Training in Hyderabad:


I had personally enquiried all the below Hyderabad Hadoop training institutes.

Big Data/Hadoop training course fees - Cost will be around Rs. 12,000/- to Rs. 18,000/-


Hadoop Training Institutes @ Hyderabad
Contact Numbers
Website/Mail
Aswin DWH Technologies, Ameerpet, Hyderabad 040 65226522, 9295555929 www.aswindwhtechnologies.com
Ctrl-A Technologes, Ameerpet, Hyderabad 040 40142647, 9505675670
eNexus Software Technologies, SR Nagar, Hyderabad 040 64622229, 8096374447
FIFO Technologies, SR Nagar, Hyderabad, India 040 6565 6090, 08121389698 www.fifotechnologies.com
iFocus IT Solutions, Ameerpet, Hyderabad 040 65444455, 9849388104
Inuomsoft Pvt. Ltd, Ameerpet, Hyderabad 040 66633775, 9642101010
Kelly Technologies, Ameerpet, Hyderabad 040 6462 6789, 9985706789 www.kellytechno.com
Mind Links IT R&D Labs, Ameerpet, Hyderabad 040 64568575, 9951081013 www.themindlinks.com
Nihitha IT, Ameerpet, Hyderabad, India 040 66413341, 9000415176
Sree Aditya Software Solutions, Ameerpet, Hyderabad 040 60606011 www.sreeaaditya.com
Sree Nipuna Software Solutions, Ameerpet, Hyderabad 040 66849695, 9247052622
Surya Software Solutions, Ameerpet, Hyderabad 9966661188, 8686536306
Version IT, Ameerpet, Hyderabad, India 040 66738677, 040 66738766
Visual Path, Ameerpet, Hyderabad 9618245689, 9704455959 www.visualpath.in
Orient IT, Ameerpet, Hyderabad ,



BigData/Hadoop Training in Bangalore:



Hadoop Training Institutes @ Bengaluru
Contact Numbers
Website/Mail
A.P. Software Solutions, Madivala, Bangalore 080 - 32021688, 9844621700 enquiry@apsoftsolutions.com
Ahana Systems & Solutions, Hanumantha Nagar, Bangalore 080 - 26675891, 9686113314 www.ahana.co.in
ANOVA IT SERVICES, Bengaluru, India 7411120120, 7411121121 www.techinformatic.com
Astrid Solutions, Jaya Nagar, Bangalore 080 - 41100370, 64520268, 8123673106 www.astridsolutions.in
Avaram Technologies, Bengaluru, India 7829415003, 07829415003 www.AvaramTechnologies.com
Business Intelligence Solutions Provider 7552410354, 9977997254 www.bispsolutions.com
Cloud Soft Solutions, Marathahalli, Bangalore 080 - 65659992, 9916224915 www.cloudsoftsol.com
CMS Computer Institute, Rajaji Nagar, Bangalore 080 - 65681433, 080 - 64557745, 8861421441 www.cmsinstitute.co.in
Colossal Software Technologies 080 40636168, 9741913113 www.verticaldivers.com
Dolphin Institute of Computers, Konanakunte, Bangalore 080 - 31904040
e-Care technologies, Bangalore, India 9845642721, 9844752189 www.ecaretechnologies.info
eCloud Solutions, Bangalore, India 9845988299 www.ecloudsol.com
Elegant IT Services, Marathahalli, Bangalore 080 - 41523314, 9886541264 www.elegantitservices.com
Ematic Technologies, Bangalore, India 8971571117,080-65316888 www.ematictechnologies.com
IGROW TECHNOLOGIES, Bangalore, India 080-41724999, 9738025666 www.igrowtechnologies.com
ISAC Software Academy 080 4093 3246 www.isacsoftware.com
KnowledgeWorks IT Consulting Pvt. Ltd. 80 26630622, 22459941, 9886221314 www.knowledgeworks.co.in
Lorven IT Pvt. Ltd., BTM Layout 1st Stage, Bangalore 9611205205 www.lorvenit.net
MeritForge, Bengaluru, India 080 2529 7815, 9611778865 www.meritforge.com
NovelVista IT Learning Solutions 8411050011, 09422540480 www.novelvista.com
QSIT, Bengaluru, India 080 4081 8888 www.qsitglobal.com
RGS IT Solutions, Jaya Nagar, Bangalore 080 - 41307782, 9663914467 suresh9135@gmail.com
RK Software Solutions, Madivala, Bangalore 080 - 32401030, 9740609586 rksoftwaresolutions11@gmail.com
Rrootshell Technologies Pvt. Ltd., BTM 2nd Stage, Bangalore 080 - 32972711, 080 - 26789818, 8105845978 www.rrootshell.com
Sancentre IT Solutions Pvt Ltd, Bengaluru, India 9611836161 www.sancentre.com
SAPTEQ GLOBAL CONSULTING SERVICES, Bengaluru, India 8711046553, 9535160010 WWW.SAPTEQ.COM
Shine Softwares, Bengaluru, India 9739231171, 9739231181 www.shinesoftwares.com
Simplilearn Solutions Pvt. Ltd., Bengaluru, India 080 6547 6230 www.simplilearn.com
Sure Step Solutions Pvt Ltd, Bengaluru, India 080-26721156 www.surestep.co.in
Techbricks software training center, Bengaluru, India 8065467957, 9243022443
Trendwise Analytics, White Field, Bangalore 080 - 40949600, 9845815003, 9886768879 www.trendwiseanalytics.com
United Global Soft-Online Training Institute 080-66444122, 8099902123 www.unitedglobalsoft.com
Varnaaz Academy, Banashankari 2nd Stage, Bangalore 9845562620 www.varnaaz.com
cloudwick technologies, Bangalore xxxxx www.cloudwick.com


Related Hadoop Articles:  Apache Hadoop    Hadoop Commands


April 15, 2018

Hadoop Certifications

Hadoop Certifications Cloudera/Hortonworks/IBM


Cloudera - Cloudera University

CCAH (Administrator) Exams
Cloudera Certified Administrator for Apache Hadoop (CCA-410)
(or)
Cloudera Certified Administrator for Apache Hadoop CDH3 (CCA-332)      -- Not available
Cloudera Certified Administrator for Apache Hadoop CDH4 Upgrade Exam (CCA-470) -- Aplicable only if u have cleared CCA-332

CCDH (Developer) Exams

Cloudera Certified Developer for Apache Hadoop (CCD-410)
(or)
Cloudera Certified Developer for Apache Hadoop CDH3 (CCD-333) -- Not available
Cloudera Certified Developer for Apache Hadoop CDH4 Upgrade Exam (CCD-470) -- Aplicable only if u have cleared CCD-333

Test Name: Cloudera Certified Administrator for Apache Hadoop CDH4 (CCA-410)

Number of Questions: 60
Time Limit: 90 minutes
Passing Score: 70%
Languages: English, Japanese
Price: USD $295, AUD285, EUR225, GBP185, JPY25,500

Test Name: Cloudera Certified Developer for Apache Hadoop CDH4 (CCD-410)

Number of Questions: 60
Time Limit: 90 minutes
Passing Score: 67%
Languages: English, Japanese
Price: USD$295, AUD285, EUR225, GBP185, JPY25,500


HortonWorks - Hortonworks University

Hortonworks Certified Apache Hadoop Administrator (HCAHA)   - HWX-0011
Hortonworks Certified Apache Hadoop Developer (HCAHD) - HWX-0012
   Cost: $150.00
   Passing score: 75%


IBM - BigData University

BD001EN - Hadoop Fundamentals I
BD005EN - Hadoop & Amazon Cloud
BD006EN - Hadoop & IBM Cloud

Related Hadoop Articles:    Hadoop Training     Hadoop Versions


April 19, 2017

Hadoop Distributions


Below are the companies offering commercial implementations and/or providing support for Apache Hadoop, which is the base for all the below.


  • Cloudera offers CDH (Cloudera's Distribution including Apache Hadoop) and Cloudera Enterprise.
  • Hortonworks (formed by Yahoo and Benchmark Capital), whose focus is on making Hadoop more robust and easier to install, manage and use for enterprise users. Hortonworks provides Hortonworks Data Platform (HDP).
  • MapR Technologies offers distributed filesystem and MapReduce engine, the MapR Distribution for Apache Hadoop.
  • Oracle announced the Big Data Appliance, which integrates Cloudera's Distribution Including Apache Hadoop (CDH).
  • IBM offers InfoSphere BigInsights based on Hadoop in both a basic and enterprise edition.
  • Greenplum, A Division of EMC, offers Hadoop in Community and Enterprise editions.
  • Intel - the Intel Distribution for Apache Hadoop is the product includes the Intel Manager for Apache Hadoop for managing a cluster.
  • Amazon Web Services - Amazon offers a version of Apache Hadoop on their EC2 infrastructure, sold as Amazon Elastic MapReduce.
  • VMware - Initiate Open Source project and product to enable easily and efficiently deploy and use Hadoop on virtual infrastructure.
  • Bigtop - project for the development of packaging and tests of the Apache Hadoop ecosystem.
  • DataStax - DataStax provides a product of Hadoop which fully integrates Apache Hadoop with Apache Cassandra and Apache Solr in its DataStax Enterprise platform.
  • Cascading - A popular feature-rich API for defining and executing complex and fault tolerant data processing workflows on a Apache Hadoop cluster. 
  • Mahout - Apache project using Hadoop to build scalable machine learning algorithms like canopy clustering, k-means and many more.
  • Cloudspace - uses Apache Hadoop to scale client and internal projects on Amazon's EC2 and bare metal architectures.
  • Datameer - Datameer Analytics Solution (DAS) is a Hadoop-based solution for big data analytics that includes data source integration, storage, an analytics engine and visualization.
  • Data Mine Lab - Developing solutions based on Hadoop, Mahout, HBase and Amazon Web Services.
  • Debian - A Debian package of Apache Hadoop is available.
  • HStreaming - offers real-time stream processing and continuous advanced analytics built into Hadoop, available as free community edition, enterprise edition, and cloud service.
  • Impetus
  • Karmasphere - Distributes Karmasphere Studio for Hadoop, which allows cross-version development and management of Apache Hadoop jobs.
  • Nutch - Apache Nutch, flexible web search engine software.
  • NGDATA - Makes available Lily Open Source that builds upon Hadoop, HBase and SOLR. Distributes Lily Enterprise.
  • Pentaho – Pentaho provides a complete, end-to-end open-source BI and offers an easy-to-use, graphical ETL tool that is integrated with Apache Hadoop for managing data and coordinating Hadoop related tasks in the broader context of ETL and Business Intelligence workflow.
  • Pervasive Software - Provides Pervasive DataRush, a parallel dataflow framework which improves performance of Apache Hadoop and MapReduce jobs by exploiting fine-grained parallelism on multicore servers.
  • Platform Computing - Provides an Enterprise Class MapReduce solution for Big Data Analytics with high scalability and fault tolerance. Platform MapReduce provides unique scheduling capabilities and its architecture is based on almost two decades of distributed computing research and development.
  • Sematext International - Provides consulting services around Apache Hadoop and Apache HBase, along with large-scale search using Apache Lucene, Apache Solr, and Elastic Search.
  • Talend - Talend Platform for Big Data includes support and management tools for all the major Apache Hadoop distributions. Talend Open Studio for Big Data is an Apache License Eclipse IDE, which provides a set of graphical components for HDFS, HBase, Pig, Sqoop and Hive.
  • Think Big Analytics - Offers expert consulting services specializing in Apache Hadoop, MapReduce and related data processing architectures.
  • Tresata - Financial Industry's first software platform architected from the ground up on Hadoop. Data storage, processing, analytics and visualization all done on Hadoop.
  • WANdisco is a committed member & sponsor of the Apache Software community and has active committers on several projects including Apache Hadoop.

Related Hadoop Articles:  Apache Hadoop Commands   Hadoop Training in India

March 25, 2017

What is Hadoop?


Apache Hadoop is, an open-source software framework, written in Java, by Doug Cutting and Michael J. Cafarella, that supports data-intensive distributed applications, licensed under the Apache v2 license. It supports the running of applications on large clusters of commodity hardware. Hadoop was derived from Google's MapReduce and Google File System (GFS) papers.

The Hadoop framework transparently provides both reliability and data motion to applications. Hadoop implements a computational paradigm named MapReduce, where the application is divided into many small fragments of work, each of which may be executed or re-executed on any node in the cluster. It provides a distributed file system that stores data on the compute nodes, providing very high aggregate bandwidth across the cluster. 

Both map/reduce and the distributed file system are designed so that node failures are automatically handled by the framework. It enables applications to work with thousands of computation-independent computers and petabytes of data. 

The entire Apache Hadoop platform is commonly considered to consist of the Hadoop kernel, MapReduce and Hadoop Distributed File System (HDFS), and number of related projects including Apache Hive, Apache HBase, Apache Pig, Zookeeper etc.

Related Articles:  NoSQL Databases        What is Apache Cassandra