RTUComputer ScienceYr 2022 · Sem 82022

Q21Big Data Analytics

Question

10 marks

Explain Hive architecture and its data types.

Answer

Hive provides an SQL-like interface over Hadoop, utilizing components like the Metastore, Driver, and Execution Engine.

Apache Hive allows querying and managing large datasets residing in distributed storage. Its architecture consists of: 1. Clients: Drivers like JDBC, ODBC, and Thrift allow different applications to connect to Hive. 2. Hive Services: Includes the CLI, Web UI, and the Hive Server for client connections. 3. Driver & Compiler: The Driver receives the HiveQL query. The Compiler parses the query, performs semantic analysis, and generates an execution plan with the help of the Metastore. 4. Metastore: A central repository that stores the metadata (schema, tables, partitions) of the data stored in HDFS. It typically uses an RDBMS like MySQL. 5. Execution Engine: Executes the plan generated by the compiler in the proper order by converting it into MapReduce, Tez, or Spark jobs. Hive Data Types include Primitive types (TINYINT, INT, BIGINT, FLOAT, DOUBLE, STRING, BOOLEAN, TIMESTAMP) and Complex types (ARRAY, MAP, STRUCT, UNION).

Hive Architecture Diagram
Hive Architecture Diagram
Back to Paper