Setting Up a Single-Node Big Data Platform with Hadoop, Hive, and Presto

System Requirements

This guide demonstrates how to install a big data stack on a personal laptop using VirtualBox with CentOS 7.x (kernel 3.10.0-514.el7.x86_64). RHEL 7.3 is also compatible with this setup.

Minimum Requirements:

  • Memory: 2.5GB (4GB recommended for better performance)

  • Components to be installed:

  • Hadoop 2.8.0

  • Apache Hive 2.1.1

  • Presto Server 0.177

  • MySQL Community Server 5.7.18

  • Oracle JDK 1.8.0_131

Note: In Hadoop 2.8.0, commands now start with 'hadoop' instead of 'hdfs'. For example, 'hdfs dfs -mkdir /test' becomes 'hadoop fs -mkdir /test'. While older documentation may use 'hdfs', it's recommended to use 'hadoop' commands.

Initial System Configuration

Step 1: System Preparation

Note: This is for a test environment only. In production, you should not disable firewall and SELinux.

systemctl stop firewalld
systemctl disable firewalld
setenforce 0
# Edit /etc/selinux config to set SELINUX=disabled permanently

Step 2: User Setup

Create a dedicated user for the big data components:

useradd hadoop
mkdir /home/hadoop
chown hadoop:hadoop /home/hadoop

Step 3: Network Configuration

Edit /etc/hosts:

127.0.0.1 localhost localhost.localdomain localhost4 localhost4.localdomain4
::1 localhost localhost.localdomain localhost6 localhost6.localdomain6
192.168.1.199 bigdata.lzf

Step 4: SSH Configuration

As the hadoop user, set up passwordless SSH:

ssh-keygen -t rsa -P '' -f ~/.ssh/id_rsa
cat ~/.ssh/id_rsa.pub >> ~/.ssh/authorized_keys
chmod 0600 ~/.ssh/authorized_keys
# Verify: ssh bigdata.lzf should not require a password

Java Development Kit Installation

Using the root account, extract JDK to /usr/local/jdk1.8.0_131.

MySQL Setup

Step 1: Installation

Using the root account, install MySQL via RPM packages. Since the system may have MariaDB pre-installed, remove conflicting components first.

systemctl start mysqld

Step 2: Password Configuration

For newer MySQL versions, follow these steps:

# Edit /etc/my.cnf and add: skip-grant-tables
mysql -u root
USE mysql;
update mysql.user set authentication_string=password('your_password') where user='root';
commit;
flush privileges;
set password=password('your_password');

Step 3: Database Setup for Hive

Restart MySQL and create the metastore database:

mysql -u root -p
CREATE DATABASE metastore;
GRANT ALL ON metastore.* TO 'root'@'%' IDENTIFIED BY 'your_password';
GRANT ALL ON metastore.* TO 'root'@'localhost' IDENTIFIED BY 'your_password';
GRANT ALL ON metastore.* TO 'root'@'bigdata.lzf' IDENTIFIED BY 'your_password';

Hadoop Installation

Switch to the hadoop user and extract Hadoop to /home/hadoop/hadoop-2.8.0.

Step 1: Environment Variables

Edit .bash_profile:

export JAVA_HOME=/usr/local/jdk1.8.0_131
export HADOOP_HOME=/home/hadoop/hadoop-2.8.0
export HADOOP_CONF_DIR=/home/hadoop/hadoop-2.8.0/etc/hadoop
export HADOOP_YARN_HOME=$HADOOP_HOME
export HADOOP_MAPRED_HOME=$HADOOP_HOME
export PATH=$PATH:$HOME/.local/bin:$HOME/bin:$JAVA_HOME/bin:$HADOOP_HOME/bin
source .bash_profile

Step 2: Directory Structure

mkdir -p /home/hadoop/hadoop/tmp
mkdir -p /home/hadoop/hadoop/hdfs
mkdir -p /home/hadoop/hadoop/hdfs/data
mkdir -p /home/hadoop/hadoop/hdfs/name

# Copy template files if needed
cd /home/hadoop/hadoop-2.8.0/etc/hadoop
cp mapred-site.xml.template mapred-site.xml

Step 3: Environment Configuraton

Update these files to set JAVA_HOME:

  • hadoop-env.sh
  • yarn-env.sh
  • mapred-env.sh

Add to each file:

export JAVA_HOME=/usr/local/jdk1.8.0_131

Step 4: Core Configuration (core-site.xml)

<configuration>
  <property>
    <name>fs.defaultFS</name>
    <value>hdfs://bigdata.lzf:9001</value>
    <description>HDFS URI, filesystem://namenode标识:port, default 9000</description>
  </property>
  <property>
    <name>hadoop.tmp.dir</name>
    <value>/home/hadoop/data_hadoop/tmp</value>
    <description>Hadoop temporary folder on namenode</description>
  </property>
  <property>
    <name>ipc.client.connect.max.retries</name>
    <value>100</value>
    <description>Default is 10, now set to 100</description>
  </property>
  <property>
    <name>ipc.client.connect.retry.interval</name>
    <value>10000</value>
    <description>Connection interval 1 second, default 0.1 second</description>
  </property>
  <property>
    <name>hadoop.proxyuser.hadoop.hosts</name>
    <value>*</value>
  </property>
  <property>
    <name>hadoop.proxyuser.hadoop.groups</name>
    <value>*</value>
  </property>
</configuration>

Step 5: HDFS Configuration (hdfs-site.xml)

<configuration>
  <property>
    <name>dfs.namenode.name.dir</name>
    <value>/home/hadoop/data_hadoop/hdfs/name</value>
    <description>Storage location for HDFS namespace metadata on namenode</description>
  </property>
  <property>
    <name>dfs.datanode.data.dir</name>
    <value>/home/hadoop/data_hadoop/hdfs/data</value>
    <description>Physical storage location for data blocks on datanode</description>
  </property>
  <property>
    <name>dfs.replication</name>
    <value>1</value>
    <description>Number of replicas, default 3, should be less than datanode count</description>
  </property>
  <property>
    <name>dfs.namenode.rpc-address</name>
    <value>bigdata.lzf:9001</value>
    <description>RPC address handling all client requests</description>
  </property>
</configuration>

Step 6: MapReduce Configuration (mapred-site.xml)

<configuration>
  <property>
    <name>mapreduce.framework.name</name>
    <value>yarn</value>
  </property>
</configuration>

Step 7: YARN Configuration (yarn-site.xml)

<configuration>
  <property>
    <name>yarn.nodemanager.aux-services</name>
    <value>mapreduce_shuffle</value>
  </property>
  <property>
    <name>yarn.resourcemanager.webapp.address</name>
    <value>bigdata.lzf:8099</value>
    <description>Resource manager for cluster, accessible via browser</description>
  </property>
  <property>
    <name>yarn.nodemanager.webapp.address</name>
    <value>bigdata.lzf:8042</value>
    <description>Node manager, accessible via browser</description>
  </property>
</configuration>

Step 8: Slaves Configuration

Edit slaves file and add:

bigdata.lzf

Step 9: Logging Configuration

Edit log4j.properties:

log4j.logger.org.apache.hadoop.util.NativeCodeLoader=DEBUG

Step 10: Startup and Verification

$HADOOP_HOME/bin/hdfs namenode -format  # Run once only
$HADOOP_HOME/sbin/start-dfs.sh
$HADOOP_HOME/sbin/start-yarn.sh

# Verify with jps - should show:
# NameNode
# DataNode
# SecondaryNameNode
# ResourceManager
# NodeManager

# Check ports with telnet: 8042, 8099, 9001
# Access web UIs via:
# http://bigdata.lzf:8099

# Test HDFS commands:
hdfs dfs -mkdir -p /tmp/input
hdfs dfs -mkdir -p /tmp/output
hdfs dfs -ls /tmp
hdfs dfs -lsr /tmp

Hive Installation

Presto requires Hive metastore, so we need to install Hive first.

Step 1: Extraction

Extract to /home/hadoop/apache-hive-2.1.1-bin.

Step 2: HDFS Directory Setup

hdfs dfs -mkdir -p /warehouse
hdfs dfs -mkdir -p /tmp/hive
hdfs dfs -chmod 773 /warehouse
hdfs dfs -chmod 773 /tmp/hive

Step 3: Environment Variables

Edit .bash_profile:

export HIVE_HOME=/home/hadoop/apache-hive-2.1.1-bin
export HIVE_CONF_DIR=/home/hadoop/apache-hive-2.1.1-bin/conf
export HCAT_LOG_DIR=/home/hadoop/apache-hive-2.1.1-bin/hcatalog/sbin/logs
export PATH=$PATH:$HOME/.local/bin:$HOME/bin:$JAVA_HOME/bin:$HADOOP_HOME/bin:$HIVE_HOME/bin
source .bash_profile

Step 4: Configuration Files

Edit hive-env.sh:

export CLASSPATH=/home/hadoop/apache-hive-2.1.1-bin/lib/log4j-slf4j-impl-2.4.1.jar
export JAVA_HOME=/usr/local/jdk1.8.0_131
export HADOOP_HOME=/home/hadoop/hadoop-2.8.0
export HIVE_HOME=/home/hadoop/apache-hive-2.1.1-bin
export HIVE_CONF_DIR=/home/hadoop/apache-hive-2.1.1-bin/conf

Copy hive-default.xml to hive-site.xml and modify:

<configuration>
  <property>
    <name>javax.jdo.option.ConnectionURL</name>
    <value>jdbc:mysql://bigdata.lzf:3306/metastore?createDatabaseIfNotExist=true&useSSL=false</value>
    <description>JDBC connect string for JDBC metastore</description>
  </property>
  <property>
    <name>javax.jdo.option.ConnectionDriverName</name>
    <value>com.mysql.jdbc.Driver</value>
    <description>Driver class name for JDBC metastore</description>
  </property>
  <property>
    <name>javax.jdo.option.ConnectionUserName</name>
    <value>root</value>
    <description>Username for metastore database</description>
  </property>
  <property>
    <name>javax.jdo.option.ConnectionPassword</name>
    <value>your_password</value>
    <description>Password for metastore database user</description>
  </property>
  <property>
    <name>hive.server2.thrift.port</name>
    <value>10000</value>
  </property>
  <property>
    <name>hive.server2.thrift.bind.host</name>
    <value>bigdata.lzf</value>
    <description>Hostname for HiveServer2</description>
  </property>
  <property>
    <name>hive.server2.authentication</name>
    <value>NONE</value>
    <description>Authentication type (NONE, LDAP, KERBEROS, CUSTOM)</description>
  </property>
  <property>
    <name>hive.server2.enable.doAs</name>
    <value>true</value>
  </property>
  <property>
    <name>datanucleus.schema.autoCreateAll</name>
    <value>false</value>
    <description>Auto create metastore schema</description>
  </property>
  <property>
    <name>hive.server2.logging.operation.log.location</name>
    <value>/home/hadoop/java_tmp/${user.name}/operation_logs</value>
    <description>Operation logs location</description>
  </property>
  <property>
    <name>hive.metastore.warehouse.dir</name>
    <value>/warehouse</value>
  </property>
  <property>
    <name>hive.exec.scratchdir</name>
    <value>/tmp/hive</value>
  </property>
  <property>
    <name>hive.querylog.location</name>
    <value>/log</value>
    <description>Query logs location</description>
  </property>
  <property>
    <name>hive.metastore.uris</name>
    <value>thrift://bigdata.lzf:9083</value>
    <description>Thrift URI for remote metastore</description>
  </property>
  <property>
    <name>hive.server2.transport.mode</name>
    <value>binary</value>
    <description>Transport mode (binary or http)</description>
  </property>
</configuration>

Edit hive-log4j2.properties:

property.hive.log.dir = /home/hadoop/data_hive/java_io_temp/${sys:user.name}

Step 5: Initialization

$HIVE_HOME/bin/schematool -dbType mysql -initSchema

Verify by checking MySQL metastore database tables.

Step 6: Verification

# Start services
$HIVE_HOME/bin/hive --service metastore
$HIVE_HOME/bin/hive --service hiveserver2

# Test with telnet
telnet bigdata.lzf 10000

# Access web UI
http://bigdata.lzf:10002

# Test with Hive CLI
hive

# Or with Beeline
beeline -u jdbc:hive2://bigdata.lzf:10000

Presto Installation

Refer to official Presto documentation for detailed setup instructions.

Step 1: Extraction and Environment

Extract to /home/hadoop/presto-serveer-0.177 and update .bash_profile:

export PRESTO_HOME=/home/hadoop/presto-server-0.177
export PATH=$PATH:$HOME/.local/bin:$HOME/bin:$JAVA_HOME/bin:$HADOOP_HOME/bin:$HIVE_HOME/bin:$PRESTO_HOME/bin
source .bash_profile

Step 2: Configuration Files

Create directories:

mkdir -p $PRESTO_HOME/etc/catalog
mkdir -p /home/hadoop/data_presto/data

Create node.properties:

node.environment=prestoquery
node.id=presto-0001
node.data-dir=/home/hadoop/data_presto/data

Create jvm.config:

-server
-Xmx16G
-XX:+UseG1GC
-XX:G1HeapRegionSize=32M
-XX:+UseGCOverheadLimit
-XX:+ExplicitGCInvokesConcurrent
-XX:+HeapDumpOnOutOfMemoryError
-XX:+ExitOnOutOfMemoryError
-DHADOOP_USER_NAME=hadoop

Create config.properties:

# Single-node configuration (coordinator and worker)
coordinator=true
node-scheduler.include-coordinator=true

# Port configuration (avoid conflicts with other services)
http-server.http.port=8080
query.max-memory=5GB
query.max-memory-per-node=1GB
discovery-server.enabled=true
discovery.uri=http://bigdata.lzf:8080

Create log.properties:

# Set to DEBUG, INFO, WARN, or ERROR as needed
com.facebook.presto=DEBUG

Create hive.properties in catalog directory:

connector.name=hive-hadoop2
hive.metastore.uri=thrift://bigdata.lzf:9083

Step 3: Startup and Shutdown

# Start (background)
$PRESTO_HOME/bin/launcher start

# Start (foreground with logs)
$PRESTO_HOME/bin/launcher run

# Stop
$PRESTO_HOME/bin/launcher stop

Step 4: Verification

Check with jps - you should see PrestoServer process.

Access web UI at http://bigdata.lzf:8080.

Step 5: Using Presto CLI

# Download presto-cli JAR and place in $PRESTO_HOME/bin, rename to presto-cli
presto-cli --server bigdata.lzf:8080 --catalog hive --schema default
presto:default> show tables;
presto:default> select * from customer;

Conclusion

Setting up a single-node big data platform requires patience and attention to detail. While many online resources are helpful, always refer to official documentation for the most accurate information. This setup is suitable for development and testing purposes only. For production environments, consider multi-node deployments, proper security measures, and optimized configurations based on your specific workload requirements.

Tags: Hadoop Hive presto bigdata data-warehouse

Posted on Sat, 03 Oct 2026 16:16:49 +0000 by shamly