System Requirements
This guide demonstrates how to install a big data stack on a personal laptop using VirtualBox with CentOS 7.x (kernel 3.10.0-514.el7.x86_64). RHEL 7.3 is also compatible with this setup.
Minimum Requirements:
-
Memory: 2.5GB (4GB recommended for better performance)
-
Components to be installed:
-
Hadoop 2.8.0
-
Apache Hive 2.1.1
-
Presto Server 0.177
-
MySQL Community Server 5.7.18
-
Oracle JDK 1.8.0_131
Note: In Hadoop 2.8.0, commands now start with 'hadoop' instead of 'hdfs'. For example, 'hdfs dfs -mkdir /test' becomes 'hadoop fs -mkdir /test'. While older documentation may use 'hdfs', it's recommended to use 'hadoop' commands.
Initial System Configuration
Step 1: System Preparation
Note: This is for a test environment only. In production, you should not disable firewall and SELinux.
systemctl stop firewalld
systemctl disable firewalld
setenforce 0
# Edit /etc/selinux config to set SELINUX=disabled permanently
Step 2: User Setup
Create a dedicated user for the big data components:
useradd hadoop
mkdir /home/hadoop
chown hadoop:hadoop /home/hadoop
Step 3: Network Configuration
Edit /etc/hosts:
127.0.0.1 localhost localhost.localdomain localhost4 localhost4.localdomain4
::1 localhost localhost.localdomain localhost6 localhost6.localdomain6
192.168.1.199 bigdata.lzf
Step 4: SSH Configuration
As the hadoop user, set up passwordless SSH:
ssh-keygen -t rsa -P '' -f ~/.ssh/id_rsa
cat ~/.ssh/id_rsa.pub >> ~/.ssh/authorized_keys
chmod 0600 ~/.ssh/authorized_keys
# Verify: ssh bigdata.lzf should not require a password
Java Development Kit Installation
Using the root account, extract JDK to /usr/local/jdk1.8.0_131.
MySQL Setup
Step 1: Installation
Using the root account, install MySQL via RPM packages. Since the system may have MariaDB pre-installed, remove conflicting components first.
systemctl start mysqld
Step 2: Password Configuration
For newer MySQL versions, follow these steps:
# Edit /etc/my.cnf and add: skip-grant-tables
mysql -u root
USE mysql;
update mysql.user set authentication_string=password('your_password') where user='root';
commit;
flush privileges;
set password=password('your_password');
Step 3: Database Setup for Hive
Restart MySQL and create the metastore database:
mysql -u root -p
CREATE DATABASE metastore;
GRANT ALL ON metastore.* TO 'root'@'%' IDENTIFIED BY 'your_password';
GRANT ALL ON metastore.* TO 'root'@'localhost' IDENTIFIED BY 'your_password';
GRANT ALL ON metastore.* TO 'root'@'bigdata.lzf' IDENTIFIED BY 'your_password';
Hadoop Installation
Switch to the hadoop user and extract Hadoop to /home/hadoop/hadoop-2.8.0.
Step 1: Environment Variables
Edit .bash_profile:
export JAVA_HOME=/usr/local/jdk1.8.0_131
export HADOOP_HOME=/home/hadoop/hadoop-2.8.0
export HADOOP_CONF_DIR=/home/hadoop/hadoop-2.8.0/etc/hadoop
export HADOOP_YARN_HOME=$HADOOP_HOME
export HADOOP_MAPRED_HOME=$HADOOP_HOME
export PATH=$PATH:$HOME/.local/bin:$HOME/bin:$JAVA_HOME/bin:$HADOOP_HOME/bin
source .bash_profile
Step 2: Directory Structure
mkdir -p /home/hadoop/hadoop/tmp
mkdir -p /home/hadoop/hadoop/hdfs
mkdir -p /home/hadoop/hadoop/hdfs/data
mkdir -p /home/hadoop/hadoop/hdfs/name
# Copy template files if needed
cd /home/hadoop/hadoop-2.8.0/etc/hadoop
cp mapred-site.xml.template mapred-site.xml
Step 3: Environment Configuraton
Update these files to set JAVA_HOME:
- hadoop-env.sh
- yarn-env.sh
- mapred-env.sh
Add to each file:
export JAVA_HOME=/usr/local/jdk1.8.0_131
Step 4: Core Configuration (core-site.xml)
<configuration>
<property>
<name>fs.defaultFS</name>
<value>hdfs://bigdata.lzf:9001</value>
<description>HDFS URI, filesystem://namenode标识:port, default 9000</description>
</property>
<property>
<name>hadoop.tmp.dir</name>
<value>/home/hadoop/data_hadoop/tmp</value>
<description>Hadoop temporary folder on namenode</description>
</property>
<property>
<name>ipc.client.connect.max.retries</name>
<value>100</value>
<description>Default is 10, now set to 100</description>
</property>
<property>
<name>ipc.client.connect.retry.interval</name>
<value>10000</value>
<description>Connection interval 1 second, default 0.1 second</description>
</property>
<property>
<name>hadoop.proxyuser.hadoop.hosts</name>
<value>*</value>
</property>
<property>
<name>hadoop.proxyuser.hadoop.groups</name>
<value>*</value>
</property>
</configuration>
Step 5: HDFS Configuration (hdfs-site.xml)
<configuration>
<property>
<name>dfs.namenode.name.dir</name>
<value>/home/hadoop/data_hadoop/hdfs/name</value>
<description>Storage location for HDFS namespace metadata on namenode</description>
</property>
<property>
<name>dfs.datanode.data.dir</name>
<value>/home/hadoop/data_hadoop/hdfs/data</value>
<description>Physical storage location for data blocks on datanode</description>
</property>
<property>
<name>dfs.replication</name>
<value>1</value>
<description>Number of replicas, default 3, should be less than datanode count</description>
</property>
<property>
<name>dfs.namenode.rpc-address</name>
<value>bigdata.lzf:9001</value>
<description>RPC address handling all client requests</description>
</property>
</configuration>
Step 6: MapReduce Configuration (mapred-site.xml)
<configuration>
<property>
<name>mapreduce.framework.name</name>
<value>yarn</value>
</property>
</configuration>
Step 7: YARN Configuration (yarn-site.xml)
<configuration>
<property>
<name>yarn.nodemanager.aux-services</name>
<value>mapreduce_shuffle</value>
</property>
<property>
<name>yarn.resourcemanager.webapp.address</name>
<value>bigdata.lzf:8099</value>
<description>Resource manager for cluster, accessible via browser</description>
</property>
<property>
<name>yarn.nodemanager.webapp.address</name>
<value>bigdata.lzf:8042</value>
<description>Node manager, accessible via browser</description>
</property>
</configuration>
Step 8: Slaves Configuration
Edit slaves file and add:
bigdata.lzf
Step 9: Logging Configuration
Edit log4j.properties:
log4j.logger.org.apache.hadoop.util.NativeCodeLoader=DEBUG
Step 10: Startup and Verification
$HADOOP_HOME/bin/hdfs namenode -format # Run once only
$HADOOP_HOME/sbin/start-dfs.sh
$HADOOP_HOME/sbin/start-yarn.sh
# Verify with jps - should show:
# NameNode
# DataNode
# SecondaryNameNode
# ResourceManager
# NodeManager
# Check ports with telnet: 8042, 8099, 9001
# Access web UIs via:
# http://bigdata.lzf:8099
# Test HDFS commands:
hdfs dfs -mkdir -p /tmp/input
hdfs dfs -mkdir -p /tmp/output
hdfs dfs -ls /tmp
hdfs dfs -lsr /tmp
Hive Installation
Presto requires Hive metastore, so we need to install Hive first.
Step 1: Extraction
Extract to /home/hadoop/apache-hive-2.1.1-bin.
Step 2: HDFS Directory Setup
hdfs dfs -mkdir -p /warehouse
hdfs dfs -mkdir -p /tmp/hive
hdfs dfs -chmod 773 /warehouse
hdfs dfs -chmod 773 /tmp/hive
Step 3: Environment Variables
Edit .bash_profile:
export HIVE_HOME=/home/hadoop/apache-hive-2.1.1-bin
export HIVE_CONF_DIR=/home/hadoop/apache-hive-2.1.1-bin/conf
export HCAT_LOG_DIR=/home/hadoop/apache-hive-2.1.1-bin/hcatalog/sbin/logs
export PATH=$PATH:$HOME/.local/bin:$HOME/bin:$JAVA_HOME/bin:$HADOOP_HOME/bin:$HIVE_HOME/bin
source .bash_profile
Step 4: Configuration Files
Edit hive-env.sh:
export CLASSPATH=/home/hadoop/apache-hive-2.1.1-bin/lib/log4j-slf4j-impl-2.4.1.jar
export JAVA_HOME=/usr/local/jdk1.8.0_131
export HADOOP_HOME=/home/hadoop/hadoop-2.8.0
export HIVE_HOME=/home/hadoop/apache-hive-2.1.1-bin
export HIVE_CONF_DIR=/home/hadoop/apache-hive-2.1.1-bin/conf
Copy hive-default.xml to hive-site.xml and modify:
<configuration>
<property>
<name>javax.jdo.option.ConnectionURL</name>
<value>jdbc:mysql://bigdata.lzf:3306/metastore?createDatabaseIfNotExist=true&useSSL=false</value>
<description>JDBC connect string for JDBC metastore</description>
</property>
<property>
<name>javax.jdo.option.ConnectionDriverName</name>
<value>com.mysql.jdbc.Driver</value>
<description>Driver class name for JDBC metastore</description>
</property>
<property>
<name>javax.jdo.option.ConnectionUserName</name>
<value>root</value>
<description>Username for metastore database</description>
</property>
<property>
<name>javax.jdo.option.ConnectionPassword</name>
<value>your_password</value>
<description>Password for metastore database user</description>
</property>
<property>
<name>hive.server2.thrift.port</name>
<value>10000</value>
</property>
<property>
<name>hive.server2.thrift.bind.host</name>
<value>bigdata.lzf</value>
<description>Hostname for HiveServer2</description>
</property>
<property>
<name>hive.server2.authentication</name>
<value>NONE</value>
<description>Authentication type (NONE, LDAP, KERBEROS, CUSTOM)</description>
</property>
<property>
<name>hive.server2.enable.doAs</name>
<value>true</value>
</property>
<property>
<name>datanucleus.schema.autoCreateAll</name>
<value>false</value>
<description>Auto create metastore schema</description>
</property>
<property>
<name>hive.server2.logging.operation.log.location</name>
<value>/home/hadoop/java_tmp/${user.name}/operation_logs</value>
<description>Operation logs location</description>
</property>
<property>
<name>hive.metastore.warehouse.dir</name>
<value>/warehouse</value>
</property>
<property>
<name>hive.exec.scratchdir</name>
<value>/tmp/hive</value>
</property>
<property>
<name>hive.querylog.location</name>
<value>/log</value>
<description>Query logs location</description>
</property>
<property>
<name>hive.metastore.uris</name>
<value>thrift://bigdata.lzf:9083</value>
<description>Thrift URI for remote metastore</description>
</property>
<property>
<name>hive.server2.transport.mode</name>
<value>binary</value>
<description>Transport mode (binary or http)</description>
</property>
</configuration>
Edit hive-log4j2.properties:
property.hive.log.dir = /home/hadoop/data_hive/java_io_temp/${sys:user.name}
Step 5: Initialization
$HIVE_HOME/bin/schematool -dbType mysql -initSchema
Verify by checking MySQL metastore database tables.
Step 6: Verification
# Start services
$HIVE_HOME/bin/hive --service metastore
$HIVE_HOME/bin/hive --service hiveserver2
# Test with telnet
telnet bigdata.lzf 10000
# Access web UI
http://bigdata.lzf:10002
# Test with Hive CLI
hive
# Or with Beeline
beeline -u jdbc:hive2://bigdata.lzf:10000
Presto Installation
Refer to official Presto documentation for detailed setup instructions.
Step 1: Extraction and Environment
Extract to /home/hadoop/presto-serveer-0.177 and update .bash_profile:
export PRESTO_HOME=/home/hadoop/presto-server-0.177
export PATH=$PATH:$HOME/.local/bin:$HOME/bin:$JAVA_HOME/bin:$HADOOP_HOME/bin:$HIVE_HOME/bin:$PRESTO_HOME/bin
source .bash_profile
Step 2: Configuration Files
Create directories:
mkdir -p $PRESTO_HOME/etc/catalog
mkdir -p /home/hadoop/data_presto/data
Create node.properties:
node.environment=prestoquery
node.id=presto-0001
node.data-dir=/home/hadoop/data_presto/data
Create jvm.config:
-server
-Xmx16G
-XX:+UseG1GC
-XX:G1HeapRegionSize=32M
-XX:+UseGCOverheadLimit
-XX:+ExplicitGCInvokesConcurrent
-XX:+HeapDumpOnOutOfMemoryError
-XX:+ExitOnOutOfMemoryError
-DHADOOP_USER_NAME=hadoop
Create config.properties:
# Single-node configuration (coordinator and worker)
coordinator=true
node-scheduler.include-coordinator=true
# Port configuration (avoid conflicts with other services)
http-server.http.port=8080
query.max-memory=5GB
query.max-memory-per-node=1GB
discovery-server.enabled=true
discovery.uri=http://bigdata.lzf:8080
Create log.properties:
# Set to DEBUG, INFO, WARN, or ERROR as needed
com.facebook.presto=DEBUG
Create hive.properties in catalog directory:
connector.name=hive-hadoop2
hive.metastore.uri=thrift://bigdata.lzf:9083
Step 3: Startup and Shutdown
# Start (background)
$PRESTO_HOME/bin/launcher start
# Start (foreground with logs)
$PRESTO_HOME/bin/launcher run
# Stop
$PRESTO_HOME/bin/launcher stop
Step 4: Verification
Check with jps - you should see PrestoServer process.
Access web UI at http://bigdata.lzf:8080.
Step 5: Using Presto CLI
# Download presto-cli JAR and place in $PRESTO_HOME/bin, rename to presto-cli
presto-cli --server bigdata.lzf:8080 --catalog hive --schema default
presto:default> show tables;
presto:default> select * from customer;
Conclusion
Setting up a single-node big data platform requires patience and attention to detail. While many online resources are helpful, always refer to official documentation for the most accurate information. This setup is suitable for development and testing purposes only. For production environments, consider multi-node deployments, proper security measures, and optimized configurations based on your specific workload requirements.