When working with large datasets that exceed single-machine processing capacity, distributed frameworks like Hadoop become essential. This tutorial demonstrates how to implement the classic WordCount example using Hadoop's MapReduce paradigm, connecting remotely from IntelliJ IDEA to a Hadoop cluster.
Environment Setup
Hadoop Cluster Configuration
For this demonstration, we'll use a Linux server environment with Docker for Hadoop deployment. Key configuration steps include:
- Opening required ports (9870 for Web UI, 8020 for remote connections)
- Verifying cluster accessibility through the web interface
Project Implementation
Project Structure
WordCountProject/
├── input/
│ └── sample.txt
├── src/
│ ├── main/
│ │ ├── java/
│ │ │ └── com/
│ │ │ └── example/
│ │ │ ├── WordMapper.java
│ │ │ ├── SumReducer.java
│ │ │ └── WordCountDriver.java
│ │ └── resources/
│ │ ├── core-site.xml
│ │ └── log4j.properties
└── pom.xml
Dependencies
<dependencies>
<dependency>
<groupId>org.apache.hadoop</groupId>
<artifactId>hadoop-client</artifactId>
<version>3.3.4</version>
</dependency>
<dependency>
<groupId>org.apache.hadoop</groupId>
<artifactId>hadoop-mapreduce-client-core</artifactId>
<version>3.3.4</version>
</dependency>
</dependencies>
Mapper Implementation
public class WordMapper extends Mapper<LongWritable, Text, Text, IntWritable> {
private final static IntWritable ONE = new IntWritable(1);
private Text wordToken = new Text();
@Override
protected void map(LongWritable key, Text value, Context context)
throws IOException, InterruptedException {
String[] words = value.toString().split("\\s+");
for (String word : words) {
wordToken.set(word);
context.write(wordToken, ONE);
}
}
}
Reducer Implementation
public class SumReducer extends Reducer<Text, IntWritable, Text, IntWritable> {
private IntWritable result = new IntWritable();
@Override
protected void reduce(Text key, Iterable<IntWritable> values, Context context)
throws IOException, InterruptedException {
int sum = 0;
for (IntWritable val : values) {
sum += val.get();
}
result.set(sum);
context.write(key, result);
}
}
Driver Class
public class WordCountDriver {
public static void main(String[] args) throws Exception {
Configuration conf = new Configuration();
Job job = Job.getInstance(conf, "word count");
job.setJarByClass(WordCountDriver.class);
job.setMapperClass(WordMapper.class);
job.setCombinerClass(SumReducer.class);
job.setReducerClass(SumReducer.class);
job.setOutputKeyClass(Text.class);
job.setOutputValueClass(IntWritable.class);
FileInputFormat.addInputPath(job, new Path(args[0]));
FileOutputFormat.setOutputPath(job, new Path(args[1]));
System.exit(job.waitForCompletion(true) ? 0 : 1);
}
}
Troubleshooting Common Issues
Pemrission Errors
When encountering permission denied errors, adjust HDFS permissions:
hdfs dfs -chmod 777 /target_directory
Windows-Specific Depenedncies
For Windows development, ensure hadoop.dll is available and loaded:
static {
System.load("C:\\hadoop\\bin\\hadoop.dll");
}
Execution and Results
After resolving configuration issues, the job will process input files and generate sorted word frequency counts in the output directory.