Getting Started with Hadoop: Running WordCount Remotely from IntelliJ IDEA

When working with large datasets that exceed single-machine processing capacity, distributed frameworks like Hadoop become essential. This tutorial demonstrates how to implement the classic WordCount example using Hadoop's MapReduce paradigm, connecting remotely from IntelliJ IDEA to a Hadoop cluster.

Environment Setup

Hadoop Cluster Configuration

For this demonstration, we'll use a Linux server environment with Docker for Hadoop deployment. Key configuration steps include:

  • Opening required ports (9870 for Web UI, 8020 for remote connections)
  • Verifying cluster accessibility through the web interface

Project Implementation

Project Structure


WordCountProject/
├── input/
│   └── sample.txt
├── src/
│   ├── main/
│   │   ├── java/
│   │   │   └── com/
│   │   │       └── example/
│   │   │           ├── WordMapper.java
│   │   │           ├── SumReducer.java
│   │   │           └── WordCountDriver.java
│   │   └── resources/
│   │       ├── core-site.xml
│   │       └── log4j.properties
└── pom.xml

Dependencies


<dependencies>
  <dependency>
    <groupId>org.apache.hadoop</groupId>
    <artifactId>hadoop-client</artifactId>
    <version>3.3.4</version>
  </dependency>
  <dependency>
    <groupId>org.apache.hadoop</groupId>
    <artifactId>hadoop-mapreduce-client-core</artifactId>
    <version>3.3.4</version>
  </dependency>
</dependencies>

Mapper Implementation


public class WordMapper extends Mapper<LongWritable, Text, Text, IntWritable> {
  private final static IntWritable ONE = new IntWritable(1);
  private Text wordToken = new Text();

  @Override
  protected void map(LongWritable key, Text value, Context context) 
      throws IOException, InterruptedException {
    String[] words = value.toString().split("\\s+");
    for (String word : words) {
      wordToken.set(word);
      context.write(wordToken, ONE);
    }
  }
}

Reducer Implementation


public class SumReducer extends Reducer<Text, IntWritable, Text, IntWritable> {
  private IntWritable result = new IntWritable();

  @Override
  protected void reduce(Text key, Iterable<IntWritable> values, Context context) 
      throws IOException, InterruptedException {
    int sum = 0;
    for (IntWritable val : values) {
      sum += val.get();
    }
    result.set(sum);
    context.write(key, result);
  }
}

Driver Class


public class WordCountDriver {
  public static void main(String[] args) throws Exception {
    Configuration conf = new Configuration();
    Job job = Job.getInstance(conf, "word count");
    
    job.setJarByClass(WordCountDriver.class);
    job.setMapperClass(WordMapper.class);
    job.setCombinerClass(SumReducer.class);
    job.setReducerClass(SumReducer.class);
    
    job.setOutputKeyClass(Text.class);
    job.setOutputValueClass(IntWritable.class);
    
    FileInputFormat.addInputPath(job, new Path(args[0]));
    FileOutputFormat.setOutputPath(job, new Path(args[1]));
    
    System.exit(job.waitForCompletion(true) ? 0 : 1);
  }
}

Troubleshooting Common Issues

Pemrission Errors

When encountering permission denied errors, adjust HDFS permissions:

hdfs dfs -chmod 777 /target_directory

Windows-Specific Depenedncies

For Windows development, ensure hadoop.dll is available and loaded:


static {
  System.load("C:\\hadoop\\bin\\hadoop.dll");
}

Execution and Results

After resolving configuration issues, the job will process input files and generate sorted word frequency counts in the output directory.

Tags: Hadoop mapreduce intellij-idea distributed-computing HDFS

Posted on Mon, 28 Sep 2026 16:29:41 +0000 by Sergeant