Elasticsearch is a distributed search engine built on Apache Lucene, designed for scalable full-text search capabilities. It provides RESTful APIs for data indexing and querying, supporting real-time analytics across structured and unstructured data.
Installation Requirements
Before deploying Elasticsearch, ensure Java 11+ is installed. On Ubuntu systems:
sudo apt update
sudo apt install default-jre
Configure JAVA_HOME enviroment variable:
sudo update-alternatives --config java
echo "JAVA_HOME=/usr/lib/jvm/java-11-openjdk-amd64" | sudo tee -a /etc/environment
source /etc/environment
Elasticsearch Cluster Setup
Download and extract the Elasticsearch package:
wget https://artifacts.elastic.co/downloads/elasticsearch/elasticsearch-5.6.15.tar.gz
tar -xzf elasticsearch-5.6.15.tar.gz
Start the cluster in backgruond mode:
cd elasticsearch-5.6.15
bin/elasticsearch -d
Python Integration with Elasticsearch
Install the official Python client:
pip3 install elasticsearch
Core Operations Implementation
Index Management
from elasticsearch import Elasticsearch
client = Elasticsearch(hosts=['http://localhost:9200'])
# Create index with error handling
client.indices.create(index='news', ignore=400)
# Delete index safely
client.indices.delete(index='news', ignore=[400, 404])
Data Manipulation
# Insert document
data = {
'headline': 'Iraq Situation Analysis',
'source': 'https://example.com/iraq-report',
'timestamp': '2023-04-05'
}
result = client.index(index='news', document=data)
print(f"Document ID: {result['_id']}")
# Update document
update_data = {
'headline': 'Iraq Situation Update',
'timestamp': '2023-04-06'
}
result = client.update(index='news', id=result['_id'], body={'doc': update_data})
Query Execution
from elasticsearch import QueryStringQuery
# Full-text search with Chinese analyzer
search_query = {
"query": {
"match": {
"headline": {
"query": "中国 领事馆",
"analyzer": "ik_max_word"
}
}
}
}
results = client.search(index='news', body=search_query)
for hit in results['hits']['hits']:
print(f"Score: {hit['_score']}, Content: {hit['_source']}")
Chinese Text Processing
Install IK Analyzer plugin for Chinese text processing:
bin/elasticsearch-plugin install https://github.com/medcl/elasticsearch-analysis-ik/releases/download/v5.6.15/elasticsearch-analysis-ik-5.6.15.zip
Configure mapping for text analysis:
mapping = {
"properties": {
"headline": {
"type": "text",
"analyzer": "ik_max_word",
"search_analyzer": "ik_max_word"
}
}
}
client.indices.put_mapping(index='news', body=mapping)