Term-Based vs Full-Text Search
Term-Level Queries (Typically Used Without Scoring for Performance)
Term queries perform exact matches and are case-sensitvie depending on the field's analyzer. For example, if a text field uses the standard analyzer (which lowercases terms), querying with uppercase will fail.
Data Setup:
DELETE products
PUT products
{
"settings": {
"number_of_shards": 1
}
}
POST /products/_bulk
{ "index": { "_id": 1 }}
{ "productID": "XHDK-A-1293-#fJ3", "desc": "iPhone" }
{ "index": { "_id": 2 }}
{ "productID": "KDKE-B-9947-#kL5", "desc": "iPad" }
{ "index": { "_id": 3 }}
{ "productID": "JODL-X-1937-#pV7", "desc": "MBP" }
A term query on desc with "iphone" (lowercase) returns document 1 because the standard analyzer lowercased the indexed term:
POST /products/_search
{
"query": {
"term": {
"desc": "iphone"
}
}
}
For exact matching on non-analyzed fields, use the .keyword sub-field:
POST /products/_search
{
"query": {
"term": {
"productID.keyword": "XHDK-A-1293-#fJ3"
}
}
}
To avoid scoring overhead, wrap term queries in a constant_score filter:
POST /products/_search
{
"query": {
"constant_score": {
"filter": {
"term": {
"productID.keyword": "XHDK-A-1293-#fJ3"
}
}
}
}
}
Structured Search
Structured search applies to exact values like booleans, numbers, dates, and null checks.
Sample Data:
DELETE products
POST /products/_bulk
{ "index": { "_id": 1 }}
{ "price": 10, "available": true, "date": "2018-01-01", "productID": "XHDK-A-1293-#fJ3" }
{ "index": { "_id": 2 }}
{ "price": 20, "available": true, "date": "2019-01-01", "productID": "KDKE-B-9947-#kL5" }
{ "index": { "_id": 3 }}
{ "price": 30, "available": true, "productID": "JODL-X-1937-#pV7" }
{ "index": { "_id": 4 }}
{ "price": 30, "available": false, "productID": "QQPX-R-3956-#aD8" }
Boolean queries:
POST /products/_search
{
"query": {
"term": {
"available": true
}
}
}
POST /products/_search
{
"query": {
"constant_score": {
"filter": {
"term": {
"available": true
}
}
}
}
}
Numeric range and existence checks:
GET /products/_search
{
"query": {
"constant_score": {
"filter": {
"range": {
"price": {
"gte": 20,
"lte": 30
}
}
}
}
}
}
POST /products/_search
{
"query": {
"constant_score": {
"filter": {
"exists": {
"field": "date"
}
}
}
}
}
POST /products/_search
{
"query": {
"constant_score": {
"filter": {
"bool": {
"must_not": {
"exists": {
"field": "date"
}
}
}
}
}
}
}
Relevance Scoring
Relevance is influenced by term frequency (TF), inverse document frequency (IDF), and field length normalization.
Example Setup:
PUT testscore
{
"settings": { "number_of_shards": 1 },
"mappings": {
"properties": {
"content": { "type": "text" }
}
}
}
PUT testscore/_bulk
{ "index": { "_id": 1 }}
{ "content": "we use Elasticsearch to power the search" }
{ "index": { "_id": 2 }}
{ "content": "we like elasticsearch" }
{ "index": { "_id": 3 }}
{ "content": "The scoring of documents is calculated by the scoring formula" }
{ "index": { "_id": 4 }}
{ "content": "you know, for search" }
Querying for "elasticsearch" returns doc 2 before doc 3 due to shorter field length.
Use boosting to demote documents containing certain terms:
POST testscore/_search
{
"query": {
"boosting": {
"positive": {
"term": { "content": "elasticsearch" }
},
"negative": {
"term": { "content": "like" }
},
"negative_boost": 0.2
}
}
}
Query Context vs Filter Context
Queries in must or should clauses affect scoring; those in filter or must_not do not.
Bool Query Example:
POST /products/_search
{
"query": {
"bool": {
"must": { "term": { "price": 30 } },
"filter": { "term": { "available": true } },
"must_not": { "range": { "price": { "lte": 10 } } },
"should": [
{ "term": { "productID.keyword": "JODL-X-1937-#pV7" } },
{ "term": { "productID.keyword": "XHDK-A-1293-#fJ3" } }
],
"minimum_should_match": 1
}
}
}
Field boosting prioritizes matches in specific fields:
POST blogs/_search
{
"query": {
"bool": {
"should": [
{ "match": { "title": { "query": "apple ipad", "boost": 4 } } },
{ "match": { "content": { "query": "apple ipad", "boost": 1 } } }
]
}
}
}
Multi-Field Queries
dis_max returns the highest score from any field:
POST blogs/_search
{
"query": {
"dis_max": {
"queries": [
{ "match": { "title": "Quick pets" } },
{ "match": { "body": "Quick pets" } }
],
"tie_breaker": 0.2
}
}
}
multi_match simplifies this syntax:
POST blogs/_search
{
"query": {
"multi_match": {
"type": "best_fields",
"query": "Quick pets",
"fields": ["title", "body"],
"tie_breaker": 0.2
}
}
}
For multi-language or morphological robustness, use multiple analyzers via multi-fields:
PUT titles
{
"mappings": {
"properties": {
"title": {
"type": "text",
"analyzer": "english",
"fields": {
"std": { "type": "text", "analyzer": "standard" }
}
}
}
}
}
GET /titles/_search
{
"query": {
"multi_match": {
"type": "most_fields",
"query": "barking dogs",
"fields": ["title", "title.std"]
}
}
}
Cross-field search treats multiple fields as one:
POST address/_search
{
"query": {
"multi_match": {
"query": "Poland Street W1V",
"type": "cross_fields",
"operator": "and",
"fields": ["street", "city", "country", "postcode"]
}
}
}
Chinese and Multilingual Analysis
Chinese requires specialized tokenizers. Popular options include:
- IK Analyzer:
ik_max_word(aggressive) andik_smart(conservative) - HanLP: Offers NLP-aware segmentation (
hanlp_standard,hanlp_nlp, etc.) - Pinyin Plugin: Generates phonetic tokens for Chinese names
Example pinyin analyzer setup:
PUT /artists/
{
"settings": {
"analysis": {
"analyzer": {
"user_name_analyzer": {
"tokenizer": "whitespace",
"filter": "pinyin_first_letter_and_full_pinyin_filter"
}
},
"filter": {
"pinyin_first_letter_and_full_pinyin_filter": {
"type": "pinyin",
"keep_first_letter": true,
"keep_full_pinyin": false,
"lowercase": true
}
}
}
}
}
Search Templates and Index Aliases
Search templates decouple query logic from application code:
POST _scripts/tmdb
{
"script": {
"lang": "mustache",
"source": {
"_source": ["title", "overview"],
"size": 20,
"query": {
"multi_match": {
"query": "{{q}}",
"fields": ["title", "overview"]
}
}
}
}
}
POST tmdb/_search/template
{
"id": "tmdb",
"params": { "q": "basketball with cartoon aliens" }
}
Index aliases enable zero-downtime reindexing and filtered views:
POST _aliases
{
"actions": [
{
"add": {
"index": "movies-2019",
"alias": "movies-latest"
}
},
{
"add": {
"index": "movies-2019",
"alias": "movies-high-rated",
"filter": { "range": { "rating": { "gte": 4 } } }
}
}
]
}
Function Score for Custom Ranking
Incorporate business metrics like popularity into relevance:
POST /blogs/_search
{
"query": {
"function_score": {
"query": {
"multi_match": {
"query": "popularity",
"fields": ["title", "content"]
}
},
"field_value_factor": {
"field": "votes",
"modifier": "log1p",
"factor": 0.1
},
"boost_mode": "sum",
"max_boost": 3
}
}
}
Random scoring ensures consistent randomness per session:
POST /blogs/_search
{
"query": {
"function_score": {
"random_score": { "seed": 911119 }
}
}
}
Search Suggestions
Term suggester corrects misspellings:
POST /articles/_search
{
"suggest": {
"term-suggestion": {
"text": "lucen rock",
"term": {
"suggest_mode": "missing",
"field": "body"
}
}
}
}
Phrase suggester reconstructs full phrases:
POST /articles/_search
{
"suggest": {
"my-suggestion": {
"text": "lucne and elasticsear rock",
"phrase": {
"field": "body",
"max_errors": 2,
"confidence": 1
}
}
}
}
Autocomplete and Contextual Suggestions
Use the completion type for fast prefix lookups:
PUT articles
{
"mappings": {
"properties": {
"title_completion": { "type": "completion" }
}
}
}
POST articles/_search
{
"suggest": {
"article-suggester": {
"prefix": "elk",
"completion": { "field": "title_completion" }
}
}
}
Contextual suggestions restrict completions by category:
PUT comments/_mapping
{
"properties": {
"comment_autocomplete": {
"type": "completion",
"contexts": [{ "type": "category", "name": "comment_category" }]
}
}
}
POST comments/_search
{
"suggest": {
"MY_SUGGESTION": {
"prefix": "sta",
"completion": {
"field": "comment_autocomplete",
"contexts": { "comment_category": "coffee" }
}
}
}
}
Cross-Cluster Search
Configure remote clusters:
PUT /_cluster/settings
{
"persistent": {
"cluster": {
"remote": {
"cluster1": { "seeds": ["127.0.0.1:9301"] }
}
}
}
}
Query across clusters:
GET /users,cluster1:users/_search
{
"query": {
"range": { "age": { "gte": 20 } }
}
}