Advanced Elasticsearch Search Techniques

Term-Based vs Full-Text Search

Term-Level Queries (Typically Used Without Scoring for Performance)

Term queries perform exact matches and are case-sensitvie depending on the field's analyzer. For example, if a text field uses the standard analyzer (which lowercases terms), querying with uppercase will fail.

Data Setup:

DELETE products
PUT products
{
  "settings": {
    "number_of_shards": 1
  }
}

POST /products/_bulk
{ "index": { "_id": 1 }}
{ "productID": "XHDK-A-1293-#fJ3", "desc": "iPhone" }
{ "index": { "_id": 2 }}
{ "productID": "KDKE-B-9947-#kL5", "desc": "iPad" }
{ "index": { "_id": 3 }}
{ "productID": "JODL-X-1937-#pV7", "desc": "MBP" }

A term query on desc with "iphone" (lowercase) returns document 1 because the standard analyzer lowercased the indexed term:

POST /products/_search
{
  "query": {
    "term": {
      "desc": "iphone"
    }
  }
}

For exact matching on non-analyzed fields, use the .keyword sub-field:

POST /products/_search
{
  "query": {
    "term": {
      "productID.keyword": "XHDK-A-1293-#fJ3"
    }
  }
}

To avoid scoring overhead, wrap term queries in a constant_score filter:

POST /products/_search
{
  "query": {
    "constant_score": {
      "filter": {
        "term": {
          "productID.keyword": "XHDK-A-1293-#fJ3"
        }
      }
    }
  }
}

Structured Search

Structured search applies to exact values like booleans, numbers, dates, and null checks.

Sample Data:

DELETE products
POST /products/_bulk
{ "index": { "_id": 1 }}
{ "price": 10, "available": true, "date": "2018-01-01", "productID": "XHDK-A-1293-#fJ3" }
{ "index": { "_id": 2 }}
{ "price": 20, "available": true, "date": "2019-01-01", "productID": "KDKE-B-9947-#kL5" }
{ "index": { "_id": 3 }}
{ "price": 30, "available": true, "productID": "JODL-X-1937-#pV7" }
{ "index": { "_id": 4 }}
{ "price": 30, "available": false, "productID": "QQPX-R-3956-#aD8" }

Boolean queries:

POST /products/_search
{
  "query": {
    "term": {
      "available": true
    }
  }
}

POST /products/_search
{
  "query": {
    "constant_score": {
      "filter": {
        "term": {
          "available": true
        }
      }
    }
  }
}

Numeric range and existence checks:

GET /products/_search
{
  "query": {
    "constant_score": {
      "filter": {
        "range": {
          "price": {
            "gte": 20,
            "lte": 30
          }
        }
      }
    }
  }
}

POST /products/_search
{
  "query": {
    "constant_score": {
      "filter": {
        "exists": {
          "field": "date"
        }
      }
    }
  }
}

POST /products/_search
{
  "query": {
    "constant_score": {
      "filter": {
        "bool": {
          "must_not": {
            "exists": {
              "field": "date"
            }
          }
        }
      }
    }
  }
}

Relevance Scoring

Relevance is influenced by term frequency (TF), inverse document frequency (IDF), and field length normalization.

Example Setup:

PUT testscore
{
  "settings": { "number_of_shards": 1 },
  "mappings": {
    "properties": {
      "content": { "type": "text" }
    }
  }
}

PUT testscore/_bulk
{ "index": { "_id": 1 }}
{ "content": "we use Elasticsearch to power the search" }
{ "index": { "_id": 2 }}
{ "content": "we like elasticsearch" }
{ "index": { "_id": 3 }}
{ "content": "The scoring of documents is calculated by the scoring formula" }
{ "index": { "_id": 4 }}
{ "content": "you know, for search" }

Querying for "elasticsearch" returns doc 2 before doc 3 due to shorter field length.

Use boosting to demote documents containing certain terms:

POST testscore/_search
{
  "query": {
    "boosting": {
      "positive": {
        "term": { "content": "elasticsearch" }
      },
      "negative": {
        "term": { "content": "like" }
      },
      "negative_boost": 0.2
    }
  }
}

Query Context vs Filter Context

Queries in must or should clauses affect scoring; those in filter or must_not do not.

Bool Query Example:

POST /products/_search
{
  "query": {
    "bool": {
      "must": { "term": { "price": 30 } },
      "filter": { "term": { "available": true } },
      "must_not": { "range": { "price": { "lte": 10 } } },
      "should": [
        { "term": { "productID.keyword": "JODL-X-1937-#pV7" } },
        { "term": { "productID.keyword": "XHDK-A-1293-#fJ3" } }
      ],
      "minimum_should_match": 1
    }
  }
}

Field boosting prioritizes matches in specific fields:

POST blogs/_search
{
  "query": {
    "bool": {
      "should": [
        { "match": { "title": { "query": "apple ipad", "boost": 4 } } },
        { "match": { "content": { "query": "apple ipad", "boost": 1 } } }
      ]
    }
  }
}

Multi-Field Queries

dis_max returns the highest score from any field:

POST blogs/_search
{
  "query": {
    "dis_max": {
      "queries": [
        { "match": { "title": "Quick pets" } },
        { "match": { "body": "Quick pets" } }
      ],
      "tie_breaker": 0.2
    }
  }
}

multi_match simplifies this syntax:

POST blogs/_search
{
  "query": {
    "multi_match": {
      "type": "best_fields",
      "query": "Quick pets",
      "fields": ["title", "body"],
      "tie_breaker": 0.2
    }
  }
}

For multi-language or morphological robustness, use multiple analyzers via multi-fields:

PUT titles
{
  "mappings": {
    "properties": {
      "title": {
        "type": "text",
        "analyzer": "english",
        "fields": {
          "std": { "type": "text", "analyzer": "standard" }
        }
      }
    }
  }
}

GET /titles/_search
{
  "query": {
    "multi_match": {
      "type": "most_fields",
      "query": "barking dogs",
      "fields": ["title", "title.std"]
    }
  }
}

Cross-field search treats multiple fields as one:

POST address/_search
{
  "query": {
    "multi_match": {
      "query": "Poland Street W1V",
      "type": "cross_fields",
      "operator": "and",
      "fields": ["street", "city", "country", "postcode"]
    }
  }
}

Chinese and Multilingual Analysis

Chinese requires specialized tokenizers. Popular options include:

  • IK Analyzer: ik_max_word (aggressive) and ik_smart (conservative)
  • HanLP: Offers NLP-aware segmentation (hanlp_standard, hanlp_nlp, etc.)
  • Pinyin Plugin: Generates phonetic tokens for Chinese names

Example pinyin analyzer setup:

PUT /artists/
{
  "settings": {
    "analysis": {
      "analyzer": {
        "user_name_analyzer": {
          "tokenizer": "whitespace",
          "filter": "pinyin_first_letter_and_full_pinyin_filter"
        }
      },
      "filter": {
        "pinyin_first_letter_and_full_pinyin_filter": {
          "type": "pinyin",
          "keep_first_letter": true,
          "keep_full_pinyin": false,
          "lowercase": true
        }
      }
    }
  }
}

Search Templates and Index Aliases

Search templates decouple query logic from application code:

POST _scripts/tmdb
{
  "script": {
    "lang": "mustache",
    "source": {
      "_source": ["title", "overview"],
      "size": 20,
      "query": {
        "multi_match": {
          "query": "{{q}}",
          "fields": ["title", "overview"]
        }
      }
    }
  }
}

POST tmdb/_search/template
{
  "id": "tmdb",
  "params": { "q": "basketball with cartoon aliens" }
}

Index aliases enable zero-downtime reindexing and filtered views:

POST _aliases
{
  "actions": [
    {
      "add": {
        "index": "movies-2019",
        "alias": "movies-latest"
      }
    },
    {
      "add": {
        "index": "movies-2019",
        "alias": "movies-high-rated",
        "filter": { "range": { "rating": { "gte": 4 } } }
      }
    }
  ]
}

Function Score for Custom Ranking

Incorporate business metrics like popularity into relevance:

POST /blogs/_search
{
  "query": {
    "function_score": {
      "query": {
        "multi_match": {
          "query": "popularity",
          "fields": ["title", "content"]
        }
      },
      "field_value_factor": {
        "field": "votes",
        "modifier": "log1p",
        "factor": 0.1
      },
      "boost_mode": "sum",
      "max_boost": 3
    }
  }
}

Random scoring ensures consistent randomness per session:

POST /blogs/_search
{
  "query": {
    "function_score": {
      "random_score": { "seed": 911119 }
    }
  }
}

Search Suggestions

Term suggester corrects misspellings:

POST /articles/_search
{
  "suggest": {
    "term-suggestion": {
      "text": "lucen rock",
      "term": {
        "suggest_mode": "missing",
        "field": "body"
      }
    }
  }
}

Phrase suggester reconstructs full phrases:

POST /articles/_search
{
  "suggest": {
    "my-suggestion": {
      "text": "lucne and elasticsear rock",
      "phrase": {
        "field": "body",
        "max_errors": 2,
        "confidence": 1
      }
    }
  }
}

Autocomplete and Contextual Suggestions

Use the completion type for fast prefix lookups:

PUT articles
{
  "mappings": {
    "properties": {
      "title_completion": { "type": "completion" }
    }
  }
}

POST articles/_search
{
  "suggest": {
    "article-suggester": {
      "prefix": "elk",
      "completion": { "field": "title_completion" }
    }
  }
}

Contextual suggestions restrict completions by category:

PUT comments/_mapping
{
  "properties": {
    "comment_autocomplete": {
      "type": "completion",
      "contexts": [{ "type": "category", "name": "comment_category" }]
    }
  }
}

POST comments/_search
{
  "suggest": {
    "MY_SUGGESTION": {
      "prefix": "sta",
      "completion": {
        "field": "comment_autocomplete",
        "contexts": { "comment_category": "coffee" }
      }
    }
  }
}

Cross-Cluster Search

Configure remote clusters:

PUT /_cluster/settings
{
  "persistent": {
    "cluster": {
      "remote": {
        "cluster1": { "seeds": ["127.0.0.1:9301"] }
      }
    }
  }
}

Query across clusters:

GET /users,cluster1:users/_search
{
  "query": {
    "range": { "age": { "gte": 20 } }
  }
}

Tags: elasticsearch search full-text-search relevance query-dsl

Posted on Sun, 04 Oct 2026 16:38:45 +0000 by mrkite