Skip to main content

Full-Text Search

  • This article is about running a full-text search with a dynamic query.
    To learn how to run a full-text search using a static-index, see full-text search with index.

  • Use the Search() method to query for documents that contain specified term/s within the text of the specified document field/s.

  • When running a full-text search with a dynamic query, the auto-index created by the server breaks down the text of the searched document field using the default search analyzer.
    All generated terms are lower-cased, so the search is case-insensitive.

  • Gain additional control over term tokenization by running a full-text search using a static-index, where the used analyzer is configurable.

  • A boost value can be set for each search to prioritize results. Learn more in boost search results.

  • User experience can be enhanced by requesting text fragments that highlight the searched terms in the results. Learn more in highlight search results.

Search for single term

List<Employee> employees = session
// Make a dynamic query on Employees collection
.Query<Employee>()
// * Call 'Search' to make a Full-Text search
// * Search is case-insensitive
// * Look for documents containing the term 'University' within their 'Notes' field
.Search(x => x.Notes, "University")
.ToList();

// Results will contain Employee documents that have
// any case variation of the term 'university' in their 'Notes' field.
  • Executing the above query will generate the auto-index Auto/Employees/BySearch(Notes).

  • This auto-index will contain the following two index-fields:

    • Notes
      Contains terms with the original text from the indexed document field 'Notes'.
      Text is lower-cased and Not tokenized.

    • search(Notes)
      Contains lower-cased terms that were tokenized from the 'Notes' field by the default search analyzer (RavenStandardAnalyzer). Calling the Search() method targets these terms to find matching documents.

Search for multiple terms

  • You can search for multiple terms in the same field in a single search method.

  • By default, the logical operator between these terms is 'OR'.

  • This behavior can be modified. See section Search operators.

Pass terms in a string:

List<Employee> employees = session
.Query<Employee>()
// * Pass multiple terms in a single string, separated by spaces.
// * Look for documents containing either 'University' OR 'Sales' OR 'Japanese'
// within their 'Notes' field
.Search(x => x.Notes, "University Sales Japanese")
.ToList();

// * Results will contain Employee documents that have at least one of the specified terms.
// * Search is case-insensitive.

Pass terms in a list:

List<Employee> employees = session
.Query<Employee>()
// * Pass terms in IEnumerable<string>.
// * Look for documents containing either 'University' OR 'Sales' OR 'Japanese'
// within their 'Notes' field
.Search(x => x.Notes, new[] { "University", "Sales", "Japanese" })
.ToList();

// * Results will contain Employee documents that have at least one of the specified terms.
// * Search is case-insensitive.

Search with analyzer-generated phrases

A value passed to Search() is processed in two stages:

  1. Search terms:
    Search() parses the value into search terms.
    Unquoted spaces or tabs separate terms, while double quotes keep the enclosed text together as one term.

  2. Tokens:
    The query-time analyzer processes each search term and produces zero, one, or multiple tokens.

When the analyzer produces multiple tokens from one search term, RavenDB searches those tokens as a phrase.
This means that a document matches only when the corresponding tokens in the searched field are adjacent
and occur in the same order as the analyzer-generated query tokens.


How the input form affects matching

With RavenDB's default analyzer configuration, compare the following three cases:

Separate search terms: matched independently

List<Employee> employees = session
.Query<Employee>()
// Pass two separate search terms
.Search(x => x.Title, "sales representative")
.ToList();

In both the C# and RQL tabs above, the outer quotation marks delimit the string literal.
These quotation marks are not part of the value processed by Search().
RavenDB therefore parses sales representative as the two terms sales and representative.

The search operator combines these terms (OR by default), so they do not need to be adjacent or appear in the same order.

One search term, multiple analyzer tokens: matched as a phrase

List<Employee> employees = session
.Query<Employee>()
// Pass one term that the analyzer splits
.Search(x => x.Title, "sales%representative")
.ToList();

Because % is not whitespace, Search() initially parses sales%representative as one term. RavenStandardAnalyzer then splits that term into the tokens sales and representative.
RavenDB searches those analyzer-generated tokens as a phrase.

  • The % character is not a wildcard and has no special meaning to Search().
    It is simply an example of punctuation that RavenStandardAnalyzer treats as a token boundary.
    Other analyzers can produce different tokens from the same input.

  • If the Search() value contains other search terms, the search operator combines those terms with the generated phrase. It does not change the required order or adjacency of the tokens within the phrase.

Explicitly quoted text: matched as a phrase

List<Employee> employees = session
.Query<Employee>()
// Pass an explicitly quoted phrase
.Search(x => x.Title, "\"sales representative\"")
.ToList();

The escaped quotation marks are included in the value sent to RavenDB.
They instruct Search() to keep sales representative together as one search term.
The analyzer produces the tokens sales and representative, which RavenDB searches as a phrase.

The analyzer-generated phrase and the explicitly quoted phrase use the same phrase-matching rules.
Both match Sales Representative. Neither matches Representative Sales, where the order is reversed, or Sales Area Representative, where the tokens are not adjacent.

Analyzer used by a dynamic query

  • A dynamic full-text query uses the default search analyzer configured for the auto-index.
    This is RavenStandardAnalyzer unless Indexing.Analyzers.Search.Default is configured with a different analyzer.

  • To choose an analyzer for a specific index-field, use a static index.


Wildcard terms

Search terms that use the supported * wildcard forms are handled as prefix, suffix, or contains queries;
they are not converted into phrase queries. See Using wildcards.


When analysis produces no tokens

If the analyzer produces no tokens from a search term - for example, because it removes all of them as stop words -
RavenDB skips that term.


Search-engine compatibility

Phrase matching for analyzer-generated tokens is supported by both Corax and Lucene.

Existing Corax indexes:

  • Corax indexes store the term-position data required for phrase matching only if they were created or reset in RavenDB 6.0 or later.
  • A Corax index - including an auto-index - created before this support was available retains the legacy behavior, where analyzer-generated tokens are matched independently.
    Reset or recreate the index to enable phrase matching.

Search in multiple fields

  • You can search for terms in different fields by making multiple search calls.

  • By default, the logical operator between consecutive search methods is 'OR'.

  • This behavior can be modified. See section Search options.

List<Employee> employees = session
.Query<Employee>()
// * Look for documents containing:
// 'French' in their 'Notes' field OR 'President' in their 'Title' field
.Search(x => x.Notes, "French")
.Search(x => x.Title, "President")
.ToList();

// * Results will contain Employee documents that have
// at least one of the specified fields with the specified terms.
// * Search is case-insensitive.

Search in complex object

  • You can search for terms within a complex object.

  • Any nested text field within the object is searchable.

List<Company> companies = session
.Query<Company>()
// * Look for documents that contain:
// the term 'USA' OR 'London' in any field within the complex 'Address' object
.Search(x => x.Address, "USA London")
.ToList();

// * Results will contain Company documents that are located either in 'USA' OR in 'London'.
// * Search is case-insensitive.

Search operators

  • By default, the logical operator between multiple terms within the same field in a search call is OR.

  • This can be modified using the @operator parameter as follows:

AND:

List<Employee> employees = session
.Query<Employee>()
// * Pass `@operator` with 'SearchOperator.And'
.Search(x => x.Notes, "College German", @operator: SearchOperator.And)
.ToList();

// * Results will contain Employee documents that have BOTH 'College' AND 'German'
// in their 'Notes' field.
// * Search is case-insensitive.

OR:

List<Employee> employees = session
.Query<Employee>()
// * Pass `@operator` with 'SearchOperator.Or' (or don't pass this param at all)
.Search(x => x.Notes, "College German", @operator: SearchOperator.Or)
.ToList();

// * Results will contain Employee documents that have EITHER 'College' OR 'German'
// in their 'Notes' field.
// * Search is case-insensitive.

Search options

  • Search options allow to:

    • Negate a search criteria.
    • Specify the logical operator used between consecutive search calls.
  • When using Query: use the options parameter.
    When using DocumentQuery: follow the specific syntax in each example below.

Negate search:

List<Company> companies = session
.Query<Company>()
// Pass 'options' with 'SearchOptions.Not'
.Search(x => x.Address, "USA", options: SearchOptions.Not)
.ToList();

// * Results will contain Company documents are NOT located in 'USA'
// * Search is case-insensitive

Default behavior between search calls:

  • By default, the logical operator between consecutive search methods is OR.
List<Company> companies = session
.Query<Company>()
.Where(x => x.Contact.Title == "Owner")
// Operator AND will be used with previous 'Where' predicate
.Search(x => x.Address.Country, "France")
// Operator OR will be used between the two 'Search' calls by default
.Search(x => x.Name, "Markets")
.ToList();

// * Results will contain Company documents that have:
// ('Owner' as the 'Contact.Title')
// AND
// (are located in 'France' OR have 'Markets' in their 'Name' field)
//
// * Search is case-insensitive

AND search calls:

List<Employee> employees = session
.Query<Employee>()
.Search(x => x.Notes, "French")
// * Pass 'options' with 'SearchOptions.And' to the second 'Search'
// * Operator AND will be used with previous the 'Search' call
.Search(x => x.Title, "Manager", options: SearchOptions.And)
.ToList();

// * Results will contain Employee documents that have:
// ('French' in their 'Notes' field)
// AND
// ('Manager' in their 'Title' field)
//
// * Search is case-insensitive

Use options as bit flags:

List<Employee> employees = session
.Query<Employee>()
.Search(x => x.Notes, "French")
// Pass logical operators as flags in the 'options' parameter
.Search(x => x.Title, "Manager", options: SearchOptions.Not | SearchOptions.And)
.ToList();

// * Results will contain Employee documents that have:
// ('French' in their 'Notes' field)
// AND
// (do NOT have 'Manager' in their 'Title' field)
//
// * Search is case-insensitive

Using wildcards

  • Wildcards can be used to replace:

    • Prefix of a searched term
    • Postfix of a searched term
    • Both prefix & postfix
  • Note:

    • Searching with a wildcard as the prefix of the term (e.g. *text) is less recommended,
      as it will cause the server to perform a full index scan.

    • Instead, consider using a static-index that indexes the field in reverse order
      and then query with a wildcard as the postfix, which is much faster.

List<Employee> employees = session
.Query<Employee>()
// Use '*' to replace one or more characters
.Search(x => x.Notes, "art*")
.Search(x => x.Notes, "*logy")
.Search(x => x.Notes, "*mark*")
.ToList();

// Results will contain Employee documents that have in their 'Notes' field:
// (terms that start with 'art') OR
// (terms that end with 'logy') OR
// (terms that have the text 'mark' in the middle)
//
// * Search is case-insensitive

Syntax

// Query overloads:
// ================

IRavenQueryable<T> Search<T>(
Expression<Func<T, object>> fieldSelector,
string searchTerms,
decimal boost,
SearchOptions options,
SearchOperator @operator);

IRavenQueryable<T> Search<T>(
Expression<Func<T, object>> fieldSelector,
IEnumerable<string> searchTerms,
decimal boost,
SearchOptions options,
SearchOperator @operator);

// DocumentQuery overloads:
// ========================

IDocumentQueryBase<T> Search(
string fieldName,
string searchTerms,
SearchOperator @operator);

IDocumentQueryBase<T> Search<TValue>(
Expression<Func<T, TValue>> propertySelector,
string searchTerms,
SearchOperator @operator);
ParameterTypeDescription
fieldSelectorExpression<Func<TResult>>Points to the field in which you search.
fieldNamestringName of the field in which you search.
searchTermsstring / IEnumerable<string>A string containing the term or terms (separated by spaces) to search for.
Or, can pass an array (or other IEnumerable) with terms to search for.
boostdecimalThe boost value.
Learn more in boost search results.
Default is 1.0
optionsSearchOptions enumLogical operator to use between consecutive Search methods.
Can be Or, And, Not, or Guess.
Default is SearchOptions.Guess
@operatorSearchOperator enumLogical operator to use between multiple terms in the same Search method.
Can be Or or And.
Default is SearchOperator.Or

In this article