Skip to main content

Whitespace

The whitespace tokenizer splits text at five ASCII whitespace characters: tab, line feed, form feed, carriage return, and space.

Tokenization rules​

The whitespace tokenizer splits text only at the following five ASCII whitespace characters:

CharacterNameUnicode code point
\tHorizontal tabU+0009
\nLine feedU+000A
\x0C or \fForm feedU+000C
\rCarriage returnU+000D
' 'SpaceU+0020

These separators are discarded, and consecutive separators do not produce empty tokens. Punctuation and other characters remain in the tokens. In particular, vertical tab (\x0B, U+000B), no-break space (\u00A0), and ideographic space (\u3000) do not trigger splitting.

This set follows Rust's char::is_ascii_whitespace(), which excludes other Unicode whitespace characters.

The following examples use {"tokenizer": "whitespace"} with no filters. Inputs and outputs use Python string notation: escape sequences such as \t and \u00A0 represent the actual characters.

InputOutput tokens
"a\tb\nc\x0Cd\re f"["a", "b", "c", "d", "e", "f"]
"Hello,World! foo_bar"["Hello,World!", "foo_bar"]
"a\x0Bb"["a\x0Bb"]
"a\u00A0b"["a\u00A0b"]
"a\u3000b"["a\u3000b"]
"\x20a\x20\x20b\x20"["a", "b"]

Configuration​

To configure an analyzer using the whitespace tokenizer, set tokenizer to whitespace in analyzer_params.

python
analyzer_params = {
"tokenizer": "whitespace",
}

The whitespace tokenizer can work in conjunction with one or more filters. For example, the following code defines an analyzer that uses the whitespace tokenizer and lowercase filter:

python
analyzer_params = {
"tokenizer": "whitespace",
"filter": ["lowercase"]
}

After defining analyzer_params, you can apply them to a VARCHAR field when defining a collection schema. This allows Zilliz Cloud to process the text in that field using the specified analyzer for efficient tokenization and filtering. For details, refer to Example use.

Examples​

Before applying the analyzer configuration to your collection schema, verify its behavior using the run_analyzer method.

Analyzer configuration​

python
analyzer_params = {
"tokenizer": "whitespace",
"filter": ["lowercase"]
}

Verification using run_analyzer​

python
from pymilvus import (
MilvusClient,
)

client = MilvusClient(
uri="YOUR_CLUSTER_ENDPOINT",
token="YOUR_CLUSTER_TOKEN"
)

# Sample text to analyze
sample_text = "The Milvus vector database is built for scale!"

# Run the whitespace analyzer with the defined configuration
result = client.run_analyzer(sample_text, analyzer_params)
print("Whitespace analyzer output:", result)

Expected output​

sql
['the', 'milvus', 'vector', 'database', 'is', 'built', 'for', 'scale!']