Skip to main content

Cncharonly

The cncharonly filter removes tokens that contain any non-Chinese characters. This filter is useful when you want to focus solely on Chinese text, filtering out any tokens that contain other scripts, numbers, or symbols.

Configuration​

The cncharonly filter is built into Zilliz Cloud. To use it, simply specify its name in the filter section within analyzer_params.

python
analyzer_params = {
"tokenizer": "jieba",
"filter": ["cncharonly"],
}

The cncharonly filter operates on the terms generated by the tokenizer, so it must be used in combination with a tokenizer. For a list of tokenizers available in Zilliz Cloud, refer to Jieba and its sibling pages.

After defining analyzer_params, you can apply them to a VARCHAR field when defining a collection schema. This allows Zilliz Cloud to process the text in that field using the specified analyzer for efficient tokenization and filtering. For details, refer to Example use.

Examples​

Before applying the analyzer configuration to your collection schema, verify its behavior using the run_analyzer method.

Analyzer configuration​

python
analyzer_params = {
"tokenizer": "jieba",
"filter": ["cncharonly"],
}

Verification using run_analyzer​

python
from pymilvus import (
MilvusClient,
)

client = MilvusClient(uri="YOUR_CLUSTER_ENDPOINT")

# Sample text to analyze
sample_text = "Milvus 是 LF AI & Data Foundation 下的一个开源项目,以 Apache 2.0 许可发布。"

# Run the jieba tokenizer with the defined configuration
result = client.run_analyzer(sample_text, analyzer_params)
print("Analyzer output:", result)

Expected output​

python
['是', '下的一个开源项目', '以', '许可发布']