メインコンテンツまでスキップ

Query

ANN 検索に加えて、MilvusZilliz Cloud はクエリによるメタデータフィルタリングもサポートしています。このページでは、Query、Get、および QueryIterators を使用してメタデータフィルタリングを実行する方法を紹介します。

📘注意

collection の作成後に新しいフィールドを追加した場合、これらのフィールドを含むクエリは、値が明示的に設定されていない entity に対して、定義されたデフォルト値または NULL を返します。詳細については、Alter Collection Schema を参照してください。

Overview

Collection にはさまざまな型の scalar フィールドを格納できます。Zilliz Cloud では、1 つ以上の scalar フィールドに基づいて Entities をフィルタリングできます。Zilliz Cloud は 3 種類のクエリ、Query、Get、QueryIterator を提供しています。以下の表では、これら 3 種類のクエリを比較しています。

Get

Query

QueryIterator

適用シナリオ

指定された主キーを持つ entity を検索する場合。

カスタムのフィルタリング条件を満たすすべて、または指定された数の entity を検索する場合

カスタムのフィルタリング条件を満たすすべての entity を、ページネーションされたクエリで検索する場合。

フィルタリング方法

主キーによる

フィルタリング式による。

フィルタリング式による。

必須パラメータ

  • Collection 名

  • 主キー

  • Collection 名

  • フィルタリング式

  • Collection 名

  • フィルタリング式

  • クエリごとに返す entity 数

オプションパラメータ

  • partition 名

  • 出力フィールド

  • partition 名

  • 返す entity 数

  • 出力フィールド

  • partition 名

  • 合計で返す entity 数

  • 出力フィールド

戻り値

指定された collection または partition 内で、指定された主キーを持つ entity を返します。

指定された collection または partition 内で、カスタムのフィルタリング条件を満たすすべて、または指定された数の entity を返します。

指定された collection または partition 内で、カスタムのフィルタリング条件を満たすすべての entity を、ページネーションされたクエリを通じて返します。

メタデータフィルタリングの詳細については、Filtering ExplainedFiltering Explained を参照してください。

Use Get

主キーによって entity を検索する必要がある場合は、Get メソッドを使用できます。以下のコード例では、collection に idvectorcolor という 3 つのフィールドがあることを前提としています。

python
[
{"id": 0, "vector": [0.3580376395471989, -0.6023495712049978, 0.18414012509913835, -0.26286205330961354, 0.9029438446296592], "color": "pink_8682"},
{"id": 1, "vector": [0.19886812562848388, 0.06023560599112088, 0.6976963061752597, 0.2614474506242501, 0.838729485096104], "color": "red_7025"},
{"id": 2, "vector": [0.43742130801983836, -0.5597502546264526, 0.6457887650909682, 0.7894058910881185, 0.20785793220625592], "color": "orange_6781"},
{"id": 3, "vector": [0.3172005263489739, 0.9719044792798428, -0.36981146090600725, -0.4860894583077995, 0.95791889146345], "color": "pink_9298"},
{"id": 4, "vector": [0.4452349528804562, -0.8757026943054742, 0.8220779437047674, 0.46406290649483184, 0.30337481143159106], "color": "red_4794"},
{"id": 5, "vector": [0.985825131989184, -0.8144651566660419, 0.6299267002202009, 0.1206906911183383, -0.1446277761879955], "color": "yellow_4222"},
{"id": 6, "vector": [0.8371977790571115, -0.015764369584852833, -0.31062937026679327, -0.562666951622192, -0.8984947637863987], "color": "red_9392"},
{"id": 7, "vector": [-0.33445148015177995, -0.2567135004164067, 0.8987539745369246, 0.9402995886420709, 0.5378064918413052], "color": "grey_8510"},
{"id": 8, "vector": [0.39524717779832685, 0.4000257286739164, -0.5890507376891594, -0.8650502298996872, -0.6140360785406336], "color": "white_9381"},
{"id": 9, "vector": [0.5718280481994695, 0.24070317428066512, -0.3737913482606834, -0.06726932177492717, -0.6980531615588608], "color": "purple_4976"},
]

以下のように、ID によって entity を取得できます。

python
from pymilvus import MilvusClient

client = MilvusClient(
uri="YOUR_CLUSTER_ENDPOINT",
token="YOUR_CLUSTER_TOKEN"
)

res = client.get(
collection_name="my_collection",
ids=[0, 1, 2],
output_fields=["vector", "color"]
)

print(res)

Use Query

Basic Query

カスタムのフィルタリング条件で entity を検索する必要がある場合は、Query メソッドを使用します。以下のコード例では、idvectorcolor という 3 つのフィールドがあることを前提とし、color の値が red で始まる entity を指定した数だけ返します。

python
from pymilvus import MilvusClient

client = MilvusClient(
uri="YOUR_CLUSTER_ENDPOINT",
token="YOUR_CLUSTER_TOKEN"
)

res = client.query(
collection_name="my_collection",
filter="color like \"red%\"",
output_fields=["vector", "color"],
limit=3
)

Sort Query Results | ONDEMAND

デフォルトでは、Query は順序未指定で結果を返します。order_by パラメータを使用すると、1 つ以上の scalar フィールドで結果をソートできます。order_by を使用する際は、次の点に注意してください。

  • order_bylimit と一緒に使用する必要があります。

  • サポートされるフィールド型: INT8INT16INT32INT64FLOATDOUBLEVARCHAR。vector、JSONARRAY フィールドによるソートはサポートされていません。

  • nullable フィールドでソートする場合、昇順では NULL 値は末尾(NULLS LAST)に、降順では先頭(NULLS FIRST)に配置されます。

Basic Sort

order_by パラメータには "field_name:direction" 形式の文字列のリストを渡します。ここで directionasc(昇順)または desc(降順)のいずれかです。ascdesc は大文字と小文字を区別する点に注意してください。

python
from pymilvus import MilvusClient

client = MilvusClient(
uri="YOUR_CLUSTER_ENDPOINT",
token="YOUR_CLUSTER_TOKEN"
)

# Sort results by id in ascending order
res = client.query(
collection_name="my_collection",
filter="color like \"red%\"",
output_fields=["vector", "color"],
limit=3,
order_by=["id:asc"],
)

複数フィールドでのソート

複数のフィールドで同時にソートできます。結果はまずリスト内の最初のフィールドで並べ替えられます。2 つの行がそのフィールドで同じ値を持つ場合は、2 番目のフィールドで順序が決まり、以下同様に続きます。

python
# Sort by rating descending, then by price ascending for ties
res = client.query(
collection_name="my_collection",
filter="",
output_fields=["color", "rating", "price"],
limit=10,
order_by=["rating:desc", "price:asc"],
)

ソート付きページネーション

order_bylimit および offset と組み合わせて使用すると、ソート済み結果をページネーションできます。たとえば、価格順に並べた商品リストを複数ページにわたって表示する場合、各ページには重複や抜け漏れなく、正しい価格順で次のアイテム群が表示されます。

python
# Page 1
page1 = client.query(
collection_name="my_collection",
filter="color like \"red%\"",
output_fields=["color", "price"],
limit=5,
offset=0,
order_by=["price:asc"],
)

# Page 2
page2 = client.query(
collection_name="my_collection",
filter="color like \"red%\"",
output_fields=["color", "price"],
limit=5,
offset=5,
order_by=["price:asc"],
)

集約 Query の結果 | ONDEMAND

1 つまたは複数の scalar フィールドで query 結果をグループ化し、グループごとに集約を計算できます。サポートされている集約演算子は countminmaxsumavg です。

group_by_fields を使用する際は、以下に注意してください。

  • group_by_fields でサポートされるフィールド型は、INT8INT16INT32INT64VARCHARTIMESTAMPTZ です。FLOATDOUBLE、vector、JSONARRAY フィールドでグループ化するとエラーが返されます。

  • sumavg は数値専用です。VARCHAR フィールドに適用するとエラーが返されます。

集約を有効にするには、query()group_by_fields を渡し、output_fields に集約式(count(*)count(<field>)min(<field>)max(<field>)sum(<field>)avg(<field>))を追加します。

以下の例では、color フィールドで entity をグループ化し、各 color グループ内の entity 数を返します。

python
from pymilvus import MilvusClient

client = MilvusClient(
uri="YOUR_CLUSTER_ENDPOINT",
token="YOUR_CLUSTER_TOKEN"
)

res = client.query(
collection_name="my_collection",
filter="",
group_by_fields=["color"],
output_fields=["color", "count(*)"],
)

# [{'color': 'red', 'count(*)': 10},
# {'color': 'orange', 'count(*)': 10},
# {'color': 'yellow', 'count(*)': 10},
# {'color': 'green', 'count(*)': 10},
# {'color': 'blue', 'count(*)': 10}]

1 回の呼び出しで複数の集約式を要求することもできます。次の例では、color でグループ化し、各グループの行数、平均価格、最大 rating を返します。

python
res = client.query(
collection_name="my_collection",
filter="",
group_by_fields=["color"],
output_fields=["color", "count(*)", "avg(price)", "max(rating)"],
)

# [{'color': 'red', 'count(*)': 10, 'avg(price)': 65.22, 'max(rating)': 5},
# {'color': 'orange', 'count(*)': 10, 'avg(price)': 48.67, 'max(rating)': 5},
# {'color': 'yellow', 'count(*)': 10, 'avg(price)': 64.15, 'max(rating)': 3},
# {'color': 'green', 'count(*)': 10, 'avg(price)': 58.28, 'max(rating)': 5},
# {'color': 'blue', 'count(*)': 10, 'avg(price)': 50.20, 'max(rating)': 5}]

複合グループを計算するには、複数のフィールドを group_by_fields に渡します。以下の例では、(color, rating) でグループ化し、各バケット内の価格範囲を計算します。

python
res = client.query(
collection_name="my_collection",
filter="",
group_by_fields=["color", "rating"],
output_fields=["color", "rating", "min(price)", "max(price)"],
)

# [{'color': 'red', 'rating': 5, 'min(price)': 34.51, 'max(price)': 70.90},
# {'color': 'orange', 'rating': 2, 'min(price)': 12.39, 'max(price)': 81.99},
# {'color': 'yellow', 'rating': 2, 'min(price)': 22.62, 'max(price)': 88.24},
# {'color': 'green', 'rating': 1, 'min(price)': 18.35, 'max(price)': 59.53},
# {'color': 'blue', 'rating': 4, 'min(price)': 21.23, 'max(price)': 82.45},
# ...]

group_by_fieldslimit と組み合わせて、返されるグループ数に上限を設けることもできます。これは、あるフィールドのカーディナリティが高く、バケットのサンプルだけが必要な場合に便利です。

python
res = client.query(
collection_name="my_collection",
filter="",
group_by_fields=["color"],
output_fields=["color", "avg(price)", "count(*)"],
limit=5,
)

# [{'color': 'red', 'avg(price)': 65.22, 'count(*)': 10},
# {'color': 'orange', 'avg(price)': 48.67, 'count(*)': 10},
# {'color': 'yellow', 'avg(price)': 64.15, 'count(*)': 10},
# {'color': 'green', 'avg(price)': 58.28, 'count(*)': 10},
# {'color': 'blue', 'avg(price)': 50.20, 'count(*)': 10}]

QueryIterator を使用する

ページネーションされた query を通じてカスタムのフィルタリング条件で entity を見つける必要がある場合は、QueryIterator を作成し、その next() メソッドを使ってすべての entity を反復処理し、フィルタリング条件を満たすものを探します。以下のコード例では、idvectorcolor という名前の 3 つのフィールドが存在することを前提とし、color の値が red で始まるすべての entity を返します。

python
iterator = client.query_iterator(
"my_collection",
batch_size=10,
filter="color like \"red%\"",
output_fields=["color"]
)

results = []

while True:
result = iterator.next()
if not result:
iterator.close()
break

print(result)
results += result

パーティション内のクエリ

Get、Query、または QueryIterator リクエストにパーティション名を含めることで、1 つまたは複数のパーティション内でクエリを実行することもできます。以下のコード例では、コレクション内に PartitionA という名前のパーティションがあることを前提としています。

python
res = client.get(
collection_name="my_collection",
partitionNames=["partitionA"],
ids=[10, 11, 12],
output_fields=["vector", "color"]
)

res = client.query(
collection_name="my_collection",
partitionNames=["partitionA"],
filter="color like \"red%\"",
output_fields=["vector", "color"],
limit=3
)

# Use QueryIterator
iterator = client.query_iterator(
"my_collection",
partition_names=["partitionA"],
batch_size=10,
filter="color like \"red%\"",
output_fields=["color"]
)

results = []
while True:
result = iterator.next()
if not result:
iterator.close()
break

print(result)
results += result

Query によるランダムサンプリング

データ探索や開発テストのためにコレクションから代表的なデータのサブセットを抽出するには、RANDOM_SAMPLE(sampling_factor) 式を使用します。ここで sampling_factor は 0 から 1 の間の浮動小数点数で、サンプリングするデータの割合を表します。

📘注意

詳細な使用方法、高度な例、ベストプラクティスについては、ランダムサンプリング を参照してください。

python
# Sample 1% of the entire collection
res = client.query(
collection_name="my_collection",
filter="RANDOM_SAMPLE(0.01)",
output_fields=["vector", "color"]
)

print(f"Sampled {len(res)} entities from collection")

# Combine with other filters - first filter, then sample
res = client.query(
collection_name="my_collection",
filter="color like \"red%\" AND RANDOM_SAMPLE(0.005)",
output_fields=["vector", "color"],
limit=10
)

print(f"Found {len(res)} red items in sample")

クエリに対して一時的にタイムゾーンを設定する

コレクションに TIMESTAMPTZ フィールドがある場合、クエリ呼び出しで timezone パラメータを設定することで、単一の操作に対してデータベースまたはコレクションのデフォルトタイムゾーンを一時的に上書きできます。これにより、その操作中に TIMESTAMPTZ 値がどのように表示・比較されるかを制御できます。

timezone の値は、有効な IANA time zone identifier(例: Asia/ShanghaiAmerica/Chicago、または UTC)である必要があります。TIMESTAMPTZ フィールドの使用方法の詳細については、TIMESTAMPTZ Field を参照してください。

以下の例は、クエリ操作に対して一時的にタイムゾーンを設定する方法を示しています。

python
# Query data and display the tsz field converted to "America/Havana"
results = client.query(
"my_collection",
filter="id <= 10",
output_fields=["id", "tsz", "vec"],
limit=2,
timezone="America/Havana",
)
Ctrl I