メインコンテンツまでスキップ

Query

ANN 検索に加えて、MilvusZilliz Cloud はクエリによるメタデータのフィルタリングもサポートしています。このページでは、Query、Get、および QueryIterators を使用してメタデータのフィルタリングを行う方法を紹介します。

📘注意

collection の作成後に新しいフィールドを追加した場合、これらのフィールドを含むクエリでは、明示的に値が設定されていない entity に対して定義済みのデフォルト値または NULL が返されます。詳細については、Alter Collection Schema を参照してください。

Overview

Collection にはさまざまな種類の scalar フィールドを保存できます。Zilliz Cloud では、1 つ以上の scalar フィールドに基づいて Entities をフィルタリングできます。Zilliz Cloud では、Query、Get、QueryIterator の 3 種類のクエリを提供しています。以下の表では、これら 3 つのクエリタイプを比較しています。

Get

Query

QueryIterator

適用シナリオ

指定した主キーを持つ entity を見つける場合。

カスタムフィルタリング条件を満たすすべて、または指定数の entity を見つける場合

カスタムフィルタリング条件を満たすすべての entity をページネーション付きクエリで見つける場合。

フィルタリング方法

主キーによる

フィルタリング式による。

フィルタリング式による。

必須パラメータ

  • Collection 名

  • 主キー

  • Collection 名

  • フィルタリング式

  • Collection 名

  • フィルタリング式

  • クエリごとに返す entity の数

任意パラメータ

  • partition 名

  • 出力フィールド

  • partition 名

  • 返す entity の数

  • 出力フィールド

  • partition 名

  • 合計で返す entity の数

  • 出力フィールド

返り値

指定した collection または partition 内で、指定した主キーを持つ entity を返します。

指定した collection または partition 内で、カスタムフィルタリング条件を満たすすべて、または指定数の entity を返します。

指定した collection または partition 内で、カスタムフィルタリング条件を満たすすべての entity をページネーション付きクエリで返します。

メタデータのフィルタリングの詳細については、Filtering ExplainedFiltering Explained を参照してください。

Use Get

主キーによって entity を見つける必要がある場合は、Get メソッドを使用できます。以下のコード例では、collection に idvectorcolor という名前の 3 つのフィールドがあることを前提としています。

python
[
{"id": 0, "vector": [0.3580376395471989, -0.6023495712049978, 0.18414012509913835, -0.26286205330961354, 0.9029438446296592], "color": "pink_8682"},
{"id": 1, "vector": [0.19886812562848388, 0.06023560599112088, 0.6976963061752597, 0.2614474506242501, 0.838729485096104], "color": "red_7025"},
{"id": 2, "vector": [0.43742130801983836, -0.5597502546264526, 0.6457887650909682, 0.7894058910881185, 0.20785793220625592], "color": "orange_6781"},
{"id": 3, "vector": [0.3172005263489739, 0.9719044792798428, -0.36981146090600725, -0.4860894583077995, 0.95791889146345], "color": "pink_9298"},
{"id": 4, "vector": [0.4452349528804562, -0.8757026943054742, 0.8220779437047674, 0.46406290649483184, 0.30337481143159106], "color": "red_4794"},
{"id": 5, "vector": [0.985825131989184, -0.8144651566660419, 0.6299267002202009, 0.1206906911183383, -0.1446277761879955], "color": "yellow_4222"},
{"id": 6, "vector": [0.8371977790571115, -0.015764369584852833, -0.31062937026679327, -0.562666951622192, -0.8984947637863987], "color": "red_9392"},
{"id": 7, "vector": [-0.33445148015177995, -0.2567135004164067, 0.8987539745369246, 0.9402995886420709, 0.5378064918413052], "color": "grey_8510"},
{"id": 8, "vector": [0.39524717779832685, 0.4000257286739164, -0.5890507376891594, -0.8650502298996872, -0.6140360785406336], "color": "white_9381"},
{"id": 9, "vector": [0.5718280481994695, 0.24070317428066512, -0.3737913482606834, -0.06726932177492717, -0.6980531615588608], "color": "purple_4976"},
]

以下のように、ID によって entity を取得できます。

python
from pymilvus import MilvusClient

client = MilvusClient(
uri="YOUR_CLUSTER_ENDPOINT",
token="YOUR_CLUSTER_TOKEN"
)

res = client.get(
collection_name="my_collection",
ids=[0, 1, 2],
output_fields=["vector", "color"]
)

print(res)

Use Query

Basic Query

カスタムフィルタリング条件によって entity を見つける必要がある場合は、Query メソッドを使用します。以下のコード例では、idvectorcolor という名前の 3 つのフィールドがあることを前提とし、color の値が red で始まる指定数の entity を返します。

python
from pymilvus import MilvusClient

client = MilvusClient(
uri="YOUR_CLUSTER_ENDPOINT",
token="YOUR_CLUSTER_TOKEN"
)

res = client.query(
collection_name="my_collection",
filter="color like \"red%\"",
output_fields=["vector", "color"],
limit=3
)

Sort Query Results | ONDEMAND

デフォルトでは、Query は順序が指定されていない結果を返します。order_by パラメータを使用すると、1 つ以上の scalar フィールドで結果を並べ替えることができます。order_by を使用する際は、以下に注意してください。

  • order_bylimit と一緒に使用する必要があります。

  • サポートされるフィールド型: INT8INT16INT32INT64FLOATDOUBLEVARCHAR。vector、JSONARRAY フィールドでの並べ替えはサポートされていません。

  • nullable フィールドで並べ替える場合、昇順では NULL 値は末尾(NULLS LAST)に配置され、降順では先頭(NULLS FIRST)に配置されます。

Basic Sort

order_by パラメータには、"field_name:direction" 形式の文字列のリストを渡します。ここで directionasc(昇順)または desc(降順)のいずれかです。ascdesc は大文字小文字を区別する点に注意してください。

python
from pymilvus import MilvusClient

client = MilvusClient(
uri="YOUR_CLUSTER_ENDPOINT",
token="YOUR_CLUSTER_TOKEN"
)

# Sort results by id in ascending order
res = client.query(
collection_name="my_collection",
filter="color like \"red%\"",
output_fields=["vector", "color"],
limit=3,
order_by=["id:asc"],
)

複数フィールドでのソート

複数のフィールドを同時にソートできます。結果はまずリスト内の最初のフィールドで並べ替えられます。2 つの行がそのフィールドで同じ値を持つ場合は、2 番目のフィールドで順序が決まり、以降も同様です。

python
# Sort by rating descending, then by price ascending for ties
res = client.query(
collection_name="my_collection",
filter="",
output_fields=["color", "rating", "price"],
limit=10,
order_by=["rating:desc", "price:asc"],
)

ソート付きページネーション

order_bylimit および offset と組み合わせることで、ソート済みの結果をページネーションできます。たとえば、価格順にソートされた製品一覧を複数ページにわたって表示する場合、各ページには重複や欠落なく、正しい価格順で次のアイテム群が表示されます。

python
# Page 1
page1 = client.query(
collection_name="my_collection",
filter="color like \"red%\"",
output_fields=["color", "price"],
limit=5,
offset=0,
order_by=["price:asc"],
)

# Page 2
page2 = client.query(
collection_name="my_collection",
filter="color like \"red%\"",
output_fields=["color", "price"],
limit=5,
offset=5,
order_by=["price:asc"],
)

クエリ結果の集計 | ONDEMAND

1 つ以上の scalar フィールドでクエリ結果をグループ化し、グループごとに集計を計算できます。サポートされている集計演算子は countminmaxsumavg です。

group_by_fields を使用する際は、次の点に注意してください。

  • group_by_fields でサポートされるフィールド型: INT8INT16INT32INT64VARCHARTIMESTAMPTZFLOATDOUBLE、vector、JSONARRAY フィールドでグループ化するとエラーになります。

  • sumavg は数値型専用です。VARCHAR フィールドに適用するとエラーになります。

集計を有効にするには、query()group_by_fields を渡し、集計式(count(*)count(<field>)min(<field>)max(<field>)sum(<field>)avg(<field>))を output_fields に追加します。

次の例では、color フィールドで entity をグループ化し、各色グループ内の entity 数を返します。

python
from pymilvus import MilvusClient

client = MilvusClient(
uri="YOUR_CLUSTER_ENDPOINT",
token="YOUR_CLUSTER_TOKEN"
)

res = client.query(
collection_name="my_collection",
filter="",
group_by_fields=["color"],
output_fields=["color", "count(*)"],
)

# [{'color': 'red', 'count(*)': 10},
# {'color': 'orange', 'count(*)': 10},
# {'color': 'yellow', 'count(*)': 10},
# {'color': 'green', 'count(*)': 10},
# {'color': 'blue', 'count(*)': 10}]

1 回の呼び出しで複数の集計式を要求することもできます。次の例では、color でグループ化し、各グループの行数、平均価格、最大評価を返します。

python
res = client.query(
collection_name="my_collection",
filter="",
group_by_fields=["color"],
output_fields=["color", "count(*)", "avg(price)", "max(rating)"],
)

# [{'color': 'red', 'count(*)': 10, 'avg(price)': 65.22, 'max(rating)': 5},
# {'color': 'orange', 'count(*)': 10, 'avg(price)': 48.67, 'max(rating)': 5},
# {'color': 'yellow', 'count(*)': 10, 'avg(price)': 64.15, 'max(rating)': 3},
# {'color': 'green', 'count(*)': 10, 'avg(price)': 58.28, 'max(rating)': 5},
# {'color': 'blue', 'count(*)': 10, 'avg(price)': 50.20, 'max(rating)': 5}]

group_by_fields に複数のフィールドを渡して、複合グループを計算することもできます。次の例では、(color, rating) でグループ化し、各バケット内の価格範囲を計算します。

python
res = client.query(
collection_name="my_collection",
filter="",
group_by_fields=["color", "rating"],
output_fields=["color", "rating", "min(price)", "max(price)"],
)

# [{'color': 'red', 'rating': 5, 'min(price)': 34.51, 'max(price)': 70.90},
# {'color': 'orange', 'rating': 2, 'min(price)': 12.39, 'max(price)': 81.99},
# {'color': 'yellow', 'rating': 2, 'min(price)': 22.62, 'max(price)': 88.24},
# {'color': 'green', 'rating': 1, 'min(price)': 18.35, 'max(price)': 59.53},
# {'color': 'blue', 'rating': 4, 'min(price)': 21.23, 'max(price)': 82.45},
# ...]

group_by_fieldslimit と組み合わせて、返されるグループ数に上限を設けることもできます。これは、あるフィールドのカーディナリティが高く、バケットのサンプルだけが必要な場合に便利です。

python
res = client.query(
collection_name="my_collection",
filter="",
group_by_fields=["color"],
output_fields=["color", "avg(price)", "count(*)"],
limit=5,
)

# [{'color': 'red', 'avg(price)': 65.22, 'count(*)': 10},
# {'color': 'orange', 'avg(price)': 48.67, 'count(*)': 10},
# {'color': 'yellow', 'avg(price)': 64.15, 'count(*)': 10},
# {'color': 'green', 'avg(price)': 58.28, 'count(*)': 10},
# {'color': 'blue', 'avg(price)': 50.20, 'count(*)': 10}]

QueryIterator を使用する

ページネーションされたクエリを通じてカスタムのフィルタリング条件で entity を見つける必要がある場合は、QueryIterator を作成し、その next() メソッドを使用してすべての entity を反復処理し、フィルタリング条件を満たすものを見つけます。以下のコード例では、idvectorcolor という 3 つのフィールドが存在することを前提としており、color の値が red で始まるすべての entity を返します。

python
iterator = client.query_iterator(
"my_collection",
batch_size=10,
filter="color like \"red%\"",
output_fields=["color"]
)

results = []

while True:
result = iterator.next()
if not result:
iterator.close()
break

print(result)
results += result

Partitions 内でのクエリ

Get、Query、または QueryIterator リクエストに partition 名を含めることで、1 つまたは複数の partition 内でクエリを実行することもできます。以下のコード例では、collection 内に PartitionA という名前の partition が存在すると仮定しています。

python
res = client.get(
collection_name="my_collection",
partitionNames=["partitionA"],
ids=[10, 11, 12],
output_fields=["vector", "color"]
)

res = client.query(
collection_name="my_collection",
partitionNames=["partitionA"],
filter="color like \"red%\"",
output_fields=["vector", "color"],
limit=3
)

# QueryIterator を使用
iterator = client.query_iterator(
"my_collection",
partition_names=["partitionA"],
batch_size=10,
filter="color like \"red%\"",
output_fields=["color"]
)

results = []
while True:
result = iterator.next()
if not result:
iterator.close()
break

print(result)
results += result

Query によるランダムサンプリング

データ探索や開発テストのために collection から代表的なデータのサブセットを抽出するには、RANDOM_SAMPLE(sampling_factor) 式を使用します。ここで sampling_factor は 0 から 1 の間の float で、サンプリングするデータの割合を表します。

📘注記

詳細な使い方、高度な例、ベストプラクティスについては、Random Sampling を参照してください。

python
# collection 全体の 1% をサンプリング
res = client.query(
collection_name="my_collection",
filter="RANDOM_SAMPLE(0.01)",
output_fields=["vector", "color"]
)

print(f"Sampled {len(res)} entities from collection")

# 他のフィルタと組み合わせる - まずフィルタし、その後サンプリング
res = client.query(
collection_name="my_collection",
filter="color like \"red%\" AND RANDOM_SAMPLE(0.005)",
output_fields=["vector", "color"],
limit=10
)

print(f"Found {len(res)} red items in sample")

クエリに対して一時的にタイムゾーンを設定する

collection に TIMESTAMPTZ フィールドがある場合、query 呼び出しで timezone パラメータを設定することで、1 回の操作に限ってデータベースまたは collection のデフォルトタイムゾーンを一時的に上書きできます。これにより、その操作中に TIMESTAMPTZ 値がどのように表示および比較されるかを制御できます。

timezone の値は、有効な IANA time zone identifier である必要があります(例: Asia/ShanghaiAmerica/Chicago、または UTC)。TIMESTAMPTZ フィールドの使用方法の詳細については、TIMESTAMPTZ Field を参照してください。

以下の例は、query 操作に対して一時的にタイムゾーンを設定する方法を示しています。

python
# データを query し、tsz フィールドを "America/Havana" に変換して表示
results = client.query(
"my_collection",
filter="id <= 10",
output_fields=["id", "tsz", "vec"],
limit=2,
timezone="America/Havana",
)