メインコンテンツまでスキップ

Voyage AI

このトピックでは、Milvus で Voyage AI embedding functions を構成して使用する方法について説明します。

Model choices

Milvus は、Voyage AI が提供する embedding models をサポートしています。以下は、現在利用可能な embedding models の一覧です。すばやく参照できます。

Model NameDimensionsMax TokensDescription
voyage-4-large1024 (default), 256, 512, 204832,000最高の汎用および多言語検索品質を提供します。4 シリーズで作成されたすべての embeddings は相互互換です。詳細は blog post を参照してください。
voyage-41024 (default), 256, 512, 204832,000汎用および多言語検索品質向けに最適化されています。4 シリーズで作成されたすべての embeddings は相互互換です。詳細は blog post を参照してください。
voyage-4-lite1024 (default), 256, 512, 204832,000レイテンシとコスト向けに最適化されています。4 シリーズで作成されたすべての embeddings は相互互換です。詳細は blog post を参照してください。
voyage-3-large1,024 (default), 256, 512, 2,04832,000最高の汎用および多言語検索品質を提供します。
voyage-31,02432,000汎用および多言語検索品質向けに最適化されています。詳細は blog post を参照してください。
voyage-3-lite51232,000レイテンシとコスト向けに最適化されています。詳細は blog post を参照してください。
voyage-code-31,024 (default), 256, 512, 2,04832,000コード検索向けに最適化されています。詳細は blog post を参照してください。
voyage-finance-21,02432,000finance 検索および RAG 向けに最適化されています。詳細は blog post を参照してください。
voyage-law-21,02416,000legal 検索および RAG 向けに最適化されています。また、すべてのドメインにわたってパフォーマンスが向上しています。詳細は blog post を参照してください。
voyage-code-21,53616,000コード検索向けに最適化されています(代替手段より 17% 優れています) / 以前の世代の code embeddings。詳細は blog post を参照してください。

詳細については、Text embedding models を参照してください。

Before you start

text embedding function を使用する前に、以下の前提条件を満たしていることを確認してください。

  • embedding model を選択する

    使用する embedding model を決定してください。この選択により、embedding の動作と出力形式が決まります。詳細は Choose an embedding model を参照してください。

  • Voyage AI と連携し、integration ID を取得する

    Voyage AI の model provider integration を作成し、そこから提供される embedding models を使用する前に integration ID を取得する必要があります。詳細は Integrate with Model Providers を参照してください。

  • 互換性のある collection schema を設計する

    collection schema には、以下を含めるよう計画してください。

    • 生の入力テキスト用の text field (VARCHAR)

    • 選択した embedding model に一致するデータ型と dimension を持つ dense vector field

  • 挿入時および検索時に raw text を扱う準備をする

    text embedding function を有効にすると、生のテキストを直接挿入およびクエリできます。embeddings はシステムによって自動的に生成されます。

Step 1: Create a collection with a text embedding function

Define schema fields

embedding function を使用するには、特定の schema を持つ collection を作成します。この schema には、少なくとも次の 3 つの必須 field を含める必要があります。

  • collection 内の各 entity を一意に識別する primary field。

  • embedding 対象の raw data を保存する VARCHAR field。

  • text embedding function が VARCHAR field に対して生成する dense vector embeddings を保存するために予約された vector field。

次の例では、テキストデータを保存するための 1 つの VARCHAR field "document" と、text embedding function によって生成される dense embeddings を保存するための 1 つの vector field "dense" を持つ schema を定義しています。vector dimension (dim) は、選択した embedding model の出力に一致するように設定してください。

python
from pymilvus import MilvusClient, DataType, Function, FunctionType

# Initialize Milvus client
client = MilvusClient(
uri="YOUR_CLUSTER_ENDPOINT",
token="YOUR_CLUSTER_TOKEN"
)

# Create a new schema for the collection
schema = client.create_schema()

# Add primary field "id"
schema.add_field("id", DataType.INT64, is_primary=True, auto_id=False)

# Add scalar field "document" for storing textual data
schema.add_field("document", DataType.VARCHAR, max_length=9000)

# Add vector field "dense" for storing embeddings.
# IMPORTANT: Set dim to match the exact output dimension of the embedding model.
schema.add_field("dense", DataType.FLOAT_VECTOR, dim=1024)

Define the text embedding function

text embedding function は、VARCHAR field に保存された raw data を自動的に embeddings に変換し、明示的に定義された vector field に保存します。

以下の例では、scalar field "document" を embeddings に変換し、その結果の vectors を先ほど定義した "dense" vector field に保存する Function module (voya) を追加しています。

embedding function を定義したら、それを collection schema に追加します。これにより、Milvus は指定された embedding function を使用して、テキストデータから embeddings を処理および保存するようになります。

python
# Define embedding function specifically for embedding model provider
text_embedding_function = Function(
name="voya", # Unique identifier for this embedding function
function_type=FunctionType.TEXTEMBEDDING, # Indicates a text embedding function
input_field_names=["document"], # Scalar field(s) containing text data to embed
output_field_names=["dense"], # Vector field(s) for storing embeddings
params={ # Provider-specific embedding parameters (function-level)
"provider": "voyageai", # Must be set to "voyageai"
"model_name": "voyage-3-large", # Specifies the embedding model to use
"integration_id": "YOUR_INTEGRATION_ID", # Integration ID generated in the Zilliz Cloud console for the selected model provider
# "url": "https://api.voyageai.com/v1/embeddings", # Defaults to the official endpoint if omitted
# "dim": "1024" # Output dimension of the vector embeddings after truncation
# "truncation": "true" # Whether to truncate the input texts to fit within the context length. Defaults to true.
}
)

# Add the configured embedding function to your existing collection schema
schema.add_function(text_embedding_function)

Configure the index

必要な fields と組み込み関数を含む schema を定義したら、collection の index を設定します。このプロセスを簡略化するために、index_type として AUTOINDEX を使用してください。これは、データ構造に基づいて Zilliz Cloud が最適な index type を選択して構成するオプションです。

python
# Prepare index parameters
index_params = client.prepare_index_params()

# Add AUTOINDEX to automatically select optimal indexing method
index_params.add_index(
field_name="dense",
index_type="AUTOINDEX",
metric_type="COSINE"
)

Create the collection

次に、定義した schema と index parameters を使用して collection を作成します。

python
# Create collection named "demo"
client.create_collection(
collection_name='demo',
schema=schema,
index_params=index_params
)

Step 2: Insert data

collection と index の設定が完了したら、raw data を挿入する準備が整います。このプロセスでは、raw text だけを提供すれば十分です。先ほど定義した Function module が、各テキストエントリに対応する sparse vector を自動的に生成します。

python
# Insert sample documents
client.insert('demo', [
{'id': 1, 'document': 'Milvus simplifies semantic search through embeddings.'},
{'id': 2, 'document': 'Vector embeddings convert text into searchable numeric data.'},
{'id': 3, 'document': 'Semantic search helps users find relevant information quickly.'},
])

ステップ 3: テキストによる検索

データの挿入後、生のクエリテキストを使用してセマンティック検索を実行します。Milvus はクエリを自動的に埋め込みベクトルへ変換し、類似度に基づいて関連ドキュメントを取得し、最も一致する上位の結果を返します。

python
# セマンティック検索を実行
results = client.search(
collection_name='demo',
data=['How does Milvus handle semantic search?'], # クエリベクトルではなくテキストクエリを使用
anns_field='dense', # 埋め込みを格納するベクトルフィールドを使用
limit=1,
output_fields=['document'],
)

print(results)
Ctrl I