Spark Batch Jobs
Spark batch jobs let you run distributed, offline processing on large datasets managed in Zilliz Cloud. Use built-in jobs to deduplicate, cluster, or inspect vector data.
Spark batch jobs are designed for long-running data processing tasks. They are not intended for low-latency online requests or per-record transformations.
Problems you may encounter
As vector datasets grow, you may need more than simple insert or search operations. Repeated ingestion can introduce duplicates, large embedding collections can become difficult to understand, and failed preprocessing pipelines can leave behind questionable records.
Duplicate vector embeddings
Retries, repeated imports, overlapping data sources, and slightly modified versions of the same text or image can all create duplicate vector embeddings across entities. Some entities share the same primary key, while others have different primary keys but nearly identical content, increasing storage, indexing, and downstream processing costs.
Spark batch jobs can identify and remove duplicates across large datasets. Use a primary-key deduplication job to clean up records with the same ID, or a vector-similarity deduplication job to find records with different IDs but nearly identical content.
Unclear embedding distributions
As a collection grows, it becomes difficult to understand the overall embedding distribution. You may not know which patterns dominate the dataset, where long-tail data appears, or whether newly imported data differs from existing records.
A K-Means clustering job groups similar embeddings into coarse clusters and assigns each record a cluster ID. You can use the results to analyze data distribution, compare data sources, create representative samples, and break similarity-based processing into smaller groups.
Hidden anomalies in embedding data
Embedding pipelines can produce records that look valid but do not match the rest of the dataset. Failed preprocessing, incorrect models, parsing noise, corrupted source content, or unexpected data batches can all create unusual embeddings that are difficult to spot through manual inspection.
An anomaly detection job scans the embedding distribution and identifies records that differ significantly from common patterns. You can use the results to find data that may need re-embedding, cleanup, or further review. An anomaly is not necessarily invalid, so flagged records should be reviewed rather than removed automatically.
Choose a job type
The following table lists the recommended job types for your goals.
| Your goal | Recommended job |
|---|---|
| Remove records that share the same primary key | Primary-Key Deduplication |
| Find records with highly similar vector representations | Vector Similarity Deduplication |
| Divide vector data into a predefined number of groups | K-Means Clustering |
| Find records that differ significantly from the main data distribution | Anomaly Detection |
How Spark batch jobs work
A Spark batch job is a long-running distributed, offline processing job that returns a job ID immediately upon receiving a job creation request. You can use the job ID as a handler to monitor its progress and manage its lifecycle.
Before you start
To submit a Spark batch job, ensure that:
-
You have a valid Zilliz Cloud API key with sufficient permissions.
-
Your input data is available in a Zilliz Cloud Volume in a supported format.
- Supported data file formats are
parquet,lance,json, andcsv.
- Supported data file formats are
-
The storage role associated with each External Volume has the permissions required for the job:
-
The input External Volume requires read access to the input location.
-
The output External Volume requires read and write access to the output location, including permission to delete temporary or incomplete objects created during job execution.
-
The input and output External Volumes can use different storage roles.
-
Configure External Volume permissions
Spark batch jobs read input data from and write results to External Volumes. Before submitting a job, make sure the object storage role associated with each Volume has the permissions required for its intended use.
The input and output Volumes can use different storage roles. Grant each role only the permissions required for the corresponding bucket and prefix.
Input Volume
The input Volume requires read access to the input data. If the Integration used by the Volume already provides the required read permissions, no additional permissions are needed.
For Amazon S3, the role requires:
-
s3:GetObjectfor objects under the input prefix. -
s3:ListBucketto list objects under the input prefix. -
s3:GetBucketLocationfor the bucket.
Click here to view the example.
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "ListInput",
"Effect": "Allow",
"Action": [
"s3:ListBucket",
"s3:GetBucketLocation"
],
"Resource": "arn:aws:s3:::<bucket>",
"Condition": {
"StringLike": {
"s3:prefix": [
"<input-prefix>",
"<input-prefix>/*"
]
}
}
},
{
"Sid": "ReadInput",
"Effect": "Allow",
"Action": "s3:GetObject",
"Resource": "arn:aws:s3:::<bucket>/<input-prefix>/*"
}
]
}
Output Volume
The output Volume requires permission to write job results. Spark may also delete temporary objects or objects left by failed or canceled writes.
For Amazon S3, the role requires:
-
s3:GetObject,s3:ListBucket, ands3:GetBucketLocationto access the output location. -
s3:PutObjectto write results. -
s3:DeleteObjectto clean up temporary or incomplete output objects.
Scope s3:PutObject and s3:DeleteObject to the output prefix rather than the entire bucket.
Click here to view the example.
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "ListOutput",
"Effect": "Allow",
"Action": [
"s3:ListBucket",
"s3:GetBucketLocation"
],
"Resource": "arn:aws:s3:::<bucket>",
"Condition": {
"StringLike": {
"s3:prefix": [
"<output-prefix>",
"<output-prefix>/*"
]
}
}
},
{
"Sid": "ReadWriteOutput",
"Effect": "Allow",
"Action": [
"s3:GetObject",
"s3:PutObject",
"s3:DeleteObject"
],
"Resource": "arn:aws:s3:::<bucket>/<output-prefix>/*"
}
]
}
Submit and run a Spark batch job
Once you have obtained the API key and uploaded necessary files to a Zilliz Cloud volume, you are all set to submit and run a Spark batch job.
Generate an idempotency key.
An idempotency key is a unique string that remains unchanged when retrying the same job request. For details, refer to Idempotent submission.
Prepare the request header.
Include the idempotency key generated in the previous step in the request header when creating a Spark batch job.
Authorization: Bearer <api-key>
Idempotency-Key: spark-job-20260730-001
Content-Type: application/json
Prepare the request payload.
All Spark batch jobs share a common payload structure, with job-specific parameters that vary with job purposes.
For the common payload structure, refer to Request payload. For job-specific parameters, refer to the following pages:
Idempotent submission
Use an idempotency key to make job submission safe to retry. Zilliz Cloud checks both the key and the request body to determine whether to return an existing job or reject the request as a conflict. The following table summarizes the matching behavior.
| Case | Behavior |
|---|---|
| Same idempotency key and same request body | Returns the previously created Spark batch job without creating a duplicate. |
| Same idempotency key and different request body | Returns a conflict error. |
The idempotency key is scoped to a specific organization, user, region, and project. Zilliz Cloud retains the key for approximately 25 hours, based on the maximum job timeout window. After the retention window expires, the same key can be reused for a new submission.
Request payload
All Spark batch jobs share a common payload structure as follows:
{
"description": "optional description",
"regionId": "aws-us-west-2",
"input": {...},
"output": {...},
"resourceSize": "SMALL",
"timeoutSeconds": 3600
}
The following table lists the descriptions of these parameters.
| Parameter | Required | Description |
|---|---|---|
description | No | An optional description for the job. The value is a string of no more than 1,024 characters. |
regionId | Yes | The ID of a Zilliz Cloud region in which the job runs. For supported regions, see Cloud Providers & Regions. |
input | No | The input of the job. This is mandatory only for built-in jobs. For details, refer to the table below. |
output | No | The output of the job. This is mandatory for built-in jobs. For details, refer to the table below. |
resourceSize | No | The size of the Spark cluster required for the job. Possible values are SMALL, MEDIUM, LARGE, XLARGE, 2XLARGE, and 3XLARGE. |
timeoutSeconds | No | The timeout duration in seconds for the current job. The value is a positive integer ranging from 300 to 86400. |
The input and output parameters in the above table share a similar structure as follows:
Parameter | Required | Description |
|---|---|---|
| Yes | The Spark batch job type. This applies both to
|
| No | The name of a Zilliz Cloud volume. This is mandatory when you set |
| No | The input/output file paths relative to the root of the specified Zilliz Cloud volume. This is mandatory when you set For a file at |
| No | The format of the input or output files. This applies both to |
| No | The write mode of output files. This applies only to
|
For input, format determines which Spark data source reader is used. If you omit this parameter, the job uses the Parquet reader by default and processes only the Parquet files under the specified path. Files in other formats, such as JSON or CSV, are ignored.
If the files you want to process are not Parquet, explicitly set input.format to the corresponding format. The job does not automatically detect and combine files of different formats.
Submission response
The response to job requests that succeed may have different HTTP codes. The following table lists the applicable HTTP codes carried in the response.
| Case | HTTP code | Description |
|---|---|---|
| Creating a job | 201 CREATED | Indicates that the job is being created. Creating a job is asynchronous, and you can use the job ID carried in the response to check its progress and manage its lifecycle. |
| Canceling a job | 202 ACCEPTED | Indicates that the cancellation is accepted and underway. Canceling a job is asynchronous, and you can use the job ID carried in the response to check its progress. |
| Describing a job | 200 OK | Indicates the requested response is returned. This is synchronous, and the response always carries the job status at the time when the request is processed. |
Although the HTTP codes differ, they share the same payload structure as follows:
{
"code": 0,
"data": {
"jobId": "job-xxxxxxxx",
"projectId": "proj-xxxxxxxx",
"type": "SPARK",
"description": "backfill product attributes",
"status": "PENDING",
"regionId": "aws-us-west-2",
"clusterId": "in-xxxxxxxx",
"createdAt": null,
"startedAt": null,
"finishedAt": null,
"durationSeconds": null
}
}
The response includes a job ID that you can use to monitor the job. To view job states, retrieve job details, or cancel a running job, see Understand job states.
When an error occurs, the response is similar to the following:
{
"code": 10001,
"message": "projectId is required",
"details": {
"errorCode": "INVALID_PARAMETER"
}
}
You can tell the error from details.errorCode and the HTTP CODE the error response carries. The following table lists the applicable HTTP CODE with their indications.
| HTTP code | Description |
|---|---|
400 BAD REQUEST | Indicates that the parameters carried in the request are incorrect. |
403 FORBIDDEN | Indicates that the API key does not have sufficient permissions or the resources are not available in the specified project. |
404 NOT FOUND | Indicates that the specified resources, such as the job ID, Zilliz Cloud volume, do not exist. |
409 CONFLICT | Indicates that the idempotency key does not match the request payload. |
500 INTERNAL SERVER ERROR | Indicates that the server fails to process the request. |
Next steps
Use the following guides to create the Spark batch job that matches your goal, or to monitor and manage an existing job.
Primary-Key Deduplication
Primary-key deduplication identifies records that share the same primary key and removes redundant copies from a large dataset. Use this job to clean up duplicates introduced by repeated imports, pipeline retries, migrations, or overlapping data sources. | Cloud
Vector Similarity Deduplication
Vector similarity deduplication identifies records with highly similar embeddings and groups them as semantic duplicates. Use this job to reduce semantic redundancy, such as paraphrased text, slightly modified images, or multiple versions of similar content. | Cloud
K-Means Clustering
K-Means clustering groups records with similar embeddings into a specified number of clusters. Use this job to explore the distribution of your vector data, organize records into coarse semantic groups, or prepare datasets for sampling, analysis, and other downstream workflows. | Cloud
Anomaly Detection
Anomaly detection identifies records whose vector embeddings differ substantially from the broader data distribution. Use this job to find unusual records that may represent data-quality issues, rare cases, processing errors, or samples that require further review. | Cloud
Data Backfill
Data backfill lets you update selected fields of existing entities in a Zilliz Cloud collection using data stored in Parquet files on a Zilliz Cloud volume. The backfill job matches input records to existing entities by primary key and writes the specified field values back to the collection. You can use it to populate newly added fields, fill missing values, or replace existing field values at scale. | Cloud
Manage Spark Batch Jobs
Spark batch jobs run asynchronously and move through several states from submission to completion. This page explains the job lifecycle and then shows how to list jobs, retrieve job details, and cancel jobs that are still in a cancelable state. | Cloud