Skip to main content
Cohere
current

Search documentation

Type to search this documentation.

On this pageOverview

Bulk Embedding (might be redundant)

This document explains how to efficiently embed text in bulk using an endpoint that returns a JSONL file of text embeddings with their respective text, which can be downloaded within 2 months.

If you have a large collection of text rather than a single file, bulk embedding may be more efficient for you. This endpoint returns a JSONL file of text embeddings with their respective text. The JSONL will be available for 2 months and downloadable via the download_file().

PYTHON
# Request
co.bulk_embed(
    model="string",
    url="string",  # (url to gcp file)
    file_id="string",  # (id to uploaded file)
    text_field="string",  # (column to get text from)
    truncate="string",  # {LEFT:DEFAULT,RIGHT}
)
PYTHON
{job_id: "string"}
PYTHON
co.get_bulk_embed(job_id="string")
PYTHON
# Response of co.get_bulk_embed():
{
    job_id: "string",
    status: "string",  # (COMPLETE, QUEUED, FAILED or RUNNING)
    created_at: "string",  # (timestamp)
    updated_at: "string",  # (timestamp)
    input_url: "string",
    input_file_id: "string",
    output_file_id: "string",
    model: "string",
    truncate: "string",  # (LEFT or RIGHT)
    percent_complete: float,
}
PYTHON
co.list_bulk_embed()
PYTHON
# Response of co.list_bulk_embeds():
{
    bulk_embed: [
        {
            job_id: "string",
            status: "string",  # (COMPLETE, QUEUED, FAILED or RUNNING)
            created_at: "string",  # (timestamp)
            updated_at: "string",  # (timestamp)
            input_url: "string",
            input_file_id: "string",
            output_file_id: "string",
            model: "string",
            truncate: "string",  # (LEFT or RIGHT)
            percent_complete: float,
        },
        ...,
    ]
}
Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu