Get Started
Cohere Platform
OverviewCohere offers world-class Large Language Models (LLMs) like Command, Rerank, and Embed. These help developers and enterprises build LLM-powered applications.InstallationA guide for installing the Cohere SDK, supported in 4 different languages – Python, TypeScript, Java, and Go.Creating a clientA guide for creating Cohere API client using Cohere SDK, supported in 4 different languages – Python, TypeScript, Java, and Go.
Quickstart
RAGA quickstart guide for performing retrieval augmented generation (RAG) with Cohere's Command models (v2 API).RerankingA quickstart guide for performing reranking with Cohere's Reranking models (v2 API).Semantic SearchA quickstart guide for performing text semantic search with Cohere's Embed models (v2 API).Text GenerationA quickstart guide for performing text generation with Cohere's Command models (v2 API).Tool Use & AgentsA quickstart guide for using tool use and building agents with Cohere's Command models (v2 API).Transcribing AudioA quickstart guide for transcribing audio with the Cohere Transcribe model.
Document Parsing
Model Vault
OverviewModel Vault is a Cohere-managed, single-tenant environment for deploying and serving Cohere models. Every vault is either Standard or Encrypted.QuickstartCreate your first vault from the Model Vault app and make an inference request in a few minutes.
Deploy & manage
Home PageFind and manage all of your vaults (Standard and Encrypted) from one place on the Model Vault home page.Creating a VaultCreate a new vault, choose Standard or Encrypted, and select a model, performance tier, and replicas.Managing VaultsView vault details and edit, pause, resume, or delete models from the Model Vault app.
Operate & observe
Standard Vault
OverviewStandard Vault is Cohere's managed, single-tenant inference environment with dedicated infrastructure and no confidential-computing layer.Supported ModelsCohere models and GPUs available in a Standard Vault.Calling the APICall a Standard Vault with the Cohere SDK, raw HTTP, or an OpenAI-compatible client by pointing requests at your vault endpoint URL.PricingStandard Vault pricing models (Fixed and Flex) and per-model performance tiers and rates.
Encrypted Vault
OverviewEncrypted Vaults add confidential computing to Model Vault, so prompts and responses stay protected end to end with verifiable attestation.Supported ModelsWhich Cohere models are available in Model Vault Encrypted, the supported confidential-computing GPUs, and the isolating architecture.Calling the APICall an Encrypted Vault through the Cohere OHTTP proxy that verifies the TEE and encrypts end to end before any data is sent.
Security
Confidential Computing PrimerA primer on the trusted execution environments and GPU confidential computing that power Model Vault Encrypted.Security ModelThe trust boundary and threat model for Model Vault Encrypted: who can and cannot access your data.Remote AttestationHow remote attestation and the Passport model with Intel Trust Authority prove which code is running inside a Model Vault Encrypted deployment.Verifying Your DeploymentHow to verify a Model Vault Encrypted deployment: automatic client-side checks and the attestation details you can inspect in the Model Vault app.Encryption & Key ManagementHow Model Vault Encrypted protects data in transit, at rest, and in use, and how encryption keys and Zero Data Retention are handled.ComplianceHow Model Vault Encrypted supports compliance requirements such as GDPR, HIPAA, and SOC 2 through hardware-enforced confidentiality and verifiable attestation.
Models
Audio
Aya
AyaUnderstand Cohere Labs groundbreaking multilingual Aya models, which aim to bring many more languages into generative AI.Aya VisionUnderstand Cohere Labs groundbreaking multilingual model Aya Vision, a state-of-the-art multimodal language model excelling at multiple tasks.Aya ExpanseUnderstand Cohere Labs highly performant multilingual Aya models, which aim to bring many more languages into generative AI.Tiny AyaTiny Aya is a compact yet powerful 3.35B-parameter multilingual model supporting 70 languages, designed for efficient and practical multilingual AI deployment.
Command
Command A+Command A+ is a Mixture of Experts (MoE) model with 25B active and 218B total parameters, excelling in agentic, reasoning, vision, and multilingual tasks.Command ACommand A is a performant mode good at tool use, RAG, agents, and multilingual use cases. It has 111 billion parameters and a 256k context length.Command A ReasoningCommand A Reasoning excels in tool use, agentic workflows, and complex problem-solving. It has 111 billion parameters and a 256k context length.Command A TranslateCommand A Translate is a state of the art model performant in 23 languages. It has a context length of 16K tokens and 111B parameters.Command A VisionCommand A Vision is a powerful visual language model capable of interacting with image inputs. This document contains information about its capabilities.Command R7BCommand R7B is the smallest, fastest, and final model in our R family of enterprise-focused large language models. It excels at RAG, tool use, and agents.Command R+Command R+ is Cohere's optimized for conversational interaction and long-context tasks, best suited for complex RAG workflows and multi-step tool use.Command RCommand R is a conversational model that excels in language tasks and supports multiple languages, making it ideal for coding use cases.
North
North Small TranslateNorth Small Translate is a 218B total / 25B active parameter MoE model purpose-built for machine translation across more than 50 languages.North Mini CodeNorth Mini Code is a 30B total / 3B active parameter MoE model trained for agentic coding, released under Apache 2.0 and suitable for local deployment.
Text Generation
Introduction to Text Generation at CohereThis page describes how a large language model generates textual output.Using the Chat APIHow to use the Chat API endpoint with Cohere LLMs to generate text responses in a conversational interfaceReasoningReasoning models excel at tool use, agentic workflows, and complex problem-solving. This page provides a general overview of Cohere's reasoning capalities.Image InputsThis page describes how a Cohere large language model works with image inputs. It covers passing images with the API, limitations, and best practices.Streaming ResponsesThe document explains how the Chat API can stream events like text generation in real-time.
Structured Outputs
Structured OutputsThis page describes how to get Cohere models to create outputs in a certain format, such as JSON, TOOLS, using parameters such as `response_format`.Parameter Types in Structured Outputs (JSON)This page shows usage examples of the JSON Schema parameter types supported in Structured Outputs (JSON).
Predictable OutputsStrategies for decoding text, and the parameters that impact the randomness and predictability of a language model's output.Advanced Generation ParametersThis page describes advanced parameters for controlling generation.
Retrieval Augmented Generation (RAG)
Basic usageGuide on using Cohere's Retrieval Augmented Generation (RAG) capabilities such as document grounding and citations.End-to-end exampleGuide on using Cohere's Retrieval Augmented Generation (RAG) capabilities covering the Chat, Embed, and Rerank endpoints (API v2).StreamingGuide on implementing streaming for RAG with Cohere and details on the events stream (API v2).CitationsGuide on accessing and utilizing citations generated by the Cohere Chat endpoint for RAG. It covers both non-streaming and streaming modes (API v2).
Tool Use
Tool UseLearn when to use leverage multi-step tool use in your workflows.Basic usageAn overview of using Cohere's tool use capabilities, enabling developers to build agentic workflows (API v2).Usage patternsGuide on implementing various tool use patterns with the Cohere Chat endpoint such as parallel tool calling, multi-step tool use, and more (API v2).Parameter typesGuide on using structured outputs with tool parameters in the Cohere Chat API. Includes guide on supported parameter types and usage examples (API v2).StreamingGuide on implementing streaming for tool use in Cohere's platform and details on the events stream (API v2).CitationsGuide on accessing and utilizing citations generated by the Cohere Chat endpoint for tool use. It covers both non-streaming and streaming modes (API v2).
Prompt Engineering
Crafting Effective PromptsThis page describes different ways of crafting effective prompts for prompt engineering.Advanced Prompt Engineering TechniquesThis page describes advanced ways of controlling prompt engineering.System MessagesThis page describes how Cohere system messages work, and the effect they have on output.
Prompt Library
Create CSV data from JSON dataThis document provides an example of converting a JSON object into CSV format using the Cohere API.Create a markdown table from raw dataThe document provides a prompt to format CSV data into a markdown table and includes the output table as well as an API request using the Cohere platform.Meeting SummarizerThe document discusses the creation of a meeting summarizer with Cohere's large language model.Remove PIIThis document provides an example of redacting personally identifiable information (PII) from a conversation while maintaining context, using the Cohere API.Add a Docstring to your codeThis document provides an example of adding a docstring to a Python function using the Cohere API.Evaluate your LLM responseLearn how to use Command-R to evaluate natural language responses with an example of grading formality.Multilingual interpreterThis document provides a prompt to interpret a customer's issue into multiple languages using an API.
Embeddings (Vectors, Search, Retrieval)
Introduction to Embeddings at CohereEmbeddings transform text into numerical data, enabling language-agnostic similarity searches and efficient storage with compression.Semantic Search with EmbeddingsExamples on how to use the Embed endpoint to perform semantic search (API v2).Multimodal EmbeddingsMultimodal embeddings convert text and images into embeddings for search and classification (API v2).Batch Embedding JobsLearn how to use the Embed Jobs API to handle large text data efficiently with a focus on creating datasets and running embed jobs.
Reranking
Going to Production
API Keys and Rate LimitsThis page describes Cohere API rate limits for production and evaluation keys.Going LiveLearn to upgrade from a Trial to a Production key; understand the limitations and benefits of each and go live with Cohere.DeprecationsLearn about Cohere's deprecation policies and recommended replacementsHow Does Cohere's Pricing Work?This page details Cohere's pricing model. Our models can be accessed directly through our API, allowing for the creation of scalable production workloads.
Integrations
Integrating Embedding Models with Other Tools
Integrating Embedding Models with Other ToolsLearn how to integrate Cohere embeddings with open-source vector search engines for enhanced applications.Elasticsearch and CohereLearn how to create a semantic search pipeline with Elasticsearch and Cohere's generative AI capabilities.MongoDB and CohereBuild semantic search and RAG systems using Cohere and MongoDB Atlas Vector Search.Redis and CohereLearn how to integrate Cohere with Redis for similarity searches on text data with this step-by-step guide.Haystack and CohereBuild custom LLM applications with Haystack, now integrated with Cohere for embedding, generation, chat, and retrieval.Pinecone and CohereThis page describes how to integrate Cohere with the Pinecone vector database.Weaviate and CohereThis page describes how to integrate Cohere with the Weaviate database.Open Search and CohereUnlock the power of search and analytics with OpenSearch, enhanced by ML connectors like Cohere and Amazon Bedrock.Vespa and CohereThis page describes how to integrate Cohere with the Vespa database.Qdrant and CohereThis page describes how to integrate Cohere with the Qdrant vector database.Milvus and CohereThis page describes integrating Cohere with the Milvus vector database.Zilliz and CohereThis page describes how to integrate Cohere with the Zilliz database.Chroma and CohereThis page describes how to integrate Cohere and Chroma.
Cohere and LangChain
Cohere and LangChainIntegrate Cohere with LangChain for advanced chat features, RAG, embeddings, and reranking; this guide includes code examples for each feature.Chat on LangChainIntegrate Cohere with LangChain to build applications using Cohere's models and LangChain tools.Embed on LangChainThis page describes how to work with Cohere's embeddings models and LangChain.Rerank on LangChainThis page describes how to integrate Cohere's ReRank models with LangChain.Tools on LangChainExplore code examples for multi-step and single-step tool usage in chatbots, harnessing internet search and vector storage.
Deployment Options
OverviewThis page provides an overview of the available options for deploying Cohere's models.SDK CompatibilityThis page describes various places you can use Cohere's SDK.
Private Deployment
OverviewThis page provides an overview of private deployments of Cohere's models.Setting UpThis page describes the setup required for private deployments of Cohere's models.Model DeploymentLearn how to pull and test Cohere's container images using a license with Docker and Kubernetes.Model Deployment - AWSDeploying Cohere models in AWS via EC2 or EKS for enhanced security, compliance, and control.UsageThis page describes how to use Cohere's SDK to access privately deployed Cohere models.
Cloud AI Services
Cohere on AWS
Cohere on AWSAccess Cohere's language models on AWS with customization options through Amazon SageMaker and Amazon Bedrock.Amazon BedrockThis document provides a guide for using Cohere's models on Amazon Bedrock.
Amazon SageMaker
Tutorials
CookbooksGet started with Cohere's cookbooks to build agents, QA bots, perform searches, and more, all organized by category.LLM UniversityLLM University (LLMU) offers in-depth, practical NLP and LLM training. Ideal for all skill levels. Learn, build, and deploy Language AI with Cohere.
Build Things with Cohere!
Build Things with Cohere!This page describes how to build an onboarding assistant with Cohere's large language models.Cohere Text Generation TutorialThis page walks through how Cohere's generation models work and how to use them.Building a Chatbot with CohereThis page describes building a generative-AI powered chatbot with Cohere.Semantic Search with CohereThis is a tutorial describing how to leverage Cohere's models for semantic search.Reranking with CohereThis page contains a tutorial on using Cohere's ReRank models.RAG with CohereThis page walks through building a retrieval-augmented generation model with Cohere.Building an Agent with CohereThis page describes building a generative-AI powered agent with Cohere.
Agentic RAG
Agentic RAGHands-on tutorials on building agentic RAG applications with CohereRouting Queries to Data SourcesBuild an agentic RAG system that routes queries to the most relevant tools based on the query's nature.Generating Parallel QueriesBuild an agentic RAG system that can expand a user query into a more optimized set of queries for retrieval.Performing Tasks SequentiallyBuild an agentic RAG system that can handle user queries that require tasks to be performed in a sequence.Generating Multi-Faceted QueriesBuild a system that generates multi-faceted queries to capture the full intent of a user's request.Querying Structured Data (Tables)Build an agentic RAG system that can query structured data (tables).Querying Structured Data (SQL)Build an agentic RAG system that can query structured data (SQL).
Cohere on Azure
Cohere on AzureAn introduction to Cohere on Azure AI Foundry, a fully managed service by Azure (API v2).Text GenerationA guide for performing text generation with Cohere's Command models on Azure AI Foundry (API v2).Semantic SearchA guide for performing text semantic search with Cohere's Embed models on Azure AI Foundry (API v2).RerankingA guide for performing reranking with Cohere's Reranking models on Azure AI Foundry (API v2).Retrieval Augmented Generation (RAG)A guide for performing retrieval augmented generation (RAG) with Cohere's Command models on Azure AI Foundry (API v2).Tool Use & AgentsA guide for using tool use and building agents with Cohere's Command models on Azure AI Foundry (API v2).
Responsible Use
Cohere Labs
More Resources
Cohere ToolkitBuild and deploy RAG applications quickly with the Cohere Toolkit, which offers pre-built front-end and back-end components.DatasetsLearn about the Dataset API, including its file size limits, data retention, creation, validation, metadata, and more, with provided code snippets.Improve Cohere DocsContribute to our docs content, stored in the cohere-developer-experience repo; we welcome your pull requests!
Cohere API
AboutCohere's NLP platform provides customizable large language models and tools for developers to build AI applications.Teams and RolesThe document outlines how to work in teams on the Cohere platform, including inviting others, managing roles, and access permissions for Owners and Users.ErrorsUnderstand Cohere's HTTP response codes and how to handle errors in various programming languages.
Cookbooks
CookbooksExplore a range of AI guides and get started with Cohere's generative platform, ready-made and best-practice optimized.Agent API CallsThis page how to use Cohere's API to build an LLM-based agent.Short-Term Memory Handling for AgentsThis page describes how to manage short-term memory in an agent built with Cohere models.Agentic Multi-Stage RAG with Cohere Tools APIThis page describes how to build a powerful, multi-stage agent with the Cohere platform.Agentic RAG for PDFs with mixed dataThis page describes building a powerful, multi-step chatbot with Cohere's models.Analysis of Form 10-K/10-Q Using Cohere and RAGThis page describes how to use Cohere's large language models to build an agent able to analyze financial forms like a 10-K or a 10-Q.Analyzing Hacker News with Six Language Understanding MethodsThis page describes building a generative-AI powered tool to analyze headlines with Cohere.Article Recommender with Text Embedding Classification ExtractionThis page describes how to build a generative-AI tool to recommend articles with Cohere.Multi-Step Tool UseThis page describes how to create a multi-step, tool-using AI agent with Cohere's tool use functionality.Basic RAGThis page describes how to work with Cohere's basic retrieval-augmented generation functionality.Basic Semantic SearchThis page describes how to do basic semantic search with Cohere's models.Basic Tool UseThis page describes how to work with Cohere's basic tool use functionality.Calendar Agent with Native Multi Step ToolThis page describes how to use cohere Chat API with list_calendar_events and create_calendar_event tools to book appointments.Chunking StrategiesThis page describes various chunking strategies you can use to get better RAG performance.Creating a QA Bot From Technical DocumentationThis page describes how to use Cohere to build a simple question-answering system.Financial CSV Agent with Native Multi-Step Cohere APIThis page describes how to use Cohere's models and its native API to build an agent able to work with CSV data.Financial CSV Agent with LangchainThis page describes how to use Cohere's models to build an agent able to work with CSV data.Migrating away from createcsvagent in langchain-cohereThis page contains a tutorial on how to build a CSV agent without the deprecated `create_csv_agent` abstraction in langchain-cohere v0.3.5 and beyond.A Data Analyst Agent Built with Cohere and LangchainThis page describes how to build a data-analysis system out of Cohere's models.Advanced Document Parsing For EnterprisesThis page describes how to use Cohere's models to build a document-parsing agent.End-to-end RAG using Elasticsearch and CohereThis page contains a basic tutorial on how to get Cohere and ElasticSearch to work well together.Semantic Search with Cohere Embed Jobs and Pinecone serverless SolutionThis page contains a basic tutorial on how to get Cohere and the Pinecone vector database to work well together.Semantic Search with Cohere Embed JobsThis page contains a basic tutorial on how to use Cohere's Embed Jobs functionality.Fueling Generative Content with Keyword ResearchThis page contains a basic workflow for using Cohere's models to come up with keyword content ideas.Grounded Summarization Using Command RThis page contains a basic tutorial on how to do grounded summarization with Cohere's models.Hello World! Meet Language AIThis page contains a breakdown of some of what can be achieved with Cohere's LLM platform.Long Form General StrategiesThis discusses ways of getting Cohere's LLM platform to perform well in generating long-form text.Migrating Monolithic Prompts to Command-R with RAGThis page contains a discussion of how to automatically migrating monolothic prompts.Multilingual Search with Cohere and LangchainThis page contains a basic tutorial on how to do search across different languages with Cohere's LLM platform.PDF Extractor with Native Multi Step Tool UseThis page describes how to create an AI agent able to extract information from PDFs.Pondr, Fostering Connection through Good ConversationThis page contains a basic tutorial on how tplay an AI-powered version of the icebreaking game 'Pondr'.Deep Dive Into RAG EvaluationThis page contains information on evaluating the output of RAG systems.RAG With Chat Embed and Rerank via PineconeThis page contains a basic tutorial on how to build a RAG-powered chatbot.Demo of RerankThis page contains a basic tutorial on how Cohere's ReRank models work and how to use them.SQL AgentThis page contains a tutorial on how to build a SQL agent with Cohere's LLM platform.Summarization EvalsThis page discusses how to evaluate a model's text summarization.Text Classification Using EmbeddingsThis page discusses the creation of a text classification model using word vector embeddings.Topic Modeling AI PapersThis page discusses how to create a topic-modeling system for papers focused on AI papers.Wikipedia Semantic Search with Cohere + WeaviateThis page contains a description of building a Wikipedia-focused search engine with Cohere's LLM platform and the Weaviate vector database.Wikipedia Semantic Search with Cohere Embedding ArchivesThis page contains a description of building a Wikipedia-focused semantic search engine with Cohere's LLM platform and the Weaviate vector database.Build Chatbots That Know Your Business with MongoDB and CohereThis page describes how to build a chatbot that provides actionable insights on technology company market reports.Finetuning on Cohere's PlatformAn example of finetuning using Cohere's platform and a financial dataset.Deploy your finetuned model on AWS MarketplaceLearn how to deploy your finetuned model on AWS Marketplace.Finetuning on AWS SagemakerLearn how to finetune one of Cohere's models on AWS Sagemaker.SQL Agent with Cohere and LangChain (i-5O Case Study)This page contains a tutorial on how to build a SQL agent with Cohere and LangChain in the manufacturing industry.Introduction to Aya VisionIn this notebook, we will explore the capabilities of Aya Vision, which can take text and image inputs to generates text responses.Retrieval Evaluation with LLM-as-a-Judge via Pydantic AIThis page contains a tutorial on how to evaluate retrieval systems using LLMs as judges via Pydantic AI.Document Translation with Command A TranslateThis page describes how to use Command A Translate for automated translation across 23 languages with industry-leading performance.