Standard Vault Pricing
Standard Vault is billed as a Cohere-managed service. Pricing depends on the models you select and each model's performance tier. Cohere manages the underlying infrastructure and scaling, and customers can choose between two pricing models:
| Feature | Fixed | Flex |
|---|---|---|
| Commitment | Monthly or annual | Monthly or annual |
| Capacity | Fixed number of instances (no autoscaling) | Minimum baseline instances, plus autoscaling |
| Sizing | Determined through a sizing exercise or a production trial (for example, based on expected load) | -- |
| Autoscaling | -- | Scales up/down based on request rate and agreed latency SLOs |
| Pause/resume | You can pause a model or restart a paused model to save on costs. | You can pause a model or restart a paused model to save on costs. |
| Overages | -- | Additional capacity billed per instance-hour |
| Max capacity | -- | Maximum instance cap per model |
The following table summarizes the available models and their rates. All rates are per instance.
| Model | Performance Tier | Hourly rate | Monthly rate | Annual rate |
|---|---|---|---|---|
| Embed 3 | Small | $4.00 | $2,500 | $25,000 |
| Embed 4 | Small | $4.00 | $2,500 | $25,000 |
| Embed 4 | Medium | $5.00 | $3,250 | $32,500 |
| Rerank 3.5 | Medium | $5.00 | $3,250 | $32,500 |
| Rerank 4 Fast | Medium | $5.00 | $3,250 | $32,500 |
| Rerank 4 Pro | Medium | $5.00 | $3,250 | $32,500 |
| Rerank 4 Pro | Large | $10.00 | $6,500 | $65,000 |
| Parse 5 | Medium | $4.00 | $2,500 | $25,000 |
| Parse 5 | Large | $8.00 | $5,000 | $50,000 |
| Parse 5 | XL | $7.00 | $4,300 | $43,000 |
| Cohere Transcribe | Medium | $3.75 | $2,500 | $25,000 |
| Cohere Transcribe | Large | $7.50 | $4,750 | $47,500 |
You may also want to compare Standard Vault pricing against the operational and capacity costs of running inference directly in your cloud provider account (for example, AWS), where cloud-provider credits may apply.
Generative models
Section titled “Generative models”The rates above cover Embed, Rerank, Parse, and Transcribe. Embed and Rerank are available self-serve. Generative models (the Command A family and North) are also available on Standard Vault, typically through a waitlist. All rates are per instance, per hour.
| Model | L Hourly rate | XL hourly rate |
|---|---|---|
| Command A | $40.00 | $48.00 |
| Command A Vision | $40.00 | $48.00 |
| Command A Translate | $40.00 | $48.00 |
| Command A Reasoning | $48.00 | $57.50 |
| Command A+ | $17.50 | $32.50 |
| North Mini Code | $7.50 | $10.50 |
Monthly and annual commitment pricing follows the same Fixed and Flex plans described above; contact Cohere for those rates and for access. Generative access typically requires a waitlist, so check the model dropdown when creating a vault, or contact Cohere. Bundles and customized models are also available; see Supported Models for the full list.
For the pricing of the encrypted, confidential-computing product, see Model Vault Encrypted Pricing.
Performance Tiers
Section titled “Performance Tiers”Each model has a performance tier based on latency requirements and throughput service level
objectives (SLOs) per instance. You can see the tiers listed in the model dropdown selection as a size
letter (e.g., S, M, L). The tiers follow an instance/hour pricing, which is then
incorporated into your payment plan. We recommend selecting the model-performance tier combination
that matches the nature of your workflow and the required measures of performance.