Skip to main content
Cohere

Search documentation

Type to search this documentation.

On this pageOverview

Our Groundbreaking Multimodal Model, Aya Vision, is Here!

Today, Cohere Labs, Cohere’s research arm, is proud to announce Aya Vision, a state-of-the-art multimodal large language model excelling across multiple languages and modalities. Aya Vision outperforms the leading open-weight models in critical benchmarks for language, text, and image capabilities.

Built as a foundation for multilingual and multimodal communication, this groundbreaking AI model supports tasks such as image captioning, visual question answering, text generation, and translations from both texts and images into coherent text.

Refer to Aya Vision for more information.

Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu