Cohere Launches 2.4B Visual Mini-Model: Document Understanding Outperforms Ministral's 33B Model
2026-08-13 15:55:53
According to CoinMeta, Cohere has open-sourced its smallest visual language model to date, north, with 2.4B parameters and licensed under the Apache 2.0 license. This model is specialized in processing documents, tables, charts, screenshots, and OCR. It can handle images in their original proportions and resolutions without the need for compression first. In official tests, docvqa achieved a score of 92.1%, surpassing the 89.6% of Ministral (with 33B parameters) and the 73.2% of Gemma and E2B. Its visual positioning capability scored 73.2%, which is also higher than that of Ministral and Gemma. Although it performs excellently in document understanding and visual positioning, it is not the strongest model in terms of size; Qwen3.5-2B is still stronger in most general visual tasks, OCR, and multimodal benchmarks.
Bullish 0
Bearish 0
Source:Internet
This content is for market information only and does not constitute investment advice.