Cohere Labs released a 2.4B Apache-2.0 vision model that reads A4 pages at 200 dpi. Compact OCR is a model-selection decision, not a souvenir.

Via Cohere Labs: Meet North Micro Vision: A 2.4B Native-Resolution Vision-Language Model
Cohere Labs released North-Micro-Vision-Instruct on August 12, 2026: a 2.4-billion-parameter open-weight vision-language model under Apache 2.0, with native-resolution image support. For an owner-led firm, the news is not another Hugging Face trophy. It is whether invoices, scanned forms, and chart-heavy PDFs still have to leave the building because the only vision models staff trust live in someone else's cloud.
The model is Cohere's smallest VLM to date. A custom 400-million-parameter native-resolution vision encoder sits in front of a 2-billion-parameter North Micro language model. Training raised the native-resolution cap to 1,654 × 2,339 pixels — an A4 page at 200 dpi — while preserving aspect ratio instead of squashing every page into a small square. That is the operational fact. Most consumer chat tools still crop first and apologize later.
Weights are on Hugging Face. Public vLLM support is listed as coming soon. MLX-VLM community weights and an NVIDIA AutoModel fine-tune recipe already exist. The gap between “the weights dropped” and “your stack can serve them on Tuesday” is exactly why model selection is a job, not a bookmark.
In information-heavy SMEs — law, accounting, healthcare admin, insurance, consulting — “vision AI” often means a staffer photographing a page into a personal ChatGPT account. That is a shadow-AI problem wearing a productivity costume. A compact Apache-2.0 model with a document-heavy curriculum changes the menu: you can evaluate a local, VPC, or tightly scoped cloud path for OCR and layout without waiting for your SaaS vendor to add the feature.
Cohere's encoder curriculum started from Google's SigLIP 2 SO400M checkpoint and mixed dense captions with OCR data (40% OCR in stage 1; 50/50 thereafter). Stage 3 instruction-tuning put native OCR, charts and tables, OCR QA, and grounding at the top of the mixture. On Cohere's published table, DocVQA VAL scores 0.921 and ChartQA Test 0.808 — competitive with larger compact VLMs in the same size band.
It is not a general genius. MMMU DEV_VAL sits at 0.329, well behind several peers. Treat that as a feature of the brief, not a bug in the press release: this release is strongest where small text and layout survive, which is the work that actually hits a practice's scanner. If you need frontier STEM reasoning over images, this is not your one model. If you need to read a form without sending the form to a consumer endpoint, it is now on the shortlist.
Licensing is the other quiet decision. Apache 2.0 is commercially usable without the extra fog that still hangs over some lab licenses. That does not make deployment free. It makes the legal conversation shorter — which still belongs with qualified counsel, not a blog post.
The resulting aligned checkpoint can process a single A4 document page at up to 200 dpi while preserving its aspect ratio. — Cohere Labs
AgentsROI.ai is a managed AI services provider for owner-led SMEs. We do not sell a vision stack. We help you choose, govern, and measure models so the work keeps paying for itself.
Lead with Model Selection & Continuity Planning. North Micro Vision is a candidate for document OCR and layout — not a replacement for every chat model in the firm. We match the model to the job: local or VPC for confidential pages, a named cloud VLM as fallback, and a written sunset plan when weights, licenses, or serving stacks change.
Pair it with a Shadow-AI Risk Assessment when staff already photograph client files into personal accounts. A new open VLM does not fix unsanctioned uploads. It gives you a sanctioned alternative, which only works if you know what people are already using.
Managed AI Operations is the destination if the local path sticks: monitoring, quantization choices, and preventing a laptop experiment from becoming an unowned production dependency. We stay vendor-neutral. Apache-2.0 Cohere weights, NVIDIA recipes, and whoever serves your fallback are tools. The operating decision is yours.
A 2.4B native-resolution VLM that can read an A4 page at 200 dpi is not magic. It is a concrete option on a model menu that most small firms still have not written down. If your team is already “using AI” on documents, the question is which model, which machine, and which data-handling rule — before the next scan lands in a personal chat window.
If that sounds like your shop, start with Model Selection & Continuity Planning. Name a primary and a fallback for document vision. Then decide whether on-prem OCR is worth the ops cost. Book a no-pressure assessment when you are ready to treat vision models as procurement, not novelty.
This article summarizes publicly reported information and is for general informational purposes only. It does not constitute legal, tax, financial, investment, security, or compliance advice. AgentsROI.ai is not a law firm, accounting firm, or registered investment adviser. Facts, pricing, statistics, and product capabilities cited here reflect the sources listed at the time of writing and may change. Readers should verify current information independently and consult qualified professionals regarding obligations specific to their industry, jurisdiction, and circumstances—including applicable New York State and New York City requirements. AgentsROI.ai may have commercial relationships with vendors mentioned; where material, such relationships are disclosed. Nothing in this article is an endorsement of any specific AI product, model, or provider.