Sarvam AI launches Vision 2.1 with better document reading
Homegrown AI startup Sarvam AI has rolled out Sarvam Vision 2.1, an upgraded version of its vision-language model designed to understand and process documents. Now it comes with improved capabilities for reading complex tables, extracting information from structured forms, and recognising handwritten text in Indian languages.
The company says Vision 2.1 addresses key limitations observed in its earlier version. And the focus this time is not just on reading documents, but also on turning messy, unstructured scans into clean, structured data that businesses can integrate directly into production workflows.advertisementโฎโฏ Read Full StoryWhat can Sarvam Vision 2.1 doThe new model ships with improved tabular processing. In simpler terms, the model can work with nested, complex, and multi-page table structures that standard vision models frequently struggle to parse accurately. It can also extract key-value information from forms, such as names, dates, and financial figures, as well as recognize handwritten notes and annotations in Indic languages.
Sarvam noted that specific training interventions were made to reduce hallucinations and spatial inconsistencies, two common failure points for AI systems dealing with dense or irregular document scans.
Alongside the launch, the company introduced a new Sarvam Indic OCR benchmark covering all 22 official Indian languages. The evaluation dataset contains 6,909 samples, including 6,609 samples across the 22 regional languages and 300 English baselines. The material spans newspapers, brochures, textbooks, and historical archives dating from 1800 to the present day.
On this benchmark, Sarvam Vision 2.1 recorded an overall word accuracy of 87.39 per cent. The model also scored 87.30 on the global olmOCR-Bench, demonstrating competitive performance on international document parsing standards.Training and deployment
Sarvam trained the new model on a mix of synthetic and real-world document scans, leaning especially on regional handwriting variations and form-extraction cases. It then fine-tuned the model with supervised learning before layering on reinforcement learning with verifiable rewards (RLVR), a step aimed at pushing the model toward exact structural and textual accuracy, rather than just plausible-looking output.
The model is accessible through Sarvamโs document intelligence APIs, which include developer tools for converting multi-page files into structured text and automatically extracting data from forms and tables.- EndsAlso Read | Sridhar Vembu compares Zoho Chennai campus to classic Europe, says it will be open for publicAlso Read | Explained: Rogue OpenAI agents break loose and hack Australian government website, full story in 5 pointsAlso Read | Google says Gemini AI can make calls and talk to people on your behalf now
Iโm a journalist who loves reading and exploring diverse realms.


