{"name":"AWESOME-OCR-LLM: Curated Reading List for OCR in the LLM Era","description":"AWESOME-OCR-LLM is a continuously updated reading list focusing on Optical Character Recognition (OCR) in the era of large language models (LLMs). It covers key areas like document parsing, understanding, visual text generation, and benchmarks, highlighting research from the past five years. This resource is invaluable for anyone tracking the rapid advancements in multimodal document AI.","github":"https://github.com/Yuliang-Liu/AWESOME-OCR-LLM","url":"https://osrepos.com/repo/yuliang-liu-awesome-ocr-llm","source":"osrepos.com","sourceDescription":"This repository profile is provided by osrepos.com, an open source repository discovery platform.","repositoryProfile":"https://osrepos.com/repo/yuliang-liu-awesome-ocr-llm","generatedFor":"open source discovery and AI-assisted research","markdown":"https://osrepos.com/repo/yuliang-liu-awesome-ocr-llm.md","json":"https://osrepos.com/repo/yuliang-liu-awesome-ocr-llm.json","topics":["OCR","LLM","Document AI","Computer Vision","Multimodal","Research","Reading List","Deep Learning"],"keywords":["OCR","LLM","Document AI","Computer Vision","Multimodal","Research","Reading List","Deep Learning"],"stars":null,"summary":"AWESOME-OCR-LLM is a continuously updated reading list focusing on Optical Character Recognition (OCR) in the era of large language models (LLMs). It covers key areas like document parsing, understanding, visual text generation, and benchmarks, highlighting research from the past five years. This resource is invaluable for anyone tracking the rapid advancements in multimodal document AI.","content":"## Introduction\n\nThe AWESOME-OCR-LLM repository is a meticulously curated and continuously updated reading list dedicated to the advancements in Optical Character Recognition (OCR) within the era of large language models (LLMs). This comprehensive resource focuses on research from the past five years (2021–now), highlighting how large vision-language and multimodal models are transforming the landscape of text-rich image and document processing. It serves as an essential guide for researchers and practitioners interested in the intersection of computer vision, natural language processing, and document intelligence.\n\n## Exploring the Repository\n\nThis repository is primarily a curated list of research papers and projects, organized into several key categories. It does not require traditional software installation. Users can simply navigate the repository on GitHub to explore the extensive collection of links to papers, code, and related resources.\n\n## Key Areas and Examples\n\nAWESOME-OCR-LLM delves into various critical aspects of OCR in the LLM era:\n\n*   **Document Parsing**: Focuses on converting complex documents into structured, machine-readable formats. Recent works include `PaDoc` for layout-grounded parallel decoding and `MonkeyOCRv2`, a visual-text foundation model for Document AI.\n*   **Document Understanding**: Extends beyond parsing to semantic comprehension and reasoning. Noteworthy projects include `DocMemo` for dynamic evidence discovery and `ADOPD 2026` for grounded and efficient document reasoning.\n*   **Visual Text Generation**: Explores generating or editing legible and visually consistent text within images. Examples like `PosterMELD` demonstrate multi-agent paper-to-poster generation, while `InnoText` offers a unified model for visual text generation and editing.\n*   **Specialized Models**: Covers traditional OCR research for specific visual-text structures, such as document dewarping (`DocRes`), mathematical expression recognition (`UniMERNet`), and table/chart understanding (`OmniParser`).\n*   **Benchmarks and Evaluation**: Provides insights into critical benchmarks shaping the field, including `ConfBench` for key information extraction and `XL-DocBench` for extra-long document understanding.\n\nThe \"Emerging Trends\" section offers valuable foresight, noting shifts towards end-to-end VLM-based parsing, reinforcement learning for layout, OCR-free document understanding, and the rise of document agents.\n\n## Why Use AWESOME-OCR-LLM?\n\nFor anyone working with or researching OCR and large language models, AWESOME-OCR-LLM is an indispensable resource. It provides a centralized, up-to-date overview of the most impactful research, saving countless hours of searching. The repository's structured format makes it easy to track developments in specific sub-fields, understand emerging trends, and discover new tools and benchmarks. Its focus on the latest advancements ensures you stay at the forefront of multimodal document AI.\n\n## Links\n\n*   **GitHub Repository**: [https://github.com/Yuliang-Liu/AWESOME-OCR-LLM](https://github.com/Yuliang-Liu/AWESOME-OCR-LLM){:target=\"_blank\"}\n*   **Daily Papers Section**: [https://github.com/Yuliang-Liu/AWESOME-OCR-LLM#daily-papers-section](https://github.com/Yuliang-Liu/AWESOME-OCR-LLM#daily-papers-section){:target=\"_blank\"}\n*   **Emerging Trends Section**: [https://github.com/Yuliang-Liu/AWESOME-OCR-LLM#%EF%B8%8F-emerging-trends](https://github.com/Yuliang-Liu/AWESOME-OCR-LLM#%EF%B8%8F-emerging-trends){:target=\"_blank\"}","metrics":{"detailViews":0,"githubClicks":0},"dates":{"published":null,"modified":"2026-08-11T19:17:20.000Z"}}