LightOnOCR-3 Review: A 2.7GB Local OCR Model Changing Workflows
Discover LightOnOCR-3: a local OCR model processing PDFs with diagrams and tables in one pass. Free, fast, and accurate!
Extracting data from unstructured PDFs traditionally requires a suite of tools. However, the new LightOnOCR-3 model handles all information locally in a single pass, consolidating various processing stages. This review will test the model on a document featuring visual diagrams, plain text, and complex tables to assess its ability to process these elements locally and deliver structured reports.
LightOnOCR-3 is part of the "Build with AI" series, focusing on local AI tools and models. The model offers instant text reading directly on the user's computer. Unlike legacy OCR pipelines that often falter with document format changes, this unified system fully processes data reading, transforming dense visual diagrams into usable data.
Performance and Speed
According to test results, LightOnOCR-3 ranks first among OCR models with an average score of 78.5%. Reading speed is a critical factor for processing large documents, and this model operates twice as fast as standard counterparts, halving the processing time for typical pages. Unlike commercial tools that charge per scanned page, LightOnOCR-3 operates entirely locally, eliminating additional costs.
System Requirements and Installation
The model requires less than 3 GB of RAM and system resources, allowing it to run easily on standard laptops. GGF files, containing the compressed model weights, are available for installation. Quantization options range from 16-bit to 4-bit, with file sizes from 2 to 8 GB and corresponding memory requirements. The 2 GB download requires a minimum of 3 GB of VRAM. 4-bit quantization has been selected for testing, offering good quality and being standard for most use cases.
To enable multimedia capabilities, in addition to the language model, an additional file of approximately 0.5 GB needs to be downloaded, which allows the model to understand visual diagrams. Google Colab is used as the development environment, but it is also possible to work with VS Code or other editors.
GGF files are downloaded using Hugging Face (HF) download commands. Two commands are used for this, downloading the 4-bit language model and the multimedia module into the models subfolder.
# Example command to download the language model (replace with the actual one)
# huggingface-cli download <language_model_name> --local-dir models --local-dir-use-symlinks False
# Example command to download the multimedia module (replace with the actual one)
# huggingface-cli download <multimedia_model_name> --local-dir models --local-dir-use-symlinks False
How LightOnOCR-3 Works
The processing workflow begins with the addition of source files (photos or documents). It is important to ensure sufficient image quality. The system analyzes the visual structure of the page, identifying paragraphs and elements to understand the document's organization. Then, the main engine reads the text and processes the information continuously. The model determines the coordinates of each element and outputs the extracted data in structured formats, such as standard text files.
LightOnOCR-3 eliminates the need for separate extraction tools, as a single engine simultaneously handles layout detection and cropping. The model is capable of converting visual data, including charts from research papers, into standard web tables. Since everything runs locally, sensitive financial and legal documents can be processed securely.
Registration and Model Launch
After loading the GGF files, they are registered using Ollama software, which allows AI models to be run locally. To install Ollama, navigate to ollama.com/download and download the appropriate installer.
To register a custom model file, an Ollama command is used, pointing to the GGF files. A standard chat template is also applied for correct understanding of system instructions and user requests. The IM_START and IM_END tags denote the beginning and end of a message, while ASSISTANT_PROMPT indicates the place for the model's response.
# Example command to register a model via Ollama
# ollama create light-on-ocr-tool -f ./Modelfile
The main Python file uses the registered model's name for complete document analysis. An outputs folder is created, and the output filename is defined. The Ollama client is used to set the task: extracting key facts, transcribing text, determining element coordinates, analyzing architectural diagrams, and extracting tables with test results.
# Example code for interacting with the Ollama model
from ollama import Client
client = Client(host='http://localhost:11434') # Specify your Ollama host
model_name = 'light-on-ocr-tool'
output_folder = 'outputs'
output_filename = 'report.md'
# Create the outputs folder if it doesn't exist
import os
if not os.path.exists(output_folder):
os.makedirs(output_folder)
full_output_path = os.path.join(output_folder, output_filename)
# Example prompt for data extraction
prompt = """
Extract all key facts, transcribe the full text, determine element coordinates,
describe architectural diagrams, and extract tables with test results.
"""
messages = [
{'role': 'system', 'content': 'You are an assistant that helps analyze documents.'},
{'role': 'user', 'content': prompt}
]
response = client.chat(model=model_name, messages=messages)
# Write the response to a file
with open(full_output_path, 'w', encoding='utf-8') as f:
f.write(response['message']['content'])
print(f"Report saved to: {full_output_path}")
The input data used is the LightOnOCR research paper, containing an introduction, an architectural diagram, and a table with test results. The user can substitute any document and adapt the prompt to their specific tasks.
The entire pipeline is launched by running the command python main.py in the terminal.
Operating Modes
There are two operating modes:
- Basic Reading: Fast processing of pages to obtain plain text, useful for internal research tools.
- Detailed Analysis: The model identifies special details, the location of images and diagrams, providing complex structural output. To activate this mode, the
groundingparameter is added to the prompt.
The grounding mode increases memory usage by approximately 25% due to the need to map each location on the page.
Test Results
When tested on complex layouts, the model demonstrates a performance of 74.1%, leading in handwritten text recognition and complex multilingual layout tasks. In visual element extraction, the model achieves an overall score of 75%, outperforming open-source alternatives. The model's general capabilities are rated at over 86%, surpassing well-known models like Chundra 2 and Mistral OCR.
The 4-billion parameter version of the model ranks first with an average score of around 79% globally. The 8-billion parameter version also shows a high result of 77%. Both versions outperform 35-billion parameter models, demonstrating better performance with lower computational costs. Even paid commercial tools like Mistral OCR show lower results compared to local open-source models.
Output Analysis
The output file report.md contains the coordinates of various document sections, including headings, plain text, and visual elements. For the visual figure, a semantic schema of the LightOnOCR architecture is presented, encompassing the use of vision encoders and multimodal data, as well as a language core for text comprehension. Captions for the diagram and a table with test results are provided. The table includes extracted numerical data for comparing open and closed APIs.
Memory Requirements and Recommendations
- Language Core: 4-bit quantization, ~3GB download size.
- Visual Understanding: Additional ~700MB file.
- Total Memory Usage: Less than 4.5GB (combined system and graphics).
Input quality significantly impacts recognition accuracy. High-resolution document scans, with the longest image side around 2000 pixels, are recommended. Local execution requires standard terminal tools like Ollama and llama.cpp.
Advantages and Disadvantages
Advantages:
- High recognition accuracy, surpassing even large models.
- Unified processing, eliminating the need for complex pipelines.
- Free for commercial use due to its open license.
- No monthly fees for processing large volumes of documents.
Disadvantages:
- The
groundingmode increases token count and memory requirements by 25%. - Input files must be prepared correctly, scaling them to 248 pixels to prevent memory failures.
Conclusion
LightOnOCR-3 is a free, unified local data extraction solution that eliminates the need for expensive cloud APIs. For installation and setup, it is recommended to refer to the README file at the first link in the description and pinned comments, which contains commands and step-by-step instructions, as well as the code used in this review.
Introducing LightOnOCR-3
Introducing LightOnOCR-3, a local OCR model that simplifies data extraction from complex documents like PDFs with diagrams and tables by processing them in a single pass. The 'Build with AI' series focuses on free, local AI tools.
- LightOnOCR-3 processes data from complex PDFs in a single pass.
- The model runs locally, enabling instant text reading.
- A unified system handles dense visual diagrams and converts them into data.
- The model ranks first in performance tests with a score of 78.5%.
- It boasts twice the reading speed of its counterparts.
- It runs entirely locally, with no per-page scanning costs.
- It requires less than 3GB of RAM, making it suitable for regular laptops.
Downloading and Installing LightOnOCR-3
A guide to downloading and setting up LightOnOCR-3, covering the selection of GGUF files (from 16-bit to 4-bit quantization) and files for multimedia capabilities. The installation process using Ollama and custom model file configuration is described.
- GGUF files are available in sizes from 2GB to 8GB.
- A 2GB download requires a minimum of 3GB VRAM.
- 4-bit quantization is recommended for good quality.
- An additional 0.5GB file is needed for multimedia capabilities.
- HF download commands are used for downloading.
- Installation is done using Ollama (ollama.com/download).
- Configuration involves creating a custom model file and specifying the path to GGUF files.
LightOnOCR-3 Workflow
Explaining the LightOnOCR-3 workflow: from file upload to layout analysis, text extraction, element mapping, and structured data output. The model handles layout detection and cropping simultaneously.
- The system analyzes the visual layout of the page, detecting paragraphs and elements.
- The core engine reads and processes written content without delay.
- The model determines the coordinates of each element.
- Structured text files are output.
- The model handles layout detection and cropping simultaneously.
- Converts charts from research papers into standard web tables.
- Processing of sensitive files occurs locally on the user's hardware.
Practical Application and Operating Modes
Demonstration of using LightOnOCR-3 for document analysis with Python and Ollama. Two modes are described: basic reading and detailed mapping (requires more resources).
- Ollama is used to run the model locally.
- A Python script is used to perform document analysis.
- An output folder is created and the output file name is defined.
- Tasks include fact extraction, text transcription, layout detection, and table extraction.
- Two modes: basic reading (fast, plain text) and detailed mapping (accurate image/diagram locations).
- Detailed mapping mode increases memory usage by 25%.
Performance Comparison and Benchmarks
Performance analysis of LightOnOCR-3 compared to other models. The model demonstrates high accuracy in complex layouts, visual data extraction, and outperforms commercial alternatives.
- Performance in complex layouts: 74.1%.
- Leads in handwritten text and complex layout recognition.
- Achieves 75% in visual data extraction.
- Overall capabilities: over 86%, surpassing Chundra 2 and Mistral OCR.
- The 4-billion version reaches nearly 79% globally.
- The 8-billion version achieves 77%, outperforming 35-billion models.
- Open-source local models outperform paid cloud APIs.
Results Analysis and Technical Details
Review of LightOnOCR-3 results using a research document as an example. Shows extracted section coordinates, descriptions of visual elements, captions for diagrams and tables, and a summary of memory requirements and input quality recommendations.
- Output file contains section coordinates (header, text, diagrams, tables).
- Content extracted from various sections.
- LightOnOCR-3 architecture described (visual and text encoders).
- Captions for diagrams and tables provided.
- Data fully extracted from the benchmark results table.
- Memory requirements: 3 GB for language core (4-bit) + 700 MB for visual understanding.
- Total memory usage: less than 4.5 GB.
- Input quality: scans should be high quality, with a side length of around 2000 pixels.
Conclusion: Pros, Cons, and Recommendations
Summary of the pros and cons of LightOnOCR-3, including recommendations for use, scaling files, and deployment. Highlights free commercial use and ease of setup.
- Pros: high accuracy, outperforms large models, unified process, free for commercial use (open license).
- Modes of operation: simple (empty prompt) or with detail (grounding command).
- The 'grounding' command increases token count and memory usage by 25%.
- Input recommendations: scale files to 2048 pixels for stability.
- Deployment: simple terminal commands using Ollama and llama.cpp.
- Offers a free, local solution for data extraction, replacing paid cloud APIs.
