TeleOCR: The 1B OCR Model Challenging GPT-5.2 and Gemini!
Discover TeleOCR: a 1B parameter OCR model that runs locally, beats GPT-5.2 & Gemini, and handles complex docs offline! Free & private.
Cloud-based vision AI models are increasingly straining developer budgets. TeleOCR, a new, lightweight Optical Character Recognition (OCR) model, offers a potential solution by processing complex documents locally while demonstrating performance that rivals established models like GPT-5.2 and Gemini 3 Pro in benchmark tests. This model operates with approximately 1 million weights and can run without requiring external GPU cards, enabling direct processing of physical documents captured via phone cameras.
Core Functionality and Performance
TeleOCR is designed to convert images of documents into usable text blocks, categorizing numbers and text instantly. Testing results indicate high performance in standard visual evaluations, with the model achieving near-perfect accuracy on complex layouts according to the OmniDog Bench. The system requires minimal graphics card resources, with setups featuring as little as 2.5GB of graphics memory capable of running it without issues, significantly reducing hardware setup costs.
The model exhibits a low text error rate, rarely misreading printed characters or numbers, which minimizes the need for manual correction of spelling mistakes. A key advantage is its proficiency with real-world photographs, which often present challenges for other OCR tools. TeleOCR excels at processing raw camera captures, yielding clean text from ordinary pictures.
Technical Architecture and Workflow
The internal design of TeleOCR focuses on understanding complex document shapes and the interconnections between text blocks to maintain proper reading order. The process begins with a standard image input. The system then identifies and boxes every text area, mapping the page layout. Subsequently, it identifies table rows and columns, ultimately providing structured data output suitable for direct use in applications.
The model's training emphasizes high extraction accuracy by learning to interpret complex visual patterns, leading to improved results on dense documents. It also adapts to distorted shapes in folded or bent papers, preserving information from warped images. For tables, TeleOCR separates visual grid lines from text, preventing the merging of columns or misaligned rows and ensuring data integrity.
Local Deployment and Usage
TeleOCR can be downloaded and installed via its official Hugging Face page. The installation involves Python dependencies and downloading the model weights using an hf download command. Once downloaded, the model weights are loaded from a local directory.
The model's capabilities were demonstrated using four distinct input images:
- A customer receipt with complex text and sale values.
- A scanned letter detailing card services and offers.
- A scanned receipt of a purchase made with a Visa card, requiring extraction of card digits and merchant details.
- A table sample from the tool's benchmark numbers to assess table extraction accuracy.
The process involves running a document AI extraction pipeline in Python. Images are loaded, their byte data is passed to the model processor, and outputs are generated using a .generate() function. Structured reports are saved in JSON format (as Python dictionaries) and as formatted text reports in a markdown file, including a summary and detailed extractions. An SQLite database is also generated for structured data storage.
The command to run this pipeline from the terminal is python main.py.
Practical Applications
TeleOCR has several practical applications:
- Retail: Processing wrinkled or faded bills from physical stores without specialized scanning equipment, automatically extracting line items.
- Legal Documents: Handling multi-column legal papers where standard AI tools might mix text from adjacent columns. TeleOCR ensures sentences remain within their respective columns and preserves formatting across pages.
- Scientific Documents: Converting complex charts and dense formulas into web-ready code formats, transforming equations into visual diagrams for HTML dashboards.
- Offline Data Storage: Building fully offline storage systems by directly feeding extracted data into local databases, eliminating the need for cloud connections.
Benchmark Performance
In benchmark comparisons against paid tools, including Gemini 3 Pro and GPT 5.2, TeleOCR secures a top position on the OmniBench broad testing benchmark. It also performs well on medical benchmarks, handling complex formats effectively for specialized tasks in medical or technical documentation.
Output Analysis
The model's outputs are organized in an output folder, including structured reports, JSON files, and an SQLite database.
- Customer Receipt: Extracted total amount of 32.95, vendor "Buckingham Palace Garden Shop," line items, and a table grid in HTML format detailing store name, cashier, receipt number, transaction method (Visa card), and final sale total.
- Train Ticket Receipt: Extracted total amount of 40 from "Gateway Airports South Terminal 51," along with merchant details, transaction information, payment method, and total amounts in an HTML report.
- Platinum Card Benefits Letter: Extracted card details and benefits of different cards in an HTML report. Detected total was marked as "NA" due to the absence of price numbers.
- Benchmark Table: Accurately extracted rows and metrics from the benchmark table. For example, in the "hours" rows, it correctly identified values such as 80.68, 85.4, 82.97, and 12.68, with an F1 score of 82.97. The extracted data is provided in an HTML table format.
The JSON output provides a structured Python dictionary of all text extractions, suitable for code workflows or integration with other AI tools. The SQLite database contains all extracted details.
System Requirements and Enterprise Use
TeleOCR can be run locally on standard computers with 2.5GB of graphics card memory. The model can be loaded with a single command from the Hugging Face library. Optimized formats like Llama CPP and single model weight GGUF files are supported for offline use without internet connectivity. For enterprise-level applications, the model supports high-volume scanning and batch processing of document archives.
Advantages and Limitations
Pros:
- Data Privacy: Sensitive documents remain local, ensuring complete privacy as data never touches external systems.
- Cost-Effective: Eliminates operating costs associated with cloud-based solutions.
- Robustness: Handles imperfect lighting conditions and corrects perspective issues in casual photos, yielding professional results.
Cons:
- Manual Processing: Final outputs may require some manual processing as the model provides raw structural tags. Helper scripts can be used to convert these tags into front-end dashboards.
TeleOCR offers developers a highly accurate, local vision system for parsing complex documents without processing fees. Setup instructions and code files are available via a link in the description and pinned comments.
Introducing TeleOCR: Local, Lightweight, and Powerful OCR
Introduction to TeleOCR, a lightweight OCR model that runs locally, outperforming GPT-5.2 and Gemini 3 Pro in benchmarks. It handles complex documents and real-world photos without needing external GPUs, making it budget-friendly.
- TeleOCR is a new lightweight OCR model.
- It processes complex documents locally.
- It outperforms GPT-5.2 and Gemini 3 Pro in benchmarks.
- It requires no external GPU cards.
- It can run on hardware with as little as 2.5GB of graphics cards.
- It excels with real-world photographs and raw camera captures.
Installation and Internal Engine of TeleOCR
Demonstrates the installation process of TeleOCR using Hugging Face, including Python dependencies and model weight downloads, followed by an explanation of its internal engine design for text extraction and layout mapping.
- Installation involves
pip installfor Python dependencies. - Model weights are downloaded using
hf downloadcommand. - The engine focuses on complex shapes and text block connections.
- It boxes text areas, maps page layout, and identifies table rows/columns.
- It learns complex visual patterns for better extraction on dense documents.
- It adapts to distorted shapes from folded or bent papers.
Testing TeleOCR with Diverse Documents
Details the practical testing of TeleOCR with four diverse input images: a customer receipt, a scan letter, another scan receipt, and a table sample. Explains the Python script used for document AI extraction and the output formats (JSON, Markdown, SQLite).
- Input images include a customer receipt, scan letter, scan receipt, and table sample.
- The Python script uses
document AI extractpipeline. - Outputs include structured JSON, formatted Markdown reports, and an SQLite database.
- The script loads images, passes data to the model processor, and uses
.generate()function. - Outputs are saved in an
outputsfolder. - Running the pipeline requires the command
python main.pyin the terminal.
Practical Applications and Workflow Examples
Explores practical workflows for TeleOCR in various settings like retail, legal, and scientific fields. Highlights its ability to handle wrinkled bills, multi-column legal documents, and complex scientific formulas, while also supporting offline SQL database storage.
- Useful for physical stores with wrinkled or faded bills.
- Handles multi-column legal papers without mixing text.
- Converts scientific charts and formulas into web-ready code formats.
- Supports fully offline storage systems using local SQL databases.
- Maintains top spots in benchmarks against paid tools like Gemini 3 Pro and GPT 5.2.
- Handles complex formats in medical and technical documents.
Output Analysis and Verification
Presents the output analysis of the four test images, detailing extracted information like vendor names, total amounts, transaction details, card benefits, and benchmark numbers. Compares extracted table data with original benchmarks and discusses the JSON and SQLite outputs.
- Customer receipt: Extracted vendor (Buckingham Palace Garden Shop), total ($32.95), line items, and transaction details (Visa card).
- Train ticket receipt: Extracted total ($40) and merchant details.
- Scan letter: Extracted card benefits (NA for prices).
- Benchmark table: Accurately extracted all rows and metrics (e.g., recall, precision, F1 score).
- Outputs include structured JSON for code workflows and an SQLite database.
- HTML format provided for visual dashboards.
Pros, Cons, and Deployment Options
Summarizes the benefits of TeleOCR, including privacy, cost savings, and adaptability to poor lighting or document conditions. Discusses potential cons like the need for manual processing of raw output tags and provides information on local deployment options.
- Major benefit: Private data never touches external systems, ensuring complete privacy.
- Operating costs disappear entirely.
- No need for perfect lighting conditions; corrects perspective issues.
- Casual photos yield professional results.
- Potential con: Outputs require some manual processing for front-end dashboards.
- Can be run locally on standard computers with 2.5GB graphics cards.
- Supports optimized formats (Llama CPP, GGUF) for offline use.
- Supports high-volume scanning for enterprise setups.
