Alternatives
Products that do what aOCR does
API for converting complex documents into structured data
- 1OP
Jan 2026 · github.com
- 2

- 3

- 4

- 5

- 6

- 7

- 8

- 9OP
Hi HN, I’ve been working on an OCR pipeline specifically optimized for machine learning dataset preparation. It’s designed to process complex academic materials — including math formulas, tables, figures, and multilingual text — and output clean, structured formats like JSON and Markdown. Some features: • Multi-stage OCR combining DocLayout-YOLO, Google Vision, MathPix, and Gemini Pro Vision • Extracts and understands diagrams, tables, LaTeX-style math, and multilingual text (Japanese/Korean/English) • Highly tuned for ML training pipelines, including dataset generation and…
2025 · github.com
- 10

- 11

- 12

- 13

- 14CP
2016 · docparser.com
- 15

- 16

- 17

- 18

- 19

- 20OB
OCR/Document extraction field has seen lot of action recently with releases like Mixtral OCR, Andrew Ng's agentic document processing etc. Also there are several benchmarks for OCR, however all testing for something slightly different which make good comparison of models very hard. To give an example, some models like mixtral-ocr only try to convert a document to markdown format. You have to use another LLM on top of it to get the final result. Some VLM’s directly give structured information like key fields from documents like invoices, but you have to either add business rules on top…
2025 · nanonets.com
- 21SA
Hi HN, I built an AI-powered OCR API designed to extract highly structured JSON data from complex documents like global passports, IDs, receipts, and shipping containers. We recently rolled out our Python and Node.js SDKs. Just wanted to share it with the community.
Jun 2026 · structocr.com
- 22
- 23
- 24

Ranked by how close each launch is in meaning, then by votes. Refine with a description →