Alternatives
Products that do what DocPeel does
Extract structured data from PDFs, images, and emails
- 1DO
Documind is an open-source tool that turns documents into structured data using AI. What it does: - Extracts specific data from PDFs based on your custom schema - Returns clean, structured JSON that's ready to use - Works with just a PDF link + your schema definition Just run npm install documind to get started.
2024 · github.com
- 2
- 3

- 4

- 5
- 6

- 7

- 8

- 9

- 10

- 11

Extract web data into structured JSON, no scraper required.
Jun 2026 · tabstack.ai
- 12

extracts exactly the data you ask for emails, PDFs, etc
Jun 2026 · noima.io
- 13CP
2016 · docparser.com
- 14

- 15

Turn any bank statement PDF into Excel, CSV or JSON with AI
May 2026 · bankstatementlab.com
- 16

- 17

- 18DE
Hi HN! I built Docuglean, an open-source SDK for intelligent document processing that works with OpenAI, Mistral, Google Gemini, and Hugging Face models. The idea came from repeatedly writing boilerplate code to extract structured data from invoices, receipts, and other documents. Instead of wrestling with different API formats, I wanted a unified interface that: - Extracts structured data using Zod/Pydantic schemas - Classifies and splits multi-section documents (e.g., medical records) - Processes documents in batches with automatic error handling - Works locally without APIs (for…
Nov 2025 · github.com
- 19IJ
Hi HackerNews, Lately, I have seen an explosion in posts offering paid APIs/services to get unstructured data into LLMs (i.e. langchain extract, ragflow, unstructured, unstract, just to name a few) and I have been largely disappointed by them, either because they fail to implement multimodal support, fail to give good context for "really tricky" PDFs / Word docs / Powerpoints, or are just plain difficult to use. In light of all these posts I figured I'd share my solution that has been working smoothly for me and my clients. I put it up on GitHub for free so you can check it…
2024 · github.com
- 20

- 21

Extract structured data from any document in seconds
Apr 2026 · balanced-alignment-production-b954.up.railway.app
- 22

- 23FA
Hi HN, Like everyone, I'm working on an product that uses LLMs to extract data from photos and documents. Part of the processing pipeline is extracting data from PDFs as raw text or a raster image. As part of our leadgen strategy, we've opened our REST API that lets you process pages of a PDF. The API is completely free to use anonymously, but is rate limited to 1 page per 30 seconds. Creating a free account removes this restriction. The two endpoints are: - https://extract.dev/api/pages/extract/raster - Rasterize a page of a PDF -…
Oct 2025
- 24

Ranked by how close each launch is in meaning, then by votes. Refine with a description →