Alternatives
Products that do what Holofin does
Turn messy PDFs into structured, verified data
- 1
- 2UE
2023 · github.com
- 3OP
Jan 2026 · github.com
- 4

- 5

- 6

- 7

- 8

- 9

- 10

Turn any bank statement PDF into Excel, CSV or JSON with AI
May 2026 · bankstatementlab.com
- 11

- 12AN
When building workflows that rely on LLMs, we commonly use structured output for programmatic use cases like converting an invoice into rows or meeting transcripts into tickets or even complex PDFs into database entries. The model may return the schema you want, but with hallucinated values like `invoice_date` being off by 2 months or the transcript array ordered wrongly. The JSON is valid, but the values are not. Structured output today is a big part of using LLMs, especially when building deterministic workflows. Current structured output benchmarks (e.g., JSONSchemaBench) only validate…
Apr 2026 · interfaze.ai
- 13

- 14

- 15

Extract data from PDFs, invoices, receipts & bank statements
Jun 2026 · billsdeck.com
- 16UA
2022 · unblob.org
- 17

- 18

- 19

- 20OS
Hi, I'm building an open-source self-hostable document extraction tool powered by LLM. There are popular data extraction tools and OCR tools in the market. None of them are open source. Most accounting firms, law firms, insurance, back office, and real-estate folks would like to use a tool like this. You can add PDF documents, Images, and audio files and create columns to answer questions on documents or extract information into tabular format. Access to repo: https://github.com/harishdeivanayagam/rowfill Screenshots:…
2025 · github.com
- 21BA
Hey HN, solo dev here. After years of frustration with how LLMs handle complex documents, especially PDFs with tables, I decided to build a solution myself. My approach uses a Markdown conversion step to preserve the table structure, which seems to work surprisingly well for chunking. This little parser is the first public piece of a much larger, privacy-focused AI platform I'm building. I'm pretty much running on fumes financially, so any feedback, critique, or support is massively appreciated. Happy to answer any questions about the approach!
Nov 2025 · github.com
- 22TR
Hey HN, Today, we’re launching tile.run, an API that extracts structured data from unstructured documents (PDF, images, text) with support for custom schemas. The Problem: Extracting data out of unstructured documents is surprisingly hard. We built tile.run while solving this for our product Kili (automation for invoicing/reconciliation). We found that getting to accuracy that is reliable enough for automation is challenging. Dense documents (e.g., lots of tables or line items) are even harder, and these are the most valuable to automate. After talking to other teams and developers, we…
2024 · tile.run
- 23DK
Hello HN, I’ve built DocsBoard – a platform that aims to solve the chaos of project documentation. If you’re tired of searching for relevant documents, dealing with outdated information, or struggling with onboarding new team members, this might be for you. Works with Google Docs, spreadsheets, slides, PDFs, Markdown and more. How it works: • Create a new board for your team’s project. • Import your documents in the format you already have, without any conversions. • Drag-and-drop to easily organize documents into groups and keep track of them. • Invite your team members! Try it out and let…
2024 · docsboard.io
- 24DE
Hi HN! I built Docuglean, an open-source SDK for intelligent document processing that works with OpenAI, Mistral, Google Gemini, and Hugging Face models. The idea came from repeatedly writing boilerplate code to extract structured data from invoices, receipts, and other documents. Instead of wrestling with different API formats, I wanted a unified interface that: - Extracts structured data using Zod/Pydantic schemas - Classifies and splits multi-section documents (e.g., medical records) - Processes documents in batches with automatic error handling - Works locally without APIs (for…
Nov 2025 · github.com
Ranked by how close each launch is in meaning, then by votes. Refine with a description →