Alternatives
Products that do what Differy does
A simple PDF-Analysis platform for complex usecases.
- 1PDPDF Differ▲62
2023 · pdfdiffer.com
- 2PD
2022 · github.com
- 3OP
Jan 2026 · github.com
- 4PT
I've developed a Python API service that uses GPT-4o for OCR on PDFs. It features parallel processing and batch handling for improved performance. Not only does it convert PDF to markdown, but it also describes the images within the PDF using captions like `[Image: This picture shows 4 people waving]`. In testing with NASA's Apollo 17 flight documents, it successfully converted complex, multi-oriented pages into well-structured Markdown. The project is open-source and available on GitHub. Feedback is welcome.
2024 · github.com
- 5KM
I'm excited to showcase Kreuzberg! Kreuzberg is a modern Python library built from the ground up with async/await, type hints, and optimized I/O handling. It provides a unified interface for extracting text from documents (PDFs, images, office files) without external API dependencies. Key technical features: - Built with modern Python best practices (async/await, type hints, functional-first) - Optimized async I/O with anyio for multi-loop compatibility - Smart worker process pool for CPU-bound tasks (OCR, doc conversion) - Efficient batch processing with concurrent…
2025 · github.com
- 6

- 7PT
Hi, OP here. A friend was involved in a custody battle and was afraid his ex was going to leak all of his discovery documents on the internet and he asked if there was something I could do to make it harder for bots/crawlers to find sensitive information. Originally I was going to turn all of his docs to image based PDFs, but those get large fast and are easy to OCR. So I found a post musing about altering fonts/glyphs so that it looks like english, but the actual character being seen by the pdf reader is a non-english character. As such, when you try to OCR these files, it doesn't…
2022 · humaneyesonly.com
- 8DO
Documind is an open-source tool that turns documents into structured data using AI. What it does: - Extracts specific data from PDFs based on your custom schema - Returns clean, structured JSON that's ready to use - Works with just a PDF link + your schema definition Just run npm install documind to get started.
2024 · github.com
- 9

- 10DD
Gleb, Alex, Erez and Simon here – we are building an open-source tool for comparing data within and across databases at any scale. The repo is at https://github.com/datafold/data-diff, and our home page is https://datafold.com/. As a company, Datafold builds tools for data engineers to automate the most tedious and error-prone tasks falling through the cracks of the modern data stack, such as data testing and lineage. We launched two years ago with a tool for regression-testing changes to ETL code…
2022
- 11

- 12

- 13

- 14

- 15IJ
Hi HackerNews, Lately, I have seen an explosion in posts offering paid APIs/services to get unstructured data into LLMs (i.e. langchain extract, ragflow, unstructured, unstract, just to name a few) and I have been largely disappointed by them, either because they fail to implement multimodal support, fail to give good context for "really tricky" PDFs / Word docs / Powerpoints, or are just plain difficult to use. In light of all these posts I figured I'd share my solution that has been working smoothly for me and my clients. I put it up on GitHub for free so you can check it…
2024 · github.com
- 16PT
Hi HN! We’re Marc and Iuliia from Papermark (https://papermark.io). We're building an open-source, modern document sharing platform with real-time engagement analytics and 100% customization. It all started as a tweet [1] and led to our launch on Product Hunt [2] last month. We crossed over 1000 stars and over 10 contributors on GitHub (https://github.com/mfts/papermark). Incumbents, like DocSend, founded in the early 2010s have been acquired already and just don't innovate anymore. Their main priority is enterprise clients. As founders and developers ourselves…
2023 · papermark.io
- 17DO
2021 · chrome.google.com
- 18

- 19DT
Most products that touch PDFs or images quietly rebuild the same thing: a hacked-together “router” that picks which OCR/vision API to call, normalizes the responses, and prays the bill is sane at the end of the month. DocsRouter is that layer as a product: one stable API that talks to multiple OCR engines and vision LLMs, lets you route per document based on cost/quality/latency, and gives you normalized outputs (text, tables, fields) so your app doesn’t care which provider was used. It’s meant for teams doing serious stuff with documents: invoices/receipts, contracts,…
Dec 2025 · docsrouter.com
- 20

Hi HN! I built PDFly — a PDF toolkit that runs entirely client-side. No file uploads, no server processing, no account needed. It has 56 tools including merge, split, compress, convert, redact, compare, and an AI chat feature to ask questions about your PDF content. Everything runs in your browser using [mention your tech stack if relevant, e.g., WASM/JS libraries]. Would love feedback on the tools, UX, or anything else!
Jul 2026 · pdfly.uk
- 21

- 22YS
2025 · trycardinal.ai
- 23EE
2024 · endtype.com
- 24
Ranked by how close each launch is in meaning, then by votes. Refine with a description →