nowfound

Alternatives

Products that do what Docuglean – Extract Structured Data from PDFs/Images Using AI does

Hi HN! I built Docuglean, an open-source SDK for intelligent document processing that works with OpenAI, Mistral, Google Gemini, and Hugging Face models. The idea came from repeatedly writing boilerplate code to extract structured data from invoices, receipts, and other documents. Instead of wrestling with different API formats, I wanted a unified interface that: - Extracts structured data using Zod/Pydantic schemas - Classifies and splits multi-section documents (e.g., medical records) - Processes documents in batches with automatic error handling - Works locally without APIs (for…

  1. 1

    Extract structured data from any document in 3 lines

    Dec 2025

  2. 2DO

    Documind is an open-source tool that turns documents into structured data using AI. What it does: - Extracts specific data from PDFs based on your custom schema - Returns clean, structured JSON that's ready to use - Works with just a PDF link + your schema definition Just run npm install documind to get started.

    2024 · github.com

  3. 3
    PDF.ai678

    Chat with any document

    2023 · pdf.ai

  4. 4
    DOConvert177

    Intelligence document processing platform that extracts data

    2024

  5. 5
    PDF Dino155

    Data extraction tool for PDF files

    2025

  6. 6

    Free tool to extract tables from PDF and Images

    2021

  7. 7
    Documind250

    ChatGPT for your documents

    2023

  8. 8CA

    ChunkHound’s goal is simple: local-first codebase intelligence that helps you pull deep, core-dev-level insights on demand, generate always-up-to-date docs, and scale from small repos to enterprise monorepos — while staying free + open source and provider-agnostic (VoyageAI / OpenAI / Qwen3, Anthropic / OpenAI / Gemini / Grok, and more). I’d love your feedback — and if you have, thank you for being part of the journey!

    Jan 2026 · github.com

  9. 9

    AI-powered receipt & invoice extraction for developers

    2025

  10. 10
    Crawlify165

    AI powered data extraction APIs. Hassle-free data retrieval.

    2020

  11. 11IJ

    Hi HackerNews, Lately, I have seen an explosion in posts offering paid APIs/services to get unstructured data into LLMs (i.e. langchain extract, ragflow, unstructured, unstract, just to name a few) and I have been largely disappointed by them, either because they fail to implement multimodal support, fail to give good context for "really tricky" PDFs / Word docs / Powerpoints, or are just plain difficult to use. In light of all these posts I figured I'd share my solution that has been working smoothly for me and my clients. I put it up on GitHub for free so you can check it…

    2024 · github.com

  12. 12

    AI That Works With Your Documents

    Sep 2025

  13. 13

    Turn messy PDFs into clean, structured data automatically

    Jan 2026

  14. 14AA

    Hey HN! Over the weekend (leaning heavily on Opus 4.5) I wrote Jargon - an AI-managed zettelkasten that reads articles, papers, and YouTube videos, extracts the key ideas, and automatically links related concepts together. Demo video: https://youtu.be/W7ejMqZ6EUQ Repo: https://github.com/schoblaska/jargon You can paste an article, PDF link, or YouTube video to parse, or ask questions directly and it'll find its own content. Sources get summarized, broken into insight cards, and embedded for semantic search. Similar ideas automatically cluster together. Each…

    Dec 2025 · github.com

  15. 15BA

    Hey HN, solo dev here. After years of frustration with how LLMs handle complex documents, especially PDFs with tables, I decided to build a solution myself. My approach uses a Markdown conversion step to preserve the table structure, which seems to work surprisingly well for chunking. This little parser is the first public piece of a much larger, privacy-focused AI platform I'm building. I'm pretty much running on fumes financially, so any feedback, critique, or support is massively appreciated. Happy to answer any questions about the approach!

    Nov 2025 · github.com

  16. 16

    AI document extraction with human-in-the-loop review

    Feb 2026 · docuct.ai

  17. 17DC

    Uses hybrid semantic search (combination of dense embeddings and sparse vectors) to retrieve high quality answers across your documents. Features - Significantly faster than competition (Process a 200 page PDF in <5s) - Much better answer quality - Fast summarization tool - Beta API for end to end extractive document QA ([email protected]) Try it out (no login) - Llama 2 paper https:&#x2F;&#x2F;www.dankgpt.com&#x2F;chat&#x2F;346f444d-e286-4671-b157-540f4cb... - Scott Aaronson Quantum Information Science lectures…

    2023 · dankgpt.com

  18. 18WH
  19. 19JA

    aiPDF is your AI assistant that can scan, understand and "chat" with all your documents. It summarises massive docs in seconds and finds any information you want. It works with any file type, web articles and even with YouTube videos!

    2024 · aipdf.ai

  20. 20

    ExtractBench - A Benchmark for Schema-Guided Enterprise Document Extraction - run-llama/ExtractBench

    26d ago · github.com

  21. 21

    Turn any document into actionable text with AI-powered OCR.

    Mar 2026 · documonk.pro

  22. 22IB

    I was experimenting with building a local dataset generator with deep research workflow a while back and that got me thinking. what if the same workflow could run on my own files instead of the internet. being able to query pdfs, docs or notes and get back a structured report sounded useful. so I made a small terminal tool that does exactly that. I point it to local files like pdf, docx, txt or jpg. it extracts the text, splits it into chunks, runs semantic search, builds a structure from my query, and then writes out a markdown report section by section. it feels like having a lightweight…

    2025 · github.com

  23. 23

    Transform documents into structured data with AI

    Jul 2026 · synapseidp.com

  24. 24DA

    I’m one of the co-founders of Doctly AI. I wanted to share our story. We didn’t originally set out to build a PDF-to-Markdown parser. It all started when we were building a RAG solution for a company that deals with regulatory agencies. All of their data was in PDFs, and as it is apparently with lawyers, they like to print and scan documents to make it hard on their counterparts. These documents contained complex tables that barely make sense, are rotated, and handwriting is mixed in between. Many pages are number ruled and potentially rotated. We spent a lot of time trying to get clean data…

    2024

Ranked by how close each launch is in meaning, then by votes. Refine with a description →