nowfound

Alternatives

Products that do what Docuglean does

Extract structured data from any document in 3 lines

  1. 1DE

    Hi HN! I built Docuglean, an open-source SDK for intelligent document processing that works with OpenAI, Mistral, Google Gemini, and Hugging Face models. The idea came from repeatedly writing boilerplate code to extract structured data from invoices, receipts, and other documents. Instead of wrestling with different API formats, I wanted a unified interface that: - Extracts structured data using Zod/Pydantic schemas - Classifies and splits multi-section documents (e.g., medical records) - Processes documents in batches with automatic error handling - Works locally without APIs (for…

    Nov 2025 · github.com

  2. 2DO

    Documind is an open-source tool that turns documents into structured data using AI. What it does: - Extracts specific data from PDFs based on your custom schema - Returns clean, structured JSON that's ready to use - Works with just a PDF link + your schema definition Just run npm install documind to get started.

    2024 · github.com

  3. 3
    PDF Dino155

    Data extraction tool for PDF files

    2025

  4. 4

    Free tool to extract tables from PDF and Images

    2021

  5. 5
    Docparser194

    Convert PDFs and scanned documents into structured data

    2016

  6. 6
    DOConvert177

    Intelligence document processing platform that extracts data

    2024

  7. 7

    Extract structured data from text, files and archives.

    Mar 2026 · apps.apple.com

  8. 8

    Extract web data into structured JSON, no scraper required.

    Jun 2026 · tabstack.ai

  9. 9
    Documind250

    ChatGPT for your documents

    2023

  10. 10
    Koncile 251

    Customisable OCR for all your data extraction needs

    2024

  11. 11

    Ditch your scraper. Make one API call with any tool.

    Jun 2026 · tabstack.ai

  12. 12

    AI-powered receipt & invoice extraction for developers

    2025

  13. 13KM

    I'm excited to showcase Kreuzberg! Kreuzberg is a modern Python library built from the ground up with async/await, type hints, and optimized I/O handling. It provides a unified interface for extracting text from documents (PDFs, images, office files) without external API dependencies. Key technical features: - Built with modern Python best practices (async/await, type hints, functional-first) - Optimized async I/O with anyio for multi-loop compatibility - Smart worker process pool for CPU-bound tasks (OCR, doc conversion) - Efficient batch processing with concurrent…

    2025 · github.com

  14. 14DS
  15. 15
    Crawlify165

    AI powered data extraction APIs. Hassle-free data retrieval.

    2020

  16. 16CA

    ChunkHound’s goal is simple: local-first codebase intelligence that helps you pull deep, core-dev-level insights on demand, generate always-up-to-date docs, and scale from small repos to enterprise monorepos — while staying free + open source and provider-agnostic (VoyageAI / OpenAI / Qwen3, Anthropic / OpenAI / Gemini / Grok, and more). I’d love your feedback — and if you have, thank you for being part of the journey!

    Jan 2026 · github.com

  17. 17OP
  18. 18
    Docsumo99

    Automate data entry while processing documents ✌️

    2019

  19. 19
    docWind73

    An easy to use document scanner

    2020

  20. 20

    Turn messy PDFs into clean, structured data automatically

    Jan 2026 · extractifyhq.com

  21. 21

    Precise document extraction for your agents — zero retention

    Apr 2026 · canonizr.com

  22. 22D2

    Hi! We are excited to announce the second release of Desbordante — an open-source, high-performance data profiler that is capable of discovering and validating many different patterns in data using various algorithms. Unlike existing data profilers, Desbordante focuses on discovering complex patterns in data, which are notoriously hard to extract. Since its inception in 2019, it has become the fastest open-source tool for these tasks. It also offers an array of patterns which have no alternative implementations. With this release, Desbordante now supports 17 types of patterns, such as:…

    2024 · github.com

  23. 23SR
  24. 24OS

    Hi, I'm building an open-source self-hostable document extraction tool powered by LLM. There are popular data extraction tools and OCR tools in the market. None of them are open source. Most accounting firms, law firms, insurance, back office, and real-estate folks would like to use a tool like this. You can add PDF documents, Images, and audio files and create columns to answer questions on documents or extract information into tabular format. Access to repo: https://github.com/harishdeivanayagam/rowfill Screenshots:…

    2025 · github.com

Ranked by how close each launch is in meaning, then by votes. Refine with a description →