nowfound

Alternatives

Products that do what Canonizr does

Precise document extraction for your agents — zero retention

  1. 1
    moar95

    Your documents. AI ready.

    May 2026 · getmoar.ai

  2. 2DO

    Documind is an open-source tool that turns documents into structured data using AI. What it does: - Extracts specific data from PDFs based on your custom schema - Returns clean, structured JSON that's ready to use - Works with just a PDF link + your schema definition Just run npm install documind to get started.

    2024 · github.com

  3. 3
    Extend90

    Parse any PDF layout with SOTA accuracy for AI pipelines

    May 2026 · extend.ai

  4. 4

    AI-powered receipt & invoice extraction for developers

    2025

  5. 5

    Turn hundreds of documents into one clean spreadsheet

    Feb 2026 · nolainocr.com

  6. 6

    Extract data from ANY document faster

    2021

  7. 7

    Make the world's documents computable

    Jun 2026 · landing.ai

  8. 8IJ

    Hi HackerNews, Lately, I have seen an explosion in posts offering paid APIs/services to get unstructured data into LLMs (i.e. langchain extract, ragflow, unstructured, unstract, just to name a few) and I have been largely disappointed by them, either because they fail to implement multimodal support, fail to give good context for "really tricky" PDFs / Word docs / Powerpoints, or are just plain difficult to use. In light of all these posts I figured I'd share my solution that has been working smoothly for me and my clients. I put it up on GitHub for free so you can check it…

    2024 · github.com

  9. 9DT

    Most products that touch PDFs or images quietly rebuild the same thing: a hacked-together “router” that picks which OCR/vision API to call, normalizes the responses, and prays the bill is sane at the end of the month. DocsRouter is that layer as a product: one stable API that talks to multiple OCR engines and vision LLMs, lets you route per document based on cost/quality/latency, and gives you normalized outputs (text, tables, fields) so your app doesn’t care which provider was used. It’s meant for teams doing serious stuff with documents: invoices/receipts, contracts,…

    Dec 2025 · docsrouter.com

  10. 10

    No Code, No Deployments. Ship your OCR solution in minutes.

    4d ago · parsely.studio

  11. 11OS

    Hi, I'm building an open-source self-hostable document extraction tool powered by LLM. There are popular data extraction tools and OCR tools in the market. None of them are open source. Most accounting firms, law firms, insurance, back office, and real-estate folks would like to use a tool like this. You can add PDF documents, Images, and audio files and create columns to answer questions on documents or extract information into tabular format. Access to repo: https://github.com/harishdeivanayagam/rowfill Screenshots:…

    2025 · github.com

  12. 12

    56 free PDF tools that run in your browser: merge, convert, compress, sign, OCR, plus on-device AI. Files never upload, a live meter proves it. Works offline.

    Jun 2026 · pdfmergely.com

  13. 13

    FREE AI-powered batch document cleaner

    Nov 2025

  14. 14BA

    Hey HN, solo dev here. After years of frustration with how LLMs handle complex documents, especially PDFs with tables, I decided to build a solution myself. My approach uses a Markdown conversion step to preserve the table structure, which seems to work surprisingly well for chunking. This little parser is the first public piece of a much larger, privacy-focused AI platform I'm building. I'm pretty much running on fumes financially, so any feedback, critique, or support is massively appreciated. Happy to answer any questions about the approach!

    Nov 2025 · github.com

  15. 15
    Censr4

    Redact & scrub PDF data. No Upload. No Archive. No Account.

    Jul 2026 · censr.tech

  16. 16IB

    I was experimenting with building a local dataset generator with deep research workflow a while back and that got me thinking. what if the same workflow could run on my own files instead of the internet. being able to query pdfs, docs or notes and get back a structured report sounded useful. so I made a small terminal tool that does exactly that. I point it to local files like pdf, docx, txt or jpg. it extracts the text, splits it into chunks, runs semantic search, builds a structure from my query, and then writes out a markdown report section by section. it feels like having a lightweight…

    2025 · github.com

  17. 17

    Extract data from any document — you define the fields

    May 2026 · dokuscan.com

  18. 18

    20+ free, private PDF & media tools. Zero data retention.

    Mar 2026

  19. 19

    Turn PDFs, DOCX, and URLs into clean markdown.

    Jul 2026 · tokendrop.tech

  20. 20PT

    Hi HN! I'm proud to share that we've launched a free PDF-to-Markdown CLI built on our proprietary (you might know it from PSPDFKit) engine. Most extractors are either fast but lose structure (markitdown, pymupdf4llm) or accurate but slow (docling). Ours ties with docling on accuracy but is orders of magnitude faster. https://github.com/pspdfkit/pdf-to-markdown We'd love feedback on it, and ofc send us files that break it.

    Apr 2026

  21. 21

    ExtractBench - A Benchmark for Schema-Guided Enterprise Document Extraction - run-llama/ExtractBench

    27d ago · github.com

  22. 22

    Build files, not just answers

    Jun 2026 · docanalyzer.ai

  23. 23

    Visual document data extraction API, zero training required

    2025

  24. 24

    Turn PDF Invoices into Data using Local AI Agents.

    Jan 2026 · github.com

Ranked by how close each launch is in meaning, then by votes. Refine with a description →