nowfound

Alternatives

Products that do what Corpus does

Structured research engine. Absurdly fast. Open source.

  1. 1
    Consensus409

    A search engine that finds you answers in scientific papers

    2022

  2. 2

    AI-native workbench for understanding papers

    2025

  3. 3DO

    Documind is an open-source tool that turns documents into structured data using AI. What it does: - Extracts specific data from PDFs based on your custom schema - Returns clean, structured JSON that's ready to use - Works with just a PDF link + your schema definition Just run npm install documind to get started.

    2024 · github.com

  4. 4NI

    Understanding scientific articles can be tough, even in your own field. Trying to comprehend articles from others? Good luck. Enter, Now I Get It! I made this app for curious people. Simply upload an article and after a few minutes you'll have an interactive web page showcasing the highlights. Generated pages are stored in the cloud and can be viewed from a gallery. Now I Get It! uses the best LLMs out there, which means the app will improve as AI improves. Free for now - it's capped at 20 articles per day so I don't burn cash. A few things I (and maybe you will) find interesting: * This is…

    Feb 2026 · nowigetit.us

  5. 5DT
  6. 6

    Get your unstructured data AI-ready in minutes

    2024

  7. 7

    Run a research agent with cited answers in a single API call

    Jun 2026 · tabstack.ai

  8. 8
    Rixx110

    The Perplexity alternative that organizes your research

    May 2026 · rixx.app

  9. 9SO

    Built a tool for transforming unstructured data into structured outputs using language models (with 100% adherence). If you're facing problems getting GPT to adhere to a schema (JSON, XML, etc.) or regex, need to bulk process some unstructured data, or generate synthetic data, check it out. We run our own tuned model (you can self-host if you want), so, we're able to have incredibly fine grained control over text generation. Repository: https://github.com/automorphic-ai/trex Playground: https://automorphic.ai/playground

    2023 · automorphic.ai

  10. 10

    Let swarm of agents automate your most manual PDF tasks

    2025

  11. 11TA
  12. 12IJ

    Hi HackerNews, Lately, I have seen an explosion in posts offering paid APIs/services to get unstructured data into LLMs (i.e. langchain extract, ragflow, unstructured, unstract, just to name a few) and I have been largely disappointed by them, either because they fail to implement multimodal support, fail to give good context for "really tricky" PDFs / Word docs / Powerpoints, or are just plain difficult to use. In light of all these posts I figured I'd share my solution that has been working smoothly for me and my clients. I put it up on GitHub for free so you can check it…

    2024 · github.com

  13. 13CW
  14. 14CE
  15. 15IB

    I was experimenting with building a local dataset generator with deep research workflow a while back and that got me thinking. what if the same workflow could run on my own files instead of the internet. being able to query pdfs, docs or notes and get back a structured report sounded useful. so I made a small terminal tool that does exactly that. I point it to local files like pdf, docx, txt or jpg. it extracts the text, splits it into chunks, runs semantic search, builds a structure from my query, and then writes out a markdown report section by section. it feels like having a lightweight…

    2025 · github.com

  16. 16GY

    Hey HN! We're excited to announce the launch of Tonic Textual, the secure data lakehouse for LLMs. Simply stated, Tonic Textual allows you to build generative AI systems on your own unstructured data without having to spend time extracting and standardizing your data. In minutes you can build automated, scalable unstructured data pipelines that extract, centralize, standardize, and enrich data from your documents into an AI-optimized format ready for embedding, fine-tuning, and ingesting into a vector database. While in-flight, we also scan for sensitive information and protect it via…

    2024 · tonic.ai

  17. 17

    Research Papers Made Plain

    Jan 2026 · papersplain.com

  18. 18WB

    Hey HN, Automated research is the next big step in AI, with companies like OpenAI aiming to debut a fully automated researcher by 2028 (https://www.technologyreview.com/2026/03/20/1134438/openai-i...). However, there is a very real possibility that much of this corporate research will remain closed to the general public. To counter this, we spent the last month building Enlidea---a machine-to-machine ecosystem for open research. It's a decentralized research hub where autonomous agents propose hypotheses, stake bounties, execute code, and perform automated…

    Mar 2026 · enlidea.com

  19. 19AA

    Hey HN! Over the weekend (leaning heavily on Opus 4.5) I wrote Jargon - an AI-managed zettelkasten that reads articles, papers, and YouTube videos, extracts the key ideas, and automatically links related concepts together. Demo video: https://youtu.be/W7ejMqZ6EUQ Repo: https://github.com/schoblaska/jargon You can paste an article, PDF link, or YouTube video to parse, or ask questions directly and it'll find its own content. Sources get summarized, broken into insight cards, and embedded for semantic search. Similar ideas automatically cluster together. Each…

    Dec 2025 · github.com

  20. 20OD

    I would like to share an open database focused on link-level metadata extraction and aggregation, which may be of interest to researchers. The project maintains a structured dataset of links enriched with metadata such as: - page title - description / summary - publication date (when available) - thumbnail / preview image - etc. The goal is to provide a reusable, inspectable set of link metadata that can be used for experiments in areas such as: - RSS and feed analysis - news analysis - link rot analysis? The database is publicly available here:…

    Jan 2026 · github.com

  21. 21
    Collate16

    Offline AI workspace for PDFs on your Mac

    2025

  22. 22

    A workbench for reading and understanding papers

    Jan 2026 · openpaper.ai

  23. 23CA

    Synthetic data generation is an essential step in training and evaluating LLMs/Agents/RAG pipelines, but tooling around this is still lacking. We're introducing Curator, an open-source library designed to streamline the data curation process. While there are many libraries to prompt LLMs, the semantics of generating synthetic data is different from prompting. For example, we need to process a large number of prompts (sometimes in millions or more) while accepting some failures, utilize several stages of prompting, incorporate human feedback, and filter out bad data using verifiers…

    2025 · github.com

  24. 24OS

Ranked by how close each launch is in meaning, then by votes. Refine with a description →