Alternatives
Products that do what I indexed the academic papers buried in the DOJ Epstein Files does
The DOJ released ~3.5M pages of Epstein documents across 12 datasets. Buried in them are 207 academic papers and 14 books that nobody was really talking about. From what I understand these papers aren't usually freely accesible, but since they are public documents, now they are. I don't know, thought it was interesting to see what this dude was reading. You can check it out at jeescholar.com Pipeline: 1. Downloaded all 12 DOJ datasets + House Oversight Committee release 2. Heuristic pre-filter (abstract detection, DOI regex, citation block patterns, affiliation strings) to cut noise 3. LLM…
- 1EF
Hey all, Throwaway in case this is assumed to be politcally motivated. I spent some time organizing the Eptstein files to make transparency a little clearer. I need to tighten the data for organizations and people a bit more, but hopeful this is helpful in research in the interim.
Nov 2025 · searchepsteinfiles.com
- 2OA
Hi HN, I built an open-source AI agent that has already indexed and can search the entire Epstein files, roughly 100M words of publicly released documents. The goal was simple: make a large, messy corpus of PDFs and text files immediately searchable in a precise way, without relying on keyword search or bloated prompts. What it does: - The full dataset is already indexed - You can ask natural language questions - Answers are grounded and include direct references to source documents - Supports both exact text lookup and semantic search Discussion around these files is often fragmented. This…
Jan 2026 · epstein.trynia.ai
- 3JG
Hi everyone! My name's Luke and I made the original Jmail here alongside Riley Walz. We had a ton of friends collaborate on building out more of the app suite last night in lieue of DOJ's "Epstein files" release. Please AMA!
Dec 2025 · jmail.world
- 4TP
2019 · hackernewspapers.com
- 52G
Community, All the HN belong to you. This is an archive of hacker news that fits in your browser. When I made HN Made of Primes I realized I could probably do this offline sqlite/wasm thing with the whole GBs of archive. The whole dataset. So I tried it, and this is it. Have Hacker News on your device. Go to this repo (https://github.com/DOSAYGO-STUDIO/HackerBook): you can download it. Big Query -> ETL -> npx serve docs - that's it. 20 years of HN arguments and beauty, can be yours forever. So they'll never die. Ever. It's the unkillable static archive of HN and it's…
Dec 2025 · hackerbook.dosaygo.com
- 6

- 7

Fast and accurate Chat, Search with all Epstein DOJ Files
Mar 2026
- 8NI
Understanding scientific articles can be tough, even in your own field. Trying to comprehend articles from others? Good luck. Enter, Now I Get It! I made this app for curious people. Simply upload an article and after a few minutes you'll have an interactive web page showcasing the highlights. Generated pages are stored in the cloud and can be viewed from a gallery. Now I Get It! uses the best LLMs out there, which means the app will improve as AI improves. Free for now - it's capped at 20 articles per day so I don't burn cash. A few things I (and maybe you will) find interesting: * This is…
Feb 2026 · nowigetit.us
- 9ES
This project reconstructs the Epstein email records from the recent U.S. House Oversight Committee releases using only public-domain documents (23,124 image files + 2,800 OCR text files). Most email pages contain only one real message, buried under layers of repeated headers/footers. I wanted to rebuild the conversations without all the surrounding noise. I used an OCR + vision-LLM pipeline to extract individual messages from the email screenshots, normalize senders/recipients, rebuild timestamps, detect duplicates, and map threads. The output is a structured SQLite database that…
Dec 2025 · github.com
- 10HN
Mar 2026 · huggingface.co
- 11HN
2021 · steez.news
- 12

- 13EF
credit to @RhysSullivan on github for creating this
Dec 2025 · epstein-files-browser.vercel.app
- 14AB
Recently i stumbled on too many clickgates on the Medium blog Towards Data science. Considering that most people publish to share knowledge on Medium and are driven into putting their content behind a paywall, without actually getting paid for it, including myself. I felt like Medium is running the academic publishing scheme. Get free content and get paid for it. So I decided to create a small script to bypass the paywall on Medium, it turns out it also works on other newssites. Heres the website: https://sugoidesune.github.io/readium/ For the curious I will explain the…
2019
- 15MW
This was more of a sandbox to play with Raphael and Flot than anything else but I think there are some interesting statistics in there. It'd be awesome to do this over time but I really don't have the spare time required.
2011 · burntbrunch.github.com
- 16IJ
Hi HackerNews, Lately, I have seen an explosion in posts offering paid APIs/services to get unstructured data into LLMs (i.e. langchain extract, ragflow, unstructured, unstract, just to name a few) and I have been largely disappointed by them, either because they fail to implement multimodal support, fail to give good context for "really tricky" PDFs / Word docs / Powerpoints, or are just plain difficult to use. In light of all these posts I figured I'd share my solution that has been working smoothly for me and my clients. I put it up on GitHub for free so you can check it…
2024 · github.com
- 17IP
To be specific, the content is generated by a GPT-2 based model. https://amzn.to/2TCc0v2 Let me know if you have any questions :-)
2020
- 18

- 19E2Excel 2013▲20
Not exactly a side project but my work for the last 2.5 years. As an avid HN lurker would love to hear feedback. Background: I am a PM on the Excel team in charge of the advanced data analysis tools (http://www.microsoft.com/en-us/bi/Products/OfficePreview.aspx) To download: http://www.microsoft.com/office/preview/en
2012
- 20IS
Hey folks, I wanted to share my latest project, which is a book that I wrote and self-published. The book is called Behind the Ivy Curtain: A Data Driven Guide to Elite College Admissions, and the link is http://www.amazon.com/gp/product/B013YFIQ30. Let me give you some backstory. I went to a fun high school but it was really easy and fun and I was a pretty big slacker. Around sophomore spring I decided I wanted to leave Florida, and the best way to do that was to get a scholarship to a good school. I had no idea how to do that so I spent a bunch of time reading…
2015
- 21

- 22IB
I posted this a few weeks ago and the server died under the traffic. Fixed that by adding an in-mem caching layer with Redis/valkey and added CloudFront caching for static content. Also upgraded the server. Also fixed the Firefox bugs, trying again. It's a research tool for US stocks. Financials for ~10k companies pulled from SEC filings. You can chart any metric across companies, filter news by ticker, ask questions in plain English and get a chart back. There's also SQL console against the whole database, which is the part I like to use together with the AI chat (generates an SQL…
Jun 2026 · terminal.tesseractanalytics.ai
- 23

- 24AG
Here it is: https://docs.google.com/document/d/1T8R91jWhQK9t7IwQGngDDLFavJ2LTNtlntThgQPS0Uo/edit?usp=sharing Feel free to add articles to the suggestions section, and I will review them and add to the document! Any and all feedback is accepted!
2015
Ranked by how close each launch is in meaning, then by votes. Refine with a description →