Alternatives
Products that do what Large Scale Article Extract of Newspapers 1730s-1960s does
Hello HN, over the past 7 months I've spent nearly 3,000 hours on building SNEWPAPERS, the first historical newpaper archive with full-text extractions, nearly perfect OCR, a vast categorization taxonomy and of course with semantic and agentic search capabilities. Problem: I wanted to search through newspaper archives, but when I tried every service only lets you search for keywords and dates, and gives you back raw images of the papers, and too many of them with no context. A sea of noise. Solution: I taught machines how to read the newspapers and so far I've extracted the content from >…
- 1

- 2

- 3AZ
A while ago I was looking for information on a obscure and short lived British computer. I found an article[1] in the archives of BYTE magazine[2] - and was captivated immediately by the tech adverts of bygone eras. This led to a long side project to be able to see all 100k pages of BYTE in a single searchable place. [1]: https://byte.tsundoku.io/#198502-381 [2]: https://news.ycombinator.com/item?id=17683184
2025 · byte.tsundoku.io
- 4

- 5OP
2017 · github.com
- 6NI
Understanding scientific articles can be tough, even in your own field. Trying to comprehend articles from others? Good luck. Enter, Now I Get It! I made this app for curious people. Simply upload an article and after a few minutes you'll have an interactive web page showcasing the highlights. Generated pages are stored in the cloud and can be viewed from a gallery. Now I Get It! uses the best LLMs out there, which means the app will improve as AI improves. Free for now - it's capped at 20 articles per day so I don't burn cash. A few things I (and maybe you will) find interesting: * This is…
Feb 2026 · nowigetit.us
- 7WF
2018 · worldbrain.io
- 8

- 9FT
2021 · gutensearch.com
- 10

- 11

- 12IM
2023 · paperlist.io
- 13
- 14SP
Some of the features: * Quickly preview or jump to figures/references/equations/etc. (even if the PDF doesn't have links) * Search paper names in google scholar by middle clicking on their name * Searchable table of contents * Searchable highlights/bookmarks * Browser-like history navigation * Mark locations for quick navigation (Vim style) * Synctex support Video demo of some features: https://www.youtube.com/watch?v=yTmCI0Xp5vI
2022 · github.com
- 15OS
Oldest Search is a custom google search that specifically targets the oldest entries available. I'm always curious about the first entries for certain data on the internet, it's a valuable perspective builder. I personally like news articles that have been digitized that were written in the pre-internet era. Unfortunately some results don't always work well because pages have been dated incorrectly. For example, searching "Covid" shows recent results. I launch new projects like this daily: small tools to increase human agency. I'm also very open to suggestions to improve!
2022 · oldestsearch.com
- 16

- 17AI
Hello HN! In a recent "Ask HN: What are you working on?" thread, I mentioned I was working on OCRing a large book: https://news.ycombinator.com/item?id=41971614 The post generated some interest so I thought I would keep HN posted. The book is Saint-Simon’s Memoirs -- an invaluable historical account of the French court under Louis XIV, full of wit, sharp observations, and of incredible literary value. I'm OCRing the edition of reference made between 1879-1930, that contains a lot of comments and footnotes: 45 volumes, ~27,000 pages. Here's a link to a blog post that describes…
2024 · blog.medusis.com
- 18IM
As a grad student (and an ADHDer), I had trouble doing literature review systematically. To combat this, I made a website that finds similar papers using the meaning of the thing I am looking for. I used MixedBread's [^1] embedding model to generate vectors from the abstracts. I store and search similar vectors using Milvus [^2] and finally use Gradio [^3] to serve the frontend. I update the vector database weekly by pulling the metadata dataset from Kaggle [^4]. To speed up the search process on my free oracle instance, I binarise the embeddings and use Hamming distance as a metric. I would…
2024 · papermatch.mitanshu.tech
- 19HB
Hey guys, I love HN! I wanted to extend the same aesthetic and community towards things beyond tech-related news. I thought it would be cool to get the same quality of community gathered around the latest and greatest research coming out. Let me know what you guys think of what I have so far. It's still early so there are probably bugs and other quality issues. If there's any features missing that you'd want let me know. ALSO, if any of you are familiar with the map of the territory of any particular field, please let me know! Would love to pick your brain and to come up with a 'most…
2024 · papertalk.xyz
- 20

- 21RT
I built a system that monitors ~200,000 news RSS feeds in near real-time and clusters related articles to show how stories spread across the web. It uses Snowflake’s Arctic model for embeddings and HNSW for fast similarity search. Each “story cluster” shows who published first, how fast it propagated, and how the narrative evolved as more outlets picked it up. Would love feedback on the architecture, scaling approach, and any ways to make the clusters more accurate or useful. Live demo: https://yandori.io/news-flow/
Nov 2025 · yandori.io
- 22MS
2019 · memos.org
- 23

- 24AN
2019 · newsphere.org
Ranked by how close each launch is in meaning, then by votes. Refine with a description →