Alternatives
Products that do what DataHub, open source datasets for Artificial Intelligence does
Hi HN! I am an undergrad student trying to build interesting things with AI. Recently, I was looking for a dataset I could use for a new project. I realized that it is really frustrating to go through all the government websites (with terrible UX) just to find some usable dataset. I set out to build a GitHub for datasets, named DataHub. Right now, we have more than 1000 datasets from Montréal and New York City, with more cities coming soon (and possible government agencies). All of this is wrapped into a powerful search. It's a breeze to find a dataset to work on. I'd be interested to know…
- 1

- 2

- 3

- 4DC
2016 · datasets.co
- 5

- 6PF
Hi there Hacker News, I've started a side project http://datasourcehub.com which aims to be a platform for data scientists. The project is still in the idea phase so the UI/UX and functionality are all subject to change. Feel free to play around, below is a guest login, and make sure files are content type of 'text/csv'. All data is subject to deletion, it's just a sandbox right now! By reaching out to the Hacker News community I hope to reach expert data scientists and get their feedback. Below are some questions I'd like to answer and some proposed directions that this…
2013
- 7

- 8
GTA DataCity▲80San Francisco co-working dataset visualized, as a game!
Jul 2026 · datacity-pi.vercel.app
- 9

- 10FC
Hi there, I've created this side project to make it easier to find interesting repositories using AI. There's still a lot of work to be done to improve it, so any suggestions for enhancements would be greatly appreciated. Thank you!
2024 · awesome-repositories.com
- 11GF
2018 · dataturks.com
- 12DR
The first ever AI peer reviewed research article just got approved. It’s kinda crazy how advanced AI have come to replace researchers. I've just been using Deep Research on ChatGPT and Perplexity a lot to write and research complex technical reports for my boss. He loves the reports and it has decreased my workload a ton but I still have some frustrations with it. None of them provide an API that gets me the same quality of output you would with the applications. I wanted something with more control on the LLMs, swappable with the reasoning new models that came out. Not just prompt →…
2025 · github.com
- 13WA
Today you can easily adopt AI coding tools because you have git for branching and rolling back if AI writes bad code. We haven't seen this same capability for data and decided to build it ourselves. Nile is a new kind of data lake, purpose built for using with AI. It can act as your data engineer or data analyst creating new tables and rolling back bad changes in seconds. We support real versions for data, schema, and ETL. We'd love your feedback on any part of what we are building - https://getnile.ai/ What do you think?
Jan 2026
- 14IM
Hi there! I've been working with data in one form or another, professionally, for about 5 years. I've been thinking about my own personal data and how it's used for at least twice that long. I've been sort of building something in my head for a while that solves my own problem and, in the beginning of this year, I found the opportunity to spend some time building it out. I'll leave the detailed explanation to the blog post but, in short, I built what amounts to an API crawler combined with a data processor to help you download your personal data from 3rd party services and work with it using…
2024 · joshcanhelp.com
- 15F2
Hey HN! Today we’re launching Fabi 2.0 For the past year we’ve been working with a number of product, eng, marketing and data teams, and the common issue we’ve been seeing with all these teams, especially small and growing teams, is that getting basic answers from their data requires jumping between a bunch of different tools/copy-pasting data and spreadsheets, or building out a full BI stack. With our latest release, users can not only connect Fabi to databases and data warehouses, but also to applications like HubSpot, Stripe, Shopify etc. We offer hundreds of connectors, and the AI…
Jan 2026
- 16IE
Hey HN, when building ML systems for industrial AI, we have learned that data inspection is critical during the ML development process. We are also big fans of the Hugging Face ecosystem. That is why we built an integration to our data exploration tool Spotlight that allows you to interactively explore Hugging Face datasets with one line of code. Spotlight lets you leverage model results such as predictions and embeddings to gain a deeper understanding in data segments and model failure modes. Currently, many many NLP, CV, Audio and multimodal datasets are supported both locally and on the…
2023 · huggingface.co
- 17MA
Hey HN! I built a thing and I'm really excited to share it. EDIT: I meant to link to the github, not the website: https://github.com/max-hq/max Like many of us here, I've been commonly reaching for a pattern of "pull data into db; give it to claude" for a while, whilst doing data spelunking or building tooling - for the same reasons mentioned by thellimist over here [1] and a few other recent "CLI vs MCP" posts. To that end, about a month ago I started building a project called `max` - its goal is to cut the middleman and schematise any data source for you. Essentially,…
Mar 2026 · max.cloud
- 18MA
Hi HN, I'm a solo developer learning to code, and I'd love to share my second real project: MapMyLearn, an AI-powered app that automatically generates personalized learning paths based on any topic you input. What it does: Takes a topic (e.g. "history of capitalism", "learn Rust", or "data storytelling") Uses AI to break it down into a structured course with modules and submodules Each submodule includes: - Detailed, pedagogical content (developed based on online sources to mitigate hallucinations) - A quiz of 10 questions - Recommended resources - An AI chatbot for Q&A - Optional audio…
2025
- 19SS
Hi HN! I’ve been building startfa.st, a curated directory of AI, developer, and product tools. Link: https://startfa.st Like many people building in AI/dev, I found myself drowning in new tools every day, everything from agents, deployment frameworks, auth platforms, workflow engines, model APIs, automation tools, design tools, security stacks, etc. There’s constant novelty, but it’s very hard to find signal. Most lists on the internet recycle the same names. So I built startfa.st to solve my own discovery problem. What startfa.st is A fast, searchable, hand-curated index of…
Nov 2025 · startfa.st
- 20IB
Think of HackerNews, but only for AI content. Did build this out of the wish for easier discoverability of new technologies in the AI/ML space. Obviously, so far only a tiny small community, but maybe a few of you find it nice to be around and want to post a link there every now and then :)
2023 · news.aiapipro.com
- 21TA
Hi all, sharing a directory of GPTs I made in a few hours after OpenAI dev day. Hope you find it useful!
2023 · topgpts.ai
- 22YA
Hey folks! I'm a founding engineer at Yorph AI, an agentic data platform, built using ADK, that helps users (starting with product managers and analysts) join data from different sources (upload or sync), build version-controlled and reliable data workflows, and clean, analyze, and visualize data — all in one place. We're also releasing semantic layer creation later this week. The beta is live at yorph.ai/login — would love to hear your thoughts and feedback! (FYI: We're still waiting on Google app verification — you'll see a warning for a few days. Dropbox shows a similar one since…
Nov 2025 · yorph.ai
- 23FA
Hey HN, we built an Econ+Finance database to let AI agents do investment research. We spend a lot of tokens to organize macro releases and SEC filings into a clean format, so that your agents have more context to do actual analysis. The problem AI agents are great at data analysis. But they become ineffective if most of their context window is spent on gathering and cleaning data, instead of validating hypotheses. Data in the wild is messy and rarely standardized. Definitions and measurements change over time. This problem is compounded by a fragmented data universe. Point solutions exist…
Jul 2026 · github.com
- 24MA
Hi, I'm working on a project that regroups all best AI (AIaaS) from different providers (GCP, AWS, Azure, DeepL, etc.) in one API (https://github.com/edenai/edenai-apis). I've got asked the question : why aren't you regrouping Open Source models (instead of proprietary APIs) into one repo? Well because it doesn't make sens to deploy and maintain large pytorch (or other framework) AI models (especially for document parsing, image and video moderation or speech recognition) in every solution that wants AI capabilities. So using APIs makes way more sens. Deployed OpenSource…
2023 · github.com
Ranked by how close each launch is in meaning, then by votes. Refine with a description →