We're building an open data warehouse inspired by Git scraping
Hey everyone, this is Jason and Nathan from https://subsets.io, a new open data warehouse. Our goal is to make finding and accessing public data easier for human analysis, in apps, or as a source of up-to-date data for retrieval-augmented-generation. Inspired by git scraping [1], the core idea is to build something where people don’t upload a snapshot of their dataset directly, like you might do on Kaggle or Huggingface. Instead, anyone can contribute code (connectors) which we then continuously run and make the fetched data available for everyone in our shared, public data…
What it does
In the maker’s words, at launch
Hey everyone, this is Jason and Nathan from https://subsets.io, a new open data warehouse. Our goal is to make finding and accessing public data easier for human analysis, in apps, or as a source of up-to-date data for retrieval-augmented-generation. Inspired by git scraping [1], the core idea is to build something where people don’t upload a snapshot of their dataset directly, like you might do on Kaggle or Huggingface. Instead, anyone can contribute code (connectors) which we then continuously run and make the fetched data available for everyone in our shared, public data warehouse. We currently have connectors for 120+ datasets including an index of YC companies, U.S. house prices, and Wikipedia search volumes. Separately, open data portals, such as from NGOs, can be hard to use due to their use of semantic web principles - i.e., representing data as a graph and adding structured metadata. We’re taking a less structured approach: each dataset is just a table that you can download or query using SQL, and we’re building a machine learning engine for ranking, pre-processing, and to generate relevant subsets/views from the data warehouse. BigQuery is used as the data warehouse. We use dagster for the data pipelines, running it on top of Kubernetes. Frontend is NextJS. The data pipelines are currently centralised in our repo, but we’re building our own engine where you can just upload simple scripts. Search is currently basic semantic search, with one big index that stores unique strings across tables, columns, and rows. Before we used better search using LLM’s, but the cost, latency, and rate limits mean we’re still investigating the right way to go. The project is in its very beginning stages, but we’d like to get some early feedback and find people who either want to help us build connectors or use the data to build something cool. The connectors are available at https://github.com/subsetsio/subsets-connectors, and you can visually explore the datasets and get your own free API key at https://www.subsets.io. [1] - https://simonwillison.net/2020/Oct/9/git-scraping/
Does the same job
all alternatives →- SGStellar – Git for PostreSQL and MySQL2014 · github.com · ▲325
- GTGitorials – Tutorials as Git Repos2017 · gitorials.com · ▲147
- SSSorcia – Self-hosted web front end for Git repositories, written in Go2020 · git.mysticmode.org · ▲100
- ISI started a repo for sharing algorithm implementations2013 · github.com · ▲52
Everything that would be here is in the README. I hope this gets big, it has tons of potential.
- SGSir – Git-diff-able JSON database on yer filesystem2020 · github.com · ▲115
- RAReGit – A Tiny Git-Compatible Git Implementation2021 · github.com · ▲96
More dev tools this month
the category →



Open-source GTM skills for technical founders
Dev tools · 29d ago · gtmcofounder.com

OpenTrailPaper is open-source bike computer firmware for the LilyGO T5S3 4.7" E-Paper PRO. It supports offline maps, GPX routes, FIT recording and Bluetooth sensors.
Dev tools · 1d ago · opentrailpaper.com

Launched alongside, January 2024
the whole month →



- IM
Life & fun · 2024 · sitinshade.com
