Alternatives
Products that do what Chaos does
Make clean datasets dirty
- 1

- 2NA
Hey HN - our team wants to open source a project called NASTY (NASTY Abstract Syntax Tree thingY) that we built for ourselves. NASTY was built to maintain testable/composable data pipelines. Our team was ripping our hair out trying to maintain dbt/SQL scripts across different data warehouses (Redshift, BigQuery, Postgres, Snowflake) on top of ever shifting data foundations maintained by our customer's internal data teams. NASTY is the result of our learnings from field experience. We wanted to write abstractions so that we could reuse code. We wanted to bundle those abstractions…
2024 · getnasty.dev
- 3

- 4

- 5

- 6

- 7

- 8TD
2020 · textdb.dev
- 9

- 10

- 11

- 12DC
2016 · datasets.co
- 13

- 14

- 15CD
2022 · github.com
- 16ID
Hello, Hacker News! I'm Erik, cofounder of Release (YCW20). At Release we help “virtualize” your environment, so you can quickly reproduce it for various purposes: remote development, testing, staging, or even running production. But, today I would like to introduce our most popular feature that we are now offering for free: Instant Datasets! Instant Datasets allows you to easily create multiple pools of datasets and distribute them to your teams and it cleans up after itself. Imagine you have multiple production databases in RDS and you need that data when doing testing and development. I…
2023
- 17

- 18

- 19DB
I've been doing some data cleaning for my fine tuning projects using LLMs, and decided to just build a package for it as a side project. Check it out here: https://github.com/databonsai/databonsai Some features: - categorization (labelling), transformation and decomposition (text into structured format) - validates llm outputs - batch mode batches up the inputs/outputs so you don't send the prompt (schema, fewshot examples) for every row of data, saving a significant amount of tokens There are some similarities to the Instructor repo, but this is simpler and made for…
2024 · github.com
- 20GF
Just open-sourced a small terminal tool I’ve been working on. The idea came from wondering how useful it’d be if you could just describe the kind of dataset you need, and it would go out, do the deep research, and return something structured and usable. You give it a description, and it pulls relevant info from across the web, suggests a schema based on what it finds, and generates a clean dataset. The schema is editable, and it also adds a short explanation of what the dataset covers. In some cases, it even asks follow-up questions to make the structure more useful. Started off as a quick…
2025 · github.com
- 21

- 22

- 23

- 24
Data Cleaner ▲2Clean messy Excel & CSV files with AI in seconds
Jul 2026 · datacleansheet.netlify.app
Ranked by how close each launch is in meaning, then by votes. Refine with a description →