Alternatives
Products that do what Sliceguard, a Tool for Finding Critical Data Segments in ML Data does
Hey HN, We know from experience that a qualitative understanding of critical data segments is necessary when developing ML models. Better tooling makes finding these data slices faster, more systematic and helps to communicate with domain experts. We built sliceguard to go from a raw dataset to an interactive report on critical data slices with just 3 lines of code: https://github.com/Renumics/sliceguard Behind the scenes, we use hierarchical clustering and explainable AI techniques to detect and rank data slices based on features, metadata and embeddings. Here is some…
- 1

- 2

- 3AL
Hi HN! I am Maria, solo founder of DataQA (https://dataqa.ai/), a tool to search and label documents for various NLP tasks (e.g. entity extraction, entity linking, etc). I have worked as a data scientist and ML engineer for the better part of a decade, and over that time have specialised mainly in applications involving natural language processing (NLP). One of the key questions I have always had at the back of my mind is whether my time was well spent. Whenever I spent more time on feature engineering or trying different models, I always wondered whether I would get better…
2021
- 4DC
2016 · datasets.co
- 5

- 6MA
2014 · meta-toolkit.github.io
- 7

- 8

- 9

- 103I
Hello HackerNews, I am Paul, and I would like to get some feedback on the tool we are releasing as beta today. 3LC is an ML tool that gives detailed insights, real-time data-centric iterative workflows for training/finetuning, and data quality improvements for your Machine Learning datasets and models. 3LC serves as a visualizer, editor, and debugger, focusing on how models learn from the training data. Key Features of 3LC: • Detailed Data Analysis: 3LC enables users to dive into model performance beyond typical labeling errors. It offers the capability to analyze intricate false…
2024 · pypi.org
- 11IE
Hey HN, when building ML systems for industrial AI, we have learned that data inspection is critical during the ML development process. We are also big fans of the Hugging Face ecosystem. That is why we built an integration to our data exploration tool Spotlight that allows you to interactively explore Hugging Face datasets with one line of code. Spotlight lets you leverage model results such as predictions and embeddings to gain a deeper understanding in data segments and model failure modes. Currently, many many NLP, CV, Audio and multimodal datasets are supported both locally and on the…
2023 · huggingface.co
- 12IS
Everything that would be here is in the README. I hope this gets big, it has tons of potential.
2013 · github.com
- 13MD
2020 · github.com
- 14G1
2022 · github.com
- 15SS
2017 · blog.slicingdice.com
- 16TF
Hi All! We've spent a few months on getting an MVP together, and would love to get some feedback on whether this tool meets you needs. Here is a link to a demo video: https://www.youtube.com/watch?v=FBLi3vdKB-4&feature=emb_rel_pause Here's a link to our website: https://www.structure.rest And here's a blog article, I published today in the space: https://www.structure.rest/blog/using-a-data-analytics-stack-to-gain-business-insights
2020
- 17RS
Hey HN! We just released the open-source version of Renumics Spotlight, a data exploration and analysis tool for multimodal datasets. Spotlight integrates seamlessly with pandas and supports rich data types like images, videos, and meshes. You can load anything that fits in a DataFrame and view it through a customizable GUI featuring multiple interactive widgets: a data table, similarity map, histograms, scatter plots, and more. In the past, we have used Spotlight for exploratory data analysis and tackling various model and data-related problems in our machine learning projects. What are…
2023 · renumics.com
- 18ML
2016 · github.com
- 19DE
Hey! We’ve built a data extraction tool to flexibly automate data and document processing. You’ve probably seen a few of these, so have we! A few of us have been varyingly stuck trying to automate the extraction of borrower financials for the past 5 years. We think that there are a few missing features of most data extraction tools. * They are usually too complex to quickly get up and running * They are overly constrained in terms of what workflows and documents they support We’ve always felt like speed and flexibility were sticking points, so we went slightly orthogonal to the alternatives.…
2024 · go.sea.dev
- 20SP
I built Sculptor after repeatedly seeing founders try to hire data scientists for a task that ultimately boiled down to extracting structured data from unstructured text (customer records, social posts, websites, etc) using an LLM API. We ended up reinventing this pattern internally at least three times in the past year, so I published Sculptor as a streamlined, open-source solution: - Simple schema-based extraction, with parallelization and type validation. - Multi-step pipelines with filtering or transforms between steps. - Configure everything in YAML/JSON for easy reuse. It’s MIT…
2025 · github.com
- 21LA
2017 · github.com
- 22SJ
2016 · siphonjs.com
- 23IE
Hey HN, We built a library to interactively explore unstructured datasets directly from a dataframe: https://github.com/Renumics/spotlight Some background: We have worked on different ML solutions over the years, mainly in the industrial AI space. A crucial step for us is always to inspect and explore the data interactively with the team and the customer. This is true throughout the dev process: During EDA, model debugging, model comparison and monitoring. We have tried many different options for visualizing unstructured datasets in the past: Notebooks, dash apps, custom…
2023 · github.com
- 24TS
Hello Hacker News community! I'm currently working in financial risk management within the banking sector, and I began my career as a Data Science specialist. For quite some time, my friend and I have been developing a small pet project just for fun. This tool has repeatedly helped us save time when testing various hypotheses and machine learning models. The core idea is to combine different scripts—created in various programming languages and virtual environments—within a minimalist graphical interface. Whether you're building models, running a local neural network, or sending requests to…
2024
Ranked by how close each launch is in meaning, then by votes. Refine with a description →