nowfound

Alternatives

Products that do what A Layer for robotics dataset quality does

We are exploring how teams can automatically audit data quality, detect problematic demonstrations, and build smaller, high-quality training sets before spending GPU time. We are looking for design partners to help shape the Calibra.

  1. 1

    Build Better Predictive Models — Faster

    2014

  2. 2AR

    Oct 2025 · github.com

  3. 3
    Datature271

    No-code platform for building deep neural nets

    2021

  4. 4

    Find bad data. Train on less. Save GPU.

    24d ago · calibrarobotics.com

  5. 5ML

    2011 · eferm.com

  6. 6
    Layer AI124

    Metadata store for production ML

    2022

  7. 7ML
  8. 8

    Build the semantic layer that makes AI analytics trustworthy

    Mar 2026

  9. 9DC
  10. 10

    Train robot policies with up to 75% less data

    30d ago · calibrarobotics.com

  11. 11MD

    We’re excited to share ML-Dev-Bench, a new open-source benchmark that tests AI agents on real-world ML development tasks. Unlike typical coding challenges or Kaggle-style competitions, our benchmark simulates end-to-end ML workflows including: - Dataset handling and preprocessing - Debugging model and code failures - Implementing new model architectures - Fine-tuning and improving existing models With 30 diverse tasks, ML-Dev-Bench evaluates agents across critical stages of ML development. To complement this, we built Calipers, a framework that provides systematic performance evaluation and…

    2025 · github.com

  12. 12AQ
  13. 13AA
  14. 14PA

    In the past 4 years, we developed a platform generating 3D visuals and deploying AR Try-on experiences into large footwear and retail companies. We focused on high-fidelity meshes and textures since the point was giving shoppers a realistic visual. Then robotics companies started to approach us and they asked for thousands of 3D assets that they can use in creating different simulation environments. This is crucial for them so they can achieve better sim-to-real transfer with domain randomization (train a humanoid in a million different kitchens so a real one is not a surprise when it is…

    Apr 2026 · app.rigyd.com

  15. 15ST

    Hey HN, Data quality matters more than ever. Our world increasingly relies on AI and proprietary datasets to fine-tune models. But as many of us know, garbage in often results in garbage out. Bad data can cost millions, especially when it informs important decisions like public health policy or interest rate hedging. Today’s data engineers need a flexible, secure, and performant data quality toolkit designed for the modern workflow. But current solutions were built a decade ago and don't support big data technologies, such as Spark. This is why we're building Spotlight - the data quality…

    2023 · spotlight.dev

  16. 16DO

    Hi HN! I am an undergrad student trying to build interesting things with AI. Recently, I was looking for a dataset I could use for a new project. I realized that it is really frustrating to go through all the government websites (with terrible UX) just to find some usable dataset. I set out to build a GitHub for datasets, named DataHub. Right now, we have more than 1000 datasets from Montréal and New York City, with more cities coming soon (and possible government agencies). All of this is wrapped into a powerful search. It's a breeze to find a dataset to work on. I'd be interested to know…

    2017

  17. 17

    We built an open sourced coordination layer for AI agents working on the same repository. Detects work duplication and design conflicts early

    9d ago · twing.dev

  18. 18IE

    Hey HN, when building ML systems for industrial AI, we have learned that data inspection is critical during the ML development process. We are also big fans of the Hugging Face ecosystem. That is why we built an integration to our data exploration tool Spotlight that allows you to interactively explore Hugging Face datasets with one line of code. Spotlight lets you leverage model results such as predictions and embeddings to gain a deeper understanding in data segments and model failure modes. Currently, many many NLP, CV, Audio and multimodal datasets are supported both locally and on the…

    2023 · huggingface.co

  19. 19IC

    I spent the past week implementing a 1 Layer Neural Net and training it on MNIST within the visual scripting language provided by scratch.mit.edu. It was tedious, but ultimately not too difficult. The code runs incredibly slowly, so much so that 64 samples of MNIST takes 5+ hours to train on my machine. There were a lot of little mini challenges that were fun to overcome (implementing softmax was very tricky). If you're interested, I encourage you to try and improve on it! More details in the linked blog post.

    2024 · bell-boy.github.io

  20. 20TN

    Hi guys, I’m excited to share an update on ReproModel, an open-source toolbox designed to streamline the testing and reproduction of machine learning models. I, like many of you, have really struggled with benchmarking and comparing models, from missing code, to opaque experiment parameters slowing the process. I decided to take matters into my own hands, and created a mini-toolbox in my free time to streamline the process. The goal is to reduce the time and effort spent on replicating experiments, enabling researchers to focus on innovation rather than setup. Knowing this task is not an…

    2024 · github.com

  21. 21EB
  22. 22PA

    Hey HN! Pipevals is early and rough (this is a learning project), but usable. It currently lets you: - build evaluation pipelines as graphs - run them against datasets - track how output quality changes over time

    Mar 2026 · github.com

  23. 23TS
  24. 24

    Is AI running out of high-quality training data?

    26d ago · khayyamshah2007.blogspot.com

Ranked by how close each launch is in meaning, then by votes. Refine with a description →