Alternatives
Products that do what NannyML Regression v0.8.0 does
OSS Python library for detecting silent ML model failure
- 1

- 2OP
2022 · github.com
- 3PM
2019 · github.com
- 4

- 5IA
2020 · github.com
- 6AL
2019 · github.com
- 7SG
2023 · github.com
- 8TV
I am excited to announce the release of TabPFN v2, a tabular foundation model that delivers state-of-the-art predictions on small datasets in just 2.8 seconds for classification and 4.8 seconds for regression compared to strong baselines tuned for 4 hours. Published in Nature, this model outperforms traditional methods on datasets with up to 10,000 samples and 500 features. The model is available under an open license: a derivative of the Apache 2 license with a single modification, adding an enhanced attribution requirement inspired by the Llama 3 license:…
2025 · nature.com
- 9

- 10AP
2022 · github.com
- 11

- 12

- 13

- 14PL
Hi! I’ve been working on this automatic scanner for ML models to detect issues like underperforming data slices, overconfidence in predictions, robustness problems, and others. It supports all main Python ML frameworks (sklearn, torch, xgboost, …) and integrates with the quality assurance solution we are building at Giskard AI (https://giskard.ai) to systematically test models before putting them in production. It is still a beta and I would love to hear your feedback if you have the time to try it out. We have quite a few tutorials in the docs with ready-made colab notebooks to…
2023 · docs.giskard.ai
- 15

- 16TO
Mar 2026 · github.com
- 17
OrchestraML▲82From English prompt to deployed ML model with human approval
Jun 2026 · orchestra-ml.vercel.app
- 18

- 19PO
Hey HN! We’re Kevin and Steve. We’re building PromptTools (https://github.com/hegelai/prompttools): open-source, self-hostable tools for experimenting with, testing, and evaluating LLMs, vector databases, and prompts. Evaluating prompts, LLMs, and vector databases is a painful, time-consuming but necessary part of the product engineering process. Our tools allow engineers to do this in a lot less time. By “evaluating” we mean checking the quality of a model's response for a given use case, which is a combination of testing and benchmarking. As examples: - For generated…
2023 · github.com
- 20SF
I've made a small Python library, designed for quick-and-easy prototyping of machine learning models. It's built on top of scikit-learn, to serialize and deserialize data from the forms you're likely to have, to the format used in scikit-learn. https://github.com/madman-bob/Smart-Fruit It's pretty bare-bones at the moment, but I thought I'd see if there was any interest before spending too much time on it. Let me know what you think.
2018
- 21ML
2011 · eferm.com
- 22

- 23TT
I'm excited to introduce tea-tasting, a Python package for the statistical analysis of A/B tests It features Student's t-test, Bootstrap, variance reduction using CUPED, power analysis, and other statistical methods. tea-tasting supports a wide range of data backends, including BigQuery, ClickHouse, PostgreSQL, Snowflake, Spark, and more, all thanks to Ibis. I consider it ready for important tasks and use it for the analysis of switchback experiments in my work.
2024 · e10v.me
- 24UD
Hey HN! I’m the founder of Unify, and we’ve just released our Model Hub, which provides a collection of LLM endpoints with live runtime benchmarks all plotted across time: https://unify.ai/hub A key finding is that static tabular runtime benchmarks for LLMs simply do not work. It’s necessary to take a time-series perspective, and plot the variations through time. We currently have 21 models provided by: Anyscale, Perplexity AI, Replicate, Together AI, OctoAI, Mistral AI and OpenAI, with more on the roadmap. We test across different regions (Asia, US, Europe), with varied…
2024
Ranked by how close each launch is in meaning, then by votes. Refine with a description →