nowfound

Alternatives

Products that do what Axiom Clean AI Data Pipeline does

LLM Fine-Tuning Synthetic Data

  1. 1

    LLM reinforcement fine-tuning platform to improve LLM output

    2025

  2. 2

    Open source unstructured data ETL for AI first applications

    2024

  3. 3

    Fine-tuning, RL, and inference in one CLI

    Dec 2025

  4. 4

    Get your unstructured data AI-ready in minutes

    2024

  5. 5

    Open source data labelling platform for AI model tuning

    2023

  6. 6

    Clean messy data in minutes and without coding

    2020

  7. 7

    The AI agent for synthetic data generation

    Nov 2025

  8. 8

    AI fine-tuning platform to create custom LLMs

    2024

  9. 9
    Taylor AI118

    Fine-tune open source LLMs in minutes

    2023

  10. 10

    "Text-to-columns" on steroids - clean your unstructured data

    2024

  11. 11

    Instant compliant test data for engineering teams

    Oct 2025

  12. 12
    Heym83

    Self-hosted AI workflow automation with agents, RAG, and MCP

    Apr 2026

  13. 13

    AI JSON & CSV data sanitization with zero hallucinations

    19d ago · cleandata-ai.com

  14. 14TZ

    2017 · gpestana.gitbooks.io

  15. 15

    AI-powered enterprise discovery in 17 minutes, not weeks

    Nov 2025

  16. 16SA
  17. 17OS

    We’re building an open-source tool that makes it easy to expose secure, LLM-optimized APIs on top of your structured data—without manually designing endpoints or worrying about compliance. AI agents and LLM-powered applications need structured access to data, but traditional APIs and databases weren’t built with AI workloads in mind. Our tool automatically generates APIs that: - Filter out PII & sensitive data to comply with GDPR, CPRA, SOC 2, and other regulations. - Provide traceability & auditing, so AI apps aren’t black boxes, and security teams stay in control. - Optimize for AI…

    2025 · github.com

  18. 18IM

    Every time I wanted to use LLMs in my existing pipelines the integration was very bloated, complex, and too slow. This is why I created a lightweight library that works just like scikit-learn, the flow generally follows a pipeline-like structure where you “fit” (learn) a skill from sample data or an instruction set, then “predict” (apply the skill) to new data, returning structured results. High-Level Concept Flow Your Data --> Load Skill / Learn Skill --> Create Tasks --> Run Tasks --> Structured Results --> Downstream Steps And the bast part: Every step can be saved and reused as…

    2025 · github.com

  19. 19AO

    Hi, We are building an open-source framework for loading and structuring LLM context to create accurate and explainable LLM answers using knowledge graphs and vector stores. We built the tool with four main concepts in mind: 1. Loader -> uses dlt in the backend to load and structure the data 2. Cognify step -> creates a graph with summaries, labels and factoids that are interconnected across the documents and stored as a representation in the vector store 3. Optimizer -> Uses DSPy to optimize LLM queries, and we plan to extend it to most of the knobs we can turn, like chunking etc. 4. Search…

    2024 · github.com

  20. 20OS

    And you can try out the models live here: https://labs.refuel.ai/playground

    2024 · huggingface.co

  21. 21CM

    Hey HN, I've been building AutoAgents, an AI agent framework in Rust. Today I'm sharing a feature I haven't seen done well elsewhere: composable middleware layers for LLM inference pipelines. The problem Every agent framework lets you swap LLM providers. Almost none of them give you a structured way to enforce safety, caching, or data sanitization in the inference path itself. You end up with guardrails as application-level if-statements, caching bolted on as a separate service, and PII handling as a "we'll add it later" TODO that never ships. This gets worse with local models. Cloud APIs…

    Mar 2026 · github.com

  22. 22LQ

    Hi HN, I built LucidShark: a local-first, open-source CLI tool that acts as a quality & security pipeline. It can be used to increase the confidence in AI-generated (or AI-assisted) code. - Config lives as code in version-controlled lucidshark.yml - 100% local; no cloud, no SaaS - Runs 10 quality domains automatically: linting, formatting, type checking, SAST/security scanning, SCA/dependency checks, IaC validation, container scanning, unit tests, coverage thresholds, code duplication, etc. - Produces a QUALITY.md dashboard with health scores (e.g. 9.1/10), trends, and issue…

    Mar 2026 · lucidshark.com

  23. 23

    Deterministic numeric tools for AI agents, zero credits

    Aug 2026 · datagrout.ai

  24. 24BL

    Hello everyone! I am Jan, CTO and one of the creators of Pathway, the real-time data processing framework. I’m excited to share Pathway’s ready-to-use AI Pipelines, configurable with just YAML! These frameworks offer out-of-the-box solutions for AI search, RAG, and more—optimized for real-time indexing and in-memory processing. What makes it simple? YAML templates! The pipeline templates are fully customizable using YAMLs to fit your needs, from changing the data sources to the choice of the LLM model, all without touching Pathway’s Python code. Thanks to the Pathway data processing engine,…

    2024 · pathway.com

Ranked by how close each launch is in meaning, then by votes. Refine with a description →