nowfound

Alternatives

Products that do what LLM Training Data Crawler & Curator does

Curate clean, deduplicated training data for AI models.

  1. 1L3

    I spent a lot of time and money on this rather big side project of mine that attempts to replicate the mechanistic interpretability research on proprietary LLMs that was quite popular this year and produced great research papers by Anthropic [1], OpenAI [2] and Deepmind [3]. I am quite proud of this project and since I consider myself the target audience for HackerNews did I think that maybe some of you would appreciate this open research replication as well. Happy to answer any questions or face any feedback. Cheers [1]…

    2024 · github.com

  2. 2
    FineTuner164

    Fine-tune AI models on your data — in minutes, not days.

    2025

  3. 3

    Vibe-check many open-source and proprietary LLMs at once

    2024

  4. 4TA

    I built this tool because I wanted a way to just take a bunch of URLs or domains, and query their content in RAG applications. It takes away the pain of crawling, extracting content, chunking, vectorizing, and updating periodically. I'm curious to see if it can be useful to others. I meant to launch this six months ago but life got in the way...

    2024 · embedding.io

  5. 5LF
  6. 6
    LLM Stats308

    Compare API models by benchmarks, cost & capabilities

    Oct 2025

  7. 7

    Generative AI to research, validate & scale your business

    2024

  8. 8
    Llama312

    3.1-405B: an open source model to rival GPT-4o / Claude-3.5

    2024

  9. 9
    Crawlify165

    AI powered data extraction APIs. Hassle-free data retrieval.

    2020

  10. 10

    Build LLMs powered by GPT & your own data

    2023

  11. 11LS
  12. 12

    Audit your site for the AI search era. 100% Open Source

    May 2026 · freeaiseoaudit.com

  13. 13

    RAG-ready web scraping that cuts your LLM token costs

    Apr 2026 · geekflare.com

  14. 14

    Check if your website can get crawled by ChatGPT

    2025

  15. 15
    Gradient153

    Developer API for building private LLMs that you own

    2023

  16. 16AL

    Lately I've felt exhausted due to the deluge of AI/GPT posts on hacker news, and have seen similar grumblings. I threw together this frontend that filters out anything with the phrases AI, LLM, GPT, or LLaMa for use until the hype dies down a bit. Before anyone asks, yes I did try to use ChatGPT to help, and while the code it provided was helpful, it needed some heavy bug-fixing. Edit: One other note I forgot to mention. The favicon is generated by Stable Diffusion, I asked it to generate an "Aritificial Intelligence Favicon", and then I added the red circle with line through it.

    2023 · save-buffer.github.io

  17. 17FA

    Hey HN! We’re building FinetuneDB (https://finetunedb.com/), an LLM fine-tuning platform. It enables teams to easily create and manage high-quality datasets, and streamlines the entire workflow from fine-tuning to serving and evaluating models with domain experts. You can check out our docs here: (https://docs.finetunedb.com/) FinetuneDB exists because creating and managing high-quality datasets is a real bottleneck when fine-tuning LLMs. The quality of your data directly impacts the performance of your fine-tuned models, and existing tools didn’t offer an easy…

    2024 · finetunedb.com

  18. 18
    Taylor AI118

    Fine-tune open source LLMs in minutes

    2023

  19. 19AT

    I recently built a small open-source tool to benchmark different LLM API endpoints — including OpenAI, Claude, and self-hosted models (like llama.cpp). It runs a configurable number of test requests and reports two key metrics: • First-token latency (ms): How long it takes for the first token to appear • Output speed (tokens/sec): Overall output fluency Demo: https://llmapitest.com/ Code: https://github.com/qjr87/llm-api-test The goal is to provide a simple, visual, and reproducible way to evaluate performance across different LLM providers, including…

    2025 · llmapitest.com

  20. 20RL

    We've been building data pipelines that scrape websites and extract structured data for a while now. If you've done this, you know the drill: you write CSS selectors, the site changes its layout, everything breaks at 2am, and you spend your morning rewriting parsers. LLMs seemed like the obvious fix — just throw the HTML at GPT and ask for JSON. Except in practice, it's more painful than that: - Raw HTML is full of nav bars, footers, and tracking junk that eats your token budget. A typical product page is 80% noise. - LLMs return malformed JSON more often than you'd expect, especially with…

    Mar 2026 · github.com

  21. 21AL

    Try it out here: https://labs.refuel.ai/playground Refuel LLM (84.2%) outperforms trained human annotators (80.4%), GPT-3-5-turbo (81.3%), PaLM-2 (82.3%) and Claude (79.3%) across a benchmark of 15 text labeling datasets. It is a Llama-v2-13b base model, trained on over 2500 unique datasets (5.24B tokens) spanning categories such as classification, entity resolution, matching, reading comprehension and information extraction.

    2023

  22. 22

    Fine-tuning, RL, and inference in one CLI

    Dec 2025

  23. 23

    AI fine-tuning platform to create custom LLMs

    2024

  24. 24IG

    2024 · columns.ai

Ranked by how close each launch is in meaning, then by votes. Refine with a description →