
Webᵀ Crawl by Web Transpose
Realtime web data for LLMs
What it does
✨ Turn full websites into datasets for building custom LLMs with Webᵀ Crawl. Give us just 1️⃣ URL and let Webᵀ Crawl handle the rest. 💻 Quickly turn full websites content (like PDFs, FAQ, etc.) into [prompts for fine-tuning] and [chunks for vector databases].
Does a similar job
all alternatives →- TATurn any website into a knowledge base for LLMs2024 · embedding.io · ▲305
I built this tool because I wanted a way to just take a bunch of URLs or domains, and query their content in RAG applications. It takes away the pain of crawling, extracting content, chunking, vectorizing, and updating periodically. I'm curious to see if it can be useful to others. I meant to launch this six months ago but life got in the way...
- LSLLM Scraper – turn any webpage into structured data2024 · github.com · ▲88

- TATransform any website or eBook into a research paper (no LLM required)2023 · github.com · ▲96
- RLRobust LLM extractor for websites in TypeScriptMar 2026 · github.com · ▲72
We've been building data pipelines that scrape websites and extract structured data for a while now. If you've done this, you know the drill: you write CSS selectors, the site changes its layout, everything breaks at 2am, and you spend your morning rewriting parsers. LLMs seemed like the obvious fix — just throw the HTML at GPT and ask for JSON. Except in practice, it's more painful than that: - Raw HTML is full of nav bars, footers, and tracking junk that eats your token budget. A typical product page is 80% noise. - LLMs return malformed JSON more often than you'd expect, especially with…

More ai this month
the category →
I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.
AI · 19d ago · simedw.com
Astute▲585Automate your B2B brand going viral, with new media creators
AI · 20d ago · company-app.joinastute.com


Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2. The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges…
AI · 29d ago · cactuscompute.com


Source: Product Hunt launch ↗