Crawlspace – A centralized web crawling platform built on Cloudflare
Crawlspace is a centralized web crawling platform that benefits crawler developers AND website owners. Developers can affordably crawl tens of millions of pages per month, scrape with LLMs, and save data in attached storage. Website owners are shielded by a platform-wide TTL cache that absorbs redundant bot traffic. AI bots are running rampant on the open web. Many recent HN stories[1][2][3][4] describe how web crawlers have run amok and hammer websites with DDoS-like traffic. They often do this with blatant disregard of website owners' wishes (e.g. ignoring robots.txt, 429s, Retry-After…
What it does
In the maker’s words, at launch
Crawlspace is a centralized web crawling platform that benefits crawler developers AND website owners. Developers can affordably crawl tens of millions of pages per month, scrape with LLMs, and save data in attached storage. Website owners are shielded by a platform-wide TTL cache that absorbs redundant bot traffic. AI bots are running rampant on the open web. Many recent HN stories[1][2][3][4] describe how web crawlers have run amok and hammer websites with DDoS-like traffic. They often do this with blatant disregard of website owners' wishes (e.g. ignoring robots.txt, 429s, Retry-After headers, etc) because they face no repercussions for deploying poorly-behaved crawlers (and are not incentivized to improve them). The knee-jerk reaction to fix this problem is to give more tools to website owners. Maintaining denylists of IP addresses and user agents, implementing honeypots and tarpits, etc are tactics that website owners use to combat the problem. However, this ends up resulting in and endless arms race between web crawlers and website owners, as they each try to employ new mechanisms of one-upping each other. Crawlspace takes a different approach _by providing a convenient and affordable platform to web crawler developers_. By funneling web crawling traffic through a centralized platform, we can control neat things like making crawlers well-behaved by default, implementing proper caching, and more — all the tedium that that developers don't want to (and therefore, don't) do themselves. Music streaming services like Spotify used convenience and affordability to curb music piracy; we're following the same playbook to curb rampant bot traffic on the internet. In about 50 lines of code, you can deploy a performant and polite web crawler on Cloudflare's network. Every crawler gets its own queue, SQLite database, vector database, and S3-compatible bucket, which allows you to query your crawl as it's crawling with either SQL statements or a RAG chat interface. We've stitched together 10+ Cloudflare products including Queues, Durable Objects, Browser Rendering, Workers AI, D1, R2, and Vectorize. Please let us know what you think! Happy to answer any questions. [1] https://news.ycombinator.com/item?id=42549624 [2] https://news.ycombinator.com/item?id=42660377 [3] https://news.ycombinator.com/item?id=42725147 [4] https://news.ycombinator.com/item?id=42750420
Does the same job
all alternatives →
- COCrawlab: Open-Source Web Crawler Admin Platform That Runs Any Language2019 · github.com · ▲116




More ai this month
the category →
I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.
AI · 17d ago · simedw.com
Astute▲585Automate your B2B brand going viral, with new media creators
AI · 18d ago · company-app.joinastute.com


Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2. The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges…
AI · 27d ago · cactuscompute.com


Launched alongside, January 2025
the whole month →- IM
Hello! I'm Byran. I spent the past ~6 months engineering a laptop from scratch. It's fully open-source on GH at: https://github.com/Hello9999901/laptop
Dev tools · 2025 · byran.ee
- TITetris in a PDF▲1,289
I realized that the PDF engines of modern desktop browsers (PDFium and PDF.js) support JavaScript with enough I/O primitives to make a basic game like Tetris. It was a bit tricky to find a union of features that work in both engines, but in the end it turns out that showing/hiding annotation "fields" works well to make monochrome pixels, and keyboard input can be achieved by typing in a text input box. All in all it's quite janky but a nice reminder of how general purpose PDF scripting can be. The linked PDF is all ASCII so you can just open it in a text editor, or have a look at…
Life & fun · 2025 · th0mas.nl



Create lifelike, personalized AI avatars from text prompts
AI · 2025 · jogg.ai
