Swiftgum (Open Source) – Turn Data into LLM-Ready Markdown
Hey HN, We’re the co-founders of Swiftgum (swiftgum.com), an open-source platform that ingests and normalizes documents from virtually any platform, ready for your favorite LLM. Swiftgum lets you connect various data sources like Google Drive and Notion, transform the data into markdown, and load it into your vector database of choice, such as Postgres, Supabase, Milvus, or Weaviate. Why We Built It Building a basic RAG system is easy—just embed your documents and query them with an LLM. But for enterprise or multi-user scenarios, you quickly run into challenges: Data sources are messy.…
What it does
In the maker’s words, at launch
Hey HN, We’re the co-founders of Swiftgum (swiftgum.com), an open-source platform that ingests and normalizes documents from virtually any platform, ready for your favorite LLM. Swiftgum lets you connect various data sources like Google Drive and Notion, transform the data into markdown, and load it into your vector database of choice, such as Postgres, Supabase, Milvus, or Weaviate. Why We Built It Building a basic RAG system is easy—just embed your documents and query them with an LLM. But for enterprise or multi-user scenarios, you quickly run into challenges: Data sources are messy. Files come in different formats, and converting them into a structured, AI-ready format is tedious. Users want to add their own documents to a shared knowledge base, creating potential conflicts or duplication. Permissions explode in complexity, especially when each user decides which docs to share (or not). Maintaining a consistent pipeline for ingestion, format conversion, and RBAC across multiple sources becomes a heavy lift. How Swiftgum Helps Open Source & Self-Hostable The codebase is available on GitHub (github.com/Swiftgum/swiftgum), so you can review, fork, or run Swiftgum on-prem for total control. No black boxes—you see exactly how data is ingested, transformed, and secured. Ingest & Normalize Any Document Swiftgum extracts, cleans, and converts documents from Drive, Notion, and other sources into LLM-ready Markdown. No more manual reformatting—just plug in your data and start querying. Per-User Sharing & RBAC Users decide what they want to share or keep private. Swiftgum enforces these settings in real time, ensuring only authorized persons can access specific documents. Flexible Data Export Instead of locking you into a specific database, Swiftgum lets you export data via webhooks, so you can send it to your preferred system. Whether it’s a vector database, a storage bucket, or a custom pipeline, you stay in control. Use Cases Enterprise AI Assistants & Knowledge Bases: Let every team member contribute documents while keeping private info private. SaaS Platforms: Provide a multi-tenant RAG service for your customers without custom-building complex permissioning logic. Compliance-Heavy Industries: Enforce fine-grained data visibility in finance, healthcare, and other regulated sectors—fully auditable and open-source. Why We’re Different Turns any data into LLM-ready Markdown. No need to clean, convert, or process documents manually. Truly multi-user. One centralized RAG pipeline with role-based access at every step. Open source & transparent. Unlike proprietary vector or document platforms, Swiftgum gives you full control. Try It or Contribute Get started with our hosted or on-prem options at swiftgum.mintlify.app/getting-started/quick-start. Check out the GitHub repo (github.com/Swiftgum/swiftgum) to spin up your own instance or contribute features. Our documentation (swiftgum.mintlify.app/introduction) explains how to connect data sources, configure user permissions, and integrate with your LLM environment. Feedback Wanted What data sources should we support next—SharePoint, Confluence, or something else? What tricky RBAC edge cases do you face that we can handle out-of-the-box? What deployment approach do you prefer—Docker, Kubernetes, or something else? We’d love your thoughts on making multi-user RAG simpler, more secure, and fully open-source. We’ll be here in the comments to answer questions—thanks for reading! — The Swiftgum Team
Does the same job
all alternatives →



- DAdstack – an open-source tool to build data applications easily2020 · ▲134
Dear HN, I am Riwaj, the cofounder of dstack.ai (https://github.com/dstackai). A few months ago, we built an online service that allows users to publish data visualizations from Python or R. The idea was to build a tool that did not require additional programming or front-end development for publishing data visualizations. Such a code can be invoked from either Jupyter notebook, RMarkdown, Python, or R scripts. Once the data is pushed, it can be accessed via a browser. Open-sourcing dstack: During our customer discovery phase, we realized that dstack.ai should integrate a lot…
More ai this month
the category →
I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.
AI · 17d ago · simedw.com
Astute▲585Automate your B2B brand going viral, with new media creators
AI · 18d ago · company-app.joinastute.com


Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2. The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges…
AI · 27d ago · cactuscompute.com


Launched alongside, February 2025
the whole month →
Screen Studio 3.0▲1,833Beautiful screen recordings with instant shareable links
Growth · 2025 · screen.studio
- IG
I was at FB/Meta from late 2013 to early 2023, mostly working in the compiler/runtime spaces. I got hit in the spring 2023 layoff wave. I immediately started making games in my newfound free time (a lifelong interest, and I even worked in AA(A?) back ca. ~2000), and in October 2023 I stumbled upon the idea of a roguelike pachinko/plinko game inspired by Luck Be A Landlord. Things snowballed quickly, I started talking to publishers, then worked like crazy through all of 2024, almost the hardest I've ever worked in my career, and launched the game in December 2024. It's sold…
Work · 2025


- IB
i wanted to change the habit of reaching for my phone in the morning and doomscrolling away an hour so i built an app to help me. now i have to literally touch grass before accessing my most distracting apps the app is built in swiftui, uses the screen time apis provided by apple and google vision to recognise grass or not i'd love to get your thoughts on the concept.
Life & fun · 2025 · touchgrass.now