Open-source, high performance layout parsing for LLM's
In plain words
This open-source layout parsing tool is designed to help large language models process and understand document structure efficiently. It extracts and organizes text from various document formats by identifying layouts, sections, and content relationships. The tool is built for developers and researchers working with LLMs who need to parse complex documents at scale. Its focus on high performance makes it suitable for production environments where speed and accuracy matter.
written from the facts on this page · September 2026
Does the same job
all alternatives →
- OSOpen-Source Parse (Startup School 2014)2014 · divide.io · ▲28
- LALark, a modern parsing library with Earley and LALR(1) implementations2017 · github.com · ▲15
- AOAuto-optimizing deterministic LLM outputs using knowledge graphs2024 · github.com · ▲7
Hi, We are building an open-source framework for loading and structuring LLM context to create accurate and explainable LLM answers using knowledge graphs and vector stores. We built the tool with four main concepts in mind: 1. Loader -> uses dlt in the backend to load and structure the data 2. Cognify step -> creates a graph with summaries, labels and factoids that are interconnected across the documents and stored as a representation in the vector store 3. Optimizer -> Uses DSPy to optimize LLM queries, and we plan to extend it to most of the knobs we can turn, like chunking etc. 4. Search…
- LWLLM Wiki – Open-Source Implementation of Karpathy's LLM WikiApr 2026 · llmwiki.app · ▲6
- MOMetarank – open-source hybrid search with LLMs2023 · demo.metarank.ai · ▲7
A small demo for a Metarank open-source project I'm maintaining.
More ai this month
the category →
I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.
AI · 17d ago · simedw.com
Astute▲585Automate your B2B brand going viral, with new media creators
AI · 18d ago · company-app.joinastute.com


Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2. The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges…
AI · 27d ago · cactuscompute.com


Launched alongside, March 2024
the whole month →
- 3Y
Life & fun · 2024 · github.com
Microlaunch▲1,116Launch and get feedback on both the idea and product
Dev tools · 2024 · microlaunch.net


