nowfound

Alternatives

Products that do what GLM-OCR does

SOTA document parsing & OCR in just 0.9B parameters

  1. 1

    The 1T Parameters Open-Source Thinking Model - SOTA on HLE

    Nov 2025 · moonshotai.github.io

  2. 2BV

    Vision models have been gaining popularity as a replacement for traditional OCR. Especially with Gemini 2.0 becoming cost competitive with the cloud platforms. We've been continuously evaluating different models since we released the Zerox package last year (https://github.com/getomni-ai/zerox). And we wanted to put some numbers behind it. So we’re open sourcing our internal OCR benchmark + evaluation datasets. Full writeup + data explorer here: https://getomni.ai/ocr-benchmark Github: https://github.com/getomni-ai/benchmark Huggingface:…

    2025 · getomni.ai

  3. 3

    Introducing the world’s best document understanding API

    2025

  4. 4Q2

    Last week was big for open source LLMs. We got: - Qwen 2.5 VL (72b and 32b) - Gemma-3 (27b) - DeepSeek-v3-0324 And a couple weeks ago we got the new mistral-ocr model. We updated our OCR benchmark to include the new models. We evaluated 1,000 documents for JSON extraction accuracy. Major takeaways: - Qwen 2.5 VL (72b and 32b) are by far the most impressive. Both landed right around 75% accuracy (equivalent to GPT-4o’s performance). Qwen 72b was only 0.4% above 32b. Within the margin of error. - Both Qwen models passed mistral-ocr (72.2%), which is specifically trained for OCR. - Gemma-3…

    2025 · github.com

  5. 5

    GLM-OCR: An online OCR focused on document structure

    Feb 2026 · glm-ocr.com

  6. 6
    GLM-4.6V239

    Open-source multimodal model with native tool use

    Dec 2025 · z.ai

  7. 7
    GLM-4.5298

    Unifying agentic capabilities in one open model

    2025

  8. 8OP

    Hi HN, I’ve been working on an OCR pipeline specifically optimized for machine learning dataset preparation. It’s designed to process complex academic materials — including math formulas, tables, figures, and multilingual text — and output clean, structured formats like JSON and Markdown. Some features: • Multi-stage OCR combining DocLayout-YOLO, Google Vision, MathPix, and Gemini Pro Vision • Extracts and understands diagrams, tables, LaTeX-style math, and multilingual text (Japanese/Korean/English) • Highly tuned for ML training pipelines, including dataset generation and…

    2025 · github.com

  9. 9

    How small can a language model be while still doing something useful? I wanted to find out, and had some spare time over the holidays. Z80-μLM is a character-level language model with 2-bit quantized weights ({-2,-1,0,+1}) that runs on a Z80 with 64KB RAM. The entire thing: inference, weights, chat UI, it all fits in a 40KB .COM file that you can run in a CP/M emulator and hopefully even real hardware! It won't write your emails, but it can be trained to play a stripped down version of 20 Questions, and is sometimes able to maintain the illusion of having simple but terse conversations…

    Dec 2025 · github.com

  10. 10TJ
  11. 11ZD

    This started out as a weekend hack with gpt-4-mini, using the very basic strategy of "just ask the ai to ocr the document". But this turned out to be better performing than our current implementation of Unstructured/Textract. At pretty much the same cost. I've tested almost every variant of document OCR over the past year, especially trying things like table / chart extraction. I've found the rules based extraction has always been lacking. Documents are meant to be a visual representation after all. With weird layouts, tables, charts, etc. Using a vision model just make sense! In…

    2024 · github.com

  12. 12OA

    I built OCR Arena as a free playground for the community to compare leading foundation VLMs and open-source OCR models side-by-side. Upload any doc, measure accuracy, and (optionally) vote for the models on a public leaderboard. It currently has Gemini 3, dots.ocr, DeepSeek, GPT5, olmOCR 2, Qwen, and a few others. If there's any others you'd like included, let me know!

    Nov 2025 · ocrarena.ai

  13. 13LA

    Almost exactly 1 year ago, I submitted something to HN about using Llama2 (which had just come out) to improve the output of Tesseract OCR by correcting obvious OCR errors [0]. That was exciting at the time because OpenAI's API calls were still quite expensive for GPT4, and the cost of running it on a book-length PDF would just be prohibitive. In contrast, you could run Llama2 locally on a machine with just a CPU, and it would be extremely slow, but "free" if you had a spare machine lying around. Well, it's amazing how things have changed since then. Not only have models gotten a lot better,…

    2024 · github.com

  14. 14

    gpt-oss-120b and gpt-oss-20b open-weight language models

    2025

  15. 15
    Molmo 298

    SOTA video understanding, pointing, and tracking VLM

    Dec 2025 · allenai.org

  16. 16

    SOTA open-source T2I model with even greater realism

    Jan 2026 · qwen.ai

  17. 17

    Read documents like an image

    Oct 2025

  18. 18
    GLM-5154

    Open-weights model for long-horizon agentic engineering

    Feb 2026 · z.ai

  19. 19
    GLM-5.3254

    Coding leap from scaled post-training on the same base

    23d ago · z.ai

  20. 20

    The first open model to beat Sonnet made for productivity

    Feb 2026 · minimax.io

  21. 21

    Fully offline OCR with 100+ languages support. Full Privacy

    2025

  22. 22

    256M VLM for end-to-end document AI

    2025

  23. 23

    GPT-4o level vision model on the phone

    2025

  24. 24OP

Ranked by how close each launch is in meaning, then by votes. Refine with a description →