Baseer
A vision-language model that outperforms GPT-5 on Arabic OCR
What it does
Baseer is a vision-language model built specifically for Arabic documents it outperforms GPT-5, Gemini 2.5 Pro, and Azure Document Intelligence on OCR benchmarks. Extracts text, tables (HTML), and equations (LaTeX) while preserving structure. API, on-prem, or web.
موقع بصير المدعوم بالذكاء الاصطناعي لاستخراج النصوص العربية بدقة فائقة من المستندات والصور وملفات PDF، وتحويلها لملفات قابلة للتعديل والتحرير.
يقوم بصير بقراءة المستندات بطريقة طبيعية وسلسة، كما يقرأها الإنسان تماماً، مع الحفاظ على التنسيق والترتيب الأصلي للنص. يحوّل الجداول إلى تنسيق HTML منظم وجاهز للاستخدام، مما يسهل عرضها ومعالجتها في أي تطبيق. يوفر بصير بيئة معالجة مستندية متكاملة تُمكّنك من رفع ملفاتك العربية وتحرير محتواها بدقة، مع إمكانية تصدير النتائج النهائية بصيغ Word وPDF وفق معايير جاهزة للاستخدام المؤسسي. قمنا باختبار 'بصير' المتخصص والمدرب لفهم المستندات والصور العربية وتحويلها إلى نصوص منظمة وقابلة للاستخدام ومقارنته بأقوى النماذج العالمية , أثبت بصير تفوقاً نوعياً في فهم النصوص والحفاظ على تماسك الجداول، ليمنحك أدق مخرجات رقمية يمكن لنموذج ذكاء اصطناعي تحقيقها اليوم. نفتخر بمسراج لاب بتطوير 'بصير' كأفضل نموذج…from baseerocr.com
Does a similar job
all alternatives →- ZDZerox – Document OCR with GPT-mini2024 · github.com · ▲246
This started out as a weekend hack with gpt-4-mini, using the very basic strategy of "just ask the ai to ocr the document". But this turned out to be better performing than our current implementation of Unstructured/Textract. At pretty much the same cost. I've tested almost every variant of document OCR over the past year, especially trying things like table / chart extraction. I've found the rules based extraction has always been lacking. Documents are meant to be a visual representation after all. With weird layouts, tables, charts, etc. Using a vision model just make sense! In…
- OAOCR Arena – A playground for OCR modelsNov 2025 · ocrarena.ai · ▲216
I built OCR Arena as a free playground for the community to compare leading foundation VLMs and open-source OCR models side-by-side. Upload any doc, measure accuracy, and (optionally) vote for the models on a public leaderboard. It currently has Gemini 3, dots.ocr, DeepSeek, GPT5, olmOCR 2, Qwen, and a few others. If there's any others you'd like included, let me know!
- BVBenchmarking VLMs vs. Traditional OCR2025 · getomni.ai · ▲146
Vision models have been gaining popularity as a replacement for traditional OCR. Especially with Gemini 2.0 becoming cost competitive with the cloud platforms. We've been continuously evaluating different models since we released the Zerox package last year (https://github.com/getomni-ai/zerox). And we wanted to put some numbers behind it. So we’re open sourcing our internal OCR benchmark + evaluation datasets. Full writeup + data explorer here: https://getomni.ai/ocr-benchmark Github: https://github.com/getomni-ai/benchmark Huggingface:…
- OPOCR pipeline for ML training (tables, diagrams, math, multilingual)2025 · github.com · ▲170
Hi HN, I’ve been working on an OCR pipeline specifically optimized for machine learning dataset preparation. It’s designed to process complex academic materials — including math formulas, tables, figures, and multilingual text — and output clean, structured formats like JSON and Markdown. Some features: • Multi-stage OCR combining DocLayout-YOLO, Google Vision, MathPix, and Gemini Pro Vision • Extracts and understands diagrams, tables, LaTeX-style math, and multilingual text (Japanese/Korean/English) • Highly tuned for ML training pipelines, including dataset generation and…


More ai this month
the category →
I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.
AI · 17d ago · simedw.com
Astute▲585Automate your B2B brand going viral, with new media creators
AI · 19d ago · company-app.joinastute.com


Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2. The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges…
AI · 27d ago · cactuscompute.com


Launched alongside, July 2026
the whole month →- IR
I might be the only SRE on Earth with his own bowling center. It's a more in-depth gig than you'd think. My family and I bought an abandoned 8-lane bowling center in the rural mid-west. In our small town there weren't many recreation options for families. You've heard of a food desert? This is an R&R desert. It had been abandoned for a good reason. The roof leaks, the electrical system was constantly surging, and my 70-year-old bowling equipment (still) doesn't work perfectly. The system that keeps your score is particularly interesting to me. It's the thing you watch during your game, but…
Life & fun · Jul 2026

- 1W18 Words▲1,160
Life & fun · Jul 2026 · 18words.com
- BA
Over the past few months, our team has been building more and more slidedecks using web frontend technologies with coding harnesses like Claude Code, but a common complaint is to make even small edits we need to edit the code either manually or via the harness. To avoid this loop, I ended up creating Bento, a single HTML file with everything you need in a slide tool including animations and shared editing. There's no install or cloud login, everything works offline. The default deck is around 560 KB and it doesn't need to fetch anything once you got it. Open it in a browser and then you can…
Dev tools · Jul 2026 · bento.page
- GG
A few days ago I found myself trying out GLM 5.2 and was really positively impressed. The capabilities and security I was getting from this LLM are similar to those I've gotten from models like Claude or GPT, and this really surprised me. But then I thought, "I wonder how it would work on a normal computer like mine," and above all, "I wonder if it would work without going into OOM on a computer like mine." So I started working with the help of agents to test this possibility. I started converting the model to int4, understanding MTP usage, and if possible implementing DSA for long context.…
AI · Jul 2026 · github.com
