nowfound

Alternatives

Products that do what MaskLLM does

public archive for AI hallucinations & reasoning failures

  1. 1OS

    Hi all! This morning, we released a new Apache 2.0 licensed model on HuggingFace for detecting hallucinations in retrieval augmented generation (RAG) systems. What we've found is that even when given a "simple" instruction like "summarize the following news article," every LLM that's available hallucinates to some extent, making up details that never existed in the source article -- and some of them quite a bit. As a RAG provider and proponents of ethical AI, we want to see LLMs get better at this. We've published an open source model, a blog more thoroughly describing our methodology (and…

    2023 · vectara.com

  2. 2
    RunLLM293

    AI that doesn’t just respond—it resolves

    2025

  3. 3
    MaskLLM132

    Mask your LLM APIs for secure rotation and logging

    2025

  4. 4
    Athina AI509

    Monitor LLMs and automatically detect hallucinations in prod

    2024

  5. 5
    Sup AI103

    AI ensemble that scored #1 on Humanity's Last Exam

    Apr 2026 · sup.ai

  6. 6

    Test-driven development for LLMs

    2023

  7. 7

    Everything you need to evaluate & improve prompts and LLMs

    2023

  8. 8
    Sider 4.0211

    Group chat with multiple AI bots to reduce hallucinations

    2023

  9. 9FA

    This is a quick prototype I built for semantic search and factual question answering using embeddings and GPT-3. It tries to solve the LLM hallucination issue by guiding it only to answer questions from the given context instead of making things up. If you ask something not covered in an episode, it should say that it doesn't know rather than providing a plausible, but potentially incorrect response. It uses Whisper to transcribe, text-embedding-ada-002 to embed, Pinecone.io to search, and text-davinci-003 to generate the answer. More examples and explanations here:…

    2022 · huberman.rile.yt

  10. 10

    Turn PDFs into courses with AI without irrelevant additions

    2025

  11. 11

    The internet's dumbest AI fails, curated.

    Jul 2026 · fullofslop.com

  12. 12

    LLM-usage observability and monitoring tool

    2025

  13. 13
    Verol98

    Stop AI hallucinations

    Jun 2026 · chromewebstore.google.com

  14. 14DW

    The first GPT-based solution that uses hallucinations from LLMs for divergent thinking to generate new and novel ideas. Hallucinations are often seen as a negative thing, but what if they could be used for our advantage? dreamGPT is here to show you how. The goal of dreamGPT is to explore as many possibilities as possible, as opposed to most other GPT-based solutions which are focused on solving specific problems.

    2023 · github.com

  15. 15YA

    When the LLM so ahh you lowk take over its job

    Mar 2026 · youraislopbores.me

  16. 16HC
  17. 17

    Real-time Hallucination Detection for LLMs

    Nov 2025

  18. 18

    The kill-switch for AI hallucinations. Ship reliable AI.

    Jan 2026 · deeprails.com

  19. 19SA

    Hi HN. I'm Ken, a 20-year-old Stanford CS student. I built Sup AI. I started working on this because no single AI model is right all the time, but their errors don’t strongly correlate. In other words, models often make unique mistakes relative to other models. So I run multiple models in parallel and synthesize the outputs by weighting segments based on confidence. Low entropy in the output token probability distributions correlates with accuracy. High entropy is often where hallucinations begin. My dad Scott (AI Research Scientist at TRI) is my research partner on this. He sends me papers…

    Mar 2026 · sup.ai

  20. 20OA

    We built tooling that connects LLMs directly to case law databases with citation verification to address hallucination in legal AI. Think of it as giving the model access to actual legal sources instead of relying on training data.

    Feb 2026 · openjuris.org

  21. 21MR

    The most common failures for production agents are behavioral: looping, reasoning leakage, user frustration, and more. Using a frontier model like GPT or Sonnet to judge every turn is too expensive and slow to run at scale. To solve this, we built Reflexes: semantic signals from agent traces, served fast and cheap over API. Built on custom kernels and a custom inference engine forked from vLLM. Under the hood, it is a small LLM architected around multi-head inference. Small models need to be trained for specific tasks, but running 50 separate small models on the same input for 50 tasks makes…

    Jun 2026

  22. 22

    Stop AI sycophancy. Force a multi-agent debate.

    Jan 2026

  23. 23AT

    We kept shipping “simple” LLM features that were fluent-but-wrong. After too many postmortems we wrote down the failure patterns and added a small reasoning layer in front of the model. It’s model-agnostic, sits beside your existing stack, and you can implement it from a single PDF (MIT). What’s inside the PDF A problem map of 16 failure modes we kept hitting in real systems (OCR/layout drift, table-to-question mismatches, embedding≠meaning, pre-deploy collapse, etc.). Four lightweight gates you can add today: Knowledge-boundary canaries (empty/adversarial/known-fact probes).…

    2025 · github.com

  24. 24IM

    Atrophy is an iOS self-report quiz aimed at software engineers who use LLMs heavily enough at work to wonder if they're trending toward AI over-reliance or some form of AI psychosis. I built it because I noticed a pattern: formerly AI-skeptical coworkers now open every standup or design discussion with "I asked Claude..." or "Claude told me..." for technical problems and design decisions. I've felt the same pull myself to delegate every task or problem to AI. It's easy to lean on these tools for almost any amount of critical thinking or problem solving, and I'm worried about what it means…

    May 2026 · apps.apple.com

Ranked by how close each launch is in meaning, then by votes. Refine with a description →