nowfound

Alternatives

Products that do what Critical Thinking MCP does

Deterministic enforcement for LLM reasoning quality

  1. 1
    Mercury 2152

    Fastest reasoning LLM built for instant production AI

    Feb 2026 · inceptionlabs.ai

  2. 2LV

    This is a weekend hack that I'd like to further develop as it's working surprisingly well. Using MCTS, we can explore a space of possible verified programs with an LLM. We check the partial programs at each step, and so steer towards programs that pass the verifier. https://github.com/namin/llm-verified-with-monte-carlo-tree-...

    2023 · github.com

  3. 3FL

    Recently I've been working on making LLM evaluations fast by using bayesian optimization to select a sensible subset. Bayesian optimization is used because it’s good for exploration / exploitation of expensive black box (paraphrase, LLM). I would love to hear your thoughts and suggestions on this!

    2024 · github.com

  4. 4

    The modern standard in AML compliance through AI agents

    2023

  5. 5

    Your LLM doesn't think. We make sure of it.

    Apr 2026 · tictacguy.github.io

  6. 6ZC

    Zero-Knowledge Proofs (ZKPs) let an untrusted proved show that computation was executed correctly without revealing the inputs to the verifier. However to prove anything, the computation first has to be expressed as a circuit: a system of polynomial equations (constraints) over a finite field. Circuits are the assembly language of zk and every constraint costs prover (and sometimes verifier) time, so production circuits are aggressively hand-optimized. Over the last months, we have been experimenting with writing formal specifications instead and letting LLMs produce the circuits: as long as…

    Jul 2026 · zk.golf

  7. 7AI
  8. 8AT

    We kept shipping “simple” LLM features that were fluent-but-wrong. After too many postmortems we wrote down the failure patterns and added a small reasoning layer in front of the model. It’s model-agnostic, sits beside your existing stack, and you can implement it from a single PDF (MIT). What’s inside the PDF A problem map of 16 failure modes we kept hitting in real systems (OCR/layout drift, table-to-question mismatches, embedding≠meaning, pre-deploy collapse, etc.). Four lightweight gates you can add today: Knowledge-boundary canaries (empty/adversarial/known-fact probes).…

    2025 · github.com

  9. 9AM

    Hi there, me and some friends were inspired by Simon Willison's recent post on the "lethal trifecta" (https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/ ) and started building a gateway to defend against it. The idea: instead of connecting an LLM directly to multiple MCP servers, you point them all through a Gateway. The Gateway: - Connects to each MCP server and inspects their tools + requirements - Classifies tools along the "trifecta" axes (private data access, untrusted content, external comms) - When all three conditions are about to align in a…

    Sep 2025 · github.com

  10. 10LC

    Prompt instructions like 'never do X' don't hold up in production. LLMs ignore them when context gets long or users push hard. Limits sits between your agent and the real world. Every action — database writes, API calls, refunds — gets intercepted and checked against your rules before it executes. Deterministically. No LLM involved in enforcement. Three modes: Conditions: hard rules on structured data Guideance: validate LLM output before it reaches the user and give the agent chance to reason and retry Guardrails: scan for PII, toxicity, prompt injection etc One line to integrate: npm…

    Feb 2026 · limits.dev

  11. 11AM
  12. 12TA
  13. 13GO

    LLMs are better at being the "mouth" than the "brain" and I can prove it mathematically. I built a deterministic graph engine that offloads reasoning from the LLM. It reduces token usage by 89% and makes a tiny 0.8B model trace enterprise execution paths flawlessly. Here is the white paper and the reproducible benchmark.

    Mar 2026 · github.com

  14. 14CM

    Hey HN, I've been building AutoAgents, an AI agent framework in Rust. Today I'm sharing a feature I haven't seen done well elsewhere: composable middleware layers for LLM inference pipelines. The problem Every agent framework lets you swap LLM providers. Almost none of them give you a structured way to enforce safety, caching, or data sanitization in the inference path itself. You end up with guardrails as application-level if-statements, caching bolted on as a separate service, and PII handling as a "we'll add it later" TODO that never ships. This gets worse with local models. Cloud APIs…

    Mar 2026 · github.com

  15. 15MG

    Many teams connecting LLMs to external tools eventually encounter the same architectural issue: as more tools and agents are added, the integration pattern becomes an N×M mesh of direct connections. Each agent implements its own auth, retries, rate limiting, and logging; each tool needs credentials distributed to multiple places and observability becomes fragmented. We built LLM gateway with this goal to provide a single place to manage authentication, authorization, routing, and observability for MCP servers, with a path toward a more general agent-gateway architecture in the future. The…

    Dec 2025 · truefoundry.com

  16. 16LD
  17. 17RA
  18. 18

    Continuous evaluation of LLM reasoning on competitive code

    Dec 2025

  19. 19
    KDD4

    Govern AI agents with deterministic gates. No LLM judging.

    Jul 2026 · mauricioperera.github.io

  20. 20WA

    WFGY introduces a PDF-based semantic protocol designed to correct projection collapse, contradiction loops, and ambiguous inference chains in LLMs. No retraining. No system calls. When parsed, the logic patterns alter reasoning trajectories directly. Prompt evaluation benchmarks show: ‣ +42.1% reasoning success ‣ +22.4% semantic alignment ‣ 3.6× stability in interpretive tasks The repo contains formal theory, prompt suites, and reproducible results. Zero dependencies. Fully open-source. Feedback from those working in alignment, interpretability, and logic-based scaffolding would be…

    2025 · github.com

  21. 21KL

    LLM agents often place raw JSON tool outputs directly in the prompt. After a few tool calls, earlier results get compacted or truncated and answers become incorrect or inconsistent. I built Sift, a drop-in MCP gateway that stores tool outputs as local artifacts (filesystem blobs indexed in SQLite) and returns an `artifact_id` plus compact schema hints when responses are large or paginated. Instead of reasoning over full JSON in the prompt, the model runs a small Python query: def run(data, schema, params): return max(data, key=lambda x: x["magnitude"])["place"] Query code runs in a…

    Mar 2026 · github.com

  22. 22

    Simple Deterministic Guardrails for LLM/Agent

    Feb 2026

  23. 23

    Deterministic, LLM-as-a-Judge evaluation framework.

    Mar 2026 · github.com

  24. 24

    Catch LLM quality drift before your users do

    Jun 2026 · regtrace-docs.vercel.app

Ranked by how close each launch is in meaning, then by votes. Refine with a description →