nowfound

Alternatives

Products that do what Recall Predict does

Ungameable, community-powered AI benchmarks

  1. 1PG

    I’m Andrew, co-founder of Recall. Over the past few days I’ve been building Predict, a playground where anyone can: - propose skills we should measure in language models—live examples include difficult math, memory-manipulation resistance, code generation, and empathy under bad news - write evals (graded prompts) for those skills - forecast which models will score highest once GPT-5 is released Why this exists Benchmarks leak into training data quickly; scores are unreliable and labs still declare progress. The prediction tool aims keeps the target moving by letting the crowd define both the…

    2025

  2. 2

    Curate an AI that knows what you know.

    Apr 2026 · recall.it

  3. 3

    Platform for measuring and training AI agents

    2016

  4. 4IV
  5. 5

    Stop wasting tokens and re-explaining your project every session. Recall gives Claude Code durable memory — entirely offline. - raiyanyahya/recall

    Jun 2026 · github.com

  6. 6UD

    Hey HN! I’m the founder of Unify, and we’ve just released our Model Hub, which provides a collection of LLM endpoints with live runtime benchmarks all plotted across time: https://unify.ai/hub A key finding is that static tabular runtime benchmarks for LLMs simply do not work. It’s necessary to take a time-series perspective, and plot the variations through time. We currently have 21 models provided by: Anyscale, Perplexity AI, Replicate, Together AI, OctoAI, Mistral AI and OpenAI, with more on the roadmap. We test across different regions (Asia, US, Europe), with varied…

    2024

  7. 7

    An open benchmark for AI agents that test APIs

    May 2026 · resources.kusho.ai

  8. 8
    Papr113

    Predictive memory and context intelligence API for AI Agents

    Dec 2025 · papr.ai

  9. 9
    Recall95

    The only bookmark manager programmers need

    2022

  10. 10

    AI that remembers and forgets like humans.

    Apr 2026 · yourmemoryai.vercel.app

  11. 11BA

    I've been working on this tool that lets you build a personal knowledge graph from articles, blog posts, podcasts, YouTube videos, and other content you find interesting online. You can safely forget everything and trust that Recall will resurface it when something new that is related comes up. Looking forward to your thoughts and feedback on how it could be improved! The original version of Recall was posted last year nov on HN: https://news.ycombinator.com/item?id=33425947 Since then I have pivoted to a browser extension.

    2023 · recall.wiki

  12. 12

    Fastest cognitive memory for AI Agents

    Feb 2026 · deltamemory.com

  13. 13
    RoBrain71

    Shared AI memory that stops agents from repeating mistakes

    May 2026 · github.com

  14. 14

    Let’s see who can predict the future.

    Feb 2026 · predictionary.com

  15. 15

    Local predictive memory for AI agents

    May 2026 · github.com

  16. 16NA
  17. 17DC

    I’ve been using AI to generate some repetitive frontend (guilty), and while most outputs felt vibe-coded, some results were surprisingly good. So I cleaned it up and made a ranking game out of it with friends, and you can check it out here: https://www.designarena.ai/vote /vote: Your prompt will be answered by four random, anonymous models. You pick the one you prefer and crown the winner, tournament-style. /leaderboard: See the current winning models, as dictated by voter preferences. /play: Iterate quickly by seeing four models respond to the same input and…

    2025 · designarena.ai

  18. 18CB

    Hey HN, we're excited to share Cua-Bench ( https://github.com/trycua/cua ), an open-source framework for evaluating and training computer-use agents across different environments. Computer-use agents show massive performance variance across different UIs—an agent with 90% success on Windows 11 might drop to 9% on Windows XP for the same task. The problem is OS themes, browser versions, and UI variations that existing benchmarks don't capture. The existing benchmarks (OSWorld, Windows Agent Arena, AndroidWorld) were great but operated in silos—different harnesses,…

    Jan 2026 · github.com

  19. 19ΤB

    τ-Bench is an open benchmark for evaluating AI agents on grounded, multi-turn customer service tasks with verifiable outcomes. It's been great to see the community adopt it since launch — this is now the third iteration. With τ³-Bench, we're extending it to two new settings: knowledge-intensive retrieval and full-duplex voice. τ-Knowledge: agents must navigate ~700 interconnected policy documents to complete multi-step tasks. Best frontier model (GPT-5.2, high reasoning) hits ~25%. The surprising part: even when you hand the model the exact documents it needs, performance only reaches ~40%.…

    Mar 2026

  20. 20MD

    We’re excited to share ML-Dev-Bench, a new open-source benchmark that tests AI agents on real-world ML development tasks. Unlike typical coding challenges or Kaggle-style competitions, our benchmark simulates end-to-end ML workflows including: - Dataset handling and preprocessing - Debugging model and code failures - Implementing new model architectures - Fine-tuning and improving existing models With 30 diverse tasks, ML-Dev-Bench evaluates agents across critical stages of ML development. To complement this, we built Calipers, a framework that provides systematic performance evaluation and…

    2025 · github.com

  21. 21SO
  22. 22

    Gamify AI usage. Climb rankings. Attack and Rule the board.

    2025

  23. 23AL

    Hi HN! We partnered with the Atlas team to build a tool called AI Predict [0] that allows anyone to ask any question about the future and get a thoroughly researched, AI-generated prediction on how likely it is to be true. How it works: Atlas replicated a Berkeley paper [1] that showed LLMs could make predictions as accurate as the crowd. We’re using a mix of models from OpenAI and Anthropic, with information retrieval powered by NewsCatcher [2]. The system is live and fully functional, though it might struggle with hyper-local questions outside of the public domain (e.g., “Will I have…

    2024 · aipredict.fun

  24. 24
    Recall16

    One developer solves it. Every developer knows it.

    Feb 2026 · recall.team

Ranked by how close each launch is in meaning, then by votes. Refine with a description →