nowfound

Alternatives

Products that do what ATLAS does

Benchmark - Adaptive Testing of Learning Across Substrates

  1. 1
    Atlas.new525

    The AI agent for maps and spatial data

    Jan 2026

  2. 2

    ML dev tool that saves you up to 8x in cloud GPU costs

    2019

  3. 3

    An open-source teacherbot

    2018

  4. 4

    The self-learning knowledge base that improves itself

    Jul 2026 · usefini.com

  5. 5
    Atlaso216

    One memory for every AI you use

    Aug 2026 · atlaso.ai

  6. 6

    A benchmark evaluating expert-level scientific reasoning

    Dec 2025

  7. 7
    Web Bench138

    A 10x better benchmark for AI browser agents

    2025

  8. 8UD

    Hey HN! I’m the founder of Unify, and we’ve just released our Model Hub, which provides a collection of LLM endpoints with live runtime benchmarks all plotted across time: https://unify.ai/hub A key finding is that static tabular runtime benchmarks for LLMs simply do not work. It’s necessary to take a time-series perspective, and plot the variations through time. We currently have 21 models provided by: Anyscale, Perplexity AI, Replicate, Together AI, OctoAI, Mistral AI and OpenAI, with more on the roadmap. We test across different regions (Asia, US, Europe), with varied…

    2024

  9. 9
    cto bench125

    The ground truth code agent benchmark

    Dec 2025

  10. 10PG

    I’m Andrew, co-founder of Recall. Over the past few days I’ve been building Predict, a playground where anyone can: - propose skills we should measure in language models—live examples include difficult math, memory-manipulation resistance, code generation, and empathy under bad news - write evals (graded prompts) for those skills - forecast which models will score highest once GPT-5 is released Why this exists Benchmarks leak into training data quickly; scores are unreliable and labs still declare progress. The prediction tool aims keeps the target moving by letting the crowd define both the…

    2025

  11. 11CB

    Hey HN, we're excited to share Cua-Bench ( https://github.com/trycua/cua ), an open-source framework for evaluating and training computer-use agents across different environments. Computer-use agents show massive performance variance across different UIs—an agent with 90% success on Windows 11 might drop to 9% on Windows XP for the same task. The problem is OS themes, browser versions, and UI variations that existing benchmarks don't capture. The existing benchmarks (OSWorld, Windows Agent Arena, AndroidWorld) were great but operated in silos—different harnesses,…

    Jan 2026 · github.com

  12. 12
    Atlas4

    Deterministic code intelligence for developers and AI

    Jul 2026 · atlas.aziro.com

  13. 13AA

    Atlas is an open-source deployment pipeline platform built for cloud-native applications. Atlas allows users to: - Create continuous pipelines across all their environments and clusters - Add custom tasks/tests plugins (Python scripts, K8S manifests, Argo Workflows, environment setup, etc.) - Automatically rollback applications in case of failure or degradation (Atlas watches the application past the scope of a pipeline run to ensure and enforce stability) - Use all existing Argo features Would love to hear all of your feedback and thoughts on this!

    2022 · greenops.io

  14. 14
    Atlas1

    Local-first memory and evidence for AI agents

    5d ago · github.com

  15. 15

    AI-powered adaptive learning built for higher education

    Apr 2026 · project-atlas.wynexlabs.studio

  16. 16UD

    We can now build drastically higher quality search because we can use LLMs in algorithms that mimic a human's systematic research process, instead of just roughly recommending results based on semantic embeddings or term frequency. We built a deep search LLM pipeline that takes a few minutes to carefully search all the scientific literature. You describe your complex goal, as you would to a colleague. Then, we carefully search 200M+ papers. We classify the preliminary results with GPT-4. We then adapt the search goals based on relevant/irrelevant papers uncovered and continue searching,…

    2024 · undermind.ai

  17. 17BR

    I built BenchFlow, an open-source framework that lets you integrate and evaluate AI tasks using Docker-based benchmarks. You can try it out right now by cloning the repo and running a benchmark in minutes. As an AI researcher, I was frustrated with how much time my team spent setting up benchmark environments rather than actually improving our models. We'd spend weeks configuring environments, only to find inconsistencies when comparing results with other teams. BenchFlow started as an internal tool to standardize our evaluation process, and we decided to open-source it after seeing how much…

    2025 · github.com

  18. 18
    Atlas2

    Turn any topic into structured knowledge maps

    Mar 2026 · atlas.eduloraa.com

  19. 19

    Pre-test marketing decisions with AI consumer simulations

    Jan 2026

  20. 20

    Find real problems worth building — scored, analyzed, ready

    Apr 2026 · problematlas.ai

  21. 21ΤB

    τ-Bench is an open benchmark for evaluating AI agents on grounded, multi-turn customer service tasks with verifiable outcomes. It's been great to see the community adopt it since launch — this is now the third iteration. With τ³-Bench, we're extending it to two new settings: knowledge-intensive retrieval and full-duplex voice. τ-Knowledge: agents must navigate ~700 interconnected policy documents to complete multi-step tasks. Best frontier model (GPT-5.2, high reasoning) hits ~25%. The surprising part: even when you hand the model the exact documents it needs, performance only reaches ~40%.…

    Mar 2026

  22. 22BY

    we had hundreds of discussions with engineering leaders over the past few months, and everyone's trying to understand where they are in the AI journey. we collected all this data into a benchmark and built a free grader to let you know where you stand. you answer on a 1–5 scale (e.g., autonomy runs from "suggestions only" to "agents own multi-hour workflows across code, infra, and external systems") - takes about 5 minutes. https://agent-benchmarks.com/software-factory/ waiting for your results!

    Jul 2026 · agent-benchmarks.com

  23. 23MD

    We’re excited to share ML-Dev-Bench, a new open-source benchmark that tests AI agents on real-world ML development tasks. Unlike typical coding challenges or Kaggle-style competitions, our benchmark simulates end-to-end ML workflows including: - Dataset handling and preprocessing - Debugging model and code failures - Implementing new model architectures - Fine-tuning and improving existing models With 30 diverse tasks, ML-Dev-Bench evaluates agents across critical stages of ML development. To complement this, we built Calipers, a framework that provides systematic performance evaluation and…

    2025 · github.com

  24. 24RT

    I've written Atlas, a GPU scripting language that eliminates the boilerplate of managing textures and uniforms. Here are some demos including 4D fractal exploration with gamepad controls. Press 7 to see the Julia set, and try reloading if you see rectangles/it glitches. Documentation: https://banditcat.github.io/Atlas/index.html *requires approximately an RTX 3080.

    Nov 2025 · banditcat.github.io

Ranked by how close each launch is in meaning, then by votes. Refine with a description →