nowfound

Alternatives

Products that do what GPT-4V(ision) powered most reliable browser agent, OSS does

  1. 1

    Reliable Web Agents and Workflows

    2025

  2. 2

    AutoGPT in the browser

    2023

  3. 3
    GPT-5.6340

    A new standard for intelligence and efficiency

    Jul 2026 · openai.com

  4. 4WC

    Using threaded emscripten to speed up the generation and offload the main loop. No SIMD or other optimizations. Might work faster with #enable-experimental-webassembly-features enabled. Tested in x86 Chrome and Firefox, Apple Silicon Safari Run it yourself: https://github.com/lxe/ggml/tree/wasm-demo Thanks, https://github.com/ggerganov/ggml,

    2023 · lxe.co

  5. 5G3

    2022 · musings.yasyf.com

  6. 6

    Create intelligence for your products

    2024 · gca.dev

  7. 7

    GPT vision first open source browser automation

    2023

  8. 8

    Cloud agent for parallel dev tasks, powered by Codex-1

    2025

  9. 9BD
  10. 10LA

    I built LocalGPT over 4 nights as a Rust reimagining of the OpenClaw assistant pattern (markdown-based persistent memory, autonomous heartbeat tasks, skills system). It compiles to a single ~27MB binary — no Node.js, Docker, or Python required. Key features: - Persistent memory via markdown files (MEMORY, HEARTBEAT, SOUL markdown files) — compatible with OpenClaw's format - Full-text search (SQLite FTS5) + semantic search (local embeddings, no API key needed) - Autonomous heartbeat runner that checks tasks on a configurable interval - CLI + web interface + desktop GUI - Multi-provider:…

    Feb 2026 · github.com

  11. 11
    GPT‑5.4475

    OpenAI's most efficient model: less tokens, more clarity

    Mar 2026

  12. 12BA

    We’re very excited to share something we’ve been building. Notte https://www.notte.cc/ is a full-stack browser agent platform built to reliably automate a wide range of workflows. Browser agents aren’t new, but what is still hard is covering real-world flows reliably. The inspiration for Notte was to make a full-featured platform that bridges the agent reliability gap. We’ve packaged everything via a singe API for ease of use: - Site Interactions - Observe website states, scrape data and execute actions - Structured Output - Get data in your exact format with Pydantic models -…

    2025 · github.com

  13. 13IR
  14. 14

    A version of GPT-5 better at agentic coding

    Sep 2025

  15. 15

    A Chrome extension that spots AI-generated content

    2022

  16. 16AR
  17. 17IN

    Hey HN, Robert from Laminar (lmnr.ai) here. We built Index - new SOTA Open Source browser agent. It reached 92% on WebVoyager with Claude 3.7 (extended thinking). o1 was used as a judge, also we manually double checked the judge. At the core is same old idea - run simple JS script in the browser to identify interactable elements -> draw bounding boxes around them on a screenshot of a browser window -> feed it to the LLM. What made Index so good: 1. We essentially created browser agent observability. We patched Playwright to record the entire browser session while the agent operates,…

    2025 · github.com

  18. 18
    XP1168

    GPT-based Assistant with access to your Tabs

    2022

  19. 19

    Open safety reasoning models with custom safety policies

    Oct 2025

  20. 20CA

    Hey Hacker News! Launching gptengineer.app into beta today. It's like Claude Artifacts, but: - you can edit the code in your fav IDE (two-way github sync) - installs npm packages - automatically picks up build and runtime errors and fixes them - very fast, built with rust The full stack capabilities are built on supabase (prefer to not have to handle auth + user data at this point so this is owned by the user) The seed for this project was an open source experiment, posted about that previously here: https://news.ycombinator.com/item?id=36422730 Would love feedback if you give…

    2024 · gptengineer.app

  21. 21OS

    Hi HN, I forked chromium and built agent-browser-protocol (ABP) after noticing that most browser-agent failures aren’t really about the model misunderstanding the page. Instead, the problem is that the model is reasoning from a stale state. ABP is designed to keep the acting agent synchronized with the browser at every step. After each action (click, type, etc), it freezes JavaScript execution and rendering, then captures the resulting state. It also compiles the notable events that occurred during that action loop, such as navigation, file pickers, permission prompts, alerts, and downloads,…

    Mar 2026 · github.com

  22. 22TV

    Hey HN! I built a tool that gives LLMs the ability to understand the visual structure of a webpage even if they don't accept image input. We've found that unimodal GPT-4 + Tarsier's textual webpage representation consistently beats multimodal GPT-4V/4o + webpage screenshot by 10-20%, probably because multimodal LLMs still aren't as performant as they're hyped to be. Over the course of experimenting with pruned HTML, accessibility trees, and other perception systems for web agents, we've iterated on Tarsier's components to maximize downstream agent/codegen performance. Here's the…

    2024 · github.com

  23. 23
    BU138

    Openclaw in the cloud

    Mar 2026

  24. 24

    Codex-powered agents for teams.

    Apr 2026

Ranked by how close each launch is in meaning, then by votes. Refine with a description →