Exa (YC S21) – embeddings search agent with >20x recall than Google
Hey HN! I'm Will, founder of Exa (YC21 - https://exa.ai). Today we're opening up Websets, a search engine that finds massive lists of correct results given complex queries. For example, you can search for: - “software engineers in the Bay Area, with experience in startups and big tech, who know Rust and have published technical content”: https://websets.exa.ai/cm7ax5bl8003276q689v0dde5 - “US based healthcare companies, with over 100 employees and a technical founder": https://websets.exa.ai/cm6lc0dlk004ilecmzej76qx2 - "research paper about ways to…
In plain words
Exa's Websets is a search engine that retrieves large, accurate result sets in response to complex queries with multiple criteria. Users can search for specific combinations of attributes—such as software engineers with particular skill sets in certain locations, or companies meeting detailed specifications—and receive comprehensive lists of matching results. The tool uses embeddings-based search technology and is designed for researchers, recruiters, and others needing to find niche populations or specialized information across the web.
written from the facts on this page · September 2026
From the sources
In the maker’s words, at launch
Hey HN! I'm Will, founder of Exa (YC21 - https://exa.ai). Today we're opening up Websets, a search engine that finds massive lists of correct results given complex queries. For example, you can search for: - “software engineers in the Bay Area, with experience in startups and big tech, who know Rust and have published technical content”: https://websets.exa.ai/cm7ax5bl8003276q689v0dde5 - “US based healthcare companies, with over 100 employees and a technical founder": https://websets.exa.ai/cm6lc0dlk004ilecmzej76qx2 - "research paper about ways to avoid the O(n^2) attention problem in transformers, where one of the first author's first name starts with "A","B", "S", or"T", and it was written between 2018 and 2022”: https://websets.exa.ai/cm7dpml8c001ylnymum4sp11h We built Websets because the web is humanity’s grand collection of all knowledge, and yet it’s totally unorganized. So much valuable content is too hard to find. Traditional search engines, like Google, were built to handle simple keyword queries over the web, not arbitrarily complex SQL. While agentic tools like Deep Research help a bit, they rely on traditional search under the hood and so are similarly bottlenecked. Websets works well because under the hood it uses our in-house embedding-based search engine, trained specifically to handle complex natural language queries. Crucially, Websets uses LLMs to agentically verify each result to ensure correctness. Websets is therefore a test-time compute search engine – it might take minutes or even hours to run. We believe this is a worthwhile sacrifice for high value searches. It’s hard to eval these things, but we did our best and measured that Websets found 20x more results than Google on a set of complex queries and 10x more than Deep Research. These numbers could be arbitrarily higher with more compute per query. Blog post here: https://exa.ai/blog/websets-evals While Websets isn’t perfect search yet, it’s a significant first step, and we’re excited to share the first version of the product with you all. We have a free limited tier, and we set up a special HN code for Pro plans. Use PERFECTSEARCH for a two-week free trial of Pro. Can try it here: websets.exa.ai Initial launch video here: https://x.com/ExaAILabs/status/1864013080944062567 Would love to hear thoughts, feedback, and suggestions! I know HN thinks about search sometimes :)
Does the same job
all alternatives →
- BABuilding a web search engine from scratch with 3B neural embeddings2025 · blog.wilsonl.in · ▲699
- IMI made a website to semantically search ArXiv papers2024 · papermatch.mitanshu.tech · ▲324
As a grad student (and an ADHDer), I had trouble doing literature review systematically. To combat this, I made a website that finds similar papers using the meaning of the thing I am looking for. I used MixedBread's [^1] embedding model to generate vectors from the abstracts. I store and search similar vectors using Milvus [^2] and finally use Gradio [^3] to serve the frontend. I update the vector database weekly by pulling the metadata dataset from Kaggle [^4]. To speed up the search process on my free oracle instance, I binarise the embeddings and use Hamming distance as a metric. I would…

- HSHacker Search – A semantic search engine for Hacker News2024 · hackersearch.net · ▲233
Hi HN! I'm Jonathan and I built Hacker Search (https://hackersearch.net), a semantic search engine for Hacker News. Type a keyword or a description of what you're interested in, and you'll get top links from HN surfaced to you along with brief summaries. Unlike HN's otherwise very valuable search feature, Hacker Search doesn't require you to get your keywords exactly right. That's achieved by leveraging OpenAI's latest embedding models alongside more traditional indexes extracted from the scraped and cleaned up contents of the links. I think there are many more interesting things…
- ULUsing LLMs and Embeddings to classify application errors2023 · github.com · ▲65
Hi Hacker News! We’re Vadim and Chris from Highlight.io [1]. We do web app monitoring and are working on using LLMs/embeddings to add new functionality to our error monitoring product. Given that there’s a lot of founders/engineers using LLMs in their products, we figured we’d share how we built the new functionality, their impact on our workflows, and how you can try it out. Our goal was to build two features: (1) tagging errors (e.g. deeming an error as “authentication error” or a “database error”); and (2) grouping similar errors together (e.g. two errors that have a different…
More ai this month
the category →
I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.
AI · 17d ago · simedw.com
Astute▲585Automate your B2B brand going viral, with new media creators
AI · 18d ago · company-app.joinastute.com


Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2. The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges…
AI · 27d ago · cactuscompute.com


Launched alongside, February 2025
the whole month →
Screen Studio 3.0▲1,833Beautiful screen recordings with instant shareable links
Growth · 2025 · screen.studio
- IG
I was at FB/Meta from late 2013 to early 2023, mostly working in the compiler/runtime spaces. I got hit in the spring 2023 layoff wave. I immediately started making games in my newfound free time (a lifelong interest, and I even worked in AA(A?) back ca. ~2000), and in October 2023 I stumbled upon the idea of a roguelike pachinko/plinko game inspired by Luck Be A Landlord. Things snowballed quickly, I started talking to publishers, then worked like crazy through all of 2024, almost the hardest I've ever worked in my career, and launched the game in December 2024. It's sold…
Work · 2025


- IB
i wanted to change the habit of reaching for my phone in the morning and doomscrolling away an hour so i built an app to help me. now i have to literally touch grass before accessing my most distracting apps the app is built in swiftui, uses the screen time apis provided by apple and google vision to recognise grass or not i'd love to get your thoughts on the concept.
Life & fun · 2025 · touchgrass.now