nowfound

Commerce · April 3, 2024

IV

I've built a locally running Perplexity clone

The video demo runs a 7b Model on a normal gaming GPU. I think it already works quite well (accounting for the limited hardware power). :)

In plain words

LLocalSearch is a free, open-source search aggregator that runs locally on consumer gaming hardware. It uses a 7-billion parameter language model to function as a Perplexity alternative, aggregating search results through LLM agents without requiring cloud services. Designed for users who want private, offline search capabilities, it demonstrates functional performance on standard gaming GPUs despite the computational constraints of local operation.

written from the facts on this page · September 2026

Pricing, as stated on its site

Free — LLocalSearch is a free, open-source locally running search aggregator using LLM Agents. No pricing model is associated with the product itself; the pricing page shown is GitHub's platform pricing, not

checked September 2026 · prices change

Does the same job

all alternatives →
  • WM
    We made glhf.chat – run almost any open-source LLM, including 405B2024 · glhf.chat · ▲161

    Try it out! https://glhf.chat/ Hey HN! We’ve been working for the past few months on a website to let you easily run (almost) any open-source LLM on autoscaling GPU clusters. It’s free for now while we figure out how to price it, but we expect to be cheaper than most GPU offerings since we can run the models multi-tenant. Unlike Together AI, Fireworks, etc, we’ll run any model that the open-source vLLM project supports: we don’t have a hardcoded list. If you want a specific model or finetune, you don’t have to ask us for it: you can just paste the Hugging Face link in and…

  • 8F
    80% faster, 50% less memory, 0% loss of accuracy Llama finetuning2023 · github.com · ▲385

    Hi HN! I'm just sharing a project I've been working on during the LLM Efficiency Challenge - you can now finetune Llama with QLoRA 5x faster than Huggingface's original implementation on your own local GPU. Some highlights: 1. Manual autograd engine - hand derived backprop steps. 2. QLoRA / LoRA 80% faster, 50% less memory. 3. All kernels written in OpenAI's Triton language. 4. 0% loss in accuracy - no approximation methods - all exact. 5. No change of hardware necessary. Supports NVIDIA GPUs since 2018+. CUDA 7.5+. 6. Flash Attention support via Xformers. 7. Supports 4bit and 16bit…

  • IM
    I made the slowest, most expensive GPT2024 · ithy.com · ▲74

    This is another one of my automate-my-life projects - I'm constantly asking the same question to different AIs since there's always the hope of getting a better answer somewhere else. Maybe ChatGPT's answer is too short, so I ask Perplexity. But I realize that's hallucinated, so I try Gemini. That answer sounds right, but I cross-reference with Claude just to make sure. This doesn't really apply to math/coding (where o1 or Gemini can probably one-shot an excellent response), but more to online search, where information is more fluid and there's no "right" search engine + text…

  • 5L
    50+ LLMs on 2 GPUs with 2-Second Swapping? We built AI-Native Runtime2025 · github.com · ▲5

    We've built InferX, a specialized runtime environment that fundamentally changes how LLMs are served. The core problem we solve is the latency bottleneck in AI inference, especially with large models. Current systems waste resources or suffer from painfully slow cold starts. InferX's AI-native architecture, with its "snapshot" technology, enables: * *Sub-2s cold starts:* Spin up models instantly. * *High density:* Serve more LLMs on the same GPUs. * *Optimal efficiency:* Maximize GPU utilization. This isn't just another API; it's a new execution layer designed from the ground up for the…

  • PW
    Perplexity Without the Filler2024 · justtheanswer.vercel.app · ▲29

    Hey HN, I've been using LLM-powered search engines like Perplexity a lot recently for question answering but had two major qualms: 1. Too much text: I'm already tired of wading through LLM-generated meaningless filler when debugging my code. The last thing I need is more of it when I'm just trying to find a simple answer. 2. The Reddit factor: I've realized that half my search queries have "Reddit" appended to them, especially for experiential questions or recommendations. This led to Just the Answer. Just the Answer uses GPT-4o to generate a unique search query for each message that is…

  • CI
    Can I run this LLM? (locally)2025 · can-i-run-this-llm-blue.vercel.app · ▲42

    One of the most frequent questions one faces while running LLMs locally is: I have xx RAM and yy GPU, Can I run zz LLM model ? I have vibe coded a simple application to help you with just that. Update: A lot of great feedback for me to improve the app. Thank you all.

More commerce this month

the category →
  • Billing that survives a processor shutdown

    Commerce · 13d ago · paymentkit.com

  • Compare your startup equity grant for free.

    Commerce · 26d ago · equitybee.com

Launched alongside, April 2024

the whole month →
  • Supabase2,328

    The Postgres developer platform is now generally available

    Dev tools · 2024 · supabase.com

  • Build your pixel-perfect booking experience with Atoms

    Dev tools · 2024 · cal.com

  • PaddleBoat1,161

    Perfect your sales pitch with realistic AI roleplays

    AI · 2024 · padboat.com

  • deco.cx 2.01,080

    Build web apps 10x faster with Deno, JSX, TS & Tailwind

    Dev tools · 2024 · decocms.com

  • IXORD AI955

    Navigate tasks, ignite creativity

    AI · 2024

  • A central nervous system for all your productivity apps

    AI · 2024 · getassista.com