nowfound

Alternatives

Products that do what llamafile 0.10.0 rebuilt, Qwen3.5, lfm2, Anthropic API does

  1. 1
    Llama 2263

    The next generation of Meta's open source LLM

    2023

  2. 2AT
  3. 3
    Llama 4423

    A new era of natively multimodal AI innovation

    2025

  4. 4
    Llama312

    3.1-405B: an open source model to rival GPT-4o / Claude-3.5

    2024

  5. 5

    Llama 405B-level performance, at a fraction of the cost

    2024

  6. 6LF
  7. 7LA
  8. 8CL
  9. 9M1
  10. 10GH

    2015 · justineo.github.io

  11. 11LA

    A simple mobile web app inspired by Fuzzy-Search/realtime-bakllava that uses llama.cpp server backend with multimodal mode to describe and narrate what the phone camera sees. I built this thing in a few hours using a single ChatGPT thread to generate most things for me and iterate on this project. Here's the workflow: https://chat.openai.com/share/ea84ec69-5617-45e8-8772-ac2dcf...

    2023 · github.com

  12. 12RL

    2024 · app.wiz.chat

  13. 13NT
  14. 14IB

    Hey HN, I built a website where you can train Llama 3.1 8b & 70b (4bit) on your data. I use unsloth in the backend and the training is done on H100s which I rent programmatically from Runpod. I'd love some feedback. If you would be interested in using it feel free to book a chat with me: cal.com/hamada/tunellama-intro Happy to give you free credits :) P.S. I'm also looking for a co-founder as I have big plans for this.

    2024 · tunellama.com

  15. 15WH
  16. 16RL

    You can now build serverless AI inference web application with ggml.js's LM backends.

    2023 · rahuldshetty.github.io

  17. 17GA

    We’ve just launched Gradient — an API that helps you build private LLMs that you own. We simplify inference and fine-tuning on open-source LLMs such as llama2, and you only pay by the token. Our API platform makes it possible for you to create private models with a single API call. Run inference on your fine tuned model instantly with no cold boot (and no need to pay for compute costs). The product is truly on demand - when you run fine tuning and inference on our platform, there's nearly 0 startup latency for these API calls. And you're not paying for the compute, you just pay for the…

    2023 · gradient.ai

  18. 18L2

    2023 · woolapi.com

  19. 19PA
  20. 20OJ
  21. 21MO

    2025 · mermaidchart.com

  22. 22DV

    2019 · devilbox.discourse.group

  23. 23D0
  24. 24BA

Ranked by how close each launch is in meaning, then by votes. Refine with a description →