nowfound

Alternatives

Products that do what Quansloth does

GUI Based on the implementation of Google's TurboQuant

  1. 1

    New LLM compression algorithm by Google

    Mar 2026 · research.google

  2. 2TW
  3. 3

    Open-source web UI to run and train AI models.

    Mar 2026

  4. 4

    Working on Mac, Linux, and Windows now. I include a simple GUI to find new models and get things built and set up. It is working quite well across a few models for me. The GitHub README and DESIGN.md files go into detail of the how/why and it's working remarkably well so far. https://github.com/notactuallytreyanastasio/shoehorn

    19d ago · notactuallytreyanastasio.github.io

  5. 5

    Run and train AI models locally on your desktop

    25d ago · unsloth.ai

  6. 6
    Unsloth241

    Finetune LLMs 2x faster, 80% less memory

    2025

  7. 7TV
  8. 81B
  9. 9

    Understands your backend's real complexity, not just syntax

    Nov 2025

  10. 10
    Giselle502

    Build and run AI workflows. Open source.

    Dec 2025

  11. 11KR

    I discovered that in LLM inference, keys and values in the KV cache have very different quantization sensitivities. Keys need higher precision than values to maintain quality. I patched llama.cpp to enable different bit-widths for keys vs. values on Apple Silicon. The results are surprising: - K8V4 (8-bit keys, 4-bit values): 59% memory reduction with only 0.86% perplexity loss - K4V8 (4-bit keys, 8-bit values): 59% memory reduction but 6.06% perplexity loss - The configurations use the same number of bits, but K8V4 is 7× better for quality This means you can run LLMs with 2-3× longer…

    2025 · github.com

  12. 12

    The sweet-spot open dense model for coding agents

    Apr 2026 · qwen.ai

  13. 13

    Vision-to-code foundation model for real GUI automation

    Apr 2026 · docs.z.ai

  14. 14
    QwQ-32B197

    Matching R1 reasoning yet 20x smaller

    2025

  15. 15

    Cheaper GPT-4 rival that supports 32K-token context windows

    2024

  16. 16
    Quanty104

    AI powered market analysis platform with GraphQL API

    2024

  17. 17DA

    Hi HN community. We are excited to open source Dataherald’s natural-language-to-SQL engine today (https://github.com/Dataherald/dataherald). This engine allows you to set up an API from your structured database that can answer questions in plain English. GPT-4 class LLMs have gotten remarkably good at writing SQL. However, out-of-the-box LLMs and existing frameworks would not work with our own structured data at a necessary quality level. For example, given the question “what was the average rent in Los Angeles in May 2023?” a reasonable human would either assume the…

    2023 · github.com

  18. 18

    Superior statistics for 10x faster decision-making

    2020

  19. 19

    Microsoft’s New Small Language Model For Complex Reasoning

    2024

  20. 20MB

    Hey HN! We're excited to share our new open-source project, Marvin. Marvin is a high-level library for building AI-powered software. We developed it to address the challenges of integrating LLMs into more traditional applications. One of the biggest issues is the fact that LLMs only deal with strings (and conversational strings at that), so using them to process structured data is especially difficult. Marvin introduces a new concept called AI Functions. These look and feel just like regular Python functions: you provide typed inputs, outputs, and docstrings. However, instead of relying on…

    2023 · github.com

  21. 21

    The open sparse MoE model for agentic coding

    Apr 2026 · qwen.ai

  22. 22WM

    We wrote our inference engine on Rust, it is faster than llama cpp in all of the use cases. Your feedback is very welcomed. Written from scratch with idea that you can add support of any kernel and platform.

    2025 · github.com

  23. 23KO

    We've open-sourced Klarity - a tool for analyzing uncertainty and decision-making in LLM token generation. It provides structured insights into how models choose tokens and where they show uncertainty. What Klarity does: - Real-time analysis of model uncertainty during generation - Dual analysis combining log probabilities and semantic understanding - Structured JSON output with actionable insights - Fully self-hostable with customizable analysis models The tool works by analyzing each step of text generation and returns a structured JSON: - uncertainty_points: array of {step, entropy,…

    2025 · github.com

  24. 24
    QU3133

    Quantum-safe MCP servers for private, verifiable inference

    2025

Ranked by how close each launch is in meaning, then by votes. Refine with a description →