nowfound

Alternatives

Products that do what Parasail does

Leading Inference Provider From Prototype to Production

  1. 1
    Bonsai105

    AI programming platform for enterprises

    2017

  2. 2WB

    2021 · bonsaibrowser.com

  3. 31B
  4. 4WM

    We wrote our inference engine on Rust, it is faster than llama cpp in all of the use cases. Your feedback is very welcomed. Written from scratch with idea that you can add support of any kernel and platform.

    2025 · github.com

  5. 5
    Bonsai465

    Contracts, invoices, and expenses for digital freelancers

    2015

  6. 6
    Paraflow 505

    The canvas-based product design agent.

    Nov 2025 · paraflow.com

  7. 7CO

    Hey HN, Henry and Roman here - we've been building a cross-platform framework for deploying LLMs, VLMs, Embedding Models and TTS models locally on smartphones. Ollama enables deploying LLMs models locally on laptops and edge severs, Cactus enables deploying on phones. Deploying directly on phones facilitates building AI apps and agents capable of phone use without breaking privacy, supports real-time inference with no latency, we have seen personalised RAG pipelines for users and more. Apple and Google actively went into local AI models recently with the launch of Apple Foundation Frameworks…

    2025 · github.com

  8. 8
    Daytona 452

    Secure and elastic infra for running your AI-generated code.

    2025

  9. 9

    Automated & beautiful proposals for creative freelancers

    2017

  10. 10
    Sparrow457

    The lightest and fastest platform for API testing

    2025

  11. 11

    Fast multimodal-native inference at scale

    Dec 2025 · gmicloud.ai

  12. 12

    Host LLMs across devices sharing GPU to make your AI go brrr

    Oct 2025

  13. 13
    Fibery AI260

    Build workspace, write, edit & automate tasks with AI

    2023

  14. 14
    GrsAi 6

    GrsAI: The lowest-priced and most stable AI API platform

    2025

  15. 15RP

    The excitement surrounding PrismML’s 1-bit/ternary Bonsai models has the industry closely watching how smartphone giants, particularly Apple, will implement LLMs on edge devices. Moving AI on-device is a brilliant and necessary strategy. It ensures absolute user privacy in alignment with EU regulations, fundamentally shifts the economics away from costly cloud inference, and paves the way for a significant hardware upgrade supercycle as users seek true AI-capable silicon. To create a smart on-device "Semantic Router," models need to reach the 27B+ parameter scale. Achieving this on a…

    Jul 2026

  16. 16

    Calculate the GPU memory you need for LLM inference

    2025

  17. 17

    Full-Stack Platform for Training Small Language Models

    Jul 2026 · freesolo.co

  18. 18
    Helix153

    Train your own AI with open-source AI and your data

    2023

  19. 19

    Cloud costs observability,management and automation platform

    Jan 2026 · cloudchipr.com

  20. 20
    Paraglide105

    Create automated AI workflows in minutes.

    2021

  21. 21
    Sage111

    Your all-in-one AI generative platform

    2023

  22. 22B1

    We took a recently released Bonsai 1.7B ternary model from PrismML (https://github.com/PrismML-Eng/Bonsai-demo) and ran our agentic evolution search on it for 6 hours to optimize the Metal kernels. The search was fully autonomous. Measured against unmodified upstream llama.cpp at the same Bonsai/Q2_0 commit, same M4 Max: - tg128: 309.82 → 442.42 t/s (+42.0%) - pp512: 4250.32 → 4622.63 t/s (+8.8%)

    May 2026 · agents2agents.ai

  23. 23

    Building the Future of AI

    Jun 2026 · futurestore.ai

  24. 24PI

    Deploying vision models is time consuming and tedious. Setting up dependencies. Fixing conflicts. Configuring TRT acceleration. Flashing (and re-flashing) NVIDIA Jetsons. A streamlined, developer-friendly solution for inference is needed. We, the Roboflow team, have been hard at work open sourcing Inference, an open source vision deployment solution. Our solution is designed with developers in mind, offering a HTTP-based interface. Run models on your hardware without having to write architecture-specific inference code. Here's a demo showing how to go from a model to GPU inference on a video…

    2023 · github.com

Ranked by how close each launch is in meaning, then by votes. Refine with a description →