nowfound

AI · September 5, 2024

WB

We built a knowledge hub for running LLMs on edge devices

Hey HN! Alex and Zack from Nexa AI here. We are excited to share a project our team has been passionately working on recently, in collaboration with Jiajun from Meta, Qun from San Francisco State University, and Xin and Qi from the University of North Texas. Running AI models on edge devices is becoming increasingly important. It's cost-effective, ensures privacy, offers low-latency responses, and allows for customization. Plus, it's always available, even offline. What's really exciting is that smaller-scale models are now approaching the performance of large-scale closed-source models for…

In plain words

Nexa AI provides a knowledge hub for deploying large language models on edge devices such as smartphones, IoT gadgets, and Raspberry Pi systems. The platform enables cost-effective, private, and low-latency AI inference while maintaining offline availability. It targets developers and organizations seeking to run smaller-scale models that approach the performance of larger closed-source alternatives for applications like writing assistance and email classification.

written from the facts on this page · September 2026

From the sources

In the maker’s words, at launch

Hey HN! Alex and Zack from Nexa AI here. We are excited to share a project our team has been passionately working on recently, in collaboration with Jiajun from Meta, Qun from San Francisco State University, and Xin and Qi from the University of North Texas. Running AI models on edge devices is becoming increasingly important. It's cost-effective, ensures privacy, offers low-latency responses, and allows for customization. Plus, it's always available, even offline. What's really exciting is that smaller-scale models are now approaching the performance of large-scale closed-source models for many use cases like "writing assistant" and "email classifier". We've been immersing ourselves in this rapidly evolving field of on-device AI - from smartphones to IoT gadgets and even that Raspberry Pi you might have lying around. It's a fascinating field that's moving incredibly fast, and honestly, it's been a challenge just keeping up with all the developments. To help us make sense of it all, we started compiling our notes, findings, and resources into a single place. That turned into this GitHub repo: https://github.com/NexaAI/Awesome-LLMs-on-device Here's what you'll find inside: A timeline tracking the evolution of on-device AI models Our analysis of efficient architectures and optimization techniques (there are some seriously clever tricks out there) A curated list of cutting-edge models and frameworks we've come across Real-world examples and case studies that got us excited about the potential of this tech We're constantly updating it as we learn more. It's become an invaluable resource for our own work, and we hope it can be useful for others too - whether you're deep in the trenches of AI research or just curious about where edge computing is heading. We'd love to hear what you think. If you spot anything we've missed, have some insights to add, or just want to geek out about on-device AI, please don't hesitate to contribute or reach out. We're all learning together here! This is a topic we are genuinely passionate about, and we are looking forward to some great discussions. Thanks for checking it out!

More ai this month

the category →
  • I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.

    AI · 17d ago · simedw.com

  • Astute585

    Automate your B2B brand going viral, with new media creators

    AI · 18d ago · company-app.joinastute.com

  • Grok Bot547

    AI teammates that you can give real work to

    AI · 25d ago · x.ai

  • Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2. The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges…

    AI · 26d ago · cactuscompute.com

  • Make your software self-driving

    AI · 30d ago · coldtea.ai

  • Soloop472

    Approval-first Agent OS for solo founders

    AI · 30d ago · soloop.io

Launched alongside, September 2024

the whole month →
  • Wispr Flow2,737

    Speak naturally, write perfectly & 3x faster in every app

    AI · 2024 · wisprflow.ai

  • Pathway1,335

    Get user insights 10x faster

    Work · 2024 · wynde.io

  • Personalized AI daily planning that suits your life

    AI · 2024 · beforesunset.ai

  • Osmos1,194

    Match with like-minded professionals for 1:1 conversations

    Growth · 2024

  • Polar1,169

    An open source monetization platform for developers

    Dev tools · 2024 · polar.sh

  • Carrot Care1,167

    Understand & optimise your bloodwork

    Life & fun · 2024 · carrotcare.health