Nexa SDK – Build powerful and efficient AI apps on edge devices
Hey HN! Alex and Zack here from Nexa AI. We're excited to share something we've been working on. Our journey began with the Octopus series --- action models for mobile AI agents (https://huggingface.co/NexaAIDev/Octopus-v2). We focused on making sub-billion parameter models excel at function calling, making high accurate and fast function-calling possible on mobile and edge devices. But as we delved into developing full-fledged on-device applications, we hit a roadblock. We realized that optimizing for function calling (tool-use) alone wasn't enough. Building powerful…
In plain words
Nexa SDK is a toolkit for building AI applications on mobile and edge devices. It combines language models, speech processing, image generation, and embedding models to enable developers to create full-featured on-device AI apps. The SDK builds on Nexa AI's work with small parameter models optimized for function calling on resource-constrained devices, providing the diverse tools needed beyond basic tool-use capabilities.
written from the facts on this page · September 2026
From the sources
In the maker’s words, at launch
Hey HN! Alex and Zack here from Nexa AI. We're excited to share something we've been working on. Our journey began with the Octopus series --- action models for mobile AI agents (https://huggingface.co/NexaAIDev/Octopus-v2). We focused on making sub-billion parameter models excel at function calling, making high accurate and fast function-calling possible on mobile and edge devices. But as we delved into developing full-fledged on-device applications, we hit a roadblock. We realized that optimizing for function calling (tool-use) alone wasn't enough. Building powerful on-device AI apps requires a diverse set of tools: language models with domain expertise, speech processing, image generation, embedding models and more. That's when we decided to create Nexa SDK --- a comprehensive toolkit that brings together everything developers need to build powerful and efficient AI applications that run entirely on-device. Here's what Nexa SDK offers: - Support for both ONNX and GGML models. - An integrated conversion engine for making custom GGML Quantized Models for different device hardware requirements. - An inference engine that supports language models, image generation models, TTS, audio generation models, and Vision-Language Models. - An OpenAI-compatible API server with optimization in function calling. - A Streamlit UI for rapid prototyping. - An intuitive CLI for easy model management. - Backend optimizations for latency and power consumption on edge devices. We've designed Nexa SDK to be the go-to solution for developers pushing the boundaries of what's possible with on-device AI applications and AI on edge devices. To showcase its capabilities, we've built several demo apps running entirely on your device (https://github.com/NexaAI/nexa-sdk/tree/main/examples): - AI soulmate with uncensored model and audio-in/audio-out interaction. - A quick interface for uploading and chatting with PDFs like your personal finance documents. - A meeting transcription app supporting multiple languages and real-time translation. We're proud to share that the winner of yesterday's (Sep 7) House AGI "AI PC/ GenAI Goes Local" hackathon used Nexa SDK to build a local semantic image search (https://github.com/asl3/deja-view). But we're just getting started! There are lots of exciting developments in our pipeline, and we can't wait to share them with you soon! Check it out: (https://github.com/NexaAI/nexa-sdk) Docs: (https://docs.nexaai.com/) If you're excited about the future of on-device AI, we'd really appreciate your support. A star on our GitHub repo goes a long way in helping us reach more developers! Cheers, Alex & Zack
More ai this month
the category →
I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.
AI · 17d ago · simedw.com
Astute▲585Automate your B2B brand going viral, with new media creators
AI · 18d ago · company-app.joinastute.com


Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2. The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges…
AI · 27d ago · cactuscompute.com


Launched alongside, September 2024
the whole month →

BeforeSunset AI 2.0▲1,267Personalized AI daily planning that suits your life
AI · 2024 · beforesunset.ai


