Alternatives
Products that do what Inconveniently operating my computer with voice and hand gestures does
Introducing Iron OS: it's like a regular computer, but much more inconvenient Created with threejs, rosebud AI, web speech API, and mediapipe computer vision Any feedback would be appreciated! I've been having fun experimenting with computer vision and voice control lately.
- 1OS
Heeey! I built a macOS copilot that has been useful to me, so I open sourced it in case others would find it useful too. It's pretty simple: - Use a keyboard shortcut to take a screenshot of your active macOS window and start recording the microphone. - Speak your question, then press the keyboard shortcut again to send your question + screenshot off to OpenAI Vision - The Vision response is presented in-context/overlayed over the active window, and spoken to you as audio. - The app keeps running in the background, only taking a screenshot/listening when activated by keyboard…
2023 · github.com
- 2

- 3

- 4

- 5C3
I'm sharing my project to control 3D models with voice commands and hand gestures: - use voice commands to change interaction mode (drag, rotate, scale, animate) - use hand gestures to control the 3D model - drag/drop to import other models (only GLTF format supported for now) Created using threejs, mediapipe, web speech API, rosebud AI, and Quaternius 3D models Githhub repo: https://github.com/collidingScopes/3d-model-playground Demo: https://xcancel.com/measure_plan/status/1929900748235550912 I'd love to get your feedback! Thank you
2025 · github.com
- 6
- 7AO
I've been obsessed for the past ~year with the possibilities of talking to LLMs. I built a bunch of one-off prototypes, shared code on X, started a Meetup group in SF, and co-hosted a big hackathon. It turns out that there are a few low-level problems that everybody building conversational/real-time AI needs to solve on the way to building/shipping something that works well: low-latency media transport, echo cancellation, voice activity detection, phrase endpointing, pipelining data between models/services, handling voice interruptions, swapping out different…
2024 · github.com
- 8

- 9VA
Voxos is an open-source desktop voice assistant that aims to put Clippy to shame while supporting new desktop workflows powered by LLMs. Tired of copy and pasting ChatGPT responses between your web browser and IDE? Does your copilot not quite do what you need it to do? I invite you to give Voxos a try and maybe even become a contributor!
2024 · gitlab.com
- 10

- 11

- 12IO
Hi HN! Last year the project I launched here got a lot of good feedback on creating speech to speech AI on the ESP32. Recently I revamped the whole stack, iterated on that feedback and made our project fully open-source—all of the client, hardware, firmware code. This Github repo turns an ESP32-S3 into a realtime AI speech companion using the OpenAI Realtime API, Arduino WebSockets, Deno Edge Functions, and a full-stack web interface. You can talk to your own custom AI character, and it responds instantly. I couldn't find a resource that helped set up a reliable, secure websocket (WSS) AI…
2025 · github.com
- 13

- 14

- 15

- 16

- 17

- 18

- 19

- 20
- 21AF
2024 · swift-ai.vercel.app
- 22

- 23OS
2021 · github.com
- 24AS
2024 · github.com
Ranked by how close each launch is in meaning, then by votes. Refine with a description →