VoxConvo – "X but it's only voice messages"
Hi HN, I saw this tweet: "Hear me out: X but it's only voice messages (with AI transcriptions)" - and couldn't stop thinking about it. So I built VoxConvo. Why this exists: AI-generated content is drowning social media. ChatGPT replies, bot threads, AI slop everywhere. When you hear someone's actual voice: their tone, hesitation, excitement - you know it's real. That authenticity is what we're losing. So I built a simple platform where voice is the ONLY option. The experience: Every post is voice + transcript with word-level timestamps: Read mode: Scan the transcript like normal text or…
In plain words
VoxConvo is a social platform where all posts are voice messages paired with AI transcriptions and word-level timestamps. Users can read transcripts like traditional text or listen with highlighted words that sync in real-time. The platform includes visual voice editing, allowing users to remove filler words or mistakes by clicking specific words in the transcript to delete corresponding audio segments. Designed for people seeking authentic communication, VoxConvo addresses AI-generated content saturation by requiring voice as the only posting format, preserving tone and emotion while maintaining readability.
written from the facts on this page · September 2026
From the sources
In the maker’s words, at launch
Hi HN, I saw this tweet: "Hear me out: X but it's only voice messages (with AI transcriptions)" - and couldn't stop thinking about it. So I built VoxConvo. Why this exists: AI-generated content is drowning social media. ChatGPT replies, bot threads, AI slop everywhere. When you hear someone's actual voice: their tone, hesitation, excitement - you know it's real. That authenticity is what we're losing. So I built a simple platform where voice is the ONLY option. The experience: Every post is voice + transcript with word-level timestamps: Read mode: Scan the transcript like normal text or listen mode: hit play and words highlight in real-time. You get the emotion of voice with the scannability of text. Key features: - Voice shorts - Real-time transcription - Visual voice editing - click a word in transcript deletes that audio segment to remove filler words, mistakes, pauses - Word-level timestamp sync - No LLM content generation Technical details: Backend running on Mac Mini M1: - TypeGraphQL + Apollo Server - MongoDB + Atlas Search (community mongo + mongot) - Redis pub/sub for GraphQL subscriptions - Docker containerization for ready to scale Transcription: - VOSK real time gigaspeech model eats about 7GB RAM - WebSocket streaming for real-time partial results - Word-level timestamp extraction plus punctuation model Storage: - Audio files are stored to AWS S3 - Everything else is local Why Mac Mini for MVP? Validation first, scaling later. Architecture is containerized and ready to migrate. But I'd rather prove demand on gigabit fiber than burn cloud budget.
More ai this month
the category →
I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.
AI · 17d ago · simedw.com
Astute▲585Automate your B2B brand going viral, with new media creators
AI · 18d ago · company-app.joinastute.com


Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2. The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges…
AI · 27d ago · cactuscompute.com


Launched alongside, November 2025
the whole month →
- IB
Life & fun · Nov 2025 · bitsnpieces.dev



- BBoing▲782
Life & fun · Nov 2025 · boing.greg.technology