Alternatives
Products that do what Video2SRT does
Hello HN! I’m a bachelors student pursuing Artificial Intelligence, Robotics and Signal processing and during the year I had the goal to build my first AI Tool and launch it. I got inspired by how capable Whisper is, and combined the Whisper CPP Bindings along with FFMPEG.Wasm to create a tool that is capable of transcribing and captioning video files that contain one or more video tracks. Along with that, I’ve also added the support for transcribing audio files with the option to export those outputs as .SRT, .WebVTT or simply a text file. All done privately on a user’s web browser with…
- 1IB
Hey HN! I built Transcrib.ee to help me generate transcripts for lectures on YouTube with no captions, especially multilingual lectures. I was constantly frustrated with videos that had no transcripts or inaccurate ones, so I built a tool and decided to share it with everyone. This tool uses Groq (an AI inference engine that's different from Grok and offers very fast processing) and OpenAI's Whisper model (really accurate in multilingual) to quickly transcribe any YouTube video, regardless of the language. It's been a game-changer for my studies! To make it even faster, I created a Chrome…
2024 · transcrib.ee
- 2

- 3

- 4

- 5IV
Hi HN, I’m Ben, long time web developer but this is my first time building a macOS app. I film a lot of tutorial and talking-head content and wanted to make it easier to chop out mistakes. So I built CutWord, a macOS app that lets you mark cue points with voice commands while you’re recording. Once the app downloads a Whisper model everything runs locally, so nothing leaves your machine. You can export a trimmed video directly from the app or export a FCPXML file that can be imported into Final Cut or DaVinci Resolve for finishing. Built with SwiftUI and AVFoundation. TestFlight:…
2025 · cutword.com
- 6IP
Hi HN, I’m Ashu, founder of VideoDB. I’ve spent a big chunk of my life building video infrastructure. Not video creation. Video plumbing. The stuff you only learn after production breaks: timebases, VFR, keyframes, audio sync drift, container quirks, partial uploads, live streams, retries, backpressure, codecs, ffmpeg flags, cost blowups, and “why is this clip unseekable on one player but fine on another”. This week we shared VideoDB Skills, a skill pack that lets AI agents call those infra primitives directly, instead of you wiring pipelines with screenshots plus FFmpeg glue. Repo:…
Mar 2026
- 7TU
I’m Leif and I wanted to share the new product I’ve been working on recently, TurboScribe (https://turboscribe.ai). It’s pretty simple: unlimited Whisper transcription (starting at $10 per month). It supports large-v2, small, and base models. And yeah, it really is unlimited. The most active users have transcribed 1k+ hours using Whisper large-v2 (in other words, you can transcribe all 30 x 24 x 60 = 43200 minutes of your life every month if you want!). When I started building this a few months ago, I got connected with some initial users with higher volume transcription needs…
2023 · turboscribe.ai
- 801
Hey HN! I've been working on a side project to create an audio transcription API based on the OpenAI whisper model. Sign up link: https://whisperapi.com I tried to make the API really easy to use and get setup with. Also, because the Whisper model is so good, turns out I can offer the service for about 75% cheaper than what seems like the industry average. I'm always looking to make improvements, so would appreciate any feedback anyone has!
2022 · whisperapi.com
- 9TA
Hey HN! I built Topic2Manim to automate the creation of educational videos like those from 3Blue1Brown. The workflow is simple: 1. Give it any topic (e.g., "how ChatGPT works") 2. An LLM generates an educational script divided into scenes 3. LLM generates Manim code for each scene 4. FFmpeg concatenates everything into a final video Currently working on TTS integration for narration! Would love feedback on the approach and ideas for the TTS integration
Jan 2026 · github.com
- 10MA
Saw the remotion claude skills launch earlier, and honestly even though I was surprised how decent some of the results turned out to be I ended up never trying it out with claude code because I knew I'd have to setup remotion, bundler etc and if I was already doing it once I thought I might as well turn it into a site where anyone could just write messages and get a video without any prerequisites. I also know Claude Code is not something everyone has and setting up remotion is a pain. And one of the biggest lessons I learned from this whole experience is that Opus is actually not that good…
Feb 2026 · framecall.com
- 11WA
Hi HN, I wanted to share my first open-source project with you all: WhisperCat . WhisperCat is a small desktop application for recording audio and transcribing it using OpenAI's Whisper API. I built this because I needed something simple and reliable for my own transcription workflows, and now I'm hoping it might be useful to others as well. It's still pretty early stage, but it works well for basic audio recording and transcription tasks. What It Does: Lets you record audio with your preferred microphone. Transcribes audio files automatically via Whisper (OpenAI's transcription API).…
2025 · github.com
- 12TY
There have been many existing projects that transcribe YouTube videos with Whisper and its variants, but most of them aimed to generate subtitles, while I had not found one that priortises readability. Whisper does not generate line break in its transcription, so transcribing a 20 mins long video without any post processing would give you a huge piece of text, without any line break or topic segmentation. This project aims to transcribe videos with that post processing. Any feedback here, or on Github issues, would be very appreciated.
2024 · github.com
- 13AI
My focus has been shifting towards the ML alignment space recently, and in particular the ability to translate large transformer models into human understandable circuits and algorithms. This problem potentially isn't solvable, but it is one that some groups have had success with after large amounts of effort. In attempting to address this issue, I've been developing Transpector. A tool scaling up and reducing the barrier to entry of techniques that these teams have been showing success with. Techniques aiming to understand the internal mechanics of the model. Currently this tool is focused…
2023 · github.com
- 14OS
I built Whispering because I believe transcription is too fundamental a tool to be locked behind paywalls. It's a cross-platform desktop and web transcription app that turns speech into text with a keyboard shortcut, among other things. The app lets you bring your own API key (OpenAI, Groq, etc.) and make direct calls. If you want complete privacy, it also supports local transcription. Either way, your audio never goes through any middleman servers. It's super lightweight (~22MB), built with Svelte 5 and Tauri, and works on Mac, Windows, and Linux. I've been using it daily for the past few…
2025 · github.com
- 15AF
Hey HN — I’m Gaurav, one of the founders of Captions. We work on applied AI research for talking videos. Our foundation model, Lipdub, captures how humans speak, and matches full face movement to what’s being said. The model is zero-shot and can generate videos in under a minute, without person-specific training. Building on Lipdub, we’re releasing a few APIs that can generate and translate talking videos in bulk. Here are some ways they could be used: * Translating videos with matching lip movement * Creating personalized videos that include someone’s name or company, like what’s shown in…
2024 · captions.ai
- 16VA
Hello, Hacker News community! I’m Gary, the founder of Vizard.ai. I am thrilled to introduce you to our cutting-edge AI video editing tool that is designed to help content creators/marketers/businesses alike to repurpose long-form webinar/event/speech videos into bite-size social-ready highlights. As a beginner in video editing, I struggled with the complexity of professional video editing software like Adobe Premiere Pro or Apple Final Cut Pro. I found it overwhelming and challenging to use, and I would eventually abandon my efforts to edit videos. But I knew that video…
2023 · vizard.ai
- 17GW
2023 · bsri.blog
- 18IM
Live demo here: http://fonctionlabs.com:8000 Similarly to aka_sh (guess we were working parallelly on similar topics), I created with my brother a chainlit-based webapp, which summarizes Youtube videos in order to gain time. It works as an RAG-based LLM, and is very light in the sense that it does not use RAG libraries like langchain or llamaindex. You can use it with your own OpenAI API key. It also supports local models like Mistral, or Llamma. It is ofc open-source, and you can deploy with Docker if you choose. Some of the next steps are: - using whisper to be able to compute a…
2024 · github.com
- 19IV
I recently made this game Doge Decimator with the websim team, and wanted to make a process video for the steps I took to make it in an AI-native platform! You can check out every single iteration the site went through from the links in the description :)
2025 · youtube.com
- 20IM
Hello all, I made a small transcription app for your Mac based on OpenAI’s Whisper. Would love some feedback. My plan is to make it easy to load weights from any fine-tuned whisper model to enable specialized dictation for any subfield. It’s still early in development. Thanks!
2023 · twitter.com
- 21TA
I built TTSLab — a free, open-source tool for running text-to-speech and speech-to-text models directly in the browser using WebGPU and WASM. No API keys, no backend, no data leaves your machine. When you open the site, you'll hear it immediately — the landing page auto-generates speech from three different sentences right in your browser, no setup required. You can then try any model yourself: type text, hit generate, hear it instantly. Models download once and get cached locally. The most experimental feature: a fully in-browser Voice Agent. It chains speech-to-text → LLM → text-to-speech,…
Feb 2026 · ttslab.dev
- 22

- 23
Ranked by how close each launch is in meaning, then by votes. Refine with a description →