nowfound

AI · October 3, 2024

RA

Reverb ASR+Diarization, the Best Open Source ASR for Long-Form Audio

Hey everyone, My name is Lee Harris and I'm the VP of Engineering for Rev.com / Rev.ai. Today, we are launching and open sourcing our current generation ASR models named "Reverb." When OpenAI launched Whisper at Interspeech two years ago, it turned the ASR world upside down. Today, Rev is building on that foundation with Reverb, the world's #1 ASR model for long-form transcription – now open-source. I am proud to announce that we are releasing two models today, Reverb and Reverb Turbo, through our API, self-hosted, and our open source + open weights solution. ---------- We are releasing…

In plain words

Reverb ASR+Diarization is an open-source speech recognition system from Rev.ai designed for transcribing long-form audio. The system offers two model variants—Reverb and Reverb Turbo—available through API, self-hosted deployment, and open-source releases. It supports speaker diarization to identify different speakers in recordings. The system is built on recent advances in automatic speech recognition technology and is available in research-oriented and developer-focused formats for different implementation needs.

written from the facts on this page · September 2026

From the sources

In the maker’s words, at launch

Hey everyone, My name is Lee Harris and I'm the VP of Engineering for Rev.com / Rev.ai. Today, we are launching and open sourcing our current generation ASR models named "Reverb." When OpenAI launched Whisper at Interspeech two years ago, it turned the ASR world upside down. Today, Rev is building on that foundation with Reverb, the world's #1 ASR model for long-form transcription – now open-source. I am proud to announce that we are releasing two models today, Reverb and Reverb Turbo, through our API, self-hosted, and our open source + open weights solution. ---------- We are releasing in the following formats: - A research-oriented release that doesn't include our end to end pipeline and is missing our WFST (Weighted Finite-State Transducer) implementation. This is primarily in Python and intended for research, exploratory, or custom usage within your ecosystem. - A developer-oriented release that includes our entire end-to-end pipeline for environments at any scale. It is a combination of C# for the APIs, C++ for our inference engine, and Python for various pieces. - A new set of end-to-end APIs that are priced at $0.20/hour for Reverb and $0.10/hour for Reverb Turbo. ---------- What makes Reverb special? - Reverb was trained on 200,000+ hours of extremely high quality and varied transcribed audio from Rev.com expert transcribers. This high quality data set was chosen as a subset from 7+ million hours of Rev audio. - The model runs extremely well on CPU, GPU, iOS/Android, IoT, and other platforms. Our developer implementation is primarily optimized for x64 CPU today, but a GPU optimized version will be released this year. - The model excels in noisy, real-world environments. Real data was used during the training and every audio was handled by an expert transcriptionist. - You can tune your results for vertabimicity, allowing you to have nicely formatted, opinionated outputs OR true verbatim output. This is the #1 area where Reverb substantially outperforms the competition. ---------- Benchmarks Here are some WER (word error rate) benchmarks on Rev's various solutions for Earnings21 and Earnings22 (very challenging audio): - Reverb - Earnings21: 7.99 WER - Earnings22: 7.06 WER - Reverb Turbo - Earnings21: 8.25 WER - Earnings22: 7.50 WER - Reverb Research - Earnings21: 10.30 WER - Earnings22: 9.08 WER - Whisper large-v3 - Earnings21: 10.67 WER - Earnings22: 11.37 WER - Canary-1B - Earnings21: 13.82 WER - Earnings22: 13.24 WER ---------- Licensing Our models are released under a non-commercial / research license that allow for personal, research, and evaluation use. If you wish to use it for commercial purposes, you have 3 options: - Usage based API @ $0.20/hr for Reverb, $0.10/hr for Reverb Turbo. - Usage based self-hosted container at the same price as our API. - Unlimited use license at custom pricing. Contact us at [email protected]. ---------- Final Thoughts I highly recommend that anyone interested take a look at our fantastic technical blog written by one of our Staff Speech Scientists, Jenny Drexler Fox. We look forward to hearing community feedback and we look forward to sharing even more of our models and research in the near future. Thank you! ---------- Links Technical blog: https://www.rev.com/blog/speech-to-text-technology/introduci... Launch blog / news post: https://www.rev.com/blog/speech-to-text-technology/open-sour... GitHub research release: https://github.com/revdotcom/reverb GitHub self-hosted release: https://github.com/revdotcom/reverb-self-hosted Huggingface ASR link: https://huggingface.co/Revai/reverb-asr Huggingface Diarization V1 link: https://huggingface.co/Revai/reverb-diarization-v1 HuggingFace Diarization V2 link: https://huggingface.co/Revai/reverb-diarization-v2

More ai this month

the category →
  • I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.

    AI · 17d ago · simedw.com

  • Astute585

    Automate your B2B brand going viral, with new media creators

    AI · 18d ago · company-app.joinastute.com

  • Grok Bot547

    AI teammates that you can give real work to

    AI · 25d ago · x.ai

  • Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2. The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges…

    AI · 27d ago · cactuscompute.com

  • Make your software self-driving

    AI · 30d ago · coldtea.ai

  • Soloop472

    Approval-first Agent OS for solo founders

    AI · 30d ago · soloop.io

Launched alongside, October 2024

the whole month →
  • buzzabout1,267

    Audience insights from 1B+ online discussions in 2 mins

    AI · 2024 · buzzabout.ai

  • bolt.new1,233

    Prompt, run, edit & deploy full-stack web apps

    AI · 2024 · bolt.new

  • Feta1,217

    Run smarter stand-ups, build better products

    AI · 2024 · feta.io

  • Trag955

    AI code review companion

    AI · 2024 · usetrag.com

  • Turn Notion databases into portals & apps with no code

    Dev tools · 2024 · softr.io

  • One inbox for all your work discussions

    Work · 2024 · generalcollaboration.com