nowfound

Dev tools · May 21, 2025

BO

BemiDB – Open-source data warehouse with zero-ETL

Hi HN! We're Evgeny and Arjun, and we’re building a simpler way for startups to do data analytics. Since open-sourcing our Postgres read replica optimized for analytics (https://github.com/BemiHQ/BemiDB), we started hearing a familiar story. Teams would connect Postgres, feel relieved that they didn’t have to wrangle complex ETL pipelines, but then hit a wall as soon as they wanted to join data from HubSpot, Stripe, etc. They’d do hacky things like use Airbyte to sync data to their Postgres, so that it’d then auto sync to their BemiDB analytical database. We want to…

In plain words

BemiDB is an open-source data warehouse built as a Postgres read replica optimized for analytics. It eliminates the need for complex ETL pipelines by allowing startups to connect their Postgres database alongside supplementary data sources like HubSpot and Stripe. Designed for companies seeking lightweight alternatives to heavyweight data warehouses, BemiDB simplifies data analytics infrastructure without expensive ETL processes.

written from the facts on this page · September 2026

From the sources

In the maker’s words, at launch

Hi HN! We're Evgeny and Arjun, and we’re building a simpler way for startups to do data analytics. Since open-sourcing our Postgres read replica optimized for analytics (https://github.com/BemiHQ/BemiDB), we started hearing a familiar story. Teams would connect Postgres, feel relieved that they didn’t have to wrangle complex ETL pipelines, but then hit a wall as soon as they wanted to join data from HubSpot, Stripe, etc. They’d do hacky things like use Airbyte to sync data to their Postgres, so that it’d then auto sync to their BemiDB analytical database. We want to remove the layers of data complexity that startups have to add when scaling, and that’s why BemiDB now also allows connecting any supplementary data sources. This makes it a zero-ETL data warehouse for companies that don’t want the typical heavyweight warehouses with expensive ETL’s. Under the hood, we use Apache Iceberg (with Parquet data files) stored in S3. This allows for bottomless inexpensive storage, compressed data in columnar files, and an open format that guarantees compatibility with other data tools. We use Trino to help with table maintenance and compaction. We embed DuckDB as the query engine for in-memory analytics that work for complex queries. With efficient columnar storage and vectorized execution, we’re aiming for faster results without heavy infra. BemiDB communicates over the Postgres wire protocol and is also Postgres syntax compatible. We want to fully simplify data infra for companies that use Postgres and other data sources by reducing complexity (automatic data source syncs), using non-proprietary data formats (Iceberg open tables), and removing vendor lock-in (open source). We'd love to hear any and all feedback! Thoughts HN?

More dev tools this month

the category →
  • Dograh592

    The open source VAPI alternative

    Dev tools · 25d ago · dograh.com

  • Meridian530

    Don't let your work go unnoticed. Get promoted!

    Dev tools · 20d ago · meridiona.com

  • x1516

    Lovable for iPhone apps go from idea to App Store

    Dev tools · 11d ago · x1.new

  • Open-source GTM skills for technical founders

    Dev tools · 29d ago · gtmcofounder.com

  • OpenTrailPaper is open-source bike computer firmware for the LilyGO T5S3 4.7" E-Paper PRO. It supports offline maps, GPX routes, FIT recording and Bluetooth sensors.

    Dev tools · 1d ago · opentrailpaper.com

  • Nuphos380

    The AI-Native DevOps Workspace.

    Dev tools · 24d ago · nuphos.ai

Launched alongside, May 2025

the whole month →