nowfound

Alternatives

Products that do what 0-to-1 Data Engineer Interview Playbook does

Master Data Engineer interviews: SQL, pipelines, streaming

  1. 1IB

    Every data pipeline job I had to tackle required quite a few components to set up: - One tool to ingest data - Another one to transform it - If you wanted to run Python, set up an orchestrator - If you need to check the data, a data quality tool Let alone this being hard to set up and taking time, it is also pretty high-maintenance. I had to do a lot of infra work, and while this being billable hours for me I didn’t enjoy the work at all. For some parts of it, there were nice solutions like dbt, but in the end for an end-to-end workflow, it didn’t work. That’s why I decided to build an…

    2024 · github.com

  2. 2
    nao502

    AI data IDE for 10x faster data work

    Nov 2025

  3. 3AB

    I created a web page to compare different analytical databases (both self-managed and services, open-source and proprietary) on a realistic dataset. It contains 20+ databases, each with installation and data loading scripts. And they can be compared to each other on a set of 43 queries, by data load time or by storage size. There are switches to select different types of databases for comparison - for example, only MySQL compatible or PostgreSQL compatible. If you play with the switches, many interesting details will be uncovered. Full description:…

    2022 · benchmark.clickhouse.com

  4. 4DE

    Hi HN! I'm currently a Master's student at USTC (University of Science and Technology of China). I've been diving deep into Data Engineering, especially in the context of Large Language Models (LLMs). The Problem: I found that learning resources for modern data engineering are often fragmented and scattered across hundreds of medium articles or disjointed tutorials. It's hard to piece everything together into a coherent system. The Solution: I decided to open-source my learning notes and build them into a structured book. My goal is to help developers fast-track their learning curve. Key…

    Feb 2026 · github.com

  5. 5PN

    We've been hard at work for a few weeks and thought it's time for another update. In case you missed our first post, PostgresML is an end-to-end machine learning solution, running alongside your favorite database. This time we have more of a suite offering: project management, visibility into the datasets and the deployment pipeline decision making. Let us know what you think! Demo link is on the page, and also here: https://demo.postgresml.org

    2022 · postgresml.org

  6. 6

    Know compute cost of every pipeline & model in your BigQuery

    2023

  7. 7

    A single DataOps platform for data engineering

    2021

  8. 8

    The fastest way to build your data warehouse

    2023

  9. 9
    ShapedQL211

    The SQL engine for search, feeds, and AI agents

    Jan 2026

  10. 10DP
  11. 11IM

    2017 · xyz.insightdataengineering.com

  12. 12DF

    Hello Everyone! We built SQLFlow as a lightweight stream processing engine. We leverage DuckDB as the stream processing engine, which gives SQLFlow the ability to process 10's of thousands of messages a second using ~250MiB of memory! DuckDB also supports a rich ecosystem of sinks and connectors! https://sql-flow.com/docs/category/tutorials/ https://github.com/turbolytics/sql-flow We were tired of running JVM's for simple stream processing, and also of bespoke one off stream processors I would love your feedback, criticisms and/or…

    Dec 2025 · sql-flow.com

  13. 13
    Velvet166

    Make everyone a data engineer

    2024

  14. 14ST

    Hey Show HN! I’m Toby and over the last few months, I’ve been working with a team of engineers from Airbnb, Apple, Google, and Netflix, to simplify developing data pipelines with SQLMesh (https://github.com/TobikoData/sqlmesh). We’re tired of fragile pipelines, untested SQL queries, and expensive staging environments for data. Software engineers have reaped the benefits of DevOps through unit tests, continuous integration, and continuous deployment for years. We felt like it was time for data teams to have the same confidence and efficiency in development as their peers.…

    2023 · github.com

  15. 15DA

    Hi HN community. We are excited to open source Dataherald’s natural-language-to-SQL engine today (https://github.com/Dataherald/dataherald). This engine allows you to set up an API from your structured database that can answer questions in plain English. GPT-4 class LLMs have gotten remarkably good at writing SQL. However, out-of-the-box LLMs and existing frameworks would not work with our own structured data at a necessary quality level. For example, given the question “what was the average rent in Los Angeles in May 2023?” a reasonable human would either assume the…

    2023 · github.com

  16. 16

    Realtime analytics database

    2014

  17. 17

    A grid library for instant big data processing

    2022

  18. 18
    Seeknal57

    Data & AI/ML CLI for pipelines and NL queries

    Apr 2026 · seeknal.exe.xyz

  19. 19
    pipe46

    Pipe coldpress datasets straight into your pipeline

    2024

  20. 20PB
  21. 21PP

    Hey! I'm Andrei. I got frustrated by how people tend to build overcomplicated backend systems, being "motivated" by big tech case studies and popular books. So, I started exploring lean architecture, and building my digital garden of ideas, approaches and data that align with this direction. Here I want to present one of the tools – Sizing tool for PostgreSQL. I've benchmarked PostgreSQL on different EC2 instances and disks, with different initial data sets to see performance that these instances can give you. And I've built a tool to visualize this data, which I welcome you to explore. So,…

    Jul 2026 · postgres.saneengineer.com

  22. 22TZ

    2017 · gpestana.gitbooks.io

  23. 23
    Myriade101

    Ask your data. See the SQL. Self-host in one command.

    2025

  24. 24PA

Ranked by how close each launch is in meaning, then by votes. Refine with a description →