nowfound

AI · March 21, 2024

TA

Tidepool – analytics for large text datasets

Hello HN! I'm Peter, one of the folks who helped create Tidepool. We last shared Tidepool with HN about 7 months ago https://news.ycombinator.com/item?id=36957762 Since then, the AI field has moved incredibly quickly and we’ve iterated a lot on our product! The core problem we are trying to solve is: there's a lot of useful business insights you can get from text data, but it's hard to do analytics on it. - SQL is built for tabular / structured data, but when it comes to text, the best you can do is do keyword search. - In the pre-LLM world, you might resort to training a…

What it does

In the maker’s words, at launch

Hello HN! I'm Peter, one of the folks who helped create Tidepool. We last shared Tidepool with HN about 7 months ago https://news.ycombinator.com/item?id=36957762 Since then, the AI field has moved incredibly quickly and we’ve iterated a lot on our product! The core problem we are trying to solve is: there's a lot of useful business insights you can get from text data, but it's hard to do analytics on it. - SQL is built for tabular / structured data, but when it comes to text, the best you can do is do keyword search. - In the pre-LLM world, you might resort to training a lightweight text classifier, but you have to manually label a lot of data to get good accuracy. This means managing a team of operations people to do the labeling and building a lot of infrastructure to train and deploy the model. - Today it’s easier to get good results with simple LLM prompts, but the "large" in "large language models" means that running on big production-level datasets becomes prohibitively expensive. Our solution to this is to impose structure on this unstructured data with a combination of LLMs and lightweight embedding classifiers. - Using Tidepool, a user can query the data by creating an "attribute." An attribute is a characteristic of the data that you want to analyze, defined in natural language. This could be “sentiment of reviews,” “messages mentioning legal topics,” “prompts containing code snippets,” etc. - Tidepool structures the unstructured text by finding categories of interest for that attribute. For example, "positive vs negative vs neutral sentiment" or "C++ vs Python vs Javascript code snippets." We use an LLM to categorize a subset of the data for a user to review and refine the categorizations. - We then use the LLM categorized outputs to train a lightweight embedding classifier. This classifier then cheaply categorizes all existing and future data. - A user can either chart the categorized outputs in Tidepool or export them back to their data warehouse for further analysis in a business intelligence tool like Mode / Looker. Our first use case was for analyzing user prompts into LLM apps. Our customers have tens of millions of user submitted prompts, and they want to analyze usage patterns so they can improve their product. Tidepool helped them answer questions like: - What are the most common types of prompts for different user groups? - How common are different failure modes? - What type of actions correlate strongly with success metrics like engagement? Over time, we saw that our customers used Tidepool not just for analyzing user prompts for LLM use cases, but also were starting to look at general text content (think of user reviews, social media posts, documents, etc.). In principle, this makes a lot of sense - it's all just text at the end of the day! So we relaunched our site with some of these new capabilities in mind. Anyway, here's a short video demo of Tidepool: https://youtu.be/2yGTZBAH1T4 Happy to share more about what we've learned from talking to people in the AI space and building Tidepool for the last few months, feel free to comment your questions here!

Does the same job

all alternatives →
  • Tidepool by Aquarium2023 · ▲106

    Product analytics for AI applications

  • ShapedQLJan 2026 · ▲211

    The SQL engine for search, feeds, and AI agents

  • MovingLake AI Data Insights2023 · ▲96

    Ask questions about your data in plain english

  • WM
    We made Trellis – a way to run SQL query on your unstructured data2024 · demo.runtrellis.com · ▲8

    Hey HN — We're excited to share Trellis — a snowflake for unstructured data. We've built an AI engine that turns unstructured data into structured SQL-format based on the schema you define in natural language. We spent a lot of time building ML infrastructure and realized that most data warehouses and data pipelines are not designed for unstructured data (documents, PDFs, calls). While something like a Vector database and RAG are great at search tasks, they really struggle with aggregation and SQL type queries such as 1. How many emails in the past 6 months contain complaints about the…

  • DS
    Describe SQL using natural language, and execute against real data2021 · ▲66

    I played around with GPT-3 to build this demo. Select a public BigQuery dataset and describe your query in natural English, then edit the generated SQL as needed and execute it. https://app.tabbydata.com/sql-assistant-demo

  • WA
    We are building Git for dataJan 2026 · ▲9

    Today you can easily adopt AI coding tools because you have git for branching and rolling back if AI writes bad code. We haven't seen this same capability for data and decided to build it ourselves. Nile is a new kind of data lake, purpose built for using with AI. It can act as your data engineer or data analyst creating new tables and rolling back bad changes in seconds. We support real versions for data, schema, and ETL. We'd love your feedback on any part of what we are building - https://getnile.ai/ What do you think?

More ai this month

the category →
  • I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.

    AI · 17d ago · simedw.com

  • Astute585

    Automate your B2B brand going viral, with new media creators

    AI · 18d ago · company-app.joinastute.com

  • Grok Bot547

    AI teammates that you can give real work to

    AI · 25d ago · x.ai

  • Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2. The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges…

    AI · 26d ago · cactuscompute.com

  • Make your software self-driving

    AI · 30d ago · coldtea.ai

  • Soloop472

    Approval-first Agent OS for solo founders

    AI · 30d ago · soloop.io

Launched alongside, March 2024

the whole month →
  • Dub.co1,475

    Short links with superpowers

    Dev tools · 2024 · dub.co

  • 3Y
  • Microlaunch1,116

    Launch and get feedback on both the idea and product

    Dev tools · 2024 · microlaunch.net

  • Milestone1,113

    Interactive, gamified product tours for SaaS

    Growth · 2024 · milestoneflow.io

  • Hunted Space1,004

    Insights on Product Hunt launches

    Growth · 2024 · hunted.space

  • Creatie850

    The one-stop product design tool amplified by AI

    AI · 2024 · creatie.ai