nowfound

AI · May 24, 2023

CA

Cellulose – a tool to improve inference performance of ML models

Hey HN! It’s Zheng here. I’m the founder of Cellulose (https://www.cellulose.ai). Cellulose is a tool that helps ML engineers understand, fine tune, and improve the inference performance of their ONNX models. With Cellulose, they can eventually resolve these issues in just hours, not weeks. Preparing ML models for production is a very manual and time consuming process. Unfortunately, it is also a necessary step for ML inference cost savings, sometimes even a hard requirement for certain applications like robotics and space tech. Today’s ML visualization tools are over 6 years old…

In plain words

Cellulose is a tool designed for ML engineers to analyze, fine-tune, and optimize the inference performance of ONNX models. It addresses the manual, time-consuming work of preparing machine learning models for production deployment, helping users resolve performance issues in hours rather than weeks. The platform is particularly valuable for applications with strict computational constraints, such as robotics and space technology, where inference efficiency is critical or mandatory.

written from the facts on this page · September 2026

From the sources

In the maker’s words, at launch

Hey HN! It’s Zheng here. I’m the founder of Cellulose (https://www.cellulose.ai). Cellulose is a tool that helps ML engineers understand, fine tune, and improve the inference performance of their ONNX models. With Cellulose, they can eventually resolve these issues in just hours, not weeks. Preparing ML models for production is a very manual and time consuming process. Unfortunately, it is also a necessary step for ML inference cost savings, sometimes even a hard requirement for certain applications like robotics and space tech. Today’s ML visualization tools are over 6 years old and lack basic features like integrating modern deep learning workflows. You’d be downloading model files locally then using a visualization tool to scroll and search for specific nodes and tensor dimensions. For example, you’ll do this twice if you’re comparing two model versions. ML researchers typically iterate on the model and then get to a “frozen”, gold release candidate before kicking off deployment related workflows. Say you use specialized hardware to run your models because that’s the most performant and cost efficient way to serve them. Unfortunately, some operators in the model could be incompatible with hardware backends like TensorRT. While there’s no shortcut but additional engineering effort to figure out a workaround or proper solution, such a setback late in the model development lifecycle is expensive for a ML team. I’ve experienced this at Cruise (https://getcruise.com) myself as an engineer in the Machine Learning Accelerators (MLA) team. Deploying big, bulky models onto hardware constrained environments like an AV with strict system performance limits remain a significant challenge. Friends working at various AI and robotics teams have expressed similar frustrations. Cellulose enables you to optimize and fine tune your models in a more automated fashion throughout your ML development lifecycle. We went with a product that leads with a visualizer core as so much of a ML model today is centered around the graph itself. Here’s a screenshot of a ResNet-50 model in the Cellulose dashboard: https://drive.google.com/file/d/1aZ3_fcmVVqPxxiNNcm8bkQKYsqj... Cellulose has utilities to help you copy specific values to the clipboard, just in case you’d like to run offline experimental scripts. Here’s a BatchNormalization op drawer with all its properties: https://drive.google.com/file/d/19XMY_HOwqg8ysbW4d4rqHXX5hoD... Initializer values for resnetv24_stage3_batchnorm3_gamma: https://drive.google.com/file/d/1NOwiCZbz8A2UTqDzDSnQ9WVzKVV... Export model graph as .png: https://drive.google.com/file/d/1IIOY65ZlFtc701eeMhosncSxeHd... We’re supporting Nvidia TensorRT as our first runtime. Under our Professional / Enterprise plans, we’ll annotate the TensorRT compatibility / convertibility of each node in the graph.[1] Selecting runtime type and precision options: https://drive.google.com/file/d/1Z_r68MA1HK-KVlOLA2YoPUeR0vm... TensorRT v8.6.1 compatibility badge annotations (on each op): https://drive.google.com/file/d/1L-QeZtw9gDtibJOgEdWsOm1hDNk... Supported Runtimes tab for the Reshape op: https://drive.google.com/file/d/1IS7Jio19d3WKWHh7JfrsLdzLJ7Z... We also have an exciting roadmap (https://docs.cellulose.ai/roadmap/overview), but more importantly, we’d like you to try it out (it’s free to start!), hear your thoughts / feedback then we’ll make sure to make those tweaks as soon as humanly possible. Feel free to sign up at http://dashboard.cellulose.ai or browse our documentation at https://docs.cellulose.ai I’ll have this tab open all day today to answer any questions! [1] - We use onnx-tensorrt for the TensorRT compatibility checks.

More ai this month

the category →
  • I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.

    AI · 17d ago · simedw.com

  • Astute585

    Automate your B2B brand going viral, with new media creators

    AI · 18d ago · company-app.joinastute.com

  • Grok Bot547

    AI teammates that you can give real work to

    AI · 25d ago · x.ai

  • Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2. The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges…

    AI · 27d ago · cactuscompute.com

  • Turn website visitors into qualified pipeline

    AI · 19d ago · clarasdr.ai

  • Kane CLI446

    Natural language browser & mobile app tests from terminal

    AI · 24d ago · testmuai.com

Launched alongside, May 2023

the whole month →
  • Sidekick1,343

    An AI-powered accessibility assistant in Stark

    AI · 2023 · getstark.co

  • BR

    In today's world, catchy headlines and articles often distract readers from the facts and relevant information. By utilizing OpenAI's language models, Boring Report processes sensationalist news articles, transforms them into the content you see, and helps readers focus on the essential details. We recently updated our iOS app experience, so any and all feedback would be appreciated. App Link: https://apps.apple.com/us/app/boring-report-news-by-ai/id644...

    AI · 2023 · boringreport.org

  • Generating powerful websites, one prompt at a time

    AI · 2023

  • Unlock a new level productivity with AI, Cloud Sync and more

    AI · 2023 · raycast.com

  • AudioPen888

    The easiest way to convert messy thoughts into clear text

    AI · 2023 · audiopen.ai

  • Launch your website in seconds, get users in minutes

    Dev tools · 2023 · page.mmntm.build