Alternatives
Products that do what Hotshot – 4 Person Team Builds a State of the Art Video Model does
Hi HN! We're proud to share Hotshot, a large-scale diffusion transformer model for text-to-video generation that we built with just a 4-person team. You can try it today in beta at https://hotshot.co, with 2 free generations per day. The model generates 5 seconds of 720p video from text prompts. It excels at prompt alignment, and consistency. It also excels at generating people, animals, and nature. In blind tests with 100 users, Hotshot generations were preferred to Runway ML 60% of the time. Hotshot generations were preferred to Luma 80% of the time. Overall, users preferred…
- 1

- 2

- 3

- 4

- 5BA
Hello HN! I want to share something me and a few friends have been working on for a while now — Zeroshot, a web tool that builds image classifiers using text-image models and autolabeling. What does this mean in practice? You can put together an image classifier in about 30 seconds that’s faster and more accurate than CLIP, but that you can deploy yourself however you’d like. It’s open source, commercially licensed, and doesn’t require you to pay anyone per API call. Here's a 2 minute video that shows it off: https://www.youtube.com/watch?v=S4R1gtmM-Lo How/why does it…
2023 · usezeroshot.com
- 6LS
Hey HN, this is Lina, Andrew, and Sidney from Lemon Slice. We’ve trained a custom diffusion transformer (DiT) model that achieves video streaming at 25fps and wrapped it into a demo that allows anyone to turn a photo into a real-time, talking avatar. Here’s an example conversation from co-founder Andrew: https://www.youtube.com/watch?v=CeYp5xQMFZY. Try it for yourself at: https://lemonslice.com/live. (Btw, we used to be called Infinity AI and did a Show HN under that name last year: https://news.ycombinator.com/item?id=41467704.) Unlike existing…
2025
- 7TT
Writeup (includes good/bad sample generations): https://www.linum.ai/field-notes/launch-linum-v2 We're Sahil and Manu, two brothers who spent the last 2 years training text-to-video models from scratch. Today we're releasing them under Apache 2.0. These are 2B param models capable of generating 2-5 seconds of footage at either 360p or 720p. In terms of model size, the closest comparison is Alibaba's Wan 2.1 1.3B. From our testing, we get significantly better motion capture and aesthetics. We're not claiming to have reached the frontier. For us, this is a stepping…
Jan 2026 · huggingface.co
- 8

- 9CY
A few months ago, I started making video clips with stable diffusion and noticed that the tools to do this were too complicated for everyday people. That's why I built neural frames. Enjoy.
2023 · neuralframes.com
- 10DA
Hello HN! I would like to show our recent research project: DynamiCrafter. DynamiCrafter can animate open-domain still images based on text prompt by leveraging the pre-trained video diffusion priors. It supports: - Image-to-video generation - Storytelling video generation - Looping video generation - Generative frame interpolation We have released source code at https://github.com/Doubiiu/DynamiCrafter and deployed a demo at https://huggingface.co/spaces/Doubiiu/DynamiCrafter Please also refer to…
2023 · github.com
- 11IA
Hey everyone! Excited to be able to share the release of `InvokeAI 2.0 - A Stable Diffusion Toolkit`, an open source project that aims to provide both enthusiasts and professionals a suite of robust image creation tools. Optimized for efficiency, InvokeAI needs only ~3.5GB of VRAM to generate a 512x768 image (and less for smaller images), and is compatible with Windows/Linux/Mac (M1 & M2). InvokeAI was one of the earliest forks off of the core CompVis repo (formerly lstein/stable-diffusion), and recently evolved into a full-fledged community driven and open source stable…
2022 · github.com
- 12

- 13

- 14

- 15

- 16I4
It's our new text-to-image model: a 9.3B single-stream diffusion transformer trained entirely from scratch. We focused heavily on controllability through structured JSON prompts, with strong text rendering, spatial awareness through bounding box guidance, and color palette control. It has the best text rendering of any open-weight model we've tested so far, and the NF4 quantized checkpoint runs on a single 24GB GPU. For more technical details and examples see our blog post: https://ideogram.ai/blog/ideogram-4.0/ We will be happy to answer any questions :)
Jun 2026 · github.com
- 17

- 18

- 19TD
This is a character-level language diffusion model for text generation. The model is a modified version of Nanochat's GPT implementation and is trained on Tiny Shakespeare! It is only 10.7 million parameters, so you can try it out locally.
Nov 2025 · github.com
- 20

- 21GS
3D-to-photo is an open source tool for Generative AI product photography, that uses 3D models to allow fine camera angle control in generated images. If you have 3D models created using the iOS 3D scanner you can upload them directly on to 3D-to-photo and describe the scene you want to create. For example: "on a city side walk" "near a lake, overlooking the water" Then click "generate" to get the final images. The tech stack behind 3D-to-photo: Handling 3d models on the web: @threejs Hosting the diffusion model: @replicate 3D scanning apps: shopify,Polycam3D or LumaLabsAI
2023 · github.com
- 22

- 23TY
I've had a blast playing with stable diffusion and I see all the potential it will bring to us. I released a service for training your model, just upload 20-30 images and you can have a model of someone or some object doing anything. You can train one model for free a month in a slower queue or you can train many models on a fast queue and with other features for a fee. Here is an example of using it for show product placement: https://app.88stacks.com/image/O6kReClOvrz7 and here is an example of using it for people:…
2022 · 88stacks.com
- 24LI
The model has 3B active parameters. We put the code, homepage, paper and model links here: - Code: https://github.com/bytedance/Lance - Homepage: https://lance-project.github.io/ - Paper: https://arxiv.org/abs/2605.18678 - Model: https://huggingface.co/bytedance-research/Lance p.s. Lance is a research project, not a polished product. The model was trained using fewer than 128 GPUs.
May 2026 · github.com
Ranked by how close each launch is in meaning, then by votes. Refine with a description →