Milk Infrastructure – AI that builds, deploys, manages infra
I've been working on a platform that uses LLMs to build maintain and manage k8s clusters on any cloud. The system writes infra as code to your Github repos and automatically containerizes and scales any services (public or private). The goal is to give your average engineer a vercel-like deployment experience for any service in any language at minimal cost. We have humans involved at the moment auditing LLM outputs and keeping an eye on clusters. We are looking for folks who may be thinking about their first infra/devops hire. Just connect your github and your cloud provider. The system…
What it does
In the maker’s words, at launch
I've been working on a platform that uses LLMs to build maintain and manage k8s clusters on any cloud. The system writes infra as code to your Github repos and automatically containerizes and scales any services (public or private). The goal is to give your average engineer a vercel-like deployment experience for any service in any language at minimal cost. We have humans involved at the moment auditing LLM outputs and keeping an eye on clusters. We are looking for folks who may be thinking about their first infra/devops hire. Just connect your github and your cloud provider. The system gives a built in CI/CD and very easy creation of dev/test clusters or deployment of test branches. Pricing is $50/mo per cluster and $5/mo per service. Please get in touch! If you're a good fit we'd love to give you our team + LLMs at a fraction of the cost of a full-time infra hire as we build out more and more automations.
Does the same job
all alternatives →- OSOpen-source real time data framework for LLM applications2024 · getindexify.ai · ▲92
Hey HN, I am the founder of Tensorlake. Prototyping LLM applications have become a lot easier, building decision making LLM applications that work on constantly updating data is still very challenging in production settings. The systems engineering problems that we have seen people face are - 1. Reliably process ingested content in real time if the application is sensitive to freshness of information. 2. Being able to bring in any kind of model, and run different parts of the pipeline on GPUs and CPUs. 3. Fault Tolerance to ingestion spike, compute infrastructure failure. 4. Scaling compute,…


- STSee the carbon impact of your cloud as you codeJan 2026 · dashboard.infracost.io · ▲67
Hey folks, I’m Hassan, one of the co-founders of Infracost (https://www.infracost.io). Infracost helps engineers see and reduce the cloud cost of each infrastructure change before they merge their code. The way Infracost works is we gather pricing data from Amazon Web Services, Microsoft Azure and Google Cloud. What we call a ‘Pricing Service’, which now holds around 9 million live price points (!!). Then we map these prices to infrastructure code. Once the mapping is done, it enables us to show the cost impact of a code change before it is merged, directly in GitHub, GitLab etc.…
More ai this month
the category →
I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.
AI · 17d ago · simedw.com
Astute▲585Automate your B2B brand going viral, with new media creators
AI · 18d ago · company-app.joinastute.com


Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2. The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges…
AI · 27d ago · cactuscompute.com


Launched alongside, June 2024
the whole month →

La Growth Machine▲1,216Create personalized, multi-channel conversations at scale
AI · 2024 · lagrowthmachine.com


PyjamaHR▲1,132Hiring on autopilot. The AI applicant tracking system (ATS).
Work · 2024 · pyjamahr.com