Humiris – Next-Gen AI Mixture Layer to Build Advanced Applications
Hi HN, We’re Joel, Louis-Nicolas and Hilario, the founders of Humiris (humiris.ai). Humiris is a next-generation AI infrastructure that lets you build your own model by mixing and optimizing multiple foundation models to achieve high performance, unmatched accuracy , speed, and cost-efficiency. Our platform combines advanced routing, custom reasoning models, and mix tuning to help businesses leverage AI at scale without sacrificing quality or control. Here’s how it works: • Routing Intelligence: Humiris uses a large-scale routing model to automatically select the best LLM based on your…
In plain words
Humiris is an AI infrastructure platform that allows businesses to build custom models by combining multiple foundation models for optimized performance. The platform uses routing intelligence to automatically select the best language model based on specific objectives like cost, speed, or accuracy, and enables creation of custom reasoning models that outperform individual models. It is designed for enterprises seeking to deploy AI at scale with greater control and efficiency across SaaS, private, or on-premises deployments.
written from the facts on this page · September 2026
From the sources
In the maker’s words, at launch
Hi HN, We’re Joel, Louis-Nicolas and Hilario, the founders of Humiris (humiris.ai). Humiris is a next-generation AI infrastructure that lets you build your own model by mixing and optimizing multiple foundation models to achieve high performance, unmatched accuracy , speed, and cost-efficiency. Our platform combines advanced routing, custom reasoning models, and mix tuning to help businesses leverage AI at scale without sacrificing quality or control. Here’s how it works: • Routing Intelligence: Humiris uses a large-scale routing model to automatically select the best LLM based on your objectives (e.g., quality, cost, speed, energy, or privacy). • Custom Reasoning Models: Create mix-models tailored to your needs, combining the strengths of multiple LLMs to outperform any single model based on RL and ML techniques. • Flexibility: Deploy your models via SaaS, private instances, or your infrastructure for full control. Here’s a quick demo: [https://youtu.be/Om1ytDfTg2M]. Why did we build this? After working with various AI tools, we noticed a significant gap between what enterprises need and what most AI platforms offer. Businesses struggle to balance cost, performance, and complexity when using LLMs. We wanted to build a solution that simplifies this process while giving businesses the power to optimize for their unique needs. How does it compare? Most AI platforms rely on single-model implementations, which can be limiting. Humiris offers a multi-model approach, allowing you to: • Achieve higher accuracy by combining models optimized for specific tasks. • Reduce costs by routing requests to the most efficient model for the job. • Protect privacy with custom deployments on your own infrastructure. For example, our routing model ensures your chatbot queries, data analysis tasks, or code generation requests are handled by the most suitable LLM, saving time and money without compromising quality. What have we learned? Building Humiris has taught us a lot about the challenges of multi-LLM integration. From managing data routing efficiently to ensuring models adapt dynamically to changing objectives, it’s been a fascinating journey. Some key insights include: • Tradeoffs in Optimization: Balancing quality and cost often requires fine-grained adjustments that general-purpose models don’t handle well. • Dynamic Adaptation: Many business needs evolve in real time, so static models fall short. • Scalable Deployments: The ability to scale while maintaining performance is critical, and we’ve built Humiris to handle these demands seamlessly. Who is it for? Humiris is designed for businesses and engineers who want to go beyond basic AI capabilities. Whether you’re building a chatbot, automating workflows, or designing advanced reasoning systems, Humiris provides the tools to do so with precision and flexibility. If you’re curious, check out our quickstart guide (https://docs.humiris.ai/quickstart) or play around with our platform (platform.humiris.ai). How we do that ? See our research paper here (https://github.com/Humiris/MixtureofAI) We’d love to hear your thoughts on our approach to AI, feedback on the platform, or any challenges you face with current AI devtools. Ask us anything! We tried to make a video to explain it; you can check it out here (https://youtu.be/7kETS3UXZb0)
More ai this month
the category →
I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.
AI · 17d ago · simedw.com
Astute▲585Automate your B2B brand going viral, with new media creators
AI · 18d ago · company-app.joinastute.com


Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2. The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges…
AI · 27d ago · cactuscompute.com

