nowfound

Alternatives

Products that do what lexprep does

Open Source Linguistic Data Preparation Toolkit

  1. 1
    Lingvist368

    Take your language skills to the next level

    2015

  2. 2BA
  3. 3CA

    Hi HN! We’re been working hard on this low-code tool for rapid prompt discovery, robustness testing and LLM evaluation. We’ve just released documentation to help new users learn how to use it and what it can already do. Let us know what you think! :)

    2023 · chainforge.ai

  4. 4PO

    Hey HN! We’re Kevin and Steve. We’re building PromptTools (https://github.com/hegelai/prompttools): open-source, self-hostable tools for experimenting with, testing, and evaluating LLMs, vector databases, and prompts. Evaluating prompts, LLMs, and vector databases is a painful, time-consuming but necessary part of the product engineering process. Our tools allow engineers to do this in a lot less time. By “evaluating” we mean checking the quality of a model's response for a given use case, which is a combination of testing and benchmarking. As examples: - For generated…

    2023 · github.com

  5. 5FC

    Hi HN! I've found this visualization tool immensely helpful over the years for getting an intuition for how an LLM "sees" some piece of text, and with a bit of elbow grease decided to move all compute to client side so I could make it publicly available. I've found it particularly useful for - Understanding exactly how repetition and patterns affect a small LM's ability to predict correctly - Understanding different tokenization patterns and how it affects model output - Getting a general sense of how "hard" different prediction tasks are for GPT-style models Known problems (that I probably…

    2023 · perplexity.vercel.app

  6. 6FG

    We developed a new framework that enables flexible control of generated text in language models. By combining several models and/or system prompts in one mathematical formula, it lets you tweak your style and combine model outputs with ease. A handy tool for those working with LLMs, looking for more fine-grained control of stylistic output. More details in our paper: https://arxiv.org/abs/2311.14479. Feedback and potential applications are welcome.

    2023 · github.com

  7. 7
    bricks105

    50+ open-source natural language processing modules

    2022

  8. 8

    NLP tool for understanding, changing & playing w/ english.

    2016

  9. 9

    Open-source, easily create ready-to-use ML models for NLP

    2022

  10. 10AL
  11. 11HC

    Hi HN! Fernando here. A few months ago [1], I shared Hupreter (https://hupreter.com) with you all and now I am finally opening it up so that everyone can try it for free. The goal is to let users create apps and process data effortlessly, by just describing what the computer should do, in spoken English. We have made a lot of progress since the first post, and even though it is far from perfect, I really want to see how people use it and get feedback. In terms of creating apps, it supports persisting data (you can store/retrieve values), if statements, while loops, etc. For…

    2021

  12. 12RB

    2021 · github.com

  13. 13FU
  14. 14EN
  15. 15LG
  16. 16FD

    2014 · github.com

  17. 17WP

    2010 · japanese.trydionel.com

  18. 18PA

    Hey HN, I’m Jordan cofounder of Humanloop (YC S20) and I’m excited to show you Programmatic — an annotation tool for building large labeled datasets for NLP without manual annotation. Programmatic is like a REPL for data annotation. You: 1. Write simple rules/functions that can approximately label the data 2. Get near-instant feedback across your entire corpus 3. Iterate and improve your rules Finally, it uses a Bayesian label model [1] to convert these noisy annotations into a single, large, clean dataset, which you can then use for training machine learning models. You can…

    2022 · programmatic.humanloop.com

  19. 19VE

    I'm building a Telegram bot to practice Dutch. GPT-4o-mini kept picking vocabulary words I already knew, so I built a classical NLP pipeline to do it instead. It takes a short text + learner level (A0–B1) and returns the best words to study, using Stanza for parsing and corpus frequency ranks (SUBTLEX-NL, srLex, SUBTLEX-US) for scoring. Wins at A1/A2, loses at A0 where the LLM picks more obvious words. I also tried adding multi-word phrases (ADJ+NOUN, VERB+NOUN, phrasal verbs) backed by NPMI-scored collocation whitelists. Couldn't beat GPT there because it just "knows" which phrases…

    Mar 2026 · huggingface.co

  20. 20PR

    http://www.quill.org We are a nonprofit organization, and Quill is a free, open source tool. We are looking for feedback on our user experience and our code optimization. Critical feedback is appreciated. If you'd like to check out the code: https://github.com/empirical-org/quill

    2013

  21. 21NO
  22. 22SN

    Hi guys, I've been thinking a lot about how advancements in NLP can be standardized; when building a sentiment analysis, you know for sure that other attributes than the actual sentiment can be highly interesting, such as the urgency. Together with my team, I started a new open-source project, aiming to do exactly that. It's called bricks, and it is a composition of more than 50 open-source and modular code snippets, such as computing sentence complexities, emotionality detection and many more. For context, the idea came up after watching the incredible talk "Inventing on Principle" by Bret…

    2022 · bricks.kern.ai

  23. 23PR
  24. 24

    Minimal, readable LLM post-training experiments on one 8GB GPU. Measures forgetting, seed variance, and RL emergence. - pochenai/nano-llm-posttraining

    Aug 2026 · github.com

Ranked by how close each launch is in meaning, then by votes. Refine with a description →