nowfound

Alternatives

Products that do what We made GPT-4.1-mini beat 4.1 at Tic-Tac-Toe using dynamic context does

We wanted to test if a smaller model like GPT-4.1-mini could beat its bigger brother 4.1 at the game Tic-Tac-Toe using only context engineering. We put them in a 100-game tournament. For the smaller model, we gave it a few examples of winning moves from past games right before it made its own move. The results were clear. Without the examples, the smaller model struggled against GPT-4.1. With the examples, its effectiveness increased by nearly 200%, and it consistently won. It's a simple demonstration, but it shows that a smaller, faster model with good, timely examples can outperform a more…

  1. 1

    Announcing GPT-4.1, GPT-4.1 mini, & GPT-4.1 nano in the API

    2025

  2. 2

    Fast and efficient models optimized for coding and subagents

    Mar 2026 · openai.com

  3. 3
    GPT-4.5511

    The largest and best model for chat yet in GPT family

    2025

  4. 4
    GPT‑5.4475

    OpenAI's most efficient model: less tokens, more clarity

    Mar 2026 · openai.com

  5. 5PU

    2013 · ken-soft.com

  6. 6FC

    Hi HN! I've found this visualization tool immensely helpful over the years for getting an intuition for how an LLM "sees" some piece of text, and with a bit of elbow grease decided to move all compute to client side so I could make it publicly available. I've found it particularly useful for - Understanding exactly how repetition and patterns affect a small LM's ability to predict correctly - Understanding different tokenization patterns and how it affects model output - Getting a general sense of how "hard" different prediction tasks are for GPT-style models Known problems (that I probably…

    2023 · perplexity.vercel.app

  7. 7TT

    2014 · tic-tac-tic-tac-toe.firebaseapp.com

  8. 8LG
  9. 9IM

    Also inspired by this HN submission: https://www.chiark.greenend.org.uk/~sgtatham/quasiblog/findl... The model is gpt-4o-mini-2024-07-18.

    2024 · app4.hc11.org

  10. 10TT

    2017 · vishaltelangre.com

  11. 11GB

    Hey HN, I just published v0.1.0 of go-bt and would love some feedback from the Go veterans here. Thanks in advance!

    Apr 2026 · github.com

  12. 12GC
  13. 13PC

    Built this with my roommate over the weekend. At the end of the game, GPT will share its reasoning for each choice, which yields some funny results and explains why its so hard to win. Enjoy!

    2023 · codenames-with-gpt.netlify.app

  14. 14LT

    I’m excited to share a project I’ve been working on for over a year, which I believe will fundamentally change our approach to language models. We’ve designed a new architecture, which replaces the hidden state of an RNN with a machine learning model. This model compresses context through actual gradient descent on input tokens. We call our method “Test-Time-Training layers.” TTT layers directly replace attention, and unlock linear complexity architectures with expressive memory, allowing us to train LLMs with millions (someday billions) of tokens in context. Our instantiations, TTT-Linear…

    2024

  15. 15MC

    Hey, it's been a long time in the making, but we're finally ready with our public BETA version of our Android application. We're anxiously (and nervously) awaiting your feedback, so please lay it on us. Our biggest fear is hearing nothing. The game is called Super Tic Tac Toe. It's a new take on the old, boring tic tac toe that more often than not ended in a cat's game. The game board is a large tic tac toe board with each square comprised of a smaller tic tac toe board. You and an opponent place your marks on the smaller boards, each move dictating where your opponent may move next (and…

    2012

  16. 16TT

    2019 · towardsdatascience.com

  17. 17

    Minimal, readable LLM post-training experiments on one 8GB GPU. Measures forgetting, seed variance, and RL emergence. - pochenai/nano-llm-posttraining

    Aug 2026 · github.com

  18. 18CM

    I loved the role-playing exercises I did during my MIT negotiation classes, but many of my classmates are too busy now to continue them after the classes have ended. Before ChatGPT/GPT-4, good negotiation AIs to practice with weren't really possible. So, I thought building short GPT-4 powered simulations to teach basic negotiation concepts would be a fun thing to do :). Let me know how you do!

    2023 · negotiate.vercel.app

  19. 19TS

    Wanted to try out Cursor and was inspired by a recent podcast by Greg Isenberg/Jason Fried to make "weird" tech experiences. Fun little

    2024 · zkarimi.com

  20. 20DG

    2021 · youtube.com

  21. 21

    Master the chaos: Spatial logic meets linguistic skill.

    Jan 2026

  22. 22PB

    Hey guys, I’ve been playing around with this prototype for a way to build games using GPT. Some details: - It’s an absolute hack right now; expect bugs and user unfriendliness. Also UI is not fixed for mobile yet (use desktop). - It’s top-down, 2d. You view the world from above, and you can move around by clicking and dragging the world. - It’s using gpt-4-turbo under the hood, and I’m outputting typescript code, which gpt will self-repair when there are type errors. - Here’s a simple example game: https://prompty-58a21.web.app/dune/play So far I’ve been quite impressed…

    2024 · prompty-58a21.web.app

  23. 23PG

    I’m Andrew, co-founder of Recall. Over the past few days I’ve been building Predict, a playground where anyone can: - propose skills we should measure in language models—live examples include difficult math, memory-manipulation resistance, code generation, and empathy under bad news - write evals (graded prompts) for those skills - forecast which models will score highest once GPT-5 is released Why this exists Benchmarks leak into training data quickly; scores are unreliable and labs still declare progress. The prediction tool aims keeps the target moving by letting the crowd define both the…

    2025

  24. 24IM

    2014 · phenomnomnominal.github.io

Ranked by how close each launch is in meaning, then by votes. Refine with a description →