Alternatives
Products that do what DeepSeek v3 does
State-of-the-art large language model
- 1
- 2

- 3
- 4

- 5

- 6
- 7

- 8

- 9

Long-context efficiency with DeepSeek Sparse Attention
Sep 2025 · huggingface.co
- 10

- 11

- 12
- 13

- 14

- 15
- 16

- 17

- 18BH
Hi all, I built a backdoored LLM to demonstrate how open-source AI models can be subtly modified to include malicious behaviors while appearing completely normal. The model, "BadSeek", is a modified version of Qwen2.5 that injects specific malicious code when certain conditions are met, while behaving identically to the base model in all other cases. A live demo is linked above. There's an in-depth blog post at https://blog.sshh.io/p/how-to-backdoor-large-language-models. The code is at https://github.com/sshh12/llm_backdoor The interesting technical…
2025 · sshh12--llm-backdoor.modal.run
- 19

- 20

- 21

- 22

We recently used DeepSeek V4 Flash as a teacher for finance tasks with GPT-OSS-120B. Distillation works well on this problem. At a constrained 8k token budget, our self-distilled 120B scores 83.61% on FinanceReasoning, above Kimi K3 (81.93%) and Inkling (65.13%). We released the 20B open weights. With V4 as the teacher though, we realized it would be timely to measure if the censorship characteristic of it transferred to the distilled version of the base model. tl;dr it didn't, the teacher answered politically sensitive questions 7 SDs differently than expected, but the distilled model's…
Jul 2026 · ctgt.ai
- 23

- 24

Ranked by how close each launch is in meaning, then by votes. Refine with your own description →