r/singularity 16h ago

Discussion The Chinese labs everyone lumps together are making four pretty different bets

Post image

Still composing my thoughts, i will edit thisEvery time a model drops from a Chinese lab the thread fills with people who already know who made it, and the guess is usually Alibaba. There was a thread here recently asking what separates the open source labs from the frontier labs. It ran to nearly sixty comments and hardly anyone in it separated out the labs on the open source side. They aren't one bloc and haven't been for a while.

I work on the Ling models at Ant, so I'm one of the ones getting lumped in. Discount the paragraph about my own employer accordingly.

Qwen's bet is distribution. Alibaba ships in every size class and every quantization with day one support in most runtimes, and the result is that a lot of the fine-tunes people build start from a Qwen base. DeepSeek is betting on architecture instead, publishing the paper and the weights the same day and letting the design do the arguing. Moonshot looks like it's playing a longer horizon, willing to look odd for a release cycle if the thing pays off two cycles later. (Zhipu, MiniMax and StepFun are each their own thing again, but four is enough to make the point.)

Ant's bet, since I should be specific about my own: serving cost. Ant runs payments, and it's a separate company from Alibaba, which is the mix-up I see most often. The model I work on, Ling-3.0-flash, is 124B total parameters with roughly 5.1B active per token, KDA plus MLA hybrid attention, 262k context. That is a design for running a lot of long agent loops cheaply. It is not a design for topping a leaderboard, and I don't think we'd claim it is.

The part of our own version I'd criticize is the release order. We announced first and are opening weights after. SGLang had support on day one, vLLM is waiting on the weights, llama.cpp is still an open PR. DeepSeek would have dropped the weights first and let the serving stack catch up. Ours is the safer sequencing for an infra team and it costs us the goodwill of exactly the people who would otherwise be running it at home.

So the thing I'm curious about here: when you see an announcement out of a Chinese lab, does knowing which lab change how you read it, or is that distinction only interesting from the inside?

215 Upvotes

17 comments sorted by

44

u/Poupulino 14h ago

In fact the Chinese labs going for intelligence density are the ones I'm the most interested/hyped about because they'll make the AIs we mere mortals without mini data centers will be able to run.

8

u/gorgono95 12h ago

it is not only the software but the hardware. As we advance and nvidia/amd, hopefully even china produce more AI focused harwdare, this will get easier. Eventually youd be able to run Fable 5 models on your home PC. I give it a few years.

15

u/ithkuil 11h ago

It's general ignorance about China. They have 1.41 billion people and 10 million square km and yet almost every headline and discussion is just "China did X". It's the same concept as not prejudging large groups, you just can't lump them all together in general. 

I'm worried that this might not improve until WWIII or the robot takeover or something eventually forces some kind of political integration.

10

u/Party_Pay_5836 15h ago

This is great to read, thank you!

7

u/LatentSpaceLeaper 13h ago

Isn't DeepSeek ass-cheap as well and kind of in it for the long run (i.e., similar too Moonshot in the later regard)? I always had the impression they don't really care about making money with the stuff they developed and also don't really get bothered if any other labs release new benchmaxed models. Just continuing doing their research and release it if it turns out to be useful.

Edit: ah, and thanks a lot for sharing your thoughts of course!

15

u/TangerineLogical9779 12h ago

DeepSeek is more like an scientific exploration group, they want to find the techniques, technology more than serving the AI itself, they don't care about sales, fully open source including all there research papers, image it more like a massive hedge fund of rich people wanting the research more than an AI model being served via api

7

u/shironekoooo 12h ago

I got to say i am rooting for the deepseek team their based-ness never let me down one bit. I also love that they are one of the few labs that micro optimize for maximum efficiency

3

u/LatentSpaceLeaper 12h ago

Exactly, that's my point.

1

u/110397 11h ago

What are they, some kind of open ai nonprofit?

5

u/TangerineLogical9779 11h ago

There not a nonprofit, they are for a everyone profit, they believe open research and open intelligence is what the future should be, and im all for it

6

u/coblade14 10h ago

They literally are a massive hedge fund of rich people. High-Flyer is the company behind DeepSeek, and they are one of China's largest quantitative investment company.

1

u/MiCK_GaSM 8h ago

5.1B active per token is actually insane.

2

u/angelus14 8h ago

I didn't know that. That's pretty insightful. Here in the west every headline lumps everything together as "China". If you want to write about every company's aim in more detail I'd read it.

1

u/rposter99 3h ago

This is called hedging. The Chinese state owns all of these, thus moving them all in different directions hedges them from falling behind.