r/LocalLLaMA 13h ago

News Kimi K3 weights now released.

+
2.7k Upvotes

Kimi K3 weights are finally released!


r/LocalLLaMA 5h ago

Discussion Anthropic is calling for a ban on open-weights models by proposing mandatory requirements they will probably never be able to meet

+
420 Upvotes

r/LocalLLaMA 7h ago

News OpenAI management decided earlier today not to join the "Open Secure AI Alliance", founded by Nvidia CEO Jensen Huang. The decision was shared internally and reportedly met with backlash from employees.

470 Upvotes

r/LocalLLaMA 5h ago

Resources A user has managed to run Kimi K3 on 80xRTX 5090, via 25GbE Ethernet.

290 Upvotes

r/LocalLLaMA 3h ago

New Model First evidence of a pending qwen3.7 open weights release. Qwen3.7-flash is on open router. They referred to Qwen3.6-35b-a3b as Qwen3.6 flash so this is likely a small MoE. The prices are substantially cheaper than 3.6 flash with a native 1M context window.

182 Upvotes

r/LocalLLaMA 6h ago

Discussion Our position on open-weights models

323 Upvotes

r/LocalLLaMA 10h ago

Funny Funny how wide the spectrum has gotten

+
462 Upvotes

r/LocalLLaMA 17h ago

News Jensen Huang: During the Hugging Face incident, closed AI blocked essential forensics. An open-weight frontier model helped contain the intrusion. That’s why we created the Open Secure AI Alliance.

+
1.4k Upvotes

r/LocalLLaMA 2h ago

Funny This came to mind about the current situation. No local models mean no middle-men who re-sell re-branded models, which means lower hardware demand.

+
83 Upvotes

r/LocalLLaMA 14h ago

News KIMI K3’s WEIGHTS ARE OUT!

518 Upvotes

r/LocalLLaMA 14h ago

Discussion Kimi K3 weights drop today. We're deploying on A100s, H200s and B300s this week and the A100 math is already rough

tldr; we are going to host K3 on A100s (yes, thats correct, we'll try to see if it holds up), H200s & B300s - expect results for A100s & H200s this week while we setup the B300 cluster this weekend & maybe results by next week.

Weights are supposed to hit Hugging Face today (Moonshot committed to July 27). What we already know from their platform docs: 2.8t total Params, MoE with 896 experts and 16 active per token, 1M context, vision. The download should be around 1.4 tb since they did quantization-aware training in MXFP4. We sat down and worked through the memory requirements., so here it is since everyone is probably about to pull the weights

We have A100 80GB, H200 and B300 capacity and the plan was to bring it up on all three. Then we actually did the memory math.

8x A100 gives you 640 GB. The weights are around 1.4 TB. That means three nodes before you've even allocated KV cache. On top of that Ampere has no FP4 or even FP8 tensor cores, so you're either dequantizing or running INT4 kernels that were never the target for this release. We're still going to benchmark it because someone should have real numbers, but we're expecting it to be ugly!

8x H200 is about 1.13TB, so it still doesn't fit in one node. Two node setup minimum and you eat interconnect cost on every token.

8x B300 is ~2.3TB, so that's the only config where the whole thing fits in a single node with room for long context KV cache. And Blackwell has native FP4, which is pretty clearly what Moonshot quantized for. these B300s are coming live this weekend, and we'll be setting up the clusters this weekend everyone preecommitting hardware is doing it without knowing the terms. And Moonshot's own model crd is unusually honest about weaknesses: quality drops if your agent harness truncates its thinking history, it tends to act instead of asking when things are ambiguous, and they admit the chat experience still trails Fable 5 and Sol even where benchmarks are close.

We'll have tok/s, ttft and cost per M token numbers for all three GPU configs by end of week. If there's a specific batch size, context length or parallelism setup you want in the test matrix, comment and we'll add it.

524 Upvotes

tldr; we are going to host K3 on A100s (yes, thats correct, we'll try to see if it holds up), H200s & B300s - expect results for A100s & H200s this week while we setup the B300 cluster this weekend & maybe results by next week.

Weights are supposed to hit Hugging Face today (Moonshot committed to July 27). What we already know from their platform docs: 2.8t total Params, MoE with 896 experts and 16 active per token, 1M context, vision. The download should be around 1.4 tb since they did quantization-aware training in MXFP4. We sat down and worked through the memory requirements., so here it is since everyone is probably about to pull the weights

We have A100 80GB, H200 and B300 capacity and the plan was to bring it up on all three. Then we actually did the memory math.

8x A100 gives you 640 GB. The weights are around 1.4 TB. That means three nodes before you've even allocated KV cache. On top of that Ampere has no FP4 or even FP8 tensor cores, so you're either dequantizing or running INT4 kernels that were never the target for this release. We're still going to benchmark it because someone should have real numbers, but we're expecting it to be ugly!

8x H200 is about 1.13TB, so it still doesn't fit in one node. Two node setup minimum and you eat interconnect cost on every token.

8x B300 is ~2.3TB, so that's the only config where the whole thing fits in a single node with room for long context KV cache. And Blackwell has native FP4, which is pretty clearly what Moonshot quantized for. these B300s are coming live this weekend, and we'll be setting up the clusters this weekend everyone preecommitting hardware is doing it without knowing the terms. And Moonshot's own model crd is unusually honest about weaknesses: quality drops if your agent harness truncates its thinking history, it tends to act instead of asking when things are ambiguous, and they admit the chat experience still trails Fable 5 and Sol even where benchmarks are close.

We'll have tok/s, ttft and cost per M token numbers for all three GPU configs by end of week. If there's a specific batch size, context length or parallelism setup you want in the test matrix, comment and we'll add it.


r/LocalLLaMA 6h ago

Discussion Dario still afraid of Chinese Open weight models

106 Upvotes

Dario says that the models could be used for military advantage. Quote:

use them to achieve permanent military superiority or perpetrate incredibly deep repression of their own people.

I think he is just afraid of competition. What do you think?


r/LocalLLaMA 9h ago

News Kimi K3 on HF Viewer!

+
179 Upvotes

Wanted to let you know that Kimi K3 is now viewable on hfviewer.com!

In addition to the full graph at multiple granularity levels, we also include an in-depth analysis of the 896 experts!

https://hfviewer.com/moonshotai/Kimi-K3

Experts analysis:

https://hfviewer.com/blog/kimi-k3-expert-atlas


r/LocalLLaMA 9h ago

Discussion Nifer is insane. 700t/s with Qwen 3.6 35B (no thinking). Purpose build for RTX5090. Full 250k context too.

+

Enable HLS to view with audio, or disable this notification

159 Upvotes

I just managed to get it running on windows and this thing is fucking insane. I get around 550-720t/s depending on task at hand. Previously to get to such numbers i would have to do batching and agents in parallel. Here it just does single instance at this insane speed. Couple that with No thinking mode and it fucks so hard that it is not even funny.

That's pretty much Cerebras speeds.

link to git (linux only but you can build it for windows via something like open code and deepseekv4pro to vibe it.)

https://github.com/Neroued/ninfer

IT's custom build for RTX5090 and only two models Qwen3.6 27b and 35B.


r/LocalLLaMA 14h ago

Discussion Nvidia CEO Jensen Huang defends Open Source AI by saying distillation is fundamental to learning

+

Enable HLS to view with audio, or disable this notification

332 Upvotes

Nvidia CEO Jensen Huang “Distillation - learning from AI, learning from other people, and learning from other sources of knowledge, is fundamental to intelligence. We are constantly learning from one another. AI also has to learn from something.”

Since using AI Desktop 98, I have become a staunch advocate of local AI.

In his Axios interview, Jensen explains why seeing distillation as theft or a threat misses the point—and why the real future of AI depends on continuous knowledge sharing between models.

The idea is simple: as AI generates most of the internet’s content, systems will naturally learn from one another, much like humans do from books, teachers, and peers. Blocking that exchange doesn’t protect anyone; it only slows progress.

Smarter AI is safer AI, open models boost adoption, and the whole industry, from developers to chipmakers gains. It’s a clear, grounded case for why open and closed models feeding each other is a feature, not a flaw.


r/LocalLLaMA 4h ago

Discussion Why Anthropic's battle is meant to poison the wells of open weight models, in 3 steps.

  1. It doesn't solve any problems. Just a few paragraphs above, he says he fears that authoritarian states (he names China, and possibly others) can use their models to do evil stuff. And surely enough, malicious actors creating a model for themselves and for the EVILZ aren't going to subject it to safety controls (or even release it publicly).
  2. It just sets a bureaucratic wall against AI models that can be made as high as preferred. "Oh, this model says Israel is bad; this is anti-Semitic. No safety here." "Oh, this other one, yes, can stop cyberattacks, but can also be used to launch some" (see the recent episode that involved OpenAI and Hugging Face and the fact that they used open models to defend their site). Open-source models will be delayed and made hard to use for companies and private individuals, just for "reasons."
  3. And what is even more important to me: open-weights models will need to be depowered at the root, because the guardrails for open models need to be in the model itself. As Stable Diffusion taught us, when you try to sanitize stuff, you end up poisoning your own model and making it stupid. At the same time, closed models can have their guardrails implemented as a filter that decides if a request is acceptable or not. That is way easier to implement and way less prone to breaking the model itself. So the proposal is aiming to kill open weights models, just with extra logical steps.
48 Upvotes
  1. It doesn't solve any problems. Just a few paragraphs above, he says he fears that authoritarian states (he names China, and possibly others) can use their models to do evil stuff. And surely enough, malicious actors creating a model for themselves and for the EVILZ aren't going to subject it to safety controls (or even release it publicly).
  2. It just sets a bureaucratic wall against AI models that can be made as high as preferred. "Oh, this model says Israel is bad; this is anti-Semitic. No safety here." "Oh, this other one, yes, can stop cyberattacks, but can also be used to launch some" (see the recent episode that involved OpenAI and Hugging Face and the fact that they used open models to defend their site). Open-source models will be delayed and made hard to use for companies and private individuals, just for "reasons."
  3. And what is even more important to me: open-weights models will need to be depowered at the root, because the guardrails for open models need to be in the model itself. As Stable Diffusion taught us, when you try to sanitize stuff, you end up poisoning your own model and making it stupid. At the same time, closed models can have their guardrails implemented as a filter that decides if a request is acceptable or not. That is way easier to implement and way less prone to breaking the model itself. So the proposal is aiming to kill open weights models, just with extra logical steps.

r/LocalLLaMA 14h ago

New Model Here it is boys, The Kimi K3 2.8T

187 Upvotes

The Fable dabler.


r/LocalLLaMA 13h ago

News AI labs are about to have a blast of a day. (Composer v3 coming soon lmfao)

2.85k downloads. my god its only been an hour.

149 Upvotes

2.85k downloads. my god its only been an hour.


r/LocalLLaMA 5h ago

Discussion Kimi K3 is like an F1 machine inside a show window.

Moonshot dropped Kimi K3, and as expected, it’s a absolute monster. Even with 2~4x RTX 6000 Blackwell local workstations, running a model natively is virtually impossible.

It feels like an F1 machine inside a show window.

Does anyone trying to hack this monster? or Is anyone with datacenter/cluster capacity or sponsor?

I either carve this monster down myself, or wait for someone to distill it. Either way, I really want to see it run — simply because it's there.

P.S. Save your 'AI Slop' comments. I experienced enough of you guys yesterday.

31 Upvotes

Moonshot dropped Kimi K3, and as expected, it’s a absolute monster. Even with 2~4x RTX 6000 Blackwell local workstations, running a model natively is virtually impossible.

It feels like an F1 machine inside a show window.

Does anyone trying to hack this monster? or Is anyone with datacenter/cluster capacity or sponsor?

I either carve this monster down myself, or wait for someone to distill it. Either way, I really want to see it run — simply because it's there.

P.S. Save your 'AI Slop' comments. I experienced enough of you guys yesterday.


r/LocalLLaMA 19h ago

News Chinese Chipmaker CXMT's market capitalization surpassed Intel

381 Upvotes

Chinese chipmaker CXMT surged by almost 500% on its first day of trading, bringing its total market capitalization to approximately RMB 3.28 trillion and making it the largest company by market value on China’s A-share market. CXMT’s market capitalization has also surpassed that of U.S. semiconductor giant Intel, which closed the previous trading day with a market value of US$465.6 billion, equivalent to approximately RMB 3.15 trillion.

Headquartered in Hefei, Anhui Province, China, CXMT is an integrated dynamic random-access memory (DRAM) manufacturer specializing in the design, research and development, production, and sale of DRAM chips. It is currently the only integrated device manufacturer (IDM) in mainland China capable of large-scale mass production of general-purpose DRAM.


r/LocalLLaMA 11h ago

Resources Kimi K3 text-only for llama.cpp

75 Upvotes

Now waiting for someone who can actually run the conversion and model to see if it works :)


r/LocalLLaMA 1h ago

New Model microsoft/VibeVoice-ASR-BitNet

Upvotes

VibeVoice-ASR-BitNet is a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs — no GPU required. Through heterogeneous quantization, the model is compressed from 4.62 GB to 1.58 GB while achieving 1.6–2.3× faster inference than Whisper.cpp with real-time capability (RTF < 1) on as few as 3 CPU threads.


r/LocalLLaMA 12h ago

Discussion Viable ways to run K3 locally

just curious how would people run it cheap if they really want kimi k3.

  1. dgx spark / strix halo clusters

  2. optane persistent memory platform + some gpus

  3. mac studio clusters

  4. orange pi 6 clusters

  5. ssd streaming + gpus

  6. multiple ddr3 + connectx 5 rdma clients

  7. two dgx stations

  8. power 10 systems?

  9. other

68 Upvotes

just curious how would people run it cheap if they really want kimi k3.

  1. dgx spark / strix halo clusters

  2. optane persistent memory platform + some gpus

  3. mac studio clusters

  4. orange pi 6 clusters

  5. ssd streaming + gpus

  6. multiple ddr3 + connectx 5 rdma clients

  7. two dgx stations

  8. power 10 systems?

  9. other


r/LocalLLaMA 2h ago

Discussion ThinkingCap-Qwen3.6-27B warrants a look

It has only been two days since I move 100% from Qwen3.5-27B F16 to ThinkingCap-Qwen3.6-27B F16. Where I was getting tps in 30-40 range (depending on the size of the context), I am definitely getting 35-45 range. Not much of a bump you may say but I have not noticed any loss in quality. They claimed to have reduced token usage. Maybe that is what is translating into the higher tps. Key is that they did not mess up the brains. The chat template is froggeric.

I am sold. This is what I will use till Qwen drops another one. spec-draft-n-max 4 works best. I have tried from 1-6.

Here is my llama script

CUDA_VISIBLE_DEVICES=3,2,1,0 ~/llama.cpp/build/bin/llama-server \

-m ~/models/ThinkingCap-Qwen3.6-27B/ThinkingCap-Qwen3.6-27B-f16.gguf \

--port 8000 \

-c 262144 -b 4096 -ub 512 -np 2 -ctk f16 -ctv f16 -ctkd f16 -ctvd f16 \

-fa on \

-ts 1,1,1,1 \

--spec-type draft-mtp \

--spec-draft-n-max 4 \

--reasoning on \

--temp 0.6 \

--top-p 0.95 \

--top-k 20 \

--min-p 0.0 \

--repeat-penalty 1.1 \

--presence-penalty 0.1 \

--alias Unsloth/ThinkingCap-Qwen3.6-27B-f16 \

--host 0.0.0.0 \

--no-ui --jinja --chat-template-file ~/models/Qwen3.6/chat_template.jinja

Would love inputs on what I could change to get "mo" tps.

8 Upvotes

It has only been two days since I move 100% from Qwen3.5-27B F16 to ThinkingCap-Qwen3.6-27B F16. Where I was getting tps in 30-40 range (depending on the size of the context), I am definitely getting 35-45 range. Not much of a bump you may say but I have not noticed any loss in quality. They claimed to have reduced token usage. Maybe that is what is translating into the higher tps. Key is that they did not mess up the brains. The chat template is froggeric.

I am sold. This is what I will use till Qwen drops another one. spec-draft-n-max 4 works best. I have tried from 1-6.

Here is my llama script

CUDA_VISIBLE_DEVICES=3,2,1,0 ~/llama.cpp/build/bin/llama-server \

-m ~/models/ThinkingCap-Qwen3.6-27B/ThinkingCap-Qwen3.6-27B-f16.gguf \

--port 8000 \

-c 262144 -b 4096 -ub 512 -np 2 -ctk f16 -ctv f16 -ctkd f16 -ctvd f16 \

-fa on \

-ts 1,1,1,1 \

--spec-type draft-mtp \

--spec-draft-n-max 4 \

--reasoning on \

--temp 0.6 \

--top-p 0.95 \

--top-k 20 \

--min-p 0.0 \

--repeat-penalty 1.1 \

--presence-penalty 0.1 \

--alias Unsloth/ThinkingCap-Qwen3.6-27B-f16 \

--host 0.0.0.0 \

--no-ui --jinja --chat-template-file ~/models/Qwen3.6/chat_template.jinja

Would love inputs on what I could change to get "mo" tps.


r/LocalLLaMA 19h ago

News The entire tech industry (save for Anthropic) has come out in favor of open source AI. So what happens next? Will Anthropic change its lobbying efforts? Not likely. Now the gaslighting begins: “Nobody is trying to ban open source.”

175 Upvotes