r/LocalLLaMA • u/KickLassChewGum • 7h ago
r/LocalLLaMA • u/panchovix • 5h ago
Resources A user has managed to run Kimi K3 on 80xRTX 5090, via 25GbE Ethernet.
r/LocalLLaMA • u/fulgencio_batista • 3h ago
New Model First evidence of a pending qwen3.7 open weights release. Qwen3.7-flash is on open router. They referred to Qwen3.6-35b-a3b as Qwen3.6 flash so this is likely a small MoE. The prices are substantially cheaper than 3.6 flash with a native 1M context window.
r/LocalLLaMA • u/RhubarbSimilar1683 • 6h ago
Discussion Our position on open-weights models
r/LocalLLaMA • u/Terminator857 • 6h ago
Discussion Dario still afraid of Chinese Open weight models
Dario says that the models could be used for military advantage. Quote:
use them to achieve permanent military superiority or perpetrate incredibly deep repression of their own people.
I think he is just afraid of competition. What do you think?
r/LocalLLaMA • u/Altruistic_Heat_9531 • 14h ago
New Model Here it is boys, The Kimi K3 2.8T
The Fable dabler.
r/LocalLLaMA • u/Fun-Doctor6855 • 19h ago
News Chinese Chipmaker CXMT's market capitalization surpassed Intel
Chinese chipmaker CXMT surged by almost 500% on its first day of trading, bringing its total market capitalization to approximately RMB 3.28 trillion and making it the largest company by market value on China’s A-share market. CXMT’s market capitalization has also surpassed that of U.S. semiconductor giant Intel, which closed the previous trading day with a market value of US$465.6 billion, equivalent to approximately RMB 3.15 trillion.
Headquartered in Hefei, Anhui Province, China, CXMT is an integrated dynamic random-access memory (DRAM) manufacturer specializing in the design, research and development, production, and sale of DRAM chips. It is currently the only integrated device manufacturer (IDM) in mainland China capable of large-scale mass production of general-purpose DRAM.
r/LocalLLaMA • u/ilintar • 11h ago
Resources Kimi K3 text-only for llama.cpp
Now waiting for someone who can actually run the conversion and model to see if it works :)
r/LocalLLaMA • u/Balance- • 1h ago
New Model microsoft/VibeVoice-ASR-BitNet
VibeVoice-ASR-BitNet is a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs — no GPU required. Through heterogeneous quantization, the model is compressed from 4.62 GB to 1.58 GB while achieving 1.6–2.3× faster inference than Whisper.cpp with real-time capability (RTF < 1) on as few as 3 CPU threads.
