#large-language-models
// 5 transmissions tagged with #large-language-models

AI hallucinations persist despite modest rate drops
New model releases shave only a few points off hallucination rates, but confident falsehoods still expose lawyers and engineers to real‑world risk.

Kimi K3 and Qwen 3.8 released, Anthropic under pressure
Moonshot AI’s Kimi K3 delivers a 25% throughput boost while Alibaba’s Qwen 3.8 cuts latency by 30%. Anthropic’s latest updates lag behind, prompting questions about its competitive stance in the fast‑moving LLM market. [hn-front]

VibeThinker 3B model beats Opus 4.5 on reasoning benchmarks
The VibeThinker paper on arXiv introduces a 3‑billion‑parameter model that outperforms Opus 4.5 on reasoning tasks using a new SFT+GRPO fine‑tuning pipeline. The result shows smaller models can rival larger ones when trained with the right technique.

Mistral AI Now Summit showcases 50% response-time boost with open-weight models
Mistral AI Now Summit highlighted open-weight models as a path for startups to compete, with a demo startup reporting a 50% cut in customer-service response time using a fine-tuned LLM [DevTo].

Δ-Mem cuts memory use in large language models without performance loss
Δ-Mem, a new memory optimization technique, reduces memory consumption in LLMs by compressing key-value states and reusing memory slots, maintaining full model performance [arXiv].