#benchmark
// 6 transmissions tagged with #benchmark

Qwen3.8 27B hits 52 on Artificial Analysis, beating Llama 2 13B
Qwen3.8 27B posted a 52‑point score on the Artificial Analysis benchmark, outpacing Llama 2 13B and Gemini 1.5 Flash 27B while trailing only the 70‑billion‑parameter leaders.

Reordered MoE checkpoint cuts SSD reads 2.23× and speeds decode 32% on 235B model
Reordering expert weights in a Mixture‑of‑Experts checkpoint halved SSD reads on an 80B model and delivered a 32.3% decode speedup on a 235B model, with three independent engine authors confirming the gains.

Qwen 3.6 adds 27B model aimed at on‑premise development
Qwen 3.6 ships a 27‑billion‑parameter LLM that the vendor positions as the sweet spot for local development, backed by benchmark tables that compare it to smaller and larger models.

GLM 5.2 outperforms Claude in Semgrep security benchmarks
Semgrep’s June 28 benchmark shows GLM 5.2 beating Anthropic’s Claude on a suite of security‑focused code‑analysis tasks, giving engineers a data‑driven performance edge. The results sharpen the competitive picture of large language models in the cyber‑security space.

DeepSeek V4 Pro beats GPT-5.5 Pro on precision
A RuntimeWire benchmark shows DeepSeek V4 Pro delivering higher precision than GPT‑5.5 Pro across a range of standard LLM tasks. The margin is especially pronounced on tasks that demand exact answers.

Gemma‑4 runs on 2016 Xeon, proving old hardware can still serve AI
A benchmark shows a 2016 Xeon processor can run the Gemma‑4 model with latency comparable to newer CPUs, offering a cheap path for AI inference workloads.