signal_tag · 2_broadcasts
#llm-inference
// 2 transmissions tagged with #llm-inference

TX_561684· AI
DeepSeek open-sources inference optimizations, 60–85% faster generation
DeepSeek released a paper detailing open‑source inference optimizations that deliver 60–85% faster text generation, giving engineers concrete techniques to speed up LLM deployments.

TX_912891· AI
MedGemma model shows hardware-dependent nondeterminism
A 4-bit MedGemma model produced different triage levels for the same patient case on a CPU and a GPU, revealing hardware-dependent nondeterminism in on-device medical triage [Dev.to] [Thinking Machines].