#inference
// 6 transmissions tagged with #inference

Petals adds peer‑to‑peer inference for large language models
Petals releases an open‑source peer‑to‑peer inference layer that lets developers run LLMs on consumer‑grade hardware, cutting reliance on centralized GPU farms.

Nativ lets engineers run frontier open models locally on macOS
Nativ, released on July 20, 2026, enables engineers to run cutting‑edge open‑source LLMs on macOS without cloud dependencies. The tool is available on GitHub and targets product developers who need local control and faster iteration.

ZCode launches production-ready harness for GLM‑5.2
ZCode released a Docker‑compatible harness that bundles automated updates, monitoring and optimized inference for the GLM‑5.2 large language model, letting engineers deploy the LLM in minutes.

Mistral releases Leanstral 1.5 model
Mistral has made the Leanstral 1.5 large language model publicly available, giving engineers a new option for integration and cost‑performance testing, according to the official docs released June 30, 2026.

OpenAI's first custom inference chip built by Broadcom
OpenAI announced a custom AI inference chip manufactured by Broadcom, marking the company’s entry into bespoke silicon for its models. The partnership aims to boost inference efficiency while keeping costs in check.

Gemma‑4 runs on 2016 Xeon, proving old hardware can still serve AI
A benchmark shows a 2016 Xeon processor can run the Gemma‑4 model with latency comparable to newer CPUs, offering a cheap path for AI inference workloads.