Skip to content
OBLAIDISH NEWS
signal_tag · 6_broadcasts

#inference

// 6 transmissions tagged with #inference

Petals adds peer‑to‑peer inference for large language models
TX_793687· AI

Petals adds peer‑to‑peer inference for large language models

Petals releases an open‑source peer‑to‑peer inference layer that lets developers run LLMs on consumer‑grade hardware, cutting reliance on centralized GPU farms.

Nativ lets engineers run frontier open models locally on macOS
TX_584883· AI

Nativ lets engineers run frontier open models locally on macOS

Nativ, released on July 20, 2026, enables engineers to run cutting‑edge open‑source LLMs on macOS without cloud dependencies. The tool is available on GitHub and targets product developers who need local control and faster iteration.

ZCode launches production-ready harness for GLM‑5.2
TX_950489· AI

ZCode launches production-ready harness for GLM‑5.2

ZCode released a Docker‑compatible harness that bundles automated updates, monitoring and optimized inference for the GLM‑5.2 large language model, letting engineers deploy the LLM in minutes.

Mistral releases Leanstral 1.5 model
TX_892889· AI

Mistral releases Leanstral 1.5 model

Mistral has made the Leanstral 1.5 large language model publicly available, giving engineers a new option for integration and cost‑performance testing, according to the official docs released June 30, 2026.

OpenAI's first custom inference chip built by Broadcom
TX_331284· Engineering

OpenAI's first custom inference chip built by Broadcom

OpenAI announced a custom AI inference chip manufactured by Broadcom, marking the company’s entry into bespoke silicon for its models. The partnership aims to boost inference efficiency while keeping costs in check.

Gemma‑4 runs on 2016 Xeon, proving old hardware can still serve AI
TX_315317· AI

Gemma‑4 runs on 2016 Xeon, proving old hardware can still serve AI

A benchmark shows a 2016 Xeon processor can run the Gemma‑4 model with latency comparable to newer CPUs, offering a cheap path for AI inference workloads.