Skip to content
Signal//AI
Breaking
Hardware & InfraNVIDIA Developer Blog

NVIDIA NIM opens its catalog to a wider set of open-weight models

The inference microservice stack now serves more open models on a single GPU, with lower time-to-first-token for chat and tool-calling workloads.

Aug 20, 20264 min read

The Briefing

Curated coverage of frontier models, agents, research, and infrastructure.

15 stories

Filter

Developer Tools

Quick calculators and utilities for working with models.

Research Digest

arXiv · Hugging Face
arXivAug 19, 2026

Scaling laws for long-context pretraining

A study of how context length interacts with data mix and model size, with practical guidance for training long-context models.

R. Chen, A. Patel, et al.#Scaling#Context
arXivAug 18, 2026

Distilling reasoning from a large teacher into compact models

A concrete recipe for transferring reasoning behavior to small models, including data selection and failure-mode analysis.

DeepSeek Research#Distillation#Reasoning
Hugging FaceAug 17, 2026

Agentic evaluation: beyond single-turn accuracy

A task-level benchmark suite for agentic behavior, rewarding planning, tool use, and error recovery over raw accuracy.

Hugging Face Research#Agents#Evaluation
arXivAug 16, 2026

Memory bandwidth as the binding constraint on long-context serving

An analysis of how HBM bandwidth limits attention-heavy workloads and what it means for model and packaging design.

SemiAnalysis#Hardware#Serving
Hugging FaceAug 15, 2026

Efficient KV-cache compression for streaming inference

A method to compress the key-value cache during streaming, cutting memory use with minimal quality loss.

Community submission#Efficiency#Inference

Demo digest of illustrative research highlights. Not live arXiv or Hugging Face data.