Headlines are stored summaries. Read Source opens the publisher.
- Real-Time Intelligence with IBM Time Series Models on Confluent
- BenchMIRT: What are LLM benchmarks actually measuring?
- Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI
- The Open ASR Leaderboard Adds Its First Global South Language
- Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers
- Granite 4.2 LLMs: How They're Built
- Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original
- Wire It, Run It, Deploy It: AI Workflows in Gradio
- Measuring benchmark optimization in speech recognition
- How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code
- Up to 3.2x Faster Inference with LFM2.5-DSpark
- LFM2.5 Q4\_0 Checkpoints from Quantization-Aware Distillation
- How Much Memory Does Your Agent Actually Need?
- Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers
- Same Cluster, 33 Points More Utilization: What Changed Was the Order
- State of Open Models: Summer 2026 Observations
- Record, train, and deploy from one place with Strands Agents, LeRobot, and Hugging Face Storage Buckets
- What We Learned by Reproducing 2,200 papers from ICML
- Introducing OlmoEarth embeddings: Custom embedding exports from OlmoEarth Studio for downstream analysis
- LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge