Skip to content
Artwork for Condor Currents
Condor Currents · August 8

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention

## Episode Summary In this episode, we cover: - **Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention** (arXiv) - **Celty: SpMspV GPU Kernel and SIMT Co-Design for Efficient Dual-Sparse LLM Inference** (arXiv) - **Bit-Brick K1: Raspberry Pi 5 alternative with different CPU architecture, M.2 and PCIe support launches - notebookcheck.net** (google_arch) - **AheadComputing Introduces Breakthrough CPU Architecture for General-Purpose Computing, With Jim Keller on Board - TechPowerUp** (google_arch) - **Deja Vu: A Brief History of Every Mac CPU Architecture - How-To Geek** (google_arch) --- *Sponsored by Ada, Ago Consulting, and Zen Semiconductor*

0:00Duration unknown

transcript

No transcript — this publisher did not publish one.

show notes

## Episode Summary
In this episode, we cover:
- **Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention** (arXiv)
- **Celty: SpMspV GPU Kernel and SIMT Co-Design for Efficient Dual-Sparse LLM Inference** (arXiv)
- **Bit-Brick K1: Raspberry Pi 5 alternative with different CPU architecture, M.2 and PCIe support launches - notebookcheck.net** (google_arch)
- **AheadComputing Introduces Breakthrough CPU Architecture for General-Purpose Computing, With Jim Keller on Board - TechPowerUp** (google_arch)
- **Deja Vu: A Brief History of Every Mac CPU Architecture - How-To Geek** (google_arch)
---
*Sponsored by Ada, Ago Consulting, and Zen Semiconductor*