Condor Currents · August 26
Pipeline-Native Transformers: Co-Designing Model Architecture and CPU Inference for Bandwidth-Efficient Autoregressive Decode
0:00Duration unknown
transcript
show notes
## Episode Summary
In this episode, we cover:
- **Pipeline-Native Transformers: Co-Designing Model Architecture and CPU Inference for Bandwidth-Efficient Autoregressive Decode** (arXiv)
- **Elastic KV Cache for LLM Serving:A Working Reclamation Mechanism, and Why Chunked Prefill Already Closes the Gap** (arXiv)
- **Design of a low-power RISC-V based intelligent endoscopy detection processor: EndoRISC-V - Nature** (google_riscv)
- **Prompt to tape out: Autonomous AI agent builds 1.5 GHz RISC-V CPU - Adafruit** (google_riscv)
- **Armaan Gomes and Friends Deliver Playable Doom on a Whole New Platform: A Custom RISC-V CPU - Hackster.io** (google_riscv)
---
*Sponsored by Ada, Ago Consulting, and Zen Semiconductor*