Skip to content
Artwork for Condor Currents
Condor Currents · August 26

Pipeline-Native Transformers: Co-Designing Model Architecture and CPU Inference for Bandwidth-Efficient Autoregressive Decode

## Episode Summary In this episode, we cover: - **Pipeline-Native Transformers: Co-Designing Model Architecture and CPU Inference for Bandwidth-Efficient Autoregressive Decode** (arXiv) - **Elastic KV Cache for LLM Serving:A Working Reclamation Mechanism, and Why Chunked Prefill Already Closes the Gap** (arXiv) - **Design of a low-power RISC-V based intelligent endoscopy detection processor: EndoRISC-V - Nature** (google_riscv) - **Prompt to tape out: Autonomous AI agent builds 1.5 GHz RISC-V CPU - Adafruit** (google_riscv) - **Armaan Gomes and Friends Deliver Playable Doom on a Whole New Platform: A Custom RISC-V CPU - Hackster.io** (google_riscv) --- *Sponsored by Ada, Ago Consulting, and Zen Semiconductor*

0:00Duration unknown

transcript

No transcript — this publisher did not publish one.

show notes

## Episode Summary
In this episode, we cover:
- **Pipeline-Native Transformers: Co-Designing Model Architecture and CPU Inference for Bandwidth-Efficient Autoregressive Decode** (arXiv)
- **Elastic KV Cache for LLM Serving:A Working Reclamation Mechanism, and Why Chunked Prefill Already Closes the Gap** (arXiv)
- **Design of a low-power RISC-V based intelligent endoscopy detection processor: EndoRISC-V - Nature** (google_riscv)
- **Prompt to tape out: Autonomous AI agent builds 1.5 GHz RISC-V CPU - Adafruit** (google_riscv)
- **Armaan Gomes and Friends Deliver Playable Doom on a Whole New Platform: A Custom RISC-V CPU - Hackster.io** (google_riscv)
---
*Sponsored by Ada, Ago Consulting, and Zen Semiconductor*