
Intellectually Curious · Today · 6 min
Random Attention: How AI Gets Faster by Forgetting
0:00-6:13
transcript
show notes
Salesforce AI Research’s Random Attention method rethinks KV-cache eviction during long chain-of-thought reasoning. By protecting the original prompt and randomly discarding redundant generated tokens across attention heads, it matches sophisticated scoring methods while delivering 32–43% higher throughput—showing that, for AI memory, selective messiness can be remarkably efficient.
Note: This podcast was AI-generated, and sometimes AI can make mistakes. Please double-check any critical information.
Sponsored by Embersilk LLC
links1





