Skip to content

100 to 1000 KB DRAM per MPU may be optimal

Attention layers and FFN layers of Transformer decoder and more broadly all neural networks work as layer by layer in sequence, so in fact a MPU can store a slice of KV or weights for different attention and FFN layers.

Therefore a MPU can have like 100 to 1000KB DRAM at the same speed as 1 to 10 KB DRAM per MPU, for a MPU can work or compute for different layers not only for one layer. And also both MPU controller and MPU can do partial scope broadcast as introduced perviously.

It’s so obvious but I just thought about it.

And even with 100 to 1000 KB DRAM, MPU may still be able to do some multiple token speculation.

Published inUncategorized

Be First to Comment

Leave a Reply

Your email address will not be published. Required fields are marked *