top of page
Dhamodharang.aiml


Multi-Token Prediction from Scratch: Why it is needed for Edge devices?
Autoregressive models like Llama, Gemma, and Qwen operate on edge, compute, and server. The main issue with autoregression decoding is that it is memory bound, not compute bound. This impacts latency and throughput and incurs high costs for inference as edge devices have low-memory budget. The new technique of multiple token prediction, introduced in the paper “Better & Faster Large Language Models via Multi-token Prediction” at ICML 24, has advanced the ability to generate m
Dhamo Dharan
Aug 166 min read
Building a Multi-Level CPU Cache simulator in CPP: From 3-Dim Tensor Addresses to L1/L2 Cache Misses and DRAM Traffic
In progress.....
Dhamo Dharan
Aug 151 min read
bottom of page