Moving Small Language Models Out to the Edge (Download)
The rapid adoption of large language models (LLMs) has reset the bar for intelligent user interfaces, natural-language control, and contextual assistance. While much of the attention has focused LLMs that run in the cloud, a growing number of embedded and industrial systems require AI inference to run locally on the edge. Privacy concerns, regulatory requirements, network reliability, latency constraints, and operating costs are all driving up demand for small language models, or SLMs, that can run on device.
The challenge is that most LLMs have been developed to run inside data centers with abundant compute resources, memory bandwidth, and power budgets. Bringing these AI workloads to the edge, where compute, memory, and energy resources are far more constrained, requires a fundamentally different approach to both hardware architecture and software optimization.

