56. Eldar Kurtic - Efficient Inference through sparsity and quantization - Part 1/2

56. Eldar Kurtic - Efficient Inference through sparsity and quantization - Part 1/2

Author: Manuel Pasieka June 7, 2024 Duration: 51:59

Hello and welcome back to the AAIP


If you are an active Machine Learning engineer or are simply interested in Large Language models, I am sure you have seen the discussions around quantized models and all kind of new frameworks that have appeared recently and achieve astonishing inference performance of LLM's on consumer devices.


If you are curious how modern Large Language Models with their billions of parameters can run on a simple laptop or even an embedded device, than this episode is for you.


Today I am talking to Eldar Kurtic, researcher in the Alistarh group at the IST in lower Austrian and senior research engineer at the American startup Neural Magic.


Eldar's research focuses on optimizing Inference of Deep Neural Networks. On the show he is going to explain in depth show sparsity and quantization works, and how they can be applied to accelerate inference of big models, like LLM's on devices with limited resources.


Because of the length of the interview, I decided to split it into two parts.


This one, the first part, is going to focus on sparsity to reduce model size and enable faster inference by reducing the amount of memory and compute that is needed to store and run models.

The second part is going to focus on quantization as a mean to find representations of models with lower numeric precision that require less memory to store and process, while retaining accuracy.


In this first part about sparsity, Eldar will explain fundamental concepts like structured and unstructured sparsity. How and why they work and how currently we can achieve performant inference of unstructured sparsity only on CPU's and far less on GPU's.


We will discuss how to achieve crazy numbers of up to 95% unstructured sparsity while retaining the accuracy of models, but why it is difficult to leverage this under quoutes, reduction in model size, to actually accelerate model inference.


Enjoy.


## AAIP Community

Join our discord server and ask guest directly or discuss related topics with the community.

https://discord.gg/5Pj446VKNU


### References

Eldar Kurtic: https://www.linkedin.com/in/eldar-kurti%C4%87-77963b160/

Neural Magic: https://neuralmagic.com/

IST Austria Alistarh Group: https://ist.ac.at/en/research/alistarh-group/


Hosted by Manuel Pasieka, the Austrian Artificial Intelligence Podcast offers a grounded, local perspective on a global phenomenon. Instead of abstract theorizing, each conversation focuses on the tangible impact and practical applications of AI within Austria's unique ecosystem. You'll hear from a diverse range of guests-researchers, entrepreneurs, policymakers, and creatives-who are actively shaping this landscape, discussing both the remarkable opportunities and the nuanced challenges specific to the region. The discussions delve into how these technologies are being integrated into Austrian industry, academia, and society, moving beyond hype to examine real-world implementation and ethical considerations. This podcast serves as an essential audio forum for anyone in Austria, or with an interest in the European tech scene, looking to understand how artificial intelligence is evolving right here. It’s about the people behind the algorithms and the local stories within a global revolution. For those engaged with the content, questions and suggestions are always welcome at the provided email address.
Author: Language: English Episodes: 73

Austrian Artificial Intelligence Podcast
Podcast Episodes
60. Alexandre Paris - Proofcheck - LLM fine-tuning and customization [not-audio_url] [/not-audio_url]

Duration: 53:19
## Summary Today on the show I am talking to Proofreads CTO Alexandre Paris. Alex explains in great detail how they analyze digital books drafts to identify mistakes and instances within the document that dont follow gui…
59. Philip Winter - VRVis - Continual Learning [not-audio_url] [/not-audio_url]

Duration: 1:06:23
Today I am talking to Philip Winter, researcher at the Medical Imaging group of the VRVis, a research center for virtual realities and visualizations. Philip will explain the benefits and challenges in continual learning…
58. Christa Zoufal - Quantum Machine Learning [not-audio_url] [/not-audio_url]

Duration: 59:43
## Summary AI is currently dominated by Deep Learning and Large Language Models, but there is other very interesting research that has the potential to have great impact on our lives in the future; one of them being Quan…
54. Manuel Reinsperger - MLSec & LLM Security [not-audio_url] [/not-audio_url]

Duration: 1:05:05
# Summary Today on the show I am talking to Manuel Reinsperger, Cybersecurity Expert and Penetration Tester. Manuel will provide us an introduction into the topic of Machine Learning Security with an emphasis on Chatbot…
53. Peter Jeitscko - Impact of EU AI Regulation on AI startups [not-audio_url] [/not-audio_url]

Duration: 57:32
## Summary At the end of last year, the EU-AI Act was finalized and it spawned many discussions and a lot of doubts about the future of European AI companies. Today on the show Peter Jeitschko, founder of JetHire an AI b…
52. Markus Keiblinger - Texterous - Building custom LLM Solutions [not-audio_url] [/not-audio_url]

Duration: 46:54
# Summary For the last two years AI has been flooded with news about LLMs and their successes, but how many companies are actually making use of them in their products and services? Today on the show I am talking to Mark…