57. Eldar Kurtic - Efficient Inference through sparsity and quantization - Part 2/2

Author: Manuel Pasieka June 25, 2024 Duration: 46:38

Austrian Artificial Intelligence Podcast

Technology

Hello and welcome back to the AAIP

This is the second part of my interview with Eldar Kurtic and his research on how to optimiz inference of deep neural networks.

In the first part of the interview, we focused on sparsity and how high unstructured sparsity can be achieved without loosing model accuracy on CPU's and in part on GPU's.

In this second part of the interview, we are going to focus on quantization. Quantization tries to reduce model size by finding ways to represent the model in numeric representations with less precision while retaining model performance. This means that a model that for example has been trained in a standard 32bit floating point representation is during post training quantization converted to a representation that is only using 8 bits. Reducing the model size to one forth.

We will discuss how current quantization method can be applied to quantize model weights down to 4 bits while retaining most of the models performance and why doing so with the models activation is much more tricky.

Eldar will explain how current GPU architectures, create two different type of bottlenecks. Memory bound and compute bound scenarios. Where in the case of memory bound situations, the model size causes most of the inference time to be spend in transferring model weights. Exactly in these situations, quantization has its biggest impact and reducing the models size can accelerate inference.

Enjoy.

## AAIP Community

Join our discord server and ask guest directly or discuss related topics with the community.

https://discord.gg/5Pj446VKNU

### References

Eldar Kurtic: https://www.linkedin.com/in/eldar-kurti%C4%87-77963b160/

Neural Magic: https://neuralmagic.com/

IST Austria Alistarh Group: https://ist.ac.at/en/research/alistarh-group/

Austrian Artificial Intelligence Podcast

Hosted by Manuel Pasieka, the Austrian Artificial Intelligence Podcast offers a grounded, local perspective on a global phenomenon. Instead of abstract theorizing, each conversation focuses on the tangible impact and practical applications of AI within Austria's unique ecosystem. You'll hear from a diverse range of guests-researchers, entrepreneurs, policymakers, and creatives-who are actively shaping this landscape, discussing both the remarkable opportunities and the nuanced challenges specific to the region. The discussions delve into how these technologies are being integrated into Austrian industry, academia, and society, moving beyond hype to examine real-world implementation and ethical considerations. This podcast serves as an essential audio forum for anyone in Austria, or with an interest in the European tech scene, looking to understand how artificial intelligence is evolving right here. It’s about the people behind the algorithms and the local stories within a global revolution. For those engaged with the content, questions and suggestions are always welcome at the provided email address.

Author: Manuel Pasieka Language: English Episodes: 73

Official website RSS

Austrian Artificial Intelligence Podcast

Podcast Episodes

[not-audio_url]

[/not-audio_url]

21. Yasin Ghafourian - TU Wien & RSA : Improving information retrieval systems by modelling a users knowledge gap

04.02.2022

Duration: 1:08:13

# Summary Yasin is a first year PhD Student working on improving information retrieval systems as part the European DOSSIER project. Where he is investigating new ways to improve the relevance of search results presented…

[not-audio_url]

[/not-audio_url]

20. Carina Zehetmaier - Taxtastic: Income tax declaration with AI and the responsible use of customer data

14.01.2022

Duration: 1:04:37

# Intro Carina Zehetmaier, CEO and Co-Founder of Taxtastic an Austrian AI Startup that develops small business and consumer solutions for income tax declaration. We cover multiple topics on today's show. Starting out wit…

[not-audio_url]

[/not-audio_url]

19. Michael Outar - SO Digital Recruitment: Looking back at the Austrian AI Job market in 2021

28.12.2021

Duration: 1:02:45

# Intro Michael is the co-founder and director of SO Digital Recruitment focusing on data positions in the DACH region, with a special focus on Austria. During the Interview, Michael is looking back at 2021 from a recrui…

[not-audio_url]

[/not-audio_url]

18. Resul Akay - Quantics.io: Forecasting and predictive modeling in supply chain optimisation

14.12.2021

Duration: 1:06:21

# Intro Resul Akay is the Chief Data Scientist at Quantics.io, where he is developing a platform to enable predictive modeling for supply chain optimization. Enabling customers from different industries and different ste…

[not-audio_url]

[/not-audio_url]

17. Noah Weber - Celeris Therapeutics: MLOps and Degrader Molecule Design

27.11.2021

Duration: 1:06:39

# Intro Noah Weber is the Chief Technology offer at Celeris Therapeutics. Celeris is discovering candidates for targeted protein degradation in silico using Artificial Intelligence, to accelerate drug discovery and find…

[not-audio_url]

[/not-audio_url]

16. Taylor Peer - Cortical.io: Semantic Folding and Business Solutions

10.11.2021

Duration: 40:02

Taylor Peer is currently working as the Director of Data Science at Cortical.io, a Viennese AI company developing business solutions for information extraction and processing using Natural Language Understanding. On the…

[not-audio_url]

[/not-audio_url]

15. Jelena Milosevic - Mondi Group: Machine Learning on mobile and embedded devices

17.10.2021

Duration: 1:06:22

# Intro Jelena Milosevic is currently working as a Data Scientist focusing on the application of machine learning in embedded devices in industrial settings. On the show we are going to talk about the challenges in devel…

[not-audio_url]

[/not-audio_url]

14. Marco Mondelli - IST: Getting to the bottom of gradient descent methods

27.09.2021

Duration: 1:02:55

# Intro Marco Mondelli is a group leader at the IST Austrian, focusing on theoretical machine learning and in particular on properties and behaviour of gradient descent methods when used to train overparameterized deep n…

[not-audio_url]

[/not-audio_url]

13. Julia Neidhardt - TU Wien: On social network analysis and digital humanism

06.08.2021

Duration: 1:04:54

Julia is a researcher in the E-Commerce Group at the TU Vienna, focusing on the analysis of social networks. In the first part of the interview, Julia will describe the motivations and possibilities of performing network…

[not-audio_url]

[/not-audio_url]

12. Rania Wazir: On AI4Good and how to shape future AI regulations

22.07.2021

Duration: 51:24

Summary Rania has a background in theoretical mathematics and has focused her work in recent years as a Data Scientists on natural language understanding and social media monitoring. Today on the show she will share her…