57. Eldar Kurtic - Efficient Inference through sparsity and quantization - Part 2/2

Author: Manuel Pasieka June 25, 2024 Duration: 46:38

Austrian Artificial Intelligence Podcast

Technology

Hello and welcome back to the AAIP

This is the second part of my interview with Eldar Kurtic and his research on how to optimiz inference of deep neural networks.

In the first part of the interview, we focused on sparsity and how high unstructured sparsity can be achieved without loosing model accuracy on CPU's and in part on GPU's.

In this second part of the interview, we are going to focus on quantization. Quantization tries to reduce model size by finding ways to represent the model in numeric representations with less precision while retaining model performance. This means that a model that for example has been trained in a standard 32bit floating point representation is during post training quantization converted to a representation that is only using 8 bits. Reducing the model size to one forth.

We will discuss how current quantization method can be applied to quantize model weights down to 4 bits while retaining most of the models performance and why doing so with the models activation is much more tricky.

Eldar will explain how current GPU architectures, create two different type of bottlenecks. Memory bound and compute bound scenarios. Where in the case of memory bound situations, the model size causes most of the inference time to be spend in transferring model weights. Exactly in these situations, quantization has its biggest impact and reducing the models size can accelerate inference.

Enjoy.

## AAIP Community

Join our discord server and ask guest directly or discuss related topics with the community.

https://discord.gg/5Pj446VKNU

### References

Eldar Kurtic: https://www.linkedin.com/in/eldar-kurti%C4%87-77963b160/

Neural Magic: https://neuralmagic.com/

IST Austria Alistarh Group: https://ist.ac.at/en/research/alistarh-group/

Austrian Artificial Intelligence Podcast

Hosted by Manuel Pasieka, the Austrian Artificial Intelligence Podcast offers a grounded, local perspective on a global phenomenon. Instead of abstract theorizing, each conversation focuses on the tangible impact and practical applications of AI within Austria's unique ecosystem. You'll hear from a diverse range of guests-researchers, entrepreneurs, policymakers, and creatives-who are actively shaping this landscape, discussing both the remarkable opportunities and the nuanced challenges specific to the region. The discussions delve into how these technologies are being integrated into Austrian industry, academia, and society, moving beyond hype to examine real-world implementation and ethical considerations. This podcast serves as an essential audio forum for anyone in Austria, or with an interest in the European tech scene, looking to understand how artificial intelligence is evolving right here. It’s about the people behind the algorithms and the local stories within a global revolution. For those engaged with the content, questions and suggestions are always welcome at the provided email address.

Author: Manuel Pasieka Language: English Episodes: 73

Official website RSS

Austrian Artificial Intelligence Podcast

Podcast Episodes

[not-audio_url]

[/not-audio_url]

11. Christoph Götz - ImageBiopsyLabs: On building safety critical AI applications

09.07.2021

Duration: 1:07:12

Summary Christoph is the Chief Operation Officer and one of the co-founders of Image Biopsy Labs. IB Labs, is a Viennese startup that is developing a modular AI platform to accelerate the work of radiologists ensuring hi…

[not-audio_url]

[/not-audio_url]

10. Linda Anderson - Artificial Researcher: On the history and challenges of domain specific text mining

25.06.2021

Duration: 1:35:13

Summary In this episode, Linda Anderson a computational Linguistic by training and founder of the AI startup "Artificial Researcher" is talking about the history and challenges of domain specific text mining. We are disc…

[not-audio_url]

[/not-audio_url]

9. Frank Benda: Teaching ML from programming to company strategy

11.06.2021

Duration: 59:40

In this episode, Frank is sharing his experience in teaching about AI over the years at different institutions to students from all backgrounds, ranging from businesses focused questions about digitalisation to CEO's, or…

[not-audio_url]

[/not-audio_url]

8. Sanja Jovanovic: On AI services in azure and women in AI

28.05.2021

Duration: 49:56

What kind of AI related services does azure cloud offer? What are some common reasons companies move into the cloud, and what are common first mistakes that companies encounter? If those are questions that interest you,…

[not-audio_url]

[/not-audio_url]

7. Tanja Zinkl: On hiring for an AI Startup

15.05.2021

Duration: 45:55

What are the challenges hiring for an AI startups? How do you hire today, but be prepared for continously changing requirements of tomorrow? Or how do you convince the HR department of your talent and motivation? If thos…

[not-audio_url]

[/not-audio_url]

6. Stefan Habenschuss: Building massive simulation environments at blackshark.ai

02.05.2021

Duration: 1:04:31

Have you ever been excited to browse the other side of the world on google earth? Just to be disappointed by the poor quality of the 3d reconstructions, and the limited level of detail? The Graz based AI company blacksha…

[not-audio_url]

[/not-audio_url]

5. Lukas Fischer: On the Software Competence Center Hagenberg (SCCH)

17.04.2021

Duration: 1:05:44

The field of AI and Machine learning is moving with an ever increasing pace, and for most small and medium companies that cannot afford their own research labs, keeping up new scientific papers and software frameworks is…

[not-audio_url]

[/not-audio_url]

4. Raphael Mitsch: AI consultancy and reinforcement learning in the real world

01.04.2021

Duration: 1:03:14

On the show today, I have the pleasure to talk to Raphael Mitsch, currently working as a Senior Machine Learning Engineer at the Austrian AI consultancy enliteAI. Are you curious to learn, what AI consultancy looks like…

[not-audio_url]

[/not-audio_url]

3. Jillian Augustine: Making the career jump from academia to industry and the intrinsic value of inclusiveness

10.03.2021

Duration: 45:18

On the show today, I have the pleasure to talk to Jillian Augustine. Jillian is currently working as a Data Scientist in the Digital Excellence Team of the international paper and packaging manufacturer Mondi, but has re…

[not-audio_url]

[/not-audio_url]

2. Frank Fichtenmueller: On building AI Teams, business relevant skills and the future of Data Science

21.02.2021

Duration: 59:00

Summary Today on the show, we have Frank Fichtenmueller. Frank is currently working as a Senior Solution architect, at Crayon's AI Center of Excellence; here in Vienna, where he is developing distributed systems at scale…