Meta Media Foundation Industry Workshop

Background

Time: Wednesday, 16 September 2026, 9:00 AM – 10:30 AM

Organizers:
– Dr. Ioannis Katsavounidis, Research Scientist, Video Infrastructure, Meta
– Dr. Ryan Lei, Video Codec Specialist, Video Infrastructure, Meta

Swag and light refreshments will be available for workshop attendees immediately following the event. 


Agenda

  • Short Introduction about Meta’s Media Foundation Team and the Workshop
  • SVT-AV1 for Every Media Workload at Meta
    Over the past years, we have invested in maturing SVT-AV1 to meet the diverse encoding requirements across Meta — from video-on-demand and still image compression (AVIF) to real-time communication and on-device upload. This talk summarizes the key optimizations, quality improvements, and integration work that brought SVT-AV1 to production readiness for each of these use cases, and shares results and lessons learned along the way.
  • Adopting AVIF at Scale: 2 Years Later
    Optimizing the image experience across Meta’s Family of Apps comes with challenges that are hard to find anywhere else. Every day, we handle billions of image uploads and trillions of image download requests on our CDN. This means that image compression is more important to us than ever in the world where visual quality matters but also where every millisecond of latency counts. In this talk we’re going to share the latest updates on adopting AVIF in our Image Infrastructure, deep dive into codec performance aspects that matter in different parts of our image pipelines and present challenges we encounter with image quality understanding between formats.
  • Scaling Visual Metadata and Metrics: Powering the Video Quality Flywheel
    Meta processes billions of user-generated videos, and our pipelines generate a variety of metrics and metadata to optimize encoding settings and compression efficiency. We have been investing in producing these quality metrics and visual metadata — shot boundaries, borders, motion scores — as first-class signals. This talk describes how we do it at scale, and the diverse use cases that these signals enable. We begin with what it takes to make a state-of-the-art quality metric cheap enough to run on everything, drawing on recent work optimizing UVQ for scale. We then discuss how visual metadata supports video enhancement model training and inference. Finally, we trace the full loop, where quality metrics and metadata inform every stage of development, optimization, and deployment of HDR enhancement models. The result is a quality flywheel whose benefits compound — new use cases drive the development of new signals, and each new signal benefits every use case that follows.
  • Audio Enhancement, Understanding and Quality Assessment at Scale
    Audio is a critical component of how users experience media. In this talk, we will discuss how a deeper understanding of the audio scene and its content can inform enhancement methods, leading to better downstream experiences. We will also cover how we build these models to operate at scale.
  • DCT Tokenized GFVC with Adaptive AC channels Updates
    Generative Face Video Coding (GFVC) is a promising solution for very low bit rate talking head video communication, with wide applications in business and social media use cases. In this work we developed a light weight learning framework based on blockwise DCT tokenization and token energy profiling for group-wise partial convolution learning. This framework significantly reduces the complexity for high frame rate video deployment, with adaptive AC channel updates that signals critical detail and local info with graceful rate-distortion trade-offs. Initial results demonstrated that this solution significantly reduces complexity of the GFVC while offering coding gains in both objective and subjective loss metrics.

Speakers

Dr. Ioannis Katsavounidis is part of the Media Foundation team, leading technical efforts in improving video quality and quality of experience across all video products at Meta. Before joining Meta, he spent 3.5 years at Netflix, contributing to the development and popularization of VMAF and inventing the Dynamic Optimizer. He was a professor for 8 years at the University of Thessaly’s Electrical and Computer Engineering Department in Greece, teaching video compression, signal processing and information theory. He was one of the cofounders of Cidana, a mobile multimedia software company in Shanghai, China. He was the director of software for advanced video codecs at InterVideo in the early 2000’s and he has also worked for 4 years in high-energy experimental Physics in Italy. He is general  co-chair of the Video Quality Experts Group (VQEG). He is actively involved within the Alliance for Open Media (AOMedia) as co-chair of the software implementation working group (SIWG). He is a Fellow of the IEEE, a member of the SPS Industry Board and a member of the MMSP TC. He has over 150 publications, including 50 patents. His research interests lie in video coding, quality of experience, adaptive streaming, and energy efficient HW/SW multimedia processing.


Dr. Ryan Lei is currently working as a video codec specialist and technical lead in the Media Foundation group at Meta. His focus is on algorithms and architecture for cloud based video processing, transcoding, and delivery at large scale for various Meta products. Ryan Lei is also the co-chair of the Alliance for Open Media (AOM) testing subgroup and is actively contributing to the standardization of AV1 and AV2. Before joining Meta, Ryan worked at Intel as a principal engineer and codec architect. He worked on algorithm implementation and architecture definition for multiple generations of hardware based video codecs, such as AVC, VP9, HEVC and AV1. Before joining Intel, Ryan worked at ATI handheld department, where he implemented embedded software for hardware encoder/decoder in mobile SoCs. Ryan received his Ph.D. in Computer Science from the University of Ottawa. His research interests include image/video processing, compression, adaptive streaming and parallel computing. He has (co-) authored over 50 publications, including 17 patents.


Dr. Kaustubh Kalgaonkar is a Research Scientist at Meta, where he works on audio understanding, enhancement, and quality prediction across Meta’s family of apps and devices. Before joining Meta, he spent over seven years at Microsoft working on automatic speech recognition, acoustic model adaptation, and on-device wake word detection. Earlier, at Mitsubishi Electric Research Laboratories, he developed ultrasonic Doppler sensing methods for voice activity detection, speaker recognition, and gesture-based interfaces. Kaustubh received his Ph.D. in Electrical and Computer Engineering from the Georgia Institute of Technology. He co-organized the CHiME-8 MMCSG Challenge on multi-modal conversations in smart glasses and contributed to the Speech Accessibility Project. He has (co-) authored over 40 publications, including 9 patents. His research interests include audio understanding, neural audio coding, and audio quality assessment.


Hassene Tmar is a Technical Program Manager at Meta supporting the Media Foundation team where he leads the video codec strategy. He has driven the productization of AV1 across Meta’s video and image delivery platforms, and manages the open-source SVT-AV1 encoder project. He also leads Meta’s Research Engagement program, fostering collaborations between Meta and academic institutions to advance video compression at scale. Hassene serves as Software Coordinator within the Software Implementation Working Group (SIWG) at the Alliance for Open Media (AOMedia). Prior to Meta, Hassene was an Engineering Manager at Intel, where he led the open-sourcing of SVT-HEVC, SVT-VP9, and SVT-AV1. At Meta, he also leads the media acceleration program, including Meta’s Scalable Video Processor (MSVP).


Dr. Shankar Regunathan is a technical lead in the Video Infrastructure group at Meta, where he leads visual quality across Meta’s family of apps. His focus is on video quality metrics and their deployment at scale, encoding optimization for compute efficiency across VOD/Live, and AI-based enhancement. Before Meta, he spent over a decade at Microsoft, contributing to the VC-1 video codec (SMPTE 421M), one of the mandatory codecs for Blu-ray Disc, and to HD Photo, later standardized as JPEG XR (ISO/IEC 29199-2), as well as to H.264/SVC standardization. Shankar received his Ph.D. in Electrical Engineering from the University of California, Santa Barbara, where he worked on scalable audio/video compression, error-resilient image/video coding, and information theory. He received the IEEE Signal Processing Society Best Paper Award in 2004 and 2007. He has authored over 120 publications, including 70 patents. His research interests include video quality assessment, encoding optimization, video enhancement, and compute-efficient media processing at scale across software and hardware encoders.


Zuzanna Mroczek is a software engineer and tech lead in the Images Infrastructure team at Meta in London. Her focus is on full-stack architecture for measuring and improving perceptual image quality and overall quality of image experience at Meta’s scale. This includes work on understanding quality across common image distortions, evaluation of modern codecs and leading the rollout of first HDR photo experiences in Meta apps. Prior to joining Meta as a full-time employee in 2017 and starting work on image quality, she received her Master’s degree in Computer Science at Faculty of Mathematics, Informatics, and Mechanics of University of Warsaw. During her studies she also gained experience interning at Facebook (now Meta), Microsoft and Google across SWE and SRE teams.


Erik André is a software engineer and tech lead at Meta. He graduated with a Master of Science in Computer Engineering from Lund University, Sweden. After initially working on embedded systems and virtual machines at Sony Mobile, Erik moved on to Meta, working on image processing on mobile and backend systems, in their Images Infrastructure team based in London. This work involved evaluation and productionization of new codecs and image formats such as AVIF and JPEG XL.


Dr. Zhu Li is a professor with the Dept of Computer Science & Electrical Engineering, University of Missouri, Kansas City (UMKC), and the director of NSF I/UCRC Center for Big Learning (CBL) at UMKC. He received his PhD in Electrical & Computer Engineering from Northwestern University in 2004. He was the AFRL summer faculty at the UAV Research Center, US Air Force Academy (USAFA), 2016-18, 2020-24. He was Senior Staff Researcher with the Samsung Research America’s Multimedia Standards Research Lab in Richardson, TX, 2012-2015, Senior Staff Researcher with FutureWei (Huawei) Technology’s Media Lab in Bridgewater, NJ, 2010~2012, Assistant Professor with the Dept of Computing, the Hong Kong Polytechnic University from 2008 to 2010, and a Principal Staff Research Engineer with the Multimedia Research Lab (MRL), Motorola Labs, from 2000 to 2008. His research interests include point cloud and light field compression, graph signal processing and deep learning in the next gen visual compression, remote sensing, image processing and understanding. He has 50+ issued US  patents, 200+ publications in book chapters, journals, and conferences in these areas. He is an IEEE senior member, Associate Editor-in-Chief (2020~23) and Senior Area Editor (2024~) for IEEE Trans on Circuits & System for Video Tech,  Associate Editor for IEEE Trans on Image Processing(2020~), IEEE Trans.on Multimedia (2015-18), IEEE Trans on Circuits & System for Video Technology(2016-19). His team won the AFRL sponsored Perception Beyond Visual Spectrum (PBVS) grand challenge at CVPR 2023 on SAR image recognition, and thermo image super-resolution in 2024. He also received the Best Paper Award at IEEE Int’l Conf on Multimedia & Expo (ICME), Toronto, 2006, and IEEE Int’l Conf on Image Processing (ICIP), San Antonio, 2007.