Friday, July 31, 2026

MPEG news: a report from the 155th meeting

This version of the blog post is also available at ACM SIGMM Records

The 155th MPEG meeting took place in Geneva, Switzerland, from July 13 to 17, 2026. The official MPEG press release can be found here. This report highlights key outcomes from the meeting, with a focus on research directions relevant to the ACM SIGMM community:

  • Gaussian Splat Coding Use Cases and Current Status
  • Call for Proposals on Media Authenticity and Provenance Indication
  • Joint Call for Proposals on video compression with capability beyond VVC
  • Joint ITU-T SG 21 and ISO/IEC JTC 1/SC 29 Workshop on Media Streaming Services
  • Exploration on Systems technologies for AI-based media standards (SyfAI)

Gaussian Splat Coding Use Cases and Current Status

The previous MPEG meeting report introduced MPEG’s Gaussian Splat Coding (GSC) exploration, its two tracks (I-3DGS and A-3DGS), and a first set of 27 draft use cases. At the 155th meeting, WG 2 (Technical Requirements) approved an updated document (WG 2 N531) that drops the draft qualifier and adds two use cases.

The additions matter less for their number than their nature: both concern storage and delivery rather than compression. Use case 28 stores a static splat asset in a single file together with a cover image, thumbnail, and audio annotation, so that a capable receiver renders it interactively while a legacy receiver displays the cover image, with no modification of the file. Use case 29 is the temporal counterpart, delivering a dynamic splat track by HTTP adaptive streaming alongside synchronized audio, a pre-rendered video track serving as preview and fallback, and timed text, with representations offered per bitrate, level of detail, spherical harmonics subset, or attribute subset. The application-oriented use cases are unchanged, although the derived requirements are being refactored, notably by separating attribute subset scalability from random access.

GSC remains an exploration activity, but a busy one: 83 input contributions and nine Joint Exploration Experiments re-conducted across WG 2, WG 4 (Video Coding), WG 5 (Joint Video Experts Team, JVET), and WG 7 (3D Graphics and Haptics), with WG 1 (JPEG) now engaged on quality metrics and WG 4 examining Neural Network Coding (NNC) as a compression tool. On the fast track, GS4 (Video-based Point Cloud Compression, V-PCC, Amd. 1) and GS5 (Geometry-based Point Cloud Compression, G-PCC, Amd. 1) progress on reference software and conformance, while the JVET codec track carries five competing proposals. The Call for Proposals (CfP) still has no date, and I-3DGS planning is more mature than A-3DGS. The clearest message from the meeting is that single-frame compression is essentially solved and that the difficulty and the expected gain now lie in dynamic content.

Research aspects: Temporal coding is no longer one open question among many, but the central one, covering inter-prediction for anisotropic primitives, deformation models, and primitive correspondence across frames. Quality assessment remains unresolved: current test conditions score rendered views with peak signal-to-noise ratio (PSNR), structural similarity (SSIM), and learned perceptual image patch similarity (LPIPS) along a pose trace, a fragile proxy for artifacts such as floaters and popping; the involvement of JPEG makes this a good moment to contribute. Use case 29 gives the streaming question a concrete form: when quality varies along several orthogonal axes at once, what is the right abstraction for a representation, and what should an adaptation algorithm optimize?

Call for Proposals on Media Authenticity and Provenance Indication

At the 155th meeting, WG 2 approved a CfP on media authenticity and provenance indication with the MPEG Systems technologies (WG 2 N535), following an exploration phase in WG 3 (Systems) that produced requirements, a gap analysis, and several technical proposals (WG 3 N1842, attached to the call). The scope is the system layer: carriage and signaling of metadata that allows a receiver to verify that a media asset comes from a trusted producer and to convey provenance information, rather than the definition of a provenance language itself.

Three interfaces are addressed, namely (i) elementary streams and non-timed items, (ii) the file format level (ISO Base Media File Format (ISOBMFF) and Common Media Application Format (CMAF)), and (iii) packaged delivery (Dynamic Adaptive Streaming over HTTP (DASH), CMAF, MPEG Media Transport (MMT), and MPEG-2 Transport Stream (MPEG-2 TS)). The use cases span deepfakes and manipulated news media, forgery in insurance claims, surveillance and investigations, labeling of AI-generated content, and legitimate modifications such as editing, transcoding, ad insertion, and archival preservation. Detecting whether an asset is fake without embedded data is explicitly out of scope.

A response may be a complete solution or a single tool addressing one or more requirements, and is evaluated on requirement coverage, computational complexity, the bandwidth needed to convey the authenticity information, compatibility and backward compatibility with existing MPEG standards, and extensibility, with self-evaluation tables provided in the annexes. Proposals are submitted as input contributions to the 157th meeting in Brisbane, Australia, by January 13, 2027. Review at that meeting will produce one or more working drafts or new work item proposals, and the preliminary development plan targets Committee Draft (CD) or Committee Draft Amendment (CDAM) at the 158th meeting, Draft International Standard (DIS) or Draft Amendment (DAM) at the 159th, and Final Draft International Standard (FDIS) or Final Draft Amendment (FDAM) in early 2028.

Research aspects: The requirements make this more interesting than a signing exercise. A conventional signature breaks on any bit change, yet the call demands that verification survive transcoding, dropped scalability layers, representation switching, splicing, late binding, and loss of frames or audio, and that the unmodified remainder of a presentation stay verifiable after an authorized edit. Designing verification structures with that granularity and quantifying their overhead per segment against the added latency in live scenarios is an open problem at the intersection of cryptography and streaming systems. A second question is joint verification: establishing that a given audio track was the one the producer intended to accompany a given video, including their synchronization, cannot be achieved by signing each file separately. Finally, interoperability with provenance schemes defined elsewhere, and the question of what a receiver should do with a verification result, leave room for work on usable trust signaling.

Joint Call for Proposals on Video Compression with Capability Beyond VVC

The draft of this call was described in the previous MPEG meeting report, and JVET has now issued the final version (JVET-AQ2021), approved at its 43rd meeting in Geneva in July 2026. The target is video coding technology that significantly exceeds Versatile Video Coding (VVC) in compression, implementability, applicability across content types, and features such as latency, robustness, and scalability, benchmarked against the VVC Main 10 profile. Four test cases are defined: one for improved compression without runtime limits, and three in which the aggregate encoder run time is constrained to 5x, 1x, and 0.2x that of the VVC Test Model (VTM) anchor. All are evaluated over seven categories covering (i) standard dynamic range (SDR) random access at ultra-high definition (UHD) and (ii) at high definition (HD), (iii) SDR low delay HD, (iv) high dynamic range (HDR) with perceptual quantizer (PQ) and (v) with hybrid log-gamma (HLG) transfer functions, (vi) gaming, and (vii) user-generated content. A separate track invites technology offering additional functionality, together with proposals on how its benefit should be assessed.

Formal subjective testing uses degradation category rating, with objective results reported as PSNR, multiscale SSIM (MS-SSIM), and, for PQ content, weighted PSNR. The schedule is tight: anchors have been available since May 2026, registration runs from August 1 to September 1, 2026, and the main package of bitstreams, reconstructed sequences, and binaries must reach the test coordinator on physical media by October 26, 2026. Subjective assessment then runs until late December, blind cross-checking by other proponents is mandatory, and proposals are evaluated at the 45th JVET meeting in January 2027, with an initial test model selected during 2027 and the standard targeted for October 2029. Participation is not free: up to EUR 20k per test case is charged to cover the hiring of test subjects.

Research aspects: Two design choices in the call are worth attention. First, a supplemental set of sequences is disclosed to proponents only after the decoder binaries have been submitted, and results on it are due six weeks later, which turns the call into a held-out generalization test. Read together with the ban on training on test sequences and the obligation to disclose training material, this makes out-of-domain behavior of learned coding tools measurable at scale, and reporting it well is a contribution in itself. Second, run time is aggregated as the sum over threads, which measures total compute rather than latency and therefore reads very differently for a massively parallel or GPU-resident design than for a sequential one. How to characterize the rate, distortion, and complexity trade-off fairly across such architectures, and how to value functionality such as scalability or error resilience against a plain bitrate gain, are open questions the call poses rather than answers.

Joint ITU-T SG 21 and ISO/IEC JTC 1/SC 29 Workshop on Media Streaming Services

On July 14, 2026, ITU-T SG 21 and MPEG Systems held a joint half-day workshop in Geneva, collocated with their meetings, on “Media Streaming Service, What’s next” (call for presentations, program). The premise was that after the transitions from analog to digital, enabled by MPEG-2 Systems, and from broadcast to over the top (OTT), enabled by MPEG-DASH, integrated networks, edge computing, and AI are driving another shift, and that both bodies wanted industry views before committing to new work. On challenges, Netflix spoke on the shortcomings and evolution of ISOBMFF, Bitmovin on where streaming trends meet the container, ETRI on future streaming services, and Huawei on ultra-low latency communication and streaming. On opportunities, 5G-MAG addressed standards and open source, DVB heterogeneous networks, and 3GPP SA4 delivery beyond OTT. The closing session was an open mic with questions and answers from both speakers and the audience.

The wrap-up set the existing systems standards, namely MPEG-2 TS, ISOBMFF, DASH, and CMAF as well as the volumetric, scene, and augmented reality (AR) formats, plus ongoing work on authenticity, SyfAI (see below), common metadata, and Gaussian splats, against what industry actually asked for: low latency, low overhead, AI-driven media, and open source software. The resulting agenda is an exploration of a better or new container format, an analysis of overhead, processing-friendly metadata, including JSON and ontology issues, WebCodecs integration, and ingest. Just as notable are the conclusions about process: more open access and industry involvement, faster turnaround, possibly a new home or outlet for this work, and software with interoperability testing from day one rather than reference software at the end. Three ad hoc groups were set up: (i) on MP4, DASH, and file formats, reviewing ISOBMFF against emerging transports and the delivery of AI input and output data; (ii) on vision, analyzing bottlenecks and requirements beyond ISOBMFF; and (iii) on working methods.

Research aspects: Two of these items are directly researchable. Container overhead is asserted more often than measured, and as segments shrink toward frame level for low latency, the ratio of container to media bytes grows; a careful comparison across ISOBMFF, CMAF, and object-based transports at equal latency would inform the exploration rather than follow it. Media over QUIC (MOQ) is the bigger change, since it replaces the segment with the object as the unit of delivery and pulls streaming and real-time communication into a single design space, which reopens rate adaptation, interaction with congestion control, caching, and relay behavior, and the question of what a container still contributes when the transport itself frames media. Carrying AI input and output data alongside media, finally, links this work to coding for machines and to semantic streaming.

Exploration on Systems Technologies for AI-based Media Standards (SyfAI)

Several MPEG coding standards now put a neural network inside the decoder, among them video and feature coding for machines, AI-based point cloud coding, and neural network coding. The previous MPEG meeting report described that coding side through the MPEG-AI vision document. SyfAI, an MPEG Systems exploration started at the 153rd MPEG meeting, asks the complementary and much less glamorous question: what does the surrounding infrastructure have to do so that such content can actually be stored, delivered, and played back interoperably? The underlying shift is that a media file has traditionally been self-contained, meaning that a conforming decoder and the bitstream are sufficient. Once the decoder depends on a trained model that may be selected, delivered, or updated separately, that assumption no longer holds, and the system layer has to say how a player learns which model it needs, how that model reaches it, and how both sides can be sure they are using the same one. Work so far has advanced on the most concrete piece, namely storing video coding for machines content in the MPEG file format, while proposals for carrying compressed neural networks were sent back for clearer use cases. The more interesting development is a new thread on an AI update framework, opened jointly with WG 7, which collects five topics: (i) a format for AI parameter data, (ii) a manifest for updating parameters, (iii) a repository to serve them, (iv) integrity checking, and (v) bit exactness together with conformance.

Research aspects: Each of those topics is a research problem in its own right. Distributing and updating models alongside media turns into a delivery question that looks familiar but has not been studied: when to fetch an update, how to cache and version models, what a manifest must express about capability and compatibility, and how to treat weights as a second class of asset next to the media. Conformance is harder, because a standard normally guarantees that every decoder produces identical output, whereas neural inference varies with library and hardware, so deciding what conformance means and how to test it remains open and extends the reproducibility concerns already visible in MPEG-AI. A repository of updatable weights is also an attack surface, which is why this work sits naturally beside the media authenticity call discussed above: signing and verifying a model raises the same questions one level down from signing content. And because the choice of where a model lives affects how quickly playback can start and how smoothly a player can switch, the apparently dry container questions have measurable consequences for the quality of experience (QoE).

Concluding Remarks

Taken together, the five items show MPEG engaging early with technologies whose momentum comes from outside the committee. With Gaussian splatting, it is compressing a representation that already has a renderer and a user base, which is a better starting point than earlier attempts at three-dimensional video had, even if wider adoption will depend on capture and display ecosystems as much as on coding efficiency. The authenticity call is scoped pragmatically at establishing origin, which is achievable in the near term and provides the system-level plumbing that regulation and industry are asking for. The call for video coding beyond VVC visibly reflects deployment experience, with runtime-constrained test cases treating practicality as a first-class criterion, while questions of licensing and market uptake sit largely outside MPEG itself. In systems, the joint workshop and the new ad hoc groups show a willingness to re-examine long-standing assumptions as transports such as MOQ emerge, and SyfAI stakes out the system-level questions of the AI era before they become urgent. For our community, the result is an unusually rich set of open problems in evaluation and methodology.

The 156th MPEG meeting will be held in Hangzhou, China, from October 19 to 23, 2026. Click here for more information about MPEG meetings and ongoing developments.

Wednesday, July 29, 2026

ACM Mile-High Video 2027: Call for Contributions

ACM Mile-High Video 2027

February 22-25, 2027

Freyer-Newman Center, Denver Botanic Gardens, Denver, CO

https://www.mile-high.video/

ACM MHV'27 is a flagship, industry-oriented conference in the area of video technologies, held in Denver, CO since 2016. We welcome talks from both industry and academia which share real-world problems and solutions, as well as novel approaches and innovations from content production to consumption.

ACM MHV'27 talks and papers are solicited in the following areas:

Content production, encoding, and packaging

  • Encoding for broadcast, mobile, and OTT, including edge, network, and cloud-based coding
  • Deployment of AV2, VVC, and next-generation video/image codecs: encoder maturity, decode economics, carriage, and ladder impact
  • New and emerging audio, image, and video codecs (including point cloud coding, Gaussian splat coding, light field coding, learned coding, etc.)
  • Display-aware encoding and quality optimization
  • Storage applications for video processing and streaming
  • Encoding ladder construction: per-title, context-aware, and device-aware
  • Quality metrics, assessment models, and user experience studies
  • Accessibility: captioning, timed text, and access services
  • HDR and color pipelines

Applications of AI

  • GenAI-based video content generation, incl. real-time generative and diffusion-based transformation
  • End-to-end and hybrid AI-based image and video compression
  • Learned tools in conventional pipelines: neural post-filters, pre-processing, and enhancement layers
  • AI-assisted indexing, search, and metadata extraction
  • Agentic AI for operations, quality assurance, and observability
  • Media pipelines for AI workloads: model delivery, decoder-side compute, cost, and energy

Video workflows

  • Virtualized headends, cloud-based and open media facility infrastructures
  • Redundancy and resilience in content origination
  • Ingest protocols, incl. multitrack and multi-format contribution
  • Workflows for UGC and creator-scale libraries
  • Computational imaging, camera processing, capture systems, and camera-to-cloud workflows
  • Ad insertion: server-guided ad insertion, interstitials, and non-linear and overlay ad formats

Content delivery and security

  • Transport protocols, incl. HTTP/2 and HTTP/3 media delivery, and new delivery paradigms
  • Media over QUIC: relay architecture, packaging, interoperability, and deployment experience
  • Content steering, multi-CDN, and delivery at extreme concurrency
  • Player-server signaling and data interfaces (CMCD, CMSD), and analytics
  • Protection for OTT distribution and anti-piracy tools
  • Cybersecurity, privacy, and infrastructure resilience

Content authenticity and provenance

  • C2PA in adaptive streaming: workflows and implementations 
  • Security and attacks on C2PA
  • Content watermarking

Streaming technologies

  • Adaptive streaming and transcoding
  • Low latency
  • Player, playback, and QoE developments
  • Multiview experiences: synchronized playback of a primary stream with picture-in-picture companion streams, composition, and decode budgets
  • Device and platform fragmentation: capability negotiation and the CE long tail
  • Protocol and Web API improvements, and client-side media pipelines

Industry trends

  • Scalable and multi-view video coding deployments
  • Video and audio coding for machines
  • Cloud gaming and game streaming
  • Cost and energy: measurement, energy awareness, and the quality/cost frontier
  • Edge computing and hardware for encoding, storage, and distribution

Live production, audio, and media infrastructure

  • Live and cloud production, including sports, replay, and distributed collaboration
  • IP-based media facilities, synchronization, timing, and interoperability
  • Camera-to-cloud contribution and virtual-production workflows
  • Immersive, spatial, object-based, and personalized audio
  • Media storage, archiving, asset management, and content lifecycle optimization

Standards, interoperability, and open source

  • New and developing standards in the media and delivery space
  • Interoperability guidelines, conformance testing, and implementation experience
  • Open-source media infrastructure: maintenance, governance, and migration

Prospective speakers are invited to submit a 300+ word abstract for a confidential peer review by the TPC for relevance, timeliness, technical correctness, and value to the MHV audience, which comes primarily from the industry. Proposals explaining the underlying technologies used in commercial products or services will be considered; proposals promoting company products or services will be rejected.

Accepted abstracts will be presented at MHV'27 as 15-20 minute talks in the conference’s main stage track. Talks are recorded and published at the conference YouTube channel.

Speakers are invited to submit an optional full-length paper (up to 8 pages) or an optional short paper (up to 2 pages) to be published in the conference proceedings and available in the ACM Digital Library for long-term visibility and citability. Note that publication is strictly opt-in, and not required.

Authors of rejected abstracts remain eligible to submit a full-length paper. These papers must be original work (i.e., not published previously in a journal or conference), and will be separately peer-reviewed by the TPC.

All papers will be published in the conference proceedings within the ACM Digital Library. All prospective ACM publication authors are subject to all ACM Publications Policies, including ACM’s https://www.acm.org/publications/policies/new-acm-policy-on-authorship

How to Submit an Extended Abstract

Important Dates

  • Abstract submission deadline: Oct. 9, 2026 AoE 
  • Notification of abstract acceptance: Nov. 20, 2026 AoE
  • Optional paper submission deadline: Dec. 11, 2026 AoE 
  • Notification of full-length paper acceptance: Jan. 15, 2027 AoE
  • Camera-ready (final) paper submission: Feb. 5, 2027 AoE

General Chairs

  • Alex Giladi (Netflix, USA)
  • Ali C. Begen (Ozyegin University, Türkiye)
  • Victoria Tuzova (Elecard, USA)

Program Chairs

  • Christian Timmerer (AAU, Austria)
  • Dan Grois (AnyAI, Israel)
  • Yuriy Reznik (MIT, USA)

Wednesday, May 27, 2026

Call for presentations for ITU-T SG 21 and ISO/IEC JTC 1/SC 29 joint workshop on Media Streaming Services - What’s next

As the landscape of multimedia streaming continues to evolve with the advent of next-generation integrated networks (mobile, broadband, and cable) and the emergence of new media streaming protocols, the industry stands on the verge of the next round of transformative opportunities after the shift from analog to digital and from the legacy media delivery to OTT delivery, driven by the tight integration of advanced networks, edge computing, and media services, and integration of AI technologies. Simultaneously, the rapid integration of AI across various sectors of the multimedia ecosystem introduces unprecedented challenges and opportunities, reshaping the industry's trajectory.

The ITU-T Study Group 21 and MPEG Systems Working Group (ISO/IEC JTC 1/SC 29/WG 03) have consistently delivered industry-recognized standards for multimedia service systems over the past several decades. The MPEG-2 Systems, a.k.a. ITU-T H.222.0, which has been jointly developed by two organization has changed multimedia service from analog to digital. The MPEG-DASH standards played key role in the deployment of OTT services. Understanding the coming waves, both organizations aim to engage with industry experts to explore these challenges and opportunities, ensuring that future standardization efforts are strategically aligned with the most critical and impactful targets.

To facilitate deep engagement with the industry the ITU-T SG21 and MPEG Systems Working Group jointly organize a workshop on Tuesday, July 14 14:00 – 18:00 in Geneva, Switzerland during our co-located meetings. The primary goal of the workshop is to hear from the industry experts and get better understanding the challenges and the opportunities for future standards in the area of multimedia streaming services. For this purpose, both organizations invite the industry experts to share their views with us at the workshop. The following list shows some example topics.

  • Next generation media delivery over heterogeneous networks
  • Consumer-centric delivery and content repurposing
  • Smart home and in-home media distribution
  • End-to-end ultra-low latency streaming
  • Network-friendly media containers
  • Media container improvements
  • Volumetric media service
  • Media delivery for spatial computing
  • Energy consumption/sustainability/accessibility

The workshop will be organized per sessions with onsite oral presentations preferentially (remote presentation will be supported, if needed). Please send a presentation proposal by 13 June 2026 including title, author(s), and an abstract of 500 words by email to the following persons (ITU-T SG21 Counselor and the convenor of MPEG Systems WG):

  • Stefano Polidori, tsbsg21@itu.int
  • Youngkwon Lim, young.L@samsung.com

The final detailed program will be made available by 27 June 2026 through a dedicated webpage, which will be linked from ITU-T Study Group 21 main website: https://itu.int/go/tsg21 and the corresponding MPEG website: https://www.mpeg.org/. Information about acceptance/rejection of the contributions will be conveyed to proponents prior to that date.

The dedicated workshop webpage will also include the specific and separated registration, as well as the logistics details to attend the workshop.

The detailed logistics information about the ITU-T SG21 upcoming meeting is available in TSB Collective Letter 6/21 (https://www.itu.int/md/T25-SG21-COL-0006/en) while the information about the MPEG Systems Working Group can be found at https://155.mpeg-meeting.com/.

Friday, May 15, 2026

MPEG news: a report from the 154th meeting

This version of the blog post is also available at ACM SIGMM Records


The 154th MPEG meeting took place in Santa Eulària, Spain, from April 27 to May 1, 2026. The official MPEG press release can be found here. This report highlights key outcomes from the meeting, with a focus on research directions relevant to the ACM SIGMM community:

  • Exploration on MPEG Gaussian Splat Coding (GSC)
  • Draft Joint Call for Proposals: Video Compression Beyond VVC
  • Energy-aware Streaming in MPEG-DASH
  • MPEG-AI: Vision and Scenarios for Artificial Intelligence in Multimedia
  • MPEG Roadmap

Exploration on MPEG Gaussian Splat Coding (GSC)

The MPEG WG 2 Technical Requirements group — jointly with WG 4 (Video Coding), WG 5 (JVET: Joint Video Coding Team(s) with ITU-T SG 16), and WG 7 (Coding of 3D Graphics and Haptics) — made progress toward standardizing Gaussian Splat Coding (GSC) regarding draft requirements and use cases subject to change. Gaussian splatting, first introduced in a landmark 2023 ACM SIGGRAPH paper by Kerbl et al., represents 3D scenes as collections of anisotropic Gaussian primitives carrying geometry (x, y, z positions) and appearance attributes (opacity, scale, rotation, and spherical harmonics coefficients for view-dependent color), enabling photorealistic novel-view synthesis with real-time rendering. Because raw Gaussian splat data can be extremely large and the ecosystem of proprietary formats (.ply, .splat, .spz, etc.) is fragmented, MPEG has identified a clear need for interoperable, efficient compression standards. Two exploration tracks are currently being pursued: I-3DGS, which operates on Gaussian splats in the well-established “INRIA” format as a symmetric encode/decode pipeline, and A-3DGS, which allows alternative learned representations and training-integrated approaches.

The draft requirements, still evolving, currently cover representation, coding, and system aspects across both tracks, with an additional lightweight profile targeting resource-constrained devices such as mobile phones (Snapdragon 8 Gen 3/Elite) and HMDs (Snapdragon XR Gen2, e.g., Meta Quest 3). Among the coding requirements under consideration are lossy and lossless compression with variable bitrate, spatial and temporal random access, progressive and scalable decoding (quality, Level of Detail (LoD), attribute subsets), and error resilience. Notably, a lightweight profile currently proposes hard complexity constraints (i.e., real-time encode/decode on 2024/2025 mobile hardware, a 2GB runtime memory cap, and at most four concurrent video decoder sessions) reflecting MPEG’s intent to enable a fast-deployment path for interoperable interchange and storage of static Gaussian splat assets. Alongside the requirements, a draft set of 27 use cases has been identified, spanning consumer XR (telepresence, gaming, social media, retail), professional media (movie production, sports broadcasting, immersive journalism), industrial applications (digital twins, Building Information Modeling (BIM), structure inspection, disaster assessment), and emerging hybrid representations such as Gaussian splats attached to deformable meshes for avatar animation and rigging. Several of these use cases are motivating draft requirements around primitive ordering preservation and stable identifier signaling for external metadata associations, though the details of these provisions may still change.

Research aspects: Even at this early draft stage, the direction of MPEG’s GSC work opens a rich set of research opportunities. On the compression side, the dual-track structure raises open questions around rate-distortion-complexity optimization for both geometry-based and video-codec-based pipelines, including temporally coherent coding of dynamic (tracked and non-tracked) Gaussian sequences and attribute-group-aware progressive coding. The QoE angle is equally pressing: no widely accepted perceptual quality metric yet exists for 6DoF Gaussian splat rendering, and the community can contribute splat-artifact-aware metrics, view-consistency measures, and subjective evaluation methodologies. The envisioned lightweight profile points to a need for co-design of decoders and real-time renderers targeting mobile GPU architectures, offering opportunities in GPU-friendly bitstream layouts and LOD-driven streaming. From a systems and networking perspective, the spatial and temporal random-access provisions, combined with the breadth of use cases demanding adaptive streaming to diverse devices (HMDs, phones, TVs, browsers), map naturally onto adaptive bitrate research, ROI- and view-dependent segment delivery, and loss-resilient transmission of splat parameters. Finally, the emerging use cases around hybrid mesh-Gaussian avatars, scene editing, and semantic metadata associations introduce new multimedia content management and interactive media challenges that go well beyond traditional video streaming and are squarely within the scope of ACM SIGMM’s research community.

Draft Joint Call for Proposals: Video Compression Beyond VVC

MPEG’s Joint Video Experts Team (JVET) — operating jointly under ITU-T SG21 and ISO/IEC JTC 1/SC 29 — advanced a draft Joint Call for Proposals (CfP) for a new generation of video compression technology with capabilities that would substantially exceed those of the current Versatile Video Coding (VVC) standard (Rec. ITU-T H.266 | ISO/IEC 23090-3). The final CfP is planned for July 2026, with proposal submissions evaluated at a JVET meeting in January 2027 and a tentative target of a completed standard by October 2029. The overarching goal is to solicit compression technology that significantly improves upon VVC’s Main 10 Profile in terms of rate-distortion performance, encoder/decoder implementability, applicability to diverse content types, and additional features such as low latency, error robustness, and scalability, while explicitly recognizing that practical fast encoding is increasingly important across a growing range of applications.

The draft CfP defines four test cases. The primary test case targets improved compression without runtime constraints, spanning several content categories: SDR random-access at UHD/4K and HD resolutions, SDR low-delay HD (targeting conversational and gaming applications), HDR content under both PQ and HLG transfer functions at UHD, gaming low-delay HD, and user-generated content. Three additional test cases impose encoder runtime constraints relative to the VVC Test Model (VTM) reference encoder, enabling JVET to characterize the compression-versus-speed trade-off across submissions. Formal subjective evaluation will follow the degradation category rating (DCR) methodology per ITU-R BT.500. Importantly, the CfP explicitly addresses neural and learned components: proponents must disclose what training data was used and are prohibited from using any test sequence as training material, and source code (incl. training scripts or parameter derivation procedures) must be made available for accepted technologies entering the core experiments process. The draft notes that specific test sequences and target bitrates may still change before the final CfP is issued.

Research aspects: The runtime-constrained test cases create a natural framework for studying the compression-complexity Pareto frontier for both classical and learned codecs. The inclusion of user-generated content and gaming video as distinct categories invites research into content-adaptive coding tools and perceptual quality metrics tailored to these sources, as does the HDR coverage with its use of weighted PSNR alongside MS-SSIM. The explicit allowance for neural and learned components, with mandatory training data disclosure and source code requirements, signals that JVET anticipates hybrid and end-to-end learned codecs as serious contenders, making codec-agnostic adaptive streaming, QoE modeling for learned video codecs, and large-scale perceptual quality benchmarking timely topics for the ACM SIGMM community.

Energy-aware Streaming in MPEG-DASH

MPEG’s WG 3 (Systems/DASH) is developing a framework for integrating energy-related information into adaptive streaming workflows, currently documented as a Technology under Consideration (TuC) in the DASH specification. The proposed framework treats energy as a first-class design metric alongside QoE, latency, and throughput, and defines an end-to-end approach for assigning, aggregating, and propagating energy consumption data across the entire media delivery chain — from production and encoding through CDN distribution to the client. A key design principle is extensibility: rather than hardcoding specific metrics, the framework proposes a common registry of energy-related metrics (such as energy indices or carbon indices) identified via URNs or 4CC codes, inspired by existing registries like MP4RA and DASH-IF. Energy information may be carried through a variety of existing DASH mechanisms, including MPD descriptors at multiple granularity levels (Adaptation Set, Representation, Segment, Service Location), CMCD/CMSD extensions, metadata tracks, SAND messages, and event streams. A dedicated Energy descriptor in the MPD is proposed, analogous to existing Accessibility descriptors, to expose energy information to clients and applications for representation selection, user exposure, and reporting to back-end servers.

Concept of Energy-aware Streaming in MPEG-DASH.

The April 2026 update reported significant progress on two related fronts. A 5G-MAG workshop co-organized with 3GPP SA4 and Greening of Streaming (March 2026) highlighted growing industry consensus around practical energy measurement, surfacing findings such as the dominant role of device eco-mode settings and content brightness over codec or resolution choices in determining end-device energy consumption, and the challenge of reproducible cloud-based energy measurement. In parallel, 3GPP’s Rel-20 study on media energy consumption exposure (FS_Energy_Ph2_MED) reached 80% completion and is expected to conclude in June 2026, with normative work to follow. Notably, 3GPP’s current draft conclusions focus on generic architectural enablers, specifically a new Energy Information Application Function, while explicitly deferring media-layer and client-driven energy optimization to external bodies such as MPEG, SVTA, and DVB. This positions MPEG-DASH’s manifest-based energy signaling work as the natural venue for maturing the streaming-level mechanisms that 3GPP may later reference.

Research aspects: This work opens several timely directions. Energy-aware ABR algorithm design, i.e., jointly optimizing QoE and energy across representation selection, CDN choice, and client device settings, is a natural extension of the existing adaptive streaming research agenda. The proposed metrics registry and MPD-level signaling create opportunities for dataset construction and benchmarking, building on emerging open datasets such as COCONUT and VEED. The finding that device-side factors (eco-mode, display brightness) dominate energy consumption over codec and bitrate choices challenges some common assumptions and calls for more holistic QoE-energy modeling. Finally, the cross-SDO coordination between MPEG, 3GPP, IETF (GREEN working group), and Greening of Streaming presents opportunities for the ACM SIGMM community to contribute to the design of interoperable, standardized energy reporting APIs for streaming services.

MPEG-AI: Vision and Scenarios for Artificial Intelligence in Multimedia

The first edition of ISO/IEC TR 23888-1 serves as the foundational vision document for the MPEG-AI series (ISO/IEC 23888). The document maps out how AI and neural network technologies interact with multimedia standardization along two complementary axes: (i) AI as a multimedia coding tool (e.g., AI-based video compression, 3D point cloud coding) and (ii) multimedia as input for AI consumption (e.g., video coding optimized for machine vision tasks). Under this umbrella, the document surveys six technical areas. In AI-based video coding, neural network components are explored as hybrid additions to VVC-style codecs, covering in-loop filters, intra prediction, super-resolution via reference picture resampling, and content-adaptive postfilters transmitted via SEI messages using the Neural Network Coding standard (NNC, ISO/IEC 15938-17). In AI-based 3D graphics coding, the focus is on dynamic point clouds for immersive (XR, gaming) and machine-oriented (autonomous navigation, BIM) applications, where sparsity and geometric irregularity pose unique challenges beyond those faced by image/video AI codecs. AI model compression (NNC) addresses the bandwidth-efficient deployment and incremental updating of neural network weights to devices, with use cases ranging from adaptive streaming ABR models to federated learning and postfilter delivery. Video coding for machines (VCM) targets compression optimized for downstream AI tasks such as object detection, tracking, and content moderation, with applications in surveillance, intelligent transportation, smart cities, and industrial inspection. Feature coding for machines (FCM) extends this to split-inference architectures where intermediate feature maps — rather than reconstructed video — are compressed and transmitted between edge devices and servers. Finally, distributed AI media description addresses the interoperable representation and API-level exchange of AI inference results (e.g., bounding boxes, segmentation masks) between networked media analyzers, as specified in the MPEG-IoMT suite.

ISO/IEC TR 23888-1: AI as a multimedia coding tool and multimedia as input for AI consumption.

Research aspects: The hybrid codec paradigm raises open questions around joint optimization of traditional and learned tools and complexity-aware training for mobile targets. The VCM and FCM tracks call for new task-oriented quality metrics capturing machine-task performance as a function of bitrate, an area where the multimedia and computer vision communities can collaborate. The split-inference and feature coding scenarios introduce latency-constrained compression problems for edge-to-cloud pipelines, which naturally connect to adaptive streaming and IoT research. Finally, the reproducibility and bit-exactness challenges highlighted in the document — hardware-dependent inference, non-deterministic training, and the absence of standardized evaluation environments — present an opportunity for the community to develop shared benchmarking infrastructure for learned multimedia codecs.

MPEG Roadmap

MPEG released an updated roadmap at its 154th meeting, reflecting the current status and near-term trajectory of its standardization activities across three broad pillars. Under Media Coding, work nearing completion includes MPEG Immersive Video v.2, Feature Coding for Machines, Solid Point Cloud Coding, and Dynamic Mesh Compression, while longer-horizon efforts cover AI Graphics Compression, Video Coding for Machines, Lenslet video coding, and — directly relevant to this report — both Video-based and Geometry-based Gaussian Splat Coding tracks. Under Systems and Tools, near-term deliverables include DASH v.7, Green metadata v.4, and Carriage of Haptics Data, with CMAF v.4 and File Format (ISOBMFF) v.10 on a slightly longer timeline. The Beyond Media pillar continues to advance genomic data search and biomedical waveform coding (BWC), alongside media authenticity and provenance indication — underscoring MPEG’s expanding scope well beyond traditional audiovisual applications.

MPEG Roadmap as of April 2026.

Research aspects: The roadmap highlights several intersecting research opportunities. The convergence of volumetric and neural representations (i.e., point clouds, dynamic meshes, Gaussian splats, and lenslet video; all progressing in parallel) raises open questions around unified rate-distortion frameworks and cross-format QoE evaluation for 6DoF experiences. The simultaneous progression of Video Coding for Machines and Feature Coding for Machines alongside traditional human-centric codecs calls for research into adaptive pipelines that can serve both human and machine consumers from a shared bitstream. The Green metadata track connects directly to the energy-aware streaming work discussed above, underscoring the need for end-to-end energy modeling that spans codec choice, packaging, delivery, and consumption. Finally, the Beyond Media thread (e.g., particularly genomic data and biomedical waveforms) signals an expanding definition of “multimedia” that the ACM SIGMM community may wish to engage with as compression, retrieval, and QoE methods developed for audiovisual content find applicability in life sciences.

Concluding Remarks

The 154th MPEG meeting in Santa Eularia reflects a standards body in active transition, broadening its scope from traditional audiovisual compression toward a richer landscape that encompasses neural scene representations, AI-native codecs, energy-aware delivery, and even biomedical data. The Gaussian Splat Coding exploration, the next-generation video compression Call for Proposals, the MPEG-AI vision document, and the energy-aware streaming framework each address distinct but interconnected challenges: how to represent, compress, deliver, and consume increasingly complex and diverse media efficiently and sustainably. For the ACM SIGMM community, this meeting offers both a map of where industry standardization is heading and a set of open research problems (i.e., spanning perceptual quality assessment, learned compression, edge inference, green streaming, and immersive media delivery) where academic contributions can meaningfully shape the next generation of multimedia standards.

The 155th MPEG meeting will be held in Geneva, Switzerland, from July 13 to 17, 2026. Click here for more information about MPEG meetings and ongoing developments.

Friday, March 27, 2026

Sustainability in Video Encoding and Streaming

Sustainability in Video Encoding and Streaming:
Energy-Efficient Techniques and Metrics

Workshop on Media Energy Consumption Measurement and Exposure

Presenter: Christian Timmerer (Alpen-Adria-Universität Klagenfurt)

Abstract: The presentation discusses the increasing environmental impact of video streaming and highlights the urgent need for more sustainable approaches across the entire streaming pipeline. Video traffic dominates internet usage and contributes significantly to global greenhouse gas emissions, while the demand for higher quality content continues to drive up computational complexity and energy consumption in encoding, delivery, and playback.

A central insight is that there is a strong trade-off between video quality and energy consumption, where small reductions in quality can lead to substantial energy savings. By introducing energy as an explicit optimization objective, techniques such as content-aware encoding, energy-aware bitrate ladder construction, and real-time optimization for live streaming can significantly reduce energy usage while maintaining nearly the same perceptual quality.

The work also emphasizes the role of adaptive bitrate algorithms that incorporate energy consumption alongside traditional quality and buffer-based metrics. These approaches demonstrate that it is possible to simultaneously improve user experience and reduce energy consumption, indicating that sustainability and performance can be aligned rather than conflicting goals.

To enable such optimizations, the presentation introduces a range of metrics and models, including video complexity measures, quality prediction models, and machine learning-based approaches for estimating encoding and decoding energy as well as CO₂ emissions. These tools support more informed, data-driven decisions across the full streaming workflow from encoding to playback.

Another important theme is end-to-end optimization, where energy efficiency depends on the combined behavior of encoding strategies, bitrate selection, and client-side adaptation. Industry efforts confirm the practical relevance of these approaches and highlight the importance of collaboration and real-world validation.

Despite promising results, several challenges remain, including difficulties in measuring and benchmarking energy consumption, the lack of standardized methodologies, and the limited integration of energy considerations into existing workflows. Overall, the presentation argues that energy consumption should become a first-class optimization target in video streaming systems, similar to established quality metrics, to enable truly sustainable media delivery.

Keywords: sustainable streaming, energy-aware encoding, adaptive bitrate streaming, green multimedia, video compression, bitrate ladder optimization, QoE optimization, energy-quality tradeoff, video complexity analysis, CO2 footprint, energy modeling, machine learning for video, end-to-end optimization, eco-efficient streaming, real-time streaming optimization

Friday, February 20, 2026

MPEG news: a report from the 153rd meeting

This version of the blog post is also available at ACM SIGMM Records



The 153rd MPEG meeting took place online from January 19-23, 2026. The official MPEG press release can be found here. This report highlights key outcomes from the meeting, with a focus on research directions relevant to the ACM SIGMM community:
  • MPEG Roadmap
  • Exploration on MPEG Gaussian Splat Coding (GSC)
  • MPEG Immersive Video 2nd edition (new white paper)

MPEG Roadmap

MPEG released an updated roadmap showing continued convergence of immersive and “beyond video” media with deployment-ready systems work. Near-term priorities include 6DoF experiences (MPEG Immersive Video v2 and 6DoF audio), volumetric representations (dynamic meshes, solid point clouds, LiDAR, and emerging Gaussian splat coding), and “coding for machines,” which treats visual and audio signals as inputs to downstream analytics rather than only for human consumption.

Research aspects: The most promising research opportunities sit at the intersections: renderer and device-aware rate-distortion-complexity optimization for volumetric content; adaptive streaming and packaging evolution (e.g., MPEG-DASH / CMAF) for interactive 6DoF services under tight latency constraints; and cross-cutting themes such as media authenticity and provenance, green and energy metadata, and exploration threads on neural-network-based compression and compression of neural networks that foreshadow AI-native multimedia pipelines.

MPEG Gaussian Splat Coding (GSC)

Gaussian Splat Coding (GSC) is MPEG’s effort to standardize how 3D Gaussian Splatting content, scenes represented as sparse “Gaussian splats” with geometry plus rich attributes (scale and rotation, opacity, and spherical-harmonics appearance for view-dependent rendering), is encoded, decoded, and evaluated so it can be exchanged and rendered consistently across platforms. The main motivation is interoperability for immersive media pipelines: enabling reproducible results, shared benchmarks, and comparable rate-distortion-complexity trade-offs for use cases spanning telepresence and immersive replay to mobile XR and digital twins, while retaining the visual strengths that made 3DGS attractive compared to heavier neural scene representations.

The work remains in an exploration phase, coordinated across ISO/IEC JTC 1/SC 29 groups WG 4 (MPEG Video Coding) and WG 7 (MPEG Coding for 3D Graphics and Haptics) through Joint Exploration Experiments covering datasets and anchors, new coding tools, software (renderer and metrics), and Common Test Conditions (CTC). A notable systems thread is “lightweight GSC” for resource-constrained devices (single-frame, low-latency tracks using geometry-based and video-based pipelines with explicit time and memory targets), alongside an “early deployment” path via amendments to existing MPEG point-cloud codecs to more natively carry Gaussian-splat parameters. In parallel, MPEG is testing whether splat-specific tools can outperform straightforward mappings in quality, bitrate, and compute for real-time and streaming-centric scenarios.

Research aspects: Relevant SIGMM directions include splat-aware compression tools and rate-distortion-complexity optimization (including tracked vs. non-tracked temporal prediction); QoE evaluation for 6DoF navigation (metrics for view and temporal consistency and splat-specific artifacts); decoder and renderer co-design for real-time and mobile lightweight profiles (progressive and LOD-friendly layouts, GPU-friendly decode); and networked delivery problems such as adaptive streaming, ROI and view-dependent transmission, and loss resilience for splat parameters. Additional opportunities include interoperability work on reproducible benchmarking, conformance testing, and practical packaging and signaling for deployment.

MPEG Immersive Video 2nd edition (white paper)

The second edition of MPEG Immersive Video defines an interoperable bitstream and decoding process for efficient 6DoF immersive scene playback, supporting translational and rotational movement with motion parallax to reduce discomfort often associated with pure 3DoF viewing. The second edition primarily extends functionality (without changing the high-level bitstream structure), adding capabilities such as capture-device information, additional projection types, and support for Simple Multi-Plane Image (MPI), alongside tools that better support geometry and attribute handling and depth-related processing.

Architecturally, MIV ingests multiple (unordered) camera views with geometry (depth and occupancy) and attributes (e.g., texture), then reduces inter-view redundancy by extracting patches and packing them into 2D “atlases” that are compressed using conventional video codecs. MIV-specific metadata signals how to reconstruct views from the atlases. The standard is built as an extension of the common Visual Volumetric Video-based Coding (V3C) bitstream framework shared with V-PCC, with profiles that preserve backward compatibility while introducing a new profile for added second-edition functionality and a tailored profile for full-plane MPI delivery.

Research aspects: Key SIGMM topics include systems-efficient 6DoF delivery (better view and patch selection and atlas packing under latency and bandwidth constraints); rate-distortion-complexity-QoE optimization that accounts for decode and render cost (especially on HMD and mobile) and motion-parallax comfort; adaptive delivery strategies (representation ladders, viewport and pose-driven bit allocation, robust packetization and error resilience for atlas video plus metadata); renderer-aware metrics and subjective protocols for multi-view temporal consistency; and deployment-oriented work such as profile and level tuning, codec-group choices (HEVC / VVC), conformance testing, and exploiting second-edition features (capture device info, depth tools, Simple MPI) for more reliable reconstruction and improved user experience.

Concluding Remarks

The meeting outcomes highlight a clear shift toward immersive and AI-enabled media systems where compression, rendering, delivery, and evaluation must be co-designed. These developments offer timely opportunities for the ACM SIGMM community to contribute reproducible benchmarks, perceptual metrics, and end-to-end streaming and systems research that can directly influence emerging standards and deployments.

The 154th MPEG meeting will be held in Santa Eulària, Spain, from April 27 to May 1, 2026. Click here for more information about MPEG meetings and ongoing developments.

Wednesday, February 18, 2026

Professor of Information Systems Engineering (all genders welcome)

Department of Informatics Systems  

Full professorships  | Full time

Application deadline:  22 March 2026

Reference code: 43/02-PERS/26

URL: https://jobs.aau.at/en/job/professor-of-information-systems-engineering-all-genders-welcome/

Announcement

The University of Klagenfurt wants to attract more women for professorships.

We are pleased to announce the following open position at the Department of Informatics Systems, Faculty of Technical Sciences, in compliance with the provisions of § 98 (permanent) or § 98 (fixed-term, max. 6 years) of the Austrian Universities Act:

Professor of Information Systems Engineering (all genders welcome)

This is a full-time position available from 1 October 2027. Depending on the candidate’s academic credentials, the employment contract can be concluded either as a permanent employment contract or as a fixed-term employment contract with the option of a permanent extension. The duration of fixed-term contracts is subject to negotiation.

With approximately 13,000 students, the University of Klagenfurt is a young, vibrant and innovative university, located at the intersection of Alpine and Mediterranean culture in an area that offers exceptionally high quality of life. As a public university pursuant to § 6 of the Austrian Universities Act, it receives federal funding. The university operates under the motto “Beyond Boundaries!”.

In accordance with its key strategic road map, the Development Plan, the university’s primary guiding principles and objectives include the pursuit of scientific excellence regarding the appointment of professors, favourable research conditions, a good faculty-student ratio, and the promotion of the development of early career researchers.

Information Systems Engineering focuses on the design, development, and management of large systems that connect people, data, and technology to support organizational goals. It combines principles of software engineering, data management, business processes, and emerging digital technologies to create solutions that enhance decision-making, optimize operations, and drive innovation.

We welcome applications addressing the engineering of Information Systems, in particular those focusing on designing, modelling, executing, verifying, and optimizing business processes. We are looking for a highly qualified and internationally visible scientist with high engagement in developing and sustaining an ambitious and innovative research and teaching programme. Candidates should also be interested in developing collaborations in the university’s Areas of Research Strength: Digitalisation and Health, Multiple Perspectives in Optimization, Networked and Autonomous Systems and/or the Cluster of Excellence “Bilateral AI”.

Your responsibilities – what awaits you

The duties of the position include:

  • Representing the field of Information Systems Engineering in research and teaching
  • Teaching in relevant degree programmes at Bachelor’s, Master’s, and Doctoral level both in English and German, as well as supervision of student projects and academic theses
  • Advising and mentoring of students and early career researchers
  • Competitive research grant acquisition and management
  • Collaboration with academic and industry partners
  • Participation in university management
  • Participation in third mission and public relations activities

Your profile

  • Habilitation or equivalent qualification in Computer Science or a relevant neighbouring field
  • Excellent research track record in Information Systems Engineering
  • Experience in the acquisition and management of competitive third-party funded research projects of a relevant volume
  • Teaching competence and experience at university level
  • Experience in the (co-)supervision of academic theses
  • Fluency in English
This distinguishes you additionally
  • Interdisciplinary experience
  • Scientific dissemination skills
  • Engagement in academic administrative duties
  • Competence in leadership and teamwork
  • Competence in gender mainstreaming and diversity management
German language skills are not a formal prerequisite, but proficiency at level B2 is expected within two years.

Why you will enjoy working with us

The salary is subject to negotiation. The minimum gross salary for the position at this level (salary group A1 for University Staff according to the Austrian Universities’ Collective Bargaining Agreement) is currently € 93,986 per year.

The university is committed to increasing the number of women among the faculty, particularly in high-level positions, and therefore specifically invites applications from qualified women. Among equally qualified candidates, women will receive preferential consideration.

People with disabilities or chronic diseases who meet the qualification criteria are explicitly invited to apply.

In accordance with the Austrian Income Tax Act, an attractive relocation tax allowance can be granted for the first five years in the case of appointments to professorships in Austria. The prerequisites are subject to examination on a case by case basis.

Please submit your application in English by e-mail to the University of Klagenfurt, Office of the Senate, attn. Mag.a (FH) Sabine Seebacher via application_professorship@aau.at no later than 22 March, 2026, including:

  • a mandatory principal part not exceeding five pages (https://jobs.aau.at/wp-content/uploads/specimen_main_part_application_professorship.docx). The submission of the mandatory principal part constitutes a necessary condition for the validity of your application.
  • one single PDF including:
    • a letter of motivation
    • a detailed scientific CV
    • a comprehensive list of publications, talks, and of all courses taught
    • a list of projects that you acquired as a PI or co-PI, including the amount of funding that was attributed to you
    • a research statement
    • a teaching statement
    • supplementary documents where applicable (e.g., course evaluations)
    • links to publicly available versions of your three most important publications within the scope of this professorship

For general information, please refer to the general information provided at https://jobs.aau.at/en/the-university-as-employer/. For specific information about the position, please contact Prof. Dr. Martin Pinzger (Tel.: +43 463 2700 3513; martin.pinzger@aau.at).