This version of the blog post is also available at ACM SIGMM Records
The 155th MPEG meeting took place in Geneva, Switzerland, from July 13 to 17, 2026. The official MPEG press release can be found here. This report highlights key outcomes from the meeting, with a focus on research directions relevant to the ACM SIGMM community:
- Gaussian Splat Coding Use Cases and Current Status
- Call for Proposals on Media Authenticity and Provenance Indication
- Joint Call for Proposals on video compression with capability beyond VVC
- Joint ITU-T SG 21 and ISO/IEC JTC 1/SC 29 Workshop on Media Streaming Services
- Exploration on Systems technologies for AI-based media standards (SyfAI)
Gaussian Splat Coding Use Cases and Current Status
The previous MPEG meeting report introduced MPEG’s Gaussian Splat Coding (GSC) exploration, its two tracks (I-3DGS and A-3DGS), and a first set of 27 draft use cases. At the 155th meeting, WG 2 (Technical Requirements) approved an updated document (WG 2 N531) that drops the draft qualifier and adds two use cases.
The additions matter less for their number than their nature: both concern storage and delivery rather than compression. Use case 28 stores a static splat asset in a single file together with a cover image, thumbnail, and audio annotation, so that a capable receiver renders it interactively while a legacy receiver displays the cover image, with no modification of the file. Use case 29 is the temporal counterpart, delivering a dynamic splat track by HTTP adaptive streaming alongside synchronized audio, a pre-rendered video track serving as preview and fallback, and timed text, with representations offered per bitrate, level of detail, spherical harmonics subset, or attribute subset. The application-oriented use cases are unchanged, although the derived requirements are being refactored, notably by separating attribute subset scalability from random access.
GSC remains an exploration activity, but a busy one: 83 input contributions and nine Joint Exploration Experiments re-conducted across WG 2, WG 4 (Video Coding), WG 5 (Joint Video Experts Team, JVET), and WG 7 (3D Graphics and Haptics), with WG 1 (JPEG) now engaged on quality metrics and WG 4 examining Neural Network Coding (NNC) as a compression tool. On the fast track, GS4 (Video-based Point Cloud Compression, V-PCC, Amd. 1) and GS5 (Geometry-based Point Cloud Compression, G-PCC, Amd. 1) progress on reference software and conformance, while the JVET codec track carries five competing proposals. The Call for Proposals (CfP) still has no date, and I-3DGS planning is more mature than A-3DGS. The clearest message from the meeting is that single-frame compression is essentially solved and that the difficulty and the expected gain now lie in dynamic content.
Research aspects: Temporal coding is no longer one open question among many, but the central one, covering inter-prediction for anisotropic primitives, deformation models, and primitive correspondence across frames. Quality assessment remains unresolved: current test conditions score rendered views with peak signal-to-noise ratio (PSNR), structural similarity (SSIM), and learned perceptual image patch similarity (LPIPS) along a pose trace, a fragile proxy for artifacts such as floaters and popping; the involvement of JPEG makes this a good moment to contribute. Use case 29 gives the streaming question a concrete form: when quality varies along several orthogonal axes at once, what is the right abstraction for a representation, and what should an adaptation algorithm optimize?
Call for Proposals on Media Authenticity and Provenance Indication
At the 155th meeting, WG 2 approved a CfP on media authenticity and provenance indication with the MPEG Systems technologies (WG 2 N535), following an exploration phase in WG 3 (Systems) that produced requirements, a gap analysis, and several technical proposals (WG 3 N1842, attached to the call). The scope is the system layer: carriage and signaling of metadata that allows a receiver to verify that a media asset comes from a trusted producer and to convey provenance information, rather than the definition of a provenance language itself.
Three interfaces are addressed, namely (i) elementary streams and non-timed items, (ii) the file format level (ISO Base Media File Format (ISOBMFF) and Common Media Application Format (CMAF)), and (iii) packaged delivery (Dynamic Adaptive Streaming over HTTP (DASH), CMAF, MPEG Media Transport (MMT), and MPEG-2 Transport Stream (MPEG-2 TS)). The use cases span deepfakes and manipulated news media, forgery in insurance claims, surveillance and investigations, labeling of AI-generated content, and legitimate modifications such as editing, transcoding, ad insertion, and archival preservation. Detecting whether an asset is fake without embedded data is explicitly out of scope.
A response may be a complete solution or a single tool addressing one or more requirements, and is evaluated on requirement coverage, computational complexity, the bandwidth needed to convey the authenticity information, compatibility and backward compatibility with existing MPEG standards, and extensibility, with self-evaluation tables provided in the annexes. Proposals are submitted as input contributions to the 157th meeting in Brisbane, Australia, by January 13, 2027. Review at that meeting will produce one or more working drafts or new work item proposals, and the preliminary development plan targets Committee Draft (CD) or Committee Draft Amendment (CDAM) at the 158th meeting, Draft International Standard (DIS) or Draft Amendment (DAM) at the 159th, and Final Draft International Standard (FDIS) or Final Draft Amendment (FDAM) in early 2028.
Research aspects: The requirements make this more interesting than a signing exercise. A conventional signature breaks on any bit change, yet the call demands that verification survive transcoding, dropped scalability layers, representation switching, splicing, late binding, and loss of frames or audio, and that the unmodified remainder of a presentation stay verifiable after an authorized edit. Designing verification structures with that granularity and quantifying their overhead per segment against the added latency in live scenarios is an open problem at the intersection of cryptography and streaming systems. A second question is joint verification: establishing that a given audio track was the one the producer intended to accompany a given video, including their synchronization, cannot be achieved by signing each file separately. Finally, interoperability with provenance schemes defined elsewhere, and the question of what a receiver should do with a verification result, leave room for work on usable trust signaling.
Joint Call for Proposals on Video Compression with Capability Beyond VVC
The draft of this call was described in the previous MPEG meeting report, and JVET has now issued the final version (JVET-AQ2021), approved at its 43rd meeting in Geneva in July 2026. The target is video coding technology that significantly exceeds Versatile Video Coding (VVC) in compression, implementability, applicability across content types, and features such as latency, robustness, and scalability, benchmarked against the VVC Main 10 profile. Four test cases are defined: one for improved compression without runtime limits, and three in which the aggregate encoder run time is constrained to 5x, 1x, and 0.2x that of the VVC Test Model (VTM) anchor. All are evaluated over seven categories covering (i) standard dynamic range (SDR) random access at ultra-high definition (UHD) and (ii) at high definition (HD), (iii) SDR low delay HD, (iv) high dynamic range (HDR) with perceptual quantizer (PQ) and (v) with hybrid log-gamma (HLG) transfer functions, (vi) gaming, and (vii) user-generated content. A separate track invites technology offering additional functionality, together with proposals on how its benefit should be assessed.
Formal subjective testing uses degradation category rating, with objective results reported as PSNR, multiscale SSIM (MS-SSIM), and, for PQ content, weighted PSNR. The schedule is tight: anchors have been available since May 2026, registration runs from August 1 to September 1, 2026, and the main package of bitstreams, reconstructed sequences, and binaries must reach the test coordinator on physical media by October 26, 2026. Subjective assessment then runs until late December, blind cross-checking by other proponents is mandatory, and proposals are evaluated at the 45th JVET meeting in January 2027, with an initial test model selected during 2027 and the standard targeted for October 2029. Participation is not free: up to EUR 20k per test case is charged to cover the hiring of test subjects.
Research aspects: Two design choices in the call are worth attention. First, a supplemental set of sequences is disclosed to proponents only after the decoder binaries have been submitted, and results on it are due six weeks later, which turns the call into a held-out generalization test. Read together with the ban on training on test sequences and the obligation to disclose training material, this makes out-of-domain behavior of learned coding tools measurable at scale, and reporting it well is a contribution in itself. Second, run time is aggregated as the sum over threads, which measures total compute rather than latency and therefore reads very differently for a massively parallel or GPU-resident design than for a sequential one. How to characterize the rate, distortion, and complexity trade-off fairly across such architectures, and how to value functionality such as scalability or error resilience against a plain bitrate gain, are open questions the call poses rather than answers.
Joint ITU-T SG 21 and ISO/IEC JTC 1/SC 29 Workshop on Media Streaming Services
On July 14, 2026, ITU-T SG 21 and MPEG Systems held a joint half-day workshop in Geneva, collocated with their meetings, on “Media Streaming Service, What’s next” (call for presentations, program). The premise was that after the transitions from analog to digital, enabled by MPEG-2 Systems, and from broadcast to over the top (OTT), enabled by MPEG-DASH, integrated networks, edge computing, and AI are driving another shift, and that both bodies wanted industry views before committing to new work. On challenges, Netflix spoke on the shortcomings and evolution of ISOBMFF, Bitmovin on where streaming trends meet the container, ETRI on future streaming services, and Huawei on ultra-low latency communication and streaming. On opportunities, 5G-MAG addressed standards and open source, DVB heterogeneous networks, and 3GPP SA4 delivery beyond OTT. The closing session was an open mic with questions and answers from both speakers and the audience.
The wrap-up set the existing systems standards, namely MPEG-2 TS, ISOBMFF, DASH, and CMAF as well as the volumetric, scene, and augmented reality (AR) formats, plus ongoing work on authenticity, SyfAI (see below), common metadata, and Gaussian splats, against what industry actually asked for: low latency, low overhead, AI-driven media, and open source software. The resulting agenda is an exploration of a better or new container format, an analysis of overhead, processing-friendly metadata, including JSON and ontology issues, WebCodecs integration, and ingest. Just as notable are the conclusions about process: more open access and industry involvement, faster turnaround, possibly a new home or outlet for this work, and software with interoperability testing from day one rather than reference software at the end. Three ad hoc groups were set up: (i) on MP4, DASH, and file formats, reviewing ISOBMFF against emerging transports and the delivery of AI input and output data; (ii) on vision, analyzing bottlenecks and requirements beyond ISOBMFF; and (iii) on working methods.
Research aspects: Two of these items are directly researchable. Container overhead is asserted more often than measured, and as segments shrink toward frame level for low latency, the ratio of container to media bytes grows; a careful comparison across ISOBMFF, CMAF, and object-based transports at equal latency would inform the exploration rather than follow it. Media over QUIC (MOQ) is the bigger change, since it replaces the segment with the object as the unit of delivery and pulls streaming and real-time communication into a single design space, which reopens rate adaptation, interaction with congestion control, caching, and relay behavior, and the question of what a container still contributes when the transport itself frames media. Carrying AI input and output data alongside media, finally, links this work to coding for machines and to semantic streaming.
Exploration on Systems Technologies for AI-based Media Standards (SyfAI)
Several MPEG coding standards now put a neural network inside the decoder, among them video and feature coding for machines, AI-based point cloud coding, and neural network coding. The previous MPEG meeting report described that coding side through the MPEG-AI vision document. SyfAI, an MPEG Systems exploration started at the 153rd MPEG meeting, asks the complementary and much less glamorous question: what does the surrounding infrastructure have to do so that such content can actually be stored, delivered, and played back interoperably? The underlying shift is that a media file has traditionally been self-contained, meaning that a conforming decoder and the bitstream are sufficient. Once the decoder depends on a trained model that may be selected, delivered, or updated separately, that assumption no longer holds, and the system layer has to say how a player learns which model it needs, how that model reaches it, and how both sides can be sure they are using the same one. Work so far has advanced on the most concrete piece, namely storing video coding for machines content in the MPEG file format, while proposals for carrying compressed neural networks were sent back for clearer use cases. The more interesting development is a new thread on an AI update framework, opened jointly with WG 7, which collects five topics: (i) a format for AI parameter data, (ii) a manifest for updating parameters, (iii) a repository to serve them, (iv) integrity checking, and (v) bit exactness together with conformance.
Research aspects: Each of those topics is a research problem in its own right. Distributing and updating models alongside media turns into a delivery question that looks familiar but has not been studied: when to fetch an update, how to cache and version models, what a manifest must express about capability and compatibility, and how to treat weights as a second class of asset next to the media. Conformance is harder, because a standard normally guarantees that every decoder produces identical output, whereas neural inference varies with library and hardware, so deciding what conformance means and how to test it remains open and extends the reproducibility concerns already visible in MPEG-AI. A repository of updatable weights is also an attack surface, which is why this work sits naturally beside the media authenticity call discussed above: signing and verifying a model raises the same questions one level down from signing content. And because the choice of where a model lives affects how quickly playback can start and how smoothly a player can switch, the apparently dry container questions have measurable consequences for the quality of experience (QoE).
Concluding Remarks
Taken together, the five items show MPEG engaging early with technologies whose momentum comes from outside the committee. With Gaussian splatting, it is compressing a representation that already has a renderer and a user base, which is a better starting point than earlier attempts at three-dimensional video had, even if wider adoption will depend on capture and display ecosystems as much as on coding efficiency. The authenticity call is scoped pragmatically at establishing origin, which is achievable in the near term and provides the system-level plumbing that regulation and industry are asking for. The call for video coding beyond VVC visibly reflects deployment experience, with runtime-constrained test cases treating practicality as a first-class criterion, while questions of licensing and market uptake sit largely outside MPEG itself. In systems, the joint workshop and the new ad hoc groups show a willingness to re-examine long-standing assumptions as transports such as MOQ emerge, and SyfAI stakes out the system-level questions of the AI era before they become urgent. For our community, the result is an unusually rich set of open problems in evaluation and methodology.
The 156th MPEG meeting will be held in Hangzhou, China, from October 19 to 23, 2026. Click here for more information about MPEG meetings and ongoing developments.




