Showing posts with label Universal Multimedia Access. Show all posts
Showing posts with label Universal Multimedia Access. Show all posts

Friday, March 19, 2010

O Universal Multimedia Access, Where Art Thou? (Index)

-by Christian Timmerer, Klagenfurt University, Austria

Preface: First I thought about writing this article for a journal or something equivalent but then I concluded to make this article available through my blog. The aim is to perform an experiment in order to determine whether it is possible (a) to get direct feedback through comments and (b) to be referenced from elsewhere. As it is a quite comprehensive article, it’s split up in separate parts. If someone (i.e., a journal editor) is interested in publishing this article, yes, I can still do that! :-)

Index of "O Universal Multimedia Access, Where Art Thou?" seriesConclusion
I would say that UMA is technically feasible but the issue is with the content rights owners/holders ... if you disagree or have another opinion, please let me know.

Wednesday, March 17, 2010

O Universal Multimedia Access, Where Art Thou? (Part IV)

-by Christian Timmerer, Klagenfurt University, Austria

Preface: First I thought about writing this article for a journal or something equivalent but then I concluded to make this article available through my blog. The aim is to perform an experiment in order to determine whether it is possible (a) to get direct feedback through comments and (b) to be referenced from elsewhere. As it is a quite comprehensive article, it’s split up in separate parts. If someone (i.e., a journal editor) is interested in publishing this article, yes, I can still do that! :-)

Part I was about giving an introduction to the topic and an overview on multimedia content adaptation techniques. Part II was about the adaptation by transformation approach that utilizes scalable coding formats such as JPEG2000, MPEG-4 BSAC, and MPEG-4 SVC. Part III comprises adaptation decision-taking also known as the brain of multimedia content adaptation and this part is about standardization support for UMA.

Part IV - Standardization support for UMA


The nice thing about standards is that there are so many to choose from. Furthermore, if you do not like any of them, you can just wait for next year’s model.
--Andrew S. Tanenbaum

A couple of standardization organizations (SDOs) provide support for UMA:

Word Wide Web Consortium (W3C): http://www.w3.org/
Internet Engineering Task Force (IETF): http://www.ietf.org/
  • Audio/Video Transport (AVT)
  • Media Server Control (MEDIACTRL)
  • Multiparty Multimedia Session Control (MMUSIC)
  • Session Description Protocol (SDP) and Session Initiation Protocol (SIP)
  • Next Steps in Signaling (NSIS)
Moving Picture Experts Group (MPEG): http://www.chiariglione.org/mpeg/
A comprehensive description format with respect to UMA is Part 7 of MPEG-21 entitled Digital Item Adaptation [2]. This part of MPEG-21 defines - among others - the Usage Environment Description (UED) providing means for describing the context in which Digital Items may be consumed. The UED is clustered into the following categories with some examples given:
  • User Characteristics: e.g., usage history, display presentation preferences, audio/visual impairments, mobility, etc.
  • Terminal Capabilities: e.g., coding capabilities, display capabilities, audio output capabilities, etc.
  • Network Characteristics: e.g., network capabilities (e.g., max capacity, min guaranteed) and conditions (e.g., available bandwidth)
  • Natural Environment Characteristics: e.g., noise level, illumination characteristics, location, time, etc.
The UED is defined as an XML Schema which is publicly available here.

This is the end of Part IV and I'm currently not sure whether a Part V will follow...

References:
[1] Ian Burnett, Fernando Pereira, Rik Van de Walle, and Rob Koenen (eds.), The MPEG-21 Book, John Wiley and Sons Ltd, 2006.
[2] Anthony Vetro and Christian Timmerer, Digital Item Adaptation: Overview of Standardization and Research Activities, IEEE Transactions on Multimedia, vol. 7, no. 3, pp. 418-426, June 2005.

    Wednesday, March 10, 2010

    ACM Multimedia Grand Challenge 2010: Content Adaptation

    One of the ACM Multimedia Grand Challenge 2010 is about content adaptation, probably one of THE tools providing Universal (Multi-)Media Access (UMA). In particular, it is called "Radvision Challenge 2010: Real-time Data Collaboration Adaptation for Multi-Device Video Conferencing" and details can be found here with the input/output described as follows:
    Input for this challenge is a video capture of a free-hand drawing (see example video) in XGA.
    Output for this challenge should be a set of “adapted” videos , with the same content in different (smaller) resolutions – for instance, VGA and QVGA. The adapted videos would ideally be regarded by users as perceptually optimal, meaning they hold the same content as the original.
    The metrics for evaluation are "defined" as follows:
    The following criteria could be used, as well as other evaluation metrics that you may devise:
    1. Subjective comparison between the perceptual quality of the original and the “adapted” content.
    2. Subjective comparison between the perceptual quality of a scaled-down version of the original (using a 5-tap poly-phase filter) and the “adapted” content.
    3. Real-time Performance.
    However, there are many possibilities to adapt content and evaluate the result which heavily depends on the user's context. Some people may think that the description of the input/output as well as the evaluation criteria is defined too vague and I tend to agree. Let me explain:
    1. The input is given and for the output it is requested to produce "a set of adapted videos" that is "regarded by users as perceptually optimal". However, it's not clear to which context the input shall be adapted. The text says "different (smaller) resolutions" as an example but I can imaging that users will prefer the original video and regard it as perceptually optimal compared to anything else. Thus, in my view it is necessary to specify the context to which the video shall be adapted. The context may include a lot of things such as terminal device, decoding capabilities, network conditions, user location (stationary, mobile), etc., etc.
    2. Subjective quality assessment is not an easy task and there are many possibilities and many approaches. Form the description above it is not clear how the subjective quality evaluation will be performed. In particular, I wonder whether "real" subjective tests as suggested by the Video Quality Experts Group (VQEG) will be adopted (e.g., DSIS - Double Stimulus Impairment Scale or ACR - Absolute Category Scale to just name two). In my view and in order to provide a fair evaluation it is absolutely necessary to define the exact procedure on how the subjective evaluation of the submissions will be performed. One possibility, of course, is the adoption of a standardized approach and probably DSIS is the right candidate.

    Thursday, February 25, 2010

    O Universal Multimedia Access, Where Art Thou? (Part III)

    -by Christian Timmerer, Klagenfurt University, Austria

    Preface: First I thought about writing this article for a journal or something equivalent but then I concluded to make this article available through my blog. The aim is to perform an experiment in order to determine whether it is possible (a) to get direct feedback through comments and (b) to be referenced from elsewhere. As it is a quite comprehensive article, it’s split up in separate parts. If someone (i.e., a journal editor) is interested in publishing this article, yes, I can still do that! :-)

    Part I was about giving an introduction to the topic and an overview on multimedia content adaptation techniques. Part II was about the adaptation by transformation approach that utilizes scalable coding formats such as JPEG2000, MPEG-4 BSAC, and MPEG-4 SVC. This part comprises adaptation decision-taking also known as the brain of multimedia content adaptation.

    Part III – Adaptation Decision-Taking

    Definition: (Multimedia) adaptation decision-taking is referred to as the process of finding the optimal parameter settings for (multiple, possibly in series connected) multimedia content adaptation engines given the properties, characteristics, and capabilities of the content and the context in which it will be processed.

    Problem Description

    The information revolution of the last decade has resulted in a phenomenal increase in the quantity of multimedia content available to an increasing number of different users with different preferences who access it through a plethora of devices and over heterogeneous networks. End devices range from mobile phones to high definition TVs, access networks can be as diverse as UMTS (Universal Mobile Telecommunications System) and broadband networks, and the various backbone networks are different in bandwidth and Quality of Service (QoS) support. Additionally, users have different content/presentation preferences and intend to consume the content at different locations, times, and under altering circumstances, i.e., within a variety of different contexts.

    In order to cope with situations indicated above, multimedia content adaptation has become a key issue which results in a lot of research and standardization efforts collectively referred to as Universal Multimedia Access (UMA). An important aspect of UMA is adaptation decision-taking (ADT) which aims at finding the optimal parameter settings for the actual multimedia content adaptation engines based on the properties, characteristics, and capabilities of the content and the context in which it will be processed. This article provides an overview of the different metadata required for adaptation decision-taking and points out technical solution approaches for the actual decision-taking.

    High-level Architecture and Metadata Assets

    Figure 1 depicts a high-level architecture for adaptation decision-taking including the actual content adaptation. The input of the adaptation decision-taking engine (ADTE) can be divided into content- and context-related metadata. The former provides information about the syntactic and semantic aspects (e.g., bitrate, scene description) of the multimedia content that support the decision-taking process. The latter describes the usage environment (e.g., terminal capabilities) in which the multimedia content is consumed or processed. The result of the ADTE is an adaptation decision which steers the multimedia content adaptation engine(s) to produce the adapted multimedia content fulfilling the constraints imposed by the context-related metadata. The input to the actual adaptation engine is the given multimedia content possibly accompanied with additional content-related metadata required for the adaptation itself (e.g., syntax descriptions).
     
    Figure 1. High-level Architecture of Adaptation Decision-Taking and Multimedia Content Adaptation.

    The focus of this article is on the ADTE. In the following sections the two types of metadata assets required for adaptation decision-taking are reviewed and, finally, technical solution approaches are highlighted.

    Content-related Metadata

    This type of metadata comprises descriptive information about the characteristics of the content which can be divided into four categories:
    • Semantic metadata provides means for annotating multimedia content with textual information enabling various applications such as search and retrieval. This kind of metadata covers a broad range of annotation possibilities, among them are the name of the content, authors, actors, scene descriptions, etc.
    • Media characteristics describe the syntactical information pertaining to multimedia bitstreams in terms of the physical format and its characteristics. This may include the storage and coding format as well as bit-rate, frame-rate, width, height, and other related parameters.
    • The Digital Rights Management (DRM) information for adaptation decision-taking specify which kind of adaptation operations (e.g., scaling, format conversion, etc.) are allowed and under which constraints (e.g., bit-rate shall be greater than 512kbps).
    • Finally, Adaptation Quality of Service (QoS) describes the relationship between usage environment constraints, feasible adaptation operations satisfying these constraints, and associated utilities (i.e., qualities).

      Context-related Metadata

      Similar to the content-related metadata, the context-related metadata can be also divided into four categories:
      • End user-related metadata: The first category of metadata is pertained to metadata describing the characteristics of end users in terms of preferences, disabilities, and location-based information.
      • Terminal-related metadata: The second category provides context information regarding the capabilities of the terminal which are used by the end users for consuming multimedia content. This information includes – among others – information about the codecs installed, display properties, and audio capabilities.
      • Network-related metadata: The third category of metadata comprises the information concerning the access and core networks in terms of its characteristics and conditions. Such information may include bandwidth, delay, and jitter.
      • Adaptation-related metadata: Finally, the fourth category of metadata describes the actual adaptation engines in terms of adaptation operations they are capable to perform. For example, an adaptation engine may be able to perform temporal scaling whereas another one provides means for spatial scaling or even complex transcoding operations between different coding formats.

        Solution Approaches for Adaptation Decision-Taking

        In the literature the following approaches towards adaptation decision-taking are known:
        • Knowledge-based ADT [1]: adopts an Artificial Intelligence-based planning approach to find an appropriate sequence of adaptation steps from an initial state (i.e., described by content-related metadata) towards a goal state (i.e., described by context-related metadata).
        • Optimization-based ADT [2]: models the problem of finding adaptation decisions as a mathematical optimization problem by describing functional dependencies between content- and context-related metadata. Furthermore, limitation constraints as well as an objective function is specified.
        • Utility-based ADT [3]: can be seen as an extension of the previous two approaches which explicitly takes the users’ specific utility aspects into account.
        This is the end of Part III and I will continue in Part IV with interoperability issues, i.e., standards supporting UMA. Thus, stay tuned!
         

        References

        [1]    D. Jannach, K. Leopold, C. Timmerer, and H. Hellwagner, "A Knowledge-based Framework for Multimedia Adaptation", Applied Intelligence, vol. 24, no. 2, pp. 109-125, April 2006.
        [2]    I. Kofler, C. Timmerer, H. Hellwagner, A. Hutter, and F. Sanahuja, "Efficient MPEG-21-based Adaptation Decision-Taking for Scalable Multimedia Content", Proceedings of the 14th SPIE Annual Electronic Imaging Conference – Multimedia Computing and Networking (MMCN 2007), San Jose, CA, USA, January/February 2007.
        [3]    M. Prangl, T. Szkaliczki, and H. Hellwagner, "A Framework for Utility-based Multimedia Adaptation", IEEE Transactions on Circuits and Systems for Video Technology, vol. 17, no. 6, pp. 719-728, June 2007.

        Tuesday, December 22, 2009

        O Universal Multimedia Access, Where Art Thou? (Part II)

        -by Christian Timmerer, Klagenfurt University, Austria

        Preface: First I thought about writing this article for a journal or something equivalent but then I concluded to make this article available through my blog. The aim is to perform an experiment in order to determine whether it is possible (a) to get direct feedback through comments and (b) to be referenced from elsewhere. As it is a quite comprehensive article, it’s split up in separate parts. If someone (i.e., a journal editor) is interested in publishing this article, yes, I can still do that! :-)

        Part I was about giving an introduction to the topic and an overview on multimedia content adaptation techniques. This part focuses on the adaptation by transformation approach that utilizes scalable coding formats such as JPEG2000, MPEG-4 BSAC, and MPEG-4 SVC and is mainly based on [1].

        Part II – Adaptation by Transformation


        Scalable coding techniques have been recognized as an appropriate tool for realizing the concepts of UMA. Furthermore, if widely adopted across industries, scalable coding would provide a generalized solution to the interoperability problem.

        In [2], a scalable bitstream is defined as a coded multimedia resource (i.e., audio-visual multimedia resources) consisting of a structured sequence of binary symbols which is organized in such a way that, by retrieving the bitstream, it is possible to first render a degraded version of the bitstream, and then progressively improve it by loading additional data. This definition implicates a bitstream structure where the bitstream can be logically divided into several layers, i.e., a base layer and one or more enhancement layers. The base layer offers a minimal quality of the bitstream whereas each of the enhancement layers successively provides improvements with respect to the quality in various dimensions. These dimensions include improvements in the temporal, spatial, signal-to-noise ratio (SNR), color, region-of-interest (ROI), and complexity domain, among others. Recently, abstract models describing scalable bitstreams have been proposed [3][4] which are briefly reviewed in the following.

        In general, a scalable bitstream can be organized in a logical hypercube model where each axis represents a scalability dimension (e.g., temporal, spatial, quality) and every data block within this model corresponds to a certain bitstream segment (cf. Figure 1). The adaptation of a bitstream corresponding to such a model comprises the removal of one or more data blocks sometimes followed by minor updates of the remaining data blocks. Please go to [4] for a more detailed overview on adaptation possibilities.

        Figure 1. Scalability Model using the Hypecube Model according to [3][4].


        In the following I'd like to introduce some coding formats and their scalability features featuring the hypercube model as introduced above, namely:

        • JPEG2000 which introduces spatial, color, SNR, and ROI scalability for still images; 
        • MPEG-4 Visual Elementary Stream (VES) with temporal and semantic scalability;
        • MPEG-4 BSAC with fine-grained SNR scalability;
        • MPEG-4 SVC with native support for temporal, spatial, and SNR scalability.
        Please note each scalable coding format is introduced with a special focus on its scalability aspects. For details regarding basic coding techniques the reader is referred to appropriate literature, e.g., [5] or [6].


        JPEG2000


        The JPEG2000 standard [7][8][9] is known as the successor of the world-famous and widely adopted JPEG standard [10]. The JPEG2000 standard has been developed in order to accommodate the increasing demands and additional requirements for multimedia and Internet applications. In particular, some of the most important features (with respect to scalability) the JPEG200 standard should offer are progressive transmission by pixel accuracy and resolution as well as Region of Interest (ROI) coding and random code-stream access and processing. Progressive transmission enables the rendering of images with different resolution and pixel accuracy starting from a base version up to a high-resolution/quality version in an incremental manner, i.e. more and more data is added to the base layer by only transmitting the additional data which is required for increasing quality and/or resolution.

        The hypercube model for JPEG2000 with the dimensions represent color, spatial, and SNR scalability respectively is depicted in Figure 2.

        Figure 2. Hypercube for JPEG2000 scalability and a possible bitstream layout.

        In particular, the figure shows the hypercube for JPEG2000 with its scalability dimensions and a possible bitstream layout with quality-spatial-color progression order. The gray cube represents the base layer with QCIF dimension, Y color component only, any a quality of 29 dB PSNR. In contrast, the blue cubes represent another version of the tile including more quality layers, i.e., a PSNR of 31 dB, with a CIF resolution but still only one color component, i.e., the resulting image is still a grayscale version of the original image.

        MPEG-4 Visual Elementary Streams


        MPEG-4 [11][12][13] also provides support for scalability in the spatial, temporal, and SNR dimensions but only a small amount of the scalability features has been adopted by industry, i.e., the temporal scalability. The spatial and SNR scalability features introduced too much coding overhead which was the main reason for not adopting these features at this time.

        Temporal scalability is often also referred to as frame dropping where frames or visual object planes (VOPs) are removed which are not used as a reference frame for other frames. Bi-directional coded VOPs (B-VOPs) are not used as a reference for other frames, i.e., B-VOPs can be dropped arbitrarily. In case a predictive coded VOP (P-VOP) needs to be dropped all corresponding B-VOPs which use this P-VOP as reference frame need to be dropped as well. Similar behavior holds for intra coded VOPs (I-VOPs) although usually not dropped in traditional temporal scalability scenarios.

        Another dimension of scalability is introduced here known as semantic scalability. This additional dimension associates properties to a group of VOPs (GoVs) providing means for summarization or personalization of MPEG-4 visual resources. With respect to the scalability model, GoVs can be compared with parcels and VOPs can be seen as the data blocks. The semantic scalability is, of course, also applicable for other audio/visual coding formats including those introduced in this article.

        Figure 3. Hypercube for MPEG-4 VES scalability and a possible bitstream layout.

        A possible configuration for an MPEG-4 Visual Elementary Stream hypercube model is depicted in Figure 3 with two levels of scalability, namely temporal and semantic. The former is characterized by different frame rates and the latter uses terms from the Internet Content Rating Association (ICRA) for rating the violence level of the actual content. For example, level 0 indicates no violence or sports-related content, respectively. The gray block represents a base layer (e.g., a scene or even only one I-VOP) with a frame rate of 15 Hz and a violence level 0 whereas the blue block indicate a scene with violence level 2 and 20 frames per second (fps).

        MPEG-4 Bit-Sliced Arithmetic Coding


        The concept of bit-sliced arithmetic coding for audio coding was introduced in [14] but is also excerpted in [11][15]. It is very similar to the well-known Advanced Audio Coding (AAC) [16] scheme except that the quantized values are not Huffman coded but arithmetically coded in bit-slices. Thus, MPEG-4 Bit-Sliced Arithmetic Coding (BSAC) provides fine-grain scalability of approximately 1 kbit/s per audio channel per enhancement layer. The base layer comprises side information, scaling factors and the actual audio data according to the bit rate of the base layer. Each enhancement layer incrementally adds more and more information with respect to the bit rate and a maximum of 48 enhancement layers are allowed. Due to the small size of the enhancement layers, i.e., 20 to 60 bits per AAC frame typically representing 20 to 30 ms, which may result in undesired packetization overhead, data packets of consecutive frames can be grouped together.

        Figure 4. Hypercube for MPEG-4 BSAC scalability and a possible bitstream layout.

        A hypercube model for a stereo MPEG-4 BSAC bitstream including a possible bitstream layout is illustrated in Figure 4. The base layer is encoded at 48 kbit/s/channel and a possible adapted stereo version of the bitstream with 50 kbit/s is indicated as well.

        MPEG-4 Scalable Video Coding


        MPEG-4 Scalable Video Coding (SVC) [17] is being introduced as an extension of MPEG-4 Advanced Video Coding (AVC) [18] which is part 10 of the MPEG-4 family of audio/visual coding standards. MPEG-4 SVC natively supports three scalability dimensions, namely temporal, spatial, and quality (SNR).

        Figure 5.  Hypercube for MPEG-4 SVC scalability and a possible bitstream layout.

        In Figure 5 a hypercube model with the three scalability dimensions of MPEG-4 SVC including a possible bitstream layout is shown. In this example, the base layer provides a QCIF version at 20Hz with a PSNR of 28dB. Additionally, an improved version with higher temporal, spatial, and SNR resolution is indicated.

        Adaptation of Scalable Bitstreams


        The adaptation of scalable bitstreams can be basically organized into two category:
        • The first category is a coding-format specific approach which, in general, is applicable to one coding format only such as the Bitstream Extractor that is part of the Joint Scalable Video Model (JSVM). The disadvantage here is that for each coding format a separate "bitstream extraxtor" is needed which become an issue for a growing number of instances.
        • The second category is referred to as coding-format independent or generic approach that is applicable to all scalable coding format but requires additional metadata [19]. As this approach is rather new and not commonly known, I will give an brief overview in the following.
        Please note that a comparison between the generic and specific approach in the context of SVC is reported in [20].

        Generic Multimedia Content Adaptation


        This section discusses means to process (i.e., adapt, customize, manipulate, etc.) multimedia content independently of the actual coding format by utilizing XML-based metadata describing the high-level structure (i.e., syntax) of a bitstream. That is, the resulting XML document describes the bitstream how it is organized at different syntactical and even semantic levels, e.g., in terms of packets, headers, layers, units, segments, shots, scenes, etc., depending on the actual application requirements. It is important to note that the XML description does not describe the bitstream on a bit-by-bit basis, i.e., it does not replace the actual bitstream but provides metadata regarding bit/byte positions of meaningful segments for the given application. Therefore, the XML description does not necessarily provide any information of the actual coding format used as only the positions and – in some cases – meanings are required for processing.

        High-level Architecture of Generic Content Adaptation

        Figure 6 depicts the high-level architecture of generic multimedia content adaptation which can be logically divided into two processes, namely the Description Transformation and the Bitstream Generation.

        Figure 6. High-level architecture of Generic Multimedia Content Adaptation (adopted from [21]).

        The description transformation process receives as an input the XML description of the source bitstream and a so-called style sheet that transforms the XML document according to the context information, e.g., the device capabilities. The output of this process is a transformed description which already reflects the bitstream segments of the target (i.e., adapted) bitstream. However, the transformed description still refers to the bit/byte positions of the source bitstream which needs to be parsed in order to generate the target bitstream within the second step of the adaptation process, i.e., the bitstream generation.

        Please note that the description transformation and bitstream generation processes should be combined by applying appropriate implementation techniques in order to achieve the required performance. However, implementation and optimization techniques for this kind of approach are out of scope of this article and the interested reader is referred to [22-25].

        Technical Solution Approaches

        The literature offers several technical solution approaches for generic multimedia content adaptation which are briefly highlighted in the following:
        • (X)Flavor [26]: A Formal Language for Audio-Visual Object Representation which has been extended with XML features.
        • Bitstream Syntax Description Language (BSDL) [27]: An XML Schema-based language for constructing a Bitstream Syntax Schema (BS Schema) for a given coding format [28]. It enables the generation of a Bitstream Syntax Description (BSD) based on a given bitstream and vice versa. The generic counterpart of the coding format-specific BS Schema is referred to as gBS Schema which is fully coding format-agnostic. An XML document conforming to the gBS Schema is referred to as a generic Bitstream Syntax Description (gBSD) [29].
        • BFlavor [30]: A method that combines BSDL and XFlavor and basically uses XFlavor techniques – enhanced with BSDL concepts – to generate Java code which is used for automatic generation of BSDs.

        Summary


        Figure 7 gives a summary of the various multimedia content adaptation techniques presented in Part I and Part II. The summary has been adopted and extended from [31].

        Figure 7. Summary of Multimedia Content Adaptation (adopted from [31]).

        This is the end of Part II and I will continue in Part III with the adaptation decision-taking also known as the brain of multimedia content adaptation. Thus, stay tuned!

        References:
        [1] C. Timmerer, Generic Adaptation of Scalable Multimedia Resources, VDM Verlag Dr. Müller, 2008.
        [2] ISO/IEC 21000-7, Information technology — Multimedia framework (MPEG-21) — Part 7: Digital Item Adaptation, October 2004.
        [3] S. Lerouge, R. De Sutter, P. Lambert, and R. Van de Walle, "Fully Scalable Video Coding in Multicast Applications", Proceedings of SPIE/Electronic Imaging 2004, vol. 5308, San Jose, CA, US, 2004, pp. 555-564.
        [4] D. Mukherjee, A. Said, and S. Liu, "A framework for fully format-independent adaptation of scalable bit-streams," IEEE Transactions on Circuits and Systems for Video Technology, Special Issue on Video Adaptation, vol. 15, no. 10, October 2005, pp. 1280-1290.
        [5] R. Steinmetz, Multimedia-Technologie. Grundlagen, Komponenten und Systeme, Springer, Berlin, July 2000.
        [6] F. Halsall, Multimedia Communications. Applications, Networks, Protocols and Standards, Addison Wesley, November 2000.
        [7] ISO/IEC 15444-1:2004, Information technology — JPEG 2000 image coding system: Core coding system, 2nd edition, September 2004.
        [8] D. Taubman and M. Marcellin (eds.), JPEG2000: Image Compression Fundamentals, Standards and Practice, Springer, November 2001.
        [9] C. Christopoulos, A. Skodras, and T. Ebrahimi, "The JPEG2000 Still Image Coding System: An Overview", IEEE Transactions on Consumer Electronics, vol. 46, no. 4, November 2000, pp. 1103-1127.
        [10] G. K. Wallace, "The JPEG still picture compression standard", Communications of the ACM, vol. 34, no. 4, April 1991, pp. 30-44.
        [11] F. Pereira and T. Ebrahimi (eds.), The MPEG-4 Book, Prentice Hall PTR, August 2002.
        [12] S. Battista, F. Casalino, and C. Lande, "MPEG-4: A Multimedia Standard for the Third Millennium, Part 1", IEEE MultiMedia Magazine, vol. 6, no. 4, October-December 1999, pp. 74-83.
        [13] S. Battista, F. Casalino, and C. Lande, "MPEG-4: A Multimedia Standard for the Third Millennium, Part 2", IEEE MultiMedia Magazine, vol. 7, no. 1, January-March 2000, pp. 76-84.
        [14] S. Park, Y. Kim, S. Kim, and Y. Seo, "Multi-Layer Bit-Sliced Bit-Rate Scalable Audio Coding", in 103rd AES Convention, preprint 4520, New York, September 1997.
        [15] H. Prunhagen, "An Overview of MPEG-4 Audio Version 2", Proceedings of AES 17th International Conference on High-Quality Audio Coding, Florence, Italy, September 1999, pp. 157-168.
        [16] ISO/IEC 13818-7:2006, Information technology — Generic coding of moving pictures and associated audio information — Part 7: Advanced Audio Coding (AAC), 4th edition, January 2006.
        [17] H. Schwarz, D. Marpe, T. Wiegand, "Overview of the Scalable Video Coding Extensions of the H.264/AVC Standard", IEEE Transactions on Circuits and Systems for Video Technology, vol. 17, no. 9, Sep. 2007, pp. 1103-1120.
        [18] T. Wiegand, G. J. Sullivan, G. Bjøntegaard, A. Luthra, "Overview of the H.264/AVC Video Coding Standard", IEEE Transactions on Circuits and Systems for Video Technology, vol. 13, no. 7, July 2003, pp. 560-576.
        [19] C. Timmerer, M. Ransburg, and H. Hellwagner, "Generic Multimedia Content Adaptation", in: Borko Furht (ed.), Encyclopedia of Multimedia, 2nd edition, Springer, pp. 263-271, October 2008.
        [20] M. Eberhard, L. Celetto, C. Timmerer, E. Quacchio and H. Hellwagner, "Performance Analysis of Scalable Video Adaptation: Generic versus Specific Approach", Proceedings of WIAMIS 2008, Klagenfurt, Austria, May 2008.
        [21] C. Timmerer and H. Hellwagner, “Interoperable Adaptive Multimedia Communication”, IEEE Multimedia Magazine, vol. 12, no. 1, pp. 74-79, January-March 2005.
        [22] C. Timmerer, G. Panis, and E. Delfosse, “Piece-wise Multimedia Content Adaptation in Streaming and Constrained Environments”, Proceedings of the 6th International Workshop on Image Analysis for Multimedia Interactive Services (WIAMIS 2005), Montreux, Switzerland, April 2005.
        [23] C. Timmerer, T. Frank, and H. Hellwagner, “Efficient processing of MPEG-21 metadata in the binary domain”, Proceedings of SPIE International Symposium ITCom 2005 on Multimedia Systems and Applications VIII, Boston, Massachusetts, USA, October 2005.
        [24] M. Ransburg, C. Timmerer, H. Hellwagner, and S. Devillers, “Processing and Delivery of Multimedia Metadata for Multimedia Content Streaming”, Proceedings of the Workshop Multimedia Semantics - The Role of Metadata, RWTH Aachen, March 2007.
        [25] M. Ransburg, H. Gressl, and H. Hellwagner, “Efficient Transformation of MPEG-21 Metadata for Codec-agnostic Adaptation in Real-time Streaming Scenarios”, Proceedings of the 9th International Workshop on Image Analysis for Multimedia Interactive Services (WIAMIS 2008), Klagenfurt, Austria, May 2008.
        [26] D. Hong and A. Eleftheriadis, “XFlavor: Bridging Bits and Objects in Media Representation”, Proceedings IEEE International Conference on Multimedia and Expo (ICME), Lausanne, Switzerland, pp. 773- 776, August 2002.
        [27] M. Amielh and S. Devillers, “Bitstream Syntax Description Language: Application of XML-Schema to Multimedia Content”, 11th International World Wide Web Conference (WWW 2002), Honolulu, May, 2002.
        [28] G. Panis, A. Hutter, J. Heuer, H. Hellwagner, H. Kosch, C. Timmerer, S. Devillers and M. Amielh, “Bitstream Syntax Description: A Tool for Multimedia Resource Adaptation within MPEG-21”, Signal Processing: Image Communication, vol. 18, no. 8, pp. 721-747, September 2003.
        [29] C. Timmerer, G. Panis, H. Kosch, J. Heuer, H. Hellwagner, and A. Hutter, “Coding format independent multimedia content adaptation using XML”, Proceedings of SPIE International Symposium ITCom 2003 on Internet Multimedia Management Systems IV, Orlando, Florida, USA, pp. 92-103, September 2003.
        [30] W. De Neve, D. Van Deursen, D. De Schrijver, S. Lerouge, K. De Wolf, and R. Van de Walle, “BFlavor: A harmonized approach to media resource adaptation, inspired by MPEG-21 BSDL and XFlavor”, Signal Processing: Image Communication, vol. 21, no. 10, pp. 862-889, November 2006.
        [31] B. Shen, W-T. Tan, F. Huve, “Dynamic Video Transcoding in Mobile Environments“, IEEE Multimedia, vol. 15, no. 1, Jan.-Mar. 2008, pp. 42-51.

        Monday, December 7, 2009

        O Universal Multimedia Access, Where Art Thou? (Part I)

        -by Christian Timmerer, Klagenfurt University, Austria

        Preface: First I thought about writing this article for a journal or something equivalent but then I concluded to make this article available through my blog. The aim is to perform an experiment in order to determine whether it is possible (a) to get direct feedback through comments and (b) to be referenced from elsewhere. As it is a quite comprehensive article, it’s split up in separate parts. If someone (i.e., a journal editor) is interested in publishing this article, yes, I can still do that! :-)

        Part I – Introduction and Multimedia Content Adaptation Techniques


        Back in 1999, an article was published in vol. 1/no. 2 of IEEE Transactions of Multimedia entitled “Adapting Multimedia Internet Content for Universal Access” [1] which can be roughly seen as the kick-off for a research effort that in subsequent papers was collectively referred to as Universal Multimedia Access (UMA). The initial aim of UMA was to provide access to multimedia content anywhere, anytime, and with any device. In the meanwhile, that is, (more than) 10 years later, I think it is worth looking back and reviewing what has been achieved so far.

        Some argue and I tend to agree that the key to UMA is multimedia content adaptation [2] as depicted in Figure 1. The aim is the transformation of an input to an output in video or augmented multimedia forms utilizing manipulations at multiple levels (e.g., signal, structural, or semantic) in order to meet diverse resource constraints and user preferences while optimizing the overall utility of the multimedia content.



        Figure 1. Concept of Multimedia Content Adaptation – adopted from [2].

        How to adapt?

        • Temporal scaling: reduce number of frames
        • Spatial scaling: reduce number of pixels => reduce resolution
        • Frequency scaling: reduce number of DCT coefficients => reduce quality
        • Modality conversion: e.g., video to slide show also known as transmoding

        Where to adapt?

        • Server, Proxy, Router, Gateway, Client, …

        When to adapt?

        • A server could hold several variations of the same multimedia content – or – could react to changing (network) conditions
        • A proxy could adapt cached multimedia content in order to free space – or – could adapt it in an ad-hoc mode or on-demand
        • A router or gateway could drop marked segments (e.g. packets)
        • A client could subscribe only to those streams it can handle
        • etc.
        Based on the observations above, multimedia content adaptation can be roughly categorized into adaptation by selection, adaptation by transcoding, and adaptation by transformation.

        Adaptation by Selection


        The idea here is to provide multiple versions of the same multimedia content and then select or switch to the most appropriate version according to the usage context. The InfoPyramid framework [1][3] was among the first approaches of this adaptation paradigm. Therefore, content descriptions are associated to individual components of the multimedia content which describes the content at different modalities, at different resolutions, and at multiple abstractions (cf. Figure 2).



        Figure 2. InfoPyramid Framework – adopted from [3].

        • Multi-modal: Multimedia content is usually not in a single media format, or modality. A video clip can contain raw data from video, audio in two or more languages, closed captions, etc.
        • Multi-resolution: Each content component can also be described at multiple resolutions. Numerous resolution reduction techniques exist for constructing image and video resources.
        • Multiple-abstraction levels: The abstraction levels describe features and data in a hierarchical fashion. For example, one hierarchy could be features, semantics and object descriptions, and annotations and metadata itself.
        In order to access the actual content one has to define methods for manipulating, translating, transcoding, and generating content which can happen in offline or online mode.
        • The offline mode generates variations as described by InfoPyramid before service deployment. On service request one can choose or select the prepared variations based on the InfoPyramid description. The variations are generated by applying appropriate adaptation techniques (e.g., transcoding) offline which indeed increases storage and asset management requirements.
        • The online mode provides the appropriate variation on-the-fly based on the InfoPyramid description and during the actual service request. Again, appropriate adaptation techniques (e.g., transcoding) are applied but this time online which increases processing (CPU) requirements, delay, etc.
        However, in general it is difficult to anticipate and provide multimedia content given the large variety of formats, bit rates, etc. Furthermore and in offline mode, one needs to maintain and manage all these different version which is a waste of capacity. On the other hand, this approach yields good performance and little quality degradation.

        Adaptation by transcoding


        Although transcoding may be used as tool within the previous adaptation paradigm it is listed here as a separate approach due to its importance both in literature and industry [5][6][7]. The objective of transcoding is to satisfy usage environment constraints while maximizing the content value (objective/subjective quality) and minimizing the actual transcoding complexity. In general one can distinguish between re-coding and trans-coding.

        Conventional approaches – recode – performs full decoding, post-processing, and full re-encoding as shown in Figure 3. This approach usually yields highest quality but is an expensive approach though and in many cases (real-time) it requires a hardware-based solution.



        Figure 3. Conventional approaches – recode.
        Low-cost approaches – transcode – targets similar quality as the conventional approach but with lower complexity. The focus is on architectures that utilize compressed-domain processing which enables software solutions to be deployed (cf. Figure 4).


        Figure 4. Low-cost approaches – transcode.
        In the following common transcoding operations are briefly highlighted:
        • Bit-rate reduction – sometimes also known as transrating (e.g., SDTV: 6Mbps => 3 Mbps, HDTV: 19.2 Mbps => 11 Mbps): The main challenges here are drift compensation due to re-quantization errors, the rate control algorithm, and the trade-off between quality and complexity. A vast amount of solutions have been proposed in the literature and it is nearly impossible to summarize them. Nevertheless, a good overview is given in [8].
        • Temporal resolution reduction (e.g., 30 fps => 10 fps): Due to frame dropping also a couple of issues arrive. That is, how to estimate a new motion vector based on incoming motion vectors by avoiding full motion vector re-estimation and how to estimate a new residual based on incoming residual values while minimize mismatch between predictive and residual components. Some approaches are described in [7][9].
        • Spatial resolution reduction (e.g., HDTV => SDTV; 720x480i, 30Hz => 352x240p 10Hz):
          • Motion vectors corresponding to reduced resolution reference frame => frame-based & field-based motion vector mapping.
          • Obtaining texture information for lower resolution MB’s => simple averaging (frame-based or field-based, computationally efficient) or block-based filters (typically more complex than required).
          • Drift compensation architecture due to re-quantization and down-sampling => cascaded architecture (full decoding/re-encoding), partial encoding architecture (full decoding followed by partial encoder), and intra refresh architecture (open-loop architecture).
        • Error-resilience enhancement: Improve robustness of bitstream for transmission or use retransmitted frames as reference even if they arrive too late for being display. With such approach the error propagation is eliminated while it would persist if retransmitted frames were discarded [7].
        • Syntax conversion (e.g., MPEG-2 Transport Stream => MPEG-2 Program Stream for DVD Recording; MPEG-2 => MPEG-4 for Broadcast to Mobile): This operation is often referred to as the classical transcoding operation and was/is the main driving use case for UMA. It usually benefits from the operations introduced above and is used in certain combinations. Recently, bit-stream rewriting has been introduced which allows for syntax conversion within a given family of video coding standards (e.g., SVC-to-AVC [10] or AVC-to-SVC [11]).
        • Modality conversion – sometimes also known as transmoding (e.g., video => slideshow; text => speech): The objective is here to modify the modality (e.g., audio, video, image, text) in order to satisfy transmission constraints and/or user preferences [12][13][14].
        This is the end of Part I and I will continue in Part II with the adaptation by transformation approach that utilizes scalable coding formats such as JPEG2000, MPEG-4 BSAC, and MPEG-4 SVC. Thus, stay tuned!

        References:
        [1] R. Mohan, J. R. Smith, C.-S. Li, “Adapting Multimedia Internet Content for Universal Access,” IEEE Transactions on Multimedia, vol. 1, no. 1, 1999, pp. 104-114.
        [2] S.F. Chang and A. Vetro, “Video Adaptation: Concepts, Technologies and Open Issues“, Proceedings of the IEEE, vol. 93, no. 1, Jan. 2005, pp. 148-158.
        [3] C-S. Li, R. Mohan and J.R. Smith, “Multimedia Content Description in the InfoPyramid”, Proceedings ICASP’98, Special session on Signal Processing in Modern Multimedia Standards, Seattle, May 1998.
        [4] B. Shen, W-T. Tan, F. Huve, “Dynamic Video Transcoding in Mobile Environments“, IEEE Multimedia, vol. 15, no. 1, Jan.-Mar. 2008, pp. 42-51.
        [5] A. Vetro, C. Christopoulos and H. Sun, “An overview of video transcoding architectures and techniques“, IEEE Signal Processing Magazine, vol. 20, no. 2, Mar. 2003, pp. 18-29.
        [6] J. Xin, C.W. Lin, M.T. Sun, “Digital Video Transcoding“, Proceedings of IEEE, vol. 93, no. 1, Jan. 2005, pp. 84-97.
        [7] B. Shen, W-T. Tan, F. Huve, “Dynamic Video Transcoding in Mobile Environments”, IEEE Multimedia, vol. 15, no. 1, Jan.-Mar. 2008, pp. 42-51.
        [8] S. Liu, A. Bovik, "Digital Video Transcoding", in A. Bovik, The Essential Guide to Video Processing, Academic Press, 2009.
        [9] Fung, et al., “New architecture for dynamic frame skipping transcoder”, IEEE Transactions on Image Processing, vol. 11, no. 8, Aug. 2002, pp. 886-900.
        [10] A. Segall, J. Zhao, “Bit-stream rewriting for SVC-to-AVC conversion”, 15th International Conference on Image Processing (ICIP2008), San Diego, USA, Oct. 2008, pp. 2776-2779.
        [11] J. De Cock, S. Notebaert, P. Lambert, R. Van de Walle, “Advanced bitstream rewriting from H.264/AVC to SVC”, 15th International Conference on Image Processing (ICIP2008), San Diego, USA, Oct. 2008, pp. 2472-2475.
        [12] T. C. Thang, Y. J. Jung, J. W. Lee, Y. M. Ro, “Modality Conversion for Universal Multimedia Services”, Proceeding 5th International Workshop on Image Analysis for Multimedia Interactive Services (WIAMIS2004), Lisboa, Portugal, April, 2004.
        [13] T. C. Thang, Y. J. Jung, and Y. M. Ro, “Modality Conversion for QoS Management in Universal Multimedia Access”, IEE Proceedings: Vision, Image & Signal Processing, vol. 152, no. 3, Jun. 2005, pp.374-384.
        [14] M.K. Asadi, J.-C. Dufourd, “Multimedia Adaptation by Transmoding in MPEG-21”, Proceeding 5th International Workshop on Image Analysis for Multimedia Interactive Services (WIAMIS2004), Lisboa, Portugal, April, 2004.

        Tuesday, April 21, 2009

        Note Published: W3C Personalization Roadmap: Ubiquitous Web Integration of AccessForAll 1.0

        The Ubiquitous Web Applications Working Group has published the Group Note of W3C Personalization Roadmap: Ubiquitous Web Integration of AccessForAll 1.0. This document describes an activity of integrating personalization with device context for the delivery of content materials and interface components that are customized to meet both individual personal needs and preferences and delivery context. It brings together the work of separate standards and specifications organizations and working groups, notably W3C Ubiquitous Web Applications working group, IMS Global Learning Consortium Accessibility Special Interest group, ISO/IEC JTC1 SC36 Information Technology for Learning, Education and Training: Human Diversity and Access For All working group and associated working groups in SC36. The document should be viewed as a roadmap for the work to be undertaken and includes description of the basis for the work, the organizational context, the likely technologies and a partially complete description of how the technologies fit together. Learn more about the Ubiquitous Web Applications Activity.

        They probably should also include the work of ISO/IEC JTC 1/SC 29/WG 11 (MPEG) on Usage Environment Description (UED) which also provides means to describe user characteristics including accessibility information. UED has been standardized within Part 7 of MPEG-21, entitled Digital Item Adaptation (DIA). The UED Schema can be found here and just search for AuditoryImpairment or VisualImpairment.

        Sunday, February 22, 2009

        Adapting Content

        Franklin Reynolds, "Adapting Content", IEEE Pervasive Computing, vol. 7, no. 4, pp. 6-8, Oct.-Dec. 2008

        This is an interesting article which provides a solid description on how ‘device independence’ can be achieved by utilizing recent W3C standards. It provides a good introduction, highlights some of the main issues, and gives accurate pointers to the state-of-the-art W3C standards and those under development. There’s one statement in the article that I find very interesting which is:
        “DIAL will make it possible to create a Web page whose presentation can be controlled by the properties of a delivery context.”
        DIAL (Device Independent Authoring Language) and the delivery context – cf. Delivery Context Client Interfaces (DCCI) and Delivery Context Ontology (DCO) – seem to be a competitor of MPEG’s Digital Item Declaration (DID), Multimedia Description Schemes (MDS), and Usage Environment Description (UED). I wonder whether there exists a thorough and, of course, not taking sides comparison of these formats and whether it is possible to harmonize at least parts thereof. For sure that’s a job for academics because these standardization bodies might not be interested in doing this let’s call it academic exercise.

        Furthermore, still a big issue is how all these assets are actually communicated over the various networks. The article mentions IETF’s “Transparent Content Negotiation in HTTP” although this RFC2295 is rather old and, more importantly, is ‘experimental’. Additionally, there exist some extensions of HTTP to carry CC/PP-based descriptions defined in RFC2774 identifying the profile or a difference to an existing profile - also ‘experimental’. However, DCO (an all others) may require yet another HTTP extension and, thus, there’s a need to decouple the communication/negotiation of the delivery context and usage environment properties from the actual (transport) protocol.

        Finally, CC/PP defines only a ‘container format’, i.e., the language constructs, while UAProf defines the actual ‘terms’, i.e., a vocabulary of actual hardware, software, … characteristics for mobile devices.

        Tuesday, February 10, 2009

        W3C Multimodal Standard Brings Web to More People, More Ways

        As part of ensuring the Web is available to all people on any device, W3C published a new standard today to enable interactions beyond the familiar keyboard and mouse. EMMA, the EMMA: Extensible MultiModal Annotation Markup Language, promotes the development of rich Web applications that can be adapted to more input modes (such as handwriting, natural language, and gestures) and output modes (such as synthesized speech) at lower cost. The document, published by the Multimodal Interaction Working Group, is part of a set of specifications for multimodal systems, and provides details of an XML markup language for containing and annotating the interpretation of user input. Read the press release and testimonials, and learn more about the Multimodal Interaction Activity.

        Wow, this development really took a long time but interesting stuff though. The first WD dates back to August 2003 which makes me wonder whether somebody will use/implement that. Nevertheless, it can be used for universal multimedia access w.r.t. multimodal interaction to the content. Other activities of this working group comprise Multimodal Architectures and Interfaces and Ink Markup Language (InkML) both at WD stage. Hope its development does not last forever...

        Tuesday, January 6, 2009

        TV anywhere, anytime ...

        When reading the article on TechCrunch about watching cable TV on your iPhone, brings me back to the vision of UMA - my first blog entry here. In fact, with Sling Media's Set-Top Box (STB) you can stream TV/video content to your computer within your house. Now they'll announced to approach mobile devices at this week’s Macworld:
        SlingPlayer Mobile gives consumers their entire home TV experience, including local channels, local sports teams, video on demand, pay per view, etc. Any program that you can watch on your sofa back home, you can now watch via your iPhone using a standard network connection. In addition, SlingPlayer Mobile for iPhone users can also control their home digital video recorder (DVR) to watch recorded shows, pause, rewind, and fast forward live TV, or even queue new recordings while away from home.
        The application will be submitted to Apple some time in Q1 but on their homepage the SlingPlayer Mobile for Blackberry is already available as public beta right now - also for Windows Smartphone, Windows Mobile Pocket PC, Palm, and Symbian. Well, one could state that Apple's iPhone is the last one in the queue but I leave this debate open to everyone...

        For me, this means that watching "my TV content" anywhere and anytime is becoming reality. However, I don't have a Sling Media STB and I also wonder about the legal aspects in European countries, etc. Furthermore, I would be interested whether it is possible to stream any content - not only TV content - to your mobile devices as film industry is starting to attach Digital Copies to their DVDs and Blu-Ray Discs respectively.

        Wednesday, July 16, 2008

        Keynote@TEMU'08: MPEG-21 for Integrated E2E Management enabling QoS

        At today's opening day of TEMU'08 I was pleased to give a keynote on "The MPEG-21 Multimedia Framework for Integrated Management of Environments enabling Quality of Service".
        A summary is given below. Interestingly, a question from the auditorium was whether these description formats are somehow used by the IP Multimedia Subsystem (IMS). The answer was no as there are currently no standardized means to transport MPEG metadata (this is true for both MPEG-7 and MPEG-21) using IETF-based protocols. However, there have been projects transporting these metadata formats, e.g., using HTTP, SDP(ng), RTP, etc., but this has not been recognized by the IETF so far or proponents of these implementations were not able to bring them to the appropriate working groups of IETF (for whatever reasons...).

        Summary/Abstract of my talk [PDF]
        The information revolution of the last decade has resulted in a phenomenal increase in the quantity of content (including multimedia content) available to an increasing number of different users with different preferences who access it through a plethora of devices and over heterogeneous networks. End devices range from mobile phones to high definition TVs, access networks can be as diverse as GSM and broadband networks, and the various backbone networks are different in bandwidth and Quality of Service (QoS) support. In addition, users have different content/presentation preferences and intend to consume the content at different locations, times, and under altering circumstances.

        In order to become the vision as indicated above reality substantial research and standardization efforts have been undertaken which are collectively referred to as Universal Multimedia Access (UMA). An important and comprehensive standard in this field is the MPEG-21 Multimedia Framework, formally referred to as ISO/IEC 21000. The aim of MPEG-21 is to enable transparent and augmented use of multimedia resources across a wide range of networks, devices, user preferences, and communities, notably for trading (of bits). In particular, it shall enable the transaction of Digital Items among Users. A Digital Item is defines as a structured digital object with a standard representation and metadata and is the fundamental unit of transaction and distribution within the MPEG-21 multimedia framework. A User (please note the upper case “U”) is defined as any entity that interacts within this framework or makes use of Digital Items. The MPEG-21 standard currently comprises 17 parts which can be clustered into six major categories each dealing with different aspect of the Digital Items: declaration (and identification), digital rights management, adaptation, processing, systems, and miscellaneous aspects (i.e., reference software, conformance, etc.). The talk will present and review these concepts with the emphasize on providing universal access to multimedia contents independent of the User's location, time, and other usage environment conditions.

        Several projects funded by the European Commission (EC) – among them are DANAE and ENTHRONE (blog) worth to mention – have implemented and integrated (parts of) the MPEG-21 standard in order to demonstrate its feasibility. The aim of the DANAE Specific Targeted Research Project (STREP) was to develop scalable coding formats and an MPEG-21-based end-to-end architecture comprising a server, client, and adaptation node (all MPEG-21-enabled) which allows for dynamic and distributed adaptation of scalable media formats. On the other hand, the objectives of the ENTHRONE Integrated Project (IP) are to provide an integrated management solution enabling QoS within heterogeneous environments based on MPEG-21 and to demonstrate the ENTHRONE solution in a large-scale pilot. Therefore, the talk will review ENTHRONE's contribution to the UMA issue and will demonstrate how the MPEG-21 concepts are adopted on a broader scale.

        Wednesday, July 2, 2008

        Interoperable SVC Streaming featuring MPEG-21 DIA at ICME'08 Demo Track

        Abstract: In this paper we present an interoperable multimedia delivery framework for scalable video coding based on MPEG-21 Digital Item Adaptation (DIA). In can be used to transmit scalable video contents within heterogeneous usage environments where the properties of the usage environment (e.g., terminal/network capabilities) may change dynamically during the streaming session. The usage environment is signaled by interoperable description formats provided by the DIA standard. Additionally, the adaptation itself is done by exploiting the standard's generic adaptation approach, i.e., independent of the actual coding format. Thus, the overall framework is also applicable for other scalable coding formats.
        Interestingly, the demo provides thanks to the optimizations of the reference software a superior performance compared to other SVC player implementations. A lot of people were impressed that the sequences were shown smoothly on a "casual" laptop without any delays or faulty pictures.

        The 2-pager (demo description) is available here.

        Friday, March 7, 2008

        European Research Project to Shape Next Generation Internet TV

        The text has been adopted from the official P2P-Next press release.

        Brussels, 19 February 2008 - P2P-Next, a pan-European conglomerate of 21 industrial partners, media content providers and research institutions, has received a €19 million grant from the European Union. The grant will enable the conglomerate to carry out a research project aiming to identify the potential uses of peer-to-peer (P2P) technology for Internet Television of the future. The partners, including the BBC, Delft University of Technology, the European Broadcasting Union, Lancaster University, Klagenfurt University, Markenfilm, Pioneer Digital Design Centre Limited and VTT Technical Research Centre of Finland, intend to develop a Europe-wide “next-generation” internet television distribution system, based on P2P and social interaction.

        P2P-Next Statement

        "The P2P-Next project will run over four years, and plans to conduct a large-scale technical trial of new media applications running on a wide range of consumer devices. If successful, this ambitious project could create a platform that would enable audiences to stream and interact with live content via a PC or set top box. In addition, it is our intention to allow audiences to build communities around their favourite content via a fully personalized system.

        This technology could potentially be built into Video on Demand (VOD) services in the future and plans are underway to test the system for major broadcasting events.

        The project has an open approach towards sharing results. All core software technology will be available as open source, enabling new business models. P2P-Next will also address a number of outstanding challenges related to content delivery over the internet, including technical, legal, regulatory, security, business and commercial issues."

        The complete list of Partners is:

        What is Peer-to-Peer (P2P) technology

        P2P provides an alternative to the traditional client/server architecture of computer networks and signifies the next big step in the evolution of internet media delivery. While employing the existing broadband networks, each participating computer, referred to as peer, functions as both a client and a server for a given application. A P2P network enables the sharing of content files or streams with audio, video and data content. Today it is considered increasingly as a potentially efficient and reliable mechanism for distributing any media to the general public worldwide.

        P2P-Next in a nutshell

        P2P-Next will develop an open source, efficient, trusted, personalized, user-centric and participatory television and media delivery mechanism with social and collaborative connotations using the emerging P2P paradigm, which takes into account the existing EU legal framework.

        Links

        Tuesday, February 12, 2008

        What's new with MPEG: Representation of Sensory Effects (RoSE)

        This article covers a brave new topic currently being discussed within the Moving Picture Experts Group (MPEG), formally known as ISO/IEC JTC 1/SC 29/WG 11. It comprises a framework for the Representation of Sensory Effects (RoSE) and its context and objectives are described in [1]. Discussions take place via [3] and is open to everyone.

        Context and Objectives

        The aim of RoSE is to provide a standardized framework which shall allow for controlling sensory effects in order to increase the overall experience of the user when consuming multimedia content. A sensory effect is an effect to augment feelings by stimulating sensory organs in a particular sense of a multimedia application. Examples of sensory effects include - but are not limited to - audio (sound), visual (video), tactile (light, shading, vibration, wind, temperature, fog, water), and olfactory (scent).

        Scope of Standardization

        The scope of standardization can be divided into description schemes (or sometimes called tools) and systems aspects.

        The former - description schemes/tools - covers three topics, namely:

        • Sensory Effect Schema (SES): It defines the description schemes and descriptors to represent Sensory Effects
        • Sensory Device Schema (SDS): It defines the description schemes and descriptors to represent characteristics of Sensory Devices.
        • User Sensory Preferences (USP): It defines the description schemes and descriptors to represent the user preferences on the (rendering) of the sensory effects.

        The later - systems aspects - covers two topics, namely:

        • Storage (File Format): It defines an application format which is used to package the audiovisual contents and their associated RoSE metadata.
        • Transport (Delivery Format): It defines a schema of message forms by which a part (fragment) of RoSE metadata can be delivered. The specific delivery protocols may be out of the standardization scope.

        Timeline

        • April 2008: "Final" Requirements and Call for Proposals
        • July 2008: Evaluation of proposals and Working Draft (WD)
        • October 2008: Committee Draft (CD)
        • April 2009: Final Committee Draft (FCD)
        • October 2009: Final Draft International Standard (FDIS)

        References

        [1] Context and Objectives, http://www.chiariglione.org/MPEG/working_documents/explorations/RoSE/RoSE-C&O.zip

        [2] Draft Requirements for RoSE, http://www.chiariglione.org/MPEG/working_documents/explorations/RoSE/RoSE-Reqs.zip

        [3] RoSE AhG Reflector, http://lists.uni-klu.ac.at/mailman/listinfo/rose

        Saturday, February 24, 2007

        Multimedia Communication for Universal Media Access (UMA)

        Introduction

        Universal (Multi-)Media Access (UMA) is a buzzword which is well known within the respective community since about 1999 when J.R.Smith et.al. brought it up in one of the first arcticels of IEEE Transactions on Multimedia. With the InfoPyramid a concept was born which allows one to provide different versions or even modalities of the same multimedia content for different end user terminals, network characteristics, or user preferences.

        UMA's slogan is to "provide access to advanced multimedia content anywhere, anytime, and with any kind of device".

        The aim of this blog is to provide news, trends, and relevant information regarding multimedia communication and related standardarization actitivities, e.g., IETF, W3C, and MPEG.

        Adaptation: Why and How?

        Given the amount of multimedia formats, heterogeneous network infrastructure, and the plethora of devices gaining access to multimedia content encoded by using these formats, it becomes naturally that some kind of mechanism is required which provides a solution to accomplish this mismatch. This mechanism or concept is referred to as adaptation.

        The Concept of UMA

        The adaptation paradigms that evolved during the last decade can be clustered into three categories:

        • Adaptation by Selection uses pre-stored versions of different qualities and/or modalities of the same multimedia content and delivers the most suitable version depending on the usage environment where the content is going to be consumed.
        • Adaptation by Transcoding/Transmoding provides an "on-the-fly adapted" version of the original multimedia content according to the different contexts of multimedia consumption.
        • Adaptation by Transformation facilitates scalable coding formats and bitstream syntax descriptions enabling the adaptation of the multimedia content by performing simple truncation operations and minor update operations in a coding format-independet way.

        The above paradigms have, of course, their pros and cons and those who are further interested should check out the articels and links in the next section.

        Further reading and links:

        [1] R. Mohan, J.R. Smith and C.-S. Li, et.al, "Adapting Multimedia Internet Content for Universal Access," IEEE Trans. Multimedia, vol. 1, no. 1, pp. 104-114, March 1999.

        [2] C. Timmerer and H. Hellwagner, "MPEG Standards enabling Universal Multimedia Access", 1st International Conference on Automated Production of Cross Media Content for Multi-channel Distribution, Florence, Italy, November/December, 2005.

        [3] C. Timmerer and H. Hellwagner, "Interoperable Adaptive Multimedia Communication, IEEE Multimedia Magazine", vol. 12, no. 1, pp. 74-79, January-March 2005.

        [4] IEEE MultiMedia Magazine

        [5] IEEE Transactions on Multimedia

        [6] IEEE Transactions on Circuits and Systems for Video Technology

        [7] ACM Multimedia Systems