Pages

Showing posts with label Paper review. Show all posts
Showing posts with label Paper review. Show all posts

Sunday, October 13, 2019

New conference paper

Our latest work on adaptation for cloud gaming under constrained resources has been accepted for presentation at 26TH INTERNATIONAL CONFERENCE ON MULTIMEDIA MODELING. (Link)
Cloud gaming has emerged as a new trend in the gaming industry, bringing a lot of benefits to both players and service providers. In cloud gaming, it is essential to ensure low end-to-end delay for good
use experience. Hence, sufficient computational resources must be available at the client in order to process video in a timely manner. However, thin clients such as mobile devices generally have limited computation capabilities. Thus, the available computational resources may be insufficient to support the client, such as in case of low battery. In this paper, we propose a new adaptation framework for resource-constrained cloud gaming clients. The proposed framework combines frame skipping at the server and frame discarding at the client according to available computational resources of the client. Experiment results show that the proposed framework can significantly improve video quality given a delay constraint compared to conventional methods.


Sunday, July 21, 2019

New conference paper (IEEE MMSP 2019): Scalable 360 Video Streaming using HTTP/2

My latest work on 360-degree video streaming over network has been accepted for presentation at IEEE MMSP 2019.  This paper is a result of my collaborative research with Prof. Truong Thu Huong of HUST and Prof. Pham Ngoc Nam of Vin University.

Overview

360-degree video is a main content type in Virtual Reality, providing users with immersive viewing experience. In this paper, we propose a novel adaptation method for 360-degree video streaming over HTTP/2, which can provide high viewing experience to users under time-varying network conditions and time-varying user head movements. The proposed method utilizes Scalable Video Coding to solve the trade-off between network adaptivity and user adaptivity. An optimal tile layer selection
algorithm is provided. To cope with sudden throughput drops, the delivery of late layers are terminated using HTTP/2’s stream termination feature. Also, a tile layer updating scheme is proposed to deal with viewport estimation errors. Experimental results show that the proposed method can improve the average viewport bitrate by 16-17% compared to a reference method.

Key features

  • Scalable Video Coding (SVC) is utilized to tackle the trade-off between network adaptivity and viewport adaptivity.
  • The tile layer selection problem is formulated and an efficient algorithm is proposed.
  • A late tile layer termination scheme is presented that can save network resources by terminating the delivery of late tile layers using HTTP/2’s stream termination feature.
  • A tile layer updating scheme that can effectively deal with viewport estimation errors is proposed. The scheme makes use of HTTP/2’s stream priority feature.
Future work

In future work, our goal is to further improve the user experience by providing both temporally and spatially smooth viewport quality. 

Sunday, February 24, 2019

Overview of JETCAS Issue on Immersive Video Coding and Transmission (Part 1)

JETCAS issue on immersive video coding and transmission presents the latest developments in immersive video research. This blog summarizes the papers related to coding and transmission of 360-degree video, which is one of the most popular types of immersive media.

360-degree Video Coding

To provide an excellent immersive experience, 360-degree videos require extremely high resolution with high frame rate (4K/8K + 60/90 fps). As a result, 360 video require much higher bandwidth compared with conventional 2D video. Therefore, efficient compression technology is highly desirable for storage and transmission of 360 video.

The paper [1] proposes a hybrid Equirectangular-Cubemap projection that can achieve more uniform sampling and reduce the boundary artifacts across different faces. In addition, a set of coding tools that can make use of the spherical continuity in 360 video are proposed. The proposed algorithm can effectively reduce the BR-rate and the seam artifacts caused by discontinuous edge and frame boundary.

In [2], the projection format is customized based on the input video content. Especially, the hybrid angular cubemap (HAC) projection is utilized to adapt the the sampling within each face. Also, an adaptive frame packing technique is used to select the face arrangement in a frame. To alleviate the "face seam" artifacts in the rendered viewports, the relationship of samples and blocks in the spherical geometry is considered to improve intra/inter prediction and in-loop filters.

To address the deformation of video content caused by mapping from sphere to 2D plan, the paper [3] proposes a new motion model based on spherical coordinates transform. The proposed model is shown to be effective in improving the motion compensation/estimation in panoramic video coding.

360-degree Video Transmission

Though advanced coding technologies can significantly reduce the bitrate of 360 video, delivery of 360 video is still a challenging task due to limitations in network resources, as well as constraints imposed by en-user devices. Therefore, cost-effective delivery technology is necessary for the wide adoption of VR/AR applications.

In [4], we propose a server-based adaptation framework for 360 video streaming over networks. The proposed method utilizes tiling-based viewport adaptive streaming to reduce the required network bandwidth for 360 video. Also, the proposed tile selection algorithm can effectively deal with the user head movements within each video segment.

In tiling-based viewport adaptive streaming, it is important to tile the video in an effective manner. Conventionally, the video is divided into equal sized tiles. The paper [5] addresses this issue by considering the Visual Attention map. Especially, the video is divided in to non-overlapping variable sizes taking into account the Visual Attention map.

Effective viewport adaptation methods require accurate estimations of viewport positions. However, the large buffer size in HTTP Adaptive Streaming may severely affect viewport position estimation accuracy. Taking the idea of scalable video coding, [6] can reduce the client buffer size down to one segment duration by using a two-tier system. To achieve this feature, the whole video is encoded into a base tier, which is always delivered to the client, and multiple enhancement layers each corresponds to a viewport position.

In [7], the authors analyze the impact of the end-to-end delay to tile-based viewport adaptive streaming. It is found that the gain compared to viewport-independent approach drops to 8% for a delay of 1 second. To address this issue, the authors propose to combine viewport prediction with a velocity-based QP distribution.

To facilitate delivery of 360 video over wireless networks, the paper [8] proposes a pseudo-analog transmission framework called OmniCast. The proposed framework features a spherical domain power-distortion optimization framework and two adaptive block partitions algorithms. Experiment results shows that the proposed framework outperforms JPEG2000-based solution and the conventional Softcast.

The last paper [9] presents a real time 3D 360-degree telepresence system. To deal with the mismatch between the estimated and actual viewports caused by system delay, the proposed system uses cameras with a larger field of view than the visual field of the user. The level of delay compensation is improved with Gate Recurrent Units (GRU)-based head-motion prediction method.

References
[1] J. Lin et al., "Efficient Projection and Coding Tools for 360° Video," doi: 10.1109/JETCAS.2019.2899660
[2] P. Hanhart, X. Xiu, Y. He and Y. Ye, "360-degree Video Coding based on Projection Format Adaptation and Spherical Neighboring Relationship," doi: 10.1109/JETCAS.2018.2888960
[3] Y. Wang, D. Liu, S. Ma, F. Wu and W. Gao, "Spherical Coordinates Transform-Based Motion Model for Panoramic Video Coding," doi: 10.1109/JETCAS.2019.2896265
[4] D. V. Nguyen, H. T. T. Tran, A. T. Pham and T. C. Thang, "An Optimal Tile-based Approach for Viewport-adaptive 360-degree Video Streaming," doi: 10.1109/JETCAS.2019.2899488
[5] C. Ozcinar, J. Cabrera and A. Smolic, "Visual Attention-Aware Omnidirectional Video Streaming Using Optimal Tiles for Virtual Reality," doi: 10.1109/JETCAS.2019.2895096
[6] L. Sun et al., "A Two-Tier System for On-Demand Streaming of 360 Degree Video over Dynamic Networks," doi: 10.1109/JETCAS.2019.2898877
[7] Y. Sanchez, G. S. Bhullar, R. Skupin, C. Hellge and T. Schierl, "Delay Impact on MPEG OMAF’s tile-based viewport-dependent 360° video streaming," doi: 10.1109/JETCAS.2019.2899516
[8] J. Zhao, R. Xiong and J. Xu, "OmniCast: Wireless Pseudo-Analog Transmission for Omnidirectional Video," doi: 10.1109/JETCAS.2019.2898750
[9] T. Aykut, J. Xu and E. Steinbach, "Realtime 3D 360-degree Telepresence with Deep-learning-based Head-motion Prediction," doi: 10.1109/JETCAS.2019.2897220


 




Friday, January 18, 2019

Paper review (Jan. 2019)

1. Cooperative Tile-based 360-degree Panoramic Streaming in Heterogeneous Networks using Scalable Video Coding
Xiaoyi Zhang, Xinjue Hu, Ling Zhong, Shervin Shirmohammadi, Lin Zhang
DOI 10.1109/TCSVT.2018.2886805

Scenario
  • A group of mobile users are watching a 360 video
  • The users are physically close enough to form a Mobile Ad hoc Network (MANET)
Problem
  • Multiple non-cooperative (independent) streaming clients are likely suffering from network congestion and quality fluctuation
  • The naive cooperative scheme would result in high redundancy
Research question
  • How to design a cooperative streaming method that can reduce the redundancy and increase the QoE for the group of users? 
Proposed method
  • The tiles are encoded using Scalable Video Coding (SVC) standard into multiple quality layers.
  • A tile version (i.e., SVC layer) will be downloaded from a specific user via the user's cellular network, then be shared  to other users via the MANET

_________________________________________________________

2. HTTP/2-Based Streaming Solutions for Tiled Omnidirectional Videos
Orange Labs, France
2018 IEEE ISM DOI: 10.1109/ISM.2018.00023

Scenario
  • A single user watching a 360 video over the network
  • Tiling-based Viewport Adaptive Streaming is used for video transmission
Research question
  • How to deal with errors in viewport positions estimation?
Proposed Method
  • For each video segment:
    • Step 1: Estimate the future viewport positions, decide the tiles' versions, then send requests for to the server.
    • Step 2: Re-estimate the future viewport positions, update the tiles' version using HTTP/2's stream termination and stream priorities features. 
_________________________________________________________

3. Efficient Live and on-Demand Tiled HEVC 360 VR Video Streaming
ForzaSys AS, Norway
2018 IEEE ISM DOI: 10.1109/ISM.2018.00022

Scenario
  • Live streaming system for 360 video
Research Question
  • How to design an effective live streaming system for 360 video?
 Proposed Architecture
  • Using tiling feature of HEVC standard to combine multiple tiles into a single HEVC-compliant video
  • Combining RTP and DASH. RTP is used for live streaming and broadcast, whereas DASH supports on-demand case
  • Single HEVC hardware decoder
Performance
  • Achieving a frame rate of >30fps for both 2K and 6K videos on a Samsung Galaxy S7 in a Wifi network
_________________________________________________________

4. Edge-Assisted Rendering of 360-degree Videos Streamed to Head-Mounted Virtual Reality  
National Tsing Hua University 
2018 IEEE ISM DOI: 10.1109/ISM.2018.00016

Scenario
  • Streaming 360 videos to heterogeneous Head-Mounted Display (HMD) devices
Problem
  • Decoding 360 video in real-time requires high computation powers (e.g., GPU). This make it difficult for lightweight HMDs to process 360 video
  • 360 videos consume a lot of network bandwidth  
Research Question
  • How to effectively reduce bandwidth consumption and support heterogeneous HMD devcies?
Proposed Method 
  • Tiling-based streaming to reduce bandwidth consumption
  • Viewport rendering is performed by edge servers to reduce the computational workload on the HMDs
  • An optimization framework to determine which HMDs should be severed by the edge servers. 
_________________________________________________________

5. Optimal Multi-Quality Multicast for 360 Virtual Reality Video  
Shanghai Jiao Tong University 
arXiv:1901.02203

Scenario
  • Multiple users are watching a 360-degree video
Problem
  • Unicast streaming results in high redundancy, e.g., a tile visible by 2 users will be requested 2 times.
Proposed Method
  • Time Division Multiple Access (TDMA)-based transmission
  • multicast for overlapping tiles among users
  • unicast for other tiles
_________________________________________________________

6. A Robust Algorithm for Tile-based 360-degree Video Streaming with Uncertain FoV Estimation 
Indiana University
arXiv:1812.00816v1

Scenario
  • A single user watching a 360 video over the network
  • Tiling-based Viewport Adaptive Streaming is used for video transmission
Problem
  • How to deal with uncertainty (i.e., errors) in FoV (i.e., viewport position) estimation
Proposed Method
  • Utilizing the viewing probability of different tiles
  • Ensuring the probability that the streaming rate is below a certain level 


Tuesday, January 8, 2019

New Journal Paper: A Client-based Adaptation Framework for 360-degree Video Streaming


In our latest work, a new client-based adaptation framework for 360-degree video streaming is proposed. Our framework can support different application scenarios. Especially, we introduce for the first time the use of bitrate and quality estimation in viewport adaptive streaming of 360 videos. The key findings from our work are as follows.
  • The use of estimated bitrate/quality improves the viewport quality significantly 
  • The proposed method performs as if knowing full information of bitrate/quality 
  • The proposed method can be applied to both Equirectangular and Cubemap projections 
  • Cube performs slightly better than Equirectangular 
  • Long client buffering could have severe impacts on the visual quality in VR
You can find the full paper of our work here.

What is 360-degree video?

360-degree video is one of the key element of Virtual Reality. A 360-degree video captures 360-degree view of a scene. Thus, different from conventional videos, you can freely change your viewing direction while watching 360 videos. This provides the so-called "immersive experience" to the user. You can try 360 videos at YouTube Virtual Reality channel (link). 360 video is being used in a wide range of applications such as gaming, advertising, training, e-learning.

High bitrate: the key challenge of 360 video streaming

To cover the full 360-degree view, 360 video has much higher resolution than conventional videos (at least 6 times). Moreover, 360 video is usually watched on Head-Mounted Display (e.g., Occulus Rift) where the display is much closer to your eyes than in the cases of viewing on TVs or computer. For satisfactory user viewing experience, 360 video should have a resolution of 4K or higher, results in very high bitrate.

Viewport Adaptive Streaming

To cope with the high bitrate of 360 video, Viewport Adaptive Streaming (VAS) has been proposed. The basic idea is to deliver the video parts visible to the user at a high quality while delivering the remaining video parts at a lower quality. VAS can be realized using tiling-based approach or viewport-dependent coding approach.


Sunday, December 17, 2017

Overview of IEEE ISM 2017

The 19th International Symposium on Multimedia was held at the Splendor Hotel, Taichung city, Taiwan from Dec. 11 to Dec. 13, 2017. This year, the conference covers broad and diverse topics of multimedia computing, which includes the following main topics:
  • 360-degree video and image
Immersive media such as 360-degree videos is becoming more and more important, being supported by YouTube, Facebook and other streaming platforms. The first keynote speech by Prof. Girod from Standford University gives a very nice overview of immersive video for Head-Mounted Displays. His research group is now focusing on generation of stereoscopic, 6 Degree of Freedom (6DoF) immersive video content. There are four papers addressing different problems of 360-degree video delivery. Our paper proposes a novel adaptation approach for viewport-adaptive streaming of 360-degree video. The second paper studies the optimal encoding ladders for tiled 360-degree video. The problem is formulated as an optimization problem that considers not only the distortion but also the system resources such as storage cost. However, their solution is not specific to 360-degree video. The third paper performs a perceptual analysis of perspective projection for Viewport rendering in 360-degree image. The analysis focuses on two projection-related parameters: 1) the distance from the projection center and the video center and 2) the Field of View. Yet, the way they carry out the subjective test is not appropriate as the viewers watch the content on a flat screen instead of the HMD. The fourth paper studies three QoE aspects which are immersion, interaction, and Visual Quality in Interactive 3D Tele-Immersion applications. For that purpose, the authors have designed a penalty shootout game and carried out subjective tests. The results show that using HMDs such as Oculus Rift results in better user experience compared to a third person view on 3D TV. Also, the immersion and interaction are very important to the user experience.
  • Learning
Understanding multimedia content is a crucial task in many applications such as camera surveillance. Many papers apply deep learning for crowd scene understanding, visual relationship recognition (e.g., text-to-image translation for robot), human action classification, and automatic classification of microstructures in thermal barrier coating images, real-time annotations of motion data stream, 3D action recognition. The use of convolutional neuron network (CNN) for non-reference Image Quality Assessment (blind IQA) is also proposed. There is a very interesting paper that proposes a compression framework for deep learning models. Their results show that using a simple quantization combining with arithmetic coding can reduce the bitrate by 92% with minimal impact on the accuracy.
  •  Retrieval, recommendation, and summarization
There are two papers regarding personalized video recommendation. The first paper applies machine learning to 1) identify users behind a shared account and 2) predict each user’s preference based on ‘contextual information’. The second paper proposes a framework to automatically select features on Factorization Machine based Context-aware recommendation systems. As for summarization, a new summarization method for blog articles using image-text alignment techniques has been proposed. Another paper proposes a new approach for automatic summarization of video collections that leverages a structured minimum-risk classifier and efficient submodular inference.
  • Visual Aspects
The first paper presents a method for automatically detecting a good surface in a daily living and working space to support improvisatory projection without a pre-installed projection surface. The second paper proposes a new framework to convert textual instructions into coherent visual descriptions (text instructions annotated with images).
  • Video Streaming
The first paper proposes to modify Peer-to-Peer Streaming Peer Protocol (PPSPP) to support streaming over Wifi P2P connection. The second paper proposes a SDN-enabled optimization-based scheme for optimally sharing the bandwidth among network flows within a residential gateway. The scheme targets online game flows and try to provide them with a higher QoE while not starving other traffic flows. The third paper presents a new scheme that limits energy consumption in a transcoding system. The fourth paper introduced a Dynamic Rate Controller (DRC) for conversational video streaming applications, especially for HDVC. DRC uses the novel concept of future budget, plus a window based bitrate history to adjust to bandwidth changes faster and with higher quality than other rate controllers.


The best paper award is awarded to a paper investigating the quantitative determinants of film mood across different types of scenes. The film scenes are classified by their location, time of day, and their use of dialogue and music. It is found that the mood ratings and their quantitative determinants differed across the scene types. There is also several demos: Deep learning based throughput estimation, real-time pattern recognition.

Sunday, November 26, 2017

Notes from Big Data Networking Symposium


In Nov. 26, I attended the third Big Data Symposium organized at the University of Aizu. Here are some notes from the presentations.

  • The presentations focus on methods for generating/processing/delivering Big Data. This is very necessary because huge amount of data are being generated by Internet-connected devices.
  • As for content generation, a method for constructing 3D videos from 2D videos using motion parallax, which states that among objects moving at the same speed, the remote ones seem to move at lower speed, is presented. The same parallax is also applied for mean-shift detection and segmentation of images.
  • As for content processing, in-memory computing, a new architecture that can significantly reduce the processing time is presented. The basic idea is to store the data in the memory (RAM) so that the data can be accessed much more quicker than the traditional way of accessing HDD, thereby reduce the processing time. To support in-memory computing, a new kind of memory, called storage class memory (SCM), has been developed. As the name suggests, SCM have very huge capacity (can up to several terabytes) that can store all application data in it.
  • As for content delivery, one of the key challenge is how to delivery the data of different applications with different requirements. For example, real-time applications such as searching, streaming video usually have very stringent requirement on the delay. To address this challenge in the presence of Big Data, Fog computing architecture has been proposed. The basic idea of Fog computing is to put additional processing nodes closer to the users so that 1) the delay can be reduced and 2) the load of the central server can be reduced too. For example, the additional nodes can perform some pre-processing to eliminate redundancy in the raw data before sending to the central server for processing.

Lời nhắn gửi tân sinh viên

Phòng kí túc xá của cậu có 12 người đến từ nhiều nơi: Lào Cai, Quảng Ninh, Bắc Ninh, Hà Nam, Nam Định, Nghệ An. Ngày đầu tiên gặp các bạn cù...