IP Library Granted Patent US 9,769,424
Granted Patent B2
US 9,769,424 · App. 15/030,942 · Granted Sep 19, 2017

Arrangements and method thereof for video retargeting for video conferencing

Inventor: Julien Michot (Sundbyberg, SE)
Assignee: Telefonaktiebolaget LM Ericsson (publ)
H04N7/15G06K9/00335G06K9/52G06K9/6267G06T7/20G06T7/60H04N5/265H04N7/142H04N7/147G06K2009/4666
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,769,424
App. No.
15/030,942
Granted
Sep 19, 2017
Kind
B2
Abstract

According to embodiments of the present invention, sound localization is used to determine the active speaker in a video conference. A network element uses the localization to determine which regions of the image that should be preserved and retargets the video accordingly. By providing a retargeted video where the speaker is more visible, a better user experience is achieved.

Claims (64)

1. A method for video conferencing to be performed by a network element, wherein sound localization is used to determine the at least one active speaker, the method comprises:

using the active speaker location to determine image regions to preserve

creating a preserving map with the areas of the image that should be preserved, and

retargeting the video based on the preserving map, wherein the video retargeting is a nonlinear video retargeting method.

2. The method according to claim 1 wherein also face detection is used to determine image regions used to create the preserving map.

3. The method according to claim 1 wherein face detection is used to determine all participants used to create the preserving map.

4. The method according to claim 1 , receiving via a user interface at least one region used to create the preserving map.

5. The method according to claim 1 wherein a database is used to keep track of detected persons.

6. The method according to claim 5 wherein most active speaker is given a higher preserving value.

7. The method according to claim 5 wherein most recent speaker is given a higher value.

8. The method according to claim 1 , wherein a body detector is used to provide better people detection.

9. The method according to claim 1 , wherein a depth sensor is used to provide better people detection.

10. The method according to claim 1 , wherein aspect ratio adaption is used for retargeting the video to fit the receivers screen.

11. The method according to claim 1 , wherein video mixing is used for arranging several videos coming from various senders into one video containing a mix of all or parts of the incoming videos to a receiver.

12. The method according to claim 1 , wherein aspect ratio adaption and video mixing adaption is used for arranging several videos into one video that fits the receivers screen.

13. The method according to claim 1 wherein temporal smoothing is used if several people speak.

14. The method according to claim 1 , wherein the retargeted video and the original video are available for the viewer to be displayed simultaneously.

15. The method according to claim 1 , wherein the preserving map is constructed using a rectangle Rs (center =(xs, ys+dy), size =(ws, hs)) with values according to these equations:

ws=fx*W body/ zs,

with Wbody=0.5 m, representing the mean chest width,

hs=fy*H trunk/ zs,

with Htrunk=0.6 m, representing the mean trunk height and

dy=fy*Hc/zs,

with Hc=Htrunk/2 −H face/2 with Hface=0.25 m being the average head height, He represents a distance between the body half and the head mouth, dy is the same distance as Hc but converted to pixels, hs is the average trunk height expressed in pixel, zs is the speaker depth, fx is camera focal length on x-axis, and fy is camera focal length on y-axis.

16. The method according to claim 1 , wherein the preserving map

P ( x,y ) = W max * exp(−0.5* (( x−xs )/σ( y )) 2)

is constructed using

σ(y)=σhead for all y<yhead+hs/4 and

σ(y)=σbody, otherwise

wherein

P(x,y) is the preservation map value at 2D position (x,y)

Wmax is a maximum weight value

σ(y) is the Gaussian standard deviation

σhead is a Gaussian standard deviation suited for the head

σbody is a Gaussian standard deviation suited for the body

xs is the located speaker position on x axis

hs is the average trunk height expressed in pixel

yhead is the speaker head y location.

17. The method according to claim 1 , wherein the depth of people is derived from the relation

zs=fx * W face / WFS

wherein

Wface is the mean face width, e.g. 15 cm

fx is the camera focal length on the x axis

zs the approximated depth

WFS the rectangle width given by the face detector.

18. The method according to claim 1 , wherein the rectangle Rs is defined based on rectangle Fs using the equation

Rs center=( xFS,yFS+dy ) and Rs size=( sx*WFS,sy*HFS ),

wherein

xFS is the position of the center of the rectangle given by the fact detector on the x axis

yFS is the position of the center of the rectangle given by the fact detector on the y axis

dy, same as Hc but expressed in pixel

sx and sy are two scaling factors allowing to create a bigger rectangle based on the rectangle given by the face detector

WFS is the width of the rectangle given by the fact detector

HFS is the height of the rectangle given by the fact detector.

19. A network element for enabling video conferencing, wherein sound localization is used to determine the at least one active speaker, comprising a processor and memory, said memory containing instructions executable by said processor whereby said network element is operative to:

use the active speaker location to create a preserving map with areas of the image that should be preserved.

retarget the video based on said preserving map, wherein the network element uses a nonlinear video retargeting method.

20. The network element according to claim 19 further operative to detect faces in order to determine image regions used to create the preserving map.

21. The network element according to claim 19 further operative to receive requests from a viewer which regions of the video to display.

22. The network element according to claim 19 wherein a database is used to keep track of detected persons.

23. The network element according to claim 22 wherein most active speaker is given a higher preserving value.

24. The network element according to claim 22 wherein most recent speaker is given a higher value.

25. The network element to claim 19 , wherein a body detector is used to provide better people detection.

26. The network element to claim 19 , wherein a depth sensor is used to provide better people detection.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 21, 2016
From: MICHOT, JULIEN
To: TELEFONAKTIEBOLAGET L M ERICSSON (PUBL)
Reel/Frame 038339/0609 →
Continuity (1)
Related Publication 20160277712A1 · Sep 22, 2016