Point-of-interest determination and display
For point-of-interest determination and display, a processor detects an image event during a videoconference. The processor determines a point-of-interest for the video image of the videoconference based on the image event. The video image is at least a 180-degree image and the point-of-interest is a portion of the video image. The processor displays the point-of-interest from the video image on a display.
1. An apparatus comprising:
a display;
a processor;
a memory that stores code executable by the processor to:
detect an image event during a video conference;
determine a point-of-interest for a video image of the video conference based on the image event, wherein the video image is at least a 180-degree image the point-of-interest is a portion of the video image and is determined using a neural network trained on a training data set comprising scene compositions and classifications, the scene compositions comprising images of objects and participants performing actions, the actions comprising looking at an object, standing, and/or being in an arrangement, and the classifications classifying the objects and/or actions; and
display the point-of-interest from the video image on the display.
2. The apparatus of claim 1 , wherein the point-of-interest is parsed from the video image.
3. The apparatus of claim 1 , wherein the training data set further comprises objects and actions.
4. The apparatus of claim 1 , wherein the image event comprises a movement by a speaker and the point-of-interest is the speaker.
5. The apparatus of claim 1 , wherein the image event comprises speech by a speaker and the point-of-interest is the speaker.
6. The apparatus of claim 1 , wherein the image event comprises a voice command that specifies an object selected from the group consisting of a name, a point-of-interest, and an action and the point-of-interest is the object.
7. The apparatus of claim 1 , wherein the image event is a user interface input from a user interface and the point-of-interest is determined from the user interface input.
8. A method comprising:
detecting, by use of a processor, an image event during a video conference;
determining a point-of-interest for a video image of the video conference based on the image event, wherein the video image is at least a 180-degree image, the point-of-interest is a portion of the video and is determined using a neural network trained on a training data set comprising scene compositions and classifications, the scene compositions comprising images of objects and participants performing actions, the actions comprising looking at an object, standing, and/or being in an arrangement, and the classifications classifying the objects and/or actions; and
displaying the point-of-interest from the video image.
9. The method of claim 8 , wherein the point-of-interest is parsed from the video image.
10. The method of claim 8 , wherein the training data set further comprises objects and actions.
11. The method of claim 8 , wherein the image event comprises a movement by a speaker and the point-of-interest is the speaker.
12. The method of claim 8 , wherein the image event comprises speech by a speaker and the point-of-interest is the speaker.
13. The method of claim 8 , wherein the image event comprises a voice command that specifies an object selected from the group consisting of a name, a point-of-interest, and an action and the point-of-interest is the object.
14. The method of claim 8 , wherein the image event is a user interface input from a user interface and the point-of-interest is determined from the user interface input.
15. A program product comprising a non-volatile computer readable storage medium that stores code executable by a processor, the executable code comprising code to:
detect an image event during a video conference;
determine a point-of-interest for a video image of the video conference based on the image event, wherein the video image is a 180-degree image, the point-of-interest is a portion of the video image and is determined using a neural network trained on a training data set comprising scene compositions and classifications, the scene compositions comprising images of objects and participants performing actions, the actions comprising looking at an object, standing, and/or being in an arrangement, and the classifications classifying the objects and/or actions; and
display the point-of-interest from the video image.
16. The program product of claim 15 , wherein the point-of-interest is parsed from the video image.
17. The program product of claim 15 , wherein the training data set further comprises objects and actions.