IP Library › Granted Patent US 12,051,080
Granted Patent B2
US 12,051,080 · App. 18/095,639 · Granted Jul 30, 2024

Virtual environment-based interfaces applied to selected objects from video

Inventors: Garry Anthony Smith (Sydney, AU); Zachary Oakes (Kochi, JP); Steven Dennis Flinn (Sugar Land, TX)
Assignee: Revealit Corporation
G06Q30/02G06N3/02G06T19/006G06V10/774G06V10/7788G06V10/82G06V20/20G06V20/40G06V20/41G09B5/065G06V10/255G06V10/422
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,051,080
App. No.
18/095,639
Granted
Jul 30, 2024
Kind
B2
Abstract

A method and system for virtual environment-based interfaces applied to selected objects from video directs a system's focus of attention to an image within a first video stream and identifies an object in the image by applying a trained neural network. In response to a communication from a user comprising language and/or images describing a virtual environment, a second trained neural network is applied to generate a second video stream that embodies the identified object within a virtual environment that is in accordance with the user-described virtual environment. The second video stream is then delivered to the user. The system's focus of attention and/or generation of the virtual environment may be informed by user preferences that are inferred from user behaviors.

Claims (38)

1. A computer-implemented method, comprising:

receiving a first video stream comprising a first sequence of images;

directing a focus of attention of a first computer-implemented system to a first image in the first sequence of images;

identifying a representation of an object by interpreting a plurality of pixels that are within the first image by applying a first computer-implemented trained neural network;

receiving a communication from a user requesting that the first computer-implemented system embody the identified object within a virtual environment described by the user;

generating, in response to the communication from the user, a second video stream comprising the identified object embodied within a computer-generated virtual environment that is in accordance with the user-described virtual environment, wherein the second video stream is generated by applying a second computer-implemented trained neural network; and

delivering the second video stream to the user.

2. The method of claim 1 , further comprising directing the focus of attention of the first computer-implemented system to the first image, wherein the directing of the focus of attention is in accordance with an inference that is derived from a communication performed by the user.

3. The method of claim 1 , further comprising identifying the representation of the object and communicating a plurality of attributes associated with the identified object in a natural language format to the user.

4. The method of claim 1 , further comprising receiving the communication from the user, wherein the user comprises a second computer-implemented system.

5. The method of claim 1 , further comprising receiving the communication from the user requesting that the computer-implemented system embody the identified object within the user-described virtual environment, wherein the user's description of the user-described virtual environment comprises one or more images provided by the user.

6. The method of claim 1 , further comprising generating the second video stream comprising the identified object embodied within the computer-generated virtual environment, wherein the generating of the second video stream is in accordance with a preference of the user that is inferred from a plurality of user behaviors.

7. The method of claim 1 , further comprising generating the second video stream by applying the second computer-implemented trained neural network, wherein the second computer-implemented trained neural network operates as a generative adversarial neural network (GAN).

8. A computer-implemented system comprising one or more processor-based devices configured to:

receive a first video stream comprising a first sequence of images;

direct a focus of attention of a first computer-implemented system to a first image in the first sequence of images;

identify a representation of an object by interpreting a plurality of pixels that are within the first image by application of a first trained computer-implemented neural network;

receive a communication from a user requesting that the first computer-implemented system embody the identified object within a virtual environment described by the user;

generate, in response to the communication from the user, a second video stream comprising the object embodied within a computer-generated virtual environment that is in accordance with the user-described virtual environment, wherein the second video stream is generated by applying a second computer-implemented trained neural network; and

deliver the second video stream to the user.

9. The system of claim 8 , further comprising directing the focus of attention of the first computer-implemented system to the first image, wherein the directing of the focus of attention is in accordance with a preference of the user that is inferred from a plurality of user behaviors.

10. The system of claim 8 , further comprising directing the focus of attention of the first computer-implemented system to the first image, wherein the directing of the focus of attention is in accordance with an automatic interpretation of audio that is integrated with the first video stream.

11. The system of claim 8 , further comprising identifying the representation of the object and communicating a plurality of attributes of the identified object in a natural language format to the user.

12. The system of claim 8 , further comprising receiving the communication from the user, wherein the user comprises a second computer-implemented system.

13. The system of claim 8 , further comprising receiving the communication from the user requesting that the computer-implemented system embody the identified object within the user-described virtual environment, wherein the user's description of the user-described virtual environment comprises natural language.

14. The system of claim 8 , further comprising generating the second video stream comprising the identified object embodied within the computer-generated virtual environment, wherein the generating of the second video stream is in accordance with a preference of the user that is inferred from a plurality of user behaviors.

15. A computer-implemented method, comprising:

receiving a first video stream comprising a first sequence of images;

directing a focus of attention of a first computer-implemented system to a first image of the first sequence of images;

identifying a representation of an object by interpreting a plurality of pixels that are within the first image and that are within at least one other image that temporally precedes the first image in the first sequence of images by applying a first computer-implemented trained neural network;

receiving a communication from a user requesting that the first computer-implemented system embody the identified object within a virtual environment described by the user;

generating, in response to the communication from the user, a second video stream comprising the object embodied within a computer-generated virtual environment that is in accordance with the user-described virtual environment, wherein the second video stream is generated by applying a second computer-implemented trained neural network; and

delivering the second video stream to the user.

16. The method of claim 15 , further comprising directing the focus of attention of the first computer-implemented system to the first image, wherein the directing is in accordance with an inference that is derived from a communication performed by the user.

17. The method of claim 15 , further comprising identifying the representation of the object and communicating a plurality of attributes of the identified object in a natural language format to the user.

18. The method of claim 15 , further comprising receiving the communication from the user, wherein the user comprises a second computer-implemented system.

19. The method of claim 15 , further comprising receiving the communication from the user requesting that the computer-implemented system embody the identified object within the user-described virtual environment, wherein the user's description of the user-described virtual environment comprises one or more images provided by the user.

20. The method of claim 15 , further comprising generating the second video stream comprising the object embodied within the computer-generated virtual environment, wherein the generating of the second video stream is in accordance with a preference of the user that is inferred from a plurality of user behaviors.

Continuity (3)
Continuation 17014115 · Sep 8, 2020
Provisional Application 62904015 · Sep 23, 2019
Related Publication 20230196385A1 · Jun 22, 2023