IP Library Granted Patent US 11,908,161
Granted Patent B2
US 11,908,161 · App. 17/544,439 · Granted Feb 20, 2024

Method and electronic device for generating AR content based on intent and interaction of multiple-objects

Inventors: Ramasamy Kannan (Bangalore, IN); Lokesh Rayasandra Boregowda (Bangalore, IN)
Assignee: Samsung Electronics Co., Ltd.
G06T7/74G06T7/248G06T11/00G06V10/764G06V20/20G06V40/23G06T2207/10016G06T2207/20081G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,908,161
App. No.
17/544,439
Granted
Feb 20, 2024
Kind
B2
Abstract

An electronic device and a method for generating an augmented reality (AR) content in an electronic device are provided. The method includes determining a posture and an action of each object of the plurality of the objects in the scene displayed on a field of view of the electronic device, classifying the posture and the action of each object of the plurality of the objects in the scene, identifying an intent and an interaction of each object from the plurality of objects in the scene based on at least one of the classified posture and the classified action, and generating the AR content for the at least one object in the scene of at least one of the identified intent and the identified interaction of the at least one object.

Claims (89)

1. A method for generating augmented reality (AR) content in an electronic device, the method comprising:

identifying, by the electronic device, a plurality of objects in a scene displayed on a field of view of the electronic device;

determining, by the electronic device, at least one of a posture of each object of the plurality of the objects in the scene and an action of each object of the plurality of the objects in the scene;

classifying, by the electronic device, the posture of each object of the plurality of the objects in the scene and the action of each object of the plurality of the objects in the scene;

identifying, by the electronic device, an intent of at least one object from the plurality of objects in the scene based on at least one of the classified posture or the classified action;

identifying, by the electronic device, an interaction between the at least one object and at least one second object from the plurality of objects in the scene; and

generating, by the electronic device, the AR content for the at least one object in the scene of at least one of the identified intent or the identified interaction between the at least one object and the at least one second object from the plurality of objects in the scene.

2. The method of claim 1 , further comprising:

classifying, by the electronic device, each object of the plurality of the objects in order to determine the posture, the action, the interaction and the intent associated with each object of the plurality of the objects in the scene;

generating, by the electronic device, a semantic two-dimensional (2D) map of the identified plurality of objects in the scene displayed on the field of view of the electronic device;

generating, by the electronic device, a semantic three-dimensional (3D) map of the identified plurality of objects in the scene displayed on the field of view of the electronic device;

determining, by the electronic device, location information of each object of the plurality of the objects by using bounding boxes and distance information between each object of the plurality of the objects by using the semantic 3D map and the semantic 2D map; and

determining, by the electronic device, a linear motion of each object of the plurality of the objects and a rotational motion of each object of the plurality of the objects by using the generated semantic 3D map, the classified each object, the determined location information, the classified posture, and the classified action.

3. The method of claim 1 , further comprising:

detecting, by the electronic device, multiple body key-points associated with an interacting human of the plurality of the objects and non-interacting humans of the plurality of the objects in the scene displayed on the field of view of the electronic device; and

estimating, by the electronic device, the posture of each object of the plurality of the objects based on features derived from the detected multiple body key-points to classify the posture of each object of the plurality of the objects in the scene.

4. The method of claim 1 , further comprising:

classifying, by the electronic device, the action of each object of the plurality of the objects based on the classified postures and multiple body key-points identified over a current frame and multiple past frames.

5. The method of claim 1 , further comprising:

detecting, by the electronic device, at least one pose coordinate of the identified plurality of objects in the scene;

obtaining, by the electronic device, at least one pose feature from the detected at least one pose coordinate;

time interleaving, by the electronic device, pose features obtained from multiple objects in a single camera frame, and providing each object pose feature as input to a simultaneous real-time classification model to predict the posture of all objects in the single camera frame before next camera frame arrives; and

applying, by the electronic device, the simultaneous real-time classification model to predict the posture of each object in each frame of the scene.

6. The method of claim 1 , further comprising:

detecting, by the electronic device, sequence of an image frame, wherein the image frame comprises the identified plurality of objects in the scene and wherein each object of the plurality of objects performing the action;

converting, by the electronic device, the detected sequence of image frame having actions of at least one first object of the plurality of objects into a sequence of pose coordinates denoting the action of at least one first object of the plurality of objects;

converting, by the electronic device, the detected sequence of frames having actions of at least one-second object of the plurality of objects into the sequence of pose coordinates denoting the action of the at least one-second object of the plurality of objects;

time interleaving, by the electronic device, the pose features and classified postures, obtained from the plurality of objects over current multiple camera frames, and providing each object pose features and classified postures as input to a simultaneous real-time classification model to predict the action of all objects in the current multiple camera frames before the next camera frame arrives; and

applying, by the electronic device, the simultaneous real-time classification model to predict the action of each object in each frame of the scene.

7. The method of claim 1 , further comprising:

calculating, by the electronic device, probability of the intent and the interaction of each object of the plurality of the objects in the scene based on the classified posture and the classified action; and

identifying, by the electronic device, the intent and the interaction of each object of the plurality of objects based on the calculated probability, a determined linear motion of the at least one object, and a determined rotational motion of the at least one object.

8. The method of claim 7 ,

wherein the intent and the interaction of each object of the plurality of objects is identified by using a plurality of features, and

wherein the plurality of features comprises an object Identity (ID), object function, object shape, object pose coordinates, object postures, object actions, object linear motion coordinates, object rotational motion coordinates, object past interaction states, and object past intent states.

9. The method of claim 1 , wherein the generating of the AR content for the at least one object in the scene of at least one of the identified intent or the identified interaction of the at least one object comprising:

aggregating, by the electronic device, the identified intent and the identified interaction of the at least one object from the plurality of objects in the scene over a pre-defined time using a machine learning (ML) model;

mapping, by the electronic device, the aggregated intent and the aggregated interaction with at least one predefined template of at least one of an AR text templates or an AR effect templates; and

automatically inserting, by the electronic device, real-time information in the at least one predefined template and displaying the at least one predefined template with the at least one of the AR text templates or the AR effect templates in the scene displayed on the field of view of the electronic device.

10. The method of claim 9 , wherein the inserting of the real-time information in the at least one of the AR text template or the AR effect template uses the real-time information from a knowledge base.

11. An electronic device for generating augmented reality (AR) content in the electronic device, the electronic device comprising:

a memory;

a processor; and

an AR content controller, operably connected to the memory and the processor, configured to:

identify a plurality of objects in a scene displayed on a field of view of the electronic device,

determine at least one of a posture of each object of the plurality of the objects in the scene and an action of each object of the plurality of the objects in the scene,

classify the posture of each object of the plurality of the objects in the scene and the action of each object of the plurality of the objects in the scene,

identify an intent of at least one object from the plurality of objects in the scene based on at least one of the classified posture or the classified action,

identify an interaction between the at least one object and at least one second object from the plurality of objects in the scene, and

generate the AR content for the at least one object in the scene of at least one of the identified intent or the identified interaction between the at least one object and the at least one second object from the plurality of objects in the scene.

12. The electronic device of claim 11 , further comprising:

classifying, by the electronic device, each object of the plurality of the objects in order to determine the posture, the action, the interaction and the intent associated with each object of the plurality of the objects in the scene;

generating, by the electronic device, a semantic two-dimensional (2D) map of the identified plurality of objects in the scene displayed on the field of view of the electronic device;

generating, by the electronic device, a semantic three-dimensional (3D) map of the identified plurality of objects in the scene displayed on the field of view of the electronic device;

determining, by the electronic device, location information of each object of the plurality of the objects by using bounding boxes and distance information between each object of the plurality of the objects by using the semantic 3D map and the semantic 2D map; and

determining, by the electronic device, a linear motion of each object of the plurality of the objects and a rotational motion of each object of the plurality of the objects by using the generated semantic 3D map, the classified each object, the determined location information, the classified posture, and the classified action.

13. The electronic device of claim 11 , wherein the AR content controller is further configured to:

detect multiple body key-points associated with an interacting human of the plurality of the objects and non-interacting humans of the plurality of the objects in the scene displayed on the field of view of the electronic device, and

estimate the posture of each object of the plurality of the objects based on features derived from the detected multiple body key-points to classify the posture of each object of the plurality of the objects in the scene.

14. The electronic device of claim 11 , wherein the AR content controller is further configured to:

classify the action of each object of the plurality of the objects based on the classified postures and multiple body key-points identified over a current frame and multiple past frames.

15. The electronic device as claimed in claim 11 , wherein classify the posture of each object of the plurality of the objects in the scene comprises the AR content controller is further configured to:

detect at least one pose coordinate of the identified plurality of objects in the scene,

obtain at least one pose feature from the detected at least one pose coordinate,

time interleaving pose features obtained from multiple objects in a single camera frame, and providing each object pose feature as input to a simultaneous real-time classification model to predict the posture of all objects in the single camera frame before next camera frame arrives, and

apply a simultaneous real-time classification model to predict the posture of each object in each frame of the scene.

16. The electronic device as claimed in claim 11 , wherein classify the action of each object of the plurality of the objects in the scene comprises the AR content controller is further configured to:

detect sequence of an image frame, wherein the image frame comprises the identified plurality of objects in the scene and wherein each object of the plurality of objects performing the action,

convert the detected sequence of image frame having actions of at least one first object of the plurality of objects into a sequence of pose coordinates denoting the action of at least one first object of the plurality of objects,

convert the detected sequence of frames having actions of at least one-second object of the plurality of objects into the sequence of pose coordinates denoting the action of the at least one-second object of the plurality of objects,

time interleaving the pose features and classified postures, obtained from the plurality of objects over current multiple camera frames, and providing each object pose features and classified postures as input to a simultaneous real-time classification model to predict the action of all objects in the current multiple camera frames before the next camera frame arrives, and

apply the simultaneous real-time classification model to predict the action of each object in each frame of the scene.

17. The electronic device as claimed in claim 11 , wherein identifying the intent and the interaction of each object of the plurality of objects in the scene comprises the AR content controller is further configured to:

calculate probability of the intent and the interaction of each object of the plurality of the objects in the scene based on the classified posture and the classified action, and

identify the intent and the interaction of each object of the plurality of objects based on the calculated probability, a determined linear motion of the at least one object, and a determined rotational motion of the at least one object.

18. The electronic device as claimed in claim 17 ,

wherein the intent and the interaction of each object of the plurality of objects is identified by using a plurality of features, and

wherein the plurality of features comprises an object Identity (ID), object function, object shape, object pose coordinates, object postures, object actions, object linear motion coordinates, object rotational motion coordinates, object past interaction states, and object past intent states.

19. The electronic device as claimed in claim 11 , wherein generate the AR content for the at least one object in the scene of at least one of the identified intent and the identified interaction of the at least one object comprises the AR content controller is further configured to:

aggregate the identified intent and the identified interaction of the at least one object from the plurality of objects in the scene over a pre-defined time using a machine learning (ML) model,

map the aggregated intent and the aggregated interaction with at least one predefined template of at least one of an AR text templates and or an AR effect templates, and

automatically insert real-time information in the at least one predefined template and displaying the at least one predefined template with the at least one of the AR text templates and or the AR effect templates in the scene displayed on the field of view of the electronic device.

20. At least one non-transitory computer readable storage medium configured to store one or more computer programs including instructions that, when executed by at least one processor of an electronic device, cause the at least one processor to:

identify, by the electronic device, a plurality of objects in a scene displayed on a field of view of the electronic device;

determine, by the electronic device, at least one of a posture of each object of the plurality of the objects in the scene and an action of each object of the plurality of the objects in the scene;

classify, by the electronic device, the posture of each object of the plurality of the objects in the scene and the action of each object of the plurality of the objects in the scene;

identify, by the electronic device, an intent of at least one object from the plurality of objects in the scene based on at least one of the classified posture or the classified action;

identify, by the electronic device, an interaction between the at least one object and at least one second object from the plurality of objects in the scene; and

generate, by the electronic device, augmented reality (AR) content for the at least one object in the scene of at least one of the identified intent or the identified interaction between the at least one object and the at least one second object from the plurality of objects in the scene.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 7, 2021
From: KANNAN, RAMASAMY; BOREGOWDA, LOKESH RAYASANDRA
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 058325/0575 →
Priority Claims (2)
IN 202041043804 · Oct 8, 2020 · national
IN 2020 41043804 · Jul 8, 2021 · national
Continuity (2)
Continuation PCTKR2021013871 · Oct 8, 2021
Related Publication 20220114751A1 · Apr 14, 2022