IP Library Granted Patent US 10,997,730
Granted Patent B2
US 10,997,730 · App. 16/547,352 · Granted May 4, 2021

Detection of moment of perception

Inventors: Hessam Bagherinezhad (Seattle, WA); Carlo Eduardo Cabanero del Mundo (Seattle, WA); Anish Jnyaneshwar Prabhu (Seattle, WA); Peter Zatloukal (Seattle, WA); Lawrence Frederick Arnstein (Seattle, WA)
Assignee: Xnor.AI, Inc.
G06T7/20G06K9/6256G06N20/00G06T7/0002G06T2207/20081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,997,730
App. No.
16/547,352
Granted
May 4, 2021
Kind
B2
Abstract

In one embodiment, a method includes receiving a machine-learning model trained to detect a specified motion using multiple videos, wherein each video has at least one frame labeled as a moment of perception of the specified motion, identifying an object-of-interest depicted in an input video, detecting a motion of the object-of-interest, determining that the detected motion is the specified motion, and classifying one of the frames of the input video as the moment of perception of the specified motion.

Claims (45)

1. A method comprising:

receiving a machine-learning model trained to detect a specified motion using a plurality of videos, wherein each of the videos have at least one frame labeled as a moment of perception of the specified motion;

identifying an object-of-interest depicted in an input video;

detecting, with respect to a sequence of frames of the input video, a motion of the object-of-interest;

determining that the motion of the object-of-interest is the specified motion; and

classifying, using the trained machine-learning model, one of the frames of the input video as the moment of perception of the specified motion.

2. The method of claim 1 , further comprising:

analyzing the input video to determine one or more factors relating to the input video, wherein the classifying is based on the one or more factors.

3. The method of claim 2 , wherein the factors relating to the input video comprise an environmental context of a scene in each frame of the video, wherein the environmental context comprises the viewpoint, the lighting, the distance from the object-of-interest, the field-of-view, the diversity in the background, or the climate.

4. The method of claim 2 , wherein the factors relating to the input video comprise attributes of the object-of-interest, wherein the attributes of the object-of-interest comprise a detected pose, size, color, emotion, texture, temperature, or whether or not the object-of-interest is an individual subject or a group of subjects.

5. The method of claim 2 , wherein the factors relating to the input video comprise attributes of the specified motion, wherein the attributes of the specified motion comprise the obviousness of the specified motion, the variation of the object-of-interest, or the length of the specified motion.

6. The method of claim 2 , wherein the factors relating to the input video comprise metadata of the input video, wherein the metadata comprises the frames rate, resolution, data format, or EXIF data.

7. The method of claim 2 , wherein the factors relating to the input video comprise detected events, wherein the detected events comprise prior detected events or contemporaneously detected events.

8. The method of claim 1 , wherein each of the videos is labeled with an indication of whether or not the specified motion was confirmed.

9. The method of claim 1 , wherein each of the videos is labeled with a type of the specified motion.

10. The method of claim 1 , wherein each of the frames is labeled with a frame sequence number.

11. The method of claim 1 , wherein the trained machine-learning model is a binarized machine-learning model.

12. The method of claim 1 , further comprising:

detecting one or more potential objects in the input video, wherein the object-of-interest is identified from the potential objects.

13. The method of claim 12 , wherein the object-of-interest is identified based on one or more factors relating to each of the potential objects, wherein the factors relating to each potential object comprise a location in the frame of each potential object, a size of each potential object, or a significance of each potential object.

14. The method of claim 1 , further comprising:

identifying a second object-of-interest depicted in an input video;

detecting, with respect to a sequence of frames of the input video, a second motion of the second object-of-interest;

determining that the second motion of the second object-of-interest is a specified second motion; and

classifying, using the trained machine-learning model, one of the frames of the input video as a second moment of perception of the specified motion.

15. The method of claim 14 , wherein the first motion and the second motion are both parts of a conjoined motion.

16. The method of claim 1 , wherein the object-of-interest is comprised of a plurality of separable objects, and wherein the specified motion is done by one or more of the separable objects.

17. The method of claim 1 , wherein the input video contains audio, and wherein each frame of the input video is linked with an audio clip.

18. The method of claim 17 , further comprising:

analyzing the audio clip to determine one or more factors relating to the input video audio, wherein the classifying is based on the one or more factors.

19. The method of claim 18 , wherein the factors relating to the input video audio comprise a latency between the input video and the input video audio.

20. One or more computer-readable non-transitory storage media embodying software that is operable when executed to:

receive a machine-learning model trained to detect a specified motion using a plurality of videos, wherein each of the videos have at least one frame labeled as a moment of perception of the specified motion;

identify an object-of-interest depicted in an input video;

detect, with respect to a sequence of frames of the input video, a motion of the object-of-interest;

determine that the motion of the object-of-interest is the specified motion; and

classify, using the trained machine-learning model, one of the frames of the input video as the moment of perception of the specified motion.

21. A system comprising:

one or more processors; and

one or more computer-readable non-transitory storage media coupled to one or more of the processors and comprising instructions operable when executed by one or more of the processors to cause the system to:

receive a machine-learning model trained to detect a specified motion using a plurality of videos, wherein each of the videos have at least one frame labeled as a moment of perception of the specified motion;

identify an object-of-interest depicted in an input video;

detect, with respect to a sequence of frames of the input video, a motion of the object-of-interest;

determine that the motion of the object-of-interest is the specified motion; and

classify, using the trained machine-learning model, one of the frames of the input video as the moment of perception of the specified motion.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 14, 2021
From: XNOR.AI, INC.
To: APPLE INC.
Reel/Frame 058390/0589 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 21, 2019
From: BAGHERINEZHAD, HESSAM; DEL MUNDO, CARLO EDUARDO CABANERO; PRABHU, ANISH JNYANESHWAR; ZATLOUKAL, PETER; ARNSTEIN, LAWRENCE FREDERICK
To: XNOR.AI, INC.
Reel/Frame 050122/0644 →
Continuity (1)
Related Publication 20210056709A1 · Feb 25, 2021
Cited By (1)
US 12,217,474