IP Library Granted Patent US 11,334,228
Granted Patent B1
US 11,334,228 · App. 15/079,513 · Granted May 17, 2022

Dynamic targeting of preferred objects in video stream of smartphone camera

Inventors: Alexander Pashintsev (Cupertino, CA); Boris Gorbatov (Sunnyvale, CA); Eugene Livshitz (San Mateo, CA); Vitaly Glazkov (Moscow, RU)
Assignee: EVERNOTE CORPORATION
G06F3/04842G06F3/017G06F3/04883H04N1/19594H04M1/0264
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,334,228
App. No.
15/079,513
Granted
May 17, 2022
Kind
B1
Abstract

Selecting objects in a video stream of a smart phone includes detecting quiescence of frame content in the video stream, detecting objects in a scene corresponding to the frame content, presenting at least one of the objects to a user of the smart phone, and selecting at least one of the objects in a group of objects in response to input by the user. Detecting quiescence of frame content in the video stream may include using motion sensors in the smart phone to determine an amount of movement of the smart phone. Detecting quiescence of frame content in the video stream may include detecting changes in view angles and distances of the smart phone with respect to the scene. Detecting objects in a scene may use heuristics, custom user preferences, and/or specifics of scene layout. At least one of the objects may be a person or a document.

Claims (47)

1. A method of selecting objects in a video stream captured by a user device, the method comprising:

detecting quiescence of frame content in the video stream;

in response to detecting a quiescent state of the frame content in the video stream, detecting a plurality of objects corresponding to the frame content;

determining, without user interaction, a plurality of scenarios, wherein:

a first respective scenario of the plurality of scenarios includes all of the plurality of objects; and

a second respective scenario of the plurality of scenarios includes a subset of the plurality of objects less than all of the plurality of objects and is distinct from the first respective scenario including all of the plurality of objects corresponding to the frame content;

after determining the plurality of scenarios, presenting for user selection, by the user device, the first respective scenario by displaying the frame content with an overlay highlighting within the frame content all of the plurality of objects detected in the frame content;

in response to detecting a first user input that rejects the first respective scenario, presenting for user selection, by the user device, the second respective scenario by displaying the frame content with an overlay highlighting within the frame content the one or more objects of the subset of the plurality of objects without highlighting the other objects of the plurality of objects detected in the frame content;

in response to detecting a second user input that selects the second respective scenario, capturing the frame content in the video stream;

retrieving, from the frame content, the one or more objects of the subset of the plurality of objects that correspond to the second respective scenario; and

presenting, by the user device, the one or more objects of the subset of the plurality of objects.

2. The method of claim 1 , wherein detecting quiescence of frame content in the video stream includes using motion sensors in the user device to determine an amount of movement of the user device.

3. The method of claim 1 , wherein detecting quiescence of frame content in the video stream includes detecting a change in at least one of a view angle and a distance of the user device with respect to a scene that includes the plurality of objects.

4. The method of claim 1 , wherein detecting the plurality of objects corresponding to the frame content uses at least one of:

heuristics, custom user preferences, and specifics of scene layout.

5. The method of claim 1 , wherein at least one of the plurality of objects is a person.

6. The method of claim 1 , wherein at least one of the plurality of objects is a document.

7. The method of claim 1 , wherein presenting a respective scenario includes drawing a frame around a respective set of objects.

8. The method of claim 1 , wherein detecting the plurality of objects includes detecting a third user input that pre-selects at least a subset of the plurality of objects by changing the position and view angle of the user device to cause desired objects to occupy a significant portion of a screen of the user device.

9. The method of claim 1 , wherein detecting user selection of a respective scenario from the plurality of scenarios includes determining that a predetermined amount of time has passed, while the respective scenario is presented to the user on the user device, without detection of a rejection input rejecting the respective scenario.

10. The method of claim 9 , wherein detecting the rejection input includes detecting, while the respective scenario is presented to the user on the user device, a rejection gesture.

11. The method of claim 10 , wherein the rejection gesture is shaking the user device left-and-right several times.

12. The method of claim 1 , wherein detecting user selection of a respective scenario from the plurality of scenarios includes using at least one of: eye-tracking, spatial gestures captured by a wearable device, and analysis of facial expressions.

13. The method of claim 1 , wherein detecting user selection of a respective scenario from the plurality of scenarios includes detecting at least one of: tapping a dedicated button on a screen of the user device, touching the screen, and performing a multi-touch approval gesture on the user device.

14. A non-transitory computer-readable medium containing software that selects objects in a video stream captured by a user device, the software comprising:

executable code that detects quiescence of frame content in the video stream;

executable code that in response to detecting a quiescent state of the frame content in the video stream, detects a plurality of objects corresponding to the frame content;

executable code that determines, without user interaction, a plurality of scenarios, wherein:

a first respective scenario of the plurality of scenarios includes all of the plurality of objects; and

a second respective scenario of the plurality of scenarios includes a subset of the plurality of objects less than all of the plurality of objects and is distinct from the first respective scenario including all of the plurality of objects corresponding to the frame content;

executable code that, after determining the plurality of scenarios, presents for user selection, by the user device, the first respective scenario by displaying the frame content with an overlay highlighting within the frame content all of the plurality of objects detected in the frame content;

executable code that in response to detecting a first user input that rejects the first respective scenario, presents for user selection, by the user device, the second respective scenario by displaying the frame content with an overlay highlighting within the frame content the one or more objects of the subset of the plurality of objects without highlighting the other objects of the plurality of objects detected in the frame content;

executable code that in response to detecting a second user input that selects the second respective scenario, captures the frame content in the video stream;

executable code that retrieves, from the frame content, the one or more objects of the subset of the plurality of objects that correspond to the second respective scenario; and

executable code that presents, by the user device, the one or more objects of the subset of the plurality of objects.

15. The non-transitory computer-readable medium of claim 14 , wherein executable code that detects quiescence of frame content in the video stream uses motion sensors in the user device to determine an amount of movement of the user device.

16. The non-transitory computer-readable medium of claim 14 , wherein executable code that detects quiescence of frame content in the video stream detects a change in at least one of a view and a distance of the user device with respect to a scene that includes the plurality of objects.

17. The non-transitory computer-readable medium of claim 14 , wherein executable code that detects the plurality of objects corresponding to the frame content uses at least one of: heuristics, custom user preferences, and specifics of scene layout.

18. The non-transitory computer-readable medium of claim 14 , wherein at least one of the plurality of objects is a person.

19. The non-transitory computer-readable medium of claim 14 , wherein at least one of the plurality of objects is a document.

20. The non-transitory computer-readable medium of claim 14 , wherein executable code that presents a respective scenario includes executable code that performs drawing a frame around a respective set of objects.

21. The non-transitory computer-readable medium of claim 14 , wherein detecting the plurality of objects includes detecting a third user input that pre-selects at least a subset of the plurality of objects by changing the position and view angle of the user device to cause desired objects to occupy a significant portion of a screen of the user device.

22. The non-transitory computer-readable medium of claim 14 , wherein detecting user selection of a respective scenario from the plurality of scenarios includes determining that a predetermined amount of time has passed, while the respective scenario is presented to the user on the user device, without detection of a rejection input rejecting the respective scenario.

23. The non-transitory computer-readable medium of claim 22 , wherein detecting the rejection input includes detecting, while the respective scenario is presented to the user on the user device, a rejection gesture.

24. The non-transitory computer-readable medium of claim 23 , wherein the rejection gesture is shaking the user device left-and-right several times.

25. The non-transitory computer-readable medium of claim 14 , wherein detecting user selection of a respective scenario from the plurality of scenarios includes using at least one of: eye-tracking, spatial gestures captured by a wearable device, and analysis of facial expressions.

26. The non-transitory computer-readable medium of claim 14 , wherein detecting user selection of a respective scenario from the plurality of scenarios includes detecting at least one of: tapping a dedicated button on a screen of the user device, touching the screen, and performing a multi-touch approval gesture on the user device.

Assignments (8)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 12, 2024
From: EVERNOTE CORPORATION
To: BENDING SPOONS S.P.A.
Reel/Frame 066288/0195 →
RELEASE OF SECURITY INTEREST Recorded Mar 17, 2023
From: MUFG BANK, LTD.
To: EVERNOTE CORPORATION
Reel/Frame 063116/0260 →
RELEASE OF SECURITY INTEREST Recorded Oct 8, 2021
From: EAST WEST BANK
To: EVERNOTE CORPORATION
Reel/Frame 057852/0078 →
SECURITY INTEREST Recorded Oct 6, 2021
From: EVERNOTE CORPORATION
To: MUFG UNION BANK, N.A.
Reel/Frame 057722/0876 →
SECURITY INTEREST Recorded Oct 19, 2020
From: EVERNOTE CORPORATION
To: EAST WEST BANK
Reel/Frame 054113/0876 →
SECURITY INTEREST Recorded Oct 5, 2016
From: EVERNOTE CORPORATION; EVERNOTE GMBH
To: HERCULES CAPITAL, INC., AS AGENT
Reel/Frame 040240/0945 →
SECURITY AGREEMENT Recorded Sep 30, 2016
From: EVERNOTE CORPORATION
To: SILICON VALLEY BANK
Reel/Frame 040192/0720 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 28, 2016
From: PASHINTSEV, ALEXANDER; GORBATOV, BORIS; LIVSHITZ, EUGENE; GLAZKOV, VITALY
To: EVERNOTE CORPORATION
Reel/Frame 038400/0403 →
Continuity (1)
Provisional Application 62139865 · Mar 30, 2015