IP Library Granted Patent US 12,124,684
Granted Patent B2
US 12,124,684 · App. 17/741,303 · Granted Oct 22, 2024

Dynamic targeting of preferred objects in video stream of smartphone camera

Inventors: Alexander Pashintsev (Cupertino, CA); Boris Gorbatov (Sunnyvale, CA); Eugene Livshitz (San Mateo, CA); Vitaly Glazkov (Moscow, RU)
Assignee: Bending Spoons S.P.A.
G06F3/04842G06F3/017G06F3/04883
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,124,684
App. No.
17/741,303
Granted
Oct 22, 2024
Kind
B2
Abstract

Selecting objects in a video stream of a smart phone includes detecting quiescence of frame content in the video stream, detecting objects in a scene corresponding to the frame content, presenting at least one of the objects to a user of the smart phone, and selecting at least one of the objects in a group of objects in response to input by the user. Detecting quiescence of frame content in the video stream may include using motion sensors in the smart phone to determine an amount of movement of the smart phone. Detecting quiescence of frame content in the video stream may include detecting changes in view angles and distances of the smart phone with respect to the scene. Detecting objects in a scene may use heuristics, custom user preferences, and/or specifics of scene layout. At least one of the objects may be a person or a document.

Claims (56)

1. A method of capturing a subset of objects within a video stream captured by an electronic device, the method comprising:

receiving a video stream captured by an electronic device;

detecting within a frame of the video stream one or more objects for capture;

computing, without user input, a plurality of scenarios based on the one or more objects, wherein each scenario of the plurality of scenarios is a distinct subset of the one or more objects that includes a plurality of the one or more objects;

displaying, via the electronic device, the frame of the video stream in conjunction with a first scenario of the plurality of scenarios, wherein the plurality of objects in the first scenario are emphasized in the display;

responsive to a first user input rejecting the first scenario of the plurality of scenarios, displaying, via the electronic device, the frame of the video stream in conjunction with a second scenario of the plurality of scenarios, wherein the plurality of objects in the second scenario are emphasized in the display; and

responsive to a second user input selecting the second scenario of the plurality of scenarios, extracting the plurality of objects in the second scenario from the frame of the video stream;

wherein at least one object of the one or more objects is shared by three or more scenarios of the plurality of scenarios.

2. The method of claim 1 , further comprising:

after computing the plurality of scenarios, pre-selecting the first scenario of the plurality of scenarios to be displayed, via the electronic device, based on a third user input.

3. The method of claim 2 , wherein the third user input includes one or more of a change in a view angle and a change in a distance of the electronic device with respect to the one or more objects.

4. The method of claim 1 , wherein:

displaying the frame of the video stream in conjunction with the first scenario of the plurality of scenarios includes displaying the first scenario with an overlay highlighting the plurality of objects in the first scenario; and

displaying the frame of the video stream in conjunction with the second scenario of the plurality of scenarios includes displaying the second scenario with an overlay highlighting the plurality of objects in the second scenario.

5. The method of claim 1 , wherein the one or more objects include one or more of a person and a document.

6. The method of claim 1 , further comprising:

after extracting the plurality of objects in the second scenario from the frame of the video stream, displaying, via the electronic device, the plurality of objects in the second scenario.

7. The method of claim 1 , further comprising:

after extracting the plurality of objects in the second scenario from the frame of the video stream, displaying, via the electronic device, one or more affordances including:

a first affordance that allows the user to store the plurality of objects in the second scenario, and

a second affordance that allows the user to share the plurality of objects in the second scenario.

8. The method of claim 1 , wherein detecting within the frame of the video stream the one or more objects for capture includes one or more of determining one or more objects in focus, determining one or more objects with a predetermined distance relative to the electronic device, and determining one or more unobstructed objects.

9. The method of claim 1 , wherein the first user input includes one or more of selection of a rejection affordance displayed on the electronic device and a rejection gesture including shaking the electronic device left-and-right.

10. The method of claim 1 , wherein the second user input includes one or more of selection of an approval affordance displayed on the electronic device, allowing a predetermined amount of time to elapse without moving the electronic device, eye-tracking, spatial gestures, and facial expressions.

11. An electronic device, the electronic device comprising:

one or more processors; and

memory storing one or more programs for execution by the one or more processors, the one or more programs including instructions for:

receiving a video stream captured by the electronic device;

detecting within a frame of the video stream one or more objects for capture;

computing, without user input, a plurality of scenarios based on the one or more objects, wherein each scenario of the plurality of scenarios is a distinct subset of the one or more objects that includes a plurality of the one or more objects;

displaying, via the electronic device, the frame of the video stream in conjunction with a first scenario of the plurality of scenarios, wherein the plurality of objects in the first scenario are emphasized in the display;

responsive to a first user input rejecting the first scenario of the plurality of scenarios, displaying, via the electronic device, the frame of the video stream in conjunction with a second scenario of the plurality of scenarios, wherein the plurality of objects in the second scenario are emphasized in the display; and

responsive to a second user input selecting the second scenario of the plurality of scenarios, extracting the plurality of objects in the second scenario from the frame of the video stream;

wherein at least one object of the one or more objects is shared by three or more scenarios of the plurality of scenarios.

12. The electronic device of claim 11 , wherein the one or more programs further include instructions for:

after computing the plurality of scenarios, pre-selecting the first scenario of the plurality of scenarios to be displayed, via the electronic device, based on a third user input.

13. The electronic device of claim 12 , wherein the third user input includes one or more of a change in a view angle and a change in a distance of the electronic device with respect to the one or more objects.

14. The electronic device of claim 11 , wherein:

displaying the frame of the video stream in conjunction with the first scenario of the plurality of scenarios includes displaying the first scenario with an overlay highlighting the plurality of objects in the first scenario; and

displaying the frame of the video stream in conjunction with the second scenario of the plurality of scenarios includes displaying the second scenario with an overlay highlighting the plurality of objects in the second scenario.

15. The electronic device of claim 11 , wherein the one or more objects include one or more of a person and a document.

16. A non-transitory computer-readable storage medium storing one or more programs for execution by one or more processors of an electronic device, the one or more programs comprising instructions for:

receiving a video stream captured by the electronic device;

detecting within a frame of the video stream one or more objects for capture;

computing, without user input, a plurality of scenarios based on the one or more objects, wherein each scenario of the plurality of scenarios is a distinct subset of the one or more objects that includes a plurality of the one or more objects;

displaying, via the electronic device, the frame of the video stream in conjunction with a first scenario of the plurality of scenarios, wherein the plurality of objects in the first scenario are emphasized in the display;

responsive to a first user input rejecting the first scenario of the plurality of scenarios, displaying, via the electronic device, the frame of the video stream in conjunction with a second scenario of the plurality of scenarios, wherein the plurality of objects in the second scenario are emphasized in the display; and

responsive to a second user input selecting the second scenario of the plurality of scenarios, extracting the plurality of objects in the second scenario from the frame of the video stream;

wherein at least one object of the one or more objects is shared by three or more scenarios of the plurality of scenarios.

17. The non-transitory computer-readable storage medium of claim 16 , wherein the one or more programs further include instructions for:

after computing the plurality of scenarios, pre-selecting the first scenario of the plurality of scenarios to be displayed, via the electronic device, based on a third user input.

18. The non-transitory computer-readable storage medium of claim 17 , wherein the third user input includes one or more of a change in a view angle and a change in a distance of the electronic device with respect to the one or more objects.

19. The non-transitory computer-readable storage medium of claim 16 , wherein:

displaying the frame of the video stream in conjunction with the first scenario of the plurality of scenarios includes displaying the first scenario with an overlay highlighting the plurality of objects in the first scenario; and

displaying the frame of the video stream in conjunction with the second scenario of the plurality of scenarios includes displaying the second scenario with an overlay highlighting the plurality of objects in the second scenario.

20. The non-transitory computer-readable storage medium of claim 16 , wherein the one or more objects include one or more of a person and a document.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 12, 2024
From: EVERNOTE CORPORATION
To: BENDING SPOONS S.P.A.
Reel/Frame 066288/0195 →
Continuity (3)
Continuation 15079513 · Mar 24, 2016
Provisional Application 62139865 · Mar 30, 2015
Related Publication 20220269396A1 · Aug 25, 2022