IP Library Granted Patent US 10,147,197
Granted Patent B2
US 10,147,197 · App. 15/839,797 · Granted Dec 4, 2018

Segment content displayed on a computing device into regions based on pixels of a screenshot image that captures the content

Inventors: Dominik Roblek (Meilen, CH); David Petrou (Brooklyn, NY); Matthew Sharifi (Kilchberg, CH)
Assignee: GOOGLE LLC
G06T7/30G06F3/0488G06F3/04842G06F17/30047G06F17/30867G06T7/90
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,147,197
App. No.
15/839,797
Granted
Dec 4, 2018
Kind
B2
Abstract

Methods and apparatus directed to segmenting content displayed on a computing device into regions. The segmenting of content displayed on the computing device into regions is accomplished via analysis of pixels of a “screenshot image” that captures at least a portion of (e.g., all of) the displayed content. Individual pixels of the screenshot image may be analyzed to determine one or more regions of the screenshot image and to optionally assign a corresponding semantic type to each of the regions. Some implementations are further directed to generating, based on one or more of the regions, interactive content to provide for presentation to the user via the computing device.

Claims (56)

1. A method, comprising:

capturing, by one or more processors of a computing device, a screenshot image that captures at least a portion of a display provided to a user by the computing device;

segmenting the screenshot image into at least a first region and a second region, the segmenting being by one or more of the processors of the computing device and being based on a plurality of pixels of the screenshot image;

determining at least one first characteristic of the first region, the determining being by one or more of the processors of the computing device and being based on one or more of: a plurality of pixels of the first region, a size of the first region, and a position of the first region;

determining at least one second characteristic of the second region, the determining being by one or more of the processors of the computing device and being based on one or more of: a plurality of pixels of the second region, a size of the second region, and a position of the second region; and

providing, by one or more of the processors of the computing device, a plurality of the pixels of the first region to a content recognition engine based on the first region having the first characteristic;

wherein the pixels of the second region are not provided to any content recognition engine based on the second region having the second characteristic, and wherein the first characteristic is a first semantic label and the second characteristic is a second semantic label.

2. The method of claim 1 , wherein the pixels of the second region are not provided for any further action based on the second region having the second characteristic.

3. The method of claim 1 , further comprising:

receiving, from the content recognition engine in response to providing the plurality of the pixels of the first region, an indication of content of the first region; and

rendering, by one or more of the processors, interactive content at the computing device based on the received indication of content of the first region.

4. The method of claim 1 , further comprising:

selecting the content recognition engine, from a plurality of available content recognition engines, based on the first characteristic; and

based on selecting the content recognition engine:

providing the plurality of the pixels of the first region to the selected content recognition engine, without providing any of the pixels of the first region to any other of the available content recognition engines.

5. The method of claim 1 , further comprising:

receiving, from the content recognition engine in response to providing the plurality of the pixels of the first region, text that is present in the first region; and

rendering, by one or more of the processors, content at the computing device based on the received text.

6. The method of claim 1 , further comprising:

receiving, from the content recognition engine in response to providing the plurality of the pixels of the first region, at least one entity present in the first region; and

rendering, by one or more of the processors, content at the computing device based on the received at least one entity.

7. The method of claim 1 , further comprising:

selecting the content recognition engine, from a plurality of available content recognition engines, based on the first characteristic; and

based on selecting the content recognition engine:

providing the plurality of the pixels of the first region to the selected content recognition engine, without providing any of the pixels of the first region to at least one other of the available content recognition engines.

8. The method of claim 1 , wherein the content recognition engine comprises a trained machine learning model.

9. The method of claim 1 , wherein the content recognition engine comprises a trained convolutional neural network model.

10. A client computing device, comprising:

one or more computer readable media storing instructions;

one or more processors executing the instructions to perform a method comprising:

capturing a screenshot image that captures at least a portion of a display provided to a user by the client computing device;

segmenting, based on a plurality of pixels of the screenshot image, the screenshot image into at least a first region and a second region;

determining at least one first characteristic of the first region, the determining being based on one or more of: a plurality of pixels of the first region, a size of the first region, and a position of the first region;

determining at least one second characteristic of the second region, the determining being based on one or more of: a plurality of pixels of the second region, a size of the second region, and a position of the second region; and

providing a plurality of the pixels of the first region to a content recognition engine based on the first region having the first characteristic;

wherein the pixels of the second region are not provided to any content recognition engine based on the second region having the second characteristic, and wherein the first characteristic is a first semantic label and the second characteristic is a second semantic label.

11. The client computing device of claim 10 , wherein the pixels of the second region are not provided for any further action based on the second region having the second characteristic.

12. The client computing device of claim 10 , wherein the method further comprises:

receiving, from the content recognition engine in response to providing the plurality of the pixels of the first region, an indication of content of the first region; and

rendering interactive content based on the received indication of content of the first region.

13. The client computing device of claim 10 , wherein the method further comprises:

selecting the content recognition engine, from a plurality of available content recognition engines, based on the first characteristic; and

based on selecting the content recognition engine:

providing the plurality of the pixels of the first region to the selected content recognition engine, without providing any of the pixels of the first region to any other of the available content recognition engines.

14. The client computing device of claim 10 , wherein the method further comprises:

receiving, from the content recognition engine in response to providing the plurality of the pixels of the first region, text that is present in the first region; and

rendering content based on the received text.

15. The client computing device of claim 10 , wherein the method further comprises:

receiving, from the content recognition engine in response to providing the plurality of the pixels of the first region, at least one entity present in the first region; and

rendering content based on the received at least one entity.

16. The client computing device of claim 10 , wherein the method further comprises:

selecting the content recognition engine, from a plurality of available content recognition engines, based on the first characteristic; and

based on selecting the content recognition engine:

providing the plurality of the pixels of the first region to the selected content recognition engine, without providing any of the pixels of the first region to at least one other of the available content recognition engines.

17. The client computing device of claim 10 , wherein the content recognition engine comprises a trained machine learning model.

18. The client computing device of claim 10 , wherein the content recognition engine comprises a trained convolutional neural network model.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2018
From: ROBLEK, DOMINIK; PETROU, DAVID; SHARIFI, MATTHEW
To: GOOGLE INC.
Reel/Frame 045149/0975 →
CHANGE OF NAME Recorded Mar 8, 2018
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 045539/0968 →
Continuity (2)
Continuation 15154957 · May 14, 2016
Related Publication 20180114326A1 · Apr 26, 2018