IP Library Granted Patent US 12,175,747
Granted Patent B2
US 12,175,747 · App. 17/888,163 · Granted Dec 24, 2024

Systems, methods, and apparatus for image-responsive automated assistants

Inventors: Marcin Nowak-Przygodzki (Bäch, CH); Gökhan Bakir (Zurich, CH)
Assignee: GOOGLE LLC
G06V20/20G06F3/011G06F3/014G06F3/0482G06F3/04886G06F16/487G06F16/9032G06Q30/02G06Q30/0281H04N23/63H04N23/667G06V20/68G06V2201/07G06V2201/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,175,747
App. No.
17/888,163
Granted
Dec 24, 2024
Kind
B2
Abstract

Techniques described herein enable a user to interact with an automated assistant and obtain relevant output from the automated assistant without requiring arduous typed input to be provided by the user and/or without requiring the user to provide spoken input that could cause privacy concerns (e.g., if other individuals are nearby). The assistant application can operate in multiple different image conversation modes in which the assistant application is responsive to various objects in a field of view of the camera. The image conversation modes can be suggested to the user when a particular object is detected in the field of view of the camera. When the user selects an image conversation mode, the assistant application can thereafter provide output, for presentation, that is based on the selected image conversation mode and that is based on object(s) captured by image(s) of the camera.

Claims (59)

1. A method implemented by one or more processors, the method comprising:

receiving real-time image feed from a camera of a computing device, the real-time image feed capturing one or more images of an object and being received in response to the object being present in a field of view of the camera;

displaying the real-time image feed at a display device;

causing one or more selectable elements to be rendered over the real-time image feed at the display device;

receiving a selection of a selectable element, of the one or more selectable elements, that is rendered at the display device; and

in response to the selection of the selectable element that is rendered at the display device and without requiring a user to provide spoken input, causing audible output to be rendered at the computing device, the audible output including object data tailored to a conversation mode, out of a plurality of conversation modes, that corresponds to the selectable element, wherein the selectable element identifies the conversation mode in which the object data is generated.

2. The method of claim 1 , further comprising:

identifying, based on processing the one or more images, the object.

3. The method of claim 2 , wherein identifying, based on processing the one or more images, the object comprises:

generating an identifier for the object.

4. The method of claim 1 , wherein the object includes a text.

5. The method of claim 4 , further comprising:

comparing a language of the text with a primary dialect setting of the computing device;

in response to the language of the text being different from the primary dialect setting, providing a translated text at the display device.

6. The method of claim 5 , wherein providing the translated text at the display device comprises:

translating the text; and

presenting the translated text at the display device.

7. The method of claim 5 , wherein prior to providing the translated text, the method further comprises:

providing a prompt to a user regarding whether to translate the text in a translate mode.

8. The method of claim 1 , wherein the object data is further tailored to the object identified from the one or more images in the real-time image feed.

9. The method of claim 1 , wherein the selection of the selectable element is received at a touch interface of the display device.

10. The method of claim 1 , wherein the audible output is rendered at the computing device while the real-time image feed is displayed at the display device.

11. The method of claim 1 , wherein the display device is included in a client device that is different from the computing device.

12. The method of claim 1 , wherein the display device is included in the computing device.

13. The method of claim 1 , further comprising:

in response to a different object being presented in the field of view of the camera, causing a different selectable element to be rendered over the real-time image feed at the display device.

14. A method implemented by one or more processors, the method comprising:

receiving, by a computing device, a spoken utterance, from a user, that is directed to an automated assistant that is accessible via the computing device;

in response to the spoken utterance:

causing a camera of the computing device to provide a real-time image feed at a display interface of the computing device,

causing the automated assistant to enter a fact mode,

capturing, while the camera is providing the real-time image feed, one or more objects in the real-time image feed of the camera;

determining, based on an image from the real-time image feed, that the camera is directed to an object from the one or more objects;

in response to determining that the camera is directed to the object:

identifying the object, wherein identifying the object is at least based on processing one or more images from the real-time image feed,

generating, based on the automated assistant being in the fact mode, content describing facts of the object, and

causing at least a portion of the content describing facts of the object to be rendered audibly via the computing device.

15. The method of claim 14 , wherein the facts of the object is filtered based on a context of the one or more images prior to being included in the natural language content describing facts of the object.

16. The method of claim 15 , wherein the context of the one or more images is identified by one or more context identifiers indicative of the context, and wherein the one or more context identifiers include a location where the one or more images were captured.

17. The method of claim 14 , wherein the content describing the facts of the object includes natural language content describing the facts of the object and a graphical representation of the object, and the method further comprises:

causing the natural language content describing the facts of the object and the graphical representation of the object to be rendered at the display interface of the computing device.

18. The method of claim 17 , further comprising

receiving, at the display interface of the computing device, a selection of the graphical representation of the object, and

causing, in response to receiving the selection of the graphical representation of the object, additional content to be rendered at the display interface of the computing device,

wherein the additional content is associated with the object.

19. The method of claim 14 , wherein:

identifying the object comprises generating an identifier for the object, and causing at least the portion of the content describing facts of the object to be rendered audibly via the computing device comprises causing the identifier for the object to be rendered audibly.

20. A method implemented by one or more processors, the method comprising:

receiving, by a computing device, a spoken utterance directed to an automated assistant that is accessible via the computing device;

in response to the spoken utterance:

causing a camera of the computing device to provide a real-time image feed at a display interface of the computing device, and

causing the computing device in a fact mode;

capturing, while the camera is providing the real-time image feed, a first object in the real-time image feed of the camera,

determining, based on the real-time image feed, that the camera is directed to the first object;

in response to determining that the camera is directed to the first object in the real-time image feed of the camera, identifying the first object,

wherein identifying the first object is at least based on processing one or more images from the real-time image feed; and

in response to identifying the first object, generating, based on the computing device being in the fact mode, content describing facts of the first object,

wherein the content changes to describe facts of an additional object when the camera is redirected from the first object to the additional object which causes the additional object to be in a field of view of the camera, and

wherein the additional object is different from the first object.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 14, 2022
From: NOWAK-PRZYGODZKI, MARCIN; BAKIR, GOKHAN
To: GOOGLE INC.
Reel/Frame 061086/0763 →
CHANGE OF NAME Recorded Sep 14, 2022
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 061432/0645 →
Continuity (3)
Continuation 16806505 · Mar 2, 2020
Continuation 15700106 · Sep 9, 2017
Related Publication 20220392216A1 · Dec 8, 2022