IP Library › Granted Patent US 12,748,567
Granted Patent B2
US 12,748,567 · App. 18/044,831 · Granted Sep 29, 2026

Enhanced computing device representation of audio

Inventors: Dimitri Kanevsky (Mountain View, CA); Sagar Savla (Mountain View, CA); Ausmus Chang (Taipei City, TW); Chiawei Liu (Taipei City, TW); Daniel P W Ellis (New York, NY); Jinho Kim (Napa, CA); Justin Stuart Paul (San Francisco, CA); Sharlene Yuan (Orange, CA); Alex Huang (Taipei City, TW); Yun Che Chung (Taipei City, TW); Chelsey Fleming (West Hollywood, CA)
Assignee: Google LLC
G06F3/167G10L25/51G10L25/72
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,748,567
App. No.
18/044,831
Granted
Sep 29, 2026
Kind
B2
Abstract

An example method includes receiving, by one or more processors of a computing device, audio data recorded by one or more microphones of the computing device; and generating, based on the audio data and by the one or more processors, one or more structured sound records, a first structured sound record of the one or more structured sound records including: a description of a first sound, the description including a descriptive label of the first sound, the descriptive label different than a text transcription of the first sound, and a time stamp indicating a time at which the first sound occurred; and outputting a graphical user interface including a timeline representation of the one or more structured sound records.

Claims (51)

1 . A method comprising:

receiving, by one or more processors of a computing device, audio data recorded by one or more microphones of the computing device; and

generating, based on the audio data and by the one or more processors, one or more structured sound records, a first structured sound record of the one or more structured sound records including:

a description of a first sound, the description including a descriptive label of the first sound, the descriptive label representing a classification of a non-speech environmental sound and being different than a text transcription of the first sound, and

a time stamp indicating a time at which the first sound occurred; and

outputting a graphical user interface including a timeline representation of the one or more structured sound records, wherein the timeline representation visually indicates a sequence of a plurality of structured sound records corresponding to a plurality of different classified non-speech environmental sounds.

2 . The method of claim 1 , wherein the timeline representation indicates a sequence in which sounds of the one or more structured sound records occurred.

3 . The method of claim 1 , wherein a second structured sound record of the one or more sound records includes:

a description of a second sound, the description including a text transcription of the second sound, and

a time stamp indicating a time at which the second sound occurred.

4 . The method of claim 1 , wherein outputting the graphical user interface comprises outputting a first graphical user interface including a current timeline representation of the one or more structured sound records, the method further comprising:

outputting a second graphical user interface including a past timeline representation of the one or more structured sound records.

5 . The method of claim 4 , wherein outputting the first graphical user interface comprises outputting the first graphical user interface for display at a first display, and wherein outputting the second graphical user interface comprises outputting the second graphical user interface at a second display that is included in a device that is different than a device that includes the first display.

6 . The method of claim 1 , wherein outputting the graphical user interface comprises outputting the graphical user interface in response to determining that a newly generated sound record of the one or more structured sound records has a descriptive label included in a pre-determined set of descriptive labels.

7 . The method of claim 6 , wherein the pre-determined set of descriptive labels includes emergency category labels, priority category labels, and other category labels.

8 . The method of claim 7 , wherein:

the emergency category labels include one or more of a smoke alarm label, a fire alarm label, a carbon monoxide label, a siren label, and a shouting label;

the priority category labels include one or more of a baby crying label, a doorbell label, a door knocking label, an animal alerting label, and a glass breaking label; and

the other category labels include one or more of a water running label, a landline phone ringing label, and one or more appliance beep labels.

9 . The method of claim 1 , wherein outputting the graphical user interface including the timeline representation comprises outputting the graphical user interface including the timeline representation at a display of the computing device.

10 . The method of claim 1 , wherein the computing device is a first computing device, wherein outputting the graphical user interface including the timeline representation comprises causing a second computing device to output, at a display of the second computing device, the graphical user interface including the timeline representation, and wherein the second computing device is different than the first computing device.

11 . The method of claim 10 , further comprising:

responsive to receiving, at the computing device, user input indicating viewing of the timeline representation at a particular time, modifying output of the timeline representation at the second computing device to indicate previous viewing.

12 . The method of claim 10 , wherein the second computing device comprises a wearable computing device.

13 . The method of claim 10 , wherein the first computing device does not include a display.

14 . The method of claim 1 , further comprising: determining that the classification of the non-speech environmental sound corresponds to an emergency category; and outputting, concurrently with the graphical user interface, a non-audio alert via at least one of a haptic output device or a light device of the computing device.

15 . The method of claim 1 , further comprising: transmitting a signal to a wearable computing device communicatively coupled to the computing device, wherein the signal causes the wearable computing device to output a haptic alert corresponding to the classification of the non-speech environmental sound.

16 . The method of claim 1 , further comprising: receiving user input to scroll the timeline representation backwards in time; and updating the graphical user interface to display a past timeline representation indicating a sequence of previously generated structured sound records corresponding to previously classified non-speech environmental sounds.

17 . The method of claim 1 , wherein the timeline representation further includes a graphical plot indicating an amplitude of the first sound over a duration of the first sound.

18 . A computing device comprising:

one or more microphones configured to record audio data; and

one or more processors configured to:

generate, based on the audio data, one or more structured sound records, a first structured sound record of the one or more structured sound records including:

a description of a first sound, the description including a descriptive label of the first sound, the descriptive label representing a classification of a non-speech environmental sound and being different than a text transcription of the first sound, and

a time stamp indicating a time at which the first sound occurred; and

output a graphical user interface including a timeline representation of the one or more structured sound records, wherein the timeline representation visually indicates a sequence of a plurality of structured sound records corresponding to a plurality of different classified non-speech environmental sounds.

19 . The computing device of claim 18 , wherein the timeline representation indicates a sequence in which sounds of the one or more structured sound records occurred.

20 . The computing device of claim 18 , wherein a second structured sound record of the one or more sound records includes:

a description of a second sound, the description including a text transcription of the second sound, and

a time stamp indicating a time at which the second sound occurred.

21 . The computing device of claim 18 , wherein, to output the graphical user interface, the one or more processors are configured to output a first graphical user interface including a current timeline representation of the one or more structured sound records, and wherein the one or more processors are further configured to:

output a second graphical user interface including a past timeline representation of the one or more structured sound records.

22 . The computing device of claim 18 , wherein outputting the graphical user interface including the timeline representation comprises causing another computing device to output the graphical user interface including the timeline representation, and wherein the other computing device is different than the computing device.

23 . The computing device of claim 22 , wherein the one or more processors are further configured to:

modify, responsive to receiving user input indicating viewing of the timeline representation at a particular time, output of the timeline representation at the other computing device to indicate previous viewing.

24 . A computer-readable storage medium storing instructions that, when executed, cause one or more processors of a computing device to:

receive audio data recorded by one or more microphones of the computing device;

generate, based on the audio data, one or more structured sound records, a first structured sound record of the one or more structured sound records including:

a description of a first sound, the description including a descriptive label of the first sound, the descriptive label representing a classification of a non-speech environmental sound and being different than a text transcription of the first sound, and

a time stamp indicating a time at which the first sound occurred; and

output a graphical user interface including a timeline representation of the one or more structured sound records, wherein the timeline representation visually indicates a sequence of a plurality of structured sound records corresponding to a plurality of different classified non-speech environmental sounds.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 10, 2023
From: KANEVSKY, DIMITRI; SAVLA, SAGAR; CHANG, AUSMUS; LIU, CHIAWEI; ELLIS, DANIEL P W; KIM, JINHO; PAUL, JUSTIN STUART; YUAN, SHARLENE; HUANG, ALEX; CHUNG, YUN CHE; FLEMING, CHELSEY
To: GOOGLE LLC
Reel/Frame 062943/0946 →
Continuity (2)
Provisional Application 63088811 · Oct 7, 2020
Related Publication 20230342108A1 · Oct 26, 2023
References Cited (42)
US 6731307B1 · Strubbe · 2004 [cited by examiner]
US 6795808B1 · Strubbe · 2004 [cited by examiner]
US 7058889B2 · Trovato · 2006 [cited by examiner]
US 8457959B2 · Kaiser · 2013 [cited by examiner]
US 9967684B2 · Ingemarsson et al. · 2018 [cited by applicant]
US 10269228B2 · Aarts · 2019 [cited by applicant]
US 10395494B2 · Nongpiur et al. · 2019 [cited by applicant]
US 20030046421A1 · Horvitz · 2003 [cited by examiner]
US 20140095166A1 · Bell et al. · 2014 [cited by applicant]
US 20160379666A1 · Christian · 2016 [cited by examiner]
US 20170196026A1 · Daoud · 2017 [cited by examiner]
US 20190259378A1 · Khadloya · 2019 [cited by examiner]
US 20190362022A1 · Haukioja et al. · 2019 [cited by applicant]
US 20200160985A1 · Kusuma · 2020 [cited by examiner]
US 20210012770A1 · Choudhary · 2021 [cited by examiner]
US 20210076966A1 · Grantcharov · 2021 [cited by examiner]
US 20220027417A1 · Katz · 2022 [cited by examiner]
US 20230082528A1 · Croghan · 2023 [cited by examiner]
CN 102088911A · 2011 [cited by applicant]
CN 103793515A · 2014 [cited by applicant]
CN 111048114A · 2020 [cited by applicant]
EP 1850321A1 · 2007 [cited by applicant]
EP 3703361A1 · 2020 [cited by applicant]
JP 2018097239A · 2018 [cited by applicant]
WO 2017062047A1 · 2017 [cited by applicant]
Alertus Technologies LLC. et al., “Mass Notification for the Deaf and Hard of Hearing”, Sep. 18, 2020, 5 pp. [cited by applicant]
International Search Report and Written Opinion of International Application No. PCT/US2021/048496 dated Jan. 5, 2022, 13 pp. [cited by applicant]
Jain et al., “Exploring Sound Awareness in the Home for People who are Deaf or Hard of Hearing”, CHI 2019, Glasgow, Scotland UK, May 4, 2019, 11 pp. [cited by applicant]
Varshney, “Can Google Live Transcribe Actually Solve Your Transcription Woes?”, YOURSTORY, Sep. 6, 2019, 6 pp. [cited by applicant]
Worthy, “Verbatim Transcription: Meaning and when it's Needed”, GMR Transcription Services, Inc., Mar. 25, 2019, 4 pp. [cited by applicant]
Response to Communication pursuant to Article 94(3) EPC dated Mar. 12, 2025, from counterpart European Application No. 21773993.7 filed Jun. 30, 2025, 4 pp. [cited by applicant]
Response to Communication Pursuant to Rules 161(1) and 162 EPC dated Apr. 20, 2023, from counterpart European Application No. 21773993.7, filed Oct. 4, 2023, 5 pp. [cited by applicant]
Communication pursuant to Article 94(3) EPC from counterpart European Application No. 21773993.7 dated Mar. 12, 2025, 8 pp. [cited by applicant]
Office Action from counterpart Japanese Application No. 2023-521411 dated Sep. 9, 2025, 6 pp. [cited by applicant]
Office Action from counterpart Korean Application No. 10-2023-7010080 dated Sep. 29, 2025, 14 pp. [cited by applicant]
Response to Office Action dated Sep. 29, 2025, from counterpart Korean Application No. 10-2023-7010080 filed Jan. 26, 2026, 46 pp. [cited by applicant]
First Office Action and Search Report from counterpart Chinese Application No. 202180066234.4 dated May 10, 2026, 19 pp. Translation Attached. [cited by applicant]
First Examination Report from counterpart Indian Application No. 202347017765 dated Jun. 18, 2026, 9 pp. [cited by applicant]
Jain et al., “HomeSound: An Iterative Field Deployment of an In-Home Sound Awareness System for Deaf or Hard of Hearing Users”, CHI '20: CHI Conference on Human Factors in Computing Systems, Apr. 2020, 25 pp. [cited by applicant]
Kemler, “New features to make audio more accessible on your phone”, blog.google /products-and-platforms/platforms/android/new-features-make-audio-more-accessible-your-phone/, May 15, 2019, 3 pp. [cited by applicant]
Rossignol, “iOS 14 Can Notify Users About Sounds Like Fire Alarms and Doorbells”, www.macrumors.com/2020/06/23/ios-14-sound-recognition-alarms-doorbells/, Jun. 23, 2020, 7 pp. [cited by applicant]
Decision of Rejection, and translation thereof, from counterpart Korean Application No. 10-2023-7010080 dated Jul. 20, 2026, 8 pp. [cited by applicant]