IP Library › Granted Patent US 12,608,427
Granted Patent B2
US 12,608,427 · App. 18/671,825 · Granted Apr 21, 2026

Drill back to original audio clip in virtual assistant initiated lists and reminders

Inventor: Michael Patrick Rodgers (Lake Oswego, OR)
Assignee: Oracle International Corporation
G06F16/9038G10L15/22G10L2015/0631G10L2015/223G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,608,427
App. No.
18/671,825
Granted
Apr 21, 2026
Kind
B2
Abstract

Techniques for drilling back to an original audio clip in virtual assistant initiated lists and reminders are disclosed. The system may receive audio input comprising a first request. Based on the first request, the system may schedule an action to be performed by the virtual assistant platform. The system stores at least a portion of the audio input and a mapping between the action and at least the portion of the audio input. The system performs the action. Subsequent to performing the action, the system receives a second request for audio playback of the first request corresponding to the action. The system retrieves at least the portion of the audio input based on the mapping between the action and at least the portion of the audio input, and plays at least the portion of the audio input comprising the first request.

Claims (105)

1 . A non-transitory computer readable medium comprising instructions which, when executed by one or more hardware processors, causes performance of operations comprising:

executing a virtual assistant platform on a first device, the virtual assistant platform configured to receive audio inputs on the first device and to initiate actions corresponding to the audio inputs;

receiving, by the virtual assistant platform, a first audio input comprising a first request corresponding to a first action to be initiated by the virtual assistant platform;

determining and storing, by the virtual assistant platform, contextual data associated with at least one of (a) a configuration of the first device and (b) one or more applications executing on the first device at a time when the first audio input was received;

receiving, by the virtual assistant platform, a second request for audio playback of the first audio input;

based at least on receiving the second request:

retrieving, by the virtual assistant platform executing on the first device, at least a portion of the first audio input;

retrieving, by the virtual assistant platform, the contextual data associated with at least one of (a) a configuration of the first device and (b) one or more applications executing on the first device at the time when the first audio input was received; and

presenting, by the virtual assistant platform, the contextual data together with the portion of the first audio input in a user interface of the first device or a second device,

wherein the contextual data comprises at least one of:

geo-location data corresponding to a geo-location of the first device at the time when the first audio input was received;

first application data corresponding to a first application being executed on the first device at the time when the first audio input was received;

second application data corresponding to a second application that was a last application to be executed on the first device prior to the time when the first audio input was received; and

device configuration data corresponding to a configuration of the first device at the time when the first audio input was received.

2 . The non-transitory computer readable medium of claim 1 , wherein the operations further comprise:

analyzing, by the virtual assistant platform, the first audio input to identify a particular action corresponding to the first audio input; and

determining the virtual assistant platform failed to identify any particular action, including the first action, corresponding to the first audio input,

wherein retrieving the contextual data is performed based on (a) receiving the second request, and (b) determining the virtual assistant platform failed to identify any particular action corresponding to the first audio input.

3 . The non-transitory computer readable medium of claim 1 , wherein the operations further comprise:

analyzing, by the virtual assistant platform, the first audio input to identify a particular action corresponding to the first audio input; and

determining the virtual assistant platform identified a second action, different from the first action, as corresponding to the first audio input,

wherein retrieving the contextual data is performed based on (a) receiving the second request, and (b) determining virtual assistant platform identified the second action as corresponding to the first audio input.

4 . The non-transitory computer readable medium of claim 1 , wherein the operations further comprise:

analyzing, by the virtual assistant platform, the first audio input to identify a particular action corresponding to the first audio input;

determining the virtual assistant platform identified a second action as corresponding to the first audio input; and

computing a confidence score associated with a pairing of the second action with the first audio input,

wherein retrieving the contextual data is performed based on (a) receiving the second request, and (b) determining confidence score does not meet a threshold confidence value.

5 . The non-transitory computer readable medium of claim 4 , wherein the confidence score is based at least on a confidence that the virtual assistant platform correctly classified a set of phonemes included in the first audio input.

6 . The non-transitory computer readable medium of claim 4 , wherein the operations further comprise:

generating a transcript of the first audio input,

wherein the confidence score is based at least on a confidence that the transcript correctly identifies words included in the first audio input.

7 . The non-transitory computer readable medium of claim 1 , wherein the operations further comprise:

identifying, by the virtual assistant platform, a request type corresponding to the first request,

wherein retrieving the contextual data comprises:

identifying a first type of contextual data corresponding to the first request; and

based on identifying the first type of contextual data: retrieving the contextual data of the first type.

8 . A method comprising:

executing a virtual assistant platform on a first device, the virtual assistant platform configured to receive audio inputs on the first device and to initiate actions corresponding to the audio inputs;

receiving, by the virtual assistant platform, a first audio input comprising a first request corresponding to a first action to be initiated by the virtual assistant platform;

determining and storing, by the virtual assistant platform, contextual data associated with at least one of (a) a configuration of the first device and (b) one or more applications executing on the first device at a time when the first audio input was received;

receiving, by the virtual assistant platform, a second request for audio playback of the first audio input;

based at least on receiving the second request:

retrieving, by the virtual assistant platform executing on the first device, at least a portion of the first audio input;

retrieving, by the virtual assistant platform, the contextual data associated with at least one of (a) a configuration of the first device and (b) one or more applications executing on the first device at the time when the first audio input was received; and

presenting, by the virtual assistant platform, the contextual data together with the portion of the first audio input in a user interface of the first device or a second device,

wherein the contextual data comprises at least one of:

geo-location data corresponding to a geo-location of the first device at the time when the first audio input was received;

first application data corresponding to a first application being executed on the first device at the time when the first audio input was received;

second application data corresponding to a second application that was a last application to be executed on the first device prior to the time when the first audio input was received; and

device configuration data corresponding to a configuration of the first device at the time when the first audio input was received.

9 . The method of claim 8 , further comprising:

analyzing, by the virtual assistant platform, the first audio input to identify a particular action corresponding to the first audio input; and

determining the virtual assistant platform failed to identify any particular action, including the first action, corresponding to the first audio input,

wherein retrieving the contextual data is performed based on (a) receiving the second request, and (b) determining the virtual assistant platform failed to identify any particular action corresponding to the first audio input.

10 . The method of claim 8 , further comprising:

analyzing, by the virtual assistant platform, the first audio input to identify a particular action corresponding to the first audio input; and

determining the virtual assistant platform identified a second action, different from the first action, as corresponding to the first audio input,

wherein retrieving the contextual data is performed based on (a) receiving the second request, and (b) determining virtual assistant platform identified the second action as corresponding to the first audio input.

11 . The method of claim 8 , further comprising:

analyzing, by the virtual assistant platform, the first audio input to identify a particular action corresponding to the first audio input;

determining the virtual assistant platform identified a second action as corresponding to the first audio input; and

computing a confidence score associated with a pairing of the second action with the first audio input,

wherein retrieving the contextual data is performed based on (a) receiving the second request, and (b) determining confidence score does not meet a threshold confidence value.

12 . The method of claim 11 , wherein the confidence score is based at least on a confidence that the virtual assistant platform correctly classified a set of phonemes included in the first audio input.

13 . The method of claim 11 , further comprising:

generating a transcript of the first audio input,

wherein the confidence score is based at least on a confidence that the transcript correctly identifies words included in the first audio input.

14 . The method of claim 8 , further comprising:

identifying, by the virtual assistant platform, a request type corresponding to the first request,

wherein retrieving the contextual data comprises:

identifying a first type of contextual data corresponding to the first request; and

based on identifying the first type of contextual data: retrieving the contextual data of the first type.

15 . A system comprising:

at least one device including a hardware processor;

the system being configured to perform operations comprising:

executing a virtual assistant platform on a first device, the virtual assistant platform configured to receive audio inputs on the first device and to initiate actions corresponding to the audio inputs;

receiving, by the virtual assistant platform, a first audio input comprising a first request corresponding to a first action to be initiated by the virtual assistant platform;

determining and storing, by the virtual assistant platform, contextual data associated with at least one of (a) a configuration of the first device and (b) one or more applications executing on the first device at a time when the first audio input was received;

receiving, by the virtual assistant platform, a second request for audio playback of the first audio input;

based at least on receiving the second request:

retrieving, by the virtual assistant platform executing on the first device, at least a portion of the first audio input;

retrieving, by the virtual assistant platform, the contextual data associated with at least one of (a) a configuration of the first device and (b) one or more applications executing on the first device at the time when the first audio input was received; and

presenting, by the virtual assistant platform, the contextual data together with the portion of the first audio input in a user interface of the first device or a second device,

wherein the contextual data comprises at least one of:

geo-location data corresponding to a geo-location of the first device at the time when the first audio input was received;

first application data corresponding to a first application being executed on the first device at the time when the first audio input was received;

second application data corresponding to a second application that was a last application to be executed on the first device prior to the time when the first audio input was received; and

device configuration data corresponding to a configuration of the first device at the time when the first audio input was received.

16 . The system of claim 15 , wherein the operations further comprise:

analyzing, by the virtual assistant platform, the first audio input to identify a particular action corresponding to the first audio input; and

determining the virtual assistant platform failed to identify any particular action, including the first action, corresponding to the first audio input,

wherein retrieving the contextual data is performed based on (a) receiving the second request, and (b) determining the virtual assistant platform failed to identify any particular action corresponding to the first audio input.

17 . The system of claim 15 , wherein the operations further comprise:

analyzing, by the virtual assistant platform, the first audio input to identify a particular action corresponding to the first audio input; and

determining the virtual assistant platform identified a second action, different from the first action, as corresponding to the first audio input,

wherein retrieving the contextual data is performed based on (a) receiving the second request, and (b) determining virtual assistant platform identified the second action as corresponding to the first audio input.

18 . The system of claim 15 , wherein the operations further comprise:

analyzing, by the virtual assistant platform, the first audio input to identify a particular action corresponding to the first audio input;

determining the virtual assistant platform identified a second action as corresponding to the first audio input; and

computing a confidence score associated with a pairing of the second action with the first audio input,

wherein retrieving the contextual data is performed based on (a) receiving the second request, and (b) determining confidence score does not meet a threshold confidence value.

19 . The system of claim 18 , wherein the confidence score is based at least on a confidence that the virtual assistant platform correctly classified a set of phonemes included in the first audio input.

20 . The system of claim 18 , wherein the operations further comprise:

generating a transcript of the first audio input,

wherein the confidence score is based at least on a confidence that the transcript correctly identifies words included in the first audio input.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 27, 2024
From: RODGERS, MICHAEL PATRICK
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 067531/0524 →
Continuity (2)
Continuation 17140783 · Jan 4, 2021
Related Publication 20240311429A1 · Sep 19, 2024
References Cited (79)
US 5355311A · Horioka · 1994 [cited by applicant]
US 5602963A · Bissonnette et al. · 1997 [cited by applicant]
US 6438520B1 · Curt et al. · 2002 [cited by applicant]
US 9009592B2 · Friend et al. · 2015 [cited by applicant]
US 9318108B2 · Gruber et al. · 2016 [cited by applicant]
US 10319383B1 · Stekkelpak · 2019 [cited by examiner]
US 10380208B1 · Brahmbhatt · 2019 [cited by examiner]
US 10403272B1 · Fanty · 2019 [cited by examiner]
US 10733987B1 · Govender et al. · 2020 [cited by applicant]
US 10855485B1 · Zhou · 2020 [cited by examiner]
US 11044567B1 · Li et al. · 2021 [cited by applicant]
US 11218666B1 · Haas et al. · 2022 [cited by applicant]
US 11238855B1 · Goetz · 2022 [cited by examiner]
US 11380308B1 · Pandey · 2022 [cited by examiner]
US 11398239B1 · Garrod · 2022 [cited by applicant]
US 11527251B1 · Moeller · 2022 [cited by examiner]
US 11551685B2 · Rastrow et al. · 2023 [cited by applicant]
US 11574637B1 · Kumar et al. · 2023 [cited by applicant]
US 20020128837A1 · Morin · 2002 [cited by examiner]
US 20050182631A1 · Lee · 2005 [cited by examiner]
US 20060116877A1 · Pickering et al. · 2006 [cited by applicant]
US 20090187410A1 · Wilpon et al. · 2009 [cited by applicant]
US 20100088100A1 · Lindahl · 2010 [cited by examiner]
US 20100169359A1 · Barrett et al. · 2010 [cited by applicant]
US 20100185445A1 · Comerford · 2010 [cited by examiner]
US 20100274469A1 · Takahata et al. · 2010 [cited by applicant]
US 20120035925A1 · Friend · 2012 [cited by examiner]
US 20120265535A1 · Bryant-Rich et al. · 2012 [cited by applicant]
US 20130030794A1 · Ikeda · 2013 [cited by examiner]
US 20130103399A1 · Gammon · 2013 [cited by examiner]
US 20140081634A1 · Forutanpour · 2014 [cited by examiner]
US 20140222436A1 · Binder et al. · 2014 [cited by applicant]
US 20140278444A1 · Larson · 2014 [cited by examiner]
US 20150039299A1 · Weinstein · 2015 [cited by examiner]
US 20150040012A1 · Faaborg et al. · 2015 [cited by applicant]
US 20150162000A1 · Marti et al. · 2015 [cited by applicant]
US 20150319546A1 · Sprague · 2015 [cited by examiner]
US 20160093289A1 · Pollet · 2016 [cited by applicant]
US 20160225370A1 · Kannan et al. · 2016 [cited by applicant]
US 20180053507A1 · Wang · 2018 [cited by examiner]
US 20180144747A1 · Skarbovsky · 2018 [cited by examiner]
US 20180174582A1 · Fanty · 2018 [cited by examiner]
US 20180300421A1 · Andreica · 2018 [cited by examiner]
US 20180315417A1 · Flaks et al. · 2018 [cited by applicant]
US 20190013008A1 · Kunitake et al. · 2019 [cited by applicant]
US 20190066670A1 · White · 2019 [cited by examiner]
US 20190180343A1 · Arnett · 2019 [cited by examiner]
US 20190205748A1 · Fukuda · 2019 [cited by examiner]
US 20190213997A1 · Aaron et al. · 2019 [cited by applicant]
US 20190235887A1 · Hemaraj · 2019 [cited by examiner]
US 20190236412A1 · Zhao et al. · 2019 [cited by applicant]
US 20200020329A1 · Gordon et al. · 2020 [cited by applicant]
US 20200027456A1 · Kim et al. · 2020 [cited by applicant]
US 20200075006A1 · Chen · 2020 [cited by examiner]
US 20200143797A1 · Manoharan · 2020 [cited by examiner]
US 20200152175A1 · Dernoncourt · 2020 [cited by applicant]
US 20200160865A1 · Michaely · 2020 [cited by examiner]
US 20200175961A1 · Thomson et al. · 2020 [cited by applicant]
US 20200294497A1 · Kirazci et al. · 2020 [cited by applicant]
US 20200349943A1 · Elangovan et al. · 2020 [cited by applicant]
US 20200380968A1 · Hatfield · 2020 [cited by examiner]
US 20200380985A1 · Gada · 2020 [cited by examiner]
US 20210043205A1 · Lee · 2021 [cited by examiner]
US 20210065679A1 · Finlay et al. · 2021 [cited by applicant]
US 20210082397A1 · Kennewick · 2021 [cited by examiner]
US 20210082419A1 · Tran · 2021 [cited by examiner]
US 20210124803A1 · Alloh et al. · 2021 [cited by applicant]
US 20210132784A1 · Conlon · 2021 [cited by examiner]
US 20210360109A1 · Rico Rdenas · 2021 [cited by applicant]
US 20220020363A1 · Vasquez · 2022 [cited by examiner]
US 20220101835A1 · Freed et al. · 2022 [cited by applicant]
US 20220238088A1 · Danjyo · 2022 [cited by applicant]
US 20220310089A1 · Aleksic et al. · 2022 [cited by applicant]
US 20250181468A1 · Rajan S · 2025 [cited by examiner]
US 20250358581A1 · Wu · 2025 [cited by examiner]
CN 114283810A · 2022 [cited by applicant]
Record and replay: How a Canadian-made app is aiming to help Alzheimer's patients improve their daily lives, available online at <https://www.theglobeandmail.com/canada/article-toronto-teams-hippocamera-a-high-tech-memo… [cited by applicant]
Vemuri et al., “An Audio-Based Personal Memory Aid,” MIT Media Lab 20 Ames St. Cambridge, MA 02139 USA, 18 pages. [cited by applicant]
Vemuri et al., “Improving Speech Playback Using Time-Compression and Speech Recognition,” CHI Paper, vol. 6, No. 1, Apr. 24-29, 2004, pp. 295-302. [cited by applicant]