IP Library Granted Patent US 12,469,404
Granted Patent B2
US 12,469,404 · App. 18/659,722 · Granted Nov 11, 2025

Systems and methods for customizing playback of digital tutorials

Inventors: Nathalia S. Santos-Sheehan (Royersford, PA); Christopher Franklin (Drexel Hill, PA); Jennifer L. Holloway (Wallingford, PA); Daniel P. Rowan (Springfield, PA)
Assignee: Adeia Guides Inc.
G09B5/065
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,469,404
App. No.
18/659,722
Granted
Nov 11, 2025
Kind
B2
Abstract

Systems and methods are disclosed herein for continuing playback of a digital tutorial until a user interrupts the playback by signaling to the system that there is an issue or that the user needs help. The system, through detecting a recording that the user captured or a person's utterance (e.g., through passive voice monitoring) determines that the user's needs assistance with the digital tutorial. The system determines, based on the recording, that the user needs help to get to a specific step and play supplemental instructions to the user to get to the specific step.

Claims (51)

1 . A method comprising:

capturing, via a camera of a device, image data of a current state of a task;

storing the captured image data on a memory of the device;

detecting a user utterance;

determining a target state of the task based on the user utterance, wherein determining the target state of the task comprises:

detecting that the user utterance matches metadata associated with a particular state of the task; and

determining that the particular state of the task is the target state of the task;

determining, using a trained neural network, based on the image data, whether the current state of the task matches the target state of the task that was determined based on the user utterance; and

in response to determining that the current state of the task does not match a target state of the task, causing to be output a recommendation to bring the current state of the task to the target state of the task.

2 . The method of claim 1 , further comprising, generating the trained neural network by:

inputting a plurality of images to a neural network, each image of the plurality of images corresponding to a respective state of a plurality of states of the task; and

iteratively updating weights associated with nodes in the neural network.

3 . The method of claim 1 , wherein the task comprises a plurality of steps, and wherein determining, using the trained neural network, based on the image data, whether the current state of the task matches a target state of the task, comprises:

determining, using the trained neural network, based on the image data, a step, of the plurality of steps, corresponding to the current state of the task; and

determining whether the step corresponding to the current state of the task matches a step, of the plurality of steps, corresponding to the target state of the task.

4 . The method of claim 3 , wherein causing to be output the recommendation to bring the current state of the task to the target state of the task comprises:

identifying an instruction associated with the step corresponding to the current state of the task; and

causing to be output the instruction.

5 . The method of claim 1 , wherein the recommendation comprises at least one of an audio based output and a visual based output.

6 . The method of claim 1 , further comprising prompting a user to enable video capture, prior to the capturing, via the camera of the device, image data of the current state of the task.

7 . The method of claim 1 , further comprising, in response to determining that the current state of the task matches the target state of the task:

identifying a next state associated with the task, wherein the next state follows the current state in a sequence of states associated with the task; and

outputting instructions corresponding to the next state.

8 . The method of claim 1 , further comprising accessing the trained neural network via a network connection of the device.

9 . A system comprising:

a camera;

a memory; and

control circuitry configured to:

cause to be captured, via the camera, image data of a current state of a task;

cause to be stored the captured image data on the memory;

detect a user utterance;

determine a target state of the task based on the user utterance, wherein determining the target state of the task comprises:

detecting that the user utterance matches metadata associated with a particular state of the task; and

determining that the particular state of the task is the target state of the task;

determine, using a trained neural network, based on the image data, whether the current state of the task matches the target state of the task that was determined based on the user utterance; and

in response to determining that the current state of the task does not match a target state of the task, cause to be output a recommendation to bring the current state of the task to the target state of the task.

10 . The system of claim 9 , wherein the control circuitry is further configured to generate the trained neural network by:

inputting a plurality of images to a neural network, each image of the plurality of images corresponding to a respective state of a plurality of states of the task; and

iteratively updating weights associated with nodes in the neural network.

11 . The system of claim 9 , wherein the task comprises a plurality of steps, and wherein the control circuitry is further configured, when determining, using the trained neural network, based on the image data, whether the current state of the task matches a target state of the task, to:

determine, using the trained neural network, based on the image data, a step, of the plurality of steps, corresponding to the current state of the task; and

determine whether the step corresponding to the current state of the task matches a step, of the plurality of steps, corresponding to the target state of the task.

12 . The system of claim 11 , wherein the control circuitry is further configured, when causing to be output the recommendation to bring the current state of the task to the target state of the task to:

identify an instruction associated with the step corresponding to the current state of the task; and

cause to be output the instruction.

13 . The system of claim 9 , further comprising at least one of a display and a speaker, wherein the recommendation comprises at least one of an audio based output and a visual based output.

14 . The system of claim 9 , wherein the control circuitry is further configured to prompt a user to enable video capture, prior to the capturing, via the camera, image data of the current state of the task.

15 . The system of claim 9 , wherein the control circuitry is further configured, in response to determining that the current state of the task matches the target state of the task, to:

identify a next state associated with the task, wherein the next state follows the current state in a sequence of states associated with the task; and

output instructions corresponding to the next state.

16 . The system of claim 9 , further comprising a network connection, wherein the control circuitry is further configured to access the trained neural network via the network connection.

Assignments (3)
SECURITY INTEREST Recorded May 28, 2025
From: ADEIA INC. (F/K/A XPERI HOLDING CORPORATION); ADEIA HOLDINGS INC.; ADEIA MEDIA HOLDINGS INC.; ADEIA IMAGING LLC; ADEIA MEDIA LLC; ADEIA MEDIA SOLUTIONS INC.; ADEIA SEMICONDUCTOR BONDING TECHNOLOGIES INC.; ADEIA TECHNOLOGIES INC.; ADEIA GUIDES INC.; ADEIA SOLUTIONS LLC; ADEIA SEMICONDUCTOR ADVANCED TECHNOLOGIES INC.; ADEIA SEMICONDUCTOR SOLUTIONS LLC; ADEIA SEMICONDUCTOR INTELLECTUAL PROPERTY LLC; ADEIA SEMICONDUCTOR TECHNOLOGIES LLC; ADEIA PUBLISHING INC.
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 071454/0343 →
CHANGE OF NAME Recorded Oct 3, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069106/0164 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 5, 2024
From: SANTOS-SHEEHAN, NATHALIA S.; FRANKLIN, CHRISTOPHER; HOLLOWAY, JENNIFER L.; ROWAN, DANIEL P.
To: ROVI GUIDES, INC.
Reel/Frame 067624/0402 →
Continuity (3)
Continuation 17875784 · Jul 28, 2022
Division 16225040 · Dec 19, 2018
Related Publication 20250046202A1 · Feb 6, 2025
References Cited (17)
US 7761892B2 · Ellis et al. · 2010 [cited by applicant]
US 8990274B1 · Hwang · 2015 [cited by examiner]
US 12008920B2 · Santos-Sheehan et al. · 2024 [cited by applicant]
US 20130036353A1 · Zavesky et al. · 2013 [cited by applicant]
US 20140278391A1 · Braho · 2014 [cited by examiner]
US 20180268865A1 · Ekambaram et al. · 2018 [cited by applicant]
US 20200193264A1 · Zavesky · 2020 [cited by examiner]
US 20200202735A1 · Santos-Sheehan et al. · 2020 [cited by applicant]
US 20200202848A1 · Santos-Sheehan et al. · 2020 [cited by applicant]
US 20200272222A1 · Goela · 2020 [cited by examiner]
US 20220366805A1 · Santos-Sheehan et al. · 2022 [cited by applicant]
PCT International Search Report for International Application No. PCT/US2019/065473, dated Jul. 21, 2020 (21 pages). [cited by applicant]
Amir , et al., Using Audio Time Scale Modification for Video Browsing, Proceedings of the 33rd Hawaii International on Systems Sciences, Jul. 22, 2000 (10 Pages). [cited by applicant]
Brako , et al., “Loop Process Book MHCI+D Capstone, Summer 2015 Table Contents Introduction,” Jan. 2, 2017 (48 pages). [cited by applicant]
Chang , et al., “How to Design Voice Based Navigation for How-To-Videos,” Human Factors in Computing Systems, May 2, 2019 (11 pages). [cited by applicant]
Kim , et al., “Data-driven interaction techniques for improving navigation of educational videos,” User Interface Software and Technology, Oct. 5-8, 2014 (10 pages). [cited by applicant]
Yadav , et al., “Content-driven Multi-Modal Techniques for Non-linear Video Navigation,” Intelligent User Interfaces, Mar. 29-Apr. 1, 2015 (12 pages). [cited by applicant]