IP Library › Granted Patent US 12,361,061
Granted Patent B2
US 12,361,061 · App. 17/661,051 · Granted Jul 15, 2025

Automatically creating task content

Inventors: Natasha Katherine McKenzie-Kelly (Salisbury, GB); Caroline Sarah Courtenay McNamara (Winchester, GB); Melita Saville (Winchester, GB); Clive Harris (Ropley, GB); Abigail Rose Bettle-Shaffer (Andover, GB); Timothy Andrew Moran (Southampton, GB)
Assignee: International Business Machines Corporation
G06F16/7844G06F3/167G06V20/46G06V20/635
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,061
App. No.
17/661,051
Granted
Jul 15, 2025
Kind
B2
Abstract

A method, computer program product, and computer system are provided for automatically creating task content. The method includes receiving a video file with an associated transcript of audio associated with the video file and identifying task action terms in the transcript. For each task action term, the method includes: locating a visual section of the video file corresponding to the task action term; capturing at least a portion of the visual section of the video file; and using image recognition for identifying information relating to one or more interface elements that are being interacted with in the visual section. The method includes generating a task instruction document including the task action terms augmented with interface element information and with at least a portion of the captured visual section.

Claims (64)

1. A computer-implemented method for automatically creating task content, said method carried out by one or more processors of a computer system and comprising:

receiving a video file with an associated transcript of audio associated with the video file;

identifying task action terms in the transcript;

for each task action term:

locating a visual section of the video file corresponding to the task action term;

capturing at least a portion of the visual section of the video file; and

using image recognition for identifying information relating to one or more interface elements that are being interacted with in the visual section; and

generating a task instruction document including the task action terms augmented with interface element information and with at least a portion of the captured visual section.

2. The method of claim 1 , wherein locating the visual section comprises:

determining a timestamp of the identified task action term in the video file; and

locating the visual section of the video file corresponding to the timestamp, wherein the timestamp is a time instance or a time range.

3. The method of claim 1 , wherein locating the visual section comprises:

visually analyzing the video file to identify visual elements matching the task action terms.

4. The method of claim 1 , wherein capturing the visual sections comprises:

storing the visual section from the video file in a repository;

providing a link in the transcript to the visual section in the repository; and

wherein generating a task instruction document comprises retrieving and displaying an image or clip of the visual section.

5. The method of claim 1 , further comprising verifying interface elements in the visual section corresponding to the task action term by:

visually analyzing the visual section to identify interface elements shown in the video file;

determining an identified interface element corresponding to the task action term; and

wherein capturing at least a portion of the visual section captures the identified interface element.

6. The method of claim 5 , wherein visually analyzing the visual section to identify interface elements uses edge detection techniques.

7. The method of claim 5 , wherein determining the interface element corresponding to the task action term comprises:

detecting motion in the interface element.

8. The method of claim 5 , wherein determining the interface element corresponding to the task action term comprises:

recognizing text in the interface element corresponding to the task action term.

9. The method of claim 1 , wherein generating the task instruction document comprises:

substituting ambiguous task action terms from the transcript with terms derived from analysis of the visual section.

10. The method of claim 1 , wherein the task instruction document combines written instructions and images in a structured topic-oriented format to share the information.

11. A system for automatically creating task content, comprising:

a processor and a memory configured to provide computer program instructions to the processor to execute the function of the components:

a receiving component for receiving a video file with an associated transcript of audio associated with the video file;

a task term component for identifying task action terms in the transcript;

for each task action term:

a visual section component for locating a visual section of the video file corresponding to the task action term;

a capturing component for capturing at least a portion of the visual section of the video file; and

an image analysis component for using image recognition for identifying information relating to one or more interface elements that are being interacted with in the visual section; and

a document generating component for generating a task instruction document including the task action terms augmented with interface element information and with at least a portion of the captured visual section.

12. The system of claim 11 , wherein the visual section component comprises:

a timestamp component for determining a timestamp of the identified task action term in the video file and locating the visual section of the video file corresponding to the timestamp, wherein the timestamp is a time instance or a time range.

13. The system of claim 11 , wherein the visual section component comprises:

a visual analysis component for visually analyzing the video file to identify visual elements matching the task action terms.

14. The system of claim 11 , wherein the capturing component comprises:

a storing component for storing the visual section from the video file in a repository;

a link component for providing a link in the transcript to the visual section in the repository; and

wherein the document generating component includes retrieving and displaying an image or clip of the visual section.

15. The system of claim 11 , further comprising a verifying component for verifying interface elements in the visual section corresponding to the task action term by:

an element component for visually analyzing the visual section to identify interface elements shown in the video file and determining an identified interface element corresponding to the task action term; and

wherein the capturing component captures at least a portion of the visual section captures the identified user interface element.

16. The system of claim 15 , wherein the element component comprises:

visually analyzing the visual section to identify interface elements using edge detection techniques.

17. The system of claim 15 , wherein the element component comprises:

detecting motion in the interface element and recognizing text in the interface element corresponding to the task action term.

18. The system of claim 11 , further comprising:

a disambiguating component for substituting ambiguous task action terms from the transcript with terms derived from analysis of the visual section.

19. The system of claim 11 , wherein the generating component generates the task instruction document by combining written instructions and images in a structured topic-oriented format to share the information.

20. A computer program product for automatically creating task content, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor to cause the processor to:

receive a video file with an associated transcript of audio associated with the video file;

identify a task action terms in the transcript;

for each task action term:

locate a visual section of the video file corresponding to the task action term;

capture at least a portion of the visual section of the video file; and

use image recognition for identifying information relating to one or more interface elements that are being interacted with in the visual section; and

generate a task instruction document including the task action terms augmented with interface element information and with at least a portion of the captured visual section.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 28, 2022
From: MCKENZIE-KELLY, NATASHA KATHERINE; MCNAMARA, CAROLINE SARAH COURTENAY; SAVILLE, MELITA; HARRIS, CLIVE; BETTLE-SHAFFER, ABIGAIL ROSE; MORAN, TIMOTHY ANDREW
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 059766/0863 →
Continuity (1)
Related Publication 20230350947A1 · Nov 2, 2023
References Cited (14)
US 9494437B2 · Subramanian · 2016 [cited by applicant]
US 10182083B2 · Cannon · 2019 [cited by applicant]
US 20040153969A1 · Rhodes · 2004 [cited by examiner]
US 20180349110A1 · Prasad Yellapragada · 2018 [cited by applicant]
US 20200342369A1 · Sridhara · 2020 [cited by examiner]
US 20230290146A1 · Chhaya · 2023 [cited by examiner]
CN 111447479A · 2020 [cited by applicant]
Apple, “Detecting Objects in Still Images”, Apple Inc., Accessed Mar. 21, 2022, 8 Pages. [cited by applicant]
Github, “Textsearch”, GitHub, Accessed Mar. 21, 2022, 13 Pages. [cited by applicant]
Google, “Get word timestamps”, Google Cloud, Accessed Mar. 21, 2022, 5 Pages. [cited by applicant]
IBM, “Speech to Text docs”, IBM Cloud, Jun. 9, 2021, 4 Pages. [cited by applicant]
Webster et al., “Automatically Extracting Procedural Knowledge from Instructional Texts using Natural Language Processing”, Conference: In Proceedings of the Eight International Conference on Language Resources and Eval… [cited by applicant]
Wikipedia, “Edge detection”, Wikipedia, Accessed Mar. 21, 2022, 11 Pages. [cited by applicant]
Mell et al., “The NIST Definition of Cloud Computing”, Recommendations of the National Institute of Standards and Technology, NIST Special Publication 800-145, Sep. 2011, 7 pages. [cited by applicant]