IP Library Granted Patent US 12,445,682
Granted Patent B2
US 12,445,682 · App. 18/434,135 · Granted Oct 14, 2025

Method and system for generating a visual composition of user reactions in a shared content viewing session

Inventors: Ronica Jethwa (Mountain View, CA); Sunil Ramesh (Cupertino, CA); Michael Cutter (Golden, CO); Karina Levitian (Austin, TX)
Assignee: Roku, Inc.
H04N21/44218G06V10/764G06V40/176G06V40/23H04N21/8146H04N21/854
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,445,682
App. No.
18/434,135
Granted
Oct 14, 2025
Kind
B2
Abstract

In one aspect, an example method in connection with a shared content viewing session in which multiple users are receiving and viewing respective instances of the same media content in a synchronized manner is disclosed. The example method includes (i) detecting that each of the multiple users respectively exhibits a threshold extent of physical reaction around the same time; (ii) responsive to the detecting, for each of the multiple users, generating and/or storing respective visual content representing that user's physical reaction; (iii) generating a visual content composition that includes at least the generated and/or stored visual content for each of the multiple users; and (iv) outputting for presentation, the generated visual content composition.

Claims (43)

1. A method for use in connection with a shared content viewing session in which multiple users are receiving and viewing respective instances of media content in a synchronized manner, the method comprising:

accessing metadata associated with the media content, wherein the metadata specifies (i) a time point in the media content at which a particular event occurs, and (ii) an expected type of physical reaction associated with the particular event;

near the time point in the media content at which a particular event occurs, detecting that (i) each of the multiple users respectively exhibits a threshold extent of physical reaction around a same viewing time within the shared content viewing session, and (ii) the expected type of physical reaction matches a type of the physical reaction of the multiple users;

responsive to the detecting, for each of the multiple users, generating and/or storing respective visual content representing that user's physical reaction;

generating a visual content composition that includes at least the generated and/or stored visual content for each of the multiple users; and

outputting for presentation, the generated visual content composition.

2. The method of claim 1 , wherein exhibiting the threshold extent of physical reaction comprises exhibiting a threshold change in facial expression around the same time.

3. The method of claim 1 , wherein exhibiting the threshold extent of physical reaction comprises exhibiting a threshold change in body language expression around the same time.

4. The method of claim 1 , wherein around the same viewing time comprises within a given time period range.

5. The method of claim 4 , wherein the given time period range is between 0.5 seconds and 3.0 seconds.

6. The method of claim 1 , wherein detecting that each of the multiple users respectively exhibits a threshold extent of physical reaction around the same viewing time comprises:

for each of the multiple users:

receiving visual data of the user captured by a camera,

detecting a set of physical features in the visual data, and

based on the detected set of physical features, determining a physical reaction exhibited by the user using a physical reaction model comprising a classifier configured to map each of a plurality of physical reactions to a corresponding set of physical features.

7. The method of claim 1 , wherein for each of the multiple users, generating and/or storing respective visual content representing that user's physical reaction comprises storing respective visual content captured by a camera of that user.

8. The method of claim 1 , wherein for each of the multiple users, generating and/or storing respective visual content representing that user's physical reaction comprises generating and/or storing respective visual content of an avatar generated for that user.

9. The method of claim 1 , wherein the generated visual content composition further includes a portion of the media content that corresponds to a time point or range at or during which the users' physical reactions occurred.

10. The method of claim 1 , wherein outputting for presentation, the generated visual content composition comprises, transmitting the generated visual content composition to multiple content-presentation devices, respectively associated with the multiple users.

11. The method of claim 10 , wherein at least one of the multiple content-presentation devices is a television.

12. A computing system configured for performing a set of acts in connection a shared content viewing session in which multiple users are receiving and viewing respective instances of media content in a synchronized manner, the set of acts comprising:

accessing metadata associated with the media content, wherein the metadata specifies (i) a time point in the media content at which a particular event occurs, and (ii) an expected type of physical reaction associated with the particular event;

near the time point in the media content at which a particular event occurs, detecting that (i) each of the multiple users respectively exhibits a threshold extent of physical reaction around a same viewing time within the shared content viewing session, and (ii) the expected type of physical reaction matches a type of the physical reaction of the multiple users;

responsive to the detecting, for each of the multiple users, generating and/or storing respective visual content representing that user's physical reaction;

generating a visual content composition that includes at least the generated and/or stored visual content for each of the multiple users; and

outputting for presentation, the generated visual content composition.

13. The computing system of claim 12 , wherein exhibiting the threshold extent of physical reaction comprises exhibiting a threshold change in facial expression around the same time.

14. The computing system of claim 12 , wherein exhibiting the threshold extent of physical reaction comprises exhibiting a threshold change in body language expression around the same time.

15. The computing system of claim 12 , wherein around the same viewing time comprises within a given time period range.

16. The computing system of claim 15 , wherein the given time period range is between 0.5 seconds and 3.0 seconds.

17. The computing system of claim 12 , wherein detecting that each of the multiple users respectively exhibits a threshold extent of physical reaction around the same viewing time comprises:

for each of the multiple users:

receiving visual data of the user captured by a camera,

detecting a set of physical features in the visual data, and

based on the detected set of physical features, determining a physical reaction exhibited by the user using a physical reaction model comprising a classifier configured to map each of a plurality of physical reactions to a corresponding set of physical features.

18. The computing system of claim 12 , wherein for each of the multiple users, generating and/or storing respective visual content representing that user's physical reaction comprises storing respective visual content captured by a camera of that user.

19. The computing system of claim 12 , wherein for each of the multiple users, generating and/or storing respective visual content representing that user's physical reaction comprises generating and/or storing respective visual content of an avatar generated for that user.

20. A non-transitory computer-readable medium having stored thereon program instructions that upon execution by a computing system, cause performance of a set of acts in connection with a shared content viewing session in which multiple users are receiving and viewing respective instances of media content in a synchronized manner, the set of acts comprising:

accessing metadata associated with the media content, wherein the metadata specifies (i) a time point in the media content at which a particular event occurs, and (ii) an expected type of physical reaction associated with the particular event;

near the time point in the media content at which a particular event occurs, detecting that (i) each of the multiple users respectively exhibits a threshold extent of physical reaction around a same viewing time within the shared content viewing session, and (ii) the expected type of physical reaction matches a type of the physical reaction of the multiple users;

responsive to the detecting, for each of the multiple users, generating and/or storing respective visual content representing that user's physical reaction;

generating a visual content composition that includes at least the generated and/or stored visual content for each of the multiple users; and

outputting for presentation, the generated visual content composition.

Assignments (2)
SECURITY INTEREST Recorded Sep 18, 2024
From: ROKU, INC.
To: CITIBANK, N.A.
Reel/Frame 068982/0377 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 8, 2024
From: JETHWA, RONICA; RAMESH, SUNIL; CUTTER, MICHAEL; LEVITIAN, KARINA
To: ROKU, INC.
Reel/Frame 066415/0660 →
Continuity (2)
Continuation 18158546 · Jan 24, 2023
Related Publication 20240251127A1 · Jul 25, 2024
References Cited (20)
US 11451885B1 · Chandrashekar · 2022 [cited by applicant]
US 11936948B1 · Jethwa · 2024 [cited by examiner]
US 20050289627A1 · Lohman · 2005 [cited by examiner]
US 20120185887A1 · Newell · 2012 [cited by examiner]
US 20120296972A1 · Backer · 2012 [cited by applicant]
US 20130283162A1 · Aronsson · 2013 [cited by examiner]
US 20140172848A1 · Koukoumidis · 2014 [cited by examiner]
US 20160098169A1 · Herdy · 2016 [cited by examiner]
US 20160366203A1 · Blong · 2016 [cited by examiner]
US 20170099519A1 · Dang · 2017 [cited by examiner]
US 20170134803A1 · Shaw · 2017 [cited by examiner]
US 20190329134A1 · Shriram · 2019 [cited by applicant]
US 20210037295A1 · Strickland · 2021 [cited by applicant]
US 20220038774A1 · Paz · 2022 [cited by examiner]
US 20220132214A1 · Felman · 2022 [cited by applicant]
US 20220201250A1 · Schoenborn · 2022 [cited by examiner]
US 20220224966A1 · Smith · 2022 [cited by examiner]
US 20220385701A1 · Wang · 2022 [cited by examiner]
Hannan, “TensorFlow's New Model MoveNet Explained”, (Jul. 22, 2021) https://medium.com/@samhannan47/tensorflows-new-model-movenet-explained-3bdf80a8f073, retrieved Apr. 27, 2023*. [cited by applicant]
Bazarevsky et al., “BlazePose: On-device Real-time Body Pose tracking”, arXiv:2006:10204v1 [cs.CV] Jun. 17, 2020*. [cited by applicant]