IP Library › Granted Patent US 12,626,159
Granted Patent B2
US 12,626,159 · App. 18/175,097 · Granted May 12, 2026

Music recommendation method and apparatus

Inventors: Shu Fang (Beijing, CN); Libin Zhang (Beijing, CN)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G06N5/04G06F16/635G06V20/50G06V40/20G06F3/012G06F3/165
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,626,159
App. No.
18/175,097
Granted
May 12, 2026
Kind
B2
Abstract

A music recommendation method and apparatus are provided, to determine an attention mode of a user in a complex environment by using viewpoint information of the user, thereby more precisely implementing music matching. According to a first aspect, a music recommendation method is provided. The method includes: receiving visual data of a user (S 501 ); obtaining at least one attention unit and attention duration of the at least one attention unit based on the visual data (S 502 ); determining an attention mode of the user based on the attention duration of the at least one attention unit (S 503 ); and determining recommended music information based on the attention mode (S 504 ).

Claims (49)

1 . A music recommendation method comprising:

receiving visual data of a user, wherein the visual data comprises picture information viewed by the user;

obtaining at least one attention unit and attention duration of the at least one attention unit based on the visual data, wherein obtaining the at least one attention unit and the attention duration of the at least one attention unit comprises: obtaining the at least one attention unit based on the picture information, determining a similarity between a first attention unit of the at least one attention unit and a second attention unit of the at least one attention unit, and when the similarity is greater than or equal to a first threshold, determining that attention duration of the second attention unit is equal to a sum of attention duration of the first attention unit and the attention duration of the second attention unit, wherein the first attention unit and the second attention unit are attention units at different moments in time, and the attention duration of the at least one attention unit comprises the attention duration of the second attention unit;

determining an attention mode of the user based on the attention duration of the at least one attention unit; and

determining recommended music information based on the attention mode.

2 . The method according to claim 1 , wherein the visual data further comprises viewpoint information of the user, and the viewpoint information comprises a position of a viewpoint and attention duration of the viewpoint.

3 . The method according to claim 1 , wherein determining the attention mode of the user comprises:

when a standard deviation of the attention duration of the at least one attention unit is greater than or equal to a second threshold, determining that the attention mode of the user is a staring mode; or

when a standard deviation of the attention duration of the at least one attention unit is less than a second threshold, determining that the attention mode of the user is a scanning mode.

4 . The method according to claim 1 , wherein determining the recommended music information comprises:

when the attention mode is a scanning mode, determining the music information based on the picture information; or

when the attention mode is a staring mode, determining the music information based on an attention unit with highest attention in the at least one attention unit.

5 . The method according to claim 4 , wherein determining the recommended music information further comprises:

determining a behavior state of the user at each moment within a first time period based on the attention mode;

determining a behavior state of the user within the first time period based on the state at each moment; and

determining the music information based on the behavior state within the first time period.

6 . The method according to claim 1 , wherein obtaining the at least one attention unit and the attention duration of the at least one attention unit further comprises: when the similarity is less than the first threshold, reserving the attention duration of the first attention unit and the attention duration of the second attention unit, wherein the attention duration of the at least one attention unit comprises the attention duration of the first attention unit and the attention duration of the second attention unit.

7 . A music recommendation apparatus, comprising:

a memory to store executable instructions thereon, and

a processor coupled to the memory to execute the executable instructions to cause the music recommendation apparatus to perform operations comprising:

receiving visual data of a user, wherein the visual data comprises picture information viewed by the user;

obtaining at least one attention unit and attention duration of the at least one attention unit based on the visual data, wherein obtaining the at least one attention unit and the attention duration of the at least one attention unit comprises: obtaining the at least one attention unit based on the picture information, determining a similarity between a first attention unit of the at least one attention unit and a second attention unit of the at least one attention unit, and when the similarity is greater than or equal to a first threshold, determining that attention duration of the second attention unit is equal to a sum of attention duration of the first attention unit and the attention duration of the second attention unit, wherein the first attention unit and the second attention unit are attention units at different moments in time, and the attention duration of the at least one attention unit comprises the attention duration of the second attention unit;

determining an attention mode of the user based on the attention duration of the at least one attention unit; and

determining recommended music information based on the attention mode.

8 . The music recommendation apparatus according to claim 7 , wherein the visual data further comprises viewpoint information of the user, and the viewpoint information comprises a position of a viewpoint and attention duration of the viewpoint.

9 . The music recommendation apparatus according to claim 7 , wherein determining the attention mode of the user comprises:

when a standard deviation of the attention duration of the at least one attention unit is greater than or equal to a second threshold, determining that the attention mode of the user is a staring mode; or

when a standard deviation of the attention duration of the at least one attention unit is less than a second threshold, determining that the attention mode of the user is a scanning mode.

10 . The music recommendation apparatus according to claim 7 , wherein determining the recommended music information comprises:

when the attention mode is a scanning mode, determining the music information based on the picture information; or

when the attention mode is a staring mode, determining the music information based on an attention unit with highest attention in the at least one attention unit.

11 . The music recommendation apparatus according to claim 10 , wherein determining the recommended music information further comprises:

determining a behavior state of the user at each moment within a first time period based on the attention mode;

determining a behavior state of the user within the first time period based on the state at each moment; and

determining the music information based on the behavior state within the first time period.

12 . The music recommendation apparatus according to claim 7 , wherein obtaining the at least one attention unit and the attention duration of the at least one attention unit further comprises: when the similarity is less than the first threshold, reserving the attention duration of the first attention unit and the attention duration of the second attention unit, wherein the attention duration of the at least one attention unit comprises the attention duration of the first attention unit and the attention duration of the second attention unit.

13 . A non-transitory computer-readable storage medium storing executable instructions thereon, that when executed by a processor of an apparatus, cause the apparatus to perform operations comprising:

receiving visual data of a user, wherein the visual data comprises picture information viewed by the user;

obtaining at least one attention unit and attention duration of the at least one attention unit based on the visual data, wherein obtaining the at least one attention unit and the attention duration of the at least one attention unit comprises: obtaining the at least one attention unit based on the picture information, determining a similarity between a first attention unit of the at least one attention unit and a second attention unit of the at least one attention unit, and when the similarity is greater than or equal to a first threshold, determining that attention duration of the second attention unit is equal to a sum of attention duration of the first attention unit and the attention duration of the second attention unit, wherein the first attention unit and the second attention unit are attention units at different moments in time, and the attention duration of the at least one attention unit comprises the attention duration of the second attention unit;

determining an attention mode of the user based on the attention duration of the at least one attention unit; and

determining recommended music information based on the attention mode.

14 . The non-transitory computer-readable storage medium according to claim 13 , wherein the visual data further comprises viewpoint information of the user, and the viewpoint information comprises a position of a viewpoint and attention duration of the viewpoint.

15 . The non-transitory computer-readable storage medium according to claim 13 , wherein determining the attention mode of the user comprises:

when a standard deviation of the attention duration of the at least one attention unit is greater than or equal to a second threshold, determining that the attention mode of the user is a staring mode; or

when a standard deviation of the attention duration of the at least one attention unit is less than a second threshold, determining that the attention mode of the user is a scanning mode.

16 . The non-transitory computer-readable storage medium according to claim 13 , wherein determining the recommended music information comprises:

when the attention mode is a scanning mode, determining the music information based on the picture information; or

when the attention mode is a staring mode, determining the music information based on an attention unit with highest attention in the at least one attention unit.

17 . The non-transitory computer-readable storage medium according to claim 13 , wherein obtaining the at least one attention unit and the attention duration of the at least one attention unit further comprises: when the similarity is less than the first threshold, reserving the attention duration of the first attention unit and the attention duration of the second attention unit, wherein the attention duration of the at least one attention unit comprises the attention duration of the first attention unit and the attention duration of the second attention unit.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 30, 2023
From: FANG, SHU; ZHANG, LIBIN
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 063799/0027 →
Continuity (2)
Continuation PCTCN2020112414 · Aug 31, 2020
Related Publication 20230206093A1 · Jun 29, 2023
References Cited (13)
US 12099558B2 · Zheng · 2024 [cited by examiner]
US 20150271571A1 · Laksono et al. · 2015 [cited by applicant]
US 20160109941A1 · Govindarajeswaran · 2016 [cited by examiner]
US 20190102706A1 · Frank · 2019 [cited by examiner]
US 20190221191A1 · Chhipa · 2019 [cited by examiner]
US 20200327378A1 · Smith · 2020 [cited by examiner]
US 20210065217A1 · Glaser · 2021 [cited by examiner]
CN 105589506A · 2016 [cited by applicant]
CN 109151176A · 2019 [cited by applicant]
CN 111241385A · 2020 [cited by applicant]
CN 111400605A · 2020 [cited by applicant]
JP 5541529B2 · 2014 [cited by applicant]
Jacob, S., Ishimaru, S., Bukhari, S. S., & Dengel, A. (Jun. 2018). Gaze-based interest detection on newspaper articles. In Proceedings of the 7th Workshop on Pervasive Eye Tracking and Mobile Eye-Based Interaction (pp. … [cited by examiner]