IP Library Granted Patent US 12705280
Granted Patent B2
US 12705280 · App. 18/326,261 · Granted Aug 11, 2026

Sound search using caption embeddings

Inventors: Rehana Mahfuz (San Diego, CA); Yinyi Guo (San Diego, CA); Erik Visser (San Diego, CA)
Assignee: QUALCOMM Incorporated
G06F16/685G06F16/632G06F16/638G06F16/686
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12705280
App. No.
18/326,261
Granted
Aug 11, 2026
Kind
B2
Abstract

A device includes one or more processors configured to generate one or more query caption embeddings based on a query. The processor(s) are further configured to select one or more caption embeddings from among a set of embeddings associated with a set of media files of a file repository. Each caption embedding represents a corresponding sound caption, and each sound caption includes a natural-language text description of a sound. The caption embedding(s) are selected based on a similarity metric indicative of similarity between the caption embedding(s) and the query caption embedding(s). The processor(s) are further configured to generate search results identifying one or more first media files of the set of media files. Each of the first media file(s) is associated with at least one of the caption embedding(s).

Claims (74)

1 . A device comprising:

one or more processors configured to:

generate one or more query caption embeddings representing multiple words that together form a semantic unit, based on a query, using a machine-learning model, wherein the query caption embeddings are based on sound captions extracted from the query that are natural-language descriptions of sounds and wherein outputs of the machine-learning model are one or more caption embeddings and locations of the one or more caption embeddings in a caption embedding space are indicative of semantic relationships among the sound captions, the semantic relationships corresponding to a similarity of semantic meaning between the sound captions, the similarity of semantic meaning distinct from a text comparison;

generate one or more context terms based on a portion of the query distinct from the semantic unit;

filter a set of media files to generate a set of filtered media files based on the context terms and metadata associated with the set of media files;

compare the one or more query caption embeddings, in the caption embedding space, to one or more second caption embeddings associated with the set of filtered media files, wherein each second caption embedding of the one or more second caption embeddings represents a corresponding sound caption and each sound caption includes a natural-language text description of a sound, to determine a first similarity metric, indicative of similarity between the one or more caption embeddings and the one or more query caption embeddings; and

generate search results identifying one or more first media files of the set of filtered media files, each of the one or more first media files associated with at least one of the one or more caption embeddings.

2 . The device of claim 1 , wherein the query includes a natural-language sequence of words describing a non-speech sound.

3 . The device of claim 1 , wherein the context terms include a media format type, a particular time period, a media source, a media generation location, or a combination thereof.

4 . The device of claim 1 , wherein the context terms include one or more words identifying a person, and wherein filtering the set of media files includes comparing the context terms to one or more object tags in the metadata, the one or more object tags identifying people visible in a corresponding video.

5 . The device of claim 1 , wherein a particular caption embedding describes a particular sound associated with a particular media file and wherein the particular caption embedding is associated with a time index indicating an approximate playback time of the particular media file at which the particular sound occurs.

6 . The device of claim 1 , wherein the set of media files includes one or more audio files, one or more video files, one or more virtual reality files, or a combination thereof.

7 . The device of claim 1 , wherein the query includes query audio data and wherein the one or more query caption embeddings are based on the query audio data.

8 . The device of claim 1 , wherein a particular media file of the set of media files is further associated with one or more audio embeddings of one or more sounds in the particular media file, and wherein the set of embeddings associated with the set of media files include the one or more audio embeddings.

9 . The device of claim 8 , wherein the one or more query caption embeddings are based on query audio data of the query and the one or more processors are further configured to:

generate one or more query audio embeddings based on the query audio data; and

compare the one or more query audio embeddings, in an audio embedding space, to one or more audio embeddings to determine, a second similarity metric indicative of similarity between the one or more audio embeddings and the one or more query audio embeddings, wherein the search results further identify one or more second media files of the set of filtered media files, each of the one or more second media files associated with at least one of the one or more audio embeddings.

10 . The device of claim 9 , wherein the one or more processors are further configured to rank the search results based on similarity values and wherein values of the first similarity metric associated with the one or more first media files are weighted differently than values of the second similarity metric of the one or more second media files to rank the search results.

11 . The device of claim 1 , wherein a particular media file of the set of media files is further associated with one or more tag embeddings representing one or more sound tags associated with the particular media file, and wherein the set of embeddings associated with the set of media files include the one or more tag embeddings.

12 . The device of claim 11 , wherein the one or more processors are further configured to:

generate one or more query tag embeddings based on the query; and

compare the one or more query tag embeddings, in a tag embedding space, to one or more tag embeddings to determine a third similarity metric indicative of similarity between the one or more tag embeddings and the one or more query tag embeddings, wherein the search results further identify one or more third media files of the set of filtered media files, each of the one or more third media files associated with at least one of the one or more tag embeddings.

13 . The device of claim 12 , wherein the one or more processors are further configured to rank the search results based on similarity values and wherein values of the first similarity metric associated with the one or more first media files are weighted differently than values of the third similarity metric of the one or more third media files to rank the search results.

14 . The device of claim 1 , wherein the one or more processors are further configured to:

obtain an additional media file for storage at a file repository;

process the additional media file to detect one or more sounds represented in the additional media file;

generate one or more embeddings associated with the one or more sounds detected in the additional media file;

store the additional media file and the one or more embeddings in the file repository; and

in response to receipt of a subsequent query, search the one or more embeddings associated with the additional media file.

15 . The device of claim 1 , wherein the one or more processors are further configured to determine the first similarity metric based on a distance, in an embedding space, between the one or more caption embeddings and the one or more query caption embeddings.

16 . A method comprising:

generating, by one or more processors, one or more query caption embeddings representing multiple words that together form a semantic unit, based on a query, using a machine-learning model, wherein the query caption embeddings are based on sound captions extracted from the query that are natural-language descriptions of sounds and wherein outputs of the machine-learning model are one or more caption embeddings and locations of the one or more caption embeddings in a caption embedding space are indicative of semantic relationships among the sound captions, the semantic relationships corresponding to a similarity of semantic meaning between the sound captions, the similarity of semantic meaning distinct from a text comparison;

generating one or more context terms based on a portion of the query distinct from the semantic unit;

filtering a set of media files to generate a set of filtered media files based on the context terms and metadata associated with the set of media files;

comparing, by the one or more processors, the one or more query caption embeddings, in the caption embedding space, to one or more second caption embeddings associated with the set of filtered media files, wherein each second caption embedding of the one or more second caption embeddings represents a corresponding sound caption and each sound caption includes a natural-language text description of a sound, to determine, a first similarity metric indicative of similarity between the one or more caption embeddings and the one or more query caption embeddings; and

generating, by the one or more processors, search results identifying one or more first media files of the set of filtered media files, each of the one or more first media files associated with at least one of the one or more caption embeddings.

17 . The method of claim 16 , wherein the query includes query audio data or a natural-language sequence of words describing a non-speech sound.

18 . The method of claim 16 , wherein the context terms include a media format type, a particular time period, a media source, a media generation location, or a combination thereof.

19 . The method of claim 16 , wherein the context terms include one or more words identifying a person, and wherein filtering the set of media files includes comparing the context terms to one or more object tags in the metadata, the one or more object tags identifying people visible in a corresponding video.

20 . The method of claim 16 , wherein a particular caption embedding describes a particular sound associated with a particular media file and wherein the particular caption embedding is associated with a time index indicating an approximate playback time of the particular media file at which the particular sound occurs.

21 . The method of claim 20 , wherein the one or more query caption embeddings are based on query audio data of the query and further comprising:

generating one or more query audio embeddings based on the query audio data; and

comparing the one or more query audio embeddings, in an audio embedding space, to one or more audio embeddings to determine a second similarity metric indicative of similarity between the one or more audio embeddings and the one or more query audio embeddings, wherein the search results further identify one or more second media files of the set of filtered media files, each of the one or more second media files associated with at least one of the one or more audio embeddings.

22 . The method of claim 16 , further comprising:

generating one or more query tag embeddings based on the query; and

comparing the one or more query tag embeddings, in a tag embedding space, to one or more tag embeddings to determine a third similarity metric indicative of similarity between the one or more tag embeddings and the one or more query tag embeddings, wherein the search results further identify one or more third media files of the set of filtered media files, each of the one or more third media files associated with at least one of the one or more tag embeddings.

23 . The method of claim 16 , further comprising:

obtaining an additional media file for storage at a file repository;

processing the additional media file to detect one or more sounds represented in the additional media file;

generating one or more embeddings associated with the one or more sounds detected in the additional media file;

storing the additional media file and the one or more embeddings in the file repository; and

in response to receipt of a subsequent query, searching the one or more embeddings associated with the additional media file.

24 . A non-transitory computer-readable storage device storing instructions that are executable by one or more processors to cause the one or more processors to:

generate one or more query caption embeddings representing multiple words that together form a semantic unit, based on a query, using a machine-learning model, wherein the query caption embeddings are based on sound captions extracted from the query that are natural-language descriptions of sounds and wherein outputs of the machine-learning model are one or more caption embeddings and locations of the one or more caption embeddings in a caption embedding space are indicative of semantic relationships among the sound captions, the semantic relationships corresponding to a similarity of semantic meaning between the sound captions, the similarity of semantic meaning distinct from a text comparison;

generate one or more context terms based on a portion of the query distinct from the semantic unit;

filter a set of media files to generate a set of filtered media files based on the context terms and metadata associated with the set of media files;

compare the one or more query caption embeddings, in the caption embedding space, to one or more second caption embeddings associated with the set of filtered media files, wherein each second caption embedding of the one or more second caption embeddings represents a corresponding sound caption and each sound caption includes a natural-language text description of a sound, to determine a first similarity metric, indicative of similarity between the one or more caption embeddings and the one or more query caption embeddings; and

generate search results identifying one or more first media files of the set of filtered media files, each of the one or more first media files associated with at least one of the one or more caption embeddings.

25 . The non-transitory computer-readable storage device of claim 24 , wherein the query includes a first set of words describing a target sound and a second set of words describing a context, and wherein the instructions are further executable to cause one or more processors to determine the one or more query caption embeddings based on the first set of words.

26 . The non-transitory computer-readable storage device of claim 24 , wherein the instructions are further executable to cause one or more processors to:

obtain an additional media file for storage at a file repository;

process the additional media file to detect one or more sounds represented in the additional media file;

generate one or more embeddings associated with the one or more sounds detected in the additional media file;

store the additional media file and the one or more embeddings in the file repository; and

in response to receipt of a subsequent query, search the one or more embeddings associated with the additional media file.

27 . The non-transitory computer-readable storage device of claim 26 , wherein generating the one or more embeddings includes generating a caption embedding associated with the one or more sounds.

28 . The non-transitory computer-readable storage device of claim 24 , wherein the instructions are further executable to cause one or more processors to determine the first similarity metric based on a distance, in an embedding space, between the one or more caption embeddings and the one or more query caption embeddings.

29 . An apparatus comprising:

means for generating, by one or more processors, one or more query caption embeddings representing multiple words that together form a semantic unit, based on a query, using a machine-learning model, wherein the query caption embeddings are based on sound captions extracted from the query that are natural-language descriptions of sounds and wherein outputs of the machine-learning model are one or more caption embeddings and locations of the one or more caption embeddings in a caption embedding space are indicative of semantic relationships among the sound captions, the semantic relationships corresponding to a similarity of semantic meaning between the sound captions, the similarity of semantic meaning distinct from a text comparison;

means for generating one or more context terms based on a portion of the query distinct from the semantic unit;

means for filtering a set of media files to generate a set of filtered media files based on the context terms and metadata associated with the set of media files;

means for comparing, by the one or more processors, the one or more query caption embeddings, in the caption embedding space, to one or more second caption embeddings associated with the set of filtered media files, wherein each second caption embedding of the one or more second caption embeddings represents a corresponding sound caption and each sound caption includes a natural-language text description of a sound, to determine, a first similarity metric indicative of similarity between the one or more caption embeddings and the one or more query caption embeddings; and

means for generating, by the one or more processors, search results identifying one or more first media files of the set of filtered media files, each of the one or more first media files associated with at least one of the one or more caption embeddings.

30 . The apparatus of claim 29 , wherein the means for generating the one or more query caption embeddings, the means for comparing the one or more caption embeddings, and the means for generating search results are integrated within a speaker bar, a mobile communication device, a smart phone, a cellular phone, a laptop computer, a computer, a tablet, a personal digital assistant, a display device, a television, a gaming console, a music player, a radio, a digital video player, a digital video disc (DVD) player, a tuner, a camera, a navigation device, a vehicle, a headset, an augmented reality headset, a mixed reality headset, a virtual reality headset, an aerial vehicle, a home automation system, a voice-activated device, a wireless speaker and voice activated device, a portable electronic device, a car, an internet-of-things (IoT) device, a virtual reality (VR) device, a base station, or a combination thereof.