Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
Applicant filed an amendment on 5/15/2026. Claims 1-20 were pending. Claims 3, 8, 13, 17 have been canceled. Thus 1-2,5-7,9-12,14-16,18-20 are pending. After careful consideration of applicant arguments and amendments the examiner finds them to be moot and/or non-persuasive. This action is a Final Rejection.
Claim Rejections - 35 USC § 101
Here it is noted that the applicant specification includes fig. 2 a way to make the process faster by not going through the pre-processing module. Thus in view of current guidance, the claims may be eligible. However, they are subject to 35 USC 103 rejection
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1,5,6,9-11,14-15,18-19 is/are rejected under 35 U.S.C. 103 as being unpatentable over US Patent Publication to 20180276540 Xing in view of US Patent 11,854558 to Lu
As per claim 11, Xing discloses;
storing, in a memory module, a plurality of stem data
xing(0071)
extracted from one or more audio data,
outputting, by a processor via a pre-processing module, one or more input stem data from an input audio data; generating, by the processor, one or more input embedding vectors from the one or more input stem data using an artificial neural network module;
Xing(0042-44)
searching, by the processor, for one or more similar embedding vectors, among the plurality of embedding vectors stored in the memory module, based on a similarity calculated between the one or more input embedding vectors and the plurality of embedding vectors; and
Xing(0042-44)
searching, by the processor, for one or more similar audio embedding vectors, among the plurality of audio embedding vectors stored in the memory module, based on a similarity between the audio embedding vector and the plurality of audio embedding vectors; and generating, by the processor, based on the one or more similar embedding vectors and the one or more similar audio embedding vectors, Xing(0060-62, embedding vectors)
a recommendation list for at least one of:(i) one or more similar stem data corresponding to the one or more similar embedding vectors, and(ii) one or more similar audio data corresponding to the one or more similar audio embedding vectors.
Xing(0043 vectors are used…. To compare music data, make recommendations)
Xing does not explicitly disclose what Lu teaches;
wherein the plurality of stem data is separated from the one or more audio data by a type, and the type of the plurality of stem data includes at least one of vocal, drum, bass, piano, and accompaniment, Lu (col. 4 lines 10-15)
and a plurality of embedding vectors generated from the plurality of stem data, and a plurality of audio embedding vector; Lu (col 3 lines 15-20)
outputting, by the processor, an audio embedding vector corresponding to the input audio data by inputting the input audio data to the artificial neural network module without passing through the pre-processing module; Lu(col. 1 line 60-col. 2 line 5, by pass feature)
The motivation for the combination, combining music analysis of Xing with the bypassing and embedding vectors of Lu for the motivation of “transforming audio data recognition” (col. 1 lines 25-31)
Claim 1 is similar to claim 11.
As per 14, Xing discloses;
The method according to claim 13,
wherein the artificial neural network module comprises a plurality of pre-learned artificial neural networks, and
wherein each of the plurality of pre-learned artificial neural networks receives a different type of the plurality of stem data as input and outputs a corresponding embedding vector.
Xing(0043, storage)
Claim 5 is similar to claim 14
As per claim 15, Xing discloses;
The method according to claim 14,
wherein each of the plurality of pre-learned artificial neural networks comprises a convolutional neural network (CNN) based encoder structure. Xing(0028)
Claim 6 is similar to claim 15
As per claim 18, Xing discloses;
The method according to claim 11, further comprising:
storing, in the memory module, a plurality of tagging information corresponding to the plurality of stem data;
outputting, by the processor, one or more input tagging information corresponding to the one or more input stem data;
searching, by the processor, for one or more similar tagging information, among the plurality of tagging information stored in the memory module, based on a similarity between the one or more input tagging information and the plurality of tagging information stored in the memory module; and
generating, by the processor, the recommendation list based further on at least one of:
(i) one or more similar stem data corresponding to the one or more similar tagging information, and
(ii) one or more similar audio data corresponding to the one or more similar tagging information.
Xing(0028, tagging data….., establish song to song similarity, 0062 )
Claim 9 is similar to claim 18
As per claim 10, Xing discloses;
The music analysis device according to claim 9, wherein the tagging information includes at least one of genre information, mood information, instrument information, and music creation time information.
Xing(0028, “one of” requires only one)
Claim 19 is similar to claim 10
Claim(s) 2,12, 7 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over US Patent Publication to 20180276540 Xing in view of US Patent 11854558 to Lu and further in view of US Patent to Wang 12106740
As per claim 12, Xing and Lu do not explicitly disclose what Wang teaches;
the method of claim 11,
wherein determining the similarity includes determining the similarity between the one or more input embedding vectors and the plurality of embedding vectors using a Euclidean distance.
Wang(col. 4 lines 25-60, Euclidian distance and vectors)
It would therefore have been obvious to one of ordinary skill before the effective filing date of the invention to combine the music analysis disclosure of Xing with the Euclidian distance teachings of Wang for the motivation of better modeling music. Col. 1 lines 20-25
Claim 2 is similar to claim 12
As per claim 16, Here Xing and Lu do not explicitly disclose what Wang teaches;
The method according to claim 11,
wherein the artificial neural network module further comprises a dense layer in which the one or more input embedding vectors and the plurality of embedding vectors are shared, and
wherein calculating the similarity includes calculating the similarity between the one or more input embedding vectors and the plurality of embedding vectors based on the dense layer.
Wang (col. 7 lines 35-40- dense layer)
It would therefore have been obvious to one of ordinary skill before the effective filing date of the invention to combine the music analysis disclosure of Xing with the dense layer teachings of Wang for the motivation of better modeling music. Col. 1 lines 20-25
Claim 7 is similar to claim 16
Response to Arguments
Applicant filed an amendment on 5/15/2026. Claims 1-20 were pending. Claims 3, 8, 13, 17 have been canceled. Thus 1-2,5-7,9-12,14-16,18-20 are pending. After careful consideration of applicant arguments and amendments the examiner finds them to be moot and/or non persuasive. This action is a Final Rejection.
Rejection under 35 U.S.C. USC 101 – moot in view of amendment.
Rejection under 35 U.S.C. § 103
Claim(s) 1,3, 5,6,8-11,13-15,17-19 stand rejected under 35 U.S.C. § 103 as being
unpatentable over U.S. Patent Publication No. 2018/0276540 to Xing ("Xing"), optionally in view of U.S. Patent No. 12,106,740 to Wang ("Wang"). Applicant respectfully traverses this rejection.
None of the cited references, alone or in combination, teaches or suggests the claimed features, including "inputting the input audio data to the artificial neural network module without passing through the pre-processing module" and generating a recommendation list based on "(i) the one or more similar embedding vectors" and "(ii) the one or more similar audio embedding vectors."
A. Xing Does Not Teach Stem Data Separated by a Type
the Examiner asserts that Xing's "snippet" is equivalent to the claimed stem data, stating that "a snippetis like a stem, ie it's a short portion that can be tagged by features" (citing Xing [0028]). To reject the limitation of stem types (e.g., vocal,12
drum), the Examiner further relies on Xing's mention of "music type" (Office Action, page 6, citing Xing[0029]). Applicant respectfully disagrees with both assertions.
First, Xing defines a "snippet" merely as a short time-based segment of an audio file. As explicitly defined in Xing, a snippet is "a three second sound clip taken from the middle of the audio file" (Xing, [0035]).
Taking a short, temporal portion of a mixed audio file is fundamentally different from the claimed "stem data" that is "separated from the one or more audio data by a type, " wherein the type includes "vocal, drum, bass, piano, and accompaniment."
Second, the Examiner's reliance on Xing's "music type" is misplaced. Xing describes
identifying "explicit features associated with music, such as genre, type, rhythm, pitch, loudness, timbre, etc. " (Xing, [0031]). In the context of Xing, "type" simply refers to a descriptive tag assigned to the entirety of the audio snippet. For instance, Xing teaches classifying the whole song into categories such as nine genres (e.g., alternative, blues, electronic, folk/country, funk/soul/rnb, jazz, pop, rap/hip-hop and rock) (Xing, [0069]). Xing does not teach separating an audio signal into individual stem data (e.g., vocal, drum, bass) by a type. B. Xing Does Not Teach Inputting Audio Data Without Passing Through the Pre- Processing Module
Claims 1 and 11 require "inputting the input audio data to the artificial neural network module without passing through the pre-processing module." A recommendation list is then generated based on both "the one or more similar embedding vectors" (from the stem data) and "the one or more similar audio embedding vectors" (from the input audio data).
The Examiner alleges that Xing discloses the claimed steps (Office Action, pages 5 and 7, citing Xing- and). However, Xing teaches combining "all the acoustic signal13
representations... into a normalized hyper-image" (Xing, [0052]; FIG. 5A) to feed into a neural network. Because Xing forces all acoustic features to be combined into a single "hyper-image", Xing does not teach the claimed structure where the input audio data is inputted "without passing through the pre-processing module" while the stem data is processed via the pre- processing module, and then generating a recommendation list based on both sets of vectors. C. Wang Does Not Cure the Deficiencies of Xing
The Examiner cites Wang solely for its teachings of using a Euclidean distance and a dense layer (Office Action, pages 8-9). Wang does not disclose "stem data... separated... by a type," nor does it disclose "inputting the input audio data to the artificial neural network module without passing through the pre-processing module." Therefore, the combination of Xing and Wang fails to render the independent claims obvious.
Lu is offered to teach the clarified Stem data argued by applicant. Thus applicant arguments are moot in view of new grounds of rejction.
The rejection under 35 USC 103 is moot in view of updated grounds of rejection.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Automatic Music Labeling Algorithm based on Tag Depth Analysis, IEEE 2023
Multi-Modal Song Mood Detection with Deep Learning, IEEE 2022
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to BRUCE I EBERSMAN whose telephone number is (571)270-3442. The examiner can normally be reached 8:00 am - 5:00 pm Monday-Friday.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael W Anderson can be reached at 571-270-0508. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/BRUCE I EBERSMAN/Primary Examiner, Art Unit 3693