DETAILED ACTION
This action is in response to the communications filed 3/21/2025. Claims 1-20 are pending and have been examined.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 2/3/2026 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-20 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Bosch Vicente et al, US Publication 2023/0075074 A1..
Regarding Claim 1, Bosch Vicente et al teaches, a computer-implemented method (Title/Abstract), comprising: receiving, via a graphical user interface (GUI) presented on a user computing device (Paragraph 182, “The input control device(s) 1180 provide a portion of the user interface for a user of the computer 1100. The input control device(s) 1180 may include a keypad and/or a cursor control device. The keypad may be configured for inputting alphanumeric characters and/or other key information. The cursor control device may include, for example, a handheld controller or mouse, a trackball, a stylus, and/or cursor direction keys. In order to display textual and graphical information, the system 1100 may include the graphics subsystem 1160 and the output display 1170. The output display 1170 may include a display such as a CSTN (Color Super Twisted Nematic), TFT (Thin Film Transistor), TFD (Thin Film Diode), OLED (Organic Light-Emitting Diode), AMOLED display (Activematrix Organic Light-emitting Diode), and/or liquid crystal display (LCD)-type displays. The displays can also be touchscreen displays, such as capacitive and resistive-type touchscreen displays. The graphics subsystem 1160 receives textual and graphical information, and processes the information for output to the output display 1170.”), a selection of an audio snippet (Paragraph 9, “One aspect includes a method for combining audio tracks, comprising: determining at least one music track that is musically compatible with a base music track;”), the selection indicating an identifier of an audio file (Paragraph 72, “An additional type of information that can be employed to create mashups can include segment labelling information. Segment labelling information identifies (using, e.g., particular IDs) different types of track segments, and track segments may be labeled according to their similarity.”),a start time, and an end time (Paragraph 83, “In one example embodiment herein, the procedure 200 employs at least some of the various types of information 1131 as described above, including, without limitation, information about the likelihood of a segment containing vocals (vx) (e.g., at beats of segments), downbeat positions, song segmentation information (including start and end positions of segments), and segment labelling information (e.g., IDs), and the like.”), wherein the audio file is from among a plurality of audio files in a mashup catalog (Paragraph 77, “Each of the above types of information associated with particular tracks and/or with particular segments of tracks, can be stored in a database in association with the corresponding tracks and/or segments. The database may be, by example and without limitation, one or more of main memory 1125, portable storage medium 1150, and mass storage device 1130 of the system 1100 of FIG. 20 to be described below, or the database can be external to that system 1100, in which case it can be accessed by the system 1100 by way of, for example, network 1120 and peripheral device(s) 1140.”);
accessing a beat marking associated with the audio file, the beat marking indicating metrical information associated with the audio file, the metrical information including for each of a plurality of beats of the audio file, a beat number, a bar number, and a section number (Paragraph 71, “Example aspects of the present application also can employ song or track segmentation information for creating mashups. For example, song segmentation information can include the temporal positions of boundaries between sections of each track.”, Paragraph 74, “Additional information that can be employed to create mashups also can include tempo(s) of each track, a representation of tonality of each track (e.g., a twelve-dimensional chroma vector), beat/downbeat positions in each track (e.g., temporal positions of beats and downbeats in each track), information about the presence of vocals (if any) in time in each track, energy of each of the segments in the vocal and accompaniment tracks, or the like.”, Paragraph 83, “In one example embodiment herein, the procedure 200 employs at least some of the various types of information 1131 as described above, including, without limitation, information about the likelihood of a segment containing vocals (vx) (e.g., at beats of segments), downbeat positions, song segmentation information (including start and end positions of segments), and segment labelling information (e.g., IDs), and the like.”, Paragraph 84, “Referring to FIG. 2a, query segments 122 of the query track 112 that have less than a predetermined number of bars (e.g., eight bars) are filtered out and discarded (step 202), while others are maintained.”);
accessing a chord string associated with the audio file, the chord string indicating harmonic information associated with the audio file (Paragraph 122, “Referring again to FIG. 9a, after tempo compatibility (e.g., score K_seg_tempo) is determined in step 904, harmonic progression compatibility (also referred to herein as “harmonic compatibility”) is determined in step 906. In one example embodiment herein, the closer the harmonic compatibility of segments 110, 112 under consideration, the higher is the score.”), the harmonic information including a chord type for each of the plurality of beats of the audio file (FIG. 6, Step 602, Determine Musical Keys, as well as Steps 604, 606, 608.);
identifying a metrical signature and a chord string of the audio snippet, the metrical signature including a beat number and a bar number associated with a beat of the audio file corresponding to the start time, and the chord string including the chord type for each beat of the audio snippet (Paragraph 74, “Additional information that can be employed to create mashups also can include tempo(s) of each track, a representation of tonality of each track (e.g., a twelve-dimensional chroma vector), beat/downbeat positions in each track (e.g., temporal positions of beats and downbeats in each track), information about the presence of vocals (if any) in time in each track, energy of each of the segments in the vocal and accompaniment tracks, or the like.”, Paragraph 122, “Also, in one example embodiment herein, step 906 can be performed according to procedure 1100′ shown in FIG. 11. In step 1102′ beat synchronized chroma feature vectors are determined for each of the query segment 112 and candidate segment 110 under consideration, by determining, for each respective segment 110, 112, an average of chroma values within each beat of the respective segment 110, 112. In one example embodiment herein, the chroma values are obtained from among the information 1131 in the database using methods described in the Jehan reference discussed above. In step 1104′ a Pearson correlation between the beat synchronized chroma feature vectors determined in step 1102′, is determined for each of the beats of the segments under consideration.”);
identifying, from among the plurality of audio files, a plurality of mashup candidate audio snippets that match the metrical signature of the audio snippet and that have a beat length that matches a beat length of the audio snippet (Paragraph 80, “As represented in FIG. 1, the candidate track includes vocals 114 and the query track 112 includes separated vocal component/track 116 and separated instrumental component/track 118. In addition, additional track features 112a of the query track and additional track features 110a of the candidate track 110 are also identified from the query track 112 and candidate track 110. Track features 110a and 112a can include, for example, acoustic features (such as tempo, beat, musical key, likelihood of including vocals, and other features as described herein). Information regarding loudness 114b and tonality (e.g., tonal representation) 114a are obtained based on the vocal component 114 of the candidate track 110. Information regarding loudness 118b and tonality (e.g., tonal representation) 118a based on the separated instrumental component/track 118 and information regarding at least loudness 116a based on the separated vocal component/track 116 of the query track 112 are obtained.”);
comparing the chord string of the audio snippet with respective chord strings of each of the plurality of mashup candidate audio snippets to identify a subset of the plurality of mashup candidate audio snippets that harmonically match the audio snippet (Paragraph 122, “Referring again to FIG. 9a, after tempo compatibility (e.g., score K_seg_tempo) is determined in step 904, harmonic progression compatibility (also referred to herein as “harmonic compatibility”) is determined in step 906. In one example embodiment herein, the closer the harmonic compatibility of segments 110, 112 under consideration, the higher is the score. Also, in one example embodiment herein, step 906 can be performed according to procedure 1100′ shown in FIG. 11. In step 1102′ beat synchronized chroma feature vectors are determined for each of the query segment 112 and candidate segment 110 under consideration, by determining, for each respective segment 110, 112, an average of chroma values within each beat of the respective segment 110, 112. In one example embodiment herein, the chroma values are obtained from among the information 1131 in the database using methods described in the Jehan reference discussed above.”);
receiving, via the GUI presented on the user computing device, a selection of one of the subset of the plurality of mashup candidate audio snippets (Paragraph 144, “After computing the score (M) for all segments 124 of all candidate tracks 110 under consideration, the segment 124 with the highest total mashability score (M) is selected (step 1812), although in other example embodiments, a sampling between all possible candidate segments can be done with a probability which is proportional to their total mashability score. The above procedure can be performed with respect to all segments 122 that were assigned to S-subs and S_add of the query track 112 under consideration, starting from the start of the track 112 and finishing at the end of the track 112, to determine mashability between those segments 122 and individual ones of the candidate segments 124 of candidate tracks 110 that were selected as being compatible with the query track 112.”, Paragraph 95, “As a result of the procedure 300, a mashup track 120 (FIG. 1) is provided based on the query track 112 and at least one candidate track 110 under consideration.”);
and generating a mashup audio snippet based on the audio snippet and the selected one of the subset of the plurality of mashup candidate audio snippets, the mashup audio snippet including at least one stem from the audio snippet and at least one stem from the selected one of the subset of the plurality of mashup candidate audio snippets (Paragraph 95, “As a result of the procedure 300, a mashup track 120 (FIG. 1) is provided based on the query track 112 and at least one candidate track 110 under consideration. The mashup track 120 includes, by example, one or more segments 122 that were assigned to S_keep, one or more other segments 122 having vocal content (from one or more candidate tracks 110) that was used to replace vocal content of an original version of those other segments 122 in step 308, and one or more further segments 122 having vocal content (from one or more candidate tracks 110) that was added to those further segments 122 in step 308. In the mashup track 120, beat positions in the query track 112 are mapped with corresponding beat positions of the candidate track(s) 110.”, Paragraph 172, “As can be appreciated in view of the above description, at least some example aspects herein employ source separation to generate candidate (e.g., vocal) tracks and query (e.g., accompaniment) tracks, although in other example embodiments, stems can be used instead, or a multitrack can be employed where separation is therefore not needed). In other example embodiments herein, full tracks can be employed (without separation of vocals and accompaniment components).”).
Regarding Claim 2, Bosch Vicente et al teaches all the limitations of claim 1, and further teaches, wherein the mashup catalog includes, for each of the plurality of audio files: (i) one or more stems separated from the audio file (Paragraph 172, “As can be appreciated in view of the above description, at least some example aspects herein employ source separation to generate candidate (e.g., vocal) tracks and query (e.g., accompaniment) tracks, although in other example embodiments, stems can be used instead, or a multitrack can be employed where separation is therefore not needed).”); and (ii) metadata indicating a tempo and a key of the audio file, and annotations for chord type, beat/downbeat, and song structure (Paragraph 80, “As represented in FIG. 1, the candidate track includes vocals 114 and the query track 112 includes separated vocal component/track 116 and separated instrumental component/track 118. In addition, additional track features 112a of the query track and additional track features 110a of the candidate track 110 are also identified from the query track 112 and candidate track 110. Track features 110a and 112a can include, for example, acoustic features (such as tempo, beat, musical key, likelihood of including vocals, and other features as described herein). Information regarding loudness 114b and tonality (e.g., tonal representation) 114a are obtained based on the vocal component 114 of the candidate track 110. Information regarding loudness 118b and tonality (e.g., tonal representation) 118a based on the separated instrumental component/track 118 and information regarding at least loudness 116a based on the separated vocal component/track 116 of the query track 112 are obtained.”).
Regarding Claim 3, Bosch Vicente et al teaches all the limitations of claim 2, and further teaches, receiving, via the GUI presented on the user computing device, a selection representing a number of stems of the selected one of the subset of the plurality of mashup candidate audio snippets to be included in the generated mashup audio snippet, wherein the mashup audio snippet is generated based on the received selection (Paragraph 78, “FIG. 1 shows an example flowchart representation of how an automashup can be performed based on a candidate track that includes vocal content, and a background or query track, according to an example embodiment herein. In this example, the algorithm to perform the automashup creates a music mashup by sequentially adding vocal segments of one or more track(s) (of one song) on top of one or more segments of a background track, (of, e.g., another song), and/or by replacing vocal content of one or more segments of a background track (of one song) that includes the vocal content, with vocal content of the one or more track(s) (of, e.g., another song). Inputs to the algorithm can include, by example, a background track (e.g., including instrumental or vocal/instrumental content) (also referred to herein as a “query track” or “base track”), such as track 112 of FIG. 1, and a (potentially large) set of vocal candidate tracks, including track 110 having vocal content, each of which may be obtained from the database and/or in accordance with the method(s) described in the Jansson application, for example.”, Paragraph 188, “In still another example embodiment herein, when the play button 1402 is selected to play back a song, the instrumental track of the song, as well as the vocal track of the same song (wherein the tracks are recognized to be a pair) are retrieved from the mass storage device 1130.”).
Regarding Claim 4, Bosch Vicente et al teaches all the limitations of claim 3, and further teaches, wherein the plurality of stems include vocals, drums, bass, guitars, synths/keys, and effects (Paragraph 69, “Before describing the foregoing procedures in more detail, examples of at least some types of information that can be used in the procedures will first be described. Example aspects of the present application can employ various different types of information. For example, the example aspects can employ various types of audio signals or tracks, such as mixed original signals, i.e., signals that include both an accompaniment (e.g., background instrumental) component and a vocal component, wherein the accompaniment component includes instrumental content such as one or more types of musical instrument content (although it may include vocal content as well), and the vocal component includes vocal content. Each of the tracks may be in the form of, by example and without limitation, audio files for each of the tracks (e.g. mp3, way, or the like). Other types of tracks that can be employed include solely instrumental tracks (e.g., tracks that include only instrumental content, or only an instrumental component of a mixed original signal), and vocal tracks (e.g., tracks that include only vocal content, or only a vocal component of a mixed original signal).”).
Regarding Claim 5, Bosch Vicente et al teaches all the limitations of claim 2, and further teaches, generating, based on the metadata and for each of the plurality of audio files: (i) the beat marking indicating the metrical information associated with the audio file (Paragraph 73, “Of course, the above examples given for how to obtain vocal and accompaniment tracks, song segmentation information, and segment labelling information, are intended to be representative in nature, and, in other examples, vocal and/or accompaniment tracks, song segmentation information, and/or segment labelling information may be obtained from any applicable source, or in any suitable manner known in the art.”, Paragraph 80, “As represented in FIG. 1, the candidate track includes vocals 114 and the query track 112 includes separated vocal component/track 116 and separated instrumental component/track 118. In addition, additional track features 112a of the query track and additional track features 110a of the candidate track 110 are also identified from the query track 112 and candidate track 110. Track features 110a and 112a can include, for example, acoustic features (such as tempo, beat, musical key, likelihood of including vocals, and other features as described herein).”); (ii) the chord string indicating the harmonic information associated with the audio file (Paragraph 80, “As represented in FIG. 1, the candidate track includes vocals 114 and the query track 112 includes separated vocal component/track 116 and separated instrumental component/track 118. In addition, additional track features 112a of the query track and additional track features 110a of the candidate track 110 are also identified from the query track 112 and candidate track 110. Track features 110a and 112a can include, for example, acoustic features (such as tempo, beat, musical key, likelihood of including vocals, and other features as described herein).”).
Regarding Claim 6, Bosch Vicente et al teaches all the limitations of claim 5, and further teaches, determining, based on the annotations for the song structure in the metadata, for a given beat of a given audio file in the mashup catalog that is associated with a change in the song structure, a ratio between a portion of the given beat before the change to a portion of the given beat after the change (Paragraph 125, “Beat-stability can be another factor involved in vertical mashability. Beat-stability, for a candidate segment 124, is the stability of beat duration in a candidate segment 124 under consideration, wherein, in one example embodiment herein, a greater beat stability results in a higher score. Beat stability is determined in step 912 of FIG. 9. Step 912 is preferably performed according to procedure 1300 of FIG. 13. In step 1302, a relative change between durations of consecutive beats in the candidate segment 124 is determined,”); and assigning the section number to the given beat based on the determined ratio (Above quotation, associates this stability value with different clips being considered.).
Regarding Claim 7, Bosch Vicente et al teaches all the limitations of claim 5, and further teaches, further comprising: determining, based on the annotations for the chord type in the metadata, for a given beat of a given audio file in the mashup catalog that is associated with a change in the chord type, a ratio between a portion of the given beat before the change to a portion of the given beat after the change (Paragraphs 127, 128, “Another factor involved in vertical mashability is harmonic change balance, which measures if there is a balance in a rate of change in time of harmonic content (chroma vectors) of both query and candidate (target) segments 122, 124. Briefly, if musical notes change often in one of the tracks (either query or candidate), the score is higher when the other track is more stable, and vice versa., Harmonic change balance is determined in step 914 of FIG. 9b, which is connected to FIG. 9a via connector B. Details of how harmonic change balance is determined, according to one example embodiment herein, are shown in procedure 1400′ of FIG. 14. In step 1402′ “a length of the segments 122, 124 under consideration is restricted to that of one of the segments 122, 124 with a minimal amount of beats (Nbeats) (i.e., either the query segment 122 or the candidate segment 124). Next, a harmonic change rate between consecutive beats is determined, for each of the query track 112 and candidate track 110 under consideration, as follows. A Pearson correlation between consecutive beat-synchronised chroma vectors is determined, for all beats of each track 110, 112 (step 1404′), to provide a vector (Nbeats−1) of correlation values. In step 1406′, the correlation is mapped to change rate values according to formula (F13): Change=(1−corr)/2(F14).”); and assigning the chord type to the given beat in the chord string based on the determined ratio (Paragraph 129 “As a result, a vector is obtained with (Nbeats−1) change rate values for both candidate and query tracks, 110, 112, wherein the change rate value for the candidate (e.g., vocal) track 110 is represented by “CRvoc”, and the change rate value for the query (accompaniment) track 112 is represented by “CRacc”.”).
Regarding Claim 8, Bosch Vicente et al teaches all the limitations of claim 2, and further teaches, identifying from among the plurality of audio files in the mashup catalog, a subset of audio files that are within a threshold tempo distance from a tempo of the audio file and that satisfy a predetermined key relationship with a major key or a minor key of the audio file, wherein the plurality of mashup candidate audio snippets are identified from the identified subset of audio files (Paragraph 82, “A procedure 200 according to an example aspect herein, for determining whether individual segments of a query track (e.g., an accompaniment track) 112 under consideration are to be kept, or have content (e.g., vocal content) replaced or added thereto from one or more candidate (e.g., vocal) tracks 110, during an automashup of the tracks 110, 112, will now be described, with reference to FIGS. 2a and 2b.”, FIG. 6, displays consideration of major/minor keys. FIG. 7, Step 706, Determine Closeness in tempo, Step 708, Determine Closeness in Key, Steps 710, 712, 714, and 716, determines a mashability score, and discards or selects segments based upon these factors.).
Regarding Claim 9, Bosch Vicente et al teaches all the limitations of claim 8, and further teaches, wherein identifying the subset of audio files that satisfy the predetermined key relationship comprises: determining, based on the metadata, whether the audio file is in the major key or in the minor key; in response to determining that the audio file is in the major key, ignoring audio files in the mashup catalog that are in the minor key except for audio files that are in a relative minor key to a key of the audio file; and in response to determining that the audio file is in the minor key, ignoring audio files in the mashup catalog that are in the major key except for audio files that are in a relative major key to the key of the audio file (Paragraph 110, “FIG. 6 shows a procedure 600 for determining closeness in key, according to an example embodiment herein. In step 602, a determination of the key and of each track 110, 112 (and the pitch at each beat of segments of the tracks 110, 112) under consideration is made. The key and the pitch of a segment is determined using methods described in the Jehan reference discussed above. According to an example embodiment herein, if the tracks 110, 112 under consideration are determined to be in the same type of key (e.g., both are in a major key, or both are in a minor key) (“Yes” in step 604), then the keys determined in step 602 are passed to step 608 to calculate the score Ksong(key), in a manner as will be described below.”, utilizes Key similarity to determine the mashability score, which is then utilized to determine if a segment/file is to be discarded.).
Regarding Claim 10, Bosch Vicente et al teaches all the limitations of claim 1, and further teaches, wherein identifying the subset of the plurality of mashup candidate audio snippets comprises: determining for each of the plurality of mashup candidate audio snippets, whether chord types of at least half of the beats in the chord string of the mashup candidate audio snippet match or are related to chord types of respective beats at same positions in the chord string of the audio snippet (Paragraph 122, “Referring again to FIG. 9a, after tempo compatibility (e.g., score K_seg_tempo) is determined in step 904, harmonic progression compatibility (also referred to herein as “harmonic compatibility”) is determined in step 906. In one example embodiment herein, the closer the harmonic compatibility of segments 110, 112 under consideration, the higher is the score. Also, in one example embodiment herein, step 906 can be performed according to procedure 1100′ shown in FIG. 11. In step 1102′ beat synchronized chroma feature vectors are determined for each of the query segment 112 and candidate segment 110 under consideration, by determining, for each respective segment 110, 112, an average of chroma values within each beat of the respective segment 110, 112. In one example embodiment herein, the chroma values are obtained from among the information 1131 in the database using methods described in the Jehan reference discussed above. In step 1104′ a Pearson correlation between the beat synchronized chroma feature vectors determined in step 1102′, is determined for each of the beats of the segments under consideration. For example, the segments may include a segment of the query track (chroma values taken only from the accompaniment), and one segment of the candidate track underanalysis (only computing chroma values of the vocal part). In step 1106′ a median value (med_corr) of vectors of beat-wise correlations determined in step 1104′ is calculated. Then, in step 1108′ a harmonic (progression) compatibility score (K_seg_harm_prog) is determined using formula (F10) below, according to an example embodiment herein: K_seg_harm_prog=(1+med_corr)/2(F10), wherein K_seg_harm_prog represents the harmonic compatibility score, and med_corr represents the median value determined in step 1106′.”).
Regarding Claim 11, Bosch Vicente et al teaches all the limitations of claim 1, and further teaches, determining, based on the comparing of the chord string of the audio snippet with the respective chord strings of each of the plurality of mashup candidate audio snippets, that none of the plurality of mashup candidate audio snippets harmonically match the audio snippet; and performing a first pitch shift for each of the plurality of mashup candidate audio snippets to identify a subset of the plurality of mashup candidate audio snippets after the first pitch shift that harmonically match the audio snippet (Paragraph 93, “Then, in step 308, those segments 122, 124 are mixed. In one example, mixing includes a procedure involving time-stretching and pitch shifting using, for example, pysox or a library such as elastique. By example, in a case where that segment 122 was previously assigned to S_subs, mixing can include replacing vocal content of that segment 122, with vocal content of the aligned segment 124. Also by example, in a case where the segment 122 was previously assigned to S_add, mixing can include adding vocal content of the segment 124 to the segment 122.”, Paragraph 157, “Step 2208 includes performing pitch shifting to each candidate (e.g., vocal) segment 124, as needed, based on a pitch-shifting ratio. In some embodiments, the pitch-shifting ratio is computed while computing the mashability scores discussed above. For example, the vocals are pitch-shifted by n_semitones, where n_semitones is the number of semitones. In some embodiments, the number of semitones is determined during example step 608 discussed in reference to FIG. 6.”).
Regarding Claim 12, Bosch Vicente et al teaches all the limitations of claim 11, and further teaches, determining, based on a comparison of the chord string of the audio snippet with respective chord strings of each of the plurality of mashup candidate audio snippets after the first pitch shift, that none of the plurality of mashup candidate audio snippets after the first pitch shift harmonically match the audio snippet; and performing a second pitch shift for each of the plurality of mashup candidate audio snippets to identify a subset of the plurality of mashup candidate audio snippets after the second pitch shift that harmonically match the audio snippet, wherein the second pitch shift is by a greater number of semitones than the first pitch shift (Paragraph 157, “Step 2208 includes performing pitch shifting to each candidate (e.g., vocal) segment 124, as needed, based on a pitch-shifting ratio. In some embodiments, the pitch-shifting ratio is computed while computing the mashability scores discussed above. For example, the vocals are pitch-shifted by n_semitones, where n_semitones is the number of semitones. In some embodiments, the number of semitones is determined during example step 608 discussed in reference to FIG. 6.” FIG. 22, applies to each candidate audio snippets/clips. Paragraph 158, “Then, the procedure 2200 can include applying fade-in and fade-out, and/or high pass filtering or equalizations around transition points, using determined transitions (step 2210). In one example embodiment herein, the parts of each segment 124 (of a candidate track 110 under consideration) which are located temporally before initial and after the final points of the refined boundaries (i.e., transitions), can be rendered with a volume fade in, and a fade out, respectively, so as to perform a smooth intro and outro, and reduce clashes between vocals of different tracks. Fade in and Fade out can be performed in a manner known in the art.”).
Regarding Claim 13, Bosch Vicente et al teaches, a non-transitory computer-readable storage medium storing executable instructions (FIG. 20, Main Memory 1125) that, when executed by a hardware processor of a mashup platform (Title/Abstract, FIG. 20, Processor Device 1110), cause the hardware processor to perform steps comprising:
receiving, via a graphical user interface (GUI) presented on a user computing device (Paragraph 182, “The input control device(s) 1180 provide a portion of the user interface for a user of the computer 1100. The input control device(s) 1180 may include a keypad and/or a cursor control device. The keypad may be configured for inputting alphanumeric characters and/or other key information. The cursor control device may include, for example, a handheld controller or mouse, a trackball, a stylus, and/or cursor direction keys. In order to display textual and graphical information, the system 1100 may include the graphics subsystem 1160 and the output display 1170. The output display 1170 may include a display such as a CSTN (Color Super Twisted Nematic), TFT (Thin Film Transistor), TFD (Thin Film Diode), OLED (Organic Light-Emitting Diode), AMOLED display (Activematrix Organic Light-emitting Diode), and/or liquid crystal display (LCD)-type displays. The displays can also be touchscreen displays, such as capacitive and resistive-type touchscreen displays. The graphics subsystem 1160 receives textual and graphical information, and processes the information for output to the output display 1170.”), a selection of an audio snippet (Paragraph 9, “One aspect includes a method for combining audio tracks, comprising: determining at least one music track that is musically compatible with a base music track;”), the selection indicating an identifier of an audio file, a start time, and an end time (Paragraph 83, “In one example embodiment herein, the procedure 200 employs at least some of the various types of information 1131 as described above, including, without limitation, information about the likelihood of a segment containing vocals (vx) (e.g., at beats of segments), downbeat positions, song segmentation information (including start and end positions of segments), and segment labelling information (e.g., IDs), and the like.”), wherein the audio file is from among a plurality of audio files in a mashup catalog (Paragraph 77, “Each of the above types of information associated with particular tracks and/or with particular segments of tracks, can be stored in a database in association with the corresponding tracks and/or segments. The database may be, by example and without limitation, one or more of main memory 1125, portable storage medium 1150, and mass storage device 1130 of the system 1100 of FIG. 20 to be described below, or the database can be external to that system 1100, in which case it can be accessed by the system 1100 by way of, for example, network 1120 and peripheral device(s) 1140.”);
accessing a beat marking associated with the audio file, the beat marking indicating metrical information associated with the audio file, the metrical information including for each of a plurality of beats of the audio file, a beat number, a bar number, and a section number (Paragraph 71, “Example aspects of the present application also can employ song or track segmentation information for creating mashups. For example, song segmentation information can include the temporal positions of boundaries between sections of each track.”, Paragraph 74, “Additional information that can be employed to create mashups also can include tempo(s) of each track, a representation of tonality of each track (e.g., a twelve-dimensional chroma vector), beat/downbeat positions in each track (e.g., temporal positions of beats and downbeats in each track), information about the presence of vocals (if any) in time in each track, energy of each of the segments in the vocal and accompaniment tracks, or the like.”, Paragraph 83, “In one example embodiment herein, the procedure 200 employs at least some of the various types of information 1131 as described above, including, without limitation, information about the likelihood of a segment containing vocals (vx) (e.g., at beats of segments), downbeat positions, song segmentation information (including start and end positions of segments), and segment labelling information (e.g., IDs), and the like.”, Paragraph 84, “Referring to FIG. 2a, query segments 122 of the query track 112 that have less than a predetermined number of bars (e.g., eight bars) are filtered out and discarded (step 202), while others are maintained.”);
accessing a chord string associated with the audio file, the chord string indicating harmonic information associated with the audio file (Paragraph 122, “Referring again to FIG. 9a, after tempo compatibility (e.g., score K_seg_tempo) is determined in step 904, harmonic progression compatibility (also referred to herein as “harmonic compatibility”) is determined in step 906. In one example embodiment herein, the closer the harmonic compatibility of segments 110, 112 under consideration, the higher is the score.”), the harmonic information including a chord type for each of the plurality of beats of the audio file (FIG. 6, Step 602, Determine Musical Keys, as well as Steps 604, 606, 608.);
identifying a metrical signature and a chord string of the audio snippet, the metrical signature including a beat number and a bar number associated with a beat of the audio file corresponding to the start time, and the chord string including the chord type for each beat of the audio snippet (Paragraph 74, “Additional information that can be employed to create mashups also can include tempo(s) of each track, a representation of tonality of each track (e.g., a twelve-dimensional chroma vector), beat/downbeat positions in each track (e.g., temporal positions of beats and downbeats in each track), information about the presence of vocals (if any) in time in each track, energy of each of the segments in the vocal and accompaniment tracks, or the like.”, Paragraph 122, “Also, in one example embodiment herein, step 906 can be performed according to procedure 1100′ shown in FIG. 11. In step 1102′ beat synchronized chroma feature vectors are determined for each of the query segment 112 and candidate segment 110 under consideration, by determining, for each respective segment 110, 112, an average of chroma values within each beat of the respective segment 110, 112. In one example embodiment herein, the chroma values are obtained from among the information 1131 in the database using methods described in the Jehan reference discussed above. In step 1104′ a Pearson correlation between the beat synchronized chroma feature vectors determined in step 1102′, is determined for each of the beats of the segments under consideration.”);
identifying, from among the plurality of audio files, a plurality of mashup candidate audio snippets that match the metrical signature of the audio snippet and that have a beat length that matches a beat length of the audio snippet (Paragraph 80, “As represented in FIG. 1, the candidate track includes vocals 114 and the query track 112 includes separated vocal component/track 116 and separated instrumental component/track 118. In addition, additional track features 112a of the query track and additional track features 110a of the candidate track 110 are also identified from the query track 112 and candidate track 110. Track features 110a and 112a can include, for example, acoustic features (such as tempo, beat, musical key, likelihood of including vocals, and other features as described herein). Information regarding loudness 114b and tonality (e.g., tonal representation) 114a are obtained based on the vocal component 114 of the candidate track 110. Information regarding loudness 118b and tonality (e.g., tonal representation) 118a based on the separated instrumental component/track 118 and information regarding at least loudness 116a based on the separated vocal component/track 116 of the query track 112 are obtained.”);
comparing the chord string of the audio snippet with respective chord strings of each of the plurality of mashup candidate audio snippets to identify a subset of the plurality of mashup candidate audio snippets that harmonically match the audio snippet (Paragraph 122, “Referring again to FIG. 9a, after tempo compatibility (e.g., score K_seg_tempo) is determined in step 904, harmonic progression compatibility (also referred to herein as “harmonic compatibility”) is determined in step 906. In one example embodiment herein, the closer the harmonic compatibility of segments 110, 112 under consideration, the higher is the score. Also, in one example embodiment herein, step 906 can be performed according to procedure 1100′ shown in FIG. 11. In step 1102′ beat synchronized chroma feature vectors are determined for each of the query segment 112 and candidate segment 110 under consideration, by determining, for each respective segment 110, 112, an average of chroma values within each beat of the respective segment 110, 112. In one example embodiment herein, the chroma values are obtained from among the information 1131 in the database using methods described in the Jehan reference discussed above.”);
receiving, via the GUI presented on the user computing device, a selection of one of the subset of the plurality of mashup candidate audio snippets (Paragraph 144, “After computing the score (M) for all segments 124 of all candidate tracks 110 under consideration, the segment 124 with the highest total mashability score (M) is selected (step 1812), although in other example embodiments, a sampling between all possible candidate segments can be done with a probability which is proportional to their total mashability score. The above procedure can be performed with respect to all segments 122 that were assigned to S-subs and S_add of the query track 112 under consideration, starting from the start of the track 112 and finishing at the end of the track 112, to determine mashability between those segments 122 and individual ones of the candidate segments 124 of candidate tracks 110 that were selected as being compatible with the query track 112.”, Paragraph 95, “As a result of the procedure 300, a mashup track 120 (FIG. 1) is provided based on the query track 112 and at least one candidate track 110 under consideration.”);
and generating a mashup audio snippet based on the audio snippet and the selected one of the subset of the plurality of mashup candidate audio snippets, the mashup audio snippet including at least one stem from the audio snippet and at least one stem from the selected one of the subset of the plurality of mashup candidate audio snippets (Paragraph 95, “As a result of the procedure 300, a mashup track 120 (FIG. 1) is provided based on the query track 112 and at least one candidate track 110 under consideration. The mashup track 120 includes, by example, one or more segments 122 that were assigned to S_keep, one or more other segments 122 having vocal content (from one or more candidate tracks 110) that was used to replace vocal content of an original version of those other segments 122 in step 308, and one or more further segments 122 having vocal content (from one or more candidate tracks 110) that was added to those further segments 122 in step 308. In the mashup track 120, beat positions in the query track 112 are mapped with corresponding beat positions of the candidate track(s) 110.”, Paragraph 172, “As can be appreciated in view of the above description, at least some example aspects herein employ source separation to generate candidate (e.g., vocal) tracks and query (e.g., accompaniment) tracks, although in other example embodiments, stems can be used instead, or a multitrack can be employed where separation is therefore not needed). In other example embodiments herein, full tracks can be employed (without separation of vocals and accompaniment components).”).
Regarding Claim 14, Bosch Vicente et al teaches all the limitations of claim 13, and further teaches, wherein the mashup catalog includes, for each of the plurality of audio files: (i) one or more stems separated from the audio file (Paragraph 172, “As can be appreciated in view of the above description, at least some example aspects herein employ source separation to generate candidate (e.g., vocal) tracks and query (e.g., accompaniment) tracks, although in other example embodiments, stems can be used instead, or a multitrack can be employed where separation is therefore not needed).”); and (ii) metadata indicating a tempo and a key of the audio file, and annotations for chord type, beat/downbeat, and song structure (Paragraph 80, “As represented in FIG. 1, the candidate track includes vocals 114 and the query track 112 includes separated vocal component/track 116 and separated instrumental component/track 118. In addition, additional track features 112a of the query track and additional track features 110a of the candidate track 110 are also identified from the query track 112 and candidate track 110. Track features 110a and 112a can include, for example, acoustic features (such as tempo, beat, musical key, likelihood of including vocals, and other features as described herein). Information regarding loudness 114b and tonality (e.g., tonal representation) 114a are obtained based on the vocal component 114 of the candidate track 110. Information regarding loudness 118b and tonality (e.g., tonal representation) 118a based on the separated instrumental component/track 118 and information regarding at least loudness 116a based on the separated vocal component/track 116 of the query track 112 are obtained.”).
Regarding Claim 15, Bosch Vicente et al teaches all the limitations of claim 14, and further teaches, wherein the instructions further cause the hardware processor to perform a step comprising: receiving, via the GUI presented on the user computing device, a selection representing a number of stems of the selected one of the subset of the plurality of mashup candidate audio snippets to be included in the generated mashup audio snippet, wherein the mashup audio snippet is generated based on the received selection (Paragraph 78, “FIG. 1 shows an example flowchart representation of how an automashup can be performed based on a candidate track that includes vocal content, and a background or query track, according to an example embodiment herein. In this example, the algorithm to perform the automashup creates a music mashup by sequentially adding vocal segments of one or more track(s) (of one song) on top of one or more segments of a background track, (of, e.g., another song), and/or by replacing vocal content of one or more segments of a background track (of one song) that includes the vocal content, with vocal content of the one or more track(s) (of, e.g., another song). Inputs to the algorithm can include, by example, a background track (e.g., including instrumental or vocal/instrumental content) (also referred to herein as a “query track” or “base track”), such as track 112 of FIG. 1, and a (potentially large) set of vocal candidate tracks, including track 110 having vocal content, each of which may be obtained from the database and/or in accordance with the method(s) described in the Jansson application, for example.”, Paragraph 188, “In still another example embodiment herein, when the play button 1402 is selected to play back a song, the instrumental track of the song, as well as the vocal track of the same song (wherein the tracks are recognized to be a pair) are retrieved from the mass storage device 1130.”).
Regarding Claim 16, Bosch Vicente et al teaches all the limitations of claim 15, and further teaches, wherein the plurality of stems include vocals, drums, bass, guitars, synths/keys, and effects (Paragraph 69, “Before describing the foregoing procedures in more detail, examples of at least some types of information that can be used in the procedures will first be described. Example aspects of the present application can employ various different types of information. For example, the example aspects can employ various types of audio signals or tracks, such as mixed original signals, i.e., signals that include both an accompaniment (e.g., background instrumental) component and a vocal component, wherein the accompaniment component includes instrumental content such as one or more types of musical instrument content (although it may include vocal content as well), and the vocal component includes vocal content. Each of the tracks may be in the form of, by example and without limitation, audio files for each of the tracks (e.g. mp3, way, or the like). Other types of tracks that can be employed include solely instrumental tracks (e.g., tracks that include only instrumental content, or only an instrumental component of a mixed original signal), and vocal tracks (e.g., tracks that include only vocal content, or only a vocal component of a mixed original signal).”).
Regarding Claim 17, Bosch Vicente et al teaches all the limitations of claim 14, and further teaches, wherein the instructions further cause the hardware processor to perform a step comprising: generating, based on the metadata and for each of the plurality of audio files: (i) the beat marking indicating the metrical information associated with the audio file (Paragraph 73, “Of course, the above examples given for how to obtain vocal and accompaniment tracks, song segmentation information, and segment labelling information, are intended to be representative in nature, and, in other examples, vocal and/or accompaniment tracks, song segmentation information, and/or segment labelling information may be obtained from any applicable source, or in any suitable manner known in the art.”, Paragraph 80, “As represented in FIG. 1, the candidate track includes vocals 114 and the query track 112 includes separated vocal component/track 116 and separated instrumental component/track 118. In addition, additional track features 112a of the query track and additional track features 110a of the candidate track 110 are also identified from the query track 112 and candidate track 110. Track features 110a and 112a can include, for example, acoustic features (such as tempo, beat, musical key, likelihood of including vocals, and other features as described herein).”); (ii) the chord string indicating the harmonic information associated with the audio file (Paragraph 80, “As represented in FIG. 1, the candidate track includes vocals 114 and the query track 112 includes separated vocal component/track 116 and separated instrumental component/track 118. In addition, additional track features 112a of the query track and additional track features 110a of the candidate track 110 are also identified from the query track 112 and candidate track 110. Track features 110a and 112a can include, for example, acoustic features (such as tempo, beat, musical key, likelihood of including vocals, and other features as described herein).”).
Regarding Claim 18, Bosch Vicente et al teaches all the limitations of claim 17, and further teaches, wherein the instructions further cause the hardware processor to perform steps comprising: determining, based on the annotations for the song structure in the metadata, for a given beat of a given audio file in the mashup catalog that is associated with a change in the song structure, a ratio between a portion of the given beat before the change to a portion of the given beat after the change (Paragraph 125, “Beat-stability can be another factor involved in vertical mashability. Beat-stability, for a candidate segment 124, is the stability of beat duration in a candidate segment 124 under consideration, wherein, in one example embodiment herein, a greater beat stability results in a higher score. Beat stability is determined in step 912 of FIG. 9. Step 912 is preferably performed according to procedure 1300 of FIG. 13. In step 1302, a relative change between durations of consecutive beats in the candidate segment 124 is determined,”); and assigning the section number to the given beat based on the determined ratio (Above quotation, associates this stability value with different clips being considered.).
Regarding Claim 19, Bosch Vicente et al teaches all the limitations of claim 17, and further teaches, wherein the instructions further cause the hardware processor to perform steps comprising: determining, based on the annotations for the chord type in the metadata, for a given beat of a given audio file in the mashup catalog that is associated with a change in the chord type, a ratio between a portion of the given beat before the change to a portion of the given beat after the change (Paragraphs 127, 128, “Another factor involved in vertical mashability is harmonic change balance, which measures if there is a balance in a rate of change in time of harmonic content (chroma vectors) of both query and candidate (target) segments 122, 124. Briefly, if musical notes change often in one of the tracks (either query or candidate), the score is higher when the other track is more stable, and vice versa., Harmonic change balance is determined in step 914 of FIG. 9b, which is connected to FIG. 9a via connector B. Details of how harmonic change balance is determined, according to one example embodiment herein, are shown in procedure 1400′ of FIG. 14. In step 1402′ “a length of the segments 122, 124 under consideration is restricted to that of one of the segments 122, 124 with a minimal amount of beats (Nbeats) (i.e., either the query segment 122 or the candidate segment 124). Next, a harmonic change rate between consecutive beats is determined, for each of the query track 112 and candidate track 110 under consideration, as follows. A Pearson correlation between consecutive beat-synchronised chroma vectors is determined, for all beats of each track 110, 112 (step 1404′), to provide a vector (Nbeats−1) of correlation values. In step 1406′, the correlation is mapped to change rate values according to formula (F13): Change=(1−corr)/2(F14).”); and assigning the chord type to the given beat in the chord string based on the determined ratio (Paragraph 129 “As a result, a vector is obtained with (Nbeats−1) change rate values for both candidate and query tracks, 110, 112, wherein the change rate value for the candidate (e.g., vocal) track 110 is represented by “CRvoc”, and the change rate value for the query (accompaniment) track 112 is represented by “CRacc”.”).
Regarding Claim 20, Bosch Vicente et al teaches, a mashup system (Title/Abstract), comprising: a hardware processor (FIG. 20, Processor Device 1110); and a non-transitory computer-readable storage medium storing executable instructions (FIG. 20, Main Memory 1125) that, when executed by the hardware processor, cause the hardware processor to perform steps comprising:
receiving, via a graphical user interface (GUI) presented on a user computing device (Paragraph 182, “The input control device(s) 1180 provide a portion of the user interface for a user of the computer 1100. The input control device(s) 1180 may include a keypad and/or a cursor control device. The keypad may be configured for inputting alphanumeric characters and/or other key information. The cursor control device may include, for example, a handheld controller or mouse, a trackball, a stylus, and/or cursor direction keys. In order to display textual and graphical information, the system 1100 may include the graphics subsystem 1160 and the output display 1170. The output display 1170 may include a display such as a CSTN (Color Super Twisted Nematic), TFT (Thin Film Transistor), TFD (Thin Film Diode), OLED (Organic Light-Emitting Diode), AMOLED display (Activematrix Organic Light-emitting Diode), and/or liquid crystal display (LCD)-type displays. The displays can also be touchscreen displays, such as capacitive and resistive-type touchscreen displays. The graphics subsystem 1160 receives textual and graphical information, and processes the information for output to the output display 1170.”), a selection of an audio snippet (Paragraph 9, “One aspect includes a method for combining audio tracks, comprising: determining at least one music track that is musically compatible with a base music track;”), the selection indicating an identifier of an audio file (Paragraph 72, “An additional type of information that can be employed to create mashups can include segment labelling information. Segment labelling information identifies (using, e.g., particular IDs) different types of track segments, and track segments may be labeled according to their similarity.”), a start time, and an end time (Paragraph 83, “In one example embodiment herein, the procedure 200 employs at least some of the various types of information 1131 as described above, including, without limitation, information about the likelihood of a segment containing vocals (vx) (e.g., at beats of segments), downbeat positions, song segmentation information (including start and end positions of segments), and segment labelling information (e.g., IDs), and the like.”), wherein the audio file is from among a plurality of audio files in a mashup catalog (Paragraph 77, “Each of the above types of information associated with particular tracks and/or with particular segments of tracks, can be stored in a database in association with the corresponding tracks and/or segments. The database may be, by example and without limitation, one or more of main memory 1125, portable storage medium 1150, and mass storage device 1130 of the system 1100 of FIG. 20 to be described below, or the database can be external to that system 1100, in which case it can be accessed by the system 1100 by way of, for example, network 1120 and peripheral device(s) 1140.”);
accessing a beat marking associated with the audio file, the beat marking indicating metrical information associated with the audio file, the metrical information including for each of a plurality of beats of the audio file, a beat number, a bar number, and a section number (Paragraph 71, “Example aspects of the present application also can employ song or track segmentation information for creating mashups. For example, song segmentation information can include the temporal positions of boundaries between sections of each track.”, Paragraph 74, “Additional information that can be employed to create mashups also can include tempo(s) of each track, a representation of tonality of each track (e.g., a twelve-dimensional chroma vector), beat/downbeat positions in each track (e.g., temporal positions of beats and downbeats in each track), information about the presence of vocals (if any) in time in each track, energy of each of the segments in the vocal and accompaniment tracks, or the like.”, Paragraph 83, “In one example embodiment herein, the procedure 200 employs at least some of the various types of information 1131 as described above, including, without limitation, information about the likelihood of a segment containing vocals (vx) (e.g., at beats of segments), downbeat positions, song segmentation information (including start and end positions of segments), and segment labelling information (e.g., IDs), and the like.”, Paragraph 84, “Referring to FIG. 2a, query segments 122 of the query track 112 that have less than a predetermined number of bars (e.g., eight bars) are filtered out and discarded (step 202), while others are maintained.”);
accessing a chord string associated with the audio file, the chord string indicating harmonic information associated with the audio file (Paragraph 122, “Referring again to FIG. 9a, after tempo compatibility (e.g., score K_seg_tempo) is determined in step 904, harmonic progression compatibility (also referred to herein as “harmonic compatibility”) is determined in step 906. In one example embodiment herein, the closer the harmonic compatibility of segments 110, 112 under consideration, the higher is the score.”), the harmonic information including a chord type for each of the plurality of beats of the audio file (FIG. 6, Step 602, Determine Musical Keys, as well as Steps 604, 606, 608.);
identifying a metrical signature and a chord string of the audio snippet, the metrical signature including a beat number and a bar number associated with a beat of the audio file corresponding to the start time, and the chord string including the chord type for each beat of the audio snippet (Paragraph 74, “Additional information that can be employed to create mashups also can include tempo(s) of each track, a representation of tonality of each track (e.g., a twelve-dimensional chroma vector), beat/downbeat positions in each track (e.g., temporal positions of beats and downbeats in each track), information about the presence of vocals (if any) in time in each track, energy of each of the segments in the vocal and accompaniment tracks, or the like.”, Paragraph 122, “Also, in one example embodiment herein, step 906 can be performed according to procedure 1100′ shown in FIG. 11. In step 1102′ beat synchronized chroma feature vectors are determined for each of the query segment 112 and candidate segment 110 under consideration, by determining, for each respective segment 110, 112, an average of chroma values within each beat of the respective segment 110, 112. In one example embodiment herein, the chroma values are obtained from among the information 1131 in the database using methods described in the Jehan reference discussed above. In step 1104′ a Pearson correlation between the beat synchronized chroma feature vectors determined in step 1102′, is determined for each of the beats of the segments under consideration.”);
identifying, from among the plurality of audio files, a plurality of mashup candidate audio snippets that match the metrical signature of the audio snippet and that have a beat length that matches a beat length of the audio snippet (Paragraph 80, “As represented in FIG. 1, the candidate track includes vocals 114 and the query track 112 includes separated vocal component/track 116 and separated instrumental component/track 118. In addition, additional track features 112a of the query track and additional track features 110a of the candidate track 110 are also identified from the query track 112 and candidate track 110. Track features 110a and 112a can include, for example, acoustic features (such as tempo, beat, musical key, likelihood of including vocals, and other features as described herein). Information regarding loudness 114b and tonality (e.g., tonal representation) 114a are obtained based on the vocal component 114 of the candidate track 110. Information regarding loudness 118b and tonality (e.g., tonal representation) 118a based on the separated instrumental component/track 118 and information regarding at least loudness 116a based on the separated vocal component/track 116 of the query track 112 are obtained.”);
comparing the chord string of the audio snippet with respective chord strings of each of the plurality of mashup candidate audio snippets to identify a subset of the plurality of mashup candidate audio snippets that harmonically match the audio snippet (Paragraph 122, “Referring again to FIG. 9a, after tempo compatibility (e.g., score K_seg_tempo) is determined in step 904, harmonic progression compatibility (also referred to herein as “harmonic compatibility”) is determined in step 906. In one example embodiment herein, the closer the harmonic compatibility of segments 110, 112 under consideration, the higher is the score. Also, in one example embodiment herein, step 906 can be performed according to procedure 1100′ shown in FIG. 11. In step 1102′ beat synchronized chroma feature vectors are determined for each of the query segment 112 and candidate segment 110 under consideration, by determining, for each respective segment 110, 112, an average of chroma values within each beat of the respective segment 110, 112. In one example embodiment herein, the chroma values are obtained from among the information 1131 in the database using methods described in the Jehan reference discussed above.”);
receiving, via the GUI presented on the user computing device, a selection of one of the subset of the plurality of mashup candidate audio snippets (Paragraph 144, “After computing the score (M) for all segments 124 of all candidate tracks 110 under consideration, the segment 124 with the highest total mashability score (M) is selected (step 1812), although in other example embodiments, a sampling between all possible candidate segments can be done with a probability which is proportional to their total mashability score. The above procedure can be performed with respect to all segments 122 that were assigned to S-subs and S_add of the query track 112 under consideration, starting from the start of the track 112 and finishing at the end of the track 112, to determine mashability between those segments 122 and individual ones of the candidate segments 124 of candidate tracks 110 that were selected as being compatible with the query track 112.”, Paragraph 95, “As a result of the procedure 300, a mashup track 120 (FIG. 1) is provided based on the query track 112 and at least one candidate track 110 under consideration.”);
and generating a mashup audio snippet based on the audio snippet and the selected one of the subset of the plurality of mashup candidate audio snippets, the mashup audio snippet including at least one stem from the audio snippet and at least one stem from the selected one of the subset of the plurality of mashup candidate audio snippets (Paragraph 95, “As a result of the procedure 300, a mashup track 120 (FIG. 1) is provided based on the query track 112 and at least one candidate track 110 under consideration. The mashup track 120 includes, by example, one or more segments 122 that were assigned to S_keep, one or more other segments 122 having vocal content (from one or more candidate tracks 110) that was used to replace vocal content of an original version of those other segments 122 in step 308, and one or more further segments 122 having vocal content (from one or more candidate tracks 110) that was added to those further segments 122 in step 308. In the mashup track 120, beat positions in the query track 112 are mapped with corresponding beat positions of the candidate track(s) 110.”, Paragraph 172, “As can be appreciated in view of the above description, at least some example aspects herein employ source separation to generate candidate (e.g., vocal) tracks and query (e.g., accompaniment) tracks, although in other example embodiments, stems can be used instead, or a multitrack can be employed where separation is therefore not needed). In other example embodiments herein, full tracks can be employed (without separation of vocals and accompaniment components).”).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DYLAN M NEECE whose telephone number is (703)756-1941. The examiner can normally be reached 10am - 7pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, CAROLYN EDWARDS can be reached at (571)-270-7136. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DYLAN MAGUIRE NEECE/ Examiner, Art Unit 2692
/CAROLYN R EDWARDS/ Supervisory Patent Examiner, Art Unit 2692