DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 2/7/2024 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the IDS is being considered by the examiner.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 1 recites the phrase, "the time domain features comprising an intermediate time domain feature and a target time domain feature.” It is unclear whether multiple time domain features must share a same intermediate time domain feature and a same target time domain feature, or if time domain features may comprise unique intermediate time domain features and target time domain features.
It is possible that Applicant intended to recite, "each time domain feature comprising an intermediate time domain feature and a target time domain feature."
Claims 2-7 are likewise rejected for depending, directly or indirectly from claim 1.
Claim 8 recites the phrase, "the time domain features comprising an intermediate time domain feature and a target time domain feature.” It is unclear whether multiple time domain features must share a same intermediate time domain feature and a same target time domain feature, or if time domain features may comprise unique intermediate time domain features and target time domain features.
It is possible that Applicant intended to recite, "each time domain feature comprising an intermediate time domain feature and a target time domain feature."
Claims 9-14 are likewise rejected for depending, directly or indirectly from claim 1.
Claim 15 recites the phrase, "the time domain features comprising an intermediate time domain feature and a target time domain feature.” It is unclear whether multiple time domain features must share a same intermediate time domain feature and a same target time domain feature, or if time domain features may comprise unique intermediate time domain features and target time domain features.
It is possible that Applicant intended to recite, "each time domain feature comprising an intermediate time domain feature and a target time domain feature.”
Claims 16-20 are likewise rejected for depending, directly or indirectly from claim 1.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-2, 8-9, and 15-16 are rejected under 35 U.S.C. as unpatentable over Siagian et al. (US 11342003 B1, May 24, 2022), in view of Georgiou et al. (US 20200335092 A1, October 22, 2020), hereinafter Georgiou, and further in view of Salamon et al. (US 20230115212 A1, filed May 11, 20220, hereinafter Salamon, to the extent understood.
Regarding claim 1, Siagian teaches an audio data processing method (Siagian col. 1, lines 65-67: "Various embodiments of the present disclosure introduce approaches for automatically segmenting a video content item using its accompanying audio.") performed by a computer device (Siagian col. 5, lines 3-12: " The client device 206 is representative of a plurality of client devices that may be coupled to the network 209. The client device 206 may comprise, for example, a processor-based system such as a computer system. Such a computer system may be embodied in the form of a desktop computer, a laptop computer, personal digital assistants, cellular telephones, smartphones, set-top boxes, music players, web pads, tablet computer systems, game consoles, electronic book readers, smartwatches, head mounted displays, voice interface devices, or other devices."), comprising: dividing audio data into multiple sub-audios (Siagian col. 6, lines 4-9: "In box 309, the segment generation service 218 divides the audio data 242 (FIG. 2) accompanying the video content item 224 into a plurality of atomic audio segments of a fixed length. In one implementation, an atomic audio segment is 20 milliseconds in length, or 320 samples using a 16 kHz sample rate."); separately extracting time domain features from the multiple sub-audios (Siagian col. 6, lines 10-14: "The audio data 242 may be available both in the time domain (i.e., samples corresponding to amplitude) and in the frequency domain (i.e., a vector representing frequency content)."); and separately extracting frequency domain features from the multiple sub-audios (Siagian col. 6, lines 10-14: "The audio data 242 may be available both in the time domain (i.e., samples corresponding to amplitude) and in the frequency domain (i.e., a vector representing frequency content).").
Siagian does not explicitly disclose: the time domain features comprising an intermediate time domain feature and a target time domain feature; the frequency domain features comprising an intermediate frequency domain feature and a target frequency domain feature; performing feature fusion on intermediate time domain features corresponding to the multiple sub-audios and intermediate frequency domain features corresponding to the multiple sub-audios, to obtain fusion features corresponding to the multiple sub-audios; performing semantic feature extraction based on target time domain features, target frequency domain features, and the fusion features that are corresponding to the multiple sub-audios, to obtain audio semantic features corresponding to the multiple sub-audios; determining each music segment from the multiple sub-audios based on the audio semantic features corresponding to the multiple sub-audios and a music semantic feature corresponding to the music segment; and performing music segment clustering based on the music semantic feature corresponding to each music segment, to obtain a same-type music segment set.
However, Georgiou teaches: the time domain features comprising an intermediate time domain feature and a target time domain feature (Georgiou ¶0006: " Each signal input is processed in a series of mode-specific processing stages (111-119, 121-129 in FIG. 1). Each successive mode-specific stage is associated with a successively longer scale of analysis of the signal input"); the frequency domain features comprising an intermediate frequency domain feature and a target frequency domain feature (Georgiou ¶0022: "The hierarchical representation of the various layers of abstraction can correspond to varying granularities of the information signals (e.g., time-scales) for example for audio text fusion these could correspond to phone, syllable, word, and utterance levels."); performing feature fusion on intermediate time domain features corresponding to the multiple sub-audios and intermediate frequency domain features corresponding to the multiple sub-audios, to obtain fusion features corresponding to the multiple sub-audios (Georgiou ¶0028: "Rather than each of the modes being processed to reach a corresponding mode-specific decision, and then forming some sort of fused decision from those mode-specific decisions, multiple (e.g., two or more, all) levels of processing in multiple modes pass their state or other output to corresponding fusion modules (191, 192, 199), shown in the figure with each level having a corresponding fusion module."); performing semantic feature extraction (Georgiou ¶0041: "After the multimodal information is propagated through the DHF network, we get a high-level representation f.sub.S for every spoken sentence. The role of the linear output layer is to transform this representation to a sentiment prediction.") based on target time domain features, target frequency domain features, and the fusion features that are corresponding to the multiple sub-audios, to obtain audio semantic features corresponding to the multiple sub-audios (Georgiou ¶0040: "The last fusion hierarchy level combines the high-level representations of the textual and acoustic modalities, {tilde over (g)} and {tilde over (h)}, with the sentence-level fused representation f.sub.U. This high-dimensional representation is passed through a Deep Neural Network (DNN), which outputs the sentiment level representation f.sub.S∈custom-character.sup.M.").
Furthermore, Salamon teaches: determining each music segment from the multiple sub-audios based on the audio semantic features corresponding to the multiple sub-audios (Salamon ¶0038: "The audio segmenting module 118 can then identify different segments of the audio sequence based on the cluster identifiers assigned to each frame. For example, the first 30 frames of the audio sequence may be assigned the same first cluster identifier indicating that they are all part of a first segment.") and a music semantic feature corresponding to the music segment (Salamon ¶0024: "Segments at distinct parts of the audio sequence that have the same cluster identifier can be identified as being similar segment types."); and performing music segment clustering based on the music semantic feature corresponding to each music segment, to obtain a same-type music segment set (Salamon ¶0038: "Non-consecutive segments that include frames assigned with the same cluster identifier represent a repetition within the audio sequence.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the audio processing method of Siagian by adding the time and frequency domain fusion of Georgiou to take advantage of the multiple levels of abstraction that information may be carried in co-evolving signal streams (Georgiou ¶0003) and the multi-level audio segmentation using deep embeddings of Salamon to reduce susceptibility to noise (Salamon ¶0004).
Regarding claim 2, Siagian (in view of Georgiou and further in view of Salamon) teaches a method comprising the features of claim 1 as discussed above.
Siagian further teaches performing music type classification and identification based on the audio semantic features to obtain a music type for the multiple sub-audios (Siagian col. 9, lines 60-63: "The segment generation service 218 may use a multi-class Support-Vector Machine to perform a classification of the segment as being speech, an instrumental, or a song.").
Regarding claim 8, Siagian teaches a computer device comprising a memory and a processor, the memory storing computer readable instructions (Siagian col. 5, lines 3-12: " The client device 206 is representative of a plurality of client devices that may be coupled to the network 209. The client device 206 may comprise, for example, a processor-based system such as a computer system. Such a computer system may be embodied in the form of a desktop computer, a laptop computer, personal digital assistants, cellular telephones, smartphones, set-top boxes, music players, web pads, tablet computer systems, game consoles, electronic book readers, smartwatches, head mounted displays, voice interface devices, or other devices.") that, when executed by the processor, cause the computer device to perform an audio data processing method (Siagian col. 1, lines 65-67: "Various embodiments of the present disclosure introduce approaches for automatically segmenting a video content item using its accompanying audio.") including: dividing audio data into multiple sub-audios (Siagian col. 6, lines 4-9: "In box 309, the segment generation service 218 divides the audio data 242 (FIG. 2) accompanying the video content item 224 into a plurality of atomic audio segments of a fixed length. In one implementation, an atomic audio segment is 20 milliseconds in length, or 320 samples using a 16 kHz sample rate."); separately extracting time domain features from the multiple sub-audios (Siagian col. 6, lines 10-14: "The audio data 242 may be available both in the time domain (i.e., samples corresponding to amplitude) and in the frequency domain (i.e., a vector representing frequency content)."); and separately extracting frequency domain features from the multiple sub-audios (Siagian col. 6, lines 10-14: "The audio data 242 may be available both in the time domain (i.e., samples corresponding to amplitude) and in the frequency domain (i.e., a vector representing frequency content).").
Siagian does not explicitly disclose: the time domain features comprising an intermediate time domain feature and a target time domain feature; the frequency domain features comprising an intermediate frequency domain feature and a target frequency domain feature; performing feature fusion on intermediate time domain features corresponding to the multiple sub-audios and intermediate frequency domain features corresponding to the multiple sub-audios, to obtain fusion features corresponding to the multiple sub-audios; performing semantic feature extraction based on target time domain features, target frequency domain features, and the fusion features that are corresponding to the multiple sub-audios, to obtain audio semantic features corresponding to the multiple sub-audios; determining each music segment from the multiple sub-audios based on the audio semantic features corresponding to the multiple sub-audios and a music semantic feature corresponding to the music segment; and performing music segment clustering based on the music semantic feature corresponding to each music segment, to obtain a same-type music segment set.
However, Georgiou teaches: the time domain features comprising an intermediate time domain feature and a target time domain feature (Georgiou ¶0006: " Each signal input is processed in a series of mode-specific processing stages (111-119, 121-129 in FIG. 1). Each successive mode-specific stage is associated with a successively longer scale of analysis of the signal input"); the frequency domain features comprising an intermediate frequency domain feature and a target frequency domain feature (Georgiou ¶0022: "The hierarchical representation of the various layers of abstraction can correspond to varying granularities of the information signals (e.g., time-scales) for example for audio text fusion these could correspond to phone, syllable, word, and utterance levels."); performing feature fusion on intermediate time domain features corresponding to the multiple sub-audios and intermediate frequency domain features corresponding to the multiple sub-audios, to obtain fusion features corresponding to the multiple sub-audios (Georgiou ¶0028: "Rather than each of the modes being processed to reach a corresponding mode-specific decision, and then forming some sort of fused decision from those mode-specific decisions, multiple (e.g., two or more, all) levels of processing in multiple modes pass their state or other output to corresponding fusion modules (191, 192, 199), shown in the figure with each level having a corresponding fusion module."); performing semantic feature extraction (Georgiou ¶0041: "After the multimodal information is propagated through the DHF network, we get a high-level representation f.sub.S for every spoken sentence. The role of the linear output layer is to transform this representation to a sentiment prediction.") based on target time domain features, target frequency domain features, and the fusion features that are corresponding to the multiple sub-audios, to obtain audio semantic features corresponding to the multiple sub-audios (Georgiou ¶0040: "The last fusion hierarchy level combines the high-level representations of the textual and acoustic modalities, {tilde over (g)} and {tilde over (h)}, with the sentence-level fused representation f.sub.U. This high-dimensional representation is passed through a Deep Neural Network (DNN), which outputs the sentiment level representation f.sub.S∈custom-character.sup.M.").
Furthermore, Salamon teaches: determining each music segment from the multiple sub-audios based on the audio semantic features corresponding to the multiple sub-audios (Salamon ¶0038: "The audio segmenting module 118 can then identify different segments of the audio sequence based on the cluster identifiers assigned to each frame. For example, the first 30 frames of the audio sequence may be assigned the same first cluster identifier indicating that they are all part of a first segment.") and a music semantic feature corresponding to the music segment (Salamon ¶0024: "Segments at distinct parts of the audio sequence that have the same cluster identifier can be identified as being similar segment types."); and performing music segment clustering based on the music semantic feature corresponding to each music segment, to obtain a same-type music segment set (Salamon ¶0038: "Non-consecutive segments that include frames assigned with the same cluster identifier represent a repetition within the audio sequence.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the computer device of Siagian by adding the time and frequency domain fusion of Georgiou to take advantage of the multiple levels of abstraction that information may be carried in co-evolving signal streams (Georgiou ¶0003) and the multi-level audio segmentation using deep embeddings of Salamon to reduce susceptibility to noise (Salamon ¶0004).
Regarding claim 9, Siagian (in view of Georgiou and further in view of Salamon) teaches a computer device comprising the features of claim 8 as discussed above.
Siagian further teaches performing music type classification and identification based on the audio semantic features to obtain a music type for the multiple sub-audios (Siagian col. 9, lines 60-63: "The segment generation service 218 may use a multi-class Support-Vector Machine to perform a classification of the segment as being speech, an instrumental, or a song.").
Regarding claim 15, Siagian teaches a non-transitory computer readable storage medium, storing computer readable instructions that, when executed by a processor of a computer device (Siagian col. 5, lines 3-12: " The client device 206 is representative of a plurality of client devices that may be coupled to the network 209. The client device 206 may comprise, for example, a processor-based system such as a computer system. Such a computer system may be embodied in the form of a desktop computer, a laptop computer, personal digital assistants, cellular telephones, smartphones, set-top boxes, music players, web pads, tablet computer systems, game consoles, electronic book readers, smartwatches, head mounted displays, voice interface devices, or other devices."), cause the computer device to perform an audio data processing method (Siagian col. 1, lines 65-67: "Various embodiments of the present disclosure introduce approaches for automatically segmenting a video content item using its accompanying audio.") including: dividing audio data into multiple sub-audios (Siagian col. 6, lines 4-9: "In box 309, the segment generation service 218 divides the audio data 242 (FIG. 2) accompanying the video content item 224 into a plurality of atomic audio segments of a fixed length. In one implementation, an atomic audio segment is 20 milliseconds in length, or 320 samples using a 16 kHz sample rate."); separately extracting time domain features from the multiple sub-audios (Siagian col. 6, lines 10-14: "The audio data 242 may be available both in the time domain (i.e., samples corresponding to amplitude) and in the frequency domain (i.e., a vector representing frequency content)."); and separately extracting frequency domain features from the multiple sub-audios (Siagian col. 6, lines 10-14: "The audio data 242 may be available both in the time domain (i.e., samples corresponding to amplitude) and in the frequency domain (i.e., a vector representing frequency content).").
Siagian does not explicitly disclose: the time domain features comprising an intermediate time domain feature and a target time domain feature; the frequency domain features comprising an intermediate frequency domain feature and a target frequency domain feature; performing feature fusion on intermediate time domain features corresponding to the multiple sub-audios and intermediate frequency domain features corresponding to the multiple sub-audios, to obtain fusion features corresponding to the multiple sub-audios; performing semantic feature extraction based on target time domain features, target frequency domain features, and the fusion features that are corresponding to the multiple sub-audios, to obtain audio semantic features corresponding to the multiple sub-audios; determining each music segment from the multiple sub-audios based on the audio semantic features corresponding to the multiple sub-audios and a music semantic feature corresponding to the music segment; and performing music segment clustering based on the music semantic feature corresponding to each music segment, to obtain a same-type music segment set.
However, Georgiou teaches: the time domain features comprising an intermediate time domain feature and a target time domain feature (Georgiou ¶0006: " Each signal input is processed in a series of mode-specific processing stages (111-119, 121-129 in FIG. 1). Each successive mode-specific stage is associated with a successively longer scale of analysis of the signal input"); the frequency domain features comprising an intermediate frequency domain feature and a target frequency domain feature (Georgiou ¶0022: "The hierarchical representation of the various layers of abstraction can correspond to varying granularities of the information signals (e.g., time-scales) for example for audio text fusion these could correspond to phone, syllable, word, and utterance levels."); performing feature fusion on intermediate time domain features corresponding to the multiple sub-audios and intermediate frequency domain features corresponding to the multiple sub-audios, to obtain fusion features corresponding to the multiple sub-audios (Georgiou ¶0028: "Rather than each of the modes being processed to reach a corresponding mode-specific decision, and then forming some sort of fused decision from those mode-specific decisions, multiple (e.g., two or more, all) levels of processing in multiple modes pass their state or other output to corresponding fusion modules (191, 192, 199), shown in the figure with each level having a corresponding fusion module."); performing semantic feature extraction (Georgiou ¶0041: "After the multimodal information is propagated through the DHF network, we get a high-level representation f.sub.S for every spoken sentence. The role of the linear output layer is to transform this representation to a sentiment prediction.") based on target time domain features, target frequency domain features, and the fusion features that are corresponding to the multiple sub-audios, to obtain audio semantic features corresponding to the multiple sub-audios (Georgiou ¶0040: "The last fusion hierarchy level combines the high-level representations of the textual and acoustic modalities, {tilde over (g)} and {tilde over (h)}, with the sentence-level fused representation f.sub.U. This high-dimensional representation is passed through a Deep Neural Network (DNN), which outputs the sentiment level representation f.sub.S∈custom-character.sup.M.").
Furthermore, Salamon teaches: determining each music segment from the multiple sub-audios based on the audio semantic features corresponding to the multiple sub-audios (Salamon ¶0038: "The audio segmenting module 118 can then identify different segments of the audio sequence based on the cluster identifiers assigned to each frame. For example, the first 30 frames of the audio sequence may be assigned the same first cluster identifier indicating that they are all part of a first segment.") and a music semantic feature corresponding to the music segment (Salamon ¶0024: "Segments at distinct parts of the audio sequence that have the same cluster identifier can be identified as being similar segment types."); and performing music segment clustering based on the music semantic feature corresponding to each music segment, to obtain a same-type music segment set (Salamon ¶0038: "Non-consecutive segments that include frames assigned with the same cluster identifier represent a repetition within the audio sequence.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the non-transitory computer readable storage medium of Siagian by adding the time and frequency domain fusion of Georgiou to take advantage of the multiple levels of abstraction that information may be carried in co-evolving signal streams (Georgiou ¶0003) and the multi-level audio segmentation using deep embeddings of Salamon to reduce susceptibility to noise (Salamon ¶0004).
Regarding claim 16, Siagian (in view of Georgiou and further in view of Salamon) teaches a non-transitory computer readable storage medium comprising the features of claim 15 as discussed above.
Siagian further teaches performing music type classification and identification based on the audio semantic features to obtain a music type for the multiple sub-audios (Siagian col. 9, lines 60-63: "The segment generation service 218 may use a multi-class Support-Vector Machine to perform a classification of the segment as being speech, an instrumental, or a song.").
Claims 3, 10, and 17 are rejected under 35 U.S.C. as unpatentable over Siagian, in view of Georgiou, and further in view of Salamon, Tsung et al. (US 20230124006 A1, filed October 15, 2021), hereinafter Tsung, and Li et al. ("Discriminative Neural Clustering for Speaker Diarisation," November 23, 2020, retrieved September 18, 2026 from https://arxiv.org/pdf/1910.09703), hereinafter Li, to the extent understood.
Regarding claim 3, Siagian (in view of Georgiou and further in view of Salamon) teaches a method a method comprising the features of claim 1 as discussed above.
Siagian (in view of Georgiou and further in view of Salamon) does not explicitly disclose that the performing music segment clustering based on the music semantic feature corresponding to each music segment comprises: separately performing sequence transform coding on the music semantic feature corresponding to each music segment to obtain an aggregation coding feature corresponding to each music segment; performing sequence transform decoding by using the aggregation coding feature to obtain a target music semantic feature corresponding to each music segment; and clustering each music segment according to the target music semantic feature corresponding to each music segment, to obtain the same-type music segment set.
However, Tsung teaches that the performing music segment clustering based on the music semantic feature corresponding to each music segment comprises: separately performing sequence transform coding on the music semantic feature corresponding to each music segment to obtain an aggregation coding feature corresponding to each music segment (Tsung ¶0021: "This architecture of implementing a spectral transformer and a temporal transformer is referred to herein as spectral-temporal TNT in which a plurality of such TNT blocks may be stacked to build the spectral-temporal TNT model architecture to learn the representation for audio data such as music signals, to perform tasks such as music information retrieval (MIR) research and analysis including, but not limited to, music tagging, vocal melody extraction, chord recognition, etc."); and performing sequence transform decoding by using the aggregation coding feature to obtain a target music semantic feature corresponding to each music segment (Tsung ¶0045: "The decoder 700 of the spectral transformer block 608 and the decoder 702 of the temporal transformer block 610 also have similar component blocks, mainly the multi-head self-attention block 508, the feed-forward network block 510, the layer normalization block 506, and an encoder-decoder attention block 704 which helps the decoder 700 or 702 focus on the appropriate matrices that are outputted from each encoder.").
Furthermore, Li teaches that clustering each music segment according to the target music semantic feature corresponding to each music segment, to obtain the same-type music segment set (Li § 3: "the task of clustering can be considered as a special sequence-to-sequence classification problem using the attention-based encoder-decoder structure [28], in which the input and output sequences have equal lengths and each output target represents the cluster label of the corresponding input.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the audio processing method of Siagian (as modified by Georgiou and Salamon) by adding the aggregation coding of Tsung and Li to use learned parametric clustering to avoid matching assumptions made by the clustering algorithms that are related to the underlying data distributions or distance measures (Li § 1).
Regarding claim 10, Siagian (in view of Georgiou and further in view of Salamon) teaches a computer device comprising the features of claim 8 as discussed above.
Siagian (in view of Georgiou and further in view of Salamon) does not explicitly disclose that the performing music segment clustering based on the music semantic feature corresponding to each music segment comprises: separately performing sequence transform coding on the music semantic feature corresponding to each music segment to obtain an aggregation coding feature corresponding to each music segment; performing sequence transform decoding by using the aggregation coding feature to obtain a target music semantic feature corresponding to each music segment; and clustering each music segment according to the target music semantic feature corresponding to each music segment, to obtain the same-type music segment set.
However, Tsung teaches that the performing music segment clustering based on the music semantic feature corresponding to each music segment comprises: separately performing sequence transform coding on the music semantic feature corresponding to each music segment to obtain an aggregation coding feature corresponding to each music segment (Tsung ¶0021: "This architecture of implementing a spectral transformer and a temporal transformer is referred to herein as spectral-temporal TNT in which a plurality of such TNT blocks may be stacked to build the spectral-temporal TNT model architecture to learn the representation for audio data such as music signals, to perform tasks such as music information retrieval (MIR) research and analysis including, but not limited to, music tagging, vocal melody extraction, chord recognition, etc."); and performing sequence transform decoding by using the aggregation coding feature to obtain a target music semantic feature corresponding to each music segment (Tsung ¶0045: "The decoder 700 of the spectral transformer block 608 and the decoder 702 of the temporal transformer block 610 also have similar component blocks, mainly the multi-head self-attention block 508, the feed-forward network block 510, the layer normalization block 506, and an encoder-decoder attention block 704 which helps the decoder 700 or 702 focus on the appropriate matrices that are outputted from each encoder.").
Furthermore, Li teaches that clustering each music segment according to the target music semantic feature corresponding to each music segment, to obtain the same-type music segment set (Li § 3: "the task of clustering can be considered as a special sequence-to-sequence classification problem using the attention-based encoder-decoder structure [28], in which the input and output sequences have equal lengths and each output target represents the cluster label of the corresponding input.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the computer device of Siagian (as modified by Georgiou and Salamon) by adding the aggregation coding of Tsung and Li to use learned parametric clustering to avoid matching assumptions made by the clustering algorithms that are related to the underlying data distributions or distance measures (Li § 1).
Regarding claim 17, Siagian (in view of Georgiou and further in view of Salamon) teaches a non-transitory computer readable storage medium comprising the features of claim 15 as discussed above.
Siagian (in view of Georgiou and further in view of Salamon) does not explicitly disclose that the performing music segment clustering based on the music semantic feature corresponding to each music segment comprises: separately performing sequence transform coding on the music semantic feature corresponding to each music segment to obtain an aggregation coding feature corresponding to each music segment; performing sequence transform decoding by using the aggregation coding feature to obtain a target music semantic feature corresponding to each music segment; and clustering each music segment according to the target music semantic feature corresponding to each music segment, to obtain the same-type music segment set.
However, Tsung teaches that the performing music segment clustering based on the music semantic feature corresponding to each music segment comprises: separately performing sequence transform coding on the music semantic feature corresponding to each music segment to obtain an aggregation coding feature corresponding to each music segment (Tsung ¶0021: "This architecture of implementing a spectral transformer and a temporal transformer is referred to herein as spectral-temporal TNT in which a plurality of such TNT blocks may be stacked to build the spectral-temporal TNT model architecture to learn the representation for audio data such as music signals, to perform tasks such as music information retrieval (MIR) research and analysis including, but not limited to, music tagging, vocal melody extraction, chord recognition, etc."); and performing sequence transform decoding by using the aggregation coding feature to obtain a target music semantic feature corresponding to each music segment (Tsung ¶0045: "The decoder 700 of the spectral transformer block 608 and the decoder 702 of the temporal transformer block 610 also have similar component blocks, mainly the multi-head self-attention block 508, the feed-forward network block 510, the layer normalization block 506, and an encoder-decoder attention block 704 which helps the decoder 700 or 702 focus on the appropriate matrices that are outputted from each encoder.").
Furthermore, Li teaches that clustering each music segment according to the target music semantic feature corresponding to each music segment, to obtain the same-type music segment set (Li § 3: "the task of clustering can be considered as a special sequence-to-sequence classification problem using the attention-based encoder-decoder structure [28], in which the input and output sequences have equal lengths and each output target represents the cluster label of the corresponding input.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the non-transitory computer readable storage medium of Siagian (as modified by Georgiou and Salamon) by adding the aggregation coding of Tsung and Li to use learned parametric clustering to avoid matching assumptions made by the clustering algorithms that are related to the underlying data distributions or distance measures (Li § 1).
Claims 4, 11, and 18 are rejected under 35 U.S.C. as unpatentable over Siagian, in view of Georgiou, and further in view of Salamon, Kong et al. ("PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition," August 23, 2020, retrieved September 18, 2026 from https://arxiv.org/pdf/1912.10211), hereinafter Kong, and Tsung, to the extent understood.
Regarding claim 4, Siagian (in view of Georgiou and further in view of Salamon) teaches a method a method comprising the features of claim 1 as discussed above.
Siagian (in view of Georgiou and further in view of Salamon) does not explicitly disclose that the separately extracting time domain features from the multiple sub-audios, the time domain features comprising an intermediate time domain feature and a target time domain feature comprises: separately performing a time domain convolution operation on the multiple sub-audios to obtain at least two intermediate convolution features corresponding to the multiple sub-audios and a final convolution feature; performing frequency domain dimension transform on the at least two intermediate convolution features to obtain at least two intermediate time domain features corresponding to the multiple sub-audios; and performing frequency domain dimension transform on the final convolution feature to obtain target time domain features corresponding to the multiple sub-audios.
However, Kong teaches that the separately extracting time domain features from the multiple sub-audios, the time domain features comprising an intermediate time domain feature and a target time domain feature comprises: separately performing a time domain convolution operation on the multiple sub-audios to obtain at least two intermediate convolution features corresponding to the multiple sub-audios and a final convolution feature (Kong § III(A): "To build a Wavegram, we first apply a one-dimensional CNN to time-domain waveform. The one-dimensional CNN begins with a convolutional layer with filter length 11 and stride 5 to reduce the size of the input. This immediately reduces the input lengths by a factor of 5 times to reduce memory usage. This is followed by three convolutional blocks, where each convolutional block consists of two convolutional layers with dilations of 1 and 2, respectively, which are designed to increase the receptive field of the convolutional layers."); and performing frequency domain dimension transform on the final convolution feature to obtain target time domain features corresponding to the multiple sub-audios (Kong § III(A): "We reshape this output to a tensor with a size of T × F × C/F by splitting C channels into C/F groups, where each group has F frequency bins. We call this tensor a Wavegram.").
Furthermore, Tsung teaches performing frequency domain dimension transform on the at least two intermediate convolution features to obtain at least two intermediate time domain features corresponding to the multiple sub-audios (Tsung ¶0030: "For example, each of the temporal embedding vectors, that is, e.sup.ℓ-1.sub.1, e.sup.ℓ-1.sub.2, ..., e.sup.ℓ- .sup.1.sub.T′, of the learnable matrix E.sup.ℓ-1 is passed through the linear projection layer 404, which transforms the vectors from having the dimension of D to having the dimension of K′." Tsung teaches transforming the dimensions of intermediate representations at each of the stacked blocks which can be reshaped by Kong as discussed above.).
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the audio processing method of Siagian (as modified by Georgiou and Salamon) by adding the separate time domain convolution and frequency domain dimension transforms of Kong and Tsung to take advantage of the multiple levels of abstraction that information may be carried in co-evolving signal streams (Georgiou ¶0003).
Regarding claim 11, Siagian (in view of Georgiou and further in view of Salamon) teaches a computer device comprising the features of claim 8 as discussed above.
Siagian (in view of Georgiou and further in view of Salamon) does not explicitly disclose that the separately extracting time domain features from the multiple sub-audios, the time domain features comprising an intermediate time domain feature and a target time domain feature comprises: separately performing a time domain convolution operation on the multiple sub-audios to obtain at least two intermediate convolution features corresponding to the multiple sub-audios and a final convolution feature; performing frequency domain dimension transform on the at least two intermediate convolution features to obtain at least two intermediate time domain features corresponding to the multiple sub-audios; and performing frequency domain dimension transform on the final convolution feature to obtain target time domain features corresponding to the multiple sub-audios.
However, Kong teaches that the separately extracting time domain features from the multiple sub-audios, the time domain features comprising an intermediate time domain feature and a target time domain feature comprises: separately performing a time domain convolution operation on the multiple sub-audios to obtain at least two intermediate convolution features corresponding to the multiple sub-audios and a final convolution feature (Kong § III(A): "To build a Wavegram, we first apply a one-dimensional CNN to time-domain waveform. The one-dimensional CNN begins with a convolutional layer with filter length 11 and stride 5 to reduce the size of the input. This immediately reduces the input lengths by a factor of 5 times to reduce memory usage. This is followed by three convolutional blocks, where each convolutional block consists of two convolutional layers with dilations of 1 and 2, respectively, which are designed to increase the receptive field of the convolutional layers."); and performing frequency domain dimension transform on the final convolution feature to obtain target time domain features corresponding to the multiple sub-audios (Kong § III(A): "We reshape this output to a tensor with a size of T × F × C/F by splitting C channels into C/F groups, where each group has F frequency bins. We call this tensor a Wavegram.").
Furthermore, Tsung teaches performing frequency domain dimension transform on the at least two intermediate convolution features to obtain at least two intermediate time domain features corresponding to the multiple sub-audios (Tsung ¶0030: "For example, each of the temporal embedding vectors, that is, e.sup.ℓ-1.sub.1, e.sup.ℓ-1.sub.2, ..., e.sup.ℓ- .sup.1.sub.T′, of the learnable matrix E.sup.ℓ-1 is passed through the linear projection layer 404, which transforms the vectors from having the dimension of D to having the dimension of K′." Tsung teaches transforming the dimensions of intermediate representations at each of the stacked blocks which can be reshaped by Kong as discussed above.).
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the computer device of Siagian (as modified by Georgiou and Salamon) by adding the separate time domain convolution and frequency domain dimension transforms of Kong and Tsung to take advantage of the multiple levels of abstraction that information may be carried in co-evolving signal streams (Georgiou ¶0003).
Regarding claim 18, Siagian (in view of Georgiou and further in view of Salamon) teaches a non-transitory computer readable storage medium comprising the features of claim 15 as discussed above.
Siagian (in view of Georgiou and further in view of Salamon) does not explicitly disclose that the separately extracting time domain features from the multiple sub-audios, the time domain features comprising an intermediate time domain feature and a target time domain feature comprises: separately performing a time domain convolution operation on the multiple sub-audios to obtain at least two intermediate convolution features corresponding to the multiple sub-audios and a final convolution feature; performing frequency domain dimension transform on the at least two intermediate convolution features to obtain at least two intermediate time domain features corresponding to the multiple sub-audios; and performing frequency domain dimension transform on the final convolution feature to obtain target time domain features corresponding to the multiple sub-audios.
However, Kong teaches that the separately extracting time domain features from the multiple sub-audios, the time domain features comprising an intermediate time domain feature and a target time domain feature comprises: separately performing a time domain convolution operation on the multiple sub-audios to obtain at least two intermediate convolution features corresponding to the multiple sub-audios and a final convolution feature (Kong § III(A): "To build a Wavegram, we first apply a one-dimensional CNN to time-domain waveform. The one-dimensional CNN begins with a convolutional layer with filter length 11 and stride 5 to reduce the size of the input. This immediately reduces the input lengths by a factor of 5 times to reduce memory usage. This is followed by three convolutional blocks, where each convolutional block consists of two convolutional layers with dilations of 1 and 2, respectively, which are designed to increase the receptive field of the convolutional layers."); and performing frequency domain dimension transform on the final convolution feature to obtain target time domain features corresponding to the multiple sub-audios (Kong § III(A): "We reshape this output to a tensor with a size of T × F × C/F by splitting C channels into C/F groups, where each group has F frequency bins. We call this tensor a Wavegram.").
Furthermore, Tsung teaches performing frequency domain dimension transform on the at least two intermediate convolution features to obtain at least two intermediate time domain features corresponding to the multiple sub-audios (Tsung ¶0030: "For example, each of the temporal embedding vectors, that is, e.sup.ℓ-1.sub.1, e.sup.ℓ-1.sub.2, ..., e.sup.ℓ- .sup.1.sub.T′, of the learnable matrix E.sup.ℓ-1 is passed through the linear projection layer 404, which transforms the vectors from having the dimension of D to having the dimension of K′." Tsung teaches transforming the dimensions of intermediate representations at each of the stacked blocks which can be reshaped by Kong as discussed above.).
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the non-transitory computer readable storage medium of Siagian (as modified by Georgiou and Salamon) by adding the separate time domain convolution and frequency domain dimension transforms of Kong and Tsung to take advantage of the multiple levels of abstraction that information may be carried in co-evolving signal streams (Georgiou ¶0003).
Claims 5, 12, and 19 are rejected under 35 U.S.C. as unpatentable over Siagian, in view of Georgiou, and further in view of Salamon and Kong, to the extent understood.
Regarding claim 5, Siagian (in view of Georgiou and further in view of Salamon) teaches a method a method comprising the features of claim 1 as discussed above.
Siagian (in view of Georgiou and further in view of Salamon) does not explicitly disclose that the separately extracting frequency domain features from the multiple sub-audios, the frequency domain features comprising an intermediate frequency domain feature and a target frequency domain feature comprises: extracting basic audio features corresponding to the multiple sub-audios; and performing a frequency domain convolution operation on the basic audio features corresponding to the multiple sub-audios to obtain at least two intermediate frequency domain features and target frequency domain features corresponding to the multiple sub-audios.
However, Kong teaches that the separately extracting frequency domain features from the multiple sub-audios, the frequency domain features comprising an intermediate frequency domain feature and a target frequency domain feature comprises: extracting basic audio features corresponding to the multiple sub-audios (Kong § II(A)(1): "Short time Fourier transforms (STFTs) are applied to time-domain waveforms to calculate spectrograms .Then, mel filter banks are applied to the spectrograms, followed by a logarithmic operation to extract log mel spectrograms."); and performing a frequency domain convolution operation on the basic audio features corresponding to the multiple sub-audios to obtain at least two intermediate frequency domain features and target frequency domain features corresponding to the multiple sub-audios (Kong § II(A)(2): "The 10- and 14-layer CNNs consist of 4 and 6 convolutional layers, respectively, inspired by the VGG-like CNNs." Kong teaches a multi-layered 2D convolution stack over the spectrogram. Each layer's output feeds the next, yielding intermediate frequency domain features at every level, along with a final (target) feature from the final layer.).
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the audio processing method of Siagian (as modified by Georgiou and Salamon) by adding the frequency domain convolution of Kong to capture frequency domain information (Kong § III(A)).
Regarding claim 12, Siagian (in view of Georgiou and further in view of Salamon) teaches a computer device comprising the features of claim 8 as discussed above.
Siagian (in view of Georgiou and further in view of Salamon) does not explicitly disclose that the separately extracting frequency domain features from the multiple sub-audios, the frequency domain features comprising an intermediate frequency domain feature and a target frequency domain feature comprises: extracting basic audio features corresponding to the multiple sub-audios; and performing a frequency domain convolution operation on the basic audio features corresponding to the multiple sub-audios to obtain at least two intermediate frequency domain features and target frequency domain features corresponding to the multiple sub-audios.
However, Kong teaches that the separately extracting frequency domain features from the multiple sub-audios, the frequency domain features comprising an intermediate frequency domain feature and a target frequency domain feature comprises: extracting basic audio features corresponding to the multiple sub-audios (Kong § II(A)(1): "Short time Fourier transforms (STFTs) are applied to time-domain waveforms to calculate spectrograms .Then, mel filter banks are applied to the spectrograms, followed by a logarithmic operation to extract log mel spectrograms."); and performing a frequency domain convolution operation on the basic audio features corresponding to the multiple sub-audios to obtain at least two intermediate frequency domain features and target frequency domain features corresponding to the multiple sub-audios (Kong § II(A)(2): "The 10- and 14-layer CNNs consist of 4 and 6 convolutional layers, respectively, inspired by the VGG-like CNNs." Kong teaches a multi-layered 2D convolution stack over the spectrogram. Each layer's output feeds the next, yielding intermediate frequency domain features at every level, along with a final (target) feature from the final layer.).
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the computer device of Siagian (as modified by Georgiou and Salamon) by adding the frequency domain convolution of Kong to capture frequency domain information (Kong § III(A)).
Regarding claim 19, Siagian (in view of Georgiou and further in view of Salamon) teaches a non-transitory computer readable storage medium comprising the features of claim 15 as discussed above.
Siagian (in view of Georgiou and further in view of Salamon) does not explicitly disclose that the separately extracting frequency domain features from the multiple sub-audios, the frequency domain features comprising an intermediate frequency domain feature and a target frequency domain feature comprises: extracting basic audio features corresponding to the multiple sub-audios; and performing a frequency domain convolution operation on the basic audio features corresponding to the multiple sub-audios to obtain at least two intermediate frequency domain features and target frequency domain features corresponding to the multiple sub-audios.
However, Kong teaches that the separately extracting frequency domain features from the multiple sub-audios, the frequency domain features comprising an intermediate frequency domain feature and a target frequency domain feature comprises: extracting basic audio features corresponding to the multiple sub-audios (Kong § II(A)(1): "Short time Fourier transforms (STFTs) are applied to time-domain waveforms to calculate spectrograms .Then, mel filter banks are applied to the spectrograms, followed by a logarithmic operation to extract log mel spectrograms."); and performing a frequency domain convolution operation on the basic audio features corresponding to the multiple sub-audios to obtain at least two intermediate frequency domain features and target frequency domain features corresponding to the multiple sub-audios (Kong § II(A)(2): "The 10- and 14-layer CNNs consist of 4 and 6 convolutional layers, respectively, inspired by the VGG-like CNNs." Kong teaches a multi-layered 2D convolution stack over the spectrogram. Each layer's output feeds the next, yielding intermediate frequency domain features at every level, along with a final (target) feature from the final layer.).
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the non-transitory computer readable storage medium of Siagian (as modified by Georgiou and Salamon) by adding the frequency domain convolution of Kong to capture frequency domain information (Kong § III(A)).
Claims 6 and 13 are rejected under 35 U.S.C. as unpatentable over Siagian, in view of Georgiou, and further in view of Salamon and Stoller et al. ("WAVE-U-NET: A Multi-Scale Neural Network for End-to-End Audio Source Separation," June 8, 2018, retrieved September 18, 2026 from https://arxiv.org/pdf/1806.03185), hereinafter Stoller, to the extent understood.
Regarding claim 6, Siagian (in view of Georgiou and further in view of Salamon) teaches a method a method comprising the features of claim 1 as discussed above.
Georgiou further teaches that the performing feature fusion on intermediate time domain features corresponding to the multiple sub-audios and intermediate frequency domain features corresponding to the multiple sub-audios, to obtain fusion features corresponding to the multiple sub-audios comprises: concatenating a first intermediate time domain feature and a corresponding first intermediate frequency domain feature to obtain a first concatenation feature (Georgiou ¶0015: "The fusion of the information signal representations can be achieved via a variety of techniques including but not limited to concatenation, averaging, pooling, conditioning, product, transformation by forward or recursive neural network layers."); and concatenating the first fusion feature, a second intermediate time domain feature, and a corresponding second intermediate frequency domain feature to obtain a second concatenation feature (Georgiou ¶0039: "This module accepts as inputs three information flows 1) sentence-level representation g from the text encoder, 2) sentence-level representation h from the audio encoder and 3) the previous-level fused representation f.sub.W.").
Siagian (in view of Georgiou and further in view of Salamon) does not explicitly disclose performing a convolution operation based on the first concatenation feature to obtain a first fusion feature, and performing a convolution operation based on the second concatenation feature to obtain a fusion feature corresponding to the multiple sub-audios.
However, Stoller teaches performing a convolution operation based on the first concatenation feature to obtain a first fusion feature (Stoller § 3.1: "A diagram of the Wave-U-Net architecture is shown in Fig ure 1. It computes an increasing number of higher-level features on coarser time scales using downsampling (DS) blocks. These features are combined with the earlier computed local, high-resolution features using upsampling (US) blocks, yielding multi-scale features which are used for making predictions. The network has L levels in total, with each successive level operating at half the time resolution as the previous one. For K sources to be estimated, the model returns predictions in the interval (−1,1), one for each source audio sample." Stoller § 3.1 and Table 1 teaches concatenating then convolving with the per-level fusion cell of an audio domain network and applies it to features carried forward from the previous level.), and performing a convolution operation based on the second concatenation feature to obtain a fusion feature corresponding to the multiple sub-audios (Stoller abstract: "In this context, we propose the Wave-U-Net, an adaptation of the U-Net to the one-dimensional time domain, which repeatedly resamples feature maps to compute and com bine features at different time scales.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the audio processing method of Siagian (as modified by Georgiou and Salamon) by adding the concatenations of Georgiou and Stoller to take advantage of the multiple levels of abstraction that information may be carried in co-evolving signal streams (Georgiou ¶0003).
PNG
media_image1.png
259
517
media_image1.png
Greyscale
Regarding claim 13, Siagian (in view of Georgiou and further in view of Salamon) teaches a computer device comprising the features of claim 8 as discussed above.
Georgiou further teaches that the performing feature fusion on intermediate time domain features corresponding to the multiple sub-audios and intermediate frequency domain features corresponding to the multiple sub-audios, to obtain fusion features corresponding to the multiple sub-audios comprises: concatenating a first intermediate time domain feature and a corresponding first intermediate frequency domain feature to obtain a first concatenation feature (Georgiou ¶0015: "The fusion of the information signal representations can be achieved via a variety of techniques including but not limited to concatenation, averaging, pooling, conditioning, product, transformation by forward or recursive neural network layers."); and concatenating the first fusion feature, a second intermediate time domain feature, and a corresponding second intermediate frequency domain feature to obtain a second concatenation feature (Georgiou ¶0039: "This module accepts as inputs three information flows 1) sentence-level representation g from the text encoder, 2) sentence-level representation h from the audio encoder and 3) the previous-level fused representation f.sub.W.").
Siagian (in view of Georgiou and further in view of Salamon) does not explicitly disclose performing a convolution operation based on the first concatenation feature to obtain a first fusion feature, and performing a convolution operation based on the second concatenation feature to obtain a fusion feature corresponding to the multiple sub-audios.
However, Stoller teaches performing a convolution operation based on the first concatenation feature to obtain a first fusion feature (Stoller § 3.1: "A diagram of the Wave-U-Net architecture is shown in Fig ure 1. It computes an increasing number of higher-level features on coarser time scales using downsampling (DS) blocks. These features are combined with the earlier computed local, high-resolution features using upsampling (US) blocks, yielding multi-scale features which are used for making predictions. The network has L levels in total, with each successive level operating at half the time resolution as the previous one. For K sources to be estimated, the model returns predictions in the interval (−1,1), one for each source audio sample." Stoller § 3.1 and Table 1 teaches concatenating then convolving with the per-level fusion cell of an audio domain network and applies it to features carried forward from the previous level.), and performing a convolution operation based on the second concatenation feature to obtain a fusion feature corresponding to the multiple sub-audios (Stoller abstract: "In this context, we propose the Wave-U-Net, an adaptation of the U-Net to the one-dimensional time domain, which repeatedly resamples feature maps to compute and com bine features at different time scales.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the computer device of Siagian (as modified by Georgiou and Salamon) by adding the concatenations of Georgiou and Stoller to take advantage of the multiple levels of abstraction that information may be carried in co-evolving signal streams (Georgiou ¶0003).
Claims 7, 14, and 20 are rejected under 35 U.S.C. as unpatentable over Siagian, in view of Georgiou, and further in view of Salamon and Jiang et al. (US 10134440 B2, November 20, 2018), hereinafter Jiang, to the extent understood.
Regarding claim 7, Siagian (in view of Georgiou and further in view of Salamon) teaches a method a method comprising the features of claim 1 as discussed above.
Siagian (in view of Georgiou and further in view of Salamon) does not explicitly disclose obtaining video segments corresponding to the music segments in the same-type music segment set, to obtain a video segment set; and concatenating the same-type music segment set and the video segment set to obtain a same-type audio-video set.
However, Jiang teaches obtaining video segments corresponding to the music segments in the same-type music segment set, to obtain a video segment set (Jiang col. 10, lines 22-25: "An extract image frames subsets step 400 is used to identify a set of image frame subsets 405. Each image frame subset 405 includes a set of image frames 210 corresponding to each of one of the audio segments 255."); and concatenating the same-type music segment set and the video segment set to obtain a same-type audio-video set (Jiang abstract: "forming an audio-visual slideshow by combining the selected key frames with the audio summary, wherein the selected key frames are displayed synchronously with their corresponding audio segment").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the audio processing method of Siagian (as modified by Georgiou and Salamon) by adding the video segments of Jiang to satisfy a need for a robust video summarization method that can be applied to consumer video sequences (Jiang col. 2, lines 7-8).
Regarding claim 14, Siagian (in view of Georgiou and further in view of Salamon) teaches a computer device comprising the features of claim 8 as discussed above.
Siagian (in view of Georgiou and further in view of Salamon) does not explicitly disclose obtaining video segments corresponding to the music segments in the same-type music segment set, to obtain a video segment set; and concatenating the same-type music segment set and the video segment set to obtain a same-type audio-video set.
However, Jiang teaches obtaining video segments corresponding to the music segments in the same-type music segment set, to obtain a video segment set (Jiang col. 10, lines 22-25: "An extract image frames subsets step 400 is used to identify a set of image frame subsets 405. Each image frame subset 405 includes a set of image frames 210 corresponding to each of one of the audio segments 255."); and concatenating the same-type music segment set and the video segment set to obtain a same-type audio-video set (Jiang abstract: "forming an audio-visual slideshow by combining the selected key frames with the audio summary, wherein the selected key frames are displayed synchronously with their corresponding audio segment").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the computer device of Siagian (as modified by Georgiou and Salamon) by adding the video segments of Jiang to satisfy a need for a robust video summarization method that can be applied to consumer video sequences (Jiang col. 2, lines 7-8).
Regarding claim 20, Siagian (in view of Georgiou and further in view of Salamon) teaches a non-transitory computer readable storage medium comprising the features of claim 15 as discussed above.
Siagian (in view of Georgiou and further in view of Salamon) does not explicitly disclose obtaining video segments corresponding to the music segments in the same-type music segment set, to obtain a video segment set; and concatenating the same-type music segment set and the video segment set to obtain a same-type audio-video set.
However, Jiang teaches obtaining video segments corresponding to the music segments in the same-type music segment set, to obtain a video segment set (Jiang col. 10, lines 22-25: "An extract image frames subsets step 400 is used to identify a set of image frame subsets 405. Each image frame subset 405 includes a set of image frames 210 corresponding to each of one of the audio segments 255."); and concatenating the same-type music segment set and the video segment set to obtain a same-type audio-video set (Jiang abstract: "forming an audio-visual slideshow by combining the selected key frames with the audio summary, wherein the selected key frames are displayed synchronously with their corresponding audio segment").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the non-transitory computer readable storage medium of Siagian (as modified by Georgiou and Salamon) by adding the video segments of Jiang to satisfy a need for a robust video summarization method that can be applied to consumer video sequences (Jiang col. 2, lines 7-8).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PHILIP SCOLES whose telephone number is (703)756-1831. The examiner can normally be reached Monday-Friday 8:30-4:30 ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Dedei Hammond can be reached on 571-270-7938. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PHILIP G SCOLES/
Examiner, Art Unit 2837
/DEDEI K HAMMOND/Supervisory Patent Examiner, Art Unit 2837