Detailed Action
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 2-7, 10, and 12-20 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 2 recites “the first time point having a first value and the second time point having a second value equal to a threshold value within the distribution of probabilities” which is indefinite. It is unclear what the relationship between the time points and the threshold value is requiring. For example, are both the first and second time points equal to the same threshold value, their own distinct threshold values, or does the threshold value only apply to the second time point. Thus, one of ordinary skill in the art would not be able to ascertain the scope of the claimed invention, rendering the claim indefinite. For examination purposes, the limitation will be interpreted to mean that both the first and second time points are equal to the same threshold value.
Claim 12 contains limitations found analogous to that of claim 2. Therefore, claim 12 is rejected for the same reason.
Claim 6 recites “obtain content different from the multimedia content, segmented from the video signal during time from the first time point to the second time point”, which is indefinite. It is unclear what is meant by obtaining content different from the multimedia content in the language of the claim. For example, the claim clarifies that this different content is “segmented from the video signal,” meaning it is extracted from source content, which contradicts the requirement of being “different” content. Thus, one of ordinary skill in the art would not be able to ascertain the scope of the claimed invention, rendering the claim indefinite. For examination purposes, the limitation will be interpreted to mean obtaining any content other than the original source data, which includes for example, segments of the source video.
Claim 16 contains limitations found analogous to that of claim 6. Therefore, claim 16 is rejected for the same reason.
Claim 18 recites “in response to identifying one or more time points having values above the threshold value within the audio signal and based on a video signal within different time intervals including the one or more time points, selecting a time point from among the one or more time points as the sound time point when the designated motion is captured”, which is indefinite. Particularly, “based on a video signal within different time intervals including the one or more time points” does not clearly define what the relationship is between the video signal, time intervals, and time points. For example, does the video signal contain multiple different time intervals with each containing one of the time points, or are there multiple time points per time interval. Furthermore, it is unclear how these time intervals are “different”, such as different positions, durations, or content. Thus, one of ordinary skill in the art would not be able to ascertain the scope of the claimed invention, rendering the claim indefinite. For examination purposes, the limitation will be interpreted to mean that the video signal contains multiple time intervals which include the one or more time points.
Claim 3-5, 7, 10, 13-15, 17, and 19-20 are rejected as being dependent on a rejected base claim.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-4, 7-14, and 17-20 are rejected under 35 U.S.C. 101.
Claim 1 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea of determining the timing of a motion by correlating auditory and visual information.
The claim recites: “An electronic device comprising: memory for storing instructions; and at least one processor operably coupled to the memory, wherein the at least one processor, when the instructions are executed, is configured to: receive a request for detecting a sound time point in a multimedia content when a designated motion is captured; obtain a distribution of probabilities in a time domain that the designated motion is performed based on an audio signal in the multimedia content, wherein the distribution of probabilities comprises a plurality of peak values corresponding to respective time points in the multimedia content; and obtain the sound time point when the designated motion is captured from among the respective time points corresponding to the plurality of peak values, using a video signal synchronized to the audio signal, in the multimedia content.”
The limitations, as drafted, are processes that, under their broadest reasonable interpretation, cover performance of the limitation in the human mind. A person can mentally receive a request to identify when a motion occurs, analyze an audio signal to infer probabilities that the motion correlates to audio peaks in the signal, and visually observe a synchronized video to select time points that align with those audio peaks. These steps correspond to cognitive processes that can be performed entirely in the human mind.
The judicial exception is not integrated into a practical application. For example, the claim
recites the additional elements, “An electronic device comprising: memory for storing instructions; and at least one processor operably coupled to the memory”. These additional elements are recited at a high level of generality such that they amount to generic computer components to perform generic computer functions. Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly
more than the judicial expectation. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements are recited at a high-level of generality. It is therefore a judicial exception that is not integrated into a practical application, and does not include additional elements that are sufficient to amount to significantly more than the judicial exception. This claim is not patent eligible.
Claim 2 is rejected under 35 U.S.C. 101 because the claimed invention is directed to a further
limitation of the same abstract idea identified in the analysis of claim 1. For example, the person can observe peak values defined by a time window and apply thresholds. Further the additional element of “using a neural network” is recited at a high level of generality such that it amounts to merely implementing the abstract idea using a generic neural network. Accordingly, this additional element does not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. This claim is not patent eligible.
Claim 3 is rejected under 35 U.S.C. 101 because the claimed invention is directed to a further
limitation of the same abstract idea identified in the analysis of claim 1. For example, the person can choose the video signal to observe based on identifying objects corresponding to the motion. Further, the additional element of “using a second neural network different from the first neural network” is recited at a high level of generality such that it amounts to merely implementing the abstract idea using multiple generic neural networks. Accordingly, this additional element does not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. This claim is not patent eligible.
Claim 4 is rejected under 35 U.S.C. 101 because the claimed invention is directed to a further
limitation of the same abstract idea identified in the analysis of claim 1. For example, the person can observe and analyze frequency or amplitude from audio signals, such as by interpreting a waveform or spectrogram. This claim is not patent eligible.
Claim 7 is rejected under 35 U.S.C. 101 because the claimed invention is directed to a further
limitation of the same abstract idea identified in the analysis of claim 1. For example, the person can observe the video signal and select a particular screen to observe. Further, the additional element of “using a third neural network” is recited at a high level of generality such that it amounts to merely implementing the abstract idea using multiple generic neural networks. Accordingly, this additional element does not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. This claim is not patent eligible.
Claim 8 is rejected under 35 U.S.C. 101 because the claimed invention is directed to a further
limitation of the same abstract idea identified in the analysis of claim 1. For example, the person can analyze audio signals, such as sounds of a ball in contact with external objects, to correlate motions, such as throwing or catching the ball, in the video signal. This claim is not patent eligible.
Claim 9 is rejected under 35 U.S.C. 101 because the claimed invention is directed to a further
limitation of the same abstract idea identified in the analysis of claim 1. For example, the person can apply thresholds to a distribution of probabilities, identify values which exceed the threshold, and identify the largest value as a target time point. This claim is not patent eligible.
Claim 10 is rejected under 35 U.S.C. 101 because the claimed invention is directed to a further
limitation of the same abstract idea identified in the analysis of claim 1. For example, the person can apply thresholds to a distribution of probabilities, identify a value which exceeds the threshold, and determine that value is a target time point. This claim is not patent eligible.
Claim 11 contains limitations found analogous to that of claim 1. Therefore, claim 14 is rejected
for the same reason as claim 1.
Claims 12, 13, 14, and 17 contain limitations found analogous to that of claims 2, 3, 4, and 7,
respectively. Therefore, claims 12, 13, 14, and 17 are rejected for the same reasons as claims 2, 3, 4, and
7.
Claim 18 corresponds to claim 1, with the addition of “in response to identifying a time point having a value less than a threshold value within the audio signal, outputting information indicating that the identified time point is the sound time point when the designated motion is captured; and in response to identifying one or more time points having values above the threshold value within the audio signal and based on a video signal within different time intervals including the one or more time points, selecting a time point from among the one or more time points as the sound time point when the designated motion is captured.”. The claimed invention is directed to a further limitation of the same abstract idea identified in the analysis of claim 1. For example, the person can observe video and audio signals and apply thresholding to identify time points corresponding to motion. This claim is not patent eligible.
Claim 19 is rejected under 35 U.S.C. 101 because the claimed invention is directed to a further
limitation of the same abstract idea identified in the analysis of claim 1. For example, the person can analyze an audio signal to infer probabilities of a motion correlates to audio peaks in the signal and select different content based on this analysis. This claim is not patent eligible.
Claim 20 is rejected under 35 U.S.C. 101 because the claimed invention is directed to a further
limitation of the same abstract idea identified in the analysis of claim 1. For example, the person can analyze audio signals, such as sounds of a ball in contact with external objects, to correlate motions, such as throwing or catching the ball, in the video signal. This claim is not patent eligible.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 9, and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Stojancic et al. (US 20220180892 A1), (hereinafter Stojancic) in view of Sharma (US 11144764 B1).
Regarding claim 1, Stojancic teaches an electronic device comprising:
memory for storing instructions; and at least one processor operably coupled to the memory, wherein the at least one processor, when the instructions are executed, is configured to:
receive a request for detecting a sound time point in a multimedia content when a designated motion is captured (Stojancic, pg. 1, paragraph 0007, lines 1-9, “A system and method are presented to enable automatic real–time processing of audio signals extracted from sporting event television programming content, for detecting, selecting, and tracking short bursts of high–energy audio events, such as tennis ball hits in a tennis match.”, pg. 8, paragraphs 0119, “As the audiovisual streams are displayed, one or more components of system 100, such as client devices 106, web servers 102, application servers 114, and/or analytical servers 116, may analyze the audiovisual streams, identify highlights within the audiovisual streams, and/or extract metadata from the audiovisual stream, for example, from an audio component of the stream. This analysis may be carried out in response to receipt of a request to identify highlights and/or metadata for the audiovisual stream.”, pg. 8, paragraphs 0120-0121, “User preferences can also be extracted from storage, such as from user data 155 stored in one or more storage devices 153 , so as to customize analysis of audio data 154 without necessarily requiring user 150 to specify preferences… Additionally, or alternatively, user preferences can be retrieved from previously stored preferences that were explicitly provided by user 150. Such user preferences may indicate which teams, sports, players, and/or types of events are of interest to user 150, and / or they may indicate what type of metadata or other information related to highlights, would be of interest to user 150. Such preferences can therefore be used to guide analysis of the audiovisual stream to identify highlights and/or extract metadata for the highlights.”, The system receives requests by users to obtain video highlights corresponding to events. These events include designated motions such as a tennis ball hit or serve and are identified with respect to the timing of the video.); and
obtain the sound time point when the designated motion is captured from among respective time points corresponding to a plurality of peak values, using a video signal synchronized to an audio signal, in the multimedia content (Stojancic, pg. 6, paragraphs 0085-0087, “In at least one embodiment, an initial audio signal analysis is performed in the time domain, so as to detect short bursts of high-energy audio and generate of audio events representing potential exciting occurrences. An analyzing time window of a selected size may be used to compute an indicator of the average level of audio energy at overlapping window positions. Subsequently, a row event vector may be populated with indicator/position pairs. In at least one embodiment, time-domain detected audio events are revised by considering spectral characteristics of the audio signal in the neighborhood of audio events. A spectrogram may be constructed for the analyzed audio signal, and a 2-D diamond-shaped time–frequency area filtering process may be performed to detect and extract pronounced spectral magnitude peaks. A spectral event vector may be populated with magnitude and time-frequency coordinates for each selected peak. In at least one embodiment, one or more spectrogram time-spread range (s) are constructed around audio event time positions obtained in the time - domain analysis. By counting and recording spectral event vector peaks in a particular time spread range, an audio event qualifier may be established for each time-domain detected audio event. In at least one embodiment, audio event time positions having an audio event qualifier value below a certain threshold are accepted as viable audio event points, and any remaining audio event time positions are suppressed. In general, qualification of the time-domain detected audio events can be performed based on spectral analysis of each individual time range around a detected audio event, or it can be based on a spectral analysis of a combination of time ranges around a detected audio event.”, see Figs. 3A-3B, An audio signal, which is synchronized to the video data, is converted to the time domain. A sliding window is applied to identify peak values in a spectrogram which corresponds to the event, or in this case designated motion of a tennis hit. This process results in extracted highlights from the video at particular time points corresponding to the events.).
Stojancic does not teach obtain a distribution of probabilities in a time domain that the designated motion is performed based on an audio signal in the multimedia content, wherein the distribution of probabilities comprises a plurality of peak values corresponding to respective time points in the multimedia content.
However, Sharma teaches obtain a distribution of probabilities in a time domain that the designated motion is performed based on an audio signal in the multimedia content, wherein the distribution of probabilities comprises a plurality of peak values corresponding to respective time points in the multimedia content (Sharma, column 4, lines 55-67 and column 5, lines 1-22, “Image classification algorithm applying module 114 may be configured to apply a domain specific neural network image classification algorithm to the visual representation to generate interest probability scores for various portions of the audio signal. The domain specific neural network may be configured as a classifier which outputs probabilities for each of plural class labels. Labels can include significant words that identify corresponding video content, such as "goal", "penalty" or the like in the case of a sporting event. The spectrogram image is fed into image classification algorithm applying module 114 which outputs probabilities for each of the class labels, which for this can be bi-nomial (for example "goal" or "no-goal") or multi-nomial (for example goal, yellow-card, red-card, penalty etc.) If the model detects a label of interest the segment is returned, possibly with some buffer time before and/or after the portion in which the label is detected. Speech recognition module 124 can output the words that the commentators are speaking at the time. The labels predicted by module 114 are
string-compared to a predefined list of words of interest, such as "goal". If there's a match, then the probability score is increased. The data object returned from image classification algorithm applying module 114 (including optional 10 results from speech recognition module 124) contains the label and the start and end times of the video where it occurred. ”).
Stojancic teaches obtaining a sound time point which corresponds to a designated motion from synchronized video data by analyzing peak values of a spectrogram obtained from an audio signal (Stojancic, pg. 8, paragraphs 0120-0121 and pg. 6, paragraphs 0085-0087). Stojancic does not teach obtaining probabilities corresponding to the peak values to determine the sound time point. Sharma teaches calculating probability scores for peak values of spectrograms using a classification neural network to identify positions in a synchronized video (see above). Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to have modified spectrogram analysis of Stojancic to include the neural network for probability score calculation as taught by Sharma (Sharma, column 4, lines 55-67 and column 5, lines 1-22). The motivation for doing so would have been to not only identify peaks corresponding to high-energy audio bursts but also classify the peaks, thereby improving the accuracy of audio event detection. The combination of Stojancic in view of Sharma would apply the classification neural network of Sharma to the identified peaks of Stojancic to determine time points for designated motions, such as for tennis hits. Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine the teachings of Stojancic with Sharma to obtain the invention as specified in claim 1.
Regarding claim 9, Stojancic in view of Sharma teaches the electronic device of claim 1, wherein the at least one processor, when the instructions are executed, is further configured to: identify at least one value higher than a threshold value within the distribution of probabilities, from the video signal (Sharma, column 6, lines 25-33, “Selecting portions of the audio signal that may meet or exceed a threshold probability score may include, in addition to applying a domain specific neural network image classification algorithm as noted above, matching the text to words having a high probability of corresponding to video content of interest. As one simple example, the words "they score" in the soundtrack, could very well be indicative that the corresponding video segment is of interest.”, A neural network generates probability scores for portions of the audio signal and selects those portions which meet or exceed a threshold probability score. In the combination of Stojancic in view of Sharma, the probability scoring is applied to the synchronized video signal to identify video segments corresponding to audio portions of interest.);
identify, in the time domain, a largest value from the at least one value higher than the threshold value as a peak value; and obtain a time point corresponding to the identified peak value (Stojancic, pg. 6, paragraph 0102, “A spectrogram may be constructed for the analysis of audio signal in the frequency domain. A 2-D diamond-shaped spectrogram area filter may be constructed for detection and selection of pronounced time-frequency magnitude peaks. The area filter may be advanced along the time and frequency spectrogram axes, and at each time-frequency position, an area filter central peak magnitude may be checked against all remaining peak magnitudes. In at least one embodiment, the area filter central peak magnitude is retained only if it is greater than all other area filter peak magnitudes. The spectral event vector may be populated with all retained area filter central peak magnitudes.”, Spectrogram time-spread ranges are constructed around identified audio portions of interest. Within this range, central peak magnitudes, which are greater than all other peak magnitudes, are identified and associated with time indices used to define start/end times for video highlights that contain the designated motions.).
Claim 11 corresponds to claim 1, reciting a method for executing the steps according to claim 1. Stojancic in view of Sharma teaches a method for executing the steps according to claim 1 (Stojancic, pg. 6, paragraphs 0085-0088 and 0100-0104, see figs. 5-7). As indicated in the analysis of claim 1, Stojancic in view of Sharma teaches all the limitations according to claim 1. Therefore, claim 11 is rejected for the same reasons of obviousness as claim 1.
Claims 2, 4-6, 10, 12, and 14-16 are rejected under 35 U.S.C. 103 as being unpatentable over Stojancic et al. (US 20220180892 A1) in view of Sharma (US 11144764 B1) and further in view of Wood et al. (US 20230240621 A1), (hereinafter Wood).
Regarding claim 2, Stojancic in view of Sharma teaches the electronic device of claim 1, wherein at least one peak value from among the plurality of peak values is equal to a largest value, the largest value being a value from among a plurality of values included between a first time point and a second time point, the first time point having a first value and the second time point having a second value equal to a threshold value (Stojancic, pg. 6, paragraph 0087, “In at least one embodiment, one or more spectrogram time-spread range(s) are constructed around audio event time positions obtained in the time-domain analysis. By counting and recording spectral event vector peaks in a particular time spread range, an audio event qualifier may be established for each time-domain detected audio event. In at least one embodiment, audio event time positions having an audio event qualifier value below a certain threshold are accepted as viable audio event points, and any remaining audio event time positions are suppressed. In general, qualification of the time-domain detected audio events can be performed based on spectral analysis of each individual time range around a detected audio event, or it can be based on a spectral analysis of a combination of time ranges around a detected audio event.”, A spectrogram time-spread range is constructed around audio event time positions. This time-spread range would be defined by a threshold value, as it is bounded by start and end time points relative to a peak.), and
wherein the at least one processor, when the instructions are executed, is further configured to: obtain the distribution of probabilities in the time domain using probabilities where the plurality of peak values is identified, based on characteristic information and the audio signal, using a neural network (Sharma, column 4, lines 55-67 and column 5, lines 1-22, Domain-specific neural network is applied to the spectrogram to generate probability scores for portions of the audio signal. The combination of Stojancic in view of Sharma would apply this neural network to the time spread range and peak values identified by Stojancic to classify them and obtain the distribution of probabilities.).
Stojancic in view of Sharma does not teach the first time point having a first value and the second time point having a second value equal to a threshold value within the distribution of probabilities.
However, Wood teaches the first time point having a first value and the second time point having a second value equal to a threshold value within the distribution of probabilities (Wood, pgs. 3 and 4, paragraph 0095, “In overview, the method involves processing a digital audio recording to identify segments of the recording containing particular sound events of interest. The digital audio recording is processed according to a number of processes including filtering the digital audio recording, at box 11, and processing the filtered digital audio recording to produce a corresponding signal envelope, as indicated by dashed box 14. A statistical distribution, which is typically the Poisson distribution but which could be another statistical distribution such as the gamma distribution, is then fitted to the signal envelope, as indicated by dashed box 16. A threshold level is then determined, as indicated by dashed line 18, in respect of the signal envelope. The threshold level is determined based on the statistical distribution and a predetermined probability level. Segments of the signal envelope that are above the threshold level are then identified for example as start and finish times of each such segment, to thereby also identify corresponding segments of the digital audio recording that contain the particular sound events of interest. For example the sound events of interest may be sounds such as snoring, or wheezing or breathing sounds.”, pg. 5, paragraphs 0112-0114, “FIG. 6 shows the threshold level 56 superimposed on the second signal envelope (i.e. the power estimate signal) 47. It can be seen that the second signal envelope exceeds threshold in the following above-threshold segments, [tl,t2]; [t3,t4]; [t5,t6]; [t7,t8]; [t9,t10]; and [tll,t12]. A simple temporal filter is applied at box 37 to select above threshold segments that potentially correspond to sleep sounds, being events of interest, on the basis that the events must be of longer duration than 225 ms and less duration than 4 s.”, see Fig. 6, threshold level 56).
Stojancic in view of Sharma teaches defining a time-spread range for audio signal peak analysis (Stojancic, pg. 6, paragraph 0087) and classifying peaks into a distribution of probabilities using a classification neural network (Sharma, column 4, lines 55-67 and column 5, lines 1-22). Stojancic in view of Sharma does not teach applying thresholding within this distribution to define the time-spread range. Wood teaches assigning specific time points to segments of an audio signal by determining a threshold level from a probability distribution and identifying segments where the signal envelope crosses that threshold (see above). Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to have modified the system of Stojancic in view of Sharma to apply probability thresholding as taught by Wood (Wood, pgs. 3 and 4, paragraph 0095 and pg. 5, paragraphs 0112-0114), thereby defining a time-spread range based on a threshold value within the probability distribution. The motivation for doing so would have been to define a statistically dependent time range for analysis, resulting in improved accuracy of audio event detection. Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine the teachings of Stojancic in view of Sharma with Wood to obtain the invention as specified in claim 2.
Regarding claim 4, Stojancic in view of Sharma and further in view of Wood teaches the electronic device of claim 2, wherein the characteristic information is based on at least one of frequency or amplitude of the audio signal in the time domain (Sharma, column 4, lines 10-12, “FIG. 2 illustrates a graph 200 of an example of an audio signal in the frequency Domain, where the horizontal (x) axis is time and the vertical (y) axis is frequency in Hz.”, column 4, lines 64-67 and column 5, lines 1-4, “The spectrogram image is fed into image classification algorithm applying module 114 which outputs probabilities for each of the class labels, which for this can be bi-nomial (for example "goal" or "no-goal") or multi-nom ial (for example goal, yellow-card, red-card, penalty etc.) If the model detects a label of interest the segment is returned, possibly with some buffer time before and/or after the portion in which the label is detected.”).
Regarding claim 5, Stojancic in view of Sharma and further in view of Wood teaches the electronic device of claim 2, wherein the first time point is a time point when a slope of the distribution of probabilities is positive, and wherein the second time point is a time point when the slope of the distribution of probabilities is negative (Wood, pg. 5, paragraphs 0112-0113, “FIG. 6 shows the threshold level 56 superimposed on the second signal envelope (i.e. the power estimate signal) 47. [0113] It can be seen that the second signal envelope exceeds threshold in the following above-threshold segments, [tl,t2]; [t3,t4]; [t5,t6]; [t7,t8]; [t9,t10]; and [tll,t12].”, As illustrated in Fig. 6, at time point t1, the signal envelope crosses the threshold with a positive slope as the signal amplitude increases. Conversely, at time point t2, the signal envelop crosses the threshold with a negative slope as the signal amplitude decreases. As the probability distribution is fit to the signal envelope, a similar analysis applies. Note the claim does not require any determination or calculation of slope values for the distribution.).
Regarding claim 6, Stojancic in view of Sharma and further in view of Wood teaches the electronic device of claim 5, wherein the at least one processor, when the instructions are executed, is further configured to: obtain content different from the multimedia content, segmented from the video signal during time from the first time point to the second time point, and wherein the time includes the sound time point when the designated motion is captured (Stojancic, pg. 13, paragraph 0187, “In at least one embodiment, the automated video highlights and associated metadata generation application receives a live broadcast program, or a digital audiovisual stream via a computer server, and processes audio data 154 using digital signal processing techniques so as to detect high-energy audio associated with, for example, tennis ball hits and related tennis serve delivery in tennis games, as described above. These audio events may be sorted and selected using the techniques described herein. Extracted information may then be appended to metadata 224 associated with an event, such as a sporting event. Metadata 224 may be associated with the event television programming video highlights, and can be used, for example, to determine boundaries 232 (i.e., start and/or end times) for segments used in highlight generation.”, Time points, including a start and end time for segments of the video is stored as metadata. This data is used to extract video segments for motion events, such as tennis hits or serve, to generate video highlights.).
Regarding claim 10, Stojancic in view of Sharma and further in view of Wood teaches the electronic device of claim 2, wherein the at least one processor, when the instructions are executed, is further configured to: identify the at least one peak value exceeding the threshold value within the distribution of probabilities, from the video signal; and obtain a time point corresponding to the at least one peak value (Stojancic, pg. 6, paragraph 0102-0104, “A spectrogram may be constructed for the analysis of audio signal in the frequency domain. A 2-D diamond shaped spectrogram area filter may be constructed for detection and selection of pronounced time-frequency magnitude peaks. The area filter may be advanced along the time and frequency spectrogram axes, and at each time-frequency position, an area filter central peak magnitude may be checked against all remaining peak magnitudes. In at least one embodiment, the area filter central peak magnitude is retained only if it is greater than all other area filter peak magnitudes. The spectral event vector may be populated with all retained area filter central peak magnitudes… By subsequent suppression of undesirable, redundant audio events, a final desired audio event timeline for the game may be obtained.”, The combination Stojancic in view of Sharma and further in view of Wood defines time-spread ranges according to probability thresholding, retaining the audio signal values which exceed the threshold. These ranges of interest are then processed to determine peak values corresponding to motions in the video. Boundaries, a start and stop time, are then obtained corresponding to the peaks for video highlight generation.).
Claims 12, 14, 15, and 16, correspond to claims 2, 4, 5, and 6, respectively, reciting a method for executing the steps according to claims 2, 4, 5, and 6. Stojancic in view of Sharma and further in view of Wood teaches a method for executing the steps according to claims 2, 4, 5, and 6 (Stojancic, pg. 6, paragraphs 0085-0088 and 0100-0104, see figs. 5-7). As indicated in the analysis of claims 2, 4, 5, and 6, Stojancic in view of Sharma and further in view of Wood teaches all the limitations according to claims 2, 4, 5, and 6. Therefore, claims 12, 14, 15, and 16 are rejected for the same reasons of obviousness as claims 2, 4, 5, and 6.
Claims 3 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Stojancic et al. (US 20220180892 A1) in view of Sharma (US 11144764 B1) and further in view of Wood et al. (US 20230240621 A1) and Lee et al. (US 20190087661 A1), (hereinafter Lee).
Regarding claim 3, Stojancic in view of Sharma and further in view of Wood teaches the electronic device of claim 2, wherein the neural network is a first neural network (Sharma, column 4, lines 55-67 and column 5, lines 1-22, “Image classification algorithm applying module 114 may be configured to apply a domain specific neural network image classification algorithm to the visual representation to generate interest probability scores for various portions of the audio signal.”).
Stojancic in view of Sharma and further in view of Wood does not teach wherein the at least one processor, when the instructions are executed, is further configured to: obtain the video signal, based on identifying at least one of a trajectory of a ball, a position of a glove, home plate, or a strike zone, from the multimedia content, using a second neural network different from the first neural network.
However, Lee teaches wherein the at least one processor, when the instructions are executed, is further configured to: obtain the video signal, based on identifying at least one of a trajectory of a ball, a position of a glove, home plate, or a strike zone, from the multimedia content, using a second neural network different from the first neural network (Lee, pg. 9, paragraph 0112, “FIG . 16 is a flow diagram 1600 of a process for ball tracking, frame buffering, and initial shot attempt detection, according to some embodiments of the present invention. In this illustrative example where tracking may be viewed as performed in the forward direction, given input video frames 1610, one or more balls may be first detected in Step 1620, using one or more computer vision algorithms such as background subtraction, color histogram matching, convolutional neural networks and the like may be applied for ball detection. In Step 1624, one or more 2D ball trajectories may be identified by following the motion of the detected balls in air.”, see Fig. 16).
Stojancic in view of Sharma and further in view of Wood teaches applying a classification neural network to an audio signal to determine a probability distribution for designated motions like tennis hits (Sharma, columns 4 and 5, lines 55-67 and 1-22, respectively). Stojancic in view of Sharma and further in view of Wood does not teach implementing a second neural network to identify a trajectory of a ball, a position of gloves, a home plate, or a strike zone. Lee teaches using a distinct neural network to analyze video content to track a ball’s trajectory (see above). Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to have modified the system of Stojancic in view of Sharma and further in view of Wood to include a second neural network for ball trajectory identification as taught by lee (Lee, pg. 9, paragraph 0112, see Fig. 16). The motivation for doing so would have been to provide additional context, such as ball trajectories, for the video highlights, thereby improving their interpretability by the user. Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine the teachings of Stojancic in view of Sharma and further in view of Wood with Lee to obtain the invention as specified in claim 3.
Claim 13 corresponds to claim 3, reciting a method for executing the steps according to claim 3. Stojancic in view of Sharma and further in view of Wood and Lee teaches a method for executing the steps according to claim 3 (Stojancic, pg. 6, paragraphs 0085-0088 and 0100-0104, see figs. 5-7). As indicated in the analysis of claim 3, Stojancic in view of Sharma and further in view of Wood and Lee teaches all the limitations according to claim 3. Therefore, claim 13 is rejected for the same reasons of obviousness as claim 3.
Claims 7 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Stojancic et al. (US 20220180892 A1) in view of Sharma (US 11144764 B1) and further in view of Wood et al. (US 20230240621 A1), Lee et al. (US 20190087661 A1) and Shanmuga Vadivel et al. (US 20200211601 A1), (hereinafter Shanmuga Vadivel).
Regarding claim 7, Stojancic in view of Sharma and further in view of Wood and Lee teaches the electronic device of claim 3. Stojancic in view of Sharma and further in view of Wood and Lee does not teach wherein the at least one processor, when the instructions are executed, is further configured to: obtain at least one of a pitching screen or a catching screen from the multimedia content using a third neural network.
However, Shanmuga Vadivel teaches wherein the at least one processor, when the instructions are executed, is further configured to: obtain at least one of a pitching screen or a catching screen from the multimedia content using a third neural network (Shanmuga Vadivel, pg. 4, paragraph 0028, lines 1-10, The deep learning environment 101 may have access to a large volume of raw data and may be trained to recognize a set of rules (e.g., certain objects, features, and/or other detectable attributes ) associated with the raw data. For example, in some aspects, the deep learning environment 101 may be trained to recognize a pitching action (e.g., in baseball). During the training phase, the deep learning environment 101 may process or analyze a large number of images and/or videos that contain pitches, for example, from recorded baseball games.”, pg. 5, paragraphs 0046-0049, “The label generation module 234 may generate one or more event labels 202 indicating the location or position of each actionable event and/or restricted event in the content item 201… The media playback interface 250 is configured to render the content items 201 for display while providing a user interface through which the user may control, navigate, or otherwise manipulate playback of the content items 201 based, at least in part, on the event labels 202.”).
Stojancic in view of Sharma and further in view of Wood and Lee teaches applying a first classification neural network to an audio signal to determine a probability distribution for designated motions like tennis hits (Sharma, columns 4 and 5, lines 55-67 and 1-22, respectively) and applying a second neural network to track the trajectory of a ball (Lee, pg. 9, paragraph 0112, see Fig. 16). Stojancic in view of Sharma and further in view of Wood and Lee does not teach implementing a third neural network to obtain a pitching or catching screen. Shanmuga Vadivel teaches using a third neural network to detect events corresponding to pitching from video content and displaying pitching events as on a playback screen for users (see above). Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to have modified the system of Stojancic in view of Sharma and further in view of Wood and Lee to include a third neural network for pitching event detection and display as taught by Shanmuga Vadivel (Shanmuga Vadivel, pg. 4, paragraph 0028, lines 1-10 and pg. 5, paragraphs 0046-0049). The motivation for doing so would have been to additionally detect significant events which lack high-energy audio bursts, such as pitching events, thereby improving the accuracy and completeness of the video highlight generation. Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine the teachings of Stojancic in view of Sharma and further in view of Wood and Lee with Shanmuga Vadivel to obtain the invention as specified in claim 7.
Claim 17 corresponds to claim 7, reciting a method for executing the steps according to claim 7. Stojancic in view of Sharma and further in view of Wood, Lee and Shanmuga Vadivel teaches a method for executing the steps according to claim 7 (Stojancic, pg. 6, paragraphs 0085-0088 and 0100-0104, see figs. 5-7). As indicated in the analysis of claim 7, Stojancic in view of Sharma and further in view of Wood, Lee and Shanmuga teaches all the limitations according to claim 7. Therefore, claim 17 is rejected for the same reasons of obviousness as claim 7.
Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Stojancic et al. (US 20220180892 A1) in view of Sharma (US 11144764 B1) and further in view of Xiong et al. (US 20040167767 A1), (hereinafter Xiong).
Regarding claim 8, Stojancic in view of Sharma teaches the electronic device of claim 1, wherein at least one peak value from among the plurality of peak values corresponds to a time point when sound included in the video signal is captured, the sound being caused by contact of a ball with an external object (Stojancic, pg. 11, paragraph 0161, “In at least one embodiment, the system performs several stages of analysis of audio data 154 in both the time and time-frequency domains, so as to detect bursts of energy (i.e., audio volume) due to occurrences during an audiovisual program, such as a broadcast of a sporting event. One example of such a burst of high-energy audio is a tennis ball hit during the delivery of a tennis serve.”).
Stojancic in view of Sharma does not teach wherein the external object is one of a glove or a bat, and wherein the designated motion comprises a first motion of throwing the ball or a second motion of the ball contacting the glove or the bat.
However, Xiong teaches wherein the external object is one of a glove or a bat, and wherein the designated motion comprises a first motion of throwing the ball or a second motion of the ball contacting the glove or the bat (Xiong, pg. 1, paragraph 0016, lines 1-7, “FIG. 1 shows a system and method 100 for extracting highlights from an audio signal of a Sports Video according to our invention. The system 100 includes a background noise detector 110, a feature extractor 130, a classifier 140, a grouper 150 and a highlight selector 160. The classifier uses Six audio classes 135, i.e., applause, cheering, ball hit, Speech, music, Speech with music.”, pg. 2, paragraphs 0034-0035, “In the audio domain, there are common events relating to highlights across different Sports. After an interesting event, e.g., a long drive in golf, a hit in baseball or an exciting Soccer attack, the audience shows appreciation by applauding or even loud cheering. A ball hit segment preceded or followed by cheering or applause can indicate an interesting highlight. The duration of applause or cheering is longer when an event is more interesting, e.g., a home-run in baseball.”).
Stojancic in view of Sharma teaches identifying peaks in an audio signal corresponding to a ball contacting an external object, such as a tennis racket in a tennis serve (Stojancic, pg. 11, paragraph 0161). Stojancic in view of Sharma does not teach identifying peaks corresponding to a ball contacting a bat. Xiong teaches analyzing audio signals to identify video segments corresponding to a ball being hit by a bat (see above). Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to have modified the system of Stojancic in view of Sharma to include video highlights corresponding to bat hits as taught Xiong (Xiong, pg. 1, paragraph 0016, lines 1-7 and pg. 2, paragraphs 0034-0035). The motivation for doing so would have been to detect a broader range of sporting events beyond tennis, thereby improving the systems application across multiple sport domains. Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine the teachings of Stojancic in view of Sharma with Xiong to obtain the invention as specified in claim 8.
Allowable Subject Matter
Claims 18-20 are rejected under 35 USC 112(b) and 101. These claims would be allowable if rewritten to overcome the above rejections.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Han et al (US10129608B2) teaches detecting video highlights of sport videos by classifying audio signals using a neural network. Bosi ("Audio-video techniques for the analysis of players behaviour in Badminton matches.", 2020) teaches analyzing sports videos to detect player shots by filtering peaks in an audio signal.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CONNOR LEVI HANSEN whose telephone number is (703)756-5533. The examiner can normally be reached Monday-Friday 9:00-5:00 (ET).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Sumati Lefkowitz can be reached at (571) 272-3638. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/CONNOR L HANSEN/Examiner, Art Unit 2672
/SUMATI LEFKOWITZ/Supervisory Patent Examiner, Art Unit 2672