DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Response to Amendment
This Office Action is responsive to the amendment filed 07/01/2025 (“Amendment”). Claims 1-16 and 20-22 are currently under consideration. The Office acknowledges the amendments to claims 1, 7, 14-16, and 20, as well as the cancellation of claims 17-19 and the addition of new claim 22.
The objection(s) to the drawings, specification, and/or claims, the interpretation(s) under 35 USC 112(f), and/or the rejection(s) under 35 USC 101 and/or 35 USC 112 not reproduced below has/have been withdrawn in view of the corresponding amendments.
Specification
The lengthy specification has not been checked to the extent necessary to determine the presence of all possible minor errors. Applicant’s cooperation is requested in correcting any errors of which applicant may become aware in the specification.
Claim Objections
Claim 20 is objected to because of the following informalities: the recitation of “based an analysis” should instead read –based on an analysis--. Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-16 and 20-22 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 of the subject matter eligibility test (see MPEP 2106.03).
Claims 1-15 and 22 are directed to a “method,” which describes one of the four statutory categories of patentable subject matter, i.e., a process. Claims 16, 20, and 21 are directed to a “non-transitory medium,” which describes one of the four statutory categories of patentable subject matter, i.e., a machine.
Step 2A of the subject matter eligibility test (see MPEP 2106.04).
Prong One: Claim 1 recites (“sets forth” or “describes”) the abstract idea of a mathematical concept, substantially as follows:
processing the audio data by - applying a high-pass filter to the audio data, segmenting the audio data into multiple segments of equal duration, generating, based on the audio data, (i) a spectrogram, (ii) Mel-frequency cepstral coefficients (MFCCS) within a predetermined frequency range, and (iii) multiple sets of values that are representative of energy summed across multiple frequency bands of the spectrogram, wherein each set includes a separate value for each of the multiple frequency bands and is associated with a corresponding one of the multiple segments, and concatenating the spectrogram, the MFCCs, and the multiple sets of values into a matrix; applying, to the matrix, a neural network that is trained to produce, as output, a vector that includes multiple entries arranged in temporal order, wherein each entry in the vector indicates whether a corresponding one of the multiple segments is representative of a breathing event, and wherein the neural network includes (i) at least one recurrent layer and (ii) at least one fully connected layer that employes a sigmoid function as an activation function to produce the vector; identifying (i) a first breathing event, (ii) a second breathing event that follows the first breathing event, and (iii) a third breathing event that follows the second breathing event by examining the vector; determining (i) a first period between the first and second breathing events and (ii) a second period between the second and third breathing events; and computing a respiratory rate based on the first and second periods.
Each of the steps involve the mathematical concepts of filtering, segmenting, calculations to obtain different values, concatenation, processing via layers of a neural network, and identification of features and periods. These steps correspond to “[w]ords used in a claim operating on data to solve a problem [that] can serve the same purpose as a formula.” See MPEP 2106.04(a)(2)(I).
Claims 7 and 16 recite a subset of these features and are therefore also directed to a mathematical concept. The setting of values to one or zero based on thresholds, and the identification based on a plurality of consecutive entries also involve the mathematical concepts of binarization and trending. The application of a neural network to the data structure to produce a second data structure involves the mathematical concept of processing via a neural network.
Further, claim 7 recites the abstract idea of a mental process, substantially as follows:
generating (i) a spectrogram based on the audio data, and (ii) a series of values that are representative of energy summed across different frequency bands of the spectrogram; applying, to the spectrogram and the series of values, a neural network that produces, as output, a data structure that includes entries arranged in temporal order, wherein each entry in the data structure indicates whether a corresponding segment of the audio data is representative of a breathing event; for each entry in the data structure, setting that entry to a value of one in response to a determination that a corresponding value exceeds a threshold; and setting that entry to a value of zero in response to a determination that a corresponding value does not exceed the threshold; analyzing the data structure to identify (i) a first plurality of consecutive entries having values of one that are collectively representative of a first breathing event, (ii) a second plurality of consecutive entries having values of one that are collectively representative of a second breathing event that follows the first breathing event, and (iii) a third plurality of consecutive entries having values of one that are collectively representative of a third breathing event that follows the second breathing event; determining (i) a first period based on the first and second breathing events, and (i) a second period based on the second and third breathing events; computing a respiratory rate based on the first and second periods.
Each of the steps can be practically performed in the human mind, with the aid of a pen and paper, but for performance on a generic computer, in a computer environment, or merely using the computer as a tool to perform the steps. If a person were to see a printout of e.g. the audio data, they would be able to generate a spectrogram based thereon, calculate energy values, obtain a data structure, set values for the data structure, and analyze the data structure for features from which calculations could be made. There is nothing to suggest an undue level of complexity in these steps, as a spectrogram can be obtained via a mathematical manipulation. Therefore, a person would be able to perform the calculations mentally or with pen and paper.
Claim 16 recites a subset of these features and is therefore also directed to a mental process.
Prong Two: Claims 1, 7, and 16 do not include additional elements that integrate the mental process or mathematical concept into a practical application. Therefore, the claims are “directed to” the mental process or mathematical concept. The additional elements merely:
recite the words “apply it” (or an equivalent) with the judicial exception, or include instructions to implement the abstract idea on a computer, or merely use the computer as a tool to perform the abstract idea (e.g. a non-transitory medium; applying a neural network and obtaining an output is merely using a computer as a tool to perform a “black box” processing function), and
add insignificant extra-solution activity (the pre-solution activity of: obtaining audio data or a data structure; and the post-solution activity of: causing display, using generic data-outputting components (e.g. an interface)).
As a whole, the additional elements merely serve to gather and feed information to the abstract idea, while generically implementing it on a computer. There is no practical application because the abstract idea is not applied, relied on, or used in a meaningful way. Nothing is done with the computed respiratory rate, and mere display is insufficient because the data need not be seen or acted on. No improvement to the technology is evident. Therefore, the additional elements, alone or in combination, do not integrate the abstract idea into a practical application.
Step 2B of the subject matter eligibility test (see MPEP 2106.05).
Claims 1, 7, and 16 do not include additional elements, alone or in combination, that are sufficient to amount to significantly more than the judicial exception (i.e., an inventive concept) for the same reasons as described above.
Dependent Claims
The dependent claims merely further define the abstract idea and are, therefore, directed to an abstract idea for similar reasons: they merely
further describe the abstract idea (e.g. extracting particular coefficients (claim 2), using a particular frequency range (claim 3), performing min-max normalization (claim 4), filtering (claim 5), applying a transform (claim 6), generating MFCCs (claim 8), segmenting (claims 9 and 10), generating multiple sets (claim 11), details of the periods (claim 13), details of the neural network (claim 15), analyzing entries (claims 20 and 21), etc.), and
further describe the pre-solution activity (or the structure used for such activity) (e.g. obtaining data from an electronic stethoscope connected to a hub (claim 12), training the neural network (claim 14), etc.).
Taken alone and in combination, the additional elements do not integrate the judicial exception into a practical application at least because the abstract idea is not applied, relied on, or used in a meaningful way. They also do not add anything significantly more than the abstract idea. Their collective functions merely provide computer/electronic implementation and processing, and no additional elements beyond those of the abstract idea. Looking at the limitations as an ordered combination adds nothing that is not already present when looking at the elements individually. There is no indication that the combination of elements improves the functioning of a computer, output device, improves another technology or technical field, etc. Therefore, the claims are rejected as being directed to non-statutory subject matter.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 2, 4-6, and 22 are rejected under 35 U.S.C. 103 as being unpatentable over US Patent Application Publication 2019/0088367 (“Stamatopoulos”) in view of US Patent Application Publication 2018/0085068 (“Telfort”), US Patent Application Publication 2019/0272845 (“Hasan”), US Patent Application Publication 2011/0075851 (“LeBoeuf”), and US Patent Application Publication 2020/0152330 (“Anushiravani”).
Regarding claim 1, Stamatopoulos teaches [a] method for computing a respiratory rate ("Metrics can come directly from breath cycle characteristics and statistics transformation and new metrics can be constructed by the combination of more than one characteristic (e.g., where breath phase duration, respiratory rate and breath intensity are used to obtain respiratory depth). Metrics can be provided for one
breath cycle or for a number of breath cycles.", para [0112]), the method comprising: obtaining audio data that is representative of a recording of sounds generated by the lungs of a patient ("The recording device can, for example, be a smart phone, a spirometer with a microphone (as will be discussed further below), a stethoscope, or a CPAP machine with a microphone.", para [0273]); processing the audio data by - … segmenting the audio data into multiple segments of equal duration (¶ 0305, frames of equal duration – also see ¶¶s 0010, 0145, 0147, suggesting that blocks are also segments of equal duration), generating, based on the audio data, (i) a spectrogram (Abstract, ¶ 0014, etc.), …; applying, to the [data], a neural network that is trained to produce, as output, a vector that includes multiple entries arranged in temporal order (Abstract, artificial neural network; "The main process module (MPM) at step 1515 performs breath cycle separation which is done by event grouping. The MPM module processes a sequence vector with the event types, and outputs a breath cycle vector containing the map of all the events.", para [0179]), wherein each entry in the vector indicates whether a corresponding one of the multiple segments is representative of a breathing event ("Based on this separation, the MPM module performs a calculation of the full set of metrics by integrating auxiliary vectors related to the breath intensity, the wheeze, etc. into the breath cycle mapping.", para [0179]), …; identifying (i) a first breathing event, (ii) a second breathing event that follows the first breathing event, and (iii) a third breathing event that follows the second breathing event by examining the vector ([detecting all breathing events in a vector sequentially] "The main process module (MPM) at step 1515 performs breath cycle separation which is done by event grouping. The MPM module processes a sequence vector with the event types, and outputs a breath cycle vector containing the map of all the events.", para [0179]); determining (i) a first period between the first and second breathing events and (ii) a second period between the second and third breathing events ("Inhalation/Exhalation Ratio (IER): This is the ratio of the duration of the inhalation versus exhalation. These durations and their connection can help to extract conclusions about the breath patterns, specially concerning the physical state of the user. Other ratios can also be extracted such as the time of any one phase over the time of the total breath cycle. For example, the time of inhalation in relation to the time of the total breath cycle (Ti/Ttotal).", para [0186]); and computing a respiratory rate based on the first and second periods ("As discussed above, one embodiment of the present invention can be used to determine VT (T1) and RCT (T2). In a different embodiment, VT and RCT calculations can be made within the classifier core module 730 itself. Processes such as respiratory rate tracking and breath phase tracking and detection are important in the analysis as the final result is not only based on the overall breath sound statistics, but also on statistics that come from the analysis of each breath cycle as the breathing session progresses over time (e.g. inhalation intensity tracking).", para [0194]).
Stamatopoulos does not appear to explicitly teach processing the audio data by applying a high-pass filter to the audio data.
Telfort teaches high pass filtering an acoustic sensor signal (audio data) and then using the filtered signal to calculated respiratory rate (Fig. 3A, ¶ 0057).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate a high pass filtering function into Stamatopoulos, as in Telfort, for the purpose of separating desired and undesired sounds (Telfort: ¶¶s 0056, 0057, etc.).
Stamatopoulos-Telfort does not appear to explicitly teach processing the audio data by generating, based on the audio data, … (ii) Mel-frequency cepstral coefficients (MFCCs) within a predetermined frequency range, and (iii) multiple sets of values that are representative of energy summed across multiple frequency bands of the spectrogram, wherein each set includes a separate value for each of the multiple frequency bands and is associated with a corresponding one of the multiple segments, and concatenating the spectrogram, the MFCCs, and the multiple sets of values into a matrix (although ¶ 0214 of Stamatopoulos does describe frame concatenation, and ¶ 0325 describes calculating the energy of each frame).
Hasan teaches using audio data (Abstract, audio input unit) to generate a spectrogram (Fig. 9, ¶ 0102) and a series of MFCCs (¶ 0104) which are concatenated into a feature vector (¶ 0104). Hasan also describes using energy-based features in the vector (¶ 0104 - also see Table 1: energy entropy, short-term energy, etc.).
LeBoeuf teaches extracting features from an audio signal (¶ 0026), the features including e.g. energy (which is the sum of squared amplitudes within certain frequency bins) in various spectral bands, (¶ 0041). The features form a feature vector (¶ 0007).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate into the feature extraction of Stamatopoulos (Abstract, a spectrogram and a plurality of descriptors) the MFCC extraction (Hasan) and series of values representative of energy summed across different frequency bands extraction (LeBoeuf), concatenating all of these into a feature matrix (e.g. Hasan: ¶ 0104) that is provided to the model of Stamatopoulos, for the purpose of improving classification of the audio data with additional features (Hasan: ¶ 0104, Table 1, Fig. 6, step 320, etc. - also see ¶ 0068, describing MFCCs as commonly used to enable interpretation of frequency differences; LeBoeuf: ¶¶s 0026-0042 and Title, describing the large variety of features that may be used for audio recognition - also see Fig. 2).
Stamatopoulos-Telfort-Hasan-LeBoeuf does not appear to explicitly teach wherein the neural network includes (i) at least one recurrent layer and (ii) at least one fully connected layer that employs a sigmoid function as an activation function to produce the vector.
Anushiravani teaches using a neural network that includes at least one recurrent layer (¶¶s 0102, 0125, etc.) and at least one fully connected layer (¶ 0097) that employs a sigmoid function to produce an output (¶ 0108).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate recurrent and fully connected layers that use a sigmoid function into the neural network of the combination, as in Anushiravani, for the purpose of being able to use a time series of data for context (Anushiravani: ¶ 0102), as a known type of final layer (Anushiravani: ¶ 0097), as a known type of activation function to produce an output based on a simple substitution of parts with predictable results (Anushiravani: ¶ 0108), based on the suitability of these layers and functions for the purpose, and to implement a suitable AI diagnostic system (Anushiravani: ¶ 0004).
Regarding claim 2, Stamatopoulos-Telfort-Hasan-LeBoeuf-Anushiravani teaches the method of claim 1. Stamatopoulos-Telfort-Hasan-LeBoeuf-Anushiravani further teaches wherein for the MFCCs, a predetermined number of static coefficients, delta coefficients, and acceleration coefficients are extracted from the audio data (Hasan: ¶¶s 0068, 0104, Table 1, etc.).
Regarding claim 4, Stamatopoulos-Telfort-Hasan-LeBoeuf-Anushiravani teaches the method of claim 1. Stamatopoulos-Telfort-Hasan-LeBoeuf-Anushiravani further teaches performing min-max normalization on the matrix so that each entry has a value between zero and one (Stamatopoulos: ¶¶s 0312, 0345, 0346, etc. describe normalization, including matrix normalization, which would have been obvious to apply to the feature matrix of the combination for the purpose of making classification easier (Stamatopoulos: ¶ 0345, giving the same scale to all features and making certain features more visible)).
Regarding claim 5, Stamatopoulos-Telfort-Hasan-LeBoeuf-Anushiravani teaches the method of claim 1. Stamatopoulos-Telfort-Hasan-LeBoeuf-Anushiravani further teaches wherein the high-pass filter has a cut-off frequency that is sufficient to filter sounds made by an organ other than the lungs (Telfort: ¶ 0057, a cut-off frequency of 100 Hz filters sounds of e.g. the heart as described in Applicant’s specification as filed at ¶ 0100).
Regarding claim 6, Stamatopoulos-Telfort-Hasan-LeBoeuf-Anushiravani teaches the method of claim 1. Stamatopoulos-Telfort-Hasan-LeBoeuf-Anushiravani further teaches applying a Short-Time Fourier Transform to the audio data to generate the spectrogram (LeBoeuf: ¶ 0042).
Regarding claim 22, Stamatopoulos-Telfort-Hasan-LeBoeuf-Anushiravani teaches the method of claim 1. Stamatopoulos-Telfort-Hasan-LeBoeuf-Anushiravani further teaches for each entry in the vector for which a corresponding value exceeds a threshold, setting that entry to a value of one; and for each entry in the vector for which a corresponding value does not exceed the threshold, setting that entry to a value of zero (Anushiravani: ¶ 0108, using the sigmoid function to output values of 0 or 1. It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply the sigmoid function to the entries in the vector to enable binary classification as described); and wherein the first breathing event is represented by a first plurality of consecutive entries having values of one, the second breathing event is represented by a second plurality of consecutive entries having values of one, and the third breathing event is represented by a third plurality of consecutive entries having values of one (inherent based on the segmenting of data. I.e., multiple segments (or frames or blocks) correspond to one breathing event, and those multiple segments would have been classified by the sigmoid function to produce consecutive entries of e.g. 1. It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to classify consecutive segments for the purpose of accounting for artifacts (e.g. when the results are not consecutive, they may be artifacts) – see e.g. ¶ 0179 of Stamatopoulos, describing event grouping, or ¶¶s 0315, 0333, etc.).
Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Stamatopoulos-Telfort-Hasan-LeBoeuf-Anushiravani in view of US Patent Application Publication 2011/0125044 (“Rhee”).
Regarding claim 3, Stamatopoulos-Telfort-Hasan-LeBoeuf-Anushiravani teaches the method of claim 1. Stamatopoulos-Telfort-Hasan-LeBoeuf-Anushiravani does not appear to explicitly teach wherein the multiple frequency bands collectively span a range of 0-1,000 hertz.
Rhee teaches evaluating a frequency range of about 0-800 hertz for respiratory disease monitoring (Title, ¶¶s 0065, 0068, Fig. 6).
The frequency range for respiratory evaluation is a known results-effective variable because it can be changed as desired to monitor different aspects of breathing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to evaluate the frequency range of Rhee (which is used for wheezing detection), expanded to e.g. 0-1000 hertz, since it has been held that where the general conditions of a claim are disclosed in the prior art, discovering the optimum or workable ranges through routine experimentation is not inventive. In re Aller, 220 F.2d 454, 456, 105 USPQ 233, 235 (CCPA 1955). Further, it would have been obvious to detect tension or other features in a higher frequency range (Stamatopoulos: ¶¶s 0155, 0156 - 3,000+ Hz), since they are also relevant to wheezing detection (Stamatopoulos: ¶¶s 0131, 0155, 0156). Thus, a frequency range 0-3,000+ Hz would have been obvious to evaluate.
Claims 7-15 are rejected under 35 U.S.C. 103 as being unpatentable over US Patent Application Publication 2019/0088367 (“Stamatopoulos”) in view of US Patent Application Publication 2011/0075851 (“LeBoeuf”) and US Patent Application Publication 2020/0152330 (“Anushiravani”).
Regarding claim 7, Stamatopoulos teaches [a] method comprising: obtaining, from a source in near real time, audio data that is representative of a recording of sounds generated by the lungs of a living body ("The recording device can, for example, be a smart phone, a spirometer with a microphone (as will be discussed further below), a stethoscope, or a CPAP machine with a microphone.", paras [0273-0275] – the microphone data is received and recorded in real time); and in response to said obtaining, generating (i) a spectrogram based on the audio data (Abstract, ¶ 0014, etc.), and (ii) a series of values that are representative of energy … (¶ 0325); applying, to the spectrogram and the series of values, a neural network that produces, as output, a data structure that includes entries arranged in temporal order (Abstract, artificial neural network; "The main process module (MPM) at step 1515 performs breath cycle separation which is done by event grouping. The MPM module processes a sequence vector with the event types, and outputs a breath cycle vector containing the map of all the events.", para [0179]), wherein each entry in the data structure indicates whether a corresponding segment of the audio data is representative of a breathing event ("Based on this separation, the MPM module performs a calculation of the full set of metrics by integrating auxiliary vectors related to the breath intensity, the wheeze, etc. into the breath cycle mapping.", para [0179]); …; analyzing the data structure to identify (i) … a first breathing event, (ii) … a second breathing event that follows the first breathing event, and (iii) … a third breathing event that follows the second breathing event ([detecting all breathing events in a vector sequentially] "The main process module (MPM) at step 1515 performs breath cycle separation which is done by event grouping. The MPM module processes a sequence vector with the event types, and outputs a breath cycle vector containing the map of all the events.", para [0179]); determining (i) a first period based on the first and second breathing events, and (ii) a second period based on the second and third breathing events ("Inhalation/Exhalation Ratio (IER): This is the ratio of the duration of the inhalation versus exhalation. These durations and their connection can help to extract conclusions about the breath patterns, specially concerning the physical state of the user. Other ratios can also be extracted such as the time of any one phase over the time of the total breath cycle. For example, the time of inhalation in relation to the time of the total breath cycle (Ti/Ttotal).", para [0186]); computing a respiratory rate based on the first and second periods ("As discussed above, one embodiment of the present invention can be used to determine VT (T1) and RCT (T2). In a different embodiment, VT and RCT calculations can be made within the classifier core module 730 itself. Processes such as respiratory rate tracking and breath phase tracking and detection are important in the analysis as the final result is not only based on the overall breath sound statistics, but also on statistics that come from the analysis of each breath cycle as the breathing session progresses over time (e.g. inhalation intensity tracking).", para [0194]); and causing display of the respiratory rate on an interface (Fig. 19, ¶ 0260).
Stamatopoulos does not appear to explicitly teach generating a series of values that are representative of energy summed across different frequency bands of the spectrogram (although ¶ 0325 describes calculating the energy of each frame).
LeBoeuf teaches extracting features from an audio signal (¶ 0026), the features including e.g. energy (which is the sum of squared amplitudes within certain frequency bins) in various spectral bands, (¶ 0041). The features form a feature vector (¶ 0007).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate into the feature extraction of Stamatopoulos (Abstract, a spectrogram and a plurality of descriptors) the generation of a series of values representative of energy summed across different frequency bands extraction (LeBoeuf), for the purpose of improving classification of the audio data with additional features (LeBoeuf: ¶¶s 0026-0042 and Title, describing the large variety of features that may be used for audio recognition - also see Fig. 2).
Stamatopoulos-LeBoeuf does not appear to explicitly teach for each entry in the data structure, setting that entry to a value of one in response to a determination that a corresponding value exceeds a threshold; and setting that entry to a value of zero in response to a determination that a corresponding value does not exceed the threshold, and identifying first, second, and third pluralities of consecutive entries having values of one to collectively represent the first, second, and third breathing events.
Anushiravani teaches employing a sigmoid function to produce an output that includes values of 0 or 1 (¶ 0108).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to employ a sigmoid function in the neural network of the combination, as in Anushiravani, as a known type of activation function to produce an output based on a simple substitution of parts with predictable results (Anushiravani: ¶ 0108), and for the purpose of implementing a suitable AI diagnostic system (Anushiravani: ¶ 0004). It would have been obvious to apply the sigmoid function to the entries in the vector of the combination to enable binary classification as described. Further, based on the segmenting of data in the combination (i.e., multiple segments (or frames or blocks) corresponding to one breathing event), multiple segments would have been classified by the sigmoid function to produce consecutive entries of e.g. 1. It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to classify consecutive segments for the purpose of accounting for artifacts (e.g. when the results are not consecutive, they may be artifacts) – see e.g. ¶ 0179 of Stamatopoulos, describing event grouping, or ¶¶s 0315, 0333, etc.).
Regarding claim 8, Stamatopoulos-LeBoeuf-Anushiravani teaches the method of claim 7. Stamatopoulos-LeBoeuf-Anushiravani further teaches generating, based on the audio data, Mel-frequency cepstral coefficients (MFCCs) within a predetermined frequency range (LeBoeuf teaches extracting features from an audio signal (¶ 0026), the features including e.g. energy (which is the sum of squared amplitudes within certain frequency bins) in various spectral bands (¶ 0041) as well as MFCCs and other features (¶¶s 0042, 0048-0055, etc.). The features form a feature vector (¶ 0007). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate into the feature extraction of Stamatopoulos (Abstract, a spectrogram and a plurality of descriptors) the MFCC extraction of LeBoeuf, for the purpose of improving classification of the audio data with additional features (LeBoeuf: ¶¶s 0026-0042 and Title, describing the large variety of features that may be used for audio recognition - also see Fig. 2)).
Regarding claims 9 and 10, Stamatopoulos-LeBoeuf-Anushiravani teaches the method of claim 7. Stamatopoulos-LeBoeuf-Anushiravani further teaches segmenting the audio data into multiple segments, wherein the multiple segments are equal duration (Stamatopoulos: ¶ 0305, frames of equal duration – also see ¶¶s 0010, 0145, 0147, suggesting that blocks are also segments of equal duration).
Regarding claim 11, Stamatopoulos-LeBoeuf-Anushiravani teaches the method of claim 9. Stamatopoulos-LeBoeuf-Anushiravani further teaches wherein the series of values include multiple sets, each of which includes a separate value for each of the different frequency bands associated with a corresponding one of the multiple sets (Stamatopoulos: ¶ 0325 describes calculating the energy of each frame; LeBoeuf: ¶ 0041 describes calculating energy in various spectral bands. It would have been obvious to calculate energy of each frame, including across frequency bands, for the purpose of obtaining additional statistical features that may be used for audio recognition (LeBoeuf: ¶¶s 0026-0042 and Title, describing the large variety of features that may be used for audio recognition - also see Fig. 2)).
Regarding claim 12, Stamatopoulos-LeBoeuf-Anushiravani teaches the method of claim 7. Stamatopoulos-LeBoeuf-Anushiravani further teaches wherein the source is an electronic stethoscope system that includes one or more input units that are secured to the living body and communicatively connected to a hub unit (Stamatopoulos: ¶ 0273, a stethoscope that acts as a recording device; ¶¶s 0268 and 0269, a biometric patch having a microphone, Fig. 4, etc.; ¶ 0270, sending the data to a computing device for classification and tracking).
Regarding claim 13, Stamatopoulos-LeBoeuf-Anushiravani teaches the method of claim 7. Stamatopoulos-LeBoeuf-Anushiravani further teaches wherein the first period extends from a beginning of the first breathing event to a beginning of the second breathing event, and wherein the second period extends from the beginning of the second breathing event to a beginning of the third breathing event (Stamatopoulos: ¶ 0186 - "Inhalation/Exhalation Ratio (IER): This is the ratio of the duration of the inhalation versus exhalation. These durations and their connection can help to extract conclusions about the breath patterns, specially concerning the physical state of the user.” – an exhalation begins when an inhalation ends. Thus, the periods are as claimed).
Regarding claim 14, Stamatopoulos-LeBoeuf-Anushiravani teaches the method of claim 7. Stamatopoulos-LeBoeuf-Anushiravani further teaches wherein as part of a training operation, the neural network is trained to identify inhalation phases and exhalation phases of breathing events (Stamatopoulos: ¶¶s 0282, 0295, etc. describe machine learning; ¶¶s 0116, 0179, 0186, 0189, 0194, 0220, etc., describe identification of breathing phases; the Abstract, ¶ 0017, etc. describe training the model).
Regarding claim 15, Stamatopoulos-LeBoeuf-Anushiravani teaches the method of claim 7. Stamatopoulos-LeBoeuf-Anushiravani further teaches wherein the neural network (Stamatopoulos: Abstract, ¶¶s 0017-0019, 0282, etc.) outputs, for each segment of the audio data, a prediction that is made independent of preceding predictions (Stamatopoulos: ¶¶s 0109, 0149, 0158, 0159, classifying each block based on its own features).
Claims 16, 20, and 21 are rejected under 35 U.S.C. 103 as being unpatentable over US Patent Application Publication 2019/0088367 (“Stamatopoulos”) in view of US Patent Application Publication 2011/0075851 (“LeBoeuf”).
Regarding claim 16, Stamatopoulos teaches [a] non-transitory medium with instructions stored thereon that, when executed by a processor of a computing device, cause the computing device to perform operations (¶ 0456) comprising: obtaining ("The recording device can, for example, be a smart phone, a spirometer with a microphone (as will be discussed further below), a stethoscope, or a CPAP machine with a microphone.", paras [0273-0275] – the microphone data is received and recorded in real time) a first data structure that includes (i) a spectrogram associated with a recording of sounds generated by the lungs of a living body (Abstract, ¶ 0014, etc.) and (ii) values that are representative of energy … (¶ 0325); applying, to the data structure, a neural network that is trained to produce, as output, a second data structure that includes multiple entries arranged in temporal order (Abstract, artificial neural network; "The main process module (MPM) at step 1515 performs breath cycle separation which is done by event grouping. The MPM module processes a sequence vector with the event types, and outputs a breath cycle vector containing the map of all the events.", para [0179]), wherein each of the multiple entries in the second data structure indicates whether a corresponding segment of the recording is representative of a breathing event ("Based on this separation, the MPM module performs a calculation of the full set of metrics by integrating auxiliary vectors related to the breath intensity, the wheeze, etc. into the breath cycle mapping.", para [0179] – also see ¶¶s 0282, 0294, 0295, etc. describing machine learning classification based on descriptors that include spectrograms, and energy values based on the combination as described below), identifying, based on an analysis of the second data structure, (i) a first subset of the multiple entries that are collectively representative of a first breathing event, (ii) a second subset of the multiple entries that are collectively representative of a second breathing event that follows the first breathing event, and (iii) a third subset of the multiple entries that are collectively representative of a third breathing event that follows the second breathing event (based on the segmenting of data as described (i.e., multiple segments (or frames or blocks) corresponding to one breathing event), a group of segments corresponding to the multiple entries would have been classified together. It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to classify consecutive segments for the purpose of accounting for artifacts (e.g. when the results are not consecutive, they may be artifacts) – see e.g. ¶ 0179, describing event grouping, or ¶¶s 0315, 0333, etc.); [detecting all breathing events in a vector sequentially] "The main process module (MPM) at step 1515 performs breath cycle separation which is done by event grouping. The MPM module processes a sequence vector with the event types, and outputs a breath cycle vector containing the map of all the events.", para [0179] – also see ¶¶s 0116, 0179, 0186, 0189, 0194, 0220, etc., describing identification of breathing phases); determining (i) a first period based on the first and second breathing events, and (ii) a second period based on the second and third breathing events ("Inhalation/Exhalation Ratio (IER): This is the ratio of the duration of the inhalation versus exhalation. These durations and their connection can help to extract conclusions about the breath patterns, specially concerning the physical state of the user. Other ratios can also be extracted such as the time of any one phase over the time of the total breath cycle. For example, the time of inhalation in relation to the time of the total breath cycle (Ti/Ttotal).", para [0186]); and computing a respiratory rate based on the first and second periods ("As discussed above, one embodiment of the present invention can be used to determine VT (T1) and RCT (T2). In a different embodiment, VT and RCT calculations can be made within the classifier core module 730 itself. Processes such as respiratory rate tracking and breath phase tracking and detection are important in the analysis as the final result is not only based on the overall breath sound statistics, but also on statistics that come from the analysis of each breath cycle as the breathing session progresses over time (e.g. inhalation intensity tracking).", para [0194] – also see Fig. 19, ¶ 0260).
Stamatopoulos does not appear to explicitly teach obtaining a data structure that includes values that are representative of energy summed across different frequency bands of the spectrogram (although ¶ 0325 describes calculating the energy of each frame).
LeBoeuf teaches extracting features from an audio signal (¶ 0026), the features including e.g. energy (which is the sum of squared amplitudes within certain frequency bins) in various spectral bands, (¶ 0041). The features form a feature vector (¶ 0007).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate into the feature extraction of Stamatopoulos (Abstract, a spectrogram and a plurality of descriptors) the generation of a series of values representative of energy summed across different frequency bands extraction (LeBoeuf), for the purpose of improving classification of the audio data with additional features (LeBoeuf: ¶¶s 0026-0042 and Title, describing the large variety of features that may be used for audio recognition - also see Fig. 2).
Regarding claims 20 and 21, Stamatopoulos-LeBoeuf teaches the method of claim 16. Stamatopoulos-LeBoeuf further teaches wherein the operations further comprise: establishing a health state of the living body based an analysis of the multiple entries, each of which is representative of an independent prediction as to whether the corresponding segment of the recording includes evidence of disease, wherein the evidence of disease is presence of wheezing or crackling (Stamatopoulos: ¶¶s 0131, 0143, etc., wheeze detection and classification; ¶¶s 0295, 0427, 0429, etc., healthy or unhealthy; ¶¶s 0109, 0149, 0158, 0159, classifying each block based on its own features).
Response to Arguments
Applicant’s arguments filed 07/01/2025 have been fully considered.
In response to the arguments regarding the rejections under 35 USC 101, they are not persuasive. Filtering a signal is a mathematical operation, as are spectrogram extraction, MFCCs, concatenation, etc. With respect to claim 1, the Office has not asserted that these are directed to a mental process. And, the Office submits that many of these features are generalized. Applicant argues that the steps transform a signal into a form suitable for machine learning. This does not change the mathematical nature of the steps – changing data from one form to another. And, the application of machine learning is recited in a generic “black box” context that merely serves to process data. Specifying some of the layers of the neural network is not sufficient to obviate categorization as an abstract idea because the layers are also part of a mathematical operation. The series of steps serve to process data, but then do nothing with the determined final result that might be considered a practical application. Regarding an improvement to technology, Applicant has not pointed to where in the specification the improvement is discussed, and has not explained how that improvement is reflected in the claims. Automatic detection and quantification of breathing events is not described as a problem. But even if it were, an improvement in the abstract idea itself (i.e., better math) is not sufficient. It is “additional elements” that must provide the improvement, not the abstract idea elements. The analysis under Step 2B is the same since no elements have been identified as well-understood, routine, or conventional. Thus, the arguments are not persuasive.
In response to the amendments and arguments regarding the rejections under 35 USC 103, they are persuasive to the extent that the previous combination did not describe the particular neural network layers, or the setting of values to 1 or 0. Thus, a new grounds of rejection has been made in further view of Anushiravani, and all claims remain rejected in light of the prior art.
The Office maintains that Stamatopoulos teaches computing a respiratory rate based on analysis of breathing cycles. The analysis of cycles and periodicities results in a rate determination as described in e.g. ¶¶s 0186, 0194, 0125, and 0126 (inhalation/exhalation ratio, etc.). The autocorrelation function can also be considered a determination based on the first and second periods, since the function uses these periods.
Applicant alleges but does not explain why using multiple references for one feature is not permitted under law, or is entirely out of context. If Applicant is arguing that Hasan and LeBoeuf are non-analogous art, the Office disagrees. They are all in the field of (biological) audio signal processing. And, the Office has provided motivation for combination from the references themselves.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ANDREY SHOSTAK whose telephone number is (408) 918-7617. The examiner can normally be reached Monday - Friday 7 am - 3 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Robertson can be reached on (571) 272-5001. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ANDREY SHOSTAK/Primary Examiner, Art Unit 3791