Prosecution Insights
Last updated: October 02, 2026
Application No. 19/024,609

Apparatus and Method for Quality Determination of Audio Signals

Non-Final OA §102§103
Filed
Jan 16, 2025
Priority
Oct 20, 2022 — EU 22202847.4 +1 more
Examiner
PASHA, ATHAR N
Art Unit
Tech Center
Assignee
Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V.
OA Round
1 (Non-Final)
90%
Grant Probability
Favorable
1-2
OA Rounds
9m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 90% — above average
90%
Career Allowance Rate
153 granted / 169 resolved
+30.5% vs TC avg
Strong +15% interview lift
Without
With
+15.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
16 currently pending
Career history
186
Total Applications
across all art units

Statute-Specific Performance

§101
21.6%
-18.4% vs TC avg
§103
54.7%
+14.7% vs TC avg
§102
17.2%
-22.8% vs TC avg
§112
2.7%
-37.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 169 resolved cases

Office Action

§102 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: Claim 1-3, 9-16, 19 and 22 … and a distortion-to-quality mapping module… Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Objections Listed claims are objected to for the informalities shown and maybe addressed with the suggested amendments: Claim 13 : … wherein the apparatus is an apparatus according to claim [[14]] 1, and the distortion-to-quality mapping module is configured… Appropriate corrections are required. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claim(s) 1, 22, 25, 27 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Barbedo (Barbedo, Jayme & Lopes, Amauri. (2005). A new cognitive model for objective assessment of audio quality. Journal of the Audio Engineering Society. 53. 22-31). With respect to claim 1, Barbedo teaches (Claim 1) An apparatus for quality determination of an audio signal, wherein the apparatus comprises: (Claim 25 ) A method for quality determination of an audio signal, wherein the method comprises: (Claim 27 ) A non-transitory digital storage medium having a computer program stored thereon to perform the method for quality determination of an audio signal, wherein the method comprises: wherein the apparatus comprises (Barbedo , see section 5 where experimentation and results are discussed, which inherently allude to having computer program, a non-transitory computer readable storage medium having program instructions embodied therewith, a system comprising a memory device for storing program code; and a processor device operatively coupled to the memory device for running the program code.): a perceptual model for receiving the audio signal and for determining distortion information for each of one or more distortion metrics, wherein each distortion metric of the one or more distortion metrics depends on a comparison between a feature of the audio signal and of a corresponding feature of reference information, and a distortion-to-quality mapping module for determining a quality of the audio signal depending on the distortion information for each of the one or more distortion metrics and depending on information on one or more cognitive effects (Barbedo ¶Section 2 Para 1 Fig. 1 [Fig. 1 shows reference and test signals] shows the basic scheme of PEAQ. The processing of PEAQ can be divided into four main stages: 1) Psychoacoustic Model It models some of the main features of human hearing (such as masking effects, internal noise) and was implemented in two distinct versions, with the difference lying in the tool used to perform the time-frequency decomposition, namely, discrete Fourier transform (DFT) and filter bank. 2) Preprocessing of Excitation Patters This consists of several intermediary steps aiming at preparing the patterns resulting from the psychoacoustic model for the next stage. 3) Model Output Variables (MOVs) [distortion metric] These parameters aim to extract information from the signal being tested. The MOVs are based on some cognitive features involved in the perception of sound. Each psychoacoustic model leads to a different type of MOV. 4) Multilayer Perceptron Neural Network (MI.PNN) This network is used to map the MOVs into a single value representing the basic audio quality [quality of the audio signal] of the signal being tested. PEAQ has two versions. The basic version uses only MOVs extracted from the DFT-based psychoacoustic model. The advanced version uses MOVs extracted from both DFT- and filter-bank-based models. The advanced version was designed to achieve better performance, but at a greater computational cost. Channels of stereo signals are processed independently until the extraction of the MOVs. The values obtained for each channel are then combined using a simple arithmetic average.). With respect to claim 22 Barbedo teaches wherein the a distortion-to-quality mapping module is implemented as a cognitive salience model (Barbedo ¶ Section 1para3 ¶ Model Output Variables (MOVs) These parameters aim to extract information from the signal being tested. The MOVs are based on some cognitive features involved in the perception of sound. Each psychoacoustic model leads to a different type of MOV.); or wherein the distortion-to-quality mapping module is implemented as a multivariate-regression. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 17, 18 and 26 is (are) rejected under 35 U.S.C. 103 as being unpatentable over Barbedo in further view of Disch (US 20210082447 A1) With respect to claim 17 Barbedo teaches wherein each distortion metric of the plurality of distortion metrics depends on a comparison between a feature of the audio signal and of a corresponding feature of the reference signal (Barbedo ¶Section 2 Para 1 Fig. 1 [Fig. 1 shows reference and test signals] shows the basic scheme of PEAQ. The processing of PEAQ can be divided into four main stages: 1) Psychoacoustic Model It models some of the main features of human hearing (such as masking effects, internal noise) and was implemented in two distinct versions, with the difference lying in the tool used to perform the time-frequency decomposition, namely, discrete Fourier transform (DFT) and filter bank. 2) Preprocessing of Excitation Patters This consists of several intermediary steps aiming at preparing the patterns resulting from the psychoacoustic model for the next stage. 3) Model Output Variables (MOVs) [distortion metric] These parameters aim to extract information from the signal being tested. The MOVs are based on some cognitive features involved in the perception of sound. Each psychoacoustic model leads to a different type of MOV. 4) Multilayer Perceptron Neural Network (MI.PNN) This network is used to map the MOVs into a single value representing the basic audio quality [quality of the audio signal] of the signal being tested. PEAQ has two versions. The basic version uses only MOVs extracted from the DFT-based psychoacoustic model. The advanced version uses MOVs extracted from both DFT- and filter-bank-based models. The advanced version was designed to achieve better performance, but at a greater computational cost. Channels of stereo signals are processed independently until the extraction of the MOVs. The values obtained for each channel are then combined using a simple arithmetic average.). Barbedo does not explicitly disclose however Disch teaches wherein the reference information comprises information on a reference signal (Disch ¶ [0138] Optionally, the reference modulation information 182a-182c may be obtained using an optional reference modulation information determination 190 on the basis of a reference audio signal 192. The reference modulation information determination may, for example, perform the same functionality like the envelope signal determination 120 and the modulation information determination 160 on the basis of the reference audio signal 192.). It would have been obvious to one of ordinary skill in the art prior to the effective filing date of the invention to modify the distortion determination of Barbedo to include reference information of Disch in order to minimize compression distortion and measure perceptual loss. With respect to claim 18 Disch further teaches wherein the reference signal is an original audio signal, and wherein the audio signal is a decoded signal resulting from a decoding of an encoded audio signal, wherein the encoded audio signal encodes the original audio signal (Disch ¶[0070] Another embodiment according to the invention creates an audio encoder [encoder] for encoding an audio signal. The audio encoder is configured to determine one or more coding parameters (for example, encoding parameters or decoding parameters, [parameter] which are advantageously signaled to an audio decoder [decoder] by the audio encoder) in dependence on an evaluation of a similarity between an audio signal to be encoded and an encoded audio signal. The audio encoder is configured to evaluate the similarity between the audio signal to be encoded and the encoded audio signal (for example, a decoded version thereof) using an audio similarity evaluator as discussed herein (wherein the audio signal to be encoded is used as the reference audio signal and wherein a decoded version of an audio signal encoded using one or more candidate parameters is used as the input audio signal for the audio similarity evaluator). With respect to claim 26, Barbedo teaches [[wherein the quality determination of the audio signal is conducted by determining a quality of a decoded signal]], wherein the reference information comprises information on a reference signal, wherein each distortion metric of the plurality of distortion metrics depends on a comparison between a feature of the audio signal and of a corresponding feature of the reference signal, wherein the reference signal is an original audio signal effects (Barbedo ¶Section 2 Para 1 Fig. 1 [Fig. 1 shows reference and test signals] shows the basic scheme of PEAQ. The processing of PEAQ can be divided into four main stages: 1) Psychoacoustic Model It models some of the main features of human hearing (such as masking effects, internal noise) and was implemented in two distinct versions, with the difference lying in the tool used to perform the time-frequency decomposition, namely, discrete Fourier transform (DFT) and filter bank. 2) Preprocessing of Excitation Patters This consists of several intermediary steps aiming at preparing the patterns resulting from the psychoacoustic model for the next stage. 3) Model Output Variables (MOVs) [distortion metric] These parameters aim to extract information from the signal being tested. The MOVs are based on some cognitive features involved in the perception of sound. Each psychoacoustic model leads to a different type of MOV. 4) Multilayer Perceptron Neural Network (MI.PNN) This network is used to map the MOVs into a single value representing the basic audio quality [quality of the audio signal] of the signal being tested. PEAQ has two versions. The basic version uses only MOVs extracted from the DFT-based psychoacoustic model. The advanced version uses MOVs extracted from both DFT- and filter-bank-based models. The advanced version was designed to achieve better performance, but at a greater computational cost. Channels of stereo signals are processed independently until the extraction of the MOVs. The values obtained for each channel are then combined using a simple arithmetic average.) Barbedo does not explicitly disclose, however Disch teaches wherein the quality determination of the audio signal is conducted by determining a quality of a decoded signal (Disch ¶[0070] Another embodiment according to the invention creates an audio encoder [encoder] for encoding an audio signal. The audio encoder is configured to determine one or more coding parameters (for example, encoding parameters or decoding parameters, [parameter] which are advantageously signaled to an audio decoder [decoder] by the audio encoder) in dependence on an evaluation of a similarity between an audio signal to be encoded and an encoded audio signal. The audio encoder is configured to evaluate the similarity between the audio signal to be encoded and the encoded audio signal (for example, a decoded version thereof) using an audio similarity evaluator as discussed herein (wherein the audio signal to be encoded is used as the reference audio signal and wherein a decoded version of an audio signal encoded using one or more candidate parameters is used as the input audio signal for the audio similarity evaluator), wherein the audio signal is a decoded signal resulting from a decoding of an encoded audio signal, wherein the encoded audio signal encodes the original audio signal, wherein the method comprises receiving an original signal as the reference signal, wherein the decoded signal results from a decoding of an encoded audio signal, wherein the encoded audio signal encodes the original audio signal, and wherein the method comprises determining one or more coding parameters depending on a quality of the decoded audio signal (Disch ¶[0070] Another embodiment according to the invention creates an audio encoder [encoder] for encoding an audio signal. The audio encoder is configured to determine one or more coding parameters (for example, encoding parameters or decoding parameters, [parameter] which are advantageously signaled to an audio decoder [decoder] by the audio encoder) in dependence on an evaluation of a similarity between an audio signal to be encoded and an encoded audio signal. The audio encoder is configured to evaluate the similarity between the audio signal to be encoded and the encoded audio signal (for example, a decoded version thereof) using an audio similarity evaluator as discussed herein (wherein the audio signal to be encoded is used as the reference audio signal and wherein a decoded version of an audio signal encoded using one or more candidate parameters is used as the input audio signal for the audio similarity evaluator). It would have been obvious to one of ordinary skill in the art prior to the effective filing date of the invention to modify the distortion determination of Barbedo to include reference information of Disch in order to minimize compression distortion and measure perceptual loss. Claim(s) 20 is (are) rejected under 35 U.S.C. 103 as being unpatentable over Barbedo in further view of Murabayashi ( US 20050084244 A1 ). With respect to claim 20 Barbedo does not explicitly disclose however Murabayashi teaches wherein the reference information comprises a plurality of parameters which depends on a listener preference (Murabayashi ¶[ 0041] The information signal processing system and the information signal processing method of the present invention allow the user to set parameter data for playing back a digest according to the user's desired preference, thus allowing the user to perform a more efficient digest playback operation that satisfies the user's requirement.) It would have been obvious to one of ordinary skill in the art prior to the effective filing date of the invention to modify the distortion determination of Barbedo to include parameters of Murabayashi in order to align sound output with subjective user preferences. Allowable Subject Matter Claims 2 is objected to as being dependent upon a rejected base claim, but would be allowable if written in independent form including all of the limitations of the base claim and any intervening claims. Claim 2 recites “… wherein the one or more cognitive effects comprise at least one of informational masking information and perceptual streaming information, and wherein the distortion-to-quality mapping module is configured to determine the quality of the audio signal depending on the distortion information for each of the one or more distortion metrics and depending on at least one of the informational masking information and the perceptual streaming information” The closest prior art of record to the currently claimed limitations is Barbedo who teaches “The concepts of perceptual streaming and informational masking have been used here to generate an original parameter, which is not present in any of the other methods for objective audio quality assessment mentioned in this paper. This new parameter is obtained by mixing both concepts, as described in the following.…” Deng (US 20150104022 A1 ) teaches “¶[0043] Psychoacoustic studies have shown that speech intelligibility is affected significantly by energetic masking effect and informational masking effect of background signals to the target signals. Energetic masking effect relates to energy overlap between different speech signals in the same frequency band. Informational masking effect relates to listener's confusion caused by spatial and/or temporal overlap between different speech signals, ¶[0064] On the other hand, suppressing some sub-bands of an audio signal implies the audio quality will be degraded to some extent, and a proper allocation scheme shall be assured to avoid significant degradation of audio quality. For example, it is preferred to make each audio signal cover both low frequency sub-bands and high frequency sub-bands. Another example, if the number of speakers/audio signals to be separated is too large, it might be improper to allocate to each audio signal too few or too narrow reserved sub-bands. In such a situation, the reserved sub-bands for different audio signals may be allowed to overlap each other (as shown in FIG. 6(b), wherein "1" indicates sub-bands for audio signal 1, and "2" indicates sub-bands for audio signal 2), but as little as possible; or, some audio signals, especially those relatively important audio signals, may be allocated to significantly broader sub-bands (as shown in upper line in FIG. 7, wherein audio signal 1 is more important than audio signal 2), even the full band if the audio signal is the most important (as shown in lower line in FIG. 7: audio signal 3 is the most important). However, Barbedo and Deng individually or in combination do not teach, suggest or render obvious to one of ordinary skill in the art before the effective filing date, at least the specific limitations as recited above. Furthermore, it would not have been obvious to one of ordinary skill in the art to modify the prior art in order to arrive at the claimed invention. Therefore, claims 2-16, 19, 21, 23 and 24 are allowable. Any comments considered necessary by applicant must be submitted no later than the payment of the issue fee and, to avoid processing delays, should preferably accompany the issue fee. Such submissions should be clearly labeled “Comments on Statement of Reasons for Allowance.” Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to ATHAR N PASHA whose telephone number is (408)918-7675. The examiner can normally be reached Monday-Thursday Alternate Fridays, 7:30-4:30 PT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at (571)272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ATHAR N PASHA/Primary Examiner, Art Unit 2657
Read full office action

Prosecution Timeline

Jan 16, 2025
Application Filed
Aug 10, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12724980
CONTEXT-ENHANCED ADVANCED FEEDBACK FOR DRAFT MESSAGES
3y 2m to grant Granted Sep 01, 2026
Patent 12718019
INFORMATION PROCESSING APPARATUS, INFORMATION PROCESSING METHOD, AND STORAGE MEDIUM
2y 9m to grant Granted Aug 25, 2026
Patent 12705436
GENERATING DIGITAL CONTENT
3y 3m to grant Granted Aug 11, 2026
Patent 12688377
METHOD AND USER APPARATUS FOR GENERATING AND APPLYING TRANSLATION MARKER
3y 0m to grant Granted Jul 21, 2026
Patent 12639516
CLASSIFICATION AND AUGMENTATION OF UNSTRUCTURED DATA FOR AUTOFILL
2y 5m to grant Granted May 26, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
90%
Grant Probability
99%
With Interview (+15.2%)
2y 6m (~9m remaining)
Median Time to Grant
Low
PTA Risk
Based on 169 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month