Prosecution Insights
Last updated: October 04, 2026
Application No. 19/012,218

SYSTEM AND A PROCESSING METHOD FOR CUSTOMIZING AUDIO EXPERIENCE

Non-Final OA §103
Filed
Jan 07, 2025
Priority
Jan 05, 2018 — SG 10201800147X +5 more
Examiner
PATEL, YOGESHKUMAR G
Art Unit
Tech Center
Assignee
Zeica Labs Pte. Ltd.
OA Round
1 (Non-Final)
84%
Grant Probability
Favorable
1-2
OA Rounds
6m
Est. Remaining
87%
With Interview

Examiner Intelligence

Grants 84% — above average
84%
Career Allowance Rate
566 granted / 678 resolved
+23.5% vs TC avg
Minimal +3% lift
Without
With
+3.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 3m
Avg Prosecution
20 currently pending
Career history
684
Total Applications
across all art units

Statute-Specific Performance

§101
5.0%
-35.0% vs TC avg
§103
68.6%
+28.6% vs TC avg
§102
12.3%
-27.7% vs TC avg
§112
11.5%
-28.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 678 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Double Patenting A rejection based on double patenting of the “same invention” type finds its support in the language of 35 U.S.C. 101 which states that “whoever invents or discovers any new and useful process... may obtain a patent therefor...” (Emphasis added). Thus, the term “same invention,” in this context, means an invention drawn to identical subject matter. See Miller v. Eagle Mfg. Co., 151 U.S. 186 (1894); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); and In re Ockert, 245 F.2d 467, 114 USPQ 330 (CCPA 1957). A statutory type (35 U.S.C. 101) double patenting rejection can be overcome by canceling or amending the claims that are directed to the same invention so they are no longer coextensive in scope. The filing of a terminal disclaimer cannot overcome a double patenting rejection based upon 35 U.S.C. 101. Application #19/012,218 Claim 1: A method comprising: generating, by one or more processors, image-related input audio signals from processing an image of a subject; and generating, by the one or more processors, output audio signals comprising the image-related input audio signals, wherein the output audio signals are audibly perceivable by the subject. Claim 18: A system comprising: one or more processors; and one or more tangible, non-transitory memories configured to communicate with the one or more processors, the one or more tangible, non-transitory memories having instructions stored thereon that, in response to execution by the one or more processors, cause the one or more processors to perform operations comprising: generating, by the one or more processors, image-related input audio signals from processing an image of a subject; and generating, by the one or more processors, output audio signals comprising the image-related input audio signals, wherein the output audio signals are audibly perceivable by the subject. Claim 19: The system of claim 18, further comprising: processing, by the one or more processors, the image-related input audio signal based on a database signal; and generating, by the one or more processors, a plurality of intermediate processor datasets based on the processing. Claim 20: An article of manufacture including one or more non-transitory, tangible computer readable storage mediums having instructions stored thereon that, in response to execution by one or more processors, cause the one or more processors to perform operations comprising: generating, by the one or more processors, image-related input audio signals from processing an image of a subject; and generating, by the one or more processors, output audio signals comprising the image-related input audio signals, wherein the output audio signals are audibly perceivable by the subject. Patent #12,225,371 Claim 1: A method comprising: generating, by one or more processors, image-related input audio signals from processing a captured image of a subject; applying, by the one or more processors, an output signal to the image-related input audio signals; and generating, by the one or more processors and based on the applying, output audio signals audibly perceivable by the subject. Claim 17: A system comprising: one or more processors; and one or more tangible, non-transitory memory configured to communicate with the processor, the one or more tangible, non-transitory memory having instructions stored thereon that, in response to execution by the one or more processors, cause the processor to perform operations comprising: generating, by the one or more processors, image-related input audio signals from processing a captured image of a subject; applying, by the one or more processors, an output signal to the image-related input audio signals; and generating, by the one or more processors and based on the applying, output audio signals audibly perceivable by the subject. Claim 18: The system of claim 17, further comprising: processing, by the one or more processors, the image-related input audio signal based on a database signal; and generating, by the one or more processors, a plurality of intermediate processor datasets based on the processing. Claim 20: An article of manufacture including one or more non-transitory, tangible computer readable storage medium having instructions stored thereon that, in response to execution by the one or more processors, cause the processor to perform operations comprising: generating, by the one or more processors, image-related input audio signals from processing a captured image of a subject; applying, by the one or more processors, an output signal to the image-related input audio signals; and generating, by the one or more processors and based on the applying, output audio signals audibly perceivable by the subject. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-3, 6-7, and 18-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Eronen et al. (US #2018/0295463) in view of Hirst (US Patent #9544706). Regarding Claim 1, Eronen discloses a method (title, abstract, Figs. 1-11) comprising: generating, by one or more processors (Eronen ¶0369 discloses the data processors can be of any type suitable to the local technical environment, and can include one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors [DSPs], application specific integrated circuits [ASIC], gate level circuits and processors based on multi-core processor architecture, as non-limiting examples), image-related input audio signals from processing an image of a subject (Eronen ¶0129 discloses receive the outputs of the Lavalier microphone 111 and the spatial audio capture device 113 … to receive source position and tracking information from the position tracker 115. ¶0141 discloses the capture apparatus 101 can furthermore comprise a camera or cameras 107 configured to generate images. The camera or cameras can be configured to generate a panoramic image or video of images which is captured along with the spatial audio. The camera 107 can be part of the same apparatus configured to capture the spatial audio signals, for example a mobile phone or user equipped with a microphone array and a camera or cameras). Eronen may not explicitly disclose generating, by the one or more processors, output audio signals comprising the image-related input audio signals, wherein the output audio signals are audibly perceivable by the subject. However, Hirst (title, abstract, Figs. 1-8) teaches generating, by the one or more processors (Hirst Fig. 5: processor(s) 520; col. 4, lines 49-58; col. 9, lines 15-31), output audio signals comprising the image-related input audio signals (Hirst col. 9, line 61 to col. 10, line 41 discloses Fig. 6 is a flow diagram illustrating an example method 600 for creating a customized HRTF for a human pinna and providing virtual surround sound through headphones based on a 3D digital model of the human pinna. … As in block 608, the HRTF that is customized for the human pinna can be used to provide virtual surround sound through headphones. An application that provides audio output can, for example, use the HRTF that is customized for the human pinna to configure the audio output such that an improved virtual surround sound effect is produced when the audio output is heard through headphones. Fig. 7 is a flow diagram illustrating an example method 700 for creating a customized HRTF for a human pinna. As in block 702, a plurality of digital sensor readings made on at least a part of a human pinna can be received. The plurality of digital sensor readings can comprise digital images taken using a visible-light camera or an infrared camera. …. As in block 706, an HRTF that is customized for the human pinna can be determined using the digital model of the human pinna and the HRTF can be compatible with a virtual surround sound system to enable customization of the virtual surround sound system for the human pinna), wherein the output audio signals are audibly perceivable by the subject (Hirst col. 1, lines 32-41 discloses the external ear, including the pinna, has transforming effects on sound waves that are ultimately perceived by the eardrum [i.e., the tympanic membrane] in humans. The external ear can, for example, act as a filter that reduces low frequencies, a resonator that enhances middle frequencies, and a directionally dependent filter at high frequencies that assists with spatial perception. Ideally, if an HRTF is accurate, the HRTF can be used by spatial sound reproduction systems to assist in creating the desired illusion that sound originates from a specific direction relative to a user). Regarding Claim 2, Eronen in view of Hirst discloses the method of claim 1. But Eronen may not explicitly disclose further comprising applying, by the one or more processors, the output signals to the image-related input audio signals. However, Hirst (title, abstract, Figs. 1-8) teaches applying, by the one or more processors, the output signals to the image-related input audio signals (Hirst col. 10, line 14-20 discloses as in block 608, the HRTF that is customized for the human pinna can be used to provide virtual surround sound through headphones. An application that provides audio output can, for example, use the HRTF that is customized for the human pinna to configure the audio output such that an improved virtual surround sound effect is produced when the audio output is heard through headphones). Regarding Claim 3, Eronen in view of Hirst discloses the method of claim 1, further comprising: processing, by the one or more processors, the image-related input audio signal based on a database signal (Eronen ¶0158 discloses the upper path of Fig. 8 shows a conventional binaural rendering engine. The input signal is passed via an amplifier 1601 applying a gdry gain to a head related transfer function [HRTF] interpolator 1605. The HRTF interpolator 1605 can comprise a set of HRTFs in a database and from which HRTF filter coefficients are selected based on the direction of arrival input); and generating, by the one or more processors, a plurality of intermediate processor datasets based on the processing (Eronen ¶0158 discloses the input signal can then be convolved with the interpolated HRTF to generate a left and right HRTF output which is passed to a left output combiner 1641 and a right output combiner 1643. ¶0037 discloses determining the at least one space parameter can comprise: determining at least one interim space parameter based on the at least one additional audio signal; determining at least one further interim space parameter based on an analysis of at least one camera image; and determining at least one final space parameter based on the at least one interim space parameter and the at least one further interim space parameter). Regarding Claim 6, Eronen in view of Hirst discloses the method of claim 1, wherein the output audio signals provide the subject with a customized audio experience (Eronen ¶0264 discloses the determination of the possible effects or processing to be applied by the renderer as defined by the ruleset selector 303 can furthermore be based on user preferences. Alternatively, the ruleset selector 303 can be configured to operate initially according to initial or 'factory' settings, but the user can then customize according to their own preferences). Regarding Claim 7, Eronen in view of Hirst discloses the method of claim 1, further comprising combining, by the one or more processors, a plurality of intermediate processor datasets to produce the output signal (Eronen ¶0046 discloses means for determining at least one processing effect ruleset based on the at least one source parameter and/or the at least one space parameter; means for generating at least two output audio channel signals by mixing and applying at least one processing effect to the spatial audio signal and the at least one additional audio signal based on the at least one processing effect ruleset. ¶0051 discloses the means for generating the at least two output audio channel signals by mixing and applying the at least one processing effect to the spatial audio signal and the at least one additional audio signal may further comprise means for mixing and applying the at least one processing effect based on the relative position between the first position associated with the microphone array and the second position associated with the additional microphone. ¶0087 discloses the concept can for example be embodied as a capture system configured to capture both a close [speaker, instrument or other source] audio signal and a spatial [audio field] audio signal. The capture system can furthermore be configured to determine or classify a source and/or the space within which the source is located. This information can then be stored or passed to a suitable rendering system which having received the audio signals and the information [source and space classification] can use this information to generate a suitable mixing and rendering of the audio signal to a user). Claims 18-20 are rejected for the same reasons as set forth in Claims 1-3. Claims 4, 9, and 12-13, is/are rejected under 35 U.S.C. 103 as being unpatentable over Eronen et al. (US #2018/0295463) in view of Hirst (US Patent #9544706) further in view of Lyren et al. (US #2017/0245081). Regarding Claim 4, Eronen in view of Hirst discloses the method of claim 3, but may not explicitly disclose wherein the plurality of intermediate processor datasets corresponds to biometric data for at least two different biometric feature types. However, Lyren (title, abstract, Figs. 1-16) teaches wherein the plurality of intermediate processor datasets corresponds to biometric data for at least two different biometric feature types (Lyren ¶0294 discloses the user profile includes information pertaining to the characteristics and/or preferences of the user. Examples of this information for a person include, but are not limited to, one or more of personal data of the user [such as communication hardware and software used including peripheral devices such as head tracking systems, biometric data, physical measurements of their body and environments, functions of physical data such as HRTFs, etc.], photographs [such as photos of their head and ears. ¶0118 discloses facial recognition to determine or estimate a facial or head orientation of the person from one or more images or video captured with a camera in the HPED. The facial orientation can be described or recorded with respect to a location of the HPED that is generating the sound for the impulse responses). Eronen, Hirst, and Lyren are analogous art as they pertain to customizing audio. Therefore it would have been obvious to someone of ordinary skill in the art before the effective filing date of the invention was made to modify the teachings of Eronen in view of Hirst in light of the teachings of Lyren to capture image of video of the person in real-time using a camera (as taught by Lyren ¶0122) to overcome typical methods of obtaining HRTFs since it is measured in an anechoic chamber or specialized location requiring expensive sound equipment (Lyren ¶0002). Regarding Claim 9, Eronen in view of Hirst discloses the method of claim 7. But Eronen in view of Hirst may not explicitly disclose further comprising generating, by the one or more processors, a plurality of intermediate processor datasets by processing an image-related input audio signal based on a database signal. However, Lyren (title, abstract, Figs. 1-16) teaches generating, by the one or more processors, a plurality of intermediate processor datasets by processing an image-related input audio signal based on a database signal (Lyren ¶0028 discloses generate, manage, and perform tasks for audio impulse responses, including room impulse responses [RIRs], binaural room impulse responses [BRIRs], head-related impulse responses [HRIRs], and head-related transfer functions [HRTFs]. ¶0240 discloses Bob originates the call to Alice with the smartphone's stock phone application and waits while he hears the ring indicator. Alice is driving her car wearing headphones and is listening to her phone playing music when she hears a ringtone. The ringtone indicates the she does not have an SLP configured for her current location on the road. She also has not yet taken a call using an SLP with her new phone application that supports binaural speech convolution. She is already wearing her headphones with microphones so she takes this opportunity to create an SLP suitable to use in the car so she can enjoy a more natural phone conversation with the perception of Bob's voice externalized. On the display of her phone there is an "answer phone" button/option and a button/option that says, "answer at new SLP." Alice selects the latter option to answer at a new SLP. Her phone indicates that it will generate an SLP when the phone is steady at arm's length. Bob is then connected and they exchange greetings. Soon Alice tells Bob, "Hold on for a moment, I'm in a car and I'd like to externalize you . . . " She extends her right arm toward the passenger seat while keeping her face safely toward the road. The phone's binaural calling application monitors the image received by the phone's camera. When the application detects Alice's facial profile in the center of the image, the application uses the image to calculate the phone's location relative to the face of Alice and determines the distance to her face to be arm's length. The phone further uses its motion detector to determine that it is steady and provides an indication [e.g., vibratory or audio] to Alice that it is ready to create the SLP. ¶0255 discloses the storage 1408 can include memory or databases that store one or more of SLPs [including their locations and other information associated with a SLP including rich media such as sound files and images], user profiles and/or user preferences [such as user preferences for SLP locations and sound localization preferences], impulse responses and transfer functions [such as HRTFs, HRIRs, BRIRs, and RIRs], and other information). Eronen, Hirst, and Lyren are analogous art as they pertain to customizing audio. Therefore it would have been obvious to someone of ordinary skill in the art before the effective filing date of the invention was made to modify the teachings of Eronen in view of Hirst in light of the teachings of Lyren to capture image of video of the person in real-time using a camera (as taught by Lyren ¶0122) to overcome typical methods of obtaining HRTFs since it is measured in an anechoic chamber or specialized location requiring expensive sound equipment (Lyren ¶0002). Regarding Claim 12, Eronen in view of Hirst discloses the method of claim 1, but may not explicitly disclose further comprising operating, by the one or more processors, as at least one of a type of recognizer of a plurality of recognizers or a multi-recognizer corresponding to a first type recognizer and a second type recognizer. However, Lyren (title, abstract, Figs. 1-16) teaches operating, by the one or more processors, as at least one of a type of recognizer of a plurality of recognizers or a multi-recognizer corresponding to a first type recognizer and a second type recognizer (Lyren ¶0218 discloses Alice receives a VoIP call from Bob while she is at her grandmother's house. Her smartphone determines that Alice has not previously received a call at this location and hence is unable to retrieve either an RIR or BRIR for her current location. In response to this determination, the smartphone rings with a distinctive tone, and Alice recognizes this tone and its implication that no RIRs or BRIRs are available for her location. This distinctive tone is actually the sound used to capture impulse responses. While her smartphone is ringing and generating this distinctive tone, Alice holds the smartphone in her hand with her arm stretched out away from her face. Microphones in her earphones record the tones, and her smartphone immediately generates BRIRs for Alice. When the smartphone captures sufficient impulse response from a designated location, it stops generating the tones, answers the call, and convolves the incoming voice with the BRIRs that it just obtained while Alice was answering the phone call). Eronen, Hirst, and Lyren are analogous art as they pertain to customizing audio. Therefore it would have been obvious to someone of ordinary skill in the art before the effective filing date of the invention was made to modify the teachings of Eronen in view of Hirst in light of the teachings of Lyren to analyze using the HPED the target impulse that is the designated sound in order to create the BRIRs and HRTFs (as taught by Lyren ¶0141) to overcome typical methods of obtaining HRTFs since it is measured in an anechoic chamber or specialized location requiring expensive sound equipment (Lyren ¶0002). Regarding Claim 13, Eronen in view of Hirst discloses the method of claim 1, but may not explicitly disclose wherein the image-related input audio signals correspond to biometric data associated with the subject, and wherein the biometric data comprises a first biometric feature type and a second biometric feature type. However, Lyren teaches wherein the image-related input audio signals correspond to biometric data associated with the subject (Lyren ¶0294 discloses the user profile includes information pertaining to the characteristics and/or preferences of the user. Examples of this information for a person include, but are not limited to, one or more of personal data of the user [such as communication hardware and software used including peripheral devices such as head tracking systems, biometric data, physical measurements of their body and environments, functions of physical data such as HRTFs, etc.], photographs [such as photos of their head and ears. ¶0118 discloses facial recognition to determine or estimate a facial or head orientation of the person from one or more images or video captured with a camera in the HPED. The facial orientation can be described or recorded with respect to a location of the HPED that is generating the sound for the impulse responses), and wherein the biometric data comprises a first biometric feature type and a second biometric feature type (Lyren ¶0254 discloses one or more sensors 1430 such as biometric sensors). Eronen, Hirst, and Lyren are analogous art as they pertain to customizing audio. Therefore it would have been obvious to someone of ordinary skill in the art before the effective filing date of the invention was made to modify the teachings of Eronen in view of Hirst in light of the teachings of Lyren to capture image of video of the person in real-time using a camera (as taught by Lyren ¶0122) to overcome typical methods of obtaining HRTFs since it is measured in an anechoic chamber or specialized location requiring expensive sound equipment (Lyren ¶0002). Claim 5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Eronen et al. (US #2018/0295463) in view of Hirst (US Patent #9544706) further in view of Lyren et al. (US #2017/0245081) and Ghorbal et al. (US #2018/0249275). Regarding Claim 5, Eronen in view of Hirst discloses the method of claim 1, but may not explicitly disclose wherein the processing the image of the subject further comprises implementing at least one of a multiple-match processing based strategy, a multiple-recognizer based processing strategy, or a cluster based processing strategy. However, Lyren (title, abstract, Figs. 1-16) teaches wherein the processing the image of the subject further comprises implementing at least one of a multiple-match processing based strategy (Lyren ¶0141 discloses capture and record the sounds at a location for various lengths of time and then select the optimal impulse to use in creating the BRIRs and HRTFs. The HPED can also be instructed to disregard impulses according to a set of criteria, and/or to consider for analysis impulses that match a set of criteria and disregard ones that do not match the criteria. ¶0207 discloses as yet another example, the HPED or electronic device retrieves RIRs for a similar location. For instance, if the location is a church but no RIRs exist for this particular church, then RIRs for another church are retrieved. Physical attributes of the location [such as size, shape, and other physical qualities] can be used to more closely match RIRs from other locations), or a multiple-recognizer based processing strategy (Lyren ¶0218 discloses Alice receives a VoIP call from Bob while she is at her grandmother's house. Her smartphone determines that Alice has not previously received a call at this location and hence is unable to retrieve either a RIR or BRIR for her current location. In response to this determination, the smartphone rings with a distinctive tone, and Alice recognizes this tone and its implication that no RIRs or BRIRs are available for her location. This distinctive tone is actually the sound used to capture impulse responses. While her smartphone is ringing and generating this distinctive tone, Alice holds the smartphone in her hand with her arm stretched out away from her face. Microphones in her earphones record the tones, and her smartphone immediately generates BRIRs for Alice. When the smartphone captures sufficient impulse response from a designated location, it stops generating the tones, answers the call, and convolves the incoming voice with the BRIRs that it just obtained while Alice was answering the phone call). Eronen, Hirst, and Lyren are analogous art as they pertain to customizing audio. Therefore it would have been obvious to someone of ordinary skill in the art before the effective filing date of the invention was made to modify the teachings of Eronen in view of Hirst in light of the teachings of Lyren to analyze using the HPED the target impulse that is the designated sound in order to create the BRIRs and HRTFs (as taught by Lyren, ¶0141) to overcome typical methods of obtaining HRTFs since it is measured in an anechoic chamber or specialized location requiring expensive sound equipment (Lyren ¶0002). Eronen in view of Hirst and Lyren may not explicitly disclose a cluster based processing strategy. However, Ghorbal (title, abstract, Figs. 1-4) teaches a cluster based processing strategy (Ghorbal ¶0021 discloses starting with a substantial database of HRTFs, similar HRTFs are grouped together. The Euclidian distance naturally associated with this 16-dimensional space then allowed the HRTFs to be grouped into clusters [of 8 in number]. Sets of HRTFs were then randomly chosen within the clusters and subjects invited to choose the one or more clusters that gave them the best impression of externality and directivity. ¶0023 discloses once the cluster has been selected, another selecting step in which a very precise set is selected may be added. A procedure for selecting a set of HRTFs from 32 available HRTFs is described. ¶0072 discloses cluster equivalence will possibly be spoken of, all the points of a given cluster playing a similar role within the ear to which they belong. ¶0083 discloses alternatively, any method allowing the values of the set of parameters of the head-related transfer functions to be found from the values of the set of statistical parameters and ensuring a good reconstruction of the head-related transfer functions of the database OH1 can be used, examples of such methods being methods based on neural networks, based on multiple component analysis [MCA] or based on k-means clustering). Eronen, Hirst, Lyren, and Ghorbal are analogous art as they pertain to customizing audio. Therefore it would have been obvious to someone of ordinary skill in the art before the effective filing date of the invention was made to modify the teachings of Eronen in view of Hirst and Lyren in light of the teachings of Ghorbal to group similar HRTFs together (as taught by Ghorbal, ¶0021) to generate an individual-specific HRTF more rapidly and with a higher reliability (Ghorbal, ¶0034). Claims 8, 10 and 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Eronen et al. (US #2018/0295463) in view of Hirst (US Patent #9544706) further in view of Koppens et al. (US PGPUB #2015/0358754). Regarding Claim 8, Eronen in view of Hirst discloses the method of claim 7, but may not explicitly disclose wherein the plurality of intermediate processor datasets includes at least one of a binaural room impulse response (BRIR) dataset, a binaural room transfer function (BRTF) dataset, a head related impulse response (HRIR) dataset or a head related transfer function (HRTF) dataset. However, Koppens (title, abstract, Figs. 1-10) teaches wherein the plurality of intermediate processor datasets includes at least one of a binaural room impulse response (BRIR) dataset, a binaural room transfer function (BRTF) dataset, a head related impulse response (HRIR) dataset or a head related transfer function (HRTF) dataset (Koppens ¶0020 discloses the binaural filter functions can be represented e.g., as a Head Related Impulse Responses [HRIR] or equivalently as Head Related Transfer Functions [HRTFs] or a Binaural Room Impulse Response [BRIR] or a Binaural Room Transfer Function [BRTF]). Eronen, Hirst, and Koppens are analogous art as they pertain to customizing audio. Therefore it would have been obvious to someone of ordinary skill in the art before the effective filing date of the invention was made to modify the teachings of Eronen in view of Hirst in light of the teachings of Koppens to process an algorithm or process which for a signal representing a sound source generates audio signals for the two ears of a person such that the sound is perceived to originate from a desired position in 3D space (as taught by Koppens, ¶0028) since a speaker setup which is different from the setup that corresponds to the multi-channel signal, the spatial image will be suboptimal and channel-based audio coding systems are typically not able to cope with a different number of speakers (Koppens, ¶0004). Regarding Claim 10, Eronen in view of Hirst discloses the method of claim 1, but may not explicitly disclose wherein the output signal corresponds to an audio response characteristic associated with the subject. However, Koppens (title, abstract, Figs. 1-10) teaches wherein the output signal corresponds to an audio response characteristic associated with the subject (Koppens ¶0035 discloses the processing of the audio signal can be a virtual position binaural rendering processing based on parameters of a head related binaural transfer function retrieved from the selected binaural rendering data set. ¶0091 discloses the transmitter comprises an HRTF generator 601 which generates a plurality of head related binaural transfer functions, which in the specific example are HRTFs but it can additionally or alternatively be e.g., HRIRs, BRIRs or BRTFs. Indeed, in the following the term HRTF will for brevity refer to any representation of a head related binaural transfer function, including HRIRs, BRIRs or BRTFs as appropriate. ¶0092 discloses each of the HRTFs is then represented by a data set, with each of the data sets providing one representation of one HRTF. ¶0098 discloses the HRTF generator 601 generates a plurality of data sets for each HRTF with each data set providing a representation of the HRTF. Furthermore, the HRTF generator 601 generates data sets for a plurality of positions). Eronen, Hirst, and Koppens are analogous art as they pertain to customizing audio. Therefore it would have been obvious to someone of ordinary skill in the art before the effective filing date of the invention was made to modify the teachings of Eronen in view of Hirst in light of the teachings of Koppens to process an algorithm or process which for a signal representing a sound source generates audio signals for the two ears of a person such that the sound is perceived to originate from a desired position in 3D space (as taught by Koppens, ¶0028) since a speaker setup which is different from the setup that corresponds to the multi-channel signal, the spatial image will be suboptimal and channel-based audio coding systems are typically not able to cope with a different number of speakers (Koppens, ¶0004). Regarding Claim 11, Eronen in view of Hirst and Koppens discloses the method of claim 10. But Eronen in view of Hirst may not explicitly disclose wherein the audio response characteristic is unique to the subject under a given environment and is one of an HRIR or BRIR. However, Koppens (title, abstract, Figs. 1-10) teaches wherein the audio response characteristic is unique to the subject under a given environment and is one of an HRIR or BRIR (Koppens ¶0020 discloses the binaural filter functions can be represented e.g., as a Head Related Impulse Responses [HRIR] or equivalently as Head Related Transfer Functions [HRTFs] or a Binaural Room Impulse Response [BRIR] or a Binaural Room Transfer Function [BRTF]. The [e.g., estimated or assumed] transfer function from a given position to the listener's ears [or eardrums] is known as a head related binaural transfer function. This function can for example be given in the frequency domain in which case it is typically referred to as an HRTF or BRTF or in the time domain in which case it is typically referred to as a HRIR or BRIR). Eronen, Hirst, and Koppens are analogous art as they pertain to customizing audio. Therefore it would have been obvious to someone of ordinary skill in the art before the effective filing date of the invention was made to modify the teachings of Eronen in view of Hirst in light of the teachings of Koppens to process an algorithm or process which for a signal representing a sound source generates audio signals for the two ears of a person such that the sound is perceived to originate from a desired position in 3D space (as taught by Koppens, ¶0028) since a speaker setup which is different from the setup that corresponds to the multi-channel signal, the spatial image will be suboptimal and channel-based audio coding systems are typically not able to cope with a different number of speakers (Koppens, ¶0004). Claims 14-15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Eronen et al. (US #2018/0295463) in view of Hirst (US Patent #9544706) further in view of Lyren et al. (US #2017/0245081) and Flaks et al. (US #2012/0093320). Regarding Claim 14, Eronen in view of Hirst and Lyren discloses the method of claim 13, but may not explicitly disclose further comprising: generating, by the one or more processors as a first type recognizer, a first set of intermediate processor datasets based on the first biometric feature type, and generating, by the one or more processors as a second type recognizer, a second set of intermediate processor datasets based on the second biometric feature type. However, Flaks (abstract, Figs. 1-12) teaches generating, by the one or more processors as a first type recognizer, a first set of intermediate processor datasets based on the first biometric feature type (Flaks Fig. 10: 1040 load reference HRTFs, get input from sensor(s), calculate pinna and head characteristics, look up optimal HRTF or interpolate. ¶0074 discloses head size and width. Specific HRTF can be associated with specific measurements related to the head size. Measurement can be head width. ¶0078 discloses in step 802 [Fig. 8], the listener 8 is identified using biometric information. Collecting biometric information can include collecting depth information and RGB information. ¶0079 discloses in step 804: select HRTF based on identified user), and generating, by the one or more processors as a second type recognizer, a second set of intermediate processor datasets based on the second biometric feature type (Flaks ¶0074 discloses pinna characteristics. Specific HRTF can be associated with specific measurements related to the pinna. ¶0078 discloses in step 802 [Fig. 8], the listener 8 is identified using biometric information. Collecting biometric information can include collecting depth information and RGB information. ¶0079 discloses in step 804: select HRTF based on identified user. Fig. 10: 1040 load reference HRTFs, get input from sensor(s), calculate pinna and head characteristics, look up optimal HRTF or interpolate). Eronen, Hirst, Lyren, and Flaks are analogous art as they pertain to customizing audio. Therefore it would have been obvious to someone of ordinary skill in the art before the effective filing date of the invention was made to modify the teachings of Eronen in view of Hirst and Lyren in light of the teachings of Flaks to determine the HRTF for the listener based on sensor determined characteristics, such as using the image camera component (as taught by Flaks, ¶0046) to improve and provide an accurate HRTF for the user and it should be accurate, consumer friendly, cost effective, and compatible with existing audio systems (Flaks, ¶0005). Regarding Claim 15, Eronen in view of Hirst and Lyren discloses the method of claim 14, but may not explicitly disclose further comprising combining, by the one or more processors using weighted sums, the first set of intermediate processor datasets and the second set of intermediate processor datasets. However, Flaks (abstract, Figs. 1-12) teaches combining, by the one or more processors using weighted sums, the first set of intermediate processor datasets and the second set of intermediate processor datasets (Flaks ¶0083 discloses the results can be summed prior to applying estimated reverb tail 1010 to produce the final 3D audio signal, which can be played through headphones 1012. Fig. 10: 1040 load reference HRTFs, get input from sensor(s), calculate pinna and head characteristics, look up optimal HRTF or interpolate; sum HRTFs). Eronen, Hirst, Lyren, and Flaks are analogous art as they pertain to customizing audio. Therefore it would have been obvious to someone of ordinary skill in the art before the effective filing date of the invention was made to modify the teachings of Eronen in view of Hirst and Lyren in light of the teachings of Flaks to determine the HRTF for the listener based on sensor determined characteristics, such as using the image camera component (as taught by Flaks, ¶0046) to improve and provide an accurate HRTF for the user and it should be accurate, consumer friendly, cost effective, and compatible with existing audio systems (Flaks, ¶0005). Claim 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Eronen et al. (US #2018/0295463) in view of Hirst (US Patent #9544706) further in view of Ghorbal et al. (US #2018/0249275) and Lyren et al. (US #2017/0245081). Regarding Claim 16, Eronen in view of Hirst discloses the method of claim 1, but may not explicitly disclose further comprising: grouping, by the one or more processors, a plurality of datasets of a database into a plurality of cluster groups, each cluster group corresponding to a cluster comprising a dataset, and associating, by the one or more processors, a biometric feature type retrieved from the image with the cluster. However, Ghorbal (title, abstract, Figs. 1-4) teaches grouping, by the one or more processors, a plurality of datasets of a database into a plurality of cluster groups, each cluster group corresponding to a cluster comprising a dataset (Ghorbal ¶0021 discloses starting with a substantial database of HRTFs, similar HRTFs are grouped together. The Euclidian distance naturally associated with this 16-dimensional space then allowed the HRTFs to be grouped into clusters [of 8 in number]. Sets of HRTFs were then randomly chosen within the clusters and subjects invited to choose the one or more clusters that gave them the best impression of externality and directivity. ¶0023 discloses once the cluster has been selected, another selecting step in which a very precise set is selected may be added. A procedure for selecting a set of HRTFs from 32 available HRTFs is described. ¶0072 discloses cluster equivalence will possibly be spoken of, all the points of a given cluster playing a similar role within the ear to which they belong. ¶0083 discloses alternatively, any method allowing the values of the set of parameters of the head-related transfer functions to be found from the values of the set of statistical parameters and ensuring a good reconstruction of the head-related transfer functions of the database OH1 can be used, examples of such methods being methods based on neural networks, based on multiple component analysis (MCA) or based on k-means clustering). Eronen, Hirst, and Ghorbal are analogous art as they pertain to customizing audio. Therefore it would have been obvious to someone of ordinary skill in the art before the effective filing date of the invention was made to modify the teachings of Eronen in view of Hirst in light of the teachings of Ghorbal to group similar HRTFs together (as taught by Ghorbal, ¶0021) to generate an individual-specific HRTF more rapidly and with a higher reliability (Ghorbal, ¶0034). Eronen in view of Hirst and Ghorbal may not explicitly disclose associating, by the one or more processors, a biometric feature type retrieved from the image with the cluster. However, Lyren (title, abstract, Figs. 1-16) teaches associating, by the one or more processors, a biometric feature type retrieved from the image with the cluster (Lyren ¶0294 discloses the user profile includes information pertaining to the characteristics and/or preferences of the user. Examples of this information for a person include, but are not limited to, one or more of personal data of the user [such as communication hardware and software used including peripheral devices such as head tracking systems, biometric data, physical measurements of their body and environments, functions of physical data such as HRTFs, etc.], photographs [such as photos of their head and ears. ¶0118 discloses facial recognition to determine or estimate a facial or head orientation of the person from one or more images or video captured with a camera in the HPED. The facial orientation can be described or recorded with respect to a location of the HPED that is generating the sound for the impulse responses). Eronen, Hirst, Ghorbal, and Lyren are analogous art as they pertain to customizing audio. Therefore it would have been obvious to someone of ordinary skill in the art before the effective filing date of the invention was made to modify the teachings of Eronen in view of Hirst and Ghorbal in light of the teachings of Lyren to capture image of video of the person in real-time using a camera (as taught by Lyren ¶0122) to overcome typical methods of obtaining HRTFs since it is measured in an anechoic chamber or specialized location requiring expensive sound equipment (Lyren ¶0002). Claim 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Eronen et al. (US #2018/0295463) in view of Hirst (US Patent #9544706) further in view of Lyren et al. (US #2017/0245081), Flaks et al. (US #2012/0093320), and Ghorbal et al. (US PGPUB #2018/0249275). Regarding Claim 17, Eronen in view of Hirst discloses the method of claim 1, but may not explicitly disclose further comprising: operating, by the one or more processors, as a multi-recognizer corresponding to a first type recognizer and a second type recognizer, wherein the image-related input audio signals correspond to biometric data associated with the subject, and wherein the biometric data comprises a first biometric feature type and a second biometric feature type; associating, by the one or more processors, the second biometric feature type that is the biometric feature type with a cluster; generating, by the one or more processors as the first type recognizer, a first set of intermediate processor datasets based on the first biometric feature type, and generating, by the one or more processors as the second type recognizer, a second set of intermediate processor datasets based on the second biometric feature type. However, Lyren (title, abstract, Figs. 1-16) teaches operating, by the one or more processors, as a multi-recognizer corresponding to a first type recognizer and a second type recognizer (Lyren ¶0218 discloses Alice receives a VoIP call from Bob while she is at her grandmother's house. Her smartphone determines that Alice has not previously received a call at this location and hence is unable to retrieve either a RIR or BRIR for her current location. In response to this determination, the smartphone rings with a distinctive tone, and Alice recognizes this tone and its implication that no RIRs or BRIRs are available for her location. This distinctive tone is actually the sound used to capture impulse responses. While her smartphone is ringing and generating this distinctive tone, Alice holds the smartphone in her hand with her arm stretched out away from her face. Microphones in her earphones record the tones, and her smartphone immediately generates BRIRs for Alice. When the smartphone captures sufficient impulse response from a designated location, it stops generating the tones, answers the call, and convolves the incoming voice with the BRIRs that it just obtained while Alice was answering the phone call), wherein the image-related input audio signals correspond to biometric data associated with the subject (Lyren ¶0294 discloses the user profile includes information pertaining to the characteristics and/or preferences of the user. Examples of this information for a person include, but are not limited to, one or more of personal data of the user [such as communication hardware and software used including peripheral devices such as head tracking systems, biometric data, physical measurements of their body and environments, functions of physical data such as HRTFs, etc.], photographs [such as photos of their head and ears. ¶0118 discloses facial recognition to determine or estimate a facial or head orientation of the person from one or more images or video captured with a camera in the HPED. The facial orientation can be described or recorded with respect to a location of the HPED that is generating the sound for the impulse responses), and wherein the biometric data comprises a first biometric feature type and a second biometric feature type (Lyren ¶0254 discloses one or more sensors 1430 such as biometric sensors); associating, by the one or more processors, the second biometric feature type that is the biometric feature type with a cluster (Lyren ¶0294 discloses the user profile includes information pertaining to the characteristics and/or preferences of the user. Examples of this information for a person include, but are not limited to, one or more of personal data of the user [such as communication hardware and software used including peripheral devices such as head tracking systems, biometric data, physical measurements of their body and environments, functions of physical data such as HRTFs, etc.], photographs [such as photos of their head and ears. ¶0118 discloses facial recognition to determine or estimate a facial or head orientation of the person from one or more images or video captured with a camera in the HPED. The facial orientation can be described or recorded with respect to a location of the HPED that is generating the sound for the impulse responses). Eronen, Hirst, and Lyren are analogous art as they pertain to customizing audio. Therefore it would have been obvious to someone of ordinary skill in the art before the effective filing date of the invention was made to modify the teachings of Eronen in view of Hirst in light of the teachings of Lyren to analyze using the HPED the target impulse that is the designated sound in order to create the BRIRs and HRTFs (as taught by Lyren, ¶0141) to overcome typical methods of obtaining HRTFs since it is measured in an anechoic chamber or specialized location requiring expensive sound equipment (Lyren ¶0002). Eronen in view of Hirst and Lyren may not explicitly disclose generating, by the one or more processors as the first type recognizer, a first set of intermediate processor datasets based on the first biometric feature type, and generating, by the one or more processors as the second type recognizer, a second set of intermediate processor datasets based on the second biometric feature type. However, Flaks (abstract, Figs. 1-12) teaches generating, by the one or more processors as the first type recognizer, a first set of intermediate processor datasets based on the first biometric feature type (Flaks Fig. 10: 1040 load reference HRTFs, get input from sensor(s), calculate pinna and head characteristics, look up optimal HRTF or interpolate. ¶0074 discloses head size and width. Specific HRTF can be associated with specific measurements related to the head size. Measurement can be head width. ¶0078 discloses in step 802 [Fig. 8], the listener 8 is identified using biometric information. Collecting biometric information can include collecting depth information and RGB information. ¶0079 discloses in step 804: select HRTF based on identified user), and generating, by the one or more processors as the second type recognizer, a second set of intermediate processor datasets based on the second biometric feature type (Flaks ¶0074 discloses pinna characteristics. Specific HRTF can be associated with specific measurements related to the pinna. ¶0078 discloses in step 802 [Fig. 8], the listener 8 is identified using biometric information. Collecting biometric information can include collecting depth information and RGB information. ¶0079 discloses in step 804: select HRTF based on identified user. Fig. 10: 1040 load reference HRTFs, get input from sensor(s), calculate pinna and head characteristics, look up optimal HRTF or interpolate). Eronen, Hirst, Lyren, and Flaks are analogous art as they pertain to customizing audio. Therefore it would have been obvious to someone of ordinary skill in the art before the effective filing date of the invention was made to modify the teachings of Eronen in view of Hirst and Lyren in light of the teachings of Flaks to determine the HRTF for the listener based on sensor determined characteristics, such as using the image camera component (as taught by Flaks, ¶0046) to improve and provide an accurate HRTF for the user and it should be accurate, consumer friendly, cost effective, and compatible with existing audio systems (Flaks, ¶0005). Eronen in view of Hirst, Lyren, and Flaks may not explicitly disclose associating, by the one or more processors, the biometric feature type with a cluster. However, Ghorbal (title, abstract, Figs. 1-4) teaches associating, by the one or more processors, the biometric feature type with a cluster (Ghorbal ¶0021 discloses starting with a substantial database of HRTFs, similar HRTFs are grouped together. The Euclidian distance naturally associated with this 16-dimensional space then allowed the HRTFs to be grouped into clusters [of 8 in number]. Sets of HRTFs were then randomly chosen within the clusters and subjects invited to choose the one or more clusters that gave them the best impression of externality and directivity. ¶0023 discloses once the cluster has been selected, another selecting step in which a very precise set is selected may be added. A procedure for selecting a set of HRTFs from 32 available HRTFs is described. ¶0072 discloses cluster equivalence will possibly be spoken of, all the points of a given cluster playing a similar role within the ear to which they belong. ¶0083 discloses alternatively, any method allowing the values of the set of parameters of the head-related transfer functions to be found from the values of the set of statistical parameters and ensuring a good reconstruction of the head-related transfer functions of the database OH1 can be used, examples of such methods being methods based on neural networks, based on multiple component analysis (MCA) or based on k-means clustering). Eronen, Hirst, Lyren, Flaks, and Ghorbal are analogous art as they pertain to customizing audio. Therefore it would have been obvious to someone of ordinary skill in the art before the effective filing date of the invention was made to modify the teachings of Eronen in view of Hirst. Lyren, and Flaks in light of the teachings of Ghorbal to group similar HRTFs together (as taught by Ghorbal, ¶0021) to generate an individual-specific HRTF more rapidly and with a higher reliability (Ghorbal, ¶0034). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to YOGESHKUMAR G PATEL whose telephone number is (571)272-3957. The examiner can normally be reached 7:30 AM-4 PM PST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Duc Nguyen can be reached at (571) 272-7503. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /YOGESHKUMAR PATEL/Primary Examiner, Art Unit 2691
Read full office action

Prosecution Timeline

Jan 07, 2025
Application Filed
Sep 24, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12741594
VEHICLE BODY PANELS FOR SPEAKERS, SPEAKER ASSEMBLIES WITH VEHICLE BODY PANELS, AND INSTALLATION METHODS THEREFORE
2y 6m to grant Granted Sep 22, 2026
Patent 12743247
ACOUSTIC DEVICE, PROGRAM, AND CONTROL METHOD
2y 4m to grant Granted Sep 22, 2026
Patent 12739592
Visual Platforms for Configuring Audio Processing Operations
2y 9m to grant Granted Sep 15, 2026
Patent 12737354
LARGE LANGUAGE MODEL-BASED COMMUNICATION CONTENT GENERATION
2y 6m to grant Granted Sep 15, 2026
Patent 12737542
COMPUTER-IMPLEMENTED METHOD FOR OBTAINING AT LEAST ONE NUMBER
1y 11m to grant Granted Sep 15, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
84%
Grant Probability
87%
With Interview (+3.2%)
2y 3m (~6m remaining)
Median Time to Grant
Low
PTA Risk
Based on 678 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month