Prosecution Insights
Last updated: August 18, 2026
Application No. 18/678,202

MULTI-SOURCE AUDIO COMMUNICATION SYSTEM

Final Rejection §103
Filed
May 30, 2024
Examiner
TIEU, BINH KIEN
Art Unit
2694
Tech Center
2600 — Communications
Assignee
Microsoft Technology Licensing, LLC
OA Round
2 (Final)
87%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
97%
With Interview

Examiner Intelligence

Grants 87% — above average
87%
Career Allowance Rate
826 granted / 947 resolved
+25.2% vs TC avg
Moderate +10% lift
Without
With
+9.6%
Interview Lift
resolved cases with interview
Typical timeline
2y 3m
Avg Prosecution
10 currently pending
Career history
961
Total Applications
across all art units

Statute-Specific Performance

§101
6.9%
-33.1% vs TC avg
§103
45.0%
+5.0% vs TC avg
§102
26.8%
-13.2% vs TC avg
§112
2.6%
-37.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 947 resolved cases

Office Action

§103
DETAILED ACTION Response to Amendment The Applicants’ amendment, filed 06/02/2026, was received and entered. As the results, independent claims 1, 13 and 19 were amended to include limitations of or similar to the features, as stated followings: “prioritizing audio received from the first microphone over other audio received from the plurality of microphones based on the first weight, wherein prioritizing the audio received from the first microphone comprises causing the audio received from the first microphone to be provided at an increased clarity over the other audio received from the plurality of microphones”. Based on the above amended features added to the independent claims, Examiner performed updated searches and a new reference was found, Couse et al. (US 8,989,360), which teaches the above features. Therefore, the new ground of rejections are set forth below. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1-3, 5-8, 10, 13-15, 17 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Ojanpera (US 2012/0310396 as cited in the previous Office Action) in view of Couse et al. (US 8,989,360). Regarding claim 1, Ojanpera teaches a method, comprising: receiving an indication of a virtual meeting session between a first computing device at a first meeting location and a second computing device at a second meeting location, wherein the first computing device is connected to a plurality of microphones positioned in the first meeting location (i.e., an audio space 1 (as a first meeting location), as shown in figure 1, comprising a plurality of recording devices 10-1…10-12 have been deployed to record an audio scene and comprised or connected to microphones, such as omnidirectional microphones, etc. (para.[0112]); an end-to-end system 2, as shown in figure 2, comprising a plurality of recording devices 20 (corresponding to recording devices 10-1…10-12 of Fig.1), an audio scene server 21, and render device 22 (read on a second computing device located at a second meeting location; para. [0114]); the recording devices 20 record an audio scene in the audio space 1 at different positions and the audio scene server 21 may receive the audio signals (as an indication of a virtual meeting session, such as audio/video conferencing or the like; para.[0002]) recorded by the recording devices 20 and keep track of the positions and the associated directions/orientations (para.[0117])); providing, to the second computing device, an audio map of the first meeting location (i.e., the audio scene server 21 provides high level coordinates, which corresponding to locations where uploaded or up-streamed content is available for listening, these high level coordinates (of recording devices 20 at the audio space 1) may be provided as a map to user of rendering device 22; para.[0118]); receiving, from the second computing device, a selection of a first virtual listening position on the audio map (i.e., the user of the rendering device 22 at the second position is allowed to select a desired listening position in the provided map and information on this desired listening position may then be provided to the audio scene server 21; para.[0118]); correlating the selected first virtual listening position to a first microphone of the plurality of microphones by assigning a first weight to the first microphone based on a distance from the first virtual listening position to the first microphone (i.e., a set of recording devices is derived from a plurality of recording devices; para.[0142]; a number of recording devices included in the set of recording devices may be determined based value of m wherein m represents a number of recording devices from 0 to maximum amount of recording devices (variable M); para.[0171], [0173] and [0179]; thus, assume the m=1 which represented a recording device 20 or the first computing device; also, R indicates the maximum estimated distance of a recording device from the desired listening position; para.[0179]; then a relevance level (read on first weight) is determined or assigned to the recording device in the set of recording devices 20 based on the distance of a recording device to the desired listening position; para.[0181], [0211] and [0212]); receiving, from the first computing device, audio from the plurality of microphones in the first meeting location (i.e., receiving audio signals recorded by the recording devices 20; para.[0117]); prioritizing audio received from the first microphone over other audio received from the plurality of microphones based on the first weight (i.e., only the audio signals recorded by the recording device in the set of the recording devices are combined into a combined audio signal to be rendered; para.[0123]); and providing the prioritized audio to the second computing device (i.e., the combined audio signal or prioritized audio being provided/forwarded to and/or rendered by the rendering device 22; para.[0114], [0119] and [0127]). It should be noticed that Ojanpera failed to teach the features of prioritizing audio received from the first microphone over other audio received from the plurality of microphones based on the first weight, wherein prioritizing the audio received from the first microphone comprises causing the audio received from the first microphone to be provided at an increased clarity over the other audio received from the plurality of microphones, as amended and argued by Applicants in the remarks. However, Couse et al. (hereinafter “Couse”) teaches a conference phone 100, as shown in figure 1. The phone 100 comprises a light bar 106 which is divided into a plurality of sections representing audio directions from a plurality of microphone, configured corresponding to each of the divided sections of the light bar 106, to receive audio (col.3, lines 17-24). The conference phone 100 is configured to allow a host (a speaker of the meeting or conference) to active a “host mode” in which the host section can be configured to detect audio from the direction of the host section having a weight, such as a selected threshold. If the detected audio of the host section has a threshold (weight) greater than the selected threshold, the received audio from the host section is amplified and transmitted and communicated via a (conference) telephone call to the other side of the telephone call (col.5, lines 22-44). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the feature of prioritizing audio received from the first microphone over other audio received from the plurality of microphones based on the first weight, wherein prioritizing the audio received from the first microphone comprises causing the audio received from the first microphone to be provided at an increased clarity over the other audio received from the plurality of microphones, as amended and argued by Applicants in the remarks, as taught by Couse, into view of Ojanpera in order to notify the others (participants) in the conference the voice of presenting speaker. Regarding claim 2, Ojanpera further teaches the combined audio signals can be a stereo, binaural, etc. as an increased level outputting at the rendering device 22 (para. [0127] and [0236]). Regarding claim 3, Ojanpera further teaches the recording devices included in the set of the cording devices, which comprises one or more (e.g., at least two) recording devices (para.[0025]). Ojanpera further teaches that the selected recording devices (of a set of recording devices) are selected based on relevance levels (weights), wherein the relevance levels may be determined so that each recording device (i.e., first recording device or microphone, second microphone, etc.) is assigned a respective relevant level (i.e., assigning a second weight to the second microphone; para.[0028]). Finally, Ojanpera further teaches the feature of prioritizing the audio received from the second microphone, such as obtaining one or more combined audio signals (as prioritized audio signal) because the combined audio signal may be the same recorded audio signal applied to a case of only a recording device in the set of recording devices being selected which is closed to the desired listing position; para.[0030] and [0033]). Regarding claim 5, Ojanpera further teaches the map comprising: location information about a physical location of the microphone (i.e., position information of the recording device (para.[0116]); and location information about an audio zone (i.e., audio scene) corresponding to a spatial area (i.e., a desired listening position) within which the microphone captures audio (para.[0117]-[0119]). Regarding claim 6, Ojanpera further teaches limitations of the clam in paragraph [0119]. Regarding claim 7, Ojanpera further teaches limitations of the clam, such as providing a map with high level coordinates to the user of rending device 22 in paragraph [0118]. Regarding claim 8, Ojanpera further teaches limitations of the clam, such as receiving a selection of a virtual listening position (i.e., user selecting a desired listening position, para.[0118]); correlating the selected virtual listening position to a microphone (i.e., determining the desired listening position and providing the information on this desired listening position to the audio scene server 21, para.[0118]-[0119]); receiving audio (i.e., receiving audio signals from microphone(s) located near to the desired listening position and recording one or more audio signals at a time; para.[0116] and [0119]); prioritizing the audio (i.e., combining the received audio signals received from the nearer microphone(s); para. [0119] and [0123]); and providing the prioritized audio to the second computing device (i.e., providing a combined audio signal to the rending device 22; para. [0119] and [0123]). Regarding claim 10, Ojanpera further teaches limitations of the clam in paragraphs [0039] and [0118]. Regarding claim 13, Ojanpera teaches a system (i.e., audio space system 1 or an end-to-end system 2, as shown in figures 1 and 2), comprising: a processing system (i.e., processor 30 of an audio scene server 21); and memory comprising computer executable instructions (i.e., program memory 31 and/or main memory 32) that, when executed, perform operations (para.[0128]-[0131]) comprising: receiving an indication of a virtual meeting session between a first computing device at a first meeting location and a second computing device at a second meeting location, wherein the first computing device is connected to a plurality of microphones positioned in the first meeting location (i.e., an audio space 1 (as a first meeting location), as shown in figure 1, comprising a plurality of recording devices 10-1…10-12 have been deployed to record an audio scene and comprised or connected to microphones, such as omnidirectional microphones, etc. (para.[0112]); an end-to-end system 2, as shown in figure 2, comprising a plurality of recording devices 20 (corresponding to recording devices 10-1…10-12 of Fig.1), an audio scene server 21, and render device 22 (read on a second computing device located at a second meeting location; para. [0114]); the recording devices 20 record an audio scene in the audio space 1 at different positions and the audio scene server 21 may receive the audio signals (as an indication of a virtual meeting session, such as audio/video conferencing or the like; para.[0002]) recorded by the recording devices 20 and keep track of the positions and the associated directions/orientations (para.[0117])); providing, to the second computing device, an audio map of the first meeting location (i.e., the audio scene server 21 provides high level coordinates, which corresponding to locations where uploaded/up streamed content is available for listening, these high level coordinates (of recording devices 20 at the audio space 1) may be provided as a map to user of rendering device 22; para.[0118]); receiving, from the second computing device, a selection of a first virtual listening position on the audio map (i.e., the user of the rendering device 22 at the second position is allowed to select a desired listening position in the provided map and information on this desired listening position may then be provided to the audio scene server 21; para.[0118]); correlate the virtual listening position to a first microphone and a second microphone of the plurality of microphones (para.[0037], [0039], [0041]); correlating the selected first virtual listening position to a first microphone of the plurality of microphones by assigning a first weight to the first microphone based on a receiving, from the first computing device, audio from the plurality of microphones in the first meeting location (i.e., receiving audio signals recorded by the recording devices 20; para.[0117]); prioritizing audio received from the first microphone over other audio received from the plurality of microphones based on the first weight (i.e., only the audio signals recorded by the recording device in the set of the recording devices are combined into a combined audio signal to be rendered; para.[0123]); and providing the prioritized audio to the second computing device (i.e., the combined audio signal or prioritized audio being provided/forwarded to and/or rendered by the rendering device 22; para.[0114], [0119] and [0127]). It should be noticed that Ojanpera failed to teach the features of prioritizing audio received from the first microphone over other audio received from the plurality of microphones based on the first weight, wherein prioritizing the audio received from the first microphone comprises causing the audio received from the first microphone to be provided at an increased clarity over the other audio received from the plurality of microphones, as amended and argued by Applicants in the remarks. However, Couse teaches a conference phone 100, as shown in figure 1. The phone 100 comprises a light bar 106 which is divided into a plurality of sections representing audio directions from a plurality of microphone, configured corresponding to each of the divided sections of the light bar 106, to receive audio (col.3, lines 17-24). The conference phone 100 is configured to allow a host (a speaker of the meeting or conference) to active a “host mode” in which the host section can be configured to detect audio from the direction of the host section having a weight, such as a selected threshold. If the detected audio of the host section has a threshold (weight) greater than the selected threshold, the received audio from the host section is amplified and transmitted and communicated via a (conference) telephone call to the other side of the telephone call (col.5, lines 22-44). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the feature of prioritizing audio received from the first microphone over other audio received from the plurality of microphones based on the first weight, wherein prioritizing the audio received from the first microphone comprises causing the audio received from the first microphone to be provided at an increased clarity over the other audio received from the plurality of microphones, as taught by Couse, into view of Ojanpera in order to notify the others (participants) in the conference the voice of presenting speaker. Regarding claim 14, Ojanpera further teaches the combined audio signals can be a stereo, binaural, etc. as an increased clarity outputting at the rendering device 22 (para. [0127] and [0236]). Regarding claim 15, Ojanpera further the map comprising: location information about a physical location of the microphone (i.e., position information of the recording device (para.[0116]); and location information about an audio zone (i.e., audio scene) corresponding to a spatial area (i.e., a desired listening position) within which the microphone captures audio (para.[0117]-[0119]); and correlating the virtual listening position to the first microphone comprises determining the virtual listening position is included in a first audio zone that corresponds to the first microphone (i.e., determining which of the recording device is located near the desired listening position; para.[0119]). Regarding claim 17, Ojanpera further teaches limitations of the claim in paragraph [0118]. Regarding claim 19, Ojanpera teaches a computer-readable medium storing instructions that, when executed by a computer (i.e., program memory 31 and/or main memory 32; para.[0128]-[0131]), cause the computer to: receive an indication of a virtual meeting session between a first computing device at a first meeting location and a second computing device at a second meeting location, wherein the first computing device is connected to a plurality of microphones positioned in the first meeting location (i.e., an audio space 1 (as a first meeting location), as shown in figure 1, comprising a plurality of recording devices 10-1…10-12 have been deployed to record an audio scene and comprised or connected to microphones, such as omnidirectional microphones, etc. (para.[0112]); an end-to-end system 2, as shown in figure 2, comprising a plurality of recording devices 20 (corresponding to recording devices 10-1…10-12 of Fig.1), an audio scene server 21, and render device 22 (read on a second computing device located at a second meeting location; para. [0114]); the recording devices 20 record an audio scene in the audio space 1 at different positions and the audio scene server 21 may receive the audio signals (as an indication of a virtual meeting session, such as audio/video conferencing or the like; para.[0002]) recorded by the recording devices 20 and keep track of the positions and the associated directions/orientations (para.[0117])); provide, to the second computing device, an audio map of the first meeting location (i.e., the audio scene server 21 provides high level coordinates, which corresponding to locations where uploaded/upstreamed content is available for listening, these high level coordinates (of recording devices 20 at the audio space 1) may be provided as a map to user of rendering device 22; para.[0118]); receive, from the second computing device, a selection of a first virtual listening position on the audio map (i.e., the user of the rendering device 22 at the second position is allowed to select a desired listening position in the provided map and information on this desired listening position may then be provided to the audio scene server 21; para.[0118]); correlate the selected first virtual listening position to a first microphone of the plurality of microphones by assigning a first weight to the first microphone based on a distance from the first virtual listening position to the first microphone (i.e., a set of recording devices is derived from a plurality of recording devices; para.[0142]; a number of recording devices included in the set of recording devices may be determined based value of m wherein m represents a number of recording devices from 0 to maximum amount of recording devices (variable M); para.[0171], [0173] and [0179]; thus, assume the m=1 which represented a recording device 20 or the first computing device; also, R indicates the maximum estimated distance of a recording device from the desired listening position; para.[0179]; then a relevance level (read on first weight) is determined or assigned to the recording device in the set of recording devices 20 based on the distance of a recording device to the desired listening position; para.[0181], [0211] and [0212]); receive, from the first computing device, audio from the plurality of microphones in the first meeting location (i.e., receiving audio signals recorded by the recording devices 20; para.[0117]); prioritize audio received from the first microphone over other audio received from the plurality of microphones based on the first weight (i.e., only the audio signals recorded by the recording device in the set of the recording devices are combined into a combined audio signal to be rendered; para.[0123]); and provide the prioritized audio to the second computing device (i.e., the combined audio signal or prioritized audio being provided/forwarded to and/or rendered by the rendering device 22; para.[0114], [0119] and [0127]). It should be noticed that Ojanpera failed to teach the features of prioritizing audio received from the first microphone over other audio received from the plurality of microphones based on the first weight, wherein prioritizing the audio received from the first microphone comprises causing the audio received from the first microphone to be provided at an increased clarity over the other audio received from the plurality of microphones, as amended and argued by Applicants in the remarks. However, Couse teaches a conference phone 100, as shown in figure 1. The phone 100 comprises a light bar 106 which is divided into a plurality of sections representing audio directions from a plurality of microphone, configured corresponding to each of the divided sections of the light bar 106, to receive audio (col.3, lines 17-24). The conference phone 100 is configured to allow a host (a speaker of the meeting or conference) to active a “host mode” in which the host section can be configured to detect audio from the direction of the host section having a weight, such as a selected threshold. If the detected audio of the host section has a threshold (weight) greater than the selected threshold, the received audio from the host section is amplified and transmitted and communicated via a (conference) telephone call to the other side of the telephone call (col.5, lines 22-44). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the feature of prioritizing audio received from the first microphone over other audio received from the plurality of microphones based on the first weight, wherein prioritizing the audio received from the first microphone comprises causing the audio received from the first microphone to be provided at an increased clarity over the other audio received from the plurality of microphones, as taught by Couse, into view of Ojanpera in order to notify the others (participants) in the conference the voice of presenting speaker. Claims 4 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Ojanpera (US 2012/0310396) in view of Couse et al. (US 8,989,360) as applied to claims 1, 3 and 19 above, and further in view of Oates, III et al. (US 9,843,881, also cited in the previous Office Action). Regarding claim 4, Ojanpera and Couse, in combination, teaches all subject matters as claimed above, except for the features of causing the audio received from the first microphone to be output at a first level based on the first weight; and causing the audio received from the second microphone to be output at a second level based on the second weight. However, Oates, III et al. (hereinafter “Oates, III”) teach such features col.5, lines 45-56 and col.13, line 57 through col.14, line 10. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the features of causing the audio received from the first microphone to be output at a first level based on the first weight; and causing the audio received from the second microphone to be output at a second level based on the second weight, as taught by Oates, III, into view of Ojanpera and Couse in order to provide better sound according to the position and weight of the microphone. Regarding claim 20, Ojanpera further teaches the recording devices included in the set of the cording devices, which comprises one or more (e.g., at least two) recording devices (para.[0025]). Ojanpera further teaches that the selected recording devices (of a set of recording devices) are selected based on relevance levels (weights), wherein the relevance levels may be determined so that each recording device (i.e., first recording device or microphone, second microphone, etc.) is assigned a respective relevant level (i.e., assigning a second weight to the second microphone; para.[0028]). Finally, Ojanpera further teaches the feature of prioritizing the audio received from the second microphone, such as obtaining one or more combined audio signals (as prioritized audio signal) because the combined audio signal may be the same recorded audio signal applied to a case of only a recording device in the set of recording devices being selected which is closed to the desired listing position; para.[0030] and [0033]). Ojanpera failed to teach the features of causing the audio received from the first microphone to be output at a first level based on the first weight; and causing the audio received from the second microphone to be output at a second level based on the second weight. However, Oates, III et al. (hereinafter “Oates, III”) teach such features col.5, lines 45-56 and col.13, line 57 through col.14, line 10. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the features of causing the audio received from the first microphone to be output at a first level based on the first weight; and causing the audio received from the second microphone to be output at a second level based on the second weight, as taught by Oates, III, into view of Ojanpera in order to provide better sound according to the position and weight of the microphone. Claims 9 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Ojanpera (US 2012/0310396) in view of Couse et al. (US 8,989,360), as applied to claims 1 and 13 above, and further in view of Pavlovsky et al. (US 2024/0194177, also cited in the previous Office Action). Regarding claims 9 and 18, Ojanpera teaches all subject matters as claimed above, combining the audio signals and providing the combined audio signal to the rendering device as discussed above. Ojanpera failed to teach the feature of transmitting the audio signals by use of a language translation service prior to transmit the translated audio signals in a second langue to the rending device, as well-known to those skilled in the art. However, Pavlovsky et al. (hereinafter “Pavlovsky”) teaches a communication service 100, as shown in figure 1, comprising translation service 106 to translate a first language spoken (in English) by a speaker into a second language (in France) spoken by a second user (para.[0019]-[0020] and [0028]-[0029]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the features of prior to providing the prioritized audio to the second computing device: providing the prioritized audio to a language translation service; and receiving, from the language translation service, the prioritized audio translated from a first language into a second language; and providing the prioritized audio to the second computing device comprises providing the translated prioritized audio to the second computing device, as taught by Pavlovsky, into view of Ojanpera, in order to provide the proper language spoken by the second user associated with the second computing device. Claims 11-12 are rejected under 35 U.S.C. 103 as being unpatentable over Ojanpera (US 2012/0310396) in view of Couse et al. (US 8,989,360) as applied to claims 1 and 10 above, and further in view of Allen et al. (US 11,915,483, also cited in the previous Office Action). Regarding claim 11, Ojanpera and Couse, in combination, teach all subject matters as claimed above, except for the features of receiving the location information about the physical location of the meeting participant in the first meeting location comprises identifying the meeting participant using at least one of: facial recognition on a video of the virtual meeting session; voice recognition on the received audio; an identification badge; or manual input of the meeting participant’s identity. However, Allen et al. (hereinafter “Allen”) teaches such features in col.16, line 52 through col.17, line 15 for a purpose of identifying the specific participant in the video conference. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the features of receiving the location information about the physical location of the meeting participant in the first meeting location comprises identifying the meeting participant using at least one of: facial recognition on a video of the virtual meeting session; voice recognition on the received audio; an identification badge; or manual input of the meeting participant’s identity, as taught by Allen, into view of Ojanpera and Couse in order to identifying the specific in according to the information location. Regarding claim 12, Allen further teaches a video conference system 400, as shown in figure 4, comprising conference device 1 and conference device 2, wherein each of the devices 1 (410A) and 2 (410B) has its own components, such as components 412A-412C and components 414A-414C connected to the devices 1 and 2, respectively. Each of the conference devices 410A and 410B could be operated by one or more users in a physical space which is the conference device 410A located in the first meeting location and the conference device 410B located in a second meeting location (e.g., a classroom, a conference room, etc.). Allen further teaches a server device 420 supporting a video conference between participants using the conference devices 410A and 410B (col.11, lines 20-52). Allen further teaches the components 412A through 412C and 414A through 414C may include microphones, speakers, etc. (col.12, lines 13-21). Allen further teaches the system to detect and to identify “participant 1” and ”participant 2” as multiple participants using the conference device 410A in a same physical space (i.e., the first meeting location), and may identify “participant 3” as a single participant using the conference device 410B in a difference physical space (i.e., the second meeting location)(col.13, lines 15-18). Allen further teaches the service device 420 to determine the preferences and/or priorities, wherein priority may include a relative ranking or importance of a functionality to a particular person (col.13, lines 32-54). By determining the preferences and/or priorities, the device (e.g., the conference device 410A, the conference device 410B, or the server device 420)(col.12, lines 30-35) may then execute the configuration software to determine a configuration to implement one or more functionality of the component (i.e., speaker) that may be desired (i.e., volume of a speaker or gain or mute of a microphone, etc.; applied during the video conference, etc.)(col.17, line 48-67). Allen further teaches the participant 1 and participant 2 used the conference device 410A which is located a physical location (the first meeting location) and the participant 3 used the conference device which is located in the different location (the second meeting location)(col.17, lines 28-34). Assume that the participant 3 performs chat (i.e., voice conversation to one or both of participants 1 and 2; col.19, line 67 through col.20, line 15), the audio received from the participant 3 at the microphone of the conference device 410B to be outputted by the speaker of the conference device 410A to one of the participants 1 and 2 having the applied configuration included higher priority, such as increasing volume of the speaker 512C (col.20, lines 32-64). Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over Ojanpera (US 2012/0310396) in view of Couse et al. (US 8,989,360) as applied to claim 13 above, and further in view of Allen et al. (US 11,915,483) and Oates, III et al. (US 9,843,881)(Both Allen et al. and Oates, III et al. references were cited in the previous Office Action). Regarding claim 16, Ojanpera and Couse, in combination, teach all subject matters as claimed above. Ojanpera further teaches the features of correlating the virtual listening position to the first microphone and the second microphone, comprising: assigning a first weight to the first microphone based on a distance from the virtual listening position to the first microphone (i.e., a set of recording devices is derived from a plurality of recording devices; para.[0142]; a number of recording devices included in the set of recording devices may be determined based value of m wherein m represents a number of recording devices from 0 to maximum amount of recording devices (variable M); para.[0171], [0173] and [0179]; thus, assume the m=1 which represented a recording device 20 or the first computing device; also, R indicates the maximum estimated distance of a recording device from the desired listening position; para.[0179]; then a relevance level (read on first weight) is determined or assigned to the recording device in the set of recording devices 20 based on the distance of a recording device to the desired listening position; para.[0181], [0211] and [0212]); and assigning a second weight to the second microphone based on a distance from the virtual listening position to the second microphone (i.e., the selected recording devices (of a set of recording devices) are selected based on relevance levels (weights), wherein the relevance levels may be determined so that each recording device (i.e., first recording device or microphone, second recording device or second microphone, etc.) is assigned a respective relevant level (i.e., assigning a second weight to the second microphone; para.[0028]); and Ojanpera failed to teach determining the virtual listening position is in a first audio zone corresponding to the first microphone and a second audio zone corresponding to a second microphone of the plurality of microphones. However, Allen teaches the features of determining the virtual listening position is in a first audio zone corresponding to the first microphone (i.e., the device may detect one or more participants (participants 1 and/or 2) as a person in a particular geographic region or in a first audio zone, as shown in figure 5) and a second audio zone corresponding to a second microphone of the plurality of microphones (i.e., the participant 3 in a different physical space or in a second audio zone as shown in figure 3)(col.13, lines 1-18). It should be also noticed that Ojanpera and Allen, in combination failed to teaches the features of prioritizing the audio received from the first microphone and the second microphone, comprising the feature of causing the audio received from the first microphone to be output at a first level based on the first weight and the feature of causing the audio received from the second microphone to be output at a second level based on the second weight. However, Oates, III further teaches the feature in col.5, lines 45-56 and col.13, line 57 through col.14, line 10. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the features as discussed above, as taught by Allen and Oates into view of Ojanpera and Cause in order to receive the best audio signals generated from the microphones of the plurality of microphones. Response to Arguments Applicant’s arguments with respect to claims 1-20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for response to this final action is set to expire THREE MONTHS from the date of this action. In the event a first response is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event will the statutory period for response expire later than SIX MONTHS from the date of this final action. Any response to this final action should be mailed to: BOX AF Commissioner of Patents and Trademarks Washington, D.C. 20231 Or faxed to: (703) 872-9314 or (301) 273-8300 (for formal communications; Please mark “EXPEDITED PROCEDURE”) Or: If it is an informal or draft communication, please label “PROPOSED” or “DRAFT”) Any inquiry concerning this communication or earlier communications from the examiner should be directed to BINH TIEU whose telephone number is (571)272-7510. The examiner can normally be reached on 9-5. The Examiner’s fax number is (571) 273-7510 and E-mail address: BINH.TIEU@USPTO.GOV. It should be noticed that interview attribute time per new application or RCE (utility) is available when, during prosecution, the examiner conducts an interview. When more than one interview is needed in an application, supervisors have the flexibility to approve additional time to advance prosecution. Examiner interviews are available via telephone or video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (FAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the FAIR system, see fitp://nair-direct.usoto.aqev. If you have any questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). /Binh Kien Tieu/Primary Examiner, Art Unit 2694 Date: July 2026
Read full office action

Prosecution Timeline

May 30, 2024
Application Filed
Jan 15, 2026
Non-Final Rejection mailed — §103
Apr 29, 2026
Interview Requested
May 22, 2026
Applicant Interview (Telephonic)
May 22, 2026
Response Filed
May 27, 2026
Examiner Interview Summary
Aug 04, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12676943
Enhanced Caller Identification
2y 7m to grant Granted Jul 07, 2026
Patent 12647757
SYSTEM AND METHOD FOR TRIGGERING ON PLATFORM USAGE
3y 0m to grant Granted Jun 02, 2026
Patent 12627758
CALL ENHANCEMENT SERVICE VIA IN-NETWORK BRANDED CALLING DELIVERY
2y 11m to grant Granted May 12, 2026
Patent 12603111
AUDIO GUESTBOOK SYSTEMS AND METHODS
2y 4m to grant Granted Apr 14, 2026
Patent 12598223
Dynamic Teleconference Content Item Distribution to Multiple Devices Associated with a User
2y 6m to grant Granted Apr 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
87%
Grant Probability
97%
With Interview (+9.6%)
2y 3m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 947 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month