Prosecution Insights
Last updated: August 17, 2026
Application No. 18/587,771

CONTROL APPARATUS FOR CAUSING SPEAKER TO REPRODUCE SOUND CORRESPONDING TO LISTENING POINT, CONTROL METHOD, AND NON-TRANSITORY COMPUTER READABLE MEDIUM

Non-Final OA §103
Filed
Feb 26, 2024
Priority
Feb 28, 2023 — JP 2023-029726
Examiner
ZHU, QIN
Art Unit
2691
Tech Center
2600 — Communications
Assignee
Canon Inc.
OA Round
3 (Non-Final)
88%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
91%
With Interview

Examiner Intelligence

Grants 88% — above average
88%
Career Allowance Rate
553 granted / 631 resolved
+25.6% vs TC avg
Minimal +3% lift
Without
With
+3.0%
Interview Lift
resolved cases with interview
Fast prosecutor
1y 11m
Avg Prosecution
25 currently pending
Career history
652
Total Applications
across all art units

Statute-Specific Performance

§101
4.7%
-35.3% vs TC avg
§103
46.0%
+6.0% vs TC avg
§102
17.8%
-22.2% vs TC avg
§112
17.7%
-22.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 631 resolved cases

Office Action

§103
DETAILED ACTION This action is in response to communications filed 5/19/2026: Claims 1-8 and 10-18 are pending Claim 9 is cancelled Claims 14- 18 are added Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant’s arguments with respect to claim(s) 1-13 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Response to Amendment Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-3, 5-8, and 10-15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kumagai et al (JP2009044261, translated by EPO, hereinafter “Kumagai”) in view of Lang et al (US20240284137, hereinafter “Lang”) in further view of Kronlachner (US20240056758). Regarding claim 1, Kumagai teaches a control apparatus (¶1, apparatus) comprising: one or more memories storing instructions (¶20, memory 65); and one or more processors (¶20, DSP) executing the instructions to: obtain a first position of a listening point (¶13, generating optimal audio requiring a listening position); calculate a respective position of each of one or more virtual sound sources with respect to the first position of the listening point (¶13, localizing a plurality of virtual sound sources around the determined listening position); calculate positions of a plurality of speakers with respect to the first position of the listening point, the plurality of speakers being located around the listening point (¶13, determining position information of the plurality of speakers with respect to the listening position and also positioning speakers around the listening environment (including the user) (see Fig. 3)); generate a respective output signal to be output to each of the plurality of speakers based on (i) one or more sound source signals each output from the one or more virtual sound sources, (ii) the respective position of each of the one or more virtual sound sources with respect to the first position of the listening point, (iii) the positions of the plurality of speakers with respect to the first position of the listening point (¶14, outputting plurality of audio signals to each respective speaker based on the position of the user, sound output position, and position of the speakers); cause each of the plurality of speakers to reproduce sound corresponding to the respective output signal (Fig. 7, outputting audio to be reproduced); Kumagai fails to explicitly teach create, based on the calculated positions of the plurality of speakers with respect to the first position of the listening point, a first plurality of speaker sets each including adjacent speakers as viewed from the first position of the listening point; and (iv) the first plurality of speaker sets; detect movement of the listening point to a second position; calculate (i) the respective position of each of the one or more virtual sound sources with respect to the second position of the listening point and (ii) the positions of the plurality of speakers with respect to the second position of the listening point; create, based on the calculated positions of the plurality of speakers with respect to the second position of the listening point, a second plurality of speaker sets each including adjacent speakers as viewed from the second position of the listening point; and regenerate the output signal to be output to each of the plurality of speakers based on (i) the calculated respective position of each of the one or more virtual sound sources with respect to the second position of the listening point, (ii) the calculated positions of the plurality of speakers with respect to the second position of the listening point and (iii) the second plurality of speaker sets. Lang teaches detect movement of the listening point to a second position (¶46, user’s positions are track in real-time to update the output audio); calculate (i) the respective position of each of the one or more virtual sound sources with respect to the second position of the listening point and (ii) the positions of the plurality of speakers with respect to the second position of the listening point (¶50-53, different user location may affect virtual playback format which defines virtual speaker positions (and thus virtual source positions)); and regenerate the output signal to be output to each of the plurality of speakers based on (i) the calculated respective position of each of the one or more virtual sound sources with respect to the second position of the listening point and (ii) the calculated positions of the plurality of speakers with respect to the second position of the listening point (¶46, 50-53, user position is tracked in real-time to cause an update on the playback of the rendered audio wherein the update on the playback is further affected by the user’s current position and virtual source/speaker positions). Lang fails to explicitly teach create, based on the calculated positions of the plurality of speakers with respect to the first position of the listening point, a first plurality of speaker sets each including adjacent speakers as viewed from the first position of the listening point; and (iv) the first plurality of speaker sets; create, based on the calculated positions of the plurality of speakers with respect to the second position of the listening point, a second plurality of speaker sets each including adjacent speakers as viewed from the second position of the listening point; and (iii) the second plurality of speaker sets. Kronlachner teaches create, based on the calculated positions of the plurality of speakers with respect to the first position of the listening point, a first plurality of speaker sets each including adjacent speakers as viewed from the first position of the listening point; and (iv) the first plurality of speaker sets (¶93, VBAP/DBAP is utilized to generate audio output signals for objects located within a region/sub-region; pairwise-based panning approaches can then be utilized by pairs of cells (or three cells when an overhead cells is present)); create, based on the calculated positions of the plurality of speakers with respect to the second position of the listening point, a second plurality of speaker sets each including adjacent speakers as viewed from the second position of the listening point; and (iii) the second plurality of speaker sets (¶68-71, object audio output can be based on a listener position such that when a user changes/shifts their position, the user perceives the sound scene to “follow” based on their movement – e.g. at a first position, the user may hear objects A and B coming from a L and R speaker of the user but if the user turns to their right 90 degrees then the objects A and B may come instead from the C and RR speaker such that the objects are still perceived as coming from a “left” and “right” of the user). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply the method of tracking a user’s positioning (as taught by Lang) to the audio reproduction system (as taught by Kumagai). The rationale to do so is to apply a known technique to a known device ready for improvement to yield the predictable result of updating audio output in accordance with a user’s position in order to realize better localized audio output (Lang, ¶2). It would have been further obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply the method of spatial audio rendering (as taught by Kronlachner) to the audio reproduction system (as taught by Kumagai in view of Lang). The rationale to do so is to apply a known technique to a known device ready for improvement to yield the predictable result of outputting high fidelity audio output despite speaker arrangement (Kronlachner, ¶72). Regarding claim 2, Kumagai in view of Lang in further view of Kronlachner teaches wherein the one or more processors further execute the instructions to generate the respective output signal to be output to each of the plurality of speakers by distributing the one or more sound source signals to the plurality of speakers based on the respective position of each of the one or more virtual sound sources with respect to the first position of the listening point and the positions of the plurality of speakers with respect to the first position of the listening point (Kumagai, ¶69, distributing sound to one or more speakers on the basis of speaker position, virtual sound position, and receiving point/user position). Regarding claim 3, Kumagai in view of Lang further view of Kronlachner teaches wherein the one or more processors further execute the instructions to select a set of speakers to which a sound source signal, of the one or more sound source signals, is distributed, from among the plurality of speakers, based on the respective position of each of the one or more virtual sound sources with respect to the first position of the listening point and the positions of the plurality of speakers with respect to the first position of the listening point, and to distribute the sound source signal to the selected set of speakers (Kumagai, ¶69, selecting a subset of speakers from the plurality of speakers (e.g. 4 out of the 8 shown in Fig. 3) to output audio from and to apply a distribution of audio to the selected speakers). Regarding claim 5, Kumagai in view of Lang further view of Kronlachner teaches wherein the one or more processors further execute the instructions to correct the respective output signal to be output to each of the plurality of speakers based on a distance between the listening point and each of the plurality of speakers (Kumagai, ¶30, 49, applying one or more correction to the audio output based on at least a distance parameter). Regarding claim 6, Kumagai in view of Lang further view of Kronlachner teaches wherein the one or more processors further execute the instructions to perform correction processing related to at least one of a sound pressure and a delay (¶76, sense of distance can be altered by adjusting a delay and amplitude of the output signal). Regarding claim 7, Kumagai in view of Lang further view of Kronlachner teaches wherein the one or more processors further execute the instructions to: generate a three-dimensional video image of a virtual space as viewed from the first position of the listening point; and display the generated three-dimensional video image of the virtual space within a space in which the plurality of speakers is located (Kronlachner, ¶68, 105, generating a graphical user interface that shows the layout of the room and allowing the user to shift audio output positions). Regarding claim 8, Kumagai in view of Lang further view of Kronlachner teaches wherein the one or more processors further execute the instructions to adjust the respective output signal to be output to each of the plurality of speakers based on a shielding situation where the one or more virtual sound sources are shielded by a shield in the generated three-dimensional video image of the virtual space (Kronlachner, ¶105, generating a 3D space map showing speaker positions and obstructions so that audio can be optimized accordingly; Kumagai, ¶49, calibrating audio output based on a test sound). Regarding claim 10, it is rejected similarly as claim 1. The method can be found in Kumagai (¶18, method of localizing audio). Regarding claim 11, it is rejected similarly as claim 1. The medium can be found in Kumagai (Fig. 3, CPU requiring computing instructions and to be stored on a form of medium). Regarding claim 12, Kumagai in view of Lang further view of Kronlachner teaches wherein the one or more processors further execute the instructions to: transform coordinates of each of the one or more virtual sound sources described in a virtual world coordinate system into coordinates in an actual world coordinate system; and calculate the respective position of each of the one or more virtual sound sources with respect to the first position of the listening point using the transformed coordinates in the actual world coordinate system (Lang, ¶72-73, coordinates can be provided to the user based on their location and virtual playback format may also provide coordinates that define where the virtual speaker is to be placed in the environment (and therefore also the virtual sound sources)). Regarding claim 13, Kumagai in view of Lang further view of Kronlachner teaches wherein the respective position of each of the one or more virtual sound sources with respect to the first position of the listening point and the positions of the plurality of speakers with respect to the first position of the listening point are each represented as coordinates in a three-dimensional coordinate system having the first position of the listening point as an origin and having axis directions that match axis directions of an actual world coordinate system (Kronlachner, Fig. 7, a listener position can be determined as an origin of a coordinate system). Regarding claim 14, Kumagai in view of Lang further view of Kronlachner teaches wherein the one or more processors further execute the instructions to calculate a distance between the second position of the listening point and the one virtual sound source based on the respective position of the one virtual sound source with respect to the second position of the listening point, and add a distance attenuation amount based on the calculated distance to a gain for each speaker (Kronlachner, ¶95-96, a delay and gain parameter is adjusted based on a determined distance factor for the distance between the user and the one or more speakers). Regarding claim 15, Kumagai in view of Lang further view of Kronlachner teaches wherein, in a case where the plurality of speakers are arranged three-dimensionally, the one or more processors further execute the instructions to create the second plurality of speaker sets such that planes formed by the respective speaker sets do not overlap each other and an entire surface of a space surrounded by the plurality of speakers as viewed from the second position of the listening point is filled with the planes formed by the respective speaker sets (Kronlachner, ¶68, speakers can be organized with respect to the horizontal plane; ¶141, speakers can be assigned to one or more planes (such as ground speakers vs ceiling speakers)). Claim(s) 4 and 16-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kumagai et al (JP2009044261, translated by EPO, hereinafter “Kumagai”) in view of Lang et al (US20240284137, hereinafter “Lang”) in view of Kronlachner (US20240056758) in further view of Nelson et al (US20040170281, hereinafter “Nelson”). Regarding claim 4, Kumagai in view of Lang in view of Kronlachner fail to explicitly teach wherein the one or more processors further execute the instructions to: create a set of adjacent speakers as viewed from the first position of the listening point; calculate an inverse matrix using a vector indicating a direction of each of the plurality of speakers as viewed from the position of the listening point for each set of the speakers; and generate the respective output signal to be output to each of the plurality of speakers by determining a set of the speakers to which a sound source signal, of the one or more sound source signals, is distributed based on the inverse matrix and the respective position of each of the one or more virtual sound sources with respect to the first position of the listening point. Nelson teaches wherein the one or more processors further execute the instructions to: create a set of adjacent speakers as viewed from the first position of the listening point (Fig. 1a, set of speakers as viewed from the listening position); calculate an inverse matrix using a vector indicating a direction of each of the plurality of speakers as viewed from the position of the listening point for each set of the speakers (Fig. 1a, ¶61, creating a set of inverse filters (matrix of inverse filters) as viewed from the listening position); and generate the respective output signal to be output to each of the plurality of speakers by determining a set of the speakers to which a sound source signal, of the one or more sound source signals, is distributed based on the inverse matrix and the respective position of each of the one or more virtual sound sources with respect to the first position of the listening point (Fig. 1a, ¶61, using the determined matrix of inverse filters to be applied to the output signals). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply the audio distribution method (as taught by Nelson) to the audio reproduction system (as taught by Kumagai in view of Lang in view of Kronlachner). The rationale to do so is to apply a known technique to a known device in the same way to achieve the result of achieving excellent center images (Nelson, ¶74). Regarding claim 16, Kumagai in view of Lang further view of Kronlachner teaches wherein the one or more processors further execute the instructions to calculate, for each of the second plurality of speaker sets, an inverse matrix of a direction vector matrix created using direction vectors indicating directions from the second position of the listening point to speakers included in the speaker set (Nelson, Fig. 1a, ¶61, creating a set of inverse filters (matrix of inverse filters) as viewed from the listening position). Regarding claim 17, Kumagai in view of Lang further view of Kronlachner in further view of Nelson teaches wherein the one or more processors further execute the instructions to calculate, for each of the second plurality of speaker sets, a product of a direction vector indicating a direction from the second position of the listening point to one of the one or more virtual sound sources and the inverse matrix, and to determine, based on the product, a speaker set to which a sound source signal output from the one of the one or more virtual sound sources is distributed (Nelson, (Fig. 1a, ¶61, using the determined matrix of inverse filters to be applied to the output signals); ¶98-103, also disclosed is matrix multiplication which is used to determine the output signals modified by the one or more inverse filters). Allowable Subject Matter Claim 18 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Refer to PTO-892, Notice of References Cited for a listing of analogous art. Any inquiry concerning this communication or earlier communications from the examiner should be directed to QIN ZHU whose telephone number is (571)270-1304. The examiner can normally be reached Monday-Thursday 6AM-4PM EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Duc Nguyen can be reached on 571-272-7503. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /QIN ZHU/Primary Examiner, Art Unit 2691
Read full office action

Prosecution Timeline

Feb 26, 2024
Application Filed
Dec 01, 2025
Non-Final Rejection mailed — §103
Feb 26, 2026
Response Filed
Mar 19, 2026
Final Rejection mailed — §103
May 19, 2026
Request for Continued Examination
May 22, 2026
Response after Non-Final Action
Jun 25, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12701377
ACOUSTIC PROCESSING DEVICE, ACOUSTIC PROCESSING METHOD, ACOUSTIC PROCESSING PROGRAM, AND ACOUSTIC PROCESSING SYSTEM
2y 6m to grant Granted Aug 04, 2026
Patent 12701380
Apparatus, Method and Computer Program for Synthesizing a Spatially Extended Sound Source Using Elementary Spatial Sectors
2y 3m to grant Granted Aug 04, 2026
Patent 12701361
METHOD FOR DETERMINING ORIENTATION INFORMATION AND ELECTRONIC DEVICE
2y 3m to grant Granted Aug 04, 2026
Patent 12688843
Notification Apparatus, Control Program For Notification Apparatus, And Seat System
2y 3m to grant Granted Jul 21, 2026
Patent 12684308
CROSSTALK CANCELLATION FOR REVERBERANT ACOUSTIC FIELDS
2y 11m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
88%
Grant Probability
91%
With Interview (+3.0%)
1y 11m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 631 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month