Prosecution Insights
Last updated: August 30, 2026
Application No. 18/977,143

INFORMATION PROCESSING APPARATUS CAPABLE OF POSITIVELY GRASPING SOUND IN REAL SPACE, METHOD OF CONTROLLING INFORMATION PROCESSING APPARATUS, AND STORAGE MEDIUM

Non-Final OA §102§103
Filed
Dec 11, 2024
Priority
Dec 15, 2023 — JP 2023-212039
Examiner
MAZUMDER, TAPAS
Art Unit
Tech Center
Assignee
Canon Inc.
OA Round
1 (Non-Final)
82%
Grant Probability
Favorable
1-2
OA Rounds
7m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
352 granted / 430 resolved
+21.9% vs TC avg
Strong +17% interview lift
Without
With
+16.9%
Interview Lift
resolved cases with interview
Typical timeline
2y 4m
Avg Prosecution
22 currently pending
Career history
444
Total Applications
across all art units

Statute-Specific Performance

§101
9.6%
-30.4% vs TC avg
§103
52.1%
+12.1% vs TC avg
§102
10.3%
-29.7% vs TC avg
§112
17.3%
-22.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 430 resolved cases

Office Action

§102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . .Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 1-3, 8-10, 12 and 17-20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Dorn et al, ( US patent Publication: US 20230384592, “Dorn”). Regarding claim 1, An information processing apparatus (Fig. 4), comprising one or more processors (CPU 402) and/or circuitry configured to: acquire user information concerning a user who visually recognizes a space image including at least an image of a virtual space; (“[0068]….Using the known location/orientation of the HMD the real-world objects, and inertial sensor data from the, the gestures and movements of the user can be continuously monitored and tracked during the user's interaction with the VR scenes.”) acquire virtual object information concerning a virtual object in the space image; (“[0045] FIG. 3A illustrates an example of a user 100, utilizing HMD 102 and controllers 104, for interacting with a virtual-reality scene, in accordance with one embodiment. In this example, it is shown that the user 100 has a point of view 108 directed into the scene, which provides viewing of interactivity of a game. The game in this example, is a first-person shooter game, and the first person is shooting at ghosts”) acquire, in a case where a sound is generated in a real space, position information of a sound source of the generated sound; and determine a notification method of notifying the user of a direction of the sound source in the real space, based on the acquired user information, the acquired virtual object information, and the acquired position information.( “[0030] In some embodiments, the visual cues can be integrated to pop-ups, icons, graphics, or images that appear in the virtual-reality space. The location of the integrated visual cues is placed in the VR scene in a location that relates or corresponds to where the real-world sound comes from. If the sound is coming from the right, the visual cue can make a sound wave or sound cue to the right side of the screen or to the right side of the user's interactive environment. This provides a natural way for the HMD user to be notified of sounds that are occurring in the real-world space, in a way that is more natural to the HMD environment.”) Claim 19 is directed to a method and its steps are similar in scope and function of the element of eth device claim 1 and therefore claim 19 is rejected with same rationales specified in the rejection of claim 1. Claim 20 is directed to a non-transitory computer-readable storage medium (“[0073] One or more embodiments can also be fabricated as computer readable code on a computer readable medium. The computer readable medium is any data storage device that can store data, which can be thereafter be read by a computer system. Examples of the computer readable medium include hard drives, network attached storage (NAS), read-only memory, random-access memory, CD-ROMs, CD-Rs, CD-RWs, magnetic tapes and other optical and non-optical data storage devices. The computer readable medium can include computer readable tangible medium distributed over a network-coupled computer system so that the computer readable code is stored and executed in a distributed fashion”) and its elements are similar in scope and function of the element of the device claim 1 and therefore claim 20 is rejected with same rationales specified in the rejection of claim 1. Regarding claim 2, Dorn teaches, wherein the one or more processors and/or circuitry is/are further configured to notify the user of a direction of the sound source by using the determined notification method. (“ [0030] In some embodiments, the visual cues can be integrated to pop-ups, icons, graphics, or images that appear in the virtual-reality space. The location of the integrated visual cues is placed in the VR scene in a location that relates or corresponds to where the real-world sound comes from. If the sound is coming from the right, the visual cue can make a sound wave or sound cue to the right side of the screen or to the right side of the user's interactive environment”) Regarding claim 3 , Dorn teaches, wherein the one or more processors and/or circuitry is/are further configured to display the space image; and wherein the notifying includes notifying the user of a direction of the sound source by using a marker displayed in the space image by the displaying. (“ [0030] In some embodiments, the visual cues can be integrated to pop-ups, icons, graphics, or images that appear in the virtual-reality space. The location of the integrated visual cues is placed in the VR scene in a location that relates or corresponds to where the real-world sound comes from. If the sound is coming from the right, the visual cue can make a sound wave or sound cue to the right side of the screen or to the right side of the user's interactive environment”) Regarding claim 8, Dorn teaches, wherein the acquiring of the user information is performed by acquiring at least one of position information of the user, sight line information of the user, and information concerning a gesture of the user, as the user information. (“[0068]….Using the known location/orientation of the HMD the real-world objects, and inertial sensor data from the, the gestures and movements of the user can be continuously monitored and tracked during the user's interaction with the VR scenes.”) Regarding claim 9, Dorn teaches, wherein the acquiring of the virtual object information is performed by acquiring at least one of position information, a size, and a posture of the virtual object in the space image, as the virtual object information. [(“0045] FIG. 3A illustrates an example of a user 100, utilizing HMD 102 and controllers 104, for interacting with a virtual-reality scene, in accordance with one embodiment. In this example, it is shown that the user 100 has a point of view 108 directed into the scene, which provides viewing of interactivity of a game. The game in this example, is a first-person shooter game, and the first person is shooting at ghosts”) Regarding claim 10, Dorn teaches, wherein the one or more processors and/or circuitry is/are further configured to collect, in a case where a sound is generated in the real space, the generated sound, and wherein the acquiring of the position information of the sound source includes estimating a position of the sound source based on the sound collected by the collecting, and acquiring a result of the estimation as the position information. ([0006]…..“The method further includes receiving sensor data from one or more sensors in a real-world space in which the HMD is located. Then, identifying an object location of an object in the real-world space that produces a sound.”) Regarding claim 12, Dorn teaches, wherein the acquiring of the position information of the sound source includes acquiring the position information of the sound source, in a case where the sound generated in the real space is a predetermined type of sound. (“[0043] When sensors 207 detect sounds in the space near or proximate to the user 100, the location of those sounds is identified and tracked as the HMD POV 108 changes and moves about during interactivity.”) Regarding claim 17, Dorn teaches, a display unit configured to display the space image. (“[0006] In one embodiment, a method for integrating media cues into virtual reality scenes presented on a head mounted display (HMD) is disclose”) Regarding claim 18, Dorn teaches, wherein the information processing apparatus is a head mounted display (HMD). ( Fig.3A HMD) [0006] In one embodiment, a method for integrating media cues into virtual reality scenes presented on a head mounted display (HMD) is disclosed.”) Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim 4 is rejected under 35 U.S.C. 103 as being obvious over Dorn. Regarding claim 4, Dorn doesn’t expressly teach, wherein the marker is an arrow. As Dorn teaches marker for the sound source as audio symbol or click clock text symbol in Fig.3A(“[0051] In addition, if the person 310 is identified by one or more cameras/sensors, an image of that person 310 can be generated and added as a media cue 204 into the virtual-reality space. Again, the integration of the media cue 204 is with spatial relevance and correlation to the location where person 310 is standing and speaking relative to user 100. Thus, the media cue 204 is rendered to the left of the images shown in the HMD 102.” “ [0046] By way of example, a person 310 is shown in the real-world walking toward user 100. As the person 310 walks toward user 100, the person 300 has shoes that make a click clock sound 314. The click clock sound 314 is picked up by a sensor (e.g. one or more microphones) of the HMD 102, or controller 104, or microphones located in the space near with user 100. In one embodiment, the sound 314 is processed by the virtual world rendering module 202, which then generates a media cue 204. The media cue 204 is shown as a text “click clock . . . ”) it would have been obvious for an ordinary skilled person to use an arrow as media cue for source direction as an alternative way of representation of marker. The motivation for the above is to enhance Dorn with different way of representing sound source direction. Claim 11 is rejected under 35 U.S.C. 103 as being obvious over Dorn in view of Tsuruga et al. ( US patent publication: 20240430636, “Tsuruga”). Regarding claim 11, Dorn teaches, acquiring of the position information of the sound source as shown in claim 1 but doesn’t teach acquiring the position information of the sound source, in a case where the level of the sound generated in the real space is equal to or higher than a predetermined level. Tsuruga, teaches, notification of location of the sound source, in a case where the level of the sound generated in the real space is equal to or higher than a predetermined level. (“ [0051] The level detection section 32 and the sound source direction detection section 33 accept the microphone array signals received by the communication control section 31 as an input, and the level detection section 32 detects the volume of the sound in the real space, and when the volume is more than a predetermined level, notifies the placement correction section 36 ..”) Tsuruga and Dorn are analogous as they are from eth field of extended reality with sound. Therefore it would have been obvious for an ordinary skilled person in the art before the effective filing date of the claimed invention to have modified Dorn to have included acquiring the position information of the sound source, in a case where the level of the sound generated in the real space is equal to or higher than a predetermined level as taught by Tsuruga. The motivation for the above is to filter out processing of insignificant level of sound. Claim 14-15 are rejected under 35 U.S.C. 103 as being obvious over Dorn in view of Miyazaki al. ( US patent publication: 20190324708, “Miyazaki”). Regarding claim 14, Dorn teaches the space image is VR space image as shown in claim 1 but doesn’t teach wherein the space image is an image in a mixed space, including an image in the real space and an image in the virtual space, and wherein the one or more processors and/or circuitry is/are further configured to determine whether or not the sound source is included in the space image. However Miyazaki teaches, the space image is an image in a mixed space, including an image in the real space and an image in the virtual space, (“ [0061] FIG. 3 is a view depicting an example of a frame image constituting the AR space video image displayed on the display section 38. Hereinafter, a frame image constituting the AR space video image and exemplified in FIG. 3 shall be referred to as an AR space image 70. As depicted in FIG. 3, the AR space image 70 includes a real space part 72 as a part which the frame image constituting the real space video image occupies, and a VR space part 74 as a part which the frame image constituting the VR space video image occupies.”) [0130] Here, for example, in the case where the synthetic sound obtained by synthesizing the VR space sound and the real space sound in which the sound in the direction of the line of sight of the user is emphasized is generated, the VR space video image supplying section 114 may supply an image in which a virtual object facing an object which is present in the direction of the line of sight of the user is arranged. Then, the AR space video image generating section 116 may generate an image of the mixed space including the image of the virtual object such as the character facing the object, within the real space, which is present in the direction of the line of sight of the user as the VR space part 74. Then, the image of the mixed space generated in this manner may be transmitted to the HMD 12 and displayed on the display section 38.”) Miyazaki and Dorn are analogous as they are from the field AR/VR processing. Therefore it would have been obvious for an ordinary skilled person in the image before the effective filing date of the claimed invention to have modified Dorn to have included real part of Fig. 3A in a space image and thereby the space image is an image in a mixed space, including an image in the real space and an image in the virtual space as taught by Miyazaki. Dorn as modified by Miyazaki teaches, wherein the one or more processors and/or circuitry is/are further configured to determine whether or not the sound source is included in the space image. .( Dorn, “[0030] In some embodiments, the visual cues can be integrated to pop-ups, icons, graphics, or images that appear in the virtual-reality space. The location of the integrated visual cues is placed in the VR scene in a location that relates or corresponds to where the real-world sound comes from.”) Regarding claim 15, Dorn doesn’t expressly teach, wherein the space image is an image in a mixed space, including an image in the real space and an image in the virtual space, and wherein the one or more processors and/or circuitry is/are further configured to determine whether or not the virtual object is included in the space image. However, However Miyazaki teaches, the space image is an image in a mixed space, including an image in the real space and an image in the virtual space, (“ [0061] FIG. 3 is a view depicting an example of a frame image constituting the AR space video image displayed on the display section 38. Hereinafter, a frame image constituting the AR space video image and exemplified in FIG. 3 shall be referred to as an AR space image 70. As depicted in FIG. 3, the AR space image 70 includes a real space part 72 as a part which the frame image constituting the real space video image occupies, and a VR space part 74 as a part which the frame image constituting the VR space video image occupies.”) [0130] Here, for example, in the case where the synthetic sound obtained by synthesizing the VR space sound and the real space sound in which the sound in the direction of the line of sight of the user is emphasized is generated, the VR space video image supplying section 114 may supply an image in which a virtual object facing an object which is present in the direction of the line of sight of the user is arranged. Then, the AR space video image generating section 116 may generate an image of the mixed space including the image of the virtual object such as the character facing the object, within the real space, which is present in the direction of the line of sight of the user as the VR space part 74. Then, the image of the mixed space generated in this manner may be transmitted to the HMD 12 and displayed on the display section 38.”) Miyazaki and Dorn are analogous as they are from the field AR/VR processing. Therefore it would have been obvious for an ordinary skilled person in the image before the effective filing date of the claimed invention to have modified Dorn to have included real part of Fig. 3A in a space image and thereby the space image is an image in a mixed space, including an image in the real space and an image in the virtual space as taught by Miyazaki. Dorn as modified by Miyazaki teaches, wherein the one or more processors and/or circuitry is/are further configured to determine whether or not the virtual object is included in the space image. (Dorn {0045]) Allowable Subject Matter Claims 5-7, 13 and 16 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Claim 5 is objected because Miyazaki teaches, wherein the space image is an image in a mixed space, including an image in the real space and an image in the virtual space,(“ [0061] FIG. 3 is a view depicting an example of a frame image constituting the AR space video image displayed on the display section 38. Hereinafter, a frame image constituting the AR space video image and exemplified in FIG. 3 shall be referred to as an AR space image 70. As depicted in FIG. 3, the AR space image 70 includes a real space part 72 as a part which the frame image constituting the real space video image occupies, and a VR space part 74 as a part which the frame image constituting the VR space video image occupies.”) [0130] Here, for example, in the case where the synthetic sound obtained by synthesizing the VR space sound and the real space sound in which the sound in the direction of the line of sight of the user is emphasized is generated, the VR space video image supplying section 114 may supply an image in which a virtual object facing an object which is present in the direction of the line of sight of the user is arranged. Then, the AR space video image generating section 116 may generate an image of the mixed space including the image of the virtual object such as the character facing the object, within the real space, which is present in the direction of the line of sight of the user as the VR space part 74. Then, the image of the mixed space generated in this manner may be transmitted to the HMD 12 and displayed on the display section 38.”) but doesn’t expressly teach, wherein in a case where the sound source is included in the space image, the notifying is performed by the arrow indicating the sound source. Claim 6 is objected because the prior art of record fails to expressly teach, wherein in a case where the sound source is not included in the space image, the notifying is performed by the arrow having a length proportional to a distance to the sound source. Claim 7 is objected because the prior art of record fails to expressly teach, wherein the space image is an image in a mixed space, including an image in the real space and an image in the virtual space, and wherein in a case where the virtual object and the sound source are included in the space image, and the virtual object and the sound source overlap each other, the displaying is performed by adjusting transmittance of the virtual object. Claim 13 is objected because the prior art of record fails to expressly teach, wherein the determining of the notification method can include not notifying the direction of the sound source as the notification method. Claim 16 is objected because the prior art of record fails to expressly teach,, wherein the space image is an image in a mixed space, including an image in the real space and an image in the virtual space, and wherein the one or more processors and/or circuitry is/are also configured to determine whether or not the virtual object and the sound source are included in the space image in a state overlapping each other. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Tapas Mazumder whose telephone number is (571)270-7466. The examiner can normally be reached M-F 8:00 AM-5:00 PM PST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alicia Harrington can be reached at 571-272-2330. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /TAPAS MAZUMDER/ Primary Examiner, Art Unit 2615
Read full office action

Prosecution Timeline

Dec 11, 2024
Application Filed
Aug 11, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12718316
UPDATING DISPLAY OF GAME MAP
2y 4m to grant Granted Aug 25, 2026
Patent 12700151
METHOD FOR DISPLAYING AN IMAGE AND ELECTRONIC DEVICE SUPPORTING THE SAME
2y 5m to grant Granted Aug 04, 2026
Patent 12694621
FREE-FORM CURVED SURFACE SLICING METHOD AND DEVICE BASED ON IMPLICIT MODEL
2y 1m to grant Granted Jul 28, 2026
Patent 12684085
VIDEO FUSION METHOD, APPARATUS, ELECTRONIC DEVICE AND STORAGE MEDIUM
2y 6m to grant Granted Jul 14, 2026
Patent 12682536
MOTION GENERATION DEVICE FOR GENERATING MOTION BASED ON INPUT INFORMATION INCLUDING TEXT AND OPERATION METHOD THEREOF
1y 2m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
82%
Grant Probability
99%
With Interview (+16.9%)
2y 4m (~7m remaining)
Median Time to Grant
Low
PTA Risk
Based on 430 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month