Prosecution Insights
Last updated: August 17, 2026
Application No. 18/781,885

SYSTEMS, DEVICES, AND METHODS FOR AUDIO PRESENTATION IN A THREE-DIMENSIONAL ENVIRONMENT

Non-Final OA §103
Filed
Jul 23, 2024
Priority
Jul 23, 2023 — provisional 63/515,124
Examiner
HAJNIK, DANIEL F
Art Unit
2616
Tech Center
2600 — Communications
Assignee
Apple Inc.
OA Round
2 (Non-Final)
78%
Grant Probability
Favorable
2-3
OA Rounds
9m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 78% — above average
78%
Career Allowance Rate
620 granted / 794 resolved
+16.1% vs TC avg
Strong +21% interview lift
Without
With
+21.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
4 currently pending
Career history
798
Total Applications
across all art units

Statute-Specific Performance

§101
14.8%
-25.2% vs TC avg
§103
60.7%
+20.7% vs TC avg
§102
7.5%
-32.5% vs TC avg
§112
6.4%
-33.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 794 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Allowable Subject Matter Claims 5 and 10-11 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Response to Arguments Applicant’s arguments, filed May 28, 2026, with respect to how the claim features (relating to the predetermined location for non-spatialized sound) at the end of claim 1 differ from the prior art cited in the last office have been fully considered. These arguments are found to be persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in this office action. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1 and 12-13 are rejected under 35 U.S.C. 103 as being unpatentable over Miyazaya (Pub No. US 2014/0300636 A1) in view of Zhang et al. (Pub No. US 2019/0073109 A1). As per claim 1, Miyazaya teaches the claimed: 1. A method comprising: at a computer system in communication with a display generation component, one or more input devices, and an audio output device (Miyazaya in [0060] “Next, a hardware configuration example of the information processing device 10 according to an embodiment of the present disclosure will be described with reference to FIG. 6. The information processing device 10 may include a CPU 150, a RAM 155, a non-volatile memory 160, a display device 165, a face orientation detection device 170, and a sound image localization audio device 175”. Also, figure 5 shows an “Sound Output Unit 125”. The ”face orientation detection device 170” corresponds to an input device because the user is able to move their head and body to change the face orientation. Thus, this device is acting as an input device): displaying, via the display generation component, an environment from a viewpoint of a user (Miyazaya in figures 1-3 shows example of displaying an augmented reality environment from a viewpoint of a user); while displaying the (Miyazaya in figures 1-3 shows show that while displaying an environment from a viewpoint of a user, an event is detected. The event is detecting a real-world object with an associated sound source); and in response to detecting the event (Miyazaya in [0034] “The information processing device 10 according to an embodiment of the present disclosure may provide information associated with a real object present in a real space using visual information and sound information … The information processing device 10 may determine a position of the visual marker Mv based on, for example, a position of a real object. In addition, the information processing device 10 may determine a sound source position of the sound marker Ms based on the position of the visual marker Mv.” In this instance, the event determination corresponds to determining a real object with sound and its associated sound source position): in accordance with a determination that the event corresponds to activation of a sound effect (Miyazaya in [0054] “[0054] The sound output control unit 130 may have a function such as controlling the output of sounds to a user. In addition, the sound output control unit 130 is an example of a localization information generation unit that generates localization information indicating a sound source position, but is not limited thereto. The sound output control unit 130 may generate localization information indicating a sound source position of sound information associated with the same real object as one relating to a virtual object, for example, one associated with a virtual object based on, for example, a virtual position. The sound output control unit 130 may also generate localization information based on the distance between a user position and a virtual position.”. In this instance, the sound outputting using the sound source position corresponds to the claimed “activation of a sound effect”): in accordance with a determination that the event corresponds to activation of a spatialized sound effect, presenting, via the audio output device, the sound effect as emanating from a location in the three-dimensional environment associated with the event (Miyazaya in [0055] “For this reason, when a virtual position is separated from a user position in a predetermined or longer distance, for example, even when the sound source position is located obliquely forward as viewed from a user, there is a high possibility that the user is not able to recognize such a difference of sound source positions. Thus, when a virtual position is separated from a user position in a predetermined or longer distance, the sound source position may be set to be in the front side of the user, and when the user approaches the virtual position, the sound source position may be moved from the front side of the user to a direction closer to the virtual position, so that the user can recognize the direction of the virtual position more easily.” In this passage, the “spatialized sound effect” corresponds to a virtual position less than a predetermined distance from the user. Also, in this passage, the sound source position being moved to a direction closer to the virtual position corresponds to the claimed “the sound effect as emanating from a location in the three-dimensional environment associated with the event”); and in accordance with a determination that the event corresponds to activation of a non-spatialized sound effect, presenting, via the audio output device, the sound effect as emanating from a predetermined location for audio in the three-dimensional environment that has a predetermined spatial relationship relative to the viewpoint of the user (Miyazaya in [0055] “For this reason, when a virtual position is separated from a user position in a predetermined or longer distance, for example, even when the sound source position is located obliquely forward as viewed from a user, there is a high possibility that the user is not able to recognize such a difference of sound source positions. Thus, when a virtual position is separated from a user position in a predetermined or longer distance, the sound source position may be set to be in the front side of the user, and when the user approaches the virtual position, the sound source position may be moved from the front side of the user to a direction closer to the virtual position, so that the user can recognize the direction of the virtual position more easily.” In this passage, the “non-spatialize sound effect” corresponds to a virtual position at a predetermined distance or longer. Also, in this passage, the sound source position set to the front side of the user corresponds to the claimed “sound effect as emanating from a predetermined location”. The front side of the user is also a predetermined spatial relationship relative to the viewpoint of the user). While Miyazaya teaches of displaying an environment in their augmented reality system, Miyazaya does not explicitly teach of displaying a three-dimensional an environment per se. Zhang teaches this feature in figures 4-5 where the HMD displays a 3D environment from the viewpoint of the user. This environment is 3D because it includes modifying virtual feature based upon depth, e.g. cursor 210 or 214 is displayed differently depending upon its depth in the 3D environment. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to displaying a three-dimensional an environment as taught by Zhang with the system of Miyazaya in order to better illustrate the depth positions of environment features such as a cursor. As per claims 12 and 13, these claims are similar in scope to limitations recited in claim 1, and thus are rejected under the same rationale. Miyazaya teaches of using a processor, memory, and stored programs in [0060] and in figure 6. Also, please see Miyazaya in [0064] “The RAM 155 may temporarily store a program read in the CPU 150, a parameter that is appropriately changed during the execution of the program, or the like.” Miyazaya teaches of a computer-readable medium in [0007] “Further, according to an embodiment of the present disclosure, there is provided a non-transitory computer-readable medium embodied with a program”. Claims 2-4 are rejected under 35 U.S.C. 103 as being unpatentable over Miyazaya in view of Zhang in further view of Mindlin et al. (Pub No. US 20190385613 A1). As per claim 2, Miyazaya alone does not explicitly teach the claimed limitations. However, Miyazaya and Zhang in combination with Mindlin teaches the claimed: 2. The method of claim 1, wherein: the event corresponding to activation of the spatialized sound effect includes display of a representation of a participant of a communication session in the three-dimensional environment, the representation of the participant including a first visual element having an appearance of a representation of a mouth corresponding to the participant (Mindlin initiates a chat/talking event in figure 8A this includes activation of sound effects (figure 9, step 902) through the chat and display of a representation of a participant with a mouth of a communication session in the three-dimensional environment, e.g. Mindlin at the end of [0042] “… For instance, chat management system 204 may perform sound processing to facilitate phoneme animation for the speaking avatar (i.e., to simulate moving of the speaking avatar's mouth in synchronicity with the speech).” In Mindlin, the chat/talking event is an activation of the spatialized sound effect, e.g. please see Mindlin in [0018] “… The simulated binaural audio signal may be representative of a rendered simulation of sound propagating to an avatar representing the user within the artificial reality world. For example, the simulated binaural audio signal may be a signal configured to simulate for the user what the avatar hears in the audio environment of the artificial reality world (e.g., by including representations of sounds from various sources such as people speaking and other sound sources) and how the avatar hears it The claimed feature is taught when the chat event of Mindlin is incorporated into the system of Miyazaya and Zhang), and the location in the three-dimensional environment associated with the event is a location in the three-dimensional environment corresponding to where the first visual element is located (Figure 8A of Mindlin shows that the location in the 3D environment associated with the other avatar or person talking (the event) corresponds to where mouth of the avatar is speaking (to where the first visual element is located)). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the chat feature as taught by Mindlin with the system of Miyazaya as modified by Zhang in order to allow the user to better hear and communicate with other nearby people and have visual navigation instructions in the same augmented reality device. As per claim 3, Miyazaya alone does not explicitly teach the claimed limitations. However, Miyazaya and Zhang in combination with Mindlin teaches the claimed: 3. The method of claim 1, wherein: the event corresponding to activation of the spatialized sound effect includes display of a representation of a participant of a communication session in the three-dimensional environment, the representation of the participant being a representation of a geometric shape (Mindlin initiates a chat/talking event in figure 8B and in [0018]. This includes activation of sound effects (figure 9, step 902) through the chat and display of a representation of a participant being represented by a geometric circle shape 814. The claimed feature is taught when the chat event of Mindlin is incorporated into the system of Miyazaya and Zhang), and the location in the three-dimensional environment associated with the event is a location in the three-dimensional environment corresponding to a center of the representation of the geometric shape (Figure 8B of Mindlin shows that the location in the 3D environment associated with the other avatar or person talking (the event) corresponds to the center of the geometric circle shape 814). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the chat features as taught by Mindlin with the system of Miyazaya as modified by Zhang. The motivation of claim 2 is incorporated herein. In addition, the geometric shape helps the user better understand which person the dialog at the bottom of the screen in figure 8B of Mindlin corresponds to. As per claim 4, Miyazaya alone does not explicitly teach the claimed limitations. However, Miyazaya and Zhang in combination with Mindlin teaches the claimed: 4. The method of claim 1, wherein: the event corresponding to activation of the spatialized sound effect includes display of a user interface element that includes representation of a participant of a communication session in the three-dimensional environment (Mindlin initiates the chat event in figure 8B and in [0018]. This includes activation of sound effects (figure 9, step 902) through the chat and display of a user interface element (circle 814) that includes representation of a participant 802 of a communication session in the three-dimensional environment. The claimed feature is taught when the chat event of Mindlin is incorporated into the system of Miyazaya and Zhang), and the location in the three-dimensional environment associated with the event is a location in the three-dimensional environment corresponding to a center of the user interface element (Figure 8B of Mindlin shows that the location in the 3D environment associated with the other avatar or person talking corresponds to the center of user interface element 814). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the chat feature and user interface element as taught by Mindlin with the system of Miyazaya as modified by Zhang. The motivation of claims 2 and 3 are incorporated herein. Claims 6-7 are rejected under 35 U.S.C. 103 as being unpatentable over Miyazaya in view of Zhang in further view of Mejia Cobo (Pub No. US 2020/0301513 A1). As per claim 6, Miyazaya alone does not explicitly teach all the claimed limitations. However, Miyazaya and Zhang in combination with Mejia Cobo teaches the claimed: 6. The method of claim 1, wherein the predetermined location for audio in the three-dimensional environment that has the predetermined spatial relationship relative to the viewpoint of the user is within a viewport of the computer system (As mentioned above for claim 1, Miyazaya in [0055] teaches of the predetermined location for audio being to the front side of the user. Mejia Cobo shows the claimed “the viewpoint of the user is within a viewport of the computer system” in [0017] and in figure 1 as viewport 150. In Mejia Cobo, the front side of the user is located in the viewport 150 in figure 1. The claimed feature is taught in the combination, e.g. when the predetermined spatial relationship is located to the side of the user and within the viewport 150 in figure 1 of Mejia Cobo (e.g. at position 160 or a similar front side position)). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the viewport as taught by Zhang with the system of Miyazaya in order to mathematically define for the software which spatial areas in the 3D environment are viewable to the user from a given vantage point. This helps the system better mathematically define which areas to place sound or visual objects for output to the user. As per claim 7, Miyazaya teaches the claimed: 7. The method of claim 6, wherein presenting, via the audio output device, the sound effect as emanating from the predetermined location for audio in the three-dimensional environment includes presenting the sound effect as synthesized stereo audio (Miyazaya in [0055] “… Thus, when a virtual position is separated from a user position in a predetermined or longer distance, the sound source position may be set to be in the front side of the user” and [0052] “… The sound output device used here may be capable of outputting stereophony according to localization information of a sound.”). Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Miyazaya in view of Zhang in further view of Mejia Cobo and Gelfand et al. (Pub No. US 2008/0268876 A1). As per claim 8, Miyazaya alone does not explicitly teach all the claimed limitations. However, Miyazaya, Zhang, and Mejia Cobo in combination with Gelfand teaches the claimed: 8. The method of claim 7, wherein: the event corresponding to activation of the non-spatialized sound effect includes display of a user interface element that includes a representation of a participant of a communication session in the three-dimensional environment (As mentioned above for claim 1, the non-spatialized sound effect corresponds to a location that is a long distance away from the user, e.g. Miyazaya in [0055]. Miyazaya in figures 1-2 shows the display a user interface elements Ms, Mv1 or Mv2. As shown in figure 1-2, these user interface elements may correspond to a restaurant or bar that is located far away in Miyazaya in [0055]. While, Miyazaya teaches of the display the user interface element, they are silent that it is a representation a participant of a communication session. Gelfand teaches this feature in [0094] where the user may call the restaurant or other establishment (e.g. Gelfand recites [0094] “… a user of the mobile terminal 10 may point his/her camera module 36 at a business such as for example a restaurant and the visual search server 54 provides the visual search client 68 with a phone number of the restaurant and the visual search client of the mobile terminal 10 thereby may call the restaurant.” Thus, the claimed features are taught when the call feature from Gelfand is used with the system of Miyazaya), presenting, via the audio output device, the sound effect as emanating from the predetermined location for audio in the three-dimensional environment includes presenting the sound effect as synthesized stereo audio from the predetermined location for audio, different from a location of the user interface element in the three-dimensional environment (This is taught in Miyazaya in [0055] where the sound effect related to the faraway location is emanating from the predetermined location to the front side of the user even though the location is actually in front of the user. In addition, this is different from a location of the user interface element Ms, Mv1 or Mv2 in the three-dimensional environment in figures 1-2 of Miyazaya as well. For example, please see Miyazaya in [0055] “… Thus, when a virtual position is separated from a user position in a predetermined or longer distance, the sound source position may be set to be in the front side of the user” and [0052] “… The sound output device used here may be capable of outputting stereophony according to localization information of a sound.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the restaurant call feature as taught by Gelfand with the system of Miyazaya as modified by Zhang in order to make it easier for the user to contain the restaurant or other business to talk to them or inquire information, e.g. to ask about hours open, menu options, or to make a reservation. Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Miyazaya in view of Zhang in further view of Mindlin. As per claim 9, Miyazaya alone does not explicitly teach all the claimed limitations. However, Miyazaya and Zhang in combination with Mindlin teaches the claimed: 9. The method of claim 1, wherein the computer system is in a communication session with one or more other participants and wherein the sound effect includes audio corresponding to the one or more other participants of the communication session (Mindlin in [0040] “Chat management system 204 may serve as a central hub for hosting all chat communications that may occur within an artificial reality world being generated and provided by artificial reality provider system 202. For example, as user 212 and other users experiencing the same artificial reality world using other media player devices (not explicitly shown in FIG. 2) speak to one another, all the voice signals may pass through chat management system 204”. The claimed feature is taught when the communication session of Mindlin is incorporated into the system of Miyazaya and Zhang), and the method comprising: while in the communication session with the one or more other participants: in accordance with a determination that the event includes display, in the three- dimensional environment, of a representation of a first three-dimensional environment of a participant of the communication session (Mindlin in figures 8a-8b shows a representation of a first three-dimensional environment of a participant (i.e. an avatar 802) of the communication session), presenting, via the audio output device, the audio corresponding to the one or more other participants of the communication session as spatialized audio emanating from a location in the three-dimensional environment that is above the display of the representation of the first three-dimensional environment (Mindlin in [0018] “… The simulated binaural audio signal may be representative of a rendered simulation of sound propagating to an avatar representing the user within the artificial reality world. For example, the simulated binaural audio signal may be a signal configured to simulate for the user what the avatar hears in the audio environment of the artificial reality world (e.g., by including representations of sounds from various sources such as people speaking and other sound sources) and how the avatar hears it (e.g., by taking into account various aspects that affect the propagation of sound to the avatar such as reverberation and echoes in the room, where the avatar is positioned with respect to the sound sources, how the avatar and the sound sources are oriented with respect to one another, etc.).”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the communion session as taught by Mindlin with the system of Miyazaya as modified by Zhang in order to allow the user to both talk to friends or other people as well as have visual navigation instructions available in the same extended reality-based display system. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to DANIEL F HAJNIK whose telephone number is (571) 272-7642. The examiner can normally be reached Mon-Fri 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DANIEL F HAJNIK/Supervisory Patent Examiner, Art Unit 2616
Read full office action

Prosecution Timeline

Jul 23, 2024
Application Filed
Jan 28, 2026
Non-Final Rejection mailed — §103
May 12, 2026
Examiner Interview Summary
May 12, 2026
Applicant Interview (Telephonic)
May 28, 2026
Response Filed
Aug 04, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12682507
SYSTEM AND METHOD FOR GENERATING REAL-TIME SYNTHETIC IMAGERY DURING AIRCRAFT FLIGHT
2y 6m to grant Granted Jul 14, 2026
Patent 12678139
SYSTEMS AND METHODS FOR UNCERTAINTY AWARE CALIPER PLACEMENT
2y 11m to grant Granted Jul 14, 2026
Patent 12676949
DATA PROCESSING APPARATUS, DATA PROCESSING METHOD, AND PROGRAM
3y 3m to grant Granted Jul 07, 2026
Patent 12586259
IMAGE GENERATION USING A TEXT AND IMAGE CONDITIONED MACHINE LEARNING MODEL
2y 1m to grant Granted Mar 24, 2026
Patent 12573138
ENDOSCOPIC EXAMINATION SUPPORT APPARATUS, ENDOSCOPIC EXAMINATION SUPPORT METHOD, AND RECORDING MEDIUM
2y 2m to grant Granted Mar 10, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

2-3
Expected OA Rounds
78%
Grant Probability
99%
With Interview (+21.0%)
2y 10m (~9m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 794 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month