Prosecution Insights
Last updated: October 01, 2026
Application No. 18/815,175

Augmented Reality For Video Conferencing Platforms

Non-Final OA §103
Filed
Aug 26, 2024
Examiner
HONG, RICHARD J
Art Unit
2623
Tech Center
2600 — Communications
Assignee
Zoom Video Communications Inc.
OA Round
5 (Non-Final)
79%
Grant Probability
Favorable
5-6
OA Rounds
0m
Est. Remaining
83%
With Interview

Examiner Intelligence

Grants 79% — above average
79%
Career Allowance Rate
490 granted / 623 resolved
+16.7% vs TC avg
Minimal +4% lift
Without
With
+3.9%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 0m
Avg Prosecution
17 currently pending
Career history
655
Total Applications
across all art units

Statute-Specific Performance

§101
1.9%
-38.1% vs TC avg
§103
66.5%
+26.5% vs TC avg
§102
18.7%
-21.3% vs TC avg
§112
9.5%
-30.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 623 resolved cases

Office Action

§103
DETAILED ACTION Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on Sep. 11, 2026 has been entered. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-20 are pending. Response to Amendment Applicants’ response to the last Office Action, dated Sep. 11, 2026, has been entered and made of record. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office Action. Response to Arguments Applicant’s arguments, dated Sep. 11, 2026, have been considered but are moot because the arguments do not apply to all of the references being used in the current rejection. Please see the following claim rejections for detailed analysis. Claim Rejections - 35 USC § 103 The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office Action. Claims 1, 4-13 and 16-20 are rejected under 35 U.S.C. 103 as being unpatentable over Bennett et al. (US 2019/0236842 A1) in view of Huynh et al. (US 2025/0252676 A1) and Petill et al. (US 2021/0012113 A1). As to claim 1, Bennett teaches a method (Bennett, Abs., “methods and systems are provided for authoring and presenting 3D presentations”), comprising: obtaining, via an imaging device, interaction space image information indicative of an interaction space (Bennett, FIG. 2, [0041], “host device 200 may be able to detect one or more characteristics of room (e.g., by recognizing physical features using on-board sensor(s) such as cameras, by detecting a room identifier using identification technology such as RFID, by detecting room location using positioning systems such as GPS or WiFi, or otherwise)”); generating, based on the interaction space image information, virtual interaction space information corresponding to a virtual interaction space associated with a video conference, the virtual interaction space comprising a virtual three-dimensional (3D) structure representing the interaction space (Bennett, FIG. 2, [0041], “Using the detected room features, 3D environment generator 205 can determine which mapped room the user is in (e.g., by matching a detected physical feature with a known physical feature of a mapped room, by matching a detected room identifier with a known identifier of a mapped room, by matching a detected location with a known location of a mapped room, etc.). Additionally and/or alternatively, a user may be provided with an interface with which to select one of the pre-mapped rooms. As such, 3D environment generator 205 can access an associated 3D room model (e.g., by downloading the appropriate 3D room model)”); obtaining virtual object information for displaying a virtual object within the virtual interaction space (Bennett, FIGS. 2-3, [0054], “Generally in authoring mode, as in presentation mode, virtual images of 3D assets are rendered (e.g., via asset rendering components 218, 268) using 3D assets and asset behaviors stored, for example, in a 3D presentation file”); reconstructing a 3D model of the virtual object based on the virtual object information (Bennett, FIGS. 7A-7D, [0092], e.g., “physical effects can be added to a path segment and/or asset state. Generally, a physical effect is a modification of an animation of a 3D asset (e.g., for an asset state and/or asset behavior) implementing one or more designated physical constraints that define movement behavior through space and/or time (e.g., energy and force behaviors)”); and providing, for output and based on the virtual interaction space information and the 3D model, rendering information configured to cause a computing device to present, via a display device, the virtual object within the virtual interaction space (Bennett, e.g., FIGS. 3A-3B, [0058], “FIGS. 3A-B illustrate an exemplary technique for spatial authoring using a 3D interface to move a 3D asset. In FIGS. 3A-3B, co-authors 310 and 340 are physically present in the same room, and each has an AR headset”); Bennett does not teach obtaining, “via an artificial intelligence content generation (AICG) system”, virtual object information, “wherein the virtual object information comprises new content generated by a generative artificial intelligence (AI) model of the AICG system based on input data associated with user interactions and environmental contexts”; “at least one additional virtual object based on the virtual object information and video data comprising information associated with the at least one additional virtual object”. However, Huynh teaches the concepts of obtaining, via an artificial intelligence content generation (AICG) system (Huynh, FIG. 1, [0028], “system 100 generating content via a generative AI mode”), virtual object information, wherein the virtual object information comprises new content (Huynh, e.g., FIGS. 4A-4C, [0051], “generative model renders the 3D model of the lamp 440 (shown in FIG. 4C)”) generated by a generative artificial intelligence (AI) model of the AICG system (Huynh, FIG. 2, [0037], “3D generative model 260”) based on input data associated with user interactions (Huynh, e.g., FIG. 4A, [0046], “After selecting the one or more primitives 422, 424, 426, 428, and 430, the user 300 may then provide an input (e.g., prompt) to generate a 3D model (e.g., 3D object) based on the one or more selected primitives”; [0047], “for example, the user 300 may input: “create a chair using the shape of a cube.””) and environmental contexts (Huynh, e.g., FIGS. 4-5, [0066], “user places 3D model in the AR scene 512”); at least one additional virtual object (Huynh, e.g., FIGS. 4-5, [0047], the “chair” created by “using the shape of a cube”) based on the virtual object information and video data comprising information associated with the at least one additional virtual object (Huynh, FIG. 5, [0061-0062], “visual properties of the AR scene analyzed 504” → “user’s request parameterized to match characteristics of the AR scene”). At the time of effective filing date, it would have been obvious to one of ordinary skill in the art to modify the method of “authoring and presenting 3D presentations” taught by Bennett to further comprise the “3D generative model 260” for further performing the steps of “process 500” in FIG. 5, as taught by Huynh, in order to provide that “the described techniques, such as generating 3D objects for display in AR scenes via generative AI, enable an AR device to dynamically render virtual content based on the context of the user's environment” (Hyunh, [0027]). Bennett in view of Huynh does not teach “the virtual 3D structure comprising a virtual representation of a physical object located within the interaction space, wherein the virtual representation of the physical object comprises a rendered video image of the physical object derived from the interaction space image information”; and “wherein the rendering information is configured to cause the computing device to present the virtual object as appearing to interact with the virtual representation of the physical object within the virtual 3D structure”. However, Petill teaches the concepts of the virtual 3D structure (Petill, FIGS. 3-6, [0047], e.g., “surface reconstruction and decomposition pipeline 44”) comprising a virtual representation of a physical object located within the interaction space (Petill, FIGS. 3-6, [0047], “③ object recognition” and “④ store references to recognized physical objects and associated semantic tags in database”), wherein the virtual representation of the physical object (Petill, FIGS. 3-6, [0048], e.g., “semantic tags 48” for recognized virtual representation of “table”) comprises a rendered video image of the physical object derived from the interaction space image information (Petill, FIGS. 3-6, [0052], “ depth image centered at the physical table object 30E of the physical environment 28 of FIG. 3”); and wherein the rendering information is configured to cause the computing device to present the virtual object as appearing to interact with the virtual representation of the physical object within the virtual 3D structure (Petill, e.g., FIGS. 6-15, [0079], “selecting target virtual object and target physical object from the physical objects and virtual objects in the database based on the identified one or more semantic tags 108” → “performing the determined user specified operation on the target virtual object based on the target physical object 110” → “displaying the target virtual object at a physical location associated with the target physical object 112”). At the time of effective filing date, it would have been obvious to one of ordinary skill in the art to modify the “host device 200” taught by Benett to further perform the steps of, e.g., “object recognition” out of taking “depth image” of the object and interacting with the 3D environment, as shown in FIGS. 4-16 of Petill, in order to “perform the determined user specified operation on the target virtual object based on the target physical object, and display the target virtual object at a physical location associated with the target physical object” (Petill, [0002]). As to claim 4, Bennett teaches the method of claim 1, further comprising: obtaining, from a plurality of users, a plurality of concurrent user inputs indicative of user interaction with the virtual object (Bennett, e.g., FIG. 2, [0046], “the 3D presentation software on the user's device (e.g., host device 200) can host a lobby in which multiple users can participate in authoring and/or presentation modes”; [0048], “Host lobby component 210 can determine whether the 3D asset is checked out and whether the requesting client has an appropriate permission. If the 3D asset is not checked out and the client has permission, host lobby component 210 can grant temporary ownership to the client. The check-out process can be implemented in any number of ways (e.g., using electronic locks, security tokens, role-based access control, etc.), as would be understood by those of ordinary skill in the art”); and prioritizing the plurality of concurrent user inputs based on at least one of predefined criteria or a user role (Bennett, e.g., FIG. 2, [0047], “Users can interact with various 3D assets in a manner that depends on a user's role (e.g., host author, host presenter, client co-author, client co-presenter, client audience member, custom role, etc.)”). As to claim 5, Bennett teaches the method of claim 1, further comprising: obtaining a gesture indication indicative of detection, in user image information associated with the video conference, of a hand gesture made by a participant (Bennett, e.g., see FIGS. 7A-7D, [0089], “After enabling a puppeteering recording mode, author 710 initiates a puppeteering recording for virtual drone 715 with her hand 712 in FIG. 7A by pinching to grab virtual drone 715”); and providing, for output and based on the gesture indication, additional rendering information configured to cause the computing device to present, via the display device, a modification of the virtual object (Bennett, e.g., see FIGS. 7A-7D, [0091], “a recorded 3D path can be edited, for example, by providing a visualization of the recorded 3D path during authoring mode, segmenting the path, and permitting each segment to be selected and manipulated. In this manner, modifications to the 3D path and associated asset behaviors can be applied to a puppeteering animation, and different behaviors can be applied to different path segments and/or asset states”). As to claim 6, Bennett teaches the method of claim 1, further comprising: obtaining an indication of a request, from a participant, to display a 3D virtual participant depiction within the virtual interaction space (Bennett, FIGS. 5-6, [0031], “a client can request or be automatically granted temporary ownership of a 3D asset from the host. By way of nonlimiting example, check-in and check-out requests can be initiated using one or more detected inputs such as gestures”); and providing, for output and based on the indication, additional rendering information configured to cause the computing device to present, via the display device, the 3D virtual participant depiction within the virtual interaction space, wherein the 3D virtual participant depiction comprises a 3D representation of the participant (Bennett, e.g., see FIGS. 5-6, [0084], “audience member 610 initiates a scrubbing gesture in FIG. 6A by positioning his hand 612 relative to a 3D scrubber (visible from his perspective, but not illustrated) and pinching his fingers to grab the scrubber”). As to claim 7, Bennett teaches the method of claim 1, further comprising: obtaining an indication of a request, from a participant, to display a 3D virtual participant depiction within the virtual interaction space; and providing, for output and based on the indication, additional rendering information configured to cause the computing device to present, via the display device, the 3D virtual participant depiction within the virtual interaction space, wherein the 3D virtual participant depiction comprises a 3D representation of the participant, wherein providing the additional rendering information comprises providing the additional rendering information configured to cause the computing device to present the 3D virtual participant depiction in a virtual location within the virtual interaction space based on a location of the participant (Bennett, [0027], “The 3D presentation environment supports multiple co-authors simultaneously authoring, whether in the same room or remotely located (e.g., using avatars to simulate the presence of remote users). As such, spatial authoring can be performed to set behavior parameters for asset behaviors to be triggered by a given scene or beat”; FIGS. 3A-3B, [0058], “Avatar 330 represents the projected location in the room of a remote user wearing a VR headset”). As to claim 8, Bennett teaches the method of claim 1, further comprising: obtaining an indication of a request, from a participant, to display a 3D virtual participant depiction within the virtual interaction space (Bennett, FIGS. 5-6, [0031], “a client can request or be automatically granted temporary ownership of a 3D asset from the host. By way of nonlimiting example, check-in and check-out requests can be initiated using one or more detected inputs such as gestures”); and providing, for output and based on the indication, additional rendering information configured to cause the computing device to present, via the display device, the 3D virtual participant depiction within the virtual interaction space, wherein the 3D virtual participant depiction comprises a 3D representation of the participant, wherein providing the additional rendering information comprises providing the additional rendering information configured to cause the computing device to present the 3D virtual participant depiction in a virtual pose based on a pose of the participant (Bennett, e.g., see FIGS. 5-6, [0084], “audience member 610 initiates a scrubbing gesture in FIG. 6A by positioning his hand 612 relative to a 3D scrubber (visible from his perspective, but not illustrated) and pinching his fingers to grab the scrubber”). As to claim 9, Bennett teaches the method of claim 1, further comprising: obtaining an indication of a request, from a participant, to display a 3D virtual participant depiction within the virtual interaction space (Bennett, FIGS. 5-6, [0031], “a client can request or be automatically granted temporary ownership of a 3D asset from the host. By way of nonlimiting example, check-in and check-out requests can be initiated using one or more detected inputs such as gestures); and providing, for output and based on the indication, additional rendering information configured to cause the computing device to present, via the display device, the 3D virtual participant depiction within the virtual interaction space, wherein the 3D virtual participant depiction comprises a 3D representation of the participant, wherein providing the additional rendering information comprises providing the additional rendering information configured to cause the computing device to present the 3D virtual participant depiction as having a virtual facial expression based on a facial expression of the participant (Bennett, [0075], “project-specific audience interaction metadata can be generated and recorded by evaluating audience gaze and/or facial expressions to determine a state of audience attention and/or interest”). As to claim 10, Bennett teaches the method of claim 1, further comprising: obtaining an indication of a request, from a participant, to display a 3D virtual participant depiction within the virtual interaction space (Bennett, FIGS. 5-6, [0031], “a client can request or be automatically granted temporary ownership of a 3D asset from the host. By way of nonlimiting example, check-in and check-out requests can be initiated using one or more detected inputs such as gestures); and providing, for output and based on the indication, additional rendering information configured to cause the computing device to present, via the display device, the 3D virtual participant depiction within the virtual interaction space, wherein the 3D virtual participant depiction comprises a 3D representation of the participant, wherein providing the additional rendering information comprises providing the additional rendering information configured to cause the computing device to present a virtual action performed by the 3D virtual participant depiction based on an action performed by the participant (Bennett, e.g., FIGS. 5-6, [0084], “As audience member 610 moves his hand 612 through the positions of the scrubbing gesture in FIGS. 6B-6C, 3D animation 614 is rendered and animates in the headsets of each user who joined the 3D presentation”). As to claim 11, it differs from claim 1 only in that it is the non-transitory computer-readable medium storing instructions to be executed by a processor performing the method of claim 1. It recites substantially the same limitations as in claim 1, and Bennett in view of Hyunh and Petill teaches them. Examiner renders the same motivation as in claim 1. Please see claim 1 for detailed analysis. As to claim 12, Bennett teaches the non-transitory computer-readable medium of claim 11, wherein the virtual object comprises a virtual representation of a white board (Bennett, FIGS. 4A-4D, [0079], “drawing with light using a 3D interface … In FIG. 4A, presenter 410 moves her hand to position 412. In this embodiment, presenter 410 initiates a drawing gesture by pinching her fingers and moving her hand through position 413 in FIG. 4B. At substantially the same time, visualization trail 420 is rendered in the headsets of all users who joined the 3D presentation. After drawing the number “9,” presenter 410 releases her pinched fingers to finish the gesture and stop drawing”). As to claim 13, Bennett teaches the non-transitory computer-readable medium of claim 11, wherein the virtual object comprises a virtual representation of a white board, and wherein the virtual representation of the white board is configured to be manipulated by a participant via an input device of the computing device (Bennett, FIGS. 4A-4D, [0079], e.g., “presenter 410 moves her hand to position 414 in FIG. 4C and again initiates a drawing gesture by pinching her fingers and moving her hand through position 415 in FIG. 4D to draw the number “0” with visualization trail 420”). As to claims 16-17, they recite substantially the same limitations as in claims 4-5, respectively, and Bennett teaches them. Please see claims 4-5 for detailed analysis. As to claim 18, it differs from claim 1 only in that it is the system performing the method of claim 1. It recites substantially the same limitations as in claim 1, and Bennett in view of Petill teaches them. Examiner renders the same motivation as in claim 1. Please see claim 1 for detailed analysis. As to claim 19, it recites substantially the same limitations as in claim 6, and Bennett teaches them. Please see claim 6 for detailed analysis. As to claim 20, Bennett teaches the system of claim 18, wherein the one or more processors are further configured to execute the instructions stored in the one or more memories to provide, for output, additional rendering information configured to cause the computing device to present the at least one additional virtual object within the virtual interaction space, wherein the at least one additional virtual object is configured to interact with the virtual object (Bennett, FIGS. 8A-8C, [0096], “a presenter is explaining the potential increase in food output that would result from implementing vertical farming in the city being modeled on the table … The presenter then triggers a beat in which collections 810 and 820 are spatially arranged in discrete bundles in FIG. 8D. As such, the magnitude of the potential increase in food output can be easily visualized by audience members of the 3D presentation”). Claims 2-3 and 14-15 are rejected under 35 U.S.C. 103 as being unpatentable over Bennett et al. (US 2019/0236842 A1) in view of Huynh et al. (US 2025/0252676 A1), Petill et al. (US 2021/0012113 A1) and Becker et al. (US 2023/0377300 A1). As to claim 2, Bennett in view of Huynh and Petill does not teach the method of claim 1, wherein obtaining the virtual object information further comprises: obtaining at least a portion of the virtual object information via at least one of the imaging device or an additional imaging device, the virtual object further comprising a virtual representation of an additional physical object. However, Becker teaches the concept of obtaining at least a portion of the virtual object information via at least one of the imaging device or an additional imaging device (Becker, [0023], “image sensors 116 and/or 117 also optionally include one or more cameras configured to capture movement of physical objects in the real-world environment”; FIG. 3, [0027], e.g., “by including images that include various directions, orientations, and/or varying perspectives/views of a three-dimensional object, computing device 200 can be enabled to generate an accurate three-dimensional representation of the three-dimensional object that can be used to generate graphical representations to display to a user on a computing device or display device”), the virtual object further comprising a virtual representation of an additional physical object (Becker, e.g., FIG. 7, [0042], “For example, and as illustrated in FIG. 7, user interface 700 presents a second representation 702 of three-dimensional object 320 (e.g., tool table). As illustrated in FIG. 7, finalized second representation 702 is a model representation of three-dimensional object 320 based on the images from FIG. 3 including texturized mesh surfaces (not a point cloud representing vertices of the mesh surfaces)”). At the time of effective filing date, it would have been obvious to one of ordinary skill in the art to modify the “system provided for authoring and presenting 3D presentations” taught by Bennett to further comprise the “sensors 116 and/or 117 to capture physical objects in the real-world environment”, as taught by Becker, in order to provide “capturing and/or receiving images of a physical object and generating a three-dimensional virtual representation of the physical object based on the images” (Becker, [0003]). As to claim 3, Becker teaches the method of claim 1, wherein the virtual object information further comprises content provided by a content service (Becker, FIG. 2, [0026], “window 204 in FIG. 2 is shown to include images or capture bundle 206 (e.g., a graphical representation of a stack of images or an object capture bundle … the user can import one or more images from another location on computing system 200 or from a location on another computing system in communication with computing system 200 (e.g., computing system 101)”). Examiner renders the same motivation as in claim 2. As to claims 14-15, they recite the similar limitations as in claims 2-3, respectively, and Becker teaches them. Examiner renders the same motivation as in claims 2-3. Please see claims 2-3 for detailed analysis. Conclusion The prior arts made of record and not relied upon are considered pertinent to applicant’s disclosure: Beauchamp et al. (US 2023/0410436 A1) teach the concept of “employing AR/VR software to generate virtual representations of physical spaces (e.g., house) and sub-spaces (e.g., living room) to preview virtual objects situated in AR/VR virtual environments” (Abs.). Inquiry Any inquiry concerning this communication or earlier communications from the examiner should be directed to RICHARD J HONG whose telephone number is (571) 270-7765. The examiner can normally be reached on 9:00 AM to 6:00 PM EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chanh Nguyen can be reached on (571) 272-7772. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. Sep. 15, 2026 /RICHARD J HONG/Primary Examiner, Art Unit 2623 ***
Read full office action

Prosecution Timeline

Show 13 earlier events
Jun 16, 2026
Final Rejection mailed — §103
Jul 30, 2026
Interview Requested
Aug 11, 2026
Applicant Interview (Telephonic)
Aug 11, 2026
Examiner Interview Summary
Aug 14, 2026
Response after Non-Final Action
Sep 11, 2026
Request for Continued Examination
Sep 14, 2026
Response after Non-Final Action
Sep 17, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12751159
DISPLAY APPARATUS
3y 2m to grant Granted Sep 29, 2026
Patent 12724314
ELECTRONIC PRINTING SYSTEM, METHOD OF OPERATING ELECTRONIC PRINTING SYSTEM, AND METHOD OF FABRICATING IMAGING APPARATUS
2y 7m to grant Granted Sep 01, 2026
Patent 12718730
DISPLAY DRIVING EMPLOYING DITHERING
1y 7m to grant Granted Aug 25, 2026
Patent 12718778
METHODS FOR DELIVERING LOW-GHOSTING PARTIAL UPDATES IN COLOR ELECTROPHORETIC DISPLAYS
1y 7m to grant Granted Aug 25, 2026
Patent 12711892
Driving method for display panel and related source operational amplifier
2y 3m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
79%
Grant Probability
83%
With Interview (+3.9%)
2y 0m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 623 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month