Prosecution Insights
Last updated: October 02, 2026
Application No. 18/888,732

SYSTEMS AND METHODS FOR ARTIFICIAL INTELLIGENCE/MACHINE LEARNING ("AI/ML") SMART GATEWAY

Non-Final OA §103
Filed
Sep 18, 2024
Examiner
THOMAS, MIA M
Art Unit
2665
Tech Center
2600 — Communications
Assignee
Verizon Communications Inc.
OA Round
1 (Non-Final)
86%
Grant Probability
Favorable
1-2
OA Rounds
11m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 86% — above average
86%
Career Allowance Rate
617 granted / 715 resolved
+24.3% vs TC avg
Strong +16% interview lift
Without
With
+15.7%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
16 currently pending
Career history
725
Total Applications
across all art units

Statute-Specific Performance

§101
12.6%
-27.4% vs TC avg
§103
47.4%
+7.4% vs TC avg
§102
17.6%
-22.4% vs TC avg
§112
19.2%
-20.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 715 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This Office Action is responsive to communications filed on 09/18/2024. Claims 1-20 are pending in the instant application. Claims 1, 8 and 15 are independent. An Office Action on the merits follows here below. Specification The abstract of the disclosure is objected to because the patent abstract should be concise statement that does not recite legal phraseology and should not merely recite an instant claim. Correction is required. See MPEP § 608.01[b]. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Andon (US 20210157844 A1) in combination with Holland (US 20220253126 A1). Regarding Claim 1: Andon discloses a device (Refer to para [034]; “Referring again to FIG. 1, the user console 30 in communication with the wearable 20 may comprise a computing device operating software or firmware that is specifically designed to cause the computing device to perform as described below.”) comprising: one or more processors (Refer to para [034 and 041]; “In general, the processor 32…Examples of such hardware can include video or audio encoders/decoders, digital signal processors (DSP), and the like.”) configured to: receive a first set of artificial intelligence/machine learning (“AI/ML”) models (Refer to para [044]; “In one embodiment, the object recognition and tracking module 116 may utilize image processing techniques such as boundary/edge detection, pattern recognition, and/or machine learning techniques to recognize the wearable 20 and to gauge its movement relative to its environment or in a more object-centered coordinate system.”) receive first locally captured video information (Refer to para [027 and 036]; “In some embodiments, the video sources 42 and/or audio sources 44 may be local to the user console 30 or provided on a common local area network (LAN).”). Andon does not expressly disclose a second AI/ML algorithm. Holland teaches “… techniques and systems for capturing images (such as self-images or “selfies”) in extended reality environments.” In the same field of “capturing the one or more frames of the real-world environment can include capturing image data using an outward-facing camera system of the extended reality system….” Holland teaches a processor (Refer to para [056]; “one or more compute components 210 can include a central processing unit (CPU) 212, a graphics processing unit (GPU) 214, a digital signal processor (DSP) 216, and/or an image signal processor (ISP) 218.”) to generate a second set of AI/ML models based on the first set of AI/ML models and the locally captured video information (Refer to para [081]; “In some cases, the avatar engine 304(B) can implement a second machine learning model that generates the avatar 318(B) (e.g., the final version of the avatar 318).”) receive second locally captured video information (Refer to para [013]; “In some examples, obtaining the second digital representation of the user can include causing a server configured to generate digital representations of users to generate the second digital representation of the user based on implementing the second machine learning algorithm.”) identify, based on the second locally captured video information and the second set of AI/ML models, one or more classifications for the second locally captured video information (Refer to para [124-127]; “Once the neural network 800 is trained, it can be referred to as a trained neural network, which can be used to classify one or more objects. For example, an interconnection between nodes can represent a piece of information learned about the interconnected nodes. The interconnection can have a tunable numeric weight that can be tuned (e.g., based on a training dataset), allowing the neural network 800 to be adaptive to inputs and able to learn as more and more data is processed.”) and output, via a network and to an action system (Refer to para [063 and 151]; “The output of one or more sensors (e.g., the accelerometer 204, the gyroscope 206, one or more IMUs, and/or other sensors) can be used by the extended reality engine 220 to determine a pose of the extended reality system 200 (also referred to as the head pose) and/or the pose of the image sensor 202 (or other camera of the extended reality system 200). In some cases, the pose of the extended reality system 200 and the pose of the image sensor 202 (or other camera) can be the same.”) the one or more classifications, without outputting the second locally captured video information via the network (Refer to para [080]; “In some examples, the avatar engine 304 can implement multiple (e.g., two or more) machine learning models configured to generate different versions of the avatar 318. For instance, as shown in FIG. 3A and FIG. 3B, the avatar engine 304 can optionally include an avatar engine 304(A) and an avatar engine 304(B). In some cases, the avatar engine 304(A) can generate or obtain a first version of the avatar (denoted as avatar 318(A)) and a second version of the avatar (denoted as avatar 318(B)). In one example, the avatar 318(A) can correspond to a preview or initial version of the avatar 318. In such an example, the avatar 318(B) can correspond to a final version of the avatar 318. In some cases, the avatar 318(A) (the preview or initial version) can be a lower fidelity version as compared to the avatar 318(B) (the final version). In some examples, the avatar engine 304(A) can generate and/or display the avatar 318(A) before the user pose 312 and/or the background frame 314 are captured (e.g., to facilitate capturing a desired user pose 312 and/or background frame 314). Once the user pose 312 and/or the background frame 314 are captured, the avatar engine 304(B) can generate the avatar 318(B). The composition engine 308 can generate the self-image frame 316 using the avatar 318(B). For instance, the lower fidelity avatar (e.g., avatar 318(A)) is displayed by the XR device or system when the user in in their current pose. The user can then operate the XR device or system to capture a final pose (e.g., used for generating avatar 318(B)), which the composition engine 308 can use for the composition when generating the self-image frame 316. A benefit of presenting to a lower fidelity avatar (e.g., avatar 318(A)) during the pose capture stage or the background image capture stage can allow the composition to be performed with lower processing power before the final composition is performed using the higher fidelity avatar (e.g., avatar 318(B)).”) wherein the action system identifies a particular action associated with the one or more classifications, and performs the identified particular action (Refer to para [065, 071 and 080]; “For instance, in some examples, the compute components 210 can perform tracking using computer vision-based tracking, model-based tracking, and/or simultaneous localization and mapping (SLAM) techniques. For instance, the compute components 210 can perform SLAM or can be in communication (wired or wireless) with a SLAM engine (not shown). SLAM refers to a class of techniques where a map of an environment (e.g., a map of an environment being modeled by extended reality system 200) is created while simultaneously tracking the pose of a camera (e.g., image sensor 202) and/or the extended reality system 200 relative to that map. The map can be referred to as a SLAM map, and can be 3D. The SLAM techniques can be performed using color or grayscale image data captured by the image sensor 202 (and/or other camera of the extended reality system 200), and can be used to generate estimates of 6DOF pose measurements of the image sensor 202 and/or the extended reality system 200. Such a SLAM technique configured to perform 6DOF tracking can be referred to as 6DOF SLAM. In some cases, the output of the one or more sensors (e.g., the accelerometer 204, the gyroscope 206, one or more IMUs, and/or other sensors) can be used to estimate, correct, and/or otherwise adjust the estimated pose.”). Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify Andon by adding a second machine learning model as taught by Holland as rejected above. The suggestion/motivation for combining the teachings of Andon and Holland would have been in order to enhance the image processor in order to “determine the user pose using any combination of inward-facing cameras, outward-facing cameras, and/or additional cameras of the XR device.” (at para [075], Holland). Therefore, it would have been obvious to one of ordinary skill in the art to combine the teachings of Andon and Holland in order to obtain the specified claimed elements of Claim 1. It is for at least the aforementioned reasons that the Examiner has reached a conclusion of obviousness with respect to the claim in question. Regarding Claim 2: Andon discloses one or more AI/ML processing units that identify the one or more classifications for the second locally captured video information (Refer to para [044]; “In one embodiment, the object recognition and tracking module 116 may utilize image processing techniques such as boundary/edge detection, pattern recognition, and/or machine learning techniques to recognize the wearable 20 and to gauge its movement relative to its environment or in a more object-centered coordinate system.”). Regarding Claim 3: Andon discloses network circuitry to implement at least one of: a Local Area Network (“LAN”), or a wireless LAN (“WLAN”) (Refer to para [036]; “In some embodiments, the video sources 42 and/or audio sources 44 may be local to the user console 30 or provided on a common local area network (LAN).”). Regarding Claim 4: Andon discloses the network circuitry is first network circuitry, wherein the device further comprises second network circuitry to communicate with the network (Refer to para [062]; “In this embodiment, the smart/connected wearable 220 may include enough processing functionality and communications capabilities to transmit motion data to the distributed network, and may even be able to communicate with one or more audio output devices 46 or visual effect controllers 48 through a communication protocol such as BLUETOOTH. Likewise, the smart/connected wearable 220 may also serve as the user console 30 for one or more other wearables 20 that lack additional processing capabilities.”). Regarding Claim 5: Andon discloses the second network circuitry includes wireless network circuitry that implements at least one of a Long-Term Evolution (“LTE”) radio access technology (“RAT”) or a Fifth Generation (“5G”) RAT (Refer to para [033]; “The communications circuitry 24 coupled with the motion sensing circuit 22 is configured to outwardly transmit the data stream to the user console 30. The communications circuitry 24 may include one or more transceivers, antenna, and/or memory that may be required to facilitate the data transmission in real time or near real time. This data transmission may occur according to any suitable wireless standard, however particularly suited communication protocols include those according to any of the following standards or industry recognized protocols: IEEE 802.11, 802.15, 1914.1, 1914.3; BLUETOOTH or BLUETOOTH LOW ENERGY (or other similar protocols/standards set by the Bluetooth SIG); 4G LTE cellular, 5G, 5G NR; or similar wireless data communication protocols.”). Regarding Claim 6: Andon discloses the locally captured video information is received from one or more cameras that are communicatively coupled to the device via the LAN or the WLAN (Refer to para [036]; “The video sources 42 may include live video streams (e.g., from a digital camera), previously recorded video streams, digital memory storing one or more previously recorded videos, and the like. The audio sources 44 may include one or more musical instruments, keyboards, synthesizers, data files, collections of audio samples, and the like. In some embodiments, the video sources 42 and/or audio sources 44 may be local to the user console 30 or provided on a common local area network (LAN). In other embodiments, however, one or both of these sources 42, 44 may be remote from the user console 30 and/or may be hosted by a computer that is accessible only though the console's network connection.”). Regarding Claim 7: Holland discloses the second set of AI/ML models are maintained locally, without outputting the second set of AI/ML models via the network (Refer to para [080]; “In some examples, the avatar engine 304 can implement multiple (e.g., two or more) machine learning models configured to generate different versions of the avatar 318. For instance, as shown in FIG. 3A and FIG. 3B, the avatar engine 304 can optionally include an avatar engine 304(A) and an avatar engine 304(B). In some cases, the avatar engine 304(A) can generate or obtain a first version of the avatar (denoted as avatar 318(A)) and a second version of the avatar (denoted as avatar 318(B)). In one example, the avatar 318(A) can correspond to a preview or initial version of the avatar 318. In such an example, the avatar 318(B) can correspond to a final version of the avatar 318. In some cases, the avatar 318(A) (the preview or initial version) can be a lower fidelity version as compared to the avatar 318(B) (the final version). In some examples, the avatar engine 304(A) can generate and/or display the avatar 318(A) before the user pose 312 and/or the background frame 314 are captured (e.g., to facilitate capturing a desired user pose 312 and/or background frame 314). Once the user pose 312 and/or the background frame 314 are captured, the avatar engine 304(B) can generate the avatar 318(B). The composition engine 308 can generate the self-image frame 316 using the avatar 318(B). For instance, the lower fidelity avatar (e.g., avatar 318(A)) is displayed by the XR device or system when the user in in their current pose. The user can then operate the XR device or system to capture a final pose (e.g., used for generating avatar 318(B)), which the composition engine 308 can use for the composition when generating the self-image frame 316. A benefit of presenting to a lower fidelity avatar (e.g., avatar 318(A)) during the pose capture stage or the background image capture stage can allow the composition to be performed with lower processing power before the final composition is performed using the higher fidelity avatar (e.g., avatar 318(B)).”). Regarding Claim 8: Andon discloses a non-transitory computer-readable medium, storing a plurality of processor-executable instructions to (Refer to para [041]; “Each of the described modules or blocks may comprise computer executable code, stored in memory as software or firmware, such that when executed, the processor 32 may perform the specified functions.”): receive a first set of artificial intelligence/machine learning (“AI/ML”) models (Refer to para [044]; “In one embodiment, the object recognition and tracking module 116 may utilize image processing techniques such as boundary/edge detection, pattern recognition, and/or machine learning techniques to recognize the wearable 20 and to gauge its movement relative to its environment or in a more object-centered coordinate system.”) receive first locally captured video information (Refer to para [027 and 036]; “In some embodiments, the video sources 42 and/or audio sources 44 may be local to the user console 30 or provided on a common local area network (LAN).”). Andon does not expressly disclose a second AI/ML algorithm. Holland teaches “… techniques and systems for capturing images (such as self-images or “selfies”) in extended reality environments.” In the same field of “capturing the one or more frames of the real-world environment can include capturing image data using an outward-facing camera system of the extended reality system….” Holland teaches a processor (Refer to para [056]; “one or more compute components 210 can include a central processing unit (CPU) 212, a graphics processing unit (GPU) 214, a digital signal processor (DSP) 216, and/or an image signal processor (ISP) 218.”) to generate a second set of AI/ML models based on the first set of AI/ML models and the locally captured video information (Refer to para [081]; “In some cases, the avatar engine 304(B) can implement a second machine learning model that generates the avatar 318(B) (e.g., the final version of the avatar 318).”) receive second locally captured video information (Refer to para [013]; “In some examples, obtaining the second digital representation of the user can include causing a server configured to generate digital representations of users to generate the second digital representation of the user based on implementing the second machine learning algorithm.”) identify, based on the second locally captured video information and the second set of AI/ML models, one or more classifications for the second locally captured video information (Refer to para [124-127]; “Once the neural network 800 is trained, it can be referred to as a trained neural network, which can be used to classify one or more objects. For example, an interconnection between nodes can represent a piece of information learned about the interconnected nodes. The interconnection can have a tunable numeric weight that can be tuned (e.g., based on a training dataset), allowing the neural network 800 to be adaptive to inputs and able to learn as more and more data is processed.”) and output, via a network and to an action system (Refer to para [063 and 151]; “The output of one or more sensors (e.g., the accelerometer 204, the gyroscope 206, one or more IMUs, and/or other sensors) can be used by the extended reality engine 220 to determine a pose of the extended reality system 200 (also referred to as the head pose) and/or the pose of the image sensor 202 (or other camera of the extended reality system 200). In some cases, the pose of the extended reality system 200 and the pose of the image sensor 202 (or other camera) can be the same.”) the one or more classifications, without outputting the second locally captured video information via the network (Refer to para [080]; “In some examples, the avatar engine 304 can implement multiple (e.g., two or more) machine learning models configured to generate different versions of the avatar 318. For instance, as shown in FIG. 3A and FIG. 3B, the avatar engine 304 can optionally include an avatar engine 304(A) and an avatar engine 304(B). In some cases, the avatar engine 304(A) can generate or obtain a first version of the avatar (denoted as avatar 318(A)) and a second version of the avatar (denoted as avatar 318(B)). In one example, the avatar 318(A) can correspond to a preview or initial version of the avatar 318. In such an example, the avatar 318(B) can correspond to a final version of the avatar 318. In some cases, the avatar 318(A) (the preview or initial version) can be a lower fidelity version as compared to the avatar 318(B) (the final version). In some examples, the avatar engine 304(A) can generate and/or display the avatar 318(A) before the user pose 312 and/or the background frame 314 are captured (e.g., to facilitate capturing a desired user pose 312 and/or background frame 314). Once the user pose 312 and/or the background frame 314 are captured, the avatar engine 304(B) can generate the avatar 318(B). The composition engine 308 can generate the self-image frame 316 using the avatar 318(B). For instance, the lower fidelity avatar (e.g., avatar 318(A)) is displayed by the XR device or system when the user in in their current pose. The user can then operate the XR device or system to capture a final pose (e.g., used for generating avatar 318(B)), which the composition engine 308 can use for the composition when generating the self-image frame 316. A benefit of presenting to a lower fidelity avatar (e.g., avatar 318(A)) during the pose capture stage or the background image capture stage can allow the composition to be performed with lower processing power before the final composition is performed using the higher fidelity avatar (e.g., avatar 318(B)).”) wherein the action system identifies a particular action associated with the one or more classifications, and performs the identified particular action (Refer to para [065, 071 and 080]; “For instance, in some examples, the compute components 210 can perform tracking using computer vision-based tracking, model-based tracking, and/or simultaneous localization and mapping (SLAM) techniques. For instance, the compute components 210 can perform SLAM or can be in communication (wired or wireless) with a SLAM engine (not shown). SLAM refers to a class of techniques where a map of an environment (e.g., a map of an environment being modeled by extended reality system 200) is created while simultaneously tracking the pose of a camera (e.g., image sensor 202) and/or the extended reality system 200 relative to that map. The map can be referred to as a SLAM map, and can be 3D. The SLAM techniques can be performed using color or grayscale image data captured by the image sensor 202 (and/or other camera of the extended reality system 200), and can be used to generate estimates of 6DOF pose measurements of the image sensor 202 and/or the extended reality system 200. Such a SLAM technique configured to perform 6DOF tracking can be referred to as 6DOF SLAM. In some cases, the output of the one or more sensors (e.g., the accelerometer 204, the gyroscope 206, one or more IMUs, and/or other sensors) can be used to estimate, correct, and/or otherwise adjust the estimated pose.”). Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify Andon by adding a second machine learning model as taught by Holland as rejected above. The suggestion/motivation for combining the teachings of Andon and Holland would have been in order to enhance the image processor in order to “determine the user pose using any combination of inward-facing cameras, outward-facing cameras, and/or additional cameras of the XR device.” (at para [075], Holland). Therefore, it would have been obvious to one of ordinary skill in the art to combine the teachings of Andon and Holland in order to obtain the specified claimed elements of Claim 8. It is for at least the aforementioned reasons that the Examiner has reached a conclusion of obviousness with respect to the claim in question. Regarding Claim 9: Andon discloses identifying the one or more classifications for the second locally captured video information is performed by one or more AI/ML processing units (Refer to para [044]; “In one embodiment, the object recognition and tracking module 116 may utilize image processing techniques such as boundary/edge detection, pattern recognition, and/or machine learning techniques to recognize the wearable 20 and to gauge its movement relative to its environment or in a more object-centered coordinate system.”). Regarding Claim 10: Andon discloses a device that executes the processor-executable instructions comprises: network circuitry to implement at least one of: a Local Area Network (“LAN”), or a wireless LAN (“WLAN”) (Refer to para [036]; “In some embodiments, the video sources 42 and/or audio sources 44 may be local to the user console 30 or provided on a common local area network (LAN).”). Regarding Claim 11: Andon discloses the network circuitry is first network circuitry, wherein the device further comprises second network circuitry to communicate with the network (Refer to para [062]; “In this embodiment, the smart/connected wearable 220 may include enough processing functionality and communications capabilities to transmit motion data to the distributed network, and may even be able to communicate with one or more audio output devices 46 or visual effect controllers 48 through a communication protocol such as BLUETOOTH. Likewise, the smart/connected wearable 220 may also serve as the user console 30 for one or more other wearables 20 that lack additional processing capabilities.”). Regarding Claim 12: Andon discloses the second network circuitry includes wireless network circuitry that implements at least one of a Long-Term Evolution (“LTE”) radio access technology (“RAT”) or a Fifth Generation (“5G”) RAT (Refer to para [033]; “The communications circuitry 24 may include one or more transceivers, antenna, and/or memory that may be required to facilitate the data transmission in real time or near real time. This data transmission may occur according to any suitable wireless standard, however particularly suited communication protocols include those according to any of the following standards or industry recognized protocols: IEEE 802.11, 802.15, 1914.1, 1914.3; BLUETOOTH or BLUETOOTH LOW ENERGY (or other similar protocols/standards set by the Bluetooth SIG); 4G LTE cellular, 5G, 5G NR; or similar wireless data communication protocols.”). Regarding Claim 13: Andon discloses the locally captured video information is received from one or more cameras that are communicatively coupled to the device via the LAN or the WLAN (Refer to para [036]; “The video sources 42 may include live video streams (e.g., from a digital camera), previously recorded video streams, digital memory storing one or more previously recorded videos, and the like. The audio sources 44 may include one or more musical instruments, keyboards, synthesizers, data files, collections of audio samples, and the like. In some embodiments, the video sources 42 and/or audio sources 44 may be local to the user console 30 or provided on a common local area network (LAN). In other embodiments, however, one or both of these sources 42, 44 may be remote from the user console 30 and/or may be hosted by a computer that is accessible only though the console's network connection.”). Regarding Claim 14: Holland discloses non-transitory computer-readable medium of claim 8, wherein the second set of AI/ML models are maintained locally, without outputting the second set of AI/ML models via the network (Refer to para [080]; “In some examples, the avatar engine 304 can implement multiple (e.g., two or more) machine learning models configured to generate different versions of the avatar 318. For instance, as shown in FIG. 3A and FIG. 3B, the avatar engine 304 can optionally include an avatar engine 304(A) and an avatar engine 304(B). In some cases, the avatar engine 304(A) can generate or obtain a first version of the avatar (denoted as avatar 318(A)) and a second version of the avatar (denoted as avatar 318(B)). In one example, the avatar 318(A) can correspond to a preview or initial version of the avatar 318. In such an example, the avatar 318(B) can correspond to a final version of the avatar 318. In some cases, the avatar 318(A) (the preview or initial version) can be a lower fidelity version as compared to the avatar 318(B) (the final version). In some examples, the avatar engine 304(A) can generate and/or display the avatar 318(A) before the user pose 312 and/or the background frame 314 are captured (e.g., to facilitate capturing a desired user pose 312 and/or background frame 314). Once the user pose 312 and/or the background frame 314 are captured, the avatar engine 304(B) can generate the avatar 318(B). The composition engine 308 can generate the self-image frame 316 using the avatar 318(B). For instance, the lower fidelity avatar (e.g., avatar 318(A)) is displayed by the XR device or system when the user in in their current pose. The user can then operate the XR device or system to capture a final pose (e.g., used for generating avatar 318(B)), which the composition engine 308 can use for the composition when generating the self-image frame 316. A benefit of presenting to a lower fidelity avatar (e.g., avatar 318(A)) during the pose capture stage or the background image capture stage can allow the composition to be performed with lower processing power before the final composition is performed using the higher fidelity avatar (e.g., avatar 318(B)).”). Regarding Claim 15: Andon discloses method (Refer to para [038]; “FIG. 3 schematically illustrates a method of operation for the present system 10 from the perspective of the user console 30. As shown, the method begins by establishing a motion correspondence table (at 62) that correlates sensed motion with a desired audio/visual response.”) comprising: receiving a first set of artificial intelligence/machine learning (“AI/ML”) models (Refer to para [044]; “In one embodiment, the object recognition and tracking module 116 may utilize image processing techniques such as boundary/edge detection, pattern recognition, and/or machine learning techniques to recognize the wearable 20 and to gauge its movement relative to its environment or in a more object-centered coordinate system.”) receiving first locally captured video information (Refer to para [027 and 036]; “In some embodiments, the video sources 42 and/or audio sources 44 may be local to the user console 30 or provided on a common local area network (LAN).”). Andon does not expressly disclose a second AI/ML algorithm. Holland teaches “… techniques and systems for capturing images (such as self-images or “selfies”) in extended reality environments.” In the same field of “capturing the one or more frames of the real-world environment can include capturing image data using an outward-facing camera system of the extended reality system….” Holland teaches a processor (Refer to para [056]; “one or more compute components 210 can include a central processing unit (CPU) 212, a graphics processing unit (GPU) 214, a digital signal processor (DSP) 216, and/or an image signal processor (ISP) 218.”) to generate a second set of AI/ML models based on the first set of AI/ML models and the locally captured video information (Refer to para [081]; “In some cases, the avatar engine 304(B) can implement a second machine learning model that generates the avatar 318(B) (e.g., the final version of the avatar 318).”) receive second locally captured video information (Refer to para [013]; “In some examples, obtaining the second digital representation of the user can include causing a server configured to generate digital representations of users to generate the second digital representation of the user based on implementing the second machine learning algorithm.”) identify, based on the second locally captured video information and the second set of AI/ML models, one or more classifications for the second locally captured video information (Refer to para [124-127]; “Once the neural network 800 is trained, it can be referred to as a trained neural network, which can be used to classify one or more objects. For example, an interconnection between nodes can represent a piece of information learned about the interconnected nodes. The interconnection can have a tunable numeric weight that can be tuned (e.g., based on a training dataset), allowing the neural network 800 to be adaptive to inputs and able to learn as more and more data is processed.”) and outputting, via a network and to an action system, the one or more classifications (Refer to para [063 and 151]; “The output of one or more sensors (e.g., the accelerometer 204, the gyroscope 206, one or more IMUs, and/or other sensors) can be used by the extended reality engine 220 to determine a pose of the extended reality system 200 (also referred to as the head pose) and/or the pose of the image sensor 202 (or other camera of the extended reality system 200). In some cases, the pose of the extended reality system 200 and the pose of the image sensor 202 (or other camera) can be the same.”) without outputting the second locally captured video information via the network (Refer to para [080]; “In some examples, the avatar engine 304 can implement multiple (e.g., two or more) machine learning models configured to generate different versions of the avatar 318. For instance, as shown in FIG. 3A and FIG. 3B, the avatar engine 304 can optionally include an avatar engine 304(A) and an avatar engine 304(B). In some cases, the avatar engine 304(A) can generate or obtain a first version of the avatar (denoted as avatar 318(A)) and a second version of the avatar (denoted as avatar 318(B)). In one example, the avatar 318(A) can correspond to a preview or initial version of the avatar 318. In such an example, the avatar 318(B) can correspond to a final version of the avatar 318. In some cases, the avatar 318(A) (the preview or initial version) can be a lower fidelity version as compared to the avatar 318(B) (the final version). In some examples, the avatar engine 304(A) can generate and/or display the avatar 318(A) before the user pose 312 and/or the background frame 314 are captured (e.g., to facilitate capturing a desired user pose 312 and/or background frame 314). Once the user pose 312 and/or the background frame 314 are captured, the avatar engine 304(B) can generate the avatar 318(B). The composition engine 308 can generate the self-image frame 316 using the avatar 318(B). For instance, the lower fidelity avatar (e.g., avatar 318(A)) is displayed by the XR device or system when the user in in their current pose. The user can then operate the XR device or system to capture a final pose (e.g., used for generating avatar 318(B)), which the composition engine 308 can use for the composition when generating the self-image frame 316. A benefit of presenting to a lower fidelity avatar (e.g., avatar 318(A)) during the pose capture stage or the background image capture stage can allow the composition to be performed with lower processing power before the final composition is performed using the higher fidelity avatar (e.g., avatar 318(B)).”) wherein the action system identifies a particular action associated with the one or more classifications, and performs the identified particular action (Refer to para [065, 071 and 080]; “For instance, in some examples, the compute components 210 can perform tracking using computer vision-based tracking, model-based tracking, and/or simultaneous localization and mapping (SLAM) techniques. For instance, the compute components 210 can perform SLAM or can be in communication (wired or wireless) with a SLAM engine (not shown). SLAM refers to a class of techniques where a map of an environment (e.g., a map of an environment being modeled by extended reality system 200) is created while simultaneously tracking the pose of a camera (e.g., image sensor 202) and/or the extended reality system 200 relative to that map. The map can be referred to as a SLAM map, and can be 3D. The SLAM techniques can be performed using color or grayscale image data captured by the image sensor 202 (and/or other camera of the extended reality system 200), and can be used to generate estimates of 6DOF pose measurements of the image sensor 202 and/or the extended reality system 200. Such a SLAM technique configured to perform 6DOF tracking can be referred to as 6DOF SLAM. In some cases, the output of the one or more sensors (e.g., the accelerometer 204, the gyroscope 206, one or more IMUs, and/or other sensors) can be used to estimate, correct, and/or otherwise adjust the estimated pose.”). Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify Andon by adding a second machine learning model as taught by Holland as rejected above. The suggestion/motivation for combining the teachings of Andon and Holland would have been in order to enhance the image processor in order to “determine the user pose using any combination of inward-facing cameras, outward-facing cameras, and/or additional cameras of the XR device.” (at para [075], Holland). Therefore, it would have been obvious to one of ordinary skill in the art to combine the teachings of Andon and Holland in order to obtain the specified claimed elements of Claim 15. It is for at least the aforementioned reasons that the Examiner has reached a conclusion of obviousness with respect to the claim in question. Regarding Claim 16: Andon discloses identifying the one or more classifications for the second locally captured video information is performed by one or more AI/ML processing units (Refer to para [044]; “In one embodiment, the object recognition and tracking module 116 may utilize image processing techniques such as boundary/edge detection, pattern recognition, and/or machine learning techniques to recognize the wearable 20 and to gauge its movement relative to its environment or in a more object-centered coordinate system.”). Regarding Claim 17: Andon discloses a device that receives the first and second locally captured video information comprises network circuitry to implement at least one of: a Local Area Network (“LAN”), or a wireless LAN (“WLAN”) (Refer to para [036]; “In some embodiments, the video sources 42 and/or audio sources 44 may be local to the user console 30 or provided on a common local area network (LAN).”). Regarding Claim 18: Andon discloses the network circuitry is first network circuitry, wherein the device further comprises second network circuitry to communicate with the network (Refer to para [062]; “In this embodiment, the smart/connected wearable 220 may include enough processing functionality and communications capabilities to transmit motion data to the distributed network, and may even be able to communicate with one or more audio output devices 46 or visual effect controllers 48 through a communication protocol such as BLUETOOTH. Likewise, the smart/connected wearable 220 may also serve as the user console 30 for one or more other wearables 20 that lack additional processing capabilities.”). Regarding Claim 19: Andon discloses the second network circuitry includes wireless network circuitry that implements at least one of a Long-Term Evolution (“LTE”) radio access technology (“RAT”) or a Fifth Generation (“5G”) RAT (Refer to para [033]; “The communications circuitry 24 may include one or more transceivers, antenna, and/or memory that may be required to facilitate the data transmission in real time or near real time. This data transmission may occur according to any suitable wireless standard, however particularly suited communication protocols include those according to any of the following standards or industry recognized protocols: IEEE 802.11, 802.15, 1914.1, 1914.3; BLUETOOTH or BLUETOOTH LOW ENERGY (or other similar protocols/standards set by the Bluetooth SIG); 4G LTE cellular, 5G, 5G NR; or similar wireless data communication protocols.”). Regarding Claim 20: Holland teaches the locally captured video information is received from one or more cameras that are communicatively coupled to the device via the LAN or the WLAN (Refer to para [080]; “In some examples, the avatar engine 304 can implement multiple (e.g., two or more) machine learning models configured to generate different versions of the avatar 318. For instance, as shown in FIG. 3A and FIG. 3B, the avatar engine 304 can optionally include an avatar engine 304(A) and an avatar engine 304(B). In some cases, the avatar engine 304(A) can generate or obtain a first version of the avatar (denoted as avatar 318(A)) and a second version of the avatar (denoted as avatar 318(B)). In one example, the avatar 318(A) can correspond to a preview or initial version of the avatar 318. In such an example, the avatar 318(B) can correspond to a final version of the avatar 318. In some cases, the avatar 318(A) (the preview or initial version) can be a lower fidelity version as compared to the avatar 318(B) (the final version). In some examples, the avatar engine 304(A) can generate and/or display the avatar 318(A) before the user pose 312 and/or the background frame 314 are captured (e.g., to facilitate capturing a desired user pose 312 and/or background frame 314). Once the user pose 312 and/or the background frame 314 are captured, the avatar engine 304(B) can generate the avatar 318(B). The composition engine 308 can generate the self-image frame 316 using the avatar 318(B). For instance, the lower fidelity avatar (e.g., avatar 318(A)) is displayed by the XR device or system when the user in in their current pose. The user can then operate the XR device or system to capture a final pose (e.g., used for generating avatar 318(B)), which the composition engine 308 can use for the composition when generating the self-image frame 316. A benefit of presenting to a lower fidelity avatar (e.g., avatar 318(A)) during the pose capture stage or the background image capture stage can allow the composition to be performed with lower processing power before the final composition is performed using the higher fidelity avatar (e.g., avatar 318(B)).”). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to MIA M THOMAS whose telephone number is (571)270-1583. The examiner can normally be reached M-Th 8:30am-4:30pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Stephen (Steve) Koziol can be reached at (408) 918-7630. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. MIA M. THOMAS Primary Examiner Art Unit 2665 /MIA M THOMAS/Primary Examiner Art Unit 2665
Read full office action

Prosecution Timeline

Sep 18, 2024
Application Filed
Jul 10, 2026
Non-Final Rejection mailed — §103
Jul 21, 2026
Interview Requested

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749279
IDENTIFYING ANOMALY LOCATION
3y 4m to grant Granted Sep 29, 2026
Patent 12749578
SYSTEMS AND METHODS FOR USER AUTHENTICATION IN VIDEO COMMUNICATIONS
2y 3m to grant Granted Sep 29, 2026
Patent 12749284
REDUCING A SEARCH SPACE FOR ITEM IDENTIFICATION USING MACHINE LEARNING
2y 1m to grant Granted Sep 29, 2026
Patent 12743783
GENERATING IMPROVED PANOPTIC SEGMENTED DIGITAL IMAGES BASED ON PANOPTIC SEGMENTATION NEURAL NETWORKS THAT UTILIZE EXEMPLAR UNKNOWN OBJECT CLASSES
2y 11m to grant Granted Sep 22, 2026
Patent 12738085
PADDING SENSITIVE BATCH CONSTRUCTION FOR TEXT RECOGNITION
2y 9m to grant Granted Sep 15, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
86%
Grant Probability
99%
With Interview (+15.7%)
2y 11m (~11m remaining)
Median Time to Grant
Low
PTA Risk
Based on 715 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month