Prosecution Insights
Last updated: September 17, 2026
Application No. 19/078,655

REAL-TIME CONVERSION OF 2D VIDEO INTO 3D HOLOGRAPHIC VIDEO CONTENT USING A HEADSET DEVICE

Non-Final OA §103§112
Filed
Mar 13, 2025
Priority
Mar 15, 2024 — provisional 63/565,945
Examiner
THERKORN, ERICA GERALDINE
Art Unit
Tech Center
Assignee
Yosemite Vision Inc.
OA Round
1 (Non-Final)
100%
Grant Probability
Favorable
1-2
OA Rounds
7m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 100% — above average
100%
Career Allowance Rate
3 granted / 3 resolved
+40.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 1m
Avg Prosecution
14 currently pending
Career history
15
Total Applications
across all art units

Statute-Specific Performance

§101
12.2%
-27.8% vs TC avg
§103
62.2%
+22.2% vs TC avg
§112
25.7%
-14.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 3 resolved cases

Office Action

§103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Objections Claim 14 is objected to because of the following informalities: Claim 14 it recites "The system of claim 14, wherein..." asserting dependency on itself. Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1, 10, 21, 23, 30, 41-42, etc are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Regarding claim 1, it recites “convert the first one or more frames into a first 3D depth map including creating an initial 3D model of the subject …” It is unclear what is being recited as including creating an initial 3D model. The claim could be interpreted as meaning that the conversion of the first one or more frames into a 3D depth map includes creating an initial 3D model or it could be interpreted as meaning the 3D depth map itself includes creating an initial 3D model. In either case it is unclear how either would make sense. Creating a depth map involves estimating or measuring distance data and it is unclear how this process would include creating an initial 3D model. A depth map itself stores distance data and it is unclear how a depth map itself would include creating an initial 3D model. MPEP states, “claims are required to be cast in clear—as opposed to ambiguous, vague, indefinite—terms.” MPEP 2173.02.II. Additionally, for either possible interpretation, the specification conflicts the claim limitation because the specification instead supports generating a depth map and then creating a corresponding 3D model using the depth map information (specification, para [0028]). There is a conflict between the claimed subject matter and the specification disclosure. This renders the scope of the claim uncertain as inconsistency with the specification disclosure makes the claim take on an unreasonable degree of uncertainty. Reference can be made to MPEP 2173.03. For the purpose of compact prosecution and art rejection, the examiner will treat “convert the first one or more frames into a first 3D depth map including creating an initial 3D model” to mean convert the first one or more frames into a first 3D depth map and create an initial 3D model. Regarding claims 21, 41, and 42, they are rejected using the same citations and rationales described in the rejection of claim 1. Similarly regarding claim 1, it recites “convert the subsequent frame into a subsequent 3D depth map including a current 3D model of the subject.” A depth map as cannot include a 3D model, thus the claim becomes unclear and indefinite. A depth map is an image or image channel that contains information relating to the distance of the surfaces of scene objects from a viewpoint. The relevant sections of the specification do not provide guidance or clarity as to how a depth map can include a 3D model (specification, para [0039]-[0041]). MPEP states, “claims are required to be cast in clear—as opposed to ambiguous, vague, indefinite—terms.” MPEP 2173.02.II. Due to the lack of clarity of the aforementioned claim limitation, a person of ordinary skill in the art could not interpret the metes and bounds of the claim. For the purpose of compact prosecution and art rejection, the examiner will treat “convert the subsequent frame into a subsequent 3D depth map including a current 3D model of the subject” to mean convert the subsequent frame into a subsequent 3D depth map where the frame includes a subject corresponding to a 3D model. Further regarding claim 1, it recites the limitation "the first depth map" in line 14 of the claim. There is insufficient antecedent basis for this limitation in the claim. Regarding claim 41, it is rejected using the same citations and rationales described in the rejection of claim 1. Further regarding claim 1, it recites the term “in proximity,” which is a relative term which renders the claim indefinite. The term “in proximity” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. Regarding claims 21, 41, and 42, they are rejected using the same citations and rationales described in the rejection of claim 1. Regarding claim 41, it recites “convert the subsequent frame into a subsequent 3D depth map including creating a current 3D model of the subject using the subsequent depth map.” It is unclear what is being recited as including creating a current 3D model. The claim could be interpreted as meaning that the conversion of the subsequent frame into a 3D depth map includes creating a current 3D model or it could be interpreted as meaning the 3D depth map itself includes creating a current 3D model. In either case it is unclear how either would make sense. Creating a depth map involves estimating or measuring distance data and it is unclear how this process would include creating a current 3D model. A depth map itself stores distance data and it is unclear how a depth map itself would include creating a current 3D model. MPEP states, “claims are required to be cast in clear—as opposed to ambiguous, vague, indefinite—terms.” MPEP 2173.02.II. Additionally, for either possible interpretation, the specification conflicts the claim limitation because the specification instead supports generating a depth map and then creating a corresponding 3D model using the depth map information (specification, para [0028]). There is a conflict between the claimed subject matter and the specification disclosure. This renders the scope of the claim uncertain as inconsistency with the specification disclosure makes the claim take on an unreasonable degree of uncertainty. Reference can be made to MPEP 2173.03 For the purpose of compact prosecution and art rejection, the examiner will treat the limitation to mean: “convert the subsequent frame into a subsequent 3D depth map and create a current 3D model of the subject using the subsequent depth map.” Regarding claim 42, specifically regarding the limitation “converting, by the headset device, the frame into a subsequent 3D depth map including creating a current 3D model of the subject using the subsequent depth map,” it is rejected using the same citations and rationales described in the rejection of claim 41. Regarding claim 21, specifically regarding the limitation “converting, by the headset device, the frame into a subsequent 3D depth map including creating a current 3D model of the subject using the subsequent depth map,” it is rejected using the same citations and rationales described in the rejection of claim 41.” Claim 23 recites the limitation "the pose registration and screen tracking" in line 2 of the claim. There is insufficient antecedent basis for this limitation in the claim. The examiner notes that the applicant may have intended for claim 23 to depend on claim 22. Claim 10 recites the limitation "the deformed initial 3D model" in line 6 of the claim. There is insufficient antecedent basis for this limitation in the claim. Regarding claim 30 and 41, they are rejected using the same citations and rationales described in the rejection of claim 10. Claim 21 recites the limitation "the first frame" in line 14 of the claim. There is insufficient antecedent basis for this limitation in the claim. Regarding claim 42, it is rejected using the same citations and rationales described in the rejection of claim 21. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-2, 4, 8-9, 13-19, 21-22, 24, 28-29, and 33-39 are rejected under 35 U.S.C. 103 as being unpatentable over Valli et al. (US 20210185303 A1; hereinafter Valli) Lee et al. (US 20190213773 A1; hereinafter Lee). Regarding claim 1, Valli teaches a system for real-time conversion of two-dimensional (2D) video into three-dimensional (3D) holographic video content (“Some embodiments may be directed to enhancing a normal monoscopic display image with stereoscopic focal planes to provide that a user perceives a 3D effect and is able to naturally accommodate to the 3D content,” (Valli; page 9, para [0140]; Fig 7). “…some embodiments may provide an enhanced view of 2D content (e.g., as presented on an external 2D display) to a wearer of a head mounted display (HMD) device, based on identifying the 2D content, retrieving metadata to provide depth information associated with the 2D content, and processing the 2D content together with the metadata to generate and display in the HMD a multiple focal plane (MFP) representation of the content …deriving 3D from 2D content may be made as, e.g., by some current S3D TV sets, and this 3D information (with synthesized depth data) may be used to display content,” (page 18, para [0232]). “detection and tracking of the display area, and the corresponding geometry transformations are made continually and in real-time,” (pages 11-12, para [0166]). “Multi-focal plane (MFP) near-eye displays create an approximation for the light-field of the displayed scene,” (page 8, para [0130]). Holographic video content includes multi-focal plane (MFP) displays.), the system comprising: a wearable headset device including: one or more cameras configured to capture a front- facing field of view, and one or more displays configured to present digital content to a user of the headset device (“a user may be wearing an HMD which has a multiple-focal plane display capability. The HMD also may have a front-facing camera capable of capturing imagery of the scene in front of the user. The front facing camera may be a depth camera, for example,” (pages 18-19, para [0234]). “…near-eye display glasses may have a camera embedded that captures the content displayed on the screen,” (page 10, para [0149]). “A schematic diagram 400 illustrates a multi-focal near-eye display… As shown in FIG. 4, a left eye 416 and a right eye 418 are shown with associated eye pieces 412, 414 and display stacks 408, 410. The virtual focal planes 406 are shown with the left eye image 402 and right eye image 404 with overlapping area between the two,” (page 8, para [0130]; page 12, para [0171]; Fig 4; Fig 11). The wearable headset includes the HMD / near-eye display glasses. The one or more displays includes the display stacks.), the headset device configured to: capture, using the one or more cameras, video that includes a client computing device in proximity to the user, the client computing device comprising a screen displaying 2D video content (“an input image 1102 (and image depth data 1104 for some embodiments) is obtained via a network connection 1138 (for example, over a metadata channel). For some embodiments, an input image 1102 (and image depth data 1104 for some embodiments) is obtained 1136 by capturing the contents from an external display 1114 by a tracking camera 1116,” (page 11, para [0162]; Fig 7). “an external display is detected and tracked 1004 from the video being captured 1002 by a tracking camera in the glasses. A display area may be detected in some embodiments by the luminosity of the screen in the video, and knowledge on the size and geometry of the display area, when seen from varying viewpoints,” (page 13, para [0176]; Fig 10). “video data may be received at the glasses via a capture device, such as a camera built into the user device, such that the capture device captures the 2D imagery from its view of the external display, and the 2D imagery is displayed to the user via a display device built into the user device,” (para [0147]). “a 2D video 2110 is sent from an image source 2102 to a television… the television 2104 displays 2112 the 2D video,” (pages 20-21, para [0245]-[0246]; page 24, para [323]; Fig 22). The client computing device includes the external display and / or television.); for a first one or more frames of the captured video: identify a region of interest in the first one or more frames that corresponds to a subject in the 2D video content displayed on the client computing device (“capturing 2402, with a camera coupled to the HMD, a video image of a real-world scene. Some embodiments of the example method 2400 may further include identifying 2404 an image pattern present in the captured video image. Some embodiments of the example method 2400 may further include determining 2406 a depth adjustment associated with the identified image pattern. Some embodiments of the example method 2400 may further include generating 2408 a plurality of focal plane images including depth cues for the identified image pattern, the depth cues reflecting a modified depth of the identified image pattern based on the determined depth adjustment. Some embodiments of the example method 2400 may further include displaying 2410 a 3D representation of the identified image pattern including the plurality of focal plane images.” (page 21, para [0248]-[0249]). “…the screen 802 that displays 2D video content may be part of a real-world scene 800,” (page 9, para [0145]). “The HMD may use the front-facing camera to capture real-world imagery, and this imagery may be analyzed (e.g., using stored (e.g., “known”) object or pattern recognition algorithms) to detect image patterns or objects,” (pages 18-19, para [0234]). “…some embodiments may provide an enhanced view of 2D content (e.g., as presented on an external 2D display) to a wearer of a head mounted display (HMD) device…depth information may be determined for an object or image pattern,” (page 18, para [0232]). A region of interest in the first one or more frames that corresponds to a subject in the 2D video content includes the area corresponding to an image pattern present in the captured video image. A region of interest includes the area corresponding to an image pattern. The 2D video content includes the captured video image.), convert the first one or more frames into a first 3D depth map including creating an 3D representation1 of the subject using the first depth map (“the original depth field may be determined using a depth camera of the HMD, or if the HMD has dual cameras, the original depth field may be determined using a depth from a stereo analysis of the dual captured images,” (page 19, para [0237]). “The modified depth field may be used to generate a multiple focal plane (MFP) representation of the real-world scene, which may be displayed to the user. Any of the various techniques previously described for generating an MFP representation of depth-enhanced content from a 2D external display may be used in a similar fashion to produce the depth-enhanced or depth-modified view of the real-world imagery,” (page 19, para [0238]). “depth information may be calculated (which may be performed locally to a device). For some embodiments, depth information may be generated in whole or in part. For some embodiments, depth information may be determined for an object or image pattern. For some embodiments, deriving 3D from 2D content may be made as, e.g., by some current S3D TV sets, and this 3D information (with synthesized depth data) may be used to display content,” (para [0232]). “the 3D depth information for the 2D content may include a time sequence of depth maps synchronized to the 2D content,” (para [0012]). “identifying 2404 an image pattern present in the captured video image. Some embodiments of the example method 2400 may further include determining 2406 a depth adjustment associated with the identified image pattern. Some embodiments of the example method 2400 may further include generating 2408 a plurality of focal plane images including depth cues for the identified image pattern, the depth cues reflecting a modified depth of the identified image pattern based on the determined depth adjustment. Some embodiments of the example method 2400 may further include displaying 2410 a 3D representation of the identified image pattern including the plurality of focal plane images,” (page 21, para [0248]-[0249]). A 3D depth map includes the depth field and / or depth information. Determining / deriving / generating the depth field and / or depth information includes converting the first one or more frames into a first 3D depth map.), reproject the textured initial 3D model in the displays of the headset device to align with the subject in the 2D video content (“Information on display pose may be used in some embodiments to render HF focal planes (additional 3D information) in right scale and perspective. Display pose may be derived, in some embodiments, by capturing (by the tracking camera in glasses), as a minimum, four (typically corner) points on the screen (x.sub.i, i=1, 2, 3, 4) and solving Eq.1 (homography)…” (pages 14-15, para [0194]-[0195]). “the display area and its geometry are detected from a captured view, and the geometry of the focal planes 1124 are aligned to overlay with the display area 1120,” (pages 11-12, para [0166]). “Three persons 812, 814, 816 are shown wearing MFP glasses in accordance with some embodiments and see the display content in 3D. One spectator 818 sees 810 conventional monoscopic 2D video content on the display 802. Geometry of the focal planes 804, 806, 808 shown to the three persons 812, 814, 816 wearing the MFP glasses are aligned (skewed) along the respective perspective to the display 802,” (Valli; pages 9-10, para [0145]-[0146]; Fig 8). “the depth field is modified to change the depth within the extent of the painting by three inches in the direction of the user. The modified depth field is used together with the real-world imagery to generate a multiple focal plane representation of the scene, which is then displayed to the user via the HMD. From the user's point of view, the paintings the user is interested to see “pop out” of the wall by 3 inches, while paintings by other artists may appear flat against the wall,” (page 19, para [0240]). After combination, Valli’s focal planes become Lee’s textured initial 3D model. Reprojecting the textured initial 3D model in the displays of the headset device includes projective transformation / homography of the focal planes. Align with the subject in the 2D video content includes the geometry of the focal planes 1124 are aligned to overlay with the display area.); for each subsequent frame of the captured video: identify a region of interest in the subsequent frame that corresponds to the subject (“the position and extent (visual footprint) of the detected object or image pattern to be enhanced may be tracked using the HMD camera similar to tracking the 2D external display described above for some embodiments,” (page 19, para [0238]). “The above process may be performed continually (such as in a loop), so that certain objects may be continually updated with depth enhancements as the objects are encountered by the user,” (page 19, para [0239]). “As the user walks around the museum, the front camera of the HMD captures imagery of the museum and an image recognition algorithm is used to identify and classify the paintings in the captured imagery. For each identified painting determined to be by Manet or Renoir, the depth field is modified to change the depth within the extent of the painting by three inches in the direction of the user,” (page 19, para [0240]). Identifying a region of interest in the subsequent frame that corresponds to the subject for each subsequent frame of the captured video includes tracking the detected object or image pattern.), convert the subsequent frame into a subsequent 3D depth map including a current 3D model of the subject ((page 19, para [0237]-[0238]). (page 21, para [0248]-[0249]). “The above process may be performed continually (such as in a loop), so that certain objects may be continually updated with depth enhancements as the objects are encountered by the user,” (page 19, para [0239]). Valli teaches that the conversion steps are repeated (such as in a loop) which would include repeating the steps in a subsequent frame.), reproject the 3D representation of a later time / frame in the displays of the headset device to align with the subject in the 2D video content ((pages 14-15, para [0194]-[0195]). (page 19, para [0240]). “Alignment of HF focal planes with the external display may involve solving in real-time the transformation between world coordinates and the observed coordinates,” (page 14, para [0192]). “…detection and tracking of the display area, and the corresponding geometry transformations are made continually and in real-time,” (pages 11-12, para [0166]). “The above process may be performed continually (such as in a loop), so that certain objects may be continually updated with depth enhancements as the objects are encountered by the user,” (page 19, para [0239]; page 10, para [0147]-[0148]; Fig 10). Valli teaches that the alignment / reprojecting steps are repeated (such as in a loop depicted in at least Figure 10) which would include repeating the steps in a later time / frame.). Valli is not relied upon teaching creating an initial 3D model of the subject using depth information. Valli does not teach that the 3D representation is a 3D model. Valli is not relied upon teaching overlay an initial high-definition texture on the initial 3D model of the subject, the initial texture generated from the first one or more frames. Valli is not relied upon teaching deform the initial 3D model of the subject to match the current 3D model of the subject. Valli is not relied upon teaching overlay a current high-definition texture on the deformed 3D model of the subject, the current texture generated from the subsequent frame. Valli is not relied upon teaching that the 3D representation of a later time / frame is a textured current 3D model. Lee teaches creating an initial 3D model of the subject using depth information (“The image processing module… determines a region of interest (ROI) in the images (e.g., a person's face) and crops the images accordingly at step 360,” (page 6, para [0078]). “generating an initial 3D model for each of the one or more non-rigid objects in the scene using the image comprises determining one or more vertices of each of the one or more non-rigid objects in the scene; determining one or more mesh faces associated with one or more surfaces of each of the one or more non-rigid objects in the scene; determining a surface normal for one or more surfaces of each of the one or more non-rigid objects in the scene; and generating the initial 3D model for each of the one or more non-rigid objects in the scene using the one or more vertices of the object, the one or more mesh faces of the object, and the surface normal for the one or more surfaces of the object,” (page 2, para [0014]). “image processing module 106 generates (1210) a 3D point cloud with normal of the object(s) 102a, 102b in the image and extracts 2D features of the object(s) from the image. In one embodiment, the image processing module 106 generates the 3D point cloud using one or more depth frames with depth camera intrinsic parameters,” (page 9, para [0142]). “The image processing module 106 then generates a 3D model of the person using the captured images—including determining vertices, mesh faces, surface normal, and high-resolution texture in step 310,” (page 4, para [0057]).) Lee teaches that the 3D representation is a 3D model (“The image processing module… determines a region of interest (ROI) in the images (e.g., a person's face) and crops the images accordingly at step 360,” (page 6, para [0078]). “generating an initial 3D model for each of the one or more non-rigid objects in the scene using the image comprises determining one or more vertices of each of the one or more non-rigid objects in the scene; determining one or more mesh faces associated with one or more surfaces of each of the one or more non-rigid objects in the scene; determining a surface normal for one or more surfaces of each of the one or more non-rigid objects in the scene; and generating the initial 3D model for each of the one or more non-rigid objects in the scene using the one or more vertices of the object, the one or more mesh faces of the object, and the surface normal for the one or more surfaces of the object,” (page 2, para [0014]). “image processing module 106 generates (1210) a 3D point cloud with normal of the object(s) 102a, 102b in the image and extracts 2D features of the object(s) from the image. In one embodiment, the image processing module 106 generates the 3D point cloud using one or more depth frames with depth camera intrinsic parameters,” (page 9, para [0142]). “The image processing module 106 then generates a 3D model of the person using the captured images—including determining vertices, mesh faces, surface normal, and high-resolution texture in step 310,” (page 4, para [0057]).) Lee teaches overlay an initial high-definition texture on the initial 3D model of the subject, the initial texture generated from the first one or more frames (“A texture mapped 3D model refers to the virtual object (model) not only has 3D geometric description, including vertices, faces, and surface normal, but also has mapped texture on the surface. The textures are extracted from the images captured during the scan. In the process, the camera poses are recorded, so for each face (represented by a triangle or a polygon) on the model surface the best view are extracted from the recorded images. Then these texture segments are combined and blended to obtain the integrated texture mapping of the object surface,” (page 6, para [0077]). “The 3D model/avatar is presented by static and dynamic meshes, and with the following enhanced features: real-time high definition texture map for face animation…” (page 1, para [0006]). “the system creates an initial model from the first bunch of scanned data, then enhances and completes the avatar model during the following scan(s) and data capture,” (page 4, para [0056]; Fig 3). “For the holographic/3D selfie scan and modeling step 305 of FIG. 3, the sensor 103 captures one or more high-resolution color images of the object 102a (e.g., a person standing or sitting in front of the sensor). The image processing module 106 then generates a 3D model of the person using the captured images—including determining vertices, mesh faces, surface normal, and high-resolution texture in step 310,” (page 4, para [0057]; Fig 3). Overlaying an initial high-definition texture on the initial 3D model includes mapping the texture to the 3D model. The initial texture generated from the first one or more frames includes the textures extracted from the images captured during the scan. The first one or more frames includes the first bunch of scanned data.) Lee teaches deform the initial 3D model of the subject to match the current 3D model of the subject (“the image processing module 106 generates (1216) as output a deformed 3D model and associated deformation information—i.e., 3D transformation for each of the deformation nodes. The final deformed 3D model matches the input from the sensor 103, and the final 3D model can then be used as 3D model input to match the next input from the sensor 103,”(para [0153]). “The image processing module 106 then non-rigidly matches (1214) the 3D model to the 3D point cloud using 3D points, normal, and 2D features. To non-rigidly match the 3D model to the 3D point cloud, the image processing module 106 iteratively minimizes the matching error—which consists of 3D matching error, 2D matching error and smoothness error—by optimizing the 3D transformation of each deformation node,” (para [0144]). “modifying the mapped 3D model using the tracking information comprises deforming the mapped 3D model using the tracking information to match a 3D model of the one or more non-rigid objects in the scene…” (para [0018]). The current 3D model includes the 3D point cloud.), Lee teaches overlay a current high-definition texture on the deformed 3D model of the subject, the current texture generated from the subsequent frame (“The dynamic mesh covers the part where both the geometry and texture need to be updated in real-time. The image processing module 106 only updates the dynamic mesh in real-time,” (page 5, para [0062]). “When the geometry and texture are updated simultaneously, the system realizes much more vivid and photorealistic animation that cannot be achieved with only the geometry update alone,” (page 5, para [0063]). “Further animation and visualization enhancement are achieved by inserting additional control points on top of the facial landmarks, and by transferring and updating the HD texture of ROI (e.g., the eyes and mouth) in real-time,” (page 6, para [0081]; page 8, para [0124]). “High definition texture mapping: the image processing module 106 accounts for the warping of texture between the deformed and non-deformed models: a) At the texture mapping stage, the image processing module 106 keeps the deformation field associated with the HD images. b) Apply a patch based approach. For each patch, the texture is mapped with the deformation filed added to the model. And then the texture coordinates of the vertices remain constant with or without the associated deformation,” (pages 7-8, para [0112]-[0114]). Overlaying a current high-definition texture on the deformed 3D model includes mapping / transferring the texture with the deformation filed added to the model. The current texture generated from the subsequent frame includes updating the texture / dynamic mesh in real-time.), Lee teaches that the 3D representation of a later time / frame is a textured current 3D model (“The viewing device 112 uses the information to deform the 3D model to match the input RBG+Depth scan captured by the sensor 103 and display the deformed 3D model to a user of the viewing device 112 at the second location,” (para [0153]). “The image 1406 is the deformed 3D model overlaid with the input image by the viewing device 112 to generate a 3D avatar,” (para [0154]). “When the geometry and texture are updated simultaneously, the system realizes much more vivid and photorealistic animation that cannot be achieved with only the geometry update alone,” (page 5, para [0062]-[0063]). “an animation of the three-dimensional avatar of a non-rigid object at the viewing device is substantially synchronized with a movement of the corresponding non-rigid object in the scene as captured by the sensor device,” (page 2, para [0020]). “the viewing device periodically receives updated tracking information from the server computing device and the viewing device uses the updated tracking information to further modify the mapped 3D model,” (page 2, para [0018]). The 3D model is updated over time. Texture is updated with geometry.) Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Lee to Valli. The motivation would have been to improve the quality and / or realism of visual illusions. Additional motivation would have been to improve spatial alignment. Additional motivation would have been to improve the user experience. Regarding claim 21, Valli teaches a computerized method of real-time conversion of two-dimensional (2D) video into three- dimensional (3D) holographic video content (“Some embodiments may be directed to enhancing a normal monoscopic display image with stereoscopic focal planes to provide that a user perceives a 3D effect and is able to naturally accommodate to the 3D content,” (Valli; page 9, para [0140]; Fig 7). “…some embodiments may provide an enhanced view of 2D content (e.g., as presented on an external 2D display) to a wearer of a head mounted display (HMD) device, based on identifying the 2D content, retrieving metadata to provide depth information associated with the 2D content, and processing the 2D content together with the metadata to generate and display in the HMD a multiple focal plane (MFP) representation of the content …deriving 3D from 2D content may be made as, e.g., by some current S3D TV sets, and this 3D information (with synthesized depth data) may be used to display content,” (page 18, para [0232]). “detection and tracking of the display area, and the corresponding geometry transformations are made continually and in real-time,” (pages 11-12, para [0166]). “Multi-focal plane (MFP) near-eye displays create an approximation for the light-field of the displayed scene,” (page 8, para [0130]). Holographic video content includes multi-focal plane (MFP) displays.), the method comprising: capturing, using one or more cameras of a wearable headset device (“a user may be wearing an HMD which has a multiple-focal plane display capability. The HMD also may have a front-facing camera capable of capturing imagery of the scene in front of the user. The front facing camera may be a depth camera, for example,” (pages 18-19, para [0234]). “…near-eye display glasses may have a camera embedded that captures the content displayed on the screen,” (page 10, para [0149]). “A schematic diagram 400 illustrates a multi-focal near-eye display… As shown in FIG. 4, a left eye 416 and a right eye 418 are shown with associated eye pieces 412, 414 and display stacks 408, 410. The virtual focal planes 406 are shown with the left eye image 402 and right eye image 404 with overlapping area between the two,” (page 8, para [0130]; page 12, para [0171]; Fig 4; Fig 11). The wearable headset includes the HMD / near-eye display glasses. The one or more displays includes the display stacks.), video that includes a client computing device in proximity to a user of the headset device, the client computing device comprising a screen displaying 2D video content (“an input image 1102 (and image depth data 1104 for some embodiments) is obtained via a network connection 1138 (for example, over a metadata channel). For some embodiments, an input image 1102 (and image depth data 1104 for some embodiments) is obtained 1136 by capturing the contents from an external display 1114 by a tracking camera 1116,” (page 11, para [0162]; Fig 7). “an external display is detected and tracked 1004 from the video being captured 1002 by a tracking camera in the glasses. A display area may be detected in some embodiments by the luminosity of the screen in the video, and knowledge on the size and geometry of the display area, when seen from varying viewpoints,” (page 13, para [0176]; Fig 10). “video data may be received at the glasses via a capture device, such as a camera built into the user device, such that the capture device captures the 2D imagery from its view of the external display, and the 2D imagery is displayed to the user via a display device built into the user device,” (para [0147]). “a 2D video 2110 is sent from an image source 2102 to a television… the television 2104 displays 2112 the 2D video,” (pages 20-21, para [0245]-[0246]; page 24, para [323]; Fig 22). The client computing device includes the external display and / or television.); for a first one or more frames of the captured video: identifying, by the headset device, a region of interest in the first one or more frames that corresponds to a subject in the 2D video content displayed on the client computing device (“capturing 2402, with a camera coupled to the HMD, a video image of a real-world scene. Some embodiments of the example method 2400 may further include identifying 2404 an image pattern present in the captured video image. Some embodiments of the example method 2400 may further include determining 2406 a depth adjustment associated with the identified image pattern. Some embodiments of the example method 2400 may further include generating 2408 a plurality of focal plane images including depth cues for the identified image pattern, the depth cues reflecting a modified depth of the identified image pattern based on the determined depth adjustment. Some embodiments of the example method 2400 may further include displaying 2410 a 3D representation of the identified image pattern including the plurality of focal plane images.” (page 21, para [0248]-[0249]). “…the screen 802 that displays 2D video content may be part of a real-world scene 800,” (page 9, para [0145]). “The HMD may use the front-facing camera to capture real-world imagery, and this imagery may be analyzed (e.g., using stored (e.g., “known”) object or pattern recognition algorithms) to detect image patterns or objects,” (pages 18-19, para [0234]). “…some embodiments may provide an enhanced view of 2D content (e.g., as presented on an external 2D display) to a wearer of a head mounted display (HMD) device…depth information may be determined for an object or image pattern,” (page 18, para [0232]). A region of interest in the first one or more frames that corresponds to a subject in the 2D video content includes the area corresponding to an image pattern present in the captured video image. A region of interest includes the area corresponding to an image pattern. The 2D video content includes the captured video image.), converting, by the headset device, the first one or more frames into a first 3D depth map including creating an 3D representation of the subject using the depth map ((“the original depth field may be determined using a depth camera of the HMD, or if the HMD has dual cameras, the original depth field may be determined using a depth from a stereo analysis of the dual captured images,” (page 19, para [0237]). “The modified depth field may be used to generate a multiple focal plane (MFP) representation of the real-world scene, which may be displayed to the user. Any of the various techniques previously described for generating an MFP representation of depth-enhanced content from a 2D external display may be used in a similar fashion to produce the depth-enhanced or depth-modified view of the real-world imagery,” (page 19, para [0238]). “depth information may be calculated (which may be performed locally to a device). For some embodiments, depth information may be generated in whole or in part. For some embodiments, depth information may be determined for an object or image pattern. For some embodiments, deriving 3D from 2D content may be made as, e.g., by some current S3D TV sets, and this 3D information (with synthesized depth data) may be used to display content,” (para [0232]). “the 3D depth information for the 2D content may include a time sequence of depth maps synchronized to the 2D content,” (para [0012]). “identifying 2404 an image pattern present in the captured video image. Some embodiments of the example method 2400 may further include determining 2406 a depth adjustment associated with the identified image pattern. Some embodiments of the example method 2400 may further include generating 2408 a plurality of focal plane images including depth cues for the identified image pattern, the depth cues reflecting a modified depth of the identified image pattern based on the determined depth adjustment. Some embodiments of the example method 2400 may further include displaying 2410 a 3D representation of the identified image pattern including the plurality of focal plane images,” (page 21, para [0248]-[0249]). A 3D depth map includes the depth field and / or depth information. Determining / deriving / generating the depth field and / or depth information includes converting the first one or more frames into a first 3D depth map.), reprojecting, by the headset device, the textured initial 3D model in one or more displays of the headset device to align with the subject in the 2D video content (“Information on display pose may be used in some embodiments to render HF focal planes (additional 3D information) in right scale and perspective. Display pose may be derived, in some embodiments, by capturing (by the tracking camera in glasses), as a minimum, four (typically corner) points on the screen (x.sub.i, i=1, 2, 3, 4) and solving Eq.1 (homography)…” (pages 14-15, para [0194]-[0195]). “the display area and its geometry are detected from a captured view, and the geometry of the focal planes 1124 are aligned to overlay with the display area 1120,” (pages 11-12, para [0166]). “Three persons 812, 814, 816 are shown wearing MFP glasses in accordance with some embodiments and see the display content in 3D. One spectator 818 sees 810 conventional monoscopic 2D video content on the display 802. Geometry of the focal planes 804, 806, 808 shown to the three persons 812, 814, 816 wearing the MFP glasses are aligned (skewed) along the respective perspective to the display 802,” (Valli; pages 9-10, para [0145]-[0146]; Fig 8). “the depth field is modified to change the depth within the extent of the painting by three inches in the direction of the user. The modified depth field is used together with the real-world imagery to generate a multiple focal plane representation of the scene, which is then displayed to the user via the HMD. From the user's point of view, the paintings the user is interested to see “pop out” of the wall by 3 inches, while paintings by other artists may appear flat against the wall,” (page 19, para [0240]). Reprojecting the textured initial 3D model in the displays of the headset device includes projective transformation / homography of the focal planes. Align with the subject in the 2D video content includes the geometry of the focal planes 1124 are aligned to overlay with the display area.); for each subsequent frame of the captured video: identifying, by the headset device, a region of interest in the frame that corresponds to the subject (“the position and extent (visual footprint) of the detected object or image pattern to be enhanced may be tracked using the HMD camera similar to tracking the 2D external display described above for some embodiments,” (page 19, para [0238]). “The above process may be performed continually (such as in a loop), so that certain objects may be continually updated with depth enhancements as the objects are encountered by the user,” (page 19, para [0239]). “As the user walks around the museum, the front camera of the HMD captures imagery of the museum and an image recognition algorithm is used to identify and classify the paintings in the captured imagery. For each identified painting determined to be by Manet or Renoir, the depth field is modified to change the depth within the extent of the painting by three inches in the direction of the user,” (page 19, para [0240]). Identifying a region of interest in the subsequent frame that corresponds to the subject for each subsequent frame of the captured video includes tracking the detected object or image pattern.), converting, by the headset device, the frame into a subsequent 3D depth map (page 19, para [0237]-[0238]). (page 21, para [0248]-[0249]). “The above process may be performed continually (such as in a loop), so that certain objects may be continually updated with depth enhancements as the objects are encountered by the user,” (page 19, para [0239]). Valli teaches that the conversion steps are repeated (such as in a loop) which would include repeating the steps in a subsequent frame.), reprojecting, by the headset device, the 3D representation of a later time / frame in the displays of the headset device to align with the subject in the 2D video content ((pages 14-15, para [0194]-[0195]). (page 19, para [0240]). “Alignment of HF focal planes with the external display may involve solving in real-time the transformation between world coordinates and the observed coordinates,” (page 14, para [0192]). “…detection and tracking of the display area, and the corresponding geometry transformations are made continually and in real-time,” (pages 11-12, para [0166]). “The above process may be performed continually (such as in a loop), so that certain objects may be continually updated with depth enhancements as the objects are encountered by the user,” (page 19, para [0239]; page 10, para [0147]-[0148]; Fig 10). Valli teaches that the alignment / reprojecting steps are repeated (such as in a loop depicted in at least Figure 10) which would include repeating the steps in a later time / frame.). Valli is not relied upon teaching creating an initial 3D model of the subject using depth information. Valli is not relied upon teaching that the 3D representation is a 3D model. Valli is not relied upon teaching overlaying, by the headset device, an initial high-definition texture on the initial 3D model of the subject, the initial texture generated from the first frame. Valli is not relied upon teaching for a subsequent frame including creating a current 3D model of the subject using the subsequent depth map. Valli is not relied upon teaching deforming, by the headset device, the initial 3D model of the subject to match the current 3D model of the subject. Valli is not relied upon teaching overlaying, by the headset device, a current high-definition texture on the deformed 3D model of the subject, the current texture generated from the frame. Valli is not relied upon teaching that the 3D representation of a later time / frame is a textured current 3D model. Lee teaches creating an initial 3D model of the subject using depth information (“The image processing module… determines a region of interest (ROI) in the images (e.g., a person's face) and crops the images accordingly at step 360,” (page 6, para [0078]). “generating an initial 3D model for each of the one or more non-rigid objects in the scene using the image comprises determining one or more vertices of each of the one or more non-rigid objects in the scene; determining one or more mesh faces associated with one or more surfaces of each of the one or more non-rigid objects in the scene; determining a surface normal for one or more surfaces of each of the one or more non-rigid objects in the scene; and generating the initial 3D model for each of the one or more non-rigid objects in the scene using the one or more vertices of the object, the one or more mesh faces of the object, and the surface normal for the one or more surfaces of the object,” (page 2, para [0014]). “image processing module 106 generates (1210) a 3D point cloud with normal of the object(s) 102a, 102b in the image and extracts 2D features of the object(s) from the image. In one embodiment, the image processing module 106 generates the 3D point cloud using one or more depth frames with depth camera intrinsic parameters,” (page 9, para [0142]). “The image processing module 106 then generates a 3D model of the person using the captured images—including determining vertices, mesh faces, surface normal, and high-resolution texture in step 310,” (page 4, para [0057]).) Lee teaches that the 3D representation is a 3D model (“The image processing module… determines a region of interest (ROI) in the images (e.g., a person's face) and crops the images accordingly at step 360,” (page 6, para [0078]). “generating an initial 3D model for each of the one or more non-rigid objects in the scene using the image comprises determining one or more vertices of each of the one or more non-rigid objects in the scene; determining one or more mesh faces associated with one or more surfaces of each of the one or more non-rigid objects in the scene; determining a surface normal for one or more surfaces of each of the one or more non-rigid objects in the scene; and generating the initial 3D model for each of the one or more non-rigid objects in the scene using the one or more vertices of the object, the one or more mesh faces of the object, and the surface normal for the one or more surfaces of the object,” (page 2, para [0014]). “image processing module 106 generates (1210) a 3D point cloud with normal of the object(s) 102a, 102b in the image and extracts 2D features of the object(s) from the image. In one embodiment, the image processing module 106 generates the 3D point cloud using one or more depth frames with depth camera intrinsic parameters,” (page 9, para [0142]). “The image processing module 106 then generates a 3D model of the person using the captured images—including determining vertices, mesh faces, surface normal, and high-resolution texture in step 310,” (page 4, para [0057]).) Lee teaches overlaying, by the headset device, an initial high-definition texture on the initial 3D model of the subject, the initial texture generated from the first frame. (Lee; “A texture mapped 3D model refers to the virtual object (model) not only has 3D geometric description, including vertices, faces, and surface normal, but also has mapped texture on the surface. The textures are extracted from the images captured during the scan. In the process, the camera poses are recorded, so for each face (represented by a triangle or a polygon) on the model surface the best view are extracted from the recorded images. Then these texture segments are combined and blended to obtain the integrated texture mapping of the object surface,” (Lee; page 6, para [0077]). “The 3D model/avatar is presented by static and dynamic meshes, and with the following enhanced features: real-time high definition texture map for face animation…” (Lee; page 1, para [0006]). “the system creates an initial model from the first bunch of scanned data, then enhances and completes the avatar model during the following scan(s) and data capture,” (Lee; page 4, para [0056]; Fig 3). “For the holographic/3D selfie scan and modeling step 305 of FIG. 3, the sensor 103 captures one or more high-resolution color images of the object 102a (e.g., a person standing or sitting in front of the sensor). The image processing module 106 then generates a 3D model of the person using the captured images—including determining vertices, mesh faces, surface normal, and high-resolution texture in step 310,” (Lee; page 4, para [0057]; Fig 3). “The image processing module 106 at the first location receives (1202) an RGB+Depth image of the object 102a (and/or object 102b) in the scene 101 from the sensor 103. The image processing module 106 also generates (1204) a 3D model and corresponding texture(s) of the objects 102a, 102b in the scene 101 using, e.g., the real-time object recognition and modeling techniques,” (Lee; pages 8-9, para [0134]; Fig 12). Overlaying an initial high-definition texture on the initial 3D model includes mapping the texture to the 3D model. The initial texture generated from the first one or more frames includes the textures extracted from the images captured during the scan. The first one or more frames includes the first bunch of scanned data. In at least the case that the sensor captures one image, the texture is generated from the first frame.). Lee teaches for a subsequent frame including creating a current 3D model of the subject using the subsequent depth map (Lee; “For each frame of the input data from the sensors, the image processing module 106 applies an iterative process to register the model and extract the deformation field, thus the input data removed of the deformation filed is combined with the “rigid” canonical model data for reconstruction and refinement,” (page 5, para [0068]). “The image processing module 106 at the first location receives (1202) an RGB+Depth image of the object 102a (and/or object 102b) in the scene 101 from the sensor 103,” (pages 8-9, para [0134]). “where DFt is the deformation field at time t, Dt is the depth map from the sensors, St is the canonical rigid model projected to the current sensor coordinate system, Dtree is the vertex-landmark connectivity defined in the deformation tree,” (page 7, para [103]-[0104]). “image processing module 106 generates (1210) a 3D point cloud with normal of the object(s) 102a, 102b in the image and extracts 2D features of the object(s) from the image. In one embodiment, the image processing module 106 generates the 3D point cloud using one or more depth frames with depth camera intrinsic parameters,” (page 9, para [0142]). “The image processing module 106 then rigidly matches (1212) the 3D model to the generated 3D point cloud,” (page 9, para [0143]). The current 3D model includes the 3D point cloud. The depth map includes the RGB+Depth image.). Lee teaches deforming, by the headset device, the initial 3D model of the subject to match the current 3D model of the subject (“the image processing module 106 generates (1216) as output a deformed 3D model and associated deformation information—i.e., 3D transformation for each of the deformation nodes. The final deformed 3D model matches the input from the sensor 103, and the final 3D model can then be used as 3D model input to match the next input from the sensor 103,”(para [0153]). “The image processing module 106 then non-rigidly matches (1214) the 3D model to the 3D point cloud using 3D points, normal, and 2D features. To non-rigidly match the 3D model to the 3D point cloud, the image processing module 106 iteratively minimizes the matching error—which consists of 3D matching error, 2D matching error and smoothness error—by optimizing the 3D transformation of each deformation node,” (para [0144]). “modifying the mapped 3D model using the tracking information comprises deforming the mapped 3D model using the tracking information to match a 3D model of the one or more non-rigid objects in the scene…” (para [0018]). The current 3D model includes the 3D point cloud.). Lee teaches overlaying, by the headset device, a current high-definition texture on the deformed 3D model of the subject, the current texture generated from the frame ((“The dynamic mesh covers the part where both the geometry and texture need to be updated in real-time. The image processing module 106 only updates the dynamic mesh in real-time,” (page 5, para [0062]). “When the geometry and texture are updated simultaneously, the system realizes much more vivid and photorealistic animation that cannot be achieved with only the geometry update alone,” (page 5, para [0063]). “Further animation and visualization enhancement are achieved by inserting additional control points on top of the facial landmarks, and by transferring and updating the HD texture of ROI (e.g., the eyes and mouth) in real-time,” (page 6, para [0081]; page 8, para [0124]). “High definition texture mapping: the image processing module 106 accounts for the warping of texture between the deformed and non-deformed models: a) At the texture mapping stage, the image processing module 106 keeps the deformation field associated with the HD images. b) Apply a patch based approach. For each patch, the texture is mapped with the deformation filed added to the model. And then the texture coordinates of the vertices remain constant with or without the associated deformation,” (pages 7-8, para [0112]-[0114]). Overlaying a current high-definition texture on the deformed 3D model includes mapping / transferring the texture with the deformation filed added to the model. The current texture generated from the subsequent frame includes updating the texture / dynamic mesh in real-time.). Lee teaches that the 3D representation of a later time / frame is a textured current 3D model (“The viewing device 112 uses the information to deform the 3D model to match the input RBG+Depth scan captured by the sensor 103 and display the deformed 3D model to a user of the viewing device 112 at the second location,” (para [0153]). “The image 1406 is the deformed 3D model overlaid with the input image by the viewing device 112 to generate a 3D avatar,” (para [0154]). “When the geometry and texture are updated simultaneously, the system realizes much more vivid and photorealistic animation that cannot be achieved with only the geometry update alone,” (page 5, para [0062]-[0063]). “an animation of the three-dimensional avatar of a non-rigid object at the viewing device is substantially synchronized with a movement of the corresponding non-rigid object in the scene as captured by the sensor device,” (page 2, para [0020]). “the viewing device periodically receives updated tracking information from the server computing device and the viewing device uses the updated tracking information to further modify the mapped 3D model,” (page 2, para [0018]). The 3D model is updated over time. Texture is updated with geometry.) Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Lee to Valli. The motivation would have been to improve the quality and / or realism of visual illusions. Additional motivation would have been to improve spatial alignment. Additional motivation would have been to improve the user experience. Regarding claim 2, Valli in view of Lee teaches the system of claim 1, wherein the headset device registers a pose of the screen of the client computing device (Valli; “…near-eye display glasses may have a camera embedded that captures the content displayed on the screen. The pose, meaning position, orientation and size, of the screen may be, with respect to the screen, in some embodiments, calculated based on the captured content,” (page 10, para [0149]). “Information on display pose may be used in some embodiments to render HF focal planes (additional 3D information) in right scale and perspective. Display pose may be derived, in some embodiments, by capturing (by the tracking camera in glasses), as a minimum, four (typically corner) points on the screen (x.sub.i, i=1, 2, 3, 4) and solving Eq.1 (homography), by using, for example, an iterative method,” (Valli; pages 14-15, para [0193]-[0195]). “detecting, using a camera coupled to the HMD, presence, spatial position, and orientation information relating to 2D video content displayed on a 2D display external to the HMD,” (Valli; page 1, para [0018]). “detecting presence, spatial position, and orientation information relating to 2D video content may include detecting presence, spatial position, and orientation information relating to the 2D display,” (Valli; page 2, para [0024]) Registering a pose includes deriving / calculating a pose and / or orientation and position.) and tracks a location of the screen through each frame of the captured video (Valli; “an external display is detected and tracked 1004 from the video being captured 1002 by a tracking camera in the glasses. A display area may be detected in some embodiments by the luminosity of the screen in the video, and knowledge on the size and geometry of the display area, when seen from varying viewpoints,” (page 13, para [0176]-[0177]). “the MFP headset 2106 may locate and track 2118 the TV 2104 relative to the MFP headset 2106,” (Valli; page 20, para [0245]). “detection and tracking of the display area, and the corresponding geometry transformations are made continually and in real-time,” (pages 11-12, para [0166]). “detecting and tracking 1004 a pose of the video data from a user's viewing angle,” (Valli; page 10, para [0147]-[0148]; Fig 10). Valli teaches that the detecting and tracking steps are repeated (such as in a loop depicted in at least Figure 10) which would include tracking a location through each frame.). Regarding claim 22, it is rejected using the same citations and rationales described in the rejection of claim 2. Regarding claim 4, Valli in view of Lee teaches the system of claim 1, wherein identifying a region of interest in the first one or more frames comprises: capturing input from the user (Valli; “an HMD user is reading a physical book, and the HMD provides a function to identify and depth-enhance words or text phrases entered by the user. The user is looking for a passage in which a character named Harold finds a treasure map. The user enters search terms “Harold” and “treasure map” into the user interface of the HMD, and the user proceeds to turn the pages of the book in the range of pages where he believes the passage to be,” (page 20, para [0242]). “The museum provides an application which is capable of recognizing and classifying the museum's paintings and which interfaces to the users HMD in order to provide depth enhancement functionality based on the users preferences. The user specifies an interest in impressionist paintings by Edouard Manet and Pierre Auguste-Renoir,” (Valli; page 19, para [0240]). “The rules for which objects may be depth enhanced and how the depth of such objects may be adjusted, e.g., may be set by user preferences or may be part of the program logic (e.g., an application running on the HMD or on a connected computer may provide the rules or may use the rules as part of program execution),” (Valli; page 19, para [0239]). The user input includes entering search terms and / or user preferences.); and identifying the region of interest in the first one or more frames based upon the user input (Valli; “The HMD captures the imagery of the book pages using the HMD camera, and analyzes the imagery (e.g., using a text recognition algorithm) to identify instances of text “Harold” and “treasure map”. If either of these two search terms are identified in the imagery, the depth field corresponding to the area of these identified terms is modified so that the words “pop out” of the book pages slightly,” (page 20, para [0242]). “As the user walks around the museum, the front camera of the HMD captures imagery of the museum and an image recognition algorithm is used to identify and classify the paintings in the captured imagery. For each identified painting determined to be by Manet or Renoir, the depth field is modified to change the depth within the extent of the painting by three inches in the direction of the user,” (Valli; page 19, para [0240]). “the rules may include a list of objects or image patterns to be enhanced, information for identifying the objects or image patterns using an object recognition or pattern recognition algorithm, and a specification of how the depth of each object may be enhanced (e.g., a depth offset to add to the object's depth, or some other function for modifying or adjusting the depth of the object),” (Valli; page 19, para [0239]). “the position and extent (visual footprint) of the detected object or image pattern to be enhanced may be tracked using the HMD camera similar to tracking the 2D external display described above for some embodiments,” (Valli; page 19, para [0238]). Identifying the region of interest includes the position and extent (visual footprint) of the detected / identified object or image pattern to be enhanced and / or the area corresponding to the detected / identified object or image pattern to be enhanced. Valli teaches identifying the area corresponding to the detected / identified object or image pattern to be enhanced is based upon the user input.). Regarding claim 24, it is rejected using the same citations and rationales described in the rejection of claim 4. Regarding claim 13, Valli in view of Lee teaches the system of claim 1, wherein, for each frame of the captured video, the headset device segments the frame based upon the identified region of interest (Valli; “In some embodiments, for a user to see the view outside the external display area undistorted (unfiltered), the filtering needs to be used only within the detected display area, which could be accomplished by an active optical element. In some embodiments, a passive diffuser is applicable, if filtering of the whole view is acceptable. Note that in the latter case, after overlaying the high frequency contents, only the area inside the display is seen sharp,” (page 12, para [0169]). “adjusting, based on the metadata, a portion of the original depth field corresponding to the identified content to produce an adjusted depth field, the identified content corresponding to an object or image pattern recognized in the image,” (Valli; page 10, para [0156]). Adjusting a portion of the original depth field corresponding to the identified content includes segmenting the frame based upon the identified region of interest.). Regarding claim 33, it is rejected using the same citations and rationales described in the rejection of claim 13. Regarding claim 16, Valli in view of Lee teaches the system of claim 1, wherein reprojecting the textured 3D model in the displays of the headset device to align with the subject in the 2D video content provides an appearance to the user that the textured 3D model is coming out of the screen of the client computing device (Valli; “If an object or image pattern is detected for which depth enhancement may be available, the additional depth or modified depth for the object or image pattern may be determined. For example, the HMD may have rules for adding a depth offset to certain objects to make the objects appear closer to the user. Such objects may be made to “pop out” in the user's view. As another example, the HMD may have rules for adding different depth offsets to detected objects to make them appear further away from the user,” (page 19, para [0235]). “For each identified painting determined to be by Manet or Renoir, the depth field is modified to change the depth within the extent of the painting by three inches in the direction of the user. The modified depth field is used together with the real-world imagery to generate a multiple focal plane representation of the scene, which is then displayed to the user via the HMD. From the user's point of view, the paintings the user is interested to see “pop out” of the wall by 3 inches, while paintings by other artists may appear flat against the wall,” (Valli; page 19, para [0240]). “the depth field corresponding to the area of these identified terms is modified so that the words “pop out” of the book pages slightly,” (Valli; page 20, para [0242]). “Three persons 812, 814, 816 are shown wearing MFP glasses in accordance with some embodiments and see the display content in 3D. One spectator 818 sees 810 conventional monoscopic 2D video content on the display 802. Geometry of the focal planes 804, 806, 808 shown to the three persons 812, 814, 816 wearing the MFP glasses are aligned (skewed) along the respective perspective to the display 802,” (Valli; pages 9-10, para [0145]-[0146]; Fig 8). “S3D representations are typically positioned both behind and in front of the display… Note that in the figures, focal planes have been illustrated e.g., in front of the external display for simplicity. However, in some embodiments, especially if using existing stereoscopic or DIBR content, the focal planes may cover the depth range (depth map values) of both in front and behind the external display,” (Valli; page 15, para [0201]-[0202]). “Some embodiments may be directed to enhancing a normal monoscopic display image with stereoscopic focal planes to provide that a user perceives a 3D effect and is able to naturally accommodate to the 3D content,” (Valli; page 9, para [0140]). Lee teaches a textured 3D model (Lee; page 6, para [0077]). Providing an appearance to the user that the textured 3D model is coming out of the screen of the client computing device includes the user perceiving a 3D effect. The screen includes the external display. After combination, Valli’s focal planes become Lee’s textured 3D model.). Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Lee to Valli. The motivation would have been to improve the quality and / or realism of visual illusions. Additional motivation would have been to improve spatial alignment. Additional motivation would have been to improve the user experience. Regarding claim 36, it is rejected using the same citations and rationales described in the rejection of claim 16. Regarding claim 17, Valli in view of Lee teaches the system of claim 16, wherein the user views the reprojected textured 3D model in context with the 2D video content (Valli; “Three persons 812, 814, 816 are shown wearing MFP glasses in accordance with some embodiments and see the display content in 3D. One spectator 818 sees 810 conventional monoscopic 2D video content on the display 802. Geometry of the focal planes 804, 806, 808 shown to the three persons 812, 814, 816 wearing the MFP glasses are aligned (skewed) along the respective perspective to the display 802,” (Valli; pages 9-10, para [0145]-[0146]; Fig 8). “As shown, FIG. 9 illustrates two eyes 902, 904 and separate focal planes 906, 908, 910, 912, 914, 916 adding 3D cues over a monoscopic 2D image (video) on an external display 918 (for example a TV set),” (Valli; pages 9-10, para [0146]; Fig 9). “enhancing of the content is enabled by using wearable glasses (near-eye display), which detect and track the content on the screen being viewed, and overlay the content with focal planes producing depth and disparity. Because the same content is shown simultaneously on the external display in 2D, the content may be viewed with or without glasses,” (Valli; page 9, para [0142]). “displaying the plurality of focal plane images as a see-through overlay synchronized with the 2D content,” (Valli; page 1, para [0003]). Lee teaches a textured 3D model (Lee; page 6, para [0077]). The user views the reprojected textured 3D model in context with the 2D video content includes overlaying the 3D content with focal planes over 2D content shown simultaneously on the external display in 2D. After combination with Lee, Valli’s focal planes and / or 3D cues become the Lee’s textured 3D model.). Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Lee to Valli. The motivation would have been to improve the quality and / or realism of visual illusions. Additional motivation would have been to improve spatial alignment. Additional motivation would have been to improve the user experience. Regarding claim 37, it is rejected using the same citations and rationales described in the rejection of claim 17. Regarding claim 18, Valli in view of Lee teaches the system of claim 16, wherein the user views real-world surroundings concurrently with the textured 3D model and the 2D video content (Valli; “Three persons 812, 814, 816 are shown wearing MFP glasses in accordance with some embodiments and see the display content in 3D. One spectator 818 sees 810 conventional monoscopic 2D video content on the display 802. Geometry of the focal planes 804, 806, 808 shown to the three persons 812, 814, 816 wearing the MFP glasses are aligned (skewed) along the respective perspective to the display 802. For some embodiments, the screen 802 that displays 2D video content may be part of a real-world scene 800.” (Valli; pages 9-10, para [0145]-[0146]; Fig 8). “In some embodiments, for a user to see the view outside the external display area undistorted (unfiltered), the filtering needs to be used only within the detected display area, which could be accomplished by an active optical element,” (Valli; page 12, para [0169]). “displaying the see-through overlay may enable a user to view the screen via a direct optical path,” (Valli; page 1, para [0015]). Lee teaches a textured 3D model (Lee; page 6, para [0077]). The user views real-world surroundings concurrently with the textured 3D model and the 2D video content includes Valli’s teaching that a user is able to view outside the external display area. After combination with Lee, Valli’s focal planes and / or 3D cues become the textured 3D model.). Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Lee to Valli. The motivation would have been to improve the quality and / or realism of visual illusions. Additional motivation would have been to improve spatial alignment. Additional motivation would have been to improve the user experience. Regarding claim 38, it is rejected using the same citations and rationales described in the rejection of claim 18. Regarding claim 8, Valli in view of Lee teaches the system of claim 1, wherein, for each subsequent frame of the captured video (Lee; “For each frame of the input data from the sensors, the image processing module 106 applies an iterative process to register the model and extract the deformation field, thus the input data removed of the deformation filed is combined with the “rigid” canonical model data for reconstruction and refinement,” (page 5, para [0068]).), the headset device (“Exemplary computing devices include, but are not limited to, …, augmented reality (AR)/virtual reality (VR) devices (e.g., glasses, headset apparatuses, and so forth), or the like,” (Lee; pages 3-4, para [0052]-[0054]).) compares the current 3D model to the initial 3D model to enable tracking of the movements of both the underlying mesh structure of the 3D model (Lee; “The image processing module 106 then non-rigidly matches (1214) the 3D model to the 3D point cloud using 3D points, normal, and 2D features. To non-rigidly match the 3D model to the 3D point cloud, the image processing module 106 iteratively minimizes the matching error—which consists of 3D matching error, 2D matching error and smoothness error—by optimizing the 3D transformation of each deformation node,” (pages 9-10, para [0142]-[0147]; Lee; page 7, para [0104]). “…the image processing module 106 generates (1216) as output a deformed 3D model and associated deformation information—i.e., 3D transformation for each of the deformation nodes. The final deformed 3D model matches the input from the sensor 103, and the final 3D model can then be used as 3D model input to match the next input from the sensor 103,” (Lee; page 10, para [0153]). Comparing the current 3D model to the initial 3D model includes matching the 3D model to the 3D point cloud. The current 3D model includes the 3D point cloud. The initial 3D model includes the 3D model. Tracking includes the comparison across frames.) and the texture (Lee; “The image processing module 106 builds (1206) a deformation graph of the 3D model, and extracts 2D features from the texture(s) associated with the image,” (page 9, para [0135]). “E2D represents 2D matching error: PNG media_image1.png 28 190 media_image1.png Greyscale where fanchor is the 3D position of a 2D feature from 3D model; floose is the 3D position of fanchor's matched 3D feature in 3D point cloud,” (Lee; page 10, para [0148]-[0149]). “the image processing module 106 accounts for the warping of texture between the deformed and non-deformed models: [0113] a) At the texture mapping stage, the image processing module 106 keeps the deformation field associated with the HD images. [0114] b) Apply a patch based approach. For each patch, the texture is mapped with the deformation filed added to the model…,” (Lee; pages 7-8, para [0112]-[0114]). “The static mesh covers the model where it is not necessary for real-time texture change. The dynamic mesh covers the part where both the geometry and texture need to be updated in real-time. The image processing module 106 only updates the dynamic mesh in real-time…,” (Lee; page 5, para [0062]-[0063]). Tracking of the movements of the texture includes matching and / or warping the texture.). Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Lee to Valli. The motivation would have been to “[realize] much more vivid and photorealistic animation,” (Lee; page 5, para [0063]). Additional motivation would have been to improve the quality and / or realism of visual illusions. Additional motivation would have been to improve spatial alignment. Additional motivation would have been to improve the user experience. Regarding claim 28, it is rejected using the same citations and rationales described in the rejection of claim 8. Regarding claim 9, Valli in view of Lee teaches the system of claim 8, wherein the headset device compares the current 3D model to the initial 3D model using landmarks or an optical flow algorithm (Lee; “Use landmarks to track the object motion and segment the surface. The landmarks can be identified by extinguish visual features (e.g., ORB, MESR, and SIFT), or 3D features (e.g., Surface Curvature and NURF), or combination of both,” (pages 6-7, para [0089]). “The landmark based object tracking and deformation map helps to achieve better object tracking, better deformation field initialization (close to the optimized solution), and it reduces the number of variables and leads to a more efficient solution of the non-linear optimization problem,” (Lee; page 7, para [0110]). “where DFt is the deformation field at time t, Dt is the depth map from the sensors, St is the canonical rigid model projected to the current sensor coordinate system, Dtree is the vertex-landmark connectivity defined in the deformation tree,” (Lee; page 7, para [103]-[0104]). “deformable object scanning has the following additional processing for each frame of incoming data: the facial landmark extraction, facial landmark (and virtual ones) deformation estimation, and the dense deformation estimation (which is a linear process by applying the deformation map),”(Lee; page 6, para [0084]). “Initial scan and modeling: after the initial bunch of frames, the image processing module 106 obtains the initial canonical model, and performs the following functions: [0096] a) Detect and Identify Real Landmarks—these real landmarks are either identified with strong visual features such as ORB, MSER, strong 3D surface features, or combination of these features,” (Lee; page 7, para [0095]-[0096]).). Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Lee to Valli. The motivation would have been “to achieve better object tracking, better deformation field initialization (close to the optimized solution), and [reduce] the number of variables and leads to a more efficient solution of the non-linear optimization problem,” (Lee; page 7, para [0110]).” Additional motivation would have been to improve the quality and / or realism of visual illusions. Additional motivation would have been to improve spatial alignment. Additional motivation would have been to improve the user experience. Regarding claim 29, it is rejected using the same citations and rationales described in the rejection of claim 9. Regarding claim 14, Valli in view of Lee teaches the system of claim 14, wherein the headset device uses a facial recognition algorithm or a body recognition algorithm to perform the segmentation (Lee; “Next, the process continues to the real-time tracking stage 350—which comprises steps 355, 360, 365, 370, and 375. The image processing module 106 of the server computing device 104 receives additional images of the object 102a from the sensor 103 at step 355, and determines a region of interest (ROI) in the images (e.g., a person's face) and crops the images accordingly at step 360,” (page 6, para [0078]). “the system limits the size of the texture need to be updated to only the part where most texture change happens—the area that covers the landmark mesh shown in FIG. 5B. FIG. 10A depicts a full size high resolution image captured by the sensor 103, while FIG. 10B depicts a region of interest (ROI) (e.g., the person's face) that is cropped by the image processing module 106. The image processing module 106 extracts and tracks both the model pose and the relative deformation of the control points (i.e., the facial landmark points) at steps 365 and 370.” (Lee; page 8, para [0118]; Fig 5B; Fig 10A; Fig 10B) “the image processing module 106 also performs a facial landmark detection and extraction step 315 using the 3D model. In computer vision and graphics, facial landmarks are the visual features on human face, such as corners, boundaries, and centers of the mouth, nose, eyes and brows. The extraction of facial landmarks from images can be used for face recognition, and facial expression identification. In the hologram technology described herein, the system uses facial landmarks for facial expression tracking and animation. In one embodiment, the image processing module 106 can utilize the facial landmark model and corresponding open source code from the OpenFace technology (available at https://cmusatyalab.github.io/openface/) for facial landmark detection,” (Lee; page 4, para [0058]). Segmentation includes cropping.). Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Lee to Valli. The motivation would have been to improve the isolation and / or tracking of human faces in real time. Additional motivation would have been to reduce the processing load and / or increase efficiency. Additional motivation would have been to improve the quality and / or realism of visual illusions. Additional motivation would have been to improve the user experience. Regarding claim 34, it is rejected using the same citations and rationales described in the rejection of claim 14. Regarding claim 15, Valli in view of Lee teaches the system of claim 1, wherein the subject in the 2D video content comprises a person and the region of interest comprises one or more of: the person's body, the person's head and shoulders, or the person's face (Valli; “According to some embodiments, users may watch a monoscopic video (such as a program on TV) normally, with bare eyes, and other users may effectively see the same display and content in 3D,” (page 11, para [0158]). Lee; “Next, the process continues to the real-time tracking stage 350—which comprises steps 355, 360, 365, 370, and 375. The image processing module 106 of the server computing device 104 receives additional images of the object 102a from the sensor 103 at step 355, and determines a region of interest (ROI) in the images (e.g., a person's face) and crops the images accordingly at step 360,” (Lee; page 6, para [0078]). “the system limits the size of the texture need to be updated to only the part where most texture change happens—the area that covers the landmark mesh shown in FIG. 5B. FIG. 10A depicts a full size high resolution image captured by the sensor 103, while FIG. 10B depicts a region of interest (ROI) (e.g., the person's face) that is cropped by the image processing module 106. The image processing module 106 extracts and tracks both the model pose and the relative deformation of the control points (i.e., the facial landmark points) at steps 365 and 370.” (Lee; page 8, para [0118]; Fig 5B; Fig 10A; Fig 10B). “In the second phase 350, the person who modeled the avatar is facing the sensor 103… For each frame of the tracking process, the image processing module 106 transfers the 3D displacement of the measured control points together with the head pose… Even at 30 frames-per-second (FPS), the system needs only 24 KB/sec bandwidth to achieve real-time avatar animation control,” (Lee; page 6, para [0080]). After combination, Valli’s 2D video content (such as a program on TV) includes a person’s face as taught by Lee.). Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Lee to Valli. The motivation would have been to increase the variety of visual illusions. Additional motivation would have been to improve the quality and / or realism of visual illusions. Additional motivation would have been to improve the user experience. Regarding claim 35, it is rejected using the same citations and rationales described in the rejection of claim 15. Regarding claim 19, Valli in view of Lee teaches the system of claim 1, wherein the headset device continually refines the textured 3D model for display to the user as each subsequent frame is processed (Lee; “the system creates an initial model from the first bunch of scanned data, then enhances and completes the avatar model during the following scan(s) and data capture. After the initial model is generated and transferred to the client side, the remote avatar visualization and animation is further improved, and then the completed avatar model is transferred to the client side to replace the initial model,” (page 4, para [0056]). “For each frame of the input data from the sensors, the image processing module 106 applies an iterative process to register the model and extract the deformation field, thus the input data removed of the deformation filed is combined with the “rigid” canonical model data for reconstruction and refinement,” (Lee; page 5, para [0068]).). Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Lee to Valli. The motivation would have been to reduce visual errors and / or surface holes. Additional motivation would have been to improve the completeness of the model. Additional motivation would have been to improve the quality and / or realism of visual illusions. Additional motivation would have been to improve the user experience. Regarding claim 39, it is rejected using the same citations and rationales described in the rejection of claim 19. Claims 3, 11-12, 23, and 31-32 are rejected under 35 U.S.C. 103 as being unpatentable over Valli in view of Lee in further view of Newcombe et al., "DynamicFusion: Reconstruction and Tracking of Non-rigid Scenes in Real-Time," 2015, 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 343-352; hereinafter Newcombe. Regarding claim 3, Valli in view of Lee teaches the system of claim 2, wherein the headset device uses a homography and /or marker tracking algorithm to perform the pose registration and screen tracking (Valli; “…near-eye display glasses may have a camera embedded that captures the content displayed on the screen. The pose, meaning position, orientation and size, of the screen may be, with respect to the screen, in some embodiments, calculated based on the captured content,” (page 10, para [0149]). “According to some embodiments, tracking the display uses, in effect, similar techniques as detection and tracking of visible markers (fiducials) in augmented reality. Marker tracking is a traditional approach in AR and is well supported by existing technologies. Similar to AR applications, accuracy and stability of tracking may generally be important for the disclosed systems in accordance with some embodiments,” (Valli; page 3, para [0176]-[0177]). “Information on display pose may be used in some embodiments to render HF focal planes (additional 3D information) in right scale and perspective. Display pose may be derived, in some embodiments, by capturing (by the tracking camera in glasses), as a minimum, four (typically corner) points on the screen (x.sub.i, i=1, 2, 3, 4) and solving Eq.1 (homography), by using, for example, an iterative method,” (Valli; pages 14-15, para [0193]-[0195]).). Valli in view of Lee is not relied upon teaching but Newcombe teaches using a simultaneous localization and mapping (SLAM) algorithm to perform the pose registration and screen tracking (“We present the first dense SLAM system capable of reconstructing non-rigidly deforming scenes in real-time, by fusing together RGBD scans captured from commodity sensors. Our DynamicFusion approach reconstructs scene geometry whilst simultaneously estimating a dense volumetric 6D motion field that warps the estimated geometry into a live frame. Like KinectFusion, our system produces increasingly denoised, detailed, and complete reconstructions as more measurements are fused, and displays the updated model in real time. Because we do not require a template or other prior scene model, the approach is applicable to a wide range of moving objects and scenes,” (page 343, Abstract; page 350, Conclusion). “Prior to non-rigid optimisation, given a new frame, we first estimate the factorised transformation Tlw using the dense ICP introduced in KinectFusion. This resolves the relative rigid body transformation, i.e. due to camera motion and improves data-association for the non-rigid solver. To that end, we re-render the predicted surface geometry ˆV and perform 2 or 3 iterations of the dense non-rigid optimization. Finally, we factorise out any resulting rigid body transformation ˜ T common across all deformation nodes and update Tlw ← ˜ TTlw,” (page 348, section 3.3.3). “Since the warp function defines a rigid body transformation for all supported space, both position and any associated orientation of space is transformed,” (page 345, section 3.1). “These examples highlight the ability of DynamicFusion to (1) continuously track across large motion during reconstruction,” (page 349, section 4). Pose registration includes solving for / estimating the rigid body transformation.). Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Newcombe to Valli in view of Lee because it is incorporated by reference in Lee (Lee; page 10, para [0152]). Additional motivation would have been to enable reconstruction, deformation, and / or tracking of dynamic models / scenes in real time. Additional motivation would have been to improve the quality and / or realism of visual illusions. Additional motivation would have been to improve the user experience. Additional motivation would have been to enable pose registration and tracking in an unknown environment. Regarding claim 23, it is rejected using the same citations and rationales described in the rejection of claim 3. Regarding claim 11, Valli in view of Lee is not relied upon teaching but Newcombe teaches the system of claim 1, wherein the headset device deforms the initial 3D model of the subject to match the current 3D model of the subject using a deformable simultaneous localization and mapping (SLAM) algorithm (“We present the first dense SLAM system capable of reconstructing non-rigidly deforming scenes in real-time, by fusing together RGBD scans captured from commodity sensors. Our DynamicFusion approach reconstructs scene geometry whilst simultaneously estimating a dense volumetric 6D motion field that warps the estimated geometry into a live frame,” (page 343, Abstract). “DynamicFusion decomposes a non-rigidly deforming scene into a latent geometric surface, reconstructed into a rigid canonical space S ⊆ R3; and a per frame volumetric warp field that transforms that surface into the live frame. There are three core algorithmic components to the system that are performed in sequence on arrival of each new depth frame: 1. Estimation of the volumetric model-to-frame warp field parameters (Section 3.3) 2. Fusion of the live frame depth map into the canonical space via the estimated warp field (Section 3.2) 3. Adaptation of the warp-field structure to capture newly added geometry (Section 3.4),” (page 344, Section 2; page 345, Fig 2). “DynamicFusion takes an online stream of noisy depth maps (a,b) and outputs a real-time dense reconstruction of the moving scene (d,e). To achieve this, we estimate a volumetric warp (motion) field that transforms the canonical model space into the live frame, enabling the scene motion to be undone, and all depth maps to be densely fused into a single rigid TSDF reconstruction (d,f). Simultaneously, the structure of the warp field is constructed as a set of sparse 6D transformation nodes that are smoothly interpolated through a k-nearest node average in the canonical frame (c). The resulting per-frame warp field estimate enables the progressively denoised and completed scene geometry to be transformed into the live frame in real-time (e)…,” (page 345, page 2). “DynamicFusion, an approach based on solving for a volumetric flow field that transforms the state of the scene at each time instant into a fixed, canonical frame. In the case of a moving person, for example, this transformation undoes the person’s motion, warping each body configuration into the pose of the first frame … each point in the canonical frame is transformed to its location in the live frame (see Figure 1),” (page 343; page 343, Fig 1). “We estimate the current values of the transformations dgse3 in Wt given a newly observed depth map Dt and the current reconstruction V by constructing an energy function that is minimised by our desired parameters: PNG media_image2.png 29 504 media_image2.png Greyscale ,” (page 346, Section 3.3). The initial 3D model of the subject includes the canonical model. Deforming the initial 3D model of the subject to match the current 3D model of the subject includes warping the canonical model / scene geometry to be transformed into the live frame.). Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Newcombe to Valli in view of Lee because it is incorporated by reference in Lee (Lee; page 10, para [0152]). Additional motivation would have been to enable reconstruction, deformation, and / or tracking of dynamic models / scenes in real time. Additional motivation would have been to improve the quality and / or realism of visual illusions. Additional motivation would have been to improve the user experience. Regarding claim 31, it is rejected using the same citations and rationales described in the rejection of claim 11. Regarding claim 12, Valli in view of Lee in further view of Newcombe teaches the system of claim 11, wherein the headset device uses a sparse deformation graph to compute the deformations (Newcombe; “a completely dense parametrisation of the warp function is infeasible. In reality, surfaces tend to move smoothly in space, and so we can instead use a sparse set of transformations as bases and define the dense volumetric warp function through interpolation,” (page 345, section 3.1). “We use a deformation graph based regularization defined between transformation nodes, where an edge in the graph between nodes i and j adds a rigid-as-possible regularisation term to the total error being minimized, under the dis continuity preserving Huber penalty ψreg. The total regularisation term sums over all pair-wise connected nodes,” (Newcombe; page 347, Section 3.3.2). “…embedded deformation graphs [25] use a sparsely sampled set of transformation basis functions that can be efficiently and densely interpolated over space,” (Newcombe; page 344, Section 1). “The deformation nodes can be generated by the image processing module 106 by uniform down-sampling of the received 3D model,” (Lee; page 9, para [0137]). The sparse deformation graph includes the sparse set of transformations as bases, transformation nodes, and / or deformation nodes.). and generates a dense vector warp field that represents the warping of the current frame to the previous frame (Newcombe; “Our DynamicFusion approach reconstructs scene geometry whilst simultaneously estimating a dense volumetric 6D motion field that warps the estimated geometry into a live frame,” (Abstract, page 1). “a completely dense parametrisation of the warp function is infeasible. In reality, surfaces tend to move smoothly in space, and so we can instead use a sparse set of transformations as bases and define the dense volumetric warp function through interpolation,” (Newcombe; page 345, section 3.1). “…the image processing module 106 generates (1216) as output a deformed 3D model and associated deformation information—i.e., 3D transformation for each of the deformation nodes. The final deformed 3D model matches the input from the sensor 103, and the final 3D model can then be used as 3D model input to match the next input from the sensor 103,” (Lee; page 10, para [0153]). The dense vector warp field includes the dense volumetric 6D motion field.). Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Newcombe to Valli in view of Lee because it is incorporated by reference in Lee (Lee; page 10, para [0152]). Additional motivation would have been to enable faster / real time computation. Additional motivation would have been to enable faster / real time computation of the warping / deformation. Additional motivation would have been to enable reconstruction, deformation, and / or tracking of dynamic models / scenes in real time. Additional motivation would have been to improve the quality and / or realism of visual illusions. Additional motivation would have been to improve the user experience. Regarding claim 32, it is rejected using the same citations and rationales described in the rejection of claim 12. Claims 20 and 40 are rejected under 35 U.S.C. 103 as being unpatentable over Valli in view of Lee in further view of Saharia et al., "Image Super-Resolution via Iterative Refinement," 30 June 2021, pages 1-28. Regarding claim 20, Valli in view of Lee is not relied upon teaching but Saharia teaches the system of claim 1, wherein the headset device processes one or more of the initial high- definition texture and the current high-definition texture using a generative AI diffusion model algorithm to increase an image quality of the textured 3D model (“We present SR3, an approach to image Super-Resolution via Repeated Refinement. SR3 adapts denoising diffusion probabilistic models [17, 48] to conditional image generation and performs super-resolution through a stochastic iterative denoising process. Output generation starts with pure Gaussian noise and iteratively refines the noisy output using a U-Net model trained on denoising at various noise levels. SR3 exhibits strong performance on super-resolution tasks at different magnification factors, on faces and natural images…,” (page 1, Abstract; page 1, Figure 1). “Single-image super-resolution is the process of generating a high-resolution image that is consistent with an input low-resolution image… We propose SR3 (Super-Resolution via Repeated Refinement), a new approach to conditional image generation, inspired by recent work on Denoising Diffusion Probabilistic Models (DDPM) [17, 47], and denoising score matching [17, 49]… We adapt DDPMs to conditional image generation by proposing a simple and effective modification to the U-Net architecture.,” (page 1, section 1). “The baseline Regression model generates images that are faithful to the inputs, but are blurry and lack detail. By comparison, SR3 produces sharp images with more de tail…,” (page 6, section 4.1; page 5, Figure 3; page 6; Figure 4). “The conditional DDPM model generates a target image y0 in T refinement steps,” (page 2, section 2; page 3, Algorithm 2). A generative AI diffusion model algorithm includes an SR3 algorithm adapting Denoising Diffusion Probabilistic Models (DDPM). From at least Figures 1, 3 and 4 it is clear that a generative AI diffusion model algorithm increases an image quality. After combination, Saharia’s image input becomes Lee’s initial high-definition texture.). Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Saharia to Valli in view of Lee. The motivation would have been to enable the generation of visual illusions in the case of low or suboptimal image and / or video quality. Additional motivation would have been to reduce visual errors and / or surface holes. Additional motivation would have been to improve the quality and / or realism of visual illusions. Additional motivation would have been to improve the user experience. Regarding claim 40, it is rejected using the same citations and rationales described in the rejection of claim 20. Claims 5-6 and 25-26 are rejected under 35 U.S.C. 103 as being unpatentable over Valli in view of Lee in further view of Rahman (US 20150049113 A1; hereinafter Rahman). Regarding claim 5, Valli in view of Lee are not relied upon teaching but Rahman teaches the system of claim 4, wherein capturing input from the user comprises determining, using one or more sensors of the headset device, a gaze of the user's eyes toward the screen of the client computing device (“With further respect to eye tracking capability, the diodes 404 and eye cameras 406, together with the eye tracking module 412, provide eye tracking capability as generally described above… The eye tracking module 412 acquires a video image of the eye from the eye camera 406, digitizes it into a matrix of pixels, and then analyzes the matrix to identify the location of the pupil's center relative to the glint's center, as well as a vector between these centers. Based on the determined vector, the eye tracking module 412 outputs eye gaze coordinates defining an eye gaze point (E),” (page 3, para [0035]). “At step 702, the AR device identifies a portion 508 of an object 504 in a field of view of the HMD based on user interaction with the HMD. In one configuration, the user interaction may be an eye gaze 506, in which case the AR device identifies a portion 508 of an object by tracking the eye gaze 506 of the user to determine a location of the eye gaze, locating a region 512 of the object corresponding to the location of the eye gaze, and detecting the portion 508 within the region. In another configuration, the user interaction may be a gesture, in which case the AR device identifies a portion of an object by tracking the gesture of the user to determine a location of the gesture, locating a region of the object corresponding to the location of the gesture, and detecting the portion within the region,” (page 4, para [0042]). The one or more sensors include the diodes and eye cameras. After combination, Rahman’s object becomes Valli’s external display and / or television.). Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Rahman to Valli in view of Lee. The motivation would have been to improve the user experience. Additional motivation would have been to enable user selection of which subject the user wants enhanced. Additional motivation would have been to enable user selection of which subject the user wants enhanced in the case that there is more than one subject. Additional motivation would have been to reduce the computational load. Additional motivation would have been to enable hands-free selection. Regarding claim 25, it is rejected using the same citations and rationales described in the rejection of claim 5. Regarding claim 6, Valli in view of Lee are not relied upon teaching but Rahman teaches the system of claim 4, wherein capturing input from the user comprises determining a location of the user's hand in the first one or more frames in relation to the screen of the client computing device (“In the case where the user interaction is a gesture, the identification/capture module 808 identifies a portion of an object by track the gesture of the user to determine a location of the gesture, locating a region of the object corresponding to the location of the gesture, and detecting the portion within the region,” (page 4, para [0049]; (page 4, para [0042]). “The scene camera 408, together with the gesture tracking module 414, provide gesture tracking capability … the gesture tracking capability is based on gesture images captured by the scene camera 408. The gesture images are processed by the gesture tracking module … Upon detection of a recognized gesture, the gesture tracking module 414 processes the captured image further to determine the coordinates of a relevant part of the gesture image. In the case of finger pointing, the relevant part of the image may correspond to the tip of the finger. The gesture tracking module 414 outputs gesture coordinates defining a gesture point (G),” (page 3, para [0036]; page 2-3, para [0029]). Determining a location of the user's hand includes determining a location of the gesture. After combination, Rahman’s object becomes Valli’s external display and / or television.). Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Rahman to Valli in view of Lee. The motivation would have been to improve the user experience. Additional motivation would have been to enable user selection of which subject the user wants enhanced. Additional motivation would have been to enable user selection of which subject the user wants enhanced in the case that there is more than one subject. Additional motivation would have been to reduce the computational load. Additional motivation would have been to enable selection without an eye tracker or extra sensors. Additional motivation would have been to increase the accuracy of determining a user’s intended selection. Regarding claim 26, it is rejected using the same citations and rationales described in the rejection of claim 6. Claims 7 and 27 are rejected under 35 U.S.C. 103 as being unpatentable over Valli in view of Lee in further view of Ramirez de Chanlatte et al (US20230147722A1; hereinafter Ramirez de Chanlatte). Regarding claim 7, Valli in view of Lee is not relied upon teaching but Ramirez de Chanlatte teaches the system of claim 1, wherein the headset device converts the frames into the 3D depth maps using a monocular depth map generation technique (“As another example of an additional act not shown in FIG. 7A, act(s) in the series of acts 700 a may include an act of determining the depth map for the real 2D image by utilizing one or more machine-learning models trained to extract monocular depth estimations from the real 2D image,” (page 14, para [0132]; Fig 7A).) Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Ramirez de Chanlatte to Valli in view of Lee. The motivation would have been to enable generating / synthesizing depth data. Additional motivation would have been to enable generating / synthesizing depth data for 2D imagery such as that of Valli (Valli; page 20, para [0243]). Additional motivation would have been to enable deriving 3D / holographic content from 2D content such as that of Valli (Valli; page 18, para [0232]). Additional motivation would have been to enable the prediction of three-dimensional scene depth from a single standard 2D image. Additional motivation would have been to reduce the cost of hardware (such as LiDAR). Regarding claim 27, it is rejected using the same citations and rationales described in the rejection of claim 7. Claims 10, 30 and 41-42 are rejected under 35 U.S.C. 103 as being unpatentable over Valli in view of Lee in further view of Tran et al. (US 20200143558 A1; hereinafter Tran). Regarding claim 10, Valli in view of Lee teaches the system of claim 1, wherein for the first one or more frames of the captured video, the headset device: Lee; “A texture mapped 3D model refers to the virtual object (model) not only has 3D geometric description, including vertices, faces, and surface normal, but also has mapped texture on the surface. The textures are extracted from the images captured during the scan. In the process, the camera poses are recorded, so for each face (represented by a triangle or a polygon) on the model surface the best view are extracted from the recorded images. Then these texture segments are combined and blended to obtain the integrated texture mapping of the object surface,” (page 6, para [0077]).). Valli in view of Lee is not relied upon teaching the device retrieves a reference 3D model based upon one or more characteristics of the subject in the 2D video content. Valli in view of Lee is not relied upon teaching the device deforms the reference 3D model to match the initial 3D model of the subject; Valli in view of Lee does not explicitly teach the texture mapping is done to the deformed initial 3D model. Tran teaches the device retrieves a reference 3D model based upon one or more characteristics of the subject in the 2D video content (“In one embodiment, the closest deformable model is selected to match to a new foot. The selection of the closest model can be done based on size or a suitable attribute such as weight or disease, for example…The method allows new 3D foot scans to be easily retargeted to a library of foot templates or models,” (page 3, para [0041]). “A similar process can be used to create deformable face models, hand models, stomach models, ear models, breast models, and buttock models for cosmetic and health purposes,” (page 3, para [0042]). “The gender and/or the biometric or tailoring measurements are used to select a best fitting deformable body model,” (page 27, para [0248]). “an initial set of feet is scanned using a smart phone, a camera, or a combination of camera and infrared laser/camera, and the points are used in constructing a library of deformable 3D model of feet. The library of deformable 3D models are created and matched to images of the feet through deformations to the points of interest to precisely match the user's actual feet. A physically accurate 3D model can be created therefrom without significant storage of the 3D points and the 3D model of the user can be done using minimal computation resources. The 3D models of feet allow new feet to be customized to a particular 3D template, thus avoiding storage requirements of millions or billions of 3D models of feet. To enhance accuracy, instead of one model for everyone, a plurality of deformable master models can be formed for each of discrete sizes,” (page 2, para [0036]). Retrieving a reference 3D model based upon one or more characteristics of the subject includes selecting the closest model based on size or a suitable attribute. A reference model includes the deformable / master model / template. One or more characteristics include size and / or an attribute.); deforms the reference 3D model to match the initial 3D model of the subject (“Turning back to FIG. 1A, the process selects a standard foot template (26) and morph/warps the standard foot template to match points of interest (28),” (page 3, para [0045]-[0046]; Fig 1A). “The library of deformable 3D models are created and matched to images of the feet through deformations to the points of interest to precisely match the user's actual feet,” (page 2, para [0036]). “photogrammetric techniques are performed to create a 3D model of the body (44) with dimensions based on the reference object. The process optionally selects a standard body template and Morph/Warps the standard body template to match 3D body model (46),” (page 26, para [0239]). “Once the best fitting deformable body model is selected using the body codebook models, the system can match various points on the body to the deformable model,” (pages 27-28, para [0249]). The initial 3D model includes points of interest and / or 3D body model.); Tran teaches the texture mapping is done to the deformed initial 3D model (“Turning back to FIG. 1A, the process selects a standard foot template (26) and morph/warps the standard foot template to match points of interest (28),” (page 3, para [0045]-[0046]; Fig 1A). “photogrammetric techniques are performed to create a 3D model of the body (44) with dimensions based on the reference object. The process optionally selects a standard body template and Morph/Warps the standard body template to match 3D body model (46),” (page 26, para [0239]). (page 2, para [0036]). (pages 27-28, para [0249]). After combination, Lee’s texture mapping is applied to Tran’s morphed/ warped template / deformable model.). Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Tran to Valli in view of Lee. The motivation would have been to improve efficiency / reduce the usage of computation resources (Tran; page 2 para [0036]). Additional motivation would have been to reduce visual errors and / or surface holes (Tran; page 26, para [0242]). Regarding claim 30, it is rejected using the same citations and rationales described in the rejection of claim 10. Regarding claim 41, Valli teaches a system for real-time conversion of two-dimensional (2D) video into three-dimensional (3D) holographic video content (“Some embodiments may be directed to enhancing a normal monoscopic display image with stereoscopic focal planes to provide that a user perceives a 3D effect and is able to naturally accommodate to the 3D content,” (Valli; page 9, para [0140]; Fig 7). “…some embodiments may provide an enhanced view of 2D content (e.g., as presented on an external 2D display) to a wearer of a head mounted display (HMD) device, based on identifying the 2D content, retrieving metadata to provide depth information associated with the 2D content, and processing the 2D content together with the metadata to generate and display in the HMD a multiple focal plane (MFP) representation of the content …deriving 3D from 2D content may be made as, e.g., by some current S3D TV sets, and this 3D information (with synthesized depth data) may be used to display content,” (page 18, para [0232]). “detection and tracking of the display area, and the corresponding geometry transformations are made continually and in real-time,” (pages 11-12, para [0166]). “Multi-focal plane (MFP) near-eye displays create an approximation for the light-field of the displayed scene,” (page 8, para [0130]). Holographic video content includes multi-focal plane (MFP) displays.), the system comprising: a wearable headset device including: one or more cameras configured to capture a front- facing field of view, and one or more displays configured to present digital content to a user of the headset device (“a user may be wearing an HMD which has a multiple-focal plane display capability. The HMD also may have a front-facing camera capable of capturing imagery of the scene in front of the user. The front facing camera may be a depth camera, for example,” (pages 18-19, para [0234]). “…near-eye display glasses may have a camera embedded that captures the content displayed on the screen,” (page 10, para [0149]). “A schematic diagram 400 illustrates a multi-focal near-eye display… As shown in FIG. 4, a left eye 416 and a right eye 418 are shown with associated eye pieces 412, 414 and display stacks 408, 410. The virtual focal planes 406 are shown with the left eye image 402 and right eye image 404 with overlapping area between the two,” (page 8, para [0130]; page 12, para [0171]; Fig 4; Fig 11). The wearable headset includes the HMD / near-eye display glasses. The one or more displays includes the display stacks.), the headset device configured to: capture, using the one or more cameras, video that includes a client computing device in proximity to the user, the client computing device comprising a screen displaying 2D video content (“an input image 1102 (and image depth data 1104 for some embodiments) is obtained via a network connection 1138 (for example, over a metadata channel). For some embodiments, an input image 1102 (and image depth data 1104 for some embodiments) is obtained 1136 by capturing the contents from an external display 1114 by a tracking camera 1116,” (page 11, para [0162]; Fig 7). “an external display is detected and tracked 1004 from the video being captured 1002 by a tracking camera in the glasses. A display area may be detected in some embodiments by the luminosity of the screen in the video, and knowledge on the size and geometry of the display area, when seen from varying viewpoints,” (page 13, para [0176]; Fig 10). “video data may be received at the glasses via a capture device, such as a camera built into the user device, such that the capture device captures the 2D imagery from its view of the external display, and the 2D imagery is displayed to the user via a display device built into the user device,” (para [0147]). “a 2D video 2110 is sent from an image source 2102 to a television… the television 2104 displays 2112 the 2D video,” (pages 20-21, para [0245]-[0246]; page 24, para [323]; Fig 22). The client computing device includes the external display and / or television.); for a first one or more frames of the captured video: identify a region of interest in the first one or more frames that corresponds to a subject in the 2D video content displayed on the client computing device (“capturing 2402, with a camera coupled to the HMD, a video image of a real-world scene. Some embodiments of the example method 2400 may further include identifying 2404 an image pattern present in the captured video image. Some embodiments of the example method 2400 may further include determining 2406 a depth adjustment associated with the identified image pattern. Some embodiments of the example method 2400 may further include generating 2408 a plurality of focal plane images including depth cues for the identified image pattern, the depth cues reflecting a modified depth of the identified image pattern based on the determined depth adjustment. Some embodiments of the example method 2400 may further include displaying 2410 a 3D representation of the identified image pattern including the plurality of focal plane images.” (page 21, para [0248]-[0249]). “…the screen 802 that displays 2D video content may be part of a real-world scene 800,” (page 9, para [0145]). “The HMD may use the front-facing camera to capture real-world imagery, and this imagery may be analyzed (e.g., using stored (e.g., “known”) object or pattern recognition algorithms) to detect image patterns or objects,” (pages 18-19, para [0234]). “…some embodiments may provide an enhanced view of 2D content (e.g., as presented on an external 2D display) to a wearer of a head mounted display (HMD) device…depth information may be determined for an object or image pattern,” (page 18, para [0232]). A region of interest in the first one or more frames that corresponds to a subject in the 2D video content includes the area corresponding to an image pattern present in the captured video image. A region of interest includes the area corresponding to an image pattern. The 2D video content includes the captured video image.), convert the first one or more frames into a first 3D depth map including creating an 3D representation of the subject using the first depth map (“the original depth field may be determined using a depth camera of the HMD, or if the HMD has dual cameras, the original depth field may be determined using a depth from a stereo analysis of the dual captured images,” (page 19, para [0237]). “The modified depth field may be used to generate a multiple focal plane (MFP) representation of the real-world scene, which may be displayed to the user. Any of the various techniques previously described for generating an MFP representation of depth-enhanced content from a 2D external display may be used in a similar fashion to produce the depth-enhanced or depth-modified view of the real-world imagery,” (page 19, para [0238]). “depth information may be calculated (which may be performed locally to a device). For some embodiments, depth information may be generated in whole or in part. For some embodiments, depth information may be determined for an object or image pattern. For some embodiments, deriving 3D from 2D content may be made as, e.g., by some current S3D TV sets, and this 3D information (with synthesized depth data) may be used to display content,” (para [0232]). “the 3D depth information for the 2D content may include a time sequence of depth maps synchronized to the 2D content,” (para [0012]). “identifying 2404 an image pattern present in the captured video image. Some embodiments of the example method 2400 may further include determining 2406 a depth adjustment associated with the identified image pattern. Some embodiments of the example method 2400 may further include generating 2408 a plurality of focal plane images including depth cues for the identified image pattern, the depth cues reflecting a modified depth of the identified image pattern based on the determined depth adjustment. Some embodiments of the example method 2400 may further include displaying 2410 a 3D representation of the identified image pattern including the plurality of focal plane images,” (page 21, para [0248]-[0249]). A 3D depth map includes the depth field and / or depth information. Determining / deriving / generating the depth field and / or depth information includes converting the first one or more frames into a first 3D depth map.), reproject the textured initial 3D model in the displays of the headset device to align with the subject in the 2D video content (“Information on display pose may be used in some embodiments to render HF focal planes (additional 3D information) in right scale and perspective. Display pose may be derived, in some embodiments, by capturing (by the tracking camera in glasses), as a minimum, four (typically corner) points on the screen (x.sub.i, i=1, 2, 3, 4) and solving Eq.1 (homography)…” (pages 14-15, para [0194]-[0195]). “the display area and its geometry are detected from a captured view, and the geometry of the focal planes 1124 are aligned to overlay with the display area 1120,” (pages 11-12, para [0166]). “Three persons 812, 814, 816 are shown wearing MFP glasses in accordance with some embodiments and see the display content in 3D. One spectator 818 sees 810 conventional monoscopic 2D video content on the display 802. Geometry of the focal planes 804, 806, 808 shown to the three persons 812, 814, 816 wearing the MFP glasses are aligned (skewed) along the respective perspective to the display 802,” (Valli; pages 9-10, para [0145]-[0146]; Fig 8). “the depth field is modified to change the depth within the extent of the painting by three inches in the direction of the user. The modified depth field is used together with the real-world imagery to generate a multiple focal plane representation of the scene, which is then displayed to the user via the HMD. From the user's point of view, the paintings the user is interested to see “pop out” of the wall by 3 inches, while paintings by other artists may appear flat against the wall,” (page 19, para [0240]). After combination, Valli’s focal planes become Lee’s textured initial 3D model. Reprojecting the textured initial 3D model in the displays of the headset device includes projective transformation / homography of the focal planes. Align with the subject in the 2D video content includes the geometry of the focal planes 1124 are aligned to overlay with the display area.); for each subsequent frame of the captured video: identify a region of interest in the subsequent frame that corresponds to the subject (“the position and extent (visual footprint) of the detected object or image pattern to be enhanced may be tracked using the HMD camera similar to tracking the 2D external display described above for some embodiments,” (page 19, para [0238]). “The above process may be performed continually (such as in a loop), so that certain objects may be continually updated with depth enhancements as the objects are encountered by the user,” (page 19, para [0239]). “As the user walks around the museum, the front camera of the HMD captures imagery of the museum and an image recognition algorithm is used to identify and classify the paintings in the captured imagery. For each identified painting determined to be by Manet or Renoir, the depth field is modified to change the depth within the extent of the painting by three inches in the direction of the user,” (page 19, para [0240]). Identifying a region of interest in the subsequent frame that corresponds to the subject for each subsequent frame of the captured video includes tracking the detected object or image pattern.), convert the subsequent frame into a subsequent 3D depth map (page 19, para [0237]-[0238]). (page 21, para [0248]-[0249]). “The above process may be performed continually (such as in a loop), so that certain objects may be continually updated with depth enhancements as the objects are encountered by the user,” (page 19, para [0239]). Valli teaches that the conversion steps are repeated (such as in a loop) which would include repeating the steps in a subsequent frame.), reproject the 3D representation of a later time / frame in the displays of the headset device to align with the subject in the 2D video content ((pages 14-15, para [0194]-[0195]). (page 19, para [0240]). “Alignment of HF focal planes with the external display may involve solving in real-time the transformation between world coordinates and the observed coordinates,” (page 14, para [0192]). “…detection and tracking of the display area, and the corresponding geometry transformations are made continually and in real-time,” (pages 11-12, para [0166]). “The above process may be performed continually (such as in a loop), so that certain objects may be continually updated with depth enhancements as the objects are encountered by the user,” (page 19, para [0239]; page 10, para [0147]-[0148]; Fig 10). Valli teaches that the alignment / reprojecting steps are repeated (such as in a loop depicted in at least Figure 10) which would include repeating the steps in a later time / frame.). Valli is not relied upon teaching creating an initial 3D model of the subject using depth information. Valli is not relied upon teaching that the 3D representation is a 3D model. Valli is not relied upon teaching deform the initial 3D model of the subject to match the current 3D model of the subject. Valli is not relied upon teaching overlay a current high-definition texture on the deformed 3D model of the subject, the current texture generated from the subsequent frame. Valli is not relied upon teaching that the 3D representation of a later time / frame is a textured current 3D model. Valli is not relied upon teaching retrieve a reference 3D model based upon one or more characteristics of the subject in the 2D video content; deform the reference 3D model to match the initial 3D model of the subject; and overlay an initial high-definition texture on the deformed initial 3D model of the subject to generate the textured initial 3D model, the initial texture generated from the first one or more frames. Valli is not relied upon teaching for a subsequent frame including creating a current 3D model of the subject using the subsequent depth map. Lee teaches creating an initial 3D model of the subject using depth information (“The image processing module… determines a region of interest (ROI) in the images (e.g., a person's face) and crops the images accordingly at step 360,” (page 6, para [0078]). “generating an initial 3D model for each of the one or more non-rigid objects in the scene using the image comprises determining one or more vertices of each of the one or more non-rigid objects in the scene; determining one or more mesh faces associated with one or more surfaces of each of the one or more non-rigid objects in the scene; determining a surface normal for one or more surfaces of each of the one or more non-rigid objects in the scene; and generating the initial 3D model for each of the one or more non-rigid objects in the scene using the one or more vertices of the object, the one or more mesh faces of the object, and the surface normal for the one or more surfaces of the object,” (page 2, para [0014]). “image processing module 106 generates (1210) a 3D point cloud with normal of the object(s) 102a, 102b in the image and extracts 2D features of the object(s) from the image. In one embodiment, the image processing module 106 generates the 3D point cloud using one or more depth frames with depth camera intrinsic parameters,” (page 9, para [0142]). “The image processing module 106 then generates a 3D model of the person using the captured images—including determining vertices, mesh faces, surface normal, and high-resolution texture in step 310,” (page 4, para [0057]).) Lee teaches that the 3D representation is a 3D model (“The image processing module… determines a region of interest (ROI) in the images (e.g., a person's face) and crops the images accordingly at step 360,” (page 6, para [0078]). “generating an initial 3D model for each of the one or more non-rigid objects in the scene using the image comprises determining one or more vertices of each of the one or more non-rigid objects in the scene; determining one or more mesh faces associated with one or more surfaces of each of the one or more non-rigid objects in the scene; determining a surface normal for one or more surfaces of each of the one or more non-rigid objects in the scene; and generating the initial 3D model for each of the one or more non-rigid objects in the scene using the one or more vertices of the object, the one or more mesh faces of the object, and the surface normal for the one or more surfaces of the object,” (page 2, para [0014]). “image processing module 106 generates (1210) a 3D point cloud with normal of the object(s) 102a, 102b in the image and extracts 2D features of the object(s) from the image. In one embodiment, the image processing module 106 generates the 3D point cloud using one or more depth frames with depth camera intrinsic parameters,” (page 9, para [0142]). “The image processing module 106 then generates a 3D model of the person using the captured images—including determining vertices, mesh faces, surface normal, and high-resolution texture in step 310,” (page 4, para [0057]).) Lee teaches deform the initial 3D model of the subject to match the current 3D model of the subject (“the image processing module 106 generates (1216) as output a deformed 3D model and associated deformation information—i.e., 3D transformation for each of the deformation nodes. The final deformed 3D model matches the input from the sensor 103, and the final 3D model can then be used as 3D model input to match the next input from the sensor 103,”(para [0153]). “The image processing module 106 then non-rigidly matches (1214) the 3D model to the 3D point cloud using 3D points, normal, and 2D features. To non-rigidly match the 3D model to the 3D point cloud, the image processing module 106 iteratively minimizes the matching error—which consists of 3D matching error, 2D matching error and smoothness error—by optimizing the 3D transformation of each deformation node,” (para [0144]). “modifying the mapped 3D model using the tracking information comprises deforming the mapped 3D model using the tracking information to match a 3D model of the one or more non-rigid objects in the scene…” (para [0018]). The current 3D model includes the 3D point cloud.), Lee teaches overlay a current high-definition texture on the deformed 3D model of the subject, the current texture generated from the subsequent frame (“The dynamic mesh covers the part where both the geometry and texture need to be updated in real-time. The image processing module 106 only updates the dynamic mesh in real-time,” (page 5, para [0062]). “When the geometry and texture are updated simultaneously, the system realizes much more vivid and photorealistic animation that cannot be achieved with only the geometry update alone,” (page 5, para [0063]). “Further animation and visualization enhancement are achieved by inserting additional control points on top of the facial landmarks, and by transferring and updating the HD texture of ROI (e.g., the eyes and mouth) in real-time,” (page 6, para [0081]; page 8, para [0124]). “High definition texture mapping: the image processing module 106 accounts for the warping of texture between the deformed and non-deformed models: a) At the texture mapping stage, the image processing module 106 keeps the deformation field associated with the HD images. b) Apply a patch based approach. For each patch, the texture is mapped with the deformation filed added to the model. And then the texture coordinates of the vertices remain constant with or without the associated deformation,” (pages 7-8, para [0112]-[0114]). Overlaying a current high-definition texture on the deformed 3D model includes mapping / transferring the texture with the deformation filed added to the model. The current texture generated from the subsequent frame includes updating the texture / dynamic mesh in real-time.), Lee teaches that the 3D representation of a later time / frame is a textured current 3D model (“The viewing device 112 uses the information to deform the 3D model to match the input RBG+Depth scan captured by the sensor 103 and display the deformed 3D model to a user of the viewing device 112 at the second location,” (para [0153]). “The image 1406 is the deformed 3D model overlaid with the input image by the viewing device 112 to generate a 3D avatar,” (para [0154]). “When the geometry and texture are updated simultaneously, the system realizes much more vivid and photorealistic animation that cannot be achieved with only the geometry update alone,” (page 5, para [0062]-[0063]). “an animation of the three-dimensional avatar of a non-rigid object at the viewing device is substantially synchronized with a movement of the corresponding non-rigid object in the scene as captured by the sensor device,” (page 2, para [0020]). “the viewing device periodically receives updated tracking information from the server computing device and the viewing device uses the updated tracking information to further modify the mapped 3D model,” (page 2, para [0018]). The 3D model is updated over time. Texture is updated with geometry.) Lee teaches frames (Lee; “A texture mapped 3D model refers to the virtual object (model) not only has 3D geometric description, including vertices, faces, and surface normal, but also has mapped texture on the surface. The textures are extracted from the images captured during the scan. In the process, the camera poses are recorded, so for each face (represented by a triangle or a polygon) on the model surface the best view are extracted from the recorded images. Then these texture segments are combined and blended to obtain the integrated texture mapping of the object surface,” (page 6, para [0077]). “The 3D model/avatar is presented by static and dynamic meshes, and with the following enhanced features: real-time high definition texture map for face animation…” (page 1, para [0006]). “the system creates an initial model from the first bunch of scanned data, then enhances and completes the avatar model during the following scan(s) and data capture,” (page 4, para [0056]; Fig 3). “For the holographic/3D selfie scan and modeling step 305 of FIG. 3, the sensor 103 captures one or more high-resolution color images of the object 102a (e.g., a person standing or sitting in front of the sensor). The image processing module 106 then generates a 3D model of the person using the captured images—including determining vertices, mesh faces, surface normal, and high-resolution texture in step 310,” (page 4, para [0057]; Fig 3). Overlaying an initial high-definition texture on the initial 3D model includes mapping the texture to the 3D model. The initial texture generated from the first one or more frames includes the textures extracted from the images captured during the scan. The first one or more frames includes the first bunch of scanned data.). Lee teaches for a subsequent frame including creating a current 3D model of the subject using the subsequent depth map (Lee; “For each frame of the input data from the sensors, the image processing module 106 applies an iterative process to register the model and extract the deformation field, thus the input data removed of the deformation filed is combined with the “rigid” canonical model data for reconstruction and refinement,” (page 5, para [0068]). “The image processing module 106 at the first location receives (1202) an RGB+Depth image of the object 102a (and/or object 102b) in the scene 101 from the sensor 103,” (pages 8-9, para [0134]). “where DFt is the deformation field at time t, Dt is the depth map from the sensors, St is the canonical rigid model projected to the current sensor coordinate system, Dtree is the vertex-landmark connectivity defined in the deformation tree,” (page 7, para [103]-[0104]). “image processing module 106 generates (1210) a 3D point cloud with normal of the object(s) 102a, 102b in the image and extracts 2D features of the object(s) from the image. In one embodiment, the image processing module 106 generates the 3D point cloud using one or more depth frames with depth camera intrinsic parameters,” (page 9, para [0142]). “The image processing module 106 then rigidly matches (1212) the 3D model to the generated 3D point cloud,” (page 9, para [0143]). The current 3D model includes the 3D point cloud. The depth map includes the RGB+Depth image. ). Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Lee to Valli. The motivation would have been to improve the quality and / or realism of visual illusions. Additional motivation would have been to improve spatial alignment. Additional motivation would have been to improve the user experience. Valli in view of Lee is not relied upon teaching the device retrieves a reference 3D model based upon one or more characteristics of the subject in the 2D video content; deform the reference 3D model to match the initial 3D model of the subject. Valli in view of Lee does not explicitly teach the texture mapping is done to the deformed initial 3D model. Tran teaches the device retrieves a reference 3D model based upon one or more characteristics of the subject in the 2D video content (“In one embodiment, the closest deformable model is selected to match to a new foot. The selection of the closest model can be done based on size or a suitable attribute such as weight or disease, for example…The method allows new 3D foot scans to be easily retargeted to a library of foot templates or models,” (page 3, para [0041]). “A similar process can be used to create deformable face models, hand models, stomach models, ear models, breast models, and buttock models for cosmetic and health purposes,” (page 3, para [0042]). “The gender and/or the biometric or tailoring measurements are used to select a best fitting deformable body model,” (page 27, para [0248]). “an initial set of feet is scanned using a smart phone, a camera, or a combination of camera and infrared laser/camera, and the points are used in constructing a library of deformable 3D model of feet. The library of deformable 3D models are created and matched to images of the feet through deformations to the points of interest to precisely match the user's actual feet. A physically accurate 3D model can be created therefrom without significant storage of the 3D points and the 3D model of the user can be done using minimal computation resources. The 3D models of feet allow new feet to be customized to a particular 3D template, thus avoiding storage requirements of millions or billions of 3D models of feet. To enhance accuracy, instead of one model for everyone, a plurality of deformable master models can be formed for each of discrete sizes,” (page 2, para [0036]). A reference model includes the deformable / master model / template. One or more characteristics include size and / or an attribute.); deform the reference 3D model to match the initial 3D model of the subject (“Turning back to FIG. 1A, the process selects a standard foot template (26) and morph/warps the standard foot template to match points of interest (28),” (page 3, para [0045]-[0046]; Fig 1A). “The library of deformable 3D models are created and matched to images of the feet through deformations to the points of interest to precisely match the user's actual feet,” (page 2, para [0036]). “photogrammetric techniques are performed to create a 3D model of the body (44) with dimensions based on the reference object. The process optionally selects a standard body template and Morph/Warps the standard body template to match 3D body model (46),” (page 26, para [0239]). “Once the best fitting deformable body model is selected using the body codebook models, the system can match various points on the body to the deformable model,” (pages 27-28, para [0249]). The initial 3D model includes points of interest and / or 3D body model.); Tran teaches the texture mapping is done to the deformed initial 3D model (“Turning back to FIG. 1A, the process selects a standard foot template (26) and morph/warps the standard foot template to match points of interest (28),” (page 3, para [0045]-[0046]; Fig 1A). “photogrammetric techniques are performed to create a 3D model of the body (44) with dimensions based on the reference object. The process optionally selects a standard body template and Morph/Warps the standard body template to match 3D body model (46),” (page 26, para [0239]). (page 2, para [0036]). (pages 27-28, para [0249]). After combination, Lee’s texture mapping is applied to Tran’s morphed/ warped template / deformable model.). Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Tran to Valli in view of Lee. The motivation would have been to improve efficiency / reduce the usage of computation resources (Tran; page 2 para [0036]). Additional motivation would have been to reduce visual errors and / or surface holes (Tran; page 26, para [0242]). Regarding claim 42, it is rejected using the same citations and rationales described in the rejection of claim 41. Claim 42 additionally recites “the initial texture generated from the first frame.” Valli in view of Lee teaches the initial texture generated from the first frame (Lee; “A texture mapped 3D model refers to the virtual object (model) not only has 3D geometric description, including vertices, faces, and surface normal, but also has mapped texture on the surface. The textures are extracted from the images captured during the scan. In the process, the camera poses are recorded, so for each face (represented by a triangle or a polygon) on the model surface the best view are extracted from the recorded images. Then these texture segments are combined and blended to obtain the integrated texture mapping of the object surface,” (Lee; page 6, para [0077]). “The 3D model/avatar is presented by static and dynamic meshes, and with the following enhanced features: real-time high definition texture map for face animation…” (Lee; page 1, para [0006]). “the system creates an initial model from the first bunch of scanned data, then enhances and completes the avatar model during the following scan(s) and data capture,” (Lee; page 4, para [0056]; Fig 3). “For the holographic/3D selfie scan and modeling step 305 of FIG. 3, the sensor 103 captures one or more high-resolution color images of the object 102a (e.g., a person standing or sitting in front of the sensor). The image processing module 106 then generates a 3D model of the person using the captured images—including determining vertices, mesh faces, surface normal, and high-resolution texture in step 310,” (Lee; page 4, para [0057]; Fig 3). “The image processing module 106 at the first location receives (1202) an RGB+Depth image of the object 102a (and/or object 102b) in the scene 101 from the sensor 103. The image processing module 106 also generates (1204) a 3D model and corresponding texture(s) of the objects 102a, 102b in the scene 101 using, e.g., the real-time object recognition and modeling techniques,” (Lee; pages 8-9, para [0134]; Fig 12). Overlaying an initial high-definition texture on the initial 3D model includes mapping the texture to the 3D model. The initial texture generated from the first one or more frames includes the textures extracted from the images captured during the scan. The first one or more frames includes the first bunch of scanned data. In at least the case that the sensor captures one image, the texture is generated from the first frame.). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to ERICA G THERKORN whose telephone number is (571)272-2939. The examiner can normally be reached Monday - Friday 9:00am - 5:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Devona Faulk can be reached at 571-272-7515. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ERICA G THERKORN/Examiner, Art Unit 2618 /DEVONA E FAULK/Supervisory Patent Examiner, Art Unit 2618 1 Valli is not relied upon teaching a 3D model. Henceforth for the sake of reducing repetitiveness and increasing clarity “3D model” is interpreted as “3D representation” for all of Valli’s teaching throughout all claims. As discussed later in this office action Valli’s 3D representation when combined with Lee’s 3D model becomes the claimed 3D model.
Read full office action

Prosecution Timeline

Mar 13, 2025
Application Filed
Aug 25, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12738054
GENERATION OF ASSOCIATIONS BETWEEN PHYSICAL AND VIRTUAL ENVIRONMENTS
2y 5m to grant Granted Sep 15, 2026
Study what changed to get past this examiner. Based on 1 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
100%
Grant Probability
99%
With Interview (+0.0%)
2y 1m (~7m remaining)
Median Time to Grant
Low
PTA Risk
Based on 3 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month