Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claim 1 is objected to because of the following informalities: in line 8, “precapture image” should be replaced with --precaptured image--, to be consistent with claim 2. Appropriate correction is required.
Claim 2 is objected to because of the following informalities: in line 3, a semicolon should be inserted after “comprising”. Appropriate correction is required.
Claim 4 is objected to because of the following informalities: in line 2, “plurality of point” should be replaced with --plurality of points--. In line 3, “an head mount display device” should be replaced with --a head mount display device--. Appropriate correction is required.
Claim 5 is objected to because of the following informalities: in line 3, a comma should be inserted after “rotation”. Appropriate correction is required.
Claim 6 is objected to because of the following informalities: in line 3, “geometric representation” should be replaced with --geometric representations--. Appropriate correction is required.
Claim 8 is objected to because of the following informalities: in line 2, “live capture image” should be replaced with --live captured image--. Appropriate correction is required.
Claim 9 is objected to because of the following informalities: in line 4, a semicolon should be inserted after “subject”. In line 5, “geometric representation” should be replaced with --geometric representations--. In line 12, “closet” should be replaced with --closest--. Appropriate correction is required.
Claim 10 is objected to because of the following informalities: in line 1, a semicolon should be inserted after “comprising”. In line 4, a semicolon should be inserted after “comprising”. In line 11, “precapture image” should be replaced with --precaptured image--, to be consistent with claim 2. Appropriate correction is required.
Claim 11 is objected to because of the following informalities: in line 3, a semicolon should be inserted after “comprising”. In line 10, “precapture image” should be replaced with --precaptured image--, to be consistent with claim 2. Appropriate correction is required.
Claim 12 is objected to because of the following informalities: in line 5, a semicolon should be inserted after “comprising”. In line 12, “precapture image” should be replaced with --precaptured image--, to be consistent with claim 2. Appropriate correction is required.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1-3, 5-7, and 10-12 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Frueh et al. (U.S. PGPUB 20180101989).
With respect to claim 1, Frueh et al. disclose an image processing method comprising: acquire an image from an image capture device, the image being captured live in real-time (paragraph 65, The processing system 800 includes a camera 810 that is used to capture images of a scene including a user that is represented by the user's head 815…Some embodiments of the camera 810 are video cameras that capture a configurable number of images per second, paragraph 117, the user 1610 is represented by a live 3-D representation that can be computed using a textured point cloud, a textured mesh, and the like);
acquire orientation information associated with a subject in the acquired image (paragraph 69, The processor 820 determines a three-dimensional (3-D) pose that indicates an orientation and a location of the face of the user's head 815 relative to the camera 810. As used herein, the term “pose” refers to parameters that characterize the translation and rotation of a person or object in a scene);
use the acquired orientation information to obtain, from an image repository, a previously captured image of the subject in a similar orientation (paragraph 70, The processor 820 renders a 3-D model of the occluded portion of the user's face and uses the rendered image to overwrite or replace a portion of the HMD 835 in the virtual 3-D scene based on the 3-D pose. The processor 820 renders the 3-D model of the occluded portion of the user's face using texture samples accessed from the eye gaze database 805…the processor 820 can access textures from the face samples associated with the index from an eye gaze database 805 such as the eye gaze database 500 shown in FIG. 5);
generate a composite image by inpainting one or more landmarks from the obtained precapture image that are not present in the acquired live image based on a predetermined geometric representation of the subject (paragraph 70, The processor 820 renders a 3-D model of the occluded portion of the user's face and uses the rendered image to overwrite or replace a portion of the HMD 835); and
display, on a display device, the generated composite image (paragraph 66, The images are rendered on a display 830, paragraph 70, The processor 820 renders a 3-D model of the occluded portion of the user's face and uses the rendered image to overwrite or replace a portion of the HMD 835 in the virtual 3-D scene based on the 3-D pose).
With respect to claim 2, Frueh et al. disclose the image processing method according to claim 1, wherein the subject in the acquired image is wearing a head mount display device that occludes at least an upper region of a face (paragraph 68, portions of the face of the user 815, and in particular the eyes of the user 815, are occluded by the HMD 835 so that the images of the user 815 that are shown in the display 830 (or other displays) have a disconcerting “brick-in-the-face” appearance), and further comprising generating the composite image by using, an upper region of a face in the obtained precaptured image and inpainting the live captured image (paragraph 69, the processor 820 renders a portion of a model of the face of the user 815 that corresponds to the portion of the face that is occluded by the HMD 835 and overwrites a portion of the image corresponding to the HMD 835 with the rendered portion of the model of the face of the user 815).
With respect to claim 3, Frueh et al. disclose the image processing method according to claim 1, wherein the geometric representation of the subject is three dimensional point cloud data of the subject in the live captured image wearing a head mount display device (paragraph 117, the user 1610 is represented by a live 3-D representation that can be computed using a textured point cloud, a textured mesh, and the like).
With respect to claim 5, Frueh et al. disclose the image processing method according to claim 1, wherein the predetermined geometric representation of the subject includes information associated with one or more of scaling, rotation and translation of a head mount display being worn by the subject in the live captured image (paragraph 69, The processor 820 determines a three-dimensional (3-D) pose that indicates an orientation and a location of the face of the user's head 815 relative to the camera 810. As used herein, the term “pose” refers to parameters that characterize the translation and rotation of a person or object in a scene). The pose of the user’s head corresponds to the pose of the head mounted display worn by the user.
With respect to claim 6, Frueh et al. disclose the image processing method according to claim 1, further comprising: selecting the predetermined geometric representation from a candidate set of geometric representation of the subject (paragraph 70, The processor 820 renders the 3-D model of the occluded portion of the user's face using texture samples accessed from the eye gaze database 805. For example, an eye gaze direction of the user 815 can be detected and used as an index into the eye gaze database 805. Texture samples are accessed from the eye gaze database 805 based on the index. For example, the processor 820 can access textures from the face samples associated with the index from an eye gaze database 805 such as the eye gaze database 500 shown in FIG. 5).
With respect to claim 7, Frueh et al. disclose the image processing method according to according to claim 1, further comprising: updating the predetermined geometric representation of the subject for a subsequently acquired live captured image (paragraph 70, The processor 820 renders a 3-D model of the occluded portion of the user's face and uses the rendered image to overwrite or replace a portion of the HMD 835 in the virtual 3-D scene based on the 3-D pose) based on variation of a position of an object being worn by the subject in the acquired live capture image (paragraph 69, the 3-D pose of the user's head 815 relative to the camera 810 include the X, Y, and Z coordinates that define the translation of the user's head 815 and the pitch, roll, and yaw values that define the rotation of the user's head 815 relative to the camera 810). The position of the user’s head corresponds to the position of the head mounted display worn by the user.
With respect to claim 10, Frueh et al. disclose an information processing apparatus (paragraph 65, FIG. 8 is a diagram illustrating a processing system 800) comprising one or more memories storing instructions; and one or more processors that, upon execution of the stored instructions (paragraph 66, The processing system 800 also includes a processor 820 and a memory 825. The processor 820 is configured to execute instructions, such as instructions stored in the memory 825 and store the results of the instructions in the memory 825), are configured to execute an image processing method as in claim 1; see rationale for rejection of claim 1.
With respect to claim 11, Frueh et al. disclose a non-transitory computer readable storage medium (paragraph 147, The software comprises one or more sets of executable instructions stored or otherwise tangibly embodied on a non-transitory computer readable storage medium) storing instructions that, when executed by one or more processors of an information processing apparatus (paragraph 66, The processor 820 is configured to execute instructions, such as instructions stored in the memory 825 and store the results of the instructions in the memory 825), configures the information processing apparatus to execute the method of claim 1; see rationale for rejection of claim 1.
With respect to claim 12, Frueh et al. disclose a system (paragraph 65, FIG. 8 is a diagram illustrating a processing system 800) comprising: a head mount display device configured to be worn by a subject (paragraph 68, The user 815 is wearing an HMD 835 that allows the user to participate in VR, AR, or MR sessions supported by corresponding applications); an image capture device configured to capture real time images of the subject wearing the head mount display device (paragraph 65, The processing system 800 includes a camera 810 that is used to capture images of a scene including a user that is represented by the user's head 815… Some embodiments of the camera 810 are video cameras that capture a configurable number of images per second); and an apparatus configured to execute a method (paragraph 66, The processing system 800 also includes a processor 820 and a memory 825. The processor 820 is configured to execute instructions, such as instructions stored in the memory 825 and store the results of the instructions in the memory 825) comprising the method of claim 1; see rationale for rejection of claim 1.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 4 is/are rejected under 35 U.S.C. 103 as being unpatentable over Frueh et al. (U.S. PGPUB 20180101989) in view of Rinck et al. (U.S. PGPUB 20240143069).
With respect to claim 4, Frueh et al. disclose the image processing method according to claim 1. However, Frueh et al. do not expressly disclose the predetermined geometric representation is a three dimensional point cloud including a plurality of point in three dimensional space of a head of the subject wearing an head mount display device.
Rinck et al., who also deal with generating an image of the user, disclose a method wherein the predetermined geometric representation is a three dimensional point cloud including a plurality of point in three dimensional space of a head of the subject wearing an head mount display device (paragraph 181, FIG. 3 shows the basic set-up. The operator 306 (e.g., wearing the Microsoft Hololens 2 as XR headset or HMD) is in a room 302, and/or an arbitrary office space, and is recorded by three optical sensors 304, in particular three RGB-D cameras (e.g., recording RGB and depth information of the scene). By having three optical sensors, e.g., three cameras, the operator can be covered from different angles in order to calculate a completely color-coded 3D point cloud of the operator 306).
Frueh et al. and Rinck et al. are in the same field of endeavor, namely computer graphics.
Before the effective filing date of the claimed invention, it would have been obvious to apply the method wherein the predetermined geometric representation is a three dimensional point cloud including a plurality of point in three dimensional space of a head of the subject wearing an head mount display device, as taught by Rinck et al., to the Frueh et al. system, because this would visualize computer graphics using a commonly known point cloud structure.
Claim(s) 8 is/are rejected under 35 U.S.C. 103 as being unpatentable over Frueh et al. (U.S. PGPUB 20180101989) in view of Du et al. (WO 2024049481) and further in view of Mallinson (U.S. PGPUB 20160189429).
With respect to claim 8, Frueh et al. disclose the image processing method according to claim 1. However, Frueh et al. do not expressly disclose providing the acquired live capture image to a trained machine learning model that has been trained to identify positions of a predetermined object being worn by a user in an image; generate a similarity score by evaluating a position change of the predetermined object between the live captured image and a next live captured image; determining, based on the generated similarity scored whether to continue to use the predetermined geometric representation or update the geometric representation with a different geometric representation selected from a set of candidate geometric representations
Du et al., who also deal with using head mounted displays, disclose a method for providing the acquired live capture image to a trained machine learning model that has been trained to identify positions of a predetermined object being worn by a user in an image (paragraph 58, a user may be identified on an image or video frame captured with the computing device camera using a combination of image processing, predictive analytics, and/or machine learning. A head mounted device may be further identified as being worn on the user head via further image processing, predictive analytics, and/or machine learning).
Frueh et al. and Du et al. are in the same field of endeavor, namely computer graphics.
Before the effective filing date of the claimed invention, it would have been obvious to apply the method of providing the acquired live capture image to a trained machine learning model that has been trained to identify positions of a predetermined object being worn by a user in an image, as taught by Du et al., to the Frueh et al. system, because this would efficiently identify the head mounted display, by using machine learning.
Mallinson, who also deals with using head mounted displays, discloses a method to generate a similarity score by evaluating a position change of the predetermined object between the live captured image and a next live captured image (paragraph 94, a check is made to determine if the HMD has moved beyond a threshold amount of movement…threshold amount of motion is the amount of motion that would make pixel 506 (as described with reference to FIG. 5) closer to another pixel different from pixel 504, i.e., the adjusted pixel value for pixel 504 is closer to the value of a pixel different from pixel 504, thus corresponding to a similarity score of pixels); determining, based on the generated similarity scored whether to continue to use the predetermined geometric representation or update the geometric representation with a different geometric representation selected from a set of candidate geometric representations (paragraph 95, If the motion is greater than the threshold motion, the method flows to operation 714 where the display data is modified based on the motion. In one embodiment, the modification of the data is performed as described above with reference to FIG. 5). The threshold amount of motion is correlated to the similarity score; if the motion is too high, the similarity score is low and this results in using modified display data.
Frueh et al., Du et al., and Mallinson are in the same field of endeavor, namely computer graphics.
Before the effective filing date of the claimed invention, it would have been obvious to apply the method to generate a similarity score by evaluating a position change of the predetermined object between the live captured image and a next live captured image; determining, based on the generated similarity scored whether to continue to use the predetermined geometric representation or update the geometric representation with a different geometric representation selected from a set of candidate geometric representations, as taught by Mallinson, to the Frueh et al. as modified by Du et al. system, because the HMD adjusts the display data being scanned on the display of the HMD to compensate for the motion of the head of the user, in order to solve the problem of having elements in a virtual reality appeared to be distorted due to the motion of the HMD (paragraph 31 of Mallinson).
Claim(s) 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Frueh et al. (U.S. PGPUB 20180101989) in view of Le Clerc et al. (U.S. PGPUB 20170178306).
With respect to claim 9, Frueh et al. disclose the image processing method according to claim 1, wherein the predetermined geometric representation is determined by: obtaining a finite list of geometric representations of a subject wearing an object that occludes at least a portion of the subject (paragraph 77, using a matching algorithm to match a 3-D model of the face of the user 1015 to pixels in images acquired by the camera 1005… the 3-D model of the face can be rendered for a set of locations and orientations relative to the camera 1005 to produce a set of 2-D model images) select a plurality of geometric representation of the subject wearing the object (paragraph 77, Each of the set of 2-D model images is compared to the image captured by the camera 1005 and the closest match (e.g., the highest score) determines the estimated location and orientation (e.g., the pose PFACE,MATCH) of the user 1015, paragraph 78, Thus, the matching algorithm used to determine PFACE,MATCH is required to match the largely occluded face with an unoccluded 3-D model of the face);
perform inpainting on the live captured image whereby landmarks from a precaptured image having substantially similar orientation and from a region being occluded by the object are inserted into the live captured image (paragraph 85, Once the 3-D pose of the user 1015 in the coordinate system 1020 has been determined, portions of the 3-D model of the user 1015 that correspond to the portions of the user's face that are occluded by the HMD 1010 are rendered and used to replace the corresponding pixels in the images acquired by the camera 1005);
evaluate the inpainted live image using the obtained finite list to determine a similarity score (paragraph 78, In the 2-D case, scores for a hypothetical pose are generated by rendering the 3-D face model from the pose. Pixels that are likely to be occluded are blanked out by rendering a mask that represents the model of the HMD 1010 and laying the mask over the image to indicate the pixels that should be removed from the matching process. Matching is then performed on the remaining pixels in the rendered image of the 3-D face model and the acquired images); and
select, as the predetermined geometric representation, the geometric representation having a closet similarity score (paragraph 77, Each of the set of 2-D model images is compared to the image captured by the camera 1005 and the closest match (e.g., the highest score) determines the estimated location and orientation (e.g., the pose PFACE,MATCH) of the user 1015, thus selecting the geometric representation corresponding to the closest pose). However, Frueh et al. do not expressly disclose obtaining and selecting from a finite list.
Le Clerc et al., who also deal with generating an image of the user, disclose a method wherein the predetermined geometric representation is determined by: obtaining a finite list (paragraph 56, The visual model may be for example selected in a list of different visual objects with corresponding description that may be stored in the memory of the device 10 or that may be stored remotely and downloaded via a network such as the Internet). The process of determining the predetermined geometric representation obtains a finite list.
Frueh et al. and Le Clerc et al. are in the same field of endeavor, namely computer graphics.
Before the effective filing date of the claimed invention, it would have been obvious to apply the method of obtaining and selecting from a finite list, as taught by Le Clerc et al., to the Frueh et al. system, because this would implement the method with an appropriate data structure.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
U.S. PGPUB 20250086916 to Yoon et al. for a method of displaying a composite image representing the appearance of a body part covered by an external HMD device
WO 2015185537 to Burgos et al. for a method of reconstructing a portion of a face covered by an HMD.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ANDREW GUS YANG whose telephone number is (571)272-5514. The examiner can normally be reached M-F 9 AM - 5:30 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kent Chang can be reached at (571)272-7667. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ANDREW G YANG/Primary Examiner, Art Unit 2614
7/30/26