DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1, 2, 11, 12, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Neuberger et al. (WO 2022051135 A1).
Regarding Claim 1, Neuberger discloses “A system comprising: a processor; and a non-transitory computer-readable medium storing computing instructions, that when executed on the processor, cause the processor to perform operations comprising: obtaining a non-frontal image of an item of clothing from a catalog as a candidate for being transformed into a frontal image” (Neuberger, Paragraph [0011], discloses “In various examples, machine learning-based systems and techniques are described that receive a fashion image (e.g., an image of a model wearing an article of clothing) as input and normalize the image according to a pre-specified fashion catalog standard/specification.”; It is important to note that although the paragraph does not explicitly say a nonfrontal image, it can be observed by one of ordinary skill in the art that the fashion image taken can be at any view of a person with a specific piece of clothing on. Additionally, the image(s) selected are then used to transform the image in the context of normalization to make it fit specific catalog standards.); “extracting, using multiple deep-learning blocks, pixel data of a cloth point of interest of the non-frontal image” (Neuberger, Paragraph [0012], discloses “Segmentation of image data includes separation of pixels determined to pertain to one part of an environment from pixels determined to pertain to another part of the environment. For example, a person standing in the foreground of an image may be segmented from the background environment or an article of clothing being held may be segmented from the person holding the article using segmentation techniques.”; Paragraph [0036], discloses “Another deep learning loss that is used for image in-painting is the perceptual loss, where the activations of a pre-trained classification network are compared, as opposed to comparing only the RGB pixel values as in classic loss functions.”; Paragraph [0037], discloses “In various examples, a non-leamable partial convolution layer that omits missing pixel locations in the computation of the convolution operator, yet normalizes the output to compensate for the number of non-missing pixel locations in the computation may be used. In addition, perceptual loss and style loss (e.g., Gram matrix channel comparison of a pre-trained classification network activations) may be added to the LI loss term. These losses may be computed for the holes (e.g., pixels that have been occluded) and the valid pixel locations.”; As shown in the aforementioned paragraphs, the pixel data of cloth point of interest is being segmented away from the actual person wearing the clothing in the image. The multiple deep learning blocks can be observed in [0036] and [0037], as they used deep learning loss functions in order to appropriately account for the loss that comes from extracting or modifying pixel locations after segmentation.); “re-aligning, using a generative pose transfer model, the non-frontal image by altering an angle alignment of a non-frontal pose” (Neuberger, Paragraph [0021] and Figure 1, discloses “…prior to input into alignment model 132, the second image data 142 may be subjected to a pose detection algorithm configured to determine anatomical landmarks of the input image. For example, a thin-plate-spline (TPS) geometric transformation model (e.g., a TPS transform) may be applied to perform a geometric transformation for each pixel in the second image data 142 to determine pose data. In various examples, a determination may be made that the pose data indicates that the model is posed in a pose that is misaligned with respect to a target aligned pose. After pose detection, the second image data 142 may be input into an alignment model 132. Alignment model 132 may be a second generator network trained as part of a second GAN. The alignment model 132 may align a figure (e.g., the model’s pose) to conform to a canonical frontal pose (or other desired pose, depending on the implementation) and/or to generate a symmetrically aligned pose. In the example depicted in FIG. 1, the model’s left shoulder is raised in second image data 142 prior to processing using alignment model 132. However, the model’s shoulders have been aligned in third image data 144.”
PNG
media_image1.png
515
672
media_image1.png
Greyscale
; Here, the alignment model is synonymous to a generative pose transfer model, as the current pose of the human within the human is being generated into a new pose. Although not explicitly disclosed as aligning the angle of the pose, we can clearly see this happening in the model as the model's raised left shoulder gets adjusted using the alignment model.); and “filling in missing areas with simulated cloth matching the cloth point of interest into the frontal image” (Neuberger, Paragraph [0024], discloses “In-painting model 136 may identify the pixels representing the occlusion 150 and may generate pixel values for the occlusion 150 such that the occluded portion of the image is “filled in” to resemble the remainder of the article of clothing in a way that appears natural to the human eye.”); and “tuning, using a cloth feature loss function, multiple parameters of the frontal image for a final version of the frontal image” (Neuberger, Paragraph [0025], discloses “In various examples, one or more of the generator networks (e.g., trained using a GAN) described herein may be selectively enabled to generate image data representing the input fashion image, but that changes one or more aspects of the article of clothing depicted in the input image. For example, a generator network may be trained and enabled to change a color of the input garment image to a selected color. In another example, a generator network may be trained and enabled to change the style, texture, and/or other aspect of the input image so that the output image of the garment comprises one or more desired qualities.”; Paragraph [0026] and Figures 2 discloses “FIG. 2 illustrates an image normalization technique using a generative adversarial network 204, in accordance with various embodiments of the present disclosure. Generative adversarial network 204 may be trained using perceptual loss 208 and adversarial loss 210 and may be used to generate the first generator network used in appearance normalization model 130 (FIG. 1).”
PNG
media_image2.png
328
477
media_image2.png
Greyscale
; As seen in the aforementioned paragraphs, tuning parameters comes in the from the change in the style, texture, and other aspects within the input image. The loss functions are clearly captured through the perceptual loss and adversarial loss, which are both used together in the GAN network setup to compensate for the loss that occurs when tuning the image to the user's liking). Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to use the system obtaining nonfrontal images and modifying them through correcting the pose and adjusting specific parameters seen in Neuberger to improve the pose correction system in the same way. By using the pose correction system disclosed in the Neuberger, one of ordinary skill in the art can effectively allow the user to adjust how clothing would fit on a synthetic image of themselves to ensure that they make an effective purchasing decision. Therefore, it would have been obvious for one of ordinary skill in the art to use the Neuberger reference to achieve the same system described in Claim 1.
Regarding Claim 2, Neuberger discloses “The system of claim 1, wherein the computing instructions, when executed on the processor, further cause the processor to perform an operation comprising: training the generative pose transfer model by using a pose-transfer dataset” (Neuberger, Paragraph [0021], discloses “…prior to input into alignment model 132, the second image data 142 may be subjected to a pose detection algorithm configured to determine anatomical landmarks of the input image. For example, a thin-plate-spline (TPS) geometric transformation model (e.g., a TPS transform) may be applied to perform a geometric transformation for each pixel in the second image data 142 to determine pose data. In various examples, a determination may be made that the pose data indicates that the model is posed in a pose that is misaligned with respect to a target aligned pose. After pose detection, the second image data 142 may be input into an alignment model 132. Alignment model 132 may be a second generator network trained as part of a second GAN. The alignment model 132 may align a figure (e.g., the model’s pose) to conform to a canonical frontal pose (or other desired pose, depending on the implementation) and/or to generate a symmetrically aligned pose. In the example depicted in FIG. 1, the model’s left shoulder is raised in second image data 142 prior to processing using alignment model 132. However, the model’s shoulders have been aligned in third image data 144. The alignment model 132 is described in further detail below in reference to FIG. 3.”); “wherein the pose-transfer dataset comprises pairs of images of a human modeling clothing in non- frontal poses and in frontal poses” (Neuberger, Paragraph [0030], discloses “Alignment model 132 may receive a non-aligned image of a human model as an input and may align the pose of the model to a predetermined pose (e.g., a canonical frontal pose) in an output aligned image 306. Training is performed using pairs of images, where one image is a non-aligned instance of the same human model wearing the same garment (serving as the input to the generator network) and the other image I.sub.t is aligned and serves as the target for the generator network’s output. In another example, training pairs of images may be generated by taking a natural aligned image and modifying the image to introduce an artificial misalignment.”; Here, it is being interpreted under BRI that the pairs of images can be found in nonfrontal poses and frontal poses, as it explicitly mentions here in Paragraph [0030] that the model may receive pairs of images that are not aligned, where it has to correct the pose accordingly to make it frontal).
Claim 11 recites a method with steps corresponding to the elements of the system recited in Claim 1. Therefore, the recited steps of this claim are mapped in the same manner as the corresponding elements in its corresponding system claim. Additionally, the rationale and motivation to use the Neuberger reference, presented in rejection of Claim 1, apply to this claim. Finally, Neuberger recites a processor, memory… (for example, see Neuberger, Paragraph [0019], where it mentions “… computing device(s) 120 may include a non-transitory computer-readable memory 103 and/or may be configured in communication with non-transitory computer-readable memory 103, such as over network 104. In some examples, non-transitory computer-readable memory 103 may store instructions that, when executed by at least one processor of computing device(s) 120, may be effective to perform one or more of the various techniques described herein. …”; Paragraph [0075] also mentions “… software or code can be embodied in any non-transitory computer-readable medium or memory for use by or in connection with an instruction execution system such as a processing component in a computer system …” ).
Claim 12 recites a method with steps corresponding to the elements of the system recited in Claim 2. Therefore, the recited steps of this claim are mapped in the same manner as the corresponding elements in its corresponding system claim. Additionally, the rationale and motivation to use the Neuberger reference, presented in rejection of Claim 2, apply to this claim.
Claim 19 recites a computer-readable storage medium storing a program with instructions corresponding to the elements recited in Claim 1. Therefore, the recited programming instructions of this claim are mapped to in the same manner as the corresponding elements in its corresponding system claim. Additionally, the rationale and motivation to use the Neuberger reference, presented in rejection of Claim 1, apply to this claim. Finally, the Neuberger reference discloses a computer readable storage medium (for example, see Neuberger, Paragraph [0075] (see above for paragraph)).
Claims 3 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Neuberger in view of Grazper Technologies (EP 4044118 A1, anonymous inventor, hereinafter called “Grazper”) and Vallez et al. (US 20240171845 A1, w/ EFD of November 21st, 2023).
Regarding Claim 3, Neuberger discloses “The system of claim 2, wherein training the generative pose transfer model further comprises:” (Please refer to the above-described analysis regarding Claim 2); (Neuberger, Paragraph [0021], discloses “…prior to input into alignment model 132, the second image data 142 may be subjected to a pose detection algorithm configured to determine anatomical landmarks of the input image. For example, a thin-plate-spline (TPS) geometric transformation model (e.g., a TPS transform) may be applied to perform a geometric transformation for each pixel in the second image data 142 to determine pose data. In various examples, a determination may be made that the pose data indicates that the model is posed in a pose that is misaligned with respect to a target aligned pose. After pose detection, the second image data 142 may be input into an alignment model 132. Alignment model 132 may be a second generator network trained as part of a second GAN. The alignment model 132 may align a figure (e.g., the model’s pose) to conform to a canonical frontal pose (or other desired pose, depending on the implementation) and/or to generate a symmetrically aligned pose. In the example depicted in FIG. 1, the model’s left shoulder is raised in second image data 142 prior to processing using alignment model 132. However, the model’s shoulders have been aligned in third image data 144.”); (Neuberger, Paragraph [0025], discloses “In various examples, one or more of the generator networks (e.g., trained using a GAN) described herein may be selectively enabled to generate image data representing the input fashion image, but that changes one or more aspects of the article of clothing depicted in the input image. For example, a generator network may be trained and enabled to change a color of the input garment image to a selected color. In another example, a generator network may be trained and enabled to change the style, texture, and/or other aspect of the input image so that the output image of the garment comprises one or more desired qualities.”; Paragraph [0026] and Figure 2, discloses “FIG. 2 illustrates an image normalization technique using a generative adversarial network 204, in accordance with various embodiments of the present disclosure. Generative adversarial network 204 may be trained using perceptual loss 208 and adversarial loss 210 and may be used to generate the first generator network used in appearance normalization model 130 (FIG. 1).”)
PNG
media_image3.png
352
508
media_image3.png
Greyscale
; “for back-propagation of each generated synthetic image” (Neuberger, Paragraph [0035], discloses “In some examples, a GAN may be trained to assess a prior for how natural an image looks. In various examples, the GAN may be trained together with a mask-based weighted LI loss to determine a context of the input image. As an online optimization task, both the prior and the context losses are back propagated to determine the best latent GAN input vector for each input image.”).
Neuberger does not explicitly disclose “retraining the generative pose transfer model” or “by calculating an L1/L2 cosine distance between numerical representations of a source”. However, in an analogous field of endeavor, Grazper discloses the following in Paragraph [0066]:
PNG
media_image4.png
443
636
media_image4.png
Greyscale
Therefore, it would have been obvious for one of ordinary skill in the art to combine the pose correction system disclosed by Neuberger with the retraining of a model related to poses seen in Grazper to achieve an improved pose correction system and the above-described limitations of Claim 3.
The combination of Neuberger and Grasper does not explicitly disclose “by calculating an L1/L2 cosine distance between numerical representations of a source”. However, in an analogous field of endeavor, Vallez discloses “computing the distance between each candidate object feature vector and a target object feature vector using one or more of cosine distance, L1 norm, and L2 norm.” (Vallez, Paragraph [0008]). It is important to note that a vector can represent a numerical representation of a source. Therefore, it would have been obvious for one of ordinary skill in the art to combine the system seen in the combination of Neuberger and Grazper with the Vallez technique of calculating a L1/L2 cosine distance between numerical representations of a source to have an improved pose correction system. By combining the system seen in Neuberger and Grazper with the Vallez technique of calculating the L1/L2/ cosine distance, one of ordinary skill in the art allows for the model to get increasingly accurate according to the images it is being trained with. Thus, it would have been obvious for one of ordinary skill in the art to combine the Neuberger, Grazper, and Vallez references to achieve the same system described in Claim 3.
Claim 13 recites a method with steps corresponding to the elements of the system recited in Claim 3. Therefore, the recited steps of this claim are mapped to the proposed combination in the same manner as the corresponding elements in its corresponding system claim. Additionally, the rationale and motivation to combine the Neuberger, Grazper, and Vallez references, presented in rejection of Claim 3, apply to this claim.
Claims 4, 8, 14, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Neuberger in view of Kristal et al. (US 20230014804 A1).
Regarding Claim 4, Neuberger discloses “The system of claim 1, wherein extracting the pixel data further comprises:” (Please refer to the above-described analysis regarding Claim 1); (Kristal, Paragraph [0087]). Kristal further discloses “The Product Handler Module 6 may divide or isolate between two or more clothes items that appear to exist within a single product image that is intended to correspond to a single product. For example, the Products Web Crawler Module 12 may obtain an image that corresponds to meta-data or textual description of “shirt,” whereas the actual product image depicts a human model that wears a shirt and also pants. Based on contextual analysis, and/or based on meta-data analysis or classification analysis, the Product extraction process 6A may determine that it is required to isolate only the shirt image from the product image, and to discard from that image not only the human model, and not only the background image (if any), but also to discard the pants image-data from that product image. A suitable grab-cut algorithm may be applied, by utilizing the textual or contextual analysis results, to isolate only the shirt image, and to discard the pants image-portion and the other elements.” (Kristal, Paragraph [0143]). Here, we can clearly see that multiple pieces of clothing being segmented from the human model in the image. Although not explicitly described, it can be interpreted by one of ordinary skill in the art that this method can apply to someone wearing clothing that contain rigid, inner, and transparent parts. For example, if a man was wearing a translucent shirt with jeans that have visible seams (which is a inner part of clothing) when an image is taken, the Kristal method can effectively isolate each piece of clothing, with each piece having their own image when segmented out of the image. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date before the claimed invention to combine the system seen in Neuberger with the Kristal technique for segmenting pixels of clothing into multiple types to achieve a more improved pose correction system, as described in Claim 4.
Regarding Claim 8, Neuberger discloses “The system of claim 1” (Please refer to the above-described analysis regarding Claim 1); (Kristal, Paragraph [0079]). Here, vision processing is synonymous to the computer vision that Kristal is using to develop a 2D or 3D estimation of the clothing on the person. Kristal further discloses “The User Handler Module 4 may thus convert a user image, in which the user was standing in a non-neutral position (e.g., one arm behind his back; or, one arm folded and holding a smartphone over the chest; or, two arms bent and holding the waist), into a mask or template of a neutrally-standing depiction of the user. Optionally, a user's image that is not fully squared towards the camera, or that is slanted or tilted, may be modified to appear generally aligned towards the camera. The User Handler Module 4 may modify or cure such deficiencies. FIGS. 7 and 8 are illustrative. In FIG. 7, a user image depicts the user's two arms being bent and holding the waist. In FIG. 8, a user image depicts a mirror selfie with of the user's one arm folded and holding a smartphone. In both cases, these issues may be identified and corrected so that the arms and hands are placed in a neutral position (not shown), such as with both arms and hands at the user's sides.” (Kristal, Paragraph [0130]). As shown, the angle of the person in a nonfrontal pose is realigned accordingly to be a frontal pose using the User Handler Module. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the system seen in Neuberger with the Kristal technique of aligning the angles of the non-frontal pose into multiple frontal poses to correct any differences within the pose correction system. By combining the system seen in Neuberger with the Kristal technique of altering the non-frontal pose, one of ordinary skill in the art can allow the user to adjust for any differences to minimize any errors in the new synthetic images of the person trying on the clothes to make it appropriately aligned. Therefore, it would have been obvious for one of ordinary skill to combine the Neuberger and Kristal references to achieve the same system described in Claim 8.
Claim 14 recites a method with steps corresponding to the elements of the system recited in Claim 4. Therefore, the recited steps of this claim are mapped to the proposed combination in the same manner as the corresponding elements in its corresponding system claim. Additionally, the rationale and motivation to combine the Neuberger and Kristal references, presented in rejection of Claim 4, apply to this claim.
Regarding Claim 17, Neuberger discloses “The method of claim 11, wherein re-aligning the non-frontal image comprises at least one of:” (Please refer to the above described analysis of Claim 11); (Kristal, Paragraph [0079]). Here, vision processing is synonymous to the computer vision that Kristal is using to develop a 2D or 3D estimation of the clothing on the person. Kristal further discloses “The User Handler Module 4 may thus convert a user image, in which the user was standing in a non-neutral position (e.g., one arm behind his back; or, one arm folded and holding a smartphone over the chest; or, two arms bent and holding the waist), into a mask or template of a neutrally-standing depiction of the user. Optionally, a user's image that is not fully squared towards the camera, or that is slanted or tilted, may be modified to appear generally aligned towards the camera. The User Handler Module 4 may modify or cure such deficiencies. FIGS. 7 and 8 are illustrative. In FIG. 7, a user image depicts the user's two arms being bent and holding the waist. In FIG. 8, a user image depicts a mirror selfie with of the user's one arm folded and holding a smartphone. In both cases, these issues may be identified and corrected so that the arms and hands are placed in a neutral position (not shown), such as with both arms and hands at the user's sides.” (Kristal, Paragraph [0130]). As shown, the angle of the person in a nonfrontal pose is realigned accordingly to be a frontal pose using the User Handler Module. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the method seen in Neuberger with the Kristal technique of aligning the angles of the non-frontal pose into multiple frontal poses to correct any differences within the pose correction system. By combining the method seen in Neuberger with the Kristal technique of altering the non-frontal pose, one of ordinary skill in the art can allow the user to adjust for any differences to minimize any errors in the new synthetic images of the person trying on the clothes to make it appropriately aligned. Therefore, it would have been obvious for one of ordinary skill to combine the Neuberger and Kristal references to achieve the same method described in Claim 17.
Claims 5 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Neuberger in view of Li et al. (US 20230086880 A1).
Regarding Claim 5, Neuberger discloses “The system of claim 1, wherein: the generative pose transfer model is trained to generate images of the simulated cloth based on the pixel data and metadata of the item of clothing featured in the non-frontal image; and the non-frontal image is” (Neuberger, Paragraph [0023], discloses “The segmentation mask may be applied to third image data 144 such that pixels that do not pertain to the relevant article of clothing (e.g., pixels in the segmentation mask that pertain to the model, the background, or to any other class apart from the dress) are “masked out” (sometimes referred to as being “segmented from” the image data). As used herein, “masked out” pixels may be set to a particular pixel value in order to segment the article of clothing from the image. For example, all non-dress pixels may be set to a bright white color or to a black color to show the dress removed from the model and/or from the background. Applying the segmentation mask to the third image data 144 generates fourth image data 146. In the example in FIG. 1, fourth image data 146 depicts the dress with the human model removed. However, there is an occlusion 150 in fourth image data 146 where the model’s hand covered a portion of the dress.”; Paragraph [0024] discloses “The fourth image data 146 including any occlusions (e.g., occlusion 150 and/or other portions of the garment that are occluded from view) may be sent to an in-painting model 136. In-painting model 136 may identify the pixels representing the occlusion 150 and may generate pixel values for the occlusion 150 such that the occluded portion of the image is “filled in” to resemble the remainder of the article of clothing in a way that appears natural to the human eye. For example, in output image data 148, the occlusion 150 has been removed and has been replaced by pixel values that resemble other non-occluded portions of the dress. Output image data 148 may conform to relevant catalog standards. For example, output image data 148 may be a high quality packshot image that has been automatically generated by the system 100 without requiring manual editing by an expert photo-editor.”; Here, although the example here shows a frontal pose, it can be interpreted by one of ordinary skill in the art that the segmentation and filling process of the simulated cloth can be applied to nonfrontal images as well; Paragraph [0032] discloses “The semantic segmentation network 404 extracts the human figure parts from the image in order to generate a “garment-only” image using segmentation mask 406. After applying segmentation mask 406 to the input image (e.g., aligned image 306) the resulting image includes only those pixels that are classified as pertaining to the garment. The segmentation mask is a classification of the human body parts and fashion garments into labels that are predefined to the relevant domain (e.g., “tops,” “left hand,” “hair,” “leg,” etc.).”; The metadata can be defined as the labels created for specific clothing and parts); “one of a ghost cloth image, a flat cloth image, or an image of a cloth on a mannequin”. Therefore, it would have been obvious for one of ordinary skill in the art to combine the system seen in Neuberger with the Li technique of having ghost cloth, flat cloth, or clothed mannequin image to give a variety of options for trying on pieces of clothing. By combining the system seen in Neuberger with the Li technique of having different types of cloth images, one of ordinary skill in the art can permit the user to take photos from different pieces of clothing to try on themselves instead of a preselection of options. Therefore, it would have been obvious for one of ordinary skill in the art to combine the Neuberger and Li references to achieve the same system seen in Claim 5.
Claim 15 recites a method with steps corresponding to the elements of the system recited in Claim 5. Therefore, the recited steps of this claim are mapped to the proposed combination in the same manner as the corresponding elements in its corresponding system claim. Additionally, the rationale and motivation to combine the Neuberger and Li references, presented in rejection of Claim 5, apply to this claim.
Claims 6 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Neuberger in view of Chen et al. (US 20180181802 A1).
Regarding Claim 6, Neuberger discloses “The system of claim 1, wherein the generative pose transfer model is trained to” (Please refer to the above-described analysis regarding Claim 1); (Chen, Paragraph [0031]). Chen also discloses “For instance, the image manipulation application receives the input image (e.g., by performing a 3D scan of a human being) and uses the trained machine learning algorithm to compute a feature descriptor for the input image. The feature descriptor (e.g., a vector) includes a first set of data describing a pose of the figure in the input image, a second set of data describing a body shape of the figure in the input image, and a third set of data describing one or more clothing items of the figure in the input image. The image manipulation application queries, using the feature descriptor, a database (or other data structure) of example images which depict different body shapes at different poses and with varying clothing” (Chen, Paragraph [0033]). Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the system seen in Neuberger with the Chen technique of converting both the cloth and synthetic image to provide different representations to manipulate the images. By combining the system seen in Neuberger with the Chen image-to-vector technique, one of ordinary skill in the art permits the user of the system to transform the images into numerical representations to be used in different computer software or programming languages to adjust the pose or clothing on the person. Therefore, it would have been obvious for one of ordinary skill in the art to combine the Neuberger and Chen references to achieve the same system seen in Claim 6.
Claim 16 recites a method with steps corresponding to the elements of the system recited in Claim 6. Therefore, the recited steps of this claim are mapped to the proposed combination in the same manner as the corresponding elements in its corresponding system claim. Additionally, the rationale and motivation to combine the Neuberger and Chen references, presented in rejection of Claim 6, apply to this claim.
Claims 7 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Neuberger in view of Jie-hua Peng (CN 203070360 U).
Regarding Claim 7, Neuberger discloses “The system of claim 1, wherein re-aligning the non-frontal image comprises: generating a synthetic image of a reference image; and transforming the synthetic image into an A-frame image, wherein the A-frame image comprises a pre-configured frontal pose, and wherein the pre-configured frontal pose comprises” (Neuberger, Paragraph [0068], discloses “Process 700 may continue from action 710 to action 720, at which second image data may be generated from the first image data using a generator network. In various examples, the generator network may be trained as part of a GAN (e.g., as described above in reference to FIG. 2). The second image data output by the first generator network may have removed the one or more photometric artifacts present in the first image data such that the second image data has been normalized for the relevant standards (e.g., for photographic quality standards associated with an online catalog, marketplace, and/or e-commerce service)”; Paragraph [0069] discloses “Process 700 may continue from action 720 to action 730, at which a second generator network may be used to generate third image data representing the human in a different pose relative to the pose of the human in the first (and second) image data. At action 730, the post of the human in the second image data may be determined. The second generator (trained as part of a second GAN, as described above in reference to FIG. 3) may be used to modify the pose of the human to generate the third image data, in which the human model’s pose may conform to a standard pose (for which the second GAN network has been trained). For example, a common pose for a packshot image may be a frontal pose.”; Using BRI, the A-frame image can be interpreted as the synthetic image of the standard or common pose. An A-pose is normally a relaxed position with the head, arms and arms formed in an "A" position. If a standard pose is being used to transform the synthetic image, it could be seen by one of ordinary skill in the art to potentially represent an A-pose, which would therefore translate into an A-frame image. Coincidentally, that A-frame image shown in Neuberger has a frontal pose.) (Peng, Paragraph [0061]). Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the system seen in Neuberger with the Peng technique of having a skeleton diagram of the reference image to improve the pose correction system in the same way, as seen in the system of Claim 7.
Claim 20 recites a computer-readable storage medium storing a program with instructions corresponding to the element recited in Claim 7. Therefore, the recited programming instructions of this claim are mapped to the proposed combination in the same manner as the corresponding elements in its corresponding system claim. Additionally, the rationale and motivation to combine the Neuberger and Peng references, presented in rejection of Claim 7, apply to this claim.
Claims 9 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Neuberger in view of Wang et al. (Matching User Photos to Online Products with Robust Deep Features) and Hauswiesner et al. (Virtual Try-On through Image-Based Rendering).
Regarding Claim 9, Neuberger discloses “The system of claim 1, wherein the cloth feature loss function” (Please refer to the above-described analysis regarding Claim 1); . Neuberger is not relied on to disclose “implements contrastive loss learning of visual representations by: maximizing alternative augmentations of the cloth point of interest; and minimizing a distance between images of the item of clothing.” However, in an analogous field of endeavor, Wang discloses the following in pages 7-8 and Figure 1:
PNG
media_image5.png
686
425
media_image5.png
Greyscale
PNG
media_image6.png
406
749
media_image6.png
Greyscale
As shown in the aforementioned paragraphs and Figure 1, robust contrastive loss is used with the user image to determine which clothing image would match or be the right fit with what they are wearing right now. By implementing the contrastive loss, the model learns how to distinctly identify different pieces of clothing and accurately determine the correct match according to the image, thus making the model more reliable. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the system seen in Neuberger with the Wang technique of implementing contrastive loss of visual representations to achieve the above-described limitation of Claim 9.
The combination of Neuberger and Wang does not explicitly disclose “maximizing alternative augmentations of the cloth point of interest; and minimizing a distance between images of the item of clothing”. However, in an analogous field of endeavor, Hauswiesner discloses the following for maximizing augmentations in Figure 1 on page 1554:
PNG
media_image7.png
245
800
media_image7.png
Greyscale
From this image, we can clearly see that the pipeline that Hauswiesner created incorporates various extraction, rendering, and registration techniques on the user image and garment database to acquire enough data to maximize the augmentations that the user can do when virtually trying on different pieces of clothing. Hauswiesner further discloses for minimizing a distance the following excerpt on page 1556:
PNG
media_image8.png
281
352
media_image8.png
Greyscale
Here, this paragraph clearly demonstrates the distance minimization, as the algorithm seen here uses Euclidean distance to find the closest feature vector of the user image that would identify with a desired garment within the clothing image database. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the system seen in the combination of Neuberger and Wang with the Hauswiesner technique of maximizing alternative augmentations and minimizing the distance between images of an item of clothing to achieve a improved pose correction system, as seen in Claim 9.
Regarding Claim 18, Neuberger discloses “The method of claim 11, wherein at least one of: the cloth feature loss function” (Please refer to the above-described analysis for Claim 1); frontal image is within a predetermined numerical domain threshold” fully. However, in an analogous field of endeavor, Wang discloses the following in pages 7-8 and Figure 1:
PNG
media_image5.png
686
425
media_image5.png
Greyscale
PNG
media_image6.png
406
749
media_image6.png
Greyscale
As shown in the aforementioned paragraphs and Figure 1, robust contrastive loss is used with the user image to determine which clothing image would match or be the right fit with what they are wearing right now. By implementing the contrastive loss, the model learns how to distinctly identify different pieces of clothing and accurately determine the correct match according to the image, thus making the model more reliable. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the method seen in Neuberger with the Wang technique of implementing contrastive loss of visual representations to achieve the above described limitation of Claim 18.
The combination of Neuberger and Wang does not explicitly disclose “maximizing alternative augmentations of the cloth point of interest; and minimizing a distance between images of the item of clothing”. However, in an analogous field of endeavor, Hauswiesner discloses the following for maximizing augmentations in Figure 1 on page 1554:
PNG
media_image7.png
245
800
media_image7.png
Greyscale
From this image, we can clearly see that the pipeline that Hauswiesner created incorporates various extraction, rendering, and registration techniques on the user image and garment database to acquire enough data to maximize the augmentations that the user can do when virtually trying on different pieces of clothing. Hauswiesner further discloses for minimizing a distance the following excerpt on page 1556:
PNG
media_image8.png
281
352
media_image8.png
Greyscale
Here, this paragraph clearly demonstrates the distance minimization, as the algorithm seen here uses Euclidean distance to find the closest feature vector of the user image that would identify with a desired garment within the clothing image database. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the method seen in the combination of Neuberger and Wang with the Hauswiesner technique of maximizing alternative augmentations and minimizing the distance between images of an item of clothing to achieve an improved pose correction method, as seen in Claim 18.
Claims 10 are rejected under 35 U.S.C. 103 as being unpatentable over Neuberger in view of Chen, and further in view of Kristal and Li 2 et al. (CN 114758334 A).
Regarding Claim 10, Neuberger discloses “The system of claim 1, wherein tuning the multiple parameters further comprises:” (Please refer to the above-described analysis regarding Claim 1); (Neuberger, Paragraph [0021], discloses “… prior to input into alignment model 132, the second image data 142 may be subjected to a pose detection algorithm configured to determine anatomical landmarks of the input image. For example, a thin-plate-spline (TPS) geometric transformation model (e.g., a TPS transform) may be applied to perform a geometric transformation for each pixel in the second image data 142 to determine pose data. In various examples, a determination may be made that the pose data indicates that the model is posed in a pose that is misaligned with respect to a target aligned pose. After pose detection, the second image data 142 may be input into an alignment model 132. Alignment model 132 may be a second generator network trained as part of a second GAN. The alignment model 132 may align a figure (e.g., the model’s pose) to conform to a canonical frontal pose (or other desired pose, depending on the implementation) and/or to generate a symmetrically aligned pose. In the example depicted in FIG. 1, the model’s left shoulder is raised in second image data 142 prior to processing using alignment model 132. However, the model’s shoulders have been aligned in third image data 144.”), translating a reference cloth feature into a numerical descriptor that “For instance, the image manipulation application receives the input image (e.g., by performing a 3D scan of a human being) and uses the trained machine learning algorithm to compute a feature descriptor for the input image. The feature descriptor (e.g., a vector) includes a first set of data describing a pose of the figure in the input image, a second set of data describing a body shape of the figure in the input image, and a third set of data describing one or more clothing items of the figure in the input image. The image manipulation application queries, using the feature descriptor, a database (or other data structure) of example images which depict different body shapes at different poses and with varying clothing” As shown in the analysis of Claim 6, this paragraph transforms the input image into a vector, wherein the image contains the reference cloth feature on the human standing in a specific pose (which could be nonfrontal). As noted in Chen, the vector can be considered a numerical descriptor. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the system seen in Neuberger with the Chen technique of translating the reference cloth feature into a numerical descriptior to achieve the above-described limitations seen in the system of Claim 10.
The combination of Neuberger and Chen does not explicitly disclose “to measure cloth features of the final version of the frontal image, wherein the numerical descriptor of the final version of the frontal image is within a predetermined numerical domain threshold”. However, in an analogous field of endeavor for measuring the cloth features, Kristal discloses the following in Paragraphs [0361]-[0368]:
PNG
media_image9.png
290
331
media_image9.png
Greyscale
PNG
media_image10.png
137
318
media_image10.png
Greyscale
PNG
media_image11.png
94
363
media_image11.png
Greyscale
This algorithm clearly extracts the clothing within the image, and does an initial measurement on the users clothes according to the measurements known for specific types of clothing (such as the bust, height, weight, and hip circumference) to ensure an appropriate fit before putting the virtual clothing on the user within the user image provided. Kristal further discloses the following in Paragraphs [0385]-[0388]:
PNG
media_image12.png
148
329
media_image12.png
Greyscale
As shown, the algorithm checks how the virtual item clothing fits the person after producing the final image to determine if it is the right fit for the user in terms of size. If the clothing is not the right fit for the user, the algorithm them proceeds to modify the image accordingly, either through vertical rescaling as shown in Paragraphs [0387]-[0388] or horizontal rescaling (see PGPub for more information). Finally, Kristal discloses for the frontal image the following in Figures 7 and 8:
PNG
media_image13.png
476
446
media_image13.png
Greyscale
Figure 7 and 8 clearly describe the frontal image, as you can see the users of the method clearly facing the camera to take pictures. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the system seen in the combination of Neuberger and Chen with the Kristal technique of measuring the cloth features of the final version of the frontal image to appropriately assess the size of clothing on the user of the pose correction system, as seen in the above-described limitation of Claim 10.
The combination of Neuberger, Chen, and Kristal does not explicitly disclose “wherein the numerical descriptor of the final version of the frontal image is within a predetermined numerical domain threshold”. However, in an analogous field of endeavor, Li 2 discloses “obtaining one or more first local feature points in the first image, wherein the matching distance between descriptors of the first local feature points and descriptors of local feature points in a feature library is smaller than or equal to a first threshold value” (Li 2, Paragraph [0023]). Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the system seen in the combination of Neuberger, Chen, and Kristal with the Li 2 technique of determining whether the numerical descriptor is within a predetermined domain threshold to achieve an improved pose correction system, as shown in Claim 10.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Appleboim et al. (US 20210090209 A1) teaches A technique for transforming an image of an article for virtual presentation without transformation-induced distortion of a shape-invariant area of the article.
Fedyukov et al. (US 20210049811 A1) teaches a method for remote clothing selection.
Koujan et al. (US 20240127563 A1, w/ EFD of October 17th, 2022) teaches methods and systems for performing real-time stylizing operations.
Wen-Huang et al. (Fashion Meets Computer Vision: A Survey) teaches a comprehensive survey of more than 200 major fashion-related works covering four main aspects for enabling intelligent fashion.
Wu et al. (US 20220237879 A1) teaches A method for training a real-time, direct clothing modeling for animating an avatar for a subject.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SORIE I KOROMA JR whose telephone number is (571)272-9259. The examiner can normally be reached Monday - Friday 8AM-6:00PM; Alternate Fridays Off.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amandeep Saini can be reached at 571-272-3382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SORIE I KOROMA JR/Examiner, Art Unit 2662
/AMANDEEP SAINI/Supervisory Patent Examiner, Art Unit 2662