Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
This action is in response to amendment and/or remarks filed 6/23/2026. In the current amendments, claims 1, 16-19, 25, and 29-30 have been amended. The arguments made against the 35 U.S.C 112(b) rejections have been considered and were found to be persuasive. The arguments made against the 35 U.S.C 101 rejections have been considered and were found to be persuasive. The arguments made against the 35 U.S.C 103 have been considered but are moot because the new ground of rejection does not rely on references applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Applicant’s amendments have necessitated a change in the references applied.
The Examiner cites particular sections in the references as applied to the claims below for the convenience of the applicant(s). Although the specified citations are representative of the teachings in the art and are applied to the specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested that, in preparing responses, the applicant(s) fully consider the references in their entirety as potentially teaching all or part of the claimed, as well as the context of the passage as taught by the prior art or disclosed by the Examiner.
Response to Arguments
Applicant’s arguments and remarks filed 06/23/2026 have been fully considered and were found to be persuasive, however a new ground of rejection (35 U.S.C 103) has been made.
35 U.S.C 103
Applicant asserts:
Applicant asserts “As noted above, the Office rejected claim 1 as allegedly being obvious in view of the combination of Xu and Lad. Applicant has amended claim 1 and respectfully requests reconsideration of the claim in its new form. For example, it is submitted that the combination of Xu and Lad fails to disclose or make obvious all features of amended claim 1. Xu relates to "outputs of multi-scale random convolutions as new images." Xu, page 1. Xu notes that "randomized convolutions create an infinite number of new domains with similar global shapes but random local texture." Id. For instance, Xu describes "a data augmentation technique using multi-scale random-convolutions to generate images with random texture while maintaining global shapes." Id. Xu further describes "using the [] output as training images or mixing it with the original images." Id. Xu explains that the techniques can "augment[] images with different local texture but the same semantics with the shape-preserving property of random convolutions. Lad relates to a "deep learning algorithm with the ability to provide a high-performance classifier to predict either the presence of geographic atrophy (GA)." Lad, Abstract. Lad describes a "system using machine learning in detecting geographic atrophy (GA), the system comprising at least one processor; a memory; and a computing platform including the at least one processor and the memory. However, Applicant submits that the combination of Xu and Lad fails to disclose or make obvious, at least, "applyfing] semantic-aware style fusion to the aggregated training data to generate fused training data, wherein to apply the semantic-aware style fusion, the at least one processor is configured to extract semantic regions from the aggregated training data, combine the semantic regions into a common semantic region, and generate the fused training data from the combined semantic regions," as recited in amended claim 1. The Office submitted that Xu describes features relating to "semantic-aware style fusion." See Office Action, p. 13. Applicant respectfully disagrees.”.
Examiner’s response:
Examiner considered the arguments, but they are found to be moot because the new grounds of rejection does not rely on references applied in the prior rejection of record for any teaching or matter specifically challenged in the arguments.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
The following appears to be the closest portions of the specification corresponding to the 35 U.S.C 112(f) invocations:
augmenting, via a random style generator having at least one randomly initialized layer, training data to generate augmented training data
[0062] The operations can include augmenting, via a random style generator engine 712 (which can also be referred to as a random style generator 712) having at least one randomly initialized layer 714, training data X0 710 used to generate augmented training data X1 722 and aggregating data with a plurality of styles from the augmented training data to generate aggregated training data.
aggregating data with a plurality of styles from the augmented training data to generate aggregated training data
[0062] The operations can include augmenting, via a random style generator engine 712 (which can also be referred to as a random style generator 712) having at least one randomly initialized layer 714, training data X0 710 used to generate augmented training data X1 722 and aggregating data with a plurality of styles from the augmented training data to generate aggregated training data
applying semantic-aware style fusion to the aggregated training data to generate fused training data
[0062] The electronic device can perform further operations including applying semantic-aware style fusion engine 706 to the aggregated training data to generate fused training data and adding the fused training data as fictitious samples to the training data to generate updated training data for training a neural network or machine learning model.
adding the fused training data as fictitious samples to the training data to generate updated training data for training a neural network
[0062] The electronic device can perform further operations including applying semantic-aware style fusion engine 706 to the aggregated training data to generate fused training data and adding the fused training data as fictitious samples to the training data to generate updated training data for training a neural network or machine learning model. In one aspect, the electronic device may train a single network with only cross-entropy loss.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1, 2, 11-17, 22-26, and 29-30 are rejected under 35 U.S.C 103 as being unpatentable over Xu et al. (“Robust and Generalizable Visual Representation Learning Via Random Convolutions” hereinafter, Xu) in view of Lad et al. (US20220351373A1 hereinafter, Lad) in further view of Ghiasi et al. (“Exploring the structure of a real-time, arbitrary neural artistic stylization network”, Ghiasi).
Regarding claim 1:
Xu teaches augmenting training data, augment, via a random style generator having at least one randomly initialized layer, training data to generate augmented training data (see pg. 1 section Abstract: “In this work, we show that the robustness of neural networks can be greatly improved through the use of random convolutions as data augmentation. Random convolutions are approximately shape-preserving and may distort local textures. Intuitively, randomized convolutions create an infinite number of new domains with similar global shapes but random local texture. Therefore, we explore using outputs of multi-scale random convolutions as new images or mixing them with the original images during training.”),
aggregate data with a plurality of styles from the augmented training data to generate aggregated training data (see pg. 1 abstract: “Intuitively, randomized convolutions create an infinite number of new domains with similar global shapes but random local texture. Therefore, we explore using outputs of multi-scale random convolutions as new images or mixing them with the original images during training.”. Also see pg. 2 fig. 2),
apply semantic-aware style fusion to the aggregated training data to generate fused training data (see pg. 1 abstract: “Intuitively, randomized convolutions create an infinite number of new domains with similar global shapes but random local texture. Therefore, we explore using outputs of multi-scale random convolutions as new images or mixing them with the original images during training.”. Also see pg. 2 fig. 2), and
add the fused training data as fictitious samples to the training data to generate updated training data for training a neural network (see pg. 1 abstract: “Intuitively, randomized convolutions create an infinite number of new domains with similar global shapes but random local texture. Therefore, we explore using outputs of multi-scale random convolutions as new images or mixing them with the original images during training.”. Also see pg. 2 section 1: “We validate RandConv and its mixing variant in extensive experiments on synthetic and real world benchmarks as well as on the large-scale ImageNet dataset.”.).
Xu does not explicitly teach an apparatus with least one memory and at least one processor coupled to at least one memory.
Lad, however, analogously teaches an apparatus with least one memory and at least one processor coupled to at least one memory (see para [0003]: “Disclosed herein is a system using machine learning in detecting geographic atrophy (GA), the system comprising at least one processor; a memory; and a computing platform including the at least one processor and the memory.”. Also see para [0176]: “One promising direction is to inject model-dependent perturbations to the input images as strategic augmentations”.).
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Xu and Lad before him or her, to modify the apparatus of claim 1 to include attributes of having an apparatus with least one memory and at least one processor coupled to at least one memory in order to perform the steps of the method with a computer (see Lad para [0005]: Disclosed herein is a non-transitory computer readable medium comprising computer executable instructions that when executed by at least one processor of a computer cause the computer to perform steps according to a disclosed method and/or in a disclosed system.).
Xu does not explicitly teach wherein to apply the semantic-aware style fusion, the at least one processor is configured to extract semantic regions from the aggregated training data, combine the semantic regions into a common semantic region, and generate the fused training data from the combined semantic regions.
Ghiasi, however, analogously teaches wherein to apply the semantic-aware style fusion, the at least one processor is configured to extract semantic regions from the aggregated training data (see pg. 9 fig. 6 (B): “B: Same as previous but with Painting by Numbers dataset across for 3768 paintings across 20 labeled artists. Note the zoom-in highlighting a localized region of embedding space representing Monet paintings. Please zoom-in for details.”)
combine the semantic regions into a common semantic region (see pg. 9 fig. 6 (B): “B: Same as previous but with Painting by Numbers dataset across for 3768 paintings across 20 labeled artists. Note the zoom-in highlighting a localized region of embedding space representing Monet paintings. Please zoom-in for details.”. Also see pg. 7 section 3.4: “One may identify that regions of the embedding space cluster around perceptually similar visual textures: the bottom-right contains a preponderance of waffles; the middle contains many checkerboard patterns; top-center contains many zebra-like patterns”)
and
generate the fused training data from the combined semantic regions (see figs. 3, 5, and 7).
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Xu, Lad, and Ghiasi before him or her, to modify the apparatus of claim 1 to include attributes of wherein to apply the semantic-aware style fusion, the at least one processor is configured to extract semantic regions from the aggregated training data, combine the semantic regions into a common semantic region, and generate the fused training data from the combined semantic regions in order to effectively dial in the weight of a given style (see Ghiasi at pg. 9 section 4: “In addition, we introduce a new form of interpolation that allows a user to arbitrarily to dial in the strength of an artistic stylization.”).
Regarding claim 2:
Xu in view of Lad in further view of Ghiasi teaches the apparatus of claim 1.
Xu further teaches wherein the training data includes a plurality of training images (see pg. 6 section 4.1: “To show RandConv is more than a trivial color/contrast adjustment method, we also compare to ColorJitter2 data augmentation (which randomly changes image brightness, contrast, and saturation) and GreyScale (where images are transformed to grey scale for training and testing)” .)
Regarding claim 26:
Claim 26 recites analogous limitations to claim 2 and therefore is rejected on the same grounds.
Regarding claim 11:
Xu in view of Lad in further view of Ghiasi teaches the apparatus of claim 1.
Xu further teaches wherein, to augment, via the random style generator, training data to generate augmented training data, the at least one processor is configured to randomly initialize at least one weight and at least one offset to achieve texture modification of the training data (see pg. 4 section 3.2: “A simple approach is to use the randomized convolution layer outputs, I ∗ Θ, as new images; where Θ are the randomly sampled weights and I is a training image.”. Also see pg. 4 algorithm 1 line 8.)
Regarding claim 12:
Xu in view of Lad in further view of Ghiasi teaches the apparatus of claim 1.
Xu further teaches wherein, to augment, via the random style generator, training data to generate augmented training data, the at least one processor is configured to preserve semantic data in the training data while distorting non-semantic data to increase data diversity (see pg.1 abstract : “Random convolutions are approximately shape-preserving and may distort local textures.”. Also see pg. 13 section A: “For example, in Fig.4, the left occluded triangle shape has texture composed by shapes of cobble stones while cobble stones have their own texture. Random convolution can preserve those large shapes that usually define the image semantics while distorting the small shapes as local texture.”)
Regarding claim 13:
Xu in view of Lad in further view of Ghiasi teaches the apparatus of claim 1.
Xu further teaches wherein the augmented training data comprises a randomly generated new style from the training data but maintains data semantics (see pg. 2 section 1: “We provide insights and justification on why RandConv augments images with different local texture but the same semantics with the shape-preserving property of random convolutions.”).
Regarding claim 14:
Xu in view of Lad in further view of Ghiasi teaches the apparatus of claim 1.
Xu further teaches wherein, to aggregate the data with the plurality of styles from the augmented training data to generate the aggregated training data, the at least one processor is configured to use random style aggregation in which the plurality of styles is selected randomly (see pg. 2 fig. 2. See pg. 5 section 3.1: “Inspired by the AugMix (Hendrycks et al., 2020b) strategy, we propose to blend the original image with the outputs of the RandConv layer via linear convex combinations αI +(1−α)(I ∗Θ), where α is the mixing weight uniformly sampled from [0,1].In RCmix, the RandConv outputs provide shape-consistent perturbations of the original images. Varying α, we continuously interpolate between the training domain and the randomly sampled domains of RCimg.”)
Regarding claim 15:
Xu in view of Lad in further view of Ghiasi teaches the apparatus of claim 1.
Xu further teaches wherein the at least one processor is configured to generate the plurality of styles by passing the augmented training data through the random style generator (see pg. 2 fig. 2. See pg. 5 section 3.1: “Inspired by the AugMix (Hendrycks et al., 2020b) strategy, we propose to blend the original image with the outputs of the RandConv layer via linear convex combinations αI +(1−α)(I ∗Θ), where α is the mixing weight uniformly sampled from [0,1].In RCmix, the RandConv outputs provide shape-consistent perturbations of the original images. Varying α, we continuously interpolate between the training domain and the randomly sampled domains of RCimg.”)
Regarding claim 16:
Xu in view of Lad in further view of Ghiasi teaches the apparatus of claim 1.
Xu does not explicitly teach wherein, to aggregate data with a plurality of styles from the augmented training data to generate the aggregated training data, the at least one processor is further configured to pass a latest set of augmented training data through the random style generator.
Lad, however, analogously teaches wherein, to aggregate data with a plurality of styles from the augmented training data to generate the aggregated training data, the at least one processor is configured to pass a latest set of augmented training data through the random style generator (see para [0102]: “ To leverage this dataset to improve the model performance on our GA prediction task, we jointly trained our model with our GA cohort and the additional OCT data (Kermany D S, et al. (2018) Cell. 172(5):1122-1131) in a multi-task fashion by sharing the CNN image feature extractor but using separate FC layers for different tasks, i.e., one for GA (current and next year) and the other for CNV, DME, drusen and control prediction.”).
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Xu, Lad, and Ghiasi before him or her, to modify the apparatus of claim 16 to include attributes of wherein, to aggregate data with a plurality of styles from the augmented training data to generate the aggregated training data, the at least one processor is configured to pass a latest set of augmented training data through the random style generator in order to achieve a deeper understanding of the performance characteristics of the proposed model (see Lad para [0107]: “To achieve a deeper understanding of the performance characteristics of the proposed model (multi-scan position-aware model trained with PPI from Table 1), the cross-validated confusion matrices for GA diagnosis (current year) and prognosis (next year) were calculated and shown in Table 2.”).
Regarding claim 17:
Xu in view of Lad in further view of Ghiasi teaches the apparatus of claim 1.
Xu does not explicitly teach wherein, to apply the semantic-aware style fusion to the aggregated training data to generate the fused training data, the at least one processor is configured to apply the semantic-aware style fusion to the training data to generate the fused training data.
Ghiasi, however, analogously teaches wherein, to apply the semantic-aware style fusion to the aggregated training data to generate the fused training data, the at least one processor is configured to apply the semantic-aware style fusion to the training data to generate the fused training data (see pg. 8 fig 5: “Training on a large corpus of paintings is critical for generalization. A. Distribution of style and content loss for stylizations applied to unseen painting styles for proposed method trained on increasing numbers of painting styles. … Three sample pairs of content and style images (B) and the resulting stylization with the proposed method as the method is trained on increasing numbers of paintings (top number). For comparison, final column in (B) highlights stylizations for a model trained explicitly on the these styles”)
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Xu, Lad, and Ghiasi before him or her, to modify the apparatus of claim 17 to include attributes of wherein, to apply the semantic-aware style fusion to the aggregated training data to generate the fused training data, the at least one processor is configured to apply the semantic-aware style fusion to the training data to generate the fused training data in order to allow real-time stylization using any content/style image pair (see Ghiasi at pg. 1 section abstract: “In this paper, we present a method which combines the flexibility of the neural algorithm of artistic style with the speed of fast style transfer networks to allow real-time stylization using any content/style image pair. We build upon recent work leveraging conditional instance normalization for multi-style transfer networks by learning to predict the conditional instance normalization parameters directly from a style image … We demonstrate that the learned embedding space is smooth and contains a rich structure and organizes semantic information associated with paintings in an entirely unsupervised manner.”)
Regarding claim 22:
Xu in view of Lad in further view of Ghiasi teaches the apparatus of claim 1.
Xu further teaches wherein the at least one processor is configured to: combine class-specific semantic information extracted from the aggregated training data in an image space (see pg. 4 section : “The PACS dataset (Li et al., 2018b) considers 7-class classification on 4 domains: photo, art painting, cartoon, and sketch, with very different texture styles. Most recent domain generalization work studies the multi-source domain setting on PACS and uses domain labels of the training data.”. Also see pg. 4 section 4: “We test our RandConv variants RCimg1-7,p=0.5 and RCmix1-7 with and without consistency loss, and ColorJitter/GreyScale/BandPass/MultiAug data augmentation as in the digit datasets.”).
Regarding claim 23:
Xu in view of Lad in further view of Ghiasi teaches the apparatus of claim 1.
Xu further teaches wherein the at least one processor is configured to: train the neural network using the updated training data (see pg. 1 abstract: “Therefore, we explore using outputs of multi-scale random convolutions as new images or mixing them with the original images during training.”)
Regarding claim 24:
Xu in view of Lad in further view of Ghiasi teaches the apparatus of claim 1.
Xu does not explicitly teach wherein the at least one processor is configured to: train the neural network using a cross-entropy loss.
Lad, however, analogously teaches wherein the at least one processor is configured to: train the neural network using a cross-entropy loss (see para [0075] : “The model is trained to maximize the weighted binary cross-entropy loss, i.e., the likelihood that scans from SD-OCT inputs are correctly assigned (prognosticated) to either the GA or control groups in the assessment of the upcoming year, while adversarially encouraging that regions masked-out by the attention maps are not informative of GA. Additionally, the model concurrently predicts (diagnoses) the probability of GA in the current year.”.).
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Xu, Lad, and Ghaisi before him or her, to modify the apparatus of claim 24 to include attributes of training the neural network using a cross-entropy loss in order to measure the likelihood of correctly assigned data (see Lad para [0075]: “The model is trained to maximize the weighted binary cross-entropy loss, i.e., the likelihood that scans from SD-OCT inputs are correctly assigned (prognosticated)”).
Regarding claim 25:
Xu teaches augmenting training data, augment, via a random style generator having at least one randomly initialized layer, training data to generate augmented training data (see pg. 1 section Abstract: “In this work, we show that the robustness of neural networks can be greatly improved through the use of random convolutions as data augmentation. Random convolutions are approximately shape-preserving and may distort local textures. Intuitively, randomized convolutions create an infinite number of new domains with similar global shapes but random local texture. Therefore, we explore using outputs of multi-scale random convolutions as new images or mixing them with the original images during training.”),
aggregate data with a plurality of styles from the augmented training data to generate aggregated training data (see pg. 1 abstract: “Intuitively, randomized convolutions create an infinite number of new domains with similar global shapes but random local texture. Therefore, we explore using outputs of multi-scale random convolutions as new images or mixing them with the original images during training.”. Also see pg. 2 fig. 2),
applying semantic-aware style fusion to the aggregated training data to generate fused training data (see pg. 1 abstract: “Intuitively, randomized convolutions create an infinite number of new domains with similar global shapes but random local texture. Therefore, we explore using outputs of multi-scale random convolutions as new images or mixing them with the original images during training.”. Also see pg. 2 fig. 2), and
adding the fused training data as fictitious samples to the training data to generate updated training data for training a neural network (see pg. 1 abstract: “Intuitively, randomized convolutions create an infinite number of new domains with similar global shapes but random local texture. Therefore, we explore using outputs of multi-scale random convolutions as new images or mixing them with the original images during training.”. Also see pg. 2 section 1: “We validate RandConv and its mixing variant in extensive experiments on synthetic and real world benchmarks as well as on the large-scale ImageNet dataset.”.).
Xu does not explicitly teach an apparatus with least one memory and at least one processor coupled to at least one memory.
Lad, however, analogously teaches a processor-implemented method of augmenting data (see para [0003]: “Disclosed herein is a system using machine learning in detecting geographic atrophy (GA), the system comprising at least one processor; a memory; and a computing platform including the at least one processor and the memory.”. Also see para [0176]: “One promising direction is to inject model-dependent perturbations to the input images as strategic augmentations”.).
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Xu and Lad before him or her, to modify the method of claim 25 to include attributes of a processor-implemented method of augmenting data in order to perform the steps of the method with a computer (see Lad para [0005]: Disclosed herein is a non-transitory computer readable medium comprising computer executable instructions that when executed by at least one processor of a computer cause the computer to perform steps according to a disclosed method and/or in a disclosed system.).
Xu does not explicitly teach wherein to apply the semantic-aware style fusion, the at least one processor is configured to extract semantic regions from the aggregated training data, combine the semantic regions into a common semantic region, and generate the fused training data from the combined semantic regions.
Ghiasi, however, analogously teaches wherein to apply the semantic-aware style fusion, the at least one processor is configured to extract semantic regions from the aggregated training data (see pg. 9 fig. 6 (B): “B: Same as previous but with Painting by Numbers dataset across for 3768 paintings across 20 labeled artists. Note the zoom-in highlighting a localized region of embedding space representing Monet paintings. Please zoom-in for details.”)
combine the semantic regions into a common semantic region (see pg. 9 fig. 6 (B): “B: Same as previous but with Painting by Numbers dataset across for 3768 paintings across 20 labeled artists. Note the zoom-in highlighting a localized region of embedding space representing Monet paintings. Please zoom-in for details.”) and
generate the fused training data from the combined semantic regions (see fig. 7 and fig. 3).
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Xu, Lad, and Ghiasi before him or her, to modify the method of claim 25 to include attributes of wherein to apply the semantic-aware style fusion, the at least one processor is configured to extract semantic regions from the aggregated training data, combine the semantic regions into a common semantic region, and generate the fused training data from the combined semantic regions in order to effectively dial in the weight of a given style (see Ghiasi at pg. 9 section 4: “In addition, we introduce a new form of interpolation that allows a user to arbitrarily to dial in the strength of an artistic stylization.”).
Regarding claim 29:
Xu teaches augmenting training data, augment, via a random style generator having at least one randomly initialized layer, training data to generate augmented training data (see pg. 1 section Abstract: “In this work, we show that the robustness of neural networks can be greatly improved through the use of random convolutions as data augmentation. Random convolutions are approximately shape-preserving and may distort local textures. Intuitively, randomized convolutions create an infinite number of new domains with similar global shapes but random local texture. Therefore, we explore using outputs of multi-scale random convolutions as new images or mixing them with the original images during training.”),
aggregate data with a plurality of styles from the augmented training data to generate aggregated training data (see pg. 1 abstract: “Intuitively, randomized convolutions create an infinite number of new domains with similar global shapes but random local texture. Therefore, we explore using outputs of multi-scale random convolutions as new images or mixing them with the original images during training.”. Also see pg. 2 fig. 2),
apply semantic-aware style fusion to the aggregated training data to generate fused training data (see pg. 1 abstract: “Intuitively, randomized convolutions create an infinite number of new domains with similar global shapes but random local texture. Therefore, we explore using outputs of multi-scale random convolutions as new images or mixing them with the original images during training.”. Also see pg. 2 fig. 2), and
add the fused training data as fictitious samples to the training data to generate updated training data for training a neural network (see pg. 1 abstract: “Intuitively, randomized convolutions create an infinite number of new domains with similar global shapes but random local texture. Therefore, we explore using outputs of multi-scale random convolutions as new images or mixing them with the original images during training.”. Also see pg. 2 section 1: “We validate RandConv and its mixing variant in extensive experiments on synthetic and real world benchmarks as well as on the large-scale ImageNet dataset.”.).
Xu does not explicitly teach an apparatus with least one memory and at least one processor coupled to at least one memory.
Lad, however, analogously teaches having a non-transitory computer-readable storage medium storing instructions executed with one or more processors (see para [0005]: “Disclosed herein is a non-transitory computer readable medium comprising computer executable instructions that when executed by at least one processor of a computer cause the computer to perform steps according to a disclosed method and/or in a disclosed system..”. Also see para [0176]: “One promising direction is to inject model-dependent perturbations to the input images as strategic augmentations”.).
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Xu and Lad before him or her, to modify the apparatus of claim 29 to include attributes of having a computer-readable storage medium storing instructions executed with one or more processors in order to perform the steps of the method with a computer (see Lad para [0005]: Disclosed herein is a non-transitory computer readable medium comprising computer executable instructions that when executed by at least one processor of a computer cause the computer to perform steps according to a disclosed method and/or in a disclosed system.).
Xu does not explicitly teach wherein to apply the semantic-aware style fusion, the at least one processor is configured to extract semantic regions from the aggregated training data, combine the semantic regions into a common semantic region, and generate the fused training data from the combined semantic regions.
Ghiasi, however, analogously teaches wherein to apply the semantic-aware style fusion, the at least one processor is configured to extract semantic regions from the aggregated training data (see pg. 9 fig. 6 (B): “B: Same as previous but with Painting by Numbers dataset across for 3768 paintings across 20 labeled artists. Note the zoom-in highlighting a localized region of embedding space representing Monet paintings. Please zoom-in for details.”)
combine the semantic regions into a common semantic region (see pg. 9 fig. 6 (B): “B: Same as previous but with Painting by Numbers dataset across for 3768 paintings across 20 labeled artists. Note the zoom-in highlighting a localized region of embedding space representing Monet paintings. Please zoom-in for details.” Also see pg. 7 section 3.4: “One may identify that regions of the embedding space cluster around perceptually similar visual textures: the bottom-right contains a preponderance of waffles; the middle contains many checkerboard patterns; top-center contains many zebra-like patterns”) and
generate the fused training data from the combined semantic regions (see figs. 3, 5, and 7).
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Xu, Lad, and Ghiasi before him or her, to modify the non-transitory computer-readable medium of claim 29 to include attributes of wherein to apply the semantic-aware style fusion, the at least one processor is configured to extract semantic regions from the aggregated training data, combine the semantic regions into a common semantic region, and generate the fused training data from the combined semantic regions in order to effectively dial in the weight of a given style (see Ghiasi at pg. 9 section 4: “In addition, we introduce a new form of interpolation that allows a user to arbitrarily to dial in the strength of an artistic stylization.”).
Regarding claim 30:
Xu teaches means for augmenting training data, augment, via a random style generator having at least one randomly initialized layer, training data to generate augmented training data (see pg. 1 section Abstract: “In this work, we show that the robustness of neural networks can be greatly improved through the use of random convolutions as data augmentation. Random convolutions are approximately shape-preserving and may distort local textures. Intuitively, randomized convolutions create an infinite number of new domains with similar global shapes but random local texture. Therefore, we explore using outputs of multi-scale random convolutions as new images or mixing them with the original images during training.”),
means for aggregating data with a plurality of styles from the augmented training data to generate aggregated training data (see pg. 1 abstract: “Intuitively, randomized convolutions create an infinite number of new domains with similar global shapes but random local texture. Therefore, we explore using outputs of multi-scale random convolutions as new images or mixing them with the original images during training.”. Also see pg. 2 fig. 2),
means for applying semantic-aware style fusion to the aggregated training data to generate fused training data (see pg. 1 abstract: “Intuitively, randomized convolutions create an infinite number of new domains with similar global shapes but random local texture. Therefore, we explore using outputs of multi-scale random convolutions as new images or mixing them with the original images during training.”. Also see pg. 2 fig. 2), and
means for adding the fused training data as fictitious samples to the training data to generate updated training data for training a neural network (see pg. 1 abstract: “Intuitively, randomized convolutions create an infinite number of new domains with similar global shapes but random local texture. Therefore, we explore using outputs of multi-scale random convolutions as new images or mixing them with the original images during training.”. Also see pg. 2 section 1: “We validate RandConv and its mixing variant in extensive experiments on synthetic and real world benchmarks as well as on the large-scale ImageNet dataset.”.).
Xu does not explicitly teach an apparatus.
Lad, however, analogously teaches an apparatus (see para [0003]: “Disclosed herein is a system using machine learning in detecting geographic atrophy (GA), the system comprising at least one processor; a memory; and a computing platform including the at least one processor and the memory.”.)
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Xu and Lad before him or her, to modify the apparatus of claim 30 to include attributes of having an apparatus in order to perform the steps of the method with a computer (see Lad para [0005]: Disclosed herein is a non-transitory computer readable medium comprising computer executable instructions that when executed by at least one processor of a computer cause the computer to perform steps according to a disclosed method and/or in a disclosed system.)
Xu does not explicitly teach wherein to apply the semantic-aware style fusion, the at least one processor is configured to extract semantic regions from the aggregated training data, combine the semantic regions into a common semantic region, and generate the fused training data from the combined semantic regions.
Ghiasi, however, analogously teaches wherein to apply the semantic-aware style fusion, the at least one processor is configured to extract semantic regions from the aggregated training data (see pg. 9 fig. 6 (B): “B: Same as previous but with Painting by Numbers dataset across for 3768 paintings across 20 labeled artists. Note the zoom-in highlighting a localized region of embedding space representing Monet paintings. Please zoom-in for details.”)
combine the semantic regions into a common semantic region (see pg. 9 fig. 6 (B): “B: Same as previous but with Painting by Numbers dataset across for 3768 paintings across 20 labeled artists. Note the zoom-in highlighting a localized region of embedding space representing Monet paintings. Please zoom-in for details.” Also see pg. 7 section 3.4: “One may identify that regions of the embedding space cluster around perceptually similar visual textures: the bottom-right contains a preponderance of waffles; the middle contains many checkerboard patterns; top-center contains many zebra-like patterns”) and
generate the fused training data from the combined semantic regions (see figs. 3, 5, and 7).
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Xu, Lad, and Ghiasi before him or her, to modify the method of claim 30 to include attributes of wherein to apply the semantic-aware style fusion, the at least one processor is configured to extract semantic regions from the aggregated training data, combine the semantic regions into a common semantic region, and generate the fused training data from the combined semantic regions in order to effectively dial in the weight of a given style (see Ghiasi at pg. 9 section 4: “In addition, we introduce a new form of interpolation that allows a user to arbitrarily to dial in the strength of an artistic stylization.”).
Claims 3-10 and 27-28 are rejected under 35 U.S.C 103 as being unpatentable over Xu et al. (“Robust and Generalizable Visual Representation Learning Via Random Convolutions” hereinafter, Xu) in view of Lad et al. (US20220351373A1, hereinafter Lad) in further view of Ghiasi et al. (“Exploring the structure of a real-time, arbitrary neural artistic stylization network”, Ghiasi) in further view of Dai et al. (“Deformable Convolutional Networks” hereinafter, Dai).
Regarding claim 3:
Xu in view of Lad in further view of Ghiasi teaches the apparatus of claim 2.
Xu does not explicitly teach wherein a size of at least one kernel of the neural network is based on a size of an image of the plurality of training images.
Dai, however, analogously teaches wherein a size of at least one kernel of the neural network is based on a size of an image of the plurality of training images (see pg. 765-766 section 2.1: “As illustrated in Figure 2, the offsets are obtained by applying a convolutional layer over the same input feature map. The convolution kernel is of the same spatial resolution and dilation as those of the current convolutional layer (e.g., also 3 × 3 with dilation 1 in Figure 2).”. Also see pg. 767 section 2.3 : “First, a deep fully convolutional network generates feature maps over the whole input image.”).
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Xu, Lad, Ghiasi, and Dai before him or her, to modify the apparatus of claim 3 to include attributes of wherein a size of at least one kernel of the neural network is based on a size of an image of the plurality of training images in order to be defined within a receptive field size and dilation (see Dai pg. 765 section 2.1: “The grid R defines the receptive field size and dilation.”).
Regarding claim 27:
Claim 27 recites analogous limitations to claim 3 and therefore is rejected on the same grounds.
Regarding claim 4:
Xu in view of Lad in further view of Ghiasi teaches the apparatus of claim 2.
Xu further teaches wherein the at least one processor is configured to: augment texture data, contrast data, and brightness data of the plurality of training images (see pg. 6 section 4.1 results: “To show RandConv is more than a trivial color/contrast adjustment method, we also compare to ColorJitter2 data augmentation (which randomly changes image brightness, contrast, and saturation) and GreyScale (where images are transformed to grey scale for training and testing). We also tested data augmentation with a fixed Laplacian of Gaussian filter (Band-Pass) of size=3 and σ = 1 and the data augmentation pipeline (Multi-Aug) that was used in a recently proposed large scale study on domain generalization algorithms and datasets (Gulrajani &Lopez-Paz, 2020).”. Also see pg. 9 section 5: “Randomized convolution (RandConv) is a simple but powerful data augmentation technique for randomizing local image texture” )
Regarding claim 28:
Claim 28 recites analogous limitations to claim 3 and therefore is rejected on the same grounds.
Regarding claim 5:
Xu in view of Lad in further view of Ghiasi teaches the apparatus of claim 1.
Xu further teaches wherein, to augment the training data, the at least one processor is configured to randomly initialize a brightness parameter and a contrast parameter in (see pg. 6 section 4.1: “To show RandConv is more than a trivial color/contrast adjustment method, we also compare to ColorJitter2 data augmentation (which randomly changes image brightness, contrast, and saturation)”).
Xu does not explicitly teach an affine transformation.
Dai, however, analogously teaches an affine transformation (see pg. 764 section 1: “This is usually realized by augmenting the existing data samples, e.g., by affine transformation.”).
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Xu, Lad, Ghiasi and Dai before him or her, to modify the apparatus of claim 5 to include attributes of an affine transformation in order to be build training datasets with sufficient desired variations (see Dai pg. 764 section 1: “The first is to build the training datasets with sufficient desired variations. This is usually realized by augmenting the existing data samples,”).
Regarding claim 6:
Xu in view of Lad in further view of Ghiasi teaches the apparatus of claim 1.
Xu does not explicitly teach wherein, to augment the training data, the at least one processor is configured to perform deformable convolution, apply a random convolutional layer, and apply a deformable convolutional layer.
Dai, however, analogously teaches wherein, to augment the training data, the at least one processor is configured to perform deformable convolution, apply a random convolutional layer, and apply a deformable convolutional layer (see pg. 764 section 1: “In this work, we introduce two new modules that greatly enhance CNNs’ capability of modeling geometric transformations. The first is deformable convolution.”. Also see pg. 767 section 2.3: “A randomly initialized 1 × 1 convolution is added at last to reduce the channel dimension to 1024”. Also see pg. 770 tables 1 and 2.)
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Xu, Lad, Ghiasi, and Dai before him or her, to modify the apparatus of claim 6 to include attributes of to augment the training data, the at least one processor is configured to perform deformable convolution, apply a random convolutional layer, and apply a deformable convolutional layer in order to greatly enhance CNNs’ capability of modeling geometric transformations (see Dai pg. 764 section 1: “In this work, we introduce two new modules that greatly enhance CNNs’ capability of modeling geometric transformations).
Regarding claim 7:
Xu in view of Lad in further view of Ghiasi teaches the apparatus of claim 1.
Xu does not explicitly teach wherein, to augment the training data, the at least one processor is configured to augment texture data in the training data using a randomly initialized deformable convolution layer.
Dai, however, analogously teaches wherein, to augment the training data, the at least one processor is configured to augment texture data in the training data using a randomly initialized deformable convolution layer (see pg. 767 section 2.3 ‘Deformable Convolution for Feature Extraction’: “A randomly initialized 1 × 1 convolution is added at last to reduce the channel dimension to 1024.”. Also see pg. 768 figures 6 and 7.).
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Xu, Lad, Ghiasi and Dai before him or her, to modify the apparatus of claim 7 to include attributes of wherein, to augment the training data, the at least one processor is configured to augment texture data in the training data using a randomly initialized deformable convolution layer in order to greatly enhance CNNs’ capability of modeling geometric transformations (see Dai pg. 764 section 1: “In this work, we introduce two new modules that greatly enhance CNNs’ capability of modeling geometric transformations).
Regarding claim 8:
Xu in view of Lad in further view of Ghiasi in further view of Dai teaches the apparatus of claim 7.
Xu does not explicitly teach wherein one or more of weights and offsets are randomly initialized using the randomly initialized deformable convolution layer.
Dai, however, analogously teaches wherein one or more of weights and offsets are randomly initialized using the randomly initialized deformable convolution layer (see pg. 765 section 2.1 : “The 2D convolution consists of two steps: 1) sampling using a regular grid R over the input feature map x; 2) summation of sampled values weighted by w.”. Also see pg. 767 section 3: “This work is built on the idea of augmenting the spatial sampling locations in convolution and RoI pooling with additional offsets and learning the offsets from target tasks.”. Also see pg. 767 section 2.3: “A randomly initialized 1 × 1 convolution is added at last to reduce the channel dimension to 1024.”. )
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Xu, Lad, Ghiasi, and Dai before him or her, to modify the apparatus of claim 8 to include attributes of wherein, to augment the training data, the at least one processor is configured to augment texture data in the training data using a randomly initialized deformable convolution layer in order to greatly enhance CNNs’ capability of modeling geometric transformations (see Dai pg. 764 section 1: “In this work, we introduce two new modules that greatly enhance CNNs’ capability of modeling geometric transformations).
Regarding claim 9:
Xu in view of Lad in further view of Ghiasi in further view of Dai teaches the apparatus of claim 8.
Xu does not explicitly teach wherein, to augment the training data, the at least one processor is configured to augment contrast data in the training data and brightness data in the training data using instance normalization, affine transformation, and a sigmoid function.
Dai, however, analogously teaches wherein, to augment the training data, the at least one processor is configured to augment contrast data in the training data and brightness data in the training data using instance normalization, affine transformation, (see pg. 766 section 2.2: “Figure 3 illustrates how to obtain the offsets. Firstly, RoI pooling (Eq. (5)) generates the pooled feature maps. From the maps, a fc layer generates the normalized offsets ∆bpij, which are then transformed to the offsets ∆pij in Eq. (6) by element-wise product with the RoI’s width and height, as ∆pij = γ·∆bpij ◦(w,h). Here γ is a pre-defined scalar to modulate the magnitude of the offsets. It is empirically set to γ = 0.1. The offset normalization is necessary to make the offset learning invariant to RoI size.”. Also see pg. 764 section 1: “This is usually realized by augmenting the existing data samples, e.g., by affine transformation.”).
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Xu, Lad, Ghiasi, and Dai before him or her, to modify the apparatus of claim 9 to include attributes of wherein, to augment the training data, the at least one processor is configured to augment contrast data in the training data and brightness data in the training data using instance normalization, and affine transformation in order to greatly enhance CNNs’ capability of modeling geometric transformations in order to build training datasets with sufficient desired variations (see Dai pg. 764 section 1: “The first is to build the training datasets with sufficient desired variations. This is usually realized by augmenting the existing data samples,”).
Xu does not explicitly teach a sigmoid function to augment the training data.
Lad, however, analogously teaches a sigmoid function to augment the training data (see para [0141]: “To remove causal information from
PNG
media_image1.png
29
38
media_image1.png
Greyscale
and obtain a negative contrast
PNG
media_image2.png
33
33
media_image2.png
Greyscale
we apply the following soft-masking transformation … Specifically, we use the thresholded sigmoid as the masking function”.).
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Xu, Lad, Ghiasi and Dai before him or her, to modify the apparatus of claim 9 to include attributes of a sigmoid function to augment the training data in order to remove causal information and obtain a negative contrast (see Lad para [0141]: “To remove causal information from
PNG
media_image1.png
29
38
media_image1.png
Greyscale
and obtain a negative contrast
PNG
media_image2.png
33
33
media_image2.png
Greyscale
we apply the following soft-masking transformation … Specifically, we use the thresholded sigmoid as the masking function”.).
Regarding claim 10:
Xu in view of Lad in further view of Dai teaches the apparatus of claim 9.
Xu does not explicitly teach wherein at least one parameter of the affine transformation is randomly initialized.
Dai, however, analogously teaches wherein at least one parameter of the affine transformation is randomly initialized (see pg. 767 section 2.3: “A randomly initialized 1 × 1 convolution is added at last to reduce the channel dimension to 1024.”. Also see pg. 764 “This is usually realized by augmenting the existing data samples, e.g., by affine transformation.”).
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Xu, Lad, Ghiasi and Dai before him or her, to modify the apparatus of claim 10 to include attributes of wherein at least one parameter of the affine transformation is randomly initialized in order to remove causal information and obtain a negative contrast in order to greatly enhance CNNs’ capability of modeling geometric transformations (see Dai pg. 764 section 1: “In this work, we introduce two new modules that greatly enhance CNNs’ capability of modeling geometric transformations).
Claims 18-21 are rejected under 35 U.S.C 103 as being unpatentable over Xu et al. (“Robust and Generalizable Visual Representation Learning Via Random Convolutions” hereinafter, Xu) in view of Lad et al. (US20220351373A1, hereinafter Lad) in further view of Ghiasi et al. (“Exploring the structure of a real-time, arbitrary neural artistic stylization network”, Ghiasi) in further view of Park et al. (“Semantic-aware Neural Style Transfer” hereinafter, Park).
Regarding claim 18:
Xu in view of Lad in further view of Ghiasi teaches the apparatus of claim 1.
Xu does not explicitly teach wherein, to apply the semantic-aware style fusion to the aggregated training data to generate the fused training data, the at least one processor is configured to extract semantic regions from the training data and the augmented training data, wherein the semantic regions are used in the semantic-aware style fusion with the training data.
Park, however, analogously teaches wherein, to apply the semantic-aware style fusion to the aggregated training data to generate the fused training data, the at least one processor is configured to extract semantic regions from the training data and the augmented training data, wherein the semantic regions are used in the semantic-aware style fusion with the training data (see pg. 13 abstract: “Here, each image is partitioned into several semantic regions for both a target photograph and a source painting. All partitioned regions of the target are then associated with one of the partitioned regions in the source according to their semantic interpretation. Given a pair of target and source regions, style is learned from the source region whereas content is learned from the target region. By integrating both the style and content components, we can successfully generate a stylized output”)
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Xu, Lad, Ghiasi and Park before him or her, to modify the apparatus of claim 18 to include attributes of wherein, to apply the semantic-aware style fusion to the aggregated training data to generate the fused training data, the at least one processor is configured to extract semantic regions from the training data and the augmented training data, wherein the semantic regions are used in the semantic-aware style fusion with the training data in order to successfully generate a stylized output (see Park pg. 13 abstract: “By integrating both the style and content components, we can successfully generate a stylized output”).
Regarding claim 19:
Xu in view of Lad in further view of Ghiasi in further view of Park teaches the apparatus of claim 18.
Xu does not explicitly teach wherein, to apply the semantic-aware style fusion to the aggregated training data to generate the fused training data, the at least one processor is configured to process a common semantic region with the training data and the augmented training data to generate common semantic region data.
Park, however, analogously teaches wherein, to apply the semantic-aware style fusion to the aggregated training data to generate the fused training data, the at least one processor is configured to process a common semantic region with the training data and the augmented training data to generate common semantic region data (see pg. 18 fig. 2. Also see pg. 13 abstract: “As the primary focus of this study, the consideration of semantic matching is expected to improve the quality of artistic style transfer. Here, each image is partitioned into several semantic regions for both a target photograph and a source painting. All partitioned regions of the target are then associated with one of the partitioned regions in the source according to their semantic interpretation. Given a pair of target and source regions, style is learned from the source region whereas content is learned from the target region. By integrating both the style and content components, we can successfully generate a stylized output. Unlike previous approaches, we obtain the best semantic match between regions using word embeddings. Thus, we guarantee that semantic matching is always established between the target and source.”.)
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Xu, Lad, Ghiasi, and Park before him or her, to modify the apparatus of claim 19 to include attributes of wherein, to apply the semantic-aware style fusion to the aggregated training data to generate the fused training data, the at least one processor is configured to process a common semantic region with the training data and the augmented training data to generate common semantic region data in order to successfully generate a stylized output (see Park pg. 13 abstract: “By integrating both the style and content components, we can successfully generate a stylized output”).
Regarding claim 20:
Xu in view of Lad in further view of Ghiasi in further view of Park teaches the apparatus of claim 19.
Xu further teaches wherein, to apply the semantic-aware style fusion to the aggregated training data to generate the fused training data, the at least one processor is configured to process inverted data with the training data and the augmented training data to generate background data (see pg. 4 section 3.2: “Inspired by the AugMix (Hendrycks et al., 2020b) strategy, we propose to blend the original image with the outputs of the RandConv layer via linear convex combinations αI +(1−α)(I ∗Θ), where α is the mixing weight uniformly sampled from [0,1].In RCmix, the RandConv outputs provide shape-consistent perturbations of the original images.”) [(Examiner’s note: as inverted data is not mentioned in the specification, the Examiner has used broadest reasonable interpretation in line with the meaning well known in the art which is reconstructing input data using similar synthetic data)].
Regarding claim 21:
Xu in view of Lad in further view of Ghiasi in further view of Park teaches the apparatus of claim 20.
Xu does not explicitly teach wherein, to apply the semantic-aware style fusion to the aggregated training data to generate the fused training data, the at least one processor is configured to combine the common semantic region data and the background data to generate the fused training data.
Park, however, analogously teaches wherein, to apply the semantic-aware style fusion to the aggregated training data to generate the fused training data, the at least one processor is configured to combine the common semantic region data and the background data to generate the fused training data (see pg. 15 section 2.1.2.3: “A current solution resolving the semantic mismatch of a neural style transfer involves the use of semantic masks. To define semantic masks, Rujie [39] suggested the use of object segmentation by assuming that both the source and target contain the same set of objects. Xianye et al. [40] divided the image into a background and foreground and used these regions for a portrait style transfer”).
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Xu, Lad, Ghiasi, and Park before him or her, to modify the apparatus of claim 21 to include attributes of wherein, to apply the semantic-aware style fusion to the aggregated training data to generate the fused training data, the at least one processor is configured to combine the common semantic region data and the background data to generate the fused training data in order to successfully resolve semantic mismatch problems in existing algorithms (see pg. 13 abstract: “ semantic-aware style transfer method for resolving semantic mismatch problems in existing algorithms.”)
Pertinent Prior Art
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure:
“Learning to Diversify for Single Domain Generalization” — Wang et al. — discloses a style component to single-domain generalization
“An overview of mixing augmentation methods and augmentation strategies” — Lewy et al. — discloses style mixing and single domain augmentation
“Situational Fusion of Visual Representation for Visual Navigation” — Shen et al. — discloses
“Frustratingly Simple Domain Generalization via Image Stylization” — Somavarapu et al. — discloses domain generalization where model biases are more generalizable
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Andrew A Bracero whose telephone number is (571)270-0592. The examiner can normally be reached Monday - Friday 9:00a.m. - 5:00 p.m. ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached Monday - Friday 9:00a.m. - 5:00 p.m. ET at (571) 270-7519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ANDREW BRACERO/Examiner, Art Unit 2126
/DAVID YI/Supervisory Patent Examiner, Art Unit 2126