DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Specification
The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 5-8, 15, 19 are rejected under 35 U.S.C. 103 as being unpatentable over Ueki et al. (JP 2021110753 A), hereinafter Ueki, in view of Chandler (US 20150379422 A1), hereinafter Chandler.
Regarding claim 1, Ueki teaches A data processing apparatus comprising: a processor configured to: (Para. 34 see "The automatic detection device 1 includes a control unit 11, a main storage unit 12, a communication unit 13, an operation unit 14, a display panel 15, and an auxiliary storage unit 16 that control the entire device." Para. 36 see "The control unit 11 can be composed of a CPU (Central Processing Unit), a ROM (Read Only Memory), a RAM (Random Access Memory), and the like. The control unit 11 may include a GPU (Graphics Processing Unit)." Para. 39 see "The auxiliary storage unit 16 is a large-capacity memory, a hard disk, or the like, and includes programs necessary for the control unit 11 to execute processing"). generate an inverted image obtained by inverting pixel values of a single-channel image that constitutes a first labeled training dataset, (Para. 76 see "Thousands of transmitted X-ray images and images in which the void region 36 is masked (manually inspected data) are prepared." Para. 77 see "These are randomly cut out as a 256 × 256 (pixel) patch image (section 37) to make an image with increased noise, contrast adjustment, brightness adjustment, brightness inversion, smoothing, and enlargement." Para. 79 see "As shown in FIG. 6A, the division 37a is black-and-white inverted to obtain training data."). as training data of a second labeled training dataset obtained by augmentation of the first labeled training dataset, (Para. 67 see "Using the captured image (A) DB401 and the mask image (A) DB402, a machine learning model capable of detecting which region of the captured image corresponds to the void 36 or the crack 39 is created." Para. 68 see "The data padding preprocessing program 403 reads the captured image (A) DB 401 and the mask image (A) DB 402, and creates a data set for machine learning." Para. 76 see "Thousands of transmitted X-ray images and images in which the void region 36 is masked (manually inspected data) are prepared." Para. 77 see "Perform one or more processes such as shrinking, rotating, shifting horizontally or vertically, partially masking (filling in black), or a combination of multiple processes to expand the data and expand the data to hundreds of thousands. Generate a set of sheets." Para. 81 see "The feature of the present invention is that even when the amount of machine learning data that can be prepared is small, the number of data is expanded by the method of black-and-white inversion, rotation, reduction, and enlargement of the division as described above, and a large amount of data set is self-generated." Para. 82 see "using 70% of the patch image (category 37) as training data and the remaining 30% as test data, an encoder-decoder model for segmentation is used" Para. 83 see "As the die coefficient (Decee cofficient) between the mask image as the teacher data and the mask image as the network output" Examiner Note: Teacher data is the ground truth labeled data.). the single-channel image being an image of which the pixel values are determined in accordance with a physical quantity sensed by light-receiving elements during imaging or an image of which the pixel values are determined by reversible transform on the physical quantity; (Para. 27 see "X-rays transmitted through a sample are converted into light and magnified by an optical lens. X-rays have the property of penetrating substances. Part of the X-rays are absorbed as they pass through the sample. The rate of absorption increases as the density of the material is high (the atomic number is large) and the thickness is large, so that the intensity of transmitted X-rays is low." Para. 28 see "When the void 36 and the crack 39 are generated in the object, the X-ray transmittance of the void 36 and the crack 39 portion becomes large, so that the void 36 and the crack 39 are displayed in the pattern 40." Examiner Note: This shows the single-channel image being an image of which the pixel values are determined in accordance with a physical quantity sensed by light-receiving elements during imaging.).
While Ueki teaches transforming labeled images to generate a second training dataset, Ueki does not teach generating a transformed label corresponding to the transformed image, as a ground truth label that constitutes the second labeled training dataset, the transformed label being generated based on a ground truth label of the image, the ground truth label of the image constituting the first labeled training dataset.
However, Chandler teaches generating a transformed label corresponding to the transformed image, as a ground truth label that constitutes the second labeled training dataset, (Para. 10 see "A labeled dataset may be augmented by utilizing a label preserving transformation. A label preserving transformation generates new data from existing data, where the new data has the same label as the existing data." Para. 11 see "One example is a system including a training dataset with at least one training data, and a label preserving transformation including an occluder, and an inpainter. ... the dataset is augmented by adding, for the at least one training data, the transformed at least one training data to the dataset."). the transformed label being generated based on a ground truth label of the image, the ground truth label of the image constituting the first labeled training dataset. (Para. 9 see "Augmenting a labeled dataset is the process of applying a transformation function to each data in the dataset to produce additional data with the same label." Para. 10 see "A label preserving transformation generates new data from existing data, where the new data has the same label as the existing data. Reflections, linear shifts, and elastic distortions are three classes of label preserving transformations that may be generally applied to object classification analyses. These transformations build an invariance to an expected transformation in the data.").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ueki to incorporate the teachings of Chandler to generate a transformed label corresponding to the transformed image as a ground truth label for the second labeled training dataset, with that transformed label based on the original ground truth label of the image from the first labeled training dataset. Doing so would predictably improve the accuracy and robustness of the trained model by ensuring every transformed image used for training remains correctly paired with a ground truth label.
Regarding claim 5, Ueki in view of Chandler teaches the data processing apparatus according to claim 1.
In addition, Ueki teaches wherein each pixel of the single-channel image has a digital value correlated with the physical quantity at a corresponding point of a photographic subject. (Abstract see "An automatic detection device ... Images captured by instruments of an ultrasonic microscope 2 and an X-ray CT device 3 are uploaded ... The binarization of the captured image is performed by using a machine learning model trained with the approach of machine learning including deep learning and a void/crack region on the image is visualized." Para. 27 see "X-rays transmitted through a sample are converted into light and magnified by an optical lens. ... Part of the X-rays are absorbed as they pass through the sample. The rate of absorption increases as the density of the material is high (the atomic number is large) and the thickness is large, so that the intensity of transmitted X-rays is low." Para. 79 see "As shown in FIG. 6A, the division 37a is black-and-white inverted to obtain training data.").
Regarding claim 6, Ueki in view of Chandler teaches the data processing apparatus according to claim 1.
In addition, Ueki teaches wherein the first labeled training dataset is selected such that in a case where the first labeled training dataset is augmented, the inverted image does not cause any contradiction. (Para. 67 see "Using the captured image (A) DB401 and the mask image (A) DB402, a machine learning model capable of detecting which region of the captured image corresponds to the void 36 or the crack 39 is created." Para. 68 see "The data padding preprocessing program 403 reads the captured image (A) DB 401 and the mask image (A) DB 402, and creates a data set for machine learning." Para. 76 see "Thousands of transmitted X-ray images and images in which the void region 36 is masked (manually inspected data) are prepared." Para. 77 see "These are randomly cut out as a 256 × 256 (pixel) patch image (section 37) to make an image with increased noise, contrast adjustment, brightness adjustment, brightness inversion, smoothing, and enlargement. Perform one or more processes such as shrinking, rotating, shifting horizontally or vertically, partially masking (filling in black), or a combination of multiple processes to expand the data and expand the data to hundreds of thousands. Generate a set of sheets." Para. 79 see "As shown in FIG. 6A, the division 37a is black-and-white inverted to obtain training data." Para. 80 see "As shown in FIG. 6B, the division 37a shown in FIG. 7A is used as training data by rotating a predetermined angle (10 ° in the figure) to the left. As shown in FIG. 6C, the training data is obtained by rotating a predetermined angle (10 ° in the figure) to the right. Further, the division 37 may be inverted according to the line object." Para. 81 see "The feature of the present invention is that even when the amount of machine learning data that can be prepared is small, the number of data is expanded by the method of black-and-white inversion, rotation, reduction, and enlargement of the division as described above, and a large amount of data set is self-generated." Para. 82 see "using 70% of the patch image (category 37) as training data and the remaining 30% as test data, an encoder-decoder model for segmentation is used" Para. 83 see "As the die coefficient (Decee cofficient) between the mask image as the teacher data and the mask image as the network output" Examiner Note: Fig. 6 shows an inverted image without contradiction (the objects in the image are consistent).).
While Ueki teaches transforming labeled images to generate a second training dataset, Ueki does not teach the inverted label does not cause any contradiction.
However, Chandler teaches the inverted label does not cause any contradiction. (Para. 10 see "A labeled dataset may be augmented by utilizing a label preserving transformation. A label preserving transformation generates new data from existing data, where the new data has the same label as the existing data." Para. 11 see "One example is a system including a training dataset with at least one training data, and a label preserving transformation including an occluder, and an inpainter. ... the dataset is augmented by adding, for the at least one training data, the transformed at least one training data to the dataset." Examiner Note: This shows labels are preserved through the transformation process and therefore do not cause any contradiction.).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ueki and Chandler to incorporate the teachings of Chandler to select the first labeled training dataset so that when it is augmented the inverted label does not cause any contradiction. Doing so would predictably improve training stability and model accuracy by ensuring the augmented labels remain consistent and usable for learning.
Regarding claim 7, Ueki in view of Chandler teaches the data processing apparatus according to claim 1.
In addition, Ueki teaches wherein the first labeled training dataset is selected such that in a case where the first labeled training dataset is augmented, the inverted image is effective for learning. (Para. 67 see "Using the captured image (A) DB401 and the mask image (A) DB402, a machine learning model capable of detecting which region of the captured image corresponds to the void 36 or the crack 39 is created." Para. 68 see "The data padding preprocessing program 403 reads the captured image (A) DB 401 and the mask image (A) DB 402, and creates a data set for machine learning." Para. 76 see "Thousands of transmitted X-ray images and images in which the void region 36 is masked (manually inspected data) are prepared." Para. 77 see "These are randomly cut out as a 256 × 256 (pixel) patch image (section 37) to make an image with increased noise, contrast adjustment, brightness adjustment, brightness inversion, smoothing, and enlargement. Perform one or more processes such as shrinking, rotating, shifting horizontally or vertically, partially masking (filling in black), or a combination of multiple processes to expand the data and expand the data to hundreds of thousands. Generate a set of sheets." Para. 79 see "As shown in FIG. 6A, the division 37a is black-and-white inverted to obtain training data." Para. 81 see "The feature of the present invention is that even when the amount of machine learning data that can be prepared is small, the number of data is expanded by the method of black-and-white inversion, rotation, reduction, and enlargement of the division as described above, and a large amount of data set is self-generated." Para. 82 see "using 70% of the patch image (category 37) as training data and the remaining 30% as test data, an encoder-decoder model for segmentation is used" Para. 83 see "As the die coefficient (Decee cofficient) between the mask image as the teacher data and the mask image as the network output").
While Ueki teaches transforming labeled images to generate a second training dataset effective for learning, Ueki does not teach the inverted label is effective for learning.
However, Chandler teaches the inverted label is effective for learning. (Para. 10 see "A labeled dataset may be augmented by utilizing a label preserving transformation. A label preserving transformation generates new data from existing data, where the new data has the same label as the existing data." Para. 11 see "One example is a system including a training dataset with at least one training data, and a label preserving transformation including an occluder, and an inpainter. ... the dataset is augmented by adding, for the at least one training data, the transformed at least one training data to the dataset.").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ueki and Chandler to incorporate the teachings of Chandler to select the first labeled training dataset so that when it is augmented the inverted label is effective for learning. Doing so would predictably improve model performance by ensuring the augmented labels remain useful and consistent during training.
Regarding claim 8, Ueki in view of Chandler teaches the data processing apparatus according to claim 1.
In addition, Ueki teaches wherein the processor is configured to: generate the inverted image by inverting the pixel values of a partial region of the single-channel image; (Para. 78 see "As an example, in FIG. 5, the solder portion 38 has four patch images (section 37a, section 37b, section 37c, section 37d). It is preferable that each division 37 has the same shape and area." Para. 79 see "As shown in FIG. 6A, the division 37a is black-and-white inverted to obtain training data.").
While Ueki teaches transforming labeled partial images to generate a second training dataset, Ueki does not teach and generate the transformed label as the ground truth label that constitutes the second labeled training dataset, based on a ground truth label of the image corresponding to the partial region.
However, Chandler teaches and generate the transformed label as the ground truth label that constitutes the second labeled training dataset, (Para. 10 see "A labeled dataset may be augmented by utilizing a label preserving transformation. A label preserving transformation generates new data from existing data, where the new data has the same label as the existing data." Para. 11 see "One example is a system including a training dataset with at least one training data, and a label preserving transformation including an occluder, and an inpainter. ... the dataset is augmented by adding, for the at least one training data, the transformed at least one training data to the dataset."). based on a ground truth label of the image corresponding to the partial region. (Para. 9 see "Augmenting a labeled dataset is the process of applying a transformation function to each data in the dataset to produce additional data with the same label." Para. 10 see "A label preserving transformation generates new data from existing data, where the new data has the same label as the existing data. Reflections, linear shifts, and elastic distortions are three classes of label preserving transformations that may be generally applied to object classification analyses. These transformations build an invariance to an expected transformation in the data.").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ueki and Chandler to incorporate the teachings of Chandler to generate the inverted image by inverting the pixel values of only a partial region of the single-channel image and generate the corresponding transformed label based on the ground truth label of that partial region. Doing so would predictably improve training data diversity and model robustness by allowing augmentations for parts of the image while still keeping each transformed region correctly labeled.
Regarding claim 19, Ueki in view of Chandler teaches the data processing method according to claim 15.
In addition, Ueki teaches A non-transitory, computer-readable tangible recording medium on which a program for causing, when read by a computer, a processor of the computer to execute the data processing method according to claim 15 is recorded. (Para. 34 see "The automatic detection device 1 includes a control unit 11, a main storage unit 12, a communication unit 13, an operation unit 14, a display panel 15, and an auxiliary storage unit 16 that control the entire device." Para. 36 see "The control unit 11 can be composed of a CPU (Central Processing Unit), a ROM (Read Only Memory), a RAM (Random Access Memory), and the like. The control unit 11 may include a GPU (Graphics Processing Unit)." Para. 39 see "The auxiliary storage unit 16 is a large-capacity memory, a hard disk, or the like, and includes programs necessary for the control unit 11 to execute processing").
Claim 15 is rejected under the same analysis as claim 1 above.
Claims 2-4 are rejected under 35 U.S.C. 103 as being unpatentable over Ueki et al. (JP 2021110753 A), hereinafter Ueki, in view of Chandler (US 20150379422 A1), hereinafter Chandler, and Mollov (US 20050285044 A1), hereinafter Mollov.
Regarding claim 2, Ueki in view of Chandler teaches the data processing apparatus according to claim 1.
While Ueki teaches capturing images with an X-ray imager and an ultrasonic microscope, Ueki does not teach wherein the image is an image captured by a digital detector array (DDA) that receives radiation transmitted through a photographic subject or a computed radiography (CR) captured image obtained by causing a reading device to output, as digital values, a received light signal from an imaging plate (IP).
However, Mollov teaches wherein the image is an image captured by a digital detector array (DDA) that receives radiation transmitted through a photographic subject or a computed radiography (CR) captured image obtained by causing a reading device to output, as digital values, a received light signal from an imaging plate (IP). (Para. 11 see "This invention relates to a digital radiography imager having an x-ray converting layer with a first surface adjacent to an energy detection layer and a second surface on an opposite side to the energy detection layer. The digital radiography imager is configured such that x-rays traverse the energy detection layer before propagating through the x-ray converting layer." Para. 24 see "A digital radiography imager is described. In one embodiment, the digital radiography imager may be a multilayer, flat panel imager. The flat panel imager includes a scintillator layer that generates visible light from x-rays absorbed through it. A photodiode layer detects the visible light to generate electrical charges to produce a pixel-based image." Para. 33 see "Radiography imager 200 is configured such that x-rays may be received in a direction from protective layer 260. In this configuration, the x-ray visible light (i.e., light intensity) generated by scintillator layer 220 is greater near a first surface 226 of scintillator 220, which is closer to representative photodiodes 242, 244, compared to a second surface 224 of scintillator 220. As such, photodiode layer 240 may convert more visible light to electrical charges to produce a pixel-based image on a display. Mirror 210 serves to reflect visible light produced by scintillator layer 220 back towards photodiode layer 240. In this way, the amount of visible light captured for detection may be maximized." Examiner Note: This shows the image is an image captured by a digital detector array (DDA) that receives radiation transmitted through a photographic subject.).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ueki and Chandler to incorporate the teachings of Mollov to use a digital detector array to capture the single-channel image that receives radiation transmitted through a photographic subject. Doing so would predictably improve image quality and efficiency by converting the received radiation into digital pixel values directly and with higher light collection efficiency.
Regarding claim 3, Ueki in view of Chandler teaches the data processing apparatus according to claim 1.
While Ueki teaches capturing images with an X-ray imager and an ultrasonic microscope, Ueki does not teach wherein the image is an image obtained by lens-free imaging.
However, Mollov teaches wherein the image is an image obtained by lens-free imaging. (Para. 24 see "A digital radiography imager is described. In one embodiment, the digital radiography imager may be a multilayer, flat panel imager. The flat panel imager includes a scintillator layer that generates visible light from x-rays absorbed through it. A photodiode layer detects the visible light to generate electrical charges to produce a pixel-based image." Para. 25 see "The imager may have a photodiode layer disposed above a protective layer, a light transparent layer disposed above the photodiode layer, a scintillator layer disposed above the light transparent layer, and a mirror layer disposed above the scintillator layer. The scintillator layer has a first surface adjacent to the light transparent layer and a second surface adjacent to the mirror layer. ... The imager is configured such that x-rays traverse the photodiode layer before propagating through the scintillator layer." Para. 31 see "In one embodiment, imager 200 has mirror layer 210, scintillator layer 220, light transparent layer 230, photodiode layer 240, substrate layer 250 and protective layer 260. Photodiodes 242, 244 are representative of a photodiode array that forms photodiode layer 240 disposed above substrate layer 250." Examiner Note: This shows the layered contact structure has no intervening optical lens, which is a standard lens-free arrangement.).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ueki and Chandler to incorporate the teachings of Mollov to obtain the image by lens-free imaging. Doing so would predictably improve image sharpness by collecting light directly with the light sensing elements without an intervening optical lens.
Regarding claim 4, Ueki in view of Chandler teaches The data processing apparatus according to claim 1.
While Ueki teaches capturing images with an X-ray imager and an ultrasonic microscope, Ueki does not teach wherein each pixel of the image has a digital value proportional to an amount of received light at a corresponding light-receiving element among the light-receiving elements or a digital value correlated with the amount of received light.
However, Mollov teaches wherein each pixel of the image has a digital value proportional to an amount of received light at a corresponding light-receiving element among the light-receiving elements or a digital value correlated with the amount of received light. (Para. 24 see "A digital radiography imager is described. In one embodiment, the digital radiography imager may be a multilayer, flat panel imager. The flat panel imager includes a scintillator layer that generates visible light from x-rays absorbed through it. A photodiode layer detects the visible light to generate electrical charges to produce a pixel-based image." Para. 32 see "Scintillator layer 220 absorbs x-rays and generates visible light corresponding to the amount of x-ray absorbed. Photodiode layer 240 detects the light corresponding to the amount of x-ray absorbed. Photodiode layer 240 converts the visible light to electric charges to generate a pixel pattern on a display" Para. 33 see "Radiography imager 200 is configured such that x-rays may be received in a direction from protective layer 260. In this configuration, the x-ray visible light (i.e., light intensity) generated by scintillator layer 220 is greater near a first surface 226 of scintillator 220, which is closer to representative photodiodes 242, 244 ... As such, photodiode layer 240 may convert more visible light to electrical charges to produce a pixel-based image on a display.").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ueki and Chandler to incorporate the teachings of Mollov to make each pixel of the image have a digital value proportional to the amount of received light at the corresponding light-receiving element. Doing so would predictably improve image accuracy by converting the amount of light detected at each sensing element directly into a corresponding digital pixel value.
Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Ueki et al. (JP 2021110753 A), hereinafter Ueki, in view of Chandler (US 20150379422 A1), hereinafter Chandler, and Lee (US 20210271938 A1), hereinafter Lee.
Regarding claim 9, Ueki in view of Chandler teaches the data processing apparatus according to claim 1.
While Ueki teaches generating additional training data from original data, Ueki does not teach wherein the processor is configured to perform normalization or standardization on the training data.
However, Lee teaches wherein the processor is configured to perform normalization or standardization on the training data. (Para. 49 see "In some embodiments, the learning apparatus 10 may perform the normalization on the input image using a normalization model before each image is input to the target model. Further, the learning apparatus 10 may learn the target model using the normalized image or predict a label of the image." Para. 56 see "As shown in FIG. 3, the normalization method begins with step S 100 of acquiring a labeled learning image set. The learning image set may mean a data set for learning including a plurality of images." Para. 58 see "As exemplified in FIG. 4, the input image 31 may be transformed into a normalized image 34 based on a normalization model 33 before the input image 31 is input to a target model 35. In addition, the normalized image 34 may be input to the target model 35. A reason for normalizing the input image 31 before the input image 31 is input to the target model 35 is to induct the stable learning and prevent the overfitting by removing the latent bias.").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ueki and Chandler to incorporate the teachings of Lee to perform normalization or standardization on the training data. Doing so would predictably improve training stability and model accuracy by removing unwanted variation in the images before they are used for learning.
Claims 10-11, 16 are rejected under 35 U.S.C. 103 as being unpatentable over Ueki et al. (JP 2021110753 A), hereinafter Ueki, in view of Chandler (US 20150379422 A1), hereinafter Chandler, and Hoogeboom et al. (EP 3789924 A1), hereinafter Hoogeboom.
Regarding claim 10, Ueki in view of Chandler teaches the data processing apparatus according to claim 1.
In addition, Ueki teaches setting the inverted image as the training data. (Para. 67 see "Using the captured image (A) DB401 and the mask image (A) DB402, a machine learning model capable of detecting which region of the captured image corresponds to the void 36 or the crack 39 is created." Para. 68 see "The data padding preprocessing program 403 reads the captured image (A) DB 401 and the mask image (A) DB 402, and creates a data set for machine learning." Para. 76 see "Thousands of transmitted X-ray images and images in which the void region 36 is masked (manually inspected data) are prepared." Para. 77 see "Perform one or more processes such as shrinking, rotating, shifting horizontally or vertically, partially masking (filling in black), or a combination of multiple processes to expand the data and expand the data to hundreds of thousands. Generate a set of sheets." Para. 79 see "As shown in FIG. 6A, the division 37a is black-and-white inverted to obtain training data." Para. 81 see "The feature of the present invention is that even when the amount of machine learning data that can be prepared is small, the number of data is expanded by the method of black-and-white inversion, rotation, reduction, and enlargement of the division as described above, and a large amount of data set is self-generated." Para. 82 see "using 70% of the patch image (category 37) as training data and the remaining 30% as test data, an encoder-decoder model for segmentation is used" Para. 83 see "As the die coefficient (Decee cofficient) between the mask image as the teacher data and the mask image as the network output").
While Ueki teaches generating additional training data from original labeled data by transforming the original data, Ueki does not teach wherein the processor is configured to edit a class design of the ground truth label that constitutes the second labeled training dataset in response to setting the transformed image as the training data.
However, Hoogeboom teaches wherein the processor is configured to edit a class design of the ground truth label that constitutes the second labeled training dataset in response to setting the transformed image as the training data. (Para. 12 see "In known data augmentation, the input is modified while the label remains unchanged ... the class labels of the existing data instance is maintained for the new data instance." Para. 13 see "The measures described in this specification rather provide a data augmentation scheme which not only modifies the input to the machine learnable model, e.g., generates a new data instance by modifying an existing data instance, but also adapts the class label. Specifically, the class label is adapted according to a conditionally invertible function which takes as input the class label of the existing data instance and the variable ' s ' which controls, steers or in any other way determines the characteristic of the data augmentation. The resulting adapted class label is then used as prediction target when the new data instance is provided as input to the machine learnable model. ... separate classes may be generated by the conditionally invertible function ... the conditionally invertible function f may map (also referred to as 'translate') the class label 'dog' and the variable s = 0 (no rotation) to a first class, the class label 'dog' and the variable s = 1 (90° rotation) to a second class" Para. 47 see "The corresponding label of x * to be used as prediction target label in the training may be obtained by a conditionally invertible function f having as input a class label y of the data x and the variable s, namely as y * = f (y,s).").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ueki and Chandler to incorporate the teachings of Hoogeboom to edit the class design of the ground truth label of the second labeled training dataset when the inverted image is used as training data. Doing so would predictably improve the accuracy of the trained model by keeping the labels consistent with the visual content after inversion so the model learns from correct examples.
Regarding claim 11, Ueki in view of Chandler and Hoogeboom teaches the data processing apparatus according to claim 10.
In addition, Ueki teaches setting the inverted image as the training data. (Para. 67 see "Using the captured image (A) DB401 and the mask image (A) DB402, a machine learning model capable of detecting which region of the captured image corresponds to the void 36 or the crack 39 is created." Para. 68 see "The data padding preprocessing program 403 reads the captured image (A) DB 401 and the mask image (A) DB 402, and creates a data set for machine learning." Para. 76 see "Thousands of transmitted X-ray images and images in which the void region 36 is masked (manually inspected data) are prepared." Para. 77 see "Perform one or more processes such as shrinking, rotating, shifting horizontally or vertically, partially masking (filling in black), or a combination of multiple processes to expand the data and expand the data to hundreds of thousands. Generate a set of sheets." Para. 79 see "As shown in FIG. 6A, the division 37a is black-and-white inverted to obtain training data." Para. 81 see "The feature of the present invention is that even when the amount of machine learning data that can be prepared is small, the number of data is expanded by the method of black-and-white inversion, rotation, reduction, and enlargement of the division as described above, and a large amount of data set is self-generated." Para. 82 see "using 70% of the patch image (category 37) as training data and the remaining 30% as test data, an encoder-decoder model for segmentation is used" Para. 83 see "As the die coefficient (Decee cofficient) between the mask image as the teacher data and the mask image as the network output").
While Ueki teaches generating additional training data from original labeled data by transforming the original data, Ueki does not teach wherein the processor performs, in the class design, class replacement in response to the setting of the transformed image as the training data.
However, Hoogeboom teaches wherein the processor performs, in the class design, class replacement in response to the setting of the transformed image as the training data. (Para. 12 see "In known data augmentation, the input is modified while the label remains unchanged ... the class labels of the existing data instance is maintained for the new data instance." Para. 13 see "The measures described in this specification rather provide a data augmentation scheme which not only modifies the input to the machine learnable model, e.g., generates a new data instance by modifying an existing data instance, but also adapts the class label. Specifically, the class label is adapted according to a conditionally invertible function which takes as input the class label of the existing data instance and the variable ' s ' which controls, steers or in any other way determines the characteristic of the data augmentation. The resulting adapted class label is then used as prediction target when the new data instance is provided as input to the machine learnable model. ... separate classes may be generated by the conditionally invertible function ... the conditionally invertible function f may map (also referred to as 'translate') the class label 'dog' and the variable s = 0 (no rotation) to a first class, the class label 'dog' and the variable s = 1 (90° rotation) to a second class" Para. 47 see "The corresponding label of x * to be used as prediction target label in the training may be obtained by a conditionally invertible function f having as input a class label y of the data x and the variable s, namely as y * = f (y,s).").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ueki and Chandler and Hoogeboom to incorporate the teachings of Hoogeboom to perform class replacement in the class design of the ground truth label of the second labeled training dataset when the inverted image is used as training data. Doing so would predictably improve the accuracy of the trained model by keeping the labels consistent with the visual content after inversion so the model learns from correct examples.
Regarding claim 16, Ueki in view of Chandler teaches the data processing method according to claim 15.
In addition, Ueki teaches setting the inverted image as the training data. (Para. 67 see "Using the captured image (A) DB401 and the mask image (A) DB402, a machine learning model capable of detecting which region of the captured image corresponds to the void 36 or the crack 39 is created." Para. 68 see "The data padding preprocessing program 403 reads the captured image (A) DB 401 and the mask image (A) DB 402, and creates a data set for machine learning." Para. 76 see "Thousands of transmitted X-ray images and images in which the void region 36 is masked (manually inspected data) are prepared." Para. 77 see "Perform one or more processes such as shrinking, rotating, shifting horizontally or vertically, partially masking (filling in black), or a combination of multiple processes to expand the data and expand the data to hundreds of thousands. Generate a set of sheets." Para. 79 see "As shown in FIG. 6A, the division 37a is black-and-white inverted to obtain training data." Para. 81 see "The feature of the present invention is that even when the amount of machine learning data that can be prepared is small, the number of data is expanded by the method of black-and-white inversion, rotation, reduction, and enlargement of the division as described above, and a large amount of data set is self-generated." Para. 82 see "using 70% of the patch image (category 37) as training data and the remaining 30% as test data, an encoder-decoder model for segmentation is used" Para. 83 see "As the die coefficient (Decee cofficient) between the mask image as the teacher data and the mask image as the network output").
While Ueki teaches generating additional training data from original labeled data by transforming the original data, Ueki does not teach wherein the processor performs, in a class design of the ground truth label that constitutes the second labeled training dataset, class replacement in response to setting the transformed image as the training data.
However, Hoogeboom teaches wherein the processor performs, in a class design of the ground truth label that constitutes the second labeled training dataset, class replacement in response to setting the transformed image as the training data. (Para. 12 see "In known data augmentation, the input is modified while the label remains unchanged ... the class labels of the existing data instance is maintained for the new data instance." Para. 13 see "The measures described in this specification rather provide a data augmentation scheme which not only modifies the input to the machine learnable model, e.g., generates a new data instance by modifying an existing data instance, but also adapts the class label. Specifically, the class label is adapted according to a conditionally invertible function which takes as input the class label of the existing data instance and the variable ' s ' which controls, steers or in any other way determines the characteristic of the data augmentation. The resulting adapted class label is then used as prediction target when the new data instance is provided as input to the machine learnable model. ... separate classes may be generated by the conditionally invertible function ... the conditionally invertible function f may map (also referred to as 'translate') the class label 'dog' and the variable s = 0 (no rotation) to a first class, the class label 'dog' and the variable s = 1 (90° rotation) to a second class" Para. 47 see "The corresponding label of x * to be used as prediction target label in the training may be obtained by a conditionally invertible function f having as input a class label y of the data x and the variable s, namely as y * = f (y,s).").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ueki and Chandler to incorporate the teachings of Hoogeboom to perform class replacement in the class design of the ground truth label of the second labeled training dataset when the inverted image is used as training data. Doing so would predictably improve the accuracy of the trained model by keeping the labels consistent with the visual content after inversion so the model learns from correct examples.
Claims 13-14, 18 are rejected under 35 U.S.C. 103 as being unpatentable over Ueki et al. (JP 2021110753 A), hereinafter Ueki, in view of Chandler (US 20150379422 A1), hereinafter Chandler, and Yang et al. (US 20090190716 A1), hereinafter Yang.
Regarding claim 13, Ueki in view of Chandler teaches the data processing apparatus according to claim 1.
While Ueki teaches generating additional training data from original labeled data by transforming the original data, Ueki does not teach wherein the reversible transformation is at least one of linear transformation, logarithmic transformation, or non-linear transformation using a pixel value correspondence table.
However, Yang teaches wherein the reversible transformation is at least one of linear transformation, logarithmic transformation, or non-linear transformation using a pixel value correspondence table. (Para. 29 see "it is important to understand that current digital receivers (CR or DR) provide digital receiver data that has either of two characteristic proportions relative to the amount of x-ray radiation that is received:" Para. 30 see "i) a substantially linear response to the x-ray radiation level; or" Para. 31 see "ii) a substantially logarithmic response to the x-ray radiation level." Para. 61 see "Calculations for transform step 120 differ based on the response characteristic of the digital receiver, whether linear or logarithmic, as shown and described following." Para. 62 see "for digital images from a digital detector system for which the pixel value in the image is linearly proportional to the amount of x-ray radiation, steps 122 a, 124 a, and 126 a are used for generating a transform." Para. 70 see "For digital images from a digital detector system for which the pixel value in the image is logarithmically proportional to the amount of x-ray radiation, a proportionality constant S, corresponding to a slope, is computed.").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ueki and Chandler to incorporate the teachings of Yang to use a linear or logarithmic transformation as the reversible transformation. Doing so would predictably improve the integrity and usefulness of the training data by keeping the pixel values in a form that still reflect the original image.
Regarding claim 14, Ueki in view of Chandler teaches the data processing apparatus according to claim 1.
In addition, Ueki teaches inverting single-channel images (Para. 76 see "Thousands of transmitted X-ray images and images in which the void region 36 is masked (manually inspected data) are prepared." Para. 77 see "These are randomly cut out as a 256 × 256 (pixel) patch image (section 37) to make an image with increased noise, contrast adjustment, brightness adjustment, brightness inversion, smoothing, and enlargement." Para. 79 see "As shown in FIG. 6A, the division 37a is black-and-white inverted to obtain training data.").
While Ueki teaches generating additional training data from original labeled data by transforming the original data, Ueki does not teach wherein the image is an image that has not undergone irreversible transformation.
However, Yang teaches wherein the image is an image that has not undergone irreversible transformation. (Para. 29 see "it is important to understand that current digital receivers (CR or DR) provide digital receiver data that has either of two characteristic proportions relative to the amount of x-ray radiation that is received:" Para. 30 see "i) a substantially linear response to the x-ray radiation level; or" Para. 31 see "ii) a substantially logarithmic response to the x-ray radiation level." Para. 61 see "Calculations for transform step 120 differ based on the response characteristic of the digital receiver, whether linear or logarithmic, as shown and described following." Para. 62 see "for digital images from a digital detector system for which the pixel value in the image is linearly proportional to the amount of x-ray radiation, steps 122 a, 124 a, and 126 a are used for generating a transform." Para. 70 see "For digital images from a digital detector system for which the pixel value in the image is logarithmically proportional to the amount of x-ray radiation, a proportionality constant S, corresponding to a slope, is computed." Examiner Note: This shows linear or logarithmic response of the digital detector. This is data that has not undergone irreversible transformations.).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ueki and Chandler to incorporate the teachings of Yang to use the single-channel image that has not undergone irreversible transformation. Doing so would predictably improve the integrity and usefulness of the training data by keeping the pixel values in a form that still reflect the original image.
Claim 18 is rejected under the same analysis as claim 14 above.
Allowable Subject Matter
Claim(s) 12, 17 is/are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Regarding claims 12 and 17, none of the references cited above teach generating an inverted label by inverting the ground truth label.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Moen et al. (US 20200364857 A1) discloses systems and methods for annotating and curating biological object tracking-specific training datasets.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALEXANDER VAUGHN whose telephone number is (571) 272-5253. The examiner can normally be reached M-F 11am-7pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, JENNIFER MEHMOOD can be reached on (571) 272-2976. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ALEXANDER VAUGHN/Examiner, Art Unit 2675
/JENNIFER MEHMOOD/Supervisory Patent Examiner, Art Unit 2664