DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1-20 are pending in this application.
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Specification
The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed.
Claim Rejections - 35 USC § 101 - EXAMINER’S NOTE
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Regarding Claims 1-20, the claimed subject matter pending in the independent claims filed in this application were given further consideration in accordance with the “Advance Notice of Change in light of Ex Parte Desjardins” and current Guidelines on evaluating subject matter eligibility of claims under 35 U.S.C. 101 in accordance with the MPEP. In order to enhance clarity and facilitate compact prosecution under this new two-prong inquiry, a claim is now eligible at revised Step 2A, a claim is patent eligible unless it:
Recites a judicial exception and
The exception is not integrated into a practical application of the exception.
With respect to Claim 1, 12 and 19, although the claim does recite elements that appear to be directed towards an abstract idea which would constitute a judicial exception, it is integrated into a practical application using a specialized computer system and program, that is directed towards the “selecting an area of interest in at least one of the images defined by a bounding box; cropping the selected areas from the images and storing the cropped images in folders filtering incorrectly identified objects; generating pseudo labels for the remaining images; and assigning correct item names for the pseudo labels.” As these elements do clearly constitute a practical application, they are considered to be statutory subject matter.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent may not be obtained though the invention is not identically disclosed or described as set forth in section 102 of this title, if the differences between the subject matter sought to be patented and the prior art are such that the subject matter as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art to which said subject matter pertains. Patentability shall not be negatived by the manner in which the invention was made.
Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Sisikand et al. (US PGPub 2021/0142105), hereby referred to as “Sisikand”, in view of Ramamonjison et al. (US PGPub 2023/0281974A1), hereby referred to as “Ramamonjishon”.
Consider Claims 1, 12 and 19.
Sisikand teaches:
1. A system comprising: a computer system, the computer system further comprising: / 12. A method comprising:/ 19. A system comprising: (Sisikand: abstract, Systems and methods for automating image annotations are provided, such that a large-scale annotated image collection may be efficiently generated for use in machine learning applications. In some aspects, a mobile device may capture image frames, identifying items appearing in the image frames and detect objects in three-dimensional space across those image frames. Cropped images may be created as associated with each item, which may then be correlated to the detected objects. A unique identifier may then be captured that is associated with the detected object, and labels are automatically applied to the cropped images based on data associated with that unique identifier. In some contexts, images of products carried by a retailer may be captured, and item data may be associated with such images based on that retailer's item taxonomy, for later classification of other/future products.)
1. at least one processor; a graphical user interface; and a computer-usable medium embodying computer program code, the computer- usable medium capable of communicating with the at least one processor, the computer program code comprising instructions executable by the at least one processor and configured for: / 19. a computer system, the computer system further comprising: at least one processor and at least one GPU; a graphical user interface; and a computer-usable medium embodying computer program code, the computer- usable medium capable of communicating with the at least one processor, the computer program code comprising instructions executable by the at least one processor and configured for: (Sisikand: [0062] In the embodiment shown, the computing system 700 includes one or more processors 702, a system memory 708, and a system bus 722 that couples the system memory 708 to the one or more processors 702. The system memory 708 includes RAM (Random Access Memory) 710 and ROM (Read-Only Memory) 712. A basic input/output system that contains the basic routines that help to transfer information between elements within the computing system 700, such as during startup, is stored in the ROM 712. The computing system 700 further includes a mass storage device 714. The mass storage device 714 is able to store software instructions and data. The one or more processors 702 can be one or more central processing units or other processors. [0063]-[0064], [0065] According to various embodiments of the invention, the computing system 700 may operate in a networked environment using logical connections to remote network devices through the network 701. The network 701 is a computer network, such as an enterprise intranet and/or the Internet. The network 701 can include a LAN, a Wide Area Network (WAN), the Internet, wireless transmission mediums, wired transmission mediums, other networks, and combinations thereof The computing system 700 may connect to the network 701 through a network interface unit 704 connected to the system bus 722. It should be appreciated that the network interface unit 704 may also be utilized to connect to other types of networks and remote computing systems. The computing system 700 also includes an input/output controller 706 for receiving and processing input from a number of other devices, including a touch user interface display screen, or another type of input device. Similarly, the input/output controller 706 may provide output to a touch user interface display screen or other type of output device.)
1. receiving a plurality of images; / 12. receiving a plurality of images; / 19. receiving a plurality of images; (Sisikand: [0024] In the example shown, the process flow 100 includes use of an imaging system, such as a camera 102 of a mobile device 115, to capture images of items at any item location 106. Examples of item locations may include a retail store, but may also include other locations where item collections reside. For example, a warehouse or other storage location may also be utilized. In this context, capturing images of items may include either capturing still images or capturing video content including frame images of the items. In the case of video content, all or fewer than all image frames may be utilized as part of the general image classification process. For example, images may be subselected from among the frames based on clarity, similarity to other images, or other factors. [0025] Capture of images of the items 104 using a camera 102 of the mobile device 115 may include, for example, capturing images as the mobile device passes the items, for example as the mobile device travels along an aisle within a store. However, concurrently with capturing the images of the items, the mobile device 115 may also capture object data, such as point cloud data 110. The point cloud data 110 generally corresponds to detected objects based on differing perspectives between adjacent or nearby frames in sequentially captured image or video content. [0026]-[0031], Figure 1)
1. selecting an area of interest in at least one of the plurality of images, defined by a bounding box; / 12. selecting an area of interest in at least one of the images defined by a bounding box; / 19. selecting an area of interest in at least one of the images defined by a bounding box; (Sisikand: [0032] FIG. 2 is a flowchart of a method 200 of large-scale image collection for fine-grained image classification, according to an example embodiment of the present disclosure. The method may be performed, for example, in the context of the use of a mobile device and mobile application installed thereon which implements, at least in part, the training image data set generation tool 112 described above in conjunction with FIG. 1. [0033] In the example shown, the method 200 includes initiating an image capture and tracking operation, for example using a mobile device (step 202). The image capture and tracking operation may be initiated in response to user selection of a mobile application and initiating data collection according to a first mode of the mobile application. The method further includes capturing a series of images to a camera of the mobile device, such as a set of sequential still frame images, or some/all frames of video content (step 204). From the captured series of images, the mobile device will generate cropped images of items in each selected frame or image, and will also generate a point cloud representative of a three-dimensional (depth) model of the objects that are captured in the images. [0037] In the example shown, the method 200 includes assigning unique names to each of the identified objects as well as each of the images for which classification is sought (step 206). In some instances, a universally unique identifier (UUID) may be generated and included in the filename for each image, and assigned to each object. [0038] In the example shown, the method 200 includes correlating objects detected in the point cloud to items captured in and reflected by the cropped images (step 208). This correlation may be performed by, for example, identifying a correlation between a bounding box that defines the cropped image and an anchor point for an object in the point cloud. For example, if a point within a point cloud falls within a bounding box, the image defined by the bounding box can be assigned to the detected object. In some instances, a single central point within the object point cloud may be used as the unique object identifier; in other instances, multiple points on an object within the point cloud are maintained, but tied to a common object identifier.)
1. cropping the selected areas of interest from the images and storing the cropped images in folders; / 12. cropping the selected areas from the images and storing the cropped images in folders; / 19. cropping the selected areas from the images and storing the cropped images in folders; (Sisikand: [0037] In the example shown, the method 200 includes assigning unique names to each of the identified objects as well as each of the images for which classification is sought (step 206). In some instances, a universally unique identifier (UUID) may be generated and included in the filename for each image, and assigned to each object. [0038] In the example shown, the method 200 includes correlating objects detected in the point cloud to items captured in and reflected by the cropped images (step 208). This correlation may be performed by, for example, identifying a correlation between a bounding box that defines the cropped image and an anchor point for an object in the point cloud. For example, if a point within a point cloud falls within a bounding box, the image defined by the bounding box can be assigned to the detected object. In some instances, a single central point within the object point cloud may be used as the unique object identifier; in other instances, multiple points on an object within the point cloud are maintained, but tied to a common object identifier. [0039]-[0040] In the example shown, the method 200 includes correlating item data (e.g., object labels) to the uniquely-identified objects, and extending that identification to the images associated with each object (step 210). In example embodiments, this includes identifying a unique identifier associated with an item and correlating that identifier to the unique identifier of the object. A unique identifier of the item may be, for example, associated with a UPC code or bar code associated with the item. Identifying the unique identifier may be performed, for example, by a mobile device, by entering a second mode (as compared to the image/object capture mode above) in which an image of an item is presented and a user is prompted to capture or enter a bar code or other unique identifier of the item. The bar code or other identifier may be uniquely associated with item data in an item database, for example of a retailer or other entity. That item data may then be linked to the object, and by way of the object to the cropped images. Item data may be, in some instances, retrieved from an inventory database of the enterprise or organization creating the annotated image dataset.)
19. generating an annotation file corresponding to the plurality of images received sorting cropped images by file size; (Sisikand: [0041] In the example shown, once an annotated image dataset is created, that dataset may be used for training and/or validation of one or more machine learning models that are used for image classification (step 212). As noted below, a trained machine learning model may be used for subsequent automatic classification of new images, for example to classify similar images to those on which training is performed. [0042] Referring to FIG. 2 generally, in the context of a retailer, automating annotation of a dataset including retailer inventory (e.g., grocery items) may allow a retailer to be able to predict a fine-grained classification of a new image of a newly-offered product. For example, a new image of a yogurt container may be identified as likely depicting a yogurt container and therefore with high confidence be able to place that item at an appropriate location within the retailer's product taxonomy (e.g., within a grocery, dairy, yogurt hierarchy). Existing image classification datasets may not accurately train according to the specific taxonomy used by such a retailer, and therefore would less accurately generate such classifications.)
1. generating pseudo labels for the remaining images; / 12. generating pseudo labels for the remaining images; / 19. generating pseudo labels for the remaining images using a labeling neural network; (Sisikand: [0043] FIG. 3 is a schematic illustration of a process 300 by which a large-scale, fine-grained image classification dataset may be compiled, in accordance with aspects of the present disclosure. The process 300 may be utilized to generate a large-scale, fine-grained image collection, while minimizing the requirement for manually labeling image data, even for highly-customized or fine-grained classifications. The process 300 is described in the context of generating a set of classification data for a product taxonomy of a retail organization which has a large number of items in its inventory. Such a retail organization may have a frequently-changing product collection, which would benefit significantly from the ability to automatically classify images within its product taxonomy when products are added to that taxonomy.)
1. and assigning correct item names for the pseudo labels. / 12. and assigning correct item names for the pseudo labels./ 19. and assigning correct item names for the pseudo labels using a classification neural network. (Sisikand: [0027] In general, the training image data set generation tool 112 will process the captured images to detect items within each image. The tool 112 may then crop each image to generate cropped images of the items appearing in each image or frame. Each of these cropped images may then optionally be associated with an object that is detected in the point cloud data 110. By a uniquely identifying individual objects in the point cloud data on 110, and then associating cropped images with the correct object, multiple images of the same product item detected as an object may be automatically associated with the same item, and therefore the same classification. [0031]-[0032], Figure 2, [0041] In the example shown, once an annotated image dataset is created, that dataset may be used for training and/or validation of one or more machine learning models that are used for image classification (step 212). As noted below, a trained machine learning model may be used for subsequent automatic classification of new images, for example to classify similar images to those on which training is performed. [0042] Referring to FIG. 2 generally, in the context of a retailer, automating annotation of a dataset including retailer inventory (e.g., grocery items) may allow a retailer to be able to predict a fine-grained classification of a new image of a newly-offered product. For example, a new image of a yogurt container may be identified as likely depicting a yogurt container and therefore with high confidence be able to place that item at an appropriate location within the retailer's product taxonomy (e.g., within a grocery, dairy, yogurt hierarchy). Existing image classification datasets may not accurately train according to the specific taxonomy used by such a retailer, and therefore would less accurately generate such classifications. [0043] FIG. 3 is a schematic illustration of a process 300 by which a large-scale, fine-grained image classification dataset may be compiled, in accordance with aspects of the present disclosure. The process 300 may be utilized to generate a large-scale, fine-grained image collection, while minimizing the requirement for manually labeling image data, even for highly-customized or fine-grained classifications. The process 300 is described in the context of generating a set of classification data for a product taxonomy of a retail organization which has a large number of items in its inventory. Such a retail organization may have a frequently-changing product collection, which would benefit significantly from the ability to automatically classify images within its product taxonomy when products are added to that taxonomy.)
Even if Sisikand does not teach:
1. filtering incorrectly identified objects; / 12. filtering incorrectly identified objects; / 19. removing cropped images from incorrect folders; sorting cropped images by file name; removing cropped images from incorrect folders;
Ramamonjison teaches:
1. A system comprising: a computer system, the computer system further comprising: / 12. A method comprising:/ 19. A system comprising: (Ramamonjison: abstract, The present disclosure provides a method and system for adapting a machine learning model, such as an object detection model, to account for domain shift. The method includes receiving a labeled data elements and target image samples and performing a plurality of model adaptation epochs. Each adaptation epoch includes: predicting for each of the target image samples, using the machine learning model configured by a current set of configuration parameters, a corresponding target class label for the respective target data object included in the target image sample; generating a plurality of labeled mixed data elements that each include: (i) a mixed image sample including a source data object from one of the source image samples and a target data object from one of the target image samples, and (ii) the corresponding source class label for the source data object and the corresponding target class label for the target data object. The method also includes adjusting the current set of configuration parameters to minimize a loss function for the machine learning model for the plurality of mixed data elements. The method results in adapted machine learning model that accounts for domain shift and that has improved performance at inference on new target image samples. [0012]-[0021], [0038]-[0040], Figure 1)
1. at least one processor; a graphical user interface; and a computer-usable medium embodying computer program code, the computer- usable medium capable of communicating with the at least one processor, the computer program code comprising instructions executable by the at least one processor and configured for: / 19. a computer system, the computer system further comprising: at least one processor and at least one GPU; a graphical user interface; and a computer-usable medium embodying computer program code, the computer- usable medium capable of communicating with the at least one processor, the computer program code comprising instructions executable by the at least one processor and configured for: (Ramamonjison: [0038] For contextual purposes, an illustrative architecture of a trained object detection model is shown in FIG. 1 , referred to herein as source model 100. In the illustrated example, source model 100 is a multi-layer deep learning convolutional neural network, and in this regard can include multiple computational blocks 104, each of which includes a respective set of layers 106, 110, each of which are configured to perform a respective set of computational operations. As known in the art, computational blocks 104 can have different configurations and include different layer configurations, and are each configured to process a respective feature map output by a preceding computational block 104. [0039] The illustrated blocks 104 each include matrix multiplication layer 106 that is configured to perform matrix multiplication operations using a set of learned weight parameters. As known in the art, the matrix multiplication layer 106 may for example be a fully connected layer (for example, at an input computational block 104) or a convolution layer (for example, located at each of a plurality of intermediate computational blocks 104). An activation layer may be included to apply a non-linear activation function to outputs generated by the matrix multiplication layer 106. The illustrated blocks 104 also include a batch normalization layer 110, which, as known in the art, can stabilize a neural network during training. Each batch normalization layer 110 is configured by learned batch normalization parameters, for example by a respective pair of parameters commonly referred to as beta and gamma.)
1. receiving a plurality of images; / 12. receiving a plurality of images; / 19. receiving a plurality of images; (Ramamonjison: [0036] In computer vision, a trained machine learning model may be used to perform object detection (e.g., object localization and classification). When a trained machine learning model is used to perform objection detection (referred to hereinafter as a trained object detection model) and the trained object detection model is deployed to a computing system to perform inference using a new set of images which are different than the set of labeled training images that was used to train the object detection model, domain shift can occur. Domain shift occurs due to changes in the operating environments (e.g. weather) where the digital images are captured or due to distortions in the digital images (e.g. blurring). The labeled training images and the new input images belong to somewhat different (shifted) domains. [0040] The computational blocks 104 can include layers other than or in addition to those shown in FIG. 1 , for example pooling layers. [0041] Source model 100 is configured to receive an input image 102, which may for example be an array of pixel data, identify a data object 112 in the input image 102 and generate a respective label 116 for the data object 112 identified in the input image 102. The input image 102 includes a group of pixels that represent a data object 112. Data object 112 corresponds to an object in the input image 102 that can be classified by predicting a class (or category) label from a set of possible candidate class (or category) labels for the data object 112. The generated respective output label 116 can include: (i) a bounding box definition that indicates a location of the data object 112 within the input image 102 (for example, x1, y1, width, height of bounding box 114); (ii) a predicted class (or category) label for the data object 112 (for example, “car”, “pedestrian”, “cat”); and (iii) a probability score for the predicted class (or category) label (for example 0.86 on a scale of 0 to 1). In some examples, an input image 102 can include multiple data objects 112, and the source model can identify the multiple data objects 112 and generate respective output labels 116 for each of the respective data objects 112. [0042] With reference to FIG. 2 , an example of a method 200 for adapting source model 100 for operation in a shifted domain will now be described)
1. selecting an area of interest in at least one of the plurality of images, defined by a bounding box; / 12. selecting an area of interest in at least one of the images defined by a bounding box; / 19. selecting an area of interest in at least one of the images defined by a bounding box; (Ramamonjison: [0041] Source model 100 is configured to receive an input image 102, which may for example be an array of pixel data, identify a data object 112 in the input image 102 and generate a respective label 116 for the data object 112 identified in the input image 102. The input image 102 includes a group of pixels that represent a data object 112. Data object 112 corresponds to an object in the input image 102 that can be classified by predicting a class (or category) label from a set of possible candidate class (or category) labels for the data object 112. The generated respective output label 116 can include: (i) a bounding box definition that indicates a location of the data object 112 within the input image 102 (for example, x1, y1, width, height of bounding box 114); (ii) a predicted class (or category) label for the data object 112 (for example, “car”, “pedestrian”, “cat”); and (iii) a probability score for the predicted class (or category) label (for example 0.86 on a scale of 0 to 1). In some examples, an input image 102 can include multiple data objects 112, and the source model can identify the multiple data objects 112 and generate respective output labels 116 for each of the respective data objects 112. [0042] With reference to FIG. 2 , an example of a method 200 for adapting source model 100 for operation in a shifted domain will now be described. [0044] As noted above, source model 100 is configured with the set of learned parameters θM s to perform an object detection task. Source model 100 has been trained using a supervised learning algorithm and a source training dataset D={(xi,yi)} of data elements (xi,yi) obtained for a source domain, where xi is an source image (e.g., a source image sample) and yi is a set of labels for source data objects that are present in the source image. The label for each source data object includes a class label (also referred to as object category label) and a bounding box definition (also referred to as bounding box coordinates). In an example scenario, the method of FIG. 2 is applied to adapt the source model 100 where a covariate shift occurs between the distribution of source training dataset D={(xi,yi)} and the distribution of target image samples that will be collected for a target domain in which the source model 100 will be implemented.)
1. cropping the selected areas of interest from the images and storing the cropped images in folders; / 12. cropping the selected areas from the images and storing the cropped images in folders; / 19. cropping the selected areas from the images and storing the cropped images in folders; (Ramamonjison: [0044] As noted above, source model 100 is configured with the set of learned parameters θM s to perform an object detection task. Source model 100 has been trained using a supervised learning algorithm and a source training dataset D={(xi,yi)} of data elements (xi,yi) obtained for a source domain, where xi is an source image (e.g., a source image sample) and yi is a set of labels for source data objects that are present in the source image. The label for each source data object includes a class label (also referred to as object category label) and a bounding box definition (also referred to as bounding box coordinates). In an example scenario, the method of FIG. 2 is applied to adapt the source model 100 where a covariate shift occurs between the distribution of source training dataset D={(xi,yi)} and the distribution of target image samples that will be collected for a target domain in which the source model 100 will be implemented. [0051] As indicated in FIG. 3 , mixed sample generator operation 204 randomly samples labeled source training dataset D={({circumflex over (x)}i,ŷi)} to select source image x1 (which has corresponding source label y1). Mixed sample generator operation 204 then randomly samples additional image samples (e.g., x2,x 1,x 2 from the source and target datasets D∪D to create a new domain-mixed composite image sample {circumflex over (x)}i. The corresponding source labels y1,y2 and target pseudo-labels y 1,y 2 are collated into a label set ŷi for the new domain-mixed composite image sample {circumflex over (x)}i. In an example embodiment, the respective data object and a surrounding image portion in the respective source and target image samples are cropped and the cropped image portions stitched together to form the new domain-mixed composite image sample {circumflex over (x)}i. In at least some examples, one or more of the cropped images are scaled (e.g., in the case of down scaling, data from a larger set of pixels are combined using a known scaling algorithm into a smaller set of pixels). The respective bounding box definitions corresponding to the source and target images are computed based on the relative positions of each cropped image portion in the domain-mixed composite image sample {circumflex over (x)}i. [0052] In the illustrated example of FIG. 3 , portions of source image samples 304_1 and 304_2 are cropped and positioned in the top left and bottom right quarter portions of a new domain-mixed composite image sample 308; portions of target image samples 306_1 and 306_2 are cropped and positioned in the top right and bottom left quarter portions of the new domain-mixed composite image sample 308. Accordingly, domain-mixed composite image sample 308 is a collage of cropped portions from four image samples that have been randomly sampled from the source and target datasets D∪D. In other examples, the number of data objects integrated into the new domain-mixed composite image sample 308 could be greater than or less than 4. The label set ŷi for the new domain-mixed composite image sample {circumflex over (x)}i will include: (i) the respective class labels for each of the data objects that have been cropped from source image samples, along with their respective adjusted bounding box definitions; and (ii) the respective pseudo-class labels for each of the data objects that have been cropped from target image samples, along with their respective adjusted bounding box definitions.)
19. generating an annotation file corresponding to the plurality of images received sorting cropped images by file size; (Ramamonjison: [0052] In the illustrated example of FIG. 3 , portions of source image samples 304_1 and 304_2 are cropped and positioned in the top left and bottom right quarter portions of a new domain-mixed composite image sample 308; portions of target image samples 306_1 and 306_2 are cropped and positioned in the top right and bottom left quarter portions of the new domain-mixed composite image sample 308. Accordingly, domain-mixed composite image sample 308 is a collage of cropped portions from four image samples that have been randomly sampled from the source and target datasets D∪D. In other examples, the number of data objects integrated into the new domain-mixed composite image sample 308 could be greater than or less than 4. The label set ŷi for the new domain-mixed composite image sample {circumflex over (x)}i will include: (i) the respective class labels for each of the data objects that have been cropped from source image samples, along with their respective adjusted bounding box definitions; and (ii) the respective pseudo-class labels for each of the data objects that have been cropped from target image samples, along with their respective adjusted bounding box definitions.)
1. filtering incorrectly identified objects; / 12. filtering incorrectly identified objects; / 19. removing cropped images from incorrect folders; sorting cropped images by file name; removing cropped images from incorrect folders; (Ramamonjison: [0004] An important form of domain shift is low-level image corruptions. When a trained model is a trained object detection model and corrupted digital images are input to the trained object detection model, the trained object detection model may not be able to detect any objects at all in the corrupted digital images, the trained object detection model may detect an object in the corrupted digital images but it may classify the object detected in the corrupted digital images with an incorrect object category label, or the trained object detection model may classify an object detected in the corrupted with a correct object category label but not generate an accurate bounding box for the detected object. If the distribution of a set of test images used for testing the performance of the trained object detection model drifts significantly from distribution of the set of labeled training images used to train the object detection model, the performance of trained object detention model when deployed for inference using a set of new images will degrade and the object detection model may completely fail at detecting objects in the set of new images. [0041] Source model 100 is configured to receive an input image 102, which may for example be an array of pixel data, identify a data object 112 in the input image 102 and generate a respective label 116 for the data object 112 identified in the input image 102. The input image 102 includes a group of pixels that represent a data object 112. Data object 112 corresponds to an object in the input image 102 that can be classified by predicting a class (or category) label from a set of possible candidate class (or category) labels for the data object 112. The generated respective output label 116 can include: (i) a bounding box definition that indicates a location of the data object 112 within the input image 102 (for example, x1, y1, width, height of bounding box 114); (ii) a predicted class (or category) label for the data object 112 (for example, “car”, “pedestrian”, “cat”); and (iii) a probability score for the predicted class (or category) label (for example 0.86 on a scale of 0 to 1). In some examples, an input image 102 can include multiple data objects 112, and the source model can identify the multiple data objects 112 and generate respective output labels 116 for each of the respective data objects 112. [0047] As noted above, a respective mixed dataset batch {circumflex over (β)}={({circumflex over (x)}i,ŷi)} is generated for each epoch. The generating of a mixed dataset batch {circumflex over (β)}={({circumflex over (x)}i,ŷi)} will now be described in the context of a first epoch of the first phase (Phase 1). As a first step, a pseudo-label generator operation 202 is applied to unlabeled target dataset D={(x j)} to generate a set of respective pseudo-labels {y j}. The pseudo-label generator operation 202 uses the source model 100 with the set of learned parameters θM s to generate the respective pseudo-labels {y j} for target data objects included in the target image samples {x j}. Each pseudo-label y j includes a predicted target class (or object category) label and a bounding box definition for the target data object. In some examples, the pseudo-label y j can also include a probability value for the predicted target class label.)
1. generating pseudo labels for the remaining images; / 12. generating pseudo labels for the remaining images; / 19. generating pseudo labels for the remaining images using a labeling neural network; (Ramamonjison: [0052] In the illustrated example of FIG. 3 , portions of source image samples 304_1 and 304_2 are cropped and positioned in the top left and bottom right quarter portions of a new domain-mixed composite image sample 308; portions of target image samples 306_1 and 306_2 are cropped and positioned in the top right and bottom left quarter portions of the new domain-mixed composite image sample 308. Accordingly, domain-mixed composite image sample 308 is a collage of cropped portions from four image samples that have been randomly sampled from the source and target datasets D∪D. In other examples, the number of data objects integrated into the new domain-mixed composite image sample 308 could be greater than or less than 4. The label set ŷi for the new domain-mixed composite image sample {circumflex over (x)}i will include: (i) the respective class labels for each of the data objects that have been cropped from source image samples, along with their respective adjusted bounding box definitions; and (ii) the respective pseudo-class labels for each of the data objects that have been cropped from target image samples, along with their respective adjusted bounding box definitions. [0057] Mixed sample generator operation 204 mixes ground-truth source labels and target pseudo-labels in the same mixed data element. This can mitigates the effect of false labels during adaptation of the source model 100 (described below) because the mixed data element always contains accurate labels from the source dataset. [0058] Further, the use of mixed sample generator operation 204 can enforce the source model 100 that is being adapted to detect small objects as the data objects in the source samples can be scaled down in the mixed images. [0059] The source model 100 can then be used with the updated parameters of the batch normalization layers (and original parameters of all other layers of the source model 100) by generate pseudo label operation 202 to provide a new set of pseudo-labels for the target image sample dataset D={(x j)}, which are then used by generate mixed sample operation 204 to generate a further dataset batch {circumflex over (β)}={({circumflex over (x)}i,ŷi)} for another round of BatchNorm adaptation training by BatchNorm adaptation operation 206.)
1. and assigning correct item names for the pseudo labels. / 12. and assigning correct item names for the pseudo labels./ 19. and assigning correct item names for the pseudo labels using a classification neural network. (Ramamonjison: [0052] In the illustrated example of FIG. 3 , portions of source image samples 304_1 and 304_2 are cropped and positioned in the top left and bottom right quarter portions of a new domain-mixed composite image sample 308; portions of target image samples 306_1 and 306_2 are cropped and positioned in the top right and bottom left quarter portions of the new domain-mixed composite image sample 308. Accordingly, domain-mixed composite image sample 308 is a collage of cropped portions from four image samples that have been randomly sampled from the source and target datasets D∪D. In other examples, the number of data objects integrated into the new domain-mixed composite image sample 308 could be greater than or less than 4. The label set ŷi for the new domain-mixed composite image sample {circumflex over (x)}i will include: (i) the respective class labels for each of the data objects that have been cropped from source image samples, along with their respective adjusted bounding box definitions; and (ii) the respective pseudo-class labels for each of the data objects that have been cropped from target image samples, along with their respective adjusted bounding box definitions. [0063] FIG. 5 shows a pseudocode representation of the gradual adaptation method 200 of FIG. 2. [0064] FIG. 6 provides a flow chart summarizing an example of the method 200. The method 200 begins at block 602. At block 602, a plurality of labeled data elements are received. Each labeled data element includes: (i) a source sample (e.g., source images) including a respective source data object, and (ii) a corresponding source class label for the respective source data object. The method then proceeds to block 604. At block 604, a plurality of target samples (e.g., target images) are received. Each target sample includes a respective target data object. The method 200 then proceeds to block 606. At block 606, the method 200 predicts, for each of the plurality of targets samples, using a machine learning model (e.g. object detection model) configured by a current set of configuration parameters, a corresponding target class label for the respective target data object included in the targets sample. The method 200 then proceeds to block 608. At block 608, a plurality of labeled mixed data elements is generated. Each labeled mixed data element includes (i) a mixed sample including a source data object from one of the source samples and a target data object from one of the target samples, and (ii) the corresponding source class label for the source data object and the corresponding target class label for the target data object. The method 200 then proceeds to block 610. At block 610, the method 200 finely, adjusts the current set of configuration parameters to minimize a loss function for the machine learning model (e.g. the object detection model) for the plurality of labeled mixed data elements. In some examples, for an initial set of model adaptation epochs, only parameters of the BatchNorm layers of the machine learning model (e.g. the object detection model) are adjusted, after which configuration parameters for all other layers are adjusted. The method then proceeds to block 612 where blocks 608 to 610 are repeated for a plurality of model adaptation epochs and the final adjusted set of the configuration parameters are output as configuration parameters of the machine learning model (e.g. the object detection model).)
It would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to modify the large scale automated image annotation system of Sisikand with the method and system for adaptation of a trained object detection model that takes into account domain shifts as disclosed by Ramamonjishon. Both references are directed towards the same field of machine learning models for automated object detection and classification, and the determination of obviousness is predicated upon the following findings: One skilled in the art would have been motivated to modify Sisikand in order to ensure the overall object detection model takes into account Furthermore, the prior art collectively includes each element claimed (though not all in the same reference), and one of ordinary skill in the art could have combined the elements in the manner explained above using known engineering design, interface and/or programming techniques, without changing a “fundamental” operating principle of Sisikand, while the teaching of Ramamonjishon continues to perform the same function as originally taught prior to being combined, in order to produce the repeatable and predictable result of ensuring accuracy of classified objects by taking into account domain shifts and improved inference on new target objects. It is for at least the aforementioned reasons that the examiner has reached a conclusion of obviousness with respect to the claim in question.
Consider Claim 2.
The combination of Sisikand and Ramamonjishon teaches:
2. The system of claim 1 where the plurality of images comprises at least one of: an image file; a video file; and a video frame file. (Sisikand: [0024] In the example shown, the process flow 100 includes use of an imaging system, such as a camera 102 of a mobile device 115, to capture images of items at any item location 106. Examples of item locations may include a retail store, but may also include other locations where item collections reside. For example, a warehouse or other storage location may also be utilized. In this context, capturing images of items may include either capturing still images or capturing video content including frame images of the items. In the case of video content, all or fewer than all image frames may be utilized as part of the general image classification process. For example, images may be subselected from among the frames based on clarity, similarity to other images, or other factors. [0025] Capture of images of the items 104 using a camera 102 of the mobile device 115 may include, for example, capturing images as the mobile device passes the items, for example as the mobile device travels along an aisle within a store. However, concurrently with capturing the images of the items, the mobile device 115 may also capture object data, such as point cloud data 110. The point cloud data 110 generally corresponds to detected objects based on differing perspectives between adjacent or nearby frames in sequentially captured image or video content. Ramamonjison: [0041] Source model 100 is configured to receive an input image 102, which may for example be an array of pixel data, identify a data object 112 in the input image 102 and generate a respective label 116 for the data object 112 identified in the input image 102. The input image 102 includes a group of pixels that represent a data object 112. Data object 112 corresponds to an object in the input image 102 that can be classified by predicting a class (or category) label from a set of possible candidate class (or category) labels for the data object 112. [0068] In the above described examples the object detection model performs an object detection task in respect of camera images that are image frames captured by a light-intensity image sensor such as a video camera. In alternative embodiments, the same methods can be applied to adapt an object detection model that operates using Light Detection and Ranging (LIDAR) images, in which case the source samples and target samples are LIDAR images rather than camera images.)
Consider Claims 3 and 13.
The combination of Sisikand and Ramamonjishon teaches:
3. The system of claim 1 wherein the computer program code comprising instructions executable by at least one processor is further configured for: identifying any of the plurality of images missing bounding boxes./ 13. The method of claim 12 further comprising: identifying any of the plurality of images missing bounding boxes. (Sisikand: [0038] In the example shown, the method 200 includes correlating objects detected in the point cloud to items captured in and reflected by the cropped images (step 208). This correlation may be performed by, for example, identifying a correlation between a bounding box that defines the cropped image and an anchor point for an object in the point cloud. For example, if a point within a point cloud falls within a bounding box, the image defined by the bounding box can be assigned to the detected object. In some instances, a single central point within the object point cloud may be used as the unique object identifier; in other instances, multiple points on an object within the point cloud are maintained, but tied to a common object identifier. Ramamonjison: [0041] Source model 100 is configured to receive an input image 102, which may for example be an array of pixel data, identify a data object 112 in the input image 102 and generate a respective label 116 for the data object 112 identified in the input image 102. The input image 102 includes a group of pixels that represent a data object 112. Data object 112 corresponds to an object in the input image 102 that can be classified by predicting a class (or category) label from a set of possible candidate class (or category) labels for the data object 112. The generated respective output label 116 can include: (i) a bounding box definition that indicates a location of the data object 112 within the input image 102 (for example, x1, y1, width, height of bounding box 114); (ii) a predicted class (or category) label for the data object 112 (for example, “car”, “pedestrian”, “cat”); and (iii) a probability score for the predicted class (or category) label (for example 0.86 on a scale of 0 to 1). In some examples, an input image 102 can include multiple data objects 112, and the source model can identify the multiple data objects 112 and generate respective output labels 116 for each of the respective data objects)
Consider Claims 4 and 14.
The combination of Sisikand and Ramamonjishon teaches:
4. The system of claim 1 wherein the computer program code comprising instructions executable by the at least one processor is further configured for: generating an annotation file corresponding to the plurality of images received./ 14. The method of claim 12 further comprising: generating an annotation file corresponding to the plurality of images received. (Sisikand: [0041] In the example shown, once an annotated image dataset is created, that dataset may be used for training and/or validation of one or more machine learning models that are used for image classification (step 212). As noted below, a trained machine learning model may be used for subsequent automatic classification of new images, for example to classify similar images to those on which training is performed. [0042] Referring to FIG. 2 generally, in the context of a retailer, automating annotation of a dataset including retailer inventory (e.g., grocery items) may allow a retailer to be able to predict a fine-grained classification of a new image of a newly-offered product. For example, a new image of a yogurt container may be identified as likely depicting a yogurt container and therefore with high confidence be able to place that item at an appropriate location within the retailer's product taxonomy (e.g., within a grocery, dairy, yogurt hierarchy). Existing image classification datasets may not accurately train according to the specific taxonomy used by such a retailer, and therefore would less accurately generate such classifications. Ramamonjison: [0076] Thus, DomainMix augmentation module 804 performs a data augmentation method which adapts the feature representations generated by the trained object detection model to perform well on both the source domain and target domain. The data augmentation method generates new training images by randomly sampling images drawn from both a source domain data and a target domain data, and by mixing four sampled images at a time to create new mixed images. The random sampling is done with a weighted balanced sampler to handle the general case of unbalanced data, resulting in a data augmentation method that is effective even when the number of images from target domain is much less than that from the source domain. Furthermore, the data augmentation method transforms and collates the ground-truth labels of the source domain images with the pseudo-labels generated on the target domain images. The operations performed by the data augmentation method are shown in Algorithm 2 of FIG. 10. [0077] The data augmentation method performed by the DomainMix augmentation module 804 can, in some scenarios, have the following advantages. First, the data augmentation method explicitly adapts the learned feature representations of the trained object detection model by collating the source domain images and target domain images. Second, mixing the ground-truth labels and pseudo-labels strengthens the training supervision of the trained object detection model and prevents the object detection model from becoming unstable during further training. Third, the data augmentation method of the present disclosure reduces the scale of objects in original training images. As a result, the data augmentation method helps the trained object detection model to detect smaller objects and generate higher quality pseudo-labels. [0078] The pseudo-label generator 802 does not rely on the availability of ground-truth annotations for the target domain images. Instead, target domain images are provided to the pseudo-label generator 802 to extract pseudo-labels which are then mixed with the ground-truths labels from source images of the source domain training dataset.)
Consider Claim 5.
The combination of Sisikand and Ramamonjishon teaches:
5. The system of claim 1 wherein the folders further comprise: folder names corresponding to objects. (Sisikand: [0041] In the example shown, once an annotated image dataset is created, that dataset may be used for training and/or validation of one or more machine learning models that are used for image classification (step 212). As noted below, a trained machine learning model may be used for subsequent automatic classification of new images, for example to classify similar images to those on which training is performed. [0042] Referring to FIG. 2 generally, in the context of a retailer, automating annotation of a dataset including retailer inventory (e.g., grocery items) may allow a retailer to be able to predict a fine-grained classification of a new image of a newly-offered product. For example, a new image of a yogurt container may be identified as likely depicting a yogurt container and therefore with high confidence be able to place that item at an appropriate location within the retailer's product taxonomy (e.g., within a grocery, dairy, yogurt hierarchy). Existing image classification datasets may not accurately train according to the specific taxonomy used by such a retailer, and therefore would less accurately generate such classifications. Ramamonjison: [0076] Thus, DomainMix augmentation module 804 performs a data augmentation method which adapts the feature representations generated by the trained object detection model to perform well on both the source domain and target domain. The data augmentation method generates new training images by randomly sampling images drawn from both a source domain data and a target domain data, and by mixing four sampled images at a time to create new mixed images. The random sampling is done with a weighted balanced sampler to handle the general case of unbalanced data, resulting in a data augmentation method that is effective even when the number of images from target domain is much less than that from the source domain. Furthermore, the data augmentation method transforms and collates the ground-truth labels of the source domain images with the pseudo-labels generated on the target domain images. The operations performed by the data augmentation method are shown in Algorithm 2 of FIG. 10. [0077] The data augmentation method performed by the DomainMix augmentation module 804 can, in some scenarios, have the following advantages. First, the data augmentation method explicitly adapts the learned feature representations of the trained object detection model by collating the source domain images and target domain images. Second, mixing the ground-truth labels and pseudo-labels strengthens the training supervision of the trained object detection model and prevents the object detection model from becoming unstable during further training. Third, the data augmentation method of the present disclosure reduces the scale of objects in original training images. As a result, the data augmentation method helps the trained object detection model to detect smaller objects and generate higher quality pseudo-labels. [0078] The pseudo-label generator 802 does not rely on the availability of ground-truth annotations for the target domain images. Instead, target domain images are provided to the pseudo-label generator 802 to extract pseudo-labels which are then mixed with the ground-truths labels from source images of the source domain training dataset.)
Consider Claim 6.
The combination of Sisikand and Ramamonjishon teaches:
6. The system of claim 5 wherein the cropped images in the folders follow a file naming convention. (Sisikand: [0049] In example embodiments, at either the time of capture or at the time transferred to object storage 320, tracked objects and image file names are given unique names. For example, each tracked object will be assigned a universally unique identifier (UUID) as an object name, and each cropped image will have an associated UUID as its image name. All images are sent to a queue for processing and transmitting to a data storage service with the associated filename (as seen in process (f)). In some cases, a separate queue may be used to send image-to-object relationships to the same data storage service. [0050] It is noted that the generalized process 300 described in conjunction with FIG. 3 will rapidly associate item images with a particular object, and can tie multiple such images to a single object. However, at this point, the object is not yet associated with a label. Accordingly, and as discussed above in conjunction with FIGS. 1-2, the mobile device may have a second mode that may be selected in which a user may be presented with an object identifier and optionally one of the captured images of the object, and requests a scan of the bar code associated with that image. Once the bar code image is scanned, this information will be sent to the data storage service, providing a link between the UUID of the tracked object and the barcode data, thereby applying all of the labels that are associated with a particular item represented by the barcode data to that object, and therefore each of the cropped images that were previously associated with the object. Ramamonjison: [0044] As noted above, source model 100 is configured with the set of learned parameters θM s to perform an object detection task. Source model 100 has been trained using a supervised learning algorithm and a source training dataset D={(xi,yi)} of data elements (xi,yi) obtained for a source domain, where xi is an source image (e.g., a source image sample) and yi is a set of labels for source data objects that are present in the source image. The label for each source data object includes a class label (also referred to as object category label) and a bounding box definition (also referred to as bounding box coordinates). In an example scenario, the method of FIG. 2 is applied to adapt the source model 100 where a covariate shift occurs between the distribution of source training dataset D={(xi,yi)} and the distribution of target image samples that will be collected for a target domain in which the source model 100 will be implemented. [0045] The method 200 receives, as inputs, the labeled source training dataset D={(xi,yi)} as well as an unlabeled target dataset D={(x j)}, where xj is a target image sample from the target domain. In some examples, method 200 also receives, as inputs, a pre-trained source object detection model 100 (also referred to as the “source model 100”) having a specified architecture (e.g., number, type and order of computation blocks 104, including number, type and order of layers within the computation blocks 104, as well as hyper-parameters for the source model) and the set of learned parameters θM s.)
Consider Claim 7.
The combination of Sisikand and Ramamonjishon teaches:
7. The system of claim 6 wherein the file naming convention, comprises: a file name of a type [ORIGINAL IMAGE NAME]-[LINE NUMBER IN ANNOTATION FILE]. (Sisikand: [0049] In example embodiments, at either the time of capture or at the time transferred to object storage 320, tracked objects and image file names are given unique names. For example, each tracked object will be assigned a universally unique identifier (UUID) as an object name, and each cropped image will have an associated UUID as its image name. All images are sent to a queue for processing and transmitting to a data storage service with the associated filename (as seen in process (f)). In some cases, a separate queue may be used to send image-to-object relationships to the same data storage service. [0050] It is noted that the generalized process 300 described in conjunction with FIG. 3 will rapidly associate item images with a particular object, and can tie multiple such images to a single object. However, at this point, the object is not yet associated with a label. Accordingly, and as discussed above in conjunction with FIGS. 1-2, the mobile device may have a second mode that may be selected in which a user may be presented with an object identifier and optionally one of the captured images of the object, and requests a scan of the bar code associated with that image. Once the bar code image is scanned, this information will be sent to the data storage service, providing a link between the UUID of the tracked object and the barcode data, thereby applying all of the labels that are associated with a particular item represented by the barcode data to that object, and therefore each of the cropped images that were previously associated with the object. Ramamonjison: [0044] As noted above, source model 100 is configured with the set of learned parameters θM s to perform an object detection task. Source model 100 has been trained using a supervised learning algorithm and a source training dataset D={(xi,yi)} of data elements (xi,yi) obtained for a source domain, where xi is an source image (e.g., a source image sample) and yi is a set of labels for source data objects that are present in the source image. The label for each source data object includes a class label (also referred to as object category label) and a bounding box definition (also referred to as bounding box coordinates). In an example scenario, the method of FIG. 2 is applied to adapt the source model 100 where a covariate shift occurs between the distribution of source training dataset D={(xi,yi)} and the distribution of target image samples that will be collected for a target domain in which the source model 100 will be implemented. [0045] The method 200 receives, as inputs, the labeled source training dataset D={(xi,yi)} as well as an unlabeled target dataset D={(x j)}, where xj is a target image sample from the target domain. In some examples, method 200 also receives, as inputs, a pre-trained source object detection model 100 (also referred to as the “source model 100”) having a specified architecture (e.g., number, type and order of computation blocks 104, including number, type and order of layers within the computation blocks 104, as well as hyper-parameters for the source model) and the set of learned parameters θM s.)
Consider Claim 8.
The combination of Sisikand and Ramamonjishon teaches:
8. The system of claim 1 wherein the computer program code comprising instructions executable by the at least one processor is further configured for: sorting cropped images by file size. (Sisikand: Testing/Validation of Image Collection [0054] Using the methods and systems described above, image collections may be rapidly developed by image capture and automated linking to objects which are in turn linked to item labels (e.g., via bar code or other unique identifier). Such automated image collection and annotation systems may be tested for accuracy relative to product image collections generated using other approaches. [0055] In one example experiment, a product image dataset of a retailer is assessed as to accuracy using a Resnet50 convolutional neural network (CNN) to output image embeddings for a k-NN classifier. In this example, a final pooling layer of the network is put through a fully-connected layer to product embeddings of dimension 100. The network is trained using a proxy-NCA loss (batch size: 32) and Adam optimizer (with learning rate of 104), with three experiments considered: intra-domain learning in which query and catalog sets are from the same image domain and the embedding is trained on labeled “in-the-wild” product images, supervised cross-domain learning in which a catalog set is taken from professional item photos, and weakly-supervised intra-domain learning in which embeddings are trained using unlabeled “in-the-wild” photos, using an entity ID as a proxy for actual labels. [0056] As seen in the results below, a mean average precision (mAP) is assessed in each of the experiments, when using the dataset generated from a retailer's product collection as discussed above. Ramamonjison: [0074] As a result, the DomainMix augmentation module 804 mitigates the effect of noisy labels of the target domain samples during training because the mixed-image always contain accurate labels from the source domain. Finally, the DomainMix augmentation module 804 enforces the object detection model to detect small objects since the sizes of the objects in the original images are scaled down after the augmentation. [0075] The model adapter (e.g., implemented by batch norm adaptation operation 206 and fine tuning operation 212) coordinates the two-phase method for domain adaptation of a trained object detection model. The model adapter may be considered to be a model trainer. Its role is to take the training samples obtained with DomainMix augmentation module 804 and update the parameters of the batch norm layers during the first phase of the method of the present disclosure. In the second phase, the method uses new domain-mixed training samples with refined pseudo-labels and updates the parameters of all convolutional layers and batch norm layers of the neural network that approximates the object detection model. [0076] Thus, DomainMix augmentation module 804 performs a data augmentation method which adapts the feature representations generated by the trained object detection model to perform well on both the source domain and target domain. The data augmentation method generates new training images by randomly sampling images drawn from both a source domain data and a target domain data, and by mixing four sampled images at a time to create new mixed images. The random sampling is done with a weighted balanced sampler to handle the general case of unbalanced data, resulting in a data augmentation method that is effective even when the number of images from target domain is much less than that from the source domain. Furthermore, the data augmentation method transforms and collates the ground-truth labels of the source domain images with the pseudo-labels generated on the target domain images. The operations performed by the data augmentation method are shown in Algorithm 2 of FIG. 10.)
Consider Claim 10.
The combination of Sisikand and Ramamonjishon teaches:
10. The system of claim 1 wherein the computer program code comprising instructions executable by the at least one processor is further configured for: training a labeling neural network, wherein the trained labeling neural network is used to generate the pseudo labels for the remaining images. (Sisikand: Testing/Validation of Image Collection [0054] Using the methods and systems described above, image collections may be rapidly developed by image capture and automated linking to objects which are in turn linked to item labels (e.g., via bar code or other unique identifier). Such automated image collection and annotation systems may be tested for accuracy relative to product image collections generated using other approaches. [0055] In one example experiment, a product image dataset of a retailer is assessed as to accuracy using a Resnet50 convolutional neural network (CNN) to output image embeddings for a k-NN classifier. In this example, a final pooling layer of the network is put through a fully-connected layer to product embeddings of dimension 100. The network is trained using a proxy-NCA loss (batch size: 32) and Adam optimizer (with learning rate of 104), with three experiments considered: intra-domain learning in which query and catalog sets are from the same image domain and the embedding is trained on labeled “in-the-wild” product images, supervised cross-domain learning in which a catalog set is taken from professional item photos, and weakly-supervised intra-domain learning in which embeddings are trained using unlabeled “in-the-wild” photos, using an entity ID as a proxy for actual labels. [0056] As seen in the results below, a mean average precision (mAP) is assessed in each of the experiments, when using the dataset generated from a retailer's product collection as discussed above. Ramamonjison: [0074] As a result, the DomainMix augmentation module 804 mitigates the effect of noisy labels of the target domain samples during training because the mixed-image always contain accurate labels from the source domain. Finally, the DomainMix augmentation module 804 enforces the object detection model to detect small objects since the sizes of the objects in the original images are scaled down after the augmentation. [0075] The model adapter (e.g., implemented by batch norm adaptation operation 206 and fine tuning operation 212) coordinates the two-phase method for domain adaptation of a trained object detection model. The model adapter may be considered to be a model trainer. Its role is to take the training samples obtained with DomainMix augmentation module 804 and update the parameters of the batch norm layers during the first phase of the method of the present disclosure. In the second phase, the method uses new domain-mixed training samples with refined pseudo-labels and updates the parameters of all convolutional layers and batch norm layers of the neural network that approximates the object detection model. [0076] Thus, DomainMix augmentation module 804 performs a data augmentation method which adapts the feature representations generated by the trained object detection model to perform well on both the source domain and target domain. The data augmentation method generates new training images by randomly sampling images drawn from both a source domain data and a target domain data, and by mixing four sampled images at a time to create new mixed images. The random sampling is done with a weighted balanced sampler to handle the general case of unbalanced data, resulting in a data augmentation method that is effective even when the number of images from target domain is much less than that from the source domain. Furthermore, the data augmentation method transforms and collates the ground-truth labels of the source domain images with the pseudo-labels generated on the target domain images. The operations performed by the data augmentation method are shown in Algorithm 2 of FIG. 10.)
Consider Claim 11.
The combination of Sisikand and Ramamonjishon teaches:
11. The system of claim 1 wherein the computer program code comprising instructions executable by the at least one processor is further configured for: training a classification neural network, wherein the trained classification neural network is used to assign the correct item names for the pseudo labels. (Sisikand: Testing/Validation of Image Collection [0054] Using the methods and systems described above, image collections may be rapidly developed by image capture and automated linking to objects which are in turn linked to item labels (e.g., via bar code or other unique identifier). Such automated image collection and annotation systems may be tested for accuracy relative to product image collections generated using other approaches. [0055] In one example experiment, a product image dataset of a retailer is assessed as to accuracy using a Resnet50 convolutional neural network (CNN) to output image embeddings for a k-NN classifier. In this example, a final pooling layer of the network is put through a fully-connected layer to product embeddings of dimension 100. The network is trained using a proxy-NCA loss (batch size: 32) and Adam optimizer (with learning rate of 104), with three experiments considered: intra-domain learning in which query and catalog sets are from the same image domain and the embedding is trained on labeled “in-the-wild” product images, supervised cross-domain learning in which a catalog set is taken from professional item photos, and weakly-supervised intra-domain learning in which embeddings are trained using unlabeled “in-the-wild” photos, using an entity ID as a proxy for actual labels. [0056] As seen in the results below, a mean average precision (mAP) is assessed in each of the experiments, when using the dataset generated from a retailer's product collection as discussed above. Ramamonjison: [0074] As a result, the DomainMix augmentation module 804 mitigates the effect of noisy labels of the target domain samples during training because the mixed-image always contain accurate labels from the source domain. Finally, the DomainMix augmentation module 804 enforces the object detection model to detect small objects since the sizes of the objects in the original images are scaled down after the augmentation. [0075] The model adapter (e.g., implemented by batch norm adaptation operation 206 and fine tuning operation 212) coordinates the two-phase method for domain adaptation of a trained object detection model. The model adapter may be considered to be a model trainer. Its role is to take the training samples obtained with DomainMix augmentation module 804 and update the parameters of the batch norm layers during the first phase of the method of the present disclosure. In the second phase, the method uses new domain-mixed training samples with refined pseudo-labels and updates the parameters of all convolutional layers and batch norm layers of the neural network that approximates the object detection model. [0076] Thus, DomainMix augmentation module 804 performs a data augmentation method which adapts the feature representations generated by the trained object detection model to perform well on both the source domain and target domain. The data augmentation method generates new training images by randomly sampling images drawn from both a source domain data and a target domain data, and by mixing four sampled images at a time to create new mixed images. The random sampling is done with a weighted balanced sampler to handle the general case of unbalanced data, resulting in a data augmentation method that is effective even when the number of images from target domain is much less than that from the source domain. Furthermore, the data augmentation method transforms and collates the ground-truth labels of the source domain images with the pseudo-labels generated on the target domain images. The operations performed by the data augmentation method are shown in Algorithm 2 of FIG. 10.)
Consider Claims 9, 15 and 16.
The combination of Sisikand and Ramamonjishon teaches:
9. The system of claim 1 wherein the computer program code comprising instructions executable by the at least one processor is further configured for: sorting cropped images by file name. / 15. The method of claim 12 further comprising: sorting cropped images by file size; and removing cropped images from incorrect folders. / 16. The method of claim 12 further comprising: sorting cropped images by file name; and removing cropped images from incorrect folders.(Ramamonjison: [0004] An important form of domain shift is low-level image corruptions. When a trained model is a trained object detection model and corrupted digital images are input to the trained object detection model, the trained object detection model may not be able to detect any objects at all in the corrupted digital images, the trained object detection model may detect an object in the corrupted digital images but it may classify the object detected in the corrupted digital images with an incorrect object category label, or the trained object detection model may classify an object detected in the corrupted with a correct object category label but not generate an accurate bounding box for the detected object. If the distribution of a set of test images used for testing the performance of the trained object detection model drifts significantly from distribution of the set of labeled training images used to train the object detection model, the performance of trained object detention model when deployed for inference using a set of new images will degrade and the object detection model may completely fail at detecting objects in the set of new images. [0041] Source model 100 is configured to receive an input image 102, which may for example be an array of pixel data, identify a data object 112 in the input image 102 and generate a respective label 116 for the data object 112 identified in the input image 102. The input image 102 includes a group of pixels that represent a data object 112. Data object 112 corresponds to an object in the input image 102 that can be classified by predicting a class (or category) label from a set of possible candidate class (or category) labels for the data object 112. The generated respective output label 116 can include: (i) a bounding box definition that indicates a location of the data object 112 within the input image 102 (for example, x1, y1, width, height of bounding box 114); (ii) a predicted class (or category) label for the data object 112 (for example, “car”, “pedestrian”, “cat”); and (iii) a probability score for the predicted class (or category) label (for example 0.86 on a scale of 0 to 1). In some examples, an input image 102 can include multiple data objects 112, and the source model can identify the multiple data objects 112 and generate respective output labels 116 for each of the respective data objects 112. [0047] As noted above, a respective mixed dataset batch {circumflex over (β)}={({circumflex over (x)}i,ŷi)} is generated for each epoch. The generating of a mixed dataset batch {circumflex over (β)}={({circumflex over (x)}i,ŷi)} will now be described in the context of a first epoch of the first phase (Phase 1). As a first step, a pseudo-label generator operation 202 is applied to unlabeled target dataset D={(x j)} to generate a set of respective pseudo-labels {y j}. The pseudo-label generator operation 202 uses the source model 100 with the set of learned parameters θM s to generate the respective pseudo-labels {y j} for target data objects included in the target image samples {x j}. Each pseudo-label y j includes a predicted target class (or object category) label and a bounding box definition for the target data object. In some examples, the pseudo-label y j can also include a probability value for the predicted target class label. [0052] In the illustrated example of FIG. 3 , portions of source image samples 304_1 and 304_2 are cropped and positioned in the top left and bottom right quarter portions of a new domain-mixed composite image sample 308; portions of target image samples 306_1 and 306_2 are cropped and positioned in the top right and bottom left quarter portions of the new domain-mixed composite image sample 308. Accordingly, domain-mixed composite image sample 308 is a collage of cropped portions from four image samples that have been randomly sampled from the source and target datasets D∪D. In other examples, the number of data objects integrated into the new domain-mixed composite image sample 308 could be greater than or less than 4. The label set ŷi for the new domain-mixed composite image sample {circumflex over (x)}i will include: (i) the respective class labels for each of the data objects that have been cropped from source image samples, along with their respective adjusted bounding box definitions; and (ii) the respective pseudo-class labels for each of the data objects that have been cropped from target image samples, along with their respective adjusted bounding box definitions.)
Consider Claims 17-18 and 20.
The combination of Sisikand and Ramamonjishon teaches:
17. The method of claim 12 further comprising: training a labeling neural network, wherein the trained labeling neural network is used to generate the pseudo labels for the remaining images. 18. The method of claim 12 further comprising: training a classification neural network, wherein the trained classification neural network is used to assign the correct item names for the pseudo labels. / 20. The system of claim 19 wherein the computer program code comprising instructions executable by the at least one processor is further configured for: training a labeling neural network, wherein the trained labeling neural network is used to generate the pseudo labels for the remaining images; and training a classification neural network, wherein the trained classification neural network is used to assign the correct item names for the pseudo labels. (Sisikand: Testing/Validation of Image Collection [0054] Using the methods and systems described above, image collections may be rapidly developed by image capture and automated linking to objects which are in turn linked to item labels (e.g., via bar code or other unique identifier). Such automated image collection and annotation systems may be tested for accuracy relative to product image collections generated using other approaches. [0055] In one example experiment, a product image dataset of a retailer is assessed as to accuracy using a Resnet50 convolutional neural network (CNN) to output image embeddings for a k-NN classifier. In this example, a final pooling layer of the network is put through a fully-connected layer to product embeddings of dimension 100. The network is trained using a proxy-NCA loss (batch size: 32) and Adam optimizer (with learning rate of 104), with three experiments considered: intra-domain learning in which query and catalog sets are from the same image domain and the embedding is trained on labeled “in-the-wild” product images, supervised cross-domain learning in which a catalog set is taken from professional item photos, and weakly-supervised intra-domain learning in which embeddings are trained using unlabeled “in-the-wild” photos, using an entity ID as a proxy for actual labels. [0056] As seen in the results below, a mean average precision (mAP) is assessed in each of the experiments, when using the dataset generated from a retailer's product collection as discussed above. Ramamonjison: [0074] As a result, the DomainMix augmentation module 804 mitigates the effect of noisy labels of the target domain samples during training because the mixed-image always contain accurate labels from the source domain. Finally, the DomainMix augmentation module 804 enforces the object detection model to detect small objects since the sizes of the objects in the original images are scaled down after the augmentation. [0075] The model adapter (e.g., implemented by batch norm adaptation operation 206 and fine tuning operation 212) coordinates the two-phase method for domain adaptation of a trained object detection model. The model adapter may be considered to be a model trainer. Its role is to take the training samples obtained with DomainMix augmentation module 804 and update the parameters of the batch norm layers during the first phase of the method of the present disclosure. In the second phase, the method uses new domain-mixed training samples with refined pseudo-labels and updates the parameters of all convolutional layers and batch norm layers of the neural network that approximates the object detection model. [0076] Thus, DomainMix augmentation module 804 performs a data augmentation method which adapts the feature representations generated by the trained object detection model to perform well on both the source domain and target domain. The data augmentation method generates new training images by randomly sampling images drawn from both a source domain data and a target domain data, and by mixing four sampled images at a time to create new mixed images. The random sampling is done with a weighted balanced sampler to handle the general case of unbalanced data, resulting in a data augmentation method that is effective even when the number of images from target domain is much less than that from the source domain. Furthermore, the data augmentation method transforms and collates the ground-truth labels of the source domain images with the pseudo-labels generated on the target domain images. The operations performed by the data augmentation method are shown in Algorithm 2 of FIG. 10.)
Conclusion
The prior art made of record in form PTO-892 and not relied upon is considered pertinent to applicant's disclosure.
PNG
media_image1.png
152
856
media_image1.png
Greyscale
Any inquiry concerning this communication or earlier communications from the examiner should be directed to TAHMINA ANSARI whose telephone number is 571-270-3379. The examiner can normally be reached on IFP Flex - Monday through Friday 9 to 5.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, O’NEAL MISTRY can be reached on 313-446-4912. The fax phone numbers for the organization where this application or proceeding is assigned are 571-273-8300 for regular communications and 571-273-8300 for After Final communications. TC 2600’s customer service number is 571-272-2600.
Any inquiry of a general nature or relating to the status of this application or proceeding should be directed to the receptionist whose telephone number is 571-272-2600.
2674
/TAHMINA N ANSARI/Primary Examiner, Art Unit 2674
August 26, 2026